Gating speech processing using active acoustic sensing

Active acoustic sensing in hearables predicts speech intention using ultrasound, addressing power and hardware challenges to enhance speech processing efficiency and user experience.

WO2025193636A1PCT designated stage Publication Date: 2025-09-18GOOGLE LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/019256
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-14
Filing Date
2025-03-10
Publication Date
2025-09-18

AI Technical Summary

Technical Problem

Existing wireless hearables face challenges in efficiently managing speech processing due to power consumption, noise interference, and the need for additional hardware, which affects performance and user experience.

Method used

Implementing active acoustic sensing through audioplethysmography in hearables to predict speech intention using ultrasound signals, allowing for dynamic power management and enhanced speech processing without additional hardware.

Benefits of technology

Enables efficient power conservation and improved speech processing performance by predicting speech intention before audible sound, reducing power consumption and enhancing user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025019256_18092025_PF_FP_ABST
    Figure US2025019256_18092025_PF_FP_ABST
Patent Text Reader

Abstract

Techniques and apparatuses are described for gating speech processing using active acoustic sensing. During active acoustic sensing, a hearable (102) transmits and receives at least one ultrasound signal, which propagates within a user's (106) ear canal (116). This ultrasound signal can be modulated by muscle movements associated with speech. Some muscle movements, such as jaw movement, can indicate that the user (106) is about to speak before audible speech is generated. With active acoustic sensing, the hearable (102) can predict whether the user (106) is about to speak and use this information to gate speech processing (114). Gating speech processing (114) can conserve power and provide opportunities for enhancing speech processing.
Need to check novelty before this filing date? Find Prior Art

Description

GATING SPEECH PROCESSING USING ACTIVE ACOUSTIC SENSINGBACKGROUND

[0001] Wireless technology has become prevalent in every day life, making communication and data readily accessible to users. One ty pe of wireless technology' are wireless hearables, examples of which include wireless earbuds and wireless headphones. Wireless hearables have allowed users freedom of movement while listening to audio content from music, audio books, podcasts, and videos. With the prevalence of wireless hearables, there is a market for adding additional features to existing hearables without introducing hardw are changes.SUMMARY

[0002] Techniques and apparatuses are described for gating speech processing using active acoustic sensing. During active acoustic sensing, a hearable transmits and receives at least one ultrasound signal, which propagates within a person's ear canal. This ultrasound signal can be modulated by muscle movements associated with speech. Some muscle movements, such as jawmovement, can indicate that the person is about to speak before audible speech is generated. With active acoustic sensing, the hearable can predict that the person is about to speak and use this information to gate speech processing. Gating speech processing can conserve power and provide opportunities for enhancing speech processing.

[0003] Aspects described below include a method for gating speech processing using active acoustic sensing. The method includes transmitting, during a first time period, an ultrasound transmit signal that propagates w ithin at least a portion of an ear canal of a person. The method also includes receiving, during the first time period, an ultrasound receive signal. The ultrasound receive signal represents a version of the ultrasound transmit signal with one or more characteristics modified based on the propagation within the ear canal and based on a muscle movement made by the person prior to speaking. The method additionally includes predicting that the person is about to speak based on the ultrasound receive signal. Responsive to the predicting, the method further includes generating a control signal that causes a speech-processing mode to transition from a low-power state to a high-power state.

[0004] Aspects described below include a computer-readable storage medium comprising instructions that, responsive to execution by a processor, cause a hearable to perform any one of the methods described herein.

[0005] Aspects described below include a system with at least one transducer and at least one processor. The system is configured to perform, using the at least one transducer and the at least one processor, any one of the methods described herein.

[0006] Aspects described below include a system with means for gating speech processing using active acoustic sensing.BRIEF DESCRIPTION OF DRAWINGS

[0007] Apparatuses for and techniques that gate speech processing using active acoustic sensing are described with reference to the following drawings. The same numbers are used throughout the drawings to reference like features and components:FIG. 1 illustrates an example environment in which gating speech processing using active acoustic sensing can be implemented;FIG. 2-1 illustrates a progression of events associated with gating speech processing using active acoustic sensing;FIG. 2-2 illustrates an example timing of events associated with gating speech processing using active acoustic sensing;FIG. 3 illustrates example components of a computing device;FIG. 4 illustrates example components of a hearable;FIG. 5 illustrates example operations of two hearables;FIG. 6 illustrates an example implementation of a hearable capable of gating speech processing using active acoustic sensing;FIG. 7 illustrates an example flow diagram for gating speech processing using active acoustic sensing;FIG. 8 illustrates an example pre-processed signal for performing aspects of gating speech processing using active acoustic sensing;FIG. 9 illustrates an example method for gating speech processing using active acoustic sensing;FIG. 10 illustrates another example method for gating speech processing using active acoustic sensing; andFIG. 11 illustrates an example computing system embodying, or in which techniques may be implemented that enable use of, gating speech processing using active acoustic sensing.DETAILED DESCRIPTION

[0008] As electronic devices become more ubiquitous, users incorporate them into everyday life. A user, for example, may use an electronic device to get daily weather and traffic information,control a temperature of a home, answer a doorbell, turn on or off a light, and / or play background music. Interacting with some electronic devices, however, can be cumbersome and inefficient. An electronic device, for instance, can have a physical user interface that may require a user to navigate through one or more prompts by physically touching the electronic device. In this case, the user has to devote attention away from other primary tasks to interact with the electronic device, which can be inconvenient and disruptive.

[0009] To address this problem, some electronic devices support voice control, which enables a user to interact with the electronic device in a non-physical and less cognitively demanding way compared to other interfaces that require physical touch and / or the user’s visual attention. With voice control, the electronic device seamlessly exists in the surrounding environment and provides the user access to information and services while the user performs a primary task, such as cooking, cleaning, driving, talking with people, or reading a book. For voice control, the electronic device detects a user’s speech and recognizes a phrase (or command) that is spoken by the user. The detection of the user’s speech and the recognition of a phrase are examples of speech processing performed by the electronic device.

[0010] While voice control can provide a convenient means of interacting with an electronic device, there are several challenges associated with speech processing in general. In a noisy environment, for instance, the user’s voice can be imperceptible. Consequently, it can be challenging to detect and / or recognize phrases spoken by the user. Also, sometimes the noisy environment can cause logic that performs speech processing to incorrectly respond to a voice of another person who is not authorized to use the electronic device.

[0011] Also, some implementations of speech processing can utilize a significant amount of power to operate. For many speech-processing techniques, there is a direct relationship between performance and power consumption. Increasing performance for speech processing, for instance, can come at the cost of increasing power consumption. As such, high-performance speechprocessing techniques can consume more power compared to low-perfonnance speech-processing techniques. Consequently, power-constrained devices, such as mobile devices, may implement speech-processing techniques that consume less power at the cost of decreased performance. Sometimes this lower level of performance for speech processing can degrade the user experience.

[0012] To improve aesthetics and reduce encumbrance, it can be desirable to design hearables with smaller sizes. As space becomes limited, it can be challenging to integrate additional components, such as a voice accelerometer, within the hearables. With the prevalence of hearables, there is a market for adding additional features to existing hearables to facilitate speech processing without introducing hardware changes.

[0013] Provided according to one or more preferred embodiments is a hearable, such as an earbud, that is capable of performing a novel physiological monitoring process termed herein audioplethysmography. Audioplethysmography is an active acoustic method capable of sensing subtle physiologically -related changes observable at a user’s outer and middle ear. Instead of relying on other auxiliary sensors, such as optical or electrical sensors, audioplethysmography involves transmitting and receiving ultrasound signals that at least partially propagate within a user’s ear canal. To perform audioplethysmography, the hearable forms at least a partial seal in or around the user’s outer ear. This seal enables formation of an acoustic circuit, which includes the seal, the hearable, the ear canal, and an ear drum of the ear. By transmitting and receiving ultrasound signals, the hearable can recognize changes in the acoustic circuit associated with muscle movements made by the user as the user prepares to speak. In this manner, active acoustic sensing enables the hearable to predict whether the user is about to speak. The user’s speech can include any sound that is produced using the user’s lung’s, vocal cords, and / or mouth. Example E pes of vocalizations can involve the user speaking, whispering, shouting, whistling, singing, or making other audible utterances.

[0014] During active acoustic sensing, a hearable transmits and receives at least one ultrasound signal, which propagates within the user’s ear canal. This ultrasound signal can be modulated by the muscle movements associated with speech (e.g., jaw movement). With active acoustic sensing, the hearable can predict that the user is about to speak before audible speech is produced. Speech prediction can be used to gate (e.g., control; activate and / or deactivate; or change a state of) speech processing that is performed by a device (e.g., the hearable or another device that is communicatively coupled to the hearable). In other words, the speech prediction controls a speech-processing mode of the device. The quick reaction time associated with active acoustic sensing for speech prediction provides sufficient time for speech-processing techniques to be initialized and operational prior to the user audibly speaking.

[0015] Some techniques may attempt to predict speech using another type of sensor, such as an inertial measurement unit (IMU). The inertial measurement unit, however, can be highly susceptible to noise. Noise caused by the user performing normal motions, such as moving their head, walking, and / or running, can making it challenging for the inertial measurement unit to detect smaller muscle movements that indicate the user is about to speak. Active acoustic sensing, in contrast, is less susceptible to this noise and can therefore support gating speech processing in a variety of situations in which the user is moving.

[0016] Other techniques may attempt to gate speech processing based on detecting a user’s voice. These techniques may utilize a sensor, such as a voice accelerometer. The reliance upon detectingan audible sound, however, causes an inherent delay in gating speech processing. This delay can degrade the performance of speech processing or may require the user to preface their speech in some manner. Active acoustic sensing, in contrast, can predict that the user is about to speak and appropriately gate speech processing prior to the user producing an audible sound.

[0017] Utilizing active acoustic sensing for gating speech processing can provide several benefits. In a first aspect, the speech-processing mode can be placed in a low-power state (e.g., disabled) while the user is not speaking and dynamically placed in a high-power state (e.g., activated) once speech is predicted. This low-power state enables the device to conserve power.

[0018] In a second aspect, gating speech processing provides opportunities to enhance speech processing. As gating speech processing can allow for more relaxed power-constraints regarding speech processing, the device can utilize high-performance speech-processing techniques that consume larger amounts of power instead of low-performance speech-processing techniques. These high-performance speech-processing techniques can utilize larger and / or more computationally complex models and / or algorithms, a larger quantity of input parameters, a machine-learned model, other techniques such as beamforming, and so forth. In addition to being relatively unobtrusive, some hearables can be configured to gate speech processing using active acoustic sensing without the need for additional hardware. As such, the size, cost, and power usage of the hearable can help make speech processing accessible to a larger group of people and improve the user experience wi th hearables.Operating Environment

[0019] FIG. 1 is an illustration of an example environment 100 in which active acoustic sensing can be implemented. In the example environment 100, a hearable 102 is connected to a computing device 104 using a physical or wireless interface. The hearable 102 is a device that can play audible content provided by the computing device 104 and direct the audible content into a user 106's ear 108. The user 106 generally represents a person who is wearing the hearable 102. In some cases, the user 106 can be an authorized user of the hearable 102 and / or an authorized user of the computing device 104. The techniques for gating speech processing using active acoustic sensing can be applied regardless of whether the person wearing the hearable 102 is an authorized user of the hearable 102 and / or an authorized user of the computing device 104. In this example, the hearable 102 operates together with the computing device 104. In other examples, the hearable 102 can operate or be implemented as a stand-alone device. Although depicted as a smartphone, the computing device 104 can include other types of devices, including those described with respect to FIG. 3.

[0020] The hearable 102 is capable of performing audioplethysmography 110, which is an active acoustic method of sensing that occurs at the ear 108. The hearable 102 can perform this sensing without the use of other auxiliary sensors, such as an optical sensor or an electrical sensor. Through audioplethysmography 110, the hearable 102 can perform speech prediction 112 and can gate speech processing 114. Speech prediction 112 enables the hearable 102 (or the computing device 104) to recognize when the user 106 is preparing to speak or is about to speak prior to the user 106 uttering an audible sound. In particular, speech prediction 112 recognizes a muscle movement (e.g., ajaw movement) associated with the user 106 preparing to speak. An example muscle movement can include the user 106 opening their mouth. The speech prediction 112 can also be referred to as speech-intention detection. By predicting that the user 106 is about to speak, the speech prediction 112 is, in a general sense, detecting an intention of the user 106 to speak. The user 106’s speech can involve any type of vocalization associated with speaking, whispering, shouting, whistling, singing, or other utterances. The speech can include a single word, multiple words, or a tone.

[0021] Gating speech processing 114 enables the hearable 102 (or the computing device 104) to control aspects of speech processing based on speech prediction 112. Explained another way, gating speech processing 114 involves controlling a speech-processing mode of the device. By gating speech processing 114, the hearable 102 (or the computing device 104) can better manage power consumption compared to other devices that do not have this capability. Without gating speech processing 114, these other devices may continuously perform speech processing, which can consume a substantial amount of power. Furthermore, the power savings provided through gating speech processing 114 enables the hearable 102 (or the computing device 104) to implement speech-processing techniques that provide a higher level performance at the cost of increased power consumption.

[0022] To perform speech prediction 112, the hearable 102 uses audioplethysmography 110 to detect subtle pressure waves that propagate to the user 106’s ear canal 116. These pressure waves modify characteristics of ultrasound signals that are transmitted and received by the hearable 102 and propagate through the ear canal 116. Prior to the user 106 speaking, the ear canal 116 deforms at least in part due to the muscle movements made by the user 106 in preparation for speaking. Example muscle movements can include the user 106 positioning their jaw' in preparation for making a first utterance. The ear canal 116 also deforms while the user 106 is speaking and / or making muscle movements associated with speaking. Example muscle movements can include the user 106 moving their jaw and / or tongue to make different sounds.

[0023] To use audioplethysmography 1 10, the user 106 positions the hearable 102 in a manner that creates at least a partial seal 118 around or in the ear 108. Some parts of the ear 108 are shown in FIG. 1, including the ear canal 116 and an ear drum 120 (or ty mpanic membrane). Due to the seal 118, the hearable 102, the ear canal 116. and the ear drum 120 couple together to form an acoustic circuit. Audioplethysmography 110 involves, at least in part, measuring properties associated with this acoustic circuit. The properties of the acoustic circuit can change due to a variety of different situations or actions.

[0024] For example, consider a change that occurs in a physical structure of the ear 108. Example changes to the physical structure include a change in a geometric shape of the ear canal 116 and / or a change in a volume of the ear canal 116. This change can be caused, at least in part, by a pressure wave associated with the user 106’s speech. For instance, the tissue around the ear canal 116 and the ear drum 120 itself are slightly “squeezed’' due to the bone conduction and / or the pressure wave. This squeeze causes a volume of the ear canal 116 to be slightly reduced. As the squeezing subsides, the volume of the ear canal 116 is slightly increased. The increasing and decreasing of the volume of the ear canal 116 is indicated by the arrows in FIG. 1. The physical changes within the ear 108 can modulate an amplitude and / or phase of an ultrasound signal that propagates through the ear canal 116.

[0025] The techniques for audioplethysmography 110 can be performed while the hearable 102 is rendering (e.g., playing or transmitting) audible content and / or while the user 106 is actively moving or performing an activity. As such, active acoustic sensing enables the hearable 102 to perform speech prediction 112 in a variety7of different situations. One such situation is further described with respect to FIG. 2-1.

[0026] FIG. 2-1 illustrates a progression of events associated with gating speech processing 1 14 using active acoustic sensing. In an example environment, a user 106 wears at least one hearable 102, which is communicatively coupled to the computing device 104. The computing device 104 can perform aspects of speech processing using at least one speech processor 202. Speech processing can involve performing voice activity detection, performing voice authentication, performing speech recognition, performing conversation detection, and / or providing a voice assistant sendee. Generally speaking, speech processing can involve any type of speech-based application or processing of a speech signal (e.g., a voice signal). With speech processing, the computing device 104 can support voice control and / or speech-to-text functions. Additionally or alternatively, the hearable 102 can perform aspects of speech processing and can include the speech processor 202 (or an additional speech processor). An operation of the speech processor 202 consumes power. To manage power consumption, the hearable 102 performsspeech prediction 1 12 using active acoustic sensing and gates speech processing 114 (e.g., controls a speech-processing mode) based on the speech prediction 112, as further described below.

[0027] At time 200-1, the user 106 is not speaking. To conserve power while the user 106 is not speaking, the speech processor 202 operates in a low-power state 204. The low-power state 204 can be a disabled state, an inactive state, a sleep state, a powered-down state, or any state that consumes less power compared to a high-power state 206. In some implementations, the low- power state 204 can indicate that at least a portion of components (e.g., circuits and / or sensors) associated with speech processing (e.g., used to perform speech processing) are disabled. Example components can include a voice accelerometer or a denoising filter. Depending on the implementation, the speech processor 202 may or may not perform aspects of speech processing while in the low-power state 204. In a first example implementation, the speech processor 202 does not perform speech processing while in the low-power state 204. In a second example implementation, the speech processor 202 performs a few operations associated with speech processing while in the low-power state 204. Such operations can include sampling a received speech signal. In a third example implementation, the speech processor 202 performs a low- powered version of speech processing during the low -pow er state 204. For example, the speech processor 202 can perform a less computationally-complex version of speech processing while in the low-power state 204. As another example, the speech processor 202 can perform some aspects of speech processing that generally consume less power, such as voice activity detection, and does not perform other aspects of speech processing that generally consume more power, such as speech recognition, while in the low-power state 204. As yet another example, the speech processor 202 can sample a receive speech signal at a lower-than-normal sampling rate.

[0028] At time 200-2, the user 106 moves their muscles in preparation to speak a phrase 208. The hearable 102 predicts that the user 106 is about to speak and causes the speech processor 202 to operate in the high-power state 206. The high-power state 206 can be an enabled state, an active state, a powered-up state, or any state that consumes more power compared to the low-power state 204. The power consumption at the high-power state 206 can be on the order of approximately 15% or more (e.g., 20%, 25%, 50%, 100%, or 150%) compared to the power consumption at the low-power state 204. The term “approximately” means that the power consumption can be within 5% of a given value or less (e.g., within 3%, 2%, or 1% of the given value). Generally speaking, the high-power state 206 consumes more power than the low-power state 204. While in the high-power state 206, the speech processor 202 can perform speech processing. In comparison to the low-power state 204, the high-power state 206 enables thespeech processor 202 to perform more aspects of speech processing and / or enables the speech processor 202 to realize a higher level of performance compared to the low-power state 204.

[0029] The hearable 102 can predict that the user 106 is about to speak in sufficient time to enable the speech processor 202 to transition from the low-power state 204 to the high-power state 206. In other words, the speech prediction 112, the gating of speech processing 114, and the transition to the high-power state 206 can occur prior to the user 106 audibly speaking the phrase 208. During the transition between the low-power state 204 and the high-power state 206, the speech processor 202 may power up components and / or perform an initialization procedure.

[0030] At 200-3, the user 106 stops speaking and the hearable 102 causes the speech processor 202 to operate in the low-power state 204. The transition between the high-power state 206 and the low-power state 204 can occur shortly after the user 106 stops speaking. The events and operations that take place between times 200-1 to 200-3 are further described with respect to FIG. 2-2.

[0031] FIG. 2-2 illustrates an example timing diagram 210 associated with gating speech processing 114 using active acoustic sensing. The various actions performed by the user 106 are depicted at the top of the timing diagram 210. Operations of the hearable 102 with respect to speech prediction 112 are depicted in the middle of the timing diagram 210. The various states of the speech processor 202 are depicted at the bottom of the timing diagram 210. Durations of the actions, operations, and / or states depicted in the timing diagram 210 are for illustration purposes and are not drawTi to scale.

[0032] During the time 202-1, the user 106 does not speak. As such, no speech is present, as indicated at 212. The hearable 102 performs speech prediction 112 and predicts that the user 106 is not about to speak and / or determines that the user 106 is not speaking, as indicated at 214. The speech-processing mode associated with the speech processor 202 is in the low-powder state 204 to conserve power. Depending on the implementation, the speech processor 202 may or may not be performing some aspects of speech processing while in the low -power state 204.

[0033] During the time 202-2. the user 106 prepares to speak, as indicated at 214. The hearable 102 performs speech prediction 112 and predicts that the user 106 is about to speak, as indicated at 216. The hearable 102 gates speech processing 114 based on the prediction. In particular, the hearable 102 causes the speech-processing mode associated with the speech processor 202 to transition from the low-power state 204 to the high-power state 206. There may be some delay associated with the prediction and the speech-processing mode transitioning to the high-power state 206. Generally speaking, this delay is sufficiently small such that the speechprocessor 202 can be in the high-power state 206 prior to the user 106 speaking audibly, as indicated at 218.

[0034] While the user 106 speaks, the hearable 102 can continue performing speech prediction 112. The movement of the user 106's muscles to create the speech can cause the hearable 102 to continue to predict the occurrence of speech, as indicated at 216. The speech processing mode can continue to be in the high-power state 206 to enable the speech processor 202 to perform speech processing during this time.

[0035] Once the user 106 stops speaking at time 202-3, the hearable 102 predicts that the user 106 is not about to speak and / or detemrines that the user 106 is not speaking. As such, the hearable 102 causes the speech-processing mode to transition from the high-power state 206 to the low-power state 204. This enables the device with the speech processor 202 (e.g., the computing device 104 or the hearable 102) to conserve power.

[0036] Generally speaking, the hearable 102 has sufficient power resources to run active acoustic sensing in a continuous manner. This enables speech prediction 112 to be performed in a continuous manner, which increases the responsiveness for gating speech processing 114. The computing device 104 and the hearable 102 are further described with respect to FIGs. 3 and 4, respectively.

[0037] FIG. 3 illustrates an example implementation of the computing device 104. The computing device 104 is illustrated with various non-limiting example devices including a desktop computer 104-1, a tablet 104-2, a laptop 104-3, a television 104-4, a computing watch 104-5, computing glasses 104-6, a gaming system 104-7, a microwave 104-8, and a vehicle 104-9. Other devices may also be used, such as an augmented and / or virtual reality headset, a home service device, a smart speaker, a smart thermostat, a baby monitor, a Wi-Fi™ router, a drone, a trackpad, a drawing pad, a netbook, an e-reader, a home automation and control system, a wall display, and another home appliance. Note that the computing device 104 can be wearable, non-wearable but mobile, or relatively immobile (e.g., desktops and appliances).

[0038] The computing device 104 includes one or more computer processors 302 and at least one computer-readable medium 304, which includes memory media and storage media. Applications and / or an operating system (not shown) embodied as computer-readable instructions on the computer-readable medium 304 can be executed by the computer processor 302 to provide some of the functionalities described herein. The computer-readable medium 304 can optionally include an application 306, which can utilize some aspect of speech processing. The computer- readable medium 304 also includes a speech processor 202. Other implementations are also possible in which the speech processor 202 is implemented in hardware, software, and / or firmwarethat is considered separate from the computer-readable medium 304. Generally speaking, the application 306 can use information provided by the speech processor 202 to perform an action. Example actions can include granting the user 106 access to information stored within the application 306, enabling the user 106 to control the computing device 104 via voice commands, or providing a speech-to-text feature.

[0039] The computing device 104 can also include a network interface 308 for communicating data over wired, wireless, or optical networks. For example, the network interface 308 may communicate data over a local-area-network (LAN), a wireless local-area-network (WLAN), a personal-area-network (PAN), a wire-area-network (WAN), an intranet, the Internet, a peer-to- peer network, point-to-point network, a mesh network, Bluetooth " . and the like. The computing device 104 may also include the display 310. Although not explicitly shown, the hearable 102 can be integrated within the computing device 104, or can connect physically or wirelessly to the computing device 104. The hearable 102 is further described with respect to FIG. 4.

[0040] FIG. 4 illustrates an example hearable 102. The hearable 102 is illustrated with various non-limiting example devices, including wireless earbuds 402-1, wired earbuds 402-2, and headphones 402-3. The earbuds 402-1 and 402-2 are a type of in-ear device that fits into the ear canal 116. Each earbud 402-1 or 402-2 can represent a hearable 102. Headphones 402-3 can rest on top of or over the ears 108. The headphones 402-3 can represent closed-back headphones, open-back headphones, on-ear headphones, or over-ear headphones. Each headphone 402-2 includes two hearables 102, which are physically packaged together. In general, there is one hearable 102 for each ear 108. The headphones 402-3 may be designed in some manner or may utilize techniques, such as beamforming, to assist with directing signals used for audioplethysmography 1 10 into the ear canal 116. The hearable 102 can also represent a hearing aid (not shown).

[0041] The hearable 102 includes a communication interface 404 to communicate with the computing device 104, though this need not be used when the hearable 102 is integrated within the computing device 104. The communication interface 404 can be a wired interface or a wireless interface, in which audio content is passed from the computing device 104 to the hearable 102. The hearable 102 can also use the communication interface 404 to pass data associated with audioplethysmography 110, speech prediction 112, and / or gating speech processing 114 to the computing device 104. In general, the data provided by the communication interface 404 is in a format usable by the application 306, the speech processor 202, or another application of the computing device 104.

[0042] The communication interface 404 also enables the hearable 102 to communicate with another hearable 102. During bistatic sensing, for instance, the hearable 102 can use the communication interface 404 to coordinate with the other hearable 102 to support two-ear audioplethysmography 110, as further described with respect to FIG. 5. In particular, the transmitting hearable 102 can communicate timing and waveform information to the receiving hearable 102 to enable the receiving hearable 102 to appropriately demodulate a received ultrasound signal.

[0043] The hearable 102 includes at least one transducer 406 that can convert electrical signals into sound waves. The transducer 406 can also detect and convert sound waves into electrical signals. These sound waves may include ultrasonic frequencies, which may be used for audioplethysmography 110. In particular, a frequency spectrum (e.g., range of frequencies) that the transducer 406 uses to generate an ultrasound signal can include frequencies from the ultrasonic range, e.g., between 20 kHz to 2 megahertz (MHz). Other example frequency spectrums for audioplethysmography 110 can encompass frequencies between 20 and 60 kHz or between 30 and 40 kHz.

[0044] In an example implementation, the transducer 406 has a monostatic topology7. With this topology, the transducer 406 can convert the electrical signals into sound waves and convert sound waves into electrical signals (e.g., can transmit or receive acoustic signals). Example monostatic transducers may include piezoelectric transducers, capacitive transducers, and micro-machined ultrasonic transducers (MUTs) that use microelectromechanical systems (MEMS) technology.

[0045] Alternatively, the transducer 406 can be implemented with a bistatic topology, which includes multiple transducers that are physically separate. In this case, a first transducer converts the electrical signal into sound waves (e.g., transmits acoustic signals), and a second transducer converts sound waves into an electrical signal (e.g., receives the acoustic signals). An example bistatic topology can be implemented using at least one speaker 408 and at least one microphone 410. The speaker 408 and the microphone 410 can be dedicated for audioplethysmography 110 or can be used for both audioplethysmography 110 and other functions of the computing device 104 (e.g., passive audio sensing, presenting audible content to the user 106, capturing the user 106‘s voice for a phone call, or for voice control).

[0046] In general, the speaker 408 and the microphone 410 are directed towards the ear canal 116 (e.g., oriented towards the ear canal 116). Accordingly, the speaker 408 can direct ultrasound signals towards the ear canal 116, and the microphone 410 is responsive to receiving ultrasound signals from the direction associated with the ear canal 116.

[0047] The hearable 102 includes at least one analog circuit 412, which includes circuitry and logic for conditioning electrical signals in an analog domain. The analog circuit 412 can include analog-to-digital converters, digital-to-analog converters, amplifiers, filters, mixers, and switches for generating and modifying electrical signals. In some implementations, the analog circuit 412 includes other hardware circuitry associated with the speaker 408 or microphone 410.

[0048] The hearable 102 also includes at least one system processor 414 and at least one system medium 416 (e.g., one or more computer-readable storage media). In the depicted configuration, the system medium 416 includes a pre-processing module 418, a speech predictor 420 (or a speech onset detector), and a speech-processing gate 422 (e.g.. a speech-processing controller). The system medium 416 also optionally includes a calibration module 424. The pre-processing module 418, the speech predictor 420, the speech-processing gate 422, and the calibration module 424 can be implemented using hardware, software, firmware, or a combination thereof. In this example, the system processor 414 implements the pre-processing module 418, the speech predictor 420, the speech-processing gate 422, and the calibration module 424. In an alternative example, the computer processor 302 of the computing device 104 can implement at least a portion of the pre-processing module 418, the speech predictor 420, the speech-processing gate 422, and / or the calibration module 424. In this case, the hearable 102 can communicate digital samples of the ultrasound signals to the computing device 104 using the communication interface 404.

[0049] Operations of the pre-processing module 418, the speech predictor 420, the speechprocessing gate 422, and the calibration module 424 are further described with respect to FIG. 6. Aspects of speech prediction 112 using active acoustic sensing can be performed, at least partially, by the speech predictor 420. Aspects of gating speech processing 114 using active acoustic sensing can be performed, at least partially, by the speech-processing gate 422.

[0050] Some hearables 102 include an active-noise-cancellation circuit 426, which enables the hearables 102 to reduce background or environmental noise. In this case, the microphone 410 used for audioplethysmography 110 can be implemented using a feedback microphone of the active-noise-cancellation circuit 426. During active noise cancellation, the feedback microphone provides feedback information regarding the performance of the active noise cancellation. During audioplethysmography 110, the feedback microphone receives an ultrasound signal, which is provided to the pre-processing module 418. In some situations, active noise cancellation and audioplethysmography 1 10 are performed simultaneously using the feedback microphone. In this case, the ultrasound signal received by the feedback microphone can be provided to the pre-processing module 418 and the feedback signal for active noise cancellation can be provided to the active-noise-cancellation circuit 426.

[0051] The hearable 102 can optionally include an auxiliary' sensor 428. The auxiliary' sensor 428 can detect an audible sound produced by the user 106 to assist with gating speech processing 114. An example auxiliary sensor 428 can include a voice accelerometer.

[0052] Although not explicitly shown in FIG. 4, the system medium 416 can also include the speech processor 202 (or another speech processor). In this case, the hearable 102 uses the speech processor 202 to perform one or more aspects of speech processing (e.g., voice activity detection, voice authentication, speech recognition, conversation detection, and / or voice assistant service). Different types of audioplethysmography 110 are further described with respect to FIG. 5.Active Acoustic Sensing

[0053] FIG. 5 illustrates example operations of two hearables 102-1 and 102-2. In a first example operation, the hearables 102-1 and 102-2 perform single-ear audioplethysmography 110. This means that the hearables 102-1 and 102-2 independently perform audioplethysmography 110 on different ears 108 of the user 106. In this case, the first hearable 102-1 is proximate to (arranged at) the user 106’s right ear 108, and the second hearable 102-2 is proximate to (arranged at) the user 106's left ear 108. Each hearable 102-1 and 102-2 includes a speaker 408 and a microphone 410. The hearables 102-1 and 102-2 can operate in a monostatic manner during the same time period or during different time periods. In other words, each hearable 102-1 and 102-2 can independently transmit and receive ultrasound signals.

[0054] For example, the first hearable 102-1 uses the speaker 408 to transmit a first ultrasound transmit 502-1, which propagates within at least a portion of the user 106’s right ear canal 116. The first hearable 102-1 uses the microphone 410 to receive a first ultrasound receive signal 504-1. The first ultrasound receive signal 504-1 represents a version of the first ultrasound transmit signal 502-1 that is modified, at least in part, by the acoustic circuit associated with the right ear canal 116. This modification can change an amplitude, phase, and / or frequency of the first ultrasound receive signal 504-1 relative to the first ultrasound transmit signal 502-1.

[0055] Similarly, the second hearable 102-2 uses the speaker 408 to transmit a second ultrasound transmit signal 502-2, which propagates within at least a portion of the user 106’s left ear canal 116. The second hearable 102-2 uses the microphone 410 to receive a second ultrasound receive signal 504-2. The second ultrasound receive signal 504-2 represents a version of the second ultrasound transmit signal 502-2 that is modified by the acoustic circuit associated withthe left ear canal 116. This modification can change an amplitude, phase, and / or frequency of the second ultrasound receive signal 504-2 relative to the second ultrasound transmit signal 502-2.

[0056] The techniques of single-ear audioplethysmography 110 can be particularly beneficial as it enables the computing device 104 to compile infonnation from both hearables 102-1 and 102-2, which can further improve measurement confidence. For some aspects of audioplethysmography 1 10, it can be beneficial to analyze the acoustic channel between two ears 108, as further described below.

[0057] In a second example operation, the two hearables 102-1 and 102-2 perform two-ear audioplethysmography 110. This means that the hearables 102-1 and 102-2 jointly perform audioplethysmography 1 10 across two ears 108 of the user 106. In this case, at least one of the hearables 102 (e.g., the first hearable 102-1) includes the speaker 408, and at least one of the other hearables 102 (e.g., the second hearable 102-2) includes the microphone 410. The hearables 102-1 and 102-2 operate together in a bistatic manner during the same time period.

[0058] During operation, the first hearable 102- 1 transmits a third ultrasound transmit 502-3 using the speaker 408. The third ultrasound transmit signal 502-3 propagates through the user 106’s right ear canal 116. The third ultrasound transmit signal 502-3 also propagates through an acoustic channel that exists between the right and left ears 108. In the left ear 108, the third ultrasound transmit signal 502-3 propagates through the user 106's left ear canal 116 and is represented as a third ultrasound receive signal 504-3. The second hearable 102-2 receives the third ultrasound receive signal 504-3 using the microphone 410. The third ultrasound receive signal 504-3 represents a version of the third ultrasound transmit signal 502-3 that is modified by the acoustic circuit associated with the right ear canal 116, modified by the acoustic channel associated with the user 106’s face, and modified by the acoustic circuit associated with the left ear canal 116. This modification can change an amplitude, phase, and / or frequency of the third ultrasound receive signal 504-3 relative to the third ultrasound transmit signal 502-3. In some cases, the hearable 102-2 measures the time-of-flight (ToF) associated with the propagation from the first hearable 102-1 to the second hearable 102-2. Sometimes a combination of single-ear and two-ear audioplethysmography 1 10 are applied to further improve measurement confidence.

[0059] The ultrasound transmit signals 502 of FIG. 5 can represent a variety of different ty pes of signals as described above with respect to FIG. 4. In example implementations, the ultrasound transmit signal 502 can be a continuous-wave signal (e.g., a sinusoidal signal) or a pulsed signal. Some ultrasound transmit signals 502 can have a particular tone (or frequency). Other ultrasound transmit signals 502 can have multiple tones (or multiple frequencies). A variety7of modulations can be applied to generate the ultrasound transmit signal 502. Example modulations include linearfrequency modulations, triangular frequency modulations, stepped frequency modulations, phase modulations, or amplitude modulations. The ultrasound transmit signal 502 can be transmitted as part of a calibration procedure or a measurement procedure, as further described as part of FIG. 6.Gating Speech Processing

[0060] FIG. 6 illustrates an example implementation of the hearable 102. In the depicted configuration, the hearable 102 includes the speaker 408, the microphone 410, the analog circuit 412, the pre-processing module 418, the speech predictor 420, the speech-processing gate 422. and the calibration module 424. Other implementations of the hearable 102 are also possible in which the hearable 102 does not include the calibration module 424 to reduce processing power requirements. In this case, the pre-processing module 418 can perform aspects of frequency selection as further described below to improve the signal-to-noise ratio for audioplethysmography 110.

[0061] Outputs of the speaker 408 and the microphone 410 are coupled to inputs of the analog circuit 412. The pre-processing module 418 has inputs that are coupled to outputs of the analog circuit 412. The pre-processing module 418 also has an output that is coupled to inputs of the speech predictor 420 and the calibration module 424. In an example implementation, the preprocessing module 418 includes at least one in-phase and quadrature mixer (I / Q mixer) and at least one filter. The in-phase and quadrature mixer performs frequency down-conversion and can be implemented using at least two mixers, at least one phase shifter, and at least one combiner (e.g., a summation circuit). The filter attenuates intermodulation products that are generated by the in-phase and quadrature mixer. In an example implementation, the filter is implemented using a low-pass filter.

[0062] The pre-processing module 418 can optionally include at least one frequency selector. The frequency selector can identify and select one or more tones (or carrier frequencies) that provide a high-quality signal for later processing. The frequency selector can further pass the selected tones to other processing modules (e.g.. the speech predictor 420) and filter (or attenuate) other tones that are not selected. The frequency selector can be implemented in a similar manner as the calibration module 424, which is further described below.

[0063] An output of the speech predictor 420 is coupled to an input of the speech-processing gate 422. The speech predictor 420 can be implemented using a variety of techniques, including signal processing algorithms and / or a machine-learned model. With the speech predictor 420, the hearable 102 performs a measurement procedure that includes performing speech prediction 112 using audioplethysmography 110.

[0064] An output of the speech-processing gate 422 can be coupled (e.g., communicatively coupled) to the speech processor 202 (not shown) or another circuit (not shown) that controls pow er to the speech processor 202. The speech-processing gate 422 gates speech processing 114 based at least on an output of the speech predictor 420. In some implementations, the speechprocessing gate 422 can perform some aspect of gating speech processing 114 based on additional information provided by the auxiliary sensor 428.

[0065] The calibration module 424 has an output that is coupled to the speaker 408. The calibration module 424 includes at least one frequency selector. The frequency selector can include at least one amplitude detector, at least one phase detector, at least one quality detector, and at least one comparator. Using the frequency selector, the calibration module 424 can perform a calibration procedure that determines appropriate characteristics (e.g., waveform or signal characteristics) of ultrasound transmit signals 502 to improve audioplethysmography 110 (e.g., to enhance the performance of speech prediction 112). The calibration procedure enables audioplethysmography 1 10 to take into account the wear of the hearable 102 (e.g., the position of the hearable 102 relative to the ear canal 116) and the physical structure of the ear canal 116 to determine a transmission frequency that can increase sensitivity.

[0066] Consider an example operation of the hearable 102 in accordance with single-ear audioplethysmography 110. In this example, the hearable 102 includes the calibration module 424. With the calibration module 424, the hearable 102 can perform a calibration procedure prior to performing a measurement procedure. In some circumstances, the hearable 102 can perform on-head detection (or in-ear detection) by detecting the presence of the seal 118 and initiating the calibration procedure and / or the measurement procedure based on a determination that on-head detection is “true.” In other circumstances, the hearable 102 can initiate the calibration procedure based on a specified schedule or a timer, which can be controlled by the user 106 via the computing device 104 or the hearable 102. The calibration procedure and the measurement procedure are further described below.

[0067] During both the calibration procedure and the measurement procedure, the speaker 408 transmits the ultrasound transmit signal 502 and the microphone 410 receives the ultrasound receive signal 504. During the calibration procedure, the ultrasound transmit signal 502 and the ultrasound receive signal 504 can have tones 602-1 to 602-M, where M represents a positive integer. The multiple tones 602-1 to 602-M can be transmitted in parallel or in series over a given time interval. In this case, the ultrasound transmit signal 502 can have a particular bandwidth on the order of several kilohertz. For example, the ultrasound transmit signal 502 can have a bandwidth of approximately 4, 5, 6, 8, 10, 16, or 20 kHz. In example implementations, theultrasound transmit signal 502 is transmitted over multiple seconds, such as 2, 3, 4, 6, or more seconds. A duration of each tone 602 can be evenly divided over a total duration of the ultrasound transmit signal 502.

[0068] In an example implementation, the ultrasound transmit signal 502 for the calibration procedure can have seven tones 602 (e.g., M equals 7). In some cases, the tones 602 are evenly distributed across an interval. For example, the tones 602 can be in 1 kHz increments between 32 kHz and 38 kHz (e.g., at approximately 32, 33, 34, 35, 36, 37, and 38 kHz). The term “approximately” means that the tones 602 can be within 5% of a given value or less (e.g., within 3%, 2%, or 1% of the given value).

[0069] An amplitude of the calibration procedure’s ultrasound transmit signal 502 can be approximately the same across the tones 602-1 to 602-M. In this manner, power is evenly distributed across each tone 602. The quantity of tones 602 (e.g., AT) can be determined based on an output power of the speaker 408. Increasing the quantity of tones 602 can increase a likelihood that the hearable 102 can support speech prediction 112 across various conditions including user wear and a physical structure of the user 106’s ear canal 116. However, an amplitude of the ultrasound transmit signal 502 can be limited across these tones 602 based on the output power of the speaker 408. Thus, the quantity of tones 602 can be optimized based on an amount of output power that is available for audioplethysmography 110.

[0070] During the measurement procedure, the ultrasound transmit signal 502 and the ultrasound receive signal 504 can have selected tones 604-1 to 604-N, where N represents a positive integer that is less than or equal to M. The selected tones 604-1 to 604-N can represent a subset (sometimes a proper subset) of the tones 602-1 to 602-M. The selected tones 604 can be transmitted in parallel or in series over a given time interval.

[0071] An amplitude of the measurement procedure’s ultrasound transmit signal 502 can be approximately the same across the selected tones 604-1 to 604-N. In this manner, power is evenly distributed across each selected tone. The amplitude of the measurement procedure's ultrasound transmit signal 502 can be higher than the amplitude of the calibration procedure’s ultrasound transmit signal 502 because the available output power is distributed across fewer tones. Additionally or alternatively, a duration of each of the selected tones 604 of the measurement procedure’s ultrasound transmit signal 502 can be longer than the duration of the tones 602 of the calibration procedure’s ultrasound transmit signal 502. The higher amplitude and / or the longer duration can further improve the signal -to-noise ratio performance of the hearable 102 for audioplethysmography 110. By using a few selected tones 604 that were determined to improvesignal-to-noise ratio performance, the measurement procedure can achieve a higher level of accuracy and sensitivity for speech prediction 112.

[0072] The analog circuit 412 performs analog-to-digital conversion to generate a digital transmit signal 606 and a digital receive signal 608 based on the ultrasound transmit signal 502 and the ultrasound receive signal 504, respectively. The pre-processing module 418 performs frequency downconversion and demodulation to generate at least one pre-processed signal 610 based on the digital transmit signal 606 and the digital receive signal 608. The pre-processing module 418 can also apply filtering to generate the pre-processed signal 610.

[0073] Optionally, as part of the calibration procedure, the calibration module 424 processes the pre-processed signal 610 to determine the selected tones 604-1 to 604-N. The selected tones 604-1 to 604-N can improve performance of audioplethysmography 110 during the measurement procedure. To determine the selected tones 604-1 to 604-N, the calibration module 424 extracts the amplitude and / or phase of the pre-processed signal 610 using the amplitude detector and the phase detector, respectively. The quality detector of the calibration module 424 measures uality metrics for each tone (or frequency) of the pre-processed signal 610 and for each of the characteristics (e.g., amplitude and / or phase). Example quality metrics can include peak-to- average ratios and / or signal-to-noise ratios. The peak-to-average ratio represents a peak intensity within a frequency range of interest divided by an average intensity within this frequency range. A higher quality metric indicates a higher-quality signal, or more generally, better performance for audioplethysmography 110.

[0074] The comparator of the calibration module 424 can evaluate the quality metrics with respect to a threshold. In an example implementation, the comparator determines the selected tones 604-1 to 604-N for a subsequent measurement procedure based on the frequencies associated with the quality metrics that are greater than or equal to a threshold. Additionally or alternatively, the comparator can evaluate the quality metrics with respect to each other. In an example implementation, the comparator determines one of the selected tones based on a frequency with the highest quality metric across the amplitude. Also, the comparator can determine one of the selected tones 604-1 to 604-N based on a frequency with the highest quality metric across the phase. In other implementations, the comparator can determine a single selected tone based on a frequency having the highest quality metric associated with either the amplitude or the phase.

[0075] In general, the calibration module 424 enables the selected tones 604-1 to 604-N to be dynamically adjusted prior to the measurement procedure based on a cunent environment, which can account for a wear of the hearable 102 (e.g., a current insertion depth and / or rotation), a physical structure of the user 106’s ear canal 116, and a response characteristic of the hearable 102(e.g., speaker, microphone, and / or housing). In this manner, the calibration module 424 can improve the signal -to-noise ratio performance of the hearable 102 for the measurement procedure. The calibration module 424 can also determine which tones 604 generate ultrasound receive signals 504 with desired characteristics for speech prediction 112. In general, the calibration procedure can be performed whether or not the user 106 is speaking.

[0076] The calibration module 424 communicates the selected tones 604-1 to 604-N to the speaker 408 using a control signal. The speaker 408 accepts the control signal that identifies the selected tones 604-1 to 604-N and can transmit a subsequent ultrasound transmit signal 502 for speech prediction 112 using the selected tones 604-1 to 604-N. With the calibration procedure, the hearable 102 can dynamically adjust the transmission frequency (e.g., one or more carrier frequencies) each time the seal 118 is formed (e.g., based on the wear of the hearable 102) and based on the unique physical structure of the ear 108. Through this calibration procedure, the hearables 102 on different ears 108 may operate with one or more different ultrasound frequencies.

[0077] As part of the measurement procedure, the speech predictor 420 can perform aspects of speech prediction 112 using the pre-processed signal 610 to generate a prediction indicator 612. In an example implementation, the prediction indicator 612 is a control signal having a first state that indicates the user 106 is about to speak (e.g., speech is predicted) and a second state that indicates the user 106 is not about to speak (e.g., speech is not predicted).

[0078] To perform speech prediction 112, the speech predictor 420 analyzes the pre-processed signal 610 and detects a significant variation in an amplitude and / or phase of the pre-processed signal 610. This variation is caused by the deformation in the ear canal 116 due to muscle movements made by the user 106 in anticipation of speaking as well as during speech.

[0079] In general, the term “significantly” can mean that the values of the amplitude and / or the phase can change by 20% or more relative to a previous value (e.g., relative to an average of a set of previous values). Additionally or alternatively, a slope of the amplitude and / or the phase can vary significantly. Sometimes the slope of the amplitude and / or the phase can change signs (e.g., from a positive slope to a negative slope, or vice versa). A magnitude of the slope of the amplitude and / or the phase can sometimes change by approximately 10% or more.

[0080] The speech predictor 420 can be implemented in a variety of ways to detect the variation in the pre-processed signal 610. In a first method, the speech predictor 420 predicts that the user 106 is about to speak based on a peak of the amplitude and / or phase exceeding a predetermined threshold. In a second method, the speech predictor 420 predicts that the user 106 is about to speak based on a change in a slope of the amplitude and / or phase exceeds a predetermined threshold. In a third method, the speech predictor 420 uses a machine-learnedmodel that is trained, using supervised learning, to detect a variation in the pre-processed signal 610 that is associated with the user 106 preparing to speak. In a fourth method, the speech predictor 420 computes a variance of the pre-processed signal 610 and predicts that the user 106 is about to speak based on the variance exceeding a predetermined threshold.

[0081] Still other example implementations of the speech predictor 420 can use a light-weight (or non-computationally complex) algorithm. For example, the speech predictor 420 can use a cumulative sum (CUSUM) to predict whether or not the user 106 is about to speak. The cumulative sum can be calculated from multiple samples of at least one of the amplitude or the phase of a signal that is derived from the ultrasound receive signal 504 (e.g., the pre-processed signal 610). Consider, for instance, that the speech predictor 420 computes a cumulative sum of the pre-processed signal 610 according to Equations 1 and 2.So= 0 Equation 15„+1= max (0, Sn+ xn+1— wn) Equation 2 where Sn+i represents the cumulative sum, xn+i represents a latest sample of the pre-processed signal 610, wnrepresents a weight that is assigned to xn+i, and Snrepresents a previous value of the cumulative sum. To predict whether or not the user 106 is about to speak, the speech predictor 420 compares the cumulative sum to a predetermined threshold. If the cumulative sum is greater than or equal to the threshold, the speech predictor 420 predicts that the user 106 is about to speak. Alternatively, if the cumulative sum is less than the threshold, the speech predictor 420 predicts that the user 106 is not about to speak (e.g., predicts that the user 106 will remain silent). The prediction indicator 612, which includes information about the prediction, is passed to the speech-processing gate 422.

[0082] The speech-processing gate 422 generates a control signal 614 based on the prediction indicator 612. The control signal 614 can have a first state to cause the speech processor 202 to operate in accordance with the low-power state 204 and can have second state to cause the speech processor 202 to operate in accordance with the high-power state 206. In some cases, the speechprocessing gate 422 can generate the control signal 614 directly based on the prediction indicator 612. If the prediction indicator 612 indicates that the user 106 is about to speak, the speech-processing gate 422 generates the control signal 614 to cause the speech processor 202 to be in the high-power state 206. Alternatively, if the prediction indicator 612 indicates that the user 106 is not about to speak, the speech-processing gate 422 generates the control signal 614 to cause the speech processor 202 to be in the low-power state 204. This allows the device with the speech processor 202 to conserve power.

[0083] In another case, the speech-processing gate 422 generates the control signal 614 based on information provided by the prediction indicator 612 and the auxiliary sensor 428. Consider an example in which the auxiliary' sensor 428 is a voice accelerometer. In this example, the speechprocessing gate 422 generates the control signal 614 to cause the speech processor 202 to be in the high-power state 206 based on the prediction indicator 612 indicating that the user 106 is about to speak or based on the auxiliary sensor 428 indicating that the user 106 is speaking. In this case, the auxiliary' sensor 428 provides a back-up option for gating speech processing for cases in which active acoustic sensing unintentionally fails to predict that the user 106 is about to speak. Also, the speech processing gate 422 generates the control signal 614 to cause the speech processor 202 to be in the low-power state 204 based on the prediction indicator 612 indicating that the user 106 is not about to speak and / or based on the auxiliary sensor 428 indicating that the user 106 is not speaking.

[0084] Speech prediction 112 and the generation of the prediction indicator 612 can occur tens to hundreds of milliseconds prior to the user 106 speaking audibly. The time it takes for the speech processor 202 to transition from the low-power state 204 to the high-power state 206 can be a fraction of this time (e.g., a few tens of milliseconds).

[0085] In FIG. 6, the calibration procedure and the measurement procedure are described as individual procedures that occur at different time intervals. In particular, the calibration procedure occurs before the measurement procedure. This enables the ultrasound transmit signal 502 for the measurement procedure to be transmitted with fewer tones than the ultrasound transmit signal 502 used for the calibration procedure, which can increase signal-to-noise ratio performance for audioplethysmography 110. In some implementations, however, the hearable 102 can have sufficient output power to perform the measurement procedure with the multiple tones 602-1 to 602-M using a single ultrasound transmit signal 502. In this case, aspects of the calibration module 424 can be integrated within the pre-processing module 418 via a frequency selector. This frequency selector can effectively pass the selected tones 604-1 to 604-N to the speech predictor 420.

[0086] FIG. 7 illustrates an example flow diagram 700 for gating speech processing 114 using active acoustic sensing. At 702, the speech predictor 420 determines whether or not speech is predicted to occur. If speech is not predicted to occur, no actions are taken at 704. Alternatively, if speech is predicted to occur, operations continue at 706.

[0087] At 706, the speech-processing gate 422 initiates, within a device that performs speech processing, a transition from a low-power state 204 associated with speech processing to a high- power state 206 associated with speech processing. For example, the speech-processing gate 422causes a speech-processing mode of the device to transition from the low-power state 204 to the high-power state 206. The device can be the hearable 102 itself or another device that is communicatively coupled to the hearable 102, such as the computing device 104. The high-power state 206 can differ from the low-power state 204 based on an amount of power consumption, the quantity and / or type of components (e.g., circuits and / or sensors) that are enabled, the type of speech processing that is performed, and / or an amount of performance associated with the speech processing.

[0088] At 708, the device performs speech processing in accordance with the high-power state 206. At 710, the hearable 102 determines whether or not the user 106 is finished speaking or is not speaking. For example, the speech-processing gate 422 can determine that the user 106 is not speaking if the prediction indicator 612 transitions from a first state that indicates speech is predicted to occur to a second state that indicates speech is not predicted to occur. Additionally or alternatively, the speech-processing gate 422 can determine that the user 106 is not speaking if the auxiliary sensor 428 indicates that audible speech is not present.

[0089] If it is determined at 710 that the user 106 is still speaking, the device can continue to perform speech processing at 708. Alternatively, if it is determined at 710 that the user 106 has finished speaking, operations continue at 712.

[0090] At 712. the speech-processing gate 422 initiates, within the device that performed speech processing, the transition from the high-power state 206 to the low-power state 204. This enables the device to conserve power and utilize enhanced speech-processing techniques.

[0091] FIG. 8 illustrates an impact of a user 106's muscle movements and speech on an ultrasound receive signal 504. More specifically, FIG. 8 depicts example amplitudes and phases of pre- processed signals 610 generated by different hearables 102-1 and 102-2. As shown below, the pressure wave caused by the muscle movements associated with speech and the spoken phrase 208 can significantly impact the amplitude and / or the phase of the pre-processed signals 610. In some instances, the change in the amplitude and / or the phase can be relative to a previous state or relative to a previous trend in the amplitude and / or the phase. The previous state can refer to values of the amplitude and / or the phase during which the user 106 does not speak.

[0092] In some implementations, the speech predictor 420 can predict that the user 106 is about to speak as well as detect that the user 106 is speaking based on the amplitude of the pre-processed signal 610 provided by the hearable 102-1, the phase of the pre-processed signal 610 provided by the hearable 102-1, the amplitude of the pre-processed signal 610 provided by the hearable 102-2, the phase of the pre-processed signal 610 provided by the hearable 102-2, or some combination thereof. Generally speaking, processing a larger quantity of signals and / or tones 604 providesmore information to the speech predictor 420. This can make it easier for the speech predictor 420 to accurately perform speech prediction 112.

[0093] Graphs 800-1 and 800-2 in FIG. 8 depict amplitudes 802 and phases 804 of pre-processed signals 610 that are respectively generated by the hearables 102-1 and 102-2. Time is depicted along the horizontal axes of the graphs 800-1 and 800-2. During the time interval indicated at 806, the user 106 speaks a phrase 208 (e.g., audibly speaks, whistles, sings, or makes other utterances). This causes the amplitude 802 and / or the phase 804 of the ultrasound receive signal 504 to change significantly relative to a previous state. Prior to time 806, as indicated by 808, the user 106 makes muscle movements in anticipation of speaking. These muscle movements can include moving their j aw. With audioplethysmography 110, the speech predictor 420 can predict that the user 106 is about to speak at 808 based on the change in the amplitude 802 and / or phase 804 of the pre- processed signals 610 provided by the hearable 102-1 and / or the hearable 102-2.Example Methods

[0094] FIGs. 9 and 10 depict example methods 900 and 1000 for implementing aspects of gating speech processing 114 using active acoustic sensing. Methods 900 and 1000 are shown as sets of operations (or acts) performed but not necessarily limited to the order or combinations in which the operations are shown herein. Further, any of one or more of the operations may be repeated, combined, reorganized, or linked to provide a wide array of additional and / or alternate methods. In portions of the following discussion, reference may be made to the environment 100 of FIG. 1, the sequence of events in FIG. 2-1, and entities detailed in FIGs. 3 and 4, reference to which is made for example only. The techniques are not limited to performance by one entity or multiple entities operating on one device.

[0095] At 902, an ultrasound transmit signal is transmitted during a first time period. The ultrasound transmit signal propagates within at least a portion of an ear canal of a person. For example, the transducer 406 (or speaker 408) of the hearable 102 transmits the ultrasound transmit signal 502. The ultrasound transmit signal 502 propagates within at least a portion of the ear canal 116 of the user 106, as described with respect to FIG. 5.

[0096] At 904, an ultrasound receive signal is received. The ultrasound receive signal represents a version of the ultrasound transmit signal with one or more characteristics modified based on the propagation within the ear canal and based on a muscle movement made by the person in preparation for speaking. For example, the transducer 406 (or the microphone 410) of the hearable 102 receives the ultrasound receive signal 504. The ultrasound receive signal 504 represents a version of the ultrasound transmit signal 502 with one or more characteristicsmodified based on the propagation within the ear canal 1 16 and based on the user 106 moving one or more muscles during at least a portion of the first time period. More specifically, the ultrasound receive signal 504 is modulated based on the deformation that occurs within the ear canal 116 as caused by the user 106 moving their muscles prior to speaking. The muscle movements are made by the user 106 in preparation for or in anticipation of speaking. The muscle movements can involve the user moving their jaw and / or tongue. The user 106 can speak the phrase 208 by audibly speaking, singing, whispering, shouting, and so forth.

[0097] The hearable 102 that receives the ultrasound receive signal 504 can be a same hearable 102 that transmitted the ultrasound transmit signal 502 (e.g., the hearable 102-1 or 102-2 in FIG. 5), or another hearable 102 that did not transmit the ultrasound transmit signal 502 (e.g., the hearable 102-2 in FIG. 5). Example waveform characteristics include amplitude, phase, and / or frequency. In some implementations, a feedback microphone of an active-noise-cancellation circuit 426 can receive the ultrasound receive signal 504.

[0098] At 906, a prediction that the person is about to speak is made based on the ultrasound receive signal. For example, the speech predictor 420 predicts that the user 106 is about to speak based on the ultrasound receive signal 504, as shown in FIG. 6. More specifically, the speech predictor 420 predicts that the user 106 is about to speak based on a change in an amplitude and / or phase of a signal that is derived from the ultrasound receive signal 504 (e.g., based on the pre- processed signal 610).

[0099] At 908, a control signal that causes a speech-processing mode to transition from a low - power state to a high-power state is generated responsive to the prediction. For example, the speech-processing gate 422 generates the control signal 614 based on at least the prediction indicator 612. The control signal 614 causes a speech-processing mode of a device (e.g., the computing device 104 or the hearable 102) to transition from operating speech processing in accordance w ith the low-power state 204 to operating speech processing in accordance with the high-power state 206, as shown in FIGs. 2-1 and 7. The low-power state 204 can be a disabled state or a state that conserves power by performing fewer functions or by using less computationally complex models. In contrast, the high-power state 206 can be an enabled state or a state that consumes a larger amount of powder by performing more functions or by using more computationally complex models.

[0100] At 1002 in FIG. 10, active acoustic sensing is perfonned to detect a pressure wave that propagates within an ear canal of a person and is associated with a muscle movement made by the person prior to the person speaking. For example, the hearable 102 performs active acoustic sensing to detect a pressure wave that propagates within an ear canal 116 of a user 106 and isassociated with a muscle movement made by the user 106 prior to the user 106 speaking. The muscle movement can be ajaw movement and / or a tongue movement. To perform active acoustic sensing, the hearable 102 transmits and receives an ultrasound signal (e.g., transmits the ultrasound transmit signal 502 and receives the ultrasound receive signal 504).

[0101] At 1004, a prediction regarding whether or not the person is about to speak is made based on the active acoustic sensing. For example, the hearable 102 performs speech prediction 1 12 based on the active acoustic sensing to predict whether or not the user 106 is about to speak. More specifically, the hearable 102 performs speech prediction 112 to detect the presence or absence of a particular muscle movement that is often made by the user 106 prior to the user 106 audibly speaking.

[0102] At 1106, speech processing is gated based on the prediction. For example, the hearable 102 gates speech processing that is performed by the computing device 104 or performed by itself. Explained another way, the hearable 102 controls a speech-processing mode of a device (e.g., the computing device 104 or the hearable 102) based on the prediction. The gating is based on the prediction regarding whether or not the user 106 is about to speak. The gating can cause speech processing to be performed in accordance with a high-power state 206 if the prediction indicates the user 106 is about to speak. If the prediction indicates the user 106 is not about to speak (e.g., indicates that the user 106 is going to remain silent) the gating can cause speech processing to be performed in accordance with a low-power state 204, as described with respect to FIG. 7.Example Computing System

[0103] FIG. 11 illustrates vanous components of an example computing system 1100 that can be implemented as any type of client, server, and / or computing device as described with reference to the previous FIGs. 3 and 4 to implement aspects of active acoustic sensing using a hearable 102. The computing system 1100 includes communication devices 1102 that enable wired and / or wireless communication of device data 1104 (e.g., received data, data that is being received, data scheduled for broadcast, or data packets of the data). The communication devices 1 102 or the computing system 1100 can include one or more hearables 102. The device data 1104 or other device content can include configuration settings of the device, media content stored on the device, and / or information associated with a user of the device. Media content stored on the computing system 1100 can include any type of audio, video, and / or image data. The computing system 1100 includes one or more data inputs 1106 via which any type of data, media content, and / or inputs can be received, such as human utterances, user-selectable inputs (explicit or implicit), messages,music, television media content, recorded video content, and any other type of audio, video, and / or image data received from any content and / or data source.

[0104] The computing system 1100 also includes communication interfaces 1108, which can be implemented as any one or more of a serial and / or parallel interface, a wireless interface, any type of network interface, a modem, and as any other type of communication interface. The communication interfaces 1 108 provide a connection and / or communication links between the computing system 1100 and a communication network by which other electronic, computing, and communication devices communicate data with the computing system 1100.

[0105] The computing system 1100 includes one or more processors 1110 (e.g., any of microprocessors, controllers, and the like), which process various computer-executable instructions to control the operation of the computing system 1100. Alternatively or in addition, the computing system 1100 can be implemented with any one or combination of hardware, firmware, or fixed logic circuitry that is implemented in connection with processing and control circuits which are generally identified at 1 112. Although not shown, the computing system 1100 can include a system bus or data transfer system that couples the various components within the device. A system bus can include any one or combination of different bus structures, such as a memory bus or memory controller, a peripheral bus, a universal serial bus, and / or a processor or local bus that utilizes any of a variety of bus architectures.

[0106] The computing system 1100 also includes a computer-readable medium 1 114, such as one or more memory devices that enable persistent and / or non-transitory data storage (i.e., in contrast to mere signal transmission), examples of which include random access memory (RAM), non-volatile memory' (e.g., any one or more of a read-only memory’ (ROM), flash memory, EPROM, EEPROM, etc.), and a disk storage device. The disk storage device may be implemented as any type of magnetic or optical storage device, such as a hard disk drive, a recordable and / or rewriteable compact disc (CD), any ty pe of a digital versatile disc (DVD), and the like. The computing system 1100 can also include a mass storage medium device (storage medium) 1116.

[0107] The computer-readable medium 1114 provides data storage mechanisms to store the device data 1 104, as well as various device applications 1118 and any other types of information and / or data related to operational aspects of the computing system 1100. For example, an operating system 1120 can be maintained as a computer application with the computer-readable medium 1114 and executed on the processors 1110. The device applications 1118 may include a device manager, such as any form of a control application, software application, signal-processing and control module, code that is native to a particular device, a hardware abstraction layer for a particular device, and so on.

[0108] The device applications 1118 also include any system components, engines, or managers to implement audioplethysmography 110 for gating speech processing 114. In this example, the device applications 1118 include the pre-processing module 418, the speech predictor 420, and the speech-processing gate 422. Although not explicitly shown, the device applications 1118 can also include the calibration module 424, the application 306, and / or the speech processor 202.

[0109] Throughout this disclosure, examples are described where a computing system 1 100 (e.g., the hearable 102, the computing device 104, a client device, a server device, a computer, or another type of computing system) may analyze information (e.g., various audible and / or ultrasound signals) associated with a user 106, for example, the phrase 208 mentioned with respect to FIG. 2-1. Further to the descriptions above, a user 106 may be provided with controls allowing the user 106 to make an election as to both if and when systems, programs, and / or features described herein may enable collection of information (e.g., information about a user 106’s social network, social actions, social activities, profession, a user 106's preferences, a user 106's current location), and if the user 106 is sent content or communications from a server. The computing system 1100 can be configured to only use the information after the computing system 1100 receives explicit permission from the user 106 to use the data. For example, in situations where the hearable 102 analyzes signals for speech prediction 112 and / or speech processing, individual users 106 may be provided with an opportunity to provide input to control whether programs or features of the computing system 1100 can collect and make use of the data. Further, individual users 106 may have constant control over what programs can or cannot do with the information.

[0110] In addition, information collected may be pre-treated in one or more ways before it is transferred, stored, or otherwise used, so that personally-identifiable infonnation is removed. For example, before the computing system 1 100 shares data with another device, a user 106’s identity may be treated so that no personally identifiable information can be determined for the user 106. Thus, the user 106 may have control over whether information is collected about the user 106 and the user 106’s device, and how such information, if collected, may be used by the computing system 1100 and / or a remote computing system.Conclusion

[0111] Although techniques using, and apparatuses including, gating speech processing using active acoustic sensing have been described in language specific to features and / or methods, it is to be understood that the subject of the appended examples is not necessarily limited to the specific features or methods described. Rather, the specific features and methods are disclosed as example implementations of gating speech processing using active acoustic sensing.

[0112] Some examples are described below.

[0113] Example 1: A method comprising: transmitting, during a first time period, an ultrasound transmit signal that propagates within at least a portion of an ear canal of a person; receiving, during the first time period, an ultrasound receive signal, the ultrasound receive signal representing a version of the ultrasound transmit signal with one or more characteristics modified based on the propagation within the ear canal and based on a muscle movement made by the person prior to speaking; predicting that the person is about to speak based on the ultrasound receive signal; and responsive to the predicting, generating a control signal that causes a speech processing mode to transition from a low-power state to a high-power state.

[0114] Example 2: The method of example 1, wherein: the transmitting of the ultrasound transmit signal comprises transmitting the ultrasound transmit signal using a hearable; the receiving of the ultrasound receive signal comprises receiving the ultrasound receive signal using the hearable; and the method further comprises: processing, using the hearable, speech made by the person based on the speech processing mode being in the high-power state.

[0115] Example 3: The method of example 1, wherein: the transmitting of the ultrasound transmit signal comprises transmitting the ultrasound transmit signal using a hearable; the receiving of the ultrasound receive signal comprises receiving the ultrasound receive signal using the hearable; and the generating of the control signal comprises generating the control signal to cause a device that is coupled to the hearable to transition from operating the speech processing mode in accordance with the low-power state to operating the speech processing mode in accordance with the high-pow er state.

[0116] Example 4: The method of any previous example, wherein power consumption associated with the speech processing mode being operated in accordance with the high-power state is at least 15% higher than power consumption associated with the speech processing mode being operated in accordance with the low-power state.

[0117] Example 5: The method of any previous example, wherein: the low-power state represents a disabled state; the high-power state represents an enabled state; and the generating of the control signal comprises generating the control signal to cause the speech processing mode to transition from the disabled state to the enabled state.

[0118] Example 6: The method of any previous example, further comprising: determining that the person is not speaking prior to predicting that the person is about to speak.

[0119] Example 7: The method of any previous example, wherein predicting that the person is about to speak comprises detecting a change in a characteristic in at least one of an amplitude or a phase of a signal that is derived from the ultrasound receive signal.

[0120] Example 8: The method of example 7, wherein the detecting of the change comprises: calculating a cumulative sum of multiple samples of at least one of the amplitude or the phase of the signal that is derived from the ultrasound receive signal; and detecting the change responsive to the cumulative sum being greater than a predetermined threshold.

[0121] Example 9: The method of any previous example, further comprising: transmitting, during a second time period that occurs after the first time period, another ultrasound transmit signal that propagates within at least the portion of the ear canal of the person; receiving, during the second time period, another ultrasound receive signal, the other ultrasound receive signal representing a version of the other ultrasound transmit signal with one or more characteristics modified based on the propagation within the ear canal; predicting that the person is not about to speak based on the other ultrasound receive signal; and responsive to the predicting, generating the control signal to cause the speech processing mode to transition from the high-power state to the low-power state.

[0122] Example 10: The method of example 9, wherein predicting that the person is not about to speak comprises predicting that the person is not about to speak based on the other ultrasound receive signal and another signal provided by an auxiliary’ sensor.

[0123] Example 11 : The method of example 10, wherein the auxiliary sensor comprises a voice accelerometer.

[0124] Example 12: The method of any previous example, wherein the muscle movement comprises a jaw movement.

[0125] Example 13: The method of any previous example, wherein the speech processing mode comprises at least one of the following: voice activity detection; voice authentication; speech recognition; conversation detection; or voice assistant service.

[0126] Example 14: A non-transitory computer-readable storage medium comprising instructions that, responsive to execution by a processor, cause a hearable to perform any one of the methods of examples 1 to 13.

[0127] Example 15: A system comprising: at least one transducer; and at least one processor, the device configured to perfonn, using the at least one transducer and the at least one processor, any one of the methods of examples 1 to 13.

[0128] Example 16: The system of example 15, further comprising: a speaker; and an active-noise-cancellation circuit comprising a feedback microphone, wherein the at least one transducer comprises the speaker and the feedback microphone.

[0129] Example 17: The system of example 15, wherein: the at least one transducer comprises a speaker and a microphone; the speaker is configured to be positioned proximate to a first ear of a person; and the microphone is configured to be positioned proximate to a second ear of the person.

[0130] Example 18: The system of any one of examples 15 to 17, wherein the device comprises at least one earbud, the at least one earbud comprising the at least one transducer and the at least one processor.

Claims

CLAIMSWhat is claimed is:

1. A method comprising: transmitting, during a first time period, an ultrasound transmit signal that propagates within at least a portion of an ear canal of a person; receiving, during the first time period, an ultrasound receive signal, the ultrasound receive signal representing a version of the ultrasound transmit signal with one or more characteristics modified based on the propagation within the ear canal and based on a muscle movement made by the person prior to speaking; predicting that the person is about to speak based on the ultrasound receive signal; and responsive to the predicting, generating a control signal that causes a speech processing mode to transition from a low-power state to a high-power state.

2. The method of claim 1, wherein: the transmitting of the ultrasound transmit signal comprises transmitting the ultrasound transmit signal using a hearable; the receiving of the ultrasound receive signal comprises receiving the ultrasound receive signal using the hearable; and the method further comprises: processing, using the hearable, speech made by the person based on the speech processing mode being in the high-power state.

3. The method of claim 1, wherein: the transmitting of the ultrasound transmit signal comprises transmitting the ultrasound transmit signal using a hearable; the receiving of the ultrasound receive signal comprises receiving the ultrasound receive signal using the hearable; and the generating of the control signal comprises generating the control signal to cause a device that is coupled to the hearable to transition from operating the speech processing mode in accordance with the low-power state to operating the speech processing mode in accordance with the high-power state.

4. The method of any previous claim, wherein power consumption associated with the speech processing mode being operated in accordance with the high-power state is at least 15% higher than power consumption associated with the speech processing mode being operated in accordance with the low-power state.

5. The method of any previous claim, wherein: the low-power state represents a disabled state; the high-power state represents an enabled state; and the generating of the control signal comprises generating the control signal to cause the speech processing mode to transition from the disabled state to the enabled state.

6. The method of any previous claim, further comprising: determining that the person is not speaking prior to predicting that the person is about to speak.

7. The method of any previous claim, wherein predicting that the person is about to speak comprises detecting a change in a characteristic in at least one of an amplitude or a phase of a signal that is derived from the ultrasound receive signal.

8. The method of claim 7, wherein the detecting of the change comprises: calculating a cumulative sum of multiple samples of at least one of the amplitude or the phase of the signal that is derived from the ultrasound receive signal; and detecting the change responsive to the cumulative sum being greater than a predetermined threshold.

9. The method of any previous claim, further comprising: transmitting, during a second time period that occurs after the first time period, another ultrasound transmit signal that propagates within at least the portion of the ear canal of the person; receiving, during the second time period, another ultrasound receive signal, the other ultrasound receive signal representing a version of the other ultrasound transmit signal with one or more characteristics modified based on the propagation within the ear canal; predicting that the person is not about to speak based on the other ultrasound receive signal; and responsive to the predicting, generating the control signal to cause the speech processing mode to transition from the high-power state to the low-power state.

10. The method of claim 9, wherein predicting that the person is not about to speak comprises predicting that the person is not about to speak based on the other ultrasound receive signal and another signal provided by an auxiliary sensor.

11. The method of claim 10, wherein the auxiliary sensor comprises a voice accelerometer.

12. The method of any previous claim, wherein the muscle movement comprises a j aw movement.

13. The method of any previous claim, wherein the speech processing mode comprises at least one of the following: voice activity detection; voice authentication; speech recognition; conversation detection; or voice assistant service.

14. A non-transitory computer-readable storage medium comprising instructions that, responsive to execution by a processor, cause a hearable to perfonn any one of the methods of claims 1 to 13.

15. A system comprising: at least one transducer; and at least one processor, the device configured to perform, using the at least one transducer and the at least one processor, any one of the methods of claims 1 to 13.

16. The system of claim 15, further comprising: a speaker; and an active-noise-cancellation circuit comprising a feedback microphone, wherein the at least one transducer comprises the speaker and the feedback microphone.

17. The system of claim 15, wherein: the at least one transducer comprises a speaker and a microphone; the speaker is configured to be positioned proximate to a first ear of a person; and the microphone is configured to be positioned proximate to a second ear of the person.

18. The system of any one of claims 15 to 17, wherein the device comprises at least one earbud, the at least one earbud comprising the at least one transducer and the at least one processor.

Citation Information

Patent Citations

  • Earbud Control Using Proximity Detection

    US20170214994A1

  • In-ear liveness detection for voice user interfaces

    US20230306971A1

  • Active acoustic sensing

    WO2023240224A1