Internal noise source filtering for active acoustic sensing
Internal noise source filtering in wireless hearables addresses interference from other operations, improving the sensitivity and accuracy of audio plethysmography for speech processing.
Patent Information
- Application Number
- JP2025501492
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2026-04-30
- Estimated Expiration
- 2043-12-29
AI Technical Summary
Existing wireless hearables face challenges in performing audio plethysmography due to interference from operations like rendering audio content, active noise cancellation, or pass-through mode, which affect the sensitivity and accuracy of speech processing.
Implementing internal noise source filtering to attenuate interference within the received ultrasonic signals, allowing hearables to perform audio plethysmography while conducting these other operations.
Enhances the sensitivity and accuracy of audio plethysmography for speech processing, including speech activity detection, speech recognition, and conversation detection.
Smart Images

Figure 0007854106000001 
Figure 0007854106000002 
Figure 0007854106000003
Abstract
Description
[Background technology]
[0001] Wireless technology has permeated daily life, allowing users easy access to communication and data. One type of wireless technology is wireless hearables, such as wireless earphones and wireless headphones. Wireless hearables allow users the freedom to go about their day while listening to music, audiobooks, podcasts, and audio content from videos. The proliferation of wireless hearables has created a market for adding additional functionality to existing hearables without changing the hardware. [Overview of the project]
[0002] This paper describes techniques and apparatus for performing internal noise source filtering for active acoustic sensing. Internal noise source filtering attenuates interference caused by other operations performed by the hearable (e.g., rendering audio content, performing active noise cancellation, or operating in pass-through mode) within the received ultrasonic signal to improve the sensitivity and accuracy of audio plethysmography. This performance improvement allows audio plethysmography to be performed while the hearable is performing these other operations. Furthermore, this performance improvement enhances the capabilities of audio plethysmography used in speech processing, which may include speech activity detection, speech recognition, and / or conversation detection.
[0003] The embodiments described below include a method for performing internal noise source filtering for active acoustic sensing. This method includes transmitting an audible signal that propagates within at least a portion of the user's ear canal during a first period. The method also includes transmitting an ultrasonic transmission signal that propagates within at least a portion of the user's ear canal during a first period. The method further includes receiving an ultrasonic reception signal during a first period. The ultrasonic reception signal represents a version of the ultrasonic transmission signal having one or more characteristics modified based on propagation within the ear canal and based on vocalizations made by the user during the first period. The received ultrasonic reception signal includes internal noise components caused by interference resulting from the rendering of the audible signal. The method further includes generating a denoised signal by filtering out the internal noise components in the received ultrasonic reception signal based on the version of the audible signal. The method also includes detecting vocalizations based on the denoised signal. In one example, the version of the audible signal may represent an electrical version and / or a digital version of the audible signal.
[0004] The embodiments described below include a computer-readable storage medium containing instructions that cause a hearable to perform one of the methods described herein in response to execution by a processor.
[0005] The embodiments described below include a device having at least one transducer and at least one processor. The device is configured to use at least one transducer and at least one processor to perform one of the methods described herein.
[0006] The embodiments described below include a system having means for performing internal noise source filtering for active acoustic sensing.
[0007] Apparatus and techniques for performing internal noise source filtering for active acoustic sensing will be described with reference to the following drawings. Throughout the drawings, the same numbers will be used to refer to similar functions and components. [Brief explanation of the drawing]
[0008] [Figure 1-1] This provides an exemplary environment in which active acoustic sensing can be implemented. [Figure 1-2] This shows exemplary geometric changes in the external auditory canal that can be detected using active acoustic sensing. [Figure 2] This provides an exemplary environment in which internal noise source filtering for active acoustic sensing can be implemented. [Figure 3] This shows an example of an internal noise source that may affect audio plethysmography. [Figure 4] This shows an example signal that can be detected using a hearable. [Figure 5] This demonstrates the exemplary operation of a hearable microphone. [Figure 6] This shows an exemplary component of a computing device. [Figure 7] An example component of a hearable is shown. [Figure 8] The following demonstrates the exemplary behavior of two hearables. [Figure 9] This document illustrates an exemplary embodiment of a hearable capable of performing internal noise source filtering for active acoustic sensing. [Figure 10] An illustrative flowchart for operating a hearable is shown. [Figure 11] An example scheme implemented by the hearable calibration module is shown. [Figure 12] An exemplary embodiment of a preprocessing module for performing an audio plethysmography is shown. [Figure 13]An exemplary embodiment of a measurement module for performing internal noise source filtering is shown. [Figure 14] An exemplary embodiment of an internal noise source filter stage for performing internal noise source filtering is shown. [Figure 15] An exemplary embodiment of an external noise source filter stage is shown. [Figure 16] An exemplary method for performing internal noise source filtering is shown. [Figure 17] Another exemplary method for performing internal noise source filtering is shown. [Figure 18] An exemplary computing system that may implement a technique for embodying or enabling the use of internal noise source filtering for active acoustic sensing is shown.
Best Mode for Carrying Out the Invention
[0009] As electronic devices become more widespread, users will incorporate them into their daily lives. Users can use electronic devices, for example, to obtain daily weather and traffic information, control the temperature of their homes, respond to doorbells, turn on / off lighting, and / or play background music. However, interactions with some electronic devices can be cumbersome and inefficient. For example, an electronic device can have a physical user interface, which may require the user to navigate one or more prompts by physically touching the electronic device. In this case, the user must divert their attention from other primary tasks in order to interact with the electronic device, which is therefore inconvenient and may cause confusion.
[0010] To address this problem, some electronic devices support voice control, which enables users to interact with the electronic device in a non-physical, less cognitively demanding way compared to other interfaces that require physical contact and / or the user's visual attention. With voice control, the electronic device seamlessly exists within the surrounding environment, providing access to information and services while the user is performing major tasks such as cooking, cleaning, driving, conversing with people, or reading a book.
[0011] Voice control can provide a convenient means of interacting with an electronic device, but there are several issues associated with voice control. For example, in a noisy environment, the user's voice may not be perceived due to other external noises. As a result, it can be difficult for voice control to detect and / or recognize the voice commands spoken by the user. Also, due to a noisy environment, voice control may accidentally respond to the voice of another person who is not authorized to use the electronic device.
[0012] Some devices address these issues by integrating a voice accelerometer (VA) into the earphone. The voice accelerometer can detect when the user is speaking based on the sound transmitted by bone conduction. However, voice accelerometers can be bulky and expensive. In some cases, it may be desirable to design a smaller and wearable form factor to improve aesthetics and reduce interference. As space becomes limited, it can be difficult to integrate additional components such as a voice accelerometer within the wearable. With the popularity of wearables, there is a market for adding additional functionality to existing wearables without changing the hardware.
[0013] According to one or more preferred embodiments, a hearable, such as an earphone, is provided that can perform a novel physiological monitoring process referred herein as audioplethysmography. Audioplethysmography is an active acoustic method that can sense subtle physiologically relevant changes observable in the user's outer and middle ear. Instead of relying on other auxiliary sensors, such as optical or electrical sensors, audioplethysmography involves sending and receiving acoustic signals that propagate at least partially within the user's ear canal. To perform audioplethysmography, the hearable forms at least a partial seal inside or around the user's outer ear. This seal allows for the formation of an acoustic circuit, including the seal, the hearable, the ear canal, and the eardrum. By sending and receiving acoustic signals, the hearable can recognize changes in the acoustic circuit and detect and / or recognize vocalizations made by the user. Vocalizations can include any sound produced using the user's lungs, vocal cords, and / or mouth. Exemplary types of vocalization may include the user speaking, whispering, calling, humming, whistling, singing, or making other statements. Hearables can utilize audio plethysmography to support a variety of different features, including voice user interfaces (VUI) and / or multi-factor voice authentication.
[0014] Some hearables can offer features other than audio plethysmography. Exceptional features may include rendering (e.g., playback) of audible content, providing active noise cancellation, and / or operating in pass-through mode. Pass-through mode allows sounds from the external environment to be rendered for the user. In this way, pass-through mode enables the hearable to function as a hearing aid. While active noise cancellation can significantly attenuate sounds from the external environment, pass-through mode allows sounds from the external environment to reach the user's ears. In some cases, pass-through mode can further amplify these sounds to assist hearing impairments.
[0015] However, providing these other features while also performing audio plethysmography can be challenging. More specifically, the signals generated for rendering audio content, providing active noise cancellation, and / or operation in pass-through mode may interfere with the ultrasonic signals received for audio plethysmography. This interference can make it difficult for hearables to use audio plethysmography for speech processing.
[0016] To address this challenge, techniques and apparatus for performing internal noise source filtering for active acoustic sensing are described. Internal noise source filtering attenuates interference caused by the hearable performing other operations (e.g., rendering audio content, performing active noise cancellation, or operating in pass-through mode) within the received ultrasonic signal to improve the sensitivity and accuracy of audio plethysmography. This performance improvement makes it possible to perform audio plethysmography while the hearable is performing these other operations. Furthermore, this performance improvement enhances the capabilities of audio plethysmography used in speech processing, which may include speech activity detection, speech recognition, and / or conversation detection.
[0017] Operating environment Figure 1-1 is a diagram of an exemplary environment 100 in which active acoustic sensing can be implemented. In the exemplary environment 100, a hearable 102 is connected to a computing device 104 using a physical or wireless interface. The hearable 102 is a device that can play audible content provided by the computing device 104 and direct the audible content to the ears 108 of a user 106. In this example, the hearable 102 operates in conjunction with the computing device 104. In other examples, the hearable 102 may operate or be implemented as a standalone device. Although shown as a smartphone, the computing device 104 may include other types of devices, including those described with respect to Figure 6.
[0018] The hearable 102 can perform audio plethysmography 110, which is a method of sensing active acoustics performed in the ear 108. The hearable 102 can perform this sensing without using other auxiliary sensors such as optical or electrical sensors. Through audio plethysmography 110, the hearable 102 can utilize speech processing to detect and / or recognize utterances made by the user 106. Exemplary types of speech processing may include speech activity detection 112, speech recognition 114, and / or conversation detection 116.
[0019] In voice activity detection 112, the hearable 102 detects vocalizations made by user 106. To perform voice activity detection 112, the hearable 102 uses audio plethysmography 110 to detect minute pressure waves propagating in user 106's ear canal 118. Pressure waves can be generated by movements related to user 106's jaw, vibrations related to vocalization, or any combination thereof. Generally speaking, these pressure waves modify the characteristics of ultrasonic signals transmitted and received by the hearable 102 and propagated through the ear canal 118. Voice activity detection 112 allows the hearable 102 to enhance the performance of a voice user interface or provide multi-factor authentication.
[0020] In speech recognition 114, the hearable 102 recognizes words spoken by the user 106. The hearable 102 can perform speech recognition 114 using audio plethysmography 110, or by using audio plethysmography 110 to augment the information provided by another signal, such as an audible signal captured by the hearable 102. Speech recognition 114 can also be used to enhance the performance of a voice user interface or to enhance multi-factor authentication.
[0021] In conversation detection 116, the hearable 102 determines whether user 106 is conversing with another person. To perform conversation detection 116, the hearable 102 uses audio plethysmography 110 to filter the audible signals captured by the hearable 102. Since the signals associated with audio plethysmography 110 are not correlated with utterances from another person, audio plethysmography 110 can be used to filter and / or attenuate common signal components between the audible signals and the signals used for audio plethysmography 110. This filtering improves the signal-to-noise ratio to detect utterances made by someone other than user 106.
[0022] In response to detecting the occurrence of a conversation, the conversation detection 116 may be used to control the operation of the hearable 102 and / or computing device 104. In one example, the audio content may be paused or resumed based on whether a conversation has been detected. In another example, the volume of the audio content may be decreased or increased based on whether a conversation has been detected. In yet another example, based on the conversation detection 116 detecting a conversation, active noise cancellation may be temporarily stopped, and the hearable 102 may operate according to the transparency mode. Alternatively, based on the conversation detection 116 detecting that there is no conversation, active noise cancellation may be resumed, and the hearable 102 may exit the transparency mode.
[0023] To use audio plethysmography 110, user 106 positions the hearable 102 such that it creates at least a partial seal 120 around or inside the ear 108. Several parts of the ear 108, including the external auditory canal 118 and the ear drum 122 (or tympanic membrane), are shown in Figure 1-1. The seal 120 causes the hearable 102, the external auditory canal 118, and the ear drum 122 to connect to each other and form an acoustic circuit. Audio plethysmography 110 involves measuring, at least partially, properties associated with this acoustic circuit. The properties of the acoustic circuit can change due to a variety of different situations or actions.
[0024] For example, consider Figure 1-2, where a change occurs in the physical structure of the ear 108. Exemplary changes in physical structure include changes in the geometric shape of the external auditory canal 118 and / or changes in the volume of the external auditory canal 118. This change may be caused, at least in part, by bone conduction and / or pressure waves associated with vocalization performed by the user 106.
[0025] At 124, for example, the tissues surrounding the external auditory canal 118 and the eardrum 122 itself are slightly “compressed” by bone conduction and / or pressure waves. This compression slightly reduces the volume of the external auditory canal 118 at 124. However, at 126, the compression weakens, and the volume of the external auditory canal 118 increases slightly compared to 124. Physical changes within the ear 108 can modulate the amplitude and / or phase of the ultrasonic signal propagating through the external auditory canal 118, as will be further explained with respect to Figure 8. The audioplethysmography technique 110 can be performed while the hearable 102 is playing audible content to the user 106, as will be further explained with respect to Figures 2, 3-1, and 3-2.
[0026] Figure 2 shows exemplary environments 200-1, 200-2, and 200-3 in which internal noise source filtering 202 can be implemented for active acoustic sensing. Environments 200-1, 200-2, and 200-3 illustrate exemplary features of a hearable 102 that may be active during audio plethysmography 110. In environment 200-1, the hearable 102 renders audio content 204 to the user 106. The audio content 204 is any type of sound produced by the hearable 102, such as music, a ringtone, an alarm, or a caller's voice.
[0027] In environment 200-2, the hearable 102 provides active noise cancellation 206. Active noise cancellation 206 significantly attenuates other sounds present in the external environment, such as speech 208, music 210, and / or external noise 212. Active noise cancellation 206 can create a quiet environment for the user 106 to concentrate. Furthermore, or alternatively, active noise cancellation 206 can make audio content 204 easier for the user 106 to hear. This audio content 204 may also be played for the user 106 within environment 200-2.
[0028] In environment 200-3, the hearable 102 operates according to transparency mode 214. Transparency mode allows sounds from the external environment to be rendered for user 106. Rendering these sounds may include amplifying them to assist hearing impairments. In environment 200-3, transparency mode 214 allows user 106 to hear speech 216 from another person (for example, hearing another person speaking).
[0029] While the hearable 102 renders audio content 204, provides active noise cancellation 206, and / or operates according to the transmission mode 214, internal signals generated by the hearable 102 during any of these operations may interfere with the received ultrasonic signal. This interference may make it difficult for the audio plethysmography 110 to be used for speech processing. To address this problem, the hearable 102 performs internal noise source filtering 202. Exemplary types of internal noise sources that can be filtered using internal noise source filtering 202 are further described with reference to Figure 3.
[0030] Figure 3 shows exemplary components of the hearable 102. In the illustrated configuration, the hearable 102 includes at least one microphone 302 and circuitry 304. Microphone 302 can be used to receive ultrasonic signals for audio plethysmography 110 and / or audible signals for other operations of the hearable 102. Circuitry 304 may provide any of the features described with respect to Figure 2. For example, circuitry 304 may render audio content 204, perform active noise cancellation 206, and / or support a pass-through mode 214.
[0031] During operation, the microphone 302 receives a signal represented by the received signal 306. The circuit 304 can generate one or more signals while the microphone 302 is generating the received signal 306 and / or while the received signal 306 is being processed by the hearable 102. The hearable 102 renders (e.g., plays, emits, or transmits) the signals generated by the circuit 304. The rendered signals propagate through the ear canal 118 of the user 106.
[0032] From the perspective of any technique for processing the received signal 306, the signal generated by circuit 304 is a noise signal 308 (or self-generated noise signal 308) that may interfere with the received signal 306. For example, the noise signal 308 may affect the waveform characteristics of the received signal 306 (e.g., amplitude, frequency, and / or phase). This modification to the waveform characteristics of the received signal 306 is represented by an internal noise component 310. In some cases, the internal noise component 310 is caused by the noise signal 308 being modulated with respect to the received signal 306, or by the noise signal 308 being mixed with the received signal 306. This interference between the noise signal 308 and the received signal 306 may be due to nonlinearity in the microphone 302, mixing operation performed by the hearable 102, or some other component and / or operation of the hearable 102. This internal noise component 310 may make it difficult for the audio plethysmography 110 to be used for speech processing.
[0033] Generally speaking, circuit 304 represents an internal noise source 312. An exemplary internal noise source 312 is implemented within the hearable 102 (for example, within the housing of the hearable 102) and may include any type of circuit whose operation may interfere with the signal received for the audio plethysmography 110. In Figure 3, the internal noise source 312 may include a speaker 314, an active noise cancellation (ANC) circuit 316 (ANC circuit 316), a pass-through mode circuit 318, or any combination thereof.
[0034] Speaker 314 can be used to render audio content 204 within environment 200-1. To render audio content 204, speaker 314 generates an audible signal 320 which can represent a noise signal 308. The audible signal 320 can be associated with music 322, a human voice, or other types of sound. The audible signal 320 may also be called an audible frequency spectrum signal and generally contains frequencies between 20 hertz (Hz) and 20 kilohertz (kHz).
[0035] The active noise cancellation circuit 316 performs active noise cancellation 206 as shown in environment 200-2. To perform active noise cancellation 206, the active noise cancellation circuit 316 generates an anti-noise signal 324. The anti-noise signal 324 represents another type of noise signal 308, which may interfere with the received signal 306 and contribute to the internal noise component 310.
[0036] The pass-through mode circuit 318 allows the hearable 102 to operate according to the pass-through mode, as shown in environment 200-3. In pass-through mode, the pass-through mode circuit 318 generates a pass-through mode signal 326 that includes sound from the external environment. The pass-through mode signal 326 represents another type of noise signal 308, which may interfere with the received signal 306 and contribute to the internal noise component 310.
[0037] Designing the hearable 102 to mitigate interference caused by the noise signal 308 can be challenging. Cost and / or size constraints of the hearable 102 may make it difficult to incorporate any interference shielding techniques and / or utilize more expensive components that are less likely to generate this interference. To address this challenge, the internal noise source filtering technique 202 filters (e.g., attenuates) the internal noise component 310 in the received signal 306 to enhance the audio plethysmography 110 and / or speech processing. Exemplary techniques for utilizing audio plethysmography 110 for speech processing are further described with reference to Figures 4 and 5.
[0038] Figure 4 shows exemplary signals that can be detected by the hearable 102. In environment 400, the hearable 102 is worn by user 106. During operation, the hearable 102 can receive a variety of different signals. In one embodiment, the hearable 102 receives at least one over-the-air (OTA) signal 402 (OTA signal 402). The OTA signal 402 may include a voice component 404 and / or a noise component 406. The voice component 404 may include utterances made by user 106, such as a voiceprint phrase to activate a voice user interface. The noise component 406 may be any undesirable audible sound that may mask or interfere with the detection of utterances. The noise component 406 may represent the ambient noise 212, music 210, and / or speech 208 in Figure 2. Generally, the OTA signal 402 includes audible frequencies.
[0039] In another embodiment, the hearable 102 receives at least one bone conduction signal 408. The bone conduction signal 408 represents sound transmitted to the user's ear 108 via bone conduction. The bone conduction signal 408 also includes a speech component 404. In most situations, the bone conduction signal 408 does not include a noise component 406.
[0040] To perform active acoustic sensing, the hearable 102 transmits and receives at least one ultrasonic signal 410. The ultrasonic signal 410 propagates within the ear canal 118. The user's 106 vocalizations may alter the physical structure of the ear 108. Therefore, the ultrasonic signal 410 may also contain a speech component 404.
[0041] The ultrasonic signal 410 may not directly contain the noise component 406. However, some designs of the hearable 102, as further described below, may have the noise component 406 associated with the radio signal 402 interfering with the detection of the voice component 404 in the ultrasonic signal 410.
[0042] In some exemplary embodiments, the hearable 102 may be designed to minimize interference between the radio signal 402 and the ultrasonic signal 410. The hearable 102 may receive these signals using, for example, different microphones. In this case, the hearable 102 can directly process the ultrasonic signal 410 to detect speech. In other exemplary embodiments, the hearable 102 receives both the radio signal 402 and the ultrasonic signal 410 using the same microphone. This may be beneficial to meet the size constraints of the hearable 102, but it may result in a version of the noise component 406 being present in the ultrasonic signal 410. When the hearable 102 receives the radio signal 402, the bone conduction signal 408, and the ultrasonic signal 410, these signals may interact to make it difficult to detect the speech component 404 using audio plethysmography 110. Specifically, the radio signal 402 may be modulated with respect to the ultrasonic signal 410 or mixed with the ultrasonic signal 410, as further described with respect to Figure 5.
[0043] Figure 5 shows exemplary operation of the microphone 302 of the hearable 102. During operation, the microphone 302 receives radio signals 402, bone conduction signals 408, and ultrasonic signals 410. The microphone 302 may include filter modules that can generate separate signals related to different frequency ranges. In this example, the microphone 302 generates received audible signals 504 and received ultrasonic signals 506. These are further shown in graph 508 at the bottom of Figure 5. These signals can be downconverted to baseband frequencies. Generally, the received audible signals 504 and received ultrasonic signals 506 represent electrical signals that may be processed by other components of the hearable 102.
[0044] Graph 508 shows that the received audible signal 504 and the received ultrasonic signal 506 contain different frequencies. The received audible signal 504 may include frequencies associated with the audible frequency spectrum (e.g., frequencies between approximately 20 Hz and 20 kHz). Therefore, the received audible signal 504 is sometimes also called the received audible frequency spectrum signal. On the other hand, the received ultrasonic signal 506 may include frequencies associated with the ultrasonic frequency spectrum (e.g., frequencies between approximately 20 kHz and 2 megahertz (MHz)).
[0045] The received audible signal 504 represents the convolution of the radio signal 402 and the bone conduction signal 408. Therefore, the received audible signal 504 includes a speech component 404 and a noise component 406 related to vocalization. The speech component 404 is provided by both the radio signal 402 and the bone conduction signal 408. The received audible signal 504 can be expressed by Equation 1. Y RAS =h BC · S+h OTA (S+N) Equation 1 Here, Y RAS This represents the received audible signal 504, and h BC represents bone conduction channels, S represents vocalization, h OTA represents the radio channel, and N represents noise (e.g., external noise 212, music 210, and / or speech 208). BC·The term S represents the bone conduction component 512. h OTA ·The term S represents the voice component 404 of the received audible signal 504. h OTA ·The term N represents the noise component 406 of the received audible signal 504.
[0046] The received ultrasonic signal 506 represents the convolution of the ultrasonic signal 410 and a modulated version of the received audible signal 504 represented by the modulation component 510. The modulation component 510 can be caused by the interaction of the wireless signal 402 and the ultrasonic signal 410 of the bone conduction signal 408 when the microphone 302 receives these signals. More specifically, the modulation component 510 may be generated based on the design of the wearable 102, non - linearities within one or more components of the wearable 102, and / or the operation of the wearable 102 (e.g., hybrid operation), and may be due to interference, inter - modulation distortion, and / or harmonics. Generally speaking, the modulation component 510 represents a version of the received audible signal 504 shifted to the ultrasonic frequency. The modulation component 510 is linearly modulated with respect to the ultrasonic frequency spectrum. The received ultrasonic signal 506 can be represented by Equation 2. Y RUS =h US ·S + h MC ·Y RAS Equation 2 Here, Y RUS represents the received ultrasonic signal 506, h US represents the ultrasonic channel, S represents the voice, h MC represents the modulation channel, Y RAS represents the received audible signal 504. h US ·The term S represents the voice component 404 of the received ultrasonic signal 506. h MC ·Y RAS The term represents the modulation component 510. Due to the modulation channel, the received ultrasonic signal 506 includes a linearly modulated version of the received audible signal 504. Thus, the noise component 406 within the modulated version of the received audible signal 504 may make it difficult to directly detect the voice component 404 within the received ultrasonic signal 506.
[0047] Generally speaking, the speech component 404 is superimposed on the ultrasonic signal 410 and correlates with speech. The speech component 404 is not correlated with the noise component 406. The speech component 404 modulates the received ultrasonic signal 506 in a different way than the modulation component 510 due to bone conduction and changes in the physical structure within the ear 108.
[0048] The speech component 404 of the received ultrasonic signal 506 is frequency-dependent. In other words, different ultrasonic frequencies modulate the speech differently. However, the bone conduction component 512 does not have frequency selectivity. In other words, the bone conduction component 512 is associated with a fixed channel. The techniques of speech activity detection 112, speech recognition 114, and / or conversation detection 116 utilize the received audible signal 504 to extract the speech component 404 from the received ultrasonic signal 506, as will be further explained with reference to Figures 13 and 14.
[0049] Figure 6 shows an exemplary embodiment of computing device 104. Computing device 104 is shown with a variety of non-limiting exemplary devices, including a desktop computer 104-1, a tablet 104-2, a laptop 104-3, a television 104-4, a computing clock 104-5, computing glasses 104-6, a gaming system 104-7, a microwave oven 104-8, and a vehicle 104-9. Other devices may also be used, such as augmented reality and / or virtual reality headsets, home service devices, smart speakers, smart thermostats, baby monitors, Wi-Fi® routers, drones, trackpads, drawing pads, netbooks, e-readers, home automation and control systems, wall displays, and other household appliances. Note that computing device 104 may be wearable, non-wearable but mobile, or relatively stationary (e.g., desktops and appliances).
[0050] The computing device 104 includes one or more computer processors 602 and at least one computer-readable medium 604 including a memory medium and a storage medium. Applications and / or operating systems (not shown) embodied as computer-readable instructions on the computer-readable medium 604 may be executed by the computer processors 602 to provide some of the functionality described herein. The computer-readable medium 604 may optionally include an application 606, a voice user interface 608, and / or a voice authenticator 610. The application 606 can perform actions using information provided by the hearable 102. Exemplary actions may include displaying data related to an audio plethysmography 110 to the user 106. In the case of voice activity detection 112, the application 606 may indicate whether or not a voice has been detected. The voice user interface 608 may enable the user 106 to control the computing device 104 via voice commands. The voice authenticator 610 can authenticate user 106, and if authentication is successful, it can enable the use of the voice user interface 608. The application 606, the voice user interface 608, and / or the voice authenticator 610 can improve the performance of the computing device 104 and / or enhance security by utilizing the voice activity detection 112 and / or speech recognition 114.
[0051] The computing device 104 may also include a network interface 612 for communicating data via a wired, wireless, or optical network. For example, the network interface 612 may communicate data via a local area network (LAN), wireless local area network (WLAN), personal area network (PAN), wide area network (WAN), intranet, internet, peer-to-peer network, point-to-point network, mesh network, Bluetooth®, etc. The computing device 104 may also include a display 614. Although not expressly shown, the hearable 102 may be integrated into the computing device 104 or may be physically or wirelessly connected to the computing device 104. The hearable 102 will be further described with reference to Figure 7.
[0052] Figure 7 shows an exemplary hearable 102. The hearable 102 is shown with a variety of non-limiting exemplary devices, including wireless earphones 702-1, wired earphones 702-2, and headphones 702-3. Earphones 702-1 and 702-2 are in-ear devices that fit within the ear canal 118. Each earphone 702-1 or 702-2 can represent a hearable 102. Headphones 702-3 can be placed on or covering the ears 108. Headphones 702-3 can represent closed-back headphones, open-back headphones, on-ear headphones, or over-ear headphones. Each headphone 702-3 contains two hearables 102 that are physically packaged together. Generally, there is one hearable 102 for each ear 108.
[0053] The hearable 102 includes a communication interface 704 for communicating with the computing device 104, although this is not necessary when the hearable 102 is integrated within the computing device 104. The communication interface 704 can be a wired or wireless interface, and audio content is passed from the computing device 104 to the hearable 102. The hearable 102 can also use the communication interface 704 to pass data associated with the audio plethysmography 110 to the computing device 104. Generally, the data provided by the communication interface 704 is in a format usable by the application 606, the voice user interface 608, and / or the voice authenticator 610.
[0054] The communication interface 704 also allows one hearable 102 to communicate with another hearable 102. During bistatic sensing, for example, hearable 102 can use the communication interface 704 to coordinate with other hearables 102 to support binaural audio plethysmography 110, as further described with respect to Figure 8. Specifically, transmitting hearable 102 can communicate timing and waveform information to receiving hearable 102, enabling receiving hearable 102 to properly demodulate the received ultrasonic signal 506.
[0055] The hearable 102 includes at least one transducer 706 capable of converting electrical signals into sound waves. The transducer 706 can also detect sound waves and convert them into electrical signals. These sound waves may include ultrasonic frequencies and / or audible frequencies, either of which may be used for audio plethysmography 110. Specifically, the frequency spectrum (e.g., frequency range) used by the transducer 706 to generate acoustic signals may include frequencies from the low end of the audible range to the high end of the ultrasonic range, for example, frequencies between 20 Hz and 2 megahertz (MHz). The frequency spectra of other examples of audio plethysmography 110 may include frequencies between 20 Hz and 20 kilohertz (kHz), between 20 kHz and 2 MHz, between 20 and 96 kHz, between 20 and 60 kHz, or between 30 and 40 kHz.
[0056] In an exemplary embodiment, the transducer 706 has a monostatic topology. This topology allows the transducer 706 to convert electrical signals into sound waves and sound waves into electrical signals (for example, to transmit or receive acoustic signals). Exemplary monostatic transducers may include piezoelectric transducers, capacitive transducers, and micromachined ultrasonic transducers (MUTs) using micro-electromechanical systems (MEMS) technology.
[0057] Alternatively, the transducer 706 may be implemented in a bistatic topology including multiple physically separated transducers. In this case, the first transducer converts an electrical signal into a sound wave (e.g., transmits an acoustic signal), and the second transducer converts a sound wave into an electrical signal (e.g., receives an acoustic signal). An exemplary bistatic topology can be implemented using at least one speaker 708 and at least one microphone 710. The speaker 708 and microphone 710 may be dedicated to the audio plethysmography 110, or they may be used for both the audio plethysmography 110 and other functions of the computing device 104 (e.g., presenting audible content to user 106, making phone calls, or capturing user 106's voice for voice control). Speaker 708 may represent speaker 314 in Figure 3. Microphone 710 may represent microphone 302 in Figures 3 and 5.
[0058] Generally, the speaker 708 and microphone 710 are directed toward the ear canal 118 (e.g., oriented toward the ear canal 118). Thus, the speaker 708 can transmit an ultrasonic signal toward the ear canal 118, and the microphone 710 responds to receiving the ultrasonic signal from a direction associated with the ear canal 118. In some cases, the hearable 102 includes another microphone 710 directed toward the external environment away from the ear canal 118 (e.g., oriented toward away from the ear canal 118). This other microphone can be used to receive the radio signal 402.
[0059] The hearable 102 includes at least one analog circuit 712, which includes circuits and logic for adjusting electrical signals within the analog domain. The analog circuit 712 may include analog-to-digital converters, digital-to-analog converters, amplifiers, filters, mixers, and switches for generating and modifying electrical signals. In some embodiments, the analog circuit 712 includes other hardware circuits associated with the speaker 708 or microphone 710.
[0060] The hearable 102 also includes at least one system processor 714 and at least one system medium 716 (e.g., one or more computer-readable storage media). In the illustrated configuration, the system medium 716 includes a preprocessing module 718 and a measurement module 720. The system medium 716 also optionally includes a calibration module 722. The preprocessing module 718, the measurement module 720, and the calibration module 722 can be implemented using hardware, software, firmware, or a combination thereof. In this example, the system processor 714 implements the preprocessing module 718, the measurement module 720, and the calibration module 722. In an alternative example, the computer processor 602 of the computing device 104 may implement at least a portion of the preprocessing module 718, the measurement module 720, and / or the calibration module 722. In this case, the hearable 102 can communicate digital samples of acoustic signals to the computing device 104 using a communication interface 704.
[0061] The operation of the preprocessing module 718, the measurement module 720, and the calibration module 722 will be further described with reference to Figures 11 and 12. Embodiments of internal noise source filtering 202 can be performed, at least partially, by the measurement module 720, as will be further described with reference to Figures 13 and 14. The measurement module 720 can also perform embodiments of voice activity detection 112, speech recognition 114, and / or conversation detection 116 using active acoustic sensing, as will be further described with reference to Figure 13.
[0062] Some hearables 102 include an active noise cancellation circuit 316, which enables the hearable 102 to reduce background noise or ambient noise. In this case, the microphone 710 used for audio plethysmography 110 can be implemented using a feedback microphone for the active noise cancellation circuit 316. During active noise cancellation, the feedback microphone provides feedback information regarding the execution of active noise cancellation 206. During audio plethysmography 110, the feedback microphone receives an ultrasonic signal 410, which is provided to a pre-processing module 718. In some situations, active noise cancellation 206 and audio plethysmography 110 are performed simultaneously using a feedback microphone. In this case, the ultrasonic signal 410 received by the feedback microphone can be provided to the pre-processing module 718, and the feedback signal for active noise cancellation 206 can be provided to the active noise cancellation circuit 316. Other embodiments are also possible in which microphone 710 is implemented using a feedforward microphone of the active noise cancellation circuit 316.
[0063] Furthermore, some hearables 102 include a pass-through mode circuit 318, which enables the hearable 102 to amplify background or ambient sounds. In some embodiments, the pass-through mode circuit 318 may include at least one amplifier. The microphone 710 used in the audio plethysmography 110 can also be used in the pass-through mode 214. More specifically, the ultrasonic signal 410 received by the microphone 710 may be provided to a pre-processing module 718, and the sensed signal for the pass-through mode 214 may be provided to the pass-through mode circuit 318. The active noise cancellation circuit 316, the pass-through mode circuit 318, and / or speaker 708 represent exemplary types of circuitry 304 that may unintentionally act as internal noise sources 312 during the audio plethysmography 110.
[0064] Although not explicitly shown in Figure 7, the system medium 716 may also include a voice user interface 608 and / or a voice authenticator 610. In this case, the voice user interface 608 allows user 106 to control the operation of the hearable 102 using voice control. The voice authenticator 610 can authenticate user 106 and activate the voice user interface 608 of the hearable 102. Different types of audio plethysmography 110 will be further described with reference to Figure 8.
[0065] Active acoustic sensing Figure 8 shows exemplary operation of two hearables 102-1 and 102-2. In the first exemplary operation, hearables 102-1 and 102-2 perform unilateral audio plethysmography 110. This means that hearables 102-1 and 102-2 independently perform audio plethysmography 110 on different ears 108 of user 106. In this case, the first hearable 102-1 is positioned near user 106's right ear 108, and the second hearable 102-2 is positioned near user 106's left ear 108. Each hearable 102-1 and 102-2 includes a speaker 708 and a microphone 710. Hearables 102-1 and 102-2 can operate monostatically for the same period or for different periods. In other words, each hearable 102-1 and 102-2 can independently transmit and receive ultrasonic signals.
[0066] For example, the first hearable 102-1 transmits a first ultrasonic transmission signal 802-1 using a speaker 708, which propagates within at least a portion of the user's 106 external auditory canal 118. The first hearable 102-1 receives a first ultrasonic reception signal 804-1 using a microphone 710. The first ultrasonic reception signal 804-1 represents a version of the first ultrasonic transmission signal 802-1 that has been at least partially modified by an acoustic circuit associated with the right external auditory canal 118. This modification may alter the amplitude, phase, and / or frequency of the first ultrasonic reception signal 804-1 with respect to the first ultrasonic transmission signal 802-1.
[0067] Similarly, the second hearable 102-2 transmits a second ultrasonic transmission signal 802-2 using speaker 708, which propagates within at least a portion of the left external auditory canal 118 of user 106. The second hearable 102-2 receives a second ultrasonic reception signal 804-2 using microphone 710. The second ultrasonic reception signal 804-2 represents a version of the second ultrasonic transmission signal 802-2 modified by an acoustic circuit associated with the left external auditory canal 118. This modification may alter the amplitude, phase, and / or frequency of the second ultrasonic reception signal 804-2 with respect to the second ultrasonic transmission signal 802-2.
[0068] The technique of unilateral audio plethysmography 110 can be particularly beneficial because it allows the computing device 104 to compile information from both hearables 102-1 and 102-2, thereby further increasing the reliability of the measurement. In some embodiments of audio plethysmography 110, it may be beneficial to analyze the acoustic channels between the two ears 108, as will be further described below.
[0069] In a second exemplary operation, two hearables 102-1 and 102-2 perform binaural audio plethysmography 110. This means that hearables 102-1 and 102-2 jointly perform audio plethysmography 110 across the two ears 108 of user 106. In this case, at least one of the hearables 102 (e.g., the first hearable 102-1) includes a speaker 708, and at least one of the other hearables 102 (e.g., the second hearable 102-2) includes a microphone 710. Hearables 102-1 and 102-2 operate bistatically in cooperation during the same period.
[0070] During operation, the first hearable 102-1 transmits a third ultrasonic transmission signal 802-3 using the speaker 708. The third ultrasonic transmission signal 802-3 propagates through the user 106's right ear canal 118. The third ultrasonic transmission signal 802-3 also propagates through the acoustic channel present between the right ear 108 and the left ear 108. In the left ear 108, the third ultrasonic transmission signal 802-3 propagates through the user 106's left ear canal 118 and is represented as the third ultrasonic reception signal 804-3. The second hearable 102-2 receives the third ultrasonic reception signal 804-3 using the microphone 710. The third ultrasonic received signal 804-3 represents a version of the third ultrasonic transmitted signal 802-3 that is modified by an acoustic circuit associated with the right external auditory canal 118, modified by an acoustic channel associated with the user's face 106, and modified by an acoustic circuit associated with the left external auditory canal 118. This modification may alter the amplitude, phase, and / or frequency of the third ultrasonic received signal 804-3 with respect to the third ultrasonic transmitted signal 802-3. In some cases, the hearable 102-2 measures the time of flight (ToF) associated with propagation from the first hearable 102-1 to the second hearable 102-2. In some cases, a combination of unilateral and bilateral audio plethysmography 110 is applied to further improve the reliability of the measurement.
[0071] The ultrasonic transmission signal 802 in Figure 8 can represent various different types of signals as described above with respect to Figure 7. In an exemplary embodiment, the ultrasonic transmission signal 802 may be the ultrasonic signal 410 in Figures 4 and 5. The ultrasonic transmission signal 802 may also be a continuous wave signal (e.g., a sine wave signal) or a pulsed signal. Some ultrasonic transmission signals 802 may have a specific tone (or frequency). Other ultrasonic transmission signals 802 may have multiple tones (or multiple frequencies). Various modulations can be applied to generate the ultrasonic transmission signal 802. Exemplary modulations include linear frequency modulation, triangular frequency modulation, stepped frequency modulation, phase modulation, or amplitude modulation. The ultrasonic transmission signal 802 may be transmitted as part of a calibration procedure or measurement procedure, as further described as part of Figure 9.
[0072] Figure 9 shows an exemplary embodiment of a hearable 102 for performing audio activity detection 112. In the illustrated configuration, the hearable 102 includes a speaker 708, a microphone 710, an analog circuit 712, a pre-processing module 718, a measurement module 720, a calibration module 722, and a circuit 304. Alternatively, to reduce the need for processing power, other embodiments of the hearable 102 are also possible in which the hearable 102 does not include the calibration module 722. In this case, the pre-processing module 718 can perform a frequency selection aspect, which will be further described with respect to Figure 12, to improve the signal-to-noise ratio of the audio plethysmography 110.
[0073] The outputs of speaker 708 and microphone 710 are coupled to the inputs of analog circuit 712. Preprocessing module 718 has an input coupled to the output of analog circuit 712. Preprocessing module 718 also has an output coupled to the inputs of measurement module 720 and calibration module 722. Measurement module 720 has other inputs coupled to microphone 710 and circuit 304, respectively. Calibration module 722 has an output coupled to speaker 708.
[0074] Consider an exemplary operation of the hearable 102 using a single-ear audio plethysmography 110. If the hearable 102 includes a calibration module 722, the hearable 102 can perform a calibration process before performing the measurement process. The calibration process and the measurement process will be further described with reference to Figure 10.
[0075] During both the calibration and measurement processes, the speaker 708 transmits an ultrasonic transmission signal 802, and the microphone 710 receives an ultrasonic reception signal 804. During the calibration process, the ultrasonic transmission signal 802 and the ultrasonic reception signal 804 may have tones 902-1 to 902-M, where M represents a positive integer. During the measurement process, the ultrasonic transmission signal 802 and the ultrasonic reception signal 804 may have selected tones 904-1 to 904-N, where N represents a positive integer less than or equal to M. The selected tones 904-1 to 904-N may represent a subset (sometimes an appropriate subset) of tones 902-1 to 902-M. The microphone 710 can also receive radio signals 402 and bone conduction signals 408 during the measurement process.
[0076] The analog circuit 712 performs analog-to-digital conversion to generate a digital transmission signal 906 and a digital reception signal 908 based on the ultrasonic transmission signal 802 and the received ultrasonic signal 506, respectively. The preprocessing module 718 performs frequency downconversion and demodulation to generate at least one preprocessed signal 910 based on the digital transmission signal 906 and the digital reception signal 908. The preprocessing module 718 can also apply filtering to generate the preprocessed signal 910.
[0077] As part of the calibration procedure, the calibration module 722 processes the pre-processing signal 910 to determine the selected tones 904-1 to 904-N. The selected tones 904-1 to 904-N can improve the performance of the audio plethysmography 110 during the measurement procedure. The calibration module 722 transmits the selected tones 904-1 to 904-N to the speaker 708 using a control signal. The speaker 708 receives the control signal that identifies the selected tones 904-1 to 904-N and may use the selected tones 904-1 to 904-N to transmit a subsequent ultrasonic transmission signal 802 for the measurement procedure.
[0078] As part of the measurement procedure, the measurement module 720 can perform an action of internal noise source filtering 202 using the preprocessing signal 910 and the noise signal 308. The measurement module 720 can also perform an action of voice activity detection 112, speech recognition 114, and / or conversation detection 116 to generate voice data 912. The voice data 912 may also be called speech data or audio plethysmography data. If the environment is noisy, the measurement module 720 can also further process the preprocessing signal 910 for voice activity detection 112, speech recognition 114, and / or conversation detection 116 using the received audible signal 504 provided by the microphone 710. The voice data 912 may be communicated to the application 606, the voice user interface 608, and / or the voice authenticator 610. The voice data 912 may include an indication of whether or not the voice of user 106 has been detected, a phrase recognized based on the voice component 404 of the preprocessing signal 910, and a signal containing the voice component 404. Additionally or alternatively, the voice data 912 may include control signals for controlling the operation of the hearable 102 and / or computing device 104. Calibration and measurement procedures will be further described with reference to Figure 10.
[0079] Figure 10 shows an exemplary flowchart 1000 for operating the hearable 102. In Figure 10, the hearable 102 can optionally perform a calibration procedure in 1002 using a calibration module 722. The calibration procedure can determine appropriate characteristics (e.g., waveform or signal characteristics) of the ultrasonic transmit signal 802 to improve the audio plethysmography 110 (e.g., to enhance the performance of voice activity detection 112). The calibration procedure enables the audio plethysmography 110 to determine a transmit frequency that can increase sensitivity, taking into account the placement of the hearable 102 (e.g., the position of the hearable 102 relative to the ear canal 118) and the physical structure of the ear canal 118. The calibration procedure enables the hearable 102 to dynamically adjust the transmit frequency (e.g., one or more carrier frequencies) each time a seal 120 is formed (e.g., based on the placement of the hearable 102) and based on the unique physical structure of the ear 108. This calibration procedure allows the hearable 102 on different ears 108 to operate at one or more different ultrasonic frequencies. The steps of the calibration procedure are further described below.
[0080] In some situations, the hearable 102 can perform on-head detection by detecting the presence of the sealing 120 and initiating a calibration procedure based on the determination that on-head detection (or in-ear detection) is "true". In other situations, the hearable 102 can initiate a calibration procedure based on a specified schedule or timer that can be controlled by the user 106 via the computing device 104.
[0081] In 1004, the hearable 102 performs a calibration procedure by transmitting and receiving a first ultrasonic signal. The first ultrasonic signal propagates within at least a portion of the user's 106 external auditory canal 118 and has a plurality of tones 902-1 to 902-M (or a plurality of carrier frequencies). The plurality of tones 902-1 to 902-M are transmitted in parallel or sequentially at a given time interval. The first ultrasonic transmission signal 802 may have a specific bandwidth on the order of several kilohertz. For example, the ultrasonic transmission signal 802 may have a bandwidth of about 4, 5, 6, 8, 10, 16, or 20 kHz. In exemplary embodiments, the first ultrasonic transmission signal 802 is transmitted over a plurality of seconds, such as 2, 3, 4, 6, or 6 seconds or more. The duration of each tone 902 can be evenly divided over the total duration of the first ultrasonic transmission signal 802.
[0082] In an exemplary embodiment, the ultrasonic transmission signal 902 has seven tones 902 (e.g., M is equal to 7). In some cases, the tones 902 are evenly distributed over intervals. For example, the tones 902 may be between 32 kHz and 38 kHz (e.g., approximately 32, 33, 34, 35, 36, 37, and 38 kHz) in 1 kHz increments. The term "approximately" means that the tones 902 may be within 5% or less of a given value (e.g., within 3%, 2%, or 1% of a given value).
[0083] The amplitude of the ultrasonic transmission signal 802 can be approximately the same across tones 902-1 to 902-M. In this way, the power is evenly distributed across each tone 902. The amount of tone 902 (e.g., M) can be determined based on the output power of speaker 708. Increasing the amount of tone 902 can increase the likelihood that the hearable 102 can support voice activity detection 112 across a variety of conditions, including the user's wearing and the physical structure of the user's ear canal 118. However, the amplitude of the ultrasonic transmission signal 802 can be limited across these tones 902 based on the output of speaker 708. Therefore, the amount of tone 902 can be optimized based on the amount of output power available to the audio plethysmography 110.
[0084] In 1006, the calibration procedure selects one or more tones 904-1 to 904-N to be used in the measurement procedure based on one or more modified characteristics of the ultrasonic received signal 804. The process for selecting the tones 904 is further described with reference to Figure 11. Generally, the calibration procedure determines that the selected tones 904 improve the signal-to-noise ratio of the audio plethysmography 110 (or more specifically, the audio activity detection 112).
[0085] In 1008, the hearable 102 performs a measurement procedure using the measurement module 720. In accordance with the measurement procedure, the hearable 102 transmits a second ultrasonic transmission signal 802 that propagates within at least a portion of the user's 106 external auditory canal 118. If a calibration procedure has been performed, the second ultrasonic transmission signal 802 may have selected tones 904-1 to 904-M determined by the calibration procedure. The selected tones 904 may be transmitted in parallel or sequentially over a given time interval.
[0086] The amplitude of the second ultrasonic transmission signal 802 may be approximately the same across the selected tones 904-1 to 904-N. In this way, the power is evenly distributed across each selected tone. The amplitude of the second ultrasonic transmission signal 802 may be greater than that of the first ultrasonic transmission signal 802 because the available output power is distributed across fewer tones. Additionally or alternatively, the duration of each selected tone 904 of the second ultrasonic transmission signal 802 may be longer than the duration of the tone 902 of the first ultrasonic transmission signal 802. Higher amplitude and / or longer duration can further improve the signal-to-noise ratio performance of the hearable 102 for audio plethysmography 110. By using several selected tones 904 determined to improve signal-to-noise ratio performance, the measurement procedure can be made more accurate in detecting speech activity 112.
[0087] In 1012, the hearable 102 performs audio plethysmography (e.g., audio processing) using a second ultrasonic signal (e.g., a second ultrasonic received signal 804). One embodiment of performing audio plethysmography 110 may include performing internal noise source filtering 202, as will be further described with reference to Figures 12 and 13. The calibration module 722 will be further described with reference to Figure 11.
[0088] Figure 11 shows an exemplary scheme implemented by the calibration module 722. In the illustrated configuration, the calibration module 722 implements a frequency selector for selecting one or more tones 904 for the measurement procedure. In the exemplary embodiment, the calibration module 722 includes at least one amplitude detector 1102, at least one phase detector 1104, at least one quality detector 1106, and at least one comparator 1108. The operation of these components will be further described below.
[0089] During the calibration procedure, the calibration module 722 receives a preprocessing signal 910 from the preprocessing module 718, as previously described with respect to Figure 9. The preprocessing signal 910 may include amplitude and / or phase information related to a plurality of tones 902-1 to 902-M used to transmit the first ultrasonic signal described in 1002 of Figure 10.
[0090] In this example, the calibration module 722 uses the amplitude detector 1102 to extract the amplitude 1110 of the preprocessed signal 910 and the phase detector 1104 to extract the phase 1112 of the preprocessed signal 910. Alternatively, if the in-phase and orthogonal components of the preprocessed signal 910 are received separately, the amplitude detector 1102 and the phase detector 1104 can measure the amplitude 1110 and phase 1112, respectively, based on the in-phase and orthogonal components.
[0091] The quality detector 1106 measures quality metrics 1114-1 to 1114-2M for each of the tones 902-1 to 902-M, and for each of the characteristics (e.g., amplitude 1110 and phase 1112). Generally, quality metrics 1114 can represent a variety of different metrics, including the peak-to-average ratio and / or the signal-to-noise ratio. The peak-to-average ratio represents the peak intensity within a desired frequency range divided by the average intensity within that frequency range. Higher quality metrics 1114 indicate a higher quality signal, or more generally, better performance for the audio plethysmography 110.
[0092] In one embodiment, the comparator 1108 can evaluate quality metrics 1114-1 to 1114-2M with respect to a threshold 1116. The threshold 1116 can be set to a specific value, for example. In other cases, the calibration module 722 can dynamically determine the threshold 1116 and update the threshold 1116 over time based on the observed quality metrics 1114-1 to 1114-2M. In an exemplary embodiment, the comparator 1108 determines selected tones 904-1 to 904-N for subsequent measurement procedures based on the frequencies associated with quality metrics 1114-1 to 1114-M that are greater than or equal to the threshold 1116.
[0093] Additionally or alternatively, the comparator 1108 may evaluate quality metrics 1114-1 to 1114-2M relative to each other. In an exemplary embodiment, the comparator 1108 determines one of the selected tones 904 based on the frequency having the highest quality metric 1114 over amplitude 1110. Alternatively, the comparator 1108 can determine one of the selected tones 904 based on the frequency having the highest quality metric 1114 over phase 1112. In other embodiments, the comparator 1108 can determine a single selected tone 904 based on the frequency having the highest quality metric 1114 related to either amplitude 1110 or phase 1112.
[0094] Generally, the calibration module 722 allows the selected tones 904-1 to 904-N to be dynamically adjusted before the measurement procedure, based on the current environment which can take into account the fitting of the hearable 102 (e.g., current insertion depth and / or rotation), the physical structure of the user 106's ear canal 118, and the response characteristics of the hearable 102 (e.g., speaker, microphone, and / or housing). In this way, the calibration module 722 can improve the signal-to-noise ratio performance of the hearable 102 for the measurement procedure. The calibration module 722 can also determine which tone 904 produces an ultrasonic received signal 804 with desired characteristics for speech activity detection 112. Generally, the calibration procedure can be performed regardless of whether the user 106 is speaking or not.
[0095] In Figures 9 to 11, the calibration procedure and the measurement procedure are described as individual steps occurring at different time intervals. In particular, the calibration procedure is performed before the measurement procedure. This allows the ultrasonic transmission signal 802 for the measurement procedure to be transmitted with fewer tones than the ultrasonic transmission signal 802 used for the calibration procedure, thereby improving the signal-to-noise ratio performance of the audio plethysmography 110. However, in some embodiments, the hearable 102 may have sufficient output power to perform the measurement procedure with multiple tones 902-1 to 902-M using a single ultrasonic transmission signal 802. In this case, an embodiment of the calibration module may be integrated into the pre-processing module 718 as a frequency selector, which will be further described with respect to Figure 12. This frequency selector can effectively pass the selected tones 904-1 to 904-N for further processing. Embodiments of the measurement procedure will be further described with respect to Figure 12.
[0096] Internal noise source filtering Figure 12 shows an exemplary embodiment of the pre-processing module 718. In the illustrated configuration, the hearable 102 includes the pre-processing module 718 coupled to the measurement module 720 and the calibration module 722. The measurement module 720 is also coupled to a microphone 710 (not shown) and a circuit 304 (not shown).
[0097] The preprocessing module 718 includes at least one in-phase and quadrature mixer 1202 (I / Q mixer 1202) and at least one filter 1204. The in-phase and quadrature mixer 1202 performs frequency down-conversion. In an exemplary embodiment, the in-phase and quadrature mixer 1202 includes at least two mixers, at least one phase shifter, and at least one coupler (e.g., a summing circuit). The filter 1204 attenuates the intermodulation product generated by the in-phase and quadrature mixer 1202. In an exemplary embodiment, the filter 1204 is implemented using a low-pass filter.
[0098] The preprocessing module 718 may optionally include at least one frequency selector 1206. The frequency selector 1206 can identify and select one or more tones 904 (or carrier frequencies) that provide a high-quality signal for later processing. The frequency selector 1206 can further pass the selected tones 904 to other processing modules and filter (or attenuate) other unselected tones. The frequency selector 1206 can be implemented in a similar manner to the calibration module 722 in Figure 11. For example, the frequency selector 1206 may include an amplitude detector 1102, a phase detector 1104, a quality detector 1106, and a comparator 1108.
[0099] During operation, the in-phase and quadrature mixer 1202 uses a phase shifter and two mixers to generate in-phase and quadrature components related to the digital received signal 908. Specifically, the in-phase and quadrature mixer 1202 mixes the digital received signal 908 with a first version of the digital transmitted signal 906 having a zero-degree phase shift to generate the in-phase component. Furthermore, the in-phase and quadrature mixer 1202 mixes the digital received signal 908 with a second version of the digital transmitted signal 906 having a 180-degree phase shift to generate the quadrature signal. This mixing operation downconverts the digital received signal 908 from acoustic frequencies to baseband frequencies. Using a coupler, the in-phase and quadrature mixer 1202 combines the in-phase and quadrature components of the digital received signal 908 to generate the downconverted signal 1208. By using the in-phase and quadrature mixer 1202, it is possible to further improve the signal-to-noise ratio of the downconverted signal 1208 compared to other mixing techniques.
[0100] In this example, the down-converted signal 1208 represents a combination of the common-mode and quadrature components of the mixed-down digital received signal 908. In an alternative embodiment, the common-mode and quadrature mixer 1202 does not include a coupler and passes the common-mode and quadrature components separately to the filter 1204. In this way, the common-mode and quadrature components propagate through the filter 1204 individually.
[0101] Filter 1204 generates a filtered signal 1210 based on the down-converted signal 1208. Specifically, filter 1204 filters the down-converted signal 1208 to attenuate spurious or undesirable frequencies (e.g., intermodulation products). Some of these may relate to the operation of the common-mode and quadrature mixer 1202. In this example, the filtered signal 1210 represents a combination of the common-mode and quadrature components of the down-converted signal 1208. Alternatively, the filtered signal 1210 may represent separate or individual common-mode and quadrature components, which are passed individually to the frequency selector 1206, calibration module 722, or measurement module 720.
[0102] During the measurement procedure, the preprocessing module 718 may optionally apply a frequency selector 1206. The frequency selector 1206 passes tones that meet the quality threshold level of the performance of the audio plethysmography 110. For example, the frequency selector 1206 passes tones 904 having an amplitude 1110 and / or phase 1112 having a quality metric 1114 that is greater than or equal to the threshold 1116. The resulting signal output by the frequency selector 1206 is represented by signal 1212. In some embodiments, this signal 1212 is passed to the measurement module 720 as a preprocessing signal 910. In other embodiments where the frequency selector 1206 is not implemented, the filtered signal 1210 may be passed to the measurement module 720 and / or the calibration module 722 as a preprocessing signal 910.
[0103] Generally, the measurement module 720 can generate audio data 912 based on the preprocessed signal 910. In this case, the measurement module 720 can analyze changes in the amplitude 1110 and / or phase 1112 of the audio component 404 of the preprocessed signal 910 for audio processing. This processing technique can be used in embodiments of the hearable 102 where interference between the radio signal 402 and the ultrasonic signal 410 is minimal (if any), or in situations where audio plethysmography 110 is performed in a relatively quiet environment.
[0104] However, to cope with noisy environments, the measurement module 720 can generate audio data 912 based on the preprocessed signal 910 and the received audible signal 504. More specifically, the measurement module 720 can use the received audible signal 504 as a reference to attenuate the modulation component 510 in the preprocessed signal 910 and improve the signal-to-noise ratio associated with the audio component 404.
[0105] To address situations where an internal noise source 312 operates during audio plethysmography 110, the measurement module 720 can also apply internal noise source filtering 202. The internal noise source filtering 202 allows the measurement module 720 to attenuate the internal noise component 310 within the preprocessed signal 910 and / or the received audible signal 504, using the noise signal 308 as a reference. An exemplary embodiment of the measurement module 720 capable of performing internal noise source filtering 202 will be further described with reference to Figure 13.
[0106] Figure 13 shows an exemplary embodiment of a measurement module 720 for performing internal noise source filtering 202 and speech processing. In the illustrated configuration, the measurement module 720 includes at least one multistage filter 1302 and at least one speech processor 1304 (or utterance processor). The multistage filter 1302 is coupled between the speech processor 1304 and the preprocessing module 718. Although not explicitly shown, the speech processor 1304 may be coupled to other components of the hearable 102, such as a communication interface 704.
[0107] The speech processor 1304 can perform aspects of speech activity detection 112, speech recognition 114, and / or conversation detection 116. In exemplary embodiments, the speech processor 1304 can be implemented using a machine learning model or another module that performs signal and / or data processing. Generally, the speech processor 1304 can analyze changes in amplitude 1110 and / or phase 1112 of the speech component 404 of the preprocessed signal 910 for speech processing.
[0108] To enable the audio processor 1304 to detect and / or utilize the audio component 404, the multistage filter 1302 filters the preprocessed signal 910 to improve the signal-to-noise ratio associated with the audio component 404. In this example, the multistage filter 1302 includes at least one internal noise source filter stage 1306 and at least one external noise source filter stage 1308. Other embodiments are possible in which the multistage filter 1302 includes the internal noise source filter stage 1306 but does not include the external noise source filter stage 1308. This embodiment is possible in situations where audio plethysmography is performed in a relatively quiet environment and / or in embodiments in which the design of the hearable 102 minimizes the influence of the modulation component 510.
[0109] The internal noise source filter stage 1306 performs internal noise source filtering 202 to attenuate internal noise components 310 in the preprocessed signal 910 and / or the received audible signal 504. The internal noise source filter stage 1306 can be implemented using at least two adaptive filters 1310. An exemplary embodiment of the internal noise source filter stage 1306 is further described with reference to Figure 14. Other embodiments are also possible in which the internal noise source filter stage 1306 is implemented using a machine learning model.
[0110] The external noise source filter stage 1306 uses the received audible signal 504 to attenuate the modulation component 510 in the pre-processed signal 910. The external noise source filter stage 1308 can be implemented using at least one filter module 1312. An exemplary embodiment of the external noise source filter stage 1308 is further described with reference to Figure 15. Other embodiments are also possible in which the external noise source filter stage 1308 is implemented using a machine learning model.
[0111] During operation, the measurement module 720 receives a preprocessing signal 910 from the preprocessing module 718. The preprocessing signal 910 may include an internal noise component 310 and an audio component 404. In some embodiments, the preprocessing signal 910 also includes a modulation component 510. The internal noise component 310 and the modulation component 510 may make it difficult for the preprocessing signal 910 to be used directly for audio processing. In some cases, the internal noise component 310 may have a greater impact on the noise level of the preprocessing signal 910 compared to the modulation component 510 because the noise signal 308 is amplified within the ear canal 118.
[0112] The measurement module 720 can also receive an audible signal 504 from the microphone 710. The audible signal 504 may include an internal noise component 310, a speech component 404, and a bone conduction component 512. In noisy environments, the audible signal 504 may also include a noise component 406.
[0113] The internal noise source filter stage 1306 operates on the pre-processed signal 910 and optionally on the received audible signal 504 to attenuate the internal noise component 310. The external noise source filter stage 1308 operates at the output of the internal noise source filter stage 1308 to attenuate the modulation component 510 in the pre-processed signal 910. The output of the external noise source filter stage 1308 is represented by the de-noise signal 1314. The de-noise signal 1314 represents the filtered version of the pre-processed signal 910, including the audio component 404.
[0114] The voice processor 1304 analyzes the denoised signal 1314 to generate voice data 912. Specifically, the voice processor 1304 can generate voice data 912 related to voice activity detection 112, speech recognition 114, and / or conversation detection 116. In the case of voice activity detection 112, the voice data 912 can indicate whether or not a voice component 404 has been detected. In some examples, the voice processor 1304 can perform a signal-to-noise ratio detection process to determine whether or not the amplitude of the input signal exceeds a detection threshold. If the amplitude exceeds the detection threshold, the voice processor 1304 generates voice data 912 indicating that a utterance has been detected. Alternatively, if the amplitude does not exceed the detection threshold, the voice processor 1304 generates voice data 912 indicating that a utterance has not been detected. Voice activity detection 112, speech recognition 114, and / or conversation detection 116 can be used in a variety of different ways to control the operation of the hearable 102 and / or computing device 104. Exemplary embodiments of the multi-stage filter 1302 will be further described with reference to Figures 14 and 15.
[0115] Figure 14 shows an exemplary embodiment of the internal noise source filter stage 1306. In the illustrated configuration, the internal noise source filter stage 1306 includes two adaptive filters 1310-1 and 1310-2. Adaptive filters 1310-1 and 1310-2 can be subjected to various adaptive filtering techniques, including techniques based on least mean squares (LMS) or recursive least squares (RLS). Adaptive filter 1310-1 performs adaptive filtering using the noise signal 308 as a noise reference to attenuate the internal noise component 310 in the preprocessed signal 910. Similarly, adaptive filter 1310-2 performs adaptive filtering using the noise signal 308 as a noise reference to attenuate the internal noise component 310 in the received audible signal 504. In this example, the noise signal 308 is not correlated with the preprocessed signal 910 and the desired speech component 404 in the received audible signal 504. Therefore, the noise signal 308 can be used as a reference signal to attenuate the internal noise component 310 in the preprocessed signal 910 and the received audible signal 504 using adaptive filtering techniques.
[0116] Adaptive filters 1310-1 and 1310-2 generate a denoised preprocessed signal 1402 and a denoised audible signal 1404, respectively. The denoised preprocessed signal 1402 includes a speech component 404 and a modulation component 510. The denoised audible signal 1404 includes a speech component 404, a noise component 406, and a bone conduction component 512. The denoised audible signal 1404 may also be called the denoised signal. The term “denoised audible signal 1404” is used to represent a denoised version of the received audible signal 504 and does not necessarily mean that the denoised audible signal 1404 is audible.
[0117] Generally speaking, the denoised preprocessed signal 1402 and the denoised audible signal 1404 either do not contain the internal noise component 310, or the internal noise component 310 is substantially attenuated with respect to the corresponding input signal. To attenuate the modulation component 510 in the denoised preprocessed signal 1402, the filtering process is followed by an external noise source filter stage 1306. This will be further explained with reference to Figure 15.
[0118] Figure 15 shows an exemplary embodiment of the external noise source filter stage 1306. In the illustrated configuration, the external noise source filter stage 1306 includes at least one filter module 1502 and at least one voice enhancer 1504. The filter module 1502 can be implemented using at least one adaptive filter 1506 or at least one blind source separator 1508.
[0119] The adaptive filter 1506 performs adaptive filtering using the received audible signal 504 as a noise reference to separate the voice component 404 of the preprocessed signal 910 from the modulation component 510 (or the modulation noise component 406). Exemplary adaptive filtering techniques utilized by the adaptive filter 1506 may include techniques based on least mean squares (LMS) or recursive least squares (RLS). The blind source separator 1508 performs blind source separation (BSS) using the received audible signal 504 to separate the voice component 404 of the preprocessed signal 910 from the modulation component 510. For performing adaptive filtering or blind source separation (BSS), the preprocessed signal 910 represents a primary reference, and the received audible signal 504 represents a secondary or noise reference. Generally, the adaptive filter 1506 and the blind source separator 1508 can utilize the received audible signal 504 to significantly attenuate the modulation component 510 (or the modulation noise component 406) in the preprocessed signal 910.
[0120] The voice enhancer 1504 enhances (for example, amplifies with respect to noise levels) the speech component 404 in the output signal provided by the filter module 1502. In this way, the voice enhancer 1504 can increase sensitivity to speech processing. In an exemplary embodiment, the voice enhancer 1504 is implemented using a Wiener filter 1510.
[0121] Exemplary Method Figures 16 and 17 illustrate exemplary methods 1600 and 1700 for implementing aspects of voice activity detection 112 using active acoustic sensing. Methods 1600 and 1700 are presented as a set of actions (or behaviors) to be performed, but the actions are not necessarily limited to the order or combination shown herein. Furthermore, one or more of the actions may be repeated, combined, rearranged, or linked to provide a variety of additional and / or alternative methods. Some of the following discussions may refer to environments 100, 200-1, 200-2, 200-3 in Figures 1 and 2, and entities detailed in Figures 6 and 7, but these references are for illustrative purposes only. These techniques are not limited to being performed by one or more entities operating on a single device.
[0122] In 1602, during the first period, an audible signal is transmitted. The audible signal is intended to propagate within at least a portion of the user's ear canal. For example, circuit 304 transmits (or renders) an audible signal during the first period. The audible signal is intended to propagate within at least a portion of the user 106's ear canal 118. Circuit 304 may include a speaker 314 or 708, an active noise cancellation circuit 316, or a pass-through mode circuit 318. The audible signal may be used to play audio content 204 to the user 106 as shown in environment 200-1, to provide active noise cancellation 206 as shown in environment 200-2, or to enable the hearable 102 to operate according to a pass-through mode 214 as shown in environment 200-3 in Figure 2.
[0123] In 1604, during the first period, an ultrasonic transmission signal is transmitted. The ultrasonic transmission signal is intended to propagate within at least a portion of the user's ear canal. For example, the transducer 706 (or speaker 708) of the hearable 102 transmits an ultrasonic transmission signal 802. The ultrasonic transmission signal 802 propagates within at least a portion of the user's ear canal 118, as described with respect to Figures 4 and 8. The ultrasonic transmission signal 802 may have at least two tones 904 (or frequencies). In some embodiments, the number of tones 904 can be determined based on a calibration procedure, as described in 1002 of Figure 10. In other embodiments, the number of tones 904 can provide frequency diversity and improve the performance of the audio plethysmography 110, as described with respect to Figure 12.
[0124] At 1606, an ultrasonic received signal is received. The ultrasonic received signal represents a version of the ultrasonic transmitted signal having one or more waveform characteristics modified based on propagation within the ear canal and on vocalizations made by the user during the first period. The ultrasonic received signal contains internal noise components caused by interference resulting from the rendering of the audible signal.
[0125] For example, the transducer 706 (or microphone 710) of the hearable 102 receives an ultrasonic receiving signal 804. The ultrasonic receiving signal 804 represents a version of the ultrasonic transmitting signal 802 having one or more waveform characteristics that are modified based on propagation within the external auditory canal 118 and based on vocalizations made by the user 106 during a first period. Exemplary waveform characteristics include amplitude, phase, and / or frequency.
[0126] The ultrasonic received signal 804 includes an internal noise component 310 caused by interference resulting from the rendering of an audible signal. The internal noise component 310 may represent portions of the audible signal that are modulated with or mixed with the ultrasonic received signal 804. Generally, the internal noise component 310 represents portions of the ultrasonic received signal 804 (e.g., amplitude, phase, and / or frequency) that have been modified by interference related to the audible signal.
[0127] The hearable 102 that receives the ultrasonic receiving signal 804 may be the same hearable 102 that transmitted the ultrasonic transmitting signal 802 (e.g., hearable 102-1 or 102-2 in Figure 8), or a different hearable 102 that did not transmit the ultrasonic transmitting signal 802 (e.g., hearable 102-2 in Figure 8). In some embodiments, a feedback microphone of the active noise cancellation circuit 724 may receive the ultrasonic receiving signal 804.
[0128] In 1608, a denoised signal is generated by filtering out internal noise components within the received ultrasonic signal based on the audible signal version. For example, the measurement module 720 (e.g., multi-stage filter 1302) generates a denoised signal 1314 by filtering out internal noise components 310 within the ultrasonic signal 804 based on the audible signal version. The audible signal version is represented by the noise signal 308 shown in Figure 13.
[0129] The audible signal version can represent a digital version of the audible signal, a version of the audible signal that has been down-converted or up-converted to a specific frequency range (e.g., baseband frequency) to remove noise, a pre-processed version of the audible signal, or any combination thereof. Generally, the term “audible signal version” means that there may be some difference between playing the audible signal for user 106 and providing the audible signal as a noise signal 308 for an internal noise source filtering process. To put it another way, an audible signal version is related to the audible signal but can be modified to support the operation of hearable 102. In this case, the operation includes internal noise source filtering.
[0130] In 1610, utterances are detected based on a denoising signal. For example, the hearable 102 uses a measurement module 720 (e.g., a speech processor 1304) to detect utterances based on a denoising signal 1314. More specifically, the measurement module 720 analyzes speech components 404 present in the denoising signal 1314 to perform speech activity detection 112, speech recognition 114, and / or conversation detection 116. The measurement module 720 can generate speech data 912, which can be used to control the hearable 102 and / or computing device 104.
[0131] In Figure 17, at 1702, active acoustic sensing is performed to propagate through the user's ear canal and detect pressure waves associated with the user's vocalizations. For example, the hearable 102 performs active acoustic sensing to propagate through the user's ear canal 118 and detect pressure waves associated with the user's vocalizations. To perform active acoustic sensing, the hearable 102 transmits and receives ultrasonic signals 410 (e.g., ultrasonic transmission signal 802 and ultrasonic reception signal 804). The received ultrasonic signals 410 include a speech component 404, which enables the audio plethysmography 110 to be used for speech processing.
[0132] In circuit 1704, an audible signal that interferes with the active acoustic sensing is rendered. For example, circuit 304 renders an audible signal that interferes with the active acoustic sensing. Circuit 304 may include a speaker 314 or 708, an active noise cancellation circuit 316, and / or a pass-through mode circuit 318. The audible signal may be an audible signal 320 including audio content 204, an anti-noise signal 324, and / or a pass-through mode signal 326.
[0133] In step 1706, internal noise source filtering is performed to attenuate interference caused by rendering the audible signal. For example, the measurement module 720 performs internal noise source filtering 202 to attenuate interference (e.g., internal noise component 310) caused by rendering the audible signal. This interference can be attenuated in the preprocessed signal 910 and optionally in the received audible signal 504, as shown in Figure 14.
[0134] In 1708, speech processing is performed based on active acoustic sensing and internal noise source filtering. For example, the speech processor 1304 performs speech processing (e.g., speech activity detection 112, speech recognition 114, and / or conversation detection 116) based on active acoustic sensing and internal noise source filtering (e.g., based on the noise reduction preprocessing signal 1402).
[0135] At 1710, a signal is generated that controls the operation of at least one of the hearable or a computing device coupled to the hearable. For example, the measurement module 720 can generate voice data 912, which can be used to control the operation of the hearable 102 and / or the computing device 104.
[0136] Throughout this disclosure, the term “version of signal” is used to indicate that a second signal may be a modified version of a first signal. “Version of signal” (or second signal) may represent a digital version of the first signal, an analog version of the first signal, an electrical version of the first signal having voltage and current, an acoustic version of the first signal having acoustic properties, a down-converted version of the first signal having a lower frequency range, an up-converted version of the signal having a higher frequency range, a pre-processed version of the signal (e.g., a version in which amplitude, phase, and / or frequency are modified in some way), a filtered version of the signal, and so on. Generally, “version of signal” refers to a signal that has been modified using techniques known in the art to facilitate the operation of the hearable 102.
[0137] Exemplary computing system Figure 18 shows various components of an exemplary computing system 1800 that can be implemented as any type of client, server, and / or computing device, as described with reference to Figures 6 and 7 above, in order to implement an embodiment of active acoustic sensing using hearables.
[0138] The computing system 1800 includes a communication device 1802 that enables wired and / or wireless communication of device data 1804 (e.g., received data, data being received, data scheduled for broadcast, or data packets of data). The communication device 1802 or the computing system 1800 may include one or more hearables 102. The device data 1804 or other device content may include device configuration settings, media content stored in the device, and / or information relating to the user of the device. The media content stored in the computing system 1800 may include any type of audio, video, and / or image data. The computing system 1800 includes one or more data inputs 1806 that can receive any type of data, media content, and / or input, such as human speech, user-selectable input (explicit or implicit), messages, music, television media content, recorded video content, and any other type of audio, video, and / or image data received from any content and / or data source.
[0139] The computing system 1800 also includes a communication interface 1808, which can be implemented as one or more of the following: a serial interface and / or parallel interface, a wireless interface, any type of network interface, a modem, and any other type of communication interface. The communication interface 1808 provides a connection and / or communication link between the computing system 1800 and a communication network, through which other electronic devices, computing devices, and communication devices communicate data with the computing system 1800.
[0140] The computing system 1800 includes one or more processors 1810 (e.g., a microprocessor, a controller, etc.) that process various computer executable instructions to control the operation of the computing system 1800. Alternatively or additionally, the computing system 1800 may be implemented with one or a combination of hardware, firmware, or fixed logic circuits, which are generally implemented in relation to the processing and control circuits specified in 1812. Although not shown, the computing system 1800 may include a system bus or data transfer system that connects various components within the device. The system bus may include one or a combination of various bus structures, examples of which include a memory bus or memory controller, a peripheral bus, a universal serial bus, and / or a processor bus or local bus that utilizes one of various bus architectures.
[0141] The computing system 1800 also includes computer-readable media 1814, such as one or more memory devices, that enable persistent and / or non-temporary data storage (i.e., as opposed to mere signal transmission), examples of which include random access memory (RAM), non-volatile memory (e.g., one or more of read-only memory (ROM), flash memory, EPROM, EEPROM, etc.), and disk storage devices. The disk storage devices may be implemented as any kind of magnetic or optical storage device, examples of which include hard disk drives, recordable and / or rewritable compact discs (CDs), and any kind of digital multi-purpose discs (DVDs). The computing system 1800 may also include mass storage media devices (storage media) 1816.
[0142] The computer-readable medium 1814 provides a data storage mechanism for storing device data 1804, as well as various device applications 1818, and any other kind of information and / or data related to the operation of the computing system 1800. For example, the operating system 1820 may be maintained as a computer application using the computer-readable medium 1814 and run on the processor 1810. The device applications 1818 may include device managers such as control applications of any form, software applications, signal processing and control modules, code specific to a particular device, and hardware abstraction layers for a particular device.
[0143] The device application 1818 also includes any system components, engines, or managers to implement the internal noise source filtering 202 for active acoustic sensing. In this example, the device application 1818 includes a preprocessing module 718, a measurement module 720, and optionally a calibration module 722. Although not explicitly shown, the device application 1818 may also include an internal noise source filtering stage 1306, an application 606, a voice user interface 608, and / or a voice authenticator 610.
[0144] Through this disclosure, examples are described in which a computing system 1800 (e.g., a hearable 102, a computing device 104, a client device, a server device, a computer, or another type of computing system) can analyze information related to a user (e.g., various audible and / or ultrasonic signals), such as vocalizations. In addition to the foregoing, the systems, programs, and / or functions described herein may enable and allow the collection of information (e.g., information about the user's social networks, social actions, social activities, occupation, user preferences, and current location) and may provide the user 106 with controls that allow the user 106 to choose whether to transmit content or communications from the server to the user 106. The computing system 1800 may be configured to use the information only after it has received explicit permission from the user 106 to use the data. For example, in a situation where the hearable 102 analyzes signals for voice activity detection 112, speech recognition 114, and / or conversation detection 116, individual users 106 may be provided with the opportunity to provide input to control whether a program or function of the computing system 1800 can collect and make available the data. Furthermore, individual users 106 can always control what the program can and cannot do with the information.
[0145] Furthermore, the collected information may be preprocessed in one or more ways so that personally identifiable information is removed before it is transferred, stored, or otherwise used. For example, before the computing system 1800 shares data with other devices, the user 106's identification information may be processed so that no personally identifiable information can be determined about the user 106. Thus, the user 106 can control whether information about the user 106 and the user 106's device is collected, and if such information is collected, how it is used by the computing system 1800 and / or the remote computing system.
[0146] conclusion Techniques using internal noise source filtering for active acoustic sensing, and apparatus including the implementation thereof, have been described in language specific to their features and / or methods, but it will be understood that the subject matter of the attached examples is not necessarily limited to the specific features or methods described. Rather, specific features and methods are disclosed as exemplary embodiments of internal noise source filtering for active acoustic sensing.
[0147] Several embodiments are provided below. Example 1: A method, During the first period, an audible signal is transmitted that propagates within at least a portion of the user's ear canal. During the first period, an ultrasonic transmission signal is transmitted that propagates within at least a portion of the user's external auditory canal. The method further includes receiving an ultrasonic receiving signal during the first period, the ultrasonic receiving signal representing a version of the ultrasonic transmitting signal having one or more characteristics modified based on the propagation within the external auditory canal and based on vocalizations made by the user during the first period, the received ultrasonic receiving signal including internal noise components caused by interference resulting from the rendering of the audible signal, and the method further includes Based on the version of the audible signal, a noise-reduced signal is generated by filtering the internal noise component in the received ultrasonic signal. The process involves detecting the vocalization based on the noise reduction signal, Methods that include...
[0148] Example 2: Controlling the operation of a hearable and / or a computing device coupled to the hearable based on the detected vocalizations. The method according to Example 1, further comprising:
[0149] Example 3: Detecting the vocalization includes performing audio processing using the noise reduction signal, The method according to Example 1 or Example 2, further comprising controlling the operation of a hearable and / or a computing device coupled to the hearable based on the audio processing.
[0150] Example 4: Controlling the hearable is Pausing or resuming the rendering of audio content, Adjusting the volume of the rendered audio content, Enable or disable active noise cancellation, or Enable or disable transparency mode. The method according to Example 3, comprising at least one of the following.
[0151] Example 5: Performing the above audio processing is Perform voice activity detection, Performing speech recognition, or, Perform conversation detection, The method according to Example 3 or Example 4, comprising at least one of the following.
[0152] Example 6: The method according to any of the prior embodiments, wherein the audible signal includes the audio content requested by the user.
[0153] Example 7: The method according to Example 6, wherein the audio content includes music. Example 8: Performing active noise cancellation using the audible signal, It further includes, The audible signal is an anti-noise signal, as described in any of the prior embodiments.
[0154] Example 9: Operating the hearable according to the transmission mode, It further includes, The method according to any of the prior embodiments, wherein the audible signal includes a transmission-mode signal that provides sound from the external environment to the user's ear canal.
[0155] Example 10: The method according to any of the prior embodiments, wherein generating the noise-reduced signal comprises using an adaptive filter to filter the received version of the ultrasonic received signal based on the version of the audible signal.
[0156] Example 11: Further comprising receiving a radio signal including an audible frequency during the first period, wherein the audible frequency includes the utterance made by the user during the first period. The received ultrasonic signal includes a modulation component caused by interference resulting from the reception of the wireless signal. The method according to Embodiment 10, further comprising generating the noise reduction signal by using the received radio signal to attenuate the modulation component in the received ultrasonic received signal.
[0157] Example 12: The wireless signal includes the internal noise component, The generation of the aforementioned noise reduction signal is Using the adaptive filter, a denoised version of the received ultrasonic signal is generated by filtering the received version of the ultrasonic signal based on the version of the audible signal. A noise-reduced audible signal is generated by filtering the version of the wireless signal based on the version of the audible signal using another adaptive filter. Based on the noise-reduced audible signal, the noise-reduced signal is generated by filtering the noise-reduced version of the received ultrasonic signal. The method according to Example 11, including the method described above.
[0158] Example 13: The method according to either Example 11 or Example 12, wherein the receiving of the ultrasonic receiving signal and the receiving of the wireless signal are performed using the same microphone to receive the ultrasonic receiving signal and the wireless signal.
[0159] Example 14: The method according to any of the prior embodiments, wherein the vocalization includes speaking, humming, whistling, or singing.
[0160] Example 15: The method according to any of the prior embodiments, wherein transmitting the ultrasonic transmission signal comprises transmitting the ultrasonic transmission signal having at least two tones.
[0161] Example 16: A computer-readable storage medium comprising an instruction that causes a hearable to execute one of the methods described in Examples 1 to 15 in response to execution by a processor.
[0162] Example 17: A device, At least one transducer, Includes at least one processor, The device is configured to perform any one of the methods described in Examples 1 to 15 using the at least one transducer and the at least one processor.
[0163] Example 18: Speaker and, An active noise cancellation circuit including a feedback microphone, It further includes, The device according to Embodiment 17, wherein the at least one transducer includes the speaker and the feedback microphone.
[0164] Example 19: The at least one transducer includes a speaker and a microphone, The speaker is configured to be positioned close to the user's first ear. The device according to Embodiment 17, wherein the microphone is configured to be positioned close to the user's second ear.
[0165] Example 20: A device according to any one of Examples 17 to 19, At least one earphone, or Devices, including headphones.
Claims
1. It is a method, During the first period, an audible signal is transmitted that propagates within at least a portion of the user's ear canal. During the first period, an ultrasonic transmission signal is transmitted that propagates within at least a portion of the user's external auditory canal. The method further includes receiving an ultrasonic receiving signal during the first period, the ultrasonic receiving signal representing a version of the ultrasonic transmitting signal having one or more characteristics modified based on the propagation within the external auditory canal and based on vocalizations made by the user during the first period, the received ultrasonic receiving signal including internal noise components caused by interference resulting from the rendering of the audible signal, and the method further includes Based on the version of the audible signal, a noise-reduced signal is generated by filtering the internal noise component in the received ultrasonic signal. The process involves detecting the vocalization based on the noise reduction signal, Methods that include...
2. Based on the detected vocalizations, control the operation of the hearable and / or the operation of the computing device coupled to the hearable. The method according to claim 1, further comprising:
3. Detecting the aforementioned vocalization includes performing audio processing using the noise reduction signal, The method involves controlling the operation of a hearable and / or a computing device coupled to the hearable based on the audio processing. The method according to claim 1, further comprising:
4. Controlling the aforementioned hearable means Pausing or resuming the rendering of audio content, Adjusting the volume of the rendered audio content, Enable or disable active noise cancellation, or Enable or disable transparency mode. The method according to claim 3, comprising at least one of the following.
5. Performing the aforementioned audio processing means Perform voice activity detection, Performing speech recognition, or Perform conversation detection, The method according to claim 3, comprising at least one of the following.
6. The method according to claim 1, wherein the audible signal includes audio content requested by the user.
7. The method according to claim 6, wherein the audio content includes music.
8. Performing active noise cancellation using the aforementioned audible signal, It further includes, The method according to claim 1, wherein the audible signal includes an anti-noise signal.
9. To operate the hearable according to the transparency mode, It further includes, The method according to claim 1, wherein the audible signal includes a transmission mode signal that provides sound from the external environment to the user's ear canal.
10. The method according to claim 1, wherein generating the noise-reduced signal includes using an adaptive filter to filter the received version of the ultrasonic received signal based on the version of the audible signal.
11. The first period further includes receiving a radio signal including an audible frequency, wherein the audible frequency includes the utterance made by the user during the first period. The received ultrasonic signal includes a modulation component caused by interference resulting from the reception of the wireless signal. The noise reduction signal is generated by using the received wireless signal to attenuate the modulation component in the received ultrasonic signal. The method according to claim 10, further comprising:
12. The aforementioned wireless signal includes the aforementioned internal noise component, The generation of the aforementioned noise reduction signal is Using the adaptive filter, a denoised version of the received ultrasonic signal is generated by filtering the received version of the ultrasonic signal based on the version of the audible signal. A noise-reduced audible signal is generated by filtering the version of the wireless signal based on the version of the audible signal using another adaptive filter. Based on the noise-reduced audible signal, the noise-reduced signal is generated by filtering the noise-reduced version of the received ultrasonic signal. The method according to claim 11, including the method described in claim 11.
13. The method according to claim 11, wherein the receiving of the ultrasonic receiving signal and the receiving of the wireless signal are performed using the same microphone to receive the ultrasonic receiving signal and the wireless signal.
14. The method according to claim 1, wherein the vocalization includes speaking, humming, whistling, or singing.
15. The method according to claim 1, wherein transmitting the ultrasonic transmission signal includes transmitting the ultrasonic transmission signal having at least two tones.
16. A computer program that, in response to execution by a processor, causes a hearable to execute any one of the methods described in claims 1 to 15.
17. It is a device, At least one transducer, At least one processor, Includes, The device is configured to perform any one of the methods according to claims 1 to 15 using the at least one transducer and the at least one processor.
18. Speakers and, It further includes an active noise cancellation circuit including a feedback microphone, The device according to claim 17, wherein the at least one transducer includes the speaker and the feedback microphone.
19. The at least one transducer includes a speaker and a microphone, The speaker is configured to be positioned close to the user's first ear. The device according to claim 17, wherein the microphone is configured to be positioned close to the user's second ear.
20. The device according to claim 17, At least one earphone, or headphone, A device that includes this.
Citation Information
Patent Citations
Microphone apparatus, utterance detector, utterance detecting method, and voice outputting method
JP2006139117A
Two-way communication device with one transducer and its method
JP2007511962A
Active acoustic sensing
WO2023240224A1