Fusion of audioplethysmography and motion detection data
The fusion of audio plethysmography and motion sensing data addresses the challenges of motion artifacts in wireless hearables, improving accuracy and sensitivity for health monitoring and activity detection, making health tracking more accessible and convenient.
Patent Information
- Application Number
- JP2025500052
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2026-02-18
- Estimated Expiration
- 2043-12-29
AI Technical Summary
Health monitoring devices can be intrusive, uncomfortable, and inconvenient, leading individuals to forgo health monitoring, especially when they negatively impact movement or daily activities, and existing wireless hearables face challenges with motion artifacts that reduce accuracy and sensitivity during user movement.
Fusion of audio plethysmography and motion sensing data to attenuate motion artifacts, enabling accurate biometric monitoring and activity detection, and extending the functionality of hearables to include activity classification and control operations.
Enhances the sensitivity and accuracy of health monitoring during user movement, allowing for reliable, portable, and inexpensive health tracking, expanding the usability of hearables in various activities.
Smart Images

Figure 2026505679000001_ABST
Abstract
Description
[Background technology]
[0001] Technological advances in medicine and healthcare are enabling people to live longer, healthier lives. To further achieve this, individuals are becoming interested in tracking their personal health. Health monitoring can motivate individuals to achieve specific fitness goals by tracking incremental improvements in physical performance. Additionally, individuals can monitor the impact of various chronic conditions on their bodies. Active feedback through health monitoring enables individuals to live active and fulfilling lives with many chronic conditions and quickly recognize when it is necessary to seek treatment.
[0002] However, some devices that support health monitoring can be intrusive and uncomfortable. Thus, people may choose to forgo health monitoring if the device negatively affects their movement or causes inconvenience while performing daily activities. Therefore, it is desirable for health monitoring devices to be reliable, portable, and inexpensive to encourage more users to take advantage of these features. Summary of the Invention
[0003] Techniques and apparatus for fusing audio plethysmography and motion sensing data are described. Fusing motion sensing data with audio plethysmography expands the situations in which audio plethysmography can operate. In one aspect, the fusion of audio plethysmography and motion sensing data can be used for motion artifact filtering. Motion artifact filtering can be used to attenuate noise caused by user movement, improving the sensitivity and accuracy of audio plethysmography. This improved performance expands the ability of audio plethysmography to support use cases during situations in which the user is engaged in activity. In another aspect, the fusion of audio plethysmography and motion sensing data can expand the functionality of a hearable to include activity detection and / or activity classification. Activity detection and / or activity classification can provide additional contextual information for other use cases associated with audio plethysmography, either of which can be used to control the operation of a hearable and / or a computing device.
[0004] Aspects described below include a method for performing fusion of audio plethysmography and motion sensing data. The method includes transmitting, during a first time period, an acoustic transmit signal that propagates within at least a portion of a user's ear canal. The method also includes receiving, during the first time period, an acoustic receive signal that represents a version of the acoustic transmit signal having one or more characteristics modified based on propagation within the ear canal. The acoustic receive signal includes at least one movement artifact associated with the user moving during at least a portion of the first time period. The method further includes generating a denoised version of the acoustic receive signal using motion sensing data generated by a motion sensor during the first time period. The at least one movement artifact in the denoised version of the acoustic receive signal is attenuated relative to the at least one movement artifact in the acoustic receive signal. The method includes controlling operation of the hearable and / or a computing device coupled to the hearable based on the denoised version of the acoustic receive signal.
[0005] Aspects described below include a computer-readable storage medium containing instructions that, when executed by a processor, cause a hearable to perform any of the methods described herein.
[0006] The embodiments described below include devices having at least one transducer and at least one processor, the devices being configured to perform any of the methods described herein using the at least one transducer and the at least one processor.
[0007] Aspects described below include systems having means for performing fusion of audio plethysmography and motion sensing data.
[0008] Apparatus and techniques for performing audio plethysmography and motion sensing data fusion are described with reference to the following drawings, in which the same numbers are used throughout to refer to like features and components. [Brief explanation of the drawings]
[0009] [Figure 1-1] 1 illustrates an exemplary environment in which audio plethysmography can be implemented. [Figure 1-2] 1 illustrates exemplary geometric changes within the ear canal that can be detected using audio plethysmography. [Figure 2] 1 illustrates an exemplary environment in which audio plethysmography and motion sensing data fusion can be implemented. [Figure 3] 1 shows the effect of movement on audio plethysmography. [Figure 4] 1 illustrates exemplary components of a computing device. [Figure 5] 1 illustrates exemplary components of a hearable. [Figure 6] 1 illustrates an exemplary operation of two hearables. [Figure 7] 1 illustrates an exemplary embodiment of a hearable capable of performing fusion of audio plethysmography and motion sensing data. [Figure 8] 1 shows an exemplary flow diagram for operating a hearable. [Figure 9] 1 illustrates an exemplary scheme implemented by a calibration module of a hearable. [Figure 10] 1 illustrates an exemplary scheme that a hearable implements to perform fusion of audio plethysmography and motion sensing data. [Figure 11-1] 1 illustrates a first exemplary implementation of a motion artifact filter for performing fusion of audio plethysmography and motion sensing data. [Figure 11-2] 10 illustrates a second exemplary implementation of a motion artifact filter for performing fusion of audio plethysmography and motion sensing data. [Figure 12] 1 illustrates an exemplary implementation of a measurement module for performing fusion of audio plethysmography and motion sensing data. [Figure 13]1 illustrates an exemplary method for performing one aspect of audio plethysmography and motion sensing data fusion. [Figure 14] 10 illustrates another exemplary method for performing one aspect of audio plethysmography and motion sensing data fusion. [Figure 15] 10 illustrates yet another exemplary method for performing one aspect of audio plethysmography and motion sensing data fusion. [Figure 16] 1 illustrates an exemplary computing system that may embody or implement techniques that enable the use of fusion of audio plethysmography and motion sensing data. DETAILED DESCRIPTION OF THE INVENTION
[0010] Technological advances in medicine and healthcare are enabling people to live longer, healthier lives. To further achieve this, individuals are becoming interested in tracking their personal health. Health monitoring can motivate individuals to achieve specific fitness goals by tracking incremental improvements in physical performance. Additionally, individuals can use health monitoring to observe internal changes caused by chronic diseases. Active feedback through health monitoring allows individuals to live active and fulfilling lives with many chronic diseases and recognize when prompt medical treatment is warranted.
[0011] However, some health monitoring devices can be intrusive and uncomfortable. For example, to measure carbon dioxide levels, some devices draw blood samples from the user. Other devices may utilize auxiliary sensors, including optical or electronic sensors, which add additional weight, cost, complexity, and / or bulk. Still other devices may require constant recharging of batteries due to relatively high power usage. Thus, people may choose to forgo health monitoring if the health monitoring device negatively impacts their movement or causes inconvenience while performing daily activities. Therefore, it is desirable for health monitoring devices to be reliable, portable, efficient, and inexpensive to expand accessibility to more users.
[0012] Wireless technology has permeated everyday life, allowing users easy access to communications and data. One type of wireless technology is wireless hearables, examples of which include wireless earbuds and wireless headphones. Wireless hearables allow users freedom of movement while listening to audio content from music, audiobooks, podcasts, and videos. Due to the popularity of wireless hearables, a market exists for adding additional features to existing hearables using current hardware (e.g., without introducing any new hardware).
[0013] According to one or more preferred embodiments, provided are hearables, such as earphones, capable of performing a novel physiological monitoring process referred to herein as audio plethysmography. Audio plethysmography is an active acoustic method capable of detecting subtle, physiologically relevant changes observable in a user's outer and middle ear. Instead of relying on other auxiliary sensors, such as optical or electrical sensors, audio plethysmography involves transmitting and receiving acoustic signals that propagate at least partially within the user's ear canal. To effectively perform audio plethysmography, the hearable must form at least a partial seal in or around the user's outer ear. Such a seal allows for the formation of an acoustic circuit that includes the seal, the hearable, the ear canal, and the ear's tympanic membrane. By transmitting and receiving acoustic signals, the hearable can recognize changes in the acoustic circuit to monitor the user's biometrics, including heart rate, respiratory rate, heart rate variability, and / or blood pressure. Audio plethysmography can also be used for other use cases, including speech detection, chewing detection, gesture recognition, etc. In addition to being relatively unobtrusive, some hearables can be configured to support audio plethysmography without requiring additional hardware. Thus, the size, cost, and power usage of hearables may help make health monitoring accessible to a larger group of people and improve the user experience with hearables.
[0014] Although audio plethysmography can support a variety of different use cases, audio plethysmography can be sensitive to motion. As a user performs various activities, movement of the user's head can introduce motion artifacts (e.g., noise) into the received acoustic signal. Motion artifacts represent portions of the received acoustic signal whose amplitude, phase, and / or frequency are affected by user movement. In one aspect, these motion artifacts can make it difficult for the audio plethysmograph to process the received acoustic signal and extract the information desired for a given use case. Motion artifacts can also reduce measurement accuracy and / or cause false positives.
[0015] The challenges of performing audio plethysmography in the presence of movement may differ from the movement-based challenges experienced by other types of sensors. Because audio plethysmography is typically performed using hearables, audio plethysmography is particularly sensitive to movement of the user's head. Generally speaking, a user may be more likely to move their head in more situations than other body parts. These situations may include situations in which the user is relatively inactive (e.g., sitting, reading, or working at a desk). In contrast, other sensors placed on the user's torso or appendages (e.g., wrist) may be exposed to less movement while the user is relatively inactive. Other activities, such as running, may expose audio plethysmography and these other sensors to different types of movement. Thus, audio plethysmography may have different movement artifacts and challenges compared to other types of sensors.
[0016] To address this challenge and provide new features to existing hearables, techniques for the fusion of audio plethysmography and motion sensing data are described. Fusing motion sensing data with audio plethysmography expands the situations in which audio plethysmography can operate. In one aspect, the fusion of audio plethysmography and motion sensing data can be used for motion artifact filtering. Motion artifact filtering can be used to attenuate noise caused by user movement, improving the sensitivity and accuracy of audio plethysmography. This improved performance expands the ability of audio plethysmography to support use cases during situations in which a user is engaged in activity. One such use case includes biometric monitoring, which can be performed using audio plethysmography while a user is exercising.
[0017] In another aspect, the fusion of audio plethysmography and motion sensing data can extend the functionality of the hearable to include activity detection and / or activity classification. Activity detection and / or activity classification can provide additional contextual information for other use cases associated with audio plethysmography. Additionally or alternatively, activity detection and / or activity classification can be used to control the operation of the hearable and / or computing device.
[0018] Operating environment FIG. 1-1 is a diagram of an example environment 100 in which audio plethysmography can be implemented. In the example environment 100, a hearable 102 is connected to a computing device 104 using a physical or wireless interface. The hearable 102 is a device that can play audible content provided by the computing device 104 and direct the audible content into the ear 108 of a user 106. In this example, the hearable 102 operates in conjunction with the computing device 104. In other examples, the hearable 102 can operate or be implemented as a standalone device. While shown as a smartphone, the computing device 104 can include other types of devices, including those described with respect to FIG. 4.
[0019] The hearable 102 can perform audio plethysmography 110, an ultrasound method of sensing that occurs at the ear 108. The hearable 102 can perform this sensing without the use of other auxiliary sensors, such as optical or electrical sensors. Audio plethysmography 110 enables the hearable 102 to perform biometric monitoring 112, speech detection 114, chewing detection 116, and / or gesture recognition 118. The biometric monitoring 112 can include determining (or measuring) the heart rate, breathing rate, heart rate variability, and / or blood pressure of the user 106. Heart rate variability is the deviation in timing between heartbeats and may reflect the physiological and / or emotional state of the user 106. For example, heart rate variability may indicate a cardiac abnormality or a mental health issue, such as anxiety or depression. Biometric monitoring 112 allows users 106 to actively monitor their health and take appropriate action based on any changes in their biometrics to live longer, healthier lives.
[0020] Speech detection 114 enables the hearable 102 to determine whether the user 106 is speaking. By providing speech detection 114, audio plethysmography 110 can enhance voice control. Voice control allows a user to interact with the computing device 104 in a less physically and cognitively demanding manner compared to other interfaces that require physical contact and / or the user's visual attention. Additionally or alternatively, speech detection 114 can be used to prevent unauthorized individuals from accessing the voice control interface of the computing device 104. In particular, audio plethysmography 110 can provide an aspect of multi-factor authentication, verifying that the user 106 is speaking while the computing device 104 recognizes a voiceprint phrase. This can prevent malicious individuals from attempting to access the voice control interface by playing a recording of the user's voiceprint against the computing device 104. In this manner, speech detection 114 can enhance the security of the computing device 104 and provide robust protection from voice attacks.
[0021] Chewing detection 116 enables the hearable 102 to monitor teeth grinding. Teeth grinding is a condition in which a user 106 grinds or clenches their teeth, usually unconsciously. In severe cases, teeth grinding can cause excessive tooth wear or jaw intrusion, which can lead to tooth loss, an abnormal bite, or crooked teeth. Teeth grinding can also be an indicator of sleep apnea or stress. Chewing detection 116 enables the hearable 102 to determine the frequency and duration of teeth grinding and communicate this information to the user 106 directly via the hearable 102 or indirectly via the computing device 104. By knowing this, the user 106 can choose to seek medical advice or make lifestyle changes to reduce the occurrence of teeth grinding. In some cases, the hearable 102 can sound an alarm to wake the user 106 and stop the user 106 from grinding their teeth. Using this alarm, the user 106 can train themselves to reduce or stop grinding their teeth.
[0022] Additionally or alternatively, chewing detection 116 can capture the times when the user 106 starts and stops eating. This may be particularly useful for automatically monitoring intermittent fasting or snacking. The user 106 can later review this information to determine how well they adhered to an intermittent fasting plan or how frequently they snack. In some cases, the hearable 102 can estimate the user's 106 calorie intake based on the duration of chewing activity. The hearable 102 can also be used to train children to properly chew food before swallowing. For example, the hearable 102 can play child-friendly audio content when the child chews their food and make a sound when the child chews a target number of times before swallowing.
[0023] Gesture recognition 118 enables the hearable 102 to recognize gestures that involve the user 106 engaging different body parts or interacting with different regions of their upper body. A simple ear tug or head shake can be detected by the hearable 102 and used to control the computing device 104. More specifically, audio plethysmography 110 can detect subtle pressure waves that originate in the user's 106 upper body and propagate into the ear canal. These pressure waves are transmitted and received by the hearable 102, modifying the characteristics of the acoustic signal propagating through the ear canal. Gesture recognition 118 enables the hearable 102 to support a greater amount and variety of gesture-based control compared to the limited touch-based control of some hearables. This is because the user 106 can utilize their entire upper body region for different gesture-based control, whereas touch-based control is limited to the surface of other hearables. Gesture-based control can also be subtle so as not to attract attention, especially during social events.
[0024] To use the audio plethysmograph 110, the user 106 positions the hearable 102 to create at least a partial seal 120 around or within the ear 108. Several portions of the ear 108 are shown in FIG. 1-1 , including an ear canal 122 and a tympanic membrane 124 (or eardrum). The seal 120 couples the hearable 102, ear canal 122, and tympanic membrane 124 together to form an acoustic circuit. The audio plethysmograph 110 involves, at least in part, measuring characteristics associated with this acoustic circuit. The characteristics of the acoustic circuit can change due to a variety of different situations or actions.
[0025] 1-2, in which changes occur in the physical structure of the ear 108. Exemplary changes in physical structure include changes in the geometry of the ear canal 122 and / or changes in the volume of the ear canal 122. This change may be caused, at least in part, by subtle deformations of blood vessels within the ear canal 122 caused by the beating of the user's 106 heart. Other changes may also be caused by movement within the ear canal 124 or movement of the user's 106 jaw.
[0026] At 126, for example, the tissue surrounding the ear canal 122, and the eardrum 124 itself, are slightly "squeezed" by the deformation of blood vessels. This squeezing causes the volume of the ear canal 122 to decrease slightly at 126. However, at 128, the squeezing subsides and the volume of the ear canal 122 increases slightly compared to 126. The physical changes within the ear 108 can modulate the amplitude and / or phase of the acoustic signal propagating through the ear canal 122, as described further below.
[0027] During audio plethysmography 110, an acoustic signal propagates through at least a portion of the ear canal 122. The hearable 102 can receive an acoustic signal that represents a superposition of multiple acoustic signals propagating along different paths within the ear canal 122. Each path is associated with a delay (τ) and an amplitude (a). The delay and amplitude may change over time due to subtle changes that occur in the volume of the ear canal 122. The received acoustic signal can be expressed by Equation 1:
number
number
[0028] 2 illustrates example environments 200-1 through 200-5 in which audio plethysmography and motion sensing data fusion 202 can be implemented. In the environments 200-1 through 200-5, a user 106 performs a variety of different activities while wearing a hearable 102. The hearable 102 includes at least one motion sensor 204 or can access data generated by a motion sensor 204 that is implemented separately from the hearable 102, such as implemented within a computing device 104.
[0029] In environments 200-1, 200-2, and 200-3, user 106 engages in activities that use a relatively low amount of energy. For example, user 106 works at a desk in environment 200-1. In environment 200-2, user 106 reads a book. In environment 200-3, user 106 talks to another person. Other low-energy activities can include eating, driving, and sleeping. In general, low-energy activities may involve relatively little movement, or movement that occurs after a relatively long interval of little movement. In some cases, user 106 can be considered relatively motionless while performing an activity. While user 106 may be characterized as relatively motionless, user 106 may still move slightly due to normal bodily functions such as breathing or natural muscle movements.
[0030] For some of these low-energy activities, the user 106 may be relatively stationary with respect to Global Navigation Satellite System (GNSS) coordinates. For example, the user 106 may be sitting, standing, or lying in a particular position. For other low-energy activities, the user 106 may be relatively stationary in a vehicle (e.g., a car, train, or airplane) as the vehicle moves to different locations.
[0031] Many low-energy activities may involve users 106 minimally moving their appendages and / or bodies. Sometimes, users 106 may move their heads while engaging in low-energy activities. For example, users 106 in environment 200-1 may stop looking at a computer monitor and move their heads to look out a window. Users 106 in environment 200-2 may move their heads when scanning the pages of a book. Users 106 in environment 200-3 may move their heads when communicating with others.
[0032] In environments 200-4 and 200-5, user 106 participates in activities that involve a relatively large amount of energy compared to the low-energy activities described in environments 200-1 through 200-3. For example, in environment 200-4, user 106 moves a significant portion of his or her body. User 106 may be dancing, as shown in FIG. 2, or engaging in other activities such as walking, running, biking, skateboarding, roller skating, swimming, exercising, or playing a sport. In environment 200-5, user 106 performs chores such as washing dishes, cooking, folding laundry, mowing the lawn, gardening, and shoveling snow.
[0033] For some high-energy activities, the user 106 may be physically moving to another Global Navigation Satellite System (GNSS) coordinate system using their own power. As an example, the user 106 may be riding a bicycle to the store. For other high-energy activities, the user 106 may be stationary relative to the Global Navigation Satellite System (GNSS) coordinate system and / or may be moving their appendages or body substantially and / or rapidly. For example, at a gym, the user 106 may engage in high-energy activities by exercising on a treadmill or elliptical machine. Within the environments 200-1 and 200-5, the general movement of the user 106, and more specifically, the movement of the user's 106's head, may significantly affect the audio plethysmograph 110.
[0034] To address challenges associated with movement and provide additional features to the user 106, the hearable 102 may perform audio plethysmography and motion sensing data fusion 202. The audio plethysmography and motion sensing data fusion 202 processes audio signals associated with the audio plethysmography 110 using motion sensing data generated by the motion sensor 204. Exemplary features that may be implemented as part of the audio plethysmography and motion sensing data fusion 202 include motion artifact filtering 206, activity detection 208, and / or activity classification 210. Activity classification 210 may also be referred to as activity recognition.
[0035] Motion artifact filtering 206 enables the hearable 102 to improve the sensitivity and accuracy of the audio plethysmography 110 while the user is engaged in low-energy or high-energy activities. For example, biometric monitoring 112 can utilize motion artifact filtering 206 to accurately measure the heart rate of the user 106 while the user 106 is jogging. Motion artifact filtering 206 can also be used to reduce false prediction rates or false positives associated with other use cases of the audio plethysmography 110, such as speech detection 114, chewing detection 116, and / or gesture recognition 118. Other hearables that do not utilize motion artifact filtering 206 may not be able to accurately collect data using the audio plethysmography 110 while the user 106 is moving, which can significantly limit the usefulness of the audio plethysmography 110 and cause inconvenience to the user 106.
[0036] Activity detection 208 uses audio plethysmography and motion detection data fusion 202 to determine when the user 106 is moving. Information about when and how often the user 106 moves can provide additional data for a variety of different use cases. A sleep quality analysis can use this information, for example, to estimate how well the user 106 slept. Similar analyses can be performed to determine the frequency and / or quality of meditation. Another use case may include automatically detecting when the user 106 takes a break from sitting in front of a computer and then stands up and / or walks around, and alerting the user 106 if they have not taken a break within a predetermined time.
[0037] Activity classification 210 uses fusion 202 of audio plethysmography and motion sensing data to determine the type of activity the user 106 is engaged in. While some devices can utilize motion sensors to perform aspects of activity classification, fusing data provided by the motion sensor 204 with data generated by the audio plethysmography 110 can further disambiguate similar activities, thereby enabling a larger volume of activities to be recognized. In one aspect, data from biometric monitoring 112 via audio plethysmography 110 can be used in conjunction with the motion sensing data to accurately classify the activity the user 106 is engaged in. In another aspect, motion artifacts present in the signal associated with audio plethysmography 110 can be used to increase confidence and / or further distinguish the movements detected by the motion sensor 204.
[0038] Activity classification 210 enables hearable 102 to distinguish between low-energy activities and / or high-energy activities. For example, activity classification 210 can distinguish between, for example, whether the user 106 is mixing food by hand or by rowing. In some implementations, activity classification 210 can further distinguish between different types of low-energy and high-energy activities. For low-energy activities, hearable 102 can use activity classification 210 to identify when the user 106 is eating and when the user 106 is talking. For high-energy activities, hearable 102 can identify when the user 106 is walking and when the user 106 is running. In general, data collected using activity detection 208 and / or activity classification 210 can be used to support a variety of different use cases, including health monitoring, fitness tracking, stress tracking, etc.
[0039] Activity detection 208 and / or activity classification 210 can also be used to control the operation of hearable 102 and / or computing device 104. For example, activity detection 208 can dynamically increase the volume of audio content presented by hearable 102 when activity is detected, or decrease the volume when activity is no longer detected. Activity classification 210 can automatically set a stopwatch, play music previously selected by the user, or open a fitness tracking application on computing device 104 when it determines that user 106 is exercising.
[0040] While motion artifact filtering 206, activity detection 208, and activity classification 210 may be described with respect to audio plethysmography and motion sensing data fusion 202, these processes may also be generally associated with audio plethysmography 110. Audio plethysmography 110 involves transmitting, receiving, and processing acoustic signals, which may include performing one or more aspects of audio plethysmography and motion sensing data fusion 202. Challenges associated with performing audio plethysmography 110 while the user 106 is moving are further described with respect to FIG.
[0041] Figure 3 illustrates the effect of movement on audio plethysmography 110. At the top of Figure 3, a first graph 300-1 illustrates the amplitude and frequency of an audio plethysmography signal (APG signal 302) 302. In this case, audio plethysmography signal 302 is collected while user 106 is relatively motionless 304. User 106 may be engaged in low-energy activity, such as any of the activities described with respect to environments 200-1 to 200-3. In this case, user 106 is not moving their head (e.g., the head is relatively still). Audio plethysmography signal 302 represents a processed version of the received acoustic signal.
[0042] In this example, the audio plethysmograph 110 is used for biometric monitoring 112. The audio plethysmograph 110 can measure the heart rate of the user 106, for example, based on the maximum peak amplitude of the audio plethysmography signal 302. This measured heart rate is called the audio plethysmography-determined heart rate 306 (APG-determined heart rate 306). The audio plethysmography-determined heart rate 306 is approximately equal to the heart rate 308 of the user 106. The term “approximately” may mean that the audio plethysmography-determined heart rate 306 is within 5% of the heart rate 308 of the user 106 (e.g., within 5%, 3%, 2%, or 1% of the heart rate 308).
[0043] Because the user 106 is relatively motionless 304 in the example shown in first graph 300-1, the audio plethysmography signal 302 has relatively little noise, and the maximum peak amplitude corresponds to the heart rate 308 of the user 106. When the user 106 is moving or participating in high-energy activities, noise caused by the movement of the user 106 may adversely affect the audio plethysmography signal 302. As explained further below, this noise may make it difficult to detect the heart rate 308 of the user 106.
[0044] At the bottom of Figure 3, a second graph 300-2 shows the amplitude and frequency of another audio plethysmography signal 302 collected while the user 106 is moving 310. In this example, the user 106 is walking slowly. Other examples are possible where the user 106 is engaged in any of the high-energy activities described with respect to environments 200-4 and 200-5 of Figure 2.
[0045] Similar processing can be used to generate the audio plethysmography signal 302 shown in graphs 300-1 and 300-2. However, the audio plethysmography signal 302 in the second graph 300-2 includes several movement artifacts 312 that are not present in the first graph 300-1. The first movement artifact 312-1 represents the maximum peak amplitude of the audio plethysmography signal 302 and corresponds to the cadence of the user 106 while walking. Other movement artifacts 312-2 and 312-3 are also present in the second graph 300-2. The movement artifacts 312-2 and 312-3 may represent harmonic or intermodulation products of the movement artifact 312-1. Additionally or alternatively, the movement artifacts 312-2 and / or 312-3 may be associated with other movements, such as head or arm movements of the user 106.
[0046] If the audio plethysmograph 110 were to determine the heart rate of the user 106 based on the maximum peak amplitude of the audio plethysmograph signal 302, the audio plethysmograph 110 would erroneously identify the frequency of the movement artifact 312-1 as the heart rate of the user 106. To avoid these errors, techniques utilizing audio plethysmography and motion sensing data fusion 202 can attenuate the movement artifact 312 to enable detection of the desired information (e.g., heart rate 308). By using audio plethysmography and motion sensing data fusion 202, the audio plethysmograph signal 302 is further processed based on the motion sensing data 314, which is also shown in the second graph 300-2. In particular, the motion artifact filtering 206 performed based on the motion sensing data 314 can be used to filter one or more of the motion artifacts 312-1 through 312-3, enabling the audio plethysmograph 110 to correctly measure the heart rate 308 of the user 106. The heart rate 308 measured using the audio plethysmography and motion sensing data fusion 202 is represented by the audio plethysmography-determined heart rate 316. The frequency of the audio plethysmography signal 302 shown in the graphs 300-1 and 300-2 can be downconverted to baseband frequency to facilitate alignment with the motion sensing data 314. The computing device 104 is further described with respect to FIG. 4.
[0047] 4 illustrates an exemplary implementation of a computing device 104. Computing device 104 is shown along with various non-limiting exemplary devices, including a desktop computer 104-1, a tablet 104-2, a laptop 104-3, a television 104-4, a computing watch 104-5, computing glasses 104-6, a gaming system 104-7, a microwave oven 104-8, and a vehicle 104-9. Other devices, such as augmented reality and / or virtual reality headsets, home service devices, smart speakers, smart thermostats, baby monitors, Wi-Fi™ routers, drones, trackpads, drawing pads, netbooks, e-readers, home automation and control systems, wall displays, and other home appliances, may also be used. Note that computing device 104 may be wearable, non-wearable but mobile, or relatively fixed (e.g., desktop and appliance devices).
[0048] The computing device 104 includes one or more computer processors 402 and at least one computer-readable medium 404, including memory and storage media. Applications and / or an operating system (not shown) embodied as computer-readable instructions on the computer-readable medium 404 can be executed by the computer processor 402 to provide some of the functionality described herein. The computer-readable medium 404 also includes an audio plethysmography-based application 406 that performs actions using information provided by the hearable 102. An example action can include displaying data associated with the audio plethysmography 110 to the user 106. More specifically, the data can be associated with biometric monitoring 112, speech detection 114, chewing detection 116, gesture recognition 118, activity detection 208, and / or activity classification 210.
[0049] The computing device 104 may also include a network interface 408 for communicating data over a wired, wireless, or optical network. For example, the network interface 408 may communicate data over a local area network (LAN), a wireless local area network (WLAN), a personal area network (PAN), a wide area network (WAN), an intranet, the Internet, a peer-to-peer network, a point-to-point network, a mesh network, Bluetooth, etc. The computing device 104 may also include a display 410. Although not explicitly shown, the hearable 102 may be integrated within the computing device 104 or may be physically or wirelessly connected to the computing device 104. The hearable 102 is further described with respect to FIG. 5.
[0050] FIG. 5 illustrates an exemplary hearable 102. The hearable 102 is shown with various non-limiting exemplary devices, including a wireless earphone 502-1, a wired earphone 502-2, and a headphone 502-3. The earphones 502-1 and 502-2 are a type of in-ear device that fits within the ear canal 122. Each earphone 502-1 or 502-2 can represent a hearable 102. The headphone 502-3 can rest on top of or on the ear 108. The headphone 502-3 can represent a closed-back headphone, an open-back headphone, an on-ear headphone, or an over-ear headphone. Each headphone 502-2 includes two hearables 102 physically packaged together. Generally, there is one hearable 102 per ear 108.
[0051] The hearable 102 includes a communication interface 504 for communicating with the computing device 104, although this need not be used when the hearable 102 is integrated within the computing device 104. The communication interface 504 may be a wired or wireless interface, and audio content is passed from the computing device 104 to the hearable 102. The hearable 102 may also use the communication interface 504 to pass data associated with the audio plethysmography 110 to the computing device 104. Generally, the data provided by the communication interface 504 is in a format usable by the audio plethysmography-based application 406.
[0052] The communication interface 504 also allows the hearable 102 to communicate with other hearables 102. For example, during bistatic sensing, the hearable 102 can use the communication interface 504 to coordinate with other hearables 102 to support binaural audio plethysmography 110, as further described with respect to Figure 6. In particular, the transmitting hearable 102 can communicate timing and waveform information to the receiving hearable 102 to enable the receiving hearable 102 to properly demodulate the received acoustic signal.
[0053] The hearable 102 includes at least one transducer 506 capable of converting electrical signals into sound waves. The transducer 506 can also detect and convert sound waves into electrical signals. These sound waves may include ultrasonic and / or audible frequencies, either of which may be used in the audio plethysmograph 110. In particular, the frequency spectrum (e.g., range of frequencies) used by the transducer 506 to generate acoustic signals may include frequencies from the low end of the audible range to the high end of the ultrasonic range, e.g., 20 hertz (Hz) to 2 megahertz (MHz). Other exemplary frequency spectrums for the audio plethysmograph 110 may encompass frequencies between 20 Hz and 20 kilohertz (kHz), between 20 kHz and 2 MHz, between 20 and 96 kHz, between 20 and 60 kHz, or between 30 and 40 kHz.
[0054] In an exemplary implementation, the transducer 506 has a monostatic topology. This topology allows the transducer 506 to convert electrical signals into sound waves and convert sound waves into electrical signals (e.g., to transmit or receive acoustic signals). Exemplary monostatic transducers can include piezoelectric transducers, capacitive transducers, and micromachined ultrasonic transducers (MUTs) that use microelectromechanical systems (MEMS) technology.
[0055] Alternatively, the transducer 506 may be implemented in a bistatic topology, including multiple physically separate transducers, where a first transducer converts electrical signals into sound waves (e.g., transmits an acoustic signal) and a second transducer converts sound waves into electrical signals (e.g., receives an acoustic signal). An exemplary bistatic topology may be implemented using at least one speaker 508 and at least one microphone 510. The speaker 508 and microphone 510 may be dedicated to the audio plethysmograph 110 or may be used for both the audio plethysmograph 110 and other functions of the computing device 104 (e.g., presenting audible content to the user 106, capturing the user's 106 voice for phone calls, or for voice control).
[0056] Generally, the speaker 508 and the microphone 510 are aimed at (e.g., oriented toward) the ear canal 122. Thus, the speaker 508 can direct acoustic signals toward the ear canal 122, and the microphone 510 is responsive to receiving acoustic signals from a direction associated with the ear canal 122.
[0057] The hearable 102 includes at least one analog circuit 512, which includes circuitry and logic for conditioning electrical signals in the analog domain. The analog circuit 512 may include analog-to-digital converters, digital-to-analog converters, amplifiers, filters, mixers, and switches for generating and modifying the electrical signals. In some implementations, the analog circuit 512 includes other hardware circuitry associated with the speaker 508 or microphone 510.
[0058] The hearable 102 also includes at least one system processor 514 and at least one system medium 516 (e.g., one or more computer-readable storage media). In the illustrated configuration, the system medium 516 includes a pre-processing module 518 and a measurement module 520. The system medium 516 also optionally includes a calibration module 522. The pre-processing module 518, the measurement module 520, and the calibration module 522 can be implemented using hardware, software, firmware, or a combination thereof. In this example, the system processor 514 implements the pre-processing module 518, the measurement module 520, and the calibration module 522. In an alternative example, the computer processor 402 of the computing device 104 can implement at least a portion of the pre-processing module 518, the measurement module 520, and the calibration module 522. In this case, the hearable 102 can communicate digital samples of the acoustic signal to the computing device 104 using the communication interface 504.
[0059] The operation of the pre-processing module 518, the measurement module 520, and the calibration module 522 are further described with respect to Figures 7-10. Aspects of the fusion of audio plethysmography and motion sensing data 202 may be performed by the pre-processing module 518 and / or the measurement module 520, as further described with respect to Figures 11-1-12.
[0060] Some hearables 102 include active noise cancellation circuitry 524, which enables the hearable 102 to reduce background or environmental noise. In this case, the microphone 510 used for audio plethysmography 110 may be implemented using a feedback microphone in the active noise cancellation circuitry 524. For active noise cancellation, the feedback microphone provides feedback information about the performance of the active noise cancellation. During audio plethysmography 110, the feedback microphone receives an acoustic signal that is provided to the pre-processing module 518. In some situations, active noise cancellation and audio plethysmography 110 are performed simultaneously using the feedback microphone. In this case, the acoustic signal received by the feedback microphone may be provided to the pre-processing module 518 and the active noise cancellation circuitry 524.
[0061] The hearable 102 may also include at least one motion sensor 204. Exemplary motion sensors 204 include an inertial measurement unit (IMU), an accelerometer, an inclinometer, a gyroscope, a magnetometer, a global navigation satellite system, or some combination thereof. Generally, the motion sensor 204 can detect and / or measure one or more characteristics of motion. Some motion sensors 204 can, for example, measure linear acceleration and / or rotational rate (or angular rate), detect changes in orientation, detect changes in tilt, or some combination thereof. The linear acceleration and rotational rate can be associated with one, two, or three orthogonal axes. The motion sensor 204 generates motion sensing data 314 for audio plethysmography and motion sensing data fusion 202. The motion sensing data 314 can include time series data associated with the measured linear acceleration and / or rotational rate. Other types of motion sensing data 314 can include indications of changes in orientation and / or tilt, coordinates measured by a global navigation satellite system, etc. The motion sensing data 314 may include one or more of the movement characteristics described above.
[0062] Other implementations are possible where the motion sensor 204 is separate from the hearable 102. For example, the motion sensor 204 can be implemented within the computing device 104. As another example, the motion sensor 204 may be a separate device physically attached to the user 106 and communicatively coupled to the hearable 102 and / or the computing device 104. In this case, the audio plethysmography and motion sensing data fusion 202 can be designed to account for differences between the motion sensing data 314 provided by the motion sensor 204 and the movements observed by the hearable 102 using the audio plethysmograph 110. Different types of audio plethysmographs 110 are further described with respect to FIG. 6 .
[0063] Audioplethysmography 6 illustrates an exemplary operation of two hearables 102-1 and 102-2. In a first exemplary operation, the hearables 102-1 and 102-2 perform mono-aural audio plethysmography 110. This means that the hearables 102-1 and 102-2 independently perform audio plethysmography 110 on different ears 108 of the user 106. In this case, the first hearable 102-1 is adjacent to the right ear 108 of the user 106, and the second hearable 102-2 is adjacent to the left ear 108 of the user 106. Each hearable 102-1 and 102-2 includes a speaker 508 and a microphone 510. The hearables 102-1 and 102-2 can operate monostatically during the same or different periods. In other words, each hearable 102-1 and 102-2 can transmit and receive acoustic signals independently.
[0064] For example, the first hearable 102-1 transmits a first acoustic transmission 602-1 using the speaker 508, where the first acoustic transmission 602-1 propagates within at least a portion of the right ear canal 122 of the user 106. The first hearable 102-1 receives a first acoustic receive signal 604-1 using the microphone 510. The first acoustic receive signal 604-1 represents a version of the first acoustic transmit signal 602-1 that has been at least partially modified by acoustic circuitry associated with the right ear canal 122. This modification can change the amplitude, phase, and / or frequency of the first acoustic receive signal 604-1 compared to the first acoustic transmit signal 602-1.
[0065] Similarly, the second hearable 102-2 transmits a second acoustic transmit signal 602-2 using the speaker 508, where the second acoustic transmit signal 602-2 propagates within at least a portion of the left ear canal 122 of the user 106. The second hearable 102-2 receives the first acoustic receive signal 604-2 using the microphone 510. The second acoustic receive signal 604-2 represents a version of the second acoustic transmit signal 602-2 that has been modified by acoustic circuitry associated with the left ear canal 122. This modification can change the amplitude, phase, and / or frequency of the second acoustic receive signal 604-2 relative to the second acoustic transmit signal 602-2.
[0066] The technique of monocular audio plethysmography 110 can be particularly beneficial as it allows the computing device 104 to compile information from both hearables 102-1 and 102-2, which can further improve measurement reliability. For some aspects of audio plethysmography 110, it can be beneficial to analyze the acoustic channel between the two ears 108, as described further below.
[0067] In a second exemplary operation, the two hearables 102-1 and 102-2 perform binaural audio plethysmography 110. This means that the hearables 102-1 and 102-2 jointly perform audio plethysmography 110 across the two ears 108 of the user 106. In this case, at least one of the hearables 102 (e.g., the first hearable 102-1) includes a speaker 508, and at least one of the other hearables 102 (e.g., the second hearable 102-2) includes a microphone 510. The hearables 102-1 and 102-2 operate together bistatically during the same period.
[0068] In operation, the first hearable 102-1 transmits a third acoustic transmission 402-3 using the speaker 508. The third acoustic transmit signal 602-3 propagates through the right ear canal 122 of the user 106. The third acoustic transmit signal 602-3 also propagates through the acoustic channel that exists between the right ear 108 and the left ear 108. At the left ear 108, the third acoustic transmit signal 602-3 propagates through the left ear canal 122 of the user 106 and is represented as a third acoustic receive signal 604-3. The second hearable 102-2 receives the third acoustic receive signal 604-3 using the microphone 510. The third acoustic receive signal 604-3 represents a version of the third acoustic transmit signal 602-3 modified by acoustic circuitry associated with the right ear canal 122, modified by acoustic channels associated with the face of the user 106, and modified by acoustic circuitry associated with the left ear canal 122. This modification may change the amplitude, phase, and / or frequency of the third acoustic receive signal 604-3 relative to the third acoustic transmit signal 602-3. In some cases, the hearable 102-2 measures the time of flight (ToF) associated with propagation from the first hearable 102-1 to the second hearable 102-2. Sometimes, a combination of monaural and binaural audio plethysmography 110 is applied to improve the reliability of the measurements.
[0069] The acoustic transmit signal 602 in FIG. 6 can represent a variety of different types of signals. As described above with respect to FIG. 5, the acoustic transmit signal 602 may be an ultrasonic signal and / or an audible signal. The acoustic transmit signal 602 may also be a continuous wave signal (e.g., a sinusoidal signal) or a pulsed signal. Some acoustic transmit signals 602 may have a particular tone (or frequency). Other acoustic transmit signals 602 may have multiple tones (or multiple frequencies). Various modulations may be applied to generate the acoustic transmit signal 602. Exemplary modulations include linear frequency modulation, triangular frequency modulation, step frequency modulation, phase modulation, or amplitude modulation. The acoustic transmit signal 602 may be transmitted as part of a calibration or measurement procedure, as further described as part of FIG. 7.
[0070] FIG. 7 illustrates an exemplary implementation of a hearable 102 for performing audio plethysmography and motion sensing data fusion 202. In the illustrated configuration, the hearable 102 includes a speaker 508, a microphone 510, analog circuitry 512, a pre-processing module 518, a measurement module 520, a calibration module 522, and a motion sensor 204. However, other implementations of the hearable 102 are possible in which the hearable 102 does not include the calibration module 522 to reduce processing power requirements. In this case, the pre-processing module 518 may perform aspects of frequency selection to improve the signal-to-noise ratio of the audio plethysmography 110, as further described with respect to FIG. 10 . Some hearables 102 may not include a motion sensor 204 to save cost and / or reduce the footprint of the hearable 102. In this case, the motion sensor 204 may be communicatively coupled to the hearable 102 and implemented in another manner. For example, the motion sensor 204 may be implemented within the computing device 104 or as a separate entity that is physically attached to the user 106 .
[0071] The outputs of the speaker 508 and microphone 510 are coupled to inputs of an analog circuit 512. A pre-processing module 518 has inputs coupled to the output of the analog circuit 512 and the output of the motion sensor 204. The pre-processing module 518 also has outputs coupled to inputs of a measurement module 520 and a calibration module 522. The measurement module 520 has another input coupled to the motion sensor 204. The calibration module 522 has an output coupled to the speaker 508.
[0072] Consider an example operation of a hearable 102 with monocular audio plethysmography 110. If the hearable 102 includes a calibration module 522, the hearable 102 may perform a calibration process before performing a measurement process. The calibration process and the measurement process are further described with respect to FIG. 8.
[0073] During both the calibration and measurement processes, speaker 508 transmits acoustic transmit signal 602 and microphone 510 receives acoustic receive signal 604. During the calibration process, acoustic transmit signal 602 and acoustic receive signal 604 may include tones 702-1 through 702-M, where M represents a positive integer. During the measurement process, acoustic transmit signal 602 and acoustic receive signal 604 may include selected tones 704-1 through 704-N, where N represents a positive integer less than or equal to M. Selected tones 704-1 through 704-N may represent a subset (or even a proper subset) of tones 702-1 through 702-M.
[0074] The analog circuit 512 performs analog-to-digital conversion to generate a digital transmit signal 706 and a digital receive signal 708 based on the acoustic transmit signal 602 and the acoustic receive signal 604, respectively. The pre-processing module 518 performs frequency down-conversion and demodulation to generate at least one pre-processed signal 710 based on the digital transmit signal 706 and the digital receive signal 708. The pre-processing module 518 may also apply filtering to generate the pre-processed signal 710. For calibration and / or measurement procedures, the pre-processing module 518 may optionally apply motion artifact filtering 206 to generate the pre-processed signal 710 based on the digital receive signal 708 and the motion sensing data 314.
[0075] As part of the calibration procedure, the calibration module 522 processes the preprocessed signal 710 to determine selected tones 704-1 through 704-N. The selected tones 710-1 through 710-N can improve the performance of the audio plethysmograph 110 during a measurement procedure. The calibration module 522 communicates the selected tones 704-1 through 704-N to the speaker 508 using a control signal. The speaker 508 can accept the control signal identifying the selected tones 704-1 through 710-N and transmit a subsequent acoustic transmit signal 602 for the measurement procedure using the selected tones 704-1 through 704-N.
[0076] As part of the measurement procedure, the measurement module 520 may use the preprocessed signal 710 to perform aspects of audio plethysmography 110 to generate audio plethysmography data 712 (APG data 712). The audio plethysmography data 712 may be communicated to the audio plethysmography-based application 406. Additionally or alternatively, the measurement module 520 may use the preprocessed signal 710 and the motion sensing data 314 to perform audio plethysmography-based and motion sensing data fusion 202. In particular, the measurement module 520 may perform activity detection 208 and / or activity classification 210. Through activity detection 208 and / or activity classification 210, the audio plethysmography data 712 may include an indication of whether activity was detected and / or the type of activity detected. Additionally or alternatively, the audio plethysmography data 712 may include control signals for controlling operation of the hearable 102 and / or computing device 104. The calibration and measurement procedures are further described with respect to FIG.
[0077] FIG. 8 shows an example flow diagram 800 for operating the hearable 102. In FIG. 8 , the hearable 102 can optionally perform a calibration procedure at 802 using the calibration module 522. The calibration procedure can determine appropriate characteristics (e.g., waveform characteristics or signal characteristics) of the transmitted acoustic signal 602 to improve the audio plethysmograph 110 (e.g., to enhance the performance of the audio plethysmograph 110). The calibration procedure can determine a transmit frequency at which the audio plethysmograph 110 can improve accuracy performance, taking into account the fit of the hearable 102 (e.g., the position of the hearable 102 relative to the ear canal 122) and the physical structure of the ear canal 122. As an example, using the calibration procedure can enable the hearable 102 to determine biometrics of the user 106 with an accuracy of 95% or greater. The calibration procedure allows the hearable 102 to dynamically adjust its transmission frequency (e.g., one or more carrier frequencies) each time a seal 120 is formed (e.g., based on wearing the hearable 102) and based on the unique physical structure of the ear 108. Through this configuration procedure, hearables 102 on different ears 108 may operate at one or more different acoustic frequencies. The steps of the calibration procedure are further described below.
[0078] In some situations, the hearable 102 may perform on-head detection (or in-ear detection) by detecting the presence of the seal 120 and initiating a calibration procedure based on a determination that the on-head detection is "true." In other situations, the hearable 102 may initiate the calibration procedure based on a specified schedule or timer that the user 106 can control via the computing device 104.
[0079] At 804, the hearable 102 performs a calibration procedure by transmitting and receiving a first acoustic signal. The first acoustic signal propagates within at least a portion of the ear canal 122 of the user 106 and has multiple tones 702-1 through 702-M (or multiple carrier frequencies). The multiple tones 702-1 through 702-M are transmitted simultaneously or sequentially over a given time interval. The first acoustic transmit signal 602 may have a particular bandwidth on the order of about several kilohertz. For example, the acoustic transmit signal 602 may have a bandwidth of about 4, 5, 6, 8, 10, 16, or 20 kHz. In an exemplary implementation, the first acoustic transmit signal 602 is transmitted over multiple seconds, such as 2, 3, 4, 6, or more seconds. The duration of each tone 702 may be evenly divided over the total duration of the first acoustic transmit signal 602.
[0080] In an exemplary implementation, the acoustic transmit signal 602 has seven tones 702 (e.g., M equals 7). In some cases, the tones 702 are evenly distributed across the interval. For example, the tones 702 may be between 32 kHz and 38 kHz (e.g., approximately 32, 33, 34, 35, 36, 37, and 38 kHz) in 1 kHz increments. The term "approximately" means that the tones 702 may be within 5% of a given value or less (e.g., within 3%, 2%, or 1% of a given value).
[0081] The amplitude of the acoustic transmit signal 602 may be approximately the same across the tones 702-1 through 702-M. In this manner, power is distributed evenly across each of the tones 702. The amount (e.g., M) of the tones 702 may be determined based on the output power of the speaker 508. Increasing the amount of the tones 702 may increase the likelihood that the hearable 102 can support a given use case across a variety of conditions, including user fit and the physical structure of the ear canal 122 of the user 106. However, the amplitude of the acoustic transmit signal 602 may be limited across these tones 702 based on the output power of the speaker 508. Therefore, the amount of the tones 702 may be optimized based on the amount of output power available to the audio plethysmograph 110.
[0082] At 806, the calibration procedure selects one or more tones 704-1 through 704-N to be used in the measurement procedure based on one or more modified characteristics of the acoustic receive signal 604. The process for selecting the tones 704 is further described with respect to Figure 9. Generally, the calibration procedure determines that the selected tones 704 will improve the signal-to-noise ratio of the audio plethysmograph 110.
[0083] At 808, the hearable 102 performs a measurement procedure using the measurement module 520. In accordance with the measurement procedure, the hearable 102 transmits a second acoustic transmit signal 602 that propagates within at least a portion of the ear canal 122 of the user 106. If a calibration procedure has been performed, the second acoustic transmit signal 602 may include selected tones 704-1 through 704-N determined by the calibration procedure. The selected tones 704 may be transmitted simultaneously or sequentially over a given time interval.
[0084] The amplitude of the second acoustic transmit signal 602 may be approximately the same across the selected tones 704-1 through 704-M. In this manner, power is distributed evenly across each selected tone. Because the available output power is distributed across fewer tones, the amplitude of the second acoustic transmit signal 602 may be higher than the amplitude of the first acoustic transmit signal 602. Additionally or alternatively, the duration of each of the selected tones 704 of the second acoustic transmit signal 602 may be longer than the duration of the tones 702 of the first acoustic transmit signal 602. The higher amplitude and / or longer duration may further improve the signal-to-noise ratio performance of the hearable 102 for the audio plethysmograph 110. The measurement procedure may achieve greater accuracy for the audio plethysmograph 110 by using fewer selected tones 704 determined to improve signal-to-noise ratio performance.
[0085] At 812, the hearable 102 performs audio plethysmography and motion sensing data fusion 202 using the second acoustic signal (e.g., the second acoustic received signal 604) and the motion sensing data 314 provided by the motion sensor 204. This may include performing motion artifact filtering 206, activity detection 208, and / or activity classification 210, as further described with respect to FIG. 9. The calibration module 522 is further described with respect to FIG.
[0086] 9 illustrates an exemplary scheme implemented by the calibration module 522. In the illustrated configuration, the calibration module 522 implements a frequency selector that selects one or more tones 704 for the measurement procedure. In an exemplary implementation, the calibration module 522 includes at least one amplitude detector 902, at least one phase detector 904, at least one peak-to-average ratio (PAR) detector 906 (PAR detector 906), and at least one comparator 908. The operation of these components is further described below.
[0087] During the calibration procedure, the calibration module 522 accepts a preprocessed signal 710 from the preprocessing module 518, as described above with respect to Figure 7. The preprocessed signal 710 may include amplitude and / or phase information associated with a plurality of tones 702-1 through 702-M used to transmit a first acoustic signal, as described at 802 in Figure 8.
[0088] In this example, calibration module 522 uses amplitude detector 902 to extract amplitude 910 of preprocessed signal 710 and phase detector 904 to extract phase 912 of preprocessed signal 710. Alternatively, if the in-phase and quadrature components of preprocessed signal 710 are received separately, amplitude detector 902 and phase detector 904 can measure amplitude 910 and phase 912, respectively, based on the in-phase and quadrature components.
[0089] The peak-to-average ratio detector 906 measures peak-to-average ratios 914-1 through 914-2M for each of the tones 702-1 through 702-M and for each of the characteristics (e.g., amplitude 910 and phase 912). Generally, the peak-to-average ratio 914 represents the peak intensity within a frequency range of interest divided by the average intensity within this frequency range. For biometric monitoring 112, the frequency range may be, for example, between 0.58 and 3.3 Hz, corresponding to the likely range of human heart rates, which may be between 35 and 200 beats per minute. Other frequency ranges may be used for other use cases, including other biometrics, speech detection 114, chewing detection 116, gesture recognition 118, activity detection 208, and / or activity classification 210. A higher peak-to-average ratio 914 indicates better signal quality, or more generally, better signal-to-noise ratio performance.
[0090] In one aspect, comparator 908 can evaluate peak-to-average ratios 914-1 through 914-2M with respect to threshold 916. Threshold 916 can be set to a specific value, such as 4. In other cases, calibration module 522 can dynamically determine threshold 916 and update threshold 916 over time based on the observed peak-to-average ratios 914-1 through 914-2M. In an exemplary implementation, comparator 908 determines the tones 704-1 through 704-N selected for subsequent measurement procedures based on frequencies associated with peak-to-average ratios 914-1 through 914-2M that are greater than or equal to threshold 916.
[0091] Additionally or alternatively, comparator 908 can evaluate peak-to-average ratios 914-1 through 914-2M with respect to one another. In an exemplary embodiment, comparator 908 determines one of the selected tones 704 based on the frequency with the highest peak-to-average ratio 914 across amplitude 910. Comparator 908 can also determine one of the selected tones 704 based on the frequency with the highest peak-to-average ratio 914 across phase 912. In other embodiments, comparator 908 can determine a single selected tone 704 based on the frequency with the highest peak-to-average ratio 914 associated with either amplitude 910 or phase 912.
[0092] In general, the calibration module 522 allows for dynamically adjusting the selected tones 704-1 through 704-N prior to a measurement procedure based on the current environment, which may take into account the fit of the hearable 102 (e.g., current insertion depth and / or rotation), the physical structure of the ear canal 122 of the user 106, and the response characteristics of the hearable 102 (e.g., speaker, microphone, and / or housing). In this manner, the calibration module 522 may improve the signal-to-noise ratio performance of the hearable 102 for the measurement procedure. The calibration module 522 may also determine which tones 704 produce an acoustic receive signal 604 with characteristics desirable for a given use case. For example, in the case of biometric monitoring 112, the calibration module 522 may identify which tones 704 produce detectable cardiac modulation of amplitude 910 and / or phase 912.
[0093] 7-9, the calibration procedure and the measurement procedure are described as separate procedures occurring at different time intervals. In particular, the calibration procedure occurs before the measurement procedure. This allows the acoustic transmit signal 602 for the measurement procedure to be transmitted with fewer tones than the acoustic transmit signal 602 used for the calibration procedure, which may improve the signal-to-noise ratio performance of the audio plethysmograph 110. However, in some implementations, the hearable 102 may have sufficient output power to perform the measurement procedure with multiple tones 702-1 through 702-M using a single acoustic transmit signal 602. In this case, aspects of the calibration module may be integrated within the pre-processing module 518 as a frequency selector, as further described with respect to FIG. 10. This frequency selector may effectively pass selected tones 704-1 through 704-N for further processing. Aspects of the measurement procedure are further described with respect to FIG. 10.
[0094] Fusion of audioplethysmography and motion detection data 10 illustrates an exemplary scheme implemented by the hearable 102 to perform audio plethysmography and motion sensing data fusion 202. In the illustrated configuration, the hearable 102 includes a measurement module 520 and a pre-processing module 518 coupled to a calibration module 522. The pre-processing module 518 and / or the measurement module 520 are also coupled to the motion sensor 204 (not shown).
[0095] The pre-processing module 518 includes at least one in-phase and quadrature mixer 1002 (I / Q mixer 1002) and at least one filter 1004. The in-phase and quadrature mixer 1002 performs frequency down-conversion. In an exemplary embodiment, the in-phase and quadrature mixer 1002 includes at least two mixers: at least one phase adjuster and at least one combiner (e.g., a summing circuit). The filter 1004 attenuates intermodulation products generated by the in-phase and quadrature mixer 1002. In an exemplary embodiment, the filter 1004 is implemented using a low-pass filter.
[0096] The pre-processing module 518 optionally includes at least one frequency selector 1006 and / or at least one motion artifact filter 1008. The motion artifact filter 1008 may perform motion artifact filtering 206 to improve the performance of the audio plethysmograph 110. An exemplary implementation of the motion artifact filter 1008 is further described with respect to FIGS. 11-1 and 11-2.
[0097] The frequency selector 1006 can identify and select one or more tones 704 (or carrier frequencies) that provide a high-quality signal for subsequent processing. The frequency selector 1006 can further pass the selected tones 704 to other processing modules and filter (or attenuate) other unselected tones. The frequency selector 1006 can be implemented similarly to the calibration module 522 of FIG. 9. For example, the frequency selector 1006 can include an amplitude detector 902, a phase detector 904, a peak-to-average ratio detector 906, and a comparator 908.
[0098] The measurement module 520 may include at least one audio plethysmography module 1010 (APG module 1010), at least one activity detector 1012, at least one activity classifier 1014, or some combination thereof. The audio plethysmography module 1010 processes the information provided by the pre-processing module 518 to generate audio plethysmography data 712 for any of the described use cases, including biometric monitoring 112, speech detection 114, chewing detection 116, and / or gesture recognition 118.
[0099] For example, to measure heart rate variability as part of biometric monitoring 112, the audio plethysmography module 1010 may use peak detection estimation to identify the location of the peaks of each heartbeat within the preprocessed signal 710. This estimation may be performed across the amplitude and / or phase of the preprocessed signal 710. Exemplary peak detection estimation techniques include Z-score, local maximum, and divide and conquer. The audio plethysmography module 1010 may measure heart rate variability by calculating the root mean square of successive differences (RMSSD) between each peak (e.g., between each heartbeat).
[0100] In some cases, the audio plethysmography module 1010 may provide the audio plethysmography data 712 to other components of the measurement module 520, such as the activity detector 1012 or the activity classifier 1014. In this manner, the audio plethysmography data 712 can be used to improve activity detection 208 and / or activity classification 210. Consider an exemplary implementation in which the audio plethysmography module 1010 performs chewing detection 116 or gesture recognition 118 to determine that the user 106 is chewing or making a gesture. This information, as part of the audio plethysmography data 712, can be used by the activity detector 1012 in addition to the motion detection data 314 to determine that the user 106 is moving. In another exemplary implementation, the audio plethysmography module 1010 performs biometric monitoring 112 to measure the heart rate and / or respiratory rate of the user 106. This information, as part of the audio plethysmography data 712, can be used by the activity classifier 1014 in addition to the motion sensing data 314 to recognize the activity the user 106 is performing.
[0101] Generally, the activity detector 1012 and / or activity classifier 1014 process information provided by the pre-processing module 518 (or audio plethysmography module 1010) and the movement sensor 204 to generate audio plethysmography data 712. In particular, the activity detector 1012 performs activity detection 208, and the activity classifier 1014 performs activity classification 210.
[0102] In some cases, the audio plethysmography module 1010 may operate independently of the activity detector 1012 and / or activity classifier 1014. In other cases, the activity detector 1012 and / or activity classifier 1014 may pass information to the audio plethysmography module 1010, which may modify the operation of the audio plethysmography module 1010 or modify the audio plethysmography data 712 generated by the audio plethysmography module 1010. In still other cases, the audio plethysmography module 1010 may pass information to the activity detector 1012 and / or activity classifier 1014 to enhance activity detection 208 and / or activity classification 210. Exemplary implementations of the activity detector 1012 and activity classifier 1014 are further described with respect to FIG. 12 .
[0103] During operation, the in-phase and quadrature mixer 1002 uses a phase adjuster and two mixers to generate in-phase and quadrature components associated with the digital receive signal 708. In particular, the in-phase and quadrature mixer 1002 mixes the digital receive signal 708 with a first version of the digital transmit signal 706 having a zero-degree phase shift to generate the in-phase component. Additionally, the in-phase and quadrature mixer 1002 mixes the digital receive signal 708 with a second version of the digital transmit signal 706 having a 180-degree phase shift to generate the quadrature signal. This mixing operation downconverts the digital receive signal 708 from audio frequencies to baseband frequencies. The in-phase and quadrature mixer 1002 uses a combiner to combine the in-phase and quadrature components of the digital receive signal 708 to generate the downconverted signal 1016. The use of the in-phase and quadrature mixer 1002 can further improve the signal-to-noise ratio of the downconverted signal 1016 compared to other mixing techniques.
[0104] In this example, the downconverted signal 1016 represents a combination of the in-phase and quadrature components of the mixed-down digital received signal 708. In an alternative embodiment, the in-phase and quadrature mixer 1002 does not include a combiner and passes the in-phase and quadrature components separately to the filter 1004. In this manner, the in-phase and quadrature components propagate through the filter 1004 individually.
[0105] The filter 1004 generates a filtered signal 1018 based on the downconverted signal 1016. In particular, the filter 1004 filters the downconverted signal 1016 to attenuate spurious or undesired frequencies (e.g., intermodulation products), some of which may be associated with the operation of the in-phase and quadrature mixer 1002. In this example, the filtered signal 1018 represents a combination of the in-phase and quadrature components of the downconverted signal 1016. Alternatively, the filtered signal 1018 may represent separate or different in-phase and quadrature components that are passed separately to the frequency selector 1006 and / or the motion artifact filter 1008.
[0106] During the measurement procedure, the preprocessing module 518 may optionally apply a frequency selector 1006. The frequency selector 1006 passes tones that meet a quality threshold level for the performance of the audio plethysmograph 110. For example, the frequency selector 1006 passes tones 704 having an amplitude 910 and / or phase 912 where the peak-to-average ratio 914 is equal to or greater than a threshold 916. The resulting signal output by the frequency selector 1006 is represented by signal 1020. In some implementations, this signal 1020 is passed to the measurement module 520 and / or calibration module 522 as the preprocessed signal 710. In other implementations where the frequency selector 1006 is not implemented, the filtered signal 1018 may be passed to the measurement module 520 and / or calibration module 522 as the preprocessed signal 710.
[0107] The pre-processing module 518 may optionally apply a motion artifact filter 1008 during the measurement and / or calibration procedures. The motion artifact filter 1008 generates a denoised signal 1022 based on the input signal and the motion sensing data 314. The input signal may be the filtered signal 1018 or the signal 1020.
[0108] Generally speaking, the motion artifact filter 1008 uses the motion sensing data 314 to remove and / or significantly attenuate motion artifacts 312 present in the input signal. Due to this attenuation, in the de-noised signal 1022, the motion artifacts 312 may have amplitudes that are less than the amplitudes of desired signal components associated with the propagation of acoustic signals within the ear canal 122 of the user 106. Compared to the input signal, the de-noised signal 1022 has fewer motion artifacts 312 and / or motion artifacts 312 that are significantly smaller in amplitude.
[0109] In some implementations, the motion artifact filter 1008 can also generate a motion signal 1024. The motion signal 1024 includes the motion artifacts 312 of the input signal. The motion signal 1024 does not include signal components associated with propagation within the ear canal 122.
[0110] The audio plethysmography module 1010 may generate the audio plethysmography data 712 based on the noise-reduced signal 1022. By performing audio plethysmography 110 using the noise-reduced signal 1022 instead of the filtered signal 1018 or signal 1020, the audio plethysmography module 1010 may improve accuracy and reduce false positives in situations where the user 106 is moving or engaged in activity (e.g., low-energy activity and / or high-energy activity).
[0111] The activity detector 1012 can perform the activity detection 208 and / or the activity classifier 1014 can perform the activity classification 210 based on the preprocessed signal 710 and the motion sensing data 314 provided by the preprocessing module 518. The preprocessed signal 710 may represent a filtered signal 1018, a signal 1020 (as shown in FIG. 10 ), or a motion signal 1024 (as shown in FIG. 10 using dashed lines). As will be further described with respect to FIG. 12 , the use of the activity detector 1012 and / or the activity classifier 1014 can provide additional features to the hearable 102. Exemplary implementations of the motion artifact filter 1008 are further described with respect to FIGS. 11-1 and 11-2.
[0112] Motion Artifact Filtering FIG. 11-1 illustrates a first exemplary implementation of a motion artifact filter 1008 for performing audio plethysmography and motion sensing data fusion 202. In the illustrated configuration, the motion artifact filter 1008 includes at least one single-channel filter module 1100-1 that performs motion artifact filtering 206. The single-channel filter module 1100-1 may be implemented using at least one filter 1102 or at least one adaptive filter 1104. In operation, the single-channel filter module 1100-1 filters an input signal 1106 based on motion sensing data 314 to generate a denoised signal 1022. The input signal 1106 may include a filtered signal 1018 or a filtered signal 1020. The motion sensing data 314 includes at least three-axis linear acceleration 1108. Optionally, the motion sensing data 314 may include a rotational velocity 1110. Other implementations are possible in which other types of motion sensing data 314 are used in addition to or instead of linear acceleration 1108 .
[0113] In a first exemplary implementation, the single-channel filter module 1100-1 uses a filter 1102 to generate the denoised signal 1022. The filter 1102 can perform informed filtering based on the motion sensing data 314. Through informed filtering, the filter 1102 attenuates motion artifacts in the input signal 1106 that are also present in the motion sensing data 314. The motion artifacts 312 in the input signal 1106 can have frequencies related to frequencies associated with the motion artifacts 312 in the motion sensing data 314. Generally, these frequencies are based on frequencies associated with movement (e.g., the cadence of a user while walking). By referring to the motion sensing data 314, the filter 1102 can significantly attenuate frequencies associated with the motion artifacts 312 in the input signal 1106 and pass (or minimally attenuate) other frequencies in the input signal 1106 that are not associated with the motion artifacts 312. Stated another way, the filter 1102 represents a conventional filter, such as a low-pass filter, a high-pass filter, a band-pass filter, or some combination thereof, having at least one cutoff frequency that can be dynamically tuned (or dynamically adjusted) based on frequencies observed in the motion sensing data 314. In this case, the frequencies of interest for the audio plethysmograph 110 are different from the frequencies associated with possible motion artifacts 312. In this way, the frequencies of interest are not attenuated by the filter 1102.
[0114] In a second exemplary implementation, the single filter module 1100-1 uses an adaptive filter 1104 to generate the denoised signal 1022. Using adaptive filtering techniques, the input signal 1106 represents a primary reference and the motion sensing data 314 represents a noise reference. The adaptive filter 1104 uses the motion sensing data 314 to remove motion artifacts 312 in the input signal 1106. Various adaptive filtering techniques can be used, including techniques based on least mean squares (LMS) or recursive least squares (RLS).
[0115] The single channel filter module 1100-1 can operate with a single channel input signal, where the single channel input signal represents the filtered signal 1018 or signal 1020 associated with hearable 102-1 or hearable 102-2. Other multi-channel implementations of the motion artifact filter 1008 are also possible, as will be further described with respect to FIG.
[0116] 11-2 illustrates a second exemplary implementation of a motion artifact filter 1008 for performing audio plethysmography and motion detection data fusion 202. In the illustrated configuration, the motion artifact filter 1008 includes at least one multi-channel filter module 1100-2 that performs motion artifact filtering 206. The multi-channel filter module 1100-2 may be implemented using an adaptive filter 1104, at least one blind source separator 1112, or at least one machine-learned model 1114.
[0117] During operation, the multi-channel filter module 1100-2 accepts input signals 1106-1 and 1106-2 associated with the hearables 102-1 and 102-2, respectively. The input signals 1106-1 and 1106-2 may represent the filtered signal 1018 or the filtered signal 1020, respectively. The multi-channel filter module 1100-2 generates a noise-reduced signal 1022, a motion signal 1024, and optionally a multi-channel noise signal 1116 based on the input signals 1106-1 and 1106-2 and the motion sensing data 314. The motion sensing data 314 includes at least three-axis linear acceleration 1108. Optionally, the motion sensing data 314 may include a rotational velocity 1110. Other implementations are possible in which other types of motion sensing data 314 are used in addition to or instead of the linear acceleration 1108.
[0118] The noise-reduced signal 1022 may represent a single channel of signal components associated with the audio plethysmograph 110. For example, the noise-reduced signal 1022 may include signal components present in the input signal 1006-1 or the input signal 1006-2. The operation signal 1024 also represents a single channel of a movement artifact 312. For example, the operation signal 1024 may include a movement artifact 312 present in the input signal 1006-1 or the input signal 1006-2. The noise-reduced signal 1022 and the operation signal 1024 may be associated with the same channel (e.g., associated with the same hearable 102-1 or 102-2). The multi-channel noise signal 1116 may include other noise not associated with the movement artifact 312. In contrast to the noise-reduced signal 1022 and the operation signal 1024, the multi-channel noise signal 1116 may be associated with both the input signals 1106-1 and 1106-2.
[0119] The blind source separator 1112 applies a blind source separation (BSS) technique to separate the motion artifacts 312 from the signal components corresponding to the audio plethysmography 110. The blind source separator 1112 can use various transformation techniques, such as principal component analysis (PCA) with singular value decomposition (SVD).
[0120] The machine-learned model 1114 is implemented using one or more neural networks. A neural network includes a group of connected nodes (e.g., neurons or perceptrons) organized into one or more layers. As an example, the machine-learned model 1114 includes a deep neural network, which includes an input layer, an output layer, and one or more hidden layers disposed between the input and output layers. The nodes of a deep neural network may be partially or fully connected between layers.
[0121] In some implementations, the neural network is a recurrent neural network (e.g., a long short-term memory (LSTM) neural network) in which connections between nodes form cycles to retain information from previous portions of the input data sequence for subsequent portions of the input data sequence. In other cases, the neural network is a feedforward neural network in which connections between nodes do not form cycles. Additionally or alternatively, the machine-learned model 1114 includes another type of neural network, such as a convolutional neural network. The machine-learned model 1114 can also include one or more types of regression models, such as a single linear regression model, a multiple linear regression model, a logistic regression model, a stepwise regression model, a multivariate adaptive regression spline, a locally estimated scatterplot smoothing model, etc.
[0122] Generally, the machine-learned model 1114 is trained using supervised learning to generate the denoised signal 1022, the motion signal 1024, and the multi-channel noise signal 1116 based on the input signals 1106-1 and 1106-2. In this manner, the machine-learned model 1114 is trained to separate the movement artifacts 312 and signal components from the input signals 1106-1 and 1106-2. In some implementations, the machine-learned model 1114 may be further trained to perform one or more aspects of the audio plethysmography module 1010. Generally, supervised learning may use simulated (e.g., synthetic) data or measured (e.g., actual) data for training purposes.
[0123] In the adaptive filter 1104, blind source separator 1112, and / or machine-learned model 1114, the preprocessing module 518 may not include the frequency separator 1006, as this aspect can be handled by the motion artifact filter 1008. The measurement module 520 is further described with respect to FIG.
[0124] Activity detection and / or activity classification 12 shows an exemplary implementation of a measurement module 520 for performing audio plethysmography and motion sensing data fusion 202. In the illustrated configuration, the measurement module 520 includes an audio plethysmography module 1010, an activity detector 1012, and an activity classifier 1014. The activity detector 1012 is coupled to the audio plethysmography module 1010 and the activity classifier 1014.
[0125] The activity classifier 1014 is implemented using at least one machine-learned model 1202, which includes one or more neural networks (e.g., deep neural networks). The nodes of a deep neural network may be partially or fully connected between layers. In some implementations, the neural network is a recurrent neural network (e.g., a long short-term memory (LSTM) neural network), in which the connections between nodes form cycles to retain information from previous portions of the input data sequence for subsequent portions of the input data sequence. In other cases, the neural network is a feedforward neural network, in which the connections between nodes do not form cycles. Additionally or alternatively, the machine-learned model 1202 includes another type of neural network, such as a convolutional neural network.
[0126] The activity classifier 1014 may also include one or more types of classification models, such as a binary classification model, a multi-class classification model, a multi-label classification model, etc. Generally, the machine-learned model 1202 is trained using supervised learning to identify at least one activity type 1204 based on the preprocessed signal 710 and the motion sensing data 314.
[0127] The activity detector 1012 may be implemented using at least one filter, a comparison module, and / or a machine-learned module. In general, the activity detector 1012 may perform aspects of filtering, comparison, correlation, and / or classification to determine whether the user 106 is moving (e.g., to determine whether the user 106 is performing an activity). The machine-learned model used to implement the activity detector 1012 may be similar to the machine-learned model 1202. For example, the machine-learned model may be implemented using a classification model that classifies an input signal as either including or not including an artifact associated with movement.
[0128] In some implementations, the activity detector 1012 accepts a signal associated with the audio plethysmograph 110 (e.g., input signal 1206-2) and accepts motion sensing data 314. This may enable the activity detector 1012 to further determine whether the audio plethysmograph 110 is affected by movement of the user 106, which may be advantageous for situations where the motion sensor 204 is separate from the hearable 102. In other implementations, the activity detector 1012 may perform activity detection using only the motion sensing data 314, which may simplify processing for implementations in which the motion sensor 204 is integrated within the hearable 102.
[0129] During operation, the measurement module 520 accepts a preprocessed signal 710 from the preprocessing module 518. In FIG. 12 , the preprocessed signal 710 is represented as a first input signal 1206-1 provided to the audio plethysmography module 1010 and a second preprocessed signal 1206-2 provided to the activity detector 1012 and / or the activity classifier 1014. In various implementations, the input signals 1206-1 and 1206-2 may be the same signal or different signals, as described further below. Generally, the input signal 1206-1 includes at least a signal component associated with the propagation of an acoustic signal within the ear canal 122 of the user 106. This enables the input signal 1206-1 to be processed for any of the use cases associated with the audio plethysmography 110. In contrast, the input signal 1206-2 includes at least one motion artifact 312 associated with the movement of the user 106. This allows the input signal 1206-2 to be processed for activity detection 208 and / or activity classification 210.
[0130] In some implementations, input signals 1206-1 and 1206-2 represent the same signal, which may include filtered signal 1018 or signal 1020. In other implementations, input signals 1206-1 and 1206-2 may represent different signals. For example, first input signal 1206-1 may represent denoised signal 1022 for implementations in which pre-processing module 518 includes motion artifact filter 1008. In this case, second pre-processed signal 710-2 may represent filtered signal 1018 or signal 1020. For some implementations of motion artifact filter 1008, such as implementations including multi-channel filter module 1100-2 of FIG. 11-2, second pre-processed signal 710-2 may represent motion signal 1024.
[0131] As part of the operation of the measurement module 520, the audio plethysmography module 1010 may generate audio plethysmography data 712 (or a portion thereof) based on the input signal 1206-1. The audio plethysmography data 712 may include information associated with one or more characteristics of the acoustic receive signal 604 that is passed to the audio plethysmography module 1010 as part of the input signal 1206-1. In an exemplary implementation, the audio plethysmography data 712 may include information that supports biometric monitoring 112, speech detection 114, chewing detection 116, and / or gesture recognition 118.
[0132] Additionally or alternatively, the activity detector 1012 may generate audio plethysmography data 712 (or a portion thereof) based on the input signal 1206-2. The audio plethysmography data 712 may include information associated with one or more movement artifacts 312 present in the input signal 1206-2 and / or the motion sensing data 314. In an exemplary implementation, the audio plethysmography data 712 generated by the activity detector 1012 may include an activity detection indicator 1208 that identifies whether the user 106 has moved or engaged in an activity. The activity detection indicator 1208 may further be used to control the operation of the hearable 102 and / or the computing device 104.
[0133] In some implementations, the activity detector 1012 generates a control signal 1210 that can be provided to the audio plethysmography module 1010 and / or the activity classifier 1014. The control signal 1210 can dynamically enable and / or disable the audio plethysmography module 1010 and the activity classifier 1014 based on the activity detection indicator 1208.
[0134] Consider an example in which the preprocessing module 518 does not perform motion artifact filtering 206. In this case, the audio plethysmography module 1010 may operate on the filtered signal 1018 and / or the signal 1020, which may contain motion artifacts 312 that may adversely affect the performance of the audio plethysmography 110. To improve the performance of the audio plethysmography 110, the control signal 1210 may disable the audio plethysmography module 1010 if the activity detection indicator 1208 indicates that activity (or more generally, motion) is detected. In this case, the activity detection indicator 1208 may be communicated to the computing device 104, which may enable the computing device 104 to notify the user 106 that the audio plethysmography 110 is currently unavailable. The computing device 104 may also prompt the user 106 to reduce movement in order to enable the audio plethysmography 110. The control signal 1210 may also enable the audio plethysmography module 1010 when the activity detection indicator 1208 indicates that no activity is detected. In this manner, the activity detector 1012 controls when the audio plethysmography 110 is performed based on the input signal 1206-2 and / or the motion detection data 314.
[0135] Consider another example where the power and / or computational resources of the hearable 102 are limited. To conserve power and / or computational resources, the control signal 1210 can disable the activity classifier 1014 when the activity detection indicator 1208 indicates that no activity (or more generally, motion) is detected. The control signal 1210 can also enable the activity classifier 1014 when the activity detection indicator 1208 indicates that activity is detected. In some cases, the control signal 1210 can enable the activity classifier 1014 after the activity detection indicator 1208 indicates that activity has been occurring for a predetermined period of time.
[0136] As another part of the operation of the measurement module 520, the activity classifier 1014 may generate audio plethysmography data 712 (or a portion thereof) based on the input signal 1206-2 and the motion sensing data 314. The audio plethysmography data 712 may include information associated with one or more movement artifacts 312 present in the input signal 1206-2 and the motion sensing data 314. In exemplary implementations, the audio plethysmography data 712 generated by the activity classifier 1014 may include an activity type 1204. In some implementations, the activity type 1204 may indicate whether an activity performed by the user 106 is associated with a low-energy activity or a high-energy activity. In other implementations, the activity type 1204 may further identify a type of low-energy activity and / or a type of high-energy activity. For example, the activity type 1204 may indicate whether the user 106 is talking, reading, walking, or running.
[0137] The activity detection indicator 1208 and / or activity type 1204 can control the operation of the hearable 102 and / or the computing device 104. For example, either can appropriately adjust the volume of the hearable 102, present different audio content to the user 106, control the operation of the audio plethysmography-based application 406, etc.
[0138] Although not explicitly shown, the output of the activity classifier 1014 may, in some implementations, be coupled to an input of the audio plethysmography module 1010. In this case, the audio plethysmography module 1010 may use the activity type 1204 determined by the activity classifier 1014 to improve measurement accuracy and / or customize operation. For example, if the activity type 1204 indicates that the user 106 is engaged in high-energy activity and / or exercise, the audio plethysmography module 1010 may switch from performing gesture recognition 118 to performing biometric monitoring 112.
[0139] Additionally or alternatively, the audio plethysmography module 1010 may customize biometric measurements based on the activity type 1204 to simplify processing and / or improve accuracy. For example, if the activity type 1204 indicates that the user is exercising, the audio plethysmography module 1010 may adjust a filter to detect the user's 106 heart rate and / or breathing rate in a higher frequency range. Alternatively, if the activity type 1204 indicates that the user 106 is engaging in a low-energy activity, the audio plethysmography module 1010 may adjust a filter to detect the user's 106 heart rate and / or breathing rate in a lower frequency range.
[0140] Exemplary Methods 13, 14, and 15 illustrate exemplary methods 1300, 1400, and 1500 for implementing aspects of audio plethysmography and motion sensing data fusion 202. Methods 1300, 1400, and 1500 are illustrated as sets of operations (or acts) to be performed, but are not necessarily limited to the order or combination in which the operations are shown herein. Furthermore, any one or more of the operations may be repeated, combined, rearranged, or linked to provide various additional and / or alternative methods. In portions of the following description, reference may be made to the environments 200-1 through 200-5 of FIG. 2 and the entities detailed in FIGS. 4 and 5, with reference to FIGS. 4 and 5 being exemplary only. The techniques are not limited to being performed by one entity or multiple entities operating on one device.
[0141] 13, at 1302, an acoustic transmit signal propagating within at least a portion of a user's ear canal is transmitted and received during a first time period. The received acoustic signal represents a version of the transmitted acoustic signal having one or more characteristics modified based on propagation within the ear canal. The received acoustic signal includes at least one motion artifact associated with the user moving during at least a portion of the first time period.
[0142] For example, the hearable 102 transmits and receives acoustic signals propagating within at least a portion of the ear canal 122 of the user 106 during a first time period, as shown in FIG. 6 . Stated another way, the hearable 102 transmits an acoustic transmit signal 602 and receives an acoustic receive signal 604 during a first time period. The transmission and reception can be performed in a half-duplex or full-duplex manner, depending on the implementation. In some cases, a monaural audio plethysmograph 110 can be used to transmit and receive acoustic signals using at least one of the hearables 102-1 or 102-2 of FIG. 6 . In other cases, a binaural audio plethysmograph 110 can be used to transmit acoustic signals using hearable 102-1 and receive acoustic signals using hearable 102-2.
[0143] The acoustic signal may be referred to as the acoustic transmit signal 602 or the acoustic receive signal 604, depending on the context. The received acoustic signal (e.g., the acoustic receive signal 604) represents a version of the transmitted acoustic signal (e.g., the acoustic transmit signal 602) with one or more characteristics modified based on propagation within the ear canal 122. The one or more modified characteristics may include amplitude, phase, and / or frequency. The received acoustic signal also includes at least one motion artifact 312 associated with the user 106 moving during at least a portion of the first time period. The motion artifact 312 may represent a portion of the received acoustic signal whose amplitude, phase, and / or frequency are based on the movement of the user 106.
[0144] The version of the received acoustic signal may represent a digital version of the acoustic receive signal 604, a version of the acoustic receive signal 604 that has been downconverted to a particular frequency range for further processing (e.g., downconverted to baseband frequencies), a preprocessed version of the acoustic receive signal 604, or some combination thereof. In general, the term "version of the received acoustic signal" means that there may be some difference between the received acoustic signal and the acoustic receive signal 604 that is provided to the audio plethysmography and motion sensing data fusion process. Stated another way, the version of the acoustic receive signal 604 is related to the acoustic receive signal 604 but can be further modified to support operation of the hearable 102, where operation involves audio plethysmography and motion sensing data fusion 202.
[0145] Generally, the movements of the user 106 during the first time period correspond to the user 106 moving at least his or her head. The movement artifact 312 may be associated with the movements of the user 106's head, such as translational movements, rotational movements, tilting movements, and / or changes in orientation. Exemplary head movements may include the head moving side to side, up and down, tilting, etc. The movement artifact 312 may also be associated with other movements the user 106 makes when engaging in a particular activity, such as the low-energy activities described with respect to environments 200-1 through 200-3 or the high-energy activities described with respect to environments 200-4 and 200-5. For example, the movement artifact 312 (or another movement artifact 312) may be associated with the movements of the user 106's arms and / or legs.
[0146] At 1304, motion sensing data generated by a motion sensor during a first time period is received. For example, the motion sensor 204 generates the motion sensing data 314 during the first time period. The pre-processing module 518 and / or the measurement module 520 receives (e.g., receives) the motion sensing data 314. This enables the pre-processing module 518 and / or the measurement module 520 to perform aspects of the audio plethysmography and motion sensing data fusion 202.
[0147] At 1306, a version of the received acoustic signal is processed based on the motion sensing data. For example, the hearable 102 performs audio plethysmography and motion sensing data fusion 202 by processing a version of the received acoustic signal based on the motion sensing data 314. More specifically, the hearable 102 may perform motion artifact filtering 206, activity detection 208, and / or activity classification 210, as shown in FIGS. 2 and 10 .
[0148] At 1308, data is generated based on the processing. The data is associated with at least one of the following: one or more characteristics associated with propagation of an acoustic signal within the ear canal or at least one movement artifact associated with user movement. For example, the hearable 102 generates audio plethysmography data 712 based on the processing, as shown in FIG. 7 . If at least a portion of the audio plethysmography data 712 is generated by the audio plethysmography module 1010, the audio plethysmography data 712 can be associated with one or more characteristics of the acoustic receive signal 604, the one or more characteristics associated with propagation of the acoustic transmit signal 602 within the ear canal 122. More specifically, the audio plethysmography data 712 can include information associated with biometric monitoring 112, speech detection 114, chewing detection 116, and / or gesture recognition 118. Additionally or alternatively, if at least a portion of the audio plethysmography data 712 is generated by the activity detector 1012 or the activity classifier 1014, the audio plethysmography data 712 may be associated with at least one movement artifact 312.
[0149] 14 , at 1402, a version of a received acoustic signal having one or more characteristics associated with propagation within a user's ear canal during a first time period is accepted. The received acoustic signal also includes at least one motion artifact associated with the user moving during at least a portion of the first time period. For example, the preprocessing module 518 or the measurement module 520 accepts a version of the acoustic receive signal 604 having one or more characteristics associated with propagation within the ear canal of the user 106 during the first time period. The received acoustic signal includes at least one motion artifact 312 associated with the user 106 moving during at least a portion of the first time period.
[0150] At 1404, motion sensing data generated by the motion sensor during the first time period is accepted. For example, the pre-processing module 518 or the measurement module 520 accepts the motion sensing data 314 generated by the motion sensor 204 during the first time period, as described above with respect to 1304 of FIG.
[0151] At 1406, fusion of audio plethysmography and motion sensing data is performed by processing a version of the received acoustic signal based on the motion sensing data. For example, the hearable 102 performs audio plethysmography and motion sensing data fusion 202 by processing a version of the received acoustic signal 604 based on the motion sensing data 314. One or more of the steps described at 1408, 1410, and 1412 may be performed to perform aspects of audio plethysmography and motion sensing data fusion 202.
[0152] Optionally, at 1408, motion artifact filtering is performed based on the received acoustic signal and the motion sensing data. For example, the motion artifact filter 1008 performs motion artifact filtering 206 based on the input signal 1106 and the motion sensing data 314, as shown in FIG. 10. Various implementations of the motion artifact filter 1008 may process an input signal 1106 associated with a single channel (or a single hearable 102), as described with respect to FIG. 11-1, or may process multiple input signals 1106-1, 1106-2 associated with multiple channels (or multiple hearables 102-1, 102-2), as described with respect to FIG. 11-2.
[0153] Optionally, activity detection based on the received acoustic signal and the motion sensing data is performed at 1410. For example, the measurement module 520 performs activity detection 208 based on the preprocessed signal 710 and the motion sensing data 314, as shown in FIGS.
[0154] Optionally, activity classification is performed based on the received acoustic signal and the motion sensing data at 1412. For example, the measurement module 520 performs activity classification 210 based on the preprocessed signal 710 and the motion sensing data 314, as shown in FIGS.
[0155] In some situations, methods 1300 and / or 1400 are performed using one hearable 102 for a monocular audio plethysmograph 110. In other situations, methods 1300 and / or 1400 are performed using two hearables 102 for a binaural audio plethysmograph 110.
[0156] 15, an acoustic transmit signal propagating within at least a portion of a user's ear canal is transmitted during a first time period. For example, the hearable 102 transmits the acoustic transmit signal 602 during the first time period. The acoustic transmit signal 602 propagates within at least a portion of the ear canal 122 of the user 106, as shown in FIG.
[0157] At 1504, an acoustic receive signal representing a version of the acoustic transmit signal having one or more characteristics modified based on propagation within the ear canal is received during a first time period, the acoustic receive signal also including at least one motion artifact associated with a user moving during at least a portion of the first time period.
[0158] For example, the hearable 102 receives an acoustic receive signal 604 during a first time period. The acoustic receive signal 604 represents a version of the acoustic transmit signal 602 with one or more characteristics modified based on propagation within the ear canal 122. The one or more modified characteristics may include amplitude, phase, and / or frequency. The acoustic receive signal 604 also includes at least one movement artifact 312 associated with the user 106 moving during at least a portion of the first time period. The movement artifact 312 may represent a portion of the acoustic receive signal 604 whose amplitude, phase, and / or frequency are affected by the movement of the user 106.
[0159] In some cases, a monocular audio plethysmograph 110 can be used to transmit and receive acoustic signals using at least one of hearables 102-1 or 102-2 of Figure 6. In other cases, a binaural audio plethysmograph 110 can be used to transmit acoustic signals using hearable 102-1 and receive acoustic signals using hearable 102-2.
[0160] At 1506, a denoised version of the acoustic receive signal is generated using motion detection data generated by the motion sensor during the first time period. For example, the motion artifact filter 1008 uses the motion detection data 314 to generate the denoised signal 1022. The motion sensor 204 generates the motion detection data 314 during the first time period. The motion artifact filter 1008 may generate the denoised signal 1022 based on the motion detection data 314 using various techniques, such as informed filtering, adaptive filtering, or blind source separation, as described with respect to FIGS. 11-1 and 11-2 . The denoised signal 1022 represents a version of the acoustic receive signal 604 in which the motion artifacts 312 are attenuated (e.g., significantly attenuated, or at least 3 decibels) relative to the motion artifacts 312 in the acoustic receive signal 604.
[0161] At 1508, operation of the hearable and / or a computing device coupled to the hearable is controlled based on a denoised version of the acoustic received signal. For example, the audio plethysmography module 1010 can generate audio plethysmography data 712 based on the denoised signal 1022. The audio plethysmography data 712 can be used to control operation of the hearable 102 and / or operation of the computing device 104. Consider an example in which the audio plethysmography module 1010 performs biometric monitoring 112. In this case, the audio plethysmography module 1010 analyzes the denoised signal 1022 to measure one or more biometrics of the user 106. Exemplary biometrics may include the user's 106 heart rate, heart rate variability, respiratory rate, and / or blood pressure. The operation of the hearable 102 and / or computing device 104 can monitor the biometric and / or communicate the biometric to the user 106.
[0162] In another example, the audio plethysmography module 1010 performs speech detection 114. By analyzing the noise-reduced signal 1022, the audio plethysmography module 1010 can analyze the noise-reduced signal 1022 to determine whether the user 106 spoke during a first time period. If the user's 106 voice is detected, the hearable 102 and / or computing device 104 can enable a voice control interface and / or provide multi-factor voice authentication. If the user's 106 voice is not detected, the hearable 102 and / or computing device 104 can disable the voice control interface and / or disable voice authentication.
[0163] Consider another example in which the audio plethysmography module 1010 performs chewing detection 116. In particular, the audio plethysmography module 1010 analyzes the noise-reduced signal 1022 to determine that the user 106 is grinding their teeth. The hearable 102 and / or computing device 104 can monitor how frequently the teeth grinding is occurring and communicate this information to the user 106 and / or sound an alarm to notify the user 106.
[0164] In yet another example, the audio plethysmography module 1010 performs gesture recognition 118. In this case, the audio plethysmography module 1010 analyzes the noise-reduced signal 1022 to recognize gestures made by the user 106. The recognized gestures can be mapped to specific controls on the hearable 102 and / or computing device 104. Example controls may include changing the volume, advancing a playlist, changing the content displayed via the computing device 104, enabling or disabling voice control, etc.
[0165] Throughout this disclosure, the term "version of the signal" is used to indicate that the second signal may be a modified version of the first signal. A "version of the signal" (or second signal) may represent a digital version of the first signal, an analog version of the first signal, an electrical version of the first signal having voltage and current, an acoustic version of the first signal having acoustic characteristics, a downconverted version of the first signal having a lower frequency range, an upconverted version of the signal having a higher frequency range, a preprocessed version of the signal (e.g., a version whose amplitude, phase, and / or frequency have been modified in some way), a filtered version of the signal, etc. In general, a "version of the signal" refers to a signal that has been modified using techniques known in the art to facilitate operation of the hearable 102.
[0166] Exemplary Computing System FIG. 16 illustrates various components of an exemplary computing system 1600 that may be implemented as any type of client, server, and / or computing device as described with reference to FIGS. 4 and 5 above to implement aspects of audio plethysmography and motion sensing data fusion 202.
[0167] The computing system 1600 includes a communication device 1602 that enables wired and / or wireless communication of device data 1604 (e.g., received data, data being received, data scheduled for broadcast, or data packets of data). The communication device 1602 or computing system 1600 can include one or more hearables 102 and at least one motion sensor 204. The device data 1604 or other device content can include device configuration settings, media content stored on the device, and / or information associated with a user of the device. The media content stored on the computing system 1600 can include any type of audio, video, and / or image data. The computing system 1600 also includes one or more data inputs 1606, through which any type of data, media content, and / or input can be received, examples of which include human speech, user-selectable input (explicit or implicit), messages, music, television media content, recorded video content, and any other type of audio, video, and image data received from any content and / or data source.
[0168] Computing system 1600 also includes a communication interface 1608, which may be implemented as any one or more of a serial and / or parallel interface, a wireless interface, any type of network interface, a modem, and any other type of communication interface. Communication interface 1608 provides a connection and / or communication link between computing system 1600 and a communication network through which other electronic, computing, and communication devices communicate data with computing system 1600.
[0169] Computing system 1600 includes one or more processors 1610 (e.g., any of a microprocessor, controller, etc.) that process various computer-executable instructions to control the operation of computing system 1600. Alternatively or additionally, computing system 1600 may be implemented in any one or combination of hardware, firmware, or fixed logic circuitry implemented in association with processing and control circuitry generally identified at 1612. Although not shown, computing device 1600 may include a system bus or data transfer system that couples various components within the device. The system bus may include any one or combination of various bus structures, examples of which include a memory bus or memory controller, a peripheral bus, a universal serial bus, and / or a processor bus or local bus utilizing any of a variety of bus architectures.
[0170] Computing system 1600 also includes a computer-readable medium 1614, such as one or more memory devices that allow for persistent and / or non-transitory data storage (i.e., as opposed to mere signal transmission), examples of which include random access memory (RAM), non-volatile memory (e.g., one or more of read-only memory (ROM), flash memory, EPROM, EEPROM, etc.), and disk storage devices. The disk storage devices may be implemented as any type of magnetic or optical storage device, examples of which include hard disk drives, recordable and / or rewritable compact discs (CDs), any type of digital versatile disc (DVD), etc. Computing system 1600 also includes a mass storage media device (storage medium) 1616.
[0171] The computer-readable medium 1614 provides a data storage mechanism for storing device data 1604, as well as various device applications 1618, and any other information and / or data related to operational aspects of the computing system 1600. For example, an operating system may be maintained as a computer application using the computer-readable medium 1614 and executed on the processor 1610. The device applications 1618 may include any type of control application, software application, signal processing and control module, code specific to a particular device, a device manager, a hardware abstraction layer for a particular device, etc.
[0172] The device applications 1618 also include any system components, engines, or managers for performing the audio plethysmography and motion sensing data fusion 202. In this example, the device applications 1618 include the motion artifact filter 1008, activity detector 1012, and activity classifier 1014 of Figure 10. Although not explicitly shown, the device applications 1618 may also include the audio plethysmography-based application 406 (APG-based application 406) of Figure 4.
[0173] Throughout this disclosure, examples are described in which a computing system 1600 (e.g., a hearable 102, a computing device 104, a client device, a server device, a computer, or another type of computing system) may analyze information associated with a user (e.g., various audible and / or ultrasonic signals), such as the vocalizations 212 mentioned with respect to FIG. 2 . In addition to the above, the user 106 may be given controls that allow the user 106 to select both whether and when the systems, programs, and / or features described herein may enable collection of information (e.g., information about the user's social network, social actions, social activities, occupation, user preferences, and current location), as well as whether the user is sent content or communications from a server. The computing system 1600 may be configured to use information only after the computing system 1600 receives explicit permission from the user 106 to use the data. For example, in situations where the hearable 102 analyzes signals for biometric monitoring 112, speech detection 114, chewing detection 116, gesture recognition 118, activity detection 208, and / or activity classification 210, the individual user 106 may be given the opportunity to provide input to control whether programs or features of the computing system 1600 can collect and utilize the data. Additionally, the individual user 106 may have constant control over what the programs can or cannot do with the information.
[0174] Additionally, collected information may be preprocessed in one or more ways before being transferred, stored, or otherwise used so that personally identifiable information is removed. For example, before computing system 1600 shares data with another device, the identity of user 106 may be processed so that personally identifiable information cannot be determined about user 106. Thus, user 106 may control whether information is collected about user 106 and their device, and if collected, how such information may be used by computing system 1600 and / or a remote computing system.
[0175] conclusion Although techniques using, and devices including, audio plethysmography and motion sensing data fusion have been described in language specific to features and / or methods, it will be understood that the subject matter of the appended claims is not necessarily limited to the particular features or methods described. Rather, the particular features and methods are disclosed as exemplary implementations of audio plethysmography and motion sensing data fusion.
[0176] Some examples are provided below.
[0177] Example 1: A method comprising: transmitting and receiving, during a first time period, an acoustic signal propagating within at least a portion of a user's ear canal, the received acoustic signal representing a version of the transmitted acoustic signal having one or more characteristics modified based on the propagation within the ear canal, the received acoustic signal including at least one movement artifact associated with the user moving during at least a portion of the first time period, the method further comprising: accepting motion sensing data generated by a motion sensor during the first time period; processing a version of the received acoustic signal based on the motion sensing data; generating data associated with at least one of the one or more characteristics associated with the propagation of the acoustic signal within the ear canal or the at least one movement artifact associated with the user moving based on the processing; and A method comprising:
[0178] Example 2: The processing of the version of the received acoustic signal includes generating a de-noised signal by filtering the version of the received acoustic signal based on the motion sensing data, the de-noised signal having attenuated motion artifacts compared to the motion artifacts in the version of the received acoustic signal, and the de-noised signal including the one or more characteristics; generating the data includes generating the data based on the denoised signal. The method described in Example 1.
[0179] Example 3: The generating the noise reduction comprises: performing informed filtering to attenuate the motion artifacts in the version of the received acoustic signal that are also present in the motion sensing data; performing adaptive filtering using the received acoustic signal representing a primary criterion and the motion sensing data representing a noise criterion; The method of Example 2, comprising:
[0180] Example 4: The transmitting and receiving of the acoustic signal during the first period of time includes: transmitting and receiving a first acoustic signal using a first hearable; transmitting and receiving a second acoustic signal using a second hearable; Including, 3. The method of claim 2, wherein generating the denoised signal includes performing adaptive filtering or blind source separation based on a version of the received first acoustic signal, a version of the received second acoustic signal, and the motion sensing data to generate the denoised signal and a motion signal including the at least one motion artifact.
[0181] Example 5: The transmitting and receiving of the acoustic signal during the first period of time includes: transmitting and receiving a first acoustic signal using a first hearable; transmitting and receiving a second acoustic signal using a second hearable; Including, The receiving of the motion sensing data includes: receiving first motion sensing data generated by a first motion sensor of the first hearable during the first time period; receiving second motion sensing data generated by a second motion sensor of the second hearable during the first period; Including, generating the noise-reduced signal comprises: generating a first noise reduction signal based on the first acoustic signal and the first motion sensing data; generating a second noise reduction signal based on the second acoustic signal and the second motion sensing data; Including, generating the data includes generating the data based on the first de-noised signal and the second de-noised signal. The method described in Example 2.
[0182] Example 6: The transmitting and receiving of the acoustic signal during the first period of time includes: transmitting and receiving a first acoustic signal using a first hearable; transmitting and receiving a second acoustic signal using a second hearable; Including, generating the denoised signal includes providing a version of the received first acoustic signal, a version of the received second acoustic signal, and the motion sensing data as inputs to a machine-learned model to generate the denoised signal and the motion signal. The method described in Example 2.
[0183] Example 7: The method according to any one of Examples 2 to 6, wherein the generating the data includes measuring a biometric of the user based on the denoised signal.
[0184] Example 8: A method as in any preceding example, wherein processing the version of the received acoustic signal includes determining that the user moved during the portion of the first period based on a correlation between the at least one movement artifact in the received acoustic signal and the motion detection data.
[0185] Example 9: A method as described in any preceding example, wherein processing the version of the received acoustic signal includes classifying a type of activity the user engaged in during the first period based on the at least one movement artifact in the received acoustic signal and the motion detection data.
[0186] Example 10: The method of example 9, wherein said classifying the type of activity includes determining that the type of activity is relatively still or associated with a significant amount of movement.
[0187] Example 11: a determination that the user has moved; or A classification of the type of activity performed by the user controlling an operation of the hearable or an operation of a computing device coupled to the hearable based on at least one of The method of any one of Examples 8 to 10, further comprising:
[0188] Example 12: The motion detection data Linear accelerations associated with three orthogonal axes, or the rotational speeds associated with the three orthogonal axes 10. The method of any preceding embodiment, comprising at least one of:
[0189] Example 13: The method further includes transmitting and receiving, prior to the first period of time, another acoustic signal propagating within at least the portion of the ear canal of the user, the other acoustic signal having a plurality of tones, and the received other acoustic signal representing a version of the transmitted other acoustic signal having one or more characteristics modified by the propagation within the ear canal, the method comprising: selecting a subset of the plurality of tones based on the received other acoustic signals; further comprising said transmitting and receiving said acoustic signal during said first time period includes transmitting and receiving said acoustic signal having said subset of said plurality of tones. The method of any preceding example.
[0190] Example 14: The method of any preceding example, wherein the acoustic signal comprises an ultrasonic signal having a frequency between about 20 and 96 kilohertz.
[0191] Example 15: Transmitting audible content during at least the portion of the first time period. The method of any preceding example, further comprising:
[0192] Example 16: The method of any preceding example, wherein the generated data is used to control operation of a hearable and / or a computing device.
[0193] Example 17: The method of example 16, wherein the generated data is used to customize biometric measurements of the user.
[0194] Example 18: The method of claim 17, wherein the generated data is used to determine whether to adjust a filter for detecting the user's heart rate and / or the user's respiratory rate based on the received acoustic signal.
[0195] Example 19: A computer-readable storage medium comprising instructions that, in response to execution by a processor, cause a hearable to perform any one of the methods described in Examples 1-18.
[0196] Example 20: A device comprising: at least one transducer; at least one processor; Equipped with configured to perform any one of the methods described in Examples 1 to 18 using the at least one transducer and the at least one processor. device.
[0197] Example 21: A speaker, an active noise cancellation circuit comprising a feedback microphone; Furthermore, the at least one transducer comprises the speaker and the feedback microphone; The device of Example 20.
[0198] Example 22: The at least one transducer comprises a speaker and a microphone; the speaker is configured to be placed proximate to a first ear of a user; the microphone is configured to be placed proximate to a second ear of the user. The device of Example 20.
[0199] Example 23: The device of any one of Examples 20 to 22, further comprising a motion sensor.
[0200] Example 24: A device described in any one of Examples 20 to 23, wherein the device comprises at least one earphone.
[0201] Example 25: A method comprising: transmitting an acoustic transmit signal that propagates within at least a portion of the user's ear canal during a first period of time; and receiving, during the first time period, an acoustic receive signal representing a version of the acoustic transmit signal having one or more characteristics modified by the propagation within the ear canal, the acoustic receive signal including at least one movement artifact associated with the user moving during at least a portion of the first time period, the method further comprising: generating a denoised version of the acoustic receive signal using motion sensing data generated by a motion sensor during the first time period, wherein the at least one motion artifact in the denoised version of the acoustic receive signal is attenuated relative to the at least one motion artifact in the acoustic receive signal, the method further comprising: controlling operation of a hearable and / or a computing device coupled to the hearable based on the denoised version of the acoustic receive signal; A method comprising:
[0202] Example 26: Measuring at least one biometric of the user based on the denoised version of the acoustic received signal. further comprising wherein said controlling the operation of said hearable and / or said computing device includes communicating said at least one biometric. The method described in Example 25.
[0203] Example 27: The at least one biometric is: the user's heart rate; the user's heart rate variability; the user's respiratory rate, or blood pressure 27. The method of Example 26, comprising at least one of:
[0204] Example 28: Determining that the user has spoken during the first period based on the noise-removed signal; further comprising The controlling the operation of the hearable and / or the computing device includes: enabling a voice control interface based on said determination; or providing multi-factor voice authentication based on said determination; 10. The method of any preceding embodiment, comprising at least one of:
[0205] Example 29: Recognizing a gesture made by the user during the first period based on the noise-removed version of the acoustic received signal. further comprising controlling the operation of the hearable and / or the computing device includes controlling the operation of the hearable and / or the computing device based on the recognized gesture. The method of any preceding example.
[0206] Example 30: Determining that the user is grinding their teeth during the first period based on the noise-removed version of the acoustic receive signal. further comprising controlling the operation of the hearable and / or the computing device includes controlling the operation of the hearable and / or the computing device based on the determination. The method of any preceding example.
[0207] Example 31: A method as described in any preceding example, wherein generating the noise-reduced version of the acoustic receive signal includes adjusting a passband of a filter based on the motion detection data such that the filter passes frequencies of the version of the acoustic receive signal corresponding to biometric monitoring, speech detection, chewing detection, and / or gesture recognition.
[0208] Example 32: A method described in any one of Examples 25 to 30, wherein generating the noise-removed version of the acoustic receive signal includes performing adaptive filtering using the acoustic receive signal representing a primary reference and the motion detection data representing a noise reference.
[0209] Example 33: Transmitting a second acoustic transmit signal during the first period of time, the second acoustic transmit signal propagating within at least a portion of a second ear canal of the user; receiving, during the first time period, a second acoustic receive signal representing a version of the second acoustic transmit signal having one or more characteristics modified by the propagation within the second ear canal, the second acoustic receive signal including at least one second movement artifact associated with the user moving during at least the portion of the first time period; generating the denoised version of the acoustic receive signal includes performing adaptive filtering or blind source separation using the version of the acoustic receive signal, the second version of the acoustic receive signal, and the motion sensing data to generate the denoised signal and a motion signal including the at least one motion artifact. The method of any preceding example.
[0210] Example 34: Determining that the user moved during at least the portion of the first time period based on the movement signal. The method of Example 33, further comprising:
[0211] Example 35: Classifying a type of activity performed by the user during the first period based on the motion signal. The method of Example 33 or 34, further comprising:
[0212] Example 36: The method of Example 35, wherein classifying the type of activity includes determining that the type of activity is relatively sedentary or associated with a significant amount of movement.
[0213] Example 37: Transmitting a second acoustic transmit signal during the first period of time, the second acoustic transmit signal propagating within at least a portion of a second ear canal of the user; receiving, during the first time period, a second acoustic receive signal representing a version of the second acoustic transmit signal having one or more characteristics modified by the propagation within the second ear canal, the second acoustic receive signal including at least one second movement artifact associated with the user moving during at least the portion of the first time period, the method comprising: generating a noise-reduced version of the second received acoustic signal using other motion sensing data generated by a second motion sensor during the first time period; further comprising controlling the operation of the hearable and / or the computing device includes controlling the operation of the hearable and / or the computing device based on the noise-removed version of the acoustic receive signal and the second acoustic receive signal. The method according to any one of Examples 25 to 32.
[0214] Example 38: The motion detection data Linear accelerations associated with three orthogonal axes, or the rotational speeds associated with the three orthogonal axes The method of any one of Examples 25 to 37, comprising at least one of:
[0215] Example 39: The method further comprising transmitting, prior to the first period of time, another acoustic transmit signal that propagates within at least the portion of the ear canal of the user, the other acoustic transmit signal having a plurality of tones; receiving, prior to the first period of time, another acoustic receive signal representing a version of the other acoustic transmit signal having one or more characteristics modified by the propagation within the ear canal; selecting a subset of the plurality of tones based on a quality metric associated with the plurality of tones; further comprising transmitting the acoustic transmit signal during the first time period includes transmitting the acoustic transmit signal having the subset of the plurality of tones. The method according to any one of Examples 25 to 38.
[0216] Example 40: The method described in Examples 25 to 39, wherein the acoustic transmission signal comprises an ultrasonic signal having a frequency between about 20 and 96 kilohertz.
[0217] Example 41: The method of any preceding example, further comprising transmitting audible content during at least the portion of the first period of time.
[0218] Example 42: A computer-readable storage medium comprising instructions that, in response to execution by a processor, cause a hearable to perform any one of the methods described in Examples 25 to 41.
[0219] Example 43: A device comprising: at least one transducer; at least one processor; Equipped with configured to perform any one of the methods described in Examples 25 to 41 using the at least one transducer and the at least one processor. device.
[0220] Example 44: A speaker, an active noise cancellation circuit comprising a feedback microphone; Furthermore, the at least one transducer comprises the speaker and the feedback microphone; The device described in Example 43.
[0221] Example 45: The at least one transducer comprises a speaker and a microphone; the speaker is configured to be placed proximate to a first ear of a user; the microphone is configured to be placed proximate to a second ear of the user. The device described in Example 43.
[0222] Example 46: A device described in any one of Examples 43 to 45, further comprising a motion sensor.
[0223] Example 47: A device described in any one of Examples 43 to 46, wherein the device comprises at least one earphone.
Claims
1. 1. A method comprising: transmitting an acoustic transmit signal that propagates within at least a portion of the user's ear canal during a first period of time; and receiving, during the first time period, an acoustic receive signal representing a version of the acoustic transmit signal having one or more characteristics modified by the propagation within the ear canal, the acoustic receive signal including at least one movement artifact associated with the user moving during at least a portion of the first time period, the method further comprising: generating a denoised version of the acoustic receive signal using motion sensing data generated by a motion sensor during the first time period, wherein the at least one motion artifact in the denoised version of the acoustic receive signal is attenuated relative to the at least one motion artifact in the acoustic receive signal, the method further comprising: controlling operation of a hearable and / or a computing device coupled to the hearable based on the denoised version of the acoustic receive signal; A method comprising:
2. measuring at least one biometric of the user based on the denoised version of the acoustic received signal; further comprising and controlling the operation of the hearable and / or the computing device includes communicating the at least one biometric. The method of claim 1.
3. determining, based on the noise-removed signal, that the user has spoken during the first period; further comprising The controlling the operation of the hearable and / or the computing device may include: enabling a voice control interface based on said determination; or providing multi-factor voice authentication based on said determination; The method of claim 1 or 2, comprising at least one of:
4. recognizing a gesture made by the user during the first time period based on the noise-removed version of the acoustic receive signal; further comprising controlling the operation of the hearable and / or the computing device includes controlling the operation of the hearable and / or the computing device based on the recognized gesture.
10. A method according to any preceding claim.
5. determining that the user is grinding their teeth during the first period based on the noise-removed version of the acoustic receive signal; further comprising and controlling the operation of the hearable and / or the computing device includes controlling the operation of the hearable and / or the computing device based on the determination.
10. A method according to any preceding claim.
6. 10. The method of any preceding claim, wherein generating the denoised version of the acoustic receive signal comprises adjusting a passband of a filter based on the motion sensing data such that the filter passes frequencies of the version of the acoustic receive signal corresponding to biometric monitoring, speech detection, chewing detection, and / or gesture recognition.
7. 7. The method of claim 1, wherein generating the denoised version of the acoustic receive signal comprises performing adaptive filtering using the acoustic receive signal representing a primary reference and the motion sensing data representing a noise reference.
8. transmitting a second acoustic transmit signal during the first period of time that propagates within at least a portion of a second ear canal of the user; receiving, during the first time period, a second acoustic receive signal representing a version of the second acoustic transmit signal having one or more characteristics modified by the propagation within the second ear canal, the second acoustic receive signal including at least one second motion artifact associated with the user moving during at least the portion of the first time period; generating the denoised version of the acoustic receive signal includes performing adaptive filtering or blind source separation using the version of the acoustic receive signal, the second version of the acoustic receive signal, and the motion sensing data to generate the denoised signal and a motion signal including the at least one motion artifact.
10. A method according to any preceding claim.
9. The method of claim 8 , further comprising determining, based on the movement signal, that the user moved during at least the portion of the first time period.
10. The method of claim 8 or 9, further comprising classifying a type of activity performed by the user during the first period based on the motion signal.
11. The method of claim 10 , wherein the classifying the type of activity includes determining it as being relatively still or associated with a significant amount of movement.
12. transmitting a second acoustic transmit signal during the first period of time that propagates within at least a portion of a second ear canal of the user; and receiving, during the first time period, a second acoustic receive signal representing a version of the second acoustic transmit signal having one or more characteristics modified by the propagation within the second ear canal, the second acoustic receive signal including at least one second movement artifact associated with the user moving during at least the portion of the first time period, the method further comprising: generating a noise-reduced version of the second received acoustic signal using other motion sensing data generated by a second motion sensor during the first time period; further comprising wherein controlling the operation of the hearable and / or the computing device comprises controlling the operation of the hearable and / or the computing device based on the noise-removed version of the acoustic receive signal and the second acoustic receive signal.
10. A method according to any preceding claim.
13. The motion detection data Linear accelerations associated with three orthogonal axes, or the rotational speeds associated with the three orthogonal axes; 10. A method according to any preceding claim, comprising at least one of:
14. and transmitting, prior to the first period of time, another acoustic transmit signal propagating within at least the portion of the ear canal of the user, the other acoustic transmit signal having a plurality of tones, the method further comprising: receiving, prior to the first period of time, another acoustic receive signal representing a version of the other acoustic transmit signal having one or more characteristics modified by the propagation within the ear canal; selecting a subset of the plurality of tones based on a quality metric associated with the plurality of tones; further comprising transmitting the acoustic transmit signal during the first time period includes transmitting the acoustic transmit signal having the subset of the plurality of tones.
10. A method according to any preceding claim.
15. A computer readable storage medium comprising instructions that, when executed by a processor, cause a hearable to perform any one of the methods of claims 1 to 14.
16. A device, at least one transducer; at least one processor; Equipped with configured to perform any one of the methods according to claims 1 to 14 using said at least one transducer and said at least one processor; device.
17. A speaker and an active noise cancellation circuit comprising a feedback microphone; Furthermore, the at least one transducer comprises the speaker and the feedback microphone; 17. The device of claim 16.
18. the at least one transducer comprises a speaker and a microphone; the speaker is configured to be placed proximate to a first ear of a user; the microphone is configured to be placed proximate to a second ear of the user; 17. The device of claim 16.
19. The device of any one of claims 16 to 18, further comprising a motion sensor.
20. The device of any one of claims 16 to 19, wherein the device comprises at least one earphone.
Citation Information
Patent Citations
Infrasound Biosensor System and Method
JP2021513437A
System and Method for Leak Correction and Normalization of In-Ear Pressure Measurement for Hemodynamic Monitoring
US20210401311A1
Authentication management device, estimation method, and recording medium
WO2022195792A1
Audioplethysmography calibration
WO2023240240A1