Active self-voice normalization using bone conduction sensor
By filtering and gain adjustment of external and internal microphones and bone conduction sensors, the problem of different distortion modes between self-voice and external sound in wearable devices is solved, achieving natural audio signal processing and improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- QUALCOMM INC
- Filing Date
- 2021-10-06
- Publication Date
- 2026-07-21
Smart Images

Figure CN116491131B_ABST
Abstract
Description
[0001] Priority is claimed based on 35U.SC§119.
[0002] This patent application claims priority to non-provisional application No. 17 / 064,146, filed on October 6, 2020, entitled “ACTIVE SELF-VOICENATURALIZATION USING A BONE CONDUCTION SENSOR”, which has been assigned to the assignee of this application and is hereby expressly incorporated herein by reference. Technical Field
[0003] In general, the following relates to signal processing, and more specifically, to Active Self-Speech Naturalization (ASVN) using bone conduction sensors. Background Technology
[0004] Users may utilize wearable devices and may wish to experience direct hearing features or self-speech naturalization. In some examples, when a user speaks (e.g., generating a self-speech signal), the user's speech can propagate along two paths: an acoustic path and a bone conduction path. However, the distortion patterns from external or background signals may differ from those created by the self-speech signal. Microphones picking up input audio signals (e.g., including background noise and self-speech signals) may not be able to seamlessly process different types of signals. When using direct hearing features on wearable devices, the different distortion patterns of different signals can result in a lack of natural vocal audio input. Summary of the Invention
[0005] The described technology relates to improved methods, systems, devices, and apparatuses for supporting Active Self-Speech Naturalization (ASVN) using bone conduction sensors. Typically, as provided by the described technology, wearable devices may include an external microphone (e.g., outside the user's ear), an internal microphone (e.g., inside the user's ear), and a bone conduction sensor (e.g., inside the user's ear), each capable of picking up external sounds, such as self-speech, as input. The hearing device can determine an error associated with the bone conduction sensor input based on the difference between the input from the external microphone and the input from the internal microphone. The bone conduction input can be updated based on the error. The hearing device can perform an operation that applies a filter to the error-updated input. Furthermore, the external microphone input can be equalized according to gain. Both the error-updated, filtered bone conduction sensor input and the equalized external microphone input can be used to perform ASVN, which allows the user to perceive self-speech and additional external sounds as natural.
[0006] A method for audio signal processing at a wearable device is described. The method may include: receiving, at the wearable device including a set of microphones and a bone conduction sensor, a first input audio signal from an external microphone and a second input audio signal from an internal microphone; receiving a bone conduction signal from the bone conduction sensor, the bone conduction signal being associated with the first input audio signal and the second input audio signal; filtering the bone conduction signal based on a set of frequencies corresponding to the first input audio signal and the second input audio signal; and outputting an output audio signal to a speaker of the wearable device based on the filtering.
[0007] An apparatus for audio signal processing at a wearable device is described. The apparatus may include: a processor; a memory electrically communicating with the processor; and instructions stored in the memory. The instructions are executable by the processor to cause the apparatus to: receive a first input audio signal from an external microphone and a second input audio signal from an internal microphone at the wearable device, which includes a set of microphones and a bone conduction sensor; receive a bone conduction signal from the bone conduction sensor, the bone conduction signal being associated with the first input audio signal and the second input audio signal; filter the bone conduction signal based on a set of frequencies corresponding to the first input audio signal and the second input audio signal; and output an output audio signal to a speaker of the wearable device based on the filtering.
[0008] Another apparatus for audio signal processing at a wearable device is described. The apparatus may include units for performing the following operations: receiving a first input audio signal from an external microphone and a second input audio signal from an internal microphone at the wearable device, which includes a set of microphones and a bone conduction sensor; receiving a bone conduction signal from the bone conduction sensor, the bone conduction signal being associated with the first input audio signal and the second input audio signal; filtering the bone conduction signal based on a set of frequencies corresponding to the first input audio signal and the second input audio signal; and outputting an output audio signal to a speaker of the wearable device based on the filtering.
[0009] A non-transitory computer-readable medium is described, storing code for audio signal processing at a wearable device. The code may include instructions executable by a processor to: receive a first input audio signal from an external microphone and a second input audio signal from an internal microphone at the wearable device, which includes a set of microphones and a bone conduction sensor; receive a bone conduction signal from the bone conduction sensor, the bone conduction signal being associated with the first and second input audio signals; filter the bone conduction signal based on a set of frequencies corresponding to the first and second input audio signals; and output an output audio signal to a speaker of the wearable device based on the filtering.
[0010] Some examples of the methods, apparatuses, and non-transitory computer-readable media described herein may also include operations, features, units, or instructions for: calculating the difference between the first input audio signal and the second input audio signal; and determining an error based on the difference.
[0011] In some examples of the methods, apparatuses, and non-transitory computer-readable media described herein, filtering the bone conduction signal may further include operations, features, units, or instructions for: adjusting the first input audio signal based on the error; adjusting the second input audio signal based on the error; and applying a filter to the adjusted first input audio signal, the adjusted second input audio signal, the bone conduction signal, or a combination thereof.
[0012] Some examples of the methods, apparatuses, and non-transitory computer-readable media described herein may also include operations, features, units, or instructions for: calculating one or more power ratios corresponding to the first input audio signal, the second input audio signal, the bone conduction signal, or a combination thereof; and determining a threshold power ratio for the one or more power ratios.
[0013] Some examples of the methods, apparatuses, and non-transitory computer-readable media described herein may also include operations, features, units, or instructions for adding gain to a filtered bone conduction signal, the first input audio signal, the second input audio signal, or a combination thereof, based on the fact that one or more power ratios are below the threshold power ratio.
[0014] Some examples of the methods, apparatuses, and non-transitory computer-readable media described herein may also include operations, features, units, or instructions for updating the gain based on the filtered bone conduction signal, wherein the gain is an adjustable gain.
[0015] Some examples of the methods, apparatuses, and non-transitory computer-readable media described herein may also include operations, features, units, or instructions for equalizing the first input audio signal based on the gain and the second input audio signal.
[0016] Some examples of the methods, apparatuses, and non-transitory computer-readable media described herein may also include operations, features, units, or instructions for performing an active self-speech normalization process based on an equalized first input audio signal and the filtered bone conduction signal.
[0017] In some examples of the methods, apparatuses, and non-transitory computer-readable media described herein, performing the active self-speech normalization process may further include operations, features, units, or instructions for detecting the presence of self-speech in the first input audio signal.
[0018] In some examples of the methods, apparatuses, and non-transitory computer-readable media described herein, filtering the bone conduction signal may further include operations, features, units, or instructions for: determining that the first input audio signal and the second input audio signal comprise a set of frequencies; and filtering one or more low frequencies corresponding to the first input audio signal, the second input audio signal, or self-speech of either the first input audio signal, the second input audio signal, or both, wherein the set of frequencies includes the one or more low frequencies. Attached Figure Description
[0019] Figure 1 An example of an audio signaling scenario using a bone conduction sensor to support Active Self-Speech Naturalization (ASVN) is shown, according to aspects of this disclosure.
[0020] Figure 2 and Figure 3 An example of a signal processing scheme using a bone conduction sensor to support ASVN is shown, according to aspects of this disclosure.
[0021] Figure 4 and Figure 5 A block diagram of a wearable device supporting ASVN using a bone conduction sensor, according to aspects of this disclosure, is shown.
[0022] Figure 6 A block diagram of a signal processing manager supporting ASVN using a bone conduction sensor, according to an aspect of this disclosure, is shown.
[0023] Figure 7 A diagram of a system comprising a wearable device supporting ASVN using a bone conduction sensor, according to aspects of this disclosure, is shown.
[0024] Figures 8 to 10A flowchart illustrating a method for supporting ASVN using a bone conduction sensor, according to aspects of this disclosure, is shown. Detailed Implementation
[0025] Some users can use wearable devices (e.g., wireless communication devices, wireless headsets, earbuds, speakers, hearing aids, etc.) and can wear the devices to use them hands-free. Some wearable devices may include multiple microphones attached externally and internally to the device. These microphones can be used for a variety of purposes, such as noise detection, audio signal output, active noise cancellation, etc. When the user of the wearable device (e.g., the wearer) speaks, they can generate unique audio signals (e.g., self-speech). For example, the user's self-speech signal can propagate along an acoustic path (e.g., from the user's mouth to the microphone of the headset) and along a second sound path created by vibrations via bone conduction between the user's mouth and the microphone of the headset.
[0026] Some hearing devices, such as hearing aids or headsets, can operate in a mode that allows the user to hear external sounds. This mode can be called a transparency mode. For example, a user can activate transparency mode to determine the volume of their speech when communicating using a headset. In some cases, even when the hearing device is in transparency mode, a user's speech (e.g., self-speech) may sound different to a user without hearing devices than to a user with hearing devices. This difference may be due to changes in the acoustic path from the hearing device (e.g., lack of bone conduction acoustic path) and an imbalance in the frequency representation of self-speech within its frequency range in transparency mode (e.g., an increase in low frequencies).
[0027] As described herein, wearable devices may include bone conduction sensors for normalizing a set of frequencies of a user's speech. In some cases, hearing devices may include an external microphone (e.g., outside the user's ear), an internal microphone (e.g., inside the user's ear), and a bone conduction sensor (e.g., inside the user's ear), each capable of picking up external sounds, such as self-speech, as input. The hearing device may determine an error associated with the bone conduction sensor input based on the difference between the input from the external microphone and the input from the internal microphone. The bone conduction input may be updated based on the error and may be filtered (e.g., to suppress the overemphasis on low-frequency portions of self-speech). Furthermore, the external microphone input may be equalized according to gain. Both the updated, filtered bone conduction sensor input and the equalized external microphone input can be used to perform Active Self-Speech Naturalization (ASVN), which allows the user to perceive self-speech and additional external sounds as natural.
[0028] First, various aspects of this disclosure are described in the context of a signal processing system. Reference is also made to signal processing schemes, and various aspects of this disclosure are further illustrated and described with reference to apparatus diagrams, system diagrams, and flowcharts relating to an ASVN using a bone conduction sensor.
[0029] Figure 1 An example of an audio signaling scenario 100 supporting ASVN using a bone conduction sensor, according to aspects of this disclosure, is shown. Audio signaling scenario 100 may occur when a user 105 using wearable device 115 wishes to experience direct hearing features.
[0030] User 105 may use wearable device 115 (e.g., wireless communication device, wireless headset, earpiece, speaker, hearing aid, etc.), which may be worn by user 105 in a hands-free manner. In some cases, wearable device 115 may also be referred to as a hearing device. In some examples, user 105 may wear wearable device 115 continuously, regardless of whether wearable device 115 is currently being used (e.g., inputting audio signals, outputting audio signals, or both at one or more microphones 120). In some examples, wearable device 115 may include multiple microphones 120. For example, wearable device 115 may include one or more external microphones 120, such as external microphone 120-a and external microphone 120-b. Wearable device 115 may also include one or more internal microphones 120, such as internal microphone 120-c. Wearable device 115 may use microphones 120 for noise detection, audio signal output, active noise cancellation, etc.
[0031] When user 105 speaks, user 105 can generate a unique audio signal (e.g., self-speech). For example, user 105 can generate a self-speech signal that can propagate along acoustic path 125 (e.g., from user 105's mouth to microphone 120 of the headset). User 105 can also generate a self-speech signal that can travel along sound conduction path 130 created by vibrations of bone conduction between user 105's vocal cords or mouth and microphone 120 of wearable device 115. In some examples, wearable device 115 can perform self-speech activity detection (SVAD) based on self-speech quality. For example, wearable device 115 can identify inter-channel phase and intensity differences (e.g., interaction between external and internal microphones 120 of wearable device 115). Wearable device 115 can use the detected differences as defining features to compare self-speech signals with external signals. For example, if one or more differences between the channel phase and intensity of the internal microphone 120-c and the external microphone 120-a are detected, or if one or more differences between the channel phase and intensity of the internal microphone 120-c and the external microphone 120-a meet a threshold, the wearable device 115 can determine that a self-speech signal exists in the input audio signal.
[0032] In some examples, wearable device 115 may provide a direct hearing feature for operation in a transparent mode. The direct hearing feature allows user 105 to hear the output audio signal from wearable device 115 as if wearable device 115 were not present. The direct hearing feature allows user 105 to wear wearable device 115 hands-free, regardless of the current usage of wearable device 115 (e.g., whether wearable device 115 is using one or more microphones 120 to output audio signals, input audio signals, or both). For example, audio source 110 (e.g., a person, audio from the surrounding environment, etc.) may generate an external audio signal 135. For example, a person may speak to user 105, thus creating an external audio signal 135. Without the direct hearing feature, the external audio signal 135 may be blocked, muted, or otherwise distorted by wearable device 115. The direct listening function can use external microphone 120-a, external microphone 120-b, internal microphone 120-c, or a combination thereof to receive input audio signals (e.g., external audio signal 135), process the input audio signals, and output (e.g., via internal microphone 120-c) an audio signal that sounds natural to user 105 (e.g., sounds as if user 105 is not wearing a device).
[0033] The self-speech audio signal following acoustic path 125 and the external audio signal 135 can have different distortion modes. For example, the external audio signal 135, the self-speech audio signal following acoustic path 125, or both can have a first distortion mode. However, the self-speech following sound conduction path 130, the self-speech following acoustic path 125, or both can have a second distortion mode. The microphone 120 of the wearable device 115 can similarly detect the self-speech audio signal and the external audio signal 135. Therefore, without different processing for different signal types, the user 105 may not experience a natural-sounding input audio signal. That is, the wearable device 115 can detect input audio signals that include the external audio signal 135, self-speech via acoustic path 125, or a combination of self-speech via sound conduction path 130. The wearable device 115 can use the microphone 120 to detect the input audio signal.
[0034] In some examples, wearable device 115 may use external microphones 120-a and 120-b via acoustic path 125 to detect external audio signal 135 and self-speech. Additionally or alternatively, wearable device 115 may use one or more internal microphones 120 (e.g., internal microphone 120-c) via sound conduction path 130 to detect self-speech. Wearable device 115 may perform filtering on the received signal and may generate an output audio signal for user 105 (e.g., via internal microphone 120-c). In some cases, wearable device 115 may have difficulty generating self-speech that sounds natural without modifying external sound perception (e.g., due to different distortion patterns). For example, after performing active noise cancellation techniques to suppress low-frequency build-up, wearable device 115 may be unable to suppress the boost in the low-frequency range of self-speech, may lose the high-frequency range of self-speech, or both.
[0035] In some examples, wearable device 115 can use signals from bone conduction sensor 140 to modify the frequencies of external audio signal 135 and self-speech to achieve a natural-sounding output audio signal when wearable device 115 operates in transparency mode. For example, bone conduction sensor 140 can enable wearable device to suppress the build-up of low frequencies of self-speech, allowing equalization operations on the input audio signal to be applied to the high-frequency portions regardless of the presence of self-speech. That is, self-speech normalization can be decoupled from transparency mode (e.g., direct hearing features) at wearable device 115.
[0036] In some cases, user 105 may experience bone conduction while speaking using wearable device 115. For example, bone conduction can be the transmission of sound through the skull to the inner ear, allowing user 105 to perceive audio content using vibrations in the bones. In some examples, bones can transmit low-frequency sounds better than high-frequency sounds. Bone conduction sensor 140 may include a transducer that outputs a signal based on bone vibrations caused by audio. Additionally or alternatively, bone conduction sensor 140 may include any device (e.g., a sensor, etc.) that detects vibrations and outputs electronic signals.
[0037] In some examples, wearable device 115 may receive input audio signals from external microphones 120-a, 120-b, or both (e.g., external audio signal 135, user 105's own speech, or both) and from internal microphone 120-c. Additionally, wearable device 115 may receive bone conduction signals from bone conduction sensor 140 based on the input audio signals. Wearable device 115 may filter the bone conduction signals based on a set of frequencies of the input audio signals (e.g., the low-frequency portion of the input audio signals). For example, wearable device 115 may apply a filter to the bone conduction signals that cause errors, such as differences between input audio signals from one or more external microphones 120 and one or more internal microphones 120. In some cases, wearable device 115 may add gain to the filtered bone conduction signals and may equalize the filtered bone conduction signals based on the gain, referencing... Figure 2 and Figure 3 This will be described in further detail. Wearable device 115 can output audio signals (e.g., filtered bone conduction signals) to a speaker that can be heard by user 105.
[0038] Figure 2 An example of a signal processing scheme 200 supporting ASVN using a bone conduction sensor, according to aspects of this disclosure, is shown. In some examples, the signal processing scheme 200 can implement aspects of audio signaling scenario 100 and can include a wearable device 115-a having an external microphone 120-d, an internal microphone 120-e, and a bone conduction sensor 140-a, which may be referenced. Figure 1 Examples of the wearable device 115, microphone 120, and bone conduction sensor 140 described herein. For example, wearable device 115-a (which may be a hearing device) may use bone conduction sensor 140-a to apply direct hearing features in a transparent mode to account for self-speech.
[0039] In some cases, wearable device 115 can operate in a transparent mode where external noise is audible to user 105. Wearable device 115 can detect input audio signals from one or more external microphones 120, input audio signals from one or more internal microphones, or both. For example, wearable device 115-a can use external microphones 120-d to detect external microphone signal 205, and internal microphones 120-e to detect internal microphone signal 210, or both. External microphone signal 205 and internal microphone signal 210 may include audio signals from external sources, self-speech, or both. Self-speech audio signals and external audio signals may have different distortion modes. Wearable device 115 can perform a filtering process on the input audio signal and can generate an output audio signal for user 105. In some cases, wearable device 115 may have difficulty generating self-speech that sounds natural without modifying the perception of external sound (e.g., due to different distortion modes). For example, after performing active noise cancellation to suppress low-frequency build-up, the wearable device 115 may be unable to suppress the boost in the low-frequency range of the self-speech, may lose the high-frequency range of the self-speech, or both.
[0040] In some cases, wearable device 115 can use bone conduction sensor 140 to achieve a true transparent mode. For example, wearable device 115-a can detect bone conduction sensor signal 215 from bone conduction sensor 140-a. Wearable device 115-a can perform one or more operations on external microphone signal 205, internal microphone signal 210, bone conduction sensor signal 215, or a combination thereof to output an audio signal to the speaker of wearable device 115-a. For example, in the absence of a headset receiver, user 105 can hear the audio signal according to Formula 1:
[0041]
[0042] Among them, S ac It can be an audio signal that propagates along a purely acoustic path. It can be an audio signal that travels along the acoustic path from bone conduction, and This is the audio signal that travels along the bone conduction path. In some other examples, when using a headset, user 105 can hear the audio signal according to Equation 2:
[0043]
[0044] Where P is the passive attenuation factor and Q is the enhanced bone conduction factor. In some cases, the audio signal propagating along the bone conduction path may not be captured by the microphone 120, but can be perceived by the user 105. Therefore, the wearable device 115 can apply a filter 220 to the bone conduction sensor signal 215 based on one or more operations and frequencies of the external microphone signal 205 and the internal microphone signal 210 to address the passive attenuation and enhanced bone conduction factors.
[0045] External microphone signal 205 can be an audio signal propagating along a purely acoustic path, S ac Wearable device 115 can apply equalizer 225 to compensate for losses (e.g., passive attenuation P) caused by passive gain between external microphone 120-d and internal microphone 120-e, and to compensate for speaker distortion G. For example, equalizer 225 can adjust the input of equalizer 225 (which may be S...) ac or S ac It has an additional gain of 230g (S) ac Multiply by In some cases, wearable device 115-a can shape the additional gain g for each frequency based on user preferences for the mode. In some cases, wearable device 115-a can maintain a "headset off" state for external sounds, and then an equalizer can be applied at 235 during the ASVN process.
[0046] In some examples, at convergence point 240, wearable device 115-a may combine external microphone signal 205 (which may include additional gain 230, possibly already operated by compensator 245, or both) with internal microphone signal 210 to avoid canceling a portion of the additional playback (e.g., possibly occurring during equalization operation). In some cases, wearable device 115-a may apply compensator 245 to external microphone signal 205, or a modified external microphone signal 205 (e.g., applied to S...). ac Or with an additional gain of 230, g(S) ac ), of S ac In some cases, the compensator can be considered... To resolve noise in the bone conduction sensor signal 215. The wearable device 115-a can perform preprocessing steps on the external microphone signal 205, the bone conduction sensor signal 215, or both.
[0047] For example, wearable device 115-a can examine the power ratio between signals from bone conduction sensor 140-a and external microphone 120-d. Wearable device 115-a can suppress external microphone signal 205, bone conduction sensor signal 215, or a portion thereof with a power ratio below a threshold, which can suppress external sounds captured by bone conduction sensor 140-a. Additionally or alternatively, wearable device 115-a can measure the cross-correlation between external microphone signal 205 and bone conduction sensor signal 215, or between bone conduction sensor signal 215 and internal microphone signal 210. Wearable device 115-a can suppress uncorrelated portions of the signal (e.g., external microphone signal 205, bone conduction sensor signal 215, internal microphone signal 210, or combinations thereof), which can suppress uncorrelated noise in the signal.
[0048] In some cases, after convergence 240, the wearable device 115-a can respond to the enhanced bone conduction internal microphone signal 210. Execute error update procedure 250. For example, the error update procedure can be entered as follows: Z is the variable in Formula 4:
[0049] ||S ac -X i (Z)|| 2
[0050] Among them, X i It is the internal microphone signal 210.
[0051] In some examples, wearable device 115-a may apply filter 220 to the error-updated internal microphone signal 210, bone conduction sensor signal 215, or both. In some examples, wearable device 115-a may interpret the bone conduction sensor signal 215 as having a distortion factor T (e.g., as...). Filter 220 can be a finite impulse response (FIR) filter, an infinite impulse response (IIR) filter, or any other type of filter. In some examples, filter 220 can multiply the input (e.g., the error-updated internal microphone signal 210, the bone conduction sensor signal 215, or both) by a factor, such as... This can address distortion, T, in the bone conduction sensor signal 215; speaker distortion, G; and the enhanced bone conduction factor, Q. In some cases, the wearable device 115-a can filter one or more low frequencies of self-speech based on applying filter 220 to the error-updated internal microphone signal 210, bone conduction sensor signal 215, or both.
[0052] After applying filter 220 to the error-updated internal microphone signal 210, bone conduction sensor signal 215, or both, the wearable device 115-a can add an optional gain 255 to the output of filter 220. For example, the wearable device 115-a can add an optional gain 255 to have a small residual in the acoustically transmitted bone conduction sound. User 105 can hear a slight residual sound. If the wearable device 115-a adds an optional gain 255, this slight residual can be addressed in the ASVN procedure 235. In some cases, the selectable gain 255 can be an adjustable gain, which the wearable device 115-a can adjust. The wearable device 115-a can perform the ASVN procedure based on the equalized external microphone signal 205 and the filtered bone conduction sensor signal 215.
[0053] Figure 3 An example of a signal processing scheme 300 supporting ASVN using a bone conduction sensor, according to aspects of this disclosure, is shown. In some examples, signal processing scheme 300 may implement aspects of audio signaling scenario 100, signal processing scheme 200, or both. Signal processing scheme 300 may include a wearable device 115-b with an external microphone 120-f and an external microphone signal 305, an internal microphone 120-g with an internal microphone signal 310, and a bone conduction sensor 140-b with a bone conduction sensor signal 315, which may be referenced. Figure 1 The wearable device 115, microphone 120, and bone conduction sensor 140 described, and reference Figure 2 Examples of external microphone signal 205, internal microphone signal 210, and bone conduction sensor signal 215 are described. Signal processing scheme 300 may also include one or more operations involving filter 320, equalizer 325, additional gain 330, ASVN procedure 335, convergence of one or more signals 340, compensator 345, error update procedure 350, etc., as described in reference... Figure 2 As described. For example, wearable device 115-b may apply filter 320 to error-updated external microphone signal 305 (e.g., based on internal microphone signal 310), bone conduction sensor signal 315, or both, to interpret self-speech for direct hearing features in transparent mode.
[0054] In some cases, wearable device 115-b can operate in a transparent mode where external noise is audible to user 105. Wearable device 115-b can use external microphone 120-f to detect external microphone signal 305, and internal microphone 120-g to detect internal microphone signal 310, or both. External microphone signal 305 and internal microphone signal 310 may include audio signals from external sources, self-speech, or both. Self-speech audio signals and external audio signals may have different distortion modes. In some cases, wearable device 115-b may have difficulty producing self-speech that sounds natural without modifying the perception of external sound (e.g., due to different distortion modes). For example, after performing active noise cancellation techniques to suppress low-frequency build-up, wearable device 115-b may be unable to suppress the boost in the low-frequency range of self-speech, may lose the high-frequency range of self-speech, or both.
[0055] In some examples, wearable device 115-b may perform one or more operations to modify external microphone signal 305, bone conduction sensor signal 315, or both to interpret speech (e.g., modify reference). Figure 2 The wearable device 115-b determines whether self-speech is present in the external audio signal before determining the signal described. The wearable device 115-b can perform SVAD procedure 355 based on detecting one or more self-speech qualities. For example, the wearable device 115-b can identify inter-channel phase and intensity differences (e.g., the interaction between external microphone 120-f and internal microphone 120-g). The wearable device 115-b can use the detected differences as defining features to compare the self-speech signal with the external signal. For example, if one or more differences in channel phase and intensity are detected between internal microphone 120-g and external microphone 120-f, or if one or more differences in channel phase and intensity between internal microphone 120-g and external microphone 120-f meet a threshold, the wearable device 115-b can determine that a self-speech signal is present in the input audio signal.
[0056] In some cases, when wearable device 115-b detects self-speech during SVAD process 355, wearable device 115-b can turn on switch 360. When switch 360 is turned on, wearable device 115-b can use a filtered bone conduction sensor signal 315, an equalized external microphone signal 305, or both to perform ASVN process 335 (e.g., as referenced). Figure 2(As described in signal processing scheme 200). In some other cases, when the wearable device 115-b does not detect self-speech during the SVAD process 355, the wearable device 115-b may turn off switch 360. When switch 360 is off, the wearable device 115-b may not execute the ASVN procedure 335, but may instead output an external microphone signal 305, an internal microphone signal 310, or both without considering bone conduction (e.g., without using bone conduction sensor 140-b).
[0057] Figure 4 A block diagram 400 of a wearable device 405 supporting ASVN using a bone conduction sensor, according to an aspect of this disclosure, is shown. Wearable device 405 may be an example of an aspect of wearable device 115 described herein. Wearable device 405 may include a receiver 410, a signal processing manager 415, and a speaker 420. Wearable device 405 may also include a processor. Each of these components may communicate with each other (e.g., via one or more buses).
[0058] Receiver 410 can receive audio signals from the surrounding area (e.g., via a microphone array). The detected audio signals can be transmitted to other components of wearable device 405. Receiver 410 can utilize a single antenna or a set of antennas to communicate with other devices while providing seamless direct listening characteristics.
[0059] Signal processing manager 415 can receive a first input audio signal from an external microphone and a second input audio signal from an internal microphone at a wearable device including a set of microphones and a bone conduction sensor; receive a bone conduction signal from the bone conduction sensor, the bone conduction signal being associated with the first and second input audio signals; filter the bone conduction signal at least in part based on a set of frequencies corresponding to the first and second input audio signals; and output an output audio signal to a speaker of the wearable device based on the filtering. Signal processing manager 415 may be an example of an aspect of signal processing manager 710 described herein.
[0060] Actions performed by the signal processing manager 415 described herein can be implemented to achieve one or more potential advantages. One implementation enables a wearable device to use the signal output of a bone conduction sensor to interpret self-speech in an audio signal. The bone conduction sensor allows the wearable device to filter one or more audio signals and bone conduction sensor signals in a transparent mode, which allows naturally occurring self-speech to be used as the wearable device's output, among other advantages.
[0061] Based on the bone conduction sensor described herein, the processor of a wearable device (e.g., a processor controlling receiver 410, signal processing manager 415, speaker 420, or a combination thereof) can improve the user experience in transparent mode while ensuring relatively efficient operation. For example, the ASVN technology described herein can utilize filtering and equalization operations for microphone signals, bone conduction sensor signals, or both, based on the detection of self-speech in an external audio signal. This can enable improved transparent mode operation at the wearable device, as well as other benefits.
[0062] The signal processing manager 415 or its sub-components may be implemented in hardware, in code (e.g., software or firmware) executed by a processor, or any combination thereof. If implemented in code executed by a processor, the functionality of the signal processing manager 415 or its sub-components may be performed by a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic units, discrete hardware components, or any combination thereof designed to perform the functions described in this disclosure.
[0063] The signal processing manager 415 or its subcomponents may be physically located in various places, including distributed such that portions of the functionality are implemented by one or more physical components at different physical locations. In some examples, according to various aspects of this disclosure, the signal processing manager 415 or its subcomponents may be separate and distinct components. In some examples, according to various aspects of this disclosure, the signal processing manager 415 or its subcomponents may be combined with one or more other hardware components, including but not limited to input / output (I / O) components, transceivers, network servers, another computing device, one or more other components described in this disclosure, or combinations thereof.
[0064] Speaker 420 can provide output signals generated by other components of wearable device 405. In some examples, speaker 420 may be co-located with the internal microphone of wearable device 405. For example, speaker 420 may be a reference... Figure 7 Examples of aspects of the described speaker 725.
[0065] Figure 5A block diagram 500 of a wearable device 505 supporting ASVN using a bone conduction sensor, according to an aspect of this disclosure, is shown. Wearable device 505 may be an example of an aspect of wearable device 405 or wearable device 115 described herein. Wearable device 505 may include a receiver 510, a signal processing manager 515, and a speaker 545. Wearable device 505 may also include a processor. Each of these components may communicate with each other (e.g., via one or more buses).
[0066] Receiver 510 can receive audio signals (e.g., via a set of microphones). Information can be transmitted to other components of wearable device 505.
[0067] As described herein, signal processing manager 515 may be an example of an aspect of signal processing manager 415, signal processing manager 605, or signal processing manager 710. Signal processing manager 515 may include microphone assembly 520, bone conduction assembly 525, frequency assembly 530, and output assembly 535.
[0068] Microphone assembly 520 can receive a first input audio signal from an external microphone and a second input audio signal from an internal microphone at a wearable device including a set of microphones and a bone conduction sensor. Bone conduction assembly 525 can receive bone conduction signals from the bone conduction sensor, the bone conduction signals being associated with the first and second input audio signals. Frequency assembly 530 can filter the bone conduction signals based on a set of frequencies corresponding to the first and second input audio signals. Output assembly 535 can output an output audio signal to a speaker of the wearable device, at least in part, based on the filtering.
[0069] Speaker 545 can provide an output signal generated by other components of wearable device 505. In some examples, speaker 545 may be co-located with a microphone. For example, speaker 545 may be a reference. Figure 7 Examples of aspects of the described speaker 725.
[0070] Figure 6 A block diagram 600 of a signal processing manager 605 supporting ASVN using a bone conduction sensor, according to an aspect of this disclosure, is shown. The signal processing manager 605 may be an example of an aspect of the signal processing manager 415, signal processing manager 515, or signal processing manager 710 described herein. The signal processing manager 605 may include a microphone assembly 610, a bone conduction assembly 615, a frequency assembly 620, an output assembly 625, an error assembly 630, and a power ratio assembly 635. Each of these modules may communicate directly or indirectly with each other (e.g., via one or more buses).
[0071] Microphone assembly 610 can receive a first input audio signal from an external microphone and a second input audio signal from an internal microphone at a wearable device including a set of microphones and a bone conduction sensor. Bone conduction assembly 615 can receive bone conduction signals from the bone conduction sensor, the bone conduction signals being associated with the first and second input audio signals. Frequency assembly 620 can filter the bone conduction signals based on a set of frequencies corresponding to the first and second input audio signals, as described herein. Output assembly 625 can output an output audio signal to a speaker of the wearable device, at least in part, based on the filtering.
[0072] In some examples, the error component 630 can calculate the difference between the first input audio signal and the second input audio signal; and determine the error based on the difference. The error component 630 can adjust the first input audio signal based on the error; adjust the second input audio signal based on the error; and apply a filter to the adjusted first input audio signal, the adjusted second input audio signal, the bone conduction signal, or a combination thereof.
[0073] In some cases, the power ratio component 635 can calculate one or more power ratios corresponding to a first input audio signal, a second input audio signal, a bone conduction signal, or a combination thereof; and can determine a threshold power ratio for the one or more power ratios. The power ratio component 635 can add gain to the filtered bone conduction signal, the first input audio signal, the second input audio signal, or a combination thereof based on one or more power ratios being lower than the threshold power ratio. The power ratio component 635 can update the gain based on the filtered bone conduction signal, wherein the gain is adjustable. In some examples, the power ratio component 635 can equalize the first input audio signal based on the gain and the second input audio signal. The power ratio component 635 can perform an ASVN process based on the equalized first input audio signal and the filtered bone conduction signal. For example, the power ratio component 635 can detect the presence of self-speech in the first input audio signal.
[0074] In some cases, frequency component 620 may determine that the first input audio signal and the second input audio signal comprise a set of frequencies; and filter one or more low frequencies corresponding to the first input audio signal, the second input audio signal, or the self-speech of either the two, wherein the set of frequencies comprises one or more low frequencies.
[0075] Figure 7A diagram of a system 700 including a wearable device 705 supporting ASVN using a bone conduction sensor, according to aspects of this disclosure, is shown. Wearable device 705 may be an example of or include components of wearable device 115, wearable device 405, or wearable device 505 as described herein. Wearable device 705 may include components for bidirectional voice and data communication, including components for transmitting and receiving communications, including a signal processing manager 710, an I / O controller 715, a transceiver 720, a memory 730, and a processor 740. These components may communicate electrically via one or more buses (e.g., bus 745).
[0076] The signal processing manager 710 can receive a first input audio signal from an external microphone and a second input audio signal from an internal microphone at a wearable device including a set of microphones 750 and a bone conduction sensor 760; receive a bone conduction signal from the bone conduction sensor, the bone conduction signal being associated with the first input audio signal and the second input audio signal; filter the bone conduction signal at least in part based on a set of frequencies corresponding to the first input audio signal and the second input audio signal; and output an output audio signal to a speaker of the wearable device based on the filtering.
[0077] The I / O controller 715 can manage the input and output signals of the wearable device 705. The I / O controller 715 can also manage peripheral devices not integrated into the wearable device 705. In some cases, the I / O controller 715 can represent a physical connection or port to an external peripheral device. In some cases, the I / O controller 715 can use, for example... The operating system or other known operating systems. In other cases, the I / O controller 715 may represent or interact with a modem, keyboard, mouse, touchscreen, or similar device. In some cases, the I / O controller 715 may be implemented as part of a processor. In some cases, a user may interact with the wearable device 705 via the I / O controller 715 or via hardware components controlled by the I / O controller 715.
[0078] Transceiver 720 can communicate bidirectionally via one or more antennas, wired or wireless links. For example, transceiver 720 can represent a wireless transceiver and can communicate bidirectionally with another wireless transceiver. Transceiver 720 may also include a modem for modulating packets and providing modulated packets to the antenna for transmission, and for demodulating packets received from the antenna. In some examples, the aforementioned direct listening feature allows users to experience natural acoustic interaction with their environment when performing wireless communication or receiving data via transceiver 720.
[0079] The speaker 725 can provide the user with an output audio signal (e.g., with seamless direct listening characteristics).
[0080] Memory 730 may include random access memory (RAM) and read-only memory (ROM). Memory 730 may store computer-readable, computer-executable code 735, including instructions that, when executed, cause a processor to perform the various functions described herein. In some cases, among other things, memory 730 may contain a basic I / O system (BIOS) that controls basic hardware or software operations, such as interaction with peripheral components or devices.
[0081] Processor 740 may include intelligent hardware devices (e.g., general-purpose processors, DSPs, CPUs, microcontrollers, ASICs, FPGAs, programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or any combination thereof). In some cases, processor 740 may be configured to use a memory controller to operate a memory array. In other cases, the memory controller may be integrated into processor 740. Processor 740 may be configured to execute computer-readable instructions stored in memory (e.g., memory 730) to cause wearable device 705 to perform various functions (e.g., supporting ASVN functions or tasks using bone conduction sensors).
[0082] Code 735 may include instructions for implementing aspects of this disclosure, including instructions for supporting signal processing. In some cases, aspects of the signal processing manager 710, I / O controller 715, and / or transceiver 720 may be implemented by portions of code 735, which are executed by processor 740 or another device. Code 735 may be stored in a non-transitory computer-readable medium such as system memory or other types of memory. In some cases, code 735 may not be directly executable by processor 740, but may enable a computer (e.g., when compiled and executed) to perform the functions described herein.
[0083] Figure 8 A flowchart illustrating a method 800 for supporting ASVN using a bone conduction sensor according to aspects of this disclosure is shown. Operation of method 800 can be implemented by a wearable device or its components described herein. For example, operation of method 800 can be described by reference to... Figures 4 to 7 The described signal processing manager is used for execution. In some examples, the wearable device can execute a set of instructions to control the functional units of the wearable device to perform the functions described below. Additionally or alternatively, the wearable device may use dedicated hardware to perform aspects of the functions described below.
[0084] At point 805, the wearable device can receive a first input audio signal from an external microphone and a second input audio signal from an internal microphone at a wearable device including a set of microphones and bone conduction sensors. Operation of point 805 can be performed according to the methods described herein. In some examples, some aspects of the operation of point 805 can be found in references. Figures 4 to 7 The microphone manager described is used to perform this.
[0085] At point 810, the wearable device can receive bone conduction signals from a bone conduction sensor, which are associated with a first input audio signal and a second input audio signal. Operation of point 810 can be performed according to the methods described herein. In some examples, aspects of the operation of point 810 can be found in references. Figures 4 to 7 The described beamforming manager is used to perform this.
[0086] At point 815, the wearable device can filter the bone conduction signal based on a set of frequencies corresponding to the first and second input audio signals. Operation 815 can be performed according to the method described herein. In some examples, aspects of the operation of 815 can be found in references. Figures 4 to 7 The described signal isolation manager is used to perform this.
[0087] At point 820, the wearable device can output an audio signal to the wearable device's speaker based on filtering. Operation of point 820 can be performed according to the methods described herein. In some examples, aspects of the operation of point 820 can be found in references. Figures 4 to 7 The described filter manager is used to perform this.
[0088] Figure 9 A flowchart illustrating a method 900 for supporting ASVN using a bone conduction sensor, according to aspects of this disclosure, is shown. Operation of method 900 can be implemented by a wearable device or its components described herein. For example, operation of method 900 can be performed by reference to... Figures 4 to 7 The described signal processing manager is used for execution. In some examples, the wearable device can execute a set of instructions to control the functional units of the wearable device to perform the functions described below. Additionally or alternatively, the wearable device may use dedicated hardware to perform aspects of the functions described below.
[0089] At point 905, the wearable device can receive a first input audio signal from an external microphone and a second input audio signal from an internal microphone at a wearable device including a set of microphones and bone conduction sensors. Operation of point 905 can be performed according to the methods described herein. In some examples, some aspects of the operation of point 905 can be found in references. Figures 4 to 7 The microphone manager described is used to perform this.
[0090] At point 910, the wearable device can receive bone conduction signals from a bone conduction sensor, which are correlated with a first input audio signal and a second input audio signal. Operation of point 910 can be performed according to the methods described herein. In some examples, aspects of the operation of point 910 can be found in references. Figures 4 to 7 The described beamforming manager is used to perform this.
[0091] At point 915, the wearable device can calculate the difference between the first input audio signal and the second input audio signal. The operation of point 915 can be performed according to the methods described herein. In some examples, aspects of the operation of point 915 can be derived from, as referenced... Figures 4 to 7 The described audio scaling manager is used to perform this.
[0092] At point 920, the wearable device can determine the error based on this difference. The operation at point 920 can be performed according to the methods described herein. In some examples, aspects of the operation at point 920 can be found in references. Figures 4 to 7 The described signal isolation manager is used to perform this.
[0093] At 925, the wearable device can filter the bone conduction signal based on a set of frequencies corresponding to the first and second input audio signals. Operation at 925 can be performed according to the method described herein. In some examples, aspects of operation at 925 may be derived from, as referenced... Figures 4 to 7 The described audio scaling manager is used to perform this.
[0094] At 930, the wearable device can output an audio signal to the wearable device's speaker based on filtering. Operation 930 can be performed according to the methods described herein. In some examples, aspects of the operation of 930 can be found in references... Figures 4 to 7 The described filter manager is used to perform this.
[0095] Figure 10 A flowchart illustrating a method 1000 for supporting ASVN using a bone conduction sensor according to aspects of this disclosure is shown. Operation of method 1000 can be implemented by a wearable device or its components described herein. For example, operation of method 1000 can be performed by reference to... Figures 4 to 7 The described signal processing manager is used for execution. In some examples, the wearable device can execute a set of instructions to control the functional units of the wearable device to perform the functions described below. Additionally or alternatively, the wearable device may use dedicated hardware to perform aspects of the functions described below.
[0096] At point 1005, the wearable device can receive a first input audio signal from an external microphone and a second input audio signal from an internal microphone, comprising a set of microphones and bone conduction sensors. Operation of point 1005 can be performed according to the methods described herein. In some examples, aspects of the operation of point 1005 can be found in references... Figures 4 to 7 The microphone manager described is used to perform this.
[0097] At point 1010, the wearable device can receive bone conduction signals from a bone conduction sensor, which are associated with a first input audio signal and a second input audio signal. Operation of point 1010 can be performed according to the methods described herein. In some examples, aspects of the operation of point 1010 can be found in references. Figures 4 to 7 The described beamforming manager is used to perform this.
[0098] At point 1015, the wearable device can calculate one or more power ratios corresponding to the first input audio signal, the second input audio signal, the bone conduction signal, or a combination thereof. Operation at point 1015 can be performed according to the methods described herein. In some examples, aspects of operation at point 1015 may be derived from, as referenced... Figures 4 to 7 The described audio scaling manager is used to perform this.
[0099] At point 1020, the wearable device can determine a threshold power ratio for one or more power ratios. Operation at point 1020 can be performed according to the methods described herein. In some examples, aspects of operation at point 1020 can be found in references. Figures 4 to 7 The described signal isolation manager is used to perform this.
[0100] At position 1025, the wearable device can filter the bone conduction signal based on a set of frequencies corresponding to the first and second input audio signals. Operation at position 1025 can be performed according to the methods described herein. In some examples, aspects of operation at position 1025 may be derived from, as referenced... Figures 4 to 7 The described audio scaling manager is used to perform this.
[0101] At 1030, the wearable device can output an output audio signal to the wearable device's speaker based on filtering. Operation of 1030 can be performed according to the methods described herein. In some examples, aspects of the operation of 1030 can be found in references... Figures 4 to 7 The described filter manager is used to perform this.
[0102] It should be noted that the methods described in this paper describe possible implementations, and the operations and steps can be rearranged or otherwise modified, and other implementations are possible. Furthermore, aspects from two or more of these methods can be combined.
[0103] The techniques described in this article can be used in various signal processing systems such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal Frequency Division Multiple Access (OFDMA), Single Carrier Frequency Division Multiple Access (SC-FDMA), and others. CDMA systems can implement radio technologies such as CDMA2000 and Universal Terrestrial Radio Access (UTRA). CDMA2000 encompasses the IS-2000, IS-95, and IS-856 standards. Versions of IS-2000 are commonly referred to as CDMA2000 1X, 1X, etc. IS-856 (TIA-856) is commonly referred to as CDMA2000 1xEV-DO, High-Speed Packet Data (HRPD), etc. UTRA includes Wideband CDMA (WCDMA) and other variants of CDMA. TDMA networks can implement radio technologies such as Global System for Mobile Communications (GSM).
[0104] OFDMA systems can implement wireless technologies such as Ultra Mobile Broadband (UMB), Evolved UTRA (E-UTRA), IEEE 802.11 (Wi-Fi), IEEE 802.16 (WiMAX), IEEE 802.20, and Flash OFDM. UTRA and E-UTRA are components of the Universal Mobile Telecommunications System (UMTS). LTE, LTE-A, and LTE-A Pro are versions of UMTS that use E-UTRA. UTRA, E-UTRA, UMTS, LTE, LTE-A, LTE-A Pro, NR, and GSM are described in documents from an organization called the 3rd Generation Partnership Project (3GPP). CDMA2000 and UMB are described in documents from an organization called the 3rd Generation Partnership Project 2 (3GPP2). The technologies described herein can be used in the systems and radio technologies mentioned herein, as well as other systems and radio technologies. While some aspects of LTE, LTE-A, LTE-A Pro, or NR systems may be described for illustrative purposes, and the terms LTE, LTE-A, LTE-A Pro, or NR may be used in most of the description, the techniques described herein can be applied beyond LTE, LTE-A, LTE-A Pro, or NR applications.
[0105] Macro cells typically cover a relatively large geographic area (e.g., a radius of several kilometers) and allow unrestricted access by UEs with service subscribers to a network provider. In contrast, small cells can be associated with low-power base stations and can operate in the same or different (e.g., licensed, unlicensed, etc.) frequency bands as macro cells. Depending on the examples, small cells can include picocells, femtocells, and microcells. For example, a picocell can cover a small geographic area and allow unrestricted access by UEs with service subscribers to a network provider. A femtocell can also cover a small geographic area (e.g., a home) and provide restricted access for UEs associated with that femtocell (e.g., UEs in a Closed Subscriber Group (CSG), UEs for users at home, etc.). The eNB for a macro cell may be referred to as a macro eNB. The eNB for a small cell may be referred to as a small cell eNB, pico eNB, femtocell eNB, or home eNB. An eNB can support one or more cells (e.g., two, three, four, etc.) and can also support communication using one or more component carriers.
[0106] The signal processing system described in this paper can support synchronous or asynchronous operation. For synchronous operation, base stations can have similar frame timing sequences, and transmissions from different base stations can be approximately time-aligned. For asynchronous operation, base stations can have different frame timing sequences, and transmissions from different base stations cannot be time-aligned. The techniques described in this paper can be used for both synchronous and asynchronous operations.
[0107] The information and signals described herein can be represented using any of a variety of different techniques and methods. For example, data, instructions, commands, information, signals, bits, symbols, and chips mentioned throughout this specification can be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, light fields or light particles, or any combination thereof.
[0108] Using a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic, discrete hardware component, or any combination thereof designed to perform the functions described herein, the various illustrative blocks and modules described in connection with the disclosure herein may be implemented or executed. The general-purpose processor may be a microprocessor; however, alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices (e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such architecture).
[0109] The functions described herein can be implemented using hardware, software executed by a processor, firmware, or any combination thereof. If implemented by software executed by a processor, these functions can be stored as one or more instructions or code on or transmitted over a computer-readable medium. Other examples and implementations are within the scope of this application and the appended claims. For example, due to the nature of software, the functions described herein can be implemented using software executed by a processor, hardware, firmware, hardwiring, or any combination thereof. Features implementing the functions can also be physically placed in various locations, including portions distributed such that functions are implemented at different physical locations.
[0110] Computer-readable media includes both non-transitory computer storage media and communication media, wherein the communication media includes any medium that facilitates the transfer of a computer program from one location to another. Non-transitory storage media can be any available medium that can be accessed by a general-purpose computer or a special-purpose computer. By way of example, and not limitation, non-transitory computer-readable media can include RAM, ROM, electrically erasable programmable ROM (EEPROM), flash memory, disc-on-CD ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other non-transitory medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a general-purpose or special-purpose computer or a general-purpose or special-purpose processor. Furthermore, any connection can be appropriately referred to as computer-readable media. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of media. As used herein, disks and discs include CDs, laser discs, optical discs, DVDs, floppy disks, and Blu-ray discs, where disks typically copy data magnetically, while discs use lasers to copy data optically. The combinations above should also be included within the scope of computer-readable media.
[0111] As used herein, the word "or" as used in the claims, as in the list of entries (e.g., a list of entries preceded by phrases such as "at least one of" or "one or more of"), indicates an inclusive list such that, for example, a list of at least one of A, B, or C means A, or B, or C, or AB, or AC, or BC, or ABC (i.e., A and B and C). Furthermore, as used herein, the phrase "based on" should not be construed as a reference to a closed set of conditions. For example, an exemplary step described as "based on condition A" may be based on both condition A and condition B without departing from the scope of this disclosure. In other words, as used herein, the phrase "based on" will be interpreted in the same manner as the phrase "at least partially based on".
[0112] In the accompanying drawings, similar components or features may have the same reference numerals. Additionally, components of the same type may be distinguished by a dash followed by a second reference numeral to differentiate between similar components. If only the first reference numeral is used in this specification, the description applies to any similar component having the same first reference numeral, regardless of the second reference numeral or other subsequent reference numerals.
[0113] The specification described herein, in conjunction with the accompanying drawings, describes exemplary configurations and does not represent all examples that can be implemented or that fall within the scope of the claims. The term "exemplary" as used throughout this specification means "serving as an example, instance, or illustration," and not "preferred" or "advantageous" relative to other examples. Specific details are included in the detailed description to provide an understanding of the described techniques. However, these techniques can be implemented without using these specific details. In some cases, well-known structures and devices are shown in block diagram form to avoid obscuring the concepts of the described examples.
[0114] The description herein is provided to enable those skilled in the art to implement or use the disclosed content. Various modifications to this disclosure will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the scope of this disclosure. Therefore, this disclosure is not limited to the examples and designs described herein, but is consistent with the broadest scope of the principles and novel features disclosed herein.
Claims
1. A method for audio signal processing in a wearable device, comprising: The wearable device, which includes multiple microphones and a bone conduction sensor, receives a first input audio signal from an external microphone and a second input audio signal from an internal microphone. Receive bone conduction signals from the bone conduction sensor, the bone conduction signals being associated with the first input audio signal and the second input audio signal; The bone conduction signal is filtered at least in part based on a set of frequencies corresponding to the first input audio signal and the second input audio signal; Calculate the power ratio between the first input audio signal and the bone conduction signal; Determine the threshold power ratio; If the power ratio is lower than the threshold power ratio, gain is added to the filtered bone conduction signal; as well as The output audio signal is output to the speaker of the wearable device by combining the filtered bone conduction signal with the first input audio signal from the external microphone.
2. The method according to claim 1, further comprising: Calculate the difference between the first input audio signal and the second input audio signal; as well as The error is determined at least in part based on the difference.
3. The method according to claim 2, wherein, Filtering the bone conduction signal further includes: The first input audio signal is adjusted at least in part based on the error. The second input audio signal is adjusted at least in part based on the error; and A filter is applied to the adjusted first input audio signal, the adjusted second input audio signal, the bone conduction signal, or a combination thereof.
4. The method according to claim 1, further comprising: The gain is updated at least in part based on filtering the bone conduction signal, wherein the gain is an adjustable gain.
5. The method according to claim 1, further comprising: The first input audio signal is equalized based at least in part on the gain and the second input audio signal.
6. The method according to claim 5, further comprising: The active self-speech naturalization process is performed at least in part based on the equalized first input audio signal and the filtered bone conduction signal.
7. The method according to claim 6, wherein, The active self-speech naturalization process further includes: The presence of self-speech in the first input audio signal is detected.
8. The method according to claim 1, wherein, Filtering the bone conduction signal further includes: It is determined that the first input audio signal and the second input audio signal include multiple frequencies; and Filter one or more low frequencies corresponding to the first input audio signal, the second input audio signal, or the self-speech of both, wherein the set of frequencies includes the one or more low frequencies.
9. An apparatus for audio signal processing in a wearable device, comprising: processor; A memory that communicates electrically with the processor; as well as Instructions stored in the memory and executable by the processor to cause the device to perform the following operations: The wearable device, which includes multiple microphones and a bone conduction sensor, receives a first input audio signal from an external microphone and a second input audio signal from an internal microphone. Receive bone conduction signals from the bone conduction sensor, the bone conduction signals being associated with the first input audio signal and the second input audio signal; The bone conduction signal is filtered at least in part based on a set of frequencies corresponding to the first input audio signal and the second input audio signal; Calculate the power ratio between the first input audio signal and the bone conduction signal; Determine the threshold power ratio; If the power ratio is lower than the threshold power ratio, gain is added to the filtered bone conduction signal; as well as The output audio signal is output to the speaker of the wearable device by combining the filtered bone conduction signal with the first input audio signal from the external microphone.
10. The apparatus according to claim 9, wherein, The instructions can also be executed by the processor to cause the device to perform the following operations: Calculate the difference between the first input audio signal and the second input audio signal; and The error is determined at least in part based on the difference.
11. The apparatus according to claim 10, wherein, The instructions for filtering the bone conduction signal can also be executed by the processor to cause the device to perform the following operations: The first input audio signal is adjusted at least in part based on the error. The second input audio signal is adjusted at least in part based on the error. as well as A filter is applied to the adjusted first input audio signal, the adjusted second input audio signal, the bone conduction signal, or a combination thereof.
12. The apparatus according to claim 9, wherein, The instructions can also be executed by the processor to cause the device to perform the following operations: The gain is updated at least in part based on filtering the bone conduction signal, wherein the gain is an adjustable gain.
13. The apparatus according to claim 9, wherein, The instructions can also be executed by the processor to cause the device to perform the following operations: The first input audio signal is equalized based at least in part on the gain and the second input audio signal.
14. The apparatus according to claim 13, wherein, The instructions can also be executed by the processor to cause the device to perform the following operations: The active self-speech naturalization process is performed at least in part based on the equalized first input audio signal and the filtered bone conduction signal.
15. The apparatus according to claim 9, wherein, The instructions for filtering the bone conduction signal can also be executed by the processor to cause the device to perform the following operations: It is determined that the first input audio signal and the second input audio signal include multiple frequencies; and Filter one or more low frequencies corresponding to the first input audio signal, the second input audio signal, or the self-speech of both, wherein the set of frequencies includes the one or more low frequencies.
16. A non-transitory computer-readable medium storing code for audio signal processing at a wearable device, the code comprising instructions executable by a processor for the following operations: The wearable device, which includes multiple microphones and bone conduction sensors, receives a first input audio signal from an external microphone and a second input audio signal from an internal microphone. Receive bone conduction signals from the bone conduction sensor, the bone conduction signals being associated with the first input audio signal and the second input audio signal; The bone conduction signal is filtered at least in part based on a set of frequencies corresponding to the first input audio signal and the second input audio signal; Calculate the power ratio between the first input audio signal and the bone conduction signal; Determine the threshold power ratio; If the power ratio is lower than the threshold power ratio, gain is added to the filtered bone conduction signal; as well as The output audio signal is output to the speaker of the wearable device by combining the filtered bone conduction signal with the first input audio signal from the external microphone.