Wearable audio device with user's own voice recording

By using a VAD accelerometer and controller in a wearable audio device to detect when a user is speaking and only record the user's voice, the problem of isolating the user's voice from the ambient acoustic signals is solved, and high-quality user voice recording is achieved.

CN115699175BActive Publication Date: 2026-02-03BOSE CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202180037539.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-08
Filing Date
2021-04-26
Publication Date
2026-02-03
Estimated Expiration
2041-04-26

AI Technical Summary

Technical Problem

Existing smart devices and wearable audio devices struggle to effectively isolate a user's own voice from ambient acoustic signals, resulting in other people's conversations being mixed into the user's voice recordings.

Method used

Using a Voice Activity Detection (VAD) accelerometer and controller, the user's speech is recorded using only the VAD accelerometer signal when the user speaks. The user's speech is isolated by VAD devices such as light-based sensors, sealed roll-to-roll microphones, or feedback microphones.

Benefits of technology

This technology enables the recording of only the user's voice in wearable audio devices, reducing or eliminating interference from ambient acoustic signals and improving the privacy and quality of voice recording.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115699175B_ABST
    Figure CN115699175B_ABST
Patent Text Reader

Abstract

Various implementations include a wearable audio device configured to record a user's speech without recording other environmental acoustic signals, such as conversations of other people nearby. In some particular aspects, a wearable audio device includes a frame to contact a user's head, an electro-acoustic transducer within the frame and configured to output an audio signal, at least one microphone, a voice activity detection (VAD) accelerometer, and a controller coupled with the electro-acoustic transducer, the at least one microphone, and the VAD accelerometer and configured in a first mode to: detect that the user is speaking; and in response to detecting that the user is speaking, record the user's speech using only signals from the VAD accelerometer.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CLAIM OF PRIORITY

[0002] This application claims priority to U.S. Patent Application No. 16 / 869,759, filed May 8, 2020, which is hereby incorporated by reference in its entirety. TECHNICAL FIELD

[0003] The present disclosure relates generally to wearable audio devices. More specifically, the present disclosure relates to wearable audio devices configured to enhance recording of a user’s own voice. BACKGROUND

[0004] There are various scenarios in which a user can wish to record his or her own voice. For example, a user can wish to make a verbal to-do list, record her day or a moment in her life, or sometimes analyze her speech patterns or voice tone. Given that commonly used devices such as smart devices and wearable audio devices include microphones, it can seem logical for a user to rely on these devices to perform own voice recording. However, conventional smart devices (e.g., smart phones, smart watches) and wearable audio devices (e.g., earphones, earpieces, etc.) can not effectively isolate a user’s own voice from ambient acoustic signals, such as conversations of other people nearby. SUMMARY

[0005] All examples and features mentioned below can be combined in any technically possible manner.

[0006] Various implementations include a wearable audio device. The wearable audio device is configured to record a user’s voice without recording other ambient acoustic signals, such as conversations of other people nearby.

[0007] In some particular aspects, a wearable audio device includes a frame to contact a head of a user, an electro-acoustic transducer within the frame and configured to output an audio signal, at least one microphone, a voice activity detection (VAD) accelerometer, and a controller coupled with the electro-acoustic transducer, the at least one microphone, and the VAD accelerometer and configured, in a first mode, to: detect that the user is speaking; and in response to detecting that the user is speaking, record the user’s voice using only signals from the VAD accelerometer.

[0008] In additional particular aspects, a computer-implemented method includes, at a wearable audio device having: a frame to contact a head of a user, an electro-acoustic transducer within the frame and configured to output an audio signal, at least one microphone, and a voice activity detection (VAD) accelerometer: in a first mode: detecting that the user is speaking; and in response to detecting that the user is speaking, recording the user’s voice using only signals from the VAD accelerometer.

[0009] In further particular aspects, a wearable audio device includes a frame to contact a head of a user, an electro-acoustic transducer located within the frame and configured to output an audio signal, at least one microphone, a voice activity detection (VAD) device, and a controller coupled with the electro-acoustic transducer, the at least one microphone, and the VAD device and configured, in a first mode, to: detect that the user is speaking; and in response to detecting that the user is speaking, record speech of the user using only a signal from the VAD device, wherein the VAD device includes at least one of: an optical-based sensor, a sealed spoolie microphone, or a feedback microphone.

[0010] Implementations can include one of the following features, or any combination thereof.

[0011] In certain aspects, the VAD accelerometer remains in contact with the head of the user during recording, or is separated from the head of the user during at least a portion of the recording.

[0012] In particular cases, in the second mode, the controller is configured to adjust a directionality of audio pickup from the at least one microphone to at least one of: verify that the user is speaking or improve a quality of the recording.

[0013] In some implementations, the wearable audio device further includes an additional VAD system to verify that the user is speaking prior to initiating recording.

[0014] In certain cases, the controller is further configured to, in response to detecting that only the user is speaking, communicate with a smart device connected with the wearable audio device to initiate natural language processing (NLP) of commands in the user speech detected at the at least one microphone on the wearable audio device or a microphone array on the smart device.

[0015] In some aspects, the NLP is performed after detecting that only the user is speaking without requiring a wake word.

[0016] In particular implementations, the controller is further configured to: request feedback from the user to verify that the user is speaking; train a logic engine to recognize that the user is speaking based on a received response to the feedback request; and after training, run the logic engine to detect future instances of the user speaking so as to enable recording using only the VAD accelerometer.

[0017] In certain aspects, the wearable audio device further includes a memory to store a predetermined amount of speech recording from the user.

[0018] In some cases, a recording of the user’s voice can be accessed via the processor to perform at least one of: a) analyze the voice recording for at least one of speech patterns or voice tone; b) play back the voice recording in response to a request from the user; or c) perform a virtual personal assistant (VPA) command based on the voice recording.

[0019] In particular aspects, in the second mode, the controller activates the at least one microphone to record all detectable ambient audio.

[0020] In certain implementations, the controller is configured to switch from the first mode to the second mode in response to a user command.

[0021] In some aspects, the wearable audio device further includes a digital signal processor (DSP) coupled with the controller, wherein the controller is further configured to activate the DSP to enhance the user’s recorded voice during recording.

[0022] In particular cases, the controller is further configured to initiate playback of the recording by at least one of: a) accelerating playback of the recording; b) playing back only selected portions of the recording; or c) adjusting a playback speed of one or more selected portions of the recording.

[0023] In certain implementations, the VAD accelerometer includes a bone conduction pickup transducer.

[0024] In some aspects, the controller is further configured to use the VAD accelerometer and the at least one microphone to implement voice-over recording of a television or streaming program by: fingerprinting audio output associated with the television or streaming program using the at least one microphone while recording the user’s voice using signals from the VAD accelerometer; compiling the fingerprinted audio output with the user’s recorded voice; and providing the compiled fingerprinted audio output and the user’s recorded voice for subsequent playback in synchronization with the television or streaming program.

[0025] Two or more features described in this disclosure, including those described in the SUMMARY, can be combined to form implementations not specifically described herein.

[0026] The details of one or more implementations are discussed in the drawings and detailed description below. Other features, objects, and benefits will be apparent from the summary, drawings, and claims. BRIEF DESCRIPTION OF DRAWINGS

[0027] Figure 1 is a perspective view of an example audio device according to various implementations.

[0028] Figure 2is a perspective view of another example audio device according to various implementations.

[0029] Figure 3 is a system diagram illustrating electronics in an audio device and a smart device in communication with the electronics according to various implementations.

[0030] Figure 4 is a flow diagram illustrating a process performed by a controller according to various implementations.

[0031] It is noted that the figures of the various implementations are not necessarily drawn to scale. The figures are intended to be illustrative, and thus should not be considered to limit the scope of the implementations. In the figures, like numbering represents similar elements between the figures. DETAILED DESCRIPTION

[0032] The present disclosure is based, at least in part, on the realization that a wearable audio device can be configured to privately record a user's own voice. For example, a wearable audio device disclosed in accordance with the implementations can provide a user with the ability to record the user's own voice while excluding environmental acoustic signals such as the voice of other nearby users. In particular cases, the wearable audio device utilizes a voice activity detection (VAD) device such as a VAD accelerometer to specifically record the user's voice.

[0033] For illustrative purposes, components generally labeled in the drawings are considered substantially equivalent components and redundant discussion of those components is omitted for the sake of clarity. Numerical ranges and values described in accordance with the various implementations are merely examples of such ranges and values and are not intended to be limiting of those implementations. In some cases, the term "about" is used to modify a numerical value and, in these cases, the numerical value can be indexed by the magnitude of the + / - error such as measurement error.

[0034] The aspects and implementations disclosed herein can be applicable to a wide variety of wearable audio devices of various form factors, such as head-worn devices (e.g., headphones, earphones, earpieces, eyewear, headgear, hats, face masks), neck-worn speakers, shoulder-worn speakers, body-worn speakers (e.g., watches), and the like. Some particular aspects disclosed can be applicable to personal (wearable) audio devices such as head-worn audio devices, including earphones, earpieces, headgear, hats, face masks, eyewear, and the like. It should be noted that while particular implementations of audio devices for the purpose of acoustic output of audio are presented in some degree of detail, such presentation of particular implementations is intended to facilitate understanding by providing examples and should not be considered limiting of the scope of the disclosed content or the scope encompassed by the claims.

[0035] Aspects and implementations disclosed herein can be applicable to wearable audio devices that do or do not support bidirectional communication and to personal audio devices that do or do not support active noise reduction (ANR). For wearable audio devices that do support bidirectional communication or ANR, what is disclosed and claimed herein is intended to be applicable to speaker systems that include one or more microphones disposed on a portion of the wearable audio device that, in use, remains outside the ear (e.g., a feed-forward microphone), on a portion that, in use, is inserted into the ear (e.g., a feedback microphone), or on both such portions. Still other implementations of wearable audio devices to which what is disclosed and claimed herein is applicable will be apparent to those skilled in the art.

[0036] The wearable audio devices disclosed herein can include additional features and capabilities not explicitly described. These wearable audio devices can include additional hardware components such as one or more cameras, location tracking devices, microphones, etc., and can be capable of speech recognition, visual recognition, and other smart device functionality. The descriptions of wearable audio devices included herein are not intended to preclude these additional functionalities in such devices.

[0037] Figure 1 is a schematic illustration of a wearable audio device 10 according to various implementations. In this example implementation, the wearable audio device 10 is a pair of audio eyeglasses 20. As shown, the wearable audio device 10 can include a frame 30 having a first section (e.g., a lens section) 40 and at least one additional section (e.g., an arm section) 50 extending from the first section 40. In this example, as with conventional eyeglasses, the first (or lens) section 40 and the additional section (arm) 50 are designed to rest on a user’s head. In this example, the lens section 40 can include a set of lenses 60, which can include prescription lenses, non-prescription lenses, and / or filter lenses, and a bridge 70 (which can include padding) for resting on a user’s nose. The arm 50 can include a profile 80 for resting on a user’s respective ear.

[0038] According to particular implementations, the electronics 90 and other components for controlling the wearable audio device 10 are contained within the frame 30 (or substantially contained within the frame, such that the components can extend beyond the boundaries of the frame). In some cases, separate or duplicate sets of electronics 90 are contained in portions of the frame, such as in each of the respective arms 50 in the frame 30. However, certain components described herein can also be present in the singular.

[0039] Figure 2Another example of a wearable audio device 10 in the form of headphones 210 is depicted. In some cases, headphones 210 include on-ear headphones or over-ear headphones 210. Headphones 210 may include a frame 220 having a first section (e.g., headband) 230 and at least one additional section (e.g., earcups) 240 extending from the first section 230. In various specific embodiments, the headband 230 includes a head pad 250. Depending on a particular embodiment, electronics 90 and other components for controlling the wearable audio device 10 are stored within one or both earcups 240. It should be understood that... Figure 2 The earphone 210 shown is merely an exemplary shape factor, and in-ear headphones (also known as earpieces or earplugs), helmets, masks, etc., may include electronic devices 90 capable of performing the functions described herein.

[0040] Figure 3 It is shown that it is contained in frame 30 ( Figure 1 ) and / or frame 220 ( Figure 2 A schematic diagram of the electronic device 90 within the device. It should be understood that one or more components in the electronic device 90 may be implemented as hardware and / or software, and such components may be connected via any conventional means (e.g., hardwired and / or wireless connections). It should also be understood that any component described as being connected to or coupled to another component in the wearable audio device 10 or other systems disclosed according to a specific embodiment may communicate using any conventional hardwired connections and / or additional communication protocols. In various specific embodiments, components individually housed within the wearable audio device 10 are configured to communicate using one or more conventional wireless transceivers.

[0041] like Figure 3 As shown, it is at least partially contained in frame 20 ( Figure 1 ) or frame 210 ( Figure 2 The electronics 90 within the device may include a transducer 310 (e.g., an electroacoustic transducer), at least one microphone (e.g., a single microphone or an array of microphones) 320, and a voice activity detection (VAD) device 330. Each of the transducer 310, microphone 320, and power supply 330 is connected to a controller 340, which is configured to perform control functions according to the various specific embodiments described herein. The controller 340 may be coupled to other components in the electronics 90 via any conventional wireless and / or hardwired connection, which allows the controller 340 to send or receive signals to and control the operation of those components.

[0042] The electronics 90 can include other components not specifically shown herein, such as one or more power sources, memory and / or processors, motion / movement detection components (e.g., an inertial measurement unit, a gyroscope / magnetometer, etc.), communication components configured to communicate with one or more other electronic devices connected via one or more wireless networks (e.g., a local WiFi network, a Bluetooth connection, or a radio frequency (RF) connection), and amplification and signal processing components. It will be appreciated that these components, or functional equivalents of these components, can be connected with or form part of the controller 340. In additional optional implementations, the electronics 90 can include an interface 350 coupled with the controller 340 for enabling functions such as audio selection, powering on / off of the audio device, or voice control functions. In certain cases, the interface 350 includes buttons, a compressible interface, and / or a capacitive touch interface. Various additional functions of the electronics 90 are described in U.S. Patent Application No. 10,353,221, which is incorporated by reference herein in its entirety.

[0043] In some implementations, one or more components in the electronics 90 or functions performed by such components are located on or performed by a smart device 360, such as a smartphone, smartwatch, tablet computer, laptop computer, or other computing device. In various implementations, one or more control circuits and / or chips at the smart device 360 can be used to perform one or more functions of the controller 340. In certain cases, actions of the controller 340 are performed as software functions via one or more controllers 340. In some cases, the smart device 360 includes an interface 350 for interacting with the controller 340, however, in other cases, both the wearable audio device 10 and the smart device 350 have separate interfaces 350. In certain cases, the smart device 360 includes at least one processor 370 (e.g., one or more processors, which can include a digital signal processor) and a memory 380. In some cases, the smart device 360 also includes an additional VAD system 390, which can include one or more microphones for detecting voice activity, for example, from a user of the wearable audio device 10 and / or another user. In certain cases, the additional VAD system 390 can be used to verify that a user is speaking, as described herein, and can be used in conjunction with the VAD device 330 at the electronics 90.

[0044] In certain implementations, the controller 340 is configured to operate in one or more modes. Figure 4 is a flowchart showing an example process performed by the controller 340 in a first mode. In some cases, in the first mode, the controller 340 is configured to perform the following process:

[0045] A) detecting that the user is speaking (of the wearable audio device 10); and

[0046] B) in response to detecting that the user is speaking, recording the user’s speech using only signals from the VAD device 330 located on the wearable audio device 10.

[0047] That is, in the first mode, the controller 340 is configured to specifically record the user speech signal without capturing signals from environmental sound sources such as other users (of the speech). As described herein, the wearable audio device 10 implements this recording by using the VAD device 330 Figure 3 ) positioned to record signals indicative of the user speech without recording environmental acoustic signals. In particular cases, the controller 340 initiates recording using only signals from the VAD device 330, such that signals detected by the microphone 320 or additional environmental acoustic signal detection devices are excluded (e.g., using a logic-based VAD component).

[0048] In various additional implementations, the controller 340 is configured to improve the ability to detect that the user is speaking by training (e.g., using machine learning or other artificial intelligence components such as artificial neural networks). As optional implementations, these additional processes are shown in dashed lines in Figure 4 . In these cases, after the controller 340 detects that the user is speaking (process A above), the controller 340 is configured to:

[0049] C) request feedback from the user to verify that the user is speaking. In some cases, the controller 340 requests user feedback via one or more interfaces 350, such as via audio, haptic, and / or gesture-based interfaces to request and / or respond.

[0050] D) train a logic engine to recognize that the user is speaking based on received responses to the feedback requests. In some cases, the logic engine is contained in the controller 340, or at the wearable audio device 10 or at the smart device 360. In other implementations, the logic engine is executed at the smart device 360 or in a cloud-based platform, and can include machine learning components.

[0051] E) after training, run the logic engine to detect future instances of the user speaking in order to implement recording using only the VAD device 330. In these processes, the controller 340 includes or accesses the trained logic engine to detect that the user is speaking. This process is shown in Figure 4 as improving process A, detecting that the user is speaking.

[0052] In certain implementations, the VAD device 330 includes a VAD accelerometer. In particular cases, the VAD accelerometer is positioned on a frame (e.g., frame 20( Figure 1 ) or frame 210( Figure 2 )) such that it remains in contact with a user’s head when the wearable audio device 10 is worn by the user. That is, in particular cases, the VAD accelerometer is positioned on a portion of the frame such that it contacts the user’s head during use of the wearable audio device 10. However, in other cases, the VAD accelerometer does not always contact the user’s head such that it is physically separated from the user’s head during at least some portions of use of the wearable audio device 10. In various particular implementations, the VAD accelerometer includes a bone conduction pickup transducer configured to detect vibrations conducted via a user’s bone structure.

[0053] While the VAD device 330 is described in some cases as including a VAD accelerometer, in other cases the VAD device 330 can include one or more additional or alternative voice activity detection components, including, for example, light-based sensors, sealed coil microphones, and / or feedback microphones. In some examples, the light-based sensors can include infrared (IR) and / or laser sensors configured to detect movement of a user’s mouth. In these cases, the light-based sensors can be positioned on a frame (e.g., frame 20( Figure 1 ) or frame 210( Figure 2 )) to, for example, direct light to a user’s mouth region to detect movement of the user’s mouth. The VAD device 330 can additionally or alternatively include sealed microphones that are encapsulated to prevent detection of external acoustic signals. In these cases, the sealed coil microphones can be one or more of the microphones 320( Figure 3 ), or can be separate microphones that are dedicated as the VAD device 330 on the wearable audio device 10. In particular examples, the sealed coil microphones include microphones that are substantially enclosed except for a side that faces a user (e.g., toward a location of the user’s mouth). In some cases, the sealed coil microphones are located in an acoustically isolated housing that defines a sealed coil behind the microphone. In yet some additional implementations, the VAD device 330 can include a feedback microphone that is located in a front cavity of the wearable audio device 10, for example, in a portion of a frame (e.g., frame 20( Figure 1 ) or frame 210( Figure 2 )) that is proximate to a user’s mouth. In certain cases, the feedback microphone includes one or more of the microphones 320( Figure 3one or more microphones in the microphone array 320, or is a separate microphone dedicated to use as the VAD device 330. In some cases, the feedback microphone is also located in an acoustically isolated enclosure, e.g., similar to a sealed, spoolie microphone.

[0054] Particular implementations are described in which the VAD device 330 is a VAD accelerometer that includes a bone conduction pickup transducer. However, it should be understood that regardless of the particular form of the VAD device 330, the controller 340 is configured to isolate the signal from the VAD device 330 to record the user’s own voice without capturing the ambient acoustic signal.

[0055] As noted herein, in additional implementations, the controller 340 is configured to operate in one or more additional modes. For example, in another mode, the controller 340 is configured to adjust the directionality of audio pickup from the microphones 320. In some cases, the controller 340 is configured to adjust the directionality of audio pickup from the microphones 320 to verify that the user is speaking and / or to improve the quality of the recording. For example, in response to detecting that the user is speaking (e.g., receiving a signal from the VAD device 330 indicating that the user is speaking and / or receiving a signal from the VAD system 390( Figure 3 ) indicating that the user is speaking), the controller 340 is configured to adjust the directionality of the microphones 320. Microphone directionality can be adjusted by modifying the gain on the signal detected by one or more microphones 320, as well as performing beamforming processing to enhance the signal from one or more directions relative to other directions (e.g., creating nulls in directions other than a direction pointing towards the user’s mouth). Other methods for adjusting microphone directionality are also possible.

[0056] In certain implementations, in response to detecting that the user is speaking based on the signal from the VAD device 330, the controller 340 adjusts the microphone directivity at the microphone 320, for example, by pointing the microphone 320 at the user’s mouth and performing analysis (e.g., speech recognition) on the received signal to verify that the user is speaking. In some cases, the controller 340 is only configured to record the signal from the VAD device 330 in response to verifying that the user is speaking, for example, using confirmation from the signal received via the microphone 320. In other implementations, in response to detecting that the user is speaking based on the signal from the VAD device 330, the controller 340 adjusts the microphone directivity at the microphone 320 to improve the quality of the recording from the VAD device 330. In these cases, the controller 340 is configured to identify frequencies of the signal other than the user’s speech (e.g., low or high frequency sounds, such as an appliance humming, a motor vehicle driving nearby, or background music), and perform signal processing on the signal from the VAD device 330 to exclude those frequencies (or ranges of frequencies) from the recording. Examples of this signal processing and beamforming techniques are described in more detail in U.S. Patent No. 10,311,889, which is incorporated by reference in its entirety.

[0057] In particular cases, the controller 340 is configured to verify that the user is speaking using the additional VAD system 390 prior to initiating recording from the VAD device 330. In these cases, the controller 340 can utilize the signal from the VAD system 390 to verify that the user is speaking. In cases where the VAD system 390 includes one or more microphones, this process can be performed using those microphones in a similar manner to verifying that the user is speaking through directional adjustment of the microphone 320 (described herein). In these cases, the controller 340 is configured to not record the signal from the VAD device 330 unless the VAD system 390 verifies that the user is speaking, for example, by verifying that the signal detected by the VAD system 390 includes a user speech signal.

[0058] In various specific implementations, controller 340 is also configured to selectively implement voice control at audio device 10 using voice signals detected, such as those by microphone 320. For example, in some cases, controller 340 is configured to communicate with smart device 360 ​​in response to detecting that only a user is speaking, to initiate natural language processing (NLP) of commands in the user's voice detected at microphone 320 and / or at microphone in VAD system 390 at smart device 360. In these cases, controller 340 is configured to detect that only a user is speaking and, in response, send the detected voice signal data to smart device 360 ​​for processing (e.g., via one or more processors such as an NLP processor). In some cases, NLP is performed without a wake word after detecting that only a user is speaking. That is, controller 340 may be configured to initiate NLP based on commands from the user without a wake word (e.g., "Hey, Bose" or "Bose, please play music by Artist X"). In these specific implementations, the controller 340 enables users to voice control one or more functions of the wearable audio device 10 and / or the smart device 360 ​​without requiring a wake word.

[0059] Depending on the specific implementation, controller 340 is configured to automatically switch between modes (e.g., in response to a detected condition at audio device 10, such as detecting a user speaking, or detecting a nearby user speaking), or in response to a user command. In certain cases, controller 340 enables the user to switch between modes using user commands, for example, via interface 350. In these cases, the user can switch between recording all signals detected at VAD device 330 and not recording any signals at VAD device 330 (or disabling VAD device 330). In still other cases, controller 340 enables the user to switch to a full recording mode that activates microphone 320 to record all detectable ambient audio. In some examples, the user provides a command to interface 350 (e.g., a user interface command such as a haptic interface command, voice command, or gesture command), and controller 340 (in response to detecting the command) switches to full recording mode by activating microphone 320 and initiating recording of all signals received at microphone 320. In certain cases, these microphone signals may be recorded in addition to those from VAD device 330.

[0060] In some specific implementations, memory 380 is configured to store a predetermined amount of voice recordings from the user. Although memory 360 is shown as located in... Figure 3The smart device 380 is located in the wearable audio device 10, but one or more portions of the memory 380 may be located in the wearable audio device 10 and / or a cloud-based system (e.g., a cloud server with memory). In some cases, the recording of the user's voice is accessible (e.g., via processor 370 or another processor at wearable audio device 10 and / or smart device 360) to: i) analyze the voice recording for at least one of speech patterns or voice tone; ii) play back the voice recording in response to a request from the user; and / or iii) execute virtual personal assistant (VPA) commands based on the voice recording. In some cases, the user's voice recording is analyzed for speech patterns and / or voice tone, such as the frequency of word use like profanity, placeholders (e.g., "uh", "um", "okay"), speech rhythm, coughing, sneezing, breathing (e.g., loud breathing), hiccups, etc. In some specific implementations, controller 340 enables the user to choose how the voice recording is analyzed. For example, controller 340 may be configured to allow a user to select one or more analysis modes for speech recording, such as one or more specific speech patterns or voice tone analysis. In certain cases, the user may choose to analyze the user's speech recording using placeholder items (e.g., "um" or "uh") or rhythm (e.g., prolonged pauses or rapid sentence transitions).

[0061] In some other specific embodiments, controller 340 is configured to, for example, enhance or otherwise modulate the user's recorded speech during recording. For example, controller 340 may be coupled to one or more digital signal processors (DSPs), such as those included in processor 370 on smart device 360 ​​or in additional DSP circuitry included in electronics 90. Controller 340 may be configured to activate the DSP to enhance the user's recorded speech during recording. In some specific cases, controller 340 is configured to increase the signal-to-noise ratio (SNR) of the signal from VAD device 330 to enhance the user's recorded speech. In some cases, controller 340 filters out frequencies (or ranges) detected at VAD device 330 that are known or may be associated with sounds other than noise or user speech. For example, controller 340 may be configured to filter out low-level vibrations and / or high-frequency sounds that can be detected by VAD device 330 to enhance SNR. In a particular example, controller 340 uses a high-pass filter to remove low-frequency noise, such as from mechanical systems or wind. In other specific examples, controller 340 uses a low-pass or band-pass filter to remove other noise sources. In another example, controller 340 applies one or more filter models to the detected noise, such as filter models developed using machine learning.

[0062] In some additional embodiments, controller 340 is configured to retrieve the user's recorded audio (speech), for example, for playback. In some cases, controller 340 is configured to initiate playback of the recording (of the user's speech) at wearable audio device 10 (e.g., at transducer 310), at smart device 360 ​​(e.g., at one or more transducers), and / or at another playback device. In specific cases, controller 340 is configured to initiate playback of the user's recorded speech at transducer 310. In some cases, controller 340 is configured to initiate playback of the recording by: a) accelerating the playback of the recording; b) playing back only selected portions of the recording; and / or c) adjusting the playback speed of one or more selected portions of the recording. For example, a user might want to accelerate the playback of the recording so that they can hear a larger amount of the user's recorded speech in a shorter amount of time. In other cases, users may want to play back only selected portions of their voice recordings, such as specific conversations that occurred at a particular time of day, verbal diary / log entries, verbal annotations or reminders set at a specific time of day or location (e.g., related to location data). In still other cases, users may want to adjust the playback speed of one or more selected portions of the recording, such as speeding up or slowing down playback at a specific time of day or during a specific conversation with another user.

[0063] In some other specific implementations, controller 340 is configured to implement one or more voice-over functions. In some examples, controller 340 is configured to enable a user to record, replay, and / or share narration or other voice-over content. For example, in one process, a user may wish to save and replay and / or share a segment of narration from a television or streaming program, such as a show, movie, broadcast sports event, etc. In these examples, the television or streaming program includes audio output. In this example, a user wearing audio device 10 begins watching a television or streaming program and actuates controller 340 (e.g., via an interface command) to begin recording narration or other voice-over content. In response to a command to record narration or voice-over content, controller 340 is configured to record the user's own speech according to various methods described herein (e.g., using VAD device 330). Additionally, during the narration or voice-over recording process, controller 340 is configured to perform fingerprinting on the audio from the television or streaming program, for example, as output from a television, personal computer / laptop, smart device, or other streaming or broadcasting device. In these cases, controller 340 is configured to record timing-related data (or “fingerprinting”) from the audio of a television or streaming program, enabling narration or voice-over recording to be synchronized with individual broadcasts or playbacks of the television or streaming program by different users. Using conventional fingerprinting, controller 340 is configured to record the user’s own voice (e.g., narration) without recording the audio from the television or streaming program. This may be advantageous where the audio from the television or streaming program is subject to certain intellectual property rights (e.g., copyright). The user can use any of the interface commands described herein to pause or end recording (e.g., voice-over recording).

[0064] After recording a user's voice-over content, the content may be saved or otherwise transmitted as one or more data files (e.g., via an executed software application or otherwise accessible via controller 340) and made available for subsequent playback. In some cases, controller 340 uploads the voice-over content data to a database or other accessible platform such as a cloud-based platform. In these cases, the recording user or different users can access the previously recorded voice-over content. For example, a second user with an audio device (e.g., audio device 10) may provide an interface command detected at controller 340 for accessing the voice-over content. In these cases, controller 340 may provide the user with notifications or other indications that the voice-over content is available for one or more television or streaming programs. In any case, when controller 340 is activated, the user may initiate playback of the television or streaming program (e.g., via a television interface, streaming service, on-demand service, etc.). The microphone at audio device 10 is configured to detect audio playback associated with the television or streaming program and uses fingerprint recognition data to synchronize the voice-over content from the first user with the playback of the television or streaming program. In these cases, controller 340 is configured to mix voice-over content from the first user with audio playback of a television or streaming program for output at transducer 310.

[0065] As noted herein, the various embodiments disclosed enable private user voice recordings within wearable audio devices, compared to conventional systems and methods for recording user speech. In these embodiments, the wearable audio devices and related methods allow users to minimize or eliminate interface interaction while still recording the user's speech, for example, for subsequent playback. Users can experience numerous benefits from their own voice recordings without the negative impact of detecting the speech of other users within the recording. Furthermore, these wearable audio devices enable analysis and feedback of the user's voice recordings, for example, via transmission to one or more connected devices, such as smart devices or cloud-based devices with additional processing capabilities.

[0066] In another embodiment, audio device 10 enables two or more users to communicate wirelessly using the functions of controller 340. Specifically, controller 340 enables direct communication between two or more users, each wearing audio device 10. These embodiments can have a variety of applications. For example, these embodiments may be advantageous in environments requiring continuous low-latency communication, such as in noisy factories or dispersed work environments. These embodiments may also be advantageous in environments where internet or other internal network connections are unreliable, such as in remote environments (e.g., offshore or off-grid environments, such as oil drilling platforms). These embodiments may also be advantageous in relatively quiet environments where users may wish to speak to each other in a softer or quieter voice. In noisier environments such as concerts, nightclubs, or sporting events, these embodiments enable users of audio device 10 to communicate without significantly raising their voices. Even further, these specific implementations can help users with hearing impairments communicate with others, for example, in situations where the user of audio device 10 communicates with one or more other users wearing audio device 10 in noisy environments or environments with varying acoustic characteristics (e.g., using components in audio device 10 and / or smart device 360).

[0067] However, compared to conventional configurations, the audio device 10 disclosed according to various specific embodiments is able to isolate the user's own voice, for example, for transmission to other users. For example, in some conventional configurations, devices close to each other will detect not only the voice of a first user (e.g., when a user wearing the first device is speaking), but also the voice of a second user (e.g., when a different user wearing the second device is speaking). In this sense, when communication relies on the transmission of detected voice signals, the second user may hear not only the voice of the first user, but also the user's own voice as detected by the first device and played back after transmission between devices (e.g., an echo). This echo phenomenon can be annoying and hinder communication. In contrast, the controller 340 is configured to use only signals from the VAD device 330 (and / or the VAD system 390, where applicable) to transmit the voice pickup of the user from the wearable audio device 10, thereby avoiding or significantly reducing the pickup of the user's own voice at another device (e.g., an echo). The benefits of these implementations become even more apparent when two or more people are speaking using audio devices, such as in larger conversations, work environments, sporting events, coordinated group tasks or missions.

[0068] In some of these cases, controller 340 is also programmed to detect when users of audio devices 10 are within a defined proximity to each other. Proximity detection can be performed according to a variety of known technologies, including device detection via communication protocols (e.g., Bluetooth or BLE), public networks or cellular connections, location information (e.g., GPS data), etc. In some cases, device 10 operating controller 340 is configured in one or more operating modes to share location information with other devices 10 also operating controller 340. Other aspects of proximity detection are described in U.S. Patent Application No. 16 / 267,643 (Location-Based Personal Audio, filed February 5, 2019), the entire contents of which are incorporated herein by reference. In some cases, when audio devices 10 are close to each other but still transmitting audio detected by VAD, it is possible that a user may hear both their own voice (while speaking) and any part of the user's speech detected at a second audio device 10. In these cases, controller 340 is configured to change the volume of the transmitted audio based on the proximity between the audio devices 10. For example, it may decrease the volume of the transmitted audio when controller 340 receives an indicator that device 10 is getting closer, and increase the volume in response to an indicator that device 10 is getting farther away. In some cases, controller 340 is also configured to detect whether a speech signal received from another audio device 10 is also detectable due to proximity (e.g., picked up via microphone 320). That is, controller 340 at the first audio device 10 is configured to adjust the output at the first audio device 10 when the second user's speech (transmitted as a VAD-based signal from the second audio device 10) can also be detected in an environment close to the first user. In some cases, controller 340 modifies these settings based on whether noise cancellation is activated at the first audio device 10. For example, when controller 340 detects that noise cancellation is activated or set to cancel significant noise, controller 340 allows the signal detected by the VAD of the second user's audio device 10 to be played without modification. In other cases, when controller 340 detects that noise cancellation is not activated or is set to a low level, controller 340 stops playback of the signal detected by the VAD to avoid interfering with the outdoor voice signal heard by the first user. In any case, these dynamic adjustments by controller 340 significantly improve the user experience compared to conventional systems and methods.

[0069] The functions or portions thereof described herein, and their various modifications (hereinafter referred to as "functions") may be implemented at least in part by computer program products, such as computer programs tangibly implemented in an information carrier, such as one or more non-transitory machine-readable media, for performing or controlling the operation of one or more data processing devices, such as programmable processors, computers, multiple computers and / or programmable logic components.

[0070] Computer programs can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. Computer programs can be deployed on a single computer, distributed across one or more sites, or executed on multiple computers interconnected via a network.

[0071] The actions associated with implementing all or part of the functionality can be performed by one or more programmable processors executing one or more computer programs to perform the functions of the calibration process. All or part of the functionality can be implemented as special-purpose logic circuitry, such as FPGAs and / or ASICs (Application-Specific Integrated Circuits). Processors suitable for executing computer programs include, for example, both general-purpose microprocessors and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Generally, the processor will receive instructions and data from read-only memory or random access memory, or both. The components of a computer include a processor for executing instructions and one or more memory devices for storing instructions and data.

[0072] Additionally, one or more networked computing devices may perform actions associated with implementing all or part of the functions described herein. Networked computing devices may be connected via networks such as one or more wired and / or wireless networks such as local area networks (LANs), wide area networks (WANs), personal area networks (PANs), internet-connected devices and / or networks and / or cloud-based computing (e.g., cloud-based servers).

[0073] In various specific implementations, electronic components described as "coupled" can be linked via conventional hardwired and / or wireless devices, enabling these electronic components to transmit data to each other. Additionally, sub-components within a given component can be considered linked via conventional paths, which may not necessarily be shown.

[0074] The term “approximately” used relative to the values ​​indicated herein can be used to assign nominal variations in absolute values ​​(e.g., a few percent or less).

[0075] Several specific embodiments have been described. However, it should be understood that additional modifications may be made without departing from the scope of the inventive concept described herein, and therefore, other specific embodiments are within the scope of the following claims.

Claims

1. A wearable audio device, comprising: A frame for contacting the user's head; An electroacoustic transducer, the electroacoustic transducer being located within the frame and configured to output an audio signal; At least one microphone; Voice Activity Detection (VAD) Accelerometer; and A controller, coupled to the electroacoustic transducer, the at least one microphone, and the VAD accelerometer, and configured in a first mode: The system detects that the user is speaking; as well as In response to the detection that the user is speaking, the user's speech is recorded using only the signal from the VAD accelerometer, such that the signal detected by the at least one microphone signal is excluded.

2. The wearable audio device of claim 1, wherein the VAD accelerometer remains in contact with the user's head during the recording, or is separated from the user's head for at least a portion of the recording.

3. The wearable audio device of claim 1, wherein in the second mode, the controller is configured to adjust the directionality of audio pickup from the at least one microphone to at least one of: verifying that the user is speaking or improving the quality of the recording.

4. The wearable audio device of claim 1 further includes an additional VAD system for verifying that the user is speaking before initiating the recording.

5. The wearable audio device of claim 1, wherein the controller is further configured to, in response to detecting that only the user is speaking, communicate with a smart device connected to the wearable audio device to initiate natural language processing (NLP) of commands in the user's speech detected at the at least one microphone on the wearable audio device or a microphone array on the smart device.

6. The wearable audio device of claim 5, wherein the NLP is performed after detecting that only the user is speaking, without requiring a wake word.

7. The wearable audio device of claim 1, wherein the controller is further configured to: Request feedback from the user to verify that the user is speaking; The logic engine is trained based on the received response to the feedback request to identify that the user is speaking; as well as After the training, the logic engine is run to detect future instances of the user's speech so that recording can be achieved using only the VAD accelerometer.

8. The wearable audio device of claim 1 further includes a memory for storing a predetermined amount of voice recordings from the user.

9. The wearable audio device of claim 8, wherein the recording of the user's voice is accessible via a processor to perform at least one of the following: a) Analyze the speech recording for at least one of speech pattern or speech tone; b) Play back the speech recording in response to a request from the user; or c) Execute virtual personal assistant (VPA) commands based on the voice recording.

10. The wearable audio device of claim 1, wherein in the second mode, the controller activates the at least one microphone to record all detectable ambient audio, wherein the controller is configured to switch from the first mode to the second mode in response to a user command.

11. The wearable audio device of claim 1, wherein the controller is further configured to use the VAD accelerometer and the at least one microphone to record voice-over audio for a television or streaming program: The at least one microphone is used to fingerprint the audio output associated with the television or streaming program, while the signal from the VAD accelerometer is used to record the user's voice. The audio output of the fingerprint recognition is compiled with the user's recorded voice. as well as Provides compiled fingerprint recognition audio output and the user's recorded voice for subsequent playback synchronized with the television or streaming program.

12. The wearable audio device of claim 1, further comprising a digital signal processor (DSP) coupled to the controller, wherein the controller is further configured to activate the DSP to enhance the user's recorded speech during the recording.

13. The wearable audio device of claim 1, wherein the controller is further configured to initiate playback of the recording by at least one of the following: a) Accelerate the playback of the recorded data; b) Play back only a selected portion of the record; or c) Adjust the playback speed of one or more selected portions of the recording.

14. A computer-implemented method, comprising: In a wearable audio device comprising: a frame for contacting a user’s head; an electroacoustic transducer located within the frame and configured to output an audio signal; at least one microphone; and a voice activity detection (VAD) accelerometer; In the first mode: The system detects that the user is speaking; and In response to the detection that the user is speaking, the user's speech is recorded using only the signal from the VAD accelerometer, such that the signal detected by the at least one microphone signal is excluded.

15. The method of claim 14, further comprising, in a second mode: adjusting the directionality of audio pickup from the at least one microphone to at least one of: verifying that the user is speaking or improving the quality of the recording.

16. The method of claim 14, further comprising verifying that the user is speaking using an input signal from an attached VAD system before initiating the recording.

17. The method of claim 14, wherein in response to detecting that only the user is speaking, communication is made with a smart device connected to the wearable audio device to initiate natural language processing (NLP) of commands in the user's speech detected at the at least one microphone on the wearable audio device or a microphone array on the smart device, wherein the NLP is performed without a wake word after detecting that only the user is speaking.

18. The method of claim 14, further comprising: Request feedback from the user to verify that the user is speaking; The logic engine is trained based on the received responses to feedback requests to identify that the user is speaking; as well as After the training, the logic engine is run to detect future instances of the user's speech so that recording can be achieved using only the VAD accelerometer.

19. The method of claim 14, wherein a predetermined amount of the voice recordings from the user are stored in a memory, and wherein the recordings of the user's voices are accessible via a processor to perform at least one of the following: a) Analyze the speech recording for at least one of speech pattern or speech tone; b) Play back the speech recording in response to a request from the user; or c) Execute virtual personal assistant (VPA) commands based on the voice recording.

20. The method of claim 14, wherein the method further comprises: In the second mode, the at least one microphone is activated to record all detectable ambient audio, wherein a switch from the first mode to the second mode is performed in response to a user command.

Citation Information

Patent Citations

  • Audio signal processing for noise reduction

    US10311889B2

  • Audio eyeglasses with cable-through hinge and related flexible printed circuit

    US10353221B1

  • Location-based personal audio

    US10869154B2

  • Telephone system and telephony gateway for wireless conference call

    CN203435060U