Machine learning based self voice removal
By using machine learning-based self-speech filters and post-processing algorithms, the technical challenge of removing user speech in wearable hearing aids is solved, improving system latency tolerance and audio processing quality, and enhancing user experience.
Patent Information
- Application Number
- CN202180065647.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-25
- Filing Date
- 2021-09-13
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2041-09-13
AI Technical Summary
Existing wearable hearing aids face technical challenges in removing the user's voice signal, especially when system latency is too high, which can significantly degrade the user experience.
A machine learning-based self-speech filter is used to filter the audio signal using the inherent user vector, remove the user speech component, and combine it with post-processing algorithms such as speech enhancement, signal-to-noise ratio improvement and active noise reduction to output an audio signal with high latency tolerance.
It effectively removes user voice, improves system latency tolerance, enhances user experience, supports more complex post-processing algorithms, and improves the overall quality of audio processing.
Smart Images

Figure CN116325805B_ABST
Abstract
Description
[0001] CLAIM OF PRIORITY
[0002] This application claims priority to U.S. Patent Application No. 17 / 032,801, filed September 25, 2020, which is incorporated by reference in its entirety. TECHNICAL FIELD
[0003] The present disclosure relates generally to wearable hearing assistance devices. More specifically, the present disclosure relates to machine learning based methods for removing user speech signals in wearable hearing assistance devices. BACKGROUND
[0004] Wearable hearing assistance devices can significantly improve a user’s auditory experience. For example, such devices often employ one or more microphones and amplification components to amplify desired sounds, such as one or more sounds of other people conversing with the user. Additionally, such devices can employ techniques such as active noise reduction (ANR) to contend with unwanted environmental noise. Wearable hearing assistance devices can have various form factors, e.g., headsets, earbuds, audio eyewear, etc. However, removing unwanted acoustic signals, such as a user’s own sound signals, continues to present various technical challenges. SUMMARY
[0005] All examples and features mentioned below can be combined in any technically possible manner.
[0006] Systems and methods employing wearable hearing assistance devices that remove user speech are disclosed. Some implementations include receiving an audio signal, wherein the audio signal includes a speech component of a user and a noise component; filtering the audio signal with a self-speech filter that filters the speech component with an intrinsic user vector, wherein the intrinsic user vector is determined based on speech input of the user; and outputting a filtered audio signal in which the speech component of the user has been substantially removed from the audio signal.
[0007] In additional particular implementations, a system is provided that includes a memory and a processor coupled to the memory and configured to remove user speech for a hearing assistance device according to a method including receiving an audio signal, wherein the audio signal includes a speech component of a user and a noise component; filtering the audio signal with a self-speech filter that filters the speech component with an intrinsic user vector, wherein the intrinsic user vector is determined based on speech input of the user; and outputting a filtered audio signal in which the speech component of the user has been substantially removed from the audio signal.
[0008] Implementations can include one of the following features, or any combination thereof.
[0009] In some cases, the intrinsic user vector is determined during an offline session and includes an intrinsic compressed representation of a voice pattern of the user.
[0010] In other cases, the intrinsic compressed representation includes one of a d-vector representation or an i-vector representation.
[0011] In certain aspects, the self-voice filter includes a machine learning model trained offline according to a method including the steps of: providing a set of intrinsic user vectors and an associated mixed audio signal for each unique user, each intrinsic user vector corresponding to a unique user, wherein each associated mixed audio signal includes a speech component of the unique user and a corresponding noise component; inputting a selected intrinsic user vector and the associated mixed audio signal into the machine learning model; training the machine learning model to remove the speech component from the associated mixed audio signal based on the selected intrinsic user vector; and repeating the inputting and training steps for each intrinsic user vector.
[0012] In particular implementations, the method includes post-processing the filtered audio signal to generate an enhanced filtered audio signal; and outputting the enhanced filtered audio signal to an electro-acoustic transducer to generate an acoustic signal.
[0013] In some cases, the post-processing includes at least one of: speech enhancement, signal-to-noise ratio (SNR) improvement, beamforming, or active noise reduction.
[0014] In certain aspects, the hearing assistance device includes a head-worn device having at least one earpiece and an accessory device in communication with the head-worn device.
[0015] In some implementations, the self-voice filter is included in the accessory device.
[0016] In certain cases, the self-voice filter is included in the head-worn device.
[0017] In some implementations, the offline session is used to determine the intrinsic user vector with a voice characterization system including a machine learning model trained using a database of voice inputs from a plurality of speakers.
[0018] In certain aspects, the self-voice filter includes at least one audio processing enhancement selected from the group consisting of: signal-to-noise ratio (SNR) improvement, beamforming, or active noise reduction.
[0019] In other aspects, a system is provided that performs the following method: receiving an audio signal, where the audio signal includes a target speech component of a non-device user and a noise component; filtering the audio signal with a target speech enhancer that substantially only passes the target speech component using an intrinsic user vector, where the intrinsic user vector is determined based on speech input of the non-device user; and outputting a speech enhanced audio signal that substantially only contains the target speech component.
[0020] Two or more features described in this disclosure, including those described in the SUMMARY, can be combined to form specific implementations not specifically described herein.
[0021] The details of one or more implementations are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF DRAWINGS
[0022] Figure 1 Block diagrams of wearable hearing assistance devices are depicted in accordance with various implementations.
[0023] Figure 2 Machine learning platforms for implementing self-speech removal are depicted in accordance with various implementations.
[0024] Figure 3 Alternative machine learning platforms for implementing self-speech removal are depicted in accordance with various implementations.
[0025] Figure 4 Machine learning platforms for implementing target speech enhancement are depicted in accordance with various implementations.
[0026] Figure 5 Examples of wearable hearing assistance devices are depicted in accordance with various implementations.
[0027] It is noted that the figures of the various implementations are not necessarily drawn to scale. The figures are intended to show typical aspects of the disclosure, and therefore should not be considered limiting of the scope of the implementations. In the drawings, like numbering represents similar elements between the figures. DETAILED DESCRIPTION
[0028] Various implementations describe solutions for removing a user’s own sound in a wearable hearing assistance device (“self-speech removal”). Typically, when using a hearing assistance device, a user can be disturbed or otherwise annoyed by the amplification of the user’s own sound. However, the amplification of other people’s sounds is critical for audibility.
[0029] In hearing assistance devices, such as hearing aids, audio augmented reality systems, systems that utilize a remote microphone streamed to a headset (e.g., from a phone or other device), etc., sound is transmitted to the ear via two different paths. The first path is the “direct path,” in which sound propagates around the device or headset and directly into the ear canal. In the second “processed path,” audio that propagates through the hearing assistance device or headset is processed and then delivered to the ear canal by a driver (i.e., an electrostatic transducer or speaker).
[0030] In processing real-time audio for human use, a key factor in the performance of any hearing assistance device is the latency produced by the algorithms that process the signal along the processing path. Latency is defined as the delay between the time that audio enters the device and the time that the audio appears (typically measured in milliseconds). Too much latency in a system can significantly degrade the perceived algorithmic quality and user experience.
[0031] If the overall system latency is above a certain threshold (e.g., about 5-20 milliseconds), it can be extremely disconcerting and annoying to the user. More precisely, this threshold is driven by the human sensitivity to latency on their own voice in the processing path. However, if the user’s own voice is not present in the processing path, the user can better tolerate latency and the overall system latency threshold increases significantly (e.g., up to about 80-150 ms).
[0032] The present disclosure describes implementations that use machine learning-based processing to remove the user’s own voice (i.e., self-voice) from the processing path, effectively increasing the latency budget for audio processing.
[0033] Although generally described with reference to hearing assistance devices, the solutions disclosed herein are intended to be applicable to a wide variety of wearable audio devices, i.e., devices that are configured to be worn at least partially by a user near at least one of the user’s ears to provide amplified audio to the at least one ear. Other such implementations can include headsets, two-way communication headsets, earphones, earbuds, hearing aids, audio eyewear, wireless headsets (also known as “headsets”), and earmuffs. The presentation of particular implementations is intended to aid in understanding by the use of examples and should not be taken as limiting the scope of the disclosure or the scope covered by the claims.
[0034] Additionally, the solutions disclosed herein are applicable to wearable audio devices that provide two-way audio communication, one-way audio communication (i.e., acoustic output of audio provided electronically by another device), or no communication at all. Moreover, the content disclosed herein is applicable to wearable audio devices that are wirelessly connected to other devices, connected to other devices through conductive and / or light-transmissive cables, or not connected to any other devices at all. These teachings are applicable to wearable audio devices having physical configurations that are structured to be worn near one or both ears of a user, including and not limited to earphones having one or two earpieces, headphones, behind-the-neck earphones, headsets having a communication microphone (e.g., a boom microphone), in-ear or behind-the-ear hearing aids, wireless headsets (i.e., headsets), audio eyewear, monaural or binaural earphones, and hats, helmets, clothing, or any other physical configurations that incorporate one or two earpieces to enable audio communication and / or ear protection.
[0035] In illustrative implementations, the processed audio can include any natural or artificial sound (or acoustic signal), and the microphone(s) can include one or more microphones capable of capturing sound and converting it into an electronic signal.
[0036] In various implementations, the wearable audio devices (e.g., hearing assistance devices) described herein can incorporate active noise reduction (ANR) functionality, which can include either or both feedback-based ANR and feed-forward based ANR, in addition to possibly further providing pass-through audio and audio processed through typical hearing aid signal processing such as dynamic range compression.
[0037] Additionally, the solutions disclosed herein are intended to be applicable to a wide variety of accessory devices, i.e., devices that can communicate with the wearable audio devices and assist in processing audio signals. Illustrative accessory devices include smartphones, Internet of Things (IoT) devices, computing devices, specialized electronics, vehicles, computerized agents, carrying cases, charging cases, smartwatches, other wearable devices, and the like.
[0038] In various implementations, the wearable audio devices (e.g., hearing assistance devices) and accessory devices communicate wirelessly, e.g., using Bluetooth or other wireless protocols. In certain implementations, the wearable audio devices and accessory devices are several meters apart from each other.
[0039] Figure 1An illustrative implementation of a wearable hearing assistance device 100 is depicted that utilizes a machine learning (ML) based approach to output a processed audio signal 126 in which self-voice 118 is removed or substantially removed. As shown, the device 100 includes a set of microphones 114 configured to receive a mixed acoustic input 115 that includes a mixture of self-voice 118 and noise 120. The noise 120 generally includes all other acoustic inputs other than the self-voice 118, such as other voices, background sounds, environmental sounds, music, etc. The microphone input 116 receives the mixed audio signal from the microphones 114 and passes the mixed audio signal 128 to an audio processing system 102.
[0040] The audio processing system 102 includes a self-voice filter 104 that includes a trained machine learning (ML) model configured to process the mixed audio signal 128. More specifically, the self-voice filter 104 removes the user’s self-voice 118 using an intrinsic user vector 112 that is created based on the user’s voice, for example, during a registration process (or phase) 122. In certain implementations, the intrinsic user vector 112 that includes a compressed representation of the user’s voice pattern is determined during the registration process 122. In other implementations, the intrinsic user vector 112 can be computed by the device 100. Reference is made herein, for example, to Figure 2 The ML process for generating the intrinsic user vector 112 and the self-voice filter 104 is further described. Once the mixed audio signal 128 is processed with the self-voice filter 104 to remove the self-voice of the user (i.e., the user of the wearable hearing assistance device 100), a post-processing algorithm 106 can further enhance the filtered signal, for example, by providing speech enhancement, signal-to-noise ratio (SNR) improvement, beamforming, active noise reduction, etc. The resulting signal is then amplified by an amplifier system 108 and output via an electro-acoustic transducer 124.
[0041] Figure 2 An illustrative overview of a machine learning based process for implementing the self-voice filter 104 and computing the intrinsic user vector 112 (such as shown by the device 100 Figure 1 ) is depicted. The process can be implemented in three phases, including: an offline training phase, a registration phase, and an inference phase. Figure 1 Implementations of the inference phase are described in which the device 100 filters the self-voice of the user during operation of the device 100 by the user. The registration phase can be implemented, for example, as a one-time operation or an on-demand operation when a new user begins using the device 100. In various implementations, the offline training phase is responsible for training the two ML based models required by the other phases (i.e., the speech characterization system 103 and the self-voice filter 104).
[0042] In some implementations, the speech characterization system 103 (first model) is trained using a database (DB) of user utterances 202 and associated intrinsic user vectors DB 204, e.g., using a neural network or the like. For example, the first model can be trained with pairs of user utterances / intrinsic user vectors expected. A user utterance is input into the speech characterization system 103 to generate an intrinsic user vector. The generated intrinsic user vector can be compared to the expected intrinsic user vector from the intrinsic user vector DB 204, with the results fed back to adjust the first model. Training can continue until each user utterance from the database (DB) of user utterances 202 input into the model reliably outputs its expected intrinsic user vector. Once trained, the resulting speech characterization system 103 can be used in a registration phase to map new (or registered) user utterances to intrinsic user vectors 112 that are explicitly associated with the new user. The intrinsic user vectors 112 are typically composed of an intrinsic compressed representation of the user’s speech pattern. The intrinsic compressed representation can include, for example, a d-vector representation or an i-vector representation. As is readily understood in the art, there are various algorithms to extract d-vectors and i-vectors based on a user’s speech samples.
[0043] The auto-speech filter 104 (second model) is likewise trained offline using a neural network or the like, trained to remove a person’s speech from a mixed audio signal (i.e., an audio signal containing both auto-speech 118 and noise 120) using that person’s intrinsic user vector. For example, this model can be trained by inputting mixed audio signals from a mixed audio signal DB 206 along with associated intrinsic user vectors to obtain auto-speech filtered signals that can be compared to expected signals in an auto-speech filtered signal database 208. The results can be fed back to adjust and tune the second model. Once trained, the resulting auto-speech filter 104 can be deployed in an inference phase (e.g., on the device 100).
[0044] During the registration phase, for example, the user is instructed to record short audio clips of them saying a set of predefined utterances. This can be done on the device 100 itself or on an accessory device such as a mobile phone or co-processor. A system or guidance can be utilized to ensure that the recording contains only the user’s voice and no other interfering audio or speech. For example, the user can be instructed to speak the utterances in a quiet space, and / or the resulting audio clips can be analyzed to ensure that there is no unwanted noise. The recorded registration utterances are then passed through the speech characterization system 103 (on the device 100 itself, on a companion device, in the cloud, etc.) and the intrinsic user vectors 112 of the registered user are extracted.
[0045] Once the intrinsic user vector 112 is extracted, it is loaded and stored on, for example, the device 100, and the inference phase can be implemented. During this phase, the mixed audio 212 containing the user's own speech is first sent to the own voice filter 104 along with the user's intrinsic vector 112 to remove the user's own sound from the processing path. The resulting self-voice filtered audio 214 is then passed to any downstream post-processing algorithms 106, such as speech enhancement, SNR improvement, beamforming, ANR, etc. Finally, the self-voice filtered enhanced audio 218 is sent to the drivers to be reproduced in the user's ear canal.
[0046] Since the user's own sound has been removed from the processing path during the inference phase, the user is able to tolerate much higher processing delays than audio signals containing own voice. This higher tolerance to delay makes more complex downstream post-processing algorithms 106 possible, such as algorithms implemented on accessories such as smart phones or wireless co-processor devices, which accumulate not only delays from the algorithm processing but also delays from the wireless transmission of the audio.
[0047] Figure 3 An alternative embodiment is depicted in which Figure 2 The own voice filter 104 and post-processing algorithms 106 of FIG. 1 are combined into a single system, i.e., an own voice filter and post-processing system 105. In this way, processing enhancements such as speech enhancement, SNR improvement, beamforming, ANR, etc. are handled by a single system or machine learning trained model.
[0048] Figure 4 A further specific implementation is depicted in which speech enhanced audio 215 is provided utilizing a separate or similar machine learning platform other than the own voice filtering platform described herein. In this case, the platform creates an intrinsic user vector that is designed to pass speech from one or more non-device users (i.e., target speech). The trained model (in this case, a target speech enhancer 107) is trained to preserve the target speech of the non-device users, as opposed to the trained model being trained to remove speech corresponding to the intrinsic user vector (i.e., the user's own speech). The target speech enhancer 107 is trained by processing mixed audio signals (which include target speech and noise) along with the associated intrinsic user vector to reliably generate speech enhanced audio 209 that includes only target speech. The speech representation system 103 is trained in the same manner as the own voice filtering platform. Figure 2
[0049] Accordingly, a user of device 101 can elect to hear only speech from their companion by registering their companion with speech characterization system 103 and loading their companion’s intrinsic user vector 112 into target speech enhancer 107. When mixed audio 213 containing speech of their companion (i.e., target speech) is presented to device 101, target speech enhancer 107 generates speech enhanced audio 215 that includes substantially only speech of their companion. Processed enhanced audio 219 can then be generated from post-processing algorithm 106. In alternative embodiments, post-processing algorithm 106 can be incorporated into target speech enhancer 107.
[0050] The process can be extended to handle multiple targets, simply by registering multiple users, each of which will generate an intrinsic user vector 112, to allow speech of those targets to be passed. Multiple intrinsic user vectors 112 can then be input to target speech enhancer 107.
[0051] It will be appreciated that devices 100, 101 Figures 1 to 4 may be configured to be worn by a user to provide audio output in the vicinity of at least one ear of the user. Devices 100, 101 can have any of a variety of form factors, including configurations incorporating a single earpiece to provide audio to only one ear of the user, other configurations incorporating a pair of earpieces to provide audio to both ears of the user, and other configurations incorporating one or more independent speakers to provide audio to the environment surrounding the user. An exemplary wearable audio device is shown and described in greater detail in U.S. Patent No. 10,194,259 (Directional Audio Selection, filed February 28, 2018), which is hereby incorporated by reference in its entirety.
[0052] In illustrative implementations, acoustic input 115 Figure 1 may include any environmental acoustic signal, including acoustic signals generated by a user of a wearable hearing assistance device, including, for example, natural or artificial sounds (or acoustic signals). Microphone 114 can include one or more microphones (e.g., one or more microphone arrays including front- and / or feedback microphones) capable of capturing sound and converting it into an electronic signal.
[0053] It will be appreciated that while several examples have been provided herein relating to removal of user sound in a hearing assistance device, other methods or combinations of the described methods can also be used.
[0054] Figure 5This is a schematic diagram of an exemplary wearable hearing aid 300 (in an exemplary form factor) including a self-speech removal system contained within a housing 302. It should be understood that the exemplary wearable hearing aid 300 may include, relative to a reference... Figures 1 to 4 Some or all of the components and functions described in the depicted and described devices 100, 101. The self-speech removal system and / or targeted speech enhancer may be one of a variety of electronic devices 304 in the housing 302, or may be implemented on one or more of these electronic devices. In some embodiments, some or all of the self-speech or enhancement algorithms may be implemented in an accessory 330 configured to communicate with the wearable hearing aid device 300. In this example, the wearable audio device 300 includes an audio headset comprising two earpieces (e.g., in-ear headphones, also referred to as “earplugs”) 312, 314. While the earpieces 312, 314 are attached to the housing 302 (e.g., a neckband) configured to rest on the user's neck, other configurations, including wireless configurations, may also be utilized. Each earpiece 312, 314 is shown as including a body 316, which may include a shell formed of one or more plastics or composite materials. The main body 316 may include a mouthpiece 318 for insertion into the user's ear canal entrance and a support member 320 for holding the mouthpiece 318 in a stationary position within the user's ear. In addition to the self-voice removal system, the control unit 302 may also include other electronic devices 304, such as an amplifier, a battery, user controls, a voice activity detection (VAD) device, etc.
[0055] In some specific implementations, as described above, the separate accessory 330 may include a communication system 332 for wireless communication with device 300, and a remote processing 334 to provide some or all of the functions described herein, such as self-speech filter 104, registration process 122, post-processing algorithm 106, etc. Accessory 330 may be implemented in many embodiments. In one embodiment, accessory 330 includes a standalone device. In another embodiment, accessory 330 includes a user-supplied smartphone that utilizes a software application to enable remote processing 334 while using smartphone hardware for communication system 332. In another embodiment, accessory 330 may be implemented within the charging case of device 300. In yet another embodiment, accessory 330 may be implemented within a companion microphone accessory that also performs other functions, such as off-head beamforming and wirelessly streaming beamformed audio to device 300. Additionally, other wearable device forms may also be implemented, including earbuds, over-ear headphones, audio glasses, open-ear audio devices, etc.
[0056] refer to Figure 1 The microphone set 114 may include an in-ear microphone. (See reference)Figure 5 Such an in-ear microphone can be integrated into the earbud body 316, for example into the sound outlet 318. The in-ear microphone can also be used to perform feedback active noise reduction (ANR) and speech pickup for communication, which can be performed within the other electronics 304. Due to the bone conduction and tissue conduction of the user's own voice, the user's own voice in the ear canal is significantly different from the sound received from a microphone placed outside the ear canal. The in-ear microphone can provide unique characteristics to the mixed audio signal DB 206, the user-registered utterance 210, and the mixed audio containing the user's own voice 212 (refer to Figure 2 ) in combination with the other microphones in the set of microphones 114. The in-ear microphone and the out-of-ear microphone together provide more information to the speech characterization system 103, thereby improving the performance of the own voice filter 104.
[0057] According to various implementations, a hearing assistance device is provided that will filter the user's own voice for enhanced performance. In particular, the own voice filter 104 that utilizes the inherent user vector 112 will remove the user's voice from the mixed audio signal.
[0058] It should be understood that one or more functions of the described system can be implemented as hardware and / or software, and that various components can include communication paths that connect components by any conventional means, such as hardwired and / or wireless connections. For example, one or more non-volatile devices (e.g., centralized or distributed devices such as flash memory devices) can store and / or execute programs, algorithms, and / or parameters for one or more of the described devices. In addition, the functions described herein, or portions thereof, and various modifications thereof (hereinafter referred to as "the functions") can be implemented, at least in part, via a computer program product, for example, a computer program tangibly embodied in an information carrier, such as one or more non-transitory machine-readable media, for execution by, or to control the operation of, one or more data processing apparatus, such as a programmable processor, a computer, multiple computers, and / or a programmable logic component.
[0059] A computer program can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program can be deployed to be executed on one computer or on multiple computers that are distributed and interconnected through a network.
[0060] Actions associated with implementing all or part of the functionality can be performed by one or more programmable processors executing one or more computer programs to perform the functions of the functionality. All or part of the functionality can be implemented as, special purpose logic circuitry, e.g., an FPGA (field- programmable gate array) and / or an ASIC (application-specific integrated circuit). Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The elements of a computer include a processor for executing instructions and one or more memory devices for storing instructions and data.
[0061] It should be noted that while the specific implementations described herein utilize a microphone system to collect input signals, it should be understood that any type of sensor can be utilized, alone or with a microphone system, to collect input signals, such as an accelerometer, a thermometer, an optical sensor, a camera, etc.
[0062] Additionally, actions associated with implementing all or part of the functionality described herein can be performed by one or more networked computing devices. The networked computing devices can be connected through a network, e.g., one or more wired and / or wireless networks, such as a local area network (LAN), a wide area network (WAN), a personal area network (PAN), the Internet, Internet connection devices, and / or network- and / or cloud-based computing (e.g., cloud-based servers).
[0063] In various implementations, electronic components described as being “coupled” can be linked via conventional wired and / or wireless means so that the electronic components can communicate data with each other. Additionally, subcomponents within a given component can be considered to be linked via conventional pathways, which can not necessarily be shown.
[0064] A number of implementations have been described. Nevertheless, it will be understood that additional modifications can be made without departing from the scope of the inventive concepts described herein, and, therefore, other implementations are within the scope of the following claims.
Claims
1. A method of removing user speech for a hearing assistance device, comprising: receiving an audio signal, wherein the audio signal includes a speech component and a noise component of a user; filtering the audio signal with a self-speech filter, the self-speech filter filtering the speech component with an intrinsic user vector, wherein the intrinsic user vector is determined during an offline session based on speech input of the user and includes an intrinsic compressed representation of a speech pattern of the user, wherein the intrinsic compressed representation includes one of a d-vector representation or an i-vector representation; and outputting a filtered audio signal in which the speech component of the user has been substantially removed from the audio signal.
2. The method of claim 1, wherein the self-speech filter includes an offline trained machine learning model.
3. The method of claim 1, wherein the self-speech filter includes an offline trained machine learning model according to a method comprising: providing a set of intrinsic user vectors and associated mixed audio signals for each unique user, each intrinsic user vector corresponding to a unique user, wherein each associated mixed audio signal includes a speech component and a corresponding noise component of the unique user; inputting a selected intrinsic user vector and the associated mixed audio signal into the machine learning model; training the machine learning model to remove the speech component from the associated mixed audio signal based on the selected intrinsic user vector; and repeating the inputting step and the training step for each intrinsic user vector.
4. The method of claim 1, further comprising: post-processing the filtered audio signal to generate an enhanced filtered audio signal; and outputting the enhanced filtered audio signal to an electro-acoustic transducer to generate an acoustic signal.
5. The method of claim 4, wherein the post-processing includes at least one of speech enhancement, signal-to-noise ratio (SNR) improvement, beamforming, or active noise cancellation.
6. The method of claim 1, wherein the hearing assistance device includes a headset having at least one earpiece and an accessory device in communication with the headset.
7. The method of claim 6, wherein the self-speech filter is contained in the accessory device.
8. The method of claim 6, wherein the self-speech filter is contained in the headset.
9. The method of claim 6, wherein the hearing assistance device includes an earpiece having an in-ear microphone in communication with the accessory device.
10. The method of claim 1, wherein the offline session for determining the intrinsic user vector utilizes a speech characterization system including a machine learning model trained using a database of speech input from a plurality of speakers.
11. A system, comprising: a memory; and a processor coupled to the memory and configured to remove user speech for a hearing assistance device according to a method comprising: receiving an audio signal, wherein the audio signal includes a speech component of a user and a noise component; filtering the audio signal with a self-speech filter that filters out the speech component using an intrinsic user vector, wherein the intrinsic user vector is determined during an offline session based on speech input of the user and includes an intrinsic compressed representation of a speech pattern of the user, wherein the intrinsic compressed representation includes one of a d-vector representation or an i-vector representation; and outputting a filtered audio signal in which the speech component of the user has been substantially removed from the audio signal.
12. The system of claim 11, wherein the self-speech filter includes an offline trained machine learning model.
13. The system of claim 11, wherein the self-speech filter includes a machine learning model that is offline trained according to a method comprising: providing a set of intrinsic user vectors, each intrinsic user vector corresponding to a unique user, and an associated mixed audio signal for each unique user, wherein each associated mixed audio signal includes a speech component of the unique user and a corresponding noise component; inputting a selected intrinsic user vector and the associated mixed audio signal into the machine learning model; training the machine learning model to remove the speech component from the associated mixed audio signal based on the selected intrinsic user vector; and repeating the inputting step and the training step for each intrinsic user vector.
14. The system of claim 11, further comprising: post-processing the filtered audio signal to generate an enhanced filtered audio signal; and outputting the enhanced filtered audio signal to an electro-acoustic transducer to generate an acoustic signal.
15. The system of claim 14, wherein the post-processing includes at least one of: speech enhancement, signal-to-noise ratio (SNR) improvement, beamforming, or active noise reduction.
16. The system of claim 11, wherein the hearing assistance device includes a head-worn device having at least one earpiece and an accessory device in communication with the head-worn device.
17. The system of claim 16, wherein the self-speech filter is contained in one of the accessory device or the head-worn device.
18. The system of claim 16, wherein the earpiece includes an in-ear microphone in communication with the accessory device.
19. The system of claim 11, wherein the offline session to determine the intrinsic user vector utilizes a speech characterization system that includes a machine learning model trained using a database of speech input from a plurality of loudspeakers.
20. The system of claim 11, wherein the self-speech filter includes at least one audio processing enhancement selected from the group consisting of: signal-to-noise ratio (SNR) improvement, beamforming, or active noise reduction.
21. A system comprising: a memory; and a processor coupled to the memory and configured to perform a method comprising: receiving an audio signal, wherein the audio signal comprises a target speech component of a non-device user and a noise component; filtering the audio signal with a target speech enhancer that substantially only passes the target speech component with an intrinsic user vector, wherein the intrinsic user vector is determined during an offline session based on speech input of the non-device user and comprises an intrinsic compressed representation of a speech pattern of the user, wherein the intrinsic compressed representation comprises one of a d-vector representation or an i-vector representation; and outputting a speech-enhanced audio signal that substantially only contains the target speech component.
Citation Information
Patent Citations
Directional audio selection
US10194259B1
Hearing device or system comprising user identification unit
CN111698625A
Hearing assistance system with own voice detection
US20100260364A1