Ear-worn device with neural network-based noise modification and / or spatial focusing

By using neural networks for spatial focusing and independent noise control in ear-worn devices, the problem of ear-worn devices having difficulty distinguishing between front and rear sounds in complex environments is solved, achieving more effective noise reduction and speech enhancement.

CN121970373APending Publication Date: 2026-05-01VERTECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
VERTECH CO LTD
Filing Date
2024-08-05
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing ear-worn devices have difficulty reducing noise around the wearer that interferes with the speaker, especially in reverberant environments where it is difficult to distinguish between sounds from the front and back. Conventional beamforming patterns are affected by the wearer's head and have good high-frequency sound performance, but have limited control over noise and speech.

Method used

The system employs a neural network-based spatial focusing technique, which trains the neural network through multiple microphone inputs to distinguish the arrival direction of target speech and interference speech. Different weights are applied for sound focusing, and the volume changes of background noise and interference speech are controlled independently.

Benefits of technology

It improves the noise reduction performance of ear-worn devices in complex environments, enhances the wearer's ability to perceive target speech, and meets the personalized noise reduction needs of different wearers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121970373A_ABST
    Figure CN121970373A_ABST
Patent Text Reader

Abstract

An ear-worn device includes two or more microphones and a noise reduction circuit including a neural network circuit. The neural network circuit is configured to: receive a plurality of audio signals, where at least two of the plurality of audio signals each originate from a different one of the two or more microphones, and / or at least one of the plurality of audio signals is a beamformed audio signal originating from the two or more microphones; and implementing one or more neural network layers trained to perform background noise modification and spatial focusing based on the plurality of audio signals such that the neural network circuitry generates one or more neural network outputs based on the plurality of audio signals. The noise reduction circuit is configured to output an output audio signal based on the one or more neural network outputs, the output audio signal comprising a background noise modified and spatially focused version of a first audio signal of the plurality of audio signals.
Need to check novelty before this filing date? Find Prior Art

Description

Ear-worn devices with neural network-based noise modification and / or spatial focusing Technical Field

[0001] This disclosure relates to ear-worn devices. Some aspects relate to ear-worn devices with neural network-based noise modification and / or spatial focusing. Background Technology

[0002] Ear-worn devices (such as hearing aids) can be used to help people with hearing difficulties hear better. Typically, ear-worn devices amplify the received sound. Some ear-worn devices can also attempt to reduce noise in the received sound. Summary of the Invention

[0003] Reducing noise in the output of ear-worn devices (e.g., hearing aids, cochlear implants, and headphones) is a difficult challenge. Noise reduction is particularly challenging when the wearer is listening to one speaker while other interfering speakers are nearby. Inventors have recognized that neural networks can be used in ear-worn devices to improve noise reduction and the reduction of sounds from interfering speakers. Recently, neural networks for separating speech from noise have been developed. Further description of such neural networks for noise reduction can be found in U.S. Patent No. 11,812,225, entitled “Method, Apparatus and System for Neural Network Hearing Aid,” issued November 7, 2023, which is incorporated herein by reference in its entirety.

[0004] For background noise reduction, the inventors have recognized that if a neural network hears noise from a direction-of-arrival (DOA) at a previous time step, it can preemptively eliminate the noise from that DOA at the current time step. From another perspective, sound sources may tend to move slowly over time; therefore, if the neural network has identified a particular sound segment as speech and knows its DOA, it can reasonably infer that other sounds from the same direction are also speech.

[0005] To reduce noise from interfering speakers, conventional ear-worn devices can use beamforming to attenuate sound received from certain directions. This can involve processing sound from different microphones in different ways (e.g., applying different delays to signals received at different microphones). Conventional beamforming (both adaptive and non-adaptive) can provide improved intelligibility because it focuses sound from in front of the wearer (assuming the sound of interest originates from that direction) and attenuates sound from the wearer's sides and rear (e.g., background noise and interfering speakers).

[0006] However, conventional beamforming patterns (e.g., cardioid, supercardioid, hypercardioid, and dipole) can also have drawbacks, including: 1. The theoretical beamforming pattern can be distorted once it is implemented by a microphone placed behind the ear (e.g., in a hearing aid), at least in part due to interference from the wearer's head, torso, and ear; this can impair performance. 2. In reverberant environments, indirect paths can enter from the forward direction. For example, in a reverberant room, when a speaker is speaking directly behind the wearer, the speaker's voice can reverberate throughout the room and enter the earpiece's microphone from in front of the wearer; such sounds may not be attenuated by forward beamforming. 3. Conventional beamforming works better for high-frequency sounds than low-frequency sounds. In other words, conventional beamforming may be better at using high-frequency sounds for sound localization than low-frequency sounds. 4. Generally, there are limitations on how much sound reduction a conventional beamforming pattern can provide. 5. In quiet environments, beamforming can add noise.

[0007] The inventors have addressed these drawbacks by developing a neural network trained to perform spatial focusing, which can be implemented in certain embodiments. Spatial focusing may involve applying different weights to an audio signal based on the location of the sound source or the direction in which the audio signal is generated relative to the device. The location and / or direction of the sound can be derived by the neural network from the temporal differences of the sounds arriving at multiple microphones. The inventors have recognized that, using a single microphone, a neural network cannot adequately distinguish speakers coming from different directions; in other words, the neural network may be unable to distinguish whether the speaker is in front of or behind the wearer (or, more generally, unable to distinguish where the wearer is located). A neural network using input from multiple microphones can overcome this ambiguity. Thus, the neural network can accept multiple input audio signals from two or more microphones on an ear-worn device and be trained to perform spatial focusing. Spatial focusing helps focus on sound from a target direction and reduces sound from other directions. As a concrete example, focusing on sound originating in front of the ear-worn device wearer can help reduce interfering sounds from speakers located behind and to the sides of the ear-worn device wearer; other target directions can also be used. This method allows for the application of weights to sounds from different directions with a greater difference than that possible with conventional beamforming.

[0008] It should be understood that while some spatial focusing modes can focus on sound coming from in front of the wearer of the ear-worn device, others can focus on sound coming from other directions, such as the wearer's sides and / or rear. This type of focusing can be beneficial in certain scenarios, such as when the ear-worn device wearer is driving a vehicle and passengers are positioned to their sides and / or rear.

[0009] Some embodiments of the techniques described herein can additionally generate a target speech signal, a background noise signal, and a distractor speech signal, and these signals can be mixed together using modified levels of the target speech, distractor speech, and / or background noise. For example, the levels of distractor speech and background noise can be reduced, while the level of the target speech can be kept the same or increased. Mixing some noise and some distractor speech back into the target speech signal can help reduce distortion and improve the wearer's environmental awareness in the ear-worn device. Typically, the volume change of the background noise signal can differ from the volume change of the target speech signal by a first volume change difference, and the volume change of the distractor speech signal can differ from the volume change of the target speech signal by a second volume change difference, and the first and second volume change differences can be independently controllable.

[0010] Independent control over the volume of background noise and interfering speech can help achieve different levels of reduction, for example, based on the different preferences of different wearers. The following are non-limiting examples illustrating how independent control over the volume of background noise and interfering speech can be helpful. In a first scenario, the wearer may be sitting at a busy table with multiple conversation partners. Significantly reducing background noise but not significantly reducing any speech (i.e., not significantly reducing interfering speech) can be helpful. In a second scenario, the wearer may be sitting at a busy table with one conversation partner, but there may be loud conversations at nearby tables. Significantly reducing both background noise and interfering speech can be helpful. In a third scenario, the wearer may be sitting in a quiet café, while distracting conversations are taking place nearby. Moderately reducing background noise and significantly reducing interfering speech can be helpful. Attached Figure Description

[0011] Figure 1 shows a view of a hearing aid according to certain embodiments described herein;

[0012] Figure 2 shows the hearing aid of Figure 1 on a wearer according to some embodiments described herein;

[0013] Figure 3 illustrates eyeglasses with a built-in hearing aid according to certain embodiments described herein;

[0014] Figure 4 illustrates a system for operating an ear-worn device according to certain embodiments described herein;

[0015] Figure 5 illustrates the circuitry in an ear-worn device according to certain embodiments described herein;

[0016] Figure 6 illustrates audio signals according to certain embodiments described herein;

[0017] Figure 7 illustrates volume variations according to certain embodiments described herein;

[0018] Figure 8 illustrates a noise reduction circuit in an ear-worn device according to certain embodiments described herein;

[0019] Figure 9 illustrates neural network circuits and masking applications and subtraction circuits according to certain embodiments described herein;

[0020] Figure 10 illustrates neural network circuits and masking applications and subtraction circuits according to certain embodiments described herein;

[0021] Figure 11 illustrates neural network circuits and masking applications and subtraction circuits according to certain embodiments described herein;

[0022] Figure 12 illustrates a noise reduction circuit in an ear-worn device according to certain embodiments described herein;

[0023] Figure 13 illustrates the wide dynamic range compression (WDRC) circuit of Figure 12 in more detail according to some embodiments described herein;

[0024] Figure 14 illustrates the circuitry in an ear-worn device according to certain embodiments described herein;

[0025] Figure 15 illustrates a circuit in an ear-worn device according to certain embodiments described herein;

[0026] Figure 16 illustrates a circuit in an ear-worn device according to certain embodiments described herein;

[0027] Figure 17 illustrates a circuit in an ear-worn device according to certain embodiments described herein;

[0028] Figure 18 illustrates exemplary spatial focusing modes according to certain embodiments described herein;

[0029] Figure 19 illustrates exemplary spatial focusing modes according to certain embodiments described herein;

[0030] Figure 20 illustrates exemplary spatial focusing modes according to certain embodiments described herein;

[0031] Figure 21 illustrates a circuit for controlling spatial focusing in an ear-worn device according to certain embodiments described herein;

[0032] Figure 22 illustrates a graphical user interface (GUI) for controlling spatial focusing of an ear-worn device according to certain embodiments described herein.

[0033] Figure 23 illustrates a graphical user interface (GUI) for controlling spatial focusing of an ear-worn device according to certain embodiments described herein.

[0034] Figure 24 illustrates a graphical user interface (GUI) for controlling spatial focusing of an ear-worn device according to certain embodiments described herein.

[0035] Figure 25 illustrates a graphical user interface (GUI) for controlling spatial focusing of an ear-worn device according to certain embodiments described herein.

[0036] Figure 26 illustrates a graphical user interface (GUI) for controlling spatial focusing of an ear-worn device according to certain embodiments described herein.

[0037] Figure 27 illustrates a forward-facing supercardioid pattern according to certain embodiments described herein;

[0038] Figure 28 illustrates a rearward supercardioid pattern according to certain embodiments described herein;

[0039] Figure 29 illustrates a forward-facing super-cardioid pattern according to certain embodiments described herein;

[0040] Figure 30 illustrates a rearward supercardioid pattern according to certain embodiments described herein;

[0041] Figure 31 illustrates a forward-facing heart-shaped pattern according to certain embodiments described herein;

[0042] Figure 32 illustrates a rearward-facing heart-shaped pattern according to certain embodiments described herein; and

[0043] Figure 33 illustrates the dipole modes according to some embodiments described herein. Detailed Implementation

[0044] The above aspects and embodiments, as well as additional aspects and embodiments, are described below. These aspects and / or embodiments may be used individually, in combination, or in any combination of two or more, as this disclosure is not limited in this respect. Ear-worn device

[0045] Figure 1 shows a view of a hearing aid 100 according to certain embodiments described herein. The hearing aid 100 can be any ear-worn device or hearing aid described herein. The hearing aid 100 is a receiver-in-canal (RIC) type hearing aid (also known as a receiver-in-the-ear (RITE)). However, any other type of hearing aid (e.g., behind-the-ear, in-the-ear, canine, fully in-the-canal, open-fit, etc.) can also be used. The hearing aid 100 includes a body 111, a receiver wire 113, a receiver 106, and an earpiece 115. The body 111 is coupled to the receiver wire 113, and the receiver wire 113 is coupled to the receiver 106. The earpiece 115 is placed on top of the receiver 106. The body 111 includes a front microphone 102f, a rear microphone 102b, and a user input device 104. Body 111 also includes circuitry not shown in FIG. 1 (e.g., any circuitry described below in addition to receiver 106). When hearing aid 100 is worn, front microphone 102f may be positioned closer to the front of the wearer, and rear microphone 102b may be positioned closer to the rear of the wearer. Front microphone 102f and rear microphone 102b may be configured to receive sound signals and generate audio signals based on the sound signals. Any of the two or more microphones described herein may be front microphone 102f and rear microphone 102b of hearing aid 100. User input device 104 (e.g., a button) may be configured to control certain functions of hearing aid 100, such as leveling, activation of neural network-based noise reduction, etc.

[0046] Receiver wire 113 can be configured to transmit audio signals from body 111 to receiver 106. Receiver 106 can be configured to receive audio signals (i.e., those generated by body 111 and transmitted by receiver wire 113) and generate sound signals based on the audio signals. Earplug 115 can be configured to fit snugly in the wearer's ear and direct the sound signals generated by receiver 106 into the wearer's ear canal.

[0047] In some embodiments, the length of the body 111 may be equal to 2 cm, equal to 5 cm, or between 2 and 5 cm. In some embodiments, the weight of the hearing aid 100 may be less than 4.5 grams. In some embodiments, the spacing between the microphones may be equal to 5 mm, equal to 12 mm, or between 5 mm and 12 mm. In some embodiments, the body 111 may include a battery (not visible in FIG. 1), such as a lithium-ion rechargeable button battery.

[0048] Figure 2 illustrates a hearing aid 100 on wearer 208 according to certain embodiments described herein. Figure 2 shows wearer 208 from the rear, and as shown, the front microphone 102f is closer to the front of wearer 208 and the rear microphone 102b is closer to the rear of wearer 208. Although Figures 1 and 2 show RIC hearing aids, hearing aids with other form factors can also be used.

[0049] Figure 3 illustrates eyeglasses 300 with a built-in hearing aid according to certain embodiments described herein. Eyeglasses 300 can be any ear-worn device or hearing aid described herein. Eyeglasses 300 has a left temple 310, a right temple 312, and a front frame 314. Eyeglasses 300 also includes a receiver 306 connected to each of the left temple 310 and right temple 312. Figure 3 shows a microphone 302 disposed on the left temple 310. It should be understood that microphone 302 may also be disposed on the right temple 312 (but not visible in the figure). It should be understood that microphone 302 may also be disposed on the front frame 314 (but not visible in the figure). Although Figure 3 shows five microphones 302 on the left temple 310, more or fewer microphones may be disposed on the temples or frame. In some embodiments (such as the embodiment of Figure 3, etc.), the entrance for the microphone 302 may be disposed on the inside of the temple and / or frame (i.e., the side facing the wearer's face), thereby reducing the visibility of the entrance to other people. In some embodiments, the entrance for microphone 302 may be located on the upper side of the temple and / or frame to reduce the visibility of the entrance to other persons. In some embodiments, the entrance for microphone 302 may be located on the outer side of the temple and / or frame (i.e., the side facing away from the wearer's face). Any of the two or more microphones described herein may be any of the microphones 302 of the eyeglasses 300. It should be understood that while Figures 1 through 3 illustrate hearing aids and eyeglasses, other ear-worn devices, such as cochlear implants or headphones, may also be used.

[0050] Figure 4 illustrates a system 416 for operating an ear-worn device 400 according to certain embodiments described herein. System 416 includes an ear-worn device 400, a processing device 418, and a wireless communication link 420. The ear-worn device 400 may be, for example, a hearing aid (e.g., hearing aid 100 or glasses 300), a cochlear implant, headphones, or any other ear-worn device. The processing device 418 may be, for example, a smartphone, tablet, or laptop computer. The wireless communication link 420 may be, for example, a Bluetooth or near-field magnetic induction communication (NFMI) communication link. The processing device 418 may communicate with the ear-worn device 400 (i.e., via the wireless communication link 420). The processing device 418 may be configured to send commands to the ear-worn device 400 via the wireless communication link 420 (e.g., configure the ear-worn device 400 in a specific mode). The ear-worn device 400 may be configured to transmit information (e.g., usage data) to the processing device 418 via the wireless communication link 420. It should be understood that, although not shown, system 416 may include multiple ear-worn devices, such as an ear-worn device for wearing on the right ear and an ear-worn device for wearing on the left ear, and processing device 418 may communicate with each device via a wireless communication link.

[0051] Figure 5 illustrates circuitry in an ear-worn device 500 according to certain embodiments described herein. The ear-worn device may be, for example, a hearing aid 100, glasses 300, and / or an ear-worn device 400. The ear-worn device 500 includes a microphone 502, processing circuitry 522, noise reduction circuitry 524, processing circuitry 528, and a receiver 506. The noise reduction circuitry 524 includes neural network circuitry 526. It should be understood that the ear-worn device 500 may include more circuitry and components than shown (e.g., anti-feedback circuitry, calibration circuitry, etc.), and such circuitry and components may be positioned before, after, or between some of the circuitry and components shown in Figure 5.

[0052] In the ear-worn device 500, processing circuitry 522 is coupled between microphone 502 and noise reduction circuitry 524. Noise reduction circuitry 524 is coupled between processing circuitry 522 and processing circuitry 528. Processing circuitry 528 is coupled between noise reduction circuitry 524 and receiver 506. As mentioned herein, if element A is described as being coupled between element B and element C, other elements may exist between element A and B and / or between element A and C. It should be understood that in the ear-worn device 500, neural network circuitry 526 may be downstream of the beamforming circuitry in processing circuitry 522.

[0053] Microphone 502 may include two or more (e.g., 2, 3, 4 or more) microphones. For example, microphone 502 may include two microphones: a front microphone closer to the front of the wearer of the ear-worn device and a rear microphone closer to the rear of the wearer of the ear-worn device (e.g., microphones 102f and 102b in hearing aid 100). As another example, microphone 502 may include more than two microphones in an array (e.g., microphone 302 in glasses 300). As another example, one microphone may be in a first ear-worn device, and another microphone may be in a second ear-worn device wirelessly coupled to the first ear-worn device. Microphone 502 may be configured to receive sound signals and generate audio signals from the sound signals. The audio signals may represent multiple independent audio signals, each generated by one of microphones 502. Therefore, each of the audio signals may originate from one of microphones 502.

[0054] In some embodiments, processing circuitry 522 may include analog processing circuitry. The analog processing circuitry may be configured to perform analog processing on the audio signal received from microphone 502. For example, the analog processing circuitry may be configured to perform one or more of analog preamplification, analog filtering, and analog-to-digital conversion. Therefore, the analog processing circuitry may be configured to generate an analog-processed audio signal from the audio signal received from microphone 502. The analog-processed audio signal may include a plurality of independent signals, each an analog-processed version of one of the audio signals received from microphone 502. As mentioned herein, the analog processing circuitry may include analog-to-digital conversion circuitry, and the analog-processed signal may be a digital signal that has been converted from analog to digital by the analog-to-digital conversion circuitry.

[0055] In some embodiments, processing circuitry 522 may include digital processing circuitry. The digital processing circuitry may be configured to perform digital processing on analog processed audio signals received from analog processing circuitry. For example, the digital processing circuitry may be configured to perform one or more of wind noise reduction, input calibration, and feedback immunity processing. Therefore, the digital processing circuitry may be configured to generate digitally processed audio signals from the analog processed audio signals. The digitally processed audio signals may include multiple independent signals, each of which is a digitally processed version of one of the analog processed audio signals.

[0056] In some embodiments, processing circuitry 522 may include beamforming circuitry. Beamforming circuitry may be configured to generate one or more beamformed audio signals from two or more digitally processed audio signals. Beamformed audio signals may include one or more independent signals, each an independently beamformed version of one or more digitally processed audio signals. In some embodiments, the multiple beamformed audio signals may each have a different beamforming directional pattern. Beamforming will be described in more detail below.

[0057] The noise reduction circuit 524 includes a neural network circuit 526. The neural network circuit 526 can be configured to implement one or more neural network layers, which can be trained to perform noise modification and / or spatial focusing, as described below. The term "noise modification" may be used herein to encompass both the process of obtaining less noise in the output signal than in the input signal (i.e., noise reduction) and the process of obtaining less speech in the output signal than in the input signal. (As described below, in some embodiments, the neural network circuit can be used to obtain a speech-separated version of the signal, and in some embodiments, the neural network circuit can be used to obtain a noise-separated version of the signal.) Therefore, in some embodiments, one or more neural network layers implemented by the neural network circuit 526 can be trained to modify noise. (Further descriptions that can be considered as noise can be found below.) In such embodiments, the output from neural network circuit 526 can be: an audio signal version with less noise (or only speech, such as speech signal 603 described below) input to neural network circuit 526, configured to generate an output (e.g., a mask or spectrogram) of the audio signal version with less noise input to neural network circuit 526; an audio signal version with less speech (or only noise, such as background noise signal 601 described below) input to neural network circuit 526; or an output (e.g., a mask or spectrogram) of the audio signal version with less speech input to neural network circuit 526. In some embodiments, one or more neural network layers implemented by neural network circuit 526 can be trained to perform spatial focusing. In such embodiments, the output from neural network circuit 526 can be a spatially focused version of the audio signal input to neural network circuit 526, or an output (e.g., a mask or spectrogram) of the audio signal input to neural network circuit 526. In some embodiments, one or more neural network layers implemented by neural network circuit 526 can be trained to both modify noise and perform spatial focusing. In such embodiments, the output from neural network circuit 526 may be a noise-modified and spatially focused version of the audio signal input to neural network circuit 526 (e.g., the target speech signal 605 or interfering speech signal 607 described below), or may be configured to generate a noise-modified and spatially focused version of the audio signal input to neural network circuit 526 as the output (e.g., a mask or spectrogram). It should be understood that in some embodiments, a neural network layer may be trained to modify noise, perform spatial focusing, or both. In some embodiments, multiple neural network layers may be trained to modify noise, perform spatial focusing, or both.

[0058] This description can describe one or more neural network layers trained to perform an action or generate output for performing that action. As mentioned herein, one or more neural network layers can be considered trained to perform a specific action if they themselves perform the action, or if they generate output for performing the action. Therefore, it should be understood that one or more neural network layers can be considered trained to perform noise modification even if the neural network itself does not generate a noise-modified audio signal; a neural network that generates an output configured to generate a noise-modified audio signal can still be considered trained to perform noise modification. For example, a neural network can generate a mask configured to generate a noise-modified audio signal. It should also be understood that a neural network can be considered trained to perform spatial focusing even if the neural network itself does not generate a spatially focused audio signal; a neural network that generates an output configured to generate a spatially focused audio signal can still be considered trained to perform spatial focusing. As a non-limiting example, the output may be a mask configured to generate a spatially focused audio signal, a spectrogram, a mask configured to generate a spectrogram, or a metric computed from audio from multiple beams (each of which points to a different angle around the wearer of the ear-worn device). In some embodiments, one or more neural network layers may be configured to output a single output based on multiple input audio signals.

[0059] For example, any neural network layer described herein can be recursive, ordinary / feedforward, convolutional, generative adversarial, attention-based (e.g., Transformer), or graph-based. Typically, a neural network composed of these layers may include an input layer, multiple hidden layers, and an output layer, and these layers may consist of multiple neurons / nodes to which neural network weights can be applied.

[0060] Processing circuitry 528 can be configured to perform further processing on the output of noise reduction circuitry 524. For example, processing circuitry 52628 may include digital processing circuitry configured to perform one or more of wide dynamic range compression and output calibration.

[0061] Receiver 506 (which may be, for example, the same as receiver 106 and / or 306) may be configured to play back the output of processing circuitry 528 as sound to the user's ear. Receiver 506 may also be configured to perform digital-to-analog conversion prior to playback.

[0062] In some embodiments, multiple circuit portions of the ear-worn device 500 may be configured to process audio signals in the frequency domain. In such embodiments, processing circuitry 522 may include a short-time Fourier transform (STFT) circuit configured to transform a short window of the audio signal from the time domain to the frequency domain, and processing circuitry 528 may include an inverse STFT (iSTFT) circuit configured to transform a short window of the audio signal from the frequency domain to the time domain. In some embodiments, portions of the circuitry in the ear-worn device 500 may be configured to process audio signals in the time domain. In some embodiments, the ear-worn device may lack both STFT and iSTFT circuitry.

[0063] Deploying noise reduction techniques can introduce a delay between when a sound source emits sound and when the denoised sound is output to the user. For example, such techniques can introduce a delay between when a speaker speaks and when a listener hears the denoised speech. During face-to-face communication, long delays can create echo perception as both the original sound and the denoised version of the sound are played back to the listener. Additionally, long delays can interfere with how the listener processes incoming sound due to the disconnect between visual cues (e.g., lip movements) and the arrival of the associated sound. To achieve tolerable latency when implementing neural networks in an ear-worn device, the device may need to be capable of performing billions of operations per second. To address the power requirements of this demanding system, neural network circuitry 524 (among other circuits) can be implemented on a chip within the ear-worn device. Therefore, in some embodiments, one or more of the processing circuitry 522, noise reduction circuitry 524 (including neural network circuitry 526), ​​and processing circuitry 528 (or portions of any of these) can be implemented on a single identical chip (i.e., a single semiconductor die or substrate) within the ear-worn device. Further description of the chip (and other elements in some embodiments) for use in an ear-worn device with neural network circuitry can be found in U.S. Patent No. 11,886,974, entitled "Neural NetworkChip for Ear-Worn Device," issued January 30, 2024, the entire contents of which are incorporated herein by reference and are also hereinafter. Noise Reduction Circuitry

[0064] Figure 6 illustrates an audio signal 630a according to certain embodiments described herein. The audio signal 630a (which, as described below, may be one of the audio signals 630 input to a neural network circuit) comprises a background noise signal 601 and a speech signal 603. The speech signal 603 includes a target speech signal 605 and interfering speech signals 607.

[0065] Typically, the goal of noise reduction circuitry (e.g., any noise reduction circuitry described herein) may be to enhance the target speech signal 605. This may include, for example, amplifying the target speech signal 605 and / or attenuating the background noise signal 601 and the interfering speech signal 607. Thus, both background noise and interfering speech can be considered noise and attenuated as part of the noise reduction process. Speech signal 603 may include speech in audio signal 630a, and background noise signal 601 may include background noise in audio signal 630a. More specifically, any speech lacking features that distinguish it from the target speech (described further below) may be considered speech signal 603 of audio signal 630a. Background noise signal 601 may be considered any sound (including speech) that includes features that distinguish it from the target speech. Examples of speech types that include features that distinguish it from the target speech include noisy speech and distant speech. In some embodiments, background noise signal 601 may be equivalent to the remainder when speech signal 603 is subtracted from audio signal 630a. From another perspective, the speech signal 603 can be considered equivalent to the remainder when the background noise signal 601 is subtracted from the audio signal 630a. It should be understood that this relationship can still be true even if the subtraction process is not actually performed. For example, even if the background noise signal 601 is generated independently or through a different process, rather than through subtraction, the background noise signal 601 can still be considered equivalent to the remainder when the speech signal 603 is subtracted from the audio signal 630a.

[0066] The target speech signal 605 can be broadly and qualitatively considered as the portion of speech signal 603 that the wearer of the ear-worn device most wants to hear. The interfering speech signal 607 can be broadly and qualitatively considered as the portion of speech signal 603 that the wearer of the ear-worn device least wants to hear. However, as mentioned above, determining what the target speech signal 605 and the interfering speech signal 607 are can be difficult problems. The techniques described herein can differentiate between the target speech signal 605 and the interfering speech signal 607 based on the direction-of-arrival (DOA) relative to the wearer. More specifically, the target speech signal 605 can be a first spatially focused version of the speech signal 603 of audio signal 630a, and the interfering speech signal 607 can be a second spatially focused version of the speech signal 603 of audio signal 630a (different from the first spatially focused version). The spatially focused version of a signal can be the signal with a spatial focusing mode applied. In other words, spatial focusing can include applying a spatial focusing mode (which may also be called a spatial focusing pattern) to an audio signal. Spatial focusing modes can define different weights (also referred to as gains) applied to an audio signal based on the direction of arrival (DOA) of sound in the audio signal, where the DOA can be defined relative to the wearer of the ear-worn device. In some embodiments, the weight can be equal to 0, equal to 1, or between 0 and 1. In some embodiments, the weight can be equal to or greater than 0. In some embodiments, the weight can be greater than 0, less than 0, equal to zero, or a complex number; negative weights can flip the phase by 180 degrees, while complex weights can rotate the phase by an angle. Mapping weights to DOA provides focusing because higher weights can be applied to sounds originating from a specific direction, and lower weights can be applied to sounds originating from other directions. Applying higher weights to the direction where the target speaker is (or presumably) and lower weights to other directions helps to focus the sound from the target speaker and attenuate sounds from other interfering speakers. As an example, if the target speaker is (or presumably) in front of the wearer of the ear-worn device, the spatial focusing mode can assign higher weights to the DOA in front of the wearer than to the DOA to the sides and back of the wearer.

[0067] In some embodiments, the target speech signal 605 may be equivalent to the speech signal 603 (of audio signal 630a) that has been applied with a first spatial focusing mode. (As mentioned herein, if signal A is equivalent to signal B having a spatial focusing mode applied to signal B, then signal A may be considered to "have" a spatial focusing mode.) The first spatial focusing mode may include different weights applied to speech originating from different directions of arrival (DOA) relative to the wearer of the ear-worn device. The first spatial focusing mode may have a higher weight for the DOA where the target speaker is or is assumed to be located, and a lower weight elsewhere. For example, the first spatial focusing mode may include a higher weight applied to speech originating from a DOA facing in front of the wearer of the ear-worn device than to speech originating from DOAs facing to the side and rear of the wearer. The interfering speech signal 607 may be equivalent to the speech signal 603 of audio signal 630a that has been applied with a second spatial focusing mode. For example, the second spatial focusing mode may have a lower weight for the DOA where the target speaker is or is assumed to be located, and a higher weight elsewhere. In some embodiments, the interfering speech signal 607 can be equivalent to the remainder when the target speech signal 605 is subtracted from the speech signal 603. Alternatively, the target speech signal 605 can be equivalent to the remainder when the interfering speech signal 607 is subtracted from the speech signal 603. Yet another perspective, the second spatial focusing pattern can be the remainder when the first spatial focusing pattern is subtracted from a weighted pattern where the weights at all DOAs are equal to 1. Again, the first spatial focusing pattern can be the remainder when the second spatial focusing pattern is subtracted from a weighted pattern where the weights at all DOAs are equal to 1. Typically, the first and second spatial focusing patterns can be inverses of each other. It should be understood that the above relationships can still be true even if the subtraction process is not actually performed. For example, even if the interfering speech signal 607 is generated independently or through a different process, rather than through subtraction, the interfering speech signal 607 can still be considered equivalent to the remainder when the target speech signal 605 is subtracted from the speech signal 603. It should be understood that the interfering speech signal 607 can represent different interfering speakers weighted based on the second spatial focusing pattern. For example, if the first interfering speaker is at 45 degrees and the second interfering speaker is at 90 degrees, and the second spatial focus specifies a weight of 0.5 at 45 degrees and a weight of 0.8 at 90 degrees, then the interfering speech signal 607 can be equivalent to the audio of the first interfering speaker multiplied by 0.5 plus the audio of the second interfering speaker multiplied by 0.8.

[0068] It should be understood that, since certain spatial focusing modes may include applying weights between 0 and 1 to sound from a particular DOA, the target speech signal 605 may include speech from the speaker (weighted by a certain amount), and the interfering speech signal 607 may also include speech from the same speaker (weighted by different amounts). However, if the speaker is only located at a DOA for which a certain spatial focusing mode applies a weight of 1 or 0, then speech from said speaker may exist only in the target speech signal 605 or the interfering speech signal 607.

[0069] It should be understood that the audio signal 630a can be enhanced by reducing the volume of the background noise signal 601 and the interfering speech signal 607 and / or increasing the volume of the target speech signal 605. Noise reduction circuitry (e.g., any noise reduction circuitry described herein) can generally be configured to generate an output audio signal (e.g., output audio signal 840) comprising the target speech signal 605, the interfering speech signal 607, and the background noise signal 601, such that in the output audio signal, the volume of one or more of the background noise signal 601, the interfering speech signal 607, and the target speech signal 605 may differ from their volume in the audio signal 630a. Specifically, the volume change of the background noise signal 601 may differ from the volume change of the target speech signal 605 by a first volume change difference, and the volume change of the interfering speech signal 607 may differ from the volume change of the target speech signal 605 by a second volume change difference. The volume change between the volume in the audio signal 630a and the volume in the output signal can be measured. In some embodiments, the first volume change difference and the second volume change difference may be different. In some embodiments, the first volume change difference and the second volume change difference can be independently controlled.

[0070] Figure 7 illustrates volume variations according to certain embodiments described herein. Figure 7 shows an increase (A) in the volume of the target speech signal 605 from the audio signal 630a (referred to as the "input") to the output audio signal (referred to as the "output"). A decrease (B) in the volume of the interfering speech signal 607 from the audio signal 630a to the output audio signal. A decrease (C) in the volume of the background noise signal 601 from the audio signal 630a to the output audio signal. The first volume variation difference described above can be CA, and the second volume variation difference described above can be BA.

[0071] Figure 8 illustrates noise reduction circuitry 824 in an ear-worn device according to certain embodiments described herein. The ear-worn device can be any ear-worn device described herein (e.g., hearing aid 100, glasses 300, ear-worn device 400, and / or ear-worn device 500). Noise reduction circuitry 824 can be any noise reduction circuit described herein (e.g., noise reduction circuitry 524). Noise reduction circuitry 824 includes neural network circuitry 826 (which can be any neural network circuit described herein, e.g., neural network circuitry 526), ​​masking and subtraction circuitry 832, and mixing circuitry 834. The ear-worn device may include two or more microphones not shown in Figure 8 (e.g., microphones 102f and 102b, microphone 302, and / or microphone 502).

[0072] The neural network circuit 826 can be configured to receive a plurality of audio signals 630 (including audio signal 630a) such that (1) at least two of the plurality of audio signals 630 each originate from different microphones of two or more microphones of the ear-worn device and / or (2) at least one of the plurality of audio signals 630 is a beamformed audio signal originating from two or more microphones. (As mentioned herein, a first signal can be said to originate from a microphone if the microphone generates a first signal, or if the microphone generates a second signal and the first signal is the result of processing the second signal. In some cases, the processing can be performed on the second signal along with other signals.) Regarding option (1), as an example, one of the plurality of audio signals 630 may be the output of one of the microphones (e.g., a front microphone, such as front microphone 102f, etc.) or a processed version thereof (e.g., the output of the front microphone after processing by the processing circuit 522), and another of the plurality of audio signals 630 may be the output of another of the microphones (e.g., a rear microphone, such as rear microphone 102b, etc.) or a processed version thereof (e.g., the output of the rear microphone after processing by the processing circuit 522). Regarding option (2), as an example, one of the plurality of audio signals 630 may be the result of beamforming the outputs of two microphones (e.g., beamforming an audio signal originating from a front microphone (such as front microphone 102f, etc.) together with an audio signal originating from a rear microphone (such as rear microphone 102b, etc.). The beamforming result may have a specific directional pattern (e.g., cardioid, supercardioid, hypercardioid, or dipole, as non-limiting examples). Further description of beamforming and directional patterns can be found below. In some embodiments, at least two (or all) of the plurality of audio signals 630 may have different beamforming directional patterns. In some embodiments, at least one of the plurality of audio signals 630 may have a forward beamforming directional pattern, and at least one of the plurality of audio signals 630 may have a backward beamforming directional pattern. (As will be further described below, a backward beamforming directional pattern can typically attenuate signals from behind the wearer more than signals from in front of the wearer, and a backward beamforming directional pattern can typically attenuate signals from in front of the wearer more than signals from behind the wearer.) In some embodiments, the plurality of audio signals 630 may include two signals. In some embodiments, the plurality of audio signals 630 may include three signals. In some embodiments, the plurality of audio signals 630 may include four signals. In some embodiments, the plurality of audio signals 630 may include more than four signals. The following are non-limiting examples of groups of audio signals that may be included in the plurality of audio signals 630. In some embodiments, the plurality of audio signals 630 may include a signal having a forward supercardioid directional pattern and a signal having a backward supercardioid directional pattern.In some embodiments, the plurality of audio signals 630 may include signals having a forward cardioid directional pattern and signals having a backward cardioid directional pattern. In some embodiments, the plurality of audio signals 630 may include signals having a forward supercardioid directional pattern and signals having a backward supercardioid directional pattern. In some embodiments, the plurality of audio signals 630 may include signals having a forward cardioid directional pattern, signals having a backward cardioid directional pattern, and signals having a dipole directional pattern. In some embodiments, the plurality of audio signals 630 may include signals having a forward supercardioid directional pattern, signals having a backward supercardioid directional pattern, and signals having a dipole directional pattern. In some embodiments, the plurality of audio signals 630 may be in the frequency domain. In some embodiments, the plurality of audio signals 630 may be in the time domain. In some embodiments, the neural network circuit 826 may be configured to receive the plurality of audio signals 630 together (i.e., not one after another). In some embodiments, the neural network circuit 826 may be configured to process the plurality of audio signals 630 together (i.e., not one after another).

[0073] In some embodiments, neural network circuitry 826 may be configured to implement one or more neural network layers trained to perform noise modification and spatial focusing, such that neural network circuitry 826 generates two or more neural network outputs 836 based on a plurality of audio signals 630. (For simplicity, this description may interchangeably describe receiving signals and generating outputs based on signals performed as by neural network circuitry or one or more neural network layers implemented by neural network circuitry.) In some embodiments, noise reduction circuitry 824 may be configured to obtain a combination of at least two of the following based on two or more neural network outputs 836: (1) speech signal 603 of audio signal 630a, (2) background noise signal 601 of audio signal 630a, (3) target speech signal 105 of audio signal 630a, and (4) interference speech signal 107 of audio signal 630a. Different methods by which noise reduction circuitry 824 obtains these signals based on two or more neural network outputs 836 will be described below.

[0074] In some embodiments, two or more neural network outputs 836 may include one or more of the following: (1) an output configured to generate speech signal 603 (e.g., a mask), (2) an output configured to generate background noise signal 601 (e.g., a mask), (3) an output configured to generate target speech signal 105 (e.g., a mask), and (4) an output configured to generate interfering speech signal 107 (e.g., a mask). The mask may be a real or complex mask that varies with frequency. Therefore, when a mask is applied (e.g., multiplied or added to) an audio signal, it can operate differently on different frequency components of the audio signal. In other words, a mask can cause different frequency components of the audio signal to be multiplied by different real or complex values. A real mask may modify only the magnitude, while a complex mask may modify both the magnitude and the phase. When two or more neural network outputs 836 include two masks, these two masks may be different.

[0075] Typically, each of the two or more neural network outputs 836 may include one or more elements, such as an audio signal, multiple audio signals, a mask, multiple masks, an audio signal and a mask, multiple audio signals and multiple masks, etc.

[0076] Further regarding training, in some embodiments, one or more neural network layers implemented by neural network circuitry 826 may be trained to perform background noise modification. Training such neural network layers may include obtaining a noisy speech audio signal and a speech-separated version of the audio signal (i.e., retaining only speech). In some embodiments, a mask for obtaining a speech-separated audio signal when applied to the noisy speech audio signal may be determined. The training input data may be the noisy speech audio signal, and the training output data may be the mask. One or more neural network layers may thereby learn how to output a speech-separated mask for audio signal 630a such that when this mask is applied (e.g., multiplied or added to) audio signal 630a, the resulting output audio signal is speech signal 603, i.e., the speech-separated version of audio signal 630a. In some embodiments, a mask for obtaining a background noise-separated audio signal when applied to the noisy speech audio signal may be determined. The training input data may be the noisy speech audio signal, and the training output data may be the mask. The neural network layers can thus learn how to output a background noise separation mask for audio signal 630a, such that when this mask is applied (e.g., multiplied or added to) audio signal 630a, the resulting output audio signal is background noise signal 601, i.e., a background noise-modified version of audio signal 630a. In embodiments where one or more neural networks are trained to output either a speech separation signal or a noise separation signal itself, the output training data can be either the speech separation signal or the noise separation signal itself. Further description of the neural network trained to perform noise modification can be found in U.S. Patent No. 11,812,225, entitled "Method, Apparatus and System for Neural Network Hearing Aid," issued November 7, 2023.

[0077] In some embodiments, one or more neural network layers implemented by neural network circuitry 826 can be trained to perform spatial focusing. Spatial focusing may include applying a spatial focusing pattern to an audio signal. As described above, the spatial focusing pattern can assign different weights based on the direction of arrival (DOA) of the sound, where the DOA can be defined relative to the wearer of the ear-worn device. In some embodiments, the weights can be equal to 0, equal to 1, or between 0 and 1. In some embodiments, the weights can be equal to or greater than 0. In some embodiments, the weights can be greater than 0, less than 0, equal to zero, or complex; negative weights can flip the phase by 180 degrees, while complex weights can rotate the phase by an angle. Mapping weights to DOAs provides focusing because higher weights can be applied to sounds originating from a particular direction, and lower weights can be applied to sounds originating from other directions. To train such a neural network layer, a training audio signal can be formed from component audio signals originating from different DOAs. Multiple audio signals originating from multiple microphones can be generated from the training audio signal. When the neural network is trained to output a mask, a training mask can be determined such that when the training mask is applied to one of the multiple audio signals, what remains is each component audio signal multiplied by the weight corresponding to its originating DOA and then summed. One or more neural network layers can thereby learn how to output a mask based on multiple audio signals, such that when the mask is applied (e.g., multiplied or added to) speech signal 603, the resulting output (target speech signal 605) comprises each component of speech signal 603 multiplied by a weight corresponding to its derived DOA and then summed, i.e., a spatially focused version of speech signal 603. In embodiments where one or more neural networks are trained to output a spatially focused signal, the output training data can be the spatially focused signal itself.

[0078] In some embodiments, one or more neural network layers implemented by neural network circuitry 826 can be trained to perform background noise modification and spatial focusing. To train such neural network layers, training audio signals can be formed from component audio signals originating from different DOAs. Multiple audio signals originating from multiple microphones can be generated from the training audio signal. When the neural network is trained to output a mask, a training mask can be determined such that when the training mask is applied to one of the multiple audio signals, what remains is each component audio signal multiplied by the weight corresponding to its originating DOA and then summed. (As mentioned above, the training audio signal can include noisy speech audio signals and speech-separated versions of the audio signal, i.e., retaining only speech.) One or more neural network layers can thus learn how to output a mask based on the multiple audio signals such that when the mask is applied (e.g., multiplied or added to) audio signal 630a, the resulting output (target speech signal 605) comprises the speech of each component of audio signal 630a multiplied by the weight corresponding to its originating DOA and then summed, i.e., the background noise modified and spatially focused version of audio signal 630a. In an embodiment where one or more neural networks are trained to output a background noise-modified and spatially focused signal, the output training data may be the background noise-modified and spatially focused signal itself.

[0079] In some embodiments, the masking and subtraction circuit 832 in the noise reduction circuit 824 may be configured to obtain a combination of at least two of the following based on two or more neural network outputs 836: (1) the speech signal 603 of the audio signal 630a, (2) the background noise signal 601 of the audio signal 630a, (3) the target speech signal 105 of the audio signal 630a, and (4) the interfering speech signal 107 of the audio signal 630a. In some embodiments, the masking and subtraction circuit 832 may be configured to obtain one or more of these signals by applying a mask to the audio signal. More specifically, the neural network circuit 826 is configured to generate a mask (i.e., the mask is included in two or more neural network outputs 836). One or more neural network layers implemented by the neural network circuit 826 may be trained to generate a mask such that applying the mask to the audio signal can separate a portion of the audio signal (e.g., separate speech from background noise, or separate target speech from interfering speech). In some embodiments, the masking and subtraction circuit 832 can be configured to apply a mask to the audio signal 630a to generate a speech signal 603 of the audio signal 630a. In some embodiments, the masking and subtraction circuit 832 can be configured to apply a mask to the audio signal 630a to generate a background noise signal 601 of the audio signal 630a. In some embodiments, the masking and subtraction circuit 832 can be configured to apply a mask to the speech signal 603 of the audio signal 630a to generate a target speech signal 605 of the audio signal 630a. In some embodiments, the masking and subtraction circuit 832 can be configured to apply a mask to the speech signal 603 of the audio signal 630a to generate an interfering speech signal 607 of the audio signal 630a. In some embodiments, the masking and subtraction circuit 832 can be configured to apply a mask to the audio signal 630a to generate the target speech signal 605 of the audio signal 630a. In some embodiments, the masking and subtraction circuit 832 can be configured to apply a mask to the audio signal 630a, thereby generating an interfering speech signal 607 to the audio signal 630a. Which signal the masking and subtraction circuit 832 applies the mask to and which signal the result is can depend on how one or more neural network layers implemented by the neural network circuit 826 are trained to output the mask. (Because the masking and subtraction circuit 832 can be configured to apply a mask to the audio signal 630a, a dashed line connecting the audio signal 630a to the masking and subtraction circuit 832 is shown.)

[0080] In some embodiments, the masking and subtraction circuit 832 may be configured to obtain one or more of these signals by performing subtraction on some signals. (However, in some embodiments, other operations such as addition may be used instead.) In some embodiments, the masking and subtraction circuit 832 may be configured to subtract the speech signal 603 of the audio signal 630a from the audio signal 630 to generate a background noise signal 601 of the audio signal 630a. In some embodiments, the masking and subtraction circuit 832 may be configured to subtract the background noise signal 601 of the audio signal 630a to generate a speech signal 603 of the audio signal 630a. In some embodiments, the masking and subtraction circuit 832 may be configured to subtract a target speech signal 605 from the speech signal 603 to generate an interfering speech signal 607 of the speech signal 603 of the audio signal 630a. In some embodiments, the masking and subtraction circuit 832 may be configured to subtract the interfering speech signal 607 from the speech signal 603 to generate the target speech signal 605. In some embodiments, the masking and subtraction circuit 832 may be configured to subtract the target speech signal 605 and the interfering speech signal 607 from the audio signal 630a to generate the speech signal 603.

[0081] In some embodiments, neural network circuitry 826 may be configured to perform spatial focusing on audio signal 630a, such that one of two or more neural network outputs 836 is a first signal, or is configured to generate an output (e.g., a mask) of this first signal, which includes a spatially focused version of the speech in audio signal 630a plus unspatialized background noise in audio signal 630a. Therefore, in some embodiments, masking and subtraction circuitry 832 may be configured to apply a mask to audio signal 630a to generate the first signal. Neural network circuitry 826 may be configured to perform noise modification on the first signal, such that another of two or more neural network outputs 836 is a target speech signal 605 (i.e., the noise-removed first signal), or is configured to generate an output (e.g., a mask) of the target speech signal 605. Therefore, in some embodiments, masking and subtraction circuitry 832 may be configured to apply a mask to the first signal (or audio signal 630a) to generate the target speech signal 605. In some embodiments, the masking and subtraction circuit 832 can be configured to subtract the target speech signal 605 from the first signal to generate a background noise signal 601. In some embodiments, the masking and subtraction circuit 832 can be configured to subtract the first signal from the audio signal 630a to generate an interfering speech signal 607. In some embodiments, the above can be performed with the roles of the target speech signal 805 and the interfering speech signal 807 interchanged.

[0082] In some embodiments, neural network circuitry 826 may be configured to perform spatial focusing on audio signal 630a, such that one of two or more neural network outputs 836 is a first signal or an output (e.g., a mask) configured to generate this first signal, which includes a spatially focused version of speech in audio signal 630a plus a spatially focused version of background noise in audio signal 630a. Therefore, in some embodiments, masking and subtraction circuitry 832 may be configured to apply a mask to audio signal 630a to generate the first signal. Neural network circuitry 826 may be configured to perform noise modification on the first signal, such that another of two or more neural network outputs 836 is a target speech signal 605 (i.e., the noise-removed first signal) or an output (e.g., a mask) configured to generate the target speech signal 605. Therefore, in some embodiments, masking and subtraction circuitry 832 may be configured to apply a mask to the first signal (or audio signal 630a) to generate the target speech signal 605. In some embodiments, the masking and subtraction circuit 832 may be configured to subtract the target speech signal 605 from the first signal to generate a signal in the audio signal 603 that includes spatially focused background noise. In some embodiments, this signal including spatially focused background noise may be considered the background noise signal 601. In some embodiments, the masking and subtraction circuit 832 may be configured to subtract the first signal 605 from the audio signal 630a to generate a third signal that includes the remaining portion of the non-spatially focused background signal, consisting of interfering speech plus background noise. In some embodiments, this third signal may be considered the interfering speech signal 607. In some embodiments, the above may be performed with the roles of the target speech signal 805 and the interfering speech signal 807 interchanged. Thus, in some embodiments, the interfering speech signal 607 may include interfering speech but not background noise, while in other embodiments, the interfering speech signal 607 may include interfering speech plus some background noise. In some embodiments, the background noise signal 601 may include all (or all estimated) background noise, while in other embodiments, the background noise signal 601 may include some but not all background noise. From different perspectives, in some embodiments, the background noise signal 601 may not be spatially focused, and the interfering speech signal 607 may not include a portion of the background noise in the audio signal 630a. In some embodiments, the background noise signal 601 may include a first spatially focused version of the background noise in the audio signal 630a, and the interfering speech signal 608 may include interfering speech (i.e., a spatially focused version of the speech signal 603) plus a second spatially focused version of the background noise in the audio signal 630a.

[0083] After obtaining a combination of at least two of the following using at least two neural network outputs 836: (1) the speech signal 603 of the audio signal 630a, (2) the background noise signal 601 of the audio signal 630a, (3) the target speech signal 605 of the audio signal 630a, and (4) the interfering speech signal 107 of the audio signal 630a, the noise reduction circuit 824 may be configured to generate an output audio signal 840 including the target speech signal 605, the interfering speech signal 607, and the background noise signal 601. In some embodiments, in order to generate the output audio signal 840 including the target speech signal 605, the interfering speech signal 607, and the background noise signal 601, the noise reduction circuit 824 may need to obtain at least one of the target speech signal 605 and the interfering speech signal 607. In other words, in some embodiments, effective combinations may include speech signal 603 and target speech signal 605, speech signal 603 and interfering speech signal 607, background noise signal 601 and target speech signal 605, background noise signal 601 and interfering speech signal 607, and target speech signal 605 and interfering speech signal 607. Furthermore, in some embodiments, noise reduction 824 may be configured to obtain at least one of target speech signal 605 and interfering speech signal 607, and at least one of speech signal 603 and background noise signal 601; while in other embodiments, noise reduction 824 may be configured to obtain target speech signal 605 and interfering speech signal 607.

[0084] The noise reduction circuit 824 can be configured to generate a volume in the output audio signal 840a that differs from the volume in the audio signal 630a of one or more of the background noise signal 601, the interfering speech signal 607, and the target speech signal 605. Specifically, as shown in FIG7, the volume change of the background noise signal 601 may differ from the volume change of the target speech signal 605 by a first volume change difference, and the volume change of the interfering speech signal 607 may differ from the volume change of the target speech signal 605 by a second volume change difference. The volume change can be measured between the volume in the audio signal 630a and the volume in the output signal 840a. In some embodiments, the first volume change difference and the second volume change difference may be different. In some embodiments, the first volume change difference and the second volume change difference may be independently controllable.

[0085] More specifically, the target speech signal 605 is referred to as TS, the interfering speech signal 607 as IS, and the background noise signal 601 as BN. In some embodiments, the noise reduction circuit 834 can be configured to generate an equivalent of a TS + b IS +c The output audio signal 840 of BN. In some embodiments, the interference speech weight b and the background noise weight c can have values ​​between 0 and 1. The target speech weight a can typically have a value of 1, but other values ​​(e.g., values ​​greater than 1 or less than 1) can also be used. Therefore, in some embodiments, the output audio signal 840 can have reduced background noise and interference speech levels. Adding some background noise and some interference speech back into the target speech can help reduce distortion and improve the wearer's environmental awareness. In some embodiments, the mixing circuit 834 can be configured to generate the output audio signal 840 by mixing. As mentioned herein, mixing should be understood as any combination of different elements after applying weights. Therefore, the mixing circuit 834 can be configured to apply (e.g., multiply) different weights to the signal and sum the results. The mixing performed by the mixing circuit 834 can also be considered as interpolation.

[0086] It should be understood that an equivalent to a can be generated by multiplying the target speech signal 605 by a, the interfering speech signal 607 by b, and the background noise signal 601 by c, and then adding these intermediate products together. TS + b IS + c The output audio signal 840 of BN. However, other signals can be mixed together to obtain the same result. This is possible under the assumption that audio signal 630a is equivalent to speech signal 603 plus background noise signal 601, and speech signal 603 is equivalent to target speech signal 605 plus interference speech signal 607. As a non-limiting example, instead, we can set the target speech signal 605 to be multiplied by d, the interference speech signal 607 to be multiplied by e, and the audio signal 630a to be multiplied by f, and then add these intermediate products together. Referring to audio signal 630a as O (original) and speech signal as S, we get the following equation:

[0087] d TS + e IS + f O = d TS + e IS + f (S+BN)

[0088] = d TS + e IS + f (TS+IS)+f BN

[0089] = (d+f) TS + (e+f) IS + f BN

[0090] Then, expression a TS + b IS + c The weights a, b, and c in BN can have the following relationship with the weights d, e, and f: a = d + f, b = e + f, c = f.

[0091] As described above, the volume change of the background noise signal 601 can differ from the volume change of the target speech signal 605 by a first volume change difference, and the volume change of the interfering speech signal 607 can differ from the volume change of the target speech signal 605 by a second volume change difference. The volume change between the volume in the audio signal 630a and the volume in the output signal 840 can be measured. As will be further described below, the first volume change difference and the second volume change difference can be independently controlled by controlling one or more of the weights applied to the target speech signal 605, the background noise signal 601, and the interfering speech signal 607 in the output audio signal 840. In embodiments where the weight applied to the target speech signal 605 is always 1 or defaults to 1, the first volume change difference and the second volume change difference can be independently controlled by controlling the weights applied to the background noise signal 601 and the interfering speech signal 607 in the output audio signal 840. (Alternatively, either the weight applied to the background noise signal 601 or the weight applied to the interfering speech signal 607 can always be 1 or default to 1, and the first volume change difference and the second volume change difference can be independently controlled by controlling the weights applied to other signals.) In some embodiments, weights can be applied directly by applying weights to the background noise signal 601 and the interfering speech signal 607. For example, if the weight a of the target speech signal 605 is 1 and the weight c of the background noise signal 601 is 0.18, the volume change of the target speech signal 605 can be 0 dB, the volume change of the background noise signal 601 can be approximately -15 dB (the volume of the output audio signal 840 minus the volume in the audio signal 630a), and the difference in volume changes can be -15 dB (i.e., the first volume change difference can be -15 dB). In some embodiments, weights can be applied indirectly by applying weights to other audio signals associated with the background noise signal and the interfering speech signal, as described above. For example, if the weight f applied to the audio signal 630a is 0.18 and the weight d applied to the target speech signal 605 is 0.82, then the volume change of the target speech signal 605 can be 0 dB, the volume change of the background noise signal 601 can be about -15 dB, and the difference in volume change can be -15 dB (that is, the first volume change difference can be -15 dB).

[0092] In some embodiments, when the value selected for one volume change difference is not restricted to what value can be selected for another volume change difference, the first volume change difference and the second volume change difference can be considered independently controllable. For example, if the output audio signal 840 is in the form of a TS+b IS+c In some embodiments, the first volume change difference and the second volume change difference can be considered independently controllable when the value chosen for one of the weights a, b, or c is not restricted to what value can be chosen for the other weight.

[0093] Typically, the noise reduction circuit 824 can be configured to generate an output audio signal 840 using a combination of audio signals. In some embodiments, the combination of audio signals can be at least three signals 838, which are a group or subset of audio signal 630a (“O”), speech signal 603 (“S”), background noise signal 601 (“BN”), target speech signal 605 (“TS”), and interference speech signal 607 (“IS”). (Because in some embodiments, audio signal 630a may be used by the mixing circuit 834, audio signal 630a is shown in dashed lines as an input to the mixing circuit 834.) Valid combinations of these signals can include at least TS IS BN, TS S BN, IS S BN, TS IS O, TS SO, IS SO, TS O BN, and IS O BN. In some embodiments, the mixing circuit 834 can be configured to generate the output audio signal 840 by mixing combinations of audio signals, such as one of the combinations of at least three signals 838 described above.

[0094] As described above, the noise reduction circuit 824 can be configured to generate an output audio signal 840 comprising a target speech signal 605, an interfering speech signal 607, and a background noise signal 601 using at least two or more neural network outputs 836. It can be understood from the above that the two or more neural network outputs 836 may include one or more of the target speech signal 605, the interfering speech signal 607, and the background noise signal 601, or may include outputs from which the target speech signal 605, the interfering speech signal 607, and the background noise signal 601 can be generated. Such outputs may be masks or signals from which the target speech signal 603, the interfering speech signal 607, and / or the background noise signal 601 can be derived (e.g., by subtraction as described above).

[0095] From the above, it should be understood that the noise reduction circuit 826 can be configured to use a first audio signal 630a when generating the output audio signal 630, which can be a beamforming audio signal. For example, the masking and subtraction circuit 832 can be configured to apply a mask to the first audio signal 630a and / or the mixing circuit 834 can be configured to use the first audio signal 630a in the mixing. Regarding the mask, it should be understood that in some embodiments, the masking and subtraction circuit 832 can be configured to apply a mask to the beamforming audio signal, and in some embodiments, the masking and subtraction circuit 832 can be configured to apply a mask to a non-beamforming signal.

[0096] In some embodiments, two or more neural network outputs 836 may include one or more of the speech signal 603, background noise signal 601, target speech signal 105, and interfering speech signal 607. In other words, neural network circuit 826 may be configured to directly output one or more of these signals. In embodiments where neural network circuit 826 directly outputs signals instead of a mask, mask application and subtraction circuit 832 may alternatively include only subtraction circuitry. In some embodiments, mask application yields all signals that need to be generated. In such embodiments, mask application and subtraction circuit 832 may alternatively include only mask application circuitry. In some embodiments, neural network circuit 826 may be configured to directly output all signals that need to be generated. In such embodiments, mask application and subtraction circuit 832 may be absent.

[0097] In some embodiments, the mixing circuit 834 may be configured to mix two or more masks, and the mask application and subtraction circuit 832 may be configured to apply the mixed mask to an audio signal. This operation is equivalent to applying two or more masks independently to an audio signal and then mixing the results together. The mixed mask may include applying weights to different masks and combining (e.g., adding) the weighted masks together. In such embodiments, the mixing circuit 834 may be incorporated into the mask application and subtraction circuit 832. Therefore, in some embodiments, the mixing circuit 834 may be configured to generate an output audio signal 840 by mixing multiple (e.g., at least two) masks. The speech signal 603 is referred to as "S", the background noise signal 601 as "BN", the target speech signal 605 as "TS", and the interfering speech signal 607 as "IS". Multiple masks may include masks configured to generate: TS, IS, BN; TS, S, BN; IS, S, BN; TS and IS; TS and S; IS and S; TS and BN; IS and BN. In some embodiments, when only two masks are mixed, the result may be applied to and subsequently mixed with the audio signal 603a.

[0098] It should be understood that some or all of the speech signal 603, background noise signal 601, target speech signal 605, and interfering speech signal 607 can be generated using one or more neural network layers, and one or more neural network layers can be trained to output estimates. Therefore, in some embodiments, speech signal 603 may not necessarily include all speech in audio signal 630a, background noise signal 601 may not necessarily include all background noise in audio signal 630a, target speech signal 605 may not necessarily include all target speech in speech signal 603, and interfering speech signal 607 may not necessarily include all interfering speech in speech signal 603.

[0099] Figure 9 illustrates a neural network circuit 926 and a mask application and subtraction circuit 932 according to certain embodiments described herein. The neural network circuit 926 can be an example of any neural network circuit described herein (e.g., neural network circuits 526 and / or 826). The mask application and subtraction circuit 932 can be an example of any processing circuit described herein (e.g., mask application and subtraction circuit 832).

[0100] Neural network circuitry 926 includes circuitry configured to implement multiple neural network layers, shown in FIG. 9 as a first subset (i.e., one or more layers) 950a and a second subset (i.e., one or more layers) 950b of neural network layers. In some embodiments, such circuitry may include multiply-accumulate circuitry configured to perform a multiply-accumulate operation on an input activation vector and a neural network weight matrix as part of processing one or more neural network layers. Further description of the neural network circuitry can be found in U.S. Patent No. 11,886,974, entitled “Neural Network Chip for Ear-Worn Device,” issued January 30, 2024, the entire contents of which are incorporated herein by reference. Masking and subtraction circuitry 932 includes multipliers 952a, 956b, subtractors 954a, and 954b.

[0101] In some embodiments, neural network circuitry 926 may be configured to use a first subset 950a of neural network layers to generate one of two or more neural network outputs 836, and a second subset 950b of neural network layers to generate another of the two or more neural network outputs 836. Various options for neural network outputs 836 have been provided above, and it should be understood that different embodiments may utilize the first subset 950a and the second subset 950b of neural network layers to generate different combinations of these options. For example, neural network circuitry 926 may be configured to use the first subset 950a of neural network layers to generate speech signal 603, generate an output configured to generate speech signal 603 (e.g., a mask), background noise signal 601, and / or generate an output configured to generate background noise signal 601 (e.g., a mask). Continuing this example, neural network circuit 926 can be configured to use a second subset 950b of neural network layers to generate target speech signal 605, generate an output (e.g., a mask) configured to generate target speech signal 605, generate interfering speech signal 607, and / or generate an output (e.g., a mask) configured to generate interfering speech signal 607. In the example of Figure 9, the first of neural network outputs 836 is a mask 956a configured to generate speech signal 603, and the second of neural network outputs 836 is a mask 956b configured to generate target speech signal 605.

[0102] A first subset 950a of the neural network layer implemented by neural network circuitry 926 can be configured to receive an audio signal 630a. The audio signal 630a may originate from one or more microphones in an ear-worn device (e.g., microphones 102f and 102b, microphone 302, and / or microphone 502). For example, the audio signal 630a may be a beamformed version of a processed signal (e.g., with masking applied and subtraction circuitry 832) from two different microphones. As another example, the audio signal 630a may be a processed signal (e.g., with masking applied and subtraction circuitry 832) from a single microphone. The first subset 950a of the neural network layer implemented by neural network circuitry 926 can be configured to generate an output configured to generate a speech signal 603 based on the audio signal 630a. In the example of Figure 9, the output configured to generate the speech signal 603 is a mask 956a. A multiplier 952a can be configured to multiply the audio signal 630a by the mask 956a to produce the speech signal 603.

[0103] In some embodiments, a first subset 950a of the neural network layers implemented by neural network circuitry 926 can be trained to perform background noise modification. Further description of the neural network training can be found above. Based on this training, the first subset 950a of the neural network layers can learn how to output a speech separation mask 956a for audio signal 630a, such that when multiplier 952a multiplies audio signal 630a by mask 956a, the resulting output audio signal is speech signal 603, i.e., a speech-separated version of audio signal 630a. Further description of the neural network trained to perform noise modification can be found in U.S. Patent No. 11,812,225, entitled "Method, Apparatus and System for Neural Network Hearing Aid," issued November 7, 2023.

[0104] A second subset 950b of the neural network layer implemented by neural network circuitry 926 can be configured to receive a plurality of audio signals 630, such that at least two of the plurality of audio signals 630 each originate from different microphones in an ear-worn device (e.g., microphones 102f and 102b, microphone 302 and / or microphone 502) and / or at least one of the plurality of audio signals is a beamformed version of an audio signal originating from a microphone. In the example of FIG9, the plurality of audio signals 630 includes a speech signal 603, an audio signal 630a, and one or more other audio signals 630b. In some embodiments, some of the plurality of audio signals 630 may have a directional pattern formed by beamforming the audio signals from different microphones. In some embodiments, some (e.g., at least two) or each of the plurality of audio signals 630 may have different beamforming directional patterns. In some embodiments, at least one (or all) of the audio signals 630a and / or audio signals 630b may have a directional pattern formed by beamforming the audio signals from different microphones. In some embodiments, at least one (or all) of audio signals 630a and 630b may each have a different beamforming directivity pattern. Examples of beamforming directivity patterns include dipole, supercardioid, supercardioid, and cardioid. In some embodiments, at least one of the audio signals 630 may have a forward beamforming directivity pattern, and at least one of the audio signals 630 may have a backward beamforming directivity pattern. In some embodiments, audio signal 630a or one of the audio signals may be a dipole. In some embodiments, one of the audio signals 630a or 630b may be a forward supercardioid. In some embodiments, one of the audio signals 630a or 630b may be a forward supercardioid. In some embodiments, the plurality of audio signals 630 may include two signals. In some embodiments, the plurality of audio signals 630 may include three signals. In some embodiments, the plurality of audio signals 630 may include four signals. In some embodiments, the plurality of audio signals 630 may include more than four signals. In some embodiments, a second subset 950b of the neural network layer may receive only the speech signal 603 and the audio signal 630a as input (i.e., not the audio signal 630b). In some embodiments, certain audio signals 630a and / or audio signals 630b may be non-beamforming signals.

[0105] The following are non-limiting examples of audio signal groups that may be included or comprised of a plurality of audio signals 630. In some embodiments, the plurality of audio signals 630 may include signals having a forward supercardioid directional pattern and signals having a backward supercardioid directional pattern. In some embodiments, the plurality of audio signals 630 may include signals having a forward cardioid directional pattern and signals having a backward cardioid directional pattern. In some embodiments, the plurality of audio signals 630 may include signals having a forward supercardioid directional pattern and signals having a backward supercardioid directional pattern. In some embodiments, the plurality of audio signals 630 may include signals having a forward supercardioid directional pattern, signals having a backward supercardioid directional pattern, and voice signal 603. In some embodiments, the plurality of audio signals 630 may include signals having a forward supercardioid directional pattern, signals having a backward supercardioid directional pattern, and voice signal 603. In some embodiments, the plurality of audio signals 630 may include signals having a forward supercardioid directional pattern, signals having a backward supercardioid directional pattern, and voice signal 603. In some embodiments, the plurality of audio signals 630 may include signals having a forward supercardioid directional pattern, signals having a backward supercardioid directional pattern, signals having a dipole directional pattern, and a speech signal 603. In some embodiments, the plurality of audio signals 630 may include signals having a forward cardioid directional pattern, signals having a backward cardioid directional pattern, signals having a dipole directional pattern, and a speech signal 603. In some embodiments, the plurality of audio signals 630 may include signals having a forward supercardioid directional pattern, signals having a backward supercardioid directional pattern, signals having a dipole directional pattern, and a speech signal 603. In some embodiments, the plurality of audio signals 630 may include signals having a forward supercardioid directional pattern, signals having a backward supercardioid directional pattern, and signals having a dipole directional pattern. In some embodiments, the plurality of audio signals 630 may include signals having a forward cardioid directional pattern, signals having a backward cardioid directional pattern, and signals having a dipole directional pattern. In some embodiments, the plurality of audio signals 630 may include signals having a forward supercardioid directional pattern, signals having a backward supercardioid directional pattern, and signals having a dipole directional pattern.

[0106] A second subset 950b of the neural network layer implemented by neural network circuit 926 can be configured to generate an output configured to generate a target speech signal 605 based on multiple audio signals 630. In the example of Figure 9, the output configured to generate the target speech signal 605 is a mask 956b. Multiplier 956b can be configured to multiply the speech signal 603 by the mask 956b to produce the target speech signal 605.

[0107] In some embodiments, a second subset 950b of the neural network layer implemented by neural network circuitry 926 can be trained to perform spatial focusing. Further description of the training can be found above. Based on this training, the second subset 950b of the neural network layer can learn how to output a mask 956b based on multiple audio signals 630, such that when multiplier 956b multiplies speech signal 603 by mask 956b, the resulting output (target speech signal 605) comprises each component audio signal multiplied by a weight corresponding to its source DOA and then summed. Applying higher weights to the direction of the target speaker and lower weights to other directions helps to focus the sound from the target speaker and attenuate the sound from other interfering speakers. Therefore, the output (target speech signal 605) can be generally considered as a signal containing the target speech.

[0108] The masking application and subtraction circuit 932 can also be configured to generate a background noise signal 601 and an interfering speech signal 607. In the example of FIG9, subtractor 954a can be configured to subtract the speech signal 603 from the audio signal 630a to generate the background noise signal 601. Subtractor 954b can be configured to subtract the target speech signal 605 from the speech signal 603 to generate the interfering speech signal 607.

[0109] As described above, mask 956a can be configured to generate speech signal 603, and mask 956b can be configured to generate target speech signal 605. Masks 956a and 956b can be different. In some embodiments, mask 956b can be applied to one of audio signal 630a or audio signal 630b, instead of applying mask 956b to speech signal 603.

[0110] Typically, the audio signals and training data used to train one or more neural network layers (e.g., the second subset 950b) to perform spatial focusing may need to include spatial information (in other words, enough information for the model to infer the location of the sound source). At a minimum, the audio signals and training data may need to include multiple (i.e., at least two) audio signals, where at least two of the multiple audio signals each originate from different microphones in two or more microphones, and / or at least one of the multiple audio signals is a beamformed version of the audio signal originating from two or more microphones. Therefore, the training data used for these one or more neural network layers may typically need to include audio signals originating from different microphones. Two methods for generating multi-microphone localization training data may include 1) collecting audio from sounds originating from different DOAs using multiple microphones, and 2) synthetically creating multiple microphone signals using audio simulation as if the sound source were localized to a specific DOA. The synthetic generation of the training signal may include adding directivity to different sound sources (speech signals and noise) in a simulation, and then combining these sound sources to create a new signal with spatial audio. The neural network can be trained based on either or both of the synthetic data and the acquired data.

[0111] While the noise reduction circuit 924 includes corresponding multipliers 952a and 956b for multiplying the corresponding masks 956a and 956b by the audio signal, in some embodiments, other operations (e.g., addition) may be used to combine the masks with the signal. Additionally, one or more neural network layers implemented by neural network circuitry (e.g., neural network circuitry 926) may be configured to generate an output (e.g., a mask, such as mask 956a or 956b, etc.) configured to generate a specific signal; alternatively, in some embodiments, one or more neural network layers may be configured to generate the specific signal itself. As an example, in some embodiments, a first subset 950a of the neural network layers implemented by neural network circuitry 926 may be configured to generate the speech signal 603, and a second subset 950b of the neural network layers implemented by neural network circuitry 926 may be configured to generate the target speech signal 605.

[0112] As described above, one or more neural network layers implemented by neural network circuit 926 (and specifically, a first subset 950a of neural network layers) can be configured to generate speech signal 603, or to generate the output of speech signal 603 (e.g., mask 956a). This can include cases where speech signal 603 is obtained by applying a mask to an audio signal (e.g., audio signal 630a), and cases where a background noise signal 601 is obtained by applying a mask to an audio signal (e.g., audio signal 630a) and then a subtractor is used to generate speech signal 603 from the background noise signal 601 (e.g., by subtracting the background noise signal 601 from the audio signal 630a). Additionally, this can include cases where one or more neural network layers directly generate speech signal 603, and cases where one or more neural network layers directly generate background noise signal 601 and then a subtractor is used to generate speech signal 603 from the background noise signal 601 (e.g., by subtracting the background noise signal 601 from the audio signal 630a). Similarly, as described above, one or more neural network layers implemented by neural network circuit 926 (and specifically, a second subset 950b of neural network layers) can be configured to generate the target speech signal 605, or to generate an output (e.g., a mask 956b) configured to generate the target speech signal 605. This can include cases where the target speech signal 605 is obtained when the mask is applied to an audio signal (e.g., speech signal 603), and cases where an interfering speech signal 607 is obtained when the mask is applied to an audio signal (e.g., speech signal 603) and a subtractor is subsequently used to generate the target speech signal 605 from the interfering speech signal 607 (e.g., by subtracting the interfering speech signal 607 from the speech signal 603). Additionally, this may include the following cases: one or more neural network layers directly generate the target speech signal 605, and the following cases: one or more neural network layers directly generate the interference speech signal 607, and then a subtractor is used to generate the target speech signal 605 from the interference speech signal 607 (e.g., by subtracting the interference speech signal 607 from the speech signal 603).

[0113] As should be understood from the description in Figure 9 above, the first subset 950a of the neural network layer can be configured to generate a mask 956a, which can be configured to generate a speech signal 603, and the speech signal 603 can be the input to the second subset 950b of the neural network layer, which can be configured to generate the mask 956b. In other words, the mask 956a can also be generated before the mask 956b. Therefore, processing the first subset 950a of the neural network layer can occur before processing the second subset 950b of the neural network layer. In some embodiments, the same circuit can be used to process both the first subset 950a and the second subset 950b of the neural network layer. For example, processing the neural network layer may include using a multiplier-accumulator circuit (MAC) configured to multiply the input activation by the neural network weights. In some embodiments, the same MAC can be used to process both the first subset 950a and the second subset 950b of the neural network layer, but the neural network weights used by the MAC can be different. That is, the first subset 950a of the neural network layer can use different weights than the second subset 950b of the neural network layer. In some embodiments, different circuits can be used to process a first subset 950a of the neural network layer and a second subset 950b of the neural network layer.

[0114] Figure 10 illustrates a neural network circuit 1026 and a mask application and subtraction circuit 1032 according to certain embodiments described herein. The neural network circuit 1026 may be an example of any neural network circuit described herein (e.g., neural network circuits 526 and / or 826). The mask application and subtraction circuit 1032 may be an example of any processing circuit described herein (e.g., mask application and subtraction circuit 832).

[0115] The neural network circuit 1026 includes circuitry configured to implement one or more neural network layers. One or more neural network layers 1050 implemented by the neural network circuit 1026 can be configured to receive multiple audio signals 630. Typically, one or more neural network layers 1050 can be configured to generate a speech signal 603 or an output configured to generate a speech signal 603 based on the multiple audio signals 630, and to generate a target speech signal 605 or an output configured to generate a target speech signal 605. In a specific example of FIG. 10, based on the multiple audio signals 630, one or more neural network layers 1050 can be configured to generate masks 956a and 956b. As described above, mask 956a can be configured to generate speech signal 603 (by multiplying mask 956a by audio signal 630a using multiplier 956a). Mask 956b can be configured to generate target speech signal 605 (by multiplying mask 956b by audio signal 630a using multiplier 956b). A background noise signal 601 can be generated from the speech signal 603 (by subtracting the speech signal 603 from the audio signal 630a using subtractor 954a), and an interfering speech signal 607 can be generated from the target speech signal 605 (by subtracting the target speech signal 605 from the speech signal 603 using subtractor 954b). In some embodiments, one or more neural network layers 1050 can be configured to generate a mask that is configured to generate the background noise signal 601 and / or to generate the interfering speech signal 607. Furthermore, the speech signal 603 can be generated from the background noise signal 601, and / or the target speech signal 605 can be generated from the interfering speech signal 607. In some embodiments, masks 956a and 1056b can be generated simultaneously. In some embodiments, one or more neural network layers 1050 can be configured to directly generate some combination of the speech signal 603, the target speech signal 605, the background noise signal 601, and the interfering speech signal 607. In some embodiments, the signals generated by one or more neural network layers 1050 can be generated simultaneously. Therefore, the speech signal 603 may not necessarily be the input to the neural network layer configured to generate the mask 1056b. One or more neural network layers 1050 implemented by the neural network circuit 1026 can be considered to be trained to perform background noise modification (for generating the speech signal 603 from the audio signal 630a) and to perform background noise modification and spatial focusing (for generating the target speech signal 605 from the audio signal 630a).

[0116] The above descriptions of Figures 9 and 10 have illustrated how the speech signal 603, background noise signal 601, target speech signal 605, and interference speech signal 607 can be generated. As further described above, some combinations of these signals with the audio signal 630a can be, for example, mixed together by the mixing circuit 834.

[0117] Returning to Figure 8, the above description of Figure 8 has described the neural network circuit 826 outputting two or more neural network outputs 836. Typically, as described above, the ear-worn device may include two or more microphones, and the noise reduction circuit 824 may include the neural network circuit 826. The neural network circuit 826 may be configured to receive a plurality of audio signals 630 such that (1) at least two of the plurality of audio signals 630 each originate from different microphones of the two or more microphones of the ear-worn device and / or (2) at least one of the plurality of audio signals 630 is a beamforming audio signal originating from the two or more microphones. The neural network circuit 826 may be configured to implement one or more neural network layers trained to perform noise modification and / or spatial focusing, such that the neural network circuit 826 generates one or more neural network outputs 836 based on the plurality of audio signals 630. For example, one or more neural network outputs 836 may be one or more audio signals, configured to generate one or more outputs (e.g., masks and / or spectrograms) of output audio signals, or combinations thereof. The noise reduction circuit 824 can be configured to output an output audio signal 840 as a noise-modified and / or spatially focused version of the audio signal 630a (i.e., one of the multiple audio signals 630) based on one or more neural network outputs 836. As a specific example, one or more neural network outputs 836 may include an output audio signal that is a background noise-modified and spatially focused version of the audio signal 630a, or one or more neural network outputs 836 may be configured for the noise reduction circuit 826 to generate an output audio signal that is a background noise-modified and spatially focused version of the audio signal 630a.

[0118] In some embodiments, the noise reduction circuit 824 may be configured to perform background noise modification. In other words, one or more neural network layers may be trained to perform background noise modification, and the output audio signal 840 may include a background noise-modified version of audio signal 630a (e.g., speech signal 603 and / or background noise signal 601). In some embodiments, the noise reduction circuit 824 may be configured to perform spatial focusing. In other words, one or more neural network layers may be trained to perform spatial focusing, and the output audio signal 840 may include a spatially focused version of audio signal 630a (e.g., a spatially focused version of both speech and background noise in audio signal 630a, or a spatially focused version of speech only, i.e., the target speech signal 605). In some embodiments, the noise reduction circuit 824 may be configured to perform both background noise modification and spatial focusing. In other words, one or more neural network layers may be trained to perform both background noise modification and spatial focusing, and the output audio signal 840 may include a background noise-modified and spatially focused version of audio signal 630a (e.g., the target speech signal 605).

[0119] When the output audio signal 840 includes a spatially focused version 630a of the audio signal 630a or a background noise modified and spatially focused version of the audio signal 630a, the spatially focused portion of the output audio signal 840 may have a specific spatial focus mode. All descriptions of spatial focus modes herein, as well as controls over spatial focus modes (e.g., in the context of the target speech signal 605), may also be applied to the output audio signal 840 in such cases.

[0120] In embodiments where the neural network output 836 is a mask, the mask application and subtraction circuit 832 can be configured to apply the mask to (e.g., by multiplying or adding to) an audio signal (e.g., audio signal 630a). (Therefore, a dashed line connecting the audio signal 630a to the mask application and subtraction circuit 832 is shown). In some embodiments, when the mask is applied to the audio signal 630a, a speech signal 603 of the audio signal 630a is obtained. In some embodiments, when the mask is applied to the audio signal 630a, a background noise signal 601 of the audio signal 630a is obtained. In some embodiments, when the mask is applied to the audio signal 630a, a spatially focused version of the audio signal 630a (e.g., a target speech signal 605 or interfering speech signal 607) is obtained. In some embodiments, the mask application and subtraction circuit 832 can be configured to generate a second audio signal based on a first audio signal (which may be, for example, the neural network output 836 or an audio signal generated from the neural network output 836 when the neural network output 836 is a mask). In some embodiments, the masking and subtraction circuit 832 may be configured to subtract the speech signal 603 of the audio signal 630a from the audio signal 630a to generate a background noise signal 601 of the audio signal 630a. In some embodiments, the masking and subtraction circuit 832 may be configured to subtract the background noise signal 601 of the audio signal 630a to generate the speech signal 603 of the audio signal 630a. In some embodiments, the masking and subtraction circuit 832 may be configured to subtract a background noise-modified and spatially focused version (e.g., the target speech signal 605) of the audio signal 630a to generate an audio signal that includes background noise and interfering speech (e.g., an audio signal that includes the background noise signal 601 in addition to the interfering speech signal 607). In some embodiments, the masking and subtraction circuit 832 can be configured to subtract the target speech signal 605 from the speech signal 603 to generate an interfering speech signal 607, or it can be configured to subtract the interfering speech signal 607 from the speech signal 603 to generate the target speech signal 605.

[0121] As described above, in some embodiments, the output audio signal 840 may include a background noise modified version, a spatially focused version, or a background noise modified and spatially focused version of the audio signal 630a. In some embodiments, the output audio signal 840 may be a background noise modified version, a spatially focused version, or a background noise modified and spatially focused version of the audio signal 630a. In some embodiments, the output audio signal 840 may include a background noise modified version, a spatially focused version, or a background noise modified and spatially focused version of the audio signal 630a mixed with one or more other audio signals. In some embodiments, the mixing circuit 834 may be configured to mix two or more audio signals such that the output audio signal 840 is equivalent to a background noise modified and / or spatially focused version of the audio signal 630a (which may, as a non-limiting example, be the speech signal 603, the background noise signal 601, the target speech signal 605, and / or the interfering speech signal 607) mixed with another audio signal (which may, as a non-limiting example, be the speech signal 603, the background noise signal 601, the target speech signal 605, the interfering speech signal 607, and / or the audio signal 630a). In some embodiments, the mixing circuit 834 may be configured to mix the speech signal 603 of the audio signal 630a with the background noise signal 601 of the audio signal 630a. In some embodiments, the mixing circuit 834 may be configured to mix the speech signal 603 of the audio signal 630a with the audio signal 630a. In some embodiments, the mixing circuit 834 may be configured to mix the background noise signal 601 of the audio signal 630a with the audio signal 630a. In some embodiments, the mixing circuit 834 may be configured to mix the speech signal 603 of the audio signal 630a, the background noise signal 601 of the audio signal 630a, and the audio signal 630a. In some embodiments, the mixing circuit 834 may be configured to mix a background noise-modified and spatially focused version of the audio signal 630a (e.g., the target speech signal 605) with the audio signal 630a. In some embodiments, the mixing circuit 834 may be configured to mix a background noise-modified and spatially focused version of the audio signal 630a (e.g., the target speech signal 605) with the background noise signal 601. In some embodiments, the mixing circuit 834 may be configured to mix a first background noise-modified and spatially focused version of the audio signal 630a (e.g., the target speech signal 605) with a second background noise-modified and spatially focused version of the audio signal 630a (e.g., the interfering speech signal 607). In some embodiments, the mixing circuit 834 may be configured to mix the first background noise-modified and spatially focused version of the audio signal 630a (e.g., the target speech signal 605), the second background noise-modified and spatially focused version of the audio signal 630a (e.g., the interfering speech signal 607), and the audio signal 630a.In some embodiments, the mixing circuit 834 may be configured to mix a first background noise-modified and spatially focused version of the audio signal 630a (e.g., the target speech signal 605), a second background noise-modified and spatially focused version of the audio signal 630a (e.g., the interfering speech signal 607), and a background noise signal 601. In some embodiments, the mixing circuit 834 may be configured to mix the first background noise-modified and spatially focused version of the audio signal 630a (e.g., the target speech signal 605) with an audio signal that includes background noise and interfering speech (e.g., an audio signal that includes background noise signal 601 in addition to interfering speech signal 607). The mixing performed by the mixing circuit 834 may also be considered as interpolation.

[0122] Figure 11 illustrates a neural network circuit 1126 and a mask application and subtraction circuit 1132 according to certain embodiments described herein. The neural network circuit 1126 may be an example of any neural network circuit described herein (e.g., neural network circuits 526 and / or 826). The mask application and subtraction circuit 1132 may be an example of any processing circuit described herein (e.g., mask application and subtraction circuit 832).

[0123] The neural network circuit 1126 includes circuitry configured to implement one or more neural network layers. One or more neural network layers 1150 implemented by the neural network circuit 1126 can be configured to receive multiple audio signals 630. Typically, one or more neural network layers 1150 can be configured to generate a target speech signal 605 based on the multiple audio signals 630, or to generate an output configured to generate a target speech signal 605. In a specific example of FIG11, based on the multiple audio signals 630, one or more neural network layers 1150 can be configured to generate a mask 1056b. As described above, the mask 1056b can be configured to generate the target speech signal 605 (by multiplying the mask 1056b by the audio signal 630a using a multiplier 952b). Based on the target speech signal 605, an audio signal 609 comprising a background noise signal 601 plus an interfering speech signal 607 can be generated (by subtracting the target speech signal 605 from the audio signal 630a using a subtractor 954b). One or more neural network layers 1150 implemented by neural network circuit 1126 can be considered to be trained to perform background noise modification and spatial focusing (for generating target speech signal 603 from audio signal 630a).

[0124] The above description in Figure 11 has illustrated how the target speech signal 605 and audio signal 609 (comprising background noise signal 601 plus interfering speech signal 607) can be generated. As further described above, these signals, in a combination with audio signal 630a, can be (for example) mixed together by mixing circuitry 834. More specifically, the target speech signal 605 is referred to as TS, and the audio signal 609 as (IS+BN). In some embodiments, noise reduction circuitry 834 can be configured to generate an equivalent of a TS+b The output audio signal 840 is (IS+BN). In some embodiments, the weight b used for adding background noise to the interfering speech can have a value between 0 and 1. The target speech weight a can typically have a value of 1, but other values ​​(e.g., values ​​greater than 1 or less than 1) can also be used. Therefore, in some embodiments, the output audio signal 840 can have a reduced level of background noise and interfering speech. Adding some background noise and some interfering speech back to the target speech can help reduce distortion and improve the wearer's environmental awareness. In some embodiments, the mixing circuit 834 can be configured to generate the output audio signal 840 by mixing. Therefore, the mixing circuit 834 can be configured to apply (e.g., multiply) different weights to the signal and add the results. The mixing performed by the mixing circuit 834 can also be considered as interpolation.

[0125] It should be understood that an equivalent to a can be generated by multiplying the target speech signal 605 by a and the audio signal 609 by b. TS+b The output audio signal 840 is (IS+BN). However, other signals can be mixed together to obtain the same result. This is possible under the assumption that audio signal 630a is equivalent to speech signal 603 plus background noise signal 601, and speech signal 603 is equivalent to target speech signal 605 plus interfering speech signal 607. As a non-limiting example, instead, we can set the target speech signal 605 to be multiplied by d, audio signal 630a to be multiplied by e, and then add these intermediate products together. Referring to audio signal 630a as O (original) and speech signal as S, we get the following equation:

[0126] d TS + e O = d TS + e (S+BN)

[0127] = d TS + e (TS+IS+BN)

[0128] = (d+e) TS + e IS + e BN

[0129] Then, expression a TS + b In (IS+BN), the weights a and b can have the following relationship with the weights d and e: a = d + e, b = e.

[0130] It should be understood that in such embodiments, the first volume change difference and the second volume change difference may not be independently controllable because the background noise signal 601 and the interfering speech signal 607 can be combined together in the audio signal 609. However, the first volume change difference and the second volume change difference can still be controllable. In other words, the first volume change difference and the second volume change difference may need to have the same amount, but this amount can be controllable. In other words, the mixing circuit 834 can be configured to mix two or more audio signals such that the output audio signal includes a background noise-modified and spatially focused version of the audio signal 630a mixed with the second audio signal (i.e., the target speech signal 605). The second audio signal may include the background noise signal 601 and the interfering speech signal 607. For example, the second audio signal may be audio signal 609 or audio signal 630a. A noise reduction circuit (e.g., noise reduction circuit 824) can be configured to generate an output audio signal 840 such that, in the output audio signal, the volume change of the background noise signal 601 differs from the volume change of the target speech signal 605 by a volume change difference, the volume change of the interfering speech signal differs from the volume change of the target speech signal by the same volume change difference, and the volume change difference can be controllable. WDRC circuit

[0131] As described above, the noise reduction circuit can be configured to generate an output audio signal comprising a target speech signal 605, an interfering speech signal 607, and a background noise signal 601, such that in the output audio signal, the volume change of the background noise signal differs from the volume change of the target speech signal by a first volume change difference, and the volume change of the interfering speech signal differs from the volume change of the target speech signal by a second volume change difference, and the first and second volume change differences are independently controllable. The above description describes the generation of this output audio signal using a mixing circuit (e.g., mixing circuit 834). In some embodiments, the noise reduction circuit can be configured to generate the output audio signal using a wide dynamic range compression (WDRC) circuit. In short, some ear-worn devices can apply a non-linear, frequency-dependent gain to the incoming sound to “fit” the output sound to the wearer’s hearing profile. For example, if a wearer has significant hearing loss at higher frequencies and much less at lower frequencies, then for the same input volume, an ear-worn device can apply more gain to the lower-frequency sound compared to the higher-frequency sound to effectively equalize the audibility or perceived loudness of the different sounds across frequencies. Furthermore, because those with hearing loss typically have a narrow range of volumes they can comfortably hear (a reduced "dynamic range"), some hearing aids apply more gain to quiet sounds and less gain to loud sounds, effectively "compressing" the original signal into the wearer's dynamic range. These techniques are sometimes referred to as wide dynamic range compression (WDRC).

[0132] Figure 12 illustrates a noise reduction circuit 1224 in an ear-worn device according to certain embodiments described herein. The ear-worn device can be any ear-worn device described herein (e.g., hearing aid 100, glasses 300, ear-worn device 400, and / or ear-worn device 500). The noise reduction circuit 1224 can be any noise reduction circuit described herein (e.g., noise reduction circuit 524). The noise reduction circuit 1224 includes a neural network circuit 826, a masking and subtraction circuit 832, and a WDRC circuit 1258. The WDRC circuit 1258 can be configured to receive a plurality of audio signals 838 as input, which may include, for example, two or more of an audio signal 630a, a speech signal 603, a background noise signal 601, a target speech signal 605, and interfering speech signals 607.

[0133] Figure 13 illustrates the WDRC circuit 1258 in more detail according to certain embodiments described herein. The WDRC circuit 1258 includes an amplification pipeline 1360, one pipeline for each of a plurality of audio signals 858 received by the WDRC circuit 1258, and each amplification pipeline includes a level estimation circuit group 1362 and an amplification circuit group 1364. The plurality of audio signals 858 may include, for example, two or more of an audio signal 630a, a speech signal 603, a background noise signal 601, a target speech signal 605, and interfering speech signals 607. Each level estimation circuit group 1362 is shown as comprising multiple blocks, each block for a different frequency channel. Each amplification circuit group 1364 is shown as comprising multiple blocks, each block for a different frequency channel. (For simplicity, circuitry for converting the input signal to the frequency domain, splitting the signal into frequency channels, combining the frequency channels together, and converting to the time domain is not shown).

[0134] Typically, the WDRC circuit 1258 includes multiple hearing loss amplification (hereinafter referred to simply as "amplification") pipelines 1360. Each amplification pipeline 1360 may correspond to one of the sub-signals and includes a block of amplification circuitry 1364. Amplification circuitry 1364 may be configured to perform hearing loss amplification, i.e., to be configured to provide additional amplification to compensate for audibility loss due to hearing loss. Specifically, each corresponding block of amplification circuitry 1364 may be configured to apply amplification to the corresponding input sub-signal to produce an amplified sub-signal. The amplification applied by each block of amplification circuitry 1364 may be different. Thus, amplification circuitry 13641 in amplification pipeline 13601 may be configured to apply a first amplification to sub-signal 1, amplification circuitry 13642 in amplification pipeline 13602 may be configured to apply a second amplification to sub-signal 2, and so on, and the first amplification and the second amplification may be different. Typically, amplification can be any method used to amplify a signal to compensate for audibility loss due to hearing loss and may include, for example, one or more rules, formulas, or curves.

[0135] Each corresponding level estimation circuit group 1362 can be configured to determine the level of a corresponding sub-signal, and each corresponding amplification circuit group 1364 can be configured to amplify (e.g., apply a set of speech adaptation curves) the speech sub-signal at least in part based on the level of the speech sub-signal determined by the level estimation circuit 1362. More specifically, for the corresponding level estimation circuit group 1362 for a particular sub-signal, each block of the level estimation circuit 1362 can be configured to determine the level (e.g., power or amplitude) of the input sub-signal within a specific frequency channel and within a certain time window or on a moving average of the time window. For the corresponding amplification circuit group 1364 for a particular sub-signal, each block of the amplification circuit 1364 can be configured to apply amplification to the input sub-signal within a specific frequency channel, such that the result is an amplified sub-signal within that frequency channel, and the sum of the amplified sub-signals in different channels is the amplified sub-signal. The amplification applied by the amplification circuit group 1364 of each amplification pipeline 1360 can be different. The amplification of a specific frequency channel of the sub-signal applied by amplifier circuit 1364 may depend at least in part on the input level of that specific frequency channel of the sub-signal, as determined by level estimation circuit 1362. Input-level-dependent and frequency-dependent amplification may include applying a set of adaptation curves to the sub-signal, each adaptation curve being an output level curve relative to the input level curve for a given frequency channel (or equivalently, each adaptation curve being an output level curve relative to the frequency channel curve for a given input level). Different amplifications may include different sets of adaptation curves. Applying a set of adaptation curves to the sub-signal may include: determining the input level of the sub-signal in each frequency channel; determining the output level corresponding to the input level and frequency channel based on one of the adaptation curves; amplifying that channel of the sub-signal to that output level; and combining the results from the different frequency channels. Combiner 1374 (e.g., a summer) may be configured to combine the amplified sub-signals back into a single output signal, namely, output audio signal 840.

[0136] Therefore, the level estimation circuit 13621 in amplification pipeline 13601 can be configured to determine the level of sub-signal 1 in each frequency channel, and the amplification circuit 13641 can be configured to apply a first amplification to sub-signal 1 based on a first set of adaptation curves defining the output level according to the input level and the frequency channel. The level estimation circuit 13622 in amplification pipeline 13602 can be configured to determine the level of sub-signal 2 in each frequency channel, and the amplification circuit 13642 can be configured to apply a second amplification to sub-signal 2 based on a second set of adaptation curves defining the output level according to the input level and the frequency channel. The first amplification and the second amplification can be different; in other words, the first set of adaptation curves and the second set of adaptation curves can be different.

[0137] It should be understood that the WDRC circuit 1258 includes different level estimation circuits 1362 and different amplification circuits 1364 for different sub-signals. One sub-signal may have a block of level estimation circuit 1362 and a block of amplification circuit 1364, each block for a specific frequency channel, and another sub-signal may have a separate block of level estimation circuit 1362 and a block of amplification circuit 1364 for the same frequency channel. Therefore, each amplification pipeline 1360 can be configured to measure the input level of a different sub-signal independently. This can help avoid the "pumping effect," where a level change in one sub-signal may cause a jump in the amplification of another sub-signal that has not changed in the same way, because only a single level estimator is used for the entire signal.

[0138] It should also be understood that although Figure 13 shows more than two sub-signals, more than two amplification pipelines 1360, and more than two amplified sub-signals, in some embodiments, there may be two sub-signals, two amplification pipelines 1360, and two amplified sub-signals.

[0139] It should also be understood that horizontally correlated amplification can be configured to achieve compression, where the dynamic range of the output level is smaller than the dynamic range of the input level. Amplification including compression can be referred to as wide dynamic range compression (WDRC). Therefore, Figure 13 can illustrate multiple WDRC pipelines (i.e., amplification pipeline 1360) configured to perform WDRC.

[0140] In some embodiments, an amplification pipeline 1360 may be configured to perform amplification based on the level of its own associated sub-signal and the levels of one or more other sub-signals. For example, if one sub-signal is a speech signal 603 and another sub-signal is a background noise signal 601, the levels of the speech signal 603 and the background noise signal 601 may be used to calculate the signal-to-noise ratio (SNR), which may then be used to modify the speech and / or noise fit curves. In some embodiments, the level of the speech signal 603 may be used to set the gain of both the speech signal 603 and the background noise signal 601.

[0141] In some embodiments, the WDRC circuit 1258 may lack the level estimation circuit 1362, and therefore the amplification implemented by the amplifier circuit 1364 may not be applied based on the input level. In other words, the amplification can be independent of the input level. As an example, the amplification applied by the amplifier circuit 1364 may include a half-gain rule (the added gain is approximately equal to half the amount of hearing loss) or a quarter-gain rule (the added gain is equal to half the total hearing loss plus one-quarter of the conductive loss component of the hearing loss).

[0142] The memory 1374 can store different fitting curves and / or rule sets for different sub-signals. For example, the memory can store a set of fitting curves for the target speech signal 603, a set of fitting curves for the interfering speech signal 605, and a set of fitting curves for the background noise signal 601. In some embodiments, fitting curves for a specific sub-signal and a specific frequency channel can be stored as a set of input levels, each input level having an associated output level, thereby defining piecewise curves.

[0143] As can be understood from the above, different amplifications can be applied to different audio signals among the multiple audio signals 858. For example, different amplifications can be applied to audio signal 630a, speech signal 603, target speech signal 605, interference speech signal 607, and background noise signal 601. Therefore, the output audio signal 840 may include the target speech signal 605, interference speech signal 607, and background noise signal 601, such that in the output audio signal 601, the volume change of the background noise signal differs from the volume change of the target speech signal by a first volume change difference, and the volume change of the interference speech signal differs from the volume change of the target speech signal by a second volume change difference. Controlling the different amplifications applied to the different audio signals (such as by controlling the adaptation curve stored in memory 1372) allows the first and second volume change differences to be controlled independently. Typically, the WDRC circuit 1258 may include multiple WDRC pipelines (i.e., amplification pipeline 1360) configured to generate the output audio signal 830 by performing WDRC on a combination of audio signals. In some embodiments, the combination of audio signals may include three audio signals. In some embodiments, the three audio signals may be a group or subset of audio signal 630a (“O”), speech signal 603 (“S”), background noise signal 601 (“BN”), target speech signal 605 (“TS”), and interference speech signal 607 (“IS”). (Because in some embodiments, audio signal 630a may be used by mixing circuit 834, audio signal 630a is shown in dashed lines as an input to mixing circuit 834.) Valid combinations of these signals may include at least TS, IS, BN; TS, S, BN; IS, S, BN; TS, IS, O; TS, S, O; IS, S, O; TS, O, BN; IS, O, BN. Volume control.

[0144] As described above, the noise reduction circuit (e.g., noise reduction circuits 524, 824, and / or 1224) can be configured to use a hybrid circuit (e.g., hybrid circuit 834) and / or a WDRC circuit (e.g., WDRC circuit 1258) to output an output audio signal 840, such that the output audio signal 840 includes a target speech signal 605, an interfering speech signal 607, and a background noise signal 601, and such that in the output audio signal 840, the volume change of the background noise signal 601 can differ from the volume change of the target speech signal 605 by a first volume change difference, the volume change of the interfering speech signal 607 can differ from the volume change of the target speech signal 605 by a second volume change difference, and the first and second volume change differences can be independently controllable. The volume change between the volume in the audio signal 630a and the volume in the output signal 840 can be measured. It should also be understood that when the description involves volume change, the volume change can be an increase or a decrease in volume. The following section describes how the first volume change difference and the second volume change difference can be controlled independently.

[0145] Figure 14 illustrates circuitry in an ear-worn device according to certain embodiments described herein. The ear-worn device can be any ear-worn device described herein (e.g., hearing aid 100, glasses 300, ear-worn device 400, and / or ear-worn device 500). Figure 14 illustrates control circuitry 1442, hybrid circuitry 1434 (which may be an example of hybrid circuitry 834) and / or WDRC circuitry 1458 (which may be an example of WDRC circuitry 1258), memory 1444, and communication circuitry 1446. Memory 1444, communication circuitry 1446, and hybrid circuitry 1434 and / or WDRC circuitry 1358 are coupled to control circuitry 1442. Memory 1444 can be configured to store data. Communication circuitry 1446 can be configured to facilitate communication between the ear-worn device and other devices (e.g., smartphones, tablets, laptops, computers) via, for example, a wireless communication link (e.g., Bluetooth or NFMI).

[0146] Control circuitry 1442 can be configured to provide a first volume change control input 1448a and a second volume change control input 1448b to mixing circuitry 1434 and / or WDRC circuitry 1458. Therefore, in some embodiments, control circuitry 1442 can be configured to provide the first volume change control input 1448a and the second volume change control input 1448b to mixing circuitry 1434. In some embodiments, control circuitry 1442 can be configured to provide the first volume change control input 1448a and the second volume change control input 1448b to WDRC circuitry 1458. In some embodiments, control circuitry 1442 can be configured to provide the first volume change control input 1448a and the second volume change control input 1448b to both mixing circuitry 1434 and WDRC circuitry 1458.

[0147] In some embodiments, the mixing circuit 1434 may be configured to receive a first volume change control input 1448a and a second volume change control input 1448b, and to perform mixing using the first volume change control input 1448a and the second volume change control input 1448b, such that a first volume change difference is at least partially controlled by the first volume change control input 1448a, and a second volume change difference is at least partially controlled by the second volume change control input 1448b. For example, the mixing circuit 1434 may be configured to generate an equivalent to a by multiplying the interfering speech signal 607 by b, multiplying the background noise signal 601 by c, and then adding these intermediate products. TS+b IS+c The output audio signal 840 of BN. (For simplicity, the weight a applied to the target speech signal 605 is assumed to be 1.) The mixing circuit 1434 can be configured to control the weight c based on the first volume change input 1448a and the weight b based on the second volume change input 1448b. The weights b and c can then control the second volume change difference and the first volume change difference, respectively. For example, if the weight a is 1 and the weight c is 0.18, the volume change of the target speech signal 605 can be 0 dB and the volume change of the background noise signal 601 can be -15 dB (i.e., the first volume change difference can be -15 dB). If the weight a is 1 and the weight b is 0.5, the volume change of the target speech signal 605 can be 0 dB and the volume change of the interfering speech signal 607 can be -6 dB (i.e., the second volume change difference can be -6 dB). As another example, the mixing circuit 834 can be configured to generate an equivalent to d by multiplying the interfering speech signal 607 by e, multiplying the background noise signal 601 by f, and then summing these intermediates. TS+e IS+f The output audio signal 630 of O. (For simplicity, the weight applied to the target speech signal 605 is assumed to be 1.) The mixing circuit can be configured to make weight f control the input based on a first volume change and weight e control the input based on a second volume change. Weights e and f can together control the first volume change difference and the second volume change difference. Based on the above description a=d+f, b=e+f, c=f, for example, if weight d is 0.82, then weight e is 0.32, and weight f is 0.18, then the volume change of the target speech signal 605 can be 0 dB, the volume change of the background noise signal 601 can be -15 dB (i.e., the first volume change difference can be -15 dB), and the volume change of the interfering speech signal 807 can be -6 dB (i.e., the second volume change can be -6 dB).

[0148] In some embodiments, the first volume change control input 1448a and the second volume change control input 1448b can be numerical values. In such embodiments, the hybrid control circuit 834 can make a weight c equal to the first volume change control input 1448a and a weight b equal to the second volume change control input 1448b. In some embodiments, the hybrid control circuit 834 can derive the weight c from the first volume change control input 1448a and the weight b from the second volume change control input 1448b. For example, the first volume change control input 1448a can be an encoded version of the weight c, and the second volume change control input 1448b can be an encoded version of the weight b, and the hybrid control circuit 834 can be configured to decode the first volume change control input 1448a and the second volume change control input 1448b.

[0149] It should be understood that the first volume change control input 1448a and the second volume change control input 1448b can be different. Therefore, the first volume change difference and the second volume change difference can be different and can be controlled independently.

[0150] As should be understood from the above, in some embodiments, the first volume change control input 1448a can control the weight applied to one signal (e.g., background noise signal 601), and the second volume change control input 1448b can control the weight applied to another signal (e.g., interference speech signal 607). However, the mixing control circuit 834 can be configured to use three weights to mix the three signals together. In some embodiments, if the weight applied to the third signal (e.g., the target speech signal) always has the same value, or is by default the same value (e.g., 1), the third volume change control input may not be used. However, in some embodiments, the third volume change control input can be used to control the weight applied to the third signal using the method described below. For simplicity, this description focuses on the first volume change control input 1448a and the second volume change control input 1448b.

[0151] In some embodiments, the WDRC circuit 1458 may be configured to receive a first volume change control input 1448a and a second volume change control input 1448b, and to perform WDRC using the first volume change control input 1448a and the second volume change control input 1448b, such that a first volume change difference is controlled at least partially by the first volume change control input 1448a, and a second volume change difference is controlled at least partially by the second volume change control input 1448b. For example, the first volume change control input 1448a and the second volume change control input 1448b may be different adaptation curves or rules applied to different signals (or inputs from which adaptation curves or rules can be derived or retrieved).

[0152] In some embodiments, memory 1444 may be configured to store a first volume change control input 1448a and a second volume change control input 1448b. In some embodiments, communication circuitry 1446 may be configured to receive the first volume change control input 1448a and the second volume change control input 1448b from a processing device (i.e., an external device). Memory 1444 may be configured to store the first volume change control input 1448a and the second volume change control input 1448b. In some embodiments, control circuitry 1442 may be configured to retrieve the first volume change control input 1448a and the second volume change control input 1448b from memory 1444 and output the first volume change control input 1448a and the second volume change control input 1448b to mixing circuitry 1458 and / or WDRC circuitry 1458. In some embodiments, the communication circuit 1446 may be configured to receive a first volume change control input 1448a and a second volume change control input 1448b from an external device, and the control circuit 1442 may be configured to receive the first volume change control input 1448a and the second volume change control input 1448b from the communication circuit 1446, and provide the first volume change control input 1448a and the second volume change control input 1448b to the mixing circuit 1434 and / or the WDRC circuit 1458, without storing the data in the memory 1444.

[0153] In some embodiments, the first volume change difference and the second volume change difference can be determined as part of the fitting process. For example, a hearing expert can determine the first volume change difference and the second volume change difference during the fitting process and use their processing device (e.g., a smartphone, tablet, laptop, or computer) to transmit the first volume change control input 1448a and the second volume change control input 1448b (corresponding to the first volume change difference and the second volume change difference, respectively) to the communication circuitry 1446 of the ear-worn device. In some embodiments, the wearer's processing device (e.g., processing device 418, such as a smartphone, tablet, laptop, or computer, etc.) can run an application for communicating with the ear-worn device. In such embodiments, the application can include default values ​​for the first volume change difference and the second volume change difference, and the wearer's processing device can transmit the first volume change control input 1448a and the second volume change control input 1448b (corresponding to the default first volume change difference and the default second volume change difference, respectively) to the communication circuitry 1446 of the ear-worn device. The first volume change control input 1448a and the second volume change control input 1448b can then be stored in memory 1444. In some embodiments, the application may include different default value groups for the first volume change difference and the second volume change difference, and the wearer's processing device may transmit the group of the first volume change control input 1448a and the group of the second volume change control input 1448b (corresponding to the default first volume change difference group and the default second volume change difference group, respectively) to the communication circuitry 1446 of the ear-worn device. The group of the first volume change control input 1448a and the group of the second volume change control input 1448b can then be stored in memory 1444. When the wearer uses the application selection mode of their processing device, the processing device may transmit an indication of the selected mode to the communication circuitry 1446 of the ear-worn device, and the control circuitry 1442 may receive the mode indication and retrieve the first volume change control input 1448a and the second volume change control input 1448b corresponding to the selected mode from memory 1444. In some embodiments, the application may provide the wearer with the option to select a specific first volume change difference and a second volume change difference, or some other value related to the first volume change difference and the second volume change difference, and the wearer's processing device may transmit the first volume change control input 1448a and the second volume change control input 1448b (corresponding to the selected first volume change difference and the default second volume change difference) to the communication circuit 1446 of the ear-worn device. The control circuit 1442 may receive the selected first volume change control input 1448a and the selected second volume change control input 1448b and use them to control the mixing circuit 1434 and / or the WDRC circuit 1458.

[0154] As mentioned herein, circuitry configured to perform operations related to volume change control inputs (e.g., storing, receiving, retrieving, providing, etc.) should be understood to include circuitry configured to perform these operations using the volume change control input itself or using some other data from which the volume change control input can be obtained. For example, when reference memory 1444 is configured to store volume change control inputs, memory 1444 may be configured to store the volume change control input itself, or to obtain some other data from which the volume change control input itself can be obtained (e.g., an encoded version of the volume change control input). As another example, when reference control circuitry 1442 provides volume change control inputs to mixing circuitry 1434 and / or WDRC circuitry 1458, control circuitry 1442 may be configured to provide the volume change control input itself, or to obtain some other data from which the volume change control input itself can be obtained (e.g., an encoded version of the volume change control input).

[0155] As described above, in some embodiments, the mixing circuit 1434 can be configured to mix two signals together. For example, the mixing circuit 1434 can be configured to mix the target speech signal 605 with the audio signal 609, or to mix the target speech signal 605 with the audio signal 603a. More specifically, the mixing circuit 1434 can be configured to output a TS+b (IS+BN) or a TS+b In embodiments where weight a always has the same value or defaults to the same value (e.g., 1), the mixing circuit 1434 may be configured to receive only the volume change control input 1448a that controls the weight b, without receiving the second volume change control input. Therefore, the volume change control input 1448b is shown in the figure with a dashed line. This volume change control input controls the difference in volume change between the interfering speech and the background noise. In other words, the difference in volume change between the interfering speech and the background noise can be the same. The description of the first volume change difference herein can be applied to this volume change difference.

[0156] In other words, the mixing circuit 1434 can be configured to mix two or more audio signals such that the output audio signal includes a background noise-modified and spatially focused version of the audio signal 630a mixed with the second audio signal (i.e., the target speech signal 605). The second audio signal may include a background noise signal 601 and an interfering speech signal 607. For example, the second audio signal may be audio signal 609 or audio signal 630a. A noise reduction circuit (e.g., noise reduction circuit 824) can be configured to generate an output audio signal 840 such that in the output audio signal, the volume change of the background noise signal 601 differs from the volume change of the target speech signal 605 by a volume change difference, and the volume change of the interfering speech signal differs from the volume change of the target speech signal by the same volume change difference, and this volume change difference may be controllable. The mixing circuit 1434 can be configured to receive a volume change control input 1448a and perform mixing using the volume change control input 1448a such that the volume change difference is controlled at least partially by the volume change control input. The communication circuit 1446 can be configured to receive volume change control input 1448a from the processing device, the memory 1444 can be configured to store the volume change control input 1448a, and the control circuit 1442 can be configured to retrieve the volume change control input 1448a and output the volume change control input to the mixing circuit 1434.

[0157] In some embodiments, the amount of volume change in the background noise signal 601 may be based on the background noise level in the audio signal 603a. In some embodiments, the level of background noise may be measured using the background noise components determined by a stationary noise suppression (SNS) circuit. Figure 15 illustrates circuitry in an ear-worn device according to some embodiments described herein. The circuitry includes a neural network circuitry 826, a masking and subtraction circuitry 832, a mixing circuitry 1434, a control circuitry 1542 (which may be identical to the control circuitry 1442), and an SNS circuitry 1570. The ear-worn device may be any ear-worn device described herein (e.g., hearing aid 100, glasses 300, ear-worn device 400, and / or ear-worn device 500). In some embodiments, one or more neural network layers implemented by the neural network circuitry 826 may be particularly effective at reducing non-static noise, and separate stationary noise suppression may be achieved. Thus, the SNS circuitry 1570 may be configured to receive the audio signal 630a and generate a stationary background noise signal 1501, i.e., an estimate of the stationary background noise components of the audio signal 630a. The static background noise signal 1501 can be slowly moving. Qualitatively, the static background noise signal 1501 typically does not change significantly on a timescale of a few seconds. Quantitatively, the static background noise signal 1501 can be asymmetric, as it can be allowed to become smaller on relatively fast timescales, but can be allowed to become larger only on very long timescales (approximately 10 seconds). In some embodiments, the SNS circuit 1570 can be configured to implement a minimum statistical noise estimation algorithm to generate the static background noise signal 1501. In some embodiments, the SNS circuit 1570 can also be configured to implement algorithms other than or replacing the minimum statistical noise estimation algorithm to generate the static background noise signal 1501. In non-limiting examples, these algorithms can include spectral subtraction, Wiener filtering, and the Ephraim-Malah technique. Further description of such algorithms can be found in Chung and King's "Challenges and recent developments in hearing aids: Part I. Speech understanding in noise, microphone technologies and noise reduction algorithms", Amplify Trends, 2004, Vol. 8, No. 3, pp. 83–124, the entire contents of which are incorporated herein by reference.

[0158] Control circuitry 1542 can be configured to receive a static background noise signal 1501 (e.g., generated using SNS circuitry 1570) and generate a first volume change control input 1448a based on the level of the static background noise signal 1501. As described above, in some embodiments, mixing circuitry 1434 can be configured to mix the background noise signal 601 (generated using neural network circuitry 826) with one or more other signals to generate an output audio signal 840. In embodiments such as FIG. 15, the amount of background noise signal 601 mixed with other signals can be based on the static background noise signal 1501, rather than the background noise signal 601 itself (which can be counterintuitive). The two background noise signals do not have to be the same. Performing mixing based on the level of the background noise signal 1501 can be helpful because the static background noise signal 1501 can be a slow motion estimate of the noise, and a slow motion estimate of the noise can help reduce sudden jumps in the mixing coefficients. In some embodiments, the control circuit 1542 may be configured to perform additional smoothing on the static background noise signal 1501 before generating the first volume change control input 1448a based on the static background noise signal 1501. However, in some embodiments, the static background noise signal 1501 may move slowly enough that additional smoothing is not required. In some embodiments, the control circuit 1542 may be configured to convert the units of the static background noise signal 1501 to different units (e.g., from linear units to logarithmic units). However, in some embodiments, unit conversion may not be performed.

[0159] As further shown in Figure 15, the control circuit 1542 can be configured to receive the interfering voice signal 607 and generate a second volume change control input 1448a based on the level of the interfering voice signal 607.

[0160] Figure 16 illustrates circuitry in an ear-worn device according to certain embodiments described herein. The circuitry includes neural network circuitry 826, masking and subtraction circuitry 832, mixing circuitry 1434, and control circuitry 1642 (which may be an example of control circuitry 1442). The ear-worn device may be any ear-worn device described herein (e.g., hearing aid 100, glasses 300, ear-worn device 400, and / or ear-worn device 500). Control circuitry 1670 may be configured to receive a background noise signal 601 (e.g., generated using neural network circuitry 826) and generate a first volume change control input 1448a based on the level of the background noise signal 601. In some embodiments, control circuitry 1642 may be configured to perform additional smoothing on the background noise signal 601 before generating the first volume change control input 1448a based on the background noise signal 601. However, in some embodiments, the background noise signal 601 may move slowly enough that additional smoothing is not required. In some embodiments, control circuitry 1642 may be configured to convert the units of background noise signal 601 to different units (e.g., from linear units to logarithmic units). However, in some embodiments, unit conversion may not be performed. As further shown in FIG16, control circuitry 1670 may be configured to receive interfering speech signal 607 and generate a second volume change control input 1448a based on the level of interfering speech signal 607.

[0161] Then, typically, the control circuitry (e.g., control circuitry 1542 and / or control circuitry 1642) can be configured to generate a first volume change control input 1448a based on the level of background noise in the audio signal 630a, and to generate a second volume change control input 1448b based on the level of interfering speech in the audio signal 630a.

[0162] In some embodiments, control circuitry 1542 and / or 1642 may be configured to determine different volume change control inputs for different frequency bands based on different levels of background noise and / or interfering speech in different frequency bands, and mixing circuitry may be configured to use the different volume change control inputs to mix the different frequency bands together. However, in some embodiments, control circuitry 1542 and / or 1642 may be configured to determine a volume change control input based on a level of background noise and / or interfering speech (e.g., averaging over all frequencies), and this single volume change control input may be used to mix all frequencies together.

[0163] In some embodiments, the amount of background noise mixed back into the output signal can be reduced as the background noise level increases. However, in some embodiments, the amount of background noise mixed back can be increased again once the background noise level exceeds a certain threshold. In some embodiments, the amount of interference speech mixed back into the output signal can be reduced as the interference speech level increases. However, in some embodiments, the amount of interference speech mixed back can be increased again once the interference speech level exceeds a certain threshold.

[0164] Figure 17 illustrates circuitry in an ear-worn device according to certain embodiments described herein. The circuitry includes neural network circuitry 1726 (which may be identical to neural network circuitry 526, 826, 926, 1026, and / or 1126), control circuitry 1442, memory 1444, and communication circuitry 1446. The ear-worn device can be any ear-worn device described herein (e.g., hearing aid 100, eyeglasses 300, ear-worn device 400, and / or ear-worn device 500). As shown, control circuitry 1442 can be configured to output a first volume change control input 1448a and a second volume change control input 1448b to neural network circuitry 1726. The neural network circuit 1726 can be configured to receive a first volume change control input 1448a and a second volume change control input 1448b, and use the first volume change control input 1448a and the second volume change control input 1448b to generate a neural network output 836, such that a first volume change difference is at least partially controlled by the first volume change control input 1448a, and a second volume change difference is at least partially controlled by the second volume change control input 1448b. In some embodiments, the neural network output 836 can be an output audio signal 840, such that the first volume change difference and the second volume change difference are controlled by the first volume change control input 1448a and the second volume change control input 1448b, respectively. In some embodiments, the neural network output 836 can be an output (e.g., a mask) configured to generate (e.g., by applying to audio signal 630a) an output audio signal 840, such that the first volume change difference and the second volume change difference are controlled by the first volume change control input 1448a and the second volume change control input 1448b, respectively. Typically, one or more neural network layers implemented by neural network circuitry 1726 can be trained to perform a weighted combination of target speech signal 605, interfering speech signal 607, and background noise signal 601. In such embodiments, either or both of the processing circuitry (e.g., masking and subtraction circuitry 832) and the mixing circuitry (e.g., mixing circuitry 834 and / or 1434) may be absent. For training, each set of training data may have training input data including multiple audio signals 630 plus a first volume change control input 1448a and a second volume change control input 1448b, and training output data including an output audio signal 840, wherein the first volume change difference and the second volume change difference are as indicated by the first volume change control input 1448a and the second volume change control input 1448b, or are masks configured to generate such an output audio signal 840. Spatial Focusing Mode

[0165] As described above, the target speech signal 605 may have a spatial focus pattern. In other words, the target speech signal 607 may be equivalent to the speech signal 603 to which a specific spatial focus pattern has been applied. Typically, the output audio signal 840 may include a spatially focused signal with a spatial focus pattern. For example, the output audio signal 840 may include the target speech signal 605, or may include an audio signal 630a to which a specific spatial focus pattern has been applied. Figure 18 illustrates exemplary spatial focus patterns according to certain embodiments described herein. Figure 18 illustrates the weights of a function of DOA. (This description will use the following convention: 0 degrees is defined as in front of the wearer of the ear-worn device, 0 to 90 degrees as the wearer's left side, 0 to -90 degrees as the wearer's right side, and 90 to -90 degrees as the wearer's rear side). As shown, a weight of 1 is applied to the sound at DOAs from -30 degrees to 30 degrees, and a weight of 0 is applied to the other DOAs. Therefore, sound originating from 60 degrees in front of the wearer will be preserved, while other sounds will not be preserved. The spatial focusing pattern in Figure 18 involves a sharp transition in space from a DOA using weight 1 to a DOA using weight 0. Therefore, even a slight movement of the wearer's head can cause the sound source to transition from a spatial region using weight 1 to a spatial region using weight 0, resulting in sound cancellation.

[0166] In some embodiments, the weights can transition smoothly or approximately smoothly based on the DOA. Figure 19 illustrates exemplary spatial focusing modes according to certain embodiments described herein. Figure 19 shows the weights as a function of DOA. As shown, the weights smoothly or approximately smoothly transition from a weight of 1 at the DOA of 0 degrees (i.e., directly in front of the wearer) to a weight of 0.5 at the DOAs of 30 degrees and -30 degrees, and to a weight of 0 at the DOAs of 90 degrees and -90 degrees and above. Therefore, the entire area behind the wearer can be hidden. For example, the function shown in Figure 19 can be implemented by inputting a given DOA into a formula. For example, when 0 degrees ≤ DOA ≤ 90 degrees, the formula could be weight = (-1 / 90). DOA+1, and when -90 degrees ≤ DOA < 0 degrees, the formula can be weight = (1 / 90). DOA+1. In some embodiments, multiple DOA intervals (which can also be considered spatial regions) may be defined, and each DOA interval is associated with a weight. For example, each interval may cover 1 degree, such as for intervals for DOAs between 0 and 1 degree, intervals for DOAs between 1 and 2 degrees, etc. For a given DOA, the weight associated with the interval covering the DOA can be determined (e.g., from a lookup table).

[0167] Figure 20 illustrates exemplary spatial focusing patterns according to certain embodiments described herein. Figure 20 is the same as Figure 19, except that the weights decrease from 1 to 0.5 as the DOA increases from 0 to 25 degrees and from 0 to -25 degrees.

[0168] Typically, spatial focusing patterns can include any variation of weights with respect to the DOA (where the DOA can be defined relative to the wearer). In some embodiments, spatial focusing patterns can use weights equal to 0, equal to 1, or between 0 and 1. In some embodiments, spatial focusing patterns can use weights equal to or greater than 0. In some embodiments, weights can be greater than 0, less than 0, equal to zero, or complex; negative weights can flip the phase by 180 degrees, while complex weights can rotate the phase by an angle. As described above with reference to Figures 18-20, some spatial focusing patterns may have weights greater than 0 in a spatial region in front of the wearer and weights of 0 at other DOAs. One way to compare spatial patterns is to use weights greater than a threshold weight (such as 0.5) to determine the size of the spatial region. This spatial region can be considered as a target spatial region, or a spatial region within focus. Then, Figure 19 can be considered to show a spatial focusing pattern focused at 60 degrees in front of the wearer, while Figure 20 can be considered to show a spatial focusing pattern focused at 50 degrees in front of the wearer, and Figure 20 can be considered to show a greater amount of spatial focusing than Figure 19. However, it should be understood that in addition to the spatial focus patterns shown in Figures 18 to 20, other types of spatial focus patterns may be used, and these spatial focus patterns may include any changes in weights with DOA.

[0169] According to the spatial focusing mode, applying the spatial focusing mode to an audio signal can be equivalent to multiplying each component of the audio signal by a weight associated with the DOA from which the component originates (examples are shown in Figures 18-20). Therefore, the resulting audio signal can be the original audio signal to which the spatial focusing mode has been applied. For example, the target speech signal 605 can be the speech signal 603 of audio signal 603a to which the spatial focusing mode has been applied. In embodiments that include focusing on multiple DOAs, multiple different functions similar to those in Figures 18-20 can be generated, and the sum or union of these functions can be used.

[0170] It should be understood that a target space region with a size other than 60 degrees (e.g., as shown in Figure 19) or 50 degrees (e.g., as shown in Figure 20) relative to the wearer can also be used. In some embodiments, the target space region may have a size approximately equal to or between 10 and 180 degrees. In some embodiments, the target space region may have a size approximately equal to or between 20 and 180 degrees. In some embodiments, the target space region may have a size approximately equal to or between 30 and 180 degrees. In some embodiments, the target space region may have a size approximately equal to or between 40 and 180 degrees. In some embodiments, the target space region may have a size approximately equal to or between 50 and 180 degrees. In some embodiments, the target space region may have a size approximately equal to or between 60 and 180 degrees. In some embodiments, the target space region may have a size approximately equal to or between 10 and 150 degrees. In some embodiments, the target space region may have a size approximately equal to or between 20 and 150 degrees. In some embodiments, the target space region may have a size approximately equal to or between 30 and 150 degrees. In some embodiments, the target spatial region may have a size approximately equal to or between 40 and 150 degrees. In some embodiments, the target spatial region may have a size approximately equal to or between 50 and 150 degrees. In some embodiments, the target spatial region may have a size approximately equal to or between 60 and 150 degrees. In some embodiments, the target spatial region may have a size approximately equal to or between 10 and 120 degrees. In some embodiments, the target spatial region may have a size approximately equal to or between 20 and 120 degrees. In some embodiments, the target spatial region may have a size approximately equal to or between 30 and 120 degrees. In some embodiments, the target spatial region may have a size approximately equal to or between 40 and 120 degrees. In some embodiments, the target spatial region may have a size approximately equal to or between 50 and 120 degrees. In some embodiments, the target spatial region may have a size approximately equal to or between 60 and 120 degrees. In some embodiments, the target spatial region may have a size approximately equal to or between 10 and 90 degrees. In some embodiments, the target spatial region may have a size approximately equal to or between 20 and 90 degrees. In some embodiments, the target spatial region may have a size approximately equal to or between 30 and 90 degrees. In some embodiments, the target spatial region may have a size approximately equal to or between 40 and 90 degrees. In some embodiments, the target spatial region may have a size approximately equal to or between 50 and 90 degrees. In some embodiments, the target spatial region may have a size approximately equal to or between 60 and 90 degrees.For example, the size can be equal to or approximately equal to 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, or 180 degrees, or any other suitable angle.

[0171] In some embodiments, the spatial focusing pattern may be predetermined. In other words, the boundaries of different spatial regions and the weights associated with each spatial region can be determined during training. The spatial focusing patterns of Figures 18 to 20 may be examples of predetermined spatial focusing patterns. In some embodiments, the spatial focusing pattern may not be predetermined. In other words, the boundaries of different spatial regions and / or the weights associated with each spatial region can be determined after training time (e.g., at inference time). Further description of such spatial focusing can be found below. Spectrogram

[0172] Referring to FIG8, in some embodiments, the neural network output 836 generated by one or more neural network layers implemented by neural network circuit 826 may be a spectrogram, or an output configured to generate a spectrogram (e.g., a mask). When the neural network output 836 is a mask configured to generate a spectrogram, the mask application and subtraction circuit 832 may be configured to apply the mask to one of a plurality of audio signals 630 (e.g., by multiplication or addition) to generate a spectrogram.

[0173] More specifically, if multiple frequency intervals and multiple DOA intervals are defined, then the spectrogram can indicate the values ​​of each frequency interval derived from each DOA interval. For example, if the values ​​of the frequency intervals are n and the values ​​of the DOA intervals are m, then the spectrogram can be an n×m array. When one or more neural network layers are also trained to perform noise modification, the spectrogram can be a speech spectrogram indicating the frequency components of speech derived from each spatial region.

[0174] Each DOA interval may cover, for example, a certain angular range relative to the wearer. In embodiments where there are two microphones symmetrical about the axis connecting the two microphones (e.g., front microphone 102f and rear microphone 102b shown in FIG. 1), the ear-worn device may not be able to distinguish sound from the wearer's left or right side. In such embodiments, a spatial region can be defined as a combination of symmetrical regions on the wearer's left and right sides. For example, a region of 20-25 degrees to the wearer's left and a region of 20-25 degrees to the wearer's right can both be defined as a single spatial region. In embodiments where there are microphones that are not symmetrical about the axis connecting the microphones, the ear-worn device is able to distinguish sound from the wearer's left and right sides. In such embodiments, symmetrical regions on the wearer's left and right sides can be defined as separate regions. An example of an ear-worn device not symmetrical about the axis connecting the microphones could be eyeglasses with a built-in hearing aid (e.g., eyeglasses 300 with microphone 302). In embodiments where binaural communication is present, the ear-worn device or system of ear-worn devices is able to distinguish sound from the wearer's left and right sides. For example, a device on the wearer's left ear can detect a sound from the wearer's left side before a device on the wearer's right ear, and this earlier detection can communicate between the two ears to determine that the sound originated from the wearer's left side. Binaural communication can occur, for example, between two hearing aids, cochlear implants, or headphone systems that communicate with each other via a wireless communication link. Binaural communication can also occur in ear-worn devices such as eyeglasses with built-in hearing aids, where a device in the portion of the eyeglasses closest to one ear can communicate with a device in the portion of the eyeglasses closest to the other ear via a wired communication link within the eyeglasses.

[0175] In some embodiments, using a spectrogram, the masking and subtraction circuit 832 can be configured to apply a beam pattern to the spectrogram. The result of applying the beam pattern can be a spatially focused audio signal, thus the spectrogram can be used to generate a spatially focused audio signal. To apply the beam pattern, the masking and subtraction circuit 832 can be configured to apply different weights to sounds originating from different DOAs (as indicated by the spectrogram). These weights do not need to be predetermined.

[0176] In some embodiments of ear-worn devices, including arrays with more than two microphones (e.g., on eyeglasses such as glasses 300, etc.), beamforming circuitry can be configured to generate multiple beams (e.g., between 10 and 20 beams or equal to 10 and 20 beams) at each time step, each beam pointing at a different angle relative to the wearer within a 360-degree circle. Masking and subtraction circuitry 832 or neural network circuitry 826 can be configured to calculate a metric based on audio from the multiple beams. Thus, each beam can have a different metric. As an example, the metric could be signal-to-noise ratio (SNR) or speaker power. Masking and subtraction circuitry 832 can be configured to combine the audio from the multiple beams using the metric. For example, when the metric is SNR, masking and subtraction circuitry 832 can be configured to output a sum of audio from each beam, weighted by the SNR of each beam. Specifically, weighting can be reduced for low SNR beams and increased for high SNR beams. Therefore, the focus can be placed on those beams with the highest SNR audio.

[0177] As another example, a metric can be related to speech personalization. Further description of speech personalization can be found in U.S. Patent No. 11,818,523, entitled “System and Method for Enhancing Speech of Target Speaker from Audio Signal in an Ear-Worn Device Using Voice Signatures,” issued November 14, 2023, the entire contents of which are incorporated herein by reference. For example, a metric can indicate the presence of a particular speaker's voice in a particular beam. Masking and subtraction circuitry 832 can be configured to use the metric value to combine audio from multiple beams. In some embodiments, the output can be a slow moving average (e.g., an exponential moving average) of the audio from beams where the particular speaker's voice already exists. Thus, the output can include audio from the beam currently containing the speaker's voice, as well as sound from directions that, weighted according to an averaging function, do not currently contain the particular speaker's voice but recently did.

[0178] Therefore, the result can be a spatially focused audio signal, and thus the values ​​calculated for the metric can be used to generate the spatially focused audio signal. In some embodiments, the processing circuitry can be configured to perform a moving average across audio applications from different beams before summing the beams. Spatial Focus Control

[0179] As described above, a signal can be equivalent to another signal to which a spatial focus mode has been applied. For example, a target speech signal 605 can be equivalent to a speech signal 603 to which a spatial focus mode has been applied. As another example, an output audio signal 840 may include the target speech signal 605. As another example, an output audio signal 840 may include an audio signal 630a to which a spatial focus mode has been applied. In some embodiments, the spatial focus mode used by one or more neural network layers (i.e., by outputting a signal with a spatial focus mode or being configured to generate an output with a spatial focus mode when performing spatial focusing) can be controlled by input to a neural network circuit (e.g., any neural network circuit 826 described herein). In some embodiments, the spatial focus mode can be controlled by input to a processing circuit (e.g., any mask application and subtraction circuit 832 described herein). In some embodiments, the spatial focus mode can be controlled by input to a mixing circuit (e.g., mixing circuits 834 and / or 1434 described herein). Furthermore, as will be described below, in some embodiments, a user can control the spatial focus mode.

[0180] Figure 21 illustrates circuitry for controlling spatial focusing in an ear-worn device according to certain embodiments described herein. Figure 21 shows any one or both of communication circuitry 2146 (which may be identical to communication circuitry 1446) and sensing circuitry 2176 coupled to control circuitry (which may be identical to control circuitry 1442, 1542, and / or 1642). Figure 21 further illustrates noise reduction circuitry 2124 (which may be identical to any noise reduction circuitry described herein, such as noise reduction circuitry 524 and / or 824, etc.) including neural network circuitry 2126, processing circuitry 2132, and hybrid circuitry 2132. The ear-worn device may be any ear-worn device described herein (e.g., hearing aid 100, glasses 300, ear-worn device 400, and / or ear-worn device 500).

[0181] Communication circuitry 2146 can be configured to communicate with other devices (such as processing devices, such as smartphones or tablets, such as processing device 418, etc.) via a wireless communication link (e.g., wireless communication link 420). For example, the wireless communication link could be Bluetooth or NFMI. In some embodiments, a user can use the processing device to select a spatial focus mode (e.g., a specific spatial focus mode possessed by the target voice signal 605), or otherwise make spatial focus-related selections. Communication circuitry 2146 in the ear-worn device can be configured to receive an indication of the user's selection of the spatial focus mode from the processing device and generate one or more inputs 2166 based on the indication of the user's selection of the spatial focus mode.

[0182] Control circuitry 2142 may be configured to generate one or more spatial focus control inputs 2168 indicating a spatial focus mode, based at least in part on a user's indication of a spatial focus mode selection, and specifically at least in part on one or more inputs 2166 received from communication circuitry 2146. One or more spatial focus control inputs 2168 may indicate a spatial focus mode (i.e., selected by the user). One or more spatial focus control inputs 2168 may be inputs to one or more of the following: neural network circuitry 2126 (which may be identical to any neural network circuitry described herein, such as neural network circuitry 526, 826, 926, 1026, and / or 1126, etc.), processing circuitry 2132 (which may be identical to any processing circuitry described herein, such as mask application and subtraction circuitry 832, etc.), and hybrid circuitry 2134 (which may be any hybrid circuitry described herein, such as hybrid circuitry 834 and / or 1434, etc.). As further described below, one or more spatial focus control inputs 2168 can control spatial focus by controlling neural network circuit 2126, processing circuit 2132 and / or hybrid circuit 2134.

[0183] In embodiments where one or more spatial focus control inputs 2168 from control circuitry 2142 are input to neural network circuitry 2126, the one or more spatial focus control inputs 2168 may supplement the multiple audio signals 630 input to neural network circuitry 2126. The one or more spatial focus control inputs 2168 may indicate a spatial focus mode (e.g., as selected by a user). One or more neural network layers implemented by neural network circuitry 2126 may be trained such that the one or more spatial focus control inputs 2168 influence the mode used for spatial focus performed by the one or more neural network layers. In other words, neural network circuitry 2126 may be configured to implement one or more neural network layers trained to generate an output audio signal (e.g., a target speech signal 605 with a spatial focus mode indicated by one or more spatial focus control inputs 2168) based on multiple input audio signals 2134, or to generate an output (e.g., a mask) configured to generate an output audio signal (e.g., a target speech signal 605 with a spatial focus mode indicated by one or more spatial focus control inputs 2168). For example, referring back to FIG8, the target speech signal 605 can be equivalent to the speech signal 603 that has been applied with a specific spatial focus mode. The neural network circuit 2126 can be configured to receive one or more spatial focus control inputs 2168 indicating a specific spatial focus mode, and use one or more spatial focus control inputs 2168 to generate a neural network output 836 such that the target speech signal 605 is equivalent to the speech signal 603 that has been applied with a specific spatial focus mode.

[0184] For training such neural network layers, training can be performed as described above, except that one or more spatial focus control inputs 2168 corresponding to the spatial focus pattern can be added to the input training data, and the output training data can correspond to that spatial focus pattern. For example, input training data comprising multiple audio signals can be set up with inputs 2168 having specific values ​​corresponding to a particular spatial focus pattern. The output training data can be a mask, and when the mask is applied to one of the multiple audio signals, the spatial focus pattern is obtained by applying the mask to that audio signal.

[0185] It should be understood that one or more neural network layers may receive one or more spatial focus control inputs 2168 that indicate the spatial focus mode, rather than by circuitry operating on the output of one or more neural network layers receiving one or more spatial focus control inputs 2168.

[0186] In some embodiments, one or more neural network layers may also be configured not to receive input indicating a spatial focusing mode, and then the neural network may be configured to apply a default spatial focusing mode.

[0187] In embodiments where one or more spatial focus control inputs 2168 from control circuitry 2142 are input to processing circuitry 2132, in some such embodiments, processing circuitry 2132 may be configured to apply beam patterns to spectrograms based on one or more spatial focus control inputs 2168.

[0188] In embodiments where one or more spatial focus control inputs 2168 from control circuitry 2142 are input to mixing circuitry 2134, the target speech signal 605 is referred to as TS, the interfering speech signal 607 as IS, and the background noise signal 601 as BN. The output audio signal (not shown) from mixing circuitry 2134 can be equivalent to a TS+b IS+c BN. A relatively low value of b yields more spatial focus, while a relatively high value of b yields less spatial focus. If the mixing circuit 2134 is configured to mix a spatially focused version of the audio signal (let's call it A) with the audio signal itself (let's call it B), the output is a. A+b B, then using a value of a that is relatively higher than b will result in more spatial focus, while using a value of a that is relatively lower than b will result in less spatial focus. Therefore, based on one or more spatial focus control inputs 2168 from control circuit 2142, mixing circuit 2134 can be configured to modulate the weights used for mixing, which can actually modify the weights applied to sounds from different DOAs, thereby controlling the spatial focus.

[0189] In some embodiments, spatial focusing can be disabled. When spatial focusing is disabled, the output signal from the noise reduction circuit 2124 can be a noise-modified audio signal, and in some cases, a noise-modified beamforming audio signal. Therefore, in some embodiments, when spatial focusing is disabled, a portion of the speech signal 603 (i.e., the interfering speech signal 607) will not have a different volume variation based on spatial focusing than another portion of the speech signal 603 (i.e., the target speech signal 605). Various methods can be used to disable spatial focusing. In some embodiments, one or more spatial focusing control inputs 2168 may be present that are associated with no spatial focusing, causing the neural network circuit 2126 not to perform spatial focusing when input to it. For example, in some embodiments, one or more spatial focusing control inputs 2168 may cause the neural network circuit 2126 not to run one or more neural network layers (e.g., the second subset 950b) trained to perform spatial focusing. As another example, one or more spatial focusing control inputs 2168 may cause the neural network circuit 2126 to use a spatial focusing pattern with a weight of 1 on each DOA. In some embodiments, one or more spatial focus control inputs 2168 may cause the neural network circuit 2126 to output a mask (e.g., mask 956b) configured to generate speech signal 603 instead of target speech signal 605 (e.g., the mask may be all 1s). In some embodiments, one or more spatial focus control inputs 2168 associated with no spatial focus may be input to processing circuit 2132 and cause processing circuit 2132 not to perform spatial focus. For example, processing circuit 2132 may be configured to output speech signal 603 instead of target speech signal 605, or may be configured to replace the mask (e.g., mask 956b) with a different mask configured to generate speech signal 603. In some embodiments, the spectrogram is set to include multiple columns, each column corresponding to sounds from different spatial regions. One or more spatial focus control inputs 2168 may cause processing circuit 2132 not to modify the weights of the columns of the spectrogram, or in other words, to apply a weight of 1 to the columns of the spectrogram. In some embodiments, one or more spatial focus control inputs 2168 associated with no spatial focus can be input to the mixing circuit 2134, causing the mixing circuit 2134 to not perform spatial focus. For example, one or more spatial focus control inputs 2168 can cause the mixing circuit 2134 to remix the full interference speech signal 607 with the target speech signal 605. In other words, referring to expression a above... TS + b IS + c BN, where both a and b can be 1. As another example, the mixing circuit 2134 can weight the spatially focused signal with 0 and the non-spatially focused signal with 1 when performing mixing. In other words, referring to the expression above, a... A + b B, where a can be 0 and b can be 1. In some embodiments, the user selection can cause spatial focus to be turned off. In other words, in some embodiments, the ear-worn device (e.g., hearing aid 100, glasses 300, ear-worn device 400, and / or ear-worn device 500) can be configured to receive a user selection to turn off spatial focus. Based on receiving the user selection to turn off spatial focus, the ear-worn device can be configured to turn off spatial focus, for example, using any of the methods described above. Further description can be found below.

[0190] In some embodiments, noise modification performed using neural network circuit 2126 can be disabled. In some embodiments, to disable noise modification, one or more neural network layers trained to perform noise modification may not be run. (Graphical User Interface)

[0191] In some embodiments, spatial focusing may be based on user selection. The user can control spatial focusing using a processing device (e.g., a smartphone or tablet, such as processing device 418, etc.) that communicates with an ear-worn device (e.g., hearing aid 100, glasses 300, ear-worn device 400, and / or ear-worn device 500). Communication can be via a wireless communication link (e.g., wireless communication link 420). Any GUI described herein can be displayed on such a processing device. This user selection can be the user selection described above; in other words, an indication of the user selection of a spatial focusing mode can be received by the communication circuitry 2146 of the ear-worn device, as described above. In some embodiments, the user selection can be a selection of one of several options for a spatial focusing mode. Therefore, in some embodiments, the processing device can be configured to display a graphical user interface (GUI) including options for different spatial focusing modes, and to receive user selections for a particular spatial focusing mode.

[0192] Figure 22 illustrates a graphical user interface (GUI) 2280 for controlling spatial focusing of an ear-worn device (e.g., a hearing aid) according to certain embodiments described herein. GUI 2280 includes four options 2282a-2282d for spatial focusing modes. Typically, a GUI may include multiple options for spatial focusing modes. In the example of Figure 22, different spatial focusing modes have different spatial focusing amounts. GUI 2280 displays options 2282a-2282d as graphical representations of different spatial focusing modes. Each graphical representation in the example of Figure 22 includes a circle representing the wearer's environment and a highlighted area representing where spatial focusing occurs (e.g., the target spatial region). Option 2282a corresponds to a spatial focusing mode that does not include or barely includes spatial focusing. Option 2282b corresponds to a spatial focusing mode that focuses within 180 degrees (or approximately 180 degrees) in front of the wearer. Option 2282c corresponds to a spatial focusing mode that focuses within 90 degrees (or approximately 90 degrees) in front of the wearer. Option 2282d corresponds to a spatial focusing mode that focuses within a 45-degree (or approximately 45-degree) angle in front of the wearer. Therefore, among the displayed options, option 2282a may represent the minimum amount of spatial focusing, and option 2282d may represent the maximum amount of spatial focusing. The spatial focusing modes corresponding to options 2282a-2282d may have the form shown in Figure 19, where the mode is considered to focus on those DOAs with a weight greater than or equal to a threshold (such as 0.5). The processing device displaying GUI 2280 can be configured to receive a selection of one of options 2282a-2282d, for example, via a touch-sensitive display showing options 2282a-2282d. Based on the user selection from GUI 2280, the processing device can be configured to transmit the indication of the user selection to the communication circuitry 2146 of the ear-worn device. The communication circuit 2146 can be configured to generate one or more inputs to the control circuit 2166 based on the received user-selected instructions, and the control circuit 2142 can be configured to transmit one or more spatial focusing control inputs 2168 to the neural network circuit 2126, the processing circuit 2132, and / or the hybrid circuit 2134 based on one or more inputs received from the communication circuit 2146.

[0193] While the examples above include four options for one or more inputs to a neural network, in some embodiments there may be fewer than four options, and in some embodiments there may be more than four options. In some embodiments, there may be a large or continuous range of options for spatial focus modes (e.g., from omnidirectional to superfocus). For example, a user can use a slider on a graphical user interface to select a spatial focus mode from a continuous range of options, where the slider selects the range of DOA within the focus in front of the wearer.

[0194] For example, in an embodiment where one or more spatial focus control inputs 2168 are input to the neural network circuit 2126, if option 2282a is selected, one or more spatial focus control inputs 2168 can be 0; if option 2282b is selected, one or more spatial focus control inputs 2168 can be 1; if option 2282c is selected, one or more spatial focus control inputs 2168 can be 2; and if option 2282d is selected, one or more spatial focus control inputs 2168 can be 3. As another example (i.e., the one-hot encoding scheme), if option 2282a is selected, one or more spatial focus control inputs 2168 can be [1,0,0,0]; if option 2282b is selected, one or more spatial focus control inputs 2168 can be [0,1,0,0]; if option 2282c is selected, one or more spatial focus control inputs 2168 can be [0,0,1,0]; if option 2282d is selected, one or more spatial focus control inputs 2168 can be [0,0,0,1].

[0195] It should be understood that one or more neural network layers can receive one or more spatial focus control inputs 2168 indicating the spatial focus mode themselves, rather than the circuitry operating on the output of one or more neural network layers receiving the spatial focus control inputs 2168. Therefore, for example, one or more neural network layers can be configured to receive audio input in vector form of size 128, plus spatial focus control inputs 2168 indicating the spatial focus mode using a one-hot encoding scheme with a vector of size 4. Thus, the total size of the input received by one or more neural network layers will be 128 + 4 = 132.

[0196] To train such a neural network layer, training can be performed as described above, except that one or more spatial focus control inputs 2168 corresponding to a spatial focus pattern can be added to the input training data, and the output training data can correspond to that spatial focus pattern. For example, the input training data can be configured to include multiple audio signals formed by one or more sound signals, plus inputs [0,0,1,0] corresponding to option 2282c in Figure 22. The output training data can be a mask that, when applied to one of the multiple audio signals, focuses those sound signals originating within 90 degrees in front of the wearer and extending to +45 degrees and -45 degrees on either side of the direction corresponding to the wearer's front, such that the resulting audio signal has a spatial focus pattern.

[0197] In embodiments where one or more spatial focus control inputs 2168 are input to processing circuitry 2132, the spectrogram processed by processing circuitry 2132 is configured to include 16 spatial regions. Option 2282d can be implemented by focusing on the first two spatial regions out of the 16 total spatial regions. Option 2282c can be implemented by focusing on the first four spatial regions out of the 16 total spatial regions. Option 2282b can be implemented by focusing on the first eight spatial regions out of the 16 total spatial regions. Option 2282a can be implemented by focusing on all 16 total spatial regions. The spectrogram is configured to include 16 columns, each column corresponding to one of the 16 spatial regions. Based on one or more spatial focus control inputs 2168, the processing circuit 2132 can be configured to apply a weight of 1 to values ​​in columns corresponding to focused spatial regions and a weight of 0 to values ​​in columns corresponding to other spatial regions (e.g., to achieve the spatial focus mode shown in Figure 18), or to apply weights equal to or between 0.5 and 1 to columns corresponding to focused spatial regions and weights equal to or between 0.5 and 0 to columns corresponding to non-focused spatial regions (e.g., to achieve the spatial focus mode shown in Figure 19). Typically, the processing circuit 2132 can be configured to use higher weights for columns corresponding to focused spatial regions and lower weights for values ​​in columns not corresponding to focused spatial regions.

[0198] In some embodiments, user selection may involve choosing which spatial regions to focus on and which to not focus on. Figure 23 illustrates a graphical user interface (GUI) 2380 for controlling spatial focusing of an ear-worn device (e.g., a hearing aid) according to certain embodiments described herein. GUI 2380 includes a circle 2384 representing the wearer's environment, where the wearer is considered to be at the center of circle 2384, with the area in front of the wearer represented by the right side of circle 2384 and the area behind the wearer represented by the left side of circle 2384. Circle 2384 includes a plurality of spatial regions 2386a-2386b that the wearer can select. In the example of Figure 23, the wearer has selected spatial regions 2386a and 2386b, causing them to be highlighted. It should be understood that when using GUI 2380, the wearer is able to select one, two, or more than two spatial regions 2386a-2386f. It should also be understood that more or fewer than six spatial regions may be used in GUI 2380. Based on user selections from GUI2380, the processing device can be configured to transmit the user-selected indication to the communication circuitry 2146 of the ear-worn device. The communication circuitry 2146 can be configured to generate one or more inputs to the control circuitry 2166 based on the received user-selected indication, and the control circuitry 2142 can be configured to transmit one or more spatial focus control inputs 2168 to the neural network circuitry 2126, the processing circuitry 2132, and / or the hybrid circuitry 2134 based on one or more inputs received from the communication circuitry 2146.

[0199] In embodiments where one or more spatial focus control inputs 2168 are input to the neural network circuit 2126, the inputs 2168 may be, for example, vectors having as many elements as spatial regions 2386a-2386b, wherein an element is equal to 1 when its corresponding spatial region is selected and 0 otherwise. Further description of training neural network layers implemented by such neural network circuits can be found above.

[0200] In embodiments where one or more spatial focus control inputs 2168 are input to processing circuitry 2132, the spectrogram is configured to include multiple columns, each corresponding to a different sound from spatial regions 2386a-2386f. Based on one or more spatial focus control inputs 2168, processing circuitry 2132 can be configured to apply a weight of 1 to values ​​in columns corresponding to selected spatial regions (e.g., 2386a and 2386b in FIG. 23) and a weight of 0 to values ​​in columns corresponding to unselected spatial regions (e.g., to achieve the spatial focus mode shown in FIG. 18), or to apply weights equal to 0.5 and 1 or between 0.5 and 1 to columns corresponding to selected spatial regions and weights equal to 0.5 and 0 or between 0.5 and 0 to columns corresponding to unselected spatial regions (e.g., to achieve the spatial focus mode shown in FIG. 19). Typically, processing circuitry 2132 can be configured to use higher weights for columns corresponding to selected spatial regions and lower weights for values ​​in columns corresponding to unselected spatial regions.

[0201] In some embodiments, the user can choose whether or not to perform spatial focusing. Figure 24 illustrates a graphical user interface (GUI) 2480 for controlling spatial focusing of an ear-worn device (e.g., a hearing aid) according to certain embodiments described herein. GUI 2480 includes an option 2488 that can be triggered by the user to turn spatial focusing on or off. Based on the user selection from GUI 2380, a processing device can be configured to transmit the user-selected instruction to a communication circuitry 2146 of the ear-worn device. The communication circuitry 2146 can be configured to generate one or more inputs to a control circuitry 2166 based on the received user-selected instruction, and the control circuitry 2142 can be configured to transmit one or more spatial focusing control inputs 2168 to a neural network circuitry 2126, a processing circuitry 2132, and / or a hybrid circuitry 2134 based on one or more inputs received from the communication circuitry 2146. In some embodiments, physical user inputs on the ear-worn device (e.g., user input device 104, such as buttons, etc.) can receive the user selection to turn spatial focusing on or off. Based on user activation via physical user input on the ear-worn device, control circuitry 2142 can be configured to transmit one or more spatial focus control inputs 2168 to neural network circuitry 2126, processing circuitry 2132, and / or hybrid circuitry 2134. Further description of how one or more spatial focus control inputs 2168 can control the opening and closing of spatial focus can be found above.

[0202] In some embodiments, the user selection may be how much focus to perform. Figure 25 illustrates a graphical user interface (GUI) 2580 for controlling spatial focus of an ear-worn device (e.g., a hearing aid) according to certain embodiments described herein. GUI 2580 includes a slider option 2590 that a user can use to control the degree of focus. Slider option 2590 includes a line 2592 and a slider 2594. The wearer can control the position of slider 2594 on line 2592. 1. The ratio between the distance from the left end of line 2592 to the position of slider 2594 and 2. the length of line 2592 can be a value between 0 and 1, where a value closer to 1 indicates more spatial focus, and a value closer to 0 indicates less spatial focus. The communication circuit 2146 can be configured to generate one or more inputs to the control circuit 2166 based on the indication of the received user-selected value, and the control circuit 2142 can be configured to transmit one or more spatial focusing control inputs 2168 to the neural network circuit 2126, the processing circuit 2132, and / or the hybrid circuit 2134 based on one or more inputs received from the communication circuit 2146.

[0203] In embodiments where one or more spatial focus control inputs 2168 are input to neural network circuit 2126, one or more spatial focus control inputs 2168 indicating a user-selected value (between 0 and 1) for the spatial focus amount can have values ​​associated with the user-selected spatial focus amount. By analogy with FIG22, when the user-selected value is close to or equal to 1, one or more spatial focus control inputs 2168 can have a value associated with option 2282d; when the user-selected value is close to or equal to 0, one or more spatial focus control inputs 2168 can have a value associated with option 2282a; and when the user-selected value is between 0 and 1, one or more spatial focus control inputs 2168 can have a value associated with option 2282b or 2282c. Further description of training neural network layers implemented by such neural network circuits can be found above.

[0204] In embodiments where one or more spatial focus control inputs 2168 are input to processing circuitry 2132, the spectrogram is configured to include multiple columns, each corresponding to sounds from different spatial regions. When the user-selected value is close to or equal to 1, one or more spatial focus control inputs 2168 cause processing circuitry 2132 to apply weights to the columns of the spectrogram to achieve a spatial focus mode with a large amount of spatial focus. When the user-selected value is close to or equal to 0, one or more spatial focus control inputs 2168 cause processing circuitry 2132 to apply weights to the columns of the spectrogram to achieve a spatial focus mode with a small amount of spatial focus. When the user-selected value is between 0 and 1, one or more spatial focus control inputs 2168 cause processing circuitry 2132 to apply weights to the columns of the spectrogram to achieve a spatial focus mode with a moderate amount of spatial focus.

[0205] In embodiments where one or more spatial focus control inputs 2168 are input to the mixing circuit 2134, the mixing circuit 2134 is configured to mix a spatially focused version of the audio signal (referred to as A) with the audio signal itself (referred to as B), such that the output is a A + b B. In some embodiments, one or more spatial focus control inputs 2168 may cause the mixing circuit 2134 to use a value equal to the user-selected spatial focus amount, a. In such embodiments, b may be a constant or may be inversely proportional to a.

[0206] In some embodiments, user selection may be focusing on which speaker. Figure 26 illustrates a graphical user interface (GUI) 2680 for controlling spatial focusing of an ear-worn device (e.g., a hearing aid) according to certain embodiments described herein. GUI 2680 includes representations of the wearer 2696 in different directions relative to the wearer and representations of four speakers 2698a-2698d. In the example of Figure 26, the wearer has selected speaker 2698a, making it highlighted on GUI 2680.

[0207] In some embodiments, the ear-worn device can be configured to generate multiple tight beams using beamforming on an array of more than two microphones, and the ear-worn device or processing can be configured to calculate the power of the speech signal in the audio from each beam. Beams with power above a threshold can be considered to have a speaker in their direction. In some embodiments, a neural network can be trained to determine the speaker's direction based on the audio from one or more beams. In some embodiments, the ear-worn device can run the neural network, and information about the speaker's direction can be transmitted from the ear-worn device to a processing device. The processing device can then use this information to display a representation of the speaker in each of their respective directions in a GUI. In some embodiments, the neural network can run on the processing device itself.

[0208] Based on user selection from GUI 2380, the processing device can be configured to transmit the user-selected instruction to the communication circuitry 2146 of the ear-worn device. The communication circuitry 2146 can be configured to generate one or more inputs to the control circuitry 2166 based on the received user-selected instruction, and the control circuitry 2142 can be configured to transmit one or more spatial focusing control inputs 2168 to the neural network circuitry 2126, the processing circuitry 2132, and / or the hybrid circuitry 2134 based on one or more inputs received from the communication circuitry 2146. Based on the wearer's selection of one of the speaker representations from the GUI, the processing device can transmit the selection instruction to the ear-worn device, and the ear-worn device can use a beam focused in the direction the speaker is moving forward.

[0209] In embodiments where one or more spatial focus control inputs 2168 are input to neural network circuitry 2126, the spatial focus control inputs 2168 may be associated with a spatial focus pattern pointing in the direction of the selected speaker. In embodiments where one or more spatial focus control inputs 2168 are input to processing circuitry 2132, the spatial focus control inputs 2168 may cause processing circuitry 2132 to apply weights to columns of the spectrogram that implement the spatial focus pattern pointing in the direction of the selected speaker.

[0210] In some embodiments, the ear-worn device or processing device can be configured to determine the signal-to-noise ratio (SNR) of the acoustic environment, and the ear-worn device can be configured to generate one or more spatial focus control inputs 2168 indicating a spatial focus mode based on the SNR of the acoustic environment. If the acoustic environment has a high SNR, the ear-worn device can be configured to select fewer spatial focuses compared to a case where the acoustic environment has a low SNR. Sensing circuitry

[0211] Referring back to Figure 21, in some embodiments, sensing circuitry 2176 may include one or more of an accelerometer, a gyroscope, and a magnetometer. Sensing circuitry 2176 may be configured to generate one or more inputs 2178 based on the motion of the ear-worn device. In some embodiments, control circuitry 2142 may be configured to determine the degree of head movement (i.e., how fast the wearer's head is moving) based on one or more inputs 2178 received from sensing circuitry 2176. Further description of how sensors determine how fast the wearer's head is moving can be found in "Using Inertial Sensors to Determining Head Motion—A Review" by Ionut-Cristian S and Dan-Marius D, *Imaging Journal*, December 6, 2021; 7 (12):265, the entire contents of which are incorporated herein by reference. In some embodiments, control circuitry 2142 may be configured to generate one or more spatial focus control inputs 2168 indicating a specific spatial focus mode based on the degree of head movement. For example, control circuitry 2142 can be configured to select a spatial focus mode that has a greater spatial focus amount when the wearer's head is moving slowly or not moving at all than when the wearer's head is moving rapidly. As a specific example of using four intervals for the degree of head movement, no head movement can be associated with a spatial focus mode with a large amount of spatial focus, low-degree head movement can be associated with a spatial focus mode with a medium amount of spatial focus, medium-degree head movement can be associated with a spatial focus mode with a low amount of spatial focus, and high-degree head movement can be associated with no spatial focus. As another example, if the head movement speed is above a threshold, a spatial focus mode with a certain amount of spatial focus (which may be no spatial focus) can be used, and if the head movement speed is below the threshold, a spatial focus mode with another amount of spatial focus can be executed. basing the spatial focus amount on head movement speed can be helpful because it can potentially: 1. assist the neural network in achieving higher performance during head rotation and 2. broaden the spatial focus while assuming that head rotation is not related to the need for tight spatial focus. Typically, the control circuit 2142 can be configured to generate a first set of one or more spatial focus control inputs 2168 indicating a spatial focus mode with a first spatial focus amount based on a first head movement degree, and to generate a second set of one or more spatial focus control inputs 2168 indicating a spatial focus mode with a second spatial focus amount based on a second head movement degree, wherein the first spatial focus amount is less than the second spatial focus amount, and the first head movement degree is greater than the second head movement degree.

[0212] In some embodiments, instead of using sensing circuitry 2176 to determine how fast the wearer's head is moving, the ear-worn device may be configured to use a neural network trained to determine how fast the wearer's head is moving (e.g., using sound received by the ear-worn device as input).

[0213] In embodiments where one or more spatial focus control inputs 2168 are input to the neural network circuit 2126, one or more neural network layers implemented by the neural network circuit 2126 can be trained such that the one or more spatial focus control inputs 2168 influence the spatial focus pattern implemented by the one or more neural network layers. In other words, the neural network circuit 2126 can be configured to implement one or more neural network layers trained to generate an output audio signal having a spatial focus pattern indicated by one or more spatial focus control inputs 2168 based on a plurality of input audio signals 2130, or to generate an output configured to generate an output audio signal having a spatial focus pattern indicated by one or more spatial focus control inputs 2168.

[0214] For example, in an embodiment where one or more spatial focus control inputs 2168 are input to the neural network circuit 2126, the control circuit 2142 can be configured to determine which of four intervals the head movement velocity falls into based on one or more inputs 2178 from the sensing circuit 2176. If no head movement is detected, one or more spatial focus control inputs 2168 can be 0; if a low level of head movement is detected, one or more spatial focus control inputs 2168 can be 1; if a moderate level of head movement is detected, one or more spatial focus control inputs 2168 can be 2; and if a high level of head movement is detected, one or more spatial focus control inputs 2168 can be 3. As another example of using a one-hot encoding scheme, if no head movement is detected, one or more spatial focus control inputs 2168 can be [1,0,0,0]; if low-level head movement is detected, one or more spatial focus control inputs 2168 can be [0,1,0,0]; if moderate-level head movement is detected, one or more spatial focus control inputs 2168 can be [0,0,1,0]; and if high-level head movement is detected, one or more spatial focus control inputs 2168 can be [0,0,0,1]. Further descriptions of training such neural networks can be found above.

[0215] In embodiments where one or more spatial focus control inputs 2168 are input to processing circuitry 2132, the spatial focus control inputs 2168 may enable processing circuitry 2132 to process spectrograms to achieve spatial focus patterns associated with head motion (as indicated by one or more spatial focus control inputs 2168).

[0216] As described above, in some embodiments, the ear-worn device can generate a spectrogram indicating frequency components originating from each of multiple spatial regions. In such embodiments, applying a moving average to values ​​determined for different spatial regions can be helpful, as this averages out some errors. However, if the wearer rapidly rotates their head, this can blur the average across different spatial regions. As will be further described below, sensing circuitry 2176, configured to track head movements (e.g., using accelerometers and gyroscopes), is capable of correcting for this.

[0217] It should be understood that in an array of two microphones (e.g., on a hearing aid such as hearing aid 100), the beam that can be created using beamforming (e.g., cardioid and super-cardioid) can be wide, such that when the wearer speaks to someone in front of them and then turns their head (even by 90 degrees), the amplitude of the person's voice may only decrease by a few dB. However, using an array with more than two microphones (e.g., on glasses such as glasses 300), a narrower beam can be generated. With a narrow beam, even a slight head rotation can significantly reduce the amplitude of the sound from a person who was previously directly in front of the wearer. The sensing circuitry 2176, configured to track head movements (e.g., using an accelerometer and a gyroscope), is also able to correct for this.

[0218] More specifically, the sensing circuitry 2176, configured to track head movements (e.g., using accelerometers and gyroscopes), enables the definition of a spatial region in an absolute coordinate system rather than the coordinate system of the wearer's head (which may rotate very rapidly). The absolute coordinate system can be defined relative to the wearer's head, but on a slow time scale. Therefore, if the wearer is sitting and talking to someone, briefly turns their head to look at something, and then turns back to the person they were talking to, the coordinate system may remain in the same position and not rotate with the head (or not rotate much). However, if the wearer turns and begins talking to another person, the coordinate system may slowly (e.g., over several seconds) rotate with the head. To achieve this, an exponential moving average can be applied to the coordinate system such that it is an exponential moving average of the head's orientation. The time scale of the exponential moving average can be, for example, several seconds (e.g., 2, 3, 4 or 5, 6, 7, 8, 9 or 10 seconds, or any other suitable value). In some embodiments, during head movement, sound from a new direction (i.e., the direction the wearer turns their head toward) can be focused immediately, but focus can continue and sound from an old direction (i.e., the direction the wearer turns their head away from) can be slowly released. In other words, when the wearer turns their head toward a new direction, the aperture can be widened fairly quickly to focus on sound from the new direction, but focus continues on sound from the previous direction, and then the focus on sound from the previous direction is slowly reduced as the wearer continues to look in the new direction. This reduction can be modulated according to how long the wearer looks in the new direction so that rapid head saccades do not cause a permanent widening of the aperture. In some embodiments, this behavior can be achieved by an exponential moving average with a long time scale. In some embodiments, this behavior can be achieved by combining 1. the fully weighted sound from the new direction with 2. the sound from the old direction processed by an exponential moving average. Beamforming and Directional Patterns

[0219] As described above, in some embodiments, a neural network circuit (e.g., any neural network circuit described herein) can be configured to receive a plurality of (i.e., at least two) audio signals (e.g., audio signal 830), wherein at least two of the plurality of audio signals each originate from different microphones of two or more microphones, and / or at least one of the plurality of audio signals is a beamformed version of the audio signals originating from two or more microphones. Regarding beamforming, beamforming typically includes applying a delay (which should be understood to include zero delay) to one or more audio signals originating from different microphones, and summing the delayed signals (which should be understood to include subtraction). Different delays may be applied to signals originating from different microphones. In embodiments including only two microphones (such as a front microphone (e.g., front microphone 102f) and a rear microphone (e.g., rear microphone 102b) etc.), beamforming may include applying a delay to the signal from one of the microphones and subtracting the delay from the signal from the other microphone. The resulting signal may have a beamforming directional pattern that depends at least in part on the spacing between the front and rear microphones and the applied delay; in other words, the weight of the resulting signal can vary as a function of the angle from the microphones. Examples of beamforming directional patterns include dipole, supercardioid, supercardioid, and cardioid. Some beamforming directional patterns (which may be referred to herein as "forward") typically attenuate the signal from behind the wearer more than the signal from in front of the wearer. As will be described below, for both the front and rear microphones, a forward beamforming directional pattern can typically be generated by applying a delay larger than that of the front microphone to the rear microphone and subtracting the rear signal from the front signal. Some beamforming directional patterns (which may be referred to herein as "backward") typically attenuate the signal from in front of the wearer more than the signal from behind the wearer. As will be described below, for both the front and rear microphones, a backward beamforming directional pattern can typically be generated by applying a delay larger than that of the rear microphone to the front microphone and subtracting the front signal from the rear signal.

[0220] Figure 27 illustrates a forward supercardioid pattern 2776 according to certain embodiments described herein. The forward supercardioid pattern can be obtained by applying a delay of d / 3c to the signal from the rear microphone, where d is the distance between the front and rear microphones and c is the speed of sound. Figure 28 illustrates a backward supercardioid pattern 2876 according to certain embodiments described herein. The backward supercardioid pattern can be obtained by applying a delay of d / 3c to the signal from the front microphone. Figure 29 illustrates a forward supercardioid pattern 2976 according to certain embodiments described herein. The forward supercardioid pattern can be obtained by applying a delay of 2d / 3c to the signal from the rear microphone. Figure 30 illustrates a backward supercardioid pattern 3076 according to certain embodiments described herein. The backward supercardioid pattern can be obtained by applying a delay of 2d / 3c to the signal from the front microphone. Figure 31 illustrates a forward cardioid pattern 3176 according to certain embodiments described herein. The forward cardioid pattern can be obtained by applying a delay of d / c to the signal from the rear microphone. Figure 32 illustrates a rearward cardioid pattern 3276 according to some embodiments described herein. The rearward cardioid pattern can be obtained by applying a delay in the d / c ratio to the signal from the front microphone. Figure 33 illustrates a dipole pattern 3376 according to some embodiments described herein. The dipole pattern can be obtained by not applying a delay to either microphone. It should be understood that other patterns can be generated by applying different delays.

[0221] As described above, in some embodiments, the neural network circuit can be configured to receive multiple audio signals (e.g., audio signal 830), the multiple audio signals including at least one audio signal as a beamformed version of audio signals originating from two or more microphones. In some embodiments, the multiple audio signals may include one beamformed signal. In some embodiments, the multiple audio signals may include two beamformed signals. In some embodiments, the multiple audio signals may include three beamformed signals. In some embodiments, the multiple audio signals may include more than three beamformed signals. In some embodiments, the multiple audio signals may include multiple beamformed signals, each beamformed signal having a different beamformed directivity pattern. In some embodiments, the multiple audio signals may include a signal having a dipole mode. In some embodiments, the multiple audio signals may include a beamformed signal having a forward supercardioid directivity pattern. In some embodiments, the multiple audio signals may include a beamformed signal having a backward supercardioid directivity pattern. In some embodiments, the multiple audio signals may include a beamformed signal having a forward cardioid pattern. In some embodiments, the multiple audio signals may include a beamformed signal having a backward cardioid pattern. In some embodiments, the plurality of audio signals may include beamforming signals with a forward-facing supercardioid pattern and beamforming signals with a backward-facing supercardioid pattern. In some embodiments, the plurality of audio signals may include beamforming signals with a forward-facing supercardioid pattern and beamforming signals with a backward-facing supercardioid pattern. In some embodiments, the plurality of audio signals may include beamforming signals with a forward-facing supercardioid pattern, beamforming signals with a backward-facing supercardioid pattern, and beamforming signals with a dipole pattern. In some embodiments, the plurality of audio signals may include beamforming signals with a forward-facing supercardioid pattern, beamforming signals with a backward-facing supercardioid pattern, and beamforming signals with a dipole pattern.

[0222] As described above, binaural communication can occur in two hearing aids (e.g., two of hearing aids 100), cochlear implants, or headphone systems that communicate with each other via a wireless communication link. Binaural communication can also occur in ear-worn devices, such as eyeglasses with built-in hearing aids (e.g., eyeglasses 300), where a device in the portion of the eyeglasses closest to one ear can communicate with a device in the portion of the eyeglasses closest to the other ear via a wired communication link within the eyeglasses. In some embodiments, binaural communication can facilitate the communication of spatial information. For example, a mask or spectrogram can be transmitted from one device to another. In some embodiments, binaural communication can occur on a low-latency communication link, such as a near-field magnetic induction communication (NFMI) link.

[0223] Any neural network circuit described herein (e.g., neural network circuits 526, 826, 926, 1026, 1126, 1726, and / or 2126) may include circuitry configured to perform operations required to compute the output of a neural network layer. Such operations may be matrix-vector multiplication. In some embodiments, the neural network circuitry may include multiple identical computational tiles, each including multiple multiply-accumulate circuits configured to perform intermediate computations of matrix-vector multiplication in parallel and then compute the results of the intermediate computations into a final result. Each computational tile may additionally include memory configured to store neural network weights, registers configured to store input activation elements, and routing circuitry configured to facilitate state and data communication between computational tiles. Other types of circuitry configured to perform the processing described herein (such as any of masking and subtraction circuits (e.g., masking and subtraction circuits 832, 932, 1032, 1132, and / or 2132), hybrid circuits (e.g., hybrid circuits 834, 1434, 1534, and / or 2134), and / or WDRC circuits (e.g., WDRC circuits 1258 and / or 1458)) can be implemented as digital processing circuitry. In some embodiments, such digital processing circuitry may use a SIMD (single instruction multiple data) architecture. Any ear-worn device described herein (e.g., hearing aid 100, glasses 300, ear-worn device 400, and / or ear-worn device 500) may include a chip implementing certain portions of the circuitry. For example, any noise reduction circuitry described herein (in some embodiments, and others of other types of circuitry) may be implemented (in whole or in part) on a chip. Therefore, the chip may include the aforementioned computing units and digital processing circuitry. In some embodiments, for models with up to 10M 8-bit weights, and when operating on time-series data at 100 billion operations per second (100 GOPs / sec), the chip can achieve a power efficiency of 4 billion operations per milliwatt (4 GOPs / milliwatt), measured under the conditions of 40 degrees Celsius, using a supply voltage between 0.5 and 1.8V, and the chip performing operations without idle time. Further description of the chip (and other elements in some embodiments) for use in an ear-worn device can be found in U.S. Patent No. 11,886,974, entitled “Neural Network Chip for Ear-Worn Device,” issued January 30, 2024, the entire contents of which are incorporated herein by reference.In some embodiments, in addition to such chips including some or all of the noise reduction circuitry, any ear-worn device described herein may include a digital signal processor configured to perform other operations, such as some or all of the processes performed by processing circuitry 522 and / or processing circuitry 528. Example.

[0224] Example 1 relates to an ear-worn device including two or more microphones; and noise reduction circuitry including neural network circuitry, wherein the neural network circuitry is configured to: receive a plurality of audio signals, wherein at least two of the plurality of audio signals are each derived from different microphones of the two or more microphones, and / or at least one of the plurality of audio signals is a beamformed audio signal derived from the two or more microphones; and implement one or more neural network layers trained to perform background noise modification and spatial focusing based on the plurality of audio signals, such that the neural network circuitry generates one or more neural network outputs based on the plurality of audio signals, wherein the noise reduction circuitry is configured to output an output audio signal based on the one or more neural network outputs, the output audio signal including a background noise modified and spatially focused version of a first audio signal among the plurality of audio signals.

[0225] Example 2 relates to the ear-worn device of Example 1, wherein at least two of the plurality of audio signals have different beamforming directional patterns.

[0226] Example 3 relates to an ear-worn device of any one of Examples 1-2, wherein at least one of the plurality of audio signals has a forward beamforming directional pattern and at least one of the plurality of audio signals has a backward beamforming directional pattern.

[0227] Example 4 relates to an ear-worn device of any of Examples 1-3, wherein the output audio signal has a specific spatial focusing pattern, the specific spatial focusing pattern including different weights for speech applications in the first audio signal originating from different directions of arrival relative to the wearer of the ear-worn device.

[0228] Example 5 relates to the ear-worn device of Example 4, wherein the particular spatial focusing mode includes a higher weighting for speech applications originating from the direction of arrival toward the front of the wearer of the ear-worn device than for speech applications originating from the direction of arrival toward the sides and rear of the wearer.

[0229] Example 6 relates to an ear-worn device of any of Examples 3-5, wherein the neural network circuitry is further configured to receive one or more spatial focus control inputs indicating the particular spatial focus mode; and to use the one or more spatial focus control inputs to generate the one or more neural network outputs such that the output audio signal has the particular spatial focus mode.

[0230] Example 7 relates to an ear-worn device of any of Examples 3-6, and further includes communication circuitry configured to receive from a processing device an indication of a user selection of the particular spatial focus mode; and control circuitry configured to generate the one or more spatial focus control inputs indicating the particular spatial focus mode, based at least in part on the indication of the user selection of the particular spatial focus mode.

[0231] Example 8 relates to a system that includes the ear-worn device of Example 7 and a processing device that communicates with the ear-worn device and is configured to display a graphical user interface, the graphical user interface including options for different spatial focus modes; and receiving the user selection for the particular spatial focus mode.

[0232] Example 9 relates to an ear-worn device of any of Examples 1-8, wherein one or more neural network outputs include two or more neural network outputs; and the noise reduction circuit is configured to generate an output audio signal based on the two or more neural network outputs, the output audio signal comprising: a target speech signal including a background noise-modified and spatially focused version of the first audio signal, wherein the target speech signal includes a first spatially focused version of the speech signal and the speech signal includes speech in the first audio signal; an interfering speech signal including a second spatially focused version of the speech signal; and a background noise signal including background noise in the first audio signal; wherein the noise reduction circuit is configured to generate the output audio signal such that in the output audio signal: the volume change of the background noise signal differs from the volume change of the target speech signal by a first volume change difference; the volume change of the interfering speech signal differs from the volume change of the target speech signal by a second volume change difference; and the first volume change difference and the second volume change difference are independently controllable.

[0233] Example 10 relates to the ear-worn device of Example 9, wherein the interfering speech signal includes a remainder when the target speech signal is subtracted from the speech signal.

[0234] Example 11 relates to an ear-worn device of any one of Examples 9-10, wherein the neural network circuitry is configured to: generate a first neural network output of the two or more neural network outputs using a first subset of the one or more neural network layers; and generate a second neural network output of the two or more neural network outputs using a second subset of the one or more neural network layers; and the noise reduction circuitry is configured to obtain the speech signal and / or the background noise signal from the first neural network output of the two or more neural network outputs, and obtain the target speech signal and / or the interfering speech signal from the second neural network output of the two or more neural network outputs.

[0235] Example 12 relates to an ear-worn device of any of Examples 9-11, wherein the outputs of the two or more neural networks include two different masks.

[0236] Example 13 relates to an ear-worn device of any one of Examples 9-12, further comprising: a mixing circuit configured to generate the output audio signal by mixing a combination of audio signals; or by mixing a combination of masks; or a wide dynamic range compression (WDRC) circuit including a plurality of WDRC pipelines configured to generate the output audio signal by performing WDRC on the combination of the audio signals.

[0237] Example 14 relates to the ear-worn device of Example 13, wherein the mixing circuit is further configured to: receive a first volume change control input and a second volume change control input; and perform the mixing using the first volume change control input and the second volume change control input such that the first volume change difference is controlled at least partially by the first volume change control input, and the second volume change difference is controlled at least partially by the second volume change control input.

[0238] Example 15 relates to the ear-worn device of Example 14, and further includes: communication circuitry configured to receive a first volume change control input and a second volume change value from a processing device; a memory configured to store the first volume change control input and the second volume change control input; and control circuitry configured to retrieve the first volume change control input and the second volume change control input from the memory and output the first volume change control input and the second volume change control input to the mixing circuitry.

[0239] Example 16, for an ear-worn device according to any one of Examples 1-12, further includes a mixing circuit configured to mix two or more audio signals such that the output audio signal includes a background noise-modified and spatially focused version of the first audio signal mixed with a second audio signal.

[0240] Example 17 relates to the ear-worn device of Example 16, wherein: a background noise modified and spatially focused version of the first audio signal includes a target speech signal; the target speech signal includes a first spatially focused version of the speech signal; the speech signal includes speech in the first audio signal; the second audio signal includes: a background noise signal that includes background noise in the first audio signal; and an interfering speech signal that includes a second spatially focused version of the speech signal; and the noise reduction circuit is configured to generate the output audio signal such that in the output audio signal: the volume change of the background noise signal differs from the volume change of the target speech signal by a volume change difference; the volume change of the interfering speech signal differs from the volume change of the target speech signal by the volume change difference; and the volume change difference is controllable.

[0241] Example 18 relates to the ear-worn device of Example 17, wherein the mixing circuit is configured to: receive the volume change control input; and use the volume change control input to perform the mixing such that the volume change difference is controlled at least partially by the volume change control input.

[0242] Example 19 relates to an ear-worn device of Example 18, wherein the ear-worn device includes: communication circuitry configured to receive volume change control input from a processing device; a memory configured to store the volume change control input; and control circuitry configured to retrieve the volume change control input and output the volume change control input to the hybrid circuitry.

[0243] Example 20 relates to an ear-worn device of any of Examples 1-19, wherein the ear-worn device is further configured to receive a user selection to turn off spatial focusing.

[0244] Example 21 relates to an ear-worn device including two or more microphones and a noise reduction circuit including a neural network circuit, wherein the neural network circuit is configured to: receive a plurality of audio signals, wherein at least two of the plurality of audio signals each originate from different microphones of the two or more microphones, and / or at least one of the plurality of audio signals is a beamformed audio signal originating from the two or more microphones; and implement one or more neural network layers, the one or more neural network layers being trained to perform background noise modification and spatial focusing, such that the neural network circuit generates two or more neural network outputs based on the plurality of audio signals, wherein: the noise reduction circuit is configured to generate an output audio signal, the output audio... The audio signal includes: a target speech signal, comprising a first spatially focused version of the speech signal, wherein the speech signal includes speech in a first audio signal among the plurality of audio signals; an interfering speech signal, comprising a second spatially focused version of the speech signal; and a background noise signal, comprising background noise in the first audio signal; wherein the noise reduction circuit is configured to generate the output audio signal such that in the output audio signal: the volume change of the background noise signal differs from the volume change of the target speech signal by a first volume change difference; the volume change of the interfering speech signal differs from the volume change of the target speech signal by a second volume change difference; and the first volume change difference and the second volume change difference are independently controllable.

[0245] Example 22 relates to the ear-worn device of Example 21, wherein at least two of the plurality of audio signals have different beamforming directional patterns.

[0246] Example 23 relates to any of the ear-worn devices in Examples 21-22, wherein the target speech signal includes a speech signal to which a specific spatial focusing mode has been applied, the specific spatial focusing mode including different weights for speech applications originating from different directions of arrival relative to the wearer of the ear-worn device.

[0247] Example 24 relates to the ear-worn device of Example 23, wherein the particular spatial focusing mode includes a higher weighting for speech applications originating from the direction of arrival toward the front of the wearer of the ear-worn device than for speech applications originating from the direction of arrival toward the sides and rear of the wearer.

[0248] Example 25 relates to an ear-worn device of any of Examples 23-24, wherein the neural network circuitry is further configured to: receive one or more spatial focus control inputs indicating a particular spatial focus mode; and use the one or more spatial focus control inputs to generate the two or more neural network outputs such that the target speech signal includes the speech signal to which the particular spatial focus mode has been applied.

[0249] Example 26 relates to the ear-worn device of Example 25, and further includes communication circuitry configured to receive from a processing device an indication of a user selection of the particular spatial focus mode; and control circuitry configured to generate the one or more spatial focus control inputs indicating the particular spatial focus mode, based at least in part on the indication of the user selection of the particular spatial focus mode.

[0250] Example 27 relates to a system comprising: the ear-worn device of Example 26; and a processing device that communicates with the ear-worn device and is configured to display a graphical user interface including options for different spatial focusing modes; and to receive the user selection for the particular spatial focusing mode.

[0251] Example 28 relates to an ear-worn device as described in any one of Examples 21-27, wherein the interfering speech signal includes a remainder when the target speech signal is subtracted from the speech signal.

[0252] Example 29 relates to an ear-worn device of any of Examples 21-28, wherein the neural network circuitry is configured to: generate a first neural network output of two or more neural network outputs using a first subset of one or more neural network layers; and generate a second neural network output of the two or more neural network outputs using a second subset of the one or more neural network layers; and the noise reduction circuitry is configured to obtain the speech signal and / or the background noise signal from the first neural network output of the two or more neural network outputs, and obtain the target speech signal and / or the interfering speech signal from the second neural network output of the two or more neural network outputs.

[0253] Example 30 relates to an ear-worn device of any of Examples 21-29, wherein the outputs of the two or more neural networks include two different masks.

[0254] Example 31 relates to an ear-worn device according to any one of Examples 21-30, further comprising: a mixing circuit configured to generate the output audio signal by mixing a combination of audio signals; or by mixing a combination of masks; or a wide dynamic range compression (WDRC) circuit including a plurality of WDRC pipelines configured to generate the output audio signal by performing WDRC on the combination of the audio signals.

[0255] Example 32 relates to the ear-worn device of Example 31, wherein the mixing circuit is further configured to receive a first volume change control input and a second volume change control input; and to perform the mixing using the first volume change control input and the second volume change control input such that the first volume change difference is controlled at least partially by the first volume change control input, and the second volume change difference is controlled at least partially by the second volume change control input.

[0256] Example 33 relates to the ear-worn device of Example 32, and further includes communication circuitry configured to receive a first volume change control input and a second volume change value from a processing device; a memory configured to store the first volume change control input and the second volume change control input; and control circuitry configured to retrieve the first volume change control input and the second volume change control input from the memory and output the first volume change control input and the second volume change control input to the mixing circuitry.

[0257] Example 34 relates to an ear-worn device of any of Examples 32-33, and further includes control circuitry configured to: generate a first volume change control input based on a background noise level in the first audio signal; and generate a second volume change control input based on a level of interfering speech in the first audio signal.

[0258] Example 35 relates to an ear-worn device of any of Examples 21-34, wherein the ear-worn device is further configured to receive a user selection to turn off spatial focusing.

[0259] Example 36 relates to an ear-worn device of any one of Examples 21-35, wherein at least one of the two or more neural network outputs includes: a speech signal; a mask configured to generate the speech signal; a background noise signal; a mask configured to generate the background noise signal; a target speech signal; a mask configured to generate the target speech signal; the interfering speech signal; and a mask configured to generate the interfering speech signal.

[0260] Example 37 relates to an ear-worn device of any of Examples 21-36, wherein the background noise signal is spatially unfocused; and the interfering speech signal does not include a portion of the background noise in the first audio signal.

[0261] Example 38 relates to an ear-worn device of any of Examples 21-36, wherein the background noise signal includes a first spatially focused version of the background noise in the first audio signal; and the interfering speech signal includes a second spatially focused version of the speech signal plus a second spatially focused version of the background noise in the first audio signal.

[0262] Example 39 relates to an ear-worn device of any of Examples 21-38, wherein the ear-worn device includes a hearing aid.

[0263] Example 40 relates to an ear-worn device of any of Examples 21-39, wherein the noise reduction circuitry is implemented on a chip.

[0264] Example 41 relates to an ear-worn device including two or more microphones; and noise reduction circuitry including a neural network circuit, wherein the neural network circuitry is configured to: receive a plurality of audio signals, wherein at least two of the plurality of audio signals are each derived from different microphones of the two or more microphones, and / or at least one of the plurality of audio signals is a beamformed audio signal derived from the two or more microphones; and implement one or more neural network layers trained to perform background noise modification and / or spatial focusing based on the plurality of audio signals, such that the neural network circuitry generates one or more neural network outputs based on the plurality of audio signals, wherein the noise reduction circuitry is configured to output an output audio signal based on the one or more neural network outputs, the output audio signal including a background noise modified and / or spatially focused version of a first audio signal among the plurality of audio signals.

[0265] Example 42 relates to an ear-worn device of Example 41, wherein one or more neural network layers are trained to perform background noise modification, and the output audio signal includes a background noise modified version of the first audio signal.

[0266] Example 43 relates to an ear-worn device of Example 41, wherein the one or more neural network layers are trained to perform spatial focusing, and the output audio signal includes a spatially focused version of the first audio signal.

[0267] Example 44 relates to an ear-worn device of Example 41, wherein the one or more neural network layers are trained to perform background noise modification and spatial focusing, and the output audio signal includes a background noise modified and spatially focused version of the first audio signal.

[0268] Example 45 relates to any of the ear-worn devices in Examples 43-44, wherein the output audio signal has a specific spatial focusing pattern, the specific spatial focusing pattern including different weights for speech applications originating from different directions of arrival relative to the wearer of the ear-worn device.

[0269] Example 46 relates to the ear-worn device of Example 45, wherein the particular spatial focusing mode includes a higher weighting for speech applications originating from the direction of arrival toward the front of the wearer of the ear-worn device than for speech applications originating from the direction of arrival toward the sides and rear of the wearer.

[0270] Example 47 relates to an ear-worn device of any of Examples 45-46, wherein the neural network circuitry is further configured to receive one or more spatial focus control inputs indicating a particular spatial focus mode; and to use the one or more spatial focus control inputs to generate the one or more neural network outputs such that the output audio signal has the particular spatial focus mode.

[0271] Example 48 relates to the ear-worn device of Example 47, and further includes communication circuitry configured to receive from a processing device an indication of a user selection of the particular spatial focus mode; and control circuitry configured to generate the one or more spatial focus control inputs indicating the particular spatial focus mode, based at least in part on the indication of the user selection of the particular spatial focus mode.

[0272] Example 49 relates to a system that includes the ear-worn device of Example 48; and a processing device that communicates with the ear-worn device and is configured to: display a graphical user interface including options for different spatial focusing modes; and receive the user selection for the particular spatial focusing mode.

[0273] Example 50 relates to the system described in Example 49, wherein the multiple options include four options.

[0274] Example 51 relates to a system of any of Examples 49-50, wherein the plurality of options are graphical representations of the different spatial focusing modes.

[0275] Example 52 relates to the ear-worn device of Example 47, and further includes sensing circuitry configured to generate one or more inputs based on movement of the ear-worn device; and control circuitry configured to: determine the degree of head movement based on the one or more inputs received from the one or more sensors; and generate one or more spatial focus control inputs indicating the specific spatial focus mode based on the degree of head movement.

[0276] Example 53 relates to the ear-worn device of Example 52, wherein the control circuitry is configured to, when generating one or more inputs indicating a spatial focus mode based on the degree of head movement: generate a first set of one or more spatial focus control inputs based on a first degree of head movement, the first set of one or more spatial focus control inputs indicating a first spatial focus mode having a first spatial focus amount; and generate a second set of one or more spatial focus control inputs based on a second degree of head movement, the second set of one or more spatial focus control inputs indicating a second spatial focus mode having a second spatial focus amount; wherein the first spatial focus amount is less than the second spatial focus amount, and the first degree of head movement is greater than the second degree of head movement.

[0277] Example 54 relates to the ear-worn device of Example 47 and further includes control circuitry configured to: determine the signal-to-noise ratio (SNR) of an acoustic environment; and generate the one or more spatial focus control inputs indicating the spatial focus mode based on the SNR of the acoustic environment.

[0278] Example 55 relates to an ear-worn device of any one of Examples 43-44, wherein the output audio signal comprises: a target speech signal including a first spatially focused version of the speech signal, the speech signal including speech in the first audio signal; an interfering speech signal including a second spatially focused version of the speech signal; and a background noise signal including background noise in the first audio signal.

[0279] Example 56 relates to the ear-worn device of Example 55, wherein the noise reduction circuit is configured to generate the output audio signal such that: the volume change of the background noise signal differs from the volume change of the target speech signal by a first volume change difference; the volume change of the interfering speech signal differs from the volume change of the target speech signal by a second volume change difference; and the first volume change difference and the second volume change difference are different.

[0280] Example 57 relates to an ear-worn device of any one of Examples 55-56, wherein the noise reduction circuit is configured to generate the output audio signal such that: the volume change of the background noise signal differs from the volume change of the target speech signal by a first volume change difference; the volume change of the interfering speech signal differs from the volume change of the target speech signal by a second volume change difference; and the first volume change difference and the second volume change difference are independently controllable.

[0281] Example 58 relates to an ear-worn device of any of Examples 55-57, wherein the target speech signal includes a speech signal to which a specific spatial focusing mode has been applied, the specific spatial focusing mode including different weights for speech applications originating from different directions of arrival relative to the wearer of the ear-worn device.

[0282] Example 59 relates to the ear-worn device of Example 58, wherein the particular spatial focusing mode includes a higher weighting for speech applications originating from the direction of arrival toward the front of the wearer of the ear-worn device than for speech applications originating from the direction of arrival toward the sides and rear of the wearer.

[0283] Example 60 relates to an ear-worn device of any of Examples 58-59, wherein the neural network circuitry is further configured to receive one or more spatial focus control inputs indicative of the particular spatial focus mode; and to use the one or more spatial focus control inputs to generate the two or more neural network outputs such that the target speech signal includes the speech signal to which the particular spatial focus mode has been applied.

[0284] Example 61 relates to the ear-worn device of Example 60, and further includes: communication circuitry configured to receive from a processing device an indication of a user selection of the particular spatial focus mode; and control circuitry configured to generate the one or more spatial focus control inputs indicating the particular spatial focus mode, based at least in part on the indication of the user selection of the particular spatial focus mode.

[0285] Example 62 relates to a system that includes the ear-worn device of Example 61 and a processing device that communicates with the ear-worn device and is configured to display a graphical user interface including options for different spatial focusing modes; and to receive the user selection for the particular spatial focusing mode.

[0286] Example 63 relates to the system described in Example 62, wherein the multiple options include four options.

[0287] Example 64 relates to a system of any of Examples 62-63, wherein the plurality of options are graphical representations of the different spatial focusing modes.

[0288] Example 65 relates to the ear-worn device of Example 60, and further includes sensing circuitry configured to generate one or more inputs based on movement of the ear-worn device; and control circuitry configured to: determine the degree of head movement based on the one or more inputs received from the one or more sensors; and generate the one or more spatial focus control inputs indicating the specific spatial focus mode based on the degree of head movement.

[0289] Example 66 relates to the ear-worn device of Example 65, wherein the control circuitry is configured to, when generating one or more inputs indicating a spatial focus mode based on the degree of head movement: generate a first set of one or more spatial focus control inputs based on a first degree of head movement, the first set of one or more spatial focus control inputs indicating a first spatial focus mode having a first spatial focus amount; and generate a second set of one or more spatial focus control inputs based on a second degree of head movement, the second set of one or more spatial focus control inputs indicating a second spatial focus mode having a second spatial focus amount; wherein the first spatial focus amount is less than the second spatial focus amount, and the first degree of head movement is greater than the second degree of head movement.

[0290] Example 67 relates to the ear-worn device of Example 60, and further includes: control circuitry configured to: determine the signal-to-noise ratio (SNR) of an acoustic environment; and generate the one or more spatial focus control inputs indicating the spatial focus mode based on the SNR of the acoustic environment.

[0291] Example 68 relates to an ear-worn device according to any one of Examples 55-67, wherein the interfering speech signal includes a remainder when the target speech signal is subtracted from the speech signal.

[0292] Example 69 relates to an ear-worn device of any of Examples 55-67, wherein the one or more neural network outputs include two or more neural network outputs.

[0293] Example 70 relates to an ear-worn device of Example 69, wherein the noise reduction circuitry is configured to obtain, based on the outputs of two or more neural networks, at least one of the following: a speech signal comprising speech in a first audio signal among the plurality of audio signals; and a background noise signal comprising background noise in the first audio signal; and at least one of the following: a target speech signal comprising a first spatially focused version of the speech signal; and an interfering speech signal comprising a second spatially focused version of the speech signal.

[0294] Example 71 relates to an ear-worn device of Example 69, wherein the noise reduction circuit is configured to obtain, based on the outputs of the two or more neural networks, a target speech signal comprising a first spatially focused version of the speech signal and an interfering speech signal comprising a second spatially focused version of the speech signal.

[0295] Example 72 relates to an ear-worn device of any one of Examples 69-71, wherein the neural network circuitry is configured to: generate a first neural network output of the two or more neural network outputs using a first subset of the one or more neural network layers; and generate a second neural network output of the two or more neural network outputs using a second subset of the one or more neural network layers; and the noise reduction circuitry is configured to obtain the speech signal and / or the background noise signal from the first neural network output of the two or more neural network outputs, and obtain the target speech signal and / or the interfering speech signal from the second neural network output of the two or more neural network outputs.

[0296] Example 73 relates to an ear-worn device of any of Examples 69-72, wherein the outputs of the two or more neural networks include two different masks.

[0297] Example 74 relates to an ear-worn device of any one of Examples 69-73, wherein: at least one of the two or more neural network outputs includes: the speech signal; a mask configured to generate the speech signal; the background noise signal; a mask configured to generate the background noise signal; the target speech signal; a mask configured to generate the target speech signal; and the interfering speech signal; a mask configured to generate the interfering speech signal.

[0298] Example 75 relates to an ear-worn device of any one of Examples 55-74, further comprising: a mixing circuit configured to generate the output audio signal by mixing a combination of audio signals; or by mixing a combination of masks; or a wide dynamic range compression (WDRC) circuit including a plurality of WDRC pipelines configured to generate the output audio signal by performing WDRC on the combination of the audio signals.

[0299] Example 76 relates to an ear-worn device of Example 75, wherein the mixing circuitry is further configured to: receive a first volume change control input and a second volume change control input; and perform the mixing using the first volume change control input and the second volume change control input such that the first volume change difference is controlled at least partially by the first volume change control input, and the second volume change difference is controlled at least partially by the second volume change control input.

[0300] Example 77 relates to the ear-worn device of Example 76, and further includes: communication circuitry configured to receive a first volume change control input and a second volume change value from a processing device; a memory configured to store the first volume change control input and the second volume change control input; and control circuitry configured to retrieve the first volume change control input and the second volume change control input from the memory and output the first volume change control input and the second volume change control input to the mixing circuitry.

[0301] Example 78 relates to an ear-worn device of any of Examples 55-74, wherein the neural network circuitry is further configured to: receive a first volume change control input and a second volume change control input; and use the first volume change control input and the second volume change control input to generate the one or more neural network outputs such that the first volume change difference is controlled at least partially by the first volume change control input, and the second volume change difference is controlled at least partially by the second volume change control input.

[0302] Example 79 relates to the ear-worn device of Example 78, and further includes: communication circuitry configured to receive a first volume change control input and a second volume change value from a processing device; a memory configured to store the first volume change control input and the second volume change control input; and control circuitry configured to retrieve the first volume change control input and the second volume change control input from the memory and output the first volume change control input and the second volume change control input to the mixing circuitry.

[0303] Example 80 relates to an ear-worn device of any of Examples 76-79, and further includes: control circuitry configured to: generate a first volume change control input based on the level of background noise in the first audio signal; and generate a second volume change control input based on the level of interfering speech in the first audio signal.

[0304] Example 81 relates to an ear-worn device of any of Examples 55-80, wherein: the background noise signal is spatially unfocused; and the interfering speech signal does not include a portion of the background noise in the first audio signal.

[0305] Example 82 relates to an ear-worn device of any of Examples 55-80, wherein the background noise signal includes a first spatially focused version of the background noise in the first audio signal; and the interfering speech signal includes a second spatially focused version of the speech signal plus a second spatially focused version of the background noise in the first audio signal.

[0306] Example 83 is an ear-worn device as described in any one of Examples 43-82, wherein the ear-worn device is further configured to receive a user selection to turn off spatial focusing.

[0307] Example 84 relates to an ear-worn device of any of Examples 41-83, wherein at least two of the plurality of audio signals have different beamforming directional patterns.

[0308] Example 85 relates to the ear-worn device of Example 84, wherein at least two of the plurality of audio signals having different beamforming directional patterns include beamforming signals having dipole, supercardioid, supercardioid, or cardioid directional patterns.

[0309] Example 86 relates to an ear-worn device of any of Examples 41-74 and 78-85, and further includes mixing circuitry configured to mix two or more audio signals such that the output audio signal includes a background noise modified and / or spatially focused version of the first audio signal mixed with a second audio signal.

[0310] Example 87 relates to the ear-worn device of Example 86, wherein: the background noise modified and spatially focused version of the first audio signal includes a target speech signal; the target speech signal includes a first spatially focused version of the speech signal; the speech signal includes speech in the first audio signal; the second audio signal includes: a background noise signal that includes background noise in the first audio signal; and an interfering speech signal that includes a second spatially focused version of the speech signal; and the noise reduction circuit is configured to generate the output audio signal such that in the output audio signal: the volume change of the background noise signal differs from the volume change of the target speech signal by a volume change difference; the volume change of the interfering speech signal differs from the volume change of the target speech signal by the volume change difference; and the volume change difference is controllable.

[0311] Example 88 relates to an ear-worn device of Example 87, wherein: the mixing circuit is configured to: receive the volume change control input; and use the volume change control input to perform the mixing such that the volume change difference is controlled at least partially by the volume change control input; and the ear-worn device includes: communication circuitry configured to receive the volume change control input from a processing device; a memory configured to store the volume change control input; and control circuitry configured to retrieve the volume change control input and output the volume change control input to the mixing circuitry.

[0312] Example 89 relates to an ear-worn device of any of Examples 41-88, wherein the ear-worn device includes a hearing aid.

[0313] Example 90 relates to an ear-worn device according to any one of Examples 41-89, wherein the noise reduction circuitry is implemented on a chip.

[0314] Example 91 relates to an ear-worn device comprising: two or more microphones; and noise reduction circuitry including a neural network circuit, wherein: the neural network circuitry is configured to: receive a plurality of audio signals, wherein at least two of the plurality of audio signals are each derived from different microphones of the two or more microphones, and / or at least one of the plurality of audio signals is a beamformed audio signal derived from the two or more microphones; and implement one or more neural network layers trained to perform or generate outputs for performing background noise modification and spatial focusing based on the plurality of audio signals, such that the neural network circuitry generates one or more neural network outputs based on the plurality of audio signals, wherein: the noise reduction circuitry is configured to output an output audio signal based on the one or more neural network outputs, the output audio signal including a background noise modified and spatially focused version of a first audio signal of the plurality of audio signals.

[0315] Example 92 relates to an ear-worn device of Example 91, wherein at least two of the plurality of audio signals have different beamforming directional patterns.

[0316] Example 93 relates to an ear-worn device of any of Examples 91-92, wherein at least one of the plurality of audio signals has a forward beamforming directional pattern and at least one of the plurality of audio signals has a backward beamforming directional pattern.

[0317] Example 94 relates to an ear-worn device of any of Examples 91-93, wherein the output audio signal has a specific spatial focusing pattern, the specific spatial focusing pattern including different weights for speech applications in the first audio signal originating from different directions of arrival relative to the wearer of the ear-worn device.

[0318] Example 95 relates to an ear-worn device of Example 94, wherein the particular spatial focusing mode includes a higher weighting for speech applications originating from the direction of arrival toward the front of the wearer of the ear-worn device than for speech applications originating from the direction of arrival toward the sides and rear of the wearer.

[0319] Example 96 relates to an ear-worn device of any of Examples 93-95, wherein the neural network circuitry is further configured to: receive one or more spatial focus control inputs indicating the particular spatial focus mode; and use the one or more spatial focus control inputs to generate the one or more neural network outputs such that the output audio signal has the particular spatial focus mode.

[0320] Example 97 relates to an ear-worn device of any of Examples 93-96, further comprising: communication circuitry configured to receive from a processing device an indication of a user selection of the particular spatial focus mode; and control circuitry configured to generate the one or more spatial focus control inputs indicating the particular spatial focus mode, based at least in part on the indication of the user selection of the particular spatial focus mode.

[0321] Example 98 relates to a system comprising: an ear-worn device of Example 97; and a processing device that communicates with the ear-worn device and is configured to: display a graphical user interface including options for different spatial focus modes; and receive the user selection for the particular spatial focus mode.

[0322] Example 99 relates to an ear-worn device of any of Examples 91-98, wherein: the one or more neural network outputs include two or more neural network outputs; and the noise reduction circuit is configured to generate an output audio signal based on the two or more neural network outputs, the output audio signal including: a target speech signal including a background noise modified and spatially focused version of the first audio signal, wherein the target speech signal includes a first spatially focused version of the speech signal, and the speech signal includes speech in the first audio signal; an interfering speech signal including a second spatially focused version of the speech signal; and a background noise signal including background noise in the first audio signal; wherein the noise reduction circuit is configured to generate the output audio signal such that in the output audio signal: the volume change of the background noise signal differs from the volume change of the target speech signal by a first volume change difference; the volume change of the interfering speech signal differs from the volume change of the target speech signal by a second volume change difference; and the first volume change difference and the second volume change difference are independently controllable.

[0323] Example 100 relates to the ear-worn device of Example 99, wherein the interfering speech signal includes a remainder when the target speech signal is subtracted from the speech signal.

[0324] Example 101 relates to an ear-worn device of any of Examples 99-100, wherein the neural network circuitry is configured to: generate a first neural network output of the two or more neural network outputs using a first subset of the one or more neural network layers; and generate a second neural network output of the two or more neural network outputs using a second subset of the one or more neural network layers; and the noise reduction circuitry is configured to obtain the speech signal and / or the background noise signal from the first neural network output of the two or more neural network outputs, and obtain the target speech signal and / or the interfering speech signal from the second neural network output of the two or more neural network outputs.

[0325] Example 102 relates to an ear-worn device of any of Examples 99-101, wherein the outputs of the two or more neural networks include two different masks.

[0326] Example 103 relates to an ear-worn device of any of Examples 99-102, and further includes: a mixing circuit configured to generate the output audio signal by mixing a combination of audio signals; or by mixing a combination of masks; or a wide dynamic range compression (WDRC) circuit including a plurality of WDRC pipelines configured to generate the output audio signal by performing WDRC on the combination of the audio signals.

[0327] Example 104 relates to an ear-worn device of Example 103, wherein the mixing circuit is further configured to: receive a first volume change control input and a second volume change control input; and perform the mixing using the first volume change control input and the second volume change control input such that the first volume change difference is controlled at least partially by the first volume change control input, and the second volume change difference is controlled at least partially by the second volume change control input.

[0328] Example 105 relates to the ear-worn device of Example 104, and further includes: communication circuitry configured to receive a first volume change control input and a second volume change value from a processing device; a memory configured to store the first volume change control input and the second volume change control input; and control circuitry configured to retrieve the first volume change control input and the second volume change control input from the memory and output the first volume change control input and the second volume change control input to the mixing circuitry.

[0329] Example 106 relates to an ear-worn device of any of Examples 91-102, and further includes a mixing circuit configured to mix two or more audio signals such that the output audio signal includes a background noise-modified and spatially focused version of the first audio signal mixed with a second audio signal.

[0330] Example 107 relates to an ear-worn device of Example 106, wherein: a background noise modified and spatially focused version of the first audio signal includes a target speech signal; the target speech signal includes a first spatially focused version of the speech signal; the speech signal includes speech in the first audio signal; the second audio signal includes: a background noise signal that includes background noise in the first audio signal; and an interfering speech signal that includes a second spatially focused version of the speech signal; and the noise reduction circuit is configured to generate the output audio signal such that in the output audio signal: the volume change of the background noise signal differs from the volume change of the target speech signal by a volume change difference; the volume change of the interfering speech signal differs from the volume change of the target speech signal by the volume change difference; and the volume change difference is controllable.

[0331] Example 108 relates to an ear-worn device of Example 107, wherein: the mixing circuit is configured to: receive the volume change control input; and use the volume change control input to perform the mixing such that the volume change difference is controlled at least partially by the volume change control input.

[0332] Example 109 relates to an ear-worn device of Example 108, wherein the ear-worn device includes: communication circuitry configured to receive the volume change control input from a processing device; a memory configured to store the volume change control input; and control circuitry configured to retrieve the volume change control input and output the volume change control input to the mixing circuitry.

[0333] Example 110 relates to an ear-worn device of any of Examples 91-109, wherein the ear-worn device is further configured to receive a user selection to turn off spatial focusing.

[0334] Example 111 relates to an ear-worn device comprising: two or more microphones; noise reduction circuitry including neural network circuitry, wherein: the neural network circuitry is configured to: receive a plurality of audio signals, wherein at least two of the plurality of audio signals are each derived from different microphones of the two or more microphones, and / or at least one of the plurality of audio signals is a beamformed audio signal derived from the two or more microphones; and implement one or more neural network layers trained to generate one or more neural network outputs based on the plurality of audio signals, wherein: the one or more neural network outputs include an output audio signal, the output audio signal including a background noise-modified and spatially focused version of a first audio signal among the plurality of audio signals; or the one or more neural network outputs are configured to be used by the noise reduction circuitry to generate the output audio signal including the background noise-modified and spatially focused version of the first audio signal.

[0335] Example 112 relates to an ear-worn device of Example 111, wherein at least two of the plurality of audio signals have different beamforming directional patterns.

[0336] Example 113 relates to an ear-worn device of any of Examples 111-112, wherein at least one of the plurality of audio signals has a forward beamforming directional pattern and at least one of the plurality of audio signals has a backward beamforming directional pattern.

[0337] Example 114 relates to an ear-worn device of any of Examples 111-113, wherein the output audio signal has a specific spatial focusing pattern, the specific spatial focusing pattern including different weights on the first audio signal for speech applications originating from different directions of arrival relative to the wearer of the ear-worn device.

[0338] Example 115 relates to an ear-worn device of Example 114, wherein the particular spatial focusing mode includes a higher weighting for speech applications originating from the direction of arrival toward the front of the wearer of the ear-worn device than for speech applications originating from the direction of arrival toward the sides and rear of the wearer.

[0339] Example 116 relates to an ear-worn device of any of Examples 113-115, wherein the neural network circuitry is further configured to: receive one or more spatial focus control inputs indicating the particular spatial focus mode; and use the one or more spatial focus control inputs to generate the one or more neural network outputs such that the output audio signal has the particular spatial focus mode.

[0340] Example 117 relates to an ear-worn device of any one of Examples 113-116, further comprising: communication circuitry configured to receive from a processing device an indication of a user selection of the particular spatial focus mode; and control circuitry configured to generate the one or more spatial focus control inputs indicating the particular spatial focus mode, based at least in part on the indication of the user selection of the particular spatial focus mode.

[0341] Example 118 relates to a system comprising: an ear-worn device of Example 117; and a processing device that communicates with the ear-worn device and is configured to: display a graphical user interface including options for different spatial focusing modes; and receive the user selection for the particular spatial focusing mode.

[0342] Example 119 relates to an ear-worn device of any one of Examples 111-118, wherein: the one or more neural network outputs include two or more neural network outputs; and the noise reduction circuit is configured to generate an output audio signal based on the two or more neural network outputs, the output audio signal including: a target speech signal including a background noise modified and spatially focused version of the first audio signal, wherein the target speech signal includes a first spatially focused version of the speech signal, and the speech signal includes speech in the first audio signal; an interfering speech signal including a second spatially focused version of the speech signal; and a background noise signal including background noise in the first audio signal; wherein the noise reduction circuit is configured to generate the output audio signal such that in the output audio signal: the volume change of the background noise signal differs from the volume change of the target speech signal by a first volume change difference; the volume change of the interfering speech signal differs from the volume change of the target speech signal by a second volume change difference; and the first volume change difference and the second volume change difference are independently controllable.

[0343] Example 120 relates to the ear-worn device of Example 119, wherein the interfering speech signal includes a remainder when the target speech signal is subtracted from the speech signal.

[0344] Example 121 relates to an ear-worn device as described in any one of Examples 119-120, wherein: the neural network circuitry is configured to: generate a first neural network output of the two or more neural network outputs using a first subset of the one or more neural network layers; and generate a second neural network output of the two or more neural network outputs using a second subset of the one or more neural network layers; and the noise reduction circuitry is configured to obtain the speech signal and / or the background noise signal from the first neural network output of the two or more neural network outputs, and obtain the target speech signal and / or the interfering speech signal from the second neural network output of the two or more neural network outputs.

[0345] Example 122 relates to an ear-worn device of any of Examples 119-121, wherein the outputs of two or more neural networks include two different masks.

[0346] Example 123, for any ear-worn device as described in any of Examples 119-122, further includes: a mixing circuit configured to generate the output audio signal by mixing a combination of audio signals; or by mixing a combination of masks; or a wide dynamic range compression (WDRC) circuit including a plurality of WDRC pipelines configured to generate the output audio signal by performing WDRC on a combination of audio signals.

[0347] Example 124 relates to the ear-worn device of Example 123, wherein the mixing circuit is further configured to: receive a first volume change control input and a second volume change control input; and perform the mixing using the first volume change control input and the second volume change control input such that the first volume change difference is controlled at least partially by the first volume change control input, and the second volume change difference is controlled at least partially by the second volume change control input.

[0348] Example 125 relates to the ear-worn device of Example 124, and further includes: communication circuitry configured to receive a first volume change control input and a second volume change value from a processing device; a memory configured to store the first volume change control input and the second volume change control input; and control circuitry configured to retrieve the first volume change control input and the second volume change control input from the memory and output the first volume change control input and the second volume change control input to the mixing circuitry.

[0349] Example 126 relates to an ear-worn device of any of Examples 111-122, and further includes a mixing circuit configured to mix two or more audio signals such that the output audio signal includes a background noise-modified and spatially focused version of the first audio signal mixed with a second audio signal.

[0350] Example 127 relates to an ear-worn device of Example 126, wherein: a background noise modified and spatially focused version of the first audio signal includes a target speech signal; the target speech signal includes a first spatially focused version of the speech signal; the speech signal includes speech in the first audio signal; the second audio signal includes: a background noise signal that includes background noise in the first audio signal; and an interfering speech signal that includes a second spatially focused version of the speech signal; and the noise reduction circuit is configured to generate the output audio signal such that in the output audio signal: the volume change of the background noise signal differs from the volume change of the target speech signal by a volume change difference; the volume change of the interfering speech signal differs from the volume change of the target speech signal by the volume change difference; and the volume change difference is controllable.

[0351] Example 128 relates to an ear-worn device of Example 127, wherein: the mixing circuit is configured to: receive the volume change control input; and use the volume change control input to perform the mixing such that the volume change difference is controlled at least partially by the volume change control input.

[0352] Example 129 relates to an ear-worn device of Example 128, wherein the ear-worn device includes: communication circuitry configured to receive the volume change control input from a processing device; a memory configured to store the volume change control input; and control circuitry configured to retrieve the volume change control input and output the volume change control input to the mixing circuitry.

[0353] Example 130 relates to any of the ear-worn devices in Examples 111-129, wherein the ear-worn device is also configured to receive a user selection to turn off spatial focusing.

[0354] Several embodiments of these technologies have been described in detail, and various modifications and improvements will readily occur to those skilled in the art. Such modifications and improvements are intended to fall within the spirit and scope of the invention. Therefore, the above description is merely exemplary and not intended to be limiting. For example, any of the components described above may include hardware, software, or a combination of hardware and software.

[0355] Unless otherwise expressly stated otherwise, the indefinite articles “a” and “an” as used herein in the specification and claims shall be understood to mean “at least one”.

[0356] As used herein in the specification and claims, the phrase “and / or” should be understood to mean “any one or both” of the elements so combined, i.e., elements that are combined in some cases and separate in others. Multiple elements listed with “and / or” should be interpreted in the same way, i.e., “one or more” of the elements so combined. Optionally, other elements may exist besides those specifically identified by the “and / or” clause, whether related to or unrelated to those specifically identified.

[0357] As used herein in the specification and claims, the phrase "at least one" referring to a list of one or more elements should be understood to mean at least one element selected from any one or more elements in the list, but not necessarily including at least one of each element specifically listed in the list, and does not exclude any combination of elements in the list. This definition also allows for the optional presence of elements, whether related to or unrelated to those specifically specified elements, in addition to those specifically designated in the list of elements referred to by the phrase "at least one".

[0358] In some embodiments, the terms "about" and "approximately" may be used to refer to within ±20% of the target value, within ±10% of the target value in some embodiments, within ±5% of the target value in some embodiments, and within ±2% of the target value in some embodiments. The terms "about" and "approximately" may include the target value.

[0359] Furthermore, the wording and terminology used herein are for descriptive purposes and should not be considered restrictive. The use of “including,” “comprising,” “having,” “containing,” “involving,” and variations thereof is intended to cover the items listed thereafter and their equivalents, as well as other items.

[0360] Several aspects of at least one embodiment have been described above. It should be understood that various changes, modifications, and improvements will readily occur to those skilled in the art. Such changes, modifications, and improvements are intended for the purposes of this disclosure. Therefore, the foregoing description and figures are merely examples.

Claims

1. An ear-worn device, comprising: Two or more microphones; The noise reduction circuit includes a neural network circuit, wherein: the neural network circuit is configured to: receive a plurality of audio signals, wherein at least two of the plurality of audio signals are each derived from different microphones of the two or more microphones, and / or at least one of the plurality of audio signals is a beamformed audio signal derived from the two or more microphones; and implement one or more neural network layers trained to perform background noise modification and spatial focusing based on the plurality of audio signals, such that the neural network circuit generates one or more neural network outputs based on the plurality of audio signals, wherein: the noise reduction circuit is configured to output an output audio signal based on the one or more neural network outputs, the output audio signal including a background noise modified and spatially focused version of a first audio signal among the plurality of audio signals.

2. The ear-worn device according to claim 1, wherein, At least two of the plurality of audio signals have different beamforming directional patterns.

3. The ear-worn device according to any one of claims 1-2, wherein, At least one of the plurality of audio signals has a forward beamforming directional pattern, and at least one of the plurality of audio signals has a backward beamforming directional pattern.

4. The ear-worn device according to any one of claims 1-3, wherein: The output audio signal has a specific spatial focusing pattern, which includes different weights for voice applications in the first audio signal originating from different directions of arrival relative to the wearer of the ear-worn device.

5. The ear-worn device according to claim 4, wherein, The specific spatial focusing mode includes a higher weighting for speech applications originating from the direction of arrival toward the wearer of the ear-worn device than for speech applications originating from the direction of arrival toward the side and rear of the wearer.

6. The ear-worn device according to any one of claims 3-5, wherein, The neural network circuit is also configured to receive one or more spatial focus control inputs indicating the specific spatial focus mode; And using the one or more spatial focus control inputs to generate the one or more neural network outputs, such that the output audio signal has the specific spatial focus pattern.

7. The ear-worn device according to any one of claims 3-6, further comprising: A communication circuit is configured to receive from a processing device an indication of a user selection of the specific spatial focusing mode; And control circuitry configured to generate one or more spatial focus control inputs that indicate the particular spatial focus mode, at least in part based on a user selection of the particular spatial focus mode.

8. A system comprising: The ear-worn device as described in claim 7; The processing device, which communicates with the ear-worn device and is configured to: display a graphical user interface including options for different spatial focus modes; and receive a user selection for the specific spatial focus mode.

9. The ear-worn device according to any one of claims 1-8, wherein: The one or more neural network outputs include two or more neural network outputs; The noise reduction circuit is configured to generate an output audio signal based on the outputs of the two or more neural networks. The output audio signal includes: a target speech signal comprising a background noise-modified and spatially focused version of the first audio signal, wherein the target speech signal comprises a first spatially focused version of the speech signal, and the speech signal comprises speech from the first audio signal; an interfering speech signal comprising a second spatially focused version of the speech signal; and a background noise signal comprising background noise from the first audio signal. The noise reduction circuit is configured to generate the output audio signal such that in the output audio signal: the volume change of the background noise signal differs from the volume change of the target speech signal by a first volume change difference; the volume change of the interfering speech signal differs from the volume change of the target speech signal by a second volume change difference; and the first volume change difference and the second volume change difference are independently controllable.

10. The ear-worn device according to claim 9, wherein, The interfering speech signal includes the remainder when the target speech signal is subtracted from the speech signal.

11. The ear-worn device according to any one of claims 9-10, wherein: The neural network circuit is configured to: use a first subset of the one or more neural network layers to generate a first neural network output among the two or more neural network outputs; and use a second subset of the one or more neural network layers to generate a second neural network output among the two or more neural network outputs; and the noise reduction circuit is configured to obtain the speech signal and / or the background noise signal from the first neural network output among the two or more neural network outputs, and to obtain the target speech signal and / or the interfering speech signal from the second neural network output among the two or more neural network outputs.

12. The ear-worn device according to any one of claims 9-11, wherein, The outputs of the two or more neural networks include two different masks.

13. The ear-worn device according to any one of claims 9-12, further comprising: A mixing circuit configured to generate the output audio signal by mixing a combination of audio signals; Alternatively, the output audio signal can be generated by mixing combinations of masks; Alternatively, a wide dynamic range compression (WDRC) circuit may be used, comprising multiple WDRC pipelines configured to generate the output audio signal by performing WDRC on a combination of the audio signals.

14. The ear-worn device according to claim 13, wherein, The mixing circuit is further configured to: receive a first volume change control input and a second volume change control input; and use the first volume change control input and the second volume change control input to perform the mixing such that the first volume change difference is controlled at least partially by the first volume change control input, and the second volume change difference is controlled at least partially by the second volume change control input.

15. The ear-worn device according to claim 14, further comprising: A communication circuit configured to receive the first volume change control input and the second volume control change value from a processing device; A memory configured to store the first volume change control input and the second volume change control input; and a control circuit configured to retrieve the first volume change control input and the second volume change control input from the memory and output the first volume change control input and the second volume change control input to the mixing circuit.

16. The ear-worn device according to any one of claims 1-12, further comprising a mixing circuit configured to mix two or more audio signals such that the output audio signal includes a background noise-modified and spatially focused version of the first audio signal mixed with a second audio signal.

17. The ear-worn device according to claim 16, wherein: The first audio signal with modified background noise and spatial focus includes the target speech signal; the target speech signal includes a first spatial focus version of the speech signal. The speech signal includes the speech in the first audio signal; The second audio signal includes: a background noise signal, which includes the background noise in the first audio signal; and an interfering speech signal, which includes a second spatially focused version of the speech signal; and the noise reduction circuit is configured to generate the output audio signal such that in the output audio signal: the volume change of the background noise signal differs from the volume change of the target speech signal by a volume change difference; the volume change of the interfering speech signal differs from the volume change of the target speech signal by the volume change difference; and the volume change difference is controllable.

18. The ear-worn device according to claim 17, wherein: The mixing circuit is configured to: receive a volume change control input; and use the volume change control input to perform the mixing such that the volume change difference is controlled at least partially by the volume change control input.

19. The ear-worn device according to claim 18, wherein, The ear-worn device includes: a communication circuit configured to receive the volume change control input from a processing device; a memory configured to store the volume change control input; and a control circuit configured to retrieve the volume change control input and output the volume change control input to the hybrid circuit.

20. The ear-worn device according to any one of claims 1-19, wherein, The ear-worn device is also configured to receive a user selection to turn off spatial focusing.

21. An ear-worn device, comprising: Two or more microphones; The noise reduction circuit includes a neural network circuit, wherein: the neural network circuit is configured to: receive a plurality of audio signals, wherein at least two of the plurality of audio signals are each derived from different microphones of the two or more microphones, and / or at least one of the plurality of audio signals is a beamformed audio signal derived from the two or more microphones; and implement one or more neural network layers trained to perform background noise modification and spatial focusing, such that the neural network circuit generates two or more neural network outputs based on the plurality of audio signals, wherein: the noise reduction circuit is configured to generate an output audio signal, the output audio signal including: a target speech signal, It includes a first spatially focused version of a speech signal, wherein the speech signal includes speech in a first audio signal among a plurality of audio signals; an interfering speech signal, which includes a second spatially focused version of the speech signal; and a background noise signal, which includes background noise in the first audio signal; wherein the noise reduction circuit is configured to generate the output audio signal such that in the output audio signal: the volume change of the background noise signal differs from the volume change of the target speech signal by a first volume change difference; the volume change of the interfering speech signal differs from the volume change of the target speech signal by a second volume change difference; and the first volume change difference and the second volume change difference are independently controllable.

22. The ear-worn device according to claim 21, wherein, At least two of the plurality of audio signals have different beamforming directional patterns.

23. The ear-worn device according to any one of claims 21-22, wherein, The target speech signal includes a speech signal that has been applied with a specific spatial focusing mode, which includes different weights for speech applications originating from different directions of arrival relative to the wearer of the ear-worn device.

24. The ear-worn device according to claim 23, wherein, The specific spatial focusing mode includes a higher weighting for speech applications originating from the direction of arrival toward the wearer of the ear-worn device than for speech applications originating from the direction of arrival toward the side and rear of the wearer.

25. The ear-worn device according to any one of claims 23-24, wherein, The neural network circuit is also configured to receive one or more spatial focus control inputs indicating the specific spatial focus mode; And using the one or more spatial focus control inputs to generate the two or more neural network outputs, such that the target speech signal includes a speech signal to which the specific spatial focus pattern has been applied.

26. The ear-worn device according to claim 25, further comprising: A communication circuit is configured to receive from a processing device an indication of a user selection of the specific spatial focusing mode; And control circuitry configured to generate one or more spatial focus control inputs that indicate the particular spatial focus mode, at least in part based on a user selection of the particular spatial focus mode.

27. A system comprising: The ear-worn device of claim 26; and the processing device, which communicates with the ear-worn device and is configured to: display a graphical user interface including options for different spatial focus modes; and receive a user selection for the specific spatial focus mode.

28. The ear-worn device according to any one of claims 21-27, wherein, The interfering speech signal includes the remainder when the target speech signal is subtracted from the speech signal.

29. The ear-worn device according to any one of claims 21-28, wherein: The neural network circuit is configured to: use a first subset of the one or more neural network layers to generate a first neural network output among the two or more neural network outputs; and use a second subset of the one or more neural network layers to generate a second neural network output among the two or more neural network outputs; and the noise reduction circuit is configured to obtain the speech signal and / or the background noise signal from the first neural network output among the two or more neural network outputs, and to obtain the target speech signal and / or the interfering speech signal from the second neural network output among the two or more neural network outputs.

30. The ear-worn device according to any one of claims 21-29, wherein, The outputs of the two or more neural networks include two different masks.

31. The ear-worn device according to any one of claims 21-30, further comprising: A mixing circuit configured to generate the output audio signal by mixing a combination of audio signals; Alternatively, the output audio signal can be generated by mixing combinations of masks; Alternatively, a wide dynamic range compression (WDRC) circuit may be used, comprising multiple WDRC pipelines configured to generate the output audio signal by performing WDRC on a combination of the audio signals.

32. The ear-worn device according to claim 31, wherein, The mixing circuit is further configured to: receive a first volume change control input and a second volume change control input; and use the first volume change control input and the second volume change control input to perform the mixing such that the first volume change difference is controlled at least partially by the first volume change control input, and the second volume change difference is controlled at least partially by the second volume change control input.

33. The ear-worn device according to claim 32, further comprising: A communication circuit configured to receive the first volume change control input and the second volume control change value from a processing device; A memory configured to store the first volume change control input and the second volume change control input; and a control circuit configured to retrieve the first volume change control input and the second volume change control input from the memory and output the first volume change control input and the second volume change control input to the mixing circuit.

34. The ear-worn device according to any one of claims 32-33, further comprising: A control circuit is configured to generate the first volume change control input based on the level of background noise in the first audio signal. And generate the second volume change control input based on the level of interfering speech in the first audio signal.

35. The ear-worn device according to any one of claims 21-34, wherein, The ear-worn device is also configured to receive a user selection to turn off spatial focusing.

36. The ear-worn device according to any one of claims 21-35, wherein: At least one of the two or more neural network outputs includes: the speech signal; a mask configured to generate the speech signal; the background noise signal; a mask configured to generate the background noise signal; the target speech signal; a mask configured to generate the target speech signal; and the interfering speech signal; a mask configured to generate the interfering speech signal.

37. The ear-worn device according to any one of claims 21-36, wherein: The background noise signal is not spatially focused; and the interfering speech signal does not include a portion of the background noise in the first audio signal.

38. The ear-worn device according to any one of claims 21-36, wherein: The background noise signal includes a first spatially focused version of the background noise in the first audio signal; The interfering speech signal includes a second spatially focused version of the speech signal plus a second spatially focused version of the background noise in the first audio signal.

39. The ear-worn device according to any one of claims 21-38, wherein, The ear-worn device includes a hearing aid.

40. The ear-worn device according to any one of claims 21-39, wherein, The noise reduction circuit is implemented on a chip.

41. An ear-worn device, comprising: Two or more microphones; A noise reduction circuit includes a neural network circuit, wherein: the neural network circuit is configured to: receive a plurality of audio signals, wherein at least two of the plurality of audio signals are each derived from different microphones of two or more microphones, and / or at least one of the plurality of audio signals is a beamformed audio signal derived from the two or more microphones; and implement one or more neural network layers, the one or more neural network layers being trained to perform background noise modification and / or spatial focusing based on the plurality of audio signals, such that the neural network circuit generates one or more neural network outputs based on the plurality of audio signals, wherein: the noise reduction circuit is configured to output an output audio signal based on the one or more neural network outputs, the output audio signal including a background noise modified and / or spatially focused version of a first audio signal among the plurality of audio signals.

42. The ear-worn device according to claim 41, wherein, The one or more neural network layers are trained to perform background noise modification, and the output audio signal includes a background noise modified version of the first audio signal.

43. The ear-worn device according to claim 41, wherein, The one or more neural network layers are trained to perform spatial focusing, and the output audio signal includes a spatially focused version of the first audio signal.

44. The ear-worn device according to claim 41, wherein, The one or more neural network layers are trained to perform background noise modification and spatial focusing, and the output audio signal includes a background noise modified and spatially focused version of the first audio signal.

45. The ear-worn device according to any one of claims 43-44, wherein, The output audio signal has a specific spatial focusing pattern, which includes different weights for voice applications originating from different directions of arrival relative to the wearer of the ear-worn device.

46. ​​The ear-worn device according to claim 45, wherein, The specific spatial focusing mode includes a higher weighting for speech applications originating from the direction of arrival toward the wearer of the ear-worn device than for speech applications originating from the direction of arrival toward the side and rear of the wearer.

47. The ear-worn device according to any one of claims 45-46, wherein, The neural network circuit is also configured to receive one or more spatial focus control inputs indicating the specific spatial focus mode; And using the one or more spatial focus control inputs to generate the one or more neural network outputs, such that the output audio signal has the specific spatial focus pattern.

48. The ear-worn device according to claim 47, further comprising: A communication circuit is configured to receive from a processing device an indication of a user selection of the specific spatial focusing mode; And control circuitry configured to generate one or more spatial focus control inputs that indicate the particular spatial focus mode, at least in part based on a user selection of the particular spatial focus mode.

49. A system comprising: The ear-worn device as described in claim 48; The processing device, which communicates with the ear-worn device and is configured to: display a graphical user interface including options for different spatial focus modes; and receive a user selection for the specific spatial focus mode.

50. The system according to claim 49, wherein, Multiple options include four options.

51. The system according to any one of claims 49-50, wherein, The multiple options are graphical representations of the different spatial focusing modes.

52. The ear-worn device according to claim 47, further comprising: A sensing circuit configured to generate one or more inputs based on the movement of the ear-worn device; And control circuitry configured to: determine the degree of head movement based on one or more inputs received from one or more sensors; and generate one or more spatial focus control inputs indicating the specific spatial focus mode based on the degree of head movement.

53. The ear-worn device according to claim 52, wherein, The control circuit is configured to, when generating one or more inputs indicating a spatial focus mode based on the degree of head movement: generate a first set of one or more spatial focus control inputs based on a first degree of head movement, the first set of one or more spatial focus control inputs indicating a first spatial focus mode having a first spatial focus amount; And based on the second head movement degree, generate a second set of one or more spatial focus control inputs, the second set of one or more spatial focus control inputs indicating a second spatial focus mode with a second spatial focus amount; wherein the first spatial focus amount is less than the second spatial focus amount, and the first head movement degree is greater than the second head movement degree.

54. The ear-worn device according to claim 47, further comprising: A control circuit is configured to: determine the signal-to-noise ratio (SNR) of the acoustic environment; and generate one or more spatial focus control inputs indicating the spatial focus mode based on the SNR of the acoustic environment.

55. The ear-worn device according to any one of claims 43-44, wherein, The output audio signal includes: a target speech signal, which includes a first spatially focused version of the speech signal, the speech signal including the speech in the first audio signal; an interfering speech signal, which includes a second spatially focused version of the speech signal; and a background noise signal, which includes background noise in the first audio signal.

56. The ear-worn device according to claim 55, wherein, The noise reduction circuit is configured to generate the output audio signal such that: the volume change of the background noise signal differs from the volume change of the target speech signal by a first volume change difference; the volume change of the interfering speech signal differs from the volume change of the target speech signal by a second volume change difference; and the first volume change difference and the second volume change difference are different.

57. The ear-worn device according to any one of claims 55-56, wherein, The noise reduction circuit is configured to generate the output audio signal such that: the volume change of the background noise signal differs from the volume change of the target speech signal by a first volume change difference; the volume change of the interfering speech signal differs from the volume change of the target speech signal by a second volume change difference; and the first volume change difference and the second volume change difference are independently controllable.

58. The ear-worn device according to any one of claims 55-57, wherein, The target speech signal includes a speech signal that has been applied with a specific spatial focusing mode, which includes different weights for speech applications originating from different directions of arrival relative to the wearer of the ear-worn device.

59. The ear-worn device according to claim 58, wherein, The specific spatial focusing mode includes a higher weighting for speech applications originating from the direction of arrival toward the wearer of the ear-worn device than for speech applications originating from the direction of arrival toward the side and rear of the wearer.

60. The ear-worn device according to any one of claims 58-59, wherein, The neural network circuit is also configured to receive one or more spatial focus control inputs indicating the specific spatial focus mode; And using the one or more spatial focus control inputs to generate the two or more neural network outputs, such that the target speech signal includes a speech signal to which the specific spatial focus pattern has been applied.

61. The ear-worn device according to claim 60, further comprising: A communication circuit is configured to receive from a processing device an indication of a user selection of the specific spatial focusing mode; And control circuitry configured to generate one or more spatial focus control inputs that indicate the particular spatial focus mode, at least in part based on a user selection of the particular spatial focus mode.

62. A system comprising: The ear-worn device as described in claim 61; The processing device, which communicates with the ear-worn device and is configured to: display a graphical user interface including options for different spatial focus modes; and receive a user selection for the specific spatial focus mode.

63. The system according to claim 62, wherein, Multiple options include four options.

64. The system according to any one of claims 62-63, wherein, The multiple options are graphical representations of the different spatial focusing modes.

65. The ear-worn device according to claim 60, further comprising: A sensing circuit configured to generate one or more inputs based on the movement of the ear-worn device; And control circuitry configured to: determine the degree of head movement based on one or more inputs received from one or more sensors; and generate one or more spatial focus control inputs indicating the specific spatial focus mode based on the degree of head movement.

66. The ear-worn device according to claim 65, wherein, The control circuit is configured to, when generating one or more inputs indicating a spatial focus mode based on the degree of head movement: generate a first set of one or more spatial focus control inputs based on a first degree of head movement, the first set of one or more spatial focus control inputs indicating a first spatial focus mode having a first spatial focus amount; And based on the second head movement degree, generate a second set of one or more spatial focus control inputs, the second set of one or more spatial focus control inputs indicating a second spatial focus mode with a second spatial focus amount; wherein the first spatial focus amount is less than the second spatial focus amount, and the first head movement degree is greater than the second head movement degree.

67. The ear-worn device according to claim 60, further comprising: The control circuit is configured to: determine the signal-to-noise ratio (SNR) of the acoustic environment; and generate one or more spatial focus control inputs indicating the spatial focus mode based on the SNR of the acoustic environment.

68. The ear-worn device according to any one of claims 55-67, wherein, The interfering speech signal includes the remainder when the target speech signal is subtracted from the speech signal.

69. The ear-worn device according to any one of claims 55-67, wherein, The one or more neural network outputs include two or more neural network outputs.

70. The ear-worn device according to claim 69, wherein, The noise reduction circuit is configured to obtain, based on the output of the two or more neural networks, at least one of the following: a speech signal, which includes speech in a first audio signal among the plurality of audio signals; and background noise signal, which includes background noise in the first audio signal; And at least one of the following: a target speech signal, which includes a first spatially focused version of the speech signal; And interfering with the speech signal, which includes a second spatially focused version of the speech signal.

71. The ear-worn device according to claim 69, wherein, The noise reduction circuit is configured to obtain, based on the outputs of the two or more neural networks, a target speech signal, which includes a first spatially focused version of the speech signal; And interfering with the speech signal, including a second spatially focused version of the speech signal.

72. The ear-worn device according to any one of claims 69-71, wherein: The neural network circuit is configured to: use a first subset of the one or more neural network layers to generate a first neural network output among the two or more neural network outputs; and use a second subset of the one or more neural network layers to generate a second neural network output among the two or more neural network outputs; and the noise reduction circuit is configured to obtain the speech signal and / or the background noise signal from the first neural network output among the two or more neural network outputs, and to obtain the target speech signal and / or the interfering speech signal from the second neural network output among the two or more neural network outputs.

73. The ear-worn device according to any one of claims 69-72, wherein, The outputs of the two or more neural networks include two different masks.

74. The ear-worn device according to any one of claims 69-73, wherein: At least one of the two or more neural network outputs includes: the speech signal; a mask configured to generate the speech signal; the background noise signal; a mask configured to generate the background noise signal; the target speech signal; a mask configured to generate the target speech signal; and the interfering speech signal; a mask configured to generate the interfering speech signal.

75. The ear-worn device according to any one of claims 55-74, further comprising: A mixing circuit configured to generate the output audio signal by mixing a combination of audio signals; Alternatively, the output audio signal can be generated by mixing combinations of masks; Alternatively, a wide dynamic range compression (WDRC) circuit may be used, comprising multiple WDRC pipelines configured to generate the output audio signal by performing WDRC on a combination of the audio signals.

76. The ear-worn device according to claim 75, wherein, The mixing circuit is further configured to: receive a first volume change control input and a second volume change control input; and use the first volume change control input and the second volume change control input to perform the mixing such that the first volume change difference is controlled at least partially by the first volume change control input, and the second volume change difference is controlled at least partially by the second volume change control input.

77. The ear-worn device according to claim 76, further comprising: A communication circuit configured to receive the first volume change control input and the second volume control change value from a processing device; A memory configured to store the first volume change control input and the second volume change control input; and a control circuit configured to retrieve the first volume change control input and the second volume change control input from the memory and output the first volume change control input and the second volume change control input to the mixing circuit.

78. The ear-worn device according to any one of claims 55-74, wherein, The neural network circuit is further configured to: receive a first volume change control input and a second volume change control input; and use the first volume change control input and the second volume change control input to generate the one or more neural network outputs, such that the first volume change difference is controlled at least partially by the first volume change control input, and the second volume change difference is controlled at least partially by the second volume change control input.

79. The ear-worn device according to claim 78, further comprising: A communication circuit configured to receive the first volume change control input and the second volume control change value from a processing device; A memory configured to store the first volume change control input and the second volume change control input; and a control circuit configured to retrieve the first volume change control input and the second volume change control input from the memory and output the first volume change control input and the second volume change control input to the mixing circuit.

80. The ear-worn device according to any one of claims 76-79, further comprising: A control circuit is configured to generate the first volume change control input based on the level of background noise in the first audio signal. And generate the second volume change control input based on the level of interfering speech in the first audio signal.

81. The ear-worn device according to any one of claims 55-80, wherein: The background noise signal is not spatially focused; and the interfering speech signal does not include a portion of the background noise in the first audio signal.

82. The ear-worn device according to any one of claims 55-80, wherein: The background noise signal includes a first spatially focused version of the background noise in the first audio signal; The interfering speech signal includes a second spatially focused version of the speech signal plus a second spatially focused version of the background noise in the first audio signal.

83. The ear-worn device according to any one of claims 43-82, wherein, The ear-worn device is also configured to receive a user selection to turn off spatial focusing.

84. The ear-worn device according to any one of claims 41-83, wherein, At least two of the plurality of audio signals have different beamforming directional patterns.

85. The ear-worn device according to claim 84, wherein, At least two of the plurality of audio signals having different beamforming directional patterns include beamforming signals having dipole, supercardioid, supercardioid, or cardioid directional patterns.

86. The ear-worn device according to any one of claims 41-74 and 78-85, further comprising: A mixing circuit configured to mix two or more audio signals such that the output audio signal includes a background noise modified and / or spatially focused version of the first audio signal mixed with a second audio signal.

87. The ear-worn device according to claim 86, wherein: The first audio signal with modified background noise and spatial focus includes the target speech signal; the target speech signal includes a first spatial focus version of the speech signal. The speech signal includes the speech in the first audio signal; The second audio signal includes: a background noise signal, which includes the background noise in the first audio signal; and an interfering speech signal, which includes a second spatially focused version of the speech signal; and the noise reduction circuit is configured to generate the output audio signal such that in the output audio signal: the volume change of the background noise signal differs from the volume change of the target speech signal by a volume change difference; the volume change of the interfering speech signal differs from the volume change of the target speech signal by the volume change difference; and the volume change difference is controllable.

88. The ear-worn device according to claim 87, wherein: The mixing circuit is configured to: receive a volume change control input; and use the volume change control input to perform the mixing, such that the volume change difference is controlled at least partially by the volume change control input; The ear-worn device includes: a communication circuit configured to receive the volume change control input from a processing device; a memory configured to store the volume change control input; and a control circuit configured to retrieve the volume change control input and output the volume change control input to the mixing circuit.

89. The ear-worn device according to any one of claims 41-88, wherein, The ear-worn device includes a hearing aid.

90. The ear-worn device according to any one of claims 41-89, wherein, The noise reduction circuit is implemented on a chip.

91. An ear-worn device, comprising: Two or more microphones; The noise reduction circuit includes a neural network circuit, wherein: the neural network circuit is configured to: receive a plurality of audio signals, wherein at least two of the plurality of audio signals are each derived from different microphones of the two or more microphones, and / or at least one of the plurality of audio signals is a beamformed audio signal derived from the two or more microphones; and implement one or more neural network layers trained to perform background noise modification and spatial focusing or to generate an output for performing background noise modification and spatial focusing based on the plurality of audio signals, such that the neural network circuit generates one or more neural network outputs based on the plurality of audio signals, wherein: the noise reduction circuit is configured to output an output audio signal based on the one or more neural network outputs, the output audio signal including a background noise modified and spatially focused version of a first audio signal among the plurality of audio signals.

92. The ear-worn device according to claim 91, wherein, At least two of the plurality of audio signals have different beamforming directional patterns.

93. The ear-worn device according to any one of claims 91-92, wherein, At least one of the plurality of audio signals has a forward beamforming directional pattern, and at least one of the plurality of audio signals has a backward beamforming directional pattern.

94. The ear-worn device according to any one of claims 91-93, wherein: The output audio signal has a specific spatial focusing pattern, which includes different weights for voice applications in the first audio signal originating from different directions of arrival relative to the wearer of the ear-worn device.

95. The ear-worn device according to claim 94, wherein, The specific spatial focusing mode includes applying a higher weight to speech originating from the direction of arrival toward the wearer of the ear-worn device than to speech originating from the direction of arrival toward the side and rear of the wearer.

96. The ear-worn device according to any one of claims 93-95, wherein, The neural network circuit is also configured to receive one or more spatial focus control inputs indicating the specific spatial focus mode; And using the one or more spatial focus control inputs to generate the one or more neural network outputs, such that the output audio signal has the specific spatial focus pattern.

97. The ear-worn device according to any one of claims 93-96, further comprising: A communication circuit is configured to receive from a processing device an indication of a user selection of the specific spatial focusing mode; And control circuitry configured to generate one or more spatial focus control inputs that indicate the particular spatial focus mode, at least in part based on a user selection of the particular spatial focus mode.

98. A system comprising: The ear-worn device as described in claim 97; The processing device, which communicates with the ear-worn device and is configured to: display a graphical user interface including options for different spatial focus modes; and receive a user selection for the specific spatial focus mode.

99. The ear-worn device according to any one of claims 91-98, wherein: The one or more neural network outputs include two or more neural network outputs; The noise reduction circuit is configured to generate an output audio signal based on the outputs of the two or more neural networks. The output audio signal includes: a target speech signal comprising a background noise-modified and spatially focused version of the first audio signal, wherein the target speech signal comprises a first spatially focused version of the speech signal and the speech signal comprises speech from the first audio signal; an interfering speech signal comprising a second spatially focused version of the speech signal; and a background noise signal comprising background noise from the first audio signal. The noise reduction circuit is configured to generate the output audio signal such that in the output audio signal: the volume change of the background noise signal differs from the volume change of the target speech signal by a first volume change difference; the volume change of the interfering speech signal differs from the volume change of the target speech signal by a second volume change difference; and the first volume change difference and the second volume change difference are independently controllable.

100. The ear-worn device according to claim 99, wherein, The interfering speech signal includes the remainder when the target speech signal is subtracted from the speech signal.

101. The ear-worn device according to any one of claims 99-100, wherein: The neural network circuit is configured to: use a first subset of the one or more neural network layers to generate a first neural network output among the two or more neural network outputs; and use a second subset of the one or more neural network layers to generate a second neural network output among the two or more neural network outputs; and the noise reduction circuit is configured to obtain the speech signal and / or the background noise signal from the first neural network output among the two or more neural network outputs, and to obtain the target speech signal and / or the interfering speech signal from the second neural network output among the two or more neural network outputs.

102. The ear-worn device according to any one of claims 99-101, wherein, The outputs of the two or more neural networks include two different masks.

103. The ear-worn device according to any one of claims 99-102, further comprising: A mixing circuit configured to generate the output audio signal by mixing a combination of audio signals; Alternatively, the output audio signal can be generated by mixing combinations of masks; Alternatively, a wide dynamic range compression (WDRC) circuit may be used, comprising multiple WDRC pipelines configured to generate the output audio signal by performing WDRC on a combination of the audio signals.

104. The ear-worn device according to claim 103, wherein, The mixing circuit is further configured to: receive a first volume change control input and a second volume change control input; and use the first volume change control input and the second volume change control input to perform the mixing such that the first volume change difference is controlled at least partially by the first volume change control input, and the second volume change difference is controlled at least partially by the second volume change control input.

105. The ear-worn device according to claim 104, further comprising: A communication circuit configured to receive the first volume change control input and the second volume control change value from a processing device; A memory configured to store the first volume change control input and the second volume change control input; and a control circuit configured to retrieve the first volume change control input and the second volume change control input from the memory and output the first volume change control input and the second volume change control input to the mixing circuit.

106. The ear-worn device according to any one of claims 91-102 further includes a mixing circuit configured to mix two or more audio signals such that the output audio signal includes a background noise-modified and spatially focused version of the first audio signal mixed with a second audio signal.

107. The ear-worn device according to claim 106, wherein: The first audio signal with modified background noise and spatial focus includes the target speech signal; the target speech signal includes a first spatial focus version of the speech signal. The speech signal includes the speech in the first audio signal; The second audio signal includes: a background noise signal, which includes the background noise in the first audio signal; and an interfering speech signal, which includes a second spatially focused version of the speech signal; and the noise reduction circuit is configured to generate the output audio signal such that in the output audio signal: the volume change of the background noise signal differs from the volume change of the target speech signal by a volume change difference; the volume change of the interfering speech signal differs from the volume change of the target speech signal by the volume change difference; and the volume change difference is controllable.

108. The ear-worn device according to claim 107, wherein: The mixing circuit is configured to: receive a volume change control input; and use the volume change control input to perform the mixing such that the volume change difference is controlled at least partially by the volume change control input.

109. The ear-worn device according to claim 108, wherein, The ear-worn device includes: a communication circuit configured to receive the volume change control input from a processing device; a memory configured to store the volume change control input; and a control circuit configured to retrieve the volume change control input and output the volume change control input to the hybrid circuit.

110. The ear-worn device according to any one of claims 91-109, wherein, The ear-worn device is also configured to receive a user selection to turn off spatial focusing.

111. An ear-worn device, comprising: Two or more microphones; The noise reduction circuit includes a neural network circuit, wherein: the neural network circuit is configured to: receive a plurality of audio signals, wherein at least two of the plurality of audio signals are each derived from different microphones of the two or more microphones, and / or at least one of the plurality of audio signals is a beamformed audio signal derived from the two or more microphones; and implement one or more neural network layers trained to generate one or more neural network outputs based on the plurality of audio signals, wherein: the one or more neural network outputs include an output audio signal, the output audio signal including a background noise modified and spatially focused version of a first audio signal among the plurality of audio signals; or the one or more neural network outputs are configured to be used by the noise reduction circuit to generate an output audio signal including a background noise modified and spatially focused version of the first audio signal.

112. The ear-worn device according to claim 111, wherein, At least two of the plurality of audio signals have different beamforming directional patterns.

113. The ear-worn device according to any one of claims 111-112, wherein, At least one of the plurality of audio signals has a forward beamforming directional pattern, and at least one of the plurality of audio signals has a backward beamforming directional pattern.

114. The ear-worn device according to any one of claims 111-113, wherein: The output audio signal has a specific spatial focusing pattern, which includes different weights for voice applications in the first audio signal originating from different directions of arrival relative to the wearer of the ear-worn device.

115. The ear-worn device according to claim 114, wherein, The specific spatial focusing mode includes a higher weighting for speech applications originating from the direction of arrival toward the wearer of the ear-worn device than for speech applications originating from the direction of arrival toward the side and rear of the wearer.

116. The ear-worn device according to any one of claims 113-115, wherein, The neural network circuit is also configured to receive one or more spatial focus control inputs indicating the specific spatial focus mode; And using the one or more spatial focus control inputs to generate the one or more neural network outputs, such that the output audio signal has the specific spatial focus pattern.

117. The ear-worn device according to any one of claims 113-116, further comprising: A communication circuit is configured to receive from a processing device an indication of a user selection of the specific spatial focusing mode; And control circuitry configured to generate one or more spatial focus control inputs that indicate the particular spatial focus mode, at least in part based on a user selection of the particular spatial focus mode.

118. A system comprising: The ear-worn device as described in claim 117; The processing device, which communicates with the ear-worn device and is configured to: display a graphical user interface including options for different spatial focus modes; and receive a user selection for the specific spatial focus mode.

119. The ear-worn device according to any one of claims 111-118, wherein: The one or more neural network outputs include two or more neural network outputs; The noise reduction circuit is configured to generate an output audio signal based on the outputs of the two or more neural networks. The output audio signal includes: a target speech signal comprising a background noise-modified and spatially focused version of the first audio signal, wherein the target speech signal comprises a first spatially focused version of the speech signal, and the speech signal comprises speech from the first audio signal; an interfering speech signal comprising a second spatially focused version of the speech signal; and a background noise signal comprising background noise from the first audio signal. The noise reduction circuit is configured to generate the output audio signal such that in the output audio signal: the volume change of the background noise signal differs from the volume change of the target speech signal by a first volume change difference; the volume change of the interfering speech signal differs from the volume change of the target speech signal by a second volume change difference; and the first volume change difference and the second volume change difference are independently controllable.

120. The ear-worn device according to claim 119, wherein, The interfering speech signal includes the remainder when the target speech signal is subtracted from the speech signal.

121. The ear-worn device according to any one of claims 119-120, wherein: The neural network circuit is configured to: use a first subset of the one or more neural network layers to generate a first neural network output among the two or more neural network outputs; and use a second subset of the one or more neural network layers to generate a second neural network output among the two or more neural network outputs; and the noise reduction circuit is configured to obtain the speech signal and / or the background noise signal from the first neural network output among the two or more neural network outputs, and to obtain the target speech signal and / or the interfering speech signal from the second neural network output among the two or more neural network outputs.

122. The ear-worn device according to any one of claims 119-121, wherein, The outputs of the two or more neural networks include two different masks.

123. The ear-worn device according to any one of claims 119-122, further comprising: A mixing circuit configured to generate the output audio signal by mixing a combination of audio signals; Alternatively, the output audio signal can be generated by mixing combinations of masks; Alternatively, a wide dynamic range compression (WDRC) circuit may be used, comprising multiple WDRC pipelines configured to generate the output audio signal by performing WDRC on a combination of the audio signals.

124. The ear-worn device according to claim 123, wherein, The mixing circuit is further configured to: receive a first volume change control input and a second volume change control input; and use the first volume change control input and the second volume change control input to perform the mixing such that the first volume change difference is controlled at least partially by the first volume change control input, and the second volume change difference is controlled at least partially by the second volume change control input.

125. The ear-worn device according to claim 124, further comprising: A communication circuit configured to receive the first volume change control input and the second volume control change value from a processing device; A memory configured to store the first volume change control input and the second volume change control input; and a control circuit configured to retrieve the first volume change control input and the second volume change control input from the memory and output the first volume change control input and the second volume change control input to the mixing circuit.

126. The ear-worn device according to any one of claims 111-122, further comprising a mixing circuit configured to mix two or more audio signals such that the output audio signal includes a background noise-modified and spatially focused version of the first audio signal mixed with a second audio signal.

127. The ear-worn device according to claim 126, wherein: The first audio signal with modified background noise and spatial focus includes the target speech signal; the target speech signal includes a first spatial focus version of the speech signal. The speech signal includes the speech in the first audio signal; The second audio signal includes: a background noise signal, which includes the background noise in the first audio signal; and an interfering speech signal, which includes a second spatially focused version of the speech signal; and the noise reduction circuit is configured to generate the output audio signal such that in the output audio signal: the volume change of the background noise signal differs from the volume change of the target speech signal by a volume change difference; the volume change of the interfering speech signal differs from the volume change of the target speech signal by the volume change difference; and the volume change difference is controllable.

128. The ear-worn device according to claim 127, wherein: The mixing circuit is configured to: receive a volume change control input; and use the volume change control input to perform the mixing such that the volume change difference is controlled at least partially by the volume change control input.

129. The ear-worn device according to claim 128, wherein, The ear-worn device includes: a communication circuit configured to receive the volume change control input from a processing device; a memory configured to store the volume change control input; and a control circuit configured to retrieve the volume change control input and output the volume change control input to the hybrid circuit.

130. The ear-worn device according to any one of claims 111-129, wherein, The ear-worn device is also configured to receive a user selection to turn off spatial focusing.

Citation Information

Patent Citations

  • Method, apparatus and system for neural network hearing aid

    US11812225B2

  • System and method for enhancing speech of target speaker from audio signal in an ear-worn device using voice signatures

    US11818523B2

  • Neural network chip for ear-worn device

    US11886974B1