Ear-worn device with noise reduction and / or spatial focusing based on neural networks

A neural network-based spatial focusing system in ear-worn devices addresses noise reduction challenges by using multiple microphones to enhance target speech and attenuate background noise and interfering speech, improving performance across various acoustic conditions.

JP2026528757APending Publication Date: 2026-08-25FORTELL RESEARCH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026506184
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-08
Filing Date
2024-08-05
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Conventional ear-worn devices face challenges in noise reduction, particularly in scenarios with multiple speakers, due to distorted beamforming patterns caused by the wearer's anatomy and limitations in reverberant environments, and are less effective for low-frequency sounds, often adding noise in quiet environments.

Method used

Implementing a neural network-based spatial focusing system that uses multiple microphones to determine sound direction and apply different weights to audio signals, enhancing sounds from the target direction while attenuating others, with independent control over background noise and interfering speech volumes.

Benefits of technology

Improves noise reduction by selectively amplifying target speech and attenuating background noise and interfering speech, enhancing environmental awareness and reducing distortion, even in complex acoustic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026528757000001_ABST
    Figure 2026528757000001_ABST
Patent Text Reader

Abstract

The ear-worn device comprises two or more microphones and a noise reduction circuit including a neural network circuit. The neural network circuit is configured to receive a plurality of audio signals, each of which at least two of the plurality of audio signals is obtained from one of the two or more microphones and / or at least one of the plurality of audio signals is a beamforming audio signal obtained from the two or more microphones, and to implement one or more neural network layers that have been trained to perform background noise correction and spatial focusing based on the plurality of audio signals, such that the neural network circuit generates one or more neural network outputs based on the plurality of audio signals. The noise reduction circuit is configured to output an output audio signal that includes a background noise-corrected and spatially focused version of a first audio signal of the plurality of audio signals, based on the one or more neural network outputs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to ear-worn devices. Some aspects relate to ear-worn devices with neural network-based noise correction and / or spatial focusing.

Background Art

[0002] Ear-worn devices such as hearing aids can be used to help people with hearing loss hear better. Typically, ear-worn devices amplify the received sound. Some ear-worn devices may attempt to reduce noise in the received sound.

Summary of the Invention

[0003] Reducing noise in the output of ear-worn devices (e.g., hearing aids, cochlear implants, and earphones) is a difficult problem. Reducing noise in a scenario where a wearer is listening to one speaker while there are other interfering speakers nearby is an especially difficult problem. The inventors recognized that neural networks can be used in ear-worn devices to improve noise reduction and reduction of sound from interfering speakers. Recently, neural networks for separating speech sounds from noise have been developed. Further explanation of such neural networks for noise reduction can be found in U.S. Patent No. 11,812,225, titled "Methods, Apparatus, and Systems for Neural Network Hearing Aids," issued on November 7, 2023, the entire disclosure of which is incorporated herein by reference.

[0004] To reduce background noise, the inventors recognized that if a neural network heard noise from a specific direction of arrival (DOA) in a previous time step, the neural network may have prior information to cancel out the noise from that DOA in the current time step. From another perspective, since sound sources tend to move slowly over time, if a neural network identifies a particular sound segment as speech and knows its DOA, the neural network can reasonably estimate that other sounds from the same direction are also speech.

[0005] To reduce sound from interfering speakers, conventional in-ear devices can use beamforming to attenuate sound received from a specific direction. This may involve processing sound from different microphones in different ways (e.g., applying different delays to signals received by different microphones). Conventional beamforming (both adaptive and non-adaptive beamforming) can provide a perceptual boost, so it can focus on sound coming from the front of the wearer (the direction from which the sound of interest is expected to originate) and attenuate sound from the sides and rear of the wearer (e.g., background noise and interfering speakers).

[0006] However, conventional beamforming patterns (e.g., cardioid, supercardioid, hypercardioid, and dipole) may also have the following drawbacks: 1. Theoretical beamforming patterns can be distorted, at least partially, when implemented by microphones placed behind the ear (e.g., hearing aids) due to interference from the wearer's head, torso, and ears, which can degrade performance. 2. In reverberant environments, indirect paths can come from the front. For example, in a reverberant room, if a speaker is speaking directly behind the wearer, their voice may reverberate throughout the room and enter the microphone of an in-ear device from the front of the wearer. Such sounds may not be attenuated by a front-directional beamforming pattern. 3. Conventional beamforming may work more effectively for high-frequency sounds than for low-frequency sounds. In other words, conventional beamforming may be better suited to using high-frequency sounds for sound source localization than for low-frequency sounds. 4. Generally, there are limits to the noise reduction that conventional beamforming patterns can provide. 5. In quiet environments, beamforming may add noise.

[0007] The inventors have addressed these shortcomings by developing a neural network trained to perform spatial focusing, which can be implemented in specific embodiments. Spatial focusing may involve assigning different weights to an audio signal based on the location of the sound source or the direction from which the audio signal was generated relative to the device. The location and / or direction of the sound may be derived by the neural network from the timing differences of the sound reaching multiple microphones. The inventors have recognized that a single microphone may not adequately distinguish between speakers from different directions by a neural network. In other words, the neural network may not be able to distinguish whether the speaker is in front of or behind the wearer (or, more generally, where the wearer is located). A neural network using inputs from multiple microphones can resolve this ambiguity. Therefore, a neural network can be trained to accept multiple input audio signals from two or more microphones on an ear-worn device and perform spatial focusing. Spatial focusing can help focus on sounds coming from a target direction and reduce sounds coming from other directions. One specific example is that by focusing on sound coming from the front of the wearer of an ear-worn device, it can help reduce sound from interfering speakers located behind and to the sides of the wearer. Other target directions can be used similarly. This approach may allow for weighting of sounds coming from different directions with a larger difference than conventional beamforming.

[0008] Certain spatial focusing patterns may focus on sounds coming from the front of the wearer of the ear-mounted device, while some may focus on sounds coming from other directions, such as the sides and / or rear of the wearer. Such focusing may be beneficial in certain scenarios, for example, when the wearer of the ear-mounted device is driving a car with passengers to the sides and / or rear.

[0009] Certain embodiments of the technology described herein may additionally generate a target speech voice signal, a background noise signal, and an interfering speech voice signal, which may be mixed using modified levels of the target speech voice, interfering speech voice, and / or background noise. For example, the levels of the interfering speech voice and background noise may be reduced, while the level of the target speech voice may be maintained at a similar level or increased. Mixing some noise and some interfering speech voice with the target speech voice signal may help reduce distortion and improve the wearer's environmental awareness of the in-ear device. Generally, changes in the volume of the background noise signal differ from changes in the volume of the target speech voice signal by a first volume difference, and changes in the volume of the interfering speech voice signal differ from changes in the volume of the target speech voice signal by a second volume difference, and the first and second volume difference amounts may be independently controllable.

[0010] Independently controlling the volume of background noise and interfering speech can be useful, for example, to enable different levels of reduction of background noise and interfering speech based on different preferences of different wearers. Below are some non-limiting examples of scenarios that illustrate how independent control of background noise and interfering speech volume can be useful. In the first scenario, the wearer may be sitting at a table in a crowded restaurant with multiple conversation partners. It may be useful to significantly reduce background noise but not significantly reduce any of the speech (i.e., not significantly reduce interfering speech). In the second scenario, the wearer is sitting at a table in a crowded restaurant with one conversation partner, but there may be a loud conversation at a nearby table. It may be useful to significantly reduce both background noise and interfering speech. In the third scenario, the wearer is in a quiet cafe with a disturbing conversation taking place nearby. It may be useful to moderately reduce background noise and significantly reduce interfering speech. [Brief explanation of the drawing]

[0011] [Figure 1] This figure shows a hearing aid according to a specific embodiment described herein. [Figure 2] This figure shows the hearing aid of Figure 1 according to a specific embodiment described herein, as worn by the wearer. [Figure 3] This figure shows eyeglasses with a built-in hearing aid according to a specific embodiment described herein. [Figure 4] This figure shows a system for operating an ear-worn device according to a specific embodiment described herein. [Figure 5] This figure shows the circuitry within an ear-worn device according to a specific embodiment described herein. [Figure 6] This figure shows an audio signal according to a specific embodiment described herein. [Figure 7] This figure shows the change in volume according to a specific embodiment described herein. [Figure 8] This figure shows a noise reduction circuit in an ear-worn device according to a specific embodiment described herein. [Figure 9] This figure shows a neural network circuit and a mask application and subtraction circuit according to a specific embodiment described herein. [Figure 10] This figure shows a neural network circuit and a mask application and subtraction circuit according to a specific embodiment described herein. [Figure 11] This figure shows a neural network circuit and a mask application and subtraction circuit according to a specific embodiment described herein. [Figure 12] This figure shows a noise reduction circuit in an ear-worn device according to a specific embodiment described herein. [Figure 13] This figure shows in more detail the wide dynamic range compression (WDRC) circuit of Figure 12 according to a specific embodiment described herein. [Figure 14] This figure shows the circuitry within an ear-worn device according to a specific embodiment described herein. [Figure 15] A diagram showing a circuit within an ear-worn device according to a particular embodiment described herein. [Figure 16] A diagram showing a circuit within an ear-worn device according to a particular embodiment described herein. [Figure 17] A diagram showing a circuit within an ear-worn device according to a particular embodiment described herein. [Figure 18] A diagram showing an exemplary spatial focusing pattern according to a particular embodiment described herein. [Figure 19] A diagram showing an exemplary spatial focusing pattern according to a particular embodiment described herein. [Figure 20] A diagram showing an exemplary spatial focusing pattern according to a particular embodiment described herein. [Figure 21] A diagram showing a circuit for controlling spatial focusing in an ear-worn device according to a particular embodiment described herein. [Figure 22] A diagram showing a graphical user interface (GUI) for controlling spatial focusing of an ear-worn device according to a particular embodiment described herein. [Figure 23] A diagram showing a graphical user interface (GUI) for controlling spatial focusing of an ear-worn device according to a particular embodiment described herein. [Figure 24] A diagram showing a graphical user interface (GUI) for controlling spatial focusing of an ear-worn device according to a particular embodiment described herein. [Figure 25] A diagram showing a graphical user interface (GUI) for controlling spatial focusing of an ear-worn device according to a particular embodiment described herein. [Figure 26] A diagram showing a graphical user interface (GUI) for controlling spatial focusing of an ear-worn device according to a particular embodiment described herein. [Figure 27]A diagram showing a forward-facing hypercardioid pattern according to a specific embodiment described in this specification. [Figure 28] A diagram showing a rear-facing hypercardioid pattern according to a specific embodiment described in this specification. [Figure 29] A diagram showing a forward-facing supercardioid pattern according to a specific embodiment described in this specification. [Figure 30] A diagram showing a rear-facing hypercardioid pattern according to a specific embodiment described in this specification. [Figure 31] A diagram showing a forward-facing cardioid pattern according to a specific embodiment described in this specification. [Figure 32] A diagram showing a rear-facing cardioid pattern according to a specific embodiment described in this specification. [Figure 33] A diagram showing a dipole pattern according to a specific embodiment described in this specification.

Modes for Carrying Out the Invention

[0012] The above aspects and embodiments, as well as additional aspects and embodiments, are further described below. These aspects and / or embodiments can be used individually, all together, or in any combination of two or more because the present disclosure is not limited in this regard.

[0013] (Ear-worn device) Figure 1 shows a hearing aid 100 according to a specific embodiment described herein. The hearing aid 100 may be either an in-ear device or a hearing aid as described herein. The hearing aid 100 is a receiver-in-canal (RIC) (also known as a receiver-in-the-ear (RITE)) hearing aid. However, any other type of hearing aid (e.g., behind-the-ear, in-the-ear, intracanal, fully intracanal, open-fit, etc.) may also be used. The hearing aid 100 includes a body 111, a receiver wire 113, a receiver 106, and a dome 115. The body 111 is coupled to the receiver wire 113, and the receiver wire 113 is coupled to the receiver 106. The dome 115 is positioned on top of the receiver 106. The body 111 includes a front microphone 102f, a back microphone 102b, and a user input device 104. The main unit 111 further includes circuits not shown in Figure 1 (e.g., any of the circuits described below, excluding the receiver 106). When the hearing aid 100 is worn, the front microphone 102f is positioned closer to the front of the wearer, and the back microphone 102b is positioned closer to the rear of the wearer. The front microphone 102f and the back microphone 102b may be configured to receive sound signals and generate speech signals based on sound signals. Any of the two or more microphones described herein may be the front microphone 102f and the back microphone 102b of the hearing aid 100. A user input device 104 (e.g., a button) may be configured to control specific functions of the hearing aid 100, such as level and enabling neural network-based noise cancellation.

[0014] The receiver wire 113 may be configured to transmit an audio signal from the main unit 111 to the receiver 106. The receiver 106 may be configured to receive the audio signal (i.e., the audio signal generated by the main unit 111 and transmitted by the receiver wire 113) and to generate an audio signal based on the audio signal. The dome 115 may be configured to fit snugly inside the wearer's ear and to guide the audio signal generated by the receiver 106 into the wearer's ear canal.

[0015] In some embodiments, the length of the main body 111 may be equal to 2 cm, equal to 5 cm, or between 2 cm and 5 cm. In some embodiments, the weight of the hearing aid 100 may be less than 4.5 grams. In some embodiments, the spacing between the microphones may be equal to 5 mm, equal to 12 mm, or between 5 mm and 12 mm. In some embodiments, the main body 111 may include a battery (not shown in Figure 1), such as a lithium-ion rechargeable coin cell battery.

[0016] Figure 2 shows a hearing aid 100 fitted to a wearer 208 according to a specific embodiment described herein. Figure 2 is shown from the rear of the wearer 208, and as shown, the front microphone 102f is positioned near the front of the wearer 208, and the back microphone 102b is positioned near the rear of the wearer 208. Figures 1 and 2 show a RIC hearing aid, but hearing aids of other form factors may also be used.

[0017] Figure 3 shows eyeglasses 300 with a built-in hearing aid according to a particular embodiment described herein. The eyeglasses 300 may be either an in-ear device or a hearing aid as described herein. The eyeglasses 300 has a left temple 310, a right temple 312, and a front rim 314. The eyeglasses 300 further has a receiver 306 connected to each of the left temple 310 and the right temple 312. Figure 3 shows a microphone 302 located on the left temple 310. A microphone 302 may also be located on the right temple 312 (not shown in the figure). A microphone 302 may also be located on the front rim 314 (not shown in the figure). Figure 3 shows five microphones 302 on the left temple 310, but more or fewer microphones may be located on the temples or rims. In some embodiments (such as the embodiment in Figure 3), the entry point for the microphone 302 is located on the inside of the temple and / or rim (i.e., on the side facing the wearer's face) to make it less visible to others. In some embodiments, the entry point for the microphone 302 is located on the top of the temple and / or rim to make it less visible to others. In some embodiments, the entry point for the microphone 302 may be located on the outside of the temple and / or rim (i.e., on the side away from the wearer's face). Any of the two or more microphones described herein may be any of the microphones 302 of the eyeglasses 300. Figures 1 to 3 show hearing aids and eyeglasses, but it should be understood that other in-ear devices, such as cochlear implants or earphones, may also be used.

[0018] Figure 4 shows a system 416 for operating an in-ear device 400 according to a specific embodiment described herein. The system 416 includes an in-ear device 400, a processing device 418, and a wireless communication link 420. The in-ear device 400 may be, for example, a hearing aid (e.g., hearing aid 100 or eyeglasses 300), a cochlear implant, an earphone, or another in-ear device. The processing device 418 may be, for example, a smartphone, a tablet, or a laptop computer. The wireless communication link 420 may be, for example, Bluetooth or an NFMI communication link. The processing device 418 is able to communicate with the in-ear device 400 (i.e., via the wireless communication link 420). The processing device 418 may be configured to send commands to the in-ear device 400 via the wireless communication link 420 (e.g., to send commands to set the in-ear device 400 to a specific mode). The in-ear device 400 may be configured to send information (e.g., usage data) to the processing device 418 via the wireless communication link 420. Although not shown, the system 416 includes multiple ear-worn devices, such as an ear-worn device worn on the right ear and an ear-worn device worn on the left ear, and the processing device 418 can communicate with each device via a wireless communication link.

[0019] Figure 5 shows the circuitry of an in-ear device 500 according to a specific embodiment described herein. The in-ear device may be, for example, a hearing aid 100, eyeglasses 300, and / or an in-ear device 400. The in-ear device 500 includes a microphone 502, a processing circuit 522, a noise reduction circuit 524, a processing circuit 528, and a receiver 506. The noise reduction circuit 524 includes a neural network circuit 526. The in-ear device 500 may include more circuits and components than those shown (e.g., feedback suppression circuits, calibration circuits, etc.), and such circuits and components may be placed before, after, or between certain circuits and components among those shown in Figure 5.

[0020] In the ear-worn device 500, the processing circuit 522 is coupled between the microphone 502 and the noise reduction circuit 524. The noise reduction circuit 524 is coupled between the processing circuit 522 and the processing circuit 528. The processing circuit 528 is coupled between the noise reduction circuit 524 and the receiver 506. Where it is stated herein that element A is coupled between element B and element C, other elements may exist between element A and element B, and / or between element A and element C. In the ear-worn device 500, the neural network circuit 526 may be downstream of the beamforming circuit in the processing circuit 522.

[0021] Microphone 502 may include two or more microphones (e.g., two, three, four, or more). For example, microphone 502 may include a front microphone located closer to the front of the wearer of the in-ear device and a back microphone located closer to the rear of the wearer of the in-ear device (e.g., microphones 102f and 102b in hearing aid 100). As another example, microphone 502 may include three or more microphones in an array (e.g., microphone 302 in eyeglasses 300). As yet another example, one microphone may be in a first in-ear device and one microphone may be in a second in-ear device wirelessly coupled to the first in-ear device. Microphone 502 may be configured to receive sound signals and generate speech signals from sound signals. The speech signals represent a plurality of individual speech signals, each generated by one of microphones 502. Thus, each of the speech signals may be obtained from one of microphones 502.

[0022] In some embodiments, the processing circuit 522 may include an analog processing circuit. The analog processing circuit may be configured to perform analog processing on the audio signal received from the microphone 502. For example, the analog processing circuit may be configured to perform one or more of the following: analog preamplification, analog filtering, and analog-to-digital conversion. Thus, the analog processing circuit may be configured to generate an analog-processed audio signal from the audio signal received from the microphone 502. The analog-processed audio signal may include a plurality of individual signals, each being an analog-processed version of the audio signal received from the microphone 502. In this specification, the analog processing circuit may include an analog-to-digital conversion circuit, and the analog-processed signal may be a digital signal converted from analog to digital by the analog-to-digital conversion circuit.

[0023] In some embodiments, the processing circuit 522 may include a digital processing circuit. The digital processing circuit may be configured to perform digital processing on the analog-processed audio signal received from the analog processing circuit. For example, the digital processing circuit may be configured to perform one or more of the following processes: wind noise reduction, input calibration, and feedback suppression. Thus, the digital processing circuit may be configured to generate a digitally processed audio signal from the analog-processed audio signal. The digitally processed audio signal may include a plurality of individual signals, each being a digitally processed version of the analog-processed audio signal.

[0024] In some embodiments, the processing circuit 522 may include a beamforming circuit. The beamforming circuit may be configured to generate one or more beamforming audio signals from two or more digitally processed audio signals. The beamforming audio signals include one or more individual signals, each of which is a beamforming version of two or more digitally processed audio signals. In some embodiments, the multiple beamforming audio signals may each have a different beamforming directional pattern. Beamforming will be described in more detail below.

[0025] The noise reduction circuit 524 includes a neural network circuit 526. The neural network circuit 526 may be configured to implement one or more neural network layers that can be learned to perform noise correction and / or spatial focusing, as described below. The term “noise correction” may be used herein to encompass both the process of reducing noise in the output signal compared to the input signal (i.e., noise reduction) and the process of reducing speech in the output signal compared to the input signal. (As described below, in some embodiments, the neural network circuit may be used to obtain a speech-separated version of the signal, and in some embodiments, the neural network circuit may be used to obtain a noise-separated version of the signal). Thus, in some embodiments, one or more neural network layers implemented by the neural network circuit 526 may be learned to correct noise. (Further explanation of what may be considered noise can be found below).

[0026] In such embodiments, the output from the neural network circuit 526 may be a version of the audio signal input to the neural network circuit 526 with less noise (or only the speech voice, as described below for the speech voice signal 603), or an output configured to generate a version of the audio signal input to the neural network circuit 526 with less noise (e.g., a mask or sound map), or a version of the audio signal input to the neural network circuit 526 with reduced speech voice components (or only noise, as described below for the background noise signal 601), or an output configured to generate a version of the audio signal input to the neural network circuit 526 with reduced speech voice components (e.g., a mask or sound map).

[0027] In some embodiments, one or more neural network layers implemented by the neural network circuit 526 may be trained to perform spatial focusing. In such embodiments, the output from the neural network circuit 526 may be a spatially focused version of the audio signal input to the neural network circuit 526, or an output (e.g., a mask or soundmap) configured to produce a spatially focused version of the audio signal input to the neural network circuit 526.

[0028] In some embodiments, one or more neural network layers implemented by the neural network circuit 526 may be trained to correct noise and perform spatial focusing. In such embodiments, the output from the neural network circuit 526 may be a noise-corrected and spatially focused version of the audio signal input to the neural network circuit 526 (e.g., the target speech audio signal 605 or the interfering speech audio signal 607 described below), or an output (e.g., a mask or sound map) configured to produce a noise-corrected and spatially focused version of the audio signal input to the neural network circuit 526.

[0029] In some embodiments, one neural network layer may be trained to correct noise and perform spatial focusing, or to correct noise and perform spatial focusing. In some embodiments, multiple neural network layers may be trained to correct noise and perform spatial focusing, or to correct noise and perform spatial focusing.

[0030] This description can describe one or more neural network layers that have been trained to perform a certain action, or one or more neural network layers that have been trained to produce an output used to perform that action. As used herein, one or more neural network layers may be considered trained to perform a certain action if the one or more neural network layers perform the action themselves, or produce an output used in performing that action. Therefore, even if the neural network itself does not produce a noise-reduced audio signal, one or more neural network layers may be considered trained to perform noise reduction. A neural network that produces an output used to produce a noise-reduced audio signal may still be considered trained to perform noise reduction. For example, a neural network may produce a mask configured to produce a noise-reduced audio signal. Also, even if the neural network itself does not produce a spatially focused audio signal, a neural network may be considered trained to perform spatial focusing.

[0031] A neural network that generates an output configured to produce a spatially focused audio signal can still be considered to have been trained to perform spatial focusing. The output could, in non-limiting examples, be a mask configured to produce a spatially focused audio signal, a sound map, a mask configured to produce a sound map, or a value calculated against a metric from audio from multiple beams (each of which points to a different angle around the wearer of an in-ear device). In some embodiments, one or more neural network layers may be configured to output a single output based on multiple input audio signals.

[0032] Any neural network layers described herein may be, for example, recurrent, vanilla / feedforward, convolutional, adversarial generative, attention (e.g., transformer), or graphical. Generally, a neural network consisting of such layers includes an input layer, a number of hidden layers, and an output layer, each layer consisting of a number of neurons / nodes to which neural network weights can be applied.

[0033] The processing circuit 528 may be configured to perform further processing on the output of the noise reduction circuit 524. For example, the processing circuit 528 may include a digital processing circuit configured to perform one or more of the following: wide dynamic range compression and output calibration.

[0034] Receiver 506 (which may be identical to, for example, receivers 106 and / or 306) may be configured to reproduce the output of processing circuit 528 as sound to the user's ears. Receiver 506 may be configured to perform a digital-to-analog conversion prior to playback.

[0035] In some embodiments, a portion of the circuitry of the ear-worn device 500 may be configured to process audio signals in the frequency domain. In such embodiments, processing circuit 522 may include a short-time Fourier transform (STFT) circuit configured to convert the short window of the audio signal from the time domain to the frequency domain, and processing circuit 528 may include an inverse STFT (iSTFT) circuit configured to convert the short window of the audio signal from the frequency domain to the time domain. In some embodiments, a portion of the circuitry of the ear-worn device 500 may be configured to process audio signals in the time domain. In some embodiments, the ear-worn device may lack STFT and iSTFT circuits.

[0036] Deploying noise reduction technology can introduce a delay between the time a sound is emitted from a sound source and the time the noise-reduced sound is output to the user. For example, such technology may introduce a delay between when a speaker speaks and when a listener hears the noise-reduced speech. During face-to-face communication, long latency can cause the perception of echo because both the original sound and the noise-reduced sound are played back to the listener. Long latency can also interfere with how the listener processes the incoming sound due to a mismatch between visual cues (e.g., moving lips) and the arrival of the associated sound. To achieve acceptable latency when implementing neural networks in ear-worn devices, the ear-worn device needs to be able to perform billions of operations per second. To address such demanding power issues, the neural network circuit 524 (in addition to other circuits) can be implemented on the chip of the ear-worn device.

[0037] Accordingly, in some embodiments, one or more of the processing circuit 522, the noise reduction circuit 524 (including the neural network circuit 526), ​​and the processing circuit 528 (or any part thereof) may be mounted on a single identical chip (i.e., a single semiconductor die or substrate) in an ear-worn device. A further description of a chip incorporating a neural network circuit (or, in some embodiments, one of the other elements) used in an ear-worn device can be found in U.S. Patent No. 11,886,974, “Neural Network Chip for Ear-Worn Devices,” issued on 30 January 2024, the entire disclosure of which is incorporated herein by reference.

[0038] (Noise reduction circuit) Figure 6 shows an audio signal 630a according to a specific embodiment described herein. The audio signal 630a (which may be one of the audio signals 630 input to a neural network circuit, as described below) includes a background noise signal 601 and a spoken audio signal 603. The spoken audio signal 603 includes a target spoken audio signal 605 and an interfering spoken audio signal 607.

[0039] Generally, the purpose of a noise reduction circuit (e.g., any of the noise reduction circuits described herein) is to enhance the target speech voice signal 605. This may include, for example, amplifying the target speech voice signal 605 and / or attenuating the background noise signal 601 and the interfering speech voice signal 607. Thus, both background noise and interfering speech voices are considered noise and may be attenuated as part of the noise reduction. The speech voice signal 603 may include the speech voice in the speech signal 630a, and the background noise signal 601 may include the background noise in the speech signal 630a. More specifically, any speech voice that lacks features that can distinguish it from the target speech voice (described further below) may be considered the speech voice signal 603 in the speech signal 630a. The background noise signal 601 may be considered any sound (including speech voices) that contains features that can distinguish it from the target speech voice. Examples of types of speech voices that contain features that can distinguish them from the target speech voice include speech voices such as bubble noise and speech voices heard from a distance.

[0040] In some embodiments, the background noise signal 601 may be equivalent to the remainder when the spoken speech signal 603 is subtracted from the speech signal 630a. From another perspective, the spoken speech signal 603 may be equivalent to the remainder when the background noise signal 601 is subtracted from the speech signal 630a. The above relationship remains true even if the subtraction process is not actually performed. For example, even if the background noise signal 601 is not generated by subtraction but is generated independently or by a different procedure, the background noise signal 601 can still be considered equivalent to the remainder when the spoken speech signal 603 is subtracted from the speech signal 630a.

[0041] The target speech signal 605 can be considered, in a broad sense and qualitatively, as the portion of the speech signal 603 that the wearer of the ear-worn device is most interested in hearing. The interfering speech signal 607 can be considered, in a broad sense and qualitatively, as the portion of the speech signal 603 that the wearer of the ear-worn device is less interested in. However, as mentioned above, it is difficult to determine what the target speech signal 605 is and what the interfering speech signal 607 is. The technology described herein can distinguish between the target speech signal 605 and the interfering speech signal 607 based on the direction of arrival (DOA) to the wearer.

[0042] More specifically, the target speech signal 605 may be a first spatially focused version of the speech signal 603 of the speech signal 630a, and the interfering speech signal 607 may be a second spatially focused version of the speech signal 603 of the speech signal 630a (a different version from the first spatially focused version). The spatially focused version of a signal may be a signal to which a spatial focusing pattern has been applied. In other words, spatial focusing may include applying a spatial focusing pattern (also called a spatial focusing pattern) to the speech signal. The spatial focusing pattern can define different weights (also called gains) applied to the speech signal as a function of the direction of arrival (DOA) of the sound in the speech signal, and the DOA may be defined for the wearer of the ear-mounted device.

[0043] In some embodiments, the weights may be equal to 0, equal to 1, or between 0 and 1. In some embodiments, the weights may be greater than or equal to 0. In some embodiments, the weights may be greater than 0, less than 0, equal to zero, or complex numbers. Negative weights may invert the phase by 180 degrees, and complex number weights may rotate the phase by an angle. Focusing is performed by mapping the weights to the DOA, as higher weights are applied to sounds coming from a particular direction and lower weights are applied to sounds coming from other directions. Applying higher weights in the direction where the target speaker is located (or is assumed to be located) and lower weights in other directions can help focus on the sound from the target speaker and attenuate sounds from other interfering speakers. For example, if the target speaker is located (or is assumed to be located) in front of the wearer of the in-ear device, the spatial focusing pattern may assign higher weights to the DOA in front of the wearer than to the DOA to the sides and behind the wearer.

[0044] In some embodiments, the target speech voice signal 605 is equal to the speech voice signal 603 (of speech signal 630a) to which the first spatial focusing pattern is applied. (As used herein, if signal A is equal to signal B to which the spatial focusing pattern is applied, then signal A may be considered to "have" the spatial focusing pattern.) The first spatial focusing pattern may include different weights applied to the speech voice of the speech voice signal 603 obtained from different directions of arrival (DOA) relative to the wearer of the ear-mounted device. The first spatial focusing pattern may have higher weights for DOA where the target speaker is located or is assumed to be located, and lower weights for other locations. For example, the first spatial focusing pattern may include applying higher weights to speech voice obtained from DOA facing forward of the wearer of the ear-mounted device than to speech voice obtained from DOA facing side and rear of the wearer. The interfering speech signal 607 is equal to the speech signal 603 of speech signal 630a to which the second spatial focusing pattern has been applied. For example, the second spatial focusing pattern may have lower weights for DOA where the target speaker is located or is expected to be located, and higher weights for other locations.

[0045] In some embodiments, the interfering speech signal 607 is equal to the remainder when the target speech signal 605 is subtracted from the speech signal 603. From another viewpoint, the target speech signal 605 is equal to the remainder when the interfering speech signal 607 is subtracted from the speech signal 603. From yet another viewpoint, the second spatial focusing pattern may be the remainder when the first spatial focusing pattern is subtracted from a weighting pattern in which all DOA have a weight of 1. From yet another viewpoint, the first spatial focusing pattern may be the remainder when the second spatial focusing pattern is subtracted from a weighting pattern in which all DOA have a weight of 1. In general, the first spatial focusing pattern and the second spatial focusing pattern may be inverses of each other.

[0046] The above relationship remains true even if no subtraction is actually performed. For example, even if the interfering speech signal 607 is not generated by subtraction but is generated independently or by a different procedure, the interfering speech signal 607 can be considered equal to the remainder when the target speech signal 605 is subtracted from the speech signal 603. The interfering speech signal 607 may represent different interfering speakers weighted based on a second spatial focusing pattern. For example, if the first interfering speaker is at 45 degrees and the second interfering speaker is at 90 degrees, and the second spatial focusing specifies a weight of 0.5 at 45 degrees and a weight of 0.8 at 90 degrees, then the interfering speech signal 607 is equal to the sum of the speech from the first interfering speaker multiplied by 0.5 and the speech from the second interfering speaker multiplied by 0.8.

[0047] A particular spatial focusing pattern may include assigning weights between 0 and 1 to sounds from a particular DOA, so that the target speech signal 605 includes speech from the speaker (weighted by a certain amount), and the interfering speech signal 607 may include speech from the same speaker (weighted by different amounts). However, if the speaker is located in a DOA to which a particular spatial focusing pattern assigns a weight of 1 or 0, then speech from that speaker may only be present in the target speech signal 605 or the interfering speech signal 607.

[0048] The audio signal 630a may be enhanced by reducing the volume of the background noise signal 601 and the interfering speech audio signal 607, and / or increasing the volume of the target speech audio signal 605. The noise reduction circuit (e.g., any of the noise reduction circuits described herein) is generally configured to generate an output audio signal (e.g., output audio signal 840) which includes the target speech audio signal 605, the interfering speech audio signal 607, and the background noise signal 601, and in the output audio signal, the volume of one or more of the background noise signal 601, the interfering speech audio signal 607, and the target speech audio signal 605 may differ from their volume in the audio signal 630a.

[0049] In particular, the change in volume of the background noise signal 601 may differ from the change in volume of the target speech signal 605 by a first volume change difference, and the change in volume of the interfering speech signal 607 may differ from the change in volume of the target speech signal 605 by a second volume change difference. The change in volume can be measured between the volume of the speech signal 630a and the volume of the output signal. In some embodiments, the first volume change difference and the second volume change difference may be different. In some embodiments, the first volume change difference and the second volume change difference may be independently controllable.

[0050] Figure 7 shows volume changes according to some embodiments described herein. Figure 7 shows that the volume of the target speech signal 605 increases by A from the speech signal 630a ("input") to the output speech signal ("output"). The volume of the interfering speech signal 607 decreases by B from the speech signal 630a to the output speech signal. The volume of the background noise signal 601 decreases by C from the speech signal 630a to the output speech signal. The first volume change difference amount described above is CA, and the second volume change difference amount described above may be BA.

[0051] Figure 8 shows a noise reduction circuit 824 in an ear-mounted device according to several embodiments described herein. The ear-mounted device may be any of the ear-mounted devices described herein (e.g., hearing aid 100, eyeglasses 300, ear-mounted device 400, and / or ear-mounted device 500). The noise reduction circuit 824 may be any of the noise reduction circuits described herein (e.g., noise reduction circuit 524). The noise reduction circuit 824 includes a neural network circuit 826 (which may be any of the neural network circuits described herein, e.g., neural network circuit 526), ​​a mask application and subtraction circuit 832, and a mixing circuit 834. The ear-mounted device may include two or more microphones not shown in Figure 8 (e.g., microphones 102f and 102b, microphone 302, and / or microphone 502).

[0052] The neural network circuit 826 may be configured to receive a plurality of audio signals 630 (including audio signal 630a) such that (1) at least two of the plurality of audio signals 630 are obtained from one of two or more microphones of an ear-worn device, and / or (2) at least one of the plurality of audio signals 630 is a beamforming audio signal obtained from two or more microphones. (In this specification, a first signal is said to be obtained from a microphone if the microphone generates the first signal, or if the microphone generates a second signal and the first signal is the result of processing the second signal. In some cases, this processing may be performed on the second signal together with other signals.)

[0053] Regarding option (1), as an example, one of the multiple audio signals 630 may be the output of one of the microphones (e.g., a front microphone such as the front microphone 102f) or a processed version thereof (e.g., the output of the front microphone after processing by the processing circuit 522), and one of the multiple audio signals 630 may be the output of another of the microphones (e.g., a back microphone such as the back microphone 102b) or a processed version thereof (e.g., the output of the back microphone after processing by the processing circuit 522).

[0054] Regarding option (2), as an example, one of the multiple audio signals 630 may be the result of beamforming the outputs of two microphones (for example, beamforming the audio signal obtained from a front microphone such as the front microphone 102f and the audio signal obtained from a back microphone such as the back microphone 102b). The result of beamforming may have a specific directional pattern (for example, as an unspecified example, cardioid, supercardioid, hypercardioid, or dipole). Further explanation of beamforming and directional patterns is given below.

[0055] In some embodiments, at least two (or all) of the plurality of audio signals 630 may have different beamforming directivity patterns. In some embodiments, at least one of the plurality of audio signals 630 may have a forward-directed beamforming directivity pattern, and at least one of the plurality of audio signals 630 may have a rearward-directed beamforming directivity pattern. (As will be further described below, a forward-directed beamforming directivity pattern typically attenuates signals coming from behind the wearer more than signals coming from in front of the wearer, and a rearward-directed beamforming directivity pattern typically attenuates signals coming from in front of the wearer more than signals coming from behind the wearer.) In some embodiments, the plurality of audio signals 630 may include two signals. In some embodiments, the plurality of audio signals 630 may include three signals. In some embodiments, the plurality of audio signals 630 may include four signals. In some embodiments, the plurality of audio signals 630 may include more than four signals. The following are non-limiting examples of sets of audio signals that may be, or may be included in, the plurality of audio signals 630.

[0056] In some embodiments, the plurality of audio signals 630 may include signals having a forward-directed supercardioid directional pattern and signals having a rearward-directed supercardioid directional pattern. In some embodiments, the plurality of audio signals 630 may include signals having a forward-directed cardioid directional pattern and signals having a rearward-directed cardioid directional pattern. In some embodiments, the plurality of audio signals 630 may include signals having a forward-directed hypercardioid directional pattern and signals having a rearward-directed hypercardioid directional pattern. In some embodiments, the plurality of audio signals 630 may include signals having a forward-directed cardioid directional pattern, signals having a rearward-directed cardioid directional pattern, and signals having a dipole directional pattern.

[0057] In some embodiments, the plurality of audio signals 630 may include signals having a forward-directed hypercardioid directivity pattern, signals having a backward-directed hypercardioid directivity pattern, and signals having a dipole directivity pattern. In some embodiments, the plurality of audio signals 630 may be in the frequency domain. In some embodiments, the plurality of audio signals 630 may be in the time domain. In some embodiments, the neural network circuit 826 may be configured to receive the plurality of audio signals 630 together (i.e., not sequentially). In some embodiments, the neural network circuit 826 may be configured to process the plurality of audio signals 630 together (i.e., not sequentially).

[0058] In some embodiments, the neural network circuit 826 may be configured to implement one or more neural network layers trained to perform noise reduction and spatial focusing, and the neural network circuit 826 generates two or more neural network outputs 836 based on a plurality of speech signals 630. (For simplicity, this description may alternately describe the reception of signals and the generation of outputs based on signals performed by the neural network circuit or one or more neural network layers implemented by the neural network circuit.) In some embodiments, the noise reduction circuit 824 may be configured to obtain at least two combinations of (1) the utterance speech signal 603 of the speech signal 630a, (2) the background noise signal 601 of the speech signal 630a, (3) the target utterance speech signal 605 of the speech signal 630a, and (4) the interference utterance speech signal 607 of the speech signal 630a, based on two or more neural network outputs 836. Various ways in which the noise reduction circuit 824 can obtain these signals based on two or more neural network outputs 836 are described below.

[0059] In some embodiments, two or more neural network outputs 836 may include one or more of the following: (1) an output (e.g., a mask) configured to generate an utterance speech signal 603, (2) an output (e.g., a mask) configured to generate a background noise signal 601, (3) an output (e.g., a mask) configured to generate a target utterance speech signal 605, and (4) an output (e.g., a mask) configured to generate an interference utterance speech signal 607. The mask may be a real-valued or complex-valued mask that varies with frequency. Thus, when the mask is applied to a speech signal (e.g., multiplied by or added to the speech signal), the mask may behave differently for different frequency components of the speech signal. In other words, the mask may be such that different frequency components of the speech signal are multiplied by different real-valued or complex-valued values. A real-valued mask may modify only the magnitude, while a complex-valued mask may modify both the magnitude and the phase. If two or more neural network outputs 836 include two masks, the two masks may be different.

[0060] In general, each of the two or more neural network outputs 836 may contain one or more elements, such as one audio signal, multiple audio signals, one mask, multiple masks, one audio signal and one mask, or multiple audio signals and multiple masks.

[0061] Furthermore, with respect to learning, in some embodiments, one or more neural network layers implemented by the neural network circuit 826 may be learned to perform background noise correction. Learning such neural network layers may include obtaining the speech signal of a noisy speech and a speech-separated version of the speech signal (i.e., one in which only the speech remains). In some embodiments, a mask may be determined that, when applied to the speech signal of a noisy speech, results in a speech-separated speech signal. The learning input data may be the speech signal of a noisy speech, and the learning output data may be the mask. Thus, one or more neural network layers may learn how to output a speech-separated mask such that, when the mask is applied to the speech signal 630a (e.g., multiplied by or added to the speech signal 630a), the resulting output speech signal is the speech signal 603, i.e., the speech-separated version of the speech signal 630a.

[0062] In some embodiments, a mask may be determined that, when applied to a noisy speech audio signal, results in a background noise-separated audio signal. The training input data may be the noisy speech audio signal, and the training output data may be the mask. Thus, the neural network layer may learn how to output a background noise-separated mask of speech signal 630a such that when the mask is applied to speech signal 630a (e.g., multiplied by or added to speech signal 630a), the resulting output audio signal is the background noise signal 601, i.e., a background noise-corrected version of speech signal 630a. In embodiments in which one or more neural networks are trained to output a speech-to-speech-separated signal or a noise-separated signal, the output training data may be the speech-to-speech-separated signal or the noise-separated signal itself. Further descriptions of neural networks trained to perform noise correction can be found in U.S. Patent No. 11,812,225, entitled "Methods, Apparatus and Systems for Neural Network Hearing Aids," issued November 7, 2023.

[0063] In some embodiments, one or more neural network layers implemented by the neural network circuit 826 may be trained to perform spatial focusing. Spatial focusing may include applying a spatial focusing pattern to an audio signal. As described above, the spatial focusing pattern may specify different weights as a function of the direction of arrival (DOA) of the sound, the DOA may be defined for the wearer of the in-ear device. In some embodiments, the weights may be equal to 0, equal to 1, or between 0 and 1. In some embodiments, the weights may be greater than or equal to 0. In some embodiments, the weights may be greater than 0, less than 0, equal to zero, or complex numbers. Negative weights may invert the phase by 180 degrees, and complex number weights may rotate the phase by an angle. Focusing is performed by mapping the weights to the DOA, so that higher weights are applied to sounds coming from a particular direction and lower weights are applied to sounds coming from other directions. To train such a neural network layer, a training audio signal may be formed from component audio signals obtained from different DOA. Multiple audio signals obtained from multiple microphones may be generated from a learning audio signal.

[0064] If a neural network is trained to output a mask, the training mask may be determined such that when the training mask is applied to one of several speech signals, each of the remaining component speech signals is multiplied by the weight corresponding to the obtained DOA of that component and then summed up. This allows one or more neural network layers to learn how to output a mask based on several speech signals, and when the mask is applied to (e.g., multiplied or added to) the spoken speech signal 603, the resulting output (target spoken speech signal 605) includes each component of the spoken speech signal 603 multiplied by the weight corresponding to the obtained DOA of that component and then summed up, i.e., a spatially focused version of the spoken speech signal 603. In embodiments where one or more neural networks are trained to output a spatially focused signal, the output training data may be the spatially focused signal itself.

[0065] In some embodiments, one or more neural network layers implemented by the neural network circuit 826 may be trained to perform background noise correction and spatial focusing. To train such neural network layers, a training audio signal may be formed from component audio signals obtained from different DOAs. Multiple audio signals obtained from multiple microphones may be generated from the training audio signal. If the neural network is trained to output a mask, the training mask may be determined such that when the training mask is applied to one of the multiple audio signals, the speech of each remaining component audio signal is multiplied by the weight corresponding to the DOA from which that component was obtained, and then summed. (As described above, the training audio signal may include the audio signal of a noisy speech and a speech-separated version of the audio signal, i.e., the audio signal with only the speech remaining.)

[0066] This allows one or more neural network layers to learn how to output a mask based on multiple audio signals, and when the mask is applied to the audio signal 630a (e.g., multiplied or added), the resulting output (target utterance audio signal 605) includes the sum of the utterances of each component of the audio signal 630a, multiplied by the weight corresponding to the obtained DOA for that component, i.e., a background noise-reduced and spatially focused version of the audio signal 630a. In embodiments where one or more neural networks are trained to output a background noise-reduced and spatially focused signal, the output training data may be the background noise-reduced and spatially focused signal itself.

[0067] In some embodiments, the mask application and subtraction circuit 832 in the noise reduction circuit 824 may be configured to obtain at least two combinations of (1) the utterance speech signal 603 of the speech signal 630a, (2) the background noise signal 601 of the speech signal 630a, (3) the target utterance speech signal 605 of the speech signal 630a, and (4) the interference utterance speech signal 607 of the speech signal 630a, based on two or more neural network outputs 836. In some embodiments, the mask application and subtraction circuit 832 may be configured to obtain one or more of these signals by applying a mask to the speech signal. More specifically, consider that the neural network circuit 826 is configured to generate a mask (i.e., the mask is included in two or more neural network outputs 836).

[0068] One or more neural network layers implemented by the neural network circuit 826 may be trained to generate a mask that allows a portion of the speech signal to be separated by applying the mask to the speech signal (for example, separating speech from background noise or separating a target speech from interfering speech). In some embodiments, the mask application and subtraction circuit 832 may be configured to generate a speech speech signal 603 of speech signal 630a by applying a mask to speech signal 630a. In some embodiments, the mask application and subtraction circuit 832 may be configured to generate a background noise signal 601 of speech signal 630a by applying a mask to speech signal 630a. In some embodiments, the mask application and subtraction circuit 832 may be configured to generate a target speech speech signal 605 of speech signal 630a by applying a mask to a speech speech signal 603 of speech signal 630a. In some embodiments, the mask application and subtraction circuit 832 may be configured to generate an interference speech signal 607 of the speech signal 630a by applying a mask to the speech signal 603 of the speech signal 630a.

[0069] In some embodiments, the mask application and subtraction circuit 832 may be configured to generate a target utterance speech signal 605 of speech signal 630a by applying a mask to speech signal 630a. In some embodiments, the mask application and subtraction circuit 832 may be configured to generate an interference utterance speech signal 607 of speech signal 630a by applying a mask to speech signal 630a. Which signal the mask application and subtraction circuit 832 applies the mask to and which signal is obtained as a result depends on how one or more neural network layers implemented by the neural network circuit 826 have been trained to output the mask. (Since the mask application and subtraction circuit 832 may be configured to apply a mask to speech signal 630a, a dotted line is shown connecting speech signal 630a and the mask application and subtraction circuit 832).

[0070] In some embodiments, the masking and subtraction circuit 832 may be configured to obtain one or more of these signals by subtracting from a specific signal. (However, in some embodiments, other operations such as addition may be used.) In some embodiments, the masking and subtraction circuit 832 may be configured to generate a background noise signal 601 of the voice signal 630a by subtracting the utterance voice signal 603 of the voice signal 630a from the voice signal 630. In some embodiments, the masking and subtraction circuit 832 may be configured to generate the utterance voice signal 603 of the voice signal 630a by subtracting the background noise signal 601 of the voice signal 630a from the voice signal 630a.

[0071] In some embodiments, the masking and subtraction circuit 832 may be configured to generate an interfering speech signal 607 of speech signal 603 by subtracting the target speech signal 605 from the speech signal 603. In some embodiments, the masking and subtraction circuit 832 may be configured to generate the target speech signal 605 by subtracting the interfering speech signal 607 from the speech signal 603. In some embodiments, the masking and subtraction circuit 832 may be configured to generate the speech signal 603 by subtracting the target speech signal 605 and the interfering speech signal 607 from the speech signal 630a.

[0072] In some embodiments, the neural network circuit 826 may be configured to perform spatial focusing on the speech signal 630a such that one of two or more neural network outputs 836 is a first signal including a spatially focused version of the speech signal 630a and background noise of the speech signal 630a without spatial focusing, or an output (e.g., a mask) configured to generate this first signal. Thus, in some embodiments, the mask application and subtraction circuit 832 may be configured to generate the first signal by applying a mask to the speech signal 630a. The neural network circuit 826 may be configured to perform noise correction on the first signal such that one of two or more neural network outputs 836 is a target speech signal 605 (i.e., the first signal with the noise removed), or an output (e.g., a mask) configured to generate the target speech signal 605.

[0073] Therefore, in some embodiments, the mask application and subtraction circuit 832 may be configured to generate a target speech voice signal 605 by applying a mask to the first signal (or voice signal 630a). In some embodiments, the mask application and subtraction circuit 832 may be configured to generate a background noise signal 601 by subtracting the target speech voice signal 605 from the first signal. In some embodiments, the mask application and subtraction circuit 832 may be configured to generate an interference speech voice signal 607 by subtracting the first signal from the voice signal 630a. In some embodiments, the above may be performed with the roles of the target speech voice signal 605 and the interference speech voice signal 607 reversed.

[0074] In some embodiments, the neural network circuit 826 may be configured to perform spatial focusing on the speech signal 630a such that one of two or more neural network outputs 836 is a first signal including a spatially focused version of the speech signal 630a and a spatially focused version of the speech signal 630a with background noise, or an output (e.g., a mask) configured to generate this first signal. Thus, in some embodiments, the mask application and subtraction circuit 832 may be configured to generate the first signal by applying a mask to the speech signal 630a. The neural network circuit 826 may be configured to perform noise correction on the first signal such that one of two or more neural network outputs 836 is a target speech signal 605 (i.e., the first signal with the noise removed), or an output (e.g., a mask) configured to generate the target speech signal 605.

[0075] Accordingly, in some embodiments, the mask application and subtraction circuit 832 may be configured to generate a target speech voice signal 605 by applying a mask to the first signal (or speech signal 630a). In some embodiments, the mask application and subtraction circuit 832 may be configured to generate a signal including spatially focused background noise of the speech signal 630a by subtracting the target speech voice signal 605 from the first signal. In some embodiments, this signal including spatially focused background noise may be considered a background noise signal 601. In some embodiments, the mask application and subtraction circuit 832 may be configured to generate a third signal including the interfering speech voice and the remainder of the background noise not included in the spatially focused background signal by subtracting the first signal 605 from the speech signal 630a. In some embodiments, this third signal may be considered an interfering speech voice signal 607. In some embodiments, the above may be performed with the roles of the target speech signal 605 and the interfering speech signal 607 reversed.

[0076] Therefore, in some embodiments, the interfering speech audio signal 607 may include the interfering speech but not background noise, while in other embodiments, the interfering speech audio signal 607 may include the interfering speech and some background noise. In some embodiments, the background noise signal 601 may include all background noise (or an estimate of all background noise), while in other embodiments, the background noise signal 601 may include some background noise but not all background noise. From another perspective, in some embodiments, spatial focusing may not be applied to the background noise signal 601, and the interfering speech audio signal 607 may not include some of the background noise of the audio signal 630a. In some embodiments, the background noise signal 601 includes a first spatially focused version of the background noise of the audio signal 630a, and the interfering speech audio signal 608 includes the interfering speech (i.e., a spatially focused version of the speech audio signal 603) and a second spatially focused version of the background noise of the audio signal 630a.

[0077] The noise reduction circuit 824 may be configured to generate an output audio signal 840 that includes the target speech audio signal 605, the interfering speech audio signal 607, and the background noise signal 601, after combining at least two of the following using at least two neural network outputs 836: (1) the utterance speech signal 603 of the audio signal 630a, (2) the background noise signal 601 of the audio signal 630a, (3) the target speech audio signal 605 of the audio signal 630a, and (4) the interfering speech audio signal 607 of the audio signal 630a. In some embodiments, in order to generate the output audio signal 840 that includes the target speech audio signal 605, the interfering speech audio signal 607, and the background noise signal 601, the noise reduction circuit 824 needs to acquire at least one of the target speech audio signal 605 and the interfering speech audio signal 607.

[0078] In other words, in some embodiments, effective combinations may include the speech voice signal 603 and the target speech voice signal 605, the speech voice signal 603 and the interfering speech voice signal 607, the background noise signal 601 and the target speech voice signal 605, the background noise signal 601 and the interfering speech voice signal 607, and the target speech voice signal 605 and the interfering speech voice signal 607. To put it another way, in some embodiments, the noise reduction circuit 824 may be configured to acquire at least one of the target speech voice signal 605 and the interfering speech voice signal 607, as well as at least one of the speech voice signal 603 and the background noise signal 601, while in some embodiments, the noise reduction circuit 824 may be configured to acquire the target speech voice signal 605 and the interfering speech voice signal 607.

[0079] The noise reduction circuit 824 may be configured to generate output audio signal 840a such that the volume of one or more of the background noise signal 601, the interfering speech audio signal 607, and the target speech audio signal 605 differs from their volumes in audio signal 630a. In particular, as shown in Figure 7, the change in volume of the background noise signal 601 may differ from the change in volume of the target speech audio signal 605 by a first volume change difference, and the change in volume of the interfering speech audio signal 607 may differ from the change in volume of the target speech audio signal 605 by a second volume change difference. The volume change can be measured between the volume in audio signal 630a and the volume in output signal 840. In some embodiments, the first volume change difference and the second volume change difference may be different. In some embodiments, the first volume change difference and the second volume change difference may be independently controllable.

[0080] More specifically, if the target speech voice signal 605 is denoted as TS, the interfering speech voice signal 607 as IS, and the background noise signal 601 as BN, then in some embodiments, the noise reduction circuit 834 may be configured to generate an output voice signal 840 equivalent to a*TS+b*IS+c*BN. In some embodiments, the interfering speech voice weight b and the background noise weight c may have values ​​between 0 and 1. The target speech voice weight a typically has a value of 1, but other values ​​(e.g., values ​​greater than 1 or values ​​less than 1) may also be used. Thus, in some embodiments, the output voice signal 840 may have reduced levels of background noise and interfering speech voice. Adding some background noise and some interfering speech voice to the target speech voice may help reduce distortion and improve the wearer's environmental awareness of the in-ear device. In some embodiments, the mixing circuit 834 may be configured to generate the output voice signal 840 by mixing. In this specification, mixing should be understood as any combination of different elements after assigning weights to the different elements. Therefore, the mixing circuit 834 may be configured to assign different weights (e.g., multiply) to signals and add the results. The mixing performed by the mixing circuit 834 may be considered interpolation.

[0081] The output audio signal 840 can be generated to be equivalent to a*TS+b*IS+c*BN by multiplying the target speech audio signal 605 by a, multiplying the interfering speech audio signal 607 by b, multiplying the background noise signal 601 by c, and adding these products. However, the same result may be obtained by mixing other signals. This is true under the assumption that audio signal 630a corresponds to the sum of the speech audio signal 603 and the background noise signal 601, and that the speech audio signal 603 corresponds to the sum of the target speech audio signal 605 and the interfering speech audio signal 607. As a non-restrictive example, consider multiplying the target speech audio signal 605 by d, multiplying the interfering speech audio signal 607 by e, multiplying audio signal 630a by f, and adding these products. If audio signal 630a is O(O for Original) and the speech audio signal is S, then the following can be shown.

[0082] d*TS+e*IS+f*O=d*TS+e*IS+f*(S+BN) =d*TS+e*IS+f*(TS+IS)+f*BN =(d+f)*TS+(e+f)*IS+f*BN Next, in the equation a*TS+b*IS+c*BN, the weights a, b, and c may have the following relationships with the weights d, e, and f: a=d+f, b=e+f, and c=f.

[0083] As described above, the change in volume of the background noise signal 601 may differ from the change in volume of the target speech signal 605 by a first volume change difference, and the change in volume of the interfering speech signal 607 may differ from the change in volume of the target speech signal 605 by a second volume change difference. The change in volume can be measured between the volume of the speech signal 630a and the volume of the output signal 840. As will be further explained below, the first volume change difference and the second volume change difference can be controlled independently by controlling one or more of the weights assigned to the target speech signal 605, the background noise signal 601, and the interfering speech signal 607 in the output speech signal 840. In embodiments where the weight assigned to the target speech signal 605 is always 1, or 1 by default, the first volume change difference and the second volume change difference can be controlled independently by controlling the weight assigned to the background noise signal 601 and the weight assigned to the interfering speech signal 607 in the output speech signal 840. (Alternatively, either the weight assigned to the background noise signal 601 or the weight assigned to the interfering speech signal 607 is always 1, or 1 by default, and the first volume change difference and the second volume change difference can be controlled independently by controlling the weights assigned to other signals.)

[0084] In some embodiments, weights may be assigned by directly assigning weights to the background noise signal 601 and the interfering speech signal 607. For example, if the weight a of the target speech signal 605 is 1 and the weight c of the background noise signal 601 is 0.18, then the volume change of the target speech signal 605 may be 0 dB, the volume change of the background noise signal 601 may be approximately -15 dB (by subtracting the volume at speech signal 630a from the volume at output speech signal 840), and the difference in volume change may be -15 dB (i.e., the first volume change difference is -15 dB). In some embodiments, weights may be assigned indirectly as described above by assigning weights to other speech signals related to the background noise signal and the interfering speech signal. For example, if the weight f assigned to the speech signal 630a is 0.18 and the weight d assigned to the target speech signal 605 is 0.82, then the volume change of the target speech signal 605 is 0 dB, the volume change of the background noise signal 601 is approximately -15 dB, and the difference in volume changes may be -15 dB (i.e., the first volume change difference is -15 dB).

[0085] In some embodiments, if the value selected for the first volume change difference does not restrict the values ​​that can be selected for the second volume change difference, the first and second volume change difference can be considered independently controllable. For example, if the output audio signal 840 is of the form a*TS+b*IS+c*BN, in some embodiments, if the value a selected for one of the weights a, b, or c does not restrict the values ​​that can be selected for the other weights, the first and second volume change difference can be considered independently controllable.

[0086] Generally, the noise reduction circuit 824 may be configured to generate an output audio signal 840 using a combination of audio signals. In some embodiments, the combination of audio signals is at least three signals 838, the at least three signals 838 being a set or subset of the audio signal 630a ("O"), the speech audio signal 603 ("S"), the background noise signal 601 ("BN"), the target speech audio signal 605 ("TS"), and the interfering speech audio signal 607 ("IS"). (In some embodiments, the audio signal 630a may be used by the mixing circuit 834, so the audio signal 630a is shown as an input to the mixing circuit 834 by a dotted line.) Effective combinations of these signals may include at least TS IS BN, TS S BN, IS S BN, TS IS O, TS SO, IS SO, TS O BN, and IS O BN. In some embodiments, the mixing circuit 834 may be configured to generate an output audio signal 840 by mixing a combination of audio signals, such as one of the combinations of at least three signals 838 described above.

[0087] As described above, the noise reduction circuit 824 may be configured to generate an output audio signal 840 including a target speech audio signal 605, an interfering speech audio signal 607, and a background noise signal 601 using at least two or more neural network outputs 836. From the above, the two or more neural network outputs 836 may include one or more of the target speech audio signal 605, the interfering speech audio signal 607, and the background noise signal 601, or may include an output capable of generating the target speech audio signal 605, the interfering speech audio signal 607, and the background noise signal 601. Such an output may be a mask, or it may be an output capable of deriving the target speech audio signal 605, the interfering speech audio signal 607, and / or the background noise signal 601 (for example, by subtraction as described above).

[0088] As can be understood from the above, the noise reduction circuit 824 may be configured to use the first audio signal 630a, i.e., the beamforming audio signal, when generating the output audio signal 840. For example, the masking and subtraction circuit 832 may be configured to apply a mask to the first audio signal 630a, and / or the mixing circuit 834 may be configured to use the first audio signal 630a for mixing. With respect to masking, in some embodiments, the masking and subtraction circuit 832 may be configured to apply a mask to the beamforming audio signal, and in some embodiments, the masking and subtraction circuit 832 may be configured to apply a mask to the non-beamforming signal.

[0089] In some embodiments, two or more neural network outputs 836 may include one or more of the speech voice signal 603, background noise signal 601, target speech voice signal 605, and interfering speech voice signal 607. In other words, the neural network circuit 826 may be configured to directly output one or more of these signals. In embodiments where the neural network circuit 826 directly outputs signals rather than masks, the mask application and subtraction circuit 832 may include only the subtraction circuit. In some embodiments, mask application can result in all signals that need to be generated. In such embodiments, the mask application and subtraction circuit 832 may include only the mask application circuit. In some embodiments, the neural network circuit 826 may be configured to directly output all signals that need to be generated. In such embodiments, the mask application and subtraction circuit 832 may not be present.

[0090] The mixing circuit 834 generates all the signals used. In such embodiments, in some cases the mixing circuit 834 may be configured to mix two or more masks, and the mask application and subtraction circuit 832 may be configured to apply the mixed mask to the audio signal. Such operation may be equivalent to independently applying two or more masks to the audio signal and then mixing the results. Mixing masks may involve weighting different masks and combining (e.g., adding) the weighted masks. In such embodiments, the mixing circuit 834 may be incorporated into the mask application and subtraction circuit 832. Thus, in some embodiments, the mixing circuit 834 may be configured to generate the output audio signal 840 by mixing multiple (e.g., at least two) masks. If the utterance voice signal 603 is "S", the background noise signal 601 is "BN", the target utterance voice signal 605 is "TS", and the interfering utterance voice signal 607 is "IS", then the multiple masks may include masks configured to produce TS, IS, BN;TS, S, BN;IS, S, BN;TS and IS;TS and S;IS and S;TS and BN;IS and BN. In some embodiments, if only two masks are mixed, the result may be applied to the voice signal 630a and then mixed with the voice signal 630a.

[0091] Some or all of the utterance speech signal 603, background noise signal 601, target utterance speech signal 605, and interference utterance speech signal 607 are generated using one or more neural network layers, and one or more neural network layers can be trained to output estimations. Thus, in some embodiments, the utterance speech signal 603 does not need to include all of the utterances of speech signal 630a, the background noise signal 601 does not need to include all of the background noise of speech signal 630a, the target utterance speech signal 605 does not need to include all of the target utterances of speech signal 603, and the interference utterance speech signal 607 does not need to include all of the interference utterances of speech signal 603.

[0092] Figure 9 shows a neural network circuit 926 and a mask application and subtraction circuit 932 according to a specific embodiment described herein. The neural network circuit 926 may be an example of any neural network circuit described herein (e.g., neural network circuits 526 and / or 826). The mask application and subtraction circuit 932 may be an example of any processing circuit described herein (e.g., mask application and subtraction circuit 832).

[0093] The neural network circuit 926 includes a circuit configured to implement multiple neural network layers, as shown in Figure 9, as a first subset of neural network layers (i.e., one or more layers) 950a and a second subset of neural network layers (i.e., one or more layers) 950b. In some embodiments, such a circuit may include a multiply-accumulate circuit configured to perform multiply-accumulate operations on an input activation vector and a neural network weight matrix as part of the processing of one or more neural network layers. Further descriptions of the neural network circuit can be found in U.S. Patent No. 11,886,974, published January 30, 2024, entitled "Neural Network Chip for Ear-In-Ear Devices," the entire disclosure of which is incorporated herein by reference. The mask application and subtraction circuit 932 includes multipliers 952a, 952b, subtractor 954a, and subtractor 954b.

[0094] In some embodiments, the neural network circuit 926 may be configured to generate one of two or more neural network outputs 836 using a first subset of the neural network layer 950a, and to generate the other of two or more neural network outputs 836 using a second subset of the neural network layer 950b. Various options for the neural network output 836 have been described above, and different embodiments may generate different combinations of these options in the first subset of the neural network layer 950a and the second subset of the neural network layer 950b. For example, the neural network circuit 926 may be configured to generate a spoken speech signal 603, an output configured to generate the spoken speech signal 603 (e.g., a mask), a background noise signal 601, and / or an output configured to generate the background noise signal 601 (e.g., a mask) using the first subset of the neural network layer 950a.

[0095] Continuing this example, the neural network circuit 926 may be configured to use a second subset of the neural network layer 950b to generate the target speech signal 605, an output (e.g., a mask) configured to generate the target speech signal 605, the interfering speech signal 607, and / or an output (e.g., a mask) configured to generate the interfering speech signal 607. In the example in Figure 9, the first output of the neural network output 836 is a mask 956a configured to generate the speech signal 603, and the second output of the neural network output 836 is a mask 956b configured to generate the target speech signal 605.

[0096] A first subset of the neural network layer 950a implemented by the neural network circuit 926 may be configured to receive an audio signal 630a. The audio signal 630a may be obtained from one or more microphones of an ear-worn device (e.g., microphones 102f and 102b, microphone 302, and / or microphone 502). For example, the audio signal 630a may be a beamformed version of signals from two different microphones that have been processed (e.g., by masking and subtraction circuit 832). Alternatively, the audio signal 630a may be a processed version of a signal from one microphone that has been processed (e.g., by masking and subtraction circuit 832). A first subset of the neural network layer 950a implemented by the neural network circuit 926 may be configured to produce an output configured to generate a spoken audio signal 603 based on the audio signal 630a. In the example in Figure 9, the output configured to generate the spoken audio signal 603 is the mask 956a. The multiplier 952a may be configured to generate a spoken audio signal 603 by multiplying the audio signal 630a by a mask 956a.

[0097] In some embodiments, a first subset of the neural network layer 950a, implemented by the neural network circuit 926, may be trained to perform background noise correction. A further description of the neural network training is provided above. Based on the training, the first subset of the neural network layer 950a learns how to output a speech-to-speech separation mask 956a for the speech signal 630a, so that when the multiplier 952a multiplies the speech signal 630a by the mask 956a, the resulting output speech signal is the speech-to-speech signal 603, i.e., the speech-to-speech separation version of the speech signal 630a. A further description of the neural network trained to perform noise correction can be found in U.S. Patent No. 11,812,225, published November 7, 2023, entitled "Method, Apparatus and System for Neural Network Hearing Aids".

[0098] A second subset of the neural network layer 950b implemented by the neural network circuit 926 may be configured to receive the plurality of speech signals 630 such that each of at least two of the plurality of speech signals 630 is obtained from one different microphone of the ear-worn device (e.g., microphones 102f and 102b, microphone 302, and / or microphone 502), and / or at least one of the plurality of speech signals is a beamformed version of the speech signal obtained from the microphone. In the example in Figure 9, the plurality of speech signals 630 includes a speech signal 603, a speech signal 630a, and one or more other speech signals 630b. In some embodiments, a particular speech signal among the plurality of speech signals 630 may have a directional pattern formed by beamforming the speech signals from different microphones.

[0099] In some embodiments, certain (e.g., at least two) of the multiple audio signals 630, or each of the multiple audio signals 630, may have different beamforming directivity patterns. In some embodiments, at least one (or all) of the audio signals 630a and / or audio signals 630b may have a directivity pattern formed by beamforming audio signals from different microphones. In some embodiments, at least one (or all) of the audio signals 630a and audio signals 630b may each have different beamforming directivity patterns. Examples of beamforming directivity patterns include dipole, hypercardioid, supercardioid, and cardioid. In some embodiments, at least one of the audio signals 630 may have a forward-directed beamforming directivity pattern, and at least one of the audio signals 630 may have a backward-directed beamforming directivity pattern.

[0100] In some embodiments, the voice signal 630a or one of the voice signals may be a dipole. In some embodiments, one of the voice signals 630a or 630b may be a forward-directed hypercardioid. In some embodiments, one of the voice signals 630a or 630b may be a forward-directed supercardioid. In some embodiments, one of the voice signals 630a or 630b may be a forward-directed supercardioid. In some embodiments, the plurality of voice signals 630 may include two signals. In some embodiments, the plurality of voice signals 630 may include three signals. In some embodiments, the plurality of voice signals 630 may include four signals.

[0101] In some embodiments, the multiple audio signals 630 may include more than four signals. In some embodiments, a second subset of the neural network layer 950b may receive only the utterance audio signal 603 and audio signal 630a as input (i.e., it may not receive audio signal 630b). In some embodiments, a second subset of the neural network layer 950b may receive only the utterance audio signal 603 and audio signal 630b as input (i.e., it may not receive audio signal 630a). In some embodiments, a particular audio signal among audio signals 630a and / or audio signals 630b may be a non-beamforming signal.

[0102] The following are non-limiting examples of what the plurality of speech signals 630 may be, or what set of speech signals may be included in the plurality of speech signals 630. In some embodiments, the plurality of speech signals 630 may include signals having a forward-directed supercardioid directional pattern and signals having a backward-directed supercardioid directional pattern. In some embodiments, the plurality of speech signals 630 may include signals having a forward-directed cardioid directional pattern and signals having a backward-directed cardioid directional pattern. In some embodiments, the plurality of speech signals 630 may include signals having a forward-directed hypercardioid directional pattern and signals having a backward-directed hypercardioid directional pattern. In some embodiments, the plurality of speech signals 630 may include signals having a forward-directed supercardioid directional pattern, signals having a backward-directed supercardioid directional pattern, and a spoken speech signal 603.

[0103] In some embodiments, the plurality of voice signals 630 may include a signal having a forward-directed cardioid directional pattern, a signal having a backward-directed cardioid directional pattern, and a speech voice signal 603. In some embodiments, the plurality of voice signals 630 may include a signal having a forward-directed hypercardioid directional pattern, a signal having a backward-directed hypercardioid directional pattern, and a speech voice signal 603. In some embodiments, the plurality of voice signals 630 may include a signal having a forward-directed supercardioid directional pattern, a signal having a backward-directed supercardioid directional pattern, a signal having a dipole directional pattern, and a speech voice signal 603. In some embodiments, the plurality of voice signals 630 may include a signal having a forward-directed cardioid directional pattern, a signal having a backward-directed cardioid directional pattern, a signal having a dipole directional pattern, and a speech voice signal 603.

[0104] In some embodiments, the plurality of voice signals 630 may include a signal having a forward-directed hypercardioid directional pattern, a signal having a rearward-directed hypercardioid directional pattern, a signal having a dipole directional pattern, and a spoken voice signal 603. In some embodiments, the plurality of voice signals 630 may include a signal having a forward-directed supercardioid directional pattern, a signal having a rearward-directed supercardioid directional pattern, and a signal having a dipole directional pattern. In some embodiments, the plurality of voice signals 630 may include a signal having a forward-directed cardioid directional pattern, a signal having a rearward-directed cardioid directional pattern, and a signal having a dipole directional pattern. In some embodiments, the plurality of voice signals 630 may include a signal having a forward-directed hypercardioid directional pattern, a signal having a rearward-directed hypercardioid directional pattern, and a signal having a dipole directional pattern.

[0105] A second subset of the neural network layer 950b, implemented by the neural network circuit 926, may be configured to produce an output configured to generate a target speech signal 605 based on a plurality of speech signals 630. In the example in Figure 9, the output configured to generate the target speech signal 605 is the mask 956b. The multiplier 952b may be configured to generate the target speech signal 605 by multiplying the speech signal 603 by the mask 956b.

[0106] In some embodiments, a second subset of the neural network layer 950b, implemented by the neural network circuit 926, may be trained to perform spatial focusing. A further description of the training is provided above. Based on the training, the second subset of the neural network layer 950b may learn how to output the mask 956b based on a plurality of speech signals 630 such that when the multiplier 952b multiplies the speech signal 603 by the mask 956b, the resulting output (target speech signal 605) includes multiplying each component speech signal by a weight corresponding to the DOA from which the signal was obtained, and then summing them up. By assigning higher weights in the direction where the target speaker is located and lower weights in other directions, it may be helpful to focus on the sound from the target speaker and attenuate sounds from other interfering speakers. Thus, the output (target speech signal 605) may be broadly considered to be a signal containing the target speech.

[0107] The mask application and subtraction circuit 932 may be further configured to generate a background noise signal 601 and an interfering speech voice signal 607. In the example in Figure 9, subtractor 954a may be configured to generate the background noise signal 601 by subtracting the speech voice signal 603 from the voice signal 630a. Subtractor 954b may be configured to generate the interfering speech voice signal 607 by subtracting the target speech voice signal 605 from the speech voice signal 603.

[0108] As can be understood from the above, mask 956a is configured to generate the utterance speech signal 603, and mask 956b is configured to generate the target utterance speech signal 605, and masks 956a and 956b may be different. In some embodiments, mask 956b may be applied to one of the speech signals 630a or 630b, rather than being applied to the utterance speech signal 603.

[0109] In general, the audio signals and training data for one or more neural network layers trained to perform spatial focusing (e.g., the second subset 950b) must contain spatial information (in other words, enough information for the model to estimate the location of the sound source). At a minimum, the audio signals and training data must contain multiple (i.e., at least two) audio signals, at least two of which are obtained from one different microphone of two or more, and / or at least one of which is a beamformed version of an audio signal obtained from two or more microphones.

[0110] Therefore, the training data for one or more of these neural network layers typically needs to include audio signals obtained from different microphones. Two methods for generating training data localized by multiple microphones may include (1) collecting sounds from sound sources in different DOA using multiple microphones, and (2) using speech simulation to synthesize multiple microphone signals as if the sound source were localized in a specific DOA. Synthesizing the training signal may involve adding directionality to different sound sources (speech signals and noise) in the simulation and adding them together to create a new signal with spatial speech. The neural network can be trained on either the synthesized data or the recorded data, or both.

[0111] The noise reduction circuit 924 includes multipliers 952a and 952b for multiplying the speech signal by the respective masks 956a and 956b, although in some embodiments other operations (e.g., addition) may be used to combine the masks and the signal. Furthermore, rather than one or more neural network layers implemented by the neural network circuit (e.g., neural network circuit 926) being configured to produce an output (e.g., a mask such as mask 956a or 956b) configured to produce a specific signal, in some embodiments one or more neural network layers may be configured to produce the specific signal itself. For example, in some embodiments, a first subset 950a of the neural network layers implemented by the neural network circuit 926 may be configured to produce a speech speech signal 603, and a second subset 950b of the neural network layers implemented by the neural network circuit 926 may be configured to produce a target speech speech signal 605.

[0112] As described above, one or more neural network layers implemented by the neural network circuit 926, in particular a first subset 950a of the neural network layers, may be configured to produce a spoken speech signal 603, or an output (e.g., a mask 956a) configured to produce the spoken speech signal 603. This includes cases where, when the mask is applied to a speech signal (e.g., speech signal 630a), the spoken speech signal 603 is obtained, as well as cases where, when the mask is applied to a speech signal (e.g., speech signal 630a), a background noise signal 601 is obtained, and then a subtractor is used to generate the spoken speech signal 603 from the background noise signal 601 (e.g., by subtracting the background noise signal 601 from the speech signal 630a). This may also include cases where one or more neural network layers directly generate the spoken audio signal 603, and cases where one or more neural network layers directly generate the background noise signal 601, and then a subtractor is used to generate the spoken audio signal 603 from the background noise signal 601 (for example, by subtracting the background noise signal 601 from the audio signal 630a).

[0113] Similarly, as described above, one or more neural network layers implemented by the neural network circuit 926, in particular a second subset 950b of the neural network layers, may be configured to produce a target speech voice signal 605, or an output (e.g., a mask 956b) configured to produce the target speech voice signal 605. This may include cases where, when the mask is applied to a speech signal (e.g., a speech voice signal 603), the target speech voice signal 605 is obtained, as well as cases where, when the mask is applied to a speech signal (e.g., a speech voice signal 603), an interfering speech voice signal 607 is obtained, and then a subtractor is used to generate the target speech voice signal 605 from the interfering speech voice signal 607 (e.g., by subtracting the interfering speech voice signal 607 from the speech voice signal 603). This may also include cases where one or more neural network layers directly generate the target speech signal 605, and cases where one or more neural network layers directly generate the interference speech signal 607, and then a subtractor is used to generate the target speech signal 605 from the interference speech signal 607 (for example, by subtracting the interference speech signal 607 from the speech signal 603).

[0114] From the above description of Figure 9, the first subset 950a of the neural network layer is configured to generate a mask 956a which can be configured to generate a spoken speech signal 603, and the spoken speech signal 603 may be an input to the second subset 950b of the neural network layer which can be configured to generate a mask 956b. In other words, the mask 956a may be generated before the mask 956b. Therefore, the processing of the first subset 950a of the neural network layer may occur before the processing of the second subset 950b of the neural network layer.

[0115] In some embodiments, the same circuit may be used to process a first subset 950a of the neural network layer and a second subset 950b of the neural network layer. For example, processing the neural network layer may involve using a multiplier-adder circuit (MAC) configured to multiply input activations by neural network weights. In some embodiments, the same MAC may be used to process a first subset 950a of the neural network layer and a second subset 950b of the neural network layer, but the neural network weights used by the MAC may be different. That is, the first subset 950a of the neural network layer may use different weights than the second subset 950b of the neural network layer. In some embodiments, different circuits may be used to process a first subset 950a of the neural network layer and a second subset 950b of the neural network layer.

[0116] Figure 10 shows a neural network circuit 1026 and a mask application and subtraction circuit 1032 according to several embodiments described herein. The neural network circuit 1026 may be an example of any neural network circuit described herein (e.g., neural network circuits 526 and / or 826). The mask application and subtraction circuit 1032 may be an example of any processing circuit described herein (e.g., mask application and subtraction circuit 832).

[0117] The neural network circuit 1026 includes a circuit configured to implement one or more neural network layers. One or more neural network layers 1050 implemented by the neural network circuit 1026 may be configured to receive a plurality of speech signals 630. Generally, one or more neural network layers 1050 may be configured to generate an utterance speech signal 603, or an output configured to generate the utterance speech signal 603, and a target utterance speech signal 605, or an output configured to generate the target utterance speech signal 605, based on the plurality of speech signals 630. In the particular example of Figure 10, one or more neural network layers 1050 may be configured to generate masks 956a and 956b based on the plurality of speech signals 630. As described above, mask 956a may be configured to generate the utterance speech signal 603 (by multiplying the speech signal 630a by mask 956a using multiplier 956a). The mask 1056b may be configured to generate the target utterance audio signal 605 by multiplying the audio signal 630a by the mask 1056b using the multiplier 952b. A background noise signal 601 may be generated from the utterance audio signal 603 (by subtracting the utterance audio signal 603 from the audio signal 630a using the subtractor 954a), and an interference utterance audio signal 607 may be generated from the target utterance audio signal 605 (by subtracting the target utterance audio signal 605 from the utterance audio signal 603 using the subtractor 954b).

[0118] In some embodiments, one or more neural network layers 1050 may be configured to generate a mask configured to generate a background noise signal 601 and / or a mask configured to generate an interfering speech signal 607. Furthermore, a speech signal 603 may be generated from the background noise signal 601 and / or a target speech signal 605 may be generated from the interfering speech signal 607. In some embodiments, masks 956a and 1056b may be generated simultaneously. In some embodiments, one or more neural network layers 1050 may be configured to directly generate several combinations of the speech signal 603, the target speech signal 605, the background noise signal 601, and the interfering speech signal 607. In some embodiments, the signals generated by one or more neural network layers 1050 may be generated simultaneously.

[0119] Therefore, the utterance speech signal 603 does not necessarily have to be an input to a neural network layer configured to generate the mask 1056b. One or more neural network layers 1050 implemented by the neural network circuit 1026 can be considered to be trained to perform background noise correction (to generate the utterance speech signal 603 from the speech signal 630a) and to perform background noise correction and spatial focusing (to generate the target utterance speech signal 605 from the speech signal 630a).

[0120] The above description of Figures 9 and 10 illustrates how the speech voice signal 603, background noise signal 601, target speech voice signal 605, and interference speech voice signal 607 can be generated. Furthermore, as mentioned above, some combinations of these signals and the speech signal 630a can be mixed, for example, by a mixing circuit 834.

[0121] Returning to Figure 8, the above description of Figure 8 described the case where the neural network circuit 826 outputs two or more neural network outputs 836. Generally, as described above, the in-ear device includes two or more microphones, the noise reduction circuit 824 includes the neural network circuit 826, and the neural network circuit 826 may be configured to receive a plurality of audio signals 630 such that (1) at least two of the plurality of audio signals 630 are obtained from one of the two or more microphones of the in-ear device, and / or (2) at least one of the plurality of audio signals 630 is a beamforming audio signal obtained from the two or more microphones. The neural network circuit 826 may be configured to implement one or more neural network layers that have been trained to perform noise correction and / or spatial focusing so that the neural network circuit 826 generates one or more neural network outputs 836 based on the plurality of audio signals 630.

[0122] For example, one or more neural network outputs 836 may be one or more audio signals, one or more outputs (e.g., masks and / or sound maps) configured to generate an output audio signal, or a combination thereof. The noise reduction circuit 824 may be configured to output an output audio signal 840, which is a noise-corrected and / or spatially focused version of the audio signal 630a (i.e., one of the multiple audio signals 630) based on one or more neural network outputs 836. In particular, one or more neural network outputs 836 may include an output audio signal which is a background noise-corrected and spatially focused version of the audio signal 630a, or one or more neural network outputs 836 may be configured to be used by the noise reduction circuit 824 to generate an output audio signal which is a background noise-corrected and spatially focused version of the audio signal 630a.

[0123] In some embodiments, the noise reduction circuit 824 may be configured to perform background noise correction. In other words, one or more neural network layers may be trained to perform background noise correction, and the output audio signal 840 may include a version of the audio signal 630a with background noise corrected (e.g., the spoken audio signal 603 and / or the background noise signal 601). In some embodiments, the noise reduction circuit 824 may be configured to perform spatial focusing. In other words, one or more neural network layers may be trained to perform spatial focusing, and the output audio signal 840 may include a version of the audio signal 630a with spatial focusing applied (e.g., a version of the audio signal 630a with spatial focusing applied to both the spoken audio and background noise, or simply a version with spatial focusing applied to the spoken audio, i.e., the target spoken audio signal 605).

[0124] In some embodiments, the noise reduction circuit 824 may be configured to perform background noise correction and spatial focusing. In other words, one or more neural network layers may be trained to perform background noise correction and spatial focusing, and the output audio signal 840 may include a background noise-corrected and spatially focused version of the audio signal 630a (e.g., the target utterance audio signal 605).

[0125] If the output audio signal 840 includes a spatially focused version of the audio signal 630a, or a version of the audio signal 630a with background noise correction and spatial focusing, the spatially focused portion of the output audio signal 840 may have a specific spatial focusing pattern. All descriptions herein regarding spatial focusing patterns and control of spatial focusing patterns (for example, in the context of the target utterance audio signal 605) may also apply to the output audio signal 840 in such a scenario.

[0126] In embodiments where the neural network output 836 is a mask, the mask application and subtraction circuit 832 may be configured to apply the mask (e.g., multiply or add) to the speech signal (e.g., speech signal 630a). (Therefore, a dotted line is shown connecting the speech signal 630a and the mask application and subtraction circuit 832). In some embodiments, when the mask is applied to the speech signal 630a, an utterance speech signal 603 of the speech signal 630a is obtained. In some embodiments, when the mask is applied to the speech signal 630a, a background noise signal 601 of the speech signal 630a is obtained. In some embodiments, when the mask is applied to the speech signal 630a, a spatially focused version of the speech signal 630a is obtained (e.g., a target utterance speech signal 605 or an interference utterance speech signal 607).

[0127] In some embodiments, the masking and subtraction circuit 832 may be configured to generate a second audio signal based on a first audio signal (for example, the neural network output 836, or, if the neural network output 836 is a mask, the audio signal generated from the neural network output 836). In some embodiments, the masking and subtraction circuit 832 may be configured to generate a background noise signal 601 of the audio signal 630a by subtracting the utterance audio signal 603 of the audio signal 630a from the audio signal 630a. In some embodiments, the masking and subtraction circuit 832 may be configured to generate the utterance audio signal 603 of the audio signal 630a by subtracting the background noise signal 601 of the audio signal 630a from the audio signal 630a.

[0128] In some embodiments, the masking and subtraction circuit 832 may be configured to generate an audio signal containing background noise and interfering speech (e.g., an audio signal containing the background noise signal 601 in addition to the interfering speech signal 607) by subtracting from the audio signal 630a a version of the audio signal 630a that has been modified for background noise and spatial focusing (e.g., a target speech signal 605). In some embodiments, the masking and subtraction circuit 832 may be configured to generate the interfering speech signal 607 by subtracting from the speech signal 603 the target speech signal 605, or to generate the target speech signal 605 by subtracting from the speech signal 603 the interfering speech signal 607.

[0129] As described above, in some embodiments, the output audio signal 840 may include a version of the audio signal 630a with background noise correction, a version with spatial focusing, or a version with both background noise correction and spatial focusing. In some embodiments, the output audio signal 840 may include a version of the audio signal 630a with background noise correction, a version with spatial focusing, or a version with both background noise correction and spatial focusing. In some embodiments, the output audio signal 840 may include a version of the audio signal 630a with background noise correction, a version with spatial focusing, or a version with both background noise correction and spatial focusing, mixed with one or more other audio signals.

[0130] In some embodiments, the mixing circuit 834 may be configured to mix two or more audio signals such that the output audio signal 840 is equivalent to an audio signal 630a mixed with other audio signals (as non-limiting examples, the speech audio signal 603, the background noise signal 601, the target speech audio signal 605, the interfering speech audio signal 607, and / or audio signal 630a) in which the background noise has been corrected and / or the spatially focused version (as non-limiting examples, the speech audio signal 603, the background noise signal 601, the target speech audio signal 605, and / or interfering speech audio signal 607). In some embodiments, the mixing circuit 834 may be configured to mix the speech audio signal 603 of audio signal 630a with the background noise signal 601 of audio signal 630a. In some embodiments, the mixing circuit 834 may be configured to mix the speech audio signal 603 of audio signal 630a with audio signal 630a. In some embodiments, the mixing circuit 834 may be configured to mix the speech signal 630a with the background noise signal 601 of the speech signal 630a. In some embodiments, the mixing circuit 834 may be configured to mix the utterance speech signal 603 of the speech signal 630a, the background noise signal 601 of the speech signal 630a, and the speech signal 630a. In some embodiments, the mixing circuit 834 may be configured to mix the speech signal 630a with a version of the speech signal 630a that has undergone background noise correction and spatial focusing (e.g., a target utterance speech signal 605).

[0131] In some embodiments, the mixing circuit 834 may be configured to mix a background noise-corrected and spatially focused version of the voice signal 630a (e.g., the target utterance voice signal 605) with the background noise signal 601. In some embodiments, the mixing circuit 834 may be configured to mix a first background noise-corrected and spatially focused version of the voice signal 630a (e.g., the target utterance voice signal 605) with a second background noise-corrected and spatially focused version of the voice signal 630a (e.g., the interference utterance voice signal 607). In some embodiments, the mixing circuit 834 may be configured to mix a first background noise-corrected and spatially focused version of the voice signal 630a (e.g., the target utterance voice signal 605), a second background noise-corrected and spatially focused version of the voice signal 630a (e.g., the interference utterance voice signal 607), and the voice signal 630a.

[0132] In some embodiments, the mixing circuit 834 may be configured to mix a first version of the speech signal 630a with background noise correction and spatial focusing (e.g., the target speech signal 605), a second version of the speech signal 630a with background noise correction and spatial focusing (e.g., the interfering speech signal 607), and the background noise signal 601. In some embodiments, the mixing circuit 834 may be configured to mix the first version of the speech signal 630a with background noise correction and spatial focusing (e.g., the target speech signal 605) with a speech signal including background noise and interfering speech (e.g., a speech signal including the interfering speech signal 607 in addition to the background noise signal 601). The mixing performed by the mixing circuit 834 may be considered interpolation.

[0133] Figure 11 shows a neural network circuit 1126 and a mask application and subtraction circuit 1132 according to several embodiments described herein. The neural network circuit 1126 is an example of any neural network circuit described herein (e.g., neural network circuits 526 and / or 826). The mask application and subtraction circuit 1132 may be an example of any processing circuit described herein (e.g., mask application and subtraction circuit 832).

[0134] The neural network circuit 1126 includes a circuit configured to implement one or more neural network layers. One or more neural network layers 1150 implemented by the neural network circuit 1126 may be configured to receive a plurality of speech signals 630. Generally, one or more neural network layers 1150 may be configured to generate a target speech utterance signal 605, or an output configured to generate the target speech utterance signal 605, based on the plurality of speech signals 630. In the particular example of Figure 11, one or more neural network layers 1150 may be configured to generate a mask 1056b based on the plurality of speech signals 630. As described above, the mask 1056b may be configured to generate the target speech utterance signal 605 (by multiplying the speech signal 630a by the mask 1056b using the multiplier 952b). From the target utterance speech signal 605, a speech signal 609 including the background noise signal 601 and the interfering utterance speech signal 607 can be generated (by subtracting the target utterance speech signal 605 from the speech signal 630a using the subtractor 954b). One or more neural network layers 1150 implemented by the neural network circuit 1126 can be considered to have been trained to perform background noise correction and spatial focusing (for generating the target utterance speech signal 605 from the speech signal 630a).

[0135] The above description of Figure 11 illustrates how the target speech voice signal 605 and the voice signal 609 (including the background noise signal 601 and the interfering speech voice signal 607) may be generated. Furthermore, as described above, several combinations of these signals and the voice signal 630a may be mixed by, for example, the mixing circuit 834. More specifically, if the target speech voice signal 605 is TS and the voice signal 609 is (IS+BN), in some embodiments the noise reduction circuit 834 may be configured to generate the output voice signal 840 to be equivalent to a*TS+b*(IS+BN). In some embodiments, the weight b of the interfering speech voice and background noise may have a value between 0 and 1. The target speech voice weight a typically has a value of 1, but other values ​​(e.g., values ​​greater than 1 or values ​​less than 1) may also be used.

[0136] Therefore, in some embodiments, the output audio signal 840 may have reduced levels of background noise and interfering speech. Adding some background noise and interfering speech to the target speech can help reduce distortion and enhance the wearer's environmental awareness of the in-ear device. In some embodiments, the mixing circuit 834 may be configured to generate the output audio signal 840 by mixing. Therefore, the mixing circuit 834 may be configured to assign different weights to the signals (e.g., by multiplication) and add the results. The mixing performed by the mixing circuit 834 may be considered interpolation.

[0137] It will be understood that the output audio signal 840 can be generated to be equivalent to a*TS+b*(IS+BN) by multiplying the target speech audio signal 605 by a and the audio signal 609 by b. However, the same result may be obtained by mixing other signals. This is true under the assumption that the audio signal 630a corresponds to the sum of the speech audio signal 603 and the background noise signal 601, and that the speech audio signal 603 corresponds to the sum of the target speech audio signal 605 and the interfering speech audio signal 607. As an unrestricted example, consider multiplying the target speech audio signal 605 by d and the audio signal 630a by e, and adding these intermediate products. If the audio signal 630a is O(O for Original) and the speech audio signal is S, then it can be shown that:

[0138] d*TS+e*O=d*TS+e*(S+BN) =d*TS+e*(TS+IS+BN) =(d+e)*TS+e*IS+e*BN Next, in the equation a*TS+b*(IS+BN), the weights a and b have the following relationship with the weights d and e: a=d+e, b=e.

[0139] In such embodiments, the background noise signal 601 and the interfering speech signal 607 are combined with the speech signal 609, so the first volume change difference and the second volume change difference may not be independently controllable. However, the first volume change difference and the second volume change difference may still be controllable. In other words, the first volume change difference and the second volume change difference must be the same amount, but their amounts may be controllable. To put it another way, the mixing circuit 834 may be configured to mix two or more speech signals such that the output speech signal includes a background noise-corrected and spatially focused version of the speech signal 630a mixed with the second speech signal (i.e., the target speech signal 605). The second speech signal may include the background noise signal 601 and the interfering speech signal 607. For example, the second speech signal may be speech signal 609 or speech signal 630a. The noise reduction circuit (for example, the noise reduction circuit 824) may be configured to generate the output audio signal 840 such that, in the output audio signal, the change in volume of the background noise signal 601 differs from the change in volume of the target speech audio signal 605 by the difference in volume change, the change in volume of the interfering speech audio signal differs from the change in volume of the target speech audio signal by the same difference in volume change, and the difference in volume change is controllable.

[0140] (pyRC circuit) As described above, the noise reduction circuit may be configured to generate an output audio signal including the target speech signal 605, the interfering speech signal 607, and the background noise signal 601 such that, in the output audio signal, the change in volume of the background noise signal differs from the change in volume of the target speech signal by a first volume change difference, the change in volume of the interfering speech signal differs from the change in volume of the target speech signal by a second volume change difference, and the first and second volume change difference amounts are independently controllable. In the above description, the generation of this output audio signal was described using a mixing circuit (e.g., mixing circuit 834). In some embodiments, the noise reduction circuit may be configured to generate the output audio signal using a wide dynamic range compression (WDRC) circuit. Briefly, some in-ear devices apply a nonlinear, frequency-dependent gain to the input sound to "fit" the output sound to the wearer's hearing profile. For example, if a wearer has significant hearing loss in the high-frequency range and much less in the low-frequency range, an in-ear device may apply more gain to high-frequency sounds than to low-frequency sounds for the same input volume, in order to equalize the audibility or perceived volume of different sounds across frequencies. Furthermore, because people with hearing loss typically have a narrower range of sounds they can comfortably hear (a narrower "dynamic range"), some hearing aids apply more gain to quiet sounds and less gain to louder sounds, effectively "compressing" the original signal to the wearer's dynamic range. These techniques are sometimes called wide dynamic range compression (WDRC).

[0141] Figure 12 shows a noise reduction circuit 1224 in an in-ear device according to several embodiments described herein. The in-ear device may be any of the in-ear devices described herein (e.g., hearing aid 100, eyeglasses 300, in-ear device 400, and / or in-ear device 500). The noise reduction circuit 1224 may be any of the noise reduction circuits described herein (e.g., noise reduction circuit 524). The noise reduction circuit 1224 includes a neural network circuit 826, a mask application and subtraction circuit 832, and a WDRC circuit 1258. The WDRC circuit 1258 may be configured to receive a plurality of audio signals 838 as input, which may include, for example, two or more of the following: an audio signal 630a, an utterance audio signal 603, a background noise signal 601, a target utterance audio signal 605, and an interfering utterance audio signal 607.

[0142] Figure 13 shows in more detail a WDRC circuit 1258 according to some embodiments described herein. The WDRC circuit 1258 includes an amplification pipeline 1360 corresponding to each of a plurality of audio signals 858 received by the WDRC circuit 1258, each amplification pipeline 1360 including a set of level estimation circuits 1362 and a set of amplification circuits 1364. The plurality of audio signals 858 may include, for example, two or more of the following: an audio signal 630a, an utterance audio signal 603, a background noise signal 601, a target utterance audio signal 605, and an interfering utterance audio signal 607. Each set of level estimation circuits 1362 is shown to include multiple blocks, each block for a different frequency channel. Each set of amplification circuits 1364 is shown to include multiple blocks, each block for a different frequency channel. (Circuits that convert the input signal to the frequency domain, split the signal into frequency channels, combine the frequency channels, and convert to the time domain are not shown for simplicity.)

[0143] Generally, the WDRC circuit 1258 includes multiple hearing loss amplification (also referred to herein simply as “amplification”) pipelines 1360. Each amplification pipeline 1360 corresponds to one of the sub-signals and includes a block of amplification circuit 1364. Amplification circuit 1364 may be configured to perform hearing loss amplification, i.e., additional amplification configured to compensate for the loss of hearing due to hearing loss. In particular, each block of amplification circuit 1364 may be configured to apply amplification to each input sub-signal to produce an amplified sub-signal. The amplification applied by each block of amplification circuit 1364 may be different. For example, amplification circuit 13641 in amplification pipeline 13601 may be configured to apply a first amplification to sub-signal 1, and amplification circuit 13642 in amplification pipeline 13602 may be configured to apply a second amplification to sub-signal 2, and the first and second amplifications may be different. In general, amplification may be any method of amplifying a signal to compensate for hearing loss due to hearing loss, and may include, for example, one or more rules, formulas, or curves.

[0144] Each set of level estimation circuits 1362 is configured to determine the level of each sub-signal, and each set of amplification circuits 1364 may be configured to amplify each sub-signal (e.g., apply a set of speech-speech fitting curves) at least partially based on the level of the speech-speech sub-signal determined by the level estimation circuits 1362. More specifically, for each set of level estimation circuits 1362 for a particular sub-signal, each block of the level estimation circuits 1362 may be configured to determine the level (e.g., power or amplitude) of the input sub-signal within a particular frequency channel and within a certain time window, or on a moving average of a certain time window. For each set of amplification circuits 1364 corresponding to a particular sub-signal, each block of the amplification circuits 1364 is configured to apply amplification to the input sub-signal within a particular frequency channel, resulting in an amplified sub-signal within that frequency channel, and the sum of the amplified sub-signals in different frequency channels constitutes the amplified sub-signal. The amplification applied by each set of amplification circuits 1364 in each amplification pipeline 1360 may differ.

[0145] The amplification applied by the amplification circuit 1364 to a specific frequency channel of the sub-signal may depend, at least in part, on the input level of that specific frequency channel of the sub-signal determined by the level estimation circuit 1362. Amplification having input level dependence and frequency dependence may include applying a set of fitting curves to the sub-signal, where each fitting curve is an output level versus input level curve for a given frequency channel (or, equivalently, each fitting curve is an output level versus frequency channel curve for a given input level). Different amplifications may include different sets of fitting curves. Applying a set of fitting curves to the sub-signal may include determining the input level of the sub-signal at each frequency channel, determining the output level corresponding to that input level and frequency channel from one of the fitting curves, amplifying that channel of the sub-signal to its output level, and combining the results from different frequency channels. The combiner 1374 (e.g., an adder) may be configured to combine the amplified sub-signals into a single output signal, i.e., the output audio signal 840.

[0146] Therefore, the level estimation circuit 13621 in the amplification pipeline 13601 is configured to determine the level of sub-signal 1 in each frequency channel, and the amplification circuit 13641 may be configured to apply a first amplification to sub-signal 1 based on a first set of fitting curves that define the output level as a function of the input level and the frequency channel. The level estimation circuit 13622 in the amplification pipeline 13602 is configured to determine the level of sub-signal 2 in each frequency channel, and the amplification circuit 13642 may be configured to apply a second amplification to sub-signal 2 based on a second set of fitting curves that define the output level as a function of the input level and the frequency channel. The first and second amplifications are different, in other words, the first and second sets of fitting curves are different.

[0147] The WDRC circuit 1258 includes different level estimation circuits 1362 and different amplification circuits 1364 for different sub-signals. One sub-signal may have a block of level estimation circuit 1362 and a block of amplification circuit 1364, each block for a specific frequency channel, while other sub-signals may have a separate block of level estimation circuit 1362 and a separate block of amplification circuit 1364 for the same frequency channel. Thus, each amplification pipeline 1360 may be configured to measure the input levels separately for different sub-signals. This can help avoid the pumping effect, which occurs when a change in the level of one sub-signal causes a jump in the amplification of other sub-signals that are not changing in the same way, because only a single level estimater is used for the overall signal.

[0148] Figure 13 shows three or more sub-signals, three or more amplification pipelines 1360, and three or more amplified sub-signals, but in some embodiments there may be two sub-signals, two amplification pipelines 1360, and two amplified sub-signals.

[0149] Level-dependent amplification can be configured to perform compression such that the dynamic range of the output level is smaller than the dynamic range of the input level. Amplification that includes compression is sometimes called wide dynamic range compression (WDRC). Thus, Figure 13 may show multiple WDRC pipelines (i.e., amplification pipeline 1360) configured to perform WDRC.

[0150] In some embodiments, a single amplification pipeline 1360 may be configured to perform amplification based on the level of one or more other sub-signals in addition to the level of a sub-signal associated with the amplification pipeline. For example, if one sub-signal is a speech voice signal 603 and the other sub-signal is a background noise signal 601, the levels of the speech voice signal 603 and the background noise signal 601 are used to calculate the signal-to-noise ratio (SNR), which can be used to modify the speech voice and / or noise fitting curve. In some embodiments, the level of the speech voice signal 603 may be used to set the gain of both the speech voice signal 603 and the background noise signal 601.

[0151] In some embodiments, the WDRC circuit 1258 lacks a level estimation circuit 1362, and therefore the amplification performed by the amplification circuit 1364 may not be applied as a function of the input level. In other words, the amplification may be independent of the input level. As an example, the amplification applied by the amplification circuit 1364 may include a half-gain rule (adding a gain equal to approximately half the amount of hearing loss) or a quarter-gain rule (adding a gain equal to half the total hearing loss plus a quarter of the conductive loss component of the hearing loss).

[0152] Memory 1374 may store different fitting curves and / or fitting rules for different sub-signals. For example, the memory may store one set of fitting curves for the target speech signal 605, one set of fitting curves for the interfering speech signal 607, and one set of fitting curves for the background noise signal 601. In some embodiments, the fitting curves for specific sub-signals and specific frequency channels are stored as a set of input levels, each having an associated output level, thereby defining a piecewise curve.

[0153] As can be understood from the above, different amplifications may be applied to different audio signals of the multiple audio signals 858. For example, different amplifications may be applied to the audio signal 630a, the utterance audio signal 603, the target utterance audio signal 605, the interference utterance audio signal 607, and the background noise signal 601. Therefore, the output audio signal 840 may include the target utterance audio signal 605, the interference utterance audio signal 607, and the background noise signal 601 such that the change in volume of the background noise signal differs from the change in volume of the target utterance audio signal by a first volume change difference, and the change in volume of the interference utterance audio signal differs from the change in volume of the target utterance audio signal by a second volume change difference. By controlling the different amplifications applied to different audio signals, such as by controlling the fitting curves stored in memory 1374, the first volume change difference and the second volume change difference can be controlled independently. Generally, the WDRC circuit 1258 may include multiple WDRC pipelines (i.e., amplification pipelines 1360) configured to generate an output audio signal 830 by performing WDRC on a combination of audio signals.

[0154] In some embodiments, the combination of audio signals may include three audio signals. In some embodiments, the three audio signals may be a set or subset of the following: audio signal 630a ("O"), speech audio signal 603 ("S"), background noise signal 601 ("BN"), target speech audio signal 605 ("TS"), and interference speech audio signal 607 ("IS"). (In some embodiments, audio signal 630a may be used by the mixer circuit 834, so audio signal 630a is shown with a dotted line as an input to the mixer circuit 834.) Effective combinations of these signals may include at least TS, IS, BN; TS, S, BN; IS, S, BN; TS, IS, O; TS, S, O; IS, S, O; TS, O, BN; IS, O, BN.

[0155] (Volume control) As described above, noise reduction circuits (e.g., noise reduction circuits 524, 824, and / or 1224) may be configured to output the output audio signal 840 using a mixer (e.g., mixer circuit 834) and / or a WDRC circuit (e.g., WDRC circuit 1258) such that the output audio signal 840 includes a target speech audio signal 605, an interfering speech audio signal 607, and a background noise signal 601, and in the output audio signal 840, the change in volume of the background noise signal 601 differs from the change in volume of the target speech audio signal 605 by a first volume change difference, and the change in volume of the interfering speech audio signal 607 differs from the change in volume of the target speech audio signal 605 by a second volume change difference, and the first and second volume change difference amounts are independently controllable. The change in volume can be measured between the volume in audio signal 630a and the volume in output signal 840. Furthermore, in this specification, a change in volume may refer to either an increase or a decrease in volume. Below, we will describe a method in which these first and second volume change difference amounts can be controlled independently.

[0156] Figure 14 shows the circuitry in an ear-mounted device according to a specific embodiment described herein. The ear-mounted device may be any of the ear-mounted devices described herein (e.g., hearing aid 100, eyeglasses 300, ear-mounted device 400, and / or ear-mounted device 500). Figure 14 shows a control circuit 1442, a mixing circuit 1434 (which may be an example of a mixing circuit 834), and / or a WDRC circuit 1458 (which may be an example of a WDRC circuit 1258), a memory 1444, and a communication circuit 1446. The memory 1444, the communication circuit 1446, and the mixing circuit 1434 and / or the WDRC circuit 1458 are coupled to the control circuit 1442. The memory 1444 may be configured to store data. The communication circuit 1446 may be configured to facilitate communication between the ear-mounted device and other devices (e.g., smartphones, tablets, laptops, computers) via a wireless communication link (e.g., Bluetooth or NFMI).

[0157] The control circuit 1442 may be configured to provide a first volume change control input 1448a and a second volume change control input 1448b to the mixing circuit 1434 and / or the WDRC circuit 1458. Thus, in some embodiments, the control circuit 1442 may be configured to provide the first volume change control input 1448a and the second volume change control input 1448b to the mixing circuit 1434. In some embodiments, the control circuit 1442 may be configured to provide the first volume change control input 1448a and the second volume change control input 1448b to the WDRC circuit 1458. In some embodiments, the control circuit 1442 may be configured to provide the first volume change control input 1448a and the second volume change control input 1448b to the mixing circuit 1434 and the WDRC circuit 1458.

[0158] In some embodiments, the mixing circuit 1434 may be configured to receive a first volume change control input 1448a and a second volume change control input 1448b and perform mixing using the first volume change control input 1448a and the second volume change control input 1448b such that the first volume change difference is at least partially controlled by the first volume change control input 1448a and the second volume change difference is at least partially controlled by the second volume change control input 1448b. For example, the mixing circuit 1434 may be configured to generate an output audio signal 840 equivalent to a*TS+b*IS+c*BN by multiplying the interfering speech audio signal 607 by b and the background noise signal 601 by c and adding the products of these. (For simplicity, assume that the weight a applied to the target speech signal 605 is 1 by default.) The mixing circuit 1434 may be configured to set weight c based on a first volume change control input 1448a and weight b based on a second volume change control input 1448b. Weights b and c can control a second volume change difference and a first volume change difference, respectively.

[0159] For example, if weight a is 1 and weight c is 0.18, the volume change of the target speech signal 605 may be 0 dB and the volume change of the background noise signal 601 may be -15 dB (i.e., the first volume change difference may be -15 dB). If weight a is 1 and weight b is 0.5, the volume change of the target speech signal 605 may be 0 dB and the volume change of the interfering speech signal 607 may be -6 dB (i.e., the second volume change difference may be -6 dB). As another example, the mixer circuit 834 may be configured to generate the output speech signal 630 equivalent to d*TS+e*IS+f*O by multiplying the interfering speech signal 607 by e and the background noise signal 601 by f and adding the products of these. (For simplicity, assume that the weight applied to the target speech signal 605 is 1 by default.) The mixing circuit may be configured to set weight f based on a first volume change control input and weight e based on a second volume change control input. The weights e and f can be used to control the first and second volume change difference amounts. Based on the above explanation showing that a=d+f, b=e+f, and c=f, for example, if weight d is 0.82, weight e is 0.32, and weight f is 0.18, then the volume change of the target speech signal 605 may be 0 dB, the volume change of the background noise signal 601 may be -15 dB (i.e., the first volume change difference amount may be -15 dB), and the volume change of the interfering speech signal 607 may be -6 dB (i.e., the second volume change difference amount may be -6 dB).

[0160] In some embodiments, the first volume control input 1448a and the second volume control input 1448b may be values. In such embodiments, the mixed control circuit 834 may make weight c equal to the first volume control input 1448a and weight b equal to the second volume control input 1448. In some embodiments, the mixed control circuit 834 may derive weight c from the first volume control input 1448a and weight b from the second volume control input 1448b. For example, the first volume control input 1448a may be an encoded version of weight c, and the second volume control input 1448b may be an encoded version of weight b, and the mixed control circuit 834 may be configured to decode the first volume control input 1448a and the second volume control input 1448b.

[0161] The first volume change control input 1448a and the second volume change control input 1448b may be different. Therefore, the first volume change difference amount and the second volume change difference amount may be different and can be controlled independently.

[0162] From the above, in some embodiments, the first volume change control input 1448a may control the weight applied to one signal (e.g., background noise signal 601), and the second volume change control input 1448b may control the weight applied to another signal (e.g., interfering speech voice signal 607). However, the mixing control circuit 834 may be configured to use three weights to mix the three signals. In some embodiments, the third volume change control input may not be used if the weight applied to the third signal (e.g., target speech voice signal) is always the same value or is the same value by default (e.g., 1). However, in some embodiments, the third volume change control input may be used to control the weight applied to the third signal using the following method. For simplicity, this description will focus on the first volume change control input 1448a and the second volume change control input 1448b.

[0163] In some embodiments, the WDRC circuit 1458 may be configured to receive a first volume change control input 1448a and a second volume change control input 1448b and perform WDRC using the first volume change control input 1448a and the second volume change control input 1448b such that a first volume change difference is at least partially controlled by the first volume change control input 1448a and a second volume change difference is at least partially controlled by the second volume change control input 1448b. For example, the first volume change control input 1448a and the second volume change control input 1448b may be different fitting curves or fitting rules (or inputs from which fitting curves or fitting rules can be derived or extracted) for application to different signals.

[0164] In some embodiments, memory 1444 may be configured to store the first volume change control input 1448a and the second volume change control input 1448b. In some embodiments, communication circuit 1446 may be configured to receive the first volume change control input 1448a and the second volume change control input 1448b from a processing device (i.e., an external device). Memory 1444 may be configured to store the first volume change control input 1448a and the second volume change control input 1448b. In some embodiments, control circuit 1442 may be configured to retrieve the first volume change control input 1448a and the second volume change control input 1448b from memory 1444 and output the first volume change control input 1448a and the second volume change control input 1448b to the mixer circuit 1434 and / or WDRC circuit 1458.

[0165] In some embodiments, the communication circuit 1446 is configured to receive a first volume change control input 1448a and a second volume change control input 1448b from an external device, and the control circuit 1442 may be configured to receive the first volume change control input 1448a and the second volume change control input 1448b from the communication circuit 1446 and provide the first volume change control input 1448a and the second volume change control input 1448b to the mixing circuit 1434 and / or WDRC circuit 1458 without storing the data in the memory 1444.

[0166] In some embodiments, the first volume change difference and the second volume change difference may be determined as part of the fitting. For example, an audiotherapist may determine the first and second volume change difference during fitting and use a processing device (e.g., a smartphone, tablet, laptop, or computer) to transmit the first volume change control input 1448a and the second volume change control input 1448b (corresponding to the first and second volume change difference, respectively) to the communication circuit 1446 of the ear-worn device. In some embodiments, the processing device of the ear-worn device wearer (e.g., a processing device 418 such as a smartphone, tablet, laptop, or computer) may run an application for communicating with the ear-worn device. In such embodiments, the app includes default values ​​for a first volume change difference and a second volume change difference, and the wearer's processing device may transmit a first volume change control input 1448a and a second volume change control input 1448b (corresponding to the default first volume change difference and the default second volume change difference, respectively) to the communication circuit 1446 of the ear-worn device. The first volume change control input 1448a and the second volume change control input 1448b may then be stored in memory 1444.

[0167] In some embodiments, the app includes a set of different default values ​​for a first volume change difference and a second volume change difference, and the wearer's processing device may send a set of first volume change control inputs 1448a and second volume change control inputs 1448b (corresponding to a default set of first volume change difference amounts and a default set of second volume change difference amounts, respectively) to the communication circuit 1446 of the ear-worn device. The set of first volume change control inputs 1448a and second volume change control inputs 1448b may then be stored in memory 1444. When the wearer selects a mode using the app on the processing device, the processing device sends an instruction for the selected mode to the communication circuit 1446 of the ear-worn device, and the control circuit 1442 receives the mode instruction and may retrieve the first volume change control inputs 1448a and second volume change control inputs 1448b corresponding to the selected mode from memory 1444.

[0168] In some embodiments, the app provides the wearer with an option to select a specific first volume change difference and a second volume change difference, or other values ​​relating to the first and second volume change difference, and the wearer's processing device transmits the first volume change control input 1448a and the second volume change control input 1448b (corresponding to the selected first volume change difference and the default second volume change difference, respectively) to the communication circuit 1446 of the ear-worn device, and the control circuit 1442 receives the selected first volume change control input 1448a and the selected second volume change control input 1448b and may use them to control the mixing circuit 1434 and / or WDRC circuit 1458.

[0169] As used herein, a circuit configured to perform operations relating to a volume change control input (e.g., storing, receiving, acquiring, providing, etc.) includes a circuit configured to perform operations using the volume change control input itself or using other data from which the volume change control input can be acquired. For example, if we mean that memory 1444 is configured to store a volume change control input, then memory 1444 may be configured to store the volume change control input itself, or to store other data from which the volume change control input itself can be acquired (e.g., an encoded version of the volume change control input). As another example, if we mean that control circuit 1442 provides a volume change control input to the mixing circuit 1434 and / or WDRC circuit 1458, then control circuit 1442 may be configured to provide the volume change control input itself, or to provide other data from which the volume change control input itself can be acquired (e.g., an encoded version of the volume change control input).

[0170] As described above, in some embodiments, the mixer circuit 1434 may be configured to mix two signals. For example, the mixer circuit 1434 may be configured to mix the target speech voice signal 605 with the voice signal 609. Or, it may be configured to mix the target speech voice signal 605 with the voice signal 630a. More specifically, the mixer circuit 1434 may be configured to output a*TS+b*(IS+BN) or a*TS+b*O. In embodiments where weight a always has the same value or is the same value by default (e.g., 1), the mixer circuit 1434 may be configured to receive only the first volume change control input 1448a, which can control weight b, and not control the second volume change control input. Thus, the volume change control input 1448b is shown as a dashed line in the figure. This volume change control input may control the volume change difference amount for both the interfering speech voice and the background noise. In other words, the volume change difference amounts for the interfering speech voice and the background noise may be the same. The provisions described herein regarding the first volume change difference amount can also be applied to this volume change difference amount.

[0171] In other words, the mixing circuit 1434 may be configured to mix two or more audio signals such that the output audio signal includes a background noise-corrected and spatially focused version of the audio signal 630a mixed with the second audio signal (i.e., the target speech audio signal 605). The second audio signal may include a background noise signal 601 and an interfering speech audio signal 607. For example, the second audio signal may be audio signal 609 or audio signal 630a. A noise reduction circuit (e.g., noise reduction circuit 824) may be configured to generate the output audio signal 840 such that the volume change of the background noise signal 601 differs from the volume change of the target speech audio signal 605 by the volume change difference amount, and the volume change of the interfering speech audio signal differs from the volume change of the target speech audio signal by the same volume change difference amount, and the volume change difference amount is controllable. The mixing circuit 1434 may be configured to receive the volume change control input 1448a and perform mixing using the volume change control input 1448a such that the volume change difference is controlled at least partially by the volume change control input. The communication circuit 1446 may be configured to receive the volume change control input 1448a from the processing device, the memory 1444 may be configured to store the volume change control input 1448a, and the control circuit 1442 may be configured to take the volume change control input 1448a and output the volume change control input 1448a to the mixing circuit 1434.

[0172] In some embodiments, the amount of volume change of the background noise signal 601 may be based on the level of background noise in the speech signal 630a. In some embodiments, the level of background noise may be measured by a background noise component determined using a steady-state noise suppression (SNS) circuit. Figure 15 shows the circuit in an in-ear device according to a particular embodiment described herein. The circuit includes a neural network circuit 826, a mask application and subtraction circuit 832, a mixing circuit 1434, a control circuit 1542 (which may be the same as the control circuit 1442), and an SNS circuit 1570. The in-ear device may be any of the in-ear devices described herein (e.g., hearing aid 100, eyeglasses 300, in-ear device 400, and / or in-ear device 500).

[0173] In some embodiments, one or more neural network layers implemented by the neural network circuit 826 are particularly effective in reducing transient noise, and separate steady-state noise suppression may be implemented. Thus, the SNS circuit 1570 may be configured to receive the audio signal 630a and generate a steady-state background noise signal 1501, i.e., an estimate of the steady-state background noise component of the audio signal 630a. The steady-state background noise signal 1501 may fluctuate slowly. Qualitatively, the steady-state background noise signal 1501 generally does not need to change substantially on a time scale of a few seconds. Quantitatively, the steady-state background noise signal 1501 may be asymmetric in that it may be acceptable for it to decrease on a relatively fast time scale, but only for it to increase on a very long time scale (about 10 seconds).

[0174] In some embodiments, the SNS circuit 1570 may be configured to implement a minimum statistics noise estimation algorithm to generate a stationary background noise signal 1501. In some embodiments, the SNS circuit 1570 may be further configured to implement other algorithms in addition to, or instead of, the minimum statistics noise estimation algorithm to generate a stationary background noise signal 1501. These algorithms may include, as non-limiting examples, spectral subtraction, Wiener filtering, and the Ephraim Marach technique. Further descriptions of such algorithms can be found in Chung, King. "Challenges and recent developments in hearing aids: Part I. Speech understanding in noise, microphone technologies and noise reduction algorithms." Trends in Amplification 8.3 (2004): 83-124, which is incorporated herein by reference.

[0175] The control circuit 1542 may be configured to receive a steady background noise signal 1501 (generated using the SNS circuit 1570) and generate a first volume change control input 1448a based on the level of the steady background noise signal 1501. As described above, in some embodiments, the mixing circuit 1434 may be configured to mix the background noise signal 601 (generated using the neural network circuit 826) with one or more other signals to generate an output audio signal 840. In embodiments such as that shown in Figure 15, the amount of background noise signal 601 mixed with other signals may be based not on the background noise signal 601 itself, but on the steady background noise signal 1501 (this may be a counterintuitive approach). These two background noise signals do not necessarily have to be identical. Mixing based on the level of the background noise signal 1501 may be useful because the steady background noise signal 1501 is a slowly fluctuating estimate of noise, and a slowly fluctuating estimate of noise helps reduce sudden jumps in the mixing coefficient.

[0176] In some embodiments, the control circuit 1542 may be configured to perform additional smoothing on the steady background noise signal 1501 before generating the first volume change control input 1448a based on the steady background noise signal 1501. However, in some embodiments, the steady background noise signal 1501 may be sufficiently gradual in its variation that additional smoothing is not necessary. In some embodiments, the control circuit 1542 may be configured to convert the units of the steady background noise signal 1501 to different units (e.g., from linear units to logarithmic units). However, in some embodiments, no unit conversion is necessary.

[0177] As further shown in Figure 15, the control circuit 1542 may be configured to receive the interfering speech voice signal 607 and generate a second volume change control input 1448a based on the level of the interfering speech voice signal 607.

[0178] Figure 16 shows the circuitry in an in-ear device according to a specific embodiment described herein. The circuitry includes a neural network circuit 826, a masking and subtraction circuit 832, a mixing circuit 1434, and a control circuit 1642 (which may be an example of a control circuit 1442). The in-ear device may be any of the in-ear devices described herein (e.g., a hearing aid 100, eyeglasses 300, an in-ear device 400, and / or an in-ear device 500). The control circuit 1670 may be configured to receive a background noise signal 601 (generated using the neural network circuit 826) and generate a first volume change control input 1448a based on the level of the background noise signal 601.

[0179] In some embodiments, the control circuit 1642 may be configured to perform additional smoothing on the background noise signal 601 before generating a first volume change control input 1448a based on the background noise signal 601. However, in some embodiments, the background noise signal 601 may be sufficiently gradual in its variation that additional smoothing is not necessary. In some embodiments, the control circuit 1642 may be configured to convert the units of the background noise signal 601 to different units (e.g., from linear units to logarithmic units). However, in some embodiments, unit conversion may not be necessary. As further shown in Figure 16, the control circuit 1670 may be configured to receive the interfering speech signal 607 and generate a second volume change control input 1448a based on the level of the interfering speech signal 607.

[0180] Generally, the control circuits (e.g., control circuits 1542 and / or 1642) may be configured to generate a first volume change control input 1448a based on the level of background noise in the audio signal 630a, and a second volume change control input 1448b based on the level of interfering speech in the audio signal 630a.

[0181] In some embodiments, the control circuits 1542 and / or 1642 may be configured to determine different volume change control inputs for different frequency bands based on different levels of background noise and / or interfering speech in different frequency bands, and the mixing circuit may be configured to use different volume change control inputs to mix the different frequency bands. However, in some embodiments, the control circuits 1542 and / or 1642 may be configured to determine one volume change control input based on one level of background noise and / or interfering speech (e.g., an averaged level across all frequencies), and one volume change control input may be used to mix all frequencies.

[0182] In some embodiments, as the level of background noise increases, the amount of background noise mixed into the output signal may decrease. However, in some embodiments, if the level of background noise increases beyond a certain threshold, the amount of background noise mixed in may increase again. In some embodiments, as the level of interfering speech increases, the amount of interfering speech mixed into the output signal may decrease. However, in some embodiments, if the level of interfering speech increases beyond a certain threshold, the amount of interfering speech mixed in may increase again.

[0183] Figure 17 shows the circuitry in an ear-mounted device according to several embodiments described herein. The circuitry includes a neural network circuit 1726 (which may be the same as neural network circuits 526, 826, 926, 1026, and / or 1126), a control circuit 1442, a memory 1444, and a communication circuit 1446. The ear-mounted device may be any of the ear-mounted devices described herein (e.g., hearing aid 100, eyeglasses 300, ear-mounted device 400, and / or ear-mounted device 500). As shown, the control circuit 1442 may be configured to output a first volume change control input 1448a and a second volume change control input 1448b to the neural network circuit 1726. The neural network circuit 1726 may be configured to receive a first volume change control input 1448a and a second volume change control input 1448b, and to use the first volume change control input 1448a and the second volume change control input 1448b to generate a neural network output 836 such that a first volume change difference is controlled at least partially by the first volume change control input 1448a and a second volume change difference is controlled at least partially by the second volume change control input 1448b.

[0184] In some embodiments, the neural network output 836 may be an output audio signal 840 such that a first volume change difference and a second volume change difference are controlled by a first volume change control input 1448a and a second volume change control input 1448b, respectively. In some embodiments, the neural network output 836 may be an output (e.g., a mask) configured to generate the output audio signal 840 (e.g., by amplification of the audio signal 630a) such that a first volume change difference and a second volume change difference are controlled by a first volume change control input 1448a and a second volume change control input 1448b, respectively. Generally, one or more neural network layers implemented by the neural network circuit 1726 may be learned to perform weighted combinations of the target speech audio signal 605, the interfering speech audio signal 607, and the background noise signal 601. In such embodiments, either or both of the processing circuits (e.g., mask application and subtraction circuit 832) and the mixing circuits (e.g., mixing circuits 834 and / or 1434) may be absent. For learning, each set of learning data may have learning input data including a plurality of audio signals 630, a first volume change control input 1448a and a second volume change control input 1448b, and learning output data including an output audio signal 840 where the first volume change difference and the second volume change difference are indicated by the first volume change control input 1448a and the second volume change control input 1448b, or a mask configured to generate such an output audio signal 840.

[0185] (Spatial focusing pattern) As described above, the target speech audio signal 605 may have a spatial focusing pattern. In other words, the target speech audio signal 605 may correspond to a speech audio signal 603 to which a specific spatial focusing pattern is applied. Generally, the output audio signal 840 may include a spatially focused signal having a spatial focusing pattern. For example, the output audio signal 840 may include the target speech audio signal 605 or an audio signal 630a to which a specific spatial focusing pattern is applied. Figure 18 shows an exemplary spatial focusing pattern according to a particular embodiment described herein. Figure 18 shows the weights as a function of DOA. (In this description, we use the convention of defining 0 degrees as the front of the wearer of the in-ear device, 0 to 90 degrees as the left side of the wearer, 0 to -90 degrees as the right side of the wearer, and 90 to -90 degrees as the back of the wearer). As shown in the figure, a weight of 1 is applied to DOA sounds between -30 and 30 degrees, while a weight of 0 is applied to other DOA sounds. Therefore, sounds obtained from 60 degrees in front of the wearer are retained, but other sounds are not. The spatial focusing pattern in Figure 18 involves a spatially abrupt transition from DOA using weight 1 to DOA using weight 0. Therefore, even small movements of the wearer's head cause the sound source to transition from the spatial region using weight 1 to the spatial region using weight 0, resulting in abrupt nullification of the sound.

[0186] In some embodiments, the weights transition smoothly or approximately smoothly as a function of DOA. Figure 19 shows an exemplary spatial focusing pattern according to a particular embodiment described herein. Figure 19 shows the weights as a function of DOA. As shown, the weights transition smoothly or approximately smoothly from a weight of 1 when DOA is 0 degrees (i.e., directly in front of the wearer) to a weight of 0.5 when DOA is 30 degrees and -30 degrees, and further to a weight of 0 when DOA is 90 degrees and -90 degrees (and further out). Thus, the entire back of the wearer can be nulled. A function like that shown in Figure 19 can be realized by inputting a given DOA into the formula. For example, the formula may be weight = (-1 / 90)*DOA+1 when 0≦DOA≦90 degrees, and weight = (1 / 90)*DOA+1 when -90≦DOA<0 degrees. In some embodiments, multiple DOA bins (which can also be thought of as spatial domains) are defined, each associated with a weight. For example, each bin may contain 1 degree, such as a bin for DOA between 0 and 1 degree, a bin for DOA between 1 and 2 degrees, and so on. For a given DOA, the weight associated with the bin containing that DOA can be determined (e.g., from a lookup table).

[0187] Figure 20 shows an exemplary spatial focusing pattern according to a particular embodiment described herein. Figure 20 is similar to Figure 19, except that the weights decrease from 1 to 0.5 as the DOA transitions from 0 to 25 degrees and from 0 to -25 degrees.

[0188] In general, spatial focusing patterns can include arbitrary changes in weights relative to the DOA (where DOA can be defined relative to the wearer). In some embodiments, spatial focusing patterns can use weights equal to 0, weights equal to 1, or weights between 0 and 1. In some embodiments, spatial focusing patterns can use weights equal to 0 or greater. In some embodiments, weights may be greater than 0, less than 0, equal to zero, or complex numbers. Negative weights may invert the phase by 180 degrees, and complex number weights may rotate the phase by an angle. As described above with reference to Figures 18-20, some spatial focusing patterns may have weights greater than 0 in a particular spatial region in front of the wearer and weights of 0 in other DOA. One way to compare spatial patterns is to determine the size of the spatial region using a threshold weight, e.g., a weight greater than 0.5. This spatial region may be considered the target spatial region, or the spatial region within the focal point. Next, Figure 19 can be seen as showing a spatial focusing pattern focused 60 degrees in front of the wearer, and Figure 20 can be seen as showing a spatial focusing pattern focused 50 degrees in front of the wearer, and Figure 20 can be seen as showing a larger amount of spatial focusing pattern than Figure 19. However, other types of spatial focusing patterns other than those shown in Figures 18 to 20 may also be used, and these spatial focusing patterns may include arbitrary changes in the weights for DOA.

[0189] The result of applying a spatial focusing pattern to an audio signal may be equivalent to multiplying each component of the audio signal by a weight associated with the DOA from which that component originated, according to the spatial focusing pattern (exemplified in Figures 18-20). Therefore, the resulting audio signal may be the original audio signal to which the spatial focusing pattern was applied. For example, the target utterance audio signal 605 may be the utterance audio signal 603 of the audio signal 630a to which the spatial focusing pattern was applied. In embodiments that focus on multiple DOAs, multiple different functions may be generated, as shown in Figures 18-20, and the sum or union of these functions may be used.

[0190] Target space areas with sizes other than 60 degrees (e.g., as shown in Figure 19) or 50 degrees (e.g., as shown in Figure 20) may also be used. In some embodiments, the target space area may have a size between approximately 10 degrees and 180 degrees. In some embodiments, the target space area may have a size between approximately 20 degrees and 180 degrees. In some embodiments, the target space area may have a size between approximately 30 degrees and 180 degrees. In some embodiments, the target space area may have a size between approximately 40 degrees and 180 degrees. In some embodiments, the target space area may have a size between approximately 50 degrees and 180 degrees. In some embodiments, the target space area may have a size between approximately 60 degrees and 180 degrees. In some embodiments, the target space area may have a size between approximately 10 degrees and 150 degrees. In some embodiments, the target space area may have a size between approximately 20 degrees and 150 degrees. In some embodiments, the target space area may have a size between approximately 30 degrees and 150 degrees. In some embodiments, the target spatial region may have a size between approximately 40 and 150 degrees. In some embodiments, the target spatial region may have a size between approximately 50 and 150 degrees. In some embodiments, the target spatial region may have a size between approximately 60 and 150 degrees. In some embodiments, the target spatial region may have a size between approximately 10 and 120 degrees. In some embodiments, the target spatial region may have a size between approximately 20 and 120 degrees. In some embodiments, the target spatial region may have a size between approximately 30 and 120 degrees. In some embodiments, the target spatial region may have a size between approximately 40 and 120 degrees. In some embodiments, the target spatial region may have a size between approximately 50 and 120 degrees. In some embodiments, the target spatial region may have a size between approximately 60 and 120 degrees. In some embodiments, the target spatial region may have a size between approximately 10 and 90 degrees. In some embodiments, the target spatial region may have a size between approximately 20 and 90 degrees. In some embodiments, the target spatial region may have a size between approximately 30 and 90 degrees.In some embodiments, the target spatial region may have a size between approximately 40 and 90 degrees. In some embodiments, the target spatial region may have a size between approximately 50 and 90 degrees. In some embodiments, the target spatial region may have a size between approximately 60 and 90 degrees. For example, the size may be equal to or approximately equal to 10 degrees, 20 degrees, 30 degrees, 40 degrees, 50 degrees, 60 degrees, 70 degrees, 80 degrees, 90 degrees, 100 degrees, 110 degrees, 120 degrees, 130 degrees, 140 degrees, 150 degrees, 160 degrees, 170 degrees, or 180 degrees, or any other preferred angle.

[0191] In some embodiments, the spatial focusing pattern may be predetermined. In other words, the boundaries of various spatial regions and the weights associated with each spatial region may be determined during training. The spatial focusing patterns in Figures 18-20 may be examples of predetermined spatial focusing patterns. In some embodiments, the spatial focusing pattern may not be predetermined. In other words, the boundaries of various spatial regions and / or the weights associated with each spatial region may be determined after training time (e.g., during inference time). A further explanation of such spatial focusing is provided below.

[0192] (Sound map) Referring to Figure 8, in some embodiments, the neural network output 836 generated by one or more neural network layers implemented by the neural network circuit 826 may be a sound map, or an output configured to generate a sound map (e.g., a mask). If the neural network output 836 is a mask configured to generate a sound map, the mask application and subtraction circuit 832 may be configured to generate the sound map by applying the mask (e.g., by multiplication or addition) to one of the multiple audio signals 630.

[0193] More specifically, if multiple frequency bins and multiple DOA bins are defined, the sound map may show the values ​​of each frequency bin obtained from each DOA bin. For example, if the number of frequency bins is n and the number of DOA bins is m, the sound map may be an n × m array. If one or more neural network layers are trained to perform noise correction, the sound map may be a speech map showing the frequency components of the speech obtained from each spatial domain.

[0194] Each DOA bin may include, for example, a predetermined angular range with respect to the wearer. In embodiments where the two microphones are symmetrical with respect to the axis connecting them (for example, the front microphone 102f and back microphone 102b shown in Figure 1), the ear-mounted device may not be able to distinguish between sound coming from the wearer's left side or the wearer's right side. In such embodiments, a single spatial region may be defined by combining the symmetrical regions on the wearer's left and right sides. For example, a single spatial region may be defined for both the 20-25 degree region on the wearer's left side and the 20-25 degree region on the wearer's right side. In embodiments where the microphones are not symmetrical with respect to the axis connecting them, the ear-mounted device may be able to distinguish between sound coming from the wearer's left side and sound coming from the wearer's right side.

[0195] In such embodiments, separate regions may be defined for symmetrical areas on the left and right sides of the wearer. An example of an ear-worn device that lacks symmetry with respect to the axis connecting the microphones is eyeglasses with built-in hearing aids (e.g., eyeglasses 300 with microphone 302). In embodiments with binaural communication, an ear-worn device or system of ear-worn devices may be able to distinguish between sounds coming from the wearer's left side and sounds coming from the wearer's right side. For example, a device in the wearer's left ear may detect sounds from the wearer's left side before a device in the wearer's right ear, and this early detection can be communicated between the two ears to determine that the sound is coming from the wearer's left side. Binaural communication may occur, for example, in a system of two hearing aids, cochlear implants, or earphones communicating with each other via a wireless communication link. Binaural communication may also occur in ear-worn devices such as eyeglasses with built-in hearing aids, where a device in one ear-closer part of the eyeglasses can communicate with a device in the other ear-closer part of the eyeglasses via a wired communication link within the eyeglasses.

[0196] In some embodiments, using a sound map, the masking and subtraction circuit 832 may be configured to apply a beam pattern to the sound map. Since the result of applying the beam pattern may be a spatially focused audio signal, the sound map may be used to generate a spatially focused audio signal. To apply the beam pattern, the masking and subtraction circuit 832 may be configured to apply different weights to sounds obtained from different DOA (as shown by the sound map). These weights do not need to be predetermined.

[0197] In some embodiments, including an ear-worn device with an array of three or more microphones (e.g., eyeglasses like eyeglasses 300), the beamforming circuit may be configured to generate multiple beams (e.g., 10 to 20 beams) pointing at different angles around a 360-degree circle relative to the wearer at each time step. A masking and subtraction circuit 832 or a neural network circuit 826 may be configured to calculate a metric value from the speech from the multiple beams. Thus, each beam may have a different value for the metric. For example, the metric may be the signal-to-noise ratio (SNR) or speaker power. The masking and subtraction circuit 832 may be configured to combine the speech from the multiple beams using the metric value. For example, if the metric is SNR, the masking and subtraction circuit 832 may be configured to output a sum of the speech from each beam weighted by the SNR of each beam. In particular, low-SNR beams may be given smaller weights, and high-SNR beams may be given larger weights. Thus, it is possible to focus on the beam with the highest SNR speech.

[0198] As another example, the metric may relate to voice personalization. A further explanation of voice personalization can be found in U.S. Patent No. 11,818,523, published November 14, 2023, entitled “System and Method for Enhancing Target Speaker’s Speech from Speech Signals in an Ear-Wear Device Using Voice Signature,” the entire disclosure is incorporated herein by reference. For example, the metric may indicate whether or not a particular speaker’s voice is present in a particular beam. The masking and subtraction circuit 832 may be configured to combine speech from multiple beams using the value of the metric. In some embodiments, the output may be a slowly fluctuating average (e.g., an exponential moving average) of speech from beams that previously contained the voice of a particular speaker. Thus, the output may include, in addition to the speech from beams that currently contain the speaker’s voice, speech from directions that do not currently contain the speaker’s voice but recently did, weighted according to an averaging function.

[0199] Therefore, as a result, a spatially focused audio signal is obtained, and thus the calculated value for the metric can be used to generate a spatially focused audio signal. In some embodiments, the processing circuit may be configured to apply a moving average across audio from different beams before summing the beams.

[0200] (Control of spatial focusing) As described above, one signal may be equivalent to another signal to which a spatial focusing pattern has been applied. For example, the target speech audio signal 605 may be equivalent to the speech audio signal 603 to which a spatial focusing pattern has been applied. As another example, the output audio signal 840 may include the target speech audio signal 605. As yet another example, the output audio signal 840 may include the speech signal 630a to which a spatial focusing pattern has been applied. In some embodiments, the spatial focusing pattern used by one or more neural network layers (i.e., by outputting a signal having a spatial focusing pattern or an output configured to generate a signal having a spatial focusing pattern in order to perform spatial focusing) may be controlled via an input to a neural network circuit (e.g., any of the neural network circuits 826 described herein).

[0201] In some embodiments, the spatial focusing pattern may be controlled via an input to a processing circuit (e.g., either the mask application and subtraction circuit 832 described herein). In some embodiments, the spatial focusing pattern may be controlled via an input to a mixing circuit (e.g., the mixing circuits 834 and / or 1434 described herein). Furthermore, in some embodiments, as described below, the user can control the spatial focusing pattern.

[0202] Figure 21 shows a circuit for controlling spatial focusing in an ear-mounted device according to a particular embodiment described herein. Figure 21 shows either or both of a sensing circuit 2176 coupled to a communication circuit 2146 (which may be the same as communication circuit 1446) and a control circuit (which may be the same as control circuits 1442, 1542, and / or 1642). Figure 21 further shows a noise reduction circuit 2124 (which may be the same as any of the noise reduction circuits described herein, such as noise reduction circuits 524 and / or 824), which includes a neural network circuit 2126, a processing circuit 2132, and a mixing circuit 2134. The ear-mounted device may be any of the ear-mounted devices described herein (e.g., hearing aid 100, eyeglasses 300, ear-mounted device 400, and / or ear-mounted device 500).

[0203] The communication circuit 2146 may be configured to communicate with other devices, such as a processing device (e.g., a smartphone or tablet, such as a processing device 418), via a wireless communication link (e.g., wireless communication link 420). For example, the wireless communication link may be Bluetooth or NFMI. In some embodiments, the user can use the processing device to select a spatial focusing pattern (e.g., a specific spatial focusing pattern of the target speech signal 605) or to make selections regarding spatial focusing. In the ear-worn device, the communication circuit 2146 may be configured to receive instructions from the processing device for user selection of a spatial focusing pattern and to generate one or more inputs 2166 based on the instructions for user selection of a spatial focusing pattern.

[0204] The control circuit 2142 may be configured, at least in part, to generate one or more spatial focusing control inputs 2168 indicating a spatial focusing pattern based on user selection instructions for a spatial focusing pattern, and in particular, based on one or more inputs 2166 received from the communication circuit 2146. One or more spatial focusing control inputs 2168 may indicate a spatial focusing pattern (i.e., a spatial focusing pattern selected by the user). One or more spatial focusing control inputs 2168 may be inputs to one or more of the following: the neural network circuit 2126 (which may be the same as any of the neural network circuits described herein, such as neural network circuits 526, 826, 926, 1026, and / or 1126), the processing circuit 2132 (which may be the same as any of the processing circuits described herein, such as mask application and subtraction circuit 832), and the mixing circuit 2134 (which may be any of the mixing circuits described herein, such as mixing circuit 834 and / or 1434). One or more spatial focusing control inputs 2168 may control spatial focusing through the control of a neural network circuit 2126, a processing circuit 2132, and / or a mixing circuit 2134, as further described below.

[0205] In embodiments where one or more spatial focusing control inputs 2168 from the control circuit 2142 are input to the neural network circuit 2126, the one or more spatial focusing control inputs 2168 may be input in addition to a plurality of audio signals 630 that are input to the neural network circuit 2126. The one or more spatial focusing control inputs 2168 may represent a spatial focusing pattern (e.g., a spatial focusing pattern selected by the user). One or more neural network layers implemented by the neural network circuit 2126 may be learned so that the one or more spatial focusing control inputs 2168 influence the pattern used for spatial focusing performed by the one or more neural network layers.

[0206] In other words, the neural network circuit 2126 may be configured to implement one or more neural network layers trained to generate an output audio signal (e.g., a target utterance audio signal 605) having a spatial focusing pattern indicated by one or more spatial focusing control inputs 2168 based on a plurality of input audio signals 630a, or an output (e.g., a mask) configured to generate an output audio signal (e.g., a target utterance audio signal 605) having a spatial focusing pattern indicated by one or more spatial focusing control inputs 2168.

[0207] For example, referring back to Figure 8, the target speech signal 605 may correspond to a speech signal 603 to which a specific spatial focusing pattern has been applied. The neural network circuit 2126 may be configured to receive one or more spatial focusing control inputs 2168 indicating a specific spatial focusing pattern and to use one or more spatial focusing control inputs 2168 to generate a neural network output 836 such that the target speech signal 605 corresponds to a speech signal 603 to which a specific spatial focusing pattern has been applied.

[0208] To train such a neural network layer, training can proceed as described above, except that the input training data is modified to include one or more spatial focusing control inputs 2168 corresponding to a spatial focusing pattern, and the output training data corresponds to that spatial focusing pattern. For example, consider input training data containing multiple audio signals and inputs 2168 having specific values ​​corresponding to a particular spatial focusing pattern. The output training data may be a mask to which the spatial focusing pattern is applied when applied to one of the multiple audio signals.

[0209] Rather than one or more spatial focusing control inputs 2168 being received by a circuit operating on the output of one or more neural network layers, one or more neural network layers themselves may receive one or more spatial focusing control inputs 2168 indicating a spatial focusing pattern.

[0210] In some embodiments, one or more neural network layers are configured not to receive inputs that indicate a spatial focusing pattern, in which case the neural network may be configured to apply a default spatial focusing pattern.

[0211] In embodiments in which one or more spatial focusing control inputs 2168 from the control circuit 2142 are input to the processing circuit 2132, in some such embodiments, the processing circuit 2132 may be configured to apply a beam pattern to the sound map based on one or more spatial focusing control inputs 2168.

[0212] In embodiments where one or more spatial focusing control inputs 2168 from control circuit 2142 are input to mixer circuit 2134, if the target speech voice signal 605 is TS, the interfering speech voice signal 607 is IS, and the background noise signal 601 is BN, then the output voice signal (not shown) from mixer circuit 2134 may correspond to a*TS+b*IS+c*BN. Using a relatively low value of b may result in more spatial focusing, while using a relatively high value of b may result in less spatial focusing. Mixer circuit 2134 is configured to mix a spatially focused version of the voice signal (let's call it A) with the voice signal itself (let's call it B), and using a value of a that is relatively higher than the value of b may result in more spatial focusing, while using a value of a that is relatively lower than the value of b may result in less spatial focusing, such that the output is a*A+b*B. Therefore, based on one or more spatial focusing control inputs 2168 from the control circuit 2142, the mixing circuit 2134 is configured to adjust the weights used for mixing, thereby substantially changing the weights applied to sounds from different DOA and controlling spatial focusing.

[0213] In some embodiments, spatial focusing may be turned off. When spatial focusing is turned off, the output signal from the noise reduction circuit 2124 is a noise-corrected audio signal, and in some cases, a noise-corrected beamforming audio signal. Therefore, in some embodiments, when spatial focusing is turned off, a portion of the speech audio signal 603 (i.e., the interfering speech audio signal 607) may not have a different volume change than other portions of the speech audio signal 603 (i.e., the target speech audio signal 605) based on spatial focusing. There may be multiple ways to turn off spatial focusing.

[0214] In some embodiments, there may be one or more spatial focusing control inputs 2168 that are not associated with spatial focusing and, when input to the neural network circuit 2126, may cause the neural network circuit 2126 not to perform spatial focusing. For example, in some embodiments, one or more spatial focusing control inputs 2168 may cause the neural network circuit 2126 not to operate one or more neural network layers (e.g., a second subset 950b) that have been trained to perform spatial focusing. As another example, one or more spatial focusing control inputs 2168 may cause the neural network circuit 2126 to use a spatial focusing pattern in which each DOA has a weight of 1.

[0215] In some embodiments, one or more spatial focusing control inputs 2168 may cause the neural network circuit 2126 to output a mask (e.g., mask 956b) configured to generate the speech voice signal 603 instead of the target speech voice signal 605 (e.g., the mask may contain all 1s). In some embodiments, one or more spatial focusing control inputs 2168 that are not associated with spatial focusing may be input to the processing circuit 2132 to prevent the processing circuit 2132 from performing spatial focusing. For example, the processing circuit 2132 may be configured to output the speech voice signal 603 instead of the target speech voice signal 605, or it may be configured to replace the mask (e.g., mask 956b) with a different mask configured to generate the speech voice signal 603.

[0216] In some embodiments, the sound map includes columns corresponding to sounds coming from different spatial regions. One or more spatial focusing control inputs 2168 may cause the processing circuit 2132 not to modify the weights of the columns in the sound map, or in other words, to apply a weight of 1 to the columns in the sound map. In some embodiments, one or more spatial focusing control inputs 2168 that are not associated with spatial focusing may be input to the mixer circuit 2134, and the mixer circuit 2134 may not perform spatial focusing. For example, one or more spatial focusing control inputs 2168 may cause the mixer circuit 2134 to mix the entire interfering speech signal 607 with the target speech signal 605 and return it. In other words, referring to the above equation a*TS+b*IS+c*BN, a and b may both be 1.

[0217] As another example, the mixing circuit 2134 may, when mixing, weight the spatially focused signal with 0 and the non-spatially focused signal with 1. In other words, referring to the above equation a*A+b*B, a may be 0 and b may be 1. In some embodiments, spatial focusing may be turned off by user selection. In other words, in some embodiments, an in-ear device (e.g., hearing aid 100, eyeglasses 300, in-ear device 400, and / or in-ear device 500) may be configured to receive a user selection to turn off spatial focusing. Based on receiving a user selection to turn off spatial focusing, the in-ear device may be configured to turn off spatial focusing, for example, using one of the methods described above. Further explanation is given below.

[0218] In some embodiments, noise correction performed using the neural network circuit 2126 may be turned off. In some embodiments, in order to turn off noise correction, one or more neural network layers trained to perform noise correction may not be executed.

[0219] (Graphical User Interface) In some embodiments, spatial focusing may be based on user selection. The user may be able to control spatial focusing using a processing device (e.g., a smartphone or tablet such as processing device 418) that communicates with an in-ear device (e.g., a hearing aid 100, eyeglasses 300, an in-ear device 400, and / or an in-ear device 500). Communication may be via a wireless communication link (e.g., wireless communication link 420). Any of the GUIs described herein may be displayed on such a processing device. This user selection may be as described above, in other words, instructions for user selection of a spatial focusing pattern may be received by the communication circuit 2146 of the in-ear device, as described above. In some embodiments, the user selection may be the selection of one of several options for a spatial focusing pattern. Thus, in some embodiments, the processing device may be configured to display a graphical user interface (GUI) containing options for different spatial focusing patterns and to accept user selection of a particular spatial focusing pattern.

[0220] Figure 22 shows a graphical user interface (GUI) 2280 for controlling spatial focusing of an in-ear device (e.g., a hearing aid) according to a specific embodiment described herein. GUI 2280 includes four options 2282a to 2282d for spatial focusing patterns. In general, a GUI may include multiple options for spatial focusing patterns. In the example in Figure 22, different spatial focusing patterns have different amounts of spatial focusing. GUI 2280 displays options 2282a to 2282d as graphical representations of different spatial focusing patterns. Each graphical representation in the example in Figure 22 includes a circle representing the wearer's environment and a highlighted area representing the location where spatial focusing occurs (e.g., a target spatial region). Option 2282a corresponds to a spatial focusing pattern that does not include spatial focusing or includes very little spatial focusing. Option 2282b corresponds to a spatial focusing pattern that focuses 180 degrees (or approximately 180 degrees) in front of the wearer. Option 2282c corresponds to a spatial focusing pattern that focuses 90 degrees (or approximately 90 degrees) in front of the wearer. Option 2282d corresponds to a spatial focusing pattern that focuses 45 degrees (or approximately 45 degrees) in front of the wearer.

[0221] Therefore, option 2282a may represent the smallest amount of spatial focusing among the displayed options, and option 2282d may represent the largest amount of spatial focusing among the displayed options. The spatial focusing patterns corresponding to options 2282a to 2282d may take the form shown in Figure 19, and this pattern is considered to focus on DOA having a weight greater than or equal to a threshold such as 0.5. The processing device displaying GUI2280 may be configured to accept a selection of one of options 2282a to 2282d, for example, via a touch-sensitive display screen that displays options 2282a to 2282d. Based on the user selection from GUI2280, the processing device may be configured to transmit the user selection instruction to the communication circuit 2146 of the ear-worn device. The communication circuit 2146 is configured to generate one or more inputs to the control circuit 2166 based on the received user selection instructions, and the control circuit 2142 may be configured to transmit one or more spatial focusing control inputs 2168 to the neural network circuit 2126, the processing circuit 2132, and / or the mixing circuit 2134 based on one or more inputs received from the communication circuit 2146.

[0222] In the example above, there are four options for one or more inputs to the neural network, but in some embodiments there may be fewer than four options, and in some embodiments there may be five or more options. In some embodiments there may be a large number of options or a continuous range of options for the spatial focusing pattern (e.g., from omnidirectional to superfocus). For example, the user may use a slider on a graphical user interface to select a spatial focusing pattern from a continuous range of options, the slider selecting a range of DOA within the focal point in front of the wearer.

[0223] For example, in an embodiment in which one or more spatial focusing control inputs 2168 are input to a neural network circuit 2126, if option 2282a is selected, one or more spatial focusing control inputs 2168 may be 0; if option 2282b is selected, one or more spatial focusing control inputs 2168 may be 1; if option 2282c is selected, one or more spatial focusing control inputs 2168 may be 2; and if option 2282d is selected, one or more spatial focusing control inputs 2168 may be 3.

[0224] As another example (i.e., a one-hot scheme), when option 2282a is selected, one or more spatial focusing control inputs 2168 may be [1,0,0,0]; when option 2282b is selected, one or more spatial focusing control inputs 2168 may be [0,1,0,0]; when option 2282c is selected, one or more spatial focusing control inputs 2168 may be [0,0,1,0]; and when option 2282d is selected, one or more spatial focusing control inputs 2168 may be [0,0,0,1].

[0225] One or more neural network layers themselves may receive one or more spatial focusing control inputs 2168 indicating a spatial focusing pattern, rather than having one or more spatial focusing control inputs 2168 received by a circuit operating on the output of one or more neural network layers. For example, one or more neural network layers may be configured to receive spatial focusing control inputs 2168 indicating a spatial focusing pattern using a one-hot method with a vector of size 128 and a vector of size 4. Thus, the total size of the inputs received by one or more neural network layers is 128 + 4 = 132.

[0226] To train such a neural network layer, one or more spatial focusing control inputs 2168 corresponding to a spatial focusing pattern are added to the input training data, and the training can proceed as described above, except when the output training data corresponds to that spatial focusing pattern. For example, consider input training data containing multiple audio signals formed from one or more sound signals, and the input [0,0,1,0] corresponding to option 2282c in Figure 22. The output training data, when applied to one of the multiple audio signals, focuses on the sound signal obtained from 90 degrees in front of the wearer and extending to +45 degrees and -45 degrees on both sides of the direction corresponding to the wearer's front, and the resulting audio signal may be a mask having a spatial focusing pattern.

[0227] In an embodiment in which one or more spatial focusing control inputs 2168 are input to the processing circuit 2132, the sound map processed by the processing circuit 2132 shall include 16 spatial regions. Option 2282d may be implemented by focusing on two forward-directed spatial regions out of the 16 spatial regions. Option 2282c may be implemented by focusing on four forward-directed spatial regions out of the 16 spatial regions. Option 2282b may be implemented by focusing on eight forward-directed spatial regions out of the 16 spatial regions. Option 2282a may be implemented by focusing on all 16 spatial regions. The sound map shall include 16 columns, each corresponding to one of the 16 spatial regions. Based on one or more spatial focusing control inputs 2168, the processing circuit 2132 may be configured to apply a weight of 1 to the column values ​​corresponding to the focused spatial region and a weight of 0 to the column values ​​corresponding to other spatial regions (for example, to realize a spatial focusing pattern like that shown in Figure 18), or to apply a weight of 0.5 or more and 1 or less to the column corresponding to the focused spatial region and a weight of 0.5 or more and 0 or less to the column corresponding to the unfocused spatial region (for example, to realize a spatial focusing pattern like that shown in Figure 19). Generally, the processing circuit 2132 may be configured to use higher weights for the column corresponding to the focused spatial region and lower weights for the column values ​​that do not correspond to the focused spatial region.

[0228] In some embodiments, user selection may involve which spatial regions to focus on and which spatial regions to leave unfocused. Figure 23 shows a graphical user interface (GUI) 2380 for controlling spatial focusing of an in-ear device (e.g., a hearing aid) according to a particular embodiment described herein. The GUI 2380 includes a circle 2384 representing the wearer's environment, where the wearer is at the center of the circle 2384, the front of the wearer is represented on the right side of the circle 2384, and the back of the wearer is represented on the left side of the circle 2384. The circle 2384 includes several spatial regions 2386a to 2386b that the wearer can select. In the example in Figure 23, the wearer has selected and highlighted spatial regions 2386a and 2386b. When using the GUI 2380, the wearer may select one, two, or three or more of the spatial regions 2386a to 2386f. More or fewer than six spatial regions may be used in the GUI 2380. Based on user selection from GUI2380, the processing device may be configured to transmit user selection instructions to the communication circuit 2146 of the ear-worn device. The communication circuit 2146 is configured to generate one or more inputs to the control circuit 2166 based on the received user selection instructions, and the control circuit 2142 may be configured to transmit one or more spatial focusing control inputs 2168 to the neural network circuit 2126, the processing circuit 2132, and / or the mixing circuit 2134 based on one or more inputs received from the communication circuit 2146.

[0229] In embodiments in which one or more spatial focusing control inputs 2168 are input to the neural network circuit 2126, the input 2168 is, for example, a vector having the same number of elements as the number of spatial regions 2386a to 2386b, where the elements are equal to 1 if a corresponding spatial region is selected, and equal to 0 otherwise. A further explanation of the learning of the neural network layer implemented by such a neural network circuit can be found above.

[0230] In embodiments where one or more spatial focusing control inputs 2168 are input to the processing circuit 2132, the sound map includes columns corresponding to sounds coming from each of the spatial regions 2386a to 2386f. Based on one or more spatial focusing control inputs 2168, the processing circuit 2132 may be configured to assign a weight of 1 to the values ​​of the columns corresponding to selected spatial regions (e.g., 2386a and 2386b in Figure 23) and a weight of 0 to the values ​​of the columns corresponding to unselected spatial regions (e.g., to achieve a spatial focusing pattern like that in Figure 18), or to assign a weight of 0.5 or more to the columns corresponding to selected spatial regions and a weight of 0 or less to the columns corresponding to unselected spatial regions (e.g., to achieve a spatial focusing pattern like that in Figure 19). Generally, the processing circuit 2132 may be configured to use higher weights for the columns corresponding to selected spatial regions and lower weights for the values ​​of the columns corresponding to unselected spatial regions.

[0231] In some embodiments, the user selection may be whether or not to perform spatial focusing. Figure 24 shows a graphical user interface (GUI) 2480 for controlling spatial focusing of an in-ear device (e.g., a hearing aid) according to a particular embodiment described herein. The GUI 2480 includes an option 2488 that can be switched by the user to turn spatial focusing on or off. Based on the user selection from the GUI 2380, the processing device may be configured to transmit the user selection instruction to a communication circuit 2146 of the in-ear device. The communication circuit 2146 is configured to generate one or more inputs to a control circuit 2166 based on the received user selection instruction, and the control circuit 2142 may be configured to transmit one or more spatial focusing control inputs 2168 to a neural network circuit 2126, a processing circuit 2132, and / or a mixed circuit 2134 based on one or more inputs received from the communication circuit 2146.

[0232] In some embodiments, a physical user input on the ear-worn device (e.g., a user input device 104 such as a button) may receive a user selection to turn spatial focusing on or off. Based on the user activation of the physical user input on the ear-worn device, the control circuit 2142 may be configured to transmit one or more spatial focusing control inputs 2168 to the neural network circuit 2126, the processing circuit 2132, and / or the mixing circuit 2134. A further description of how one or more spatial focusing control inputs 2168 can control the on and off of spatial focusing can be found above.

[0233] In some embodiments, user selection may be the degree of focusing. Figure 25 shows a graphical user interface (GUI) 2580 for controlling spatial focusing of an in-ear device (e.g., a hearing aid) according to a particular embodiment described herein. The GUI 2580 includes a slider option 2590 that the user can use to control the degree of focusing. The slider option 2590 includes a line 2592 and a slider 2594. The wearer can control the position of the slider 2594 on the line 2592. The ratio of (1) the distance from the left end of the line 2592 to the position of the slider 2594 and (2) the length of the line 2592 takes a value from 0 to 1, where a value closer to 1 indicates stronger spatial focusing and a value closer to 0 indicates weaker spatial focusing. The communication circuit 2146 is configured to generate one or more inputs to the control circuit 2166 based on a user-selected value reception instruction, and the control circuit 2142 may be configured to transmit one or more spatial focusing control inputs 2168 to the neural network circuit 2126, the processing circuit 2132, and / or the mixing circuit 2134 based on one or more inputs received from the communication circuit 2146.

[0234] In embodiments in which one or more spatial focusing control inputs 2168 are input to the neural network circuit 2126, one or more spatial focusing control inputs 2168 indicating a user-selected value (between 0 and 1) for the amount of spatial focusing may have values ​​related to the amount of spatial focusing selected by the user. Similar to Figure 22, if the user-selected value is close to or equal to 1, one or more spatial focusing control inputs 2168 may have values ​​related to option 2282d; if the user-selected value is close to or equal to 0, one or more spatial focusing control inputs 2168 may have values ​​related to option 2282a; and if the user-selected value is between 0 and 1, one or more spatial focusing control inputs 2168 may have values ​​related to option 2282b or 2282c. A further explanation of the learning of the neural network layer implemented by such a neural network circuit can be found above.

[0235] In embodiments where one or more spatial focusing control inputs 2168 are input to the processing circuit 2132, the sound map includes columns corresponding to sounds coming from different spatial regions. If the user selection value is close to or equal to 1, one or more spatial focusing control inputs 2168 can cause the processing circuit 2132 to weight the columns of the sound map to implement a spatial focusing pattern with a large amount of spatial focusing. If the user selection value is close to or equal to 0, one or more spatial focusing control inputs 2168 can cause the processing circuit 2132 to weight the columns of the sound map to implement a spatial focusing pattern with a small amount of spatial focusing. If the user selection value is between 0 and 1, one or more spatial focusing control inputs 2168 can cause the processing circuit 2132 to weight the columns of the sound map to implement a spatial focusing pattern with a moderate amount of spatial focusing.

[0236] In embodiments where one or more spatial focusing control inputs 2168 are input to the mixer 2134, the mixer 2134 is configured to mix a spatially focused version of the audio signal (let's call this A) with the audio signal itself (let's call this B), with an output of a*A + b*B. In some embodiments, one or more spatial focusing control inputs 2168 cause the mixer 2134 to use a value a equal to a user-selected value for the amount of spatial focusing. In such embodiments, b may be a constant and may be inversely proportional to a.

[0237] In some embodiments, user selection is which speaker to focus on. Figure 26 shows a graphical user interface (GUI) 2680 for controlling spatial focusing of an in-ear device (e.g., a hearing aid) according to a particular embodiment described herein. The GUI 2680 includes a representation of the wearer 2696 and representations of four speakers 2698a–2698d that are oriented differently from the wearer. In the example in Figure 26, the wearer has selected speaker 2698a, which is highlighted on the GUI 2680.

[0238] In some embodiments, the ear-mounted device is configured to generate multiple tight beams using beamforming in an array of three or more microphones, and the ear-mounted device or processing may be configured to calculate the power of the speech signal in the audio from each beam. Beams with power exceeding a threshold are considered to have a speaker in the direction of the beam. In some embodiments, a neural network may be trained to determine the direction of the speaker based on the audio from one or more beams. In some embodiments, the ear-mounted device runs the neural network, and information regarding the direction of the speaker may be transmitted from the ear-mounted device to the processing device. The processing device can use this information to display a representation of the speaker in a GUI in the direction of each speaker. In some embodiments, the neural network may be run on the processing device itself.

[0239] Based on a user selection from GUI2380, the processing device may be configured to transmit a user selection instruction to a communication circuit 2146 of the ear-mounted device. The communication circuit 2146 is configured to generate one or more inputs to a control circuit 2166 based on the received user selection instruction, and the control circuit 2142 may be configured to transmit one or more spatial focusing control inputs 2168 to a neural network circuit 2126, a processing circuit 2132, and / or a mixing circuit 2134 based on one or more inputs received from the communication circuit 2146. Based on the wearer selecting one of the speaker representations from the GUI, the processing device transmits this selection instruction to the ear-mounted device, which may then use a beam focused in the direction of that speaker.

[0240] In embodiments where one or more spatial focusing control inputs 2168 are input to a neural network circuit 2126, one or more spatial focusing control inputs 2168 may be associated with a spatial focusing pattern that points in the direction of a selected speaker. In embodiments where one or more spatial focusing control inputs 2168 are input to a processing circuit 2132, one or more spatial focusing control inputs 2168 may cause the processing circuit 2132 to weight a sequence of sound maps that implement a spatial focusing pattern that points in the direction of a selected speaker.

[0241] In some embodiments, an ear-worn device or processing device is configured to determine the signal-to-noise ratio (SNR) of an acoustic environment, and the ear-worn device may be configured to generate one or more spatial focusing control inputs 2168 that indicate a spatial focusing pattern based on the SNR of the acoustic environment. When the acoustic environment has a high signal-to-noise ratio (SNR), the ear-worn device may be configured to select less spatial focusing than when the acoustic environment has a low SNR.

[0242] (Sensing circuit) Returning to Figure 21, in some embodiments, the sensing circuit 2176 may include one or more of an accelerometer, a gyroscope, and a magnetometer. The sensing circuit 2176 may be configured to generate one or more inputs 2178 based on the movement of the ear-worn device. In some embodiments, the control circuit 2142 may be configured to determine the degree of head movement (i.e., how fast the wearer's head is moving) based on one or more inputs 2178 received from the sensing circuit 2176. A further description of how sensors are used to determine how fast the wearer's head is moving can be found in Ionut-Cristian S, Dan-Marius D, Using Inertial Sensors to Determine Head Motion—A Review, J Imaging, 2021 Dec 6; 7(12): 265, the entire disclosure of which is incorporated herein by reference.

[0243] In some embodiments, the control circuit 2142 may be configured to generate one or more spatial focusing control inputs 2168 that indicate a particular spatial focusing pattern based on the degree of head movement. For example, the control circuit 2142 may be configured to select a spatial focusing pattern with a larger amount of spatial focusing when the wearer's head is moving slowly or not moving at all, compared to when the wearer's head is moving quickly. As a specific example using four bins for the degree of head movement, no head movement may be associated with a spatial focusing pattern with a large amount of spatial focusing, low head movement with a medium amount of spatial focusing, moderate head movement with a small amount of spatial focusing, and high head movement may be associated with no spatial focusing.

[0244] As another example, if the speed of head movement exceeds a threshold, a spatial focusing pattern with a certain amount of spatial focusing (including no spatial focusing) may be used, and if the speed of head movement falls below the threshold, a spatial focusing pattern with a different amount of spatial focusing may be executed. Determining the amount of spatial focusing based on the speed of head movement can be useful because (1) it may help the neural network achieve higher performance during head rotation, and (2) spatial focusing can be broadened, assuming that head rotation does not correlate with the desire for precise spatial focusing. Generally, the control circuit 2142 is configured to generate a first set of one or more spatial focusing control inputs 2168 that show a spatial focusing pattern having a first spatial focusing amount based on the degree of a first head movement, and to generate a second set of one or more spatial focusing control inputs 2168 that show a spatial focusing pattern having a second spatial focusing amount based on the degree of a second head movement, where the first spatial focusing amount is smaller than the second spatial focusing amount, and the degree of the first head movement is larger than the degree of the second head movement.

[0245] In some embodiments, instead of using the sensing circuit 2176 to determine how fast the wearer's head is moving, or in addition to using the sensing circuit 2176 to determine how fast the wearer's head is moving, the ear-worn device may be configured to determine how fast the wearer's head is moving using a neural network trained to determine how fast the wearer's head is moving (for example, using sound received by the ear-worn device as input).

[0246] In embodiments where one or more spatial focusing control inputs 2168 are input to a neural network circuit 2126, one or more neural network layers implemented by the neural network circuit 2126 may be learned so that one or more spatial focusing control inputs 2168 influence the spatial focusing pattern implemented by the one or more neural network layers. In other words, the neural network circuit 2126 may be configured to implement one or more neural network layers that have been learned to generate an output audio signal having a spatial focusing pattern indicated by one or more spatial focusing control inputs 2168 based on a plurality of input audio signals 2130, or an output configured to generate an output audio signal having a spatial focusing pattern indicated by one or more spatial focusing control inputs 2168.

[0247] For example, in an embodiment in which one or more spatial focusing control inputs 2168 are input to a neural network circuit 2126, the control circuit 2142 may be configured to determine which of the four bins the velocity of head movement belongs to, based on one or more inputs 2178 from the sensing circuit 2176. If no head movement is detected, one or more spatial focusing control inputs 2168 may be 0; if the degree of head movement is low, one or more spatial focusing control inputs 2168 may be 1; if the degree of head movement is moderate, one or more spatial focusing control inputs 2168 may be 2; and if the degree of head movement is high, one or more spatial focusing control inputs 2168 may be 3. Another example of using the one-hot method is that if no head movement is detected, one or more spatial focusing control inputs 2168 may be [1,0,0,0]; if the degree of head movement is low, one or more spatial focusing control inputs 2168 may be [0,1,0,0]; if the degree of head movement is moderate, one or more spatial focusing control inputs 2168 may be [0,0,1,0]; and if the degree of head movement is high, one or more spatial focusing control inputs 2168 may be [0,0,0,1]. A further explanation of training such a neural network is provided above.

[0248] In embodiments in which one or more spatial focusing control inputs 2168 are input to the processing circuit 2132, the one or more spatial focusing control inputs 2168 can cause the processing circuit 2132 to process the sound map so that spatial focusing patterns associated with the amount of head movement (indicated by the one or more spatial focusing control inputs 2168) are implemented.

[0249] As described above, in some embodiments, an ear-worn device may generate a sound map showing frequency components obtained from each of several spatial domains. In such embodiments, it may be useful to apply a moving average to the values ​​determined for different spatial domains, which can average out some errors. However, if the wearer rotates their head rapidly, the average across different spatial domains may become blurred. A sensing circuit 2176 configured to track head movement (e.g., using an accelerometer and a gyroscope) can compensate for this, as will be further described below.

[0250] In a two-microphone array (e.g., on a hearing aid such as hearing aid 100), the beam that can be formed using beamforming (e.g., cardioid and supercardioid) can be wide, and even if the wearer is conversing with someone in front of them and then rotates their head 90 degrees, the amplitude of that person's speech may only decrease by a few dB. However, in an array of three or more microphones (e.g., on eyeglasses such as eyeglasses 300), a narrower beam can be formed. With a narrow beam, even a slight rotation of the head can substantially reduce the amplitude of sound from a person directly in front of the wearer. A sensing circuit 2176 configured to track head movement (e.g., using an accelerometer and gyroscope) can also compensate for this.

[0251] More specifically, a sensing circuit 2176 configured to track head movement (e.g., using an accelerometer and a gyroscope) may allow defining a spatial domain in an absolute coordinate system rather than in the wearer's head coordinate system (which can rotate very quickly). The absolute coordinate system can be defined relative to the wearer's head, but on a slower time scale. Therefore, if the wearer is sitting and talking to someone, and temporarily turns their head to look at something and then turns it back to the person they are talking to, the coordinate system may remain in the same place and not rotate (or hardly rotate) with the head. However, if the wearer turns to talk to another person, the coordinate system may rotate slowly (e.g., over several seconds) with the head. To achieve this, an exponential moving average may be applied to the coordinate system so that the coordinate system becomes an exponential moving average of the head orientation. The time scale of the exponential moving average may be, for example, several seconds (e.g., 2 seconds, 3 seconds, 4 seconds, or 5 seconds, 6 seconds, 7 seconds, 8 seconds, 9 seconds, or 10 seconds, or any other preferred value).

[0252] In some embodiments, as the head moves, sounds from the new direction (i.e., the direction the wearer's head is turning) can be immediately focused on, while sounds from the old direction (i.e., the direction the wearer's head is turning away) can be gradually defocused while remaining focused. In other words, as the wearer moves their head towards a new direction, the opening can widen fairly quickly to focus on sounds from the new direction, but continue to focus on sounds from the previous direction, and as the wearer continues to look in the new direction, the focus gradually narrows to sounds from the previous direction. The gradual narrowing can be adjusted as a function of the time the wearer continues to look in the new direction, so that fast head movements may not result in a permanently wide opening. In some embodiments, this behavior can be achieved by exponential moving averages with long time scales. In some embodiments, this behavior can be achieved by (1) combining sounds from the new direction with full weights, and (2) processing sounds from the old direction with exponential moving averages.

[0253] (Beamforming and directional patterns) As described above, in some embodiments, a neural network circuit (e.g., any neural network circuit described herein) is configured to receive a plurality (i.e., at least two) audio signals (e.g., audio signal 830), where at least two of the plurality of audio signals are obtained from one different microphone of two or more microphones, and / or at least one of the plurality of audio signals is a beamformed version of the audio signals obtained from the two or more microphones. With respect to beamforming, beamforming may generally involve applying a delay (to be understood as including a delay of 0) to one or more audio signals obtained from different microphones and summing (to be understood as including a subtraction) the delayed signals. Different delays may be applied to the signals obtained from different microphones. In embodiments involving only two microphones, such as a front microphone (e.g., front microphone 102f) and a back microphone (e.g., back microphone 102b), beamforming may involve applying a delay to the signal from one of the microphones and subtracting the delayed signal from the signal from the other microphone.

[0254] The resulting signal may have a beamforming directional pattern that depends at least partially on the distance between the front and back microphones and the applied delay. In other words, the weight of the resulting signal may vary as a function of the angle from the microphone. Examples of beamforming directional patterns include dipole, hypercardioid, supercardioid, and cardioid. Certain beamforming directional patterns, also referred to herein as “front-facing,” generally attenuate signals coming from behind the wearer more than signals coming from in front of the wearer. As described below, using a front microphone and a back microphone, a front-facing beamforming directional pattern can generally be created by applying a larger delay to the back microphone than to the front microphone and subtracting the rear signal from the front signal. Certain beamforming directional patterns, also referred to herein as “back-facing,” generally attenuate signals coming from in front of the wearer more than signals coming from behind the wearer. As described below, using front and back microphones, a rearward-directed beamforming directional pattern can generally be created by applying a larger delay to the front microphone than to the back microphone and subtracting the forward signal from the directional signal.

[0255] Figure 27 shows a forward-directed hypercardioid pattern 2776 according to some embodiments described herein. A forward-directed hypercardioid pattern can result from applying a d / 3c delay to the signal from the back microphone, where d is the distance between the front and back microphones and c is the speed of sound. Figure 28 shows a rearward-directed hypercardioid pattern 2876 according to some embodiments described herein. A rearward-directed hypercardioid pattern can result from applying a d / 3c delay to the signal from the front microphone. Figure 29 shows a forward-directed supercardioid pattern 2976 according to some embodiments described herein. A forward-directed supercardioid pattern can result from applying a 2d / 3c delay to the signal from the back microphone.

[0256] Figure 30 shows a rearward-directed hypercardioid pattern 3076 according to some embodiments described herein. A rearward-directed supercardioid pattern can result from applying a 2d / 3c delay to the signal from the front microphone. Figure 31 shows a forward-directed cardioid pattern 3176 according to some embodiments described herein. A forward-directed cardioid pattern can result from applying a d / c delay to the signal from the back microphone. Figure 32 shows a rearward-directed cardioid pattern 3276 according to some embodiments described herein. A rearward-directed cardioid pattern can result from applying a d / c delay to the signal from the front microphone. Figure 33 shows a dipole pattern 3376 according to some embodiments described herein. A dipole pattern can result from not applying a delay to either microphone. Other patterns can be generated by applying different delays.

[0257] As described above, in some embodiments, the neural network circuit may be configured to receive a plurality of audio signals (e.g., audio signal 830) including at least one audio signal which is a beamformed version of an audio signal obtained from two or more microphones. In some embodiments, the plurality of audio signals may include one beamforming signal. In some embodiments, the plurality of audio signals may include two beamforming signals. In some embodiments, the plurality of audio signals may include three beamforming signals. In some embodiments, the plurality of audio signals may include more than three beamforming signals. In some embodiments, the plurality of audio signals may include a plurality of beamforming signals, each having a different beamforming directivity pattern.

[0258] In some embodiments, the multiple audio signals may include signals having a dipole pattern. In some embodiments, the multiple audio signals may include beamforming signals having a forward-directed supercardioid directional pattern. In some embodiments, the multiple audio signals may include beamforming signals having a rearward-directed supercardioid directional pattern. In some embodiments, the multiple audio signals may include beamforming signals having a forward-directed cardioid pattern. In some embodiments, the multiple audio signals may include beamforming signals having a rearward-directed cardioid pattern. In some embodiments, the multiple audio signals may include beamforming signals having a forward-directed supercardioid pattern and beamforming signals having a rearward-directed supercardioid pattern.

[0259] In some embodiments, the multiple audio signals may include a beamforming signal having a forward-directed cardioid pattern and a beamforming signal having a backward-directed cardioid pattern. In some embodiments, the multiple audio signals may include a beamforming signal having a forward-directed supercardioid pattern, a beamforming signal having a backward-directed supercardioid pattern, and a beamforming signal having a dipole pattern. In some embodiments, the multiple audio signals may include a beamforming signal having a forward-directed cardioid pattern, a beamforming signal having a backward-directed cardioid pattern, and a beamforming signal having a dipole pattern.

[0260] As described above, binaural communication can occur in a system of two hearing aids (e.g., two hearing aids 100), cochlear implants, or earphones communicating with each other via a wireless communication link. Binaural communication can also occur in ear-worn devices such as eyeglasses with built-in hearing aids (e.g., eyeglasses 300), where a device in one ear-closer part of the eyeglasses can communicate with a device in the other ear-closer part of the eyeglasses via a wired communication link within the eyeglasses. In some embodiments, binaural communication can facilitate the communication of spatial information. For example, a mask or sound map may be communicated from one device to another. In some embodiments, binaural communication may be performed via a low-latency communication link, such as a Near-Field Magnetic Induction (NFMI) communication link.

[0261] Any neural network circuit described herein (e.g., neural network circuits 526, 826, 926, 1026, 1126, 1726, and / or 2126) may include circuits configured to perform operations necessary to compute the output of the neural network layer. One such operation may be matrix-vector multiplication. In some embodiments, the neural network circuit may include a plurality of identical tiles, each containing a plurality of multiply-accumulate circuits configured to perform intermediate calculations of matrix-vector multiplication in parallel, and then aggregate the results of the intermediate calculations to compute the final result. Each tile may further include a memory configured to store neural network weights, a register configured to store input activation elements, and routing circuits configured to facilitate communication of state and data between tiles. Other types of circuits configured to perform the processing described herein, such as any of the masking and subtraction circuits (e.g., masking and subtraction circuits 832, 932, 1032, 1132, and / or 2132), mixing circuits (e.g., mixing circuits 834, 1434, 1534, and / or 2134), and / or WDRC circuits (e.g., WDRC circuits 1258 and / or 1458), may be implemented as digital processing circuits.

[0262] In some embodiments, such digital processing circuits may utilize a Single Instruction Multiple Data (SIMD) architecture. Any of the in-ear devices described herein (e.g., hearing aid 100, eyeglasses 300, in-ear device 400, and / or in-ear device 500) may include a chip that implements a particular part of the circuit. For example, any of the noise reduction circuits described herein (or, in some embodiments, one of the other types of circuits) may be implemented (whole or partially) on a chip. Thus, the chip may include the tiles and digital processing circuits described above.

[0263] In some embodiments, for a model with up to 10M 8-bit weights, operating at 100 GOPs / second on time-series data, the chip can achieve a power efficiency of 4 GOPs / milliwatt, measured at 40°C, when the chip uses a supply voltage of 0.5–1.8V and is operating without idle periods. A further description of chips incorporating neural network circuits (or, in some embodiments, one of other elements) used in ear-worn devices can be found in U.S. Patent No. 11,886,974, “Neural Network Chip for Ear-Worn Devices,” issued January 30, 2024, the entire disclosure of which is incorporated herein by reference. In some embodiments, in addition to a chip including some or all of noise reduction circuits, any ear-worn device described herein may include a digital signal processor configured to perform other operations, such as some or all of the operations performed by processing circuits 522 and / or 528.

[0264] (example) Example 1 relates to an ear-worn device comprising two or more microphones and a noise reduction circuit including a neural network circuit, wherein the neural network circuit is configured to receive a plurality of audio signals, each of which at least two of the plurality of audio signals is obtained from one of the two or more microphones and / or at least one of the plurality of audio signals is a beamforming audio signal obtained from the two or more microphones, and to implement one or more neural network layers that have been trained to perform background noise correction and spatial focusing on the plurality of audio signals, such that the neural network circuit generates one or more neural network outputs based on the plurality of audio signals, and the noise reduction circuit is configured to output an output audio signal that includes a background noise correction and spatial focusing applied version of a first audio signal of the plurality of audio signals, based on the one or more neural network outputs.

[0265] Example 2 relates to the ear-mounted device according to Example 1, wherein at least two of the plurality of audio signals have different beamforming directivity patterns.

[0266] Example 3 relates to the ear-mounted device according to Example 1 or 2, wherein at least one of the plurality of audio signals has a forward-directed beamforming directivity pattern, and at least one of the plurality of audio signals has a rearward-directed beamforming directivity pattern.

[0267] Example 4 relates to the ear-mounted device according to any one of Examples 1 to 3, wherein the output audio signal has a specific spatial focusing pattern, and the specific spatial focusing pattern includes different weights applied to speech sounds obtained from different arrival directions with respect to the wearer of the ear-mounted device in the first audio signal.

[0268] Example 5 relates to the ear-mounted device according to Example 4, wherein the weight applied to the speech sound obtained from the arrival direction with respect to the front of the wearer of the ear-mounted device is higher than the weights applied to the speech sounds obtained from the arrival directions with respect to the sides and rear of the wearer of the ear-mounted device.

[0269] Example 6 relates to the ear-mounted device according to any one of Examples 3 to 5, wherein the neural network circuit is further configured to receive one or more spatial focusing control inputs indicating the specific spatial focusing pattern, and use the one or more spatial focusing control inputs to generate the one or more neural network outputs such that the output audio signal has the specific spatial focusing pattern.

[0270] Example 7 relates to the ear-mounted device according to any one of Examples 3 to 6, further comprising a communication circuit that receives an instruction for user selection of the specific spatial focusing pattern from a processing device, and a control circuit configured to generate the one or more spatial focusing control inputs indicating the specific spatial focusing pattern based at least in part on the instruction for the user selection of the specific spatial focusing pattern.

[0271] Example 8 relates to a system comprising the ear-mounted device according to Example 7, and the processing device that communicates with the ear-mounted device, the processing device being configured to display a graphical user interface including options for different spatial focusing patterns and to receive the user selection of the specific spatial focusing pattern.

[0272] Example 9 is a target speech audio signal which includes The present invention relates to an ear-worn device according to any one of Examples 1 to 8, comprising an interfering speech audio signal including a second version of the first audio signal, and a background noise signal including background noise in the first audio signal, wherein the noise reduction circuit generates the output audio signal such that the change in volume of the background noise signal differs from the change in volume of the target speech audio signal by a first volume change difference, and the change in volume of the interfering speech audio signal differs from the change in volume of the target speech audio signal by a second volume change difference, and the first volume change difference and the second volume change difference are independently controllable.

[0273] Example 10 relates to the ear-worn device described in Example 9, wherein the interfering speech voice signal includes the remainder when the target speech voice signal is subtracted from the speech voice signal.

[0274] Example 11 relates to an ear-worn device as described in Example 9 or 10, wherein the neural network circuit is configured to use a first subset of the one or more neural network layers to generate a first neural network output from the two or more neural network outputs, and to use a second subset of the one or more neural network layers to generate a second neural network output from the two or more neural network outputs, and the noise reduction circuit is configured to acquire the speech voice signal and / or the background noise signal from the first neural network output from the two or more neural network outputs, and to acquire the target speech voice signal and / or the interfering speech voice signal from the second neural network output from the two or more neural network outputs.

[0275] Example 12 relates to an ear-worn device according to any one of Examples 9 to 11, wherein the two or more neural network outputs include two different masks.

[0276] Example 13 relates to an ear-worn device according to any one of Examples 9 to 12, further comprising a mixing circuit configured to generate the output audio signal by mixing a combination of audio signals or by mixing a combination of masks, or a wide dynamic range compression (WDRC) circuit including a plurality of WDRC pipelines configured to generate the output audio signal by performing wide dynamic range compression (WDRC) on the combination of audio signals.

[0277] Example 14 relates to an ear-worn device as described in Example 13, wherein the mixing circuit is further configured to receive a first volume change control input and a second volume change control input, and to perform the mixing using the first volume change control input and the second volume change control input such that the first volume change difference is at least partially controlled by the first volume change control input and the second volume change difference is at least partially controlled by the second volume change control input.

[0278] Example 15 relates to the ear-worn device described in Example 14, further comprising: a communication circuit configured to receive the first volume change control input and the second volume change control input from a processing device; a memory configured to store the first volume change control input and the second volume change control input; and a control circuit configured to extract the first volume change control input and the second volume change control input from the memory and output the first volume change control input and the second volume change control input to the mixing circuit.

[0279] Example 16 relates to an in-ear device according to any one of Examples 1 to 12, further comprising a mixing circuit configured to mix two or more audio signals such that the output audio signal includes a version of the first audio signal mixed with a second audio signal, which has undergone background noise correction and spatial focusing.

[0280] Example 17 relates to the ear-worn device described in Example 16, wherein the background noise-corrected and spatially focused version of the first audio signal includes a target speech audio signal, the target speech audio signal includes a first spatially focused version of the speech audio signal, the speech audio signal includes the speech of the first audio signal, and the second audio signal includes a background noise signal including background noise in the first audio signal and an interfering speech audio signal including a second spatially focused version of the speech audio signal, and the noise reduction circuit is configured to generate the output audio signal such that the volume change of the background noise signal differs from the volume change of the target speech audio signal by the volume change difference amount, and the volume change of the interfering speech audio signal differs from the volume change of the target speech audio signal by the volume change difference amount, and the volume change difference amount is controllable.

[0281] Example 18 relates to the ear-worn device described in Example 17, wherein the mixing circuit is configured to receive a volume change control input and to perform the mixing using the volume change control input such that the volume change difference is controlled at least partially by the volume change control input.

[0282] Example 19 relates to the ear-worn device described in Example 18, comprising: a communication circuit configured to receive the volume change control input from a processing device; a memory configured to store the volume change control input; and a control circuit configured to extract the volume change control input and output it to the mixing circuit.

[0283] Example 20 relates to an ear-worn device according to any one of Examples 1 to 19, wherein the ear-worn device is further configured to receive a user selection to turn off spatial focusing.

[0284] Example 21 is an ear-worn device comprising two or more microphones and a noise reduction circuit including a neural network circuit, wherein the neural network circuit is configured to receive a plurality of audio signals, each of which at least two of the plurality of audio signals is obtained from one of the two or more microphones and / or at least one of the plurality of audio signals is a beamforming audio signal obtained from the two or more microphones, and to implement one or more neural network layers that have been trained to perform background noise correction and spatial focusing to generate two or more neural network outputs based on the plurality of audio signals, wherein the noise reduction circuit is configured to generate the output audio signals. The present invention relates to an ear-worn device, wherein the output audio signal includes a target speech audio signal, which includes a first version of the speech audio signal, which is spatially focused, and an interference speech audio signal, which includes a second version of the speech audio signal, which is spatially focused, and a background noise signal, which includes background noise in the first audio signal, and the noise reduction circuit is configured to generate the output audio signal such that the change in volume of the background noise signal differs from the change in volume of the target speech audio signal by a first volume change difference, and the change in volume of the interference speech audio signal differs from the change in volume of the target speech audio signal by a second volume change difference, and the first volume change difference and the second volume change difference are independently controllable.

[0285] Example 22 relates to the ear-worn device described in Example 21, wherein at least two of the plurality of audio signals have different beamforming directional patterns.

[0286] Example 23 relates to an in-ear device as described in Example 21 or 22, wherein the target speech voice signal includes the speech voice signal to which a specific spatial focusing pattern is applied, and the specific spatial focusing pattern includes different weights applied to the speech voice obtained from different directions of arrival to the wearer of the in-ear device.

[0287] Example 24 relates to the in-ear device according to Example 23, wherein the particular spatial focusing pattern includes a weight applied to speech sounds obtained from the direction of arrival in front of the wearer of the in-ear device, which is greater than the weight applied to speech sounds obtained from the direction of arrival to the side and rear of the wearer of the in-ear device.

[0288] Example 25 relates to an ear-worn device as described in Example 23 or 24, wherein the neural network circuit receives one or more spatial focusing control inputs that indicate the specific spatial focusing pattern, and uses the one or more spatial focusing control inputs to generate the two or more neural network outputs such that the target speech audio signal includes the speech audio signal to which the specific spatial focusing pattern has been applied.

[0289] Example 26 relates to the ear-worn device according to Example 25, further comprising: a communication circuit configured to receive instructions for user selection of a particular spatial focusing pattern from a processing device; and a control circuit configured to generate the one or more spatial focusing control inputs indicating the particular spatial focusing pattern, at least in part, based on the instructions for user selection of the particular spatial focusing pattern.

[0290] Example 27 relates to a system comprising the ear-mounted device described in Example 26 and the processing device that communicates with the ear-mounted device and is configured to display a graphical user interface including options for different spatial focusing patterns and receive the user selection of the specific spatial focusing pattern.

[0291] Example 28 relates to the ear-mounted device according to any one of Examples 21 to 27, wherein the interfering speech audio signal includes the remainder when the target speech audio signal is subtracted from the speech audio signal.

[0292] Example 29 relates to the ear-mounted device according to any one of Examples 21 to 28, wherein the neural network circuit is configured to use a first subset of the one or more neural network layers to generate a first neural network output among the two or more neural network outputs, and use a second subset of the one or more neural network layers to generate a second neural network output among the two or more neural network outputs, and the noise reduction circuit is configured to obtain the speech audio signal and / or the background noise signal from the first neural network output among the two or more neural network outputs, and obtain the target speech audio signal and / or the interfering speech audio signal from the second neural network output among the two or more neural network outputs.

[0293] Example 30 relates to the ear-mounted device according to any one of Examples 21 to 29, wherein the two or more neural network outputs include two different masks.

[0294] Example 31 relates to an ear-worn device according to any one of Examples 21 to 30, further comprising a mixing circuit configured to generate the output audio signal by mixing a combination of audio signals or by mixing a combination of masks, or a wide dynamic range compression (WDRC) circuit including a plurality of WDRC pipelines configured to generate the output audio signal by performing wide dynamic range compression (WDRC) on the combination of audio signals.

[0295] Example 32 relates to an ear-worn device according to Example 31, wherein the mixing circuit is further configured to receive a first volume change control input and a second volume change control input, and to perform the mixing using the first volume change control input and the second volume change control input such that the first volume change difference is at least partially controlled by the first volume change control input and the second volume change difference is at least partially controlled by the second volume change control input.

[0296] Example 33 relates to the ear-worn device described in Example 32, further comprising: a communication circuit configured to receive the first volume change control input and the second volume change control input from a processing device; a memory configured to store the first volume change control input and the second volume change control input; and a control circuit configured to extract the first volume change control input and the second volume change control input from the memory and output the first volume change control input and the second volume change control input to the mixing circuit.

[0297] Example 34 relates to an ear-worn device according to Example 32 or 33, further comprising a control circuit configured to generate a first volume control input based on the level of background noise in the first audio signal and a second volume control input based on the level of interfering speech in the first audio signal.

[0298] Example 35 relates to an ear-worn device according to any one of Examples 21-34, wherein the ear-worn device is further configured to receive a user selection to turn off spatial focusing.

[0299] Example 36 relates to an ear-worn device according to any one of Examples 21 to 35, wherein at least one of the two or more neural network outputs includes the speech voice signal, a mask configured to generate the speech voice signal, the background noise signal, a mask configured to generate the background noise signal, the target speech voice signal, a mask configured to generate the target speech voice signal, and the interference speech voice signal, a mask configured to generate the interference speech voice signal.

[0300] Example 37 relates to an ear-worn device according to any one of Examples 21 to 36, wherein the background noise signal is not spatially focused, and the interfering speech signal does not include any of the background noise in the first speech signal.

[0301] Example 38 relates to an ear-worn device according to any one of Examples 21 to 36, wherein the background noise signal includes a first spatially focused version of the background noise in the first audio signal, and the interfering speech audio signal includes a second spatially focused version of the speech audio signal and a second spatially focused version of the background noise in the first audio signal.

[0302] Example 39 relates to an ear-worn device according to any one of Examples 21 to 38, wherein the ear-worn device includes a hearing aid.

[0303] Example 40 relates to an ear-worn device according to any one of Examples 21 to 39, wherein the noise reduction circuit is mounted on a chip.

[0304] Example 41 is an ear-worn device comprising two or more microphones and a noise reduction circuit including a neural network circuit, wherein the neural network circuit is Receiving multiple audio signals, wherein at least two of the multiple audio signals are obtained from one of the two or more microphones, and / or at least one of the multiple audio signals is a beamforming audio signal obtained from the two or more microphones. The neural network circuit is configured to implement one or more neural network layers that have been trained to perform background noise correction and spatial focusing based on the plurality of audio signals, so that it generates one or more neural network outputs based on the plurality of audio signals. The present invention relates to an ear-worn device in which the noise reduction circuit is configured to output an output audio signal that includes a version of a first audio signal among the plurality of audio signals that has been modified for background noise and / or spatially focused, based on the output of one or more neural networks.

[0305] Example 42 relates to the in-ear device described in Example 41, wherein one or more neural network layers are trained to perform background noise correction, and the output audio signal includes a background noise-corrected version of the first audio signal.

[0306] Example 43 relates to the in-ear device described in Example 41, wherein one or more neural network layers are trained to perform spatial focusing, and the output audio signal includes a spatially focused version of the first audio signal.

[0307] Example 44 relates to the in-ear device described in Example 41, wherein one or more neural network layers are trained to perform background noise correction and spatial focusing, and the output audio signal includes a background noise correction and spatial focusing applied version of the first audio signal.

[0308] Example 45 relates to an ear-mounted device as described in Example 43 or 44, wherein the output audio signal has a specific spatial focusing pattern, and the specific spatial focusing pattern includes different weights applied to spoken audio obtained from different directions of arrival to the wearer of the ear-mounted device.

[0309] Example 46 relates to the in-ear device according to Example 45, wherein the particular spatial focusing pattern includes a weight applied to speech sounds obtained from the direction of arrival in front of the wearer of the in-ear device, which is greater than the weight applied to speech sounds obtained from the direction of arrival to the side and rear of the wearer of the in-ear device.

[0310] Example 47 relates to an in-ear device as described in Example 45 or 46, wherein the neural network circuit is further configured to receive one or more spatial focusing control inputs that exhibit the particular spatial focusing pattern, and to use the one or more spatial focusing control inputs to generate the one or more neural network outputs such that the output audio signal has the particular spatial focusing pattern.

[0311] Example 48 relates to the ear-worn device according to Example 47, further comprising: a communication circuit configured to receive instructions for user selection of a particular spatial focusing pattern from a processing device; and a control circuit configured to generate the one or more spatial focusing control inputs indicating the particular spatial focusing pattern, at least in part, based on the instructions for user selection of the particular spatial focusing pattern.

[0312] Example 49 relates to a system comprising the ear-worn device described in Example 48, and the processing device communicating with the ear-worn device, the processing device being configured to display a graphical user interface including options for different spatial focusing patterns, and to receive the user selection of the particular spatial focusing pattern.

[0313] Example 50 relates to the system described in Example 49, wherein the multiple options include four options.

[0314] Example 51 relates to the system described in Example 49 or 50, wherein the multiple options are graphical representations of the different spatial focusing patterns.

[0315] Example 52 relates to the ear-worn device according to Example 47, further comprising: a sensing circuit configured to generate one or more inputs based on the movement of the ear-worn device; and a control circuit configured to determine the degree of head movement based on the one or more inputs received from one or more sensors, and to generate one or more spatial focusing control inputs indicating the particular spatial focusing pattern based on the degree of head movement.

[0316] Example 53 relates to an ear-worn device as described in Example 52, wherein the control circuit is configured to generate the one or more inputs indicating the spatial focusing pattern based on the degree of head movement, by generating a first set of one or more spatial focusing control inputs indicating a first spatial focusing pattern of a first spatial focusing amount based on a first degree of head movement, and by generating a second set of one or more spatial focusing control inputs indicating a second spatial focusing pattern of a second spatial focusing amount based on a second degree of head movement, wherein the first spatial focusing amount is smaller than the second spatial focusing amount, and the degree of the first head movement is larger than the degree of the second head movement.

[0317] Example 54 relates to the ear-worn device according to Example 47, further comprising a control circuit configured to determine the signal-to-noise ratio (SNR) of an acoustic environment and, based on the SNR of the acoustic environment, generate one or more spatial focusing control inputs that represent the spatial focusing pattern.

[0318] Example 55 relates to an ear-worn device according to Example 43 or 44, wherein the output audio signal includes a target speech audio signal, which includes a first spatially focused version of the speech audio signal, which includes the speech in the first audio signal; an interference speech audio signal, which includes a second spatially focused version of the speech; and a background noise signal, which includes background noise in the first audio signal.

[0319] Example 56 relates to the ear-worn device described in Example 55, wherein the noise reduction circuit generates the output audio signal such that the change in volume of the background noise signal differs from the change in volume of the target speech audio signal by a first volume change difference, the change in volume of the interfering speech audio signal differs from the change in volume of the target speech audio signal by a second volume change difference, and the first volume change difference and the second volume change difference are different from each other.

[0320] Example 57 relates to an ear-worn device as described in Example 55 or 56, wherein the noise reduction circuit generates the output audio signal such that the change in volume of the background noise signal differs from the change in volume of the target speech audio signal by a first volume change difference, the change in volume of the interfering speech audio signal differs from the change in volume of the target speech audio signal by a second volume change difference, and the first and second volume change difference amounts are controllable.

[0321] Example 58 relates to an ear-worn device according to any one of Examples 55-57, wherein the target speech signal includes the speech signal to which a specific spatial focusing pattern is applied, and the specific spatial focusing pattern includes different weights applied to the speech signal obtained from different directions of arrival to the wearer of the ear-worn device.

[0322] Example 59 relates to the in-ear device according to Example 58, wherein the particular spatial focusing pattern includes a weight applied to speech sounds obtained from the direction of arrival in front of the wearer of the in-ear device, which is greater than the weight applied to speech sounds obtained from the direction of arrival to the side and rear of the wearer of the in-ear device.

[0323] Example 60 relates to an ear-worn device as described in Example 58 or 59, wherein the neural network circuit receives one or more spatial focusing control inputs that indicate the particular spatial focusing pattern, and uses the one or more spatial focusing control inputs to generate the two or more neural network outputs such that the target speech audio signal includes the speech audio signal to which the particular spatial focusing pattern has been applied.

[0324] Example 61 relates to the ear-worn device according to Example 60, further comprising: a communication circuit configured to receive instructions for user selection of a particular spatial focusing pattern from a processing device; and a control circuit configured to generate the one or more spatial focusing control inputs indicating the particular spatial focusing pattern, at least in part, based on the instructions for user selection of the particular spatial focusing pattern.

[0325] Example 62 relates to a system comprising the ear-worn device described in Example 61, and the processing device communicating with the ear-worn device, the processing device configured to display a graphical user interface including options for different spatial focusing patterns, and to receive the user selection of the particular spatial focusing pattern.

[0326] Example 63 relates to the system described in Example 62, wherein the multiple options include four options.

[0327] Example 64 relates to the system described in Example 62 or 63, wherein the multiple options are graphical representations of the different spatial focusing patterns.

[0328] Example 65 relates to the ear-worn device according to Example 60, further comprising: a sensing circuit configured to generate one or more inputs based on the movement of the ear-worn device; and a control circuit configured to determine the degree of head movement based on the one or more inputs received from one or more sensors, and to generate one or more spatial focusing control inputs indicating the particular spatial focusing pattern based on the degree of head movement.

[0329] Example 66 relates to an ear-worn device as described in Example 65, wherein the control circuit is configured to generate the one or more inputs indicating the spatial focusing pattern based on the degree of head movement, by generating a first set of one or more spatial focusing control inputs indicating a first spatial focusing pattern of a first spatial focusing amount based on a first degree of head movement, and by generating a second set of one or more spatial focusing control inputs indicating a second spatial focusing pattern of a second spatial focusing amount based on a second degree of head movement, wherein the first spatial focusing amount is smaller than the second spatial focusing amount, and the degree of the first head movement is larger than the degree of the second head movement.

[0330] Example 67 relates to the ear-worn device according to Example 60, further comprising a control circuit configured to determine the signal-to-noise ratio (SNR) of an acoustic environment and, based on the SNR of the acoustic environment, generate one or more spatial focusing control inputs that represent the spatial focusing pattern.

[0331] Example 68 relates to an ear-worn device according to any one of Examples 55 to 67, wherein the interfering speech voice signal includes the remainder when the target speech voice signal is subtracted from the speech voice signal.

[0332] Example 69 relates to an ear-worn device according to any one of Examples 55 to 67, wherein the one or more neural network outputs include two or more neural network outputs.

[0333] Example 70 relates to an ear-worn device according to Example 69, wherein the noise reduction circuit is configured to obtain, based on the two or more neural network outputs, at least one of a speech speech signal including speech in a first speech signal among the plurality of speech signals, a background noise signal including background noise in the first speech signal, a target speech speech signal including a first spatially focused version of the speech speech signal, and an interference speech speech signal including a second spatially focused version of the speech speech signal.

[0334] Example 71 relates to an ear-worn device as described in Example 69, wherein the noise reduction circuit is configured to acquire, based on the two or more neural network outputs, a target speech signal including a first spatially focused version of the speech signal, and an interfering speech signal including a second spatially focused version of the speech signal.

[0335] Example 72 relates to an ear-worn device according to any one of Examples 69 to 71, wherein the neural network circuit is configured to use a first subset of the one or more neural network layers to generate a first neural network output from the two or more neural network outputs, and to use a second subset of the one or more neural network layers to generate a second neural network output from the two or more neural network outputs, and the noise reduction circuit is configured to acquire the speech voice signal and / or the background noise signal from the first neural network output from the two or more neural network outputs, and to acquire the target speech voice signal and / or the interfering speech voice signal from the second neural network output from the two or more neural network outputs.

[0336] Example 73 relates to an ear-worn device according to any one of Examples 69-72, wherein the two or more neural network outputs include two different masks.

[0337] Example 74 relates to an ear-worn device according to any one of claims 69 to 73, wherein at least one of the two or more neural network outputs includes the speech voice signal, a mask configured to generate the speech voice signal, the background noise signal, a mask configured to generate the background noise signal, the target speech voice signal, a mask configured to generate the target speech voice signal, and the interference speech voice signal, a mask configured to generate the interference speech voice signal.

[0338] Example 75 relates to an ear-worn device according to any one of Examples 55 to 74, further comprising a mixing circuit configured to generate the output audio signal by mixing a combination of audio signals or by mixing a combination of masks, or a wide dynamic range compression (WDRC) circuit including a plurality of WDRC pipelines configured to generate the output audio signal by performing wide dynamic range compression (WDRC) on the combination of audio signals.

[0339] Example 76 relates to an ear-worn device according to Example 75, wherein the mixing circuit is further configured to receive a first volume change control input and a second volume change control input, and to perform the mixing using the first volume change control input and the second volume change control input such that the first volume change difference is controlled at least partially by the first volume change control input and the second volume change difference is controlled at least partially by the second volume change control input.

[0340] Example 77 relates to the ear-worn device described in Example 76, further comprising: a communication circuit configured to receive the first volume change control input and the second volume change control input from a processing device; a memory configured to store the first volume change control input and the second volume change control input; and a control circuit configured to extract the first volume change control input and the second volume change control input from the memory and output the first volume change control input and the second volume change control input to the mixing circuit.

[0341] Example 78 relates to an ear-worn device according to any one of Examples 55 to 74, wherein the neural network circuit is further configured to receive a first volume change control input and a second volume change control input, and to use the first volume change control input and the second volume change control input to generate one or more neural network outputs such that the first volume change difference is at least partially controlled by the first volume change control input and the second volume change difference is at least partially controlled by the second volume change control input.

[0342] Example 79 relates to the ear-worn device described in Example 78, further comprising: a communication circuit configured to receive the first volume change control input and the second volume change control input from a processing device; a memory configured to store the first volume change control input and the second volume change control input; and a control circuit configured to extract the first volume change control input and the second volume change control input from the memory and output the first volume change control input and the second volume change control input to the mixing circuit.

[0343] Example 80 relates to an ear-worn device according to any one of Examples 76 to 79, further comprising a control circuit configured to generate a first volume control input based on the level of background noise in the first audio signal and a second volume control input based on the level of interfering speech in the first audio signal.

[0344] Example 81 relates to an ear-worn device according to any one of Examples 55 to 80, wherein the background noise signal is not spatially focused, and the interfering speech signal does not include any portion of the background noise in the first speech signal.

[0345] Example 82 relates to an ear-worn device according to any one of Examples 55 to 80, wherein the background noise signal includes a first spatially focused version of the background noise in the first audio signal, and the interfering speech audio signal includes a second spatially focused version of the speech audio signal and a second spatially focused version of the background noise in the first audio signal.

[0346] Example 83 relates to an ear-worn device according to any one of Examples 43 to 82, wherein the ear-worn device is further configured to receive a user selection to turn off spatial focusing.

[0347] Example 84 relates to an ear-worn device according to any one of Examples 41 to 83, wherein at least two of the plurality of audio signals have different beamforming directional patterns.

[0348] Example 85 relates to the ear-worn device described in Example 84, wherein at least two of the plurality of audio signals having different beamforming directional patterns include beamforming signals having dipole, hypercardioid, supercardioid, or cardioid directional patterns.

[0349] Example 86 relates to an in-ear device according to any one of Examples 41-74 and 78-85, further comprising a mixing circuit configured to mix two or more audio signals such that the output audio signal includes a version of the first audio signal that has been mixed with a second audio signal and subjected to background noise correction and / or spatial focusing.

[0350] Example 87 relates to an ear-worn device as described in Example 86, wherein the background noise-corrected and spatially focused version of the first audio signal includes a target speech audio signal, the target speech audio signal includes a first spatially focused version of the speech audio signal, the speech audio signal includes the speech of the first audio signal, and the second audio signal includes a background noise signal including background noise in the first audio signal and an interfering speech audio signal including a second spatially focused version of the speech audio signal, and the noise reduction circuit is configured to generate the output audio signal such that the volume change of the background noise signal differs from the volume change of the target speech audio signal by the volume change difference amount, and the volume change of the interfering speech audio signal differs from the volume change of the target speech audio signal by the volume change difference amount, and the volume change difference amount is controllable.

[0351] Example 88 relates to the ear-worn device described in Example 87, wherein the mixing circuit is configured to receive a volume change control input and to perform the mixing using the volume change control input such that the volume change difference is controlled at least partially by the volume change control input, and the ear-worn device comprises a communication circuit configured to receive the volume change control input from a processing device, a memory configured to store the volume change control input, and a control circuit configured to take out the volume change control input and output the volume change control input to the mixing circuit.

[0352] Example 89 relates to an ear-worn device according to any one of Examples 41 to 88, wherein the ear-worn device includes a hearing aid.

[0353] Example 90 relates to an ear-worn device according to any one of Examples 41 to 89, wherein the noise reduction circuit is mounted on a chip.

[0354] Example 91 is an ear-worn device comprising two or more microphones and a noise reduction circuit including a neural network circuit, wherein the neural network circuit is Receiving multiple audio signals, wherein at least two of the multiple audio signals are obtained from one of the two or more microphones, and / or at least one of the multiple audio signals is a beamforming audio signal obtained from the two or more microphones. The invention relates to an ear-worn device, wherein the neural network circuit is configured to implement one or more neural network layers that are trained to perform background noise correction and spatial focusing on the plurality of audio signals, or to produce an output used when performing background noise correction and spatial focusing, so that the neural network circuit generates one or more neural network outputs based on the plurality of audio signals, and the noise reduction circuit is configured to output an output audio signal that includes a background noise-corrected and spatially focused version of a first audio signal of the plurality of audio signals, based on the one or more neural network outputs.

[0355] Example 92 relates to the ear-worn device described in Example 91, wherein at least two of the plurality of audio signals have different beamforming directional patterns.

[0356] Example 93 relates to an ear-worn device according to Example 91 or 92, wherein at least one of the plurality of audio signals has a forward-directed beamforming directional pattern, and at least one of the plurality of audio signals has a backward-directed beamforming directional pattern.

[0357] Example 94 relates to an ear-mounted device according to any one of Examples 91 to 93, wherein the output audio signal has a specific spatial focusing pattern, the specific spatial focusing pattern includes different weights applied to the first audio signal to the spoken audio obtained from different directions of arrival to the wearer of the ear-mounted device.

[0358] Example 95 relates to the in-ear device according to Example 94, wherein the particular spatial focusing pattern includes a weight applied to speech sounds obtained from the direction of arrival in front of the wearer of the in-ear device, which is higher than the weight applied to speech sounds obtained from the direction of arrival to the side and rear of the wearer of the in-ear device.

[0359] Example 96 relates to an ear-worn device according to any one of Examples 93 to 95, wherein the neural network circuit is further configured to receive one or more spatial focusing control inputs that exhibit the particular spatial focusing pattern, and to use the one or more spatial focusing control inputs to generate the one or more neural network outputs such that the output audio signal has the particular spatial focusing pattern.

[0360] Example 97 relates to an ear-worn device according to any one of Examples 93 to 96, further comprising: a communication circuit for receiving instructions for user selection of a particular spatial focusing pattern from a processing device; and a control circuit configured to generate the one or more spatial focusing control inputs indicating the particular spatial focusing pattern, at least in part, based on the instructions for user selection of the particular spatial focusing pattern.

[0361] Example 98 relates to a system comprising the ear-worn device described in Example 97, and the processing device communicating with the ear-worn device, the processing device configured to display a graphical user interface including options for different spatial focusing patterns, and to receive the user selection of the particular spatial focusing pattern.

[0362] Example 99 is a target speech audio signal which includes a target speech audio signal which includes a target speech audio signal which includes a target speech audio signal which includes a target speech audio signal which includes a target speech audio signal which includes a target speech audio signal which includes a target speech audio signal which includes a target speech audio signal which includes a target speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a speech audio signal which includes a target The present invention relates to an ear-worn device according to any one of Examples 91 to 98, comprising an interfering speech audio signal including a second applied version, and a background noise signal including background noise in the first audio signal, wherein the noise reduction circuit generates the output audio signal such that the change in volume of the background noise signal differs from the change in volume of the target speech audio signal by a first volume change difference, and the change in volume of the interfering speech audio signal differs from the change in volume of the target speech audio signal by a second volume change difference, and the first and second volume change differences are independently controllable.

[0363] Example 100 relates to the ear-worn device described in Example 99, wherein the interfering speech voice signal includes the remainder when the target speech voice signal is subtracted from the speech voice signal.

[0364] Example 101 relates to an ear-worn device as described in Example 99 or 100, wherein the neural network circuit is configured to use a first subset of the one or more neural network layers to generate a first neural network output from the two or more neural network outputs, and to use a second subset of the one or more neural network layers to generate a second neural network output from the two or more neural network outputs, and the noise reduction circuit is configured to acquire the speech voice signal and / or the background noise signal from the first neural network output from the two or more neural network outputs, and to acquire the target speech voice signal and / or the interfering speech voice signal from the second neural network output from the two or more neural network outputs.

[0365] Example 102 relates to an ear-worn device according to any one of Examples 99 to 101, wherein the two or more neural network outputs include two different masks.

[0366] Example 103 relates to an ear-worn device according to any one of Examples 99 to 102, further comprising a mixing circuit configured to generate the output audio signal by mixing a combination of audio signals or by mixing a combination of masks, or a wide dynamic range compression (WDRC) circuit including a plurality of WDRC pipelines configured to generate the output audio signal by performing wide dynamic range compression (WDRC) on the combination of audio signals.

[0367] Example 104 relates to an ear-worn device as described in Example 103, wherein the mixing circuit is further configured to receive a first volume change control input and a second volume change control input, and to perform the mixing using the first volume change control input and the second volume change control input such that the first volume change difference is controlled at least partially by the first volume change control input and the second volume change difference is controlled at least partially by the second volume change control input.

[0368] Example 105 relates to the ear-worn device described in Example 104, further comprising: a communication circuit configured to receive the first volume change control input and the second volume change control input from a processing device; a memory configured to store the first volume change control input and the second volume change control input; and a control circuit configured to extract the first volume change control input and the second volume change control input from the memory and output the first volume change control input and the second volume change control input to the mixing circuit.

[0369] Example 106 relates to an in-ear device according to any one of Examples 91 to 102, further comprising a mixing circuit configured to mix two or more audio signals such that the output audio signal includes a version of the first audio signal mixed with a second audio signal, which has undergone background noise correction and spatial focusing.

[0370] Example 107 relates to an ear-worn device as described in Example 106, wherein the background noise-corrected and spatially focused version of the first audio signal includes a target speech audio signal, the target speech audio signal includes a first spatially focused version of the speech audio signal, the speech audio signal includes the speech of the first audio signal, and the second audio signal includes a background noise signal including background noise in the first audio signal and an interfering speech audio signal including a second spatially focused version of the speech audio signal, and the noise reduction circuit is configured to generate the output audio signal such that the volume change of the background noise signal differs from the volume change of the target speech audio signal by the volume change difference amount, and the volume change of the interfering speech audio signal differs from the volume change of the target speech audio signal by the volume change difference amount, and the volume change difference amount is controllable.

[0371] Example 108 relates to the ear-worn device described in Example 107, wherein the mixing circuit is configured to receive a volume change control input and to perform the mixing using the volume change control input such that the volume change difference is controlled at least partially by the volume change control input.

[0372] Example 109 relates to the ear-worn device described in Example 108, comprising: a communication circuit configured to receive the volume change control input from a processing device; a memory configured to store the volume change control input; and a control circuit configured to extract the volume change control input and output it to the mixing circuit.

[0373] Example 110 relates to an ear-worn device according to any one of Examples 91 to 109, wherein the ear-worn device is further configured to receive a user selection to turn off spatial focusing.

[0374] Example 111 relates to an ear-worn device comprising two or more microphones and a noise reduction circuit including a neural network circuit, wherein the neural network circuit is configured to receive a plurality of audio signals, each of which at least two of the plurality of audio signals is obtained from one of the two or more microphones and / or at least one of the plurality of audio signals is a beamforming audio signal obtained from the two or more microphones, and to implement one or more neural network layers trained to generate one or more neural network outputs based on the plurality of audio signals, wherein the one or more neural network outputs include an output audio signal that includes a background noise-corrected and spatially focused version of a first audio signal of the plurality of audio signals, or the one or more neural network outputs are configured to be used by the noise reduction circuit to generate the output audio signal that includes a background noise-corrected and spatially focused version of the first audio signal.

[0375] Example 112 relates to the ear-worn device described in Example 111, wherein at least two of the plurality of audio signals have different beamforming directional patterns.

[0376] Example 113 relates to an ear-worn device according to Example 111 or 112, wherein at least one of the plurality of audio signals has a forward-directed beamforming directional pattern, and at least one of the plurality of audio signals has a backward-directed beamforming directional pattern.

[0377] Example 114 relates to an ear-mounted device according to any one of Examples 111 to 113, wherein the output audio signal has a specific spatial focusing pattern, the specific spatial focusing pattern includes different weights applied to the first audio signal to the spoken speech obtained from different directions of arrival to the wearer of the ear-mounted device.

[0378] Example 115 relates to the in-ear device described in Example 114, wherein the particular spatial focusing pattern includes a weight applied to speech sounds obtained from the direction of arrival in front of the wearer of the in-ear device, which is greater than the weight applied to speech sounds obtained from the direction of arrival to the side and rear of the wearer of the in-ear device.

[0379] Example 116 relates to an ear-worn device according to any one of Examples 113 to 115, wherein the neural network circuit is further configured to receive one or more spatial focusing control inputs that exhibit the particular spatial focusing pattern, and to use the one or more spatial focusing control inputs to generate the one or more neural network outputs such that the output audio signal has the particular spatial focusing pattern.

[0380] Example 117 relates to an ear-worn device according to any one of Examples 113 to 116, further comprising: a communication circuit for receiving instructions for user selection of a particular spatial focusing pattern from a processing device; and a control circuit configured to generate the one or more spatial focusing control inputs indicating the particular spatial focusing pattern, at least in part, based on the instructions for user selection of the particular spatial focusing pattern.

[0381] Example 118 relates to a system comprising the ear-worn device described in Example 117, and the processing device communicating with the ear-worn device, the processing device configured to display a graphical user interface including options for different spatial focusing patterns, and to receive the user selection of the particular spatial focusing pattern.

[0382] Example 119 is a target speech audio signal which includes The present invention relates to an ear-worn device according to any one of Examples 111 to 118, comprising an interfering speech audio signal including a second version thereof, and a background noise signal including background noise in the first audio signal, wherein the noise reduction circuit generates the output audio signal such that the change in volume of the background noise signal differs from the change in volume of the target speech audio signal by a first volume change difference, and the change in volume of the interfering speech audio signal differs from the change in volume of the target speech audio signal by a second volume change difference, and the first volume change difference and the second volume change difference are independently controllable.

[0383] Example 120 relates to the ear-worn device described in Example 119, wherein the interfering speech voice signal includes the remainder when the target speech voice signal is subtracted from the speech voice signal.

[0384] Example 121 relates to an ear-worn device as described in Example 119 or 120, wherein the neural network circuit is configured to use a first subset of the one or more neural network layers to generate a first neural network output from the two or more neural network outputs, and to use a second subset of the one or more neural network layers to generate a second neural network output from the two or more neural network outputs, and the noise reduction circuit is configured to acquire the speech voice signal and / or the background noise signal from the first neural network output from the two or more neural network outputs, and to acquire the target speech voice signal and / or the interfering speech voice signal from the second neural network output from the two or more neural network outputs.

[0385] Example 122 relates to an ear-worn device as described in any one of Examples 119 to 121, wherein the two or more neural network outputs include two different masks.

[0386] Example 123 relates to an ear-worn device according to any one of Examples 119 to 122, further comprising a mixing circuit configured to generate the output audio signal by mixing a combination of audio signals or by mixing a combination of masks, or a wide dynamic range compression (WDRC) circuit including a plurality of WDRC pipelines configured to generate the output audio signal by performing wide dynamic range compression (WDRC) on the combination of audio signals.

[0387] Example 124 relates to an ear-worn device as described in Example 123, wherein the mixing circuit is further configured to receive a first volume change control input and a second volume change control input, and to perform the mixing using the first volume change control input and the second volume change control input such that the first volume change difference is at least partially controlled by the first volume change control input and the second volume change difference is at least partially controlled by the second volume change control input.

[0388] Example 125 relates to the ear-worn device described in Example 124, further comprising: a communication circuit configured to receive the first volume change control input and the second volume change control input from a processing device; a memory configured to store the first volume change control input and the second volume change control input; and a control circuit configured to extract the first volume change control input and the second volume change control input from the memory and output the first volume change control input and the second volume change control input to the mixing circuit.

[0389] Example 126 relates to an in-ear device according to any one of Examples 111 to 122, further comprising a mixing circuit configured to mix two or more audio signals such that the output audio signal includes a version of the first audio signal mixed with a second audio signal, which has undergone background noise correction and spatial focusing.

[0390] Example 127 relates to an ear-worn device as described in Example 126, wherein the background noise-corrected and spatially focused version of the first audio signal includes a target speech audio signal, the target speech audio signal includes a first spatially focused version of the speech audio signal, the speech audio signal includes the speech of the first audio signal, and the second audio signal includes a background noise signal including background noise in the first audio signal and an interfering speech audio signal including a second spatially focused version of the speech audio signal, and the noise reduction circuit is configured to generate the output audio signal such that the volume change of the background noise signal differs from the volume change of the target speech audio signal by a volume change difference amount, and the volume change of the interfering speech audio signal differs from the volume change of the target speech audio signal by a volume change difference amount, and the volume change difference amount is controllable.

[0391] Example 128 relates to the ear-worn device described in Example 127, wherein the mixing circuit is configured to receive a volume change control input and to perform the mixing using the volume change control input such that the volume change difference is controlled at least partially by the volume change control input.

[0392] Example 129 relates to the ear-worn device described in Example 128, comprising: a communication circuit configured to receive the volume change control input from a processing device; a memory configured to store the volume change control input; and a control circuit configured to extract the volume change control input and output it to the mixing circuit.

[0393] Example 130 relates to an ear-worn device according to any one of Examples 111 to 129, wherein the ear-worn device is further configured to receive a user selection to turn off spatial focusing.

[0394] After describing in detail several embodiments of the technology, those skilled in the art will readily come up with various modifications and improvements. Such modifications and improvements are intended to fall within the spirit and scope of the invention. Therefore, the foregoing description is illustrative and not intended to limit. For example, any of the above components may include hardware, software, or a combination of hardware and software.

[0395] As used herein and in the claims, the indefinite articles "a" and "an" mean "at least one" unless explicitly indicated otherwise.

[0396] As used herein and in the claims, the phrase “and / or” means “either or both” of the combined elements, that is, elements that may be present together in some cases and selectively in others. Any multiple elements listed in “and / or” should be interpreted similarly, that is, “one or more” of the combined elements. In addition to the elements specifically identified by the “and / or” clause, other elements may be present, whether or not they are related to the specifically identified elements.

[0397] As used herein and in the claims, the phrase “at least one” is understood to refer to a list of one or more elements and mean at least one element selected from any one or more elements in the list, but not necessarily including at least one of each element specifically enumerated in the list, nor does it exclude any combination of elements in the list. This definition also allows for the presence of any non-essential elements in the list of elements referred to by the phrase “at least one,” whether or not they are related to the element in question.

[0398] The terms "approximately" and "about" may mean within ±20% of the target value in some embodiments, within ±10% of the target value in some embodiments, within ±5% of the target value in some embodiments, and within ±2% of the target value in some embodiments. The terms "approximately" and "about" may include the target value.

[0399] Furthermore, the expressions and terms used herein are for illustrative purposes only and should not be considered limiting. In this specification, the use of “including,” “comprising,” “having,” “containing,” “involving,” and variations thereof is intended to encompass the items listed therein and their equivalents, as well as any additional items.

[0400] Although several aspects of at least one embodiment have been described above, various changes, modifications, and improvements will readily come to mind for those skilled in the art. Such changes, modifications, and improvements are intended to be the subject of this disclosure. Accordingly, the foregoing description and drawings are illustrative only.

Claims

1. Two or more microphones, A noise reduction circuit including a neural network circuit, An ear-worn device equipped with, The aforementioned neural network circuit is Receiving multiple audio signals, wherein at least two of the multiple audio signals are obtained from one of the two or more microphones, and / or at least one of the multiple audio signals is a beamforming audio signal obtained from the two or more microphones. The neural network circuit implements one or more neural network layers trained to perform background noise correction and spatial focusing based on the multiple audio signals, such that the neural network circuit generates one or more neural network outputs based on the multiple audio signals. It is configured to do the following: The noise reduction circuit is configured to output an output audio signal that includes a version of the first audio signal of the plurality of audio signals that has been modified for background noise correction and spatial focusing, based on the output of one or more neural networks. An ear-worn device.

2. At least two of the aforementioned plurality of audio signals have different beamforming directivity patterns. The ear-worn device according to claim 1.

3. At least one of the plurality of audio signals has a forward-directed beamforming directivity pattern, and at least one of the plurality of audio signals has a backward-directed beamforming directivity pattern. The ear-worn device according to claim 1 or 2.

4. The output audio signal has a specific spatial focusing pattern, and the specific spatial focusing pattern includes different weights applied to the first audio signal to the spoken speech obtained from different directions of arrival to the wearer of the ear-mounted device. An ear-worn device according to any one of claims 1 to 3.

5. The particular spatial focusing pattern includes the fact that the weight applied to speech sounds obtained from the direction of arrival in front of the wearer of the ear-mounted device is higher than the weight applied to speech sounds obtained from the direction of arrival to the side and rear of the wearer of the ear-mounted device. The ear-worn device according to claim 4.

6. The aforementioned neural network circuit is Receiving one or more spatial focusing control inputs that indicate the specific spatial focusing pattern, Using the one or more spatial focusing control inputs to generate the one or more neural network outputs such that the output audio signal has the specific spatial focusing pattern, It is further configured to do the following: An ear-worn device according to any one of claims 3 to 5.

7. A communication circuit that receives instructions from a processing device for the user to select a specific spatial focusing pattern, The system further comprises a control circuit configured to generate one or more spatial focusing control inputs indicating the specific spatial focusing pattern, based at least in part on the user's instruction for the specific spatial focusing pattern, An ear-worn device according to any one of claims 3 to 6.

8. The ear-worn device according to claim 7, The processing device communicates with the ear-worn device, Displaying a graphical user interface that includes options for different spatial focusing patterns, Receiving the user selection of the specific spatial focusing pattern, The processing device is configured to perform the following: A system that includes these features.

9. The aforementioned one or more neural network outputs include two or more neural network outputs, The noise reduction circuit generates an output audio signal based on the two or more neural network outputs. The output audio signal is A target speech audio signal comprising a version of the first audio signal with background noise correction and spatial focusing, wherein the target speech audio signal comprises a first version of the speech audio signal with spatial focusing, and the speech audio signal comprises the speech in the first audio signal, Interfering speech signal including a second version of the aforementioned speech signal with spatial focusing applied, A background noise signal including background noise in the first audio signal, The noise reduction circuit, in the output audio signal, The change in volume of the background noise signal differs from the change in volume of the target speech signal by a first volume change difference amount. The change in volume of the interfering speech signal differs from the change in volume of the target speech signal by a second volume change difference amount. The first volume change difference and the second volume change difference are controlled independently. The system is configured to generate the aforementioned output audio signal. An ear-worn device according to any one of claims 1 to 8.

10. The interference speech audio signal includes the remainder when the target speech audio signal is subtracted from the speech audio signal. The ear-worn device according to claim 9.

11. The neural network circuit is configured to use a first subset of the one or more neural network layers to generate a first neural network output among the two or more neural network outputs, and to use a second subset of the one or more neural network layers to generate a second neural network output among the two or more neural network outputs. The noise reduction circuit is configured to acquire the speech audio signal and / or the background noise signal from the first neural network output among the two or more neural network outputs, and to acquire the target speech audio signal and / or the interfering speech audio signal from the second neural network output among the two or more neural network outputs. The ear-worn device according to claim 9 or 10.

12. The two or more neural network outputs include two different masks, An ear-worn device according to any one of claims 9 to 11.

13. The output audio signal is generated by mixing a combination of audio signals, or A mixing circuit configured to generate the output audio signal by mixing a combination of masks, or The system further includes a wide dynamic range compression (WDRC) circuit comprising a plurality of WDRC pipelines configured to generate the output audio signal by performing wide dynamic range compression (WDRC) on the combination of audio signals. An ear-worn device according to any one of claims 9 to 12.

14. The aforementioned mixing circuit is Receiving a first volume change control input and a second volume change control input, The mixing is performed using the first volume change control input and the second volume change control input such that the first volume change difference amount is controlled at least partially by the first volume change control input and the second volume change difference amount is controlled at least partially by the second volume change control input. It is further configured to do the following: The ear-worn device according to claim 13.

15. A communication circuit configured to receive the first volume change control input and the second volume change control input from a processing device, A memory configured to store the first volume change control input and the second volume change control input, The system further comprises a control circuit configured to take the first volume change control input and the second volume change control input from the memory and output the first volume change control input and the second volume change control input to the mixing circuit. The ear-worn device according to claim 14.

16. The system further comprises a mixing circuit configured to mix two or more audio signals such that the output audio signal includes a version of the first audio signal that has been mixed with a second audio signal and subjected to background noise correction and spatial focusing. An ear-worn device according to any one of claims 1 to 12.

17. The version of the first audio signal with background noise correction and spatial focusing applied includes the target utterance audio signal, The target speech signal includes a first version of the speech signal that has been spatially focused. The aforementioned speech audio signal includes the speech of the first audio signal, The second audio signal is, The background noise signal including background noise in the first audio signal, The interfering speech signal includes a second version of the aforementioned speech signal that has undergone spatial focusing, The noise reduction circuit, in the output audio signal, The change in volume of the background noise signal differs from the change in volume of the target speech signal by a difference in volume change amount. The change in volume of the interfering speech signal differs from the change in volume of the target speech signal by the amount of the volume change difference. The amount of the volume change difference can be controlled, The system is configured to generate the aforementioned output audio signal. The ear-worn device according to claim 16.

18. The aforementioned mixing circuit is Receiving volume change control input, The mixing is performed using the volume change control input such that the volume change difference is controlled at least partially by the volume change control input, It is configured to do, The ear-worn device according to claim 17.

19. A communication circuit configured to receive the volume change control input from a processing device, A memory configured to store the volume change control input, The system includes a control circuit configured to take the volume change control input and output the volume change control input to the mixing circuit, The ear-worn device according to claim 18.

20. The ear-worn device is further configured to receive a user selection to turn off spatial focusing. An ear-worn device according to any one of claims 1 to 19.

21. Two or more microphones, A noise reduction circuit including a neural network circuit, An ear-worn device equipped with, The aforementioned neural network circuit is Receiving multiple audio signals, wherein at least two of the multiple audio signals are obtained from one of the two or more microphones, and / or at least one of the multiple audio signals is a beamforming audio signal obtained from the two or more microphones. The neural network circuit implements one or more neural network layers that have been trained to perform background noise correction and spatial focusing so as to generate two or more neural network outputs based on the plurality of audio signals. It is configured to do the following: The noise reduction circuit is configured to generate an output audio signal. The output audio signal is A target speech signal including a first version of the speech signal obtained by spatially focusing the speech signal, which includes the speech of the first speech signal among the plurality of speech signals, Interfering speech signal including a second version of the aforementioned speech signal with spatial focusing applied, A background noise signal including background noise in the first audio signal, The noise reduction circuit, in the output audio signal, The change in volume of the background noise signal differs from the change in volume of the target speech signal by a first volume change difference amount. The change in volume of the interfering speech signal differs from the change in volume of the target speech signal by a second volume change difference amount. The first volume change difference and the second volume change difference are controlled independently. The system is configured to generate the aforementioned output audio signal. An ear-worn device.

22. At least two of the aforementioned plurality of audio signals have different beamforming directivity patterns. The ear-worn device according to claim 21.

23. The target speech signal includes the speech signal to which a specific spatial focusing pattern is applied, the specific spatial focusing pattern includes different weights applied to the speech signal obtained from different directions of arrival to the wearer of the ear-mounted device. The ear-worn device according to claim 21 or 22.

24. The particular spatial focusing pattern includes the fact that the weight applied to speech sounds obtained from the direction of arrival in front of the wearer of the ear-mounted device is higher than the weight applied to speech sounds obtained from the direction of arrival to the side and rear of the wearer of the ear-mounted device. The ear-worn device according to claim 23.

25. The aforementioned neural network circuit is Receiving one or more spatial focusing control inputs that indicate the specific spatial focusing pattern, Using one or more spatial focusing control inputs to generate the two or more neural network outputs such that the target speech signal includes the speech signal to which the specific spatial focusing pattern has been applied, It is further configured to do the following: The ear-worn device according to claim 23 or 24.

26. A communication circuit configured to receive instructions from a processing device for user selection of a specific spatial focusing pattern, The system further comprises a control circuit configured to generate one or more spatial focusing control inputs indicating the specific spatial focusing pattern, based at least in part on the user's instruction for the specific spatial focusing pattern, The ear-worn device according to claim 25.

27. The ear-worn device according to claim 26, The processing device communicates with the ear-worn device, Displaying a graphical user interface that includes options for different spatial focusing patterns, Receiving the user selection of the specific spatial focusing pattern, The processing device is configured to perform the following: A system that includes these features.

28. The interference speech audio signal includes the remainder when the target speech audio signal is subtracted from the speech audio signal. An ear-worn device according to any one of claims 21 to 27.

29. The neural network circuit is configured to use a first subset of the one or more neural network layers to generate a first neural network output among the two or more neural network outputs, and to use a second subset of the one or more neural network layers to generate a second neural network output among the two or more neural network outputs. The noise reduction circuit is configured to acquire the speech audio signal and / or the background noise signal from the first neural network output among the two or more neural network outputs, and to acquire the target speech audio signal and / or the interfering speech audio signal from the second neural network output among the two or more neural network outputs. An ear-worn device according to any one of claims 21 to 28.

30. The two or more neural network outputs include two different masks, An ear-worn device according to any one of claims 21 to 29.

31. The output audio signal is generated by mixing a combination of audio signals, or A mixing circuit configured to generate the output audio signal by mixing a combination of masks, or The system further includes a wide dynamic range compression (WDRC) circuit comprising a plurality of WDRC pipelines configured to generate the output audio signal by performing wide dynamic range compression (WDRC) on the combination of audio signals. An ear-worn device according to any one of claims 21 to 30.

32. The aforementioned mixing circuit is Receiving a first volume change control input and a second volume change control input, The mixing is performed using the first volume change control input and the second volume change control input such that the first volume change difference amount is controlled at least partially by the first volume change control input and the second volume change difference amount is controlled at least partially by the second volume change control input. It is further configured to do the following: The ear-worn device according to claim 31.

33. A communication circuit configured to receive the first volume change control input and the second volume change control input from a processing device, A memory configured to store the first volume change control input and the second volume change control input, The system further comprises a control circuit configured to take the first volume change control input and the second volume change control input from the memory and output the first volume change control input and the second volume change control input to the mixing circuit. The ear-worn device according to claim 32.

34. The control circuit is further configured to generate a first volume change control input based on the level of background noise in the first audio signal and to generate a second volume change control input based on the level of interfering speech in the first audio signal. The ear-worn device according to claim 32 or 33.

35. The ear-worn device is further configured to receive a user selection to turn off spatial focusing. An ear-worn device according to any one of claims 21 to 34.

36. At least one of the two or more neural network outputs is The aforementioned speech audio signal, A mask configured to generate the aforementioned speech audio signal, The aforementioned background noise signal and, A mask configured to generate the aforementioned background noise signal, The target speech signal and, A mask configured to generate the target speech signal, The interference utterance audio signal and a mask configured to generate the interference utterance audio signal are included. An ear-worn device according to any one of claims 21 to 35.

37. The aforementioned background noise signal has not undergone spatial focusing. The interference speech signal does not include any of the background noise in the first speech signal. An ear-worn device according to any one of claims 21 to 36.

38. The background noise signal includes a first version of the background noise in the first audio signal that has been spatially focused. The interference speech signal includes a second version of the speech signal with spatial focusing applied, and a second version of the background noise in the first speech signal with spatial focusing applied. An ear-worn device according to any one of claims 21 to 36.

39. The aforementioned ear-worn device includes a hearing aid. An ear-worn device according to any one of claims 21 to 38.

40. The aforementioned noise reduction circuit is mounted on the chip. An ear-worn device according to any one of claims 21 to 39.

41. Two or more microphones, A noise reduction circuit including a neural network circuit, An ear-worn device equipped with, The aforementioned neural network circuit is Receiving multiple audio signals, wherein at least two of the multiple audio signals are obtained from one of the two or more microphones, and / or at least one of the multiple audio signals is a beamforming audio signal obtained from the two or more microphones. The neural network circuit implements one or more neural network layers trained to perform background noise correction and spatial focusing based on the multiple audio signals, such that the neural network circuit generates one or more neural network outputs based on the multiple audio signals. It is configured to do the following: The noise reduction circuit is configured to output an output audio signal that includes a version of the first audio signal among the plurality of audio signals that has been modified for background noise and / or spatially focused, based on the output of one or more neural networks. An ear-worn device.

42. The one or more neural network layers are trained to perform background noise correction, and the output audio signal includes a version of the first audio signal with background noise corrected. The ear-worn device according to claim 41.

43. The one or more neural network layers are trained to perform spatial focusing, and the output audio signal includes a spatially focused version of the first audio signal. The ear-worn device according to claim 41.

44. The one or more neural network layers are trained to perform background noise correction and spatial focusing, and the output audio signal includes a version of the first audio signal that has undergone background noise correction and spatial focusing. The ear-worn device according to claim 41.

45. The output audio signal has a specific spatial focusing pattern, and the specific spatial focusing pattern includes different weights applied to the wearer of the ear-mounted device, where the speech is obtained from different directions of arrival. The ear-worn device according to claim 43 or 44.

46. The particular spatial focusing pattern includes the fact that the weight applied to speech sounds obtained from the direction of arrival in front of the wearer of the ear-mounted device is higher than the weight applied to speech sounds obtained from the direction of arrival to the side and rear of the wearer of the ear-mounted device. The ear-worn device according to claim 45.

47. The aforementioned neural network circuit is Receiving one or more spatial focusing control inputs that indicate the specific spatial focusing pattern, Using the one or more spatial focusing control inputs to generate the one or more neural network outputs such that the output audio signal has the specific spatial focusing pattern, It is further configured to do the following: The ear-worn device according to claim 45 or 46.

48. A communication circuit configured to receive instructions from a processing device for user selection of a specific spatial focusing pattern, The system further comprises a control circuit configured to generate one or more spatial focusing control inputs indicating the specific spatial focusing pattern, based at least in part on the user's instruction for the specific spatial focusing pattern, The ear-worn device according to claim 47.

49. The ear-worn device according to claim 48, The processing device communicates with the ear-worn device, Displaying a graphical user interface that includes options for different spatial focusing patterns, Receiving the user selection of the specific spatial focusing pattern, The processing device is configured to perform the following: A system that includes these features.

50. The aforementioned multiple options include four options, The system according to claim 49.

51. The aforementioned options are graphical representations of the different spatial focusing patterns. The system according to claim 49 or 50.

52. A sensing circuit configured to generate one or more inputs based on the movement of the ear-worn device, The system further comprises a control circuit configured to determine the degree of head movement based on one or more inputs received from one or more sensors, and to generate one or more spatial focusing control inputs indicating a specific spatial focusing pattern based on the degree of head movement. The ear-worn device according to claim 47.

53. The control circuit generates the one or more inputs representing the spatial focusing pattern based on the degree of head movement, Based on a first degree of head movement, a first set of one or more spatial focusing control inputs is generated that exhibits a first spatial focusing pattern of a first spatial focusing amount, It is configured to generate a second set of one or more spatial focusing control inputs that indicate a second spatial focusing pattern of a second spatial focusing amount based on a second degree of head movement, The first spatial focusing amount is smaller than the second spatial focusing amount, and the degree of the first head movement is greater than the degree of the second head movement. The ear-worn device according to claim 52.

54. The system further comprises a control circuit configured to determine the signal-to-noise ratio (SNR) of an acoustic environment and generate one or more spatial focusing control inputs that represent the spatial focusing pattern based on the SNR of the acoustic environment. The ear-worn device according to claim 47.

55. The output audio signal is A target speech signal including a first version of the speech signal, which is obtained by spatially focusing the speech signal, including the utterance in the first speech signal, Interfering speech signal including a second version of the aforementioned speech voice with spatial focusing applied, A background noise signal including background noise in the first audio signal, The ear-worn device according to claim 43 or 44.

56. The noise reduction circuit described above is The change in volume of the background noise signal differs from the change in volume of the target speech signal by a first volume change difference amount. The change in volume of the interfering speech signal differs from the change in volume of the target speech signal by a second volume change difference amount. The first volume change difference and the second volume change difference are different from each other, The output audio signal is generated by The ear-worn device according to claim 55.

57. The noise reduction circuit described above is The change in volume of the background noise signal differs from the change in volume of the target speech signal by a first volume change difference amount. The change in volume of the interfering speech signal differs from the change in volume of the target speech signal by a second volume change difference amount. The first volume change difference and the second volume change difference can be controlled, The system is configured to generate the aforementioned output audio signal. The ear-worn device according to claim 55 or 56.

58. The target speech signal includes the speech signal to which a specific spatial focusing pattern is applied, the specific spatial focusing pattern includes different weights applied to the speech signal obtained from different directions of arrival to the wearer of the ear-mounted device. An ear-worn device according to any one of claims 55 to 57.

59. The particular spatial focusing pattern includes the fact that the weight applied to speech sounds obtained from the direction of arrival in front of the wearer of the ear-mounted device is higher than the weight applied to speech sounds obtained from the direction of arrival to the side and rear of the wearer of the ear-mounted device. The ear-worn device according to claim 58.

60. The aforementioned neural network circuit is Receiving one or more spatial focusing control inputs that indicate the specific spatial focusing pattern, Using one or more spatial focusing control inputs to generate the two or more neural network outputs such that the target speech signal includes the speech signal to which the specific spatial focusing pattern has been applied, It is further configured to do the following: The ear-worn device according to claim 58 or 59.

61. A communication circuit configured to receive instructions from a processing device for user selection of a specific spatial focusing pattern, The system further comprises a control circuit configured to generate one or more spatial focusing control inputs indicating the specific spatial focusing pattern, based at least in part on the user's instruction for the specific spatial focusing pattern, The ear-worn device according to claim 60.

62. The ear-worn device according to claim 61, The processing device communicates with the ear-worn device, Displaying a graphical user interface that includes options for different spatial focusing patterns, Receiving the user selection of the specific spatial focusing pattern, The processing device is configured to perform the following: A system that includes these features.

63. The aforementioned multiple options include four options, The system according to claim 62.

64. The aforementioned options are graphical representations of the different spatial focusing patterns. The system according to claim 62 or 63.

65. A sensing circuit configured to generate one or more inputs based on the movement of the ear-worn device, The system further comprises a control circuit configured to determine the degree of head movement based on one or more inputs received from one or more sensors, and to generate one or more spatial focusing control inputs indicating a specific spatial focusing pattern based on the degree of head movement. The ear-worn device according to claim 60.

66. The control circuit generates the one or more inputs representing the spatial focusing pattern based on the degree of head movement, Based on a first degree of head movement, a first set of one or more spatial focusing control inputs is generated that exhibits a first spatial focusing pattern of a first spatial focusing amount, Based on a second degree of head movement, a second set of one or more spatial focusing control inputs is generated that exhibits a second spatial focusing pattern of a second spatial focusing amount, It is configured to do the following: The first spatial focusing amount is smaller than the second spatial focusing amount, and the degree of the first head movement is greater than the degree of the second head movement. The ear-worn device according to claim 65.

67. The system further comprises a control circuit configured to determine the signal-to-noise ratio (SNR) of an acoustic environment and generate one or more spatial focusing control inputs that represent the spatial focusing pattern based on the SNR of the acoustic environment. The ear-worn device according to claim 60.

68. The interference speech audio signal includes the remainder when the target speech audio signal is subtracted from the speech audio signal. An ear-worn device according to any one of claims 55 to 67.

69. The aforementioned one or more neural network outputs include two or more neural network outputs. An ear-worn device according to any one of claims 55 to 67.

70. The noise reduction circuit operates based on the outputs of the two or more neural networks, At least one of the following: a speech audio signal including the spoken speech in the first audio signal among the plurality of audio signals; and a background noise signal including the background noise in the first audio signal. At least one of a target speech signal including a first version of the speech signal with spatial focusing applied, and an interfering speech signal including a second version of the speech signal with spatial focusing applied, It is configured to obtain The ear-worn device according to claim 69.

71. The noise reduction circuit operates based on the outputs of the two or more neural networks, A target speech signal including a first version of the aforementioned speech signal that has undergone spatial focusing, Interfering speech signal including a second version of the aforementioned speech signal with spatial focusing applied, It is configured to obtain The ear-worn device according to claim 69.

72. The neural network circuit is configured to use a first subset of the one or more neural network layers to generate a first neural network output among the two or more neural network outputs, and to use a second subset of the one or more neural network layers to generate a second neural network output among the two or more neural network outputs. The noise reduction circuit is configured to acquire the speech audio signal and / or the background noise signal from the first neural network output among the two or more neural network outputs, and to acquire the target speech audio signal and / or the interfering speech audio signal from the second neural network output among the two or more neural network outputs. An ear-worn device according to any one of claims 69 to 71.

73. The two or more neural network outputs include two different masks, An ear-worn device according to any one of claims 69 to 72.

74. At least one of the two or more neural network outputs is The aforementioned speech audio signal, A mask configured to generate the aforementioned speech audio signal, The aforementioned background noise signal and, A mask configured to generate the aforementioned background noise signal, The target speech signal and, A mask configured to generate the target speech signal, An ear-worn device according to any one of claims 69 to 73, comprising the interfering speech voice signal and a mask configured to generate the interfering speech voice signal.

75. The output audio signal is generated by mixing a combination of audio signals, or A mixing circuit configured to generate the output audio signal by mixing a combination of masks, or The system further includes a wide dynamic range compression (WDRC) circuit comprising a plurality of WDRC pipelines configured to generate the output audio signal by performing wide dynamic range compression (WDRC) on the combination of audio signals. An ear-worn device according to any one of claims 55 to 74.

76. The aforementioned mixing circuit is Receiving a first volume change control input and a second volume change control input, The mixing is performed using the first volume change control input and the second volume change control input such that the first volume change difference amount is controlled at least partially by the first volume change control input and the second volume change difference amount is controlled at least partially by the second volume change control input. It is further configured to do the following: The ear-worn device according to claim 75.

77. A communication circuit configured to receive the first volume change control input and the second volume change control input from a processing device, A memory configured to store the first volume change control input and the second volume change control input, The system further comprises a control circuit configured to take the first volume change control input and the second volume change control input from the memory and output the first volume change control input and the second volume change control input to the mixing circuit. The ear-worn device according to claim 76.

78. The aforementioned neural network circuit is Receiving a first volume change control input and a second volume change control input, Using the first volume change control input and the second volume change control input, one or more neural network outputs are generated such that the first volume change difference amount is at least partially controlled by the first volume change control input, and the second volume change difference amount is at least partially controlled by the second volume change control input. It is further configured to do the following: An ear-worn device according to any one of claims 55 to 74.

79. A communication circuit configured to receive the first volume change control input and the second volume change control input from a processing device, A memory configured to store the first volume change control input and the second volume change control input, The system further comprises a control circuit configured to take the first volume change control input and the second volume change control input from the memory and output the first volume change control input and the second volume change control input to the mixing circuit. The ear-worn device according to claim 78.

80. The control circuit is further configured to generate a first volume change control input based on the level of background noise in the first audio signal and to generate a second volume change control input based on the level of interfering speech in the first audio signal. An ear-worn device according to any one of claims 76 to 79.

81. The aforementioned background noise signal has not undergone spatial focusing. The interference speech signal does not include any of the background noise in the first speech signal. An ear-worn device according to any one of claims 55 to 80.

82. The background noise signal includes a first version of the background noise in the first audio signal that has been spatially focused. The interference speech signal includes a second version of the speech signal with spatial focusing applied, and a second version of the background noise in the first speech signal with spatial focusing applied. An ear-worn device according to any one of claims 55 to 80.

83. The ear-worn device is further configured to receive a user selection to turn off spatial focusing. An ear-worn device according to any one of claims 43 to 82.

84. At least two of the aforementioned plurality of audio signals have different beamforming directivity patterns. An ear-worn device according to any one of claims 41 to 83.

85. At least two of the plurality of audio signals having different beamforming directional patterns include beamforming signals having dipole, hypercardioid, supercardioid, or cardioid directional patterns. The ear-worn device according to claim 84.

86. The system further comprises a mixing circuit configured to mix two or more audio signals such that the output audio signal includes a version of the first audio signal that has been mixed with a second audio signal and subjected to background noise correction and / or spatial focusing. An ear-worn device according to any one of claims 41 to 74 and 78 to 85.

87. The version of the first audio signal with background noise correction and spatial focusing applied includes the target utterance audio signal, The target speech signal includes a first version of the speech signal that has been spatially focused. The aforementioned speech audio signal includes the speech of the first audio signal, The second audio signal is, The background noise signal including background noise in the first audio signal, The interfering speech signal includes a second version of the aforementioned speech signal that has undergone spatial focusing, The noise reduction circuit, in the output audio signal, The change in volume of the background noise signal differs from the change in volume of the target speech signal by a difference in volume change amount. The change in volume of the interfering speech signal differs from the change in volume of the target speech signal by the amount of the volume change difference. The amount of the volume change difference can be controlled, The system is configured to generate the aforementioned output audio signal. The ear-worn device according to claim 86.

88. The aforementioned mixing circuit is Receiving volume change control input, The mixing is performed using the volume change control input such that the volume change difference is controlled at least partially by the volume change control input, It is configured to do the following: The aforementioned ear-worn device is A communication circuit configured to receive the volume change control input from a processing device, A memory configured to store the volume change control input, The system includes a control circuit configured to take the volume change control input and output the volume change control input to the mixing circuit, The ear-worn device according to claim 87.

89. The aforementioned ear-worn device includes a hearing aid. An ear-worn device according to any one of claims 41 to 88.

90. The aforementioned noise reduction circuit is mounted on the chip. An ear-worn device according to any one of claims 41 to 89.

91. Two or more microphones, A noise reduction circuit including a neural network circuit, An ear-worn device equipped with, The aforementioned neural network circuit is Receiving multiple audio signals, wherein at least two of the multiple audio signals are obtained from one of the two or more microphones, and / or at least one of the multiple audio signals is a beamforming audio signal obtained from the two or more microphones. The neural network circuit implements one or more neural network layers trained to perform background noise correction and spatial focusing based on the plurality of audio signals, or to generate outputs used when performing background noise correction and spatial focusing, so that the neural network circuit generates one or more neural network outputs based on the plurality of audio signals. It is configured to do the following: The noise reduction circuit is configured to output an output audio signal that includes a version of the first audio signal of the plurality of audio signals that has been modified for background noise correction and spatial focusing, based on the output of one or more neural networks. An ear-worn device.

92. At least two of the aforementioned plurality of audio signals have different beamforming directivity patterns. The ear-worn device according to claim 91.

93. At least one of the plurality of audio signals has a forward-directed beamforming directivity pattern, and at least one of the plurality of audio signals has a backward-directed beamforming directivity pattern. The ear-worn device according to claim 91 or 92.

94. The output audio signal has a specific spatial focusing pattern, and the specific spatial focusing pattern includes different weights applied to the first audio signal to the spoken speech obtained from different directions of arrival to the wearer of the ear-mounted device. An ear-worn device according to any one of claims 91 to 93.

95. The particular spatial focusing pattern includes the fact that the weight applied to speech sounds obtained from the direction of arrival in front of the wearer of the ear-mounted device is higher than the weight applied to speech sounds obtained from the direction of arrival to the side and rear of the wearer of the ear-mounted device. The ear-worn device according to claim 94.

96. The aforementioned neural network circuit is Receiving one or more spatial focusing control inputs that indicate the specific spatial focusing pattern, Using the one or more spatial focusing control inputs to generate the one or more neural network outputs such that the output audio signal has the specific spatial focusing pattern, It is further configured to do the following: An ear-worn device according to any one of claims 93 to 95.

97. A communication circuit that receives instructions from a processing device for the user to select a specific spatial focusing pattern, The system further comprises a control circuit configured to generate one or more spatial focusing control inputs indicating the specific spatial focusing pattern, based at least in part on the user's instruction for the specific spatial focusing pattern, An ear-worn device according to any one of claims 93 to 96.

98. The ear-worn device according to claim 97, The processing device communicates with the ear-worn device, Displaying a graphical user interface that includes options for different spatial focusing patterns, Receiving the user selection of the specific spatial focusing pattern, The processing device is configured to perform the following: A system that includes these features.

99. The aforementioned one or more neural network outputs include two or more neural network outputs, The noise reduction circuit generates an output audio signal based on the two or more neural network outputs. The output audio signal is A target speech audio signal comprising a version of the first audio signal with background noise correction and spatial focusing, wherein the target speech audio signal comprises a first version of the speech audio signal with spatial focusing, and the speech audio signal comprises the speech in the first audio signal, Interfering speech signal including a second version of the aforementioned speech signal with spatial focusing applied, A background noise signal including background noise in the first audio signal, The noise reduction circuit, in the output audio signal, The change in volume of the background noise signal differs from the change in volume of the target speech signal by a first volume change difference amount. The change in volume of the interfering speech signal differs from the change in volume of the target speech signal by a second volume change difference amount. The first volume change difference and the second volume change difference are controlled independently. The system is configured to generate the aforementioned output audio signal. An ear-worn device according to any one of claims 91 to 98.

100. The interference speech audio signal includes the remainder when the target speech audio signal is subtracted from the speech audio signal. The ear-worn device according to claim 99.

101. The neural network circuit is configured to use a first subset of the one or more neural network layers to generate a first neural network output among the two or more neural network outputs, and to use a second subset of the one or more neural network layers to generate a second neural network output among the two or more neural network outputs. The noise reduction circuit is configured to acquire the speech audio signal and / or the background noise signal from the first neural network output among the two or more neural network outputs, and to acquire the target speech audio signal and / or the interfering speech audio signal from the second neural network output among the two or more neural network outputs. The ear-worn device according to claim 99 or 100.

102. The two or more neural network outputs include two different masks, An ear-worn device according to any one of claims 99 to 101.

103. The output audio signal is generated by mixing a combination of audio signals, or A mixing circuit configured to generate the output audio signal by mixing a combination of masks, or The system further includes a wide dynamic range compression (WDRC) circuit comprising a plurality of WDRC pipelines configured to generate the output audio signal by performing wide dynamic range compression (WDRC) on the combination of audio signals. An ear-worn device according to any one of claims 99 to 102.

104. The aforementioned mixing circuit is Receiving a first volume change control input and a second volume change control input, The mixing is performed using the first volume change control input and the second volume change control input such that the first volume change difference amount is controlled at least partially by the first volume change control input and the second volume change difference amount is controlled at least partially by the second volume change control input. It is further configured to do the following: The ear-worn device according to claim 103.

105. A communication circuit configured to receive the first volume change control input and the second volume change control input from a processing device, A memory configured to store the first volume change control input and the second volume change control input, The system further comprises a control circuit configured to take the first volume change control input and the second volume change control input from the memory and output the first volume change control input and the second volume change control input to the mixing circuit. The ear-worn device according to claim 104.

106. The system further comprises a mixing circuit configured to mix two or more audio signals such that the output audio signal includes a version of the first audio signal that has been mixed with a second audio signal and subjected to background noise correction and spatial focusing. An ear-worn device according to any one of claims 91 to 102.

107. The version of the first audio signal with background noise correction and spatial focusing applied includes the target utterance audio signal, The target speech signal includes a first version of the speech signal that has been spatially focused. The aforementioned speech audio signal includes the speech of the first audio signal, The second audio signal is, The background noise signal including background noise in the first audio signal, The interfering speech signal includes a second version of the aforementioned speech signal that has undergone spatial focusing, The noise reduction circuit, in the output audio signal, The change in volume of the background noise signal differs from the change in volume of the target speech signal by a difference in volume change amount. The change in volume of the interfering speech signal differs from the change in volume of the target speech signal by the amount of the volume change difference. The amount of the volume change difference can be controlled, The system is configured to generate the aforementioned output audio signal. The ear-worn device according to claim 106.

108. The aforementioned mixing circuit is Receiving volume change control input, The mixing is performed using the volume change control input such that the volume change difference is controlled at least partially by the volume change control input, It is configured to do, The ear-worn device according to claim 107.

109. A communication circuit configured to receive the volume change control input from a processing device, A memory configured to store the volume change control input, The system includes a control circuit configured to take the volume change control input and output the volume change control input to the mixing circuit, The ear-worn device according to claim 108.

110. The ear-worn device is further configured to receive a user selection to turn off spatial focusing. An ear-worn device according to any one of claims 91 to 109.

111. Two or more microphones, A noise reduction circuit including a neural network circuit, An ear-worn device equipped with, The aforementioned neural network circuit is Receiving multiple audio signals, wherein at least two of the multiple audio signals are obtained from one of the two or more microphones, and / or at least one of the multiple audio signals is a beamforming audio signal obtained from the two or more microphones. Implementing one or more neural network layers trained to generate one or more neural network outputs based on the aforementioned multiple audio signals, It is configured to do the following: The one or more neural network outputs include an output audio signal which includes a version of the first audio signal of the plurality of audio signals that has undergone background noise correction and spatial focusing, or The one or more neural network outputs are configured to be used when generating the output audio signal, which includes a version of the first audio signal that has been modified for background noise correction and spatial focusing by the noise reduction circuit. An ear-worn device.

112. At least two of the aforementioned plurality of audio signals have different beamforming directivity patterns. The ear-worn device according to claim 111.

113. At least one of the plurality of audio signals has a forward-directed beamforming directivity pattern, and at least one of the plurality of audio signals has a backward-directed beamforming directivity pattern. The ear-worn device according to claim 111 or 112.

114. The output audio signal has a specific spatial focusing pattern, and the specific spatial focusing pattern includes different weights applied to the first audio signal to the spoken speech obtained from different directions of arrival to the wearer of the ear-mounted device. An ear-worn device according to any one of claims 111 to 113.

115. The particular spatial focusing pattern includes the fact that the weight applied to speech sounds obtained from the direction of arrival in front of the wearer of the ear-mounted device is higher than the weight applied to speech sounds obtained from the direction of arrival to the side and rear of the wearer of the ear-mounted device. The ear-worn device according to claim 114.

116. The aforementioned neural network circuit is Receiving one or more spatial focusing control inputs that indicate the specific spatial focusing pattern, Using the one or more spatial focusing control inputs to generate the one or more neural network outputs such that the output audio signal has the specific spatial focusing pattern, It is further configured to do the following: An ear-worn device according to any one of claims 113 to 115.

117. A communication circuit that receives instructions from a processing device for the user to select a specific spatial focusing pattern, The system further comprises a control circuit configured to generate one or more spatial focusing control inputs indicating the specific spatial focusing pattern, based at least in part on the user's instruction for the specific spatial focusing pattern, An ear-worn device according to any one of claims 113 to 116.

118. The ear-worn device according to claim 117, The processing device communicates with the ear-worn device, Displaying a graphical user interface that includes options for different spatial focusing patterns, Receiving the user selection of the specific spatial focusing pattern, The processing device is configured to perform the following: A system that includes these features.

119. The aforementioned one or more neural network outputs include two or more neural network outputs, The noise reduction circuit generates an output audio signal based on the two or more neural network outputs. The output audio signal is A target speech audio signal comprising a version of the first audio signal with background noise correction and spatial focusing, wherein the target speech audio signal comprises a first version of the speech audio signal with spatial focusing, and the speech audio signal comprises the speech in the first audio signal, Interfering speech signal including a second version of the aforementioned speech signal with spatial focusing applied, A background noise signal including background noise in the first audio signal, The noise reduction circuit, in the output audio signal, The change in volume of the background noise signal differs from the change in volume of the target speech signal by a first volume change difference amount. The change in volume of the interfering speech signal differs from the change in volume of the target speech signal by a second volume change difference amount. The first volume change difference and the second volume change difference are controlled independently. The system is configured to generate the aforementioned output audio signal. An ear-worn device according to any one of claims 111 to 118.

120. The interference speech audio signal includes the remainder when the target speech audio signal is subtracted from the speech audio signal. The ear-worn device according to claim 119.

121. The neural network circuit is configured to use a first subset of the one or more neural network layers to generate a first neural network output among the two or more neural network outputs, and to use a second subset of the one or more neural network layers to generate a second neural network output among the two or more neural network outputs. The noise reduction circuit is configured to acquire the speech audio signal and / or the background noise signal from the first neural network output among the two or more neural network outputs, and to acquire the target speech audio signal and / or the interfering speech audio signal from the second neural network output among the two or more neural network outputs. The ear-worn device according to claim 119 or 120.

122. The two or more neural network outputs include two different masks, An ear-worn device according to any one of claims 119 to 121.

123. The output audio signal is generated by mixing a combination of audio signals, or A mixing circuit configured to generate the output audio signal by mixing a combination of masks, or The system further includes a wide dynamic range compression (WDRC) circuit comprising a plurality of WDRC pipelines configured to generate the output audio signal by performing wide dynamic range compression (WDRC) on the combination of audio signals. An ear-worn device according to any one of claims 119 to 122.

124. The aforementioned mixing circuit is Receiving a first volume change control input and a second volume change control input, The mixing is performed using the first volume change control input and the second volume change control input such that the first volume change difference amount is controlled at least partially by the first volume change control input and the second volume change difference amount is controlled at least partially by the second volume change control input. It is further configured to do the following: The ear-worn device according to claim 123.

125. A communication circuit configured to receive the first volume change control input and the second volume change control input from a processing device, A memory configured to store the first volume change control input and the second volume change control input, The system further comprises a control circuit configured to take the first volume change control input and the second volume change control input from the memory and output the first volume change control input and the second volume change control input to the mixing circuit. The ear-worn device according to claim 124.

126. The system further comprises a mixing circuit configured to mix two or more audio signals such that the output audio signal includes a version of the first audio signal that has been mixed with a second audio signal and subjected to background noise correction and spatial focusing. An ear-worn device according to any one of claims 111 to 122.

127. The version of the first audio signal with background noise correction and spatial focusing applied includes the target utterance audio signal, The target speech signal includes a first version of the speech signal that has been spatially focused. The aforementioned speech audio signal includes the speech of the first audio signal, The second audio signal is, The background noise signal including background noise in the first audio signal, The interfering speech signal includes a second version of the aforementioned speech signal that has undergone spatial focusing, The noise reduction circuit, in the output audio signal, The change in volume of the background noise signal differs from the change in volume of the target speech signal by a difference in volume change amount. The change in volume of the interfering speech signal differs from the change in volume of the target speech signal by the amount of the volume change difference. The amount of the volume change difference can be controlled, The system is configured to generate the aforementioned output audio signal. The ear-worn device according to claim 126.

128. The aforementioned mixing circuit is Receiving volume change control input, The mixing is performed using the volume change control input such that the volume change difference is controlled at least partially by the volume change control input. It is configured to do, The ear-worn device according to claim 127.

129. A communication circuit configured to receive the volume change control input from a processing device, A memory configured to store the volume change control input, The system includes a control circuit configured to take the volume change control input and output the volume change control input to the mixing circuit, The ear-worn device according to claim 128.

130. The ear-worn device is further configured to receive a user selection to turn off spatial focusing. An ear-worn device according to any one of claims 111 to 129.