System and method for enhancing hearing aid using intelligent speaker selection

By using deep neural networks and minimum variance distortionless response filters, and taking advantage of the listener's head movement and the down-thrust angle in the audio data, the system can automatically identify and enhance the expected speaker's voice in multi-speaker environments. This solves the problem of insufficient effectiveness of traditional hearing aids under non-zero down-thrust angles and improves the listener's speech intelligibility in complex environments.

CN122205337APending Publication Date: 2026-06-12NXP BV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NXP BV
Filing Date
2025-12-09
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

Traditional hearing aids struggle to effectively separate and amplify the intended speaker's voice in multi-speaker environments, especially when there is a non-zero angle of attack between the listener's head and the speaker's position.

Method used

A deep neural network is used for intelligent speaker selection. By utilizing the listener's head movement information and the down-thrust angle in the audio data, the neural network is trained to achieve beamforming to automatically identify and enhance the expected speaker's voice. A minimum variance distortionless response filter is used for signal processing.

Benefits of technology

In complex noise and reverberation environments, it can automatically identify and enhance the expected speaker's voice, alleviate cocktail party problems, and improve listener comprehension.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122205337A_ABST
    Figure CN122205337A_ABST
Patent Text Reader

Abstract

Improved hearing assistance systems and methods are disclosed herein. In one example embodiment, a hearing assistance system includes a memory device, an audio input device configured to receive an audio input signal including audio information produced from a plurality of sound sources, an audio output device, and a processing device. During an inference mode, the processing device is configured to operate in accordance with a first neural network to generate an intermediate output signal that reflects or emphasizes, to a greater extent than in the audio information, at least one intended sound source component of the audio information produced from one intended sound source of the sound sources of the plurality, the one intended sound source determined to be the one intended sound source of the sound sources based at least indirectly on a first down-chirp angle that is apparent from the audio information. The audio output device is configured to generate an audio output signal based at least indirectly on the intermediate output signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to hearing prosthetic systems and methods such as hearing aids, and more specifically, to such systems and methods employing neural networks, machine learning, or artificial intelligence. Background Technology

[0002] People experiencing hearing loss typically rely on hearing aids or other auditory prostheses or systems (or devices) to enhance sound from the front of the listener and suppress sound from other directions. Alternatively, some traditional adaptive beamforming methods can compensate for beam direction to reduce the adverse effects of reverberation.

[0003] More specifically, in some situations, such as noisy environments or when there are multiple speakers, the spectrum of human speech often overlaps (e.g., in frequency and time). People with unimpaired hearing can typically separate discrete auditory stimuli into distinct streams and determine which stream is most relevant; this can be defined as "selective attention." When a listener's brain cannot separate stimuli as described above and cannot focus auditory attention on the expected speaker (and understand what the expected speaker is saying), it is sometimes referred to as the "cocktail party problem" (or "cocktail party effect" or "cocktail party deafness"). This impairment may require the listener to wear hearing aids or other hearing systems that improve speech intelligibility and listening comfort.

[0004] At least some traditional hearing aids employ beamforming algorithms. Beamforming algorithms extract sound from the position in front of the listener. Hearing aids using this beamforming algorithm can improve intelligibility and listening comfort. However, in situations or scenarios where multiple speakers are speaking simultaneously or substantially simultaneously, these hearing aids may still fail or be insufficient to meet the listener's needs. In other words, traditional hearing aids using beamforming algorithms are still insufficient to solve hearing difficulties in multi-speaker contexts or address the aforementioned cocktail party problem.

[0005] For at least one or more reasons, it would be advantageous if new or improved hearing aid systems (or hearing aids, hearing aid devices, or hearing aids) and hearing aid methods for providing and operating such hearing aid systems could solve one or more of the aforementioned problems, or solve one or more other problems, or provide one or more benefits. Summary of the Invention

[0006] In at least some embodiments covered herein, this disclosure relates to a hearing aid system comprising one or more memory devices configured to store a first neural network, one or more audio input devices configured to receive an audio input signal comprising audio information generated from a plurality of sound sources, one or more audio output devices, and one or more processing devices at least indirectly coupled to the one or more memory devices, the one or more audio input devices, and the one or more audio output devices. During inference mode, the one or more processing devices are configured to operate according to the first neural network to generate an intermediate output signal, the intermediate output signal reflecting or emphasizing at least one expected sound source component of the audio information generated from an expected sound source among the plurality of sound sources to a greater extent than in the audio information, the expected sound source being identified at least indirectly based on a first downshoot angle clearly visible in the audio information. The one or more audio output devices are configured to generate an audio output signal at least indirectly based on the intermediate output signal.

[0007] In at least some of these embodiments, the one or more audio input devices include one or more microphones, the one or more audio output devices include one or more speakers, the one or more processing devices include at least one microprocessor or graphics processing unit (GPU), and the first neural network is a deep neural network. Furthermore, in at least some of these embodiments, the hearing aid system is a hearing aid system. Additionally, in at least some of these embodiments, the audio input device is positioned on or associated with a human listener having a listener central axis extending from it, the plurality of sound sources includes a plurality of human sound sources, and the intended sound source among the sound sources is a first human sound source among the human sound sources. Furthermore, in at least some of these embodiments, a plurality of additional axes extend between respective human sound sources and the human listener, a plurality of angular differences exist between the listener central axis and a respective additional axis among the plurality of additional axes, and the first downward angle is a first angular difference between the listener central axis and the first additional axis among the additional axes, the first angular difference being less than each of the other angular differences at a first time.

[0008] Additionally, in at least some of these embodiments, the one or more processing devices are further configured to operate to determine a second down-throw angle different from the first down-throw angle, the second down-throw angle being a second angular difference between the listener's central axis and a second additional axis among the additional axes, the second angular difference becoming smaller than the first angular difference at a second time, and at or substantially at the second time, the switching of the expected sound source among the sound sources from the first human sound source among the human sound sources to a second human sound source among the human sound sources associated with the second additional axis among the additional axes. Furthermore, in at least some of these embodiments, the intermediate output signal is a linear filter coefficient, and the one or more processing devices are further configured to operate to multiply or convolve the linear filter coefficient with the audio input signal or at least indirectly based on an additional signal of the audio input signal to generate an additional intermediate output signal, wherein the audio output signal is at least indirectly based on the additional intermediate output signal. Additionally, in at least some of these embodiments, the intermediate output signal is filter statistics, and the one or more processing devices are further configured to operate to process the statistics and the audio input signal, or at least indirectly based on additional signals of the audio input signal, through the filter to which the statistics pertain, to generate an additional intermediate output signal, wherein the audio output signal is at least indirectly based on the additional intermediate output signal. Furthermore, in at least some of these embodiments, the filter is a beamforming filter, which is a minimum variance distortion-free response (MVDR) filter.

[0009] Additionally, in at least some exemplary embodiments, this disclosure relates to a method for training a first neural network for use in a hearing aid system. The method includes providing one or more audio input devices in an area in which multiple sound sources are located, and receiving an input signal at the one or more audio input devices, the input signal including undershoot angle data as described elsewhere herein. Furthermore, the method includes providing the input signal or an intermediate signal based on the input signal to the first neural network, and generating multiple output signals by the first neural network. Moreover, the method includes, at a loss processing block, processing the output signals and expected speaker-clean speech data determined at least in part based on the undershoot angle data to determine multiple weight signals, and updating the first neural network based on the weight signals.

[0010] In at least some of these embodiments, the receiving, providing, generating, processing, and updating are repeated until the training of the first neural network is complete, and the first neural network is a deep neural network. Furthermore, in at least some of these embodiments, the one or more audio input devices include a plurality of microphones within a room simulation, each microphone being positioned to capture a sound field at a corresponding different location, and the audio input signal received by the one or more audio input devices includes clean speech data, noise data, room characteristic data, speaker / listener characteristic data, and listener random head angle data including the down-thrust angle data.

[0011] Additionally, in at least some exemplary embodiments, this disclosure relates to a method of operating a hearing aid system during inference mode, the hearing aid system including one or more memory devices configured to store a first neural network. The method includes receiving an audio input signal at one or more audio input devices, the audio input signal including audio information generated from a plurality of sound sources. Furthermore, the method includes operating one or more processing devices according to the first neural network to generate an intermediate output signal, the intermediate output signal reflecting or emphasizing at least one expected sound source component of the audio information generated from an expected sound source among the plurality of sound sources to a greater extent than in the audio information, the expected sound source being identified at least indirectly based on a first downswing angle clearly visible from the audio information. Additionally, the method includes generating an audio output signal at one or more audio output devices based at least indirectly on the intermediate output signal.

[0012] In at least some of these embodiments, the audio input device is positioned on or associated with a human listener having a listener central axis extending therefrom, the plurality of sound sources including a plurality of human sound sources, and the intended sound source among the sound sources being a first human sound source among the human sound sources. Furthermore, in at least some of these embodiments, a plurality of additional axes extend between respective human sound sources and the human listener, with a plurality of angular differences existing between the listener central axis and a respective additional axis among the plurality of additional axes, and the first downward angle is a first angular difference between the listener central axis and the first additional axis among the additional axes, the first angular difference being less than each of the other angular differences at a first time. Additionally, in at least some of these embodiments, the operation includes determining a second down-throw angle different from the first down-throw angle, the second down-throw angle being a second angular difference between the listener's central axis and a second additional axis of the additional axes, the second angular difference becoming smaller than the first angular difference at a second time, and at or substantially at the second time, the switching of the intended sound source among the sound sources from the first human sound source among the human sound sources to a second human sound source among the human sound sources associated with the second additional axis of the additional axes.

[0013] Additionally, in at least some of these embodiments, the intermediate output signal is linear filter coefficients, and the method further includes multiplying or convolving the linear filter coefficients with the audio input signal or at least indirectly based on an additional signal of the audio input signal to generate another intermediate output signal, wherein the audio output signal is at least indirectly based on the other intermediate output signal. Furthermore, in at least some of these embodiments, the intermediate output signal is filter statistics, and the method further includes processing the statistics and the audio input signal or at least indirectly based on an additional signal of the audio input signal through the filter to which the statistics pertain to generate another intermediate output signal, wherein the audio output signal is at least indirectly based on the other intermediate output signal. Furthermore, in at least some of these embodiments, the filter is a beamforming filter, and the beamforming filter is a minimum variance distortion-free response (MVDR) filter. Furthermore, in at least some of these embodiments, the first neural network is a deep neural network trained before operating in the inference mode to be able to identify the downswing angle and the corresponding expected sound source based on the received audio data.

[0014] This disclosure covers numerous embodiments, which may be advantageous in one or more aspects. In at least some embodiments of the improved hearing aid systems and methods covered herein, the improved hearing aid systems and methods (1) employ an intelligent speaker selection mechanism for training and (2) take into account down-throw angles. Using this approach, optimal spatial beamforming can be achieved in the presence of noise and reverberation, but with a down-throw angle between the listener's central axis and the speaker, without requiring any prior knowledge about the number of speakers, the individual positions of the speakers, or self-supervised (also known as unsupervised) or reinforcement learning. While the primary applications of the embodiments described herein are hearing aid-related products (e.g., chips for such devices), this disclosure also covers many other applications. For example, some such assistive applications may include other wearable devices, such as earbuds or headphones, and other applications, including teleconferencing applications, public address systems, and others. Attached Figure Description

[0015] Figure 1 This is a schematic diagram 100 showing the arrangement of a listener (the listener's head) relative to multiple speakers (the speakers' heads) in an area surrounded by walls, provided to illustrate the concept of undershot angle;

[0016] Figure 2 This is a schematic diagram illustrating an improved hearing aid system according to exemplary embodiments covered herein;

[0017] Figure 3 The provided diagram illustrates that it can be used Figure 2 The operation of deep neural networks in training mode in the improved hearing aid system;

[0018] Figure 4 and Figure 5 These are the first and second timing diagrams, respectively illustrating the process for deep neural networks (e.g., reference). Figure 3 The first and second example training routines for the described deep neural network; and

[0019] Figure 6 , Figure 7 and Figure 8 First, second, and third additional schematic diagrams are provided, illustrating first, second, and third example embodiments of the improved hearing aid system, which employ corresponding trained deep neural networks and are configured to operate in inference mode. Detailed Implementation

[0020] The inventors of this invention have recognized the aforementioned problems associated with conventional hearing aid systems and methods designed to address hearing difficulties in multi-speaker contexts. Furthermore, the inventors have particularly recognized that while conventional hearing aid systems and methods employing beamforming algorithms can improve intelligibility, such systems and methods may be insufficient, especially when the listener's head is not directly facing the intended speaker (or target speaker), resulting in a non-zero angular difference or "downward angle" between the direction the listener's head faces and the intended speaker's position. In such cases, the effectiveness of the conventional hearing aid system or method used by the listener (e.g., its real-world effectiveness in allowing the listener to hear and understand the intended speaker) may be limited when the listener's head is misaligned with the intended speaker, resulting in a non-zero downward angle.

[0021] In view of the foregoing considerations, the inventors of the present invention also recognize that a new or improved hearing aid system or method will be achieved if it can take into account any misalignment or non-zero undershoot angle between the direction the listener's head is facing and the position of the intended speaker. In this regard, the inventors of the present invention also recognize that head movement information is embedded in the audio data captured by the microphone of a hearing aid (e.g., a hearing aid), and that such head movement information extracted from the embedded phase in the microphone of the hearing aid can be used to generate a strategy for selecting the intended speaker (without utilizing any additional measurements). Furthermore, the inventors of this invention have recognized that by using a pre-learned neural network (e.g., Wave-U-Net or any other suitable network, real-valued or complex-valued), speech intelligibility in real-world environments can be improved without the use of head trackers or eye trackers in dealing with noise (e.g., unexpected human speakers, reverberation, reflections, and echoes), or alternatively by using a neural network-assisted beamforming filter (e.g., a minimum variance distortionless response (MVDR) filter, where the statistics of the signal are calculated by the neural network), such a hearing aid system or method can mitigate the cocktail party problem.

[0022] Therefore, this disclosure contemplates various embodiments employing, for example, end-to-end neural network solutions or partially using different types of beamforming. Furthermore, the inventors of this invention recognize that such novel or improved hearing aid systems or methods with misalignment or non-zero undershoot angles can be implemented using deep learning-based beamforming via an intelligent speaker selection training mechanism, which performs intelligent expected speaker selection (or "deep learning-based intelligent speaker selection beamforming") even when the listener is not facing the expected speaker (or target speaker). This intelligent expected speaker selection operation enables such hearing aid systems or methods to automatically determine which of the speakers is the expected speaker in the presence of multiple speakers. While this disclosure covers novel or improved hearing aid systems or methods particularly suitable for implementation in or as part of hearing aids, it also covers novel or improved hearing aid systems or methods suitable for various other applications and settings such as teleconferencing, public address systems, or enclosed spaces such as automobiles.

[0023] Therefore, in at least some embodiments, this disclosure relates to new or improved hearing aid systems or methods for eliminating or mitigating the cocktail party problem by implementing a deep / machine learning-based intelligent speaker selection mechanism (or a mechanism employing machine learning or artificial intelligence). At least some embodiments covered herein employ a deep learning-based intelligent speaker selection mechanism that uses a neural network model trained to learn how to determine which speaker (when several speakers are present) should be considered the intended speaker (or target speaker) at any given time. According to embodiments, any of various types of neural networks or related techniques may be employed, including, for example, artificial neural networks (ANNs), machine learning models, convolutional neural networks (CNNs), reinforcement learning models, and deep neural networks (DNNs).

[0024] Furthermore, at least some of the embodiments covered herein employ a method involving a training mechanism in which the expected speaker is modified based on the listener's head movements. This training scheme can be applied to end-to-end neural network systems, or alternatively to systems, for example, where a neural network estimates the coefficients of a linear filter or has known statistics about a beamforming filter. After training in this manner, when operating in inference mode, the neural network can determine (or assist in determining) or modify the speaker that the hearing aid system (or device) or method should focus on based on the listener's head movements. That is, the method teaches the neural network to follow spatial information embedded in the multi-input audio signals during inference in order to make intelligent selections of the expected speaker, ranging from simple cases with only two speakers to “cocktail party” situations with many (e.g., more than two) speakers. In terms of beamforming, this means that optimal beamforming can be obtained without any prior information about room size, number of speakers, noise statistics, etc.

[0025] As described above, embodiments of this disclosure specifically consider the downward angle of attack. In this respect, Figure 1 A schematic diagram 100 is provided to more clearly illustrate the concept of the down-thrust angle in the presence of two speakers. More specifically, Figure 1 A schematic representation of the listener's head 102 (i.e., the head of the listener (L) (e.g., the person listening to the sound)) relative to the heads 104 of multiple speakers (in this example, including the head 106 of the first speaker and the head 108 of the second speaker, i.e., the head of the first speaker (S1) and the head of the second speaker (S2)). As shown, at any given time, the listener's head 102 has an associated listener central axis 110, which can be defined as an axis extending directly forward from the center point 116 of the listener's head 102 and perpendicular to the ear-to-ear axis 112 extending between the ears 114 on both sides of the listener's head. The listener central axis 110 can be said to form a head angle relative to a reference axis 118. And it extends through the listener’s head 102 through the center point 116, with each of the listener’s central axis 110 and the ear-to-ear axis 112 passing through the center point 116.

[0026] Further, as shown in the figure, assuming that the corresponding sounds (e.g., vocalized sounds) are emitted from each of the heads 106 of the first speaker and 108 of the second speaker toward the head 102 of the listener, these corresponding sounds typically travel toward the head 102 of the listener along the first axis 120 and the second axis 122 (which are axes extending directly from the respective front parts of the heads of those corresponding speakers), respectively. It can be said that each of the first axis 120 and the second axis 122 has a corresponding angle associated with it relative to the reference axis 118, i.e., and These angles can be considered as the corresponding angles of arrival of sound from the head 106 of the first speaker and the head 108 of the second speaker at the head 102 of the listener. As can be seen in the illustration, the first axis 120 and the second axis 122 extend between the tip 124 of the listener's head 102 and each of the heads 106 and 108 of the first speaker, respectively. However, the first axis 120 and the second axis 122 can also be understood as extending between the center point 116 of the listener's head 102 and each of the heads 106 and 108 of the first speaker, respectively (or their respective center points). It should be understood that although sound transmitted from the head 106 of the first speaker and the head 108 generally travels along the first axis 120 and the second axis 122, sound may reach the listener's head 102 in other ways, for example, due to sound reverberation caused by the surrounding walls 126 as indicated by arrow 128. Additionally, the head 106 of the first speaker and / or the head 108 of the second speaker may optionally not face directly towards the head 102 of the listener.

[0027] exist Figure 1 In the illustrated embodiment, it can be understood that the first axis 120 passing through the head 106 of the first speaker is angularly closer to the listener's central axis 110 than the second axis 122 passing through the head 108 of the second speaker. Accordingly, the first speaker (S1) rather than the second speaker (S2) should be considered the intended speaker. In this case, the downward angle... This is defined here as the angle between the listener's central axis 110 and an axis extending between the listener's head 102 and the expected speaker's head (in this example, the axis is a first axis 120 extending between the listener's head 102 and the first speaker's head 106). That is, the down-thrust angle. Can be defined as listener ( The listener's central axis 110 is aligned with the listener's head 102 and the expected speaker (in this example, the first speaker). The angle between the axes extending from the sound sources constitutes the angle of arrival of the speaker closest to the listener's central axis. Therefore, in this example, reference 118 is used as a reference axis relative to which the angular positions of other axes can be measured, and considering that the first axis 120 (representing the angle of arrival of the sound emitted by the first speaker) is angularly closer to the listener's central axis 110 than the second axis 122 (representing the angle of arrival of the sound emitted by the second speaker), the downshoot angle can be defined as the angle of the first axis 120 relative to reference axis 118 and the head angle. The difference between (the angle between the listener's central axis 110 and the reference axis) is shown in equation (1), that is:

[0028] (1).

[0029] Go to Figure 2 This disclosure relates to improved hearing aid systems and methods utilizing the aforementioned down-thrust angle concept, including, for example, an improved hearing aid system 200. Figure 2 Specifically, a block diagram of an example hearing aid system is provided in schematic form. In this respect, the hearing aid system 200 can be considered as a hearing aid that can be worn at least partially by a person (e.g., a listener) seeking to hear or listen to sound / audio information (including speech / vocalization sounds from one or more speakers located in the surrounding environment) within that person's environment. As shown, the hearing aid system 200 specifically includes a pair of combined input / output devices 202 (e.g., one for each ear of the listener's head, such as...). Figure 1 Each ear 114 of the head 102 of the hearing person may implement one or more corresponding audio input devices 204, such as microphones, and one or more corresponding audio output devices 206, such as speakers, in each combination input / output device 202. Although the combination input / output device 202 may take the form of, for example, earplugs (or headphones), as described elsewhere herein, this disclosure is intended to cover a variety of hearing aids or other types of hearing aid systems other than those employing earplugs.

[0030] In addition to the combined input / output device 202, the hearing aid system 200 further includes a computer system 210, which is at least indirectly coupled to the combined input / output device 202, as shown by dashed line 208. The computer system 210 includes one or more processing devices 212 and one or more memory devices 214. The one or more processing devices 212 may include any one or more of, for example, a microprocessor, a controller, a graphics processing unit (GPU), a programmable logic device (PLD), an application-specific integrated circuit (ASIC), and / or other processing devices. The processing devices 212 can be operated according to various computer-executable instructions to perform any of a variety of different functions related to performing processing and taking other actions as described herein. Furthermore, the one or more memory devices 214 may include any one or more of, for example, random access memory (RAM) devices, read-only memory (ROM) devices (and various forms thereof, including electrically erasable programmable read-only memory (EEPROM) devices), and / or other memory devices. The memory devices 214 may store software, applications, or computer instructions that the one or more processing devices 212 operate according to. For example, in some embodiments, computer system 210 may employ a device that has both processing power and memory power (e.g., a processor in memory or a PIM).

[0031] Despite Figure 2 The computer system 210 is symbolically illustrated, but it should be understood that computer system 210 is intended to represent any of various embodiments of a computer system that may employ any of the various types of processing devices 212 or memory devices 214, including embodiments having multiple processing devices respectively distributed or located in different locations, and / or embodiments having multiple memory devices respectively distributed or located in different locations. Although computer system 210 may, for example, represent a mobile device such as a cellular phone, smartphone, or laptop or notebook computer, or a desktop computer, computer system 210 is also intended to represent various distributed computer devices or combined systems, such as mobile devices communicating with a cloud computing system, which in turn includes a number of processing devices and memory devices respectively located in various different corresponding locations. Communication between such multiple processing devices and memory devices, and between computer system 210 and combined input / output device 202 (as shown by dashed line 208), can occur in any of a variety of ways, such as via wired or wireless links or via the Internet. Furthermore, this disclosure covers embodiments in which one or more processing devices and / or memory devices are located within the combined input / output device 202, except for one or more computer systems that are different from those combined input / output devices, such as computer system 210.

[0032] As will be described in further detail below, according to the embodiments covered herein, one or more memory devices 214 may in particular store one or more neural networks 216, and one or more processing devices 212 may in particular execute instructions associated with such one or more neural networks. Such instructions may in particular enable the training of such one or more neural networks (e.g., during training mode), and also enable the trained one or more neural networks to perform inference operations (e.g., during inference mode).

[0033] Intelligent Speaker Selection Training System

[0034] Next reference Figure 3 In at least some embodiments, this disclosure relates to employing neural networks (e.g., by...). Figure 2 The hearing aid system 200 shown in the diagram (representing a neural network 216) is a new or improved hearing aid system or method, wherein the neural network has been trained in the manner shown in schematic diagram 300. Typically, this training approach envisions a training system with multiple microphones, which can be used to train the neural network such that the neural network optimizes the directionality of the microphone array for listening via implicit beamforming. This is achieved during training by synchronizing the listener's head movement in a multi-speaker scenario with a down-throw angle relative to the expected speaker's direction and with the expected clean speech used in the loss function. The neural network learns from spatial information embedded in the multi-array audio to optimally beamform toward the expected speaker. Most training samples should contain a non-zero down-throw angle between the listener's head and the speaker's head, but the training data should also include cases where the listener faces the expected speaker (where the down-throw angle is zero rather than non-zero), or cases with only one speaker.

[0035] More specifically, such as Figure 3 As shown, during the training of the neural network, there are n microphones 302 within a real-world setting 304 (or alternatively, within a room simulation), where The n microphones 302 can generate corresponding output signals. The corresponding output signal The sound field is captured at corresponding unique (different) locations within the real-world setting. n microphones 302 within the real-world setting 304 receive sound from their environment, as defined by several input parameters 308 (abstracted here for easier understanding and simpler relation to the simulated environment), including clean speech data 310, noise data 312, room characteristic data 314, speaker / listener characteristic data 316, and listener random head angle data 318. The listener random head angle data 318 can be the down-thrust angle as described elsewhere in this document (…). ) data. Although Figure 3 The n microphones 302 shown happen to include two microphones, as indicated by ellipsis 340, but depending on the embodiment or setup, any number of microphones (typically two or more microphones) may be present.

[0036] Additionally, in the real-world setup 304, the first microphone 320 among the n microphones 302 (e.g., provides the output signal) The microphone is close to one ear (or in one ear) (e.g., from...). Figure 1 The listener's head 102 and one ear 114). Another (or more) microphones in the n microphones 302 of the real-world setup 304, for example... Figure 3 The first of the n microphones 302 shown, and another microphone 322, can have a location elsewhere. For example, one or more such microphones among the n microphones could even be located at a remote station, such as a smartphone (or other telephone or mobile device), or at a dedicated device containing one or more microphones, or at the location of the (human) speaker, or at the location where the output signal is generated in microphone 302. or The microphones are located near (for example, at the other ear of the listener's head 102). The microphone output signal 306 ( ) is provided to the deep neural network 324 that is being trained, and in this sense, the output signal 306 ( This can also be considered as an input signal. A deep neural network 324, for example, can correspond to what is shown as stored in... Figure 2 The neural network 216 in the memory device 214, and the training of the neural network 216 can be achieved through... Figure 2 The operation of the processing device 212 is performed. The microphone 302 can be coupled to the processing device (e.g., processing device 212) in a wired or wireless manner and work together as a beamformer. Preferably, the latency of any wireless connection is lower than the latency caused by signal processing by the processing device.

[0037] In response to receiving output signal 306 ( … The output of a deep neural network 324 (which is currently being trained) 326 output signals ( The output signal 326 can be filter coefficients, a tensor with statistics, a multi-channel representation of clean speech, or an output signal corresponding to an output speaker from, for example, a wearable device. 326 output signals ( This is used to receive the loss processing block 328. Also during training, the additional processing block 330 determines and outputs the expected speaker clean speech data, as indicated by arrow 332. At the additional processing block 330 (but can also be embedded in setup / simulation 304), the listener's random head angle data 318 (which can also be the down-thrust angle) is used. The expected speaker clean speech data 332 is determined based on a combination of parts of the clean speech data 310 and (randomly defined) positional data, as indicated by arrow 334, and additionally based on a combination of parts of the clean speech data 310 and (randomly defined) positional data, as indicated by arrow 336. (The additional processing block 330 can also be viewed as representing an operation that allows finding or identifying the speaker closest to the listener's central axis during training, since such information is available.) As shown by arrow 332, during the training of the deep neural network 324, the expected speaker clean speech data (which can also be referred to as the expected speaker l) output by the additional processing block 330 is used to determine the expected speaker clean speech data 332. The pure voice) is also with 326 output signals ( Together, these are provided to the loss processing block 328. In response to receiving the expected speaker clean speech data and m output signals 326, the loss processing block 328 generates a weight update signal, indicated by arrow 338, which is provided back to the deep neural network 324 for further training of the deep neural network.

[0038] It should be understood that during the training phase, there are speech segments coming from various directions to the listener. There can be one speaker at a time, or multiple speakers simultaneously. Speakers can face the listener. (For example, such as) Figure 1 (As shown), it is also possible to look at the listener at a certain down-angle without directly facing the subject. Speech segments can be utterances with and without masks of different voices, including various male and female speakers. If more than one speaker is present at the same time, one of these speakers will be designated as the expected speaker at this time—more specifically, the speaker with the smallest (or least) angular difference relative to the listener's corresponding position relative to the listener's central axis will be selected as the expected speaker (target) at this time. However, if the listener moves, causing the down-angle relationship to change, the expected speaker can also change. Or, if one or more speakers change in terms of their position relative to the listener or in terms of who is speaking at any given time (again causing the down-angle relationship to change), the expected speaker can also change again. In addition, the acoustic signal during training may be contaminated by noise, reverberation / reflection, etc. It should also be understood that during training, the corresponding clean speech (expected speaker l ( The clean speech of the machine is used for loss calculation during training (via...). To neural networks). Ultimately, based on training, neural networks (e.g., Figure 3 The deep neural network 324 may be stored in the memory device 214 and generated by... Figure 1 The neural network 216 executed by the processing device 212 is optimized so that the output signal resembles the expected clean speech signal, but it can also output the coefficients or statistics of a given filter. Furthermore, the neural network 324 can be used to estimate only one or more angles of the current simulation / real-world setting (e.g., Figure 1 Angle This refers to information that can be used to train another neural network or processed via a beamforming filter.

[0039] It should be understood that during training, an "artificial head" can be used as both the listener's head and each speaker's head, for example, by positioning the artificial speaker at the location of the speaker's head and the microphone at the location of the listener's head. Different heads within the speaker's head can emit different sounds at various times (including at times when the listener's head may be in different positions or orientations). For example, refer to... Figure 1 Furthermore, at the first moment, the first artificial head can be located as the listener's head 102 belonging to listener L, and the second and third artificial heads can be located as belonging to the speaker, respectively. and The head of the first speaker 106 and the head of the second speaker 108. Additionally, it can be via one of the speaker heads designated as the intended speaker (e.g., for example, for...). Figure 1 The speaker shown The artificial mouth (106) of the first speaker's head is used to present pre-recorded pure speech, and can simultaneously transmit pure speech... Feeds (e.g., as indicated by arrow 332) are sent to loss processing block 328 (and thus to the neural network) for loss calculation. Other speaker heads designated as not the expected speaker (e.g., for example, ...) Figure 1 The speaker shown The second speaker's head (108) can still present other speech or noise. Due to room reverberation (e.g., as... Figure 1 (As indicated by arrow 128 in the image) and the heads of other speakers, the expected speech signal received by the microphone will be contaminated by reverberation and noise.

[0040] Alternatively, for example, in different parts of the training and at a second time, an artificial mouth can be used instead of a different one in the speaker's head designated as the intended speaker at the second time (e.g., for example, for...). Figure 1 The speaker shown The second speaker's head 108, so that the expected speech is now... (Issued) to present pre-recorded clean speech. The expected offset of the speaker to a different one in the speaker's head can be caused by changes in head positioning and the resulting different head angles of the listener's head (e.g., The change triggers the process. Through this change, the corresponding clean speech signal (e.g., as indicated by arrow 332) is fed into the loss processing block 328 (and thus into the neural network) for loss calculation. During training, for various different speakers' heads (not just...) Figure 1 The process can be repeated further, taking into account the heads of the two speakers shown, the various positions or orientations of these speakers' heads, and the various positions or angles of the listener's head. Based on the information generated from this training work, the neural network (e.g., neural network 324) now learns to select the expected speaker based on the head angle information embedded in the phase, and learns to process the microphone signal such that the output approximates the expected clean speech signal.

[0041] Now refer to each Figure 4 and Figure 5 The first timing diagram 400 and the second timing diagram 500 respectively show the first example training routine and the second example training routine for the deep neural network 324. Figure 4 The first example training routine shown in the first timing diagram 400 is a training routine that changes the expected speaker based on the downshoot angle (or head angle). In contrast, Figure 5 The second example training routine shown in the second timing diagram 500 is based on the downshoot angle (or changes in head angle) This is used to modify the training routine for the intended speaker. Both the first timing diagram 400 and the second timing diagram 500 show the changes related to... Figure 1 The example scenarios shown are consistent with the example operation for an arrangement / situation in which a listener (L) and first and second speakers (S1 and S2) are present. However, it should be understood that the first timing diagram 400 and the second timing diagram 500 are merely examples, and this disclosure envisions many other scenarios in which more than two speakers are present, for which different timing diagrams will be applied. Furthermore, it is worth noting that during training, the down-throw angle is not explicitly provided at the input of the neural network model, as this information is embedded in the phase relationship between the multiple array inputs.

[0042] For more specific reference Figure 4The first timing diagram 400 includes a first curve 402, a second curve 404, and a third curve 406, which respectively illustrate the example angular position changes of the listener's central axis 110, the first axis 120, and the second axis 122 over time (t). In the first timing diagram 400, for each of the first curve 402, the second curve 404, and the third curve 406, a change in time (t) occurs along the x-axis, and a change in angular position occurs along the y-axis. More specifically, the first curve 402 illustrates the angular position change of the listener's central axis 110, i.e., the angle... (Again, as) Figure 1 As shown, the axis extending forward from the listener's head changes with time (t). Furthermore, the second curve 404 and the third curve 406 respectively show the angular position of each of the first axis 120 and the second axis 122 corresponding to the angular position orientation of the first head 106 and the second head 108 of the first speaker (S1) and the second speaker (S2), respectively. and The positions of these corners are constant.

[0043] from Figure 1 As can be seen from this, it should be understood that the downward angle (angle) The downthrow angle can be considered as the difference between the first curve 402 and either the second curve 404 or the third curve 406 at any given time. Whether the downthrow angle constitutes the difference between the first curve 402 and the second curve or the difference between the first curve and the third curve 406 depends on whether the angular difference between the first curve 402 and the second curve is greater than or less than the difference between the first curve and the third curve. This is because, as defined herein, the downthrow angle is understood as the smallest angular difference (or, in this case, a smaller angular difference) among the angular differences between the listener's central axis and the respective axes extending between the listener's head and the various corresponding speakers. Considering this, it can be seen that, for example, during the first time period 408, the first curve 402, representing the angular position of the listener's central axis 110, is closer to the second curve 404 than to the third curve 406. Therefore, for example, at the first time period 410, the downthrow angle... The first value 412 corresponds to the difference between the first curve 402 and the second curve 404 at that first time. However, also for example during the second time period 414, the first curve 402, representing the angular position of the listener's central axis 110, is closer to the third curve 406 than to the second curve 404. Therefore, for example, at the second time 416, the downthrow angle... The second value 418 corresponds to the difference between the first curve 402 and the third curve 406 at that second time.

[0044] also, Figure 4The fourth curve 420 also illustrates how the intended speaker changes between the first speaker (S1) and the second speaker (S2) when the relative position of the listener center axis 110, represented by the first curve 402, changes relative to each of the first axis 120 and the second axis 122, represented by the second curve 404 and the third curve 406, respectively. This is based on whether the downthrow angle is determined to be between the first and second curves or between the first and third curves. More specifically, as shown, during a time period, for example, the first time period 408, when the first curve 402 is closer to the second curve 404 than to the third curve 406, the downthrow angle... Between these two curves, the expected speaker is the first speaker (S1) corresponding to the first axis 120 (and Figure 1 The first speaker's head 106), as shown in the first segment 422 of the fourth curve 420. Alternatively, as shown, during a time period such as the second time period 414, when the first curve 402 is closer to the third curve 406 than the second curve 404, the downward angle... Between these two curves, the expected speaker is the second speaker (S2) corresponding to the second axis 122 (and Figure 1 The second speaker's head (108), as shown in the second segment 424 of the fourth curve 420.

[0045] The change in the first curve 402 relative to the second curve 404 and the third curve 406 can trigger a downward angle. Variations, particularly in determining the angle of attack as measured between the listener's central axis 110 and the first axis 120, or between the listener's central axis 110 and the second axis 122 (or between the listener's central axis 110 and any other axis associated with any other speaker). The manner in which this involves determining the intended speaker can vary depending on the embodiments. For example, in at least some embodiments, when multiple speakers with similar positions relative to the listener's central axis are present, to avoid rapid, repetitive switching between or among different speakers, when the angular difference between the listener's central axis 110 and the axis of another speaker (e.g., the second axis 122) becomes smaller than the angular difference between the listener's central axis 110 and the axis of one speaker (e.g., the first axis 120), switching from one speaker (e.g., from the first speaker S1) to the other speaker (e.g., switching to the second speaker S2) as the intended speaker does not necessarily occur immediately. Instead, as... Figure 4As indicated by threshold 426, in at least some embodiments or implementations, the expected speaker is switched from the current expected speaker to the new expected speaker only after the angle difference between the listener's central axis and the axis associated with the new expected speaker decreases to a threshold amount 426 lower than the angle difference between the listener's central axis and the axis associated with the current expected speaker.

[0046] Further reference Figure 5 The second figure 500 also includes each of the first curve 402, the second curve 404, the third curve 406, and the fourth curve 420. Furthermore, in the second figure 500 shown, the angular position and downforce angle of the listener's central axis 110 relative to the first axis 120 and the second axis 122 change with time (t). The corresponding changes (including the values ​​of the downshoot angle at the first time 410 and the second time 416) and Figure 4 The same as shown. Furthermore, the expected speaker is changed between the first speaker (S1) and the second speaker (S2). Figure 4 The same applies as shown, where, for example, during the first time period 408, the expected speaker is the first speaker (S1), and during the second time period 414, the expected speaker is the second speaker (S2).

[0047] although Figure 5 and Figure 4 The aforementioned similarities exist, but Figure 5 and Figure 4 The difference is that, Figure 5 This illustrates how changes in the first curve 402 relative to the second curve 404 and the third curve 406 can trigger the determination of the downforce angle. The changes in the aspect and the alternative operating methods in determining the changes in the expected speaker that follow. More specifically, in this embodiment, it can be seen that the downward angle is tracked. (or head angle) at a specific time interval (δ or Changes on time interval 502 ( ). Figure 5 Specifically, three of the time intervals 502 are shown, namely the first time interval 504, the second time interval 506, and the third time interval 508.

[0048] As shown in the figure, each of the time intervals 502 is at the down-angle. It begins to increase from the minimum value. For example, the second time interval 506 in time interval 502 begins at the start time 510, at which time the downward angle... Starting from value 412, the value increases, which existed at the first time 410 and is most recently the downshoot angle. The minimum value. Starting at start time 510, the second time interval 506 in time interval 502 continues until finish time 512, with midpoint time 514 occurring midway between start time 510 and finish time 512. Furthermore, at midpoint time 514, the downward angle... The difference between the first curve 402 (listener center axis 110) and the second curve 404 (first axis 120) is changed to the difference between the first curve 402 and the third curve 406 (second axis 122). Similarly, the midpoint time 514 is the time expected for the speaker to change from being the first speaker (S1) as in the case of the first segment 422 to being the second speaker (S2) as in the case of the second segment 424.

[0049] Accordingly, regarding each of the first time interval 504 and the third time interval 508 in time interval 502, it can be seen that each of these time intervals includes the current downward angle. The corresponding start time, completion time, and midpoint time are determined from the most recent local minimum level. Again, regarding each of the first time interval 504 and the third time interval 508 in time interval 502, precisely at the corresponding midpoint time within each time interval, the downshoot angle changes from being determined as the difference between the first curve 402 and the third curve 406 to being determined as the difference between the first curve and the second curve 404, and accordingly, the expected speaker changes from being the second speaker (S2) to being the first speaker (S1). Despite the foregoing description, this disclosure also contemplates additional methods for determining the downshoot angle and the expected speaker, including different methods applicable to different contexts and / or different numbers of speakers.

[0050] When training the neural network (e.g., deep neural network 324) as described above, the improved hearing aid system (e.g., [missing information]) can be operated in inference mode. Figure 2 The improved hearing aid system 200. During reasoning operation mode, the improved hearing aid system 200 may employ a neural network 216 (which may also be a trained deep neural network 324) to generate sound output that specifically reflects the sound (e.g., a word or vocal expression) produced by the intended speaker at any given time. Based on the training of the neural network 324, it is determined whether the sound produced by any given speaker among two or more speakers near the improved hearing aid system 200 (and any listener wearing the improved hearing aid system) constitutes the sound of the intended speaker.

[0051] In at least some of the embodiments covered herein, during inference operation mode, one or more processing devices (e.g., processing device 212) of the improved hearing aid system (e.g., improved hearing aid system 200) operate according to a trained neural network (e.g., neural network 216) to generate an output signal (or an intermediate signal, upon which the output signal may be further generated) that reflects or emphasizes at least one expected sound source component of audio information generated from an expected sound source (e.g., an expected human speaker among multiple human speakers) of a plurality of sound sources to a greater extent than the overall audio information that may be received via an audio input device (e.g., audio input device 204). The trained neural network determines an expected sound source (and thus the expected sound source component) from the sound sources based at least indirectly on a first downthrow angle that is clearly visible from the audio information.

[0052] This disclosure envisions many different embodiments of improved hearing aid systems employing many specific forms of neural networks, which typically operate in reasoning mode as described above. Figure 6 , Figure 7 and Figure 8 First, second, and third additional schematic diagrams are provided, illustrating first, second, and third example embodiments of the improved hearing aid system, respectively shown as improved hearing aid systems 600, 700, and 800. Each of the first improved hearing aid system 600, the second improved hearing aid system 700, and the third improved hearing aid system 800 corresponds to... Figure 2 The improved hearing aid system 200 can be considered as constituting Figure 2 Different embodiments (or versions or implementations) of the improved hearing aid system 200. More specifically, as shown in the figures, the first improved hearing aid system 600, the second improved hearing aid system 700, and the third improved hearing aid system 800 respectively include a first trained deep neural network 602, a second trained deep neural network 702, and a third trained deep neural network 802, each of the first trained deep neural network 602, the second trained deep neural network 702, and the third trained deep neural network 802 corresponding to Figure 2 The neural network 216 can be considered as constituting Figure 2 Different embodiments (or versions or implementations) of the neural network 216.

[0053] Each of the first trained deep neural network 602, the second trained deep neural network 702, and the third trained deep neural network 802 respectively... Figure 6 , Figure 7 and Figure 8Symbolically shown in the middle is the training block represented by dashed line 650, which has been trained. As symbolically shown, the training of each of the first trained deep neural network 602, the second trained deep neural network 702, and the third trained deep neural network 802 specifically involves training that enables the corresponding deep neural network to determine the expected speaker selection (l) as indicated by arrow 654, as indicated by expected speaker selection block 652. As further indicated by arrow 656, this determination of expected speaker selection block 652 is based on changes in head angle or head angle information (l). As mentioned above, head angle information constitutes angle information, and the corresponding deep neural network can determine the down-thrust angle and the corresponding expected speaker based on the angle information.

[0054] Furthermore, each of the first improved hearing aid system 600, the second improved hearing aid system 700, and the third improved hearing aid system 800 is specifically shown as receiving corresponding input signals 604, 704, and 804, which can be considered as originating from a corresponding microphone / sound sensor (e.g., Figure 2 The combined input / output device 202 receives speech or other sound information signals from the corresponding audio input device in the audio input device 204. The corresponding input signals 604, 704, and 804 can be considered similar to the output signal 306 described above regarding the training of the deep neural network 324. … This applies as long as the corresponding input signals 604, 704, and 804 are input into the corresponding deep neural networks 602, 702, and 802. This is consistent with the deep neural network 324 shown during training in training mode. Figure 3 on the contrary, Figure 6 , Figure 7 and Figure 8 This is intended to represent the first improved hearing aid system 600, the second improved hearing aid system 700, and the third improved hearing aid system 800 during the inference operation mode rather than the training operation mode. However, if the simplified training (dashed line) block 650 from hearing systems 600, 700, and 800 is removed, then... Figure 6 , Figure 7 and Figure 8 The deep neural network block 324 can be represented in more detail because input 306 and output 326 can be directly mapped to input 604 and output 606, input 704 and output 706, and input 804 and output 806, respectively.

[0055] In addition, such as Figure 6 , Figure 7 and Figure 8As shown, the first, second, and third improved hearing aid systems generate corresponding output signals 606, 706, and 806 based on the received corresponding input signals 604, 704, and 804. The corresponding output signals 606, 706, and 806 can respectively constitute the corresponding speaker / sound output devices of the corresponding first improved hearing aid system 600, second improved hearing aid system 700, and third improved hearing aid system 800 (e.g., Figure 2 The audio information or signal output by the corresponding output speaker 206 of the combined input / output device 202. As will be described in further detail below, although each of the first improved hearing aid system 600, the second improved hearing aid system 700, and the third improved hearing aid system 800 generates corresponding output signals 606, 706, and 806 based on the received corresponding input signals 604, 704, and 804 through the corresponding operations of the first trained deep neural network 602, the second trained deep neural network 702, and the third trained deep neural network 802, respectively, each of the first improved hearing aid system 600, the second improved hearing aid system 700, and the third improved hearing aid system 800 operates in a correspondingly different manner.

[0056] End-to-end neural networks

[0057] For more specific reference Figure 6 The first improved hearing aid system 600 is an end-to-end neural network implementation. In this embodiment, the first trained deep neural network 602 is an end-to-end neural network that directly estimates the expected clean speech signal. In this embodiment, the signal indicated by arrow 654 is the expected speaker selection (l), representing the clean speech fed to the neural network during training and not part of the system during inference operation mode. As described above, the corresponding output signal 606 (e.g., m output signals) ) can be a corresponding speaker / sound output device of the corresponding first improved hearing aid system 600 (e.g., Figure 2 The audio information or signal output by the corresponding output speaker 206 of the combined input / output device 202 (e.g., different speakers of a hearing aid). In this embodiment, based on head angle information or changes in head angle ( — The first trained deep neural network 602 utilizes the data embedded in the corresponding input (audio) signal 604. Indirect extraction of information from the source—to complete the selection of the expected speaker.

[0058] In this embodiment, equation (2) shows that it considers a single output. A simpler example of a loss function:

[0059] (2)

[0060] in Neural networks use weights Estimated pure speech It depends on the listener's central axis and the speaker's angle of arrival (downward angle). The angle between the expected speaker's pure voice, This can be any loss function of choice, such as multi-resolution spectrogram loss or a more specific metric, such as the Hearing Aid Speech Quality Index (HASQI). For better performance, primarily in denoising, the loss function also considers phase.

[0061] Neural network estimation of linear filters

[0062] For further reference Figure 7 The second improved hearing aid system 700 is an implementation scheme in which a neural network estimation with a linear filter is present. In this embodiment, the second trained deep neural network 702 is based on the corresponding input signal 704 ( Estimate and output k coefficients (or parameters) of the linear filter used for beamforming. 708 Further, as shown in the figure, the corresponding noisy input (audio) signal 704 ( The multiplication is performed in the frequency domain with the coefficients 708 of the linear filter (or convolved in the time domain), as shown in multiplication block 710. This operation at multiplication block 710 produces the corresponding output signal 706 (e.g., m output signals). The output signal 706 may be generated by a corresponding speaker / sound output device of the corresponding first improved hearing aid system 600 (e.g., Figure 2 The combined input / output device 202 outputs audio information or signals from the corresponding output speaker 206 (e.g., different speakers of a hearing aid). More specifically, this operation at multiplication block 710 produces a corresponding output signal 706 that constitutes the expected clean speech output. Again, regarding Figure 7 In one embodiment, the concept of changing the intended speaker during training is also taken into consideration.

[0063] For considering only one output of Figure 7 For a system like this, the loss function in equation (2) can be modified into the form of equation (3):

[0064] (3)

[0065] in , , .

[0066] Neural network estimation of filter statistics

[0067] Further reference Figure 8 The third improved hearing aid system 800 is an implementation scheme in which statistical information of filters is contained in a neural network estimation. That is, in the third improved hearing aid system 800, a third trained deep neural network 802 is based on the corresponding input signal 804 ( The system estimates and outputs the statistics 808 of the minimum variance distortionless response (MVDR) filter 810. Also shown in the figure, the statistics 808 and the corresponding input signal 804 ( The signal is supplied to the MVDR filter 810. The operation of the MVDR filter then generates corresponding output signals 806 (e.g., m output signals). The output signal 806 can be generated by the corresponding speaker / sound output device of the corresponding third improved hearing aid system 800 (e.g., Figure 2 The combined input / output device 202 outputs audio information or signals from the corresponding output speaker 206 (e.g., different speakers of a hearing aid). More specifically, this operation at the MVDR filter 810 produces a corresponding output signal 806 that constitutes the desired clean speech output. Again, regarding Figure 8 In one embodiment, the concept of changing the intended speaker during training is also taken into consideration.

[0068] Regarding the third improved hearing aid system 800, when only When outputting, the possible loss function is shown in equation (4):

[0069] (4)

[0070] MVDR coefficient Directional array dependent on neural network estimation Inverse matrix related to noise ,and .

[0071] The third improved hearing aid system 800 represents various embodiments that operate by performing neural network estimation of the statistical information of the filter. Although Figure 8 A third improved hearing aid system 800 employing an MVDR filter is illustrated, but this disclosure also covers embodiments employing other filters and estimating (or generating) or providing statistical information (e.g., statistical information 808) for such other filters (including various other known filters). In practice, the third improved hearing aid system 800 can be used to extend the performance of known, reliable, and stable filter structures.

[0072] In addition to those described above, this disclosure also covers numerous embodiments and variations thereof, including various different systems and various different operating methods and implementation schemes, including methods involving training mode operation, inference mode operation, and combinations of both training mode operation and inference mode operation. For example, Figure 6 , Figure 7 and Figure 8 The corresponding input signals 604, 704 and 804 and Figure 3 Output signal 306 ( Figure 6 , Figure 7 , Figure 8 and Figure 3 Each of the ones in … These can be used directly as the microphone output, but they can also be preprocessed versions of such outputs. A common approach is to calculate the Short-Term Fourier Transform (STFT) of each microphone output and cascade its real and imaginary parts. Other types of filtering can also be applied. Additionally, multiple input features derived from the microphone outputs can be combined, for example, by cascading the STFT of the output obtained at each microphone pair with its generalized cross-correlation and phase transform (GCC-PHAT).

[0073] While the principles of the invention have been described above in conjunction with specific devices, it should be clearly understood that this description is by way of example only and not as a limitation on the scope of the invention. It is particularly intended that the invention is not limited to the embodiments and illustrations contained herein, but also includes modifications of those embodiments, including portions of embodiments falling within the scope of the appended claims and combinations of elements from different embodiments.

Claims

1. A hearing aid system, characterized in that, include: One or more memory devices, the one or more memory devices being configured to store a first neural network; One or more audio input devices, the one or more audio input devices being configured to receive an audio input signal including audio information generated from a plurality of sound sources; One or more audio output devices; as well as One or more processing devices, said one or more processing devices being at least indirectly coupled to said one or more memory devices, said one or more audio input devices and said one or more audio output devices, During inference mode, the one or more processing devices are configured to generate an intermediate output signal according to the first neural network operation. This intermediate output signal reflects or emphasizes at least one expected sound source component of the audio information generated from an expected sound source among the plurality of sound sources to a greater extent than in the audio information. The expected sound source is identified at least indirectly based on a first downshoot angle clearly visible from the audio information. The one or more audio output devices are configured to generate an audio output signal at least indirectly based on the intermediate output signal.

2. The hearing aid system according to claim 1, characterized in that, The one or more audio input devices include one or more microphones, the one or more audio output devices include one or more speakers, the one or more processing devices include at least one microprocessor or graphics processing unit (GPU), and the first neural network is a deep neural network.

3. The hearing aid system according to claim 2, characterized in that, The hearing aid system is a hearing aid system.

4. The hearing aid system according to claim 1, characterized in that, The audio input device is positioned on or associated with a human listener having a central axis extending therefrom, wherein the plurality of sound sources includes a plurality of human sound sources, and wherein the intended sound source among the sound sources is a first human sound source among the human sound sources.

5. The hearing aid system according to claim 4, characterized in that, Multiple additional axes extend between a corresponding human sound source and the human listener, wherein multiple angular differences exist between the listener's central axis and a corresponding additional axis among the multiple additional axes, and wherein the first down-angle is a first angular difference between the listener's central axis and a first additional axis among the additional axes, the first angular difference being less than each of the other angular differences at a first time.

6. The hearing aid system according to claim 5, characterized in that, The one or more processing devices are further configured to operate to determine a second downshoot angle different from the first downshoot angle, wherein the second downshoot angle is a second angular difference between the listener's central axis and a second additional axis of the additional axes, the second angular difference becoming smaller than the first angular difference at a second time, and At the second time point or substantially at the second time point, the switching of the expected sound source among the sound sources is from the first human sound source among the human sound sources to the second human sound source among the human sound sources associated with the second additional axis in the additional axis.

7. The hearing aid system according to claim 1, characterized in that, The intermediate output signal is a linear filter coefficient, and the one or more processing devices are further configured to operate to multiply or convolve the linear filter coefficient with the audio input signal or at least indirectly based on an additional signal of the audio input signal to generate an additional intermediate output signal, wherein the audio output signal is at least indirectly based on the additional intermediate output signal.

8. The hearing aid system according to claim 1, characterized in that, The intermediate output signal is statistical information of the filter, and wherein the one or more processing devices are further configured to operate to process the statistical information and the audio input signal through the filter to which the statistical information pertains, or at least indirectly based on an additional signal of the audio input signal, to generate an additional intermediate output signal, wherein the audio output signal is at least indirectly based on the additional intermediate output signal.

9. A method for training a first neural network for use in a hearing aid system, characterized in that, The method includes: One or more audio input devices are provided within an area containing multiple sound sources; Receive an input signal at one or more audio input devices, the input signal including undercut angle data as described elsewhere herein; The input signal or an intermediate signal based on the input signal is provided to the first neural network; The first neural network generates multiple output signals; At the loss processing block, the output signal and the expected speaker clean speech data determined at least in part based on the downshoot angle data are processed to determine multiple weighted signals; and The first neural network is updated based on the weight signals.

10. A method for operating a hearing aid system during reasoning mode, characterized in that, The hearing aid system includes one or more memory devices configured to store a first neural network, and the method includes: Receives audio input signals at one or more audio input devices, the audio input signals including audio information generated from multiple sound sources; One or more processing devices are operated according to the first neural network to generate an intermediate output signal, the intermediate output signal reflecting or emphasizing at least one expected sound source component of the audio information generated from an expected sound source one of the plurality of sound sources to a greater extent than in the audio information, the expected sound source being identified at least indirectly based on a first downswing angle clearly visible from the audio information; and An audio output signal is generated at least indirectly based on the intermediate output signal at one or more audio output devices.