Hearing device and corresponding method, use and computer program product

By designing a structure including input converter, preprocessor, neural network processor and signal postprocessor in the listening device, the problems of detector parameter adjustment and memory bandwidth limitation are solved, and the personalization and performance improvement of listening device parameters are achieved.

CN120128868APending Publication Date: 2025-06-10OTICON
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510084127.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-10-08
Filing Date
2020-10-09
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The detectors in existing hearing devices require parameter adjustment, and due to limited memory, neural networks require as few parameters as possible and have limited bandwidth, so as to transmit as few parameters as possible.

Method used

A hearing device is designed, including an input converter, a preprocessor, a neural network processor and a signal postprocessor. In adaptive operation mode, the device receives optimized node parameters from another part or server and implements the optimized neural network in the neural network processor.

Benefits of technology

Through the optimized neural network processor, the personalization of hearing device parameters is achieved, the performance of the detector is improved, and the memory and bandwidth limitations are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128868A_ABST
    Figure CN120128868A_ABST
Patent Text Reader

Abstract

The present application discloses a hearing device and a corresponding method, use and computer program product, wherein the hearing device comprises: an input transducer; the preprocessor is used for processing at least one electric input signal and providing a plurality of feature vectors; a neural network processor adapted to implement a neural network for implementing the detector, configured to provide an output indicative of a characteristic of the at least one electrical input signal; a post-processor configured to receive and process the output vector and provide a resulting signal; and a neural network controller connected to the neural network processor for receiving the optimized node parameters in the adaptive mode of operation and applying the optimized node parameters to nodes of the neural network to implement the optimized neural network in the neural network processor, the optimized node parameters being saved in the hearing device.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese Patent Application No. 202011074492.0, filed on October 9, 2020, with the invention title "Hearing Device Comprising a Detector and a Trained Neural Network". Technical Field

[0002] The present application relates to a hearing device, such as a hearing aid, comprising a detector, for example, for detecting an acoustic environment, said detector being, for example, a voice detector for detecting a specific keyword for a voice control interface. The present application also relates to a solution for personalizing hearing device parameters. Background Art

[0003] Detectors, such as environment detectors, self-voice detectors or keyword detectors, usually require parameter adjustment. A detector, for example, a detector that provides a decision or provides a parameter value of one or more estimates or the probability of said estimated parameter values, may be implemented using or based on supervised learning, for example, using a neural network architecture in whole or in part. Since the architecture of neural networks is very general, neural networks usually require many parameters, such as weight and bias parameters. Due to the limited memory in hearing devices, the implemented neural network is required to have as few parameters as possible. In addition, due to the limited bandwidth during programming, it is desirable to transmit as few parameters as possible to the hearing instrument. Summary of the Invention

[0004] Hearing device

[0005] In a first aspect of the present application, there is provided a hearing device configured to be located at or in a user's ear or fully or partially implanted in the user's head. The hearing device comprises:

[0006] - an input transducer comprising at least one microphone for providing at least one electrical input signal representative of sound in the environment of the hearing device;

[0007] - a pre-processor for processing said at least one electrical input signal and providing a feature vector (e.g., a plurality of feature vectors) representative of a time period of said at least one electrical input signal;

[0008] - a neural network processor adapted to implement a neural network for implementing a detector or a part thereof, for example, configured to provide an output indicating a unique characteristic of at least one electrical input signal, the neural network comprising an input layer, an output layer and a plurality of hidden layers, each layer comprising a plurality of nodes, each node being defined by a plurality of node parameters, the neural network being configured to receive said feature vector (or said plurality of feature vectors) as an input vector and provide an output vector (or corresponding output vectors) representative of the output of said detector or a part thereof according to said input vector.

[0009] The hearing device may further comprise a signal post-processor configured to receive the output vector and at least one input signal representative of a sound in the environment, and wherein the signal post-processor is configured to process the at least one input signal representative of the sound according to the output vector and provide a processed output, i.e., the resulting signal. The hearing device may further comprise a transceiver comprising a transmitter and a receiver for establishing a communication link to another part or another device or server, the communication link enabling transmission of data to and reception of data from another part or another device or server at least in an adaptive operating mode. The hearing device may further comprise a selector for sending the feature vector to the transmitter for transmission to another part or another device or server in the adaptive operating mode and for sending the feature vector to the neural network processor for use as an input to the neural network in a normal operating mode. The hearing device is configured to receive optimized node parameters from another part or another device or server in the adaptive operating mode and apply the optimized node parameters to the nodes of the neural network so as to implement an optimized neural network in the neural network processor, wherein the optimized node parameters have been determined based on the feature vector (e.g., selected among multiple sets of node parameters for corresponding candidate neural networks according to a predetermined criterion).

[0010] "Another part" may for example be another part of the hearing device (e.g., integrated with the hearing device, or physically separate from but communicating with the hearing device). "Another device" may for example be a separate (auxiliary) device, such as a separate wearable device like a remote control device or a smart phone, etc. "Server" may for example be a computer or storage medium remote from the hearing device, e.g., accessible via a network such as the Internet.

[0011] In a second aspect, the hearing device itself may comprise a plurality of pre-trained candidate neural networks such that the adaptive mode can only be executed in the hearing device, e.g., the adaptive mode runs during the daily use of the hearing device. The hearing device may comprise a plurality of pre-trained candidate neural networks (selecting an appropriate, best-performing neural network for the user involved from among the candidate neural networks for use in the hearing device), such that the adaptive mode can only be executed in the hearing device (e.g., see Figure 8 ). The hearing device may for example be configured to run the adaptive mode during the daily use of the hearing device.

[0012] Thus, an improved hearing device can be provided.

[0013] The detector can be configured, for example, to recognize characteristic features of at least one electrical input signal. The "characteristic features" can be related, for example, to features regarding the nature of the sound currently picked up by at least one microphone and represented by at least one electrical input signal. The "characteristic features" can be related, for example, to the acoustic properties of the current acoustic environment. The "characteristic features" can be related, for example, to voice components and noise components, own voice and other voices, single-speaker environments and multi-speaker environments, speech and music, the content of a given speech sequence, etc.

[0014] The predetermined criterion can be related to minimizing a cost function with respect to the output vector. When the plurality of feature vectors are extracted from a time period of the at least one electrical input signal having known properties, the predetermined criterion can be based, for example, on the performance of the neural network in terms of true positives, false positives, true rejections, and false rejections of frames of the output vector.

[0015] Multiple sets of node parameters for corresponding candidate neural networks can be optimized for different classes of people exhibiting different acoustic properties.

[0016] The neural network processor can be particularly adapted to perform the calculations of the neural network. The neural network processor can form part of a digital signal processor. The digital signal processor can include a neural network core particularly adapted to perform the operations of the neural network. The node parameters of a particular node can include, for example, weight parameters (w) and bias parameters (b). The non-linear function (f) associated with each node can be the same for all nodes, or can be different between layers or between nodes. The non-linear function (f) can be represented, for example, by a sigmoid function, a rectified linear unit (ReLU), or a softmax function. Different layers can have different non-linear functions. However, the parameters of the non-linear function can also be learned (e.g., included in the optimization process). This is the case for the rectified linear unit with respect to its parameters. In other words, the parameters of the non-linear function (f) can form part of the node parameters.

[0017] A given neural network can be defined by a given set of node parameters.

[0018] The detector (or a part thereof) implemented by a neural network can include, for example, a wake word detector, a keyword detector, or a preferred speaker detector (spouse detector). As an alternative, the detector (or a part thereof) implemented by a neural network can also include, for example, a correlation detector, a level estimator, a modulation detector, a feedback detector, a voice detector such as an own voice detector, a speech intelligibility detector of the current electrical input signal or a signal derived therefrom. The output of the detector can include an estimate of the value of a specific parameter, property, or content of the electrical input signal or the probability of such an estimated value.

[0019] The detector may include a neural network processor or form part of a neural network processor. The detector may include a pre-processor. The detector may include a post-processor.

[0020] The term "sound in the environment" may represent any sound in the environment of a user wearing a hearing device, which may reach at least one microphone and be detected as an electrical input signal. "Sound in the environment" may include, for example, any vocalizations of a person in the environment and any sounds emitted by a machine or device. "Sound in the environment" may include specific words or a plurality of specific words of a specific person (such as the user). "Sound in the environment" may include natural sounds or background noise.

[0021] The hearing device may include a neural network controller connected to the neural network processor, which is configured to receive optimized node parameters in an adaptive operating mode and apply the optimized node parameters to the nodes of the neural network to implement an optimized neural network in the neural network processor. The hearing device may include a candidate neural network. The candidate neural networks may belong to the same type, such as a feed-forward, recurrent, or convolutional neural network. The candidate neural networks may include different types of neural networks. The candidate neural networks may differ in the number of layers and nodes, for example, some candidate neural networks may have more layers (and / or a given layer may have more nodes) than other candidate neural networks. Thus, the complexity of the neural network can be adapted to a specific user.

[0022] The post-processor may be configured to provide a "decision" based on the output vector of the neural network. The post-processor may be configured to determine a resulting (estimated) detector value based on the output vector of the neural network, for example, based on the probabilities of a plurality of different (estimated) detector values. The decision or the resulting estimated detector value may be, for example, a specific wake word or command word of a voice control interface or a specific value of a voice detector such as a self-voice detector indicating the presence or absence of voice, such as the user's self-voice.

[0023] The output of the decision-making unit (post-processor), i.e., the resulting signal, may be, for example, a command word or sentence for activating the voice control interface or a wake word or sentence. As an alternative or in addition, the output of the decision-making unit (post-processor) may be fed to a transmitter (or transceiver) and transmitted to another device for further processing there, for example, to activate a personal assistant (such as a smart phone, etc.). The transceiver may receive a response from another device, such as from the personal assistant. This response may be used, for example, to control the hearing device, or it may be played to the user via the output transducer (SPK) of the hearing device.

[0024] A hearing device may include sensors for sensing characteristics of a user or the hearing device environment and providing sensor signals representing current values of the environmental characteristics. In this specification, the environmental characteristics of the hearing device may include characteristics of the user wearing the hearing device (such as parameters or states). In this specification, the environmental characteristics of the hearing device may include characteristics of the physical environment such as the acoustic environment of the hearing device. In this specification, the term "sensor" may refer to a device that provides an output signal representing the value of the current physical parameter that the sensor is adapted to measure. For example, a temperature sensor gives the current temperature of the sensor environment as an output. The hearing device may include multiple sensors. The sensors may be constituted by or include the following sensors: motion sensors (such as accelerometers), magnetometers, EEG sensors, EOG sensors, heart rate detectors, or temperature sensors, etc. A voice activity detector (VAD) or an OV detector may also be regarded as a sensor.

[0025] The hearing device may be configured such that the sensor signal is an input to a preprocessor.

[0026] The preprocessor may be configured to process at least one electrical input signal and the sensor signal to provide a feature vector. The sensor signal may be used to qualify at least one electrical input signal (for example, exclude certain time periods (or frequency bands) thereof) before generating the feature vector. As an alternative or in addition, the sensor signal (or its features) may be included in the feature vector provided by the preprocessor and used as an input vector to a neural network.

[0027] The hearing device may include an output unit for presenting the processed output signal to the user as a perceptible sound stimulus. The output unit may include an output transducer. The output unit (such as the output transducer) may include a loudspeaker of an air-conduction hearing device, a vibrator of a bone-conduction hearing device, or a multi-electrode array of a cochlear implant hearing device. The output unit may include a transmitter for transmitting the resulting (detector) signal and / or the output (audio) signal to another device.

[0028] The hearing device may include an analysis filter bank for converting a time-domain input signal into multiple sub-band signals to provide an input signal with a time-frequency representation (k, l), where k and l are frequency and time indices respectively. For each electrical input signal and / or sensor signal, the input transducer may include an analysis filter bank. The hearing device may include an analysis filter bank for converting a time-domain sensor signal into multiple sub-band signals to provide a sensor signal with a time-frequency representation (k, l), where k and l are frequency and time indices respectively. The feature vector may be provided in the time-frequency representation. If it is not necessary to reconstruct the time-domain signal (if only the feature vector is used for the detector, it is not necessary), the filter bank may be downsampled by a factor higher than the critical downsampling. We may also use a smaller subgroup of the available channels of the filter bank, and we may add the channels together. The filter bank channels may be low-pass filtered before downsampling.

[0029] The preprocessor can be configured to extract features of at least one electrical input signal and / or sensor signal. The features of at least one electrical input signal can for example include modulation, level, signal-to-noise ratio (SNR), correlation, etc. The feature vector can exhibit reduced complexity compared to the at least one electrical input signal and / or sensor signal from which the features are extracted. The feature vector can be provided according to the time-frequency representation of an audio signal (obtained by a filter bank or a warped filter bank). This time-frequency representation can be further processed into a magnitude response, and the magnitude response can be low-pass filtered and / or downsampled. In the case where the hearing device includes multiple microphones (and thus has access to multiple electrical input signals representing sound), some or all of the electrical input signals can be combined into a directional (beamformed) signal, for example a directional signal that enhances the user's own voice. The directional signal can be further enhanced by noise reduction, for example using a post-filter.

[0030] The time period of the corresponding values of at least one electrical input signal and optionally sensor signal covered by a given feature vector is used as the input to the input layer of a neural network, which includes at least one time frame of the at least one electrical input signal. The time period can include multiple time frames of the at least one electrical input signal, for example more than 3 such as more than 5 time frames, such as 2 to 50 time frames, for example corresponding to up to 0.5 to 1 s of audio such as corresponding to one or more words.

[0031] The communication link established by the transmitter and receiver circuits of the transceiver of the hearing device can be a wireless link. The wireless link can be an inductive link based on inductive coupling between an inductor of the hearing device and another device or server. The wireless link can be based on a radiation field, for example based on a Bluetooth or WLAN transceiver in the hearing device and another device or server.

[0032] The hearing device is configured to operate in multiple operating modes, including an adaptive operating mode and a normal operating mode. Whether a user is allowed to enter the adaptive mode may depend on whether the hearing device is detected as being worn on the ear. The adaptive operating mode can be, for example, the operating mode when the hearing device is connected to an auxiliary device (including a processor such as a smart phone, a PC, or a laptop computer). The auxiliary device can be configured, for example, to run the fitting software of the hearing device. The fitting software can be configured to customize (personalize) the algorithm of the hearing device for the needs of a specific user. As part of the personalization of the hearing device for the specific user involved, it includes determining the optimized node parameters of the neural network to be applied in the hearing device. This set of optimized node parameters of the neural network for a given user can be selected among multiple predetermined, optimized node parameters according to a suitable selection algorithm. The multiple predetermined, optimized node parameters of the candidate neural network used when customizing the hearing device can be stored in the hearing device. Thus, without using an auxiliary device, a suitable set of optimized node parameters can be selected, for example, by the user himself. We can also imagine a decision based on a neural network (NN) operating in one hearing device and another NN operating in another device. One NN can be used, for example, in a low SNR environment, while another NN can be used in a high SNR environment.

[0033] The detector or a part thereof implemented by the neural network can be a self-voice detector and / or a keyword detector. The self-voice detector can be implemented by the neural network. The keyword detector (such as a wake word detector) can be implemented by the neural network. The keyword detector (such as a wake word detector) specifically optimized to detect words spoken by the user can be implemented by the neural network. Additionally or alternatively, the neural network can implement an on / off detector or a preferred speaker detector such as a spouse (voice) detector.

[0034] The hearing device can be composed of or include a hearing aid, headphones, a headset, an ear protection device, or a combination thereof.

[0035] The hearing device can be adapted to provide frequency-varying gain and / or level-varying compression and / or frequency shifting (with or without frequency compression) from one or more frequency ranges to one or more other frequency ranges to compensate for the user's hearing impairment. The hearing device can include a signal processor for enhancing the input signal and providing a processed output signal.

[0036] A hearing device may include an output unit for providing a stimulus that is perceived by a user as an acoustic signal based on a processed electrical signal. The output unit may include a plurality of electrodes of a cochlear implant (for CI-type hearing devices) or a vibrator of a bone-conduction hearing device. The output unit may include an output transducer. The output transducer may include a receiver (loudspeaker) for providing the stimulus as an acoustic signal to the user (e.g., in an acoustic (air-conduction-based) hearing device). The output transducer may include a vibrator for providing the stimulus as a mechanical vibration of the skull to the user (e.g., in a bone-attached or bone-anchored hearing device).

[0037] A hearing device includes an input transducer for providing an electrical input signal representative of sound. The input transducer may include a microphone for converting an input sound into an electrical input signal. The input transducer may include a wireless receiver for receiving a wireless signal including or representative of sound and providing an electrical input signal representative of the sound. The wireless receiver may be configured, for example, to receive electromagnetic signals in the radio frequency range (3 kHz to 300 GHz). The wireless receiver may be configured, for example, to receive electromagnetic signals in the optical frequency range (e.g., infrared light from 300 GHz to 430 THz, or visible light, e.g., from 430 THz to 770 THz).

[0038] A hearing device may include a directional microphone system adapted to spatially filter sounds from the environment to enhance a target sound source among a plurality of sound sources in the local environment of a user wearing the hearing device. The directional system is adapted to detect (e.g., adaptively detect) from which direction a particular portion of the microphone signal originates. This can be achieved in a number of different ways described, for example, in the prior art. In hearing devices, microphone array beamformers are commonly used to spatially attenuate background noise sources. Many beamformer variants can be found in the literature. The minimum variance distortionless response (MVDR) beamformer is widely used in microphone array signal processing. Ideally, the MVDR beamformer leaves the signal from the target direction (also referred to as the look direction) unchanged while maximally attenuating sound signals from other directions. The generalized sidelobe canceller (GSC) structure is an equivalent representation of the MVDR beamformer that offers computational and digital representation advantages over a direct implementation of the original form. A hearing device may also include a spatially based post-filter, for example, as part of a noise reduction system.

[0039] A hearing device may include an antenna and transceiver circuitry (such as a wireless receiver) for receiving a direct electrical input signal from another device, such as from an entertainment device (e.g., a television), a communication device, a wireless microphone, or another hearing device. The direct electrical input signal may represent or include an audio signal and / or a control signal and / or an information signal. The hearing device may include demodulation circuitry for demodulating the received direct electrical input to provide a direct electrical input signal representing an audio signal and / or a control signal, for example for setting operating parameters (such as volume) and / or processing parameters of the hearing device. Generally, the wireless link established by the antenna and transceiver circuitry of the hearing device can be of any type. The wireless link is established between two devices, for example between an entertainment device (such as a TV) and the hearing device, or between two hearing devices, for example via a third intermediate device (such as a processing device, for example a remote control device, a smart phone, etc.). The wireless link is typically used under power constraints, for example since the hearing device is or includes a portable (usually battery-powered) device. The wireless link is a near-field communication-based link, for example an inductive link based on inductive coupling between antenna coils of a transmitter part and a receiver part. In another embodiment, the wireless link is based on far-field electromagnetic radiation. Communication via the wireless link is arranged according to a specific modulation scheme, for example an analog modulation scheme, such as FM (frequency modulation) or AM (amplitude modulation) or PM (phase modulation), or a digital modulation scheme, such as ASK (amplitude shift keying) such as on-off keying, FSK (frequency shift keying), PSK (phase shift keying) such as MSK (minimum shift keying) or QAM (quadrature amplitude modulation), etc.

[0040] Communication between the hearing device and another device is in the baseband (audio frequency range, such as between 0 and 20 kHz). Preferably, the frequency used for establishing a communication link between the hearing device and another device is below 70 GHz, for example in the range from 50 MHz to 70 GHz, for example above 300 MHz, for example in the ISM range above 300 MHz, for example in the 900 MHz range or in the 2.4 GHz range or in the 5.8 GHz range or in the 60 GHz range (ISM = industrial, scientific and medical, such standardized ranges are defined, for example, by the International Telecommunication Union ITU). The wireless link is based on a standardized or proprietary technology. The wireless link is based on Bluetooth technology (such as Bluetooth Low Energy technology).

[0041] The hearing device can be a portable (i.e., configured to be wearable) device or form part of it, such as a device including a native energy source such as a battery, for example a rechargeable battery. The hearing device can be, for example, a lightweight, easily wearable device, for example having a total weight of less than 100 g, such as less than 20 g.

[0042] The hearing device may include a forward or signal path between an input unit (such as an input transducer, for example a microphone or a microphone system and / or a direct electrical input such as a wireless receiver) and an output unit such as an output transducer. A signal processor is located in this forward path. The signal processor is adapted to provide frequency-dependent gain according to the specific needs of the user. The hearing device may include an analysis path having functions for analyzing the input signal (such as determining the level, modulation, signal type, acoustic feedback estimate, etc.). Part or all of the signal processing in the analysis path and / or the signal path is performed in the frequency domain. Part or all of the signal processing in the analysis path and / or the signal path is performed in the time domain.

[0043] An analog electrical signal representing an acoustic signal is converted into a digital audio signal during an analog-to-digital (AD) conversion process, where the analog signal is sampled at a predetermined sampling frequency or sampling rate f s is sampled, f s for example in a range from 8 kHz to 48 kHz (adapting to the specific needs of the application) at discrete time points t n (or n) to provide digital samples x n (or x[n]), each audio sample being represented by a predetermined N b bits representing the value of the acoustic signal at t n when, N b for example in a range from 1 to 48 bits such as 24 bits. Each audio sample is thus quantized using N b bits (resulting in 2 Nb different possible values of the audio sample). The digital sample x has a time length of 1 / f s , for f s = 20 kHz, such as 50 μs. A plurality of audio samples are arranged in time frames. A time frame includes 64 or 128 audio data samples. Other frame lengths may be used according to the actual application.

[0044] The hearing device may include an analog-to-digital (AD) converter to digitize an analog input (such as from an input transducer such as a microphone) at a predetermined sampling rate such as 20 kHz. It should be mentioned that although the (maximum) sampling rate in the forward path of the hearing device is, for example, 20 kHz (supporting an audio frequency range of 0 - 10 kHz), different, smaller frequency ranges may be used in other parts of the hearing device. It includes a digital-to-analog (DA) converter to convert the digital signal into an analog output signal, for example for presentation to the user via an output transducer.

[0045] A hearing device such as an input transducer and / or an antenna and transceiver circuitry includes a TF conversion unit for providing a time-frequency representation of an input signal. The time-frequency representation may include an array or mapping of corresponding complex-valued or real-valued of the signal involved in a specific time and frequency range. The TF conversion unit may include a filter bank for filtering the (time-varying) input signal and providing a plurality of (time-varying) output signals, each output signal including a distinct input signal frequency range. The TF conversion unit may include a Fourier transform unit for converting the time-varying input signal into a (time-)frequency domain (time-varying) signal. The frequency range considered by the hearing device, from a minimum frequency f min to a maximum frequency f max includes a part of the typical human audible frequency range from 20 Hz to 20 kHz, for example a part of the range from 20 Hz to 12 kHz. Generally, the sampling rate f s is greater than or equal to twice the maximum frequency f max , i.e., f s ≥2f max . The signals in the forward path and / or analysis path of the hearing device are split into NI (e.g., uniformly wide) frequency bands, where NI is, for example, greater than 5, such as greater than 10, such as greater than 50, such as greater than 100, such as greater than 500, and at least some of them are processed individually. The hearing device is adapted to process the signals in the forward and / or analysis path in NP different channels (NP ≤ NI). For a detector, such as a detector for analyzing the signals in the forward path, such as an electrical input signal, we may have a smaller number of channels, e.g., NP’ ≤ NP. The channels may have the same width or different widths (e.g., the width increases with frequency), overlapping or non-overlapping.

[0046] The hearing device may be configured to operate in different modes, such as a normal mode and one or more specific modes, which may be selected by the user or automatically selected, for example. The operating mode may be optimized for a specific acoustic situation or environment. The operating mode may include a low power mode, in which the functions of the hearing device are reduced (e.g., for energy saving), for example, disabling wireless communication and / or disabling specific features of the hearing device. The operating mode may include a specific adaptive mode, in which the hearing device is connected to a plurality of different candidate neural networks (e.g., located in the hearing device or an auxiliary device), and one of the candidate neural networks is planned to be selected for use in the hearing device.

[0047] The hearing device may include a plurality of detectors configured to provide status signals related to the current network environment of the hearing device (such as the current acoustic environment), and / or related to the current state of the user wearing the hearing device, and / or related to the current state or operating mode of the hearing device. As an alternative or in addition, one or more detectors may form part of an external device that communicates with the hearing device (such as wirelessly). The external device may include, for example, another hearing device, a remote control, an audio transmission device, a phone (such as a smart phone), an external sensor, etc.

[0048] One or more of the plurality of detectors may be configured to operate on the full-band signal (time domain). One or more of the plurality of detectors may be configured to operate on the frequency-band split signal ((time-)frequency domain), for example in a finite number of frequency bands.

[0049] The plurality of detectors may include level detectors for estimating the current level of the signal in the forward path. The predetermined criterion may include whether the current level of the signal in the forward path is above or below a given (L-)threshold. One or more level detectors may operate on the full-band signal (time domain). One or more level detectors may operate on the frequency-band split signal ((time-)frequency domain).

[0050] The hearing device may include a voice activity detector (VAD) for estimating whether (or with what probability) the input signal (at a particular time point) includes a voice signal. In the present specification, the voice signal includes speech signals from humans. It may also include other forms of vocalization (such as singing) produced by the human speech system. The voice activity detector unit is adapted to classify the user's current acoustic environment as a "voice" or "no voice" environment. This has the advantage that time periods of the microphone signal that include human vocalizations (such as speech) in the user's environment can be identified and thus separated from time periods that include only (or mainly) other sound sources (such as artificially generated noise). The voice activity detector is adapted to also detect the user's own voice as "voice". Alternatively, the voice detector is adapted to exclude the user's own voice from the detection of "voice".

[0051] The hearing device may include a self-voice detector for estimating whether (or with what probability) a particular input sound (such as voice, such as speech) originates from the voice of the user of the hearing system. The microphone system (or the self-voice detector) of the hearing device may be adapted to be able to distinguish the user's own voice from the voice of another person and possibly from non-voice sounds.

[0052] The plurality of detectors may include a motion detector, such as an acceleration sensor. The motion detector is configured to detect motion of the user's facial muscles and / or bones caused, for example, by speech or chewing (such as jaw movement) and provide a detector signal indicating the motion.

[0053] The hearing device may include a classification unit configured to classify the current situation based on the input signals from (at least part of) the detectors and possibly other inputs. In the present specification, the "current situation" is defined by one or more of the following:

[0054] a) The physical environment (such as including the current electromagnetic environment, for example the occurrence of planned or unplanned electromagnetic signals received by the hearing device (including audio and / or control signals), or other properties of the current environment that are different from acoustic);

[0055] b) The current acoustic situation (input level, feedback, etc.);

[0056] c) The current mode or state of the user (motion, temperature, cognitive load, etc.);

[0057] d) The current mode or state of the hearing device and / or another device communicating with the hearing device (selected program, time elapsed since the last user interaction, etc.).

[0058] The classification unit may be based on or include a neural network, such as a trained neural network.

[0059] The hearing device may also include other suitable functions for the applications involved, such as compression, noise reduction, feedback control, etc.

[0060] The hearing device may include a sound receiving device, such as a hearing aid, such as a hearing instrument, for example, a hearing instrument adapted to be located at the ear or fully or partially located in the user's ear canal, such as a headset, an earphone, an ear protection device, or a combination thereof. The hearing assistance system may include a horn loudspeaker (including a plurality of input transducers and a plurality of output transducers, for example, used in an audio conferencing scenario), for example, including a beamforming filter unit, for example, providing multiple beamforming capabilities.

[0061] Application

[0062] On the one hand, there is provided an application of the hearing device as described above, detailed in the "Detailed Description" section and defined in the claims. Applications in systems including audio distribution may be provided. Applications in systems including one or more hearing aids (such as hearing instruments), headsets, earphones, active ear protection systems, etc. may be provided, for example, uses in hands-free telephone systems, remote conferencing systems (for example, including horn loudspeakers), broadcast systems, karaoke systems, classroom amplification systems, etc.

[0063] Method

[0064] On the one hand, the present application also provides a method for selecting optimized parameters of a neural network used in a portable hearing device. The method includes:

[0065] - Providing a portable hearing device to be used by a specific user, the hearing device including a neural network processor adapted to implement a neural network including an input layer, an output layer, and a plurality of hidden layers, each layer including a plurality of nodes, each node being defined by a plurality of node parameters and a non-linear function, the neural network being configured to receive an input vector and provide an output vector of a specific non-linear function of the input vector;

[0066] - Installing the hearing device on the user;

[0067] - Provide an auxiliary device;

[0068] - Establish a communication link enabling data exchange between the hearing device and the auxiliary device;

[0069] In the hearing device, the method may further include:

[0070] - Provide at least one electrical input signal representing sounds in the user environment in which the hearing device is worn;

[0071] - Process the at least one electrical input signal and provide a feature vector representing a time period of the at least one electrical input signal;

[0072] - Transmit the feature vector via the communication link to the auxiliary device.

[0073] In the auxiliary device, the method may further include:

[0074] - Provide a plurality of pre-trained candidate neural networks, each candidate neural network having the same structure as the neural network of the hearing device, where each pre-trained network is considered a candidate network for that person, and where each pre-trained neural network has been trained based on completely or partially different training data;

[0075] - Receive the feature vector from the hearing device and provide it as an input vector to the plurality of pre-trained candidate neural networks;

[0076] - Determine corresponding output vectors corresponding to the feature vector by the plurality of pre-trained candidate neural networks;

[0077] - Compare the output vectors and select one of the plurality of candidate neural networks as the neural network optimized for the hearing device according to a predetermined criterion;

[0078] - Transmit the node parameters of the selected neural network among the plurality of candidate neural networks via the communication link to the hearing device, and

[0079] In the hearing device, the method may further include:

[0080] - Receive the node parameters and feed them to the neural network processor and apply them to the neural network.

[0081] On the other hand, the present invention provides a method for selecting optimized parameters of a neural network used in a portable hearing device. The method includes:

[0082] - Provide a portable hearing device for use by a particular user, the hearing device including a neural network processor adapted to implement a neural network including an input layer, an output layer, and a plurality of hidden layers, each layer including a plurality of nodes, each node being defined by a plurality of node parameters and a non-linear function, the neural network configured to receive an input vector and provide an output vector that is a particular non-linear function of the input vector;

[0083] - Install the hearing device on the user;

[0084] - Provide at least one electrical input signal representing sound in the environment of the user wearing the hearing device;

[0085] - Process the at least one electrical input signal and provide a plurality of feature vectors, each feature vector representing a time period of the at least one electrical input signal;

[0086] - Provide a plurality of pre-trained candidate neural networks, where each pre-trained network is considered a candidate network for the user, and where each pre-trained neural network has been trained based on completely or partially different training data;

[0087] - Receive the feature vectors and provide them as input vectors to the plurality of pre-trained candidate neural networks;

[0088] - Determine corresponding output vectors corresponding to the feature vectors by the plurality of pre-trained candidate neural networks;

[0089] - Compare the output vectors and select one of the plurality of candidate neural networks as the neural network optimized for the hearing device according to a predetermined criterion regarding the output vectors;

[0090] - Transfer the node parameters of the selected neural network among the plurality of candidate neural networks to the neural network processor of the hearing device, and

[0091] - Receive the node parameters into the neural network processor and apply them to the neural network.

[0092] When appropriately replaced by corresponding processes, some or all of the structural features of the device described above, detailed in the "Detailed Description" or defined in the claims, may be combined with the implementation of the corresponding method, and vice versa. The implementation of the method has the same advantages as the corresponding device.

[0093] The term "neural network having the same structure as the neural network of the hearing device" means that each candidate neural network has the same number of layers and each layer has the same number of nodes, but each candidate neural network has different node parameters (weights and / or biases and / or non-linear functions) determined by pre-training with different training data sets. A given neural network can be defined by a given set of node parameters.

[0094] Different candidate neural networks can be trained, for example, based on data from different populations (distinguished by age, gender), and the selected network is based on information provided about the person (such as age and / or gender).

[0095] The term "a pre-trained neural network has been trained based on completely or partially different training data" includes, for example, that the noisy part of the training data can be similar. The noise can include, for example, training data in which no relevant speech elements can be detected. The training data can also include speech data that is close to (but different from) the word expected to be detected by a detector (such as a keyword detector).

[0096] The predefined criterion can be based on the performance of the neural network in terms of true positives, false positives, true negatives, and false negatives.

[0097] Different metrics can be obtained from the four terms TP, TN, FP, and FN (true positive, true negative, false positive, false negative), for example, "accuracy" = (TP + TN) / (TP + TN + FP + FN), "precision" = TP / (TP + FP), or "recall" = TP / (TP + FN) (for example, see https: / / developers.google.com / machine- learning / crash - course / classification / Accuracy ?)

[0098] Each candidate neural network can have the same structure. At least some (such as one) candidate neural networks can have a structure different from the rest of the candidate neural networks, for example, different numbers of layers and different numbers of nodes. The candidate neural networks can include different types of neural networks. The type of neural network can be selected, for example, from feedforward, recursive, convolutional, etc.

[0099] Each candidate neural network can have been trained based on training data from different categories of people. Such a training procedure can be suitable for self-voice applications such as self-voice detection (OVD) and keyword spotting (KWS) (such as wake-word detection). Different groups of people belong to different categories. The categories can be generated, for example, from the basic populations outlined in the Figure 6 Overview. Different categories of people can exhibit different acoustic properties, such as different head-related transfer functions (HRTFs), different voices ("spectral signatures", fundamental frequencies, and / or formant frequencies, etc.), different ages, different genders, etc.

[0100] A signal (such as from a sensor) representing the current value of a characteristic of the user or the user environment can be provided in the hearing device together with at least one electrical input signal and processed to provide the feature vector.

[0101] The method of the present invention can be configured such that the neural network implements a self-voice detector (OVD) and / or a keyword detector (KWD).

[0102] The method of the present invention may include prompting the user to speak. The method of the present invention may be configured to prompt the user to say one or more (e.g., predetermined) words or sentences (e.g., including words or sentences similar to the words or sentences that the detector is planned to recognize or detect). Thus, one or more words or sentences may form the basis of at least part of the feature vectors among the plurality of feature vectors. Thus, the output vector (ground truth) of the candidate neural network can be known, which can be used to check which candidate neural network (i.e., each of their neural network parameters, e.g., parameters optimized for a specific category of people) meets the criterion (e.g., best suits the current user).

[0103] The predetermined criterion for selecting the optimized neural network parameters may be based on the prompt words or sentences spoken by the user. The optimized neural network parameters may be selected by comparing the output vectors from candidate neural networks with corresponding neural network parameters (e.g., parameters optimized for a specific category of people) based on the prompt words or sentences spoken by the user. The selected neural network parameters may be the parameters from the candidate neural network that most meets the predetermined criterion (e.g., having the highest number of correct output vector values when the user says the prompt words or sentences).

[0104] The method of the present invention may include providing several sets of a plurality of pre-trained candidate neural networks, each candidate neural network in each set having the same structure as the neural network of the hearing device, where each pre-trained network is regarded as a candidate network for the person, and where each pre-trained neural network has been trained based on completely or partially different training data, and where each set of pre-trained candidate neural networks is aimed at implementing different detectors. Thus, the selection of the optimized parameters of several different neural networks implementing different detectors can be performed simultaneously. Different detectors may include, for example, a keyword detector for detecting a limited number of keywords (such as command words), a wake word detector for detecting a specific word or combination of words for activating the voice interface, and so on.

[0105] Computer - readable medium or data carrier

[0106] The present invention further provides a tangible computer-readable medium (or data carrier) storing a computer program including program code (instructions), which, when the computer program runs on a data processing system, causes the data processing system (computer) to execute (complete) at least part (such as most or all) of the steps of the method described above, detailed in the "Detailed Description" and defined in the claims.

[0107] By way of example and not limitation, the foregoing tangible computer-readable medium may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that is accessible by a computer. As used herein, a disk includes a compact disk (CD), a laser disk, an optical disk, a digital versatile disk (DVD), a floppy disk, and a Blu-ray disk, where these disks typically reproduce data magnetically while these disks can reproduce data optically with lasers. Other storage media include storage in DNA (e.g., in synthetic DNA strands). Combinations of the above disks should also be included within the scope of computer-readable media. In addition to being stored on a tangible medium, a computer program can also be transmitted via a transmission medium such as a wired or wireless link or network such as the Internet and loaded into a data processing system to run at a location different from the tangible medium.

[0108] Computer program

[0109] In addition, this application provides a computer program (product) including instructions that, when run by a computer, cause the computer to perform the methods (steps) described above, detailed in the "Detailed Description" and defined in the claims.

[0110] Data processing system

[0111] On the one hand, the present invention further provides a data processing system including a processor and program code that causes the processor to perform at least some (such as most or all) of the steps of the methods described above, detailed in the "Detailed Description" and defined in the claims.

[0112] Hearing system

[0113] On the other hand, there is provided a hearing system including a hearing device and an auxiliary device described above, detailed in the "Detailed Description" and defined in the claims.

[0114] The hearing system is adapted to establish a communication link between the hearing device and the auxiliary device so that information (such as control and status signals, possibly audio signals) can be exchanged or forwarded from one device to another.

[0115] The auxiliary device may include a remote control, a smart phone, or other portable or wearable electronic devices such as a smart watch, etc.

[0116] The auxiliary device can be or include a remote control for controlling the functions and operations of the hearing device. The functions of the remote control are implemented in a smart phone, which may run an APP that enables the functions of controlling the audio processing device via the smart phone (the hearing device includes a suitable wireless interface to the smart phone, such as based on Bluetooth or some other standardized or proprietary solution).

[0117] The auxiliary device can be or include an audio gateway device, which is adapted to receive multiple audio signals (e.g., from entertainment devices such as TVs or music players, from telephone devices such as mobile phones, or from computers such as PCs) and is adapted to select and / or combine appropriate signals (or signal combinations) from the received audio signals for transmission to the hearing device.

[0118] Multiple parts of the processing can be performed by the auxiliary device (the division can be, for example, that the OVD and several keywords related to the functions of the hearing device (including the wake-up word for the voice control interface) are detected in the hearing device, while other keywords are detected in the auxiliary device).

[0119] The auxiliary device can be constituted by or include another hearing device. The hearing system can include two hearing devices adapted to implement a binaural hearing system such as a binaural hearing aid system.

[0120] APP

[0121] On the other hand, the present invention also provides a non-transitory application called an APP. The APP can include executable instructions configured to run on the auxiliary device to implement a user interface for the hearing device or hearing system described above, detailed in the "Detailed Description" and defined in the claims. The APP is configured to run on a mobile phone such as a smart phone or another portable device enabling communication with the hearing device or hearing system.

[0122] Definition

[0123] In this specification, a "hearing device" refers to a device suitable for improving, enhancing, and / or protecting a user's hearing ability, such as a hearing aid, for example, a hearing instrument or an active ear protection device or other audio processing device, which is achieved by receiving an acoustic signal from the user's environment, generating a corresponding audio signal, possibly modifying the audio signal, and providing the possibly modified audio signal as an audible signal to at least one ear of the user. A "hearing device" also refers to a device suitable for electronically receiving an audio signal, possibly modifying the audio signal, and providing the possibly modified audio signal as an audible signal to at least one ear of the user, such as a headset or earphone. The audible signal can be provided, for example, in the following forms: an acoustic signal radiated into the user's outer ear, an acoustic signal transmitted as a mechanical vibration through the bone structure of the user's head and / or through parts of the middle ear to the user's inner ear, and an electrical signal transmitted directly or indirectly to the user's cochlear nerve.

[0124] The hearing device can be configured to be worn in any known manner, such as a unit worn behind the ear (with a tube for guiding the radiated acoustic signal into the ear canal or with an output transducer arranged close to or in the ear canal, such as a speaker), a unit arranged wholly or partly in the auricle and / or ear canal, a unit connected to a fixed structure implanted in the skull, such as a vibrator, or a connectable or wholly or partly implantable unit, etc. The hearing device can include a single unit or several units in electronic communication with each other. The speaker can be provided in a housing together with other components of the hearing device, or it can itself be an external unit (possibly combined with a flexible guiding element, such as a dome-shaped element).

[0125] More generally, a hearing device includes an input transducer for receiving an acoustic signal from the user's environment and providing a corresponding input audio signal and / or a receiver for receiving the input audio signal electronically (i.e., wired or wirelessly), a (usually configurable) signal processing circuit for processing the input audio signal (such as a signal processor, e.g., including a configurable (programmable) processor, e.g., a digital signal processor), and an output unit for providing an audible signal to the user based on the processed audio signal. The signal processor may be adapted to process the input signal in the time domain or in multiple frequency bands. In some hearing devices, an amplifier and / or a compressor may form part of the signal processing circuit. The signal processing circuit typically includes one or more (integrated or separate) storage elements for executing programs and / or for storing parameters used (or potentially used) in the processing and / or for storing information suitable for the function of the hearing device and / or for storing information used, for example, in connection with an interface to the user and / or to a programming device (such as processed information, e.g., provided by the signal processing circuit). In some hearing devices, the output unit may include an output transducer, such as a loudspeaker for providing an air-conducted acoustic signal or a vibrator for providing a structure- or fluid-conducted acoustic signal. In some hearing devices, the output unit may include one or more output electrodes for providing an electrical signal (such as a multi-electrode array for electrically stimulating the cochlear nerve). The hearing device may include a horn loudspeaker (including multiple input transducers and multiple output transducers, e.g., used in an audio conferencing scenario).

[0126] In some hearing devices, the vibrator may be adapted to transmit a structure-conducted acoustic signal to the skull transcutaneously or percutaneously. In some hearing devices, the vibrator may be implanted in the middle ear and / or the inner ear. In some hearing devices, the vibrator may be adapted to provide a structure-conducted acoustic signal to the middle ear bones and / or the cochlea. In some hearing devices, the vibrator may be adapted to provide a fluid-conducted acoustic signal to the cochlear fluid, for example, through the oval window. In some hearing devices, the output electrodes may be implanted in the cochlea or on the inner side of the skull and may be adapted to provide an electrical signal to the hair cells of the cochlea, one or more auditory nerves, the auditory brainstem, the auditory midbrain, the auditory cortex, and / or other parts of the cerebral cortex.

[0127] A hearing device, such as a hearing aid, can be adapted to the needs of a particular user, such as hearing impairment. The configurable signal processing circuit of the hearing device may be adapted to apply compression amplification of the input signal that varies with frequency and level. Customized gain (amplification or compression) that varies with frequency and level may be determined during the fitting process by a fitting system based on the user's hearing data, such as an audiogram, using fitting principles (such as adapted to speech). The gain that varies with frequency and level may be embodied, for example, in processing parameters, e.g., uploaded to the hearing device via an interface to a programming device (fitting system) and used by a processing algorithm executed by the configurable signal processing circuit of the hearing device.

[0128] "Hearing system" refers to a system that includes one or two hearing devices. "Binaural hearing system" refers to a system that includes two hearing devices and is adapted to provide audible signals to a user's two ears in a coordinated manner. A hearing system or a binaural hearing system may also include one or more "auxiliary devices" that communicate with the hearing devices and affect and / or benefit from the functions of the hearing devices. The auxiliary device may be, for example, a remote control, an audio gateway device, a mobile phone (such as a smart phone), or a music player. A hearing device, a hearing system, or a binaural hearing system may be used, for example, to compensate for the loss of auditory ability of a hearing-impaired person, enhance or protect the auditory ability of a person with normal hearing, and / or transmit an electronic audio signal to a person. A hearing device or a hearing system may form a part of or interact with, for example, a broadcast system, an active ear protection system, a hands-free phone system, a car audio system, an entertainment (such as karaoke) system, a teleconference system, a classroom amplification system, etc.

[0129] The present invention can be used, for example, in applications such as hearing aids, headsets, or similar devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0130] The various aspects of the present invention will be best understood from the following detailed description taken in conjunction with the accompanying drawings. For clarity, these drawings are schematic and simplified, showing only the details necessary for understanding the present invention and omitting other details. Throughout the specification, the same reference numerals are used for the same or corresponding parts. Each feature of each aspect can be combined with any or all features of other aspects. These and other aspects, features, and / or technical effects will be apparent from and elucidated in conjunction with the following drawings, in which:

[0131] Figure 1 A part of a hearing instrument with a built-in detector according to the present invention is shown;

[0132] Figure 2 A hearing instrument is shown wirelessly connected to an external device, which illustrates a selection procedure of an optimized neural network used in the hearing instrument according to the present invention;

[0133] Figure 3 An exemplary personalization procedure of hearing device parameters according to the present invention is shown;

[0134] Figure 4 A flowchart of an embodiment of a proposed method for selecting a personalized neural network is shown;

[0135] Figure 5 A hearing device according to an embodiment of the present invention is shown, which uses a trained (personalized for a specific user) neural network to control the processing of a signal representing sound before the processed signal is presented to a user wearing the hearing device;

[0136] Figure 6 An exemplary procedure is shown for subdividing a basic population for providing training data for training a neural network into a plurality of subgroups for training a plurality of neural networks, thereby providing a plurality of optimized neural networks, each optimized neural network representing a different characteristic of a person under test;

[0137] Figure 7A An embodiment of a keyword detector implemented as a neural network according to the present invention is shown;

[0138] Figure 7B An audio electrical input signal context for generating an input vector of a neural network is shown; Figure 7A of the neural network is shown;

[0139] Figure 8 An embodiment of a hearing device according to the present invention is shown, which includes an adaptive unit configured to enable selection of a set of optimized parameters for a neural network among multiple sets of optimized parameters without using an external device.

[0140] The further scope of application of the present invention will become apparent from the detailed description given below. However, it should be understood that while the detailed description and the specific examples, which show preferred embodiments of the invention, are given by way of illustration only, for a person skilled in the art other embodiments of the present invention will become apparent based on the following detailed description. DETAILED DESCRIPTION

[0141] The following detailed description presented in conjunction with the accompanying drawings is used as a description of various different configurations. The detailed description includes specific details for providing a thorough understanding of the various different concepts. However, it will be apparent to a person skilled in the art that these concepts may be implemented without these specific details. Several aspects of the apparatus and method are described by means of various different blocks, functional units, modules, elements, circuits, steps, processes, algorithms, etc. (collectively referred to as "elements"). Depending on a particular application, design constraints or other reasons, these elements may be implemented using electronic hardware, computer programs, or any combination thereof.

[0142] The electronic hardware may include a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic device (PLD), gated logic, discrete hardware circuits, and other suitable hardware configured to perform the various different functions described in this specification. A computer program should be construed broadly as instructions, instruction sets, code, code segments, program code, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, executables, execution threads, programs, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.

[0143] This application relates to the field of hearing devices such as hearing aids. Consider a hearing instrument system as shown in Figure 1 and having a microphone and possibly other sensors (such as accelerometers, magnetometers, EEG sensors, and / or heart rate detectors, etc.). The hearing instrument may have a built-in detector (for example, including one or more of the sensors mentioned).

[0144] One option could be to retain only a part of the layers, for example, keeping the weights of the first layer (or the first few layers) of the neural network fixed and only updating the deeper layers. Thereby fewer parameters need to be programmed.

[0145] One way to achieve better performance is to personalize the parameters of the neural network / detector. For example, compared to a neural network that has been optimized to work well for any population, a neural network optimized for a specific person or specific population (such as male, female, or children's voices, different ages, different languages or noise environments common to a given person) may work better.

[0146] Here we propose a method for selecting a personalized neural network.

[0147] Figure 1 A part of a hearing instrument with a built-in detector is shown. Input signals (IN1, IN2, SIN) from one or more microphones (M1, M2) and / or one or more sensors SENSE are preprocessed (see unit Pre-PRO) into a feature vector FV. One or more sensors SENSE may be omitted, such that the input signals to the preprocessor Pre-PRO are only the electrical input signals (IN1, IN2) from the microphones (M1, M2), and thus the feature vector depends only on the electrical input signals (IN1, IN2) from the microphones (M1, M2). The preprocessed feature vector FV is used as the input to a neural network NN. The output of the neural network NN can be, for example, a probability or a set of probabilities p(x) for making a decision and / or detection (such as the detection of a specific word, self-voice detection, or the detection of a certain sound environment, see decision unit PostPRO and output RES). The output RES of the decision unit can be used, for example, to decide on a specific action in a hearing device, such as activating a voice control interface. The output RES of the decision unit can be passed to another device or system, for example, to initiate a service.

[0148] In Figure 1In this case, the hearing device includes a single neural network implementing a (single) detector such as a self-voice detector. The hearing device may, for example, include several neural networks operating in parallel. For example, the hearing device may include a neural network for implementing a keyword detector and another neural network for implementing a self-voice detector. In this case, the scheme for selecting an optimized neural network among multiple optimized candidate neural networks for a specific user is the same, even optimized simultaneously, for example. In this case, there are more than two groups, for example, the first and second groups of corresponding multiple optimized candidate networks, and each neural network in each group is optimized to implement a specific detector (for example, implementing a keyword detector and a self-voice detector respectively), and each candidate network in the first and second groups receives the same input vector from the user (from the hearing device) simultaneously (see Figure 2 , for a single group of candidate networks). When referring to multiple groups of networks representing different detectors, the first and second groups do not necessarily have the same input vector. A group (candidate) of NNs (each NN implementing a given type of OV detector) may have input features different from those of a group of NNs for wake-word detection.

[0149] A neural network can be regarded as a generalized non-linear function of input signals optimized for an output signal to achieve a certain behavior. While transmitting a signal through a neural network has a relatively low complexity, estimating the parameters of the neural network is computationally very burdensome and (very) time-consuming. To personalize a neural network, the neural network needs to be trained based on data from a specific person. Even if the network can be partially trained only for an individual, for example, using known transfer learning techniques, the training process and data collection are still very time-consuming. Regarding transfer learning, the hearing device can be configured to update only a part of the neural network parameters (for example, the parameters of the p last layers). Thereby, fewer parameters in the hearing instrument need to be updated.

[0150] As an alternative to optimizing a neural network for an individual, we propose to select among multiple pre-trained neural networks, where each pre-trained network is regarded as a candidate network for that person. Each pre-trained neural network can be trained based on data from a certain class of people. The number of pre-trained neural networks depends on how the actual classification of the population classes is carried out.

[0151] An exemplary neural network including input and output vectors is schematically shown in Figure 7A , 7B . However, it should be emphasized that other NN structures can also be used.

[0152] Consider Figure 2 the system shown in Figure 2Shows a hearing instrument wirelessly connected to an external device, which illustrates a selection procedure for an optimized neural network used in a hearing instrument according to the present invention. While the hearing instrument HD (due to size limitations) has limited memory and computing power, the external device ExD has much more computing power and available memory. This enables the external device to quickly evaluate (possibly pre-processed) hearing instrument sounds (in the form of feature vectors FV), which are provided as inputs FV' to several neural networks (NN 1 , …, NN K ), which neural networks have been pre-trained, for example, for different populations but implement the same detector such as a self-voice detector. Based on the sound examples FV' from the hearing device HD, see, for example, Figure 3 , different neural networks (NN 1 , …, NN K ) are evaluated. The parameters of the network with the best performance are used in the hearing instrument. The evaluation of the relative performance of the neural networks can be carried out, for example, in terms of a comparison of the numbers of true, false positive, true rejection, and false rejection. In the example of Figure 2 , the outputs of the different neural networks are the corresponding probabilities p i (OV), i = 1, …, K. These probabilities are compared, for example, with the true values (i.e., whether the feature vector represents a sound spoken by the user) to obtain the network that performs best for a given person. In another example, the neural network can be optimized to detect Q predefined keywords such as a voice interface. In this case, the output vector of each neural network will represent the probability that a given input feature vector (e.g., a word spoken by the user and picked up by the hearing aid microphone) is equal to each predefined keyword (the output vector includes p(KWi), i = 1, …, Q). The number Q of keywords can be, for example, in the range between 1 and 10. A wake word detector (Q = 1) can be implemented to detect a single keyword or key phrase, such as "Hey Siri" or "Open sesame", etc.

[0153] The system consists of a hearing device HD capable of wirelessly transmitting an audio signal or a feature vector FV derived from the audio signal (via a wireless link WLNK such as an audio link) to an external device ExD. The external device ExD (such as a smart phone or a PC used during the fitting of the hearing device) has much more memory and computing power than the hearing instrument, and it is capable of evaluating different pre-trained neural networks for neural network parameters to select a set of parameters that are most suitable for the user of the hearing device. Based on different audio examples, the external device can select the best candidate network, and the hearing device will be updated with the parameters of the best candidate network.

[0154] To save computation and transmit as little data as possible, the audio recorded at the hearing device, possibly together with sensor data such as data obtained from an accelerometer, can be preprocessed (see unit Pre-PRO) into a feature vector FV, which is used as input to the neural network. An example of a feature vector can be a time-frequency representation of the audio signal (obtained by a filter bank or a warped filter bank). The time-frequency representation can be further processed into a magnitude response, and the magnitude response can be low-pass filtered and / or downsampled. Different microphone signals can be combined into a directional microphone signal, for example a directional signal that enhances the own voice. The directional signal can be further enhanced by noise reduction using a postfilter.

[0155] In the case of relatively low complexity, the hearing device may be configured to retain candidate neural networks (node ​​parameters optimized for which may be stored in the hearing device before fitting for a specific user). Thus, the selection and installation of the candidate neural network that best suits the needs of the user may be performed entirely by the hearing device itself (the hearing device does not have to be connected to an external device, and transceiver circuitry may be omitted).

[0156] The feature vector FV may depend on the application. The feature vector FV may for example be or include the complex valued output from the filter bank, or simply the magnitude (or the square of the magnitude) of the filter bank output. Alternatively or additionally the feature vector may be cepstral coefficients such as Mel Frequency Cepstral Coefficients (MFCC) or Bark Frequency Cepstral Coefficients (BFCC). In the case of self-voice detection, the feature vector FV may contain information about the transfer function between different microphone signals.

[0157] Figure 3 An exemplary personalization procedure of hearing device parameters according to the present invention is shown.

[0158] Figure 3 An example of how the proposed training procedure can be used is shown. Via an external device ExD, the user is prompted to repeat an audio sequence such as a specific word (e.g., for keyword detection, the audio sequence may consist of keywords and confusion words). The audio sequence, here the word "cheese" (or features derived therefrom) is picked up by the hearing device (HD1, HD2) and passed to the external device ExD (possibly as a pre-processed feature vector). The data is applied to K different pre-trained neural networks (NNs). 1 ,…,NN K) For each prompted input, the external device evaluates each candidate neural network, such as the probability of the correct word (here p(cheese)), the probability of error detection (here p(tease)), the receiver operating curve, or other performance metrics. Based on the evaluated words (such as all prompted words), the network with the best performance is selected and the parameters of that network are programmed into the hearing instrument. To evaluate in different noisy environments, different types of background noise can be added to the sound recorded, for example, in the external device. Figure 4 A flowchart showing a method for selecting optimized parameters of a neural network.

[0159] Figure 4 A flowchart showing an embodiment of the proposed method for selecting a personalized neural network. The method includes the steps:

[0160] S1: Start the personalization program;

[0161] S2: Prompt words;

[0162] S3: Transmit the words spoken by the user and the words / features to the external device;

[0163] S4: Calculate and update the performance of each pre-trained neural network (see Figure 2 , 3 the NN in 1 , …, NN K );

[0164] S5: End? If no, return to step S2; if yes, proceed to the next step;

[0165] S6: Update the hearing device with the parameters of the network with the best performance (see transmitting the parameters of the neural network NNx to the hearing device and applying them to the neural network NN* in Figure 2 ).

[0166] The user can also be prompted with other words, such as typical confusing words. For example, peace-cheese, or be prompted to read text that does not contain the desired word.

[0167] In the case of self-voice (OV) detection, different networks trained for multiple groups with similar OV transfer functions between microphones can be imagined. Given the OV transfer function (TRF) measured for an individual, the gap between the measured OV TRF and the OV TRF representing each neural network can be measured. The neural network represented by the OV TRF with the highest similarity can then be selected for use in the hearing device. Alternatively, the similarity between the measured OV TRF and the OV TRFs representing different neural networks can be measured based on the neural network that provides the best OV detection.

[0168] Figure 5 illustrates a hearing device according to an embodiment of the present invention, which uses a trained (user-specific personalized) neural network to control the processing of a signal representing sound in the hearing device before the processed signal is presented to the user wearing the hearing device. The hearing device HD includes an input-detector-decision module, as Figure 1 shown. The output RES of the decision unit Post-Pro is fed to the processor PRO of the hearing device. The processor PRO receives electrical input signals (IN1, IN2) from the microphones (M1, M2) and processes these signals according to the output RES of the decision unit Post-Pro. The output RES of the decision unit Post-Pro can represent, for example, a self-voice detection control signal, a specific wake-up word or (command) keyword of a voice control interface, etc. On this basis, the processor PRO provides a processed output OUT, which is fed to an output transducer, here a loudspeaker SPK, and thus presented to the user of the hearing device. Thus, for example, a hearing device including a voice control interface (or simply, a wake-up word detector) can be implemented. The processor PRO may also include an NN-based detector (such as an OV detector). The output of the OV detector can be an input feature of other detectors such as a wake-up word detector. In Figure 5 an embodiment, the output RES of the decision unit Post-Pro is fed to the processor PRO of the hearing device. As an alternative or in addition, it can be fed to another functional part of the hearing device, for example, a voice interface for controlling the functions of the hearing device based on the recognition of multiple command words. The output of the decision unit can be, for example, a command word or sentence for activating the voice control interface or a wake-up word or sentence. As an alternative or in addition, the output RES of the decision unit Post-Pro can be fed to a transmitter (or transceiver), and thus transmitted to another device and further processed there, for example, to activate a personal assistant (such as a smart phone, etc.). The transceiver can receive, for example, a response from another device such as from the personal assistant. This response can be used to control the hearing device, or it can be played to the user via the output transducer SPK of the hearing device.

[0169] Figure 6 illustrates an exemplary procedure for dividing a basic population for providing training data for training a neural network into multiple subgroups for training multiple neural networks, so as to provide multiple optimized neural networks, each optimized neural network representing different characteristics of the test subjects. Figure 6 illustrates how the training of a neural network can be divided into different networks, each network being trained based on a subset of the data set. Thus, multiple groups with similar speech / voice characteristics can be provided. Therefore, on this basis, a corresponding number of multiple trained networks NN 1 ,…NN K .

[0170] Multiple sets of candidate networks can be generated through an iterative process. Starting with training a single NN, those with the worst NN performance are grouped, and another NN is trained for these (and the first NN is trained for the first group of people). Alternatively, the grouping can be based on age, gender, pitch, or other ways of measuring the similarity that differentiates between different speakers. Thus, it is possible that new individuals (who are not part of the training data) will perform well based on at least one trained neural network. One advantage is that the size of the neural network can remain small because the network does not have to be universal for all people.

[0171] Figure 7A An embodiment of a keyword detector implemented as a neural network according to the present invention is shown. Figure 7B An audio electrical input signal context including an input vector for generating Figure 7A the neural network is shown.

[0172] Figure 7A An embodiment of a keyword detection detector implemented as a neural network according to the present invention is shown. Figure 7A A deep neural network (DNN) is schematically shown for determining the occurrence probability p(KWq, l) of a specific keyword KWq, q = 1, …, Q at a given time point (l’) from an input vector of an electrical input signal including time-frequency representations (k, l) or its specific features (= feature vector, FV) of L time frames X(k, l), l = l’-(L - 1), …, l’, where k is the frequency index and l is the time (frame) index. The electrical input signal or its specific features (such as cepstral coefficients or spectral features, etc.) at the current time l = l’ are referred to as “feature vector FV” in Figure 1 and 2 and denoted as X(k, l’) in Figure 7A and 7B . The L (last) time frames (X(k, l)) of the input signal constitute an exemplary input vector of the neural network at the given time point l = l’ and are denoted as Figure 7A and 7B in . This “context” included in each input vector is shown in Figure 7B . The keyword detection detector can be configured such that only the parameters of the last q layers (among the NN candidate networks) are different.

[0173] The current time frame (l’) and L - 1 previous time frames are stacked as a vector and used as the input layer in the neural network (collectively denoted as See also Figure 7Bdenoted as the shaded time-frequency unit "context"). Each time frame X(k, l’) includes K values (e.g., K = 16 or K = 24 or K = 64 or K = 128) of the electrical input signal (or features extracted therefrom). The signal can be represented by its magnitude |X(k, l’)| (e.g., by ignoring its phase φ), see Figure 7B . Alternatively, the input vector can include time samples of the input signal (in the time domain) covering an appropriate time period. The appropriate number of time frames is related to the correlation inherent in speech. In an embodiment, the L - 1 previous time frames considered together with the current time frame l = l’ can, for example, correspond to a time period of more than 20 ms, such as more than 50 ms, such as more than 100 ms, such as approximately 500 ms. In an embodiment, the number of time frames considered (= L) is greater than or equal to 4, such as greater than or equal to 10, such as greater than or equal to 24, for example in the range of 10 - 100. In the present application, the width of the neural network is equal to K·L, for K = 64 and L = 10, meaning that the input layer L1 consists of N L1 = 640 nodes (representing a 32 - ms time period of the audio input signal (for a sampling frequency of 20 kHz and 64 samples per frame, and assuming no overlap of time frames)). The number of nodes (N L2 , …, N LN ) in the subsequent layers (L2, …, LN) can be greater than or less than the number of nodes N L1 of the input layer L1, and generally, is adapted to the corresponding application (considering the available number of the input data set and the number of parameters that the neural network will estimate). For applications in portable hearing devices with limited power and space, the subsequent layers (N L2 , …, N LN ) can preferably include fewer (such as significantly fewer) nodes, for example, nodes on the order of the number of output nodes. In this example, the number of nodes N LNis Q (such as ≤20, or 10 or less) as it includes Q values p(KWq,l’) (q = 1, …, Q) of a probability estimator, one value for each of the Q keywords of the speech interface. Optionally, the output layer may include Q+1 or Q+2 nodes by including a value for the detection of the user's own speech and / or for the detection of a "filler" (no keyword). In an embodiment, whenever the filter bank of the hearing device provides a new time frame of the input signal, the neural network is fed a new input feature vector (i.e., in this case, there will be a certain time frame overlap between the input vectors). However, to reduce computational complexity (and power consumption), the neural network may be executed less frequently than once per time frame, such as once every 10 time frames, or less than once every 20 time frames (such as less than once every 20 ms or less than once every 40 ms). However, preferably, the context (input feature vector) fed to the neural network at a given time point overlaps (in time) with the previous context. In an embodiment, the number of time frames ΔL between each new execution of the neural network is less than the number of time frames L in the input feature vector (ΔL < L, e.g., ΔL / L ≤ 0.5) to ensure context overlap. As an alternative to stacking time frames, a recurrent network structure (such as an LSTM or GRU network) may be utilized. Thereby, the input layer can be significantly smaller.

[0174] Figure 7A For illustrating any type of general multi-layer neural network, such as a deep neural network, which is embodied herein as a standard feed-forward neural network. The depth (number of layers) of the neural network is denoted as N in Figure 7A and can be any number, typically adapted to the application involved (e.g., limited by the size and / or power capacity of the device involved such as a portable device like a hearing aid). In an embodiment, the number of layers in the neural network is greater than or equal to 2 or 3. In an embodiment, the number of layers in the neural network is less than or equal to 10, such as in the range of 2 to 8 or in the range of 2 to 6.

[0175] Figure 7A The nodes of the neural network illustrated in v,u are used to implement the standard function of the neural network such that the values of the branches from the previous nodes to the node involved are multiplied by the weights associated with the respective branches and the contributions are added together as the sum value Y’ of node v in layer u v,u . The sum value Y’ uv then undergoes a non-linear function f, thereby providing the composite value Z of node v in layer u v,u = f(Y’ Figure 7A ). This value is fed to the next layer (u+1) via the branches connecting node v in layer u to the nodes in layer u+1. In v,u , the sum value Y’

[0176]

[0177] where w p,v (u) refers to the weight of the node v in layer u that is applied to the input from node p in layer L(u - 1), and Z p,v (u - 1) is the signal value of the p-th node in layer u - 1. The same activation function f is used for all nodes (although this is not necessarily the case). The non-linear function can be parameterized, and one or more parameters of the non-linear function can be included in the optimization of the node parameters. Additionally, the bias parameter b p,v can be associated with each node and participate in the optimization of the node parameters. An exemplary non-linear activation function Z = f(Y) is schematically shown in the illustration in Figure 7A . Typical functions used in neural networks are the rectified linear unit (ReLu), the hyperbolic tangent function (tanh), the sigmoid or softmax functions. However, other functions can also be used. As indicated, activation functions such as the ReLu function can be parameterized (e.g., to allow for different slopes).

[0178] The (possibly parameterized) activation functions f, the weights w, and the bias parameters b of the different layers of the neural network together constitute the parameters of the neural network. They represent the parameters that are (together) optimized in the corresponding iterative procedure of the neural network of the present invention. The same activation function f can be used for all nodes (such that the "parameters of the neural network" are constituted by the weights and bias parameters of each layer). In an embodiment, for at least some nodes of the neural network, the activation function f is not used.

[0179] Generally, a candidate neural network according to the present invention is optimized (trained) in an offline procedure, for example, using a model of a human head and torso (such as the Head and Torso Simulator (HATS) 4128C from Brüel& Sound&Vibration Measurement A / S), where the HATS model is "equipped" with a hearing device (or a pair of hearing devices) of the same type (style) as the hearing device that the user plans to use. The hearing device is configured to pick up (acoustically propagated) training data when located at the ears of the model (just as during normal use of the hearing device by the user). (For example, according to the Figure 6 scheme) Define multiple different categories of testers or based on tester parameters such as age, gender, weight / height ratio, occupation, "body type", etc., and optimize N x different neural networks based on the training data of (N x different) groups of people. Ideally, training data related to the user's normal behavior and acoustic environment experience should be used.

[0180] In the case of training different networks based on personalized acoustic characteristics, it is advisable to record the acoustic characteristics from different persons. The acoustic characteristics of a person can be obtained, for example, as described in [Moore et al., 2019].

[0181] For keyword detector applications, self-voice detection can be advantageously used to define where to look for keywords in the user's sentence. Thus, the self-voice detection signal can be used as an input to a pre-processor ( Figure 1 , 2 , Pre-PRO in 5) to define the electrical input signals (IN1, IN2) from the microphones (M1, M2). As an alternative, the self-voice detection signal can form part of a feature vector used as an input to a neural network. This can be advantageous since the user is unlikely to plan to trigger a keyword (such as a wake word or command word) in the middle of a sentence. The use of a self-voice presence indicator can enable keyword detection only at the start of a sentence. For example, a rule can be imposed that a keyword can be (effectively) detected only if no self-voice has been detected in the previous 0.5 seconds or the previous 1 second or the previous 2 seconds (but self-voice is detected "now").

[0182] In Figure 7A , the neural network is illustrated as a feed-forward network, but other neural network configurations can also be used, such as convolutional networks (CNNs), recurrent networks, or combinations thereof.

[0183] Figure 8 An embodiment of a hearing device according to the invention is shown, which includes an adaptive unit configured to enable the selection of a set of optimized parameters for a neural network among multiple sets of optimized parameters without using an external device. Figure 8 A self-contained hearing device HD such as a hearing aid according to the invention is shown, which includes an optimized neural network NN*, for example, for implementing processing (see control signal RES to signal processor PRO) detector DET that affects the hearing device. It can implement the same functions as shown and described in Figure 2 , but does not require an external device ( Figure 2 ExD in Figure 2 ) for the neural network node parameter optimization procedure (and thus does not require a wireless link to an external device ( 1 WLNK in 2 K K ). Optimized candidate neural networks (NN Figure 8In an embodiment, two microphones (M1, M2) are shown, each microphone providing an electrical input signal representative of sound in the environment. Other numbers of input transducers such as microphones may also be used, for example one, three or more. Input transducers (or any other body-worn microphone) from an auxiliary device such as a hearing device worn on the opposite ear may also provide input features to the neural network. A beamformer may be included in the pre-processor Pre-PRO to enable generation of a directional signal based on electrical input signals (IN1, IN2) from more than two microphones. The directional signal may for example be or include an estimate of the user's own voice (e.g., generated by an own voice beamformer directed towards the user's mouth). The beamformed signal (or its characteristic features) may be the signal fed to a neural network for implementing a detector such as an own voice detector or a keyword detector (see feature vector FV). The non-essential sensor SENSE that provides the sensor control signal SIN to the pre-processor Pre-PRO may or may not form part of the hearing device HD. The sensor may for example be a motion sensor, for example including an acceleration or gyroscope sensor. Other sensors may for example be or include a magnetometer, an electroencephalogram (EEG) sensor, a magnetoencephalogram (MEG) sensor, a heart rate detector, a photoplethysmogram (PPG) sensor, etc. The electrical input signals (IN1, IN2) are fed to the processor PRO. The processor PRO may be configured to apply a gain that varies with frequency and / or level to the electrical input signal (or its processed version, for example its spatially filtered (beamformed) version). The processor PRO provides a processed output signal OUT, which is fed to an output transducer, here a loudspeaker SPK, for presentation to the user of the hearing device.

[0184] When appropriately replaced by corresponding processes, the structural features of the apparatus described above, detailed in the "Detailed Description" and defined in the claims, may be combined with the steps of the method of the present invention.

[0185] Unless explicitly stated otherwise, the singular forms "a", "the" as used herein are intended to include the plural forms (i.e., having the meaning of "at least one"). It should be further understood that the terms "having", "including" and / or "comprising" as used in the specification indicate the presence of the stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof. It should be understood that unless explicitly stated otherwise, when an element is referred to as being "connected" or "coupled" to another element, it can be directly connected or coupled to the other element, or intervening elements may be present. As used herein, the term "and / or" includes any and all combinations of one or more of the listed related items. Unless explicitly stated otherwise, the steps of any method disclosed herein do not have to be performed in the exact order disclosed.

[0186] It should be appreciated that references to "one embodiment" or "an embodiment" or "an aspect" or "may" included in this specification mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Additionally, the particular features, structures, or characteristics may be appropriately combined in one or more embodiments of the present invention. The foregoing description is provided to enable those skilled in the art to practice the various aspects described herein. Various modifications will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects.

[0187] The claims are not limited to the aspects shown herein, but encompass the full scope consistent with the claim language, where elements recited in the singular are not meant to mean "one and only one" unless explicitly stated, but rather "one or more." Unless explicitly stated, the term "some" means one or more.

[0188] Accordingly, the scope of the present invention should be determined in accordance with the claims.

[0189] References

[0190] [Moore et al.,2019]Moore,A.H.,de Haan,J.M.,Pedersen,M.S.,Naylor,P.A.,Brookes,M.,&Jensen,J.(2019).Personalized signal-independent beamforming forbinaural hearing aids.The Journal of the Acoustical Society of America,145(5),2971-2981.

Claims

1. A hearing device configured to be located at or in a user's ear or to be fully or partially implanted in a user's head, the hearing device comprising: - an input transducer including at least one microphone for providing at least one electrical input signal representative of sound in the environment of the hearing device; - a pre-processor for processing the at least one electrical input signal and providing a plurality of feature vectors, each feature vector representing a time period of the at least one electrical input signal; - a neural network processor adapted to implement a neural network for implementing a detector or a part thereof, configured to provide an output indicating a characteristic feature of the at least one electrical input signal, the neural network including an input layer, an output layer and a plurality of hidden layers, each layer including a plurality of nodes, each node defined by a plurality of node parameters, the neural network configured to receive the plurality of feature vectors as input vectors and to provide corresponding output vectors representing the output of the detector or a part thereof according to the input vectors; - a post-processor configured to receive the output vectors, and wherein the post-processor is configured to process the output vectors and provide a resulting signal; - a neural network controller connected to the neural network processor, which is adapted to receive optimized node parameters in an adaptive operating mode and to apply the optimized node parameters to the nodes of the neural network so as to implement an optimized neural network in the neural network processor, the optimized node parameters being stored in the hearing device.

2. The hearing device according to claim 1, comprising a sensor for sensing a characteristic of the user or the environment of the hearing device and providing a sensor signal representative of a current value of the environmental characteristic, wherein the sensor signal is an input to the pre-processor.

3. The hearing device according to claim 2, wherein the pre-processor is configured to process the at least one electrical input signal and the sensor signal to provide feature vectors.

4. The hearing device according to claim 1, comprising an output transducer for presenting the processed output signal as a perceptible sound stimulus to the user.

5. The hearing device according to claim 1, comprising an analysis filter bank for converting a time-domain input signal into a plurality of sub-band signals so as to provide an input signal having a time-frequency representation (k, l), where k and l are frequency and time indices respectively.

6. The hearing device according to claim 1, wherein the pre-processor is configured to extract features of the at least one electrical input signal and / or the sensor signal.

7. The hearing device according to claim 2, wherein the time period of the at least one electrical input signal covered by a given feature vector and optionally the corresponding value of the sensor signal are used as inputs to the input layer of the neural network, which includes at least one time frame of the at least one electrical input signal.

8. The hearing device according to claim 1, wherein the detector or a part thereof implemented by the neural network is or includes a self-voice detector and / or a keyword detector.

9. The hearing device according to claim 1, being constituted by or including a hearing aid, a headset, an earphone, an ear protection device or a combination thereof.

10. The hearing device according to claim 1, wherein the predetermined criterion is related to minimizing a cost function with respect to the output vector.

11. The hearing device according to claim 1, wherein, when the plurality of feature vectors are extracted from a time period of at least one electrical input signal having known characteristics, the predetermined criterion is based on the performance of the neural network in terms of true, false positive, true rejection, and false rejection of the output vector.

12. The hearing device according to claim 1, wherein, multiple sets of node parameters for corresponding candidate neural networks are optimized for different classes of people exhibiting different acoustic properties.

13. A method for selecting optimized parameters of a neural network used in a portable hearing device, the method comprising: - providing a portable hearing device to be used by a specific user, the hearing device including a neural network processor adapted to implement a neural network including an input layer, an output layer, and a plurality of hidden layers, each layer including a plurality of nodes, each node being defined by a plurality of node parameters and a non-linear function, the neural network being configured to receive an input vector and provide an output vector that is a specific non-linear function of the input vector; - installing the hearing device on the user; - providing at least one electrical input signal representing sounds in the user's environment while wearing the hearing device; - processing the at least one electrical input signal and providing a plurality of feature vectors, each feature vector representing a time period of the at least one electrical input signal; - providing a plurality of pre-trained candidate neural networks, wherein each pre-trained network is considered a candidate network for the user, and wherein each pre-trained neural network has been trained based on completely or partially different training data; - receiving the feature vectors and providing them as input vectors to the plurality of pre-trained candidate neural networks; - determining, by the plurality of pre-trained candidate neural networks, corresponding output vectors corresponding to the feature vectors; - comparing the output vectors and selecting, according to a predetermined criterion with respect to the output vectors, one of the plurality of candidate neural networks as the neural network optimized for the hearing device; - transmitting the node parameters of the selected neural network among the plurality of candidate neural networks to the neural network processor of the hearing device, and - receiving the node parameters into the neural network processor and applying them to the neural network.

14. The method according to claim 13, wherein each candidate neural network has been trained based on training data from different classes of people exhibiting different acoustic properties.

15. The method according to claim 13, wherein a signal representing a current value of a characteristic of the user or the user's environment is provided in the hearing device together with the at least one electrical input signal and processed to provide the feature vectors.

16. The method according to claim 13, further comprising: Provide several sets of multiple pre-trained candidate neural networks, each candidate neural network in each set having the same structure as the neural network of the hearing device, where each pre-trained network is considered a candidate network for the user, and where each pre-trained neural network has been trained based on completely or partially different training data, and where each set of pre-trained candidate neural networks is aimed at implementing different detectors.

17. Use of the hearing device according to claim 1.

18. A computer program product having stored thereon a computer program comprising instructions which, when executed by a computer, cause the computer to perform the method according to claim 13.