Portable device comprising a directional system

CN113543003BActive Publication Date: 2026-09-22OTICON
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202110437844.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-22
Filing Date
2021-04-22
Publication Date
2026-09-22
Estimated Expiration
2041-04-22

Smart Images

  • Figure CN113543003B_ABST
    Figure CN113543003B_ABST
Patent Text Reader

Abstract

A portable device comprising a directional system is disclosed, which is a sound capturing device comprising an input unit comprising a plurality of input transducers, a housing, a directional noise reduction system for providing an estimate of a target sound, the directional noise reduction system comprising a beamformer unit connected to the plurality of input transducers in operation, the beamformer unit comprising a target holding, a reference beamformer and a target cancellation beamformer, the directional noise reduction system being configured to operate in at least a directional mode and an omnidirectional mode depending on a mode control signal, an antenna and a transceiver circuit for establishing an audio link to another device, wherein the sound capturing device is configured to communicate the estimate of the target sound to the other device, and a mode controller for determining the mode control signal depending on a current reference signal and a current target cancellation signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a sound capture device configured to pick up sound from the environment and transmit the processed sound to a hearing device such as a hearing aid or to another device or system. Background Technology

[0002] The sound capture device (and hearing device) can be configured to be worn by the user of the hearing device or another person. In various situations, for example:

[0003] a) The voice capture device may be worn by a user of a hearing device and configured to pick up the user's own voice and transmit it to another device, such as a telephone or any other communication device or system; or

[0004] b) The sound capture device may be configured to be worn by a person communicating with the user of the hearing device and to transmit that person's voice to the hearing device; or

[0005] c) The sound capture device may be held on a bracket such as a table and configured to pick up sound from its environment, such as the sounds of multiple (e.g., more than two) people and transmit the sound to a hearing device and / or another device or system such as a communication device.

[0006] US8391522B2 proposes using an accelerometer to alter the processing of an external microphone array. US7912237B2 proposes using a direction sensor to switch between omnidirectional and directional processing of an external microphone array. Summary of the Invention

[0007] The present invention includes a scheme for adjusting signal processing in a sound capture device based on an estimated directional performance of the microphone of the sound capture device, such as a scheme for changing the signal processing mode, for example, changing between a directional operation mode and an omnidirectional operation mode of the sound capture device.

[0008] The present invention also relates to detecting user self-voice in a sound capture device, such as a hearing device like a hearing aid, based on the estimated directional performance of the microphone of the sound capture device.

[0009] Sound capture device

[0010] In one aspect of this application, a sound capturing device is provided, configured to be worn by a person and / or placed on a surface such as a table. The sound capturing device is configured to pick up target sound from a target sound source. The sound capturing device may include:

[0011] - Contains multiple input converters IT mInput units of m = 1, 2, ..., M, where M is greater than or equal to 2, each input converter is configured to pick up sound from the environment of the sound capture device and provide a corresponding electrical input signal, each electrical input signal IN m ,m=1,…,M includes the target signal component and the noise signal component;

[0012] - A housing in which the plurality of input transducers are located, and which may include a preferred orientation;

[0013] - A directional noise reduction system for providing an estimate of a target sound s, the directional noise reduction system comprising components connected to the plurality of input converters IT during operation. m Beamformer units of m = 1, ..., M, wherein the beamformer unit may include

[0014] --Target holding and reference beamformer, configured to maintain signal components from a fixed target direction with minimal or no attenuation relative to signal components from other directions, and provide a current reference signal; and

[0015] --Target cancellation beamformer, configured to attenuate signal components from the target direction, while signal components from other directions are attenuated less than those from the target direction, and provide the current target cancellation signal.

[0016] The directional noise reduction system can be configured to operate in at least two modes based on a mode control signal:

[0017] - Directional mode, where the estimate of the target sound s is based on the target signal component from a fixed target direction; and

[0018] - Non-directional, omnidirectional mode, where the estimate of the target sound s is based on the target signal components from all directions.

[0019] The sound capture device may also include:

[0020] - Antenna and transceiver circuitry for establishing an audio link to another device, and a sound capture device configured to transmit an estimate of the target sound s to the other device.

[0021] The sound capture device may also include a mode controller for determining a mode control signal based on the current reference signal and the current target cancellation signal.

[0022] This increases the flexibility of using the sound capture device.

[0023] The target orientation of the beamformer can be aligned with the preferred orientation of the housing of the sound capture device (or can be known or estimated before the sound capture device is used). Multiple input transducers may comprise a microphone array. Preferably, the target orientation is the emission direction of the end of the microphone array, i.e., a direction parallel to the microphone array. The microphone orientation can be determined by a direction passing through the center of the microphone. The microphone array can be a linear array, where (two or more) microphones are located on a straight line (microphone orientation).

[0024] In one embodiment, the self-voice beamformer is calibrated for a preferred placement of the sound capture device on a person, such as directing the housing toward the person's mouth. The calibration process can occur in a specific calibration mode. Alternatively, calibration can occur during use, such as when detecting self-voice.

[0025] Target-keeping beamformers can be substantially omnidirectional beamformers (see, for example, see...) Figure 2A ). Target-holding beamformers can have frequency-varying attenuation (see, for example, see...). Figure 2D ).

[0026] The greatest difference between a target-keeping beamformer and a target-cancelling beamformer reflects the presence of a person wearing a sound capture device (or reflects the microphone orientation being aligned with the direction towards the current speaker, for example, when the sound capture device is located on a surface near the current speaker).

[0027] The directional noise reduction system can be configured to switch between omnidirectional and directional modes based on a mode control signal.

[0028] At least one of the input transducers may be a microphone. Most or all of the input transducers may be microphones. Multiple input transducers may consist of or include two microphones. Multiple input transducers may include a microphone array. Multiple input transducers may include MEMS microphones.

[0029] The sound capture device may include filter banks. The filter banks may be configured to enable processing within the sound capture device in the filter bank domain (frequency domain) by providing time-domain input signals in multiple sub-bands, for example, by providing multiple (K) frequency windows (k = 1, ..., K) for consecutive time intervals l, each frequency window being determined by a corresponding frequency and a time frame index (k, l). The input unit of the sound capture device may, for example, include multiple (M) analysis filter banks, each analysis filter bank connected to a different input converter among M input converters and configured to provide each of the M electrical input signals in a sub-band / time-frequency representation (k, l).

[0030] The magnitudes or processed versions of the corresponding current reference signal and current target cancellation signal can be averaged over time to provide corresponding smoothed reference and target cancellation metrics. The magnitudes (or squared magnitudes) of the current reference signal (ref(k,l)) and the current target cancellation signal (TC(k,l)) can be provided through the corresponding magnitude (or squared magnitude) calculations (see [link to relevant documentation]). Figure 3 In |ref|,(|ref| 2 ) and |TC|,(|TC| 2 The corresponding processed versions of the current reference signal and the current target cancellation signal may include, for example, a) the product of the (possibly complex) value of the current reference signal (its complex conjugate) and the current target cancellation signal (ref). * TC); and b) the square of the magnitude of the current target cancellation signal (|TC|). 2 (For example, see...) Figure 4 ).

[0031] The sound capture device may include a voice activity detector. The sound capture device may be configured such that only time-frame averaging occurs when the voice activity detector detects user voice. Voice can be detected using a voice activity detector, such as a modulation-based voice activity detector. The voice activity detector may be configured to estimate the probability (or binary value) of voice presence in separate sub-bands (e.g., in each frequency window). The smoothing magnitudes of the reference beamformer (see OMNI-BF) and the target voice cancellation beamformer (see TC-BF) may be converted to the logarithmic domain (see...). Figure 3 (The unit "log" in the text).

[0032] The sound capture device may include a combined processor configured to compare a current reference signal and a current target cancellation signal or a processed version thereof in different sub-bands and provide a corresponding sub-band comparison signal.

[0033] The sound capture device may include a decision controller configured to provide mode control signals, specifying the appropriate operating mode of the directional noise reduction system, based on sub-band comparison signals. Differences found in separate sub-bands (see...) Figure 3 The SUM cell "+" or Figure 4 The DIV unit “÷” in the text is combined across frequencies for joint decision-making (see [link]). Figure 3 , 4 (The "Decision" module within the code). The decision controller can be implemented, for example, through logical processing such as weighted summation; or through logical recursion or neural networks. Weights can be estimated based on supervised learning. Alternatively, the combination function can be manually adjusted.

[0034] The decision controller can be configured to provide a mode control signal based on a weighted sum of the comparison signals for each sub-band. When the mode control signal exhibits a first (e.g., relatively large) value across frequencies, indicating a first (relatively large) difference between the current reference signal and the current target cancellation signal or its processed version, this indicates a significant benefit of directional noise reduction, and the directional noise reduction system should be switched to (or maintained) in directional mode. Otherwise, if the mode control signal exhibits a second (e.g., relatively small) value, indicating a considerably small (e.g., less than 3 dB, less than 6 dB, or less than 9 dB) difference, the potential benefit of directional noise reduction is limited, and the directional noise reduction system should be switched to (or maintained) in omnidirectional mode. The first difference is assumed to be greater than the second difference. The directional mode can be adaptive (e.g., its noise reduction is adaptive) or fixed. The mode control signal can be a binary signal (e.g., 0 or 1). The mode control signal can also be a continuous signal (e.g., exhibiting values ​​in the interval [0,1]), and the directional noise reduction system can be used to allow for a smooth transition between different directional modes based on the mode control signal.

[0035] When the mode control signal indicates a substantial cross-frequency difference between the current reference signal and the current target cancellation signal or its processed version, the directional noise reduction system can be adapted to directional mode; and when the mode control signal indicates a relatively small cross-frequency difference between the current reference signal and the current target cancellation signal or its processed version, the directional noise reduction system can be adapted to omnidirectional mode. When the mode control signal is less than a first threshold, the directional noise reduction system can be adapted to omnidirectional mode. When the mode control signal is greater than a second threshold, the directional noise reduction system can be adapted to directional mode. When the mode control signal is between the first and second thresholds, the directional noise reduction system can be adapted to a mode between omnidirectional and directional modes.

[0036] The sound capture device may be constituted by or include a microphone device. For example, the sound capture device may be constituted by a dedicated wireless microphone device. The sound capture device may also be constituted by, or be part of, a hearing device such as a hearing aid or headphones.

[0037] On the other hand, a sound capturing device, such as a hearing device like a hearing aid, configured for wear by a user is provided. The sound capturing device includes:

[0038] - Contains multiple input converters IT m Input units of m = 1, 2, ..., M, where M is greater than or equal to 2, each input converter is configured to pick up sound from the environment of the sound capture device and provide a corresponding electrical input signal, each electrical input signal IN m ,m=1,…,M includes the target signal from the target signal source and the noise signal from one or more noise signal sources;

[0039] - Self-voice detector, configured to provide a voice control signal indicating whether or with what probability a given electrical input signal or its processed version originates from user voice.

[0040] Self-voice detectors may include:

[0041] - Connect to the multiple input converters IT during operation m Beamformer units of m = 1, ..., M, the beamformer unit includes

[0042] --Target holding and reference beamformer, configured to maintain signal components from a fixed target direction with minimal or no attenuation relative to signal components from other directions, and provide a current reference signal; and

[0043] --Target cancellation beamformer, configured to attenuate signal components from the target direction, while signal components from other directions are attenuated less than those from the target direction, and provide a current target cancellation signal;

[0044] The fixed target direction is the direction from which the sound capturing device faces the user's mouth, and the target signal is the user's own voice; and

[0045] - A controller used to determine the self-voice control signal based on the current reference signal and the current target cancellation signal.

[0046] The controller can be configured to determine a self-voice control signal based on a comparison between the current reference signal and the current target cancellation signal.

[0047] The controller can be configured to determine the self-voice control signal based on the values ​​of the reference beamformer and the target cancellation beamformer.

[0048] The target cancellation beamformer (here, the self-voice cancellation beamformer) can update its weights when self-voice is detected. This improves the performance of the self-voice cancellation beamformer (which may vary with distance (due to near field) and tilt).

[0049] Voice capture devices, such as hearing aids, may include a keyword detector for detecting one of a limited number of keywords among one of a plurality of electrical input signals or a processed version thereof, wherein the keyword detector is activated based on a self-voice control signal. The voice capture device may include a voice control interface that enables functions of the voice capture device, such as a hearing aid. The keyword detector may be connected to the voice control interface. The keyword detector may be configured to detect a wake word used to activate the voice control interface. The keyword detector may be connected to a self-voice detector.

[0050] The sound capture device includes an input unit for providing an electrical input signal representing sound. The input unit includes an input converter, such as a microphone, for converting the input sound into an electrical input signal.

[0051] Sound capture devices may include directional microphone systems adapted to spatially filter sound from the environment, thereby enhancing a target sound source among multiple sound sources in the local environment of a user wearing the sound capture device. The directional system may be adapted to detect (e.g., adaptive detection) the direction from which a specific portion of the microphone signal originates. This can be achieved, for example, in a variety of different ways described in the prior art. In sound capture devices such as hearing aids, microphone array beamformers are commonly used to spatially attenuate background noise sources. Many beamformer variations can be found in the literature, such as the Linearly Constrained Minimum Variance (LCMV) beamformer. A particular variation, the Minimum Variance Distortionless Response (MVDR) beamformer, is widely used in microphone array signal processing. Ideally, an MVDR beamformer keeps the signal from the target direction (also known as the line of sight) unchanged while attenuating sound signals from other directions to the greatest extent possible. A Generalized Sidelobe Canceller (GSC) structure is an equivalent representation of an MVDR beamformer, offering computational and digital representation advantages over a direct implementation of the original form.

[0052] The sound capture device may include an antenna and transceiver circuitry (such as a wireless transceiver or receiver) for wirelessly transmitting or receiving direct electrical input signals to or from another device, such as a communication device or another sound capture device, such as a hearing aid. The direct electrical input signal may represent or include audio signals and / or control signals and / or information signals. Communication between the hearing aid and the other device may be in baseband (audio frequency range, such as between 0 and 20 kHz). Preferably, communication between the sound capture device and the other device is based on a type of modulation at a frequency higher than 100 kHz. Preferably, the frequency used to establish a communication link between the sound capture device and the other device is below 70 GHz, for example, in the range from 50 MHz to 70 GHz, for example, above 300 MHz, for example, in the ISM range above 300 MHz, for example, in the 900 MHz range, or in the 2.4 GHz range, or in the 5.8 GHz range, or in the 60 GHz range (ISM = Industrial, Scientific and Medical, such standardized ranges are defined, for example, by the International Telecommunication Union ITU). The wireless link may be based on standardized or proprietary technologies. Wireless links can be based on Bluetooth technology (such as Bluetooth Low Energy).

[0053] Sound capture devices can have a maximum external size in the 0.15m range (e.g., a handheld mobile phone). Sound capture devices can have a maximum external size in the 0.08m range (e.g., headphones). Sound capture devices can have a maximum external size in the 0.04m range (e.g., hearing aids or hearing instruments).

[0054] The sound capture device can be a portable (i.e., configured as a wearable) device or an integral part thereof, such as a device that includes an intrinsic power source, such as a battery, for example a rechargeable battery. The sound capture device can be a lightweight, easy-to-wear device, for example having a total weight of less than 100g.

[0055] The sound capture device may include a forward or signal path between an input unit (such as an input converter, e.g., a microphone or microphone system and / or a direct electrical input, such as a wireless receiver) and an output unit such as an output converter and / or a transmitter. A signal processor may be located in this forward path. The signal processor may be adapted to provide frequency-varying gain according to the specific needs of the user. The sound capture device may include an analysis path having functionalities for analyzing the input signal (e.g., determining level, modulation, signal type, acoustic feedback estimate, etc.). Some or all of the signal processing of the analysis path and / or signal path may be performed in the frequency domain. Some or all of the signal processing of the analysis path and / or signal path may be performed in the time domain.

[0056] The sound capture device may include an analog-to-digital (AD) converter to digitize analog input (e.g., from an input converter such as a microphone) at a predetermined sampling rate such as 20 kHz. The sound capture device may also include a digital-to-analog (DA) converter to convert the digital signal into an analog output signal, for example, for presentation to a user via an output converter.

[0057] Sound capture devices, such as input units and / or antenna and transceiver circuitry, include a time-frequency (TF) conversion unit for providing a time-frequency representation of the input signal. The time-frequency representation may include an array or mapping of corresponding complex or real values ​​of the signal in question over a specific time and frequency range. The TF conversion unit may include a filter bank for filtering the (time-varying) input signal and providing multiple (time-varying) output signals, each output signal encompassing a distinctly different frequency range of the input signal. The TF conversion unit may include a Fourier transform unit for converting the time-varying input signal into a (time-varying) signal in the (time-frequency) domain. The sound capture device considers a frequency range starting from the minimum frequency f. min up to the maximum frequency f max The frequency range can include a portion of the typical human hearing range from 20Hz to 20kHz, such as a portion of the range from 20Hz to 12kHz. Typically, the sampling rate f... s Greater than or equal to the maximum frequency f max twice that, i.e., f s ≥2f maxThe signals of the forward and / or analysis paths of the sound capture device can be divided into NI (e.g., uniformly wide) frequency bands, where NI is, for example, greater than 5, greater than 10, greater than 50, greater than 100, or greater than 500, and at least some of them are processed individually. The sound capture device can be adapted to process the signals of the forward and / or analysis paths (NP≤NI) on NP different channels. The channels can be of uniform or inconsistent width (e.g., width increases with frequency), overlapping or non-overlapping.

[0058] The sound capture device can be configured to operate in different modes, such as a normal mode and one or more specific modes, which may be user-selectable or automatically selected. Operating modes can be optimized for specific acoustic conditions or environments. Operating modes may include directional and non-directional (e.g., omnidirectional) operating modes of the microphone system. Operating modes may include low-power modes, in which the functionality of the sound capture device is reduced (e.g., for energy saving), such as disabling wireless communication and / or disabling specific features of the sound capture device.

[0059] The sound capture device may include multiple detectors configured to provide status signals relating to the current network environment (e.g., the current acoustic environment) of the sound capture device, and / or the current state of the user wearing the sound capture device, and / or the current state or operating mode of the sound capture device. Alternatively or additionally, one or more detectors may form part of an external device that communicates with the sound capture device (e.g., wirelessly). External devices may include, for example, another sound capture device, a remote control, an audio transmission device, a telephone (e.g., a smartphone), external sensors, the sound capture device, etc.

[0060] One or more of a plurality of detectors can operate on a full-band signal (time domain). One or more of a plurality of detectors can operate on a band-split signal ((time-)frequency domain), for example, in a finite number of frequency bands.

[0061] Multiple detectors may include level detectors for estimating the current level of the signal in the forward path. Detectors may be configured to determine whether the current level of the signal in the forward path is above or below a given (L-) threshold. Level detectors operate on full-band signals (time domain). Level detectors operate on band-split signals ((time-)frequency domain).

[0062] The sound capture device may include a voice activity detector (VAD) for estimating whether (or with what probability) the input signal (at a specific point in time) includes a voice signal. In this specification, the voice signal may include speech signals from humans. It may also include other forms of vocalization produced by the human speech system (such as singing). The voice activity detector unit may be adapted to classify the user's current acoustic environment as a "voice" or "no-voice" environment. This has the advantage that time periods including electro-acoustic signals of human vocalizations (such as speech) in the user's environment can be identified and thus separated from time periods that only (or primarily) include other sound sources (such as artificially generated noise). The voice activity detector may be adapted to also detect the user's own voice as "voice." Alternatively, the voice activity detector may be adapted to exclude the user's own voice from the detection of "voice."

[0063] The sound capture device may include a self-voice detector for estimating whether (or with what probability) a particular input sound (such as speech) originates from the speech of a user of a hearing aid system. The microphone system of the sound capture device may be adapted to distinguish between the user's own speech and the speech of another person, and possibly with no speech.

[0064] Multiple detectors may include motion detectors, such as accelerometers. Motion detectors may be configured to detect movements of the user's facial muscles and / or bones, such as those caused by speech or chewing (e.g., jaw movements), and provide detector signals indicating such movements. Motion detectors may be configured to detect whether the involved device (e.g., a sound capture device or a hearing device) is moving or stationary. Accelerometers may be configured to detect the orientation (e.g., angle) of the device relative to gravity.

[0065] The sound capture device may include a classification unit configured to classify the current situation based on input signals from (at least partially) the detectors and possibly other inputs. In this specification, "current situation" may be defined by one or more of the following:

[0066] a) Physical environment (including the current electromagnetic environment, such as the presence of electromagnetic signals (including audio and / or control signals) that are planned or unplanned to be received by the sound capture device, or other properties of the current environment that are different from acoustics);

[0067] b) Current acoustic conditions (input level, feedback, etc.);

[0068] c) The user's current mode or state (movement, temperature, cognitive load, etc.); and

[0069] d) The current mode or state of the sound capture device and / or another device communicating with the sound capture device (selected program, time elapsed since the last user interaction, etc.).

[0070] The classification unit may be based on or include a neural network, such as a trained neural network.

[0071] The sound capture device can consist of hearing devices such as hearing aids or headphones.

[0072] Hearing devices such as hearing aids

[0073] Sound capture devices may include hearing devices such as hearing aids or constitute such devices.

[0074] Where appropriate, features of the sound capture embodiments described above and below as in the detailed embodiments, shown in the figures, or defined in the claims may be combined with features of hearing devices such as hearing aids, and vice versa.

[0075] Hearing aids may be adapted to provide frequency-varying gain and / or level-varying compression and / or frequency shifting (with or without frequency compression) from one or more frequency ranges to one or more other frequency ranges to compensate for a user's hearing loss. Hearing aids may include a signal processor for amplifying the input signal and providing a processed output signal.

[0076] Hearing aids may include an output unit for providing stimulation, perceived by the user as an acoustic signal, based on processed electrical signals. The output unit may include multiple electrodes of a cochlear implant (for CI-type hearing aids) or a vibrator of a bone conduction hearing aid. The output unit may include an output transducer. The output transducer may include a receiver (speaker) for providing the stimulation as an acoustic signal to the user (e.g., in acoustic (air conduction-based) hearing aids). The output transducer may also include a vibrator for providing the stimulation as mechanical vibrations of the skull to the user (e.g., in bone-attached or bone-anchored hearing aids).

[0077] Hearing aids may also include other suitable functions for the application in question, such as compression and feedback control.

[0078] Hearing aids may include hearing instruments, such as hearing instruments adapted to be located at the user's ear or wholly or partially in the ear canal, such as headphones, headsets, ear protection devices, or combinations thereof. Hearing aid systems may include loudspeaker amplifiers (including multiple input converters and multiple output converters, for example, for use in audio conferencing scenarios), and may include beamforming filter units, for example, providing multiple beamforming capabilities.

[0079] application

[0080] On the one hand, applications are provided for the sound capture device as described in detail in the "Detailed Description" section and as defined in the claims. Applications are available in systems including audio distribution. Applications are available in systems including one or more hearing aids (such as hearing instruments), headphones, headsets, active ear protection systems, etc., for example, in hands-free telephone systems, teleconferencing systems (e.g., including loudspeakers), etc.

[0081] method

[0082] On one hand, this application further provides a method for operating a sound capturing device configured to be worn by a person and / or placed on a surface such as a table. The sound capturing device can be configured to pick up target sound from a target sound source. The method may include one or more of the following steps, such as most or all of them:

[0083] - Provides multiple (M) electrical input signals, each electrical input signal IN m ,m=1,…,M includes the target signal component and the noise signal component;

[0084] - Provides an estimate of the target sound s;

[0085] - Provides a target holding and reference beamformer configured to attenuate signal components from other directions relative to a fixed target direction, while maintaining signal components from the fixed target direction with no or minimal attenuation relative to signal components from other directions, and provides a reference signal based on M electrical input signals;

[0086] - Provides a target cancellation beamformer configured to attenuate signal components from the target direction, while signal components from other directions are attenuated less than those from the target direction, and provides a target cancellation signal based on M electrical input signals;

[0087] - Provide at least two modes based on the mode control signal;

[0088] --Directional mode, where the estimate of the target sound s is based on the target signal component from a fixed target direction; and

[0089] --Non-directional, omnidirectional mode, where the estimate of the target sound s is based on the target signal components from all directions.

[0090] - Establish an audio link to another device;

[0091] - Transmit the estimate of the target sound s to the other device; and

[0092] - Determine the mode control signal based on the reference signal and the target elimination signal.

[0093] When appropriately replaced by a corresponding process, some or all of the structural features of the apparatus described above, in detail in the "Detailed Description," or as defined in the claims can be combined with the implementation of the method of the present invention, and vice versa. The implementation of the method has the same advantages as the corresponding apparatus.

[0094] Computer-readable media or data carrier

[0095] The present invention further provides a tangible computer-readable medium (data carrier) storing a computer program including program code (instructions), which, when the computer program is run on a data processing system (computer), causes the data processing system to perform (implement) at least some (such as most or all) of the steps of the methods described above, in detail in the "Detailed Description" and as defined in the claims.

[0096] By way of example, but not limitation, the aforementioned tangible computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to execute or store required program code in the form of instructions or data structures and is accessible by a computer. As used herein, disks include compact discs (CDs), laser discs, optical discs, digital multipurpose discs (DVDs), floppy disks, and Blu-ray discs, wherein these disks typically magnetically copy data while simultaneously being optically copied using lasers. Other storage media include those stored in DNA (e.g., in synthetic DNA strands). Combinations of the aforementioned disks should also be included within the scope of computer-readable media. In addition to being stored on tangible media, computer programs may also be transmitted via transmission media such as wired or wireless links or networks such as the Internet and loaded into data processing systems to run at locations other than tangible media.

[0097] Computer program

[0098] In addition, this application provides a computer program (product) including instructions that, when run by a computer, cause the computer to perform the steps of the methods (methods) described above, in detail in the "Detailed Description" section, and as defined in the claims.

[0099] Data processing system

[0100] In one aspect, the present invention further provides a data processing system, including a processor and program code, the program code causing the processor to perform at least some (such as most or all) of the steps of the methods described above, in detail in the "Detailed Description" section, and as defined in the claims.

[0101] Hearing system

[0102] On the other hand, a hearing system is provided that includes the sound capture device and another device described above, in detail in the "Detailed Description" section, and as defined in the claims.

[0103] The hearing system may be adapted to establish a communication link between the sound capture device and "another device" so that information (such as control and / or status signals, and / or audio signals) can be exchanged or forwarded from one device to another.

[0104] The sound capture device may include a remote control, a smartphone, or other portable electronic device with sound capture and communication capabilities, such as a wireless microphone unit, or be part of such a device.

[0105] "Another device" can be a hearing device such as a hearing aid. Hearing devices may include air conduction hearing aids, bone conduction hearing aids, cochlear implant hearing aids, or combinations thereof.

[0106] The hearing system can be adapted to allow the sound capture device to transmit an estimate of the target sound to "another device".

[0107] definition

[0108] In this specification, "hearing aid" as a hearing instrument refers to a device suitable for improving, enhancing, and / or protecting a user's hearing ability, which achieves this by receiving sound signals from the user's environment, generating corresponding audio signals, possibly modifying the audio signals, and providing the possibly modified audio signals as audible signals to at least one of the user's ears. The audible signals may be provided, for example, as sound signals radiating into the user's outer ear, sound signals transmitted as mechanical vibrations through the bone structures of the user's head and / or through parts of the middle ear to the user's inner ear, and electrical signals transmitted directly or indirectly to the user's cochlear nerve.

[0109] Hearing aids can be configured to be worn in any known manner, such as as a unit worn behind the ear (having a tube that directs radiated sound signals into the ear canal or having an output transducer, such as a speaker, arranged close to or located within the ear canal), as a unit wholly or partially arranged in the auricle and / or ear canal, as a unit connected to a fixed structure implanted in the skull, such as a vibrator, or as a connectable unit that is wholly or partially implanted. Hearing aids may include a single unit or several units that communicate with each other (e.g., acoustically, electrically, or optically). The speaker may be housed within the housing along with other components of the hearing aid, or it may be an external unit (possibly combined with a flexible guiding element such as a dome-shaped element).

[0110] More generally, a hearing aid includes an input transducer for receiving sound signals from the user's environment and providing a corresponding input audio signal, and / or a receiver for receiving the input audio signal electronically (i.e., wired or wirelessly); signal processing circuitry (typically configurable) for processing the input audio signal (such as a signal processor, for example including a configurable (programmable) processor, such as a digital signal processor); and an output unit for providing an audible signal to the user based on the processed audio signal. The signal processor may be adapted to process the input signal in the time domain or in multiple frequency bands. In some hearing aids, amplifiers and / or compressors may constitute the signal processing circuitry. The signal processing circuitry typically includes one or more (integrated or separate) storage elements for executing programs and / or for storing parameters used (or potentially used) in the processing and / or for storing information suitable for the hearing aid's functionality and / or for storing information used, for example, in conjunction with an interface to the user and / or an interface to a programming device (such as processed information, for example, provided by the signal processing circuitry). In some hearing aids, the output unit may include an output transducer, such as a loudspeaker for providing airborne sound signals or a vibrator for providing sound signals propagating through structures or fluids. In some hearing aids, the output unit may include one or more output electrodes for providing electrical signals that electrically stimulate the cochlear nerve (e.g., to a multi-electrode array) (cochlear implant hearing aids).

[0111] In some hearing aids, the vibrator may be adapted to transmit structurally propagated sound signals to the skull transdermally or through the skin. In some hearing aids, the vibrator may be implanted in the middle ear and / or inner ear. In some hearing aids, the vibrator may be adapted to provide structurally propagated sound signals to the middle ear bones and / or cochlea. In some hearing aids, the vibrator may be adapted to provide fluid-propagated sound signals to the cochlear fluid, for example, through the oval window. In some hearing aids, the output electrode may be implanted in the cochlea or on the medial side of the skull and may be adapted to provide electrical signals to the hair cells of the cochlea, one or more auditory nerves, the auditory brainstem, the auditory midbrain, the auditory cortex, and / or other parts of the cerebral cortex.

[0112] Hearing aids can be adapted to the specific needs of users, such as those with hearing loss. The configurable signal processing circuitry of a hearing aid can be adapted to apply frequency- and level-variable compression and amplification of the input signal. Customized frequency- and level-variable gain (amplification or compression) can be determined during the fitting process by the fitting system based on the user's hearing data, such as an audiogram, using basic fitting principles (e.g., speech adaptation). This frequency- and level-variable gain can be reflected, for example, in processing parameters, uploaded to the hearing aid via an interface to a programming device (fitting system), and used by a processing algorithm executed by the hearing aid's configurable signal processing circuitry.

[0113] A “hearing system” refers to a system that includes one or two hearing aids. A “binaural hearing system” refers to a system that includes two hearing aids and is adapted to work together to provide audible signals to both of a user’s ears. A hearing system or a binaural hearing system may also include one or more “assistive devices” that communicate with the hearing aids and influence and / or benefit from the functionality of the hearing aids. The aforementioned assistive devices may include at least one of the following: a remote control, a remote microphone, an audio gateway device, an entertainment device such as a music player, a wireless communication device such as a mobile phone (e.g., a smartphone), or a tablet computer, or another device, such as one that includes a graphical interface. Hearing aids, hearing systems, or binaural hearing systems may be used, for example, to compensate for hearing loss in individuals with hearing impairments, enhance or protect the hearing ability of individuals with normal hearing, and / or transmit electronic audio signals to individuals. Hearing aids or hearing systems may, for example, be part of or interact with broadcasting systems, active ear protection systems, hands-free telephone systems, car audio systems, entertainment systems (such as TV, music playback, or karaoke), teleconferencing systems, classroom amplification systems, etc.

[0114] The invention can be applied, for example, to assistive devices such as hearing aids or hearing aid systems. Attached Figure Description

[0115] Various aspects of the invention will be best understood from the following detailed description taken in conjunction with the accompanying drawings. For clarity, these drawings are schematic and simplified, showing only the details necessary for understanding the invention while omitting other details. Throughout the specification, the same reference numerals are used for the same or corresponding parts. Features of each aspect may be combined with any or all features of other aspects. These and other aspects, features, and / or technical effects will be apparent from and illustrated in the following figures, wherein:

[0116] Figure 1A A sound capture device is shown in an ideal position, attached to a person's shirt and configured to pick up the wearer's voice.

[0117] Figure 1B A sound capture device is shown that is poorly positioned, with the microphone axis pointing away from the wearer's mouth;

[0118] Figure 1C A sound capture device used as a desktop microphone is shown;

[0119] Figure 2A A perfect target cancellation beamformer is shown;

[0120] Figure 2B The image shows a case where the sound capture device is tilted, so that the null direction of the target cancellation beamformer is not directly pointed at the user's mouth;

[0121] Figure 2C The image shows the sound capture device placed on a table;

[0122] Figure 2D The example shows a reference beam pattern that is cardioid, with its null direction pointing away from the user's voice.

[0123] Figure 3 A first embodiment of the input stage of a sound capture device, such as a microphone unit or a hearing device, according to the present invention is shown;

[0124] Figure 4 A second embodiment of the input stage of a sound capture device, such as a microphone unit, according to the present invention is shown;

[0125] Figure 5A An embodiment of the sound capture device according to the invention is shown, which includes a light indicator (LED) for indicating the correct (optimal) position / direction;

[0126] Figure 5B An embodiment of the sound capture device according to the invention is shown, which includes a light indicator (LED) for marking incorrect (non-optimal) positions / directions;

[0127] Figure 6 An adaptive beamformer configuration is shown, wherein the adaptive beamformer Y(k) of the k-th subband is created by subtracting the (e.g., fixed) target cancellation beamformer C2(k) transformed by the adaptive factor β(k) from the (e.g., fixed) omnidirectional beamformer C1(k);

[0128] Figure 7 It shows the relationship with Figure 6 The adaptive beamformer configuration shown is similar, wherein the adaptive beam pattern Y(k) is created by subtracting another fixed beam pattern C1(k) from the target cancellation beamformer C2(k) which is converted by the adaptive factor β(k).

[0129] Figure 8 An embodiment of a hearing device according to the present invention is shown, which includes a BTE portion and an ITE portion;

[0130] Figure 9 An embodiment of the self-voice detector according to the present invention is shown;

[0131] Figure 10 A voice control interface connected to a self-voice detector according to the present invention is shown;

[0132] Figure 11 A block diagram of a hearing device including a self-voice detector according to the present invention is shown;

[0133] Figure 12A block diagram of a sound capture device including a pattern detector according to the present invention is shown.

[0134] The further applicability of the invention will become apparent from the detailed description given below. However, it should be understood that while the detailed description and specific examples illustrate preferred embodiments of the invention, they are given for illustrative purposes only. Other embodiments of the invention will become apparent to those skilled in the art based on the following detailed description. Detailed Implementation

[0135] The detailed description below, taken in conjunction with the accompanying drawings, serves as a description of various different configurations. This detailed description includes specific details to provide a thorough understanding of several different concepts. However, it will be apparent to those skilled in the art that these concepts can be implemented without these specific details. Several aspects of the apparatus and method are described by various different blocks, functional units, modules, elements, circuits, steps, processes, algorithms, etc. (collectively, “elements”). Depending on the specific application, design constraints, or other reasons, these elements may be implemented using electronic hardware, computer programs, or any combination thereof.

[0136] Electronic hardware may include microelectromechanical systems (MEMS), (e.g., application-specific integrated circuits), microprocessors, microcontrollers, digital signal processors (DSPs), field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), gating logic, discrete hardware circuits, printed circuit boards (PCBs) (e.g., flexible PCBs), and other suitable hardware configured to perform the various functions described in this specification, such as sensors for sensing and / or recording the physical properties of the environment, devices, users, etc. Computer programs should be interpreted broadly as instructions, instruction sets, code, code segments, program code, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, executables, threads of execution, programs, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description languages, or other names.

[0137] This application relates to the field of audio communication, and more particularly to sound capture devices such as hearing aids. On one hand, it relates to the interaction between a hearing aid (hearing instrument) and an external (assistive) device. The assistive device may take the form of a (e.g., wireless) sound capture device, for example, including a microphone array configured to communicate with the hearing aid. The wireless sound capture device may be adapted for wear by a person, such as the user of the hearing aid or another person, and / or adapted to be located where sounds of interest to the hearing aid user can be picked up, such as on a supporting structure like a table or shelf. The wireless sound capture device may include at least two microphones and be configured to apply directional processing to enhance desired sound signals picked up by the microphones of the sound capture device. Directional processing is necessary when sounds of interest always originate from the same desired direction. When the sound capture device is attached to a person, that person's speech is (presumably) of interest. Given that the sound capture device is correctly installed, the microphone array (such as a linear array) is always pointed towards the person's mouth. Thereby, directional processing can be applied to enhance the person's own speech while attenuating background noise.

[0138] The sound capture device can thus intercept sounds of interest and transmit the captured sounds directly to, for example, a user of a hearing aid. This results in a better signal-to-noise ratio compared to sound picked up directly by the microphone of a hearing aid.

[0139] However, sound capture devices may not always be used to pick up the voice of a single speaker. Sometimes, a sound capture device might be placed on a table to pick up the voice of anyone around the table. In this case, an omnidirectional response of the microphone may be more suitable than a directional response. Figure 1A-1C Different usage scenarios for sound capture devices are illustrated. A sound capture device such as a microphone unit (MICU) includes a housing containing two microphones (M1, M2). The two microphones form a microphone orientation (M-DIR). The microphone orientation (in...) Figure 1A-1C In the embodiments, the target direction is parallel to the longitudinal (“preferred”) direction of the housing formation. The target direction of the target-holding beamformer can be determined relative to the microphone direction or relative to the preferred direction of the housing of the sound capture device.

[0140] Figure 1A The MICU, a sound capture device, is shown in an ideal position, attached to the shirt of a person named MICU-W and configured to pick up the wearer's voice. Figure 1A This illustrates the planned use of a clip-on microphone unit for self-voice pickup. The microphone array (M1, M2) is pointed (M-DIR) at the user's mouth (the signal of interest), thereby enabling efficient directional attenuation of background noise. Background noise can be attenuated using directional processing, and the background noise is attenuated while the user's mouth direction OV-DIR remains unchanged (see dashed beam diagram DIR). If the sound capture device MICU is not installed correctly, for example... Figure 1BAs shown, the user's voice may be attenuated by the directional system. Figure 1B A sound capture device with suboptimal positioning is shown, where the microphone axis M-DIR is pointed away from the wearer's mouth. In this case, to ensure that the target speaker MICU-W is not attenuated by the directional noise reduction system, the directional noise reduction system should be turned off, making the sensor array sensitivity omnidirectional (switch to omnidirectional mode, see dashed circular beam diagram OMNI). Figure 1C The MICU, a sound capture device used as a desktop microphone, is shown. Figure 1C In this scenario, the sound capture device is placed on a support structure such as a table to pick up speech from people sitting around the table. In this case, directional microphone modes may attenuate parts of the speech of interest. Therefore, omnidirectional microphone sensitivity is preferred (see the hemispherical beam diagram OMNI).

[0141] The sound capture device according to the invention, for example Figure 1A-1C The different use cases of the microphone unit MICU shown in the diagram Figure 2A-2D The image shows a beam diagram focused on an exemplary beam pattern used to control the operating modes of a directional system.

[0142] This invention proposes a quality estimate based on potential directional benefits in a sound capture device (MICU) that switches between directional and omnidirectional modes. The quality of the directional beamformer can be evaluated based on an estimate of how well the zero-axis is manipulated toward the target speaker, relative to a reference beammap such as an omnidirectional beammap. In many adaptive noise reduction algorithms, a useful module is the target cancellation beamformer. The target cancellation beamformer is a directional beammap with its zero-axis pointing toward the signal of interest, ideally completely removing the target signal, thereby obtaining an estimate of the background noise in the absence of a target signal. The target cancellation beamformer can be pre-calibrated for a specific target location / direction, such as (ideally) the direction of the user's own speech (OV-DIR). The target cancellation beamformer in... Figure 2A The diagram shows (see the solid-line cardioid denoted as "DIR"). In this case, we expect to fully benefit from the directional noise reduction system because we see a large difference between the target cancellation beamformer DIR and the reference beammap (dashed circular pattern, denoted as OMNI-REF). The zero axis of the cardioid pattern points directly to the user's mouth (OV-DIR), thereby canceling the user's voice (MICU-W). The dashed beammap shows the omnidirectional reference beammap OMNI-REF. Considering the difference between the reference beammap and the beammap of the target cancellation beamformer, we see that the highest difference is obtained when the zero axis of the target beamformer points directly to the user's mouth (OV-DIR). In this case, when the sound capture device ("clamp array", MICU) is tilted ( Figure 2BThe difference between the target cancellation beamformer (solid line, DIR) and the reference beammap (dashed line, OMNI-REF) becomes smaller, and the user's voice is not completely canceled by the target cancellation beamformer. In this case, a smaller difference is seen between the target cancellation beamformer and the reference omnidirectional beammap (dashed line). Similarly, when the sound capture device (“microphone array”, MICU) is placed on a table (see...) Figure 2C When the target beamformer (SURF) is in the target direction (M-DIR), the sound of interest cannot arrive solely from the intended target direction. The sound of interest may arrive from any direction around the table (depending on the practical situation). Therefore, a high mean difference cannot be observed between the target beamformer (solid line, DIR) and the reference beammap (dashed line, OMNI-REF). The reference beammap does not necessarily have to be omnidirectional; for example, a cardioid pointing in the opposite direction to the target beamformer (a solid line cardioid denoted as DIR) can be used as the reference beammap. This is in Figure 2D It is shown in the figure (see the dashed heart shape denoted as REF). Figure 2A-2D The situation and Figure 1A-1C The configurations are similar, and the same reference names are used for the same components.

[0143] The term “beammap” (as used throughout this application) may also be called a “sensitivity map”, which indicates the spatial sensitivity (e.g., angular coherence) of a (directional) microphone system.

[0144] The following discussion Figure 3 and 4 The text outlines the use of... Figure 2A-2D The principle illustrated in this invention includes a pattern detector (see...). Figure 3 , 4 An embodiment of the sound capture device MICU (referred to as MODE-DET in the middle). Figure 3 and 4The ideal microphone orientation (equivalent to the direction toward the wearer's mouth, OV-DIR) of the microphones (M1, M2) of the sound capture device MICU-W and the input unit IU of the sound capture device is shown. The first and second microphones (M1, M2) respectively provide (time-domain, e.g., digitized) electrical input signals x1, x2. The sound capture device includes corresponding analysis filter banks for providing the first and second electrical input signals (x1, x2) in time-frequency representation (X1, X2, respectively). The (time-frequency domain) first and second electrical input signals (X1, X2) are fed to the pattern detector MODE-DET, and more particularly to the beamforming unit F-BF. The beamforming unit is configured to provide multiple fixed beamformers, including a reference beamformer ref and a target cancellation beamformer TC, each beamformer being a linear combination of the first and second electrical input signals (X1, X2), wherein the weights w of the respective beamformers are... ij The difference between the (e.g., omnidirectional) beamformer (OMNI-BF, signal ref) and the target voice cancellation beamformer (TC-BF, signal TC) across frequency bands is used for decision-making. A high difference indicates optimal conditions for the directional noise reduction system, enabling directional enhancement of user voice. A smaller difference between the two beamformers indicates suboptimal conditions for the directional noise reduction system. For the difference between the first and second thresholds, a gradient between omnidirectional and directional modes can be implemented. The first threshold may be lower than the second threshold. These thresholds can vary with frequency, for example, in different sub-bands. Preferably, the difference between the two directional signals is updated only when user voice is present. User voice can be detected using a voice activity detector. The sound capture device may be embodied, for example, in a microphone unit, and is adapted to communicate with another device, such as a hearing aid. The sound capture device may be embodied, for example, in a hearing device, such as a hearing aid.

[0145] Figure 3A first embodiment of the input stage of a sound capture device, such as a microphone unit or a hearing device, according to the present invention is shown. The magnitudes of the reference beamformer (see OMNI-BF, signal ref) and the target voice cancellation beamformer (see TC-BF, signal TC), i.e., signals |ref| and |TC|, are averaged (e.g., smoothed across time frames using a first-order low-pass filter (see corresponding unit LP)) to obtain stable estimates, see signals <|ref|> and <|TC|>, thereby avoiding fluctuating decisions. Preferably, smoothing occurs only when user voice is detected. Voice can be detected using a voice activity detector (see VAD), such as a modulation-based voice activity detector. The smoothed magnitudes of the reference beamformer OMNI-BF and the target voice cancellation beamformer TC-BF are respectively converted to the logarithmic domain (see unit log), see signals log(<|ref|>) and log(<|TC|>). The difference found in the separate channels (see...) Figure 3 The SUM unit "+" in the module is combined across frequencies into a joint decision (see module COMB-F). The COMB-F combination unit can be implemented, for example, through weighted summation, logistic regression, or a neural network. The weights can be estimated based on supervised learning. Alternatively, the combination function can be manually adjusted. When the difference between the estimated reference directional signal and the target speech cancellation signal is high, it indicates a high benefit of the directional noise reduction system, and the microphone unit MICU should switch to directional noise reduction. Otherwise, if the difference is small (e.g., less than 3 dB, less than 6 dB, or less than 9 dB), the potential benefit of directional noise reduction is limited, and the microphone unit should switch to omnidirectional mode. The directional mode can be adaptive or fixed. The decision (see the "Decision" module) can be a smooth transition between different directional modes (see... Figure 3 The illustration shows a smooth transition from omnidirectional mode to directional mode (represented by the signal M-CTR) as the difference between the omnidirectional beamformer and the target cancellation beamformer increases (represented by the signal COMP). Alternatively, the decision can be a binary transition between directional and omnidirectional. Hysteresis can be incorporated into the decision. In addition to switching only between directional and omnidirectional modes, the frequency shaping of the audio signal can also be altered based on the detected mode. The output of the mode detector MODE-DET, here the decision module, is the mode control signal M-CTR.

[0146] Figure 4 Another embodiment of the input stage of the sound capture device according to the present invention is shown. Figure 4 The embodiment includes an input unit IU that provides electrical input signals (X1, X2) and a beamforming unit F-BF that provides a reference beamformer ref and a target cancellation beamformer TC. Figure 3 The implementation is the same. However, compared with the consideration Figure 3In the embodiment, the difference between the reference beammap (ref) and the target speech beammap (TC) is opposite. Figure 4 The embodiment provides a normalized correlation coefficient β between two directional signals:

[0147]

[0148] See module TC, which provides the same signal. * ref and |TC| 2 And (controlled by the Voice Activity Detector (VAD)) a smoothed version that provides these signals. <TC * ref> and <|TC| 2 The low-pass filter LP is composed of a multiplication unit (division unit ÷), and the final combination unit provides β. This coefficient can also be used as an adaptive coefficient in an adaptive beamformer, see, for example, [Elko and Pong; 1995] or EP3588981A1 or EP3253075A1. In cases where the target voice is dominant (and the target cancellation beamformer can cancel the target signal), the value of β will increase. If β frequently has high values, we can therefore detect the user's self-voice (self-voice detection). If high values ​​of β occur frequently, we can therefore apply directional processing. Preferably, β is updated only when voice activity is detected, see... Figure 4 The VAD unit (for other applications such as noise reduction, it can be averaged based on the absence of voice). Since β can be calculated across channels, these values ​​should be combined into a single decision across frequencies (see COMB-F and the “Decision” unit). The decision (see the “Decision” module) can be a smooth transition between different directional modes (see... Figure 4 The illustration shows the smooth transition from omnidirectional mode to directional mode (represented by the mode control signal M-CTR) as the absolute value of parameter β increases (see |β| on the horizontal axis). (As...) Figure 3 As in the previous embodiment, the decision can be a binary transition between directional and omnidirectional modes. Hysteresis can be incorporated into the decision. Besides switching only between directional and omnidirectional modes, the frequency shaping of the audio signal can also be altered based on the detected mode. The output of the mode detector MODE-DET, in this case, is the mode control signal M-CTR, which is part of the decision module. As... Figure 3 As in the previous embodiment, the combined unit COMB-F (and / or decision unit) can be implemented, for example, through weighted summation or through logistic regression or through a neural network. The weights can be estimated based on supervised learning or through manual adjustment.

[0149] In combination Figure 3 and 4In the described embodiments, different candidate self-voice cancellation beamformers can be provided (e.g., based on predetermined beamformer weights, such as those stored in memory). The advantage of having multiple (e.g., several) candidate self-voice beamformers is that covering a range of mouth-to-sound device distances becomes possible, as the optimal self-voice cancellation beamformer varies with distance. Possible candidate self-voice beamformers can, for example, cover a range of 10-30 cm from the mouth. The beamformer with the deepest null direction can be selected at a given time point.

[0150] Joint decision-making across different frequency bands can be obtained through cross-frequency combination differences (or parameter β). Decisions can be based on a trained neural network. The COMB-F module, or the "decision-making" module, can be implemented through a trained neural network. The decision result in the decision module is the mode control signal M-CTR, which can be provided as the output vector of the trained neural network, where the input vector is the corresponding comparison unit (…). Figure 3 The "+" and Figure 4 The combination of "÷" in the signal (which varies with frequency). Figure 3 In this context, the output of the comparison unit (+) and the input of the "cross-frequency combination" unit COMB-F are log<|ref(k,l)|>-log<|TC(k,l)|>. Figure 4 In the comparison unit (÷), the output and the input of the "cross-frequency combination" unit COMB-F are β(k,l), where k and l are the frequency index and the time frame index, respectively.

[0151] Since the user of the MICU-W only wears and does not hear the sound capture device, for example, when implemented as a microphone unit MICU, it may require indication of the orientation quality and / or how well the sound capture device is installed. Indications may be provided, for example, via visual indicators such as LEDs or displays with information, or tactile indicators such as vibrators, or acoustic indicators. This is in Figure 5A , 5B The text shows (it shows the same as) Figure 1A and 1B (The same conditions apply to each case). The indication can be based on the orientation pattern estimated by the aforementioned detector. Alternatively, the indication can be based on orientation / azimuth sensors such as accelerometers or magnetometers. Figure 5A and 5B An embodiment of the sound capture device MICU according to the present invention is shown, which includes marking the correct (optimal) sound capture device on the wearer MICU-W. Figure 5A ) and errors (non-optimal) Figure 5BA position / direction indicator LED. The direction of the detected orientation quality or sound capture device can be conveyed to the user, for example, through a color change, such as from green to red (e.g., via yellow as an intermediate level), or via a constant flashing pattern, etc.

[0152] Figure 6 and 7 A corresponding embodiment of an adaptive beamformer configuration is shown, which can be used to implement a self-voice beamformer in a sound capture device according to the invention. Figure 6 and 7 Both diagrams show a dual-microphone configuration, commonly used in hearing devices such as hearing aids (or other sound capture devices) at the current level of technological advancement. However, these beamformers can be based on more than two microphones, for example, more than three microphones (e.g., as a linear array, or possibly configured in a non-linear manner). For a given frequency band k, the adaptive beammap Y(k) is obtained by linearly combining two beamformers C1(k) and C2(k). Each of C1(k) and C2(k) (time exponents have been omitted for simplicity) represents a different (possibly fixed) linear combination of the first and second electrical input signals X1 and X2 from the first and second microphones M1 and M2, respectively. The first and second electrical input signals X1 and X2 are provided by corresponding analytical filter banks (“filter banks”). Frequency domain signals (downstream of the corresponding analytical filter banks) are indicated by thick arrows, while the time domain properties of the outputs of the first and second microphones (M1, M2) are indicated by thin arrows. Figure 3 and 4 The F-BF module, which provides fixed beamformer ref and TC, is equivalent to Figure 6 and 7 The module F-BF is marked by a solid-line rectangle. Figure 3 and 4 The signals ref and TC are respectively equivalent to Figure 6 Signals C1(k) and C2(k). In another embodiment, Figure 3 and 4 The signals ref and TC can respectively be equivalent to Figure 7 The signals C1(k) and C2(k).

[0153] Figure 6 An adaptive beamformer configuration is shown, wherein the adaptive beamformer Y(k) of the k-th subband is created by subtracting the (e.g., fixed) omnidirectional beamformer C1(k) from the (e.g., fixed) target cancellation beamformer C2(k) calculated with an adaptive factor β(k). The adaptive factor β can be determined, for example, as:

[0154]

[0155] Figure 6 The two beamformers C1 and C2 are, for example, orthogonal. However, this is not actually necessary. Figure 7 The beamformers are not orthogonal. When beamformers C1 and C2 are orthogonal, uncorrelated noise will be attenuated when β = 0.

[0156] exist Figure 6 The (reference) beam pattern C1(k) in the image is an omnidirectional beam pattern (see, for example, [reference image]). Figure 2A At the same time, Figure 7 The (reference) beamformation C1(k) is a beamformer with the zero direction pointing in the opposite direction to C2(k) (see, for example, [reference]). Figure 2D Other fixed beam patterns C1(k) and C2(k) can also be used.

[0157] Figure 7 It shows the relationship with Figure 6 A similar adaptive beamformer configuration is shown, where the adaptive beammap Y(k) is created by subtracting another fixed beammap C1(k) from a target-cancelling beamformer C2(k) calculated with an adaptive factor β(k). This set of beamformers is not orthogonal. Figure 6 and 7 C2 in the context indicates that in the case of a self-voice cancellation beamformer, β will increase when self-voice is present.

[0158] A beam pattern can be, for example, a combination of an omnidirectional delay and summation beamformer C1(k) and its null-direction delay and subtraction beamformer C2(k) pointing towards the target direction (such as the mouth of a person wearing a sound capture device, i.e., a target cancellation beamformer), such as... Figure 6 As shown; or, it can be two delay-and-subtract beamformers, such as Figure 7 As shown, one of the beamformers, C1(k), has maximum gain in the target direction, while the other beamformer, C2(k), is a target cancellation beamformer. Other combinations of beamformers can also be applied. Preferably, the beamformers should be orthogonal, i.e., [w 11 w 12 [w] 21 w 22 ] H =0. The adaptive beammap is obtained by subtracting the target elimination beamformer C2(k) from C1(k) after calculating the target elimination beamformer C2(k) using a complex-valued, frequency-varying, adaptively updated conversion factor β(k), i.e.:

[0159]

[0160] in According to Figure 6 or Figure 7 The complex beamformer weights, and x = [x1, x2]T The input signals at the two microphones (after processing by the filter bank).

[0161] exist Figure 6 and 7 In the context, Figure 3 and 4 The fixed reference beamformer ref is therefore equivalent to And a fixed target cancellation beamformer TC is equivalent to in and Let x be the complex beamformer weights, for example, predetermined and stored in memory (or updated occasionally during use), and x = [x1, x2]. T This represents the (current) electrical input signal at both microphones (after filter bank processing).

[0162] Figure 8 An embodiment of the hearing device according to the present invention is shown, which includes a BTE portion and an ITE portion. Figure 8 An embodiment of a hearing device according to the invention is shown, which includes at least two input transducers such as microphones located in the BTE section and / or ITE section. Figure 8 Hearing devices such as hearing aids include a BTE section adapted to be placed at or behind the user's ear and an ITE section adapted to be placed in or within the user's ear canal. The BTE and ITE sections are connected (e.g., electrically connected) via a connecting element IC and internal wiring within the ITE and BTE sections (see, for example, the wiring schematically shown as Wx in the BTE section). Each of the BTE and ITE sections may respectively include an input transducer such as a microphone (M... BTE and M ITE This device is used to pick up sound from the environment of a user wearing a hearing device, and, in certain operating modes, to pick up the user's voice. The ITE section may include an earmold for delivering a relatively large sound pressure level to the eardrum of the user (e.g., a user with severe to profound hearing loss). Output transducers such as speakers may be located in the BTE section, and connecting element ICs may include tubes for acoustically propagating sound to and from the earmold to the user's eardrum.

[0163] The hearing device HD includes an input unit comprising two or more input transducers (such as microphones) (each input transducer is used to provide an electrical input audio signal representing an input sound signal). The input unit also includes two (individually selectable) wireless receivers (WLR1, WLR2) for providing corresponding directly received auxiliary audio input and / or control or information signals. The BTE section includes a substrate SUB on which multiple electronic components (MEM, FE, DSP) are mounted. The BTE section includes a configurable signal processor DSP and a memory MEM accessible therefrom. In embodiments, the signal processor DSP is formed as part of an integrated circuit, such as (primarily) a digital integrated circuit, while the front-end chip FE primarily includes analog circuitry and / or mixed analog-digital circuitry (including interfaces to microphones and speakers).

[0164] The hearing device HD includes an output converter SPK, which provides an enhanced output signal as a stimulus that can be perceived as sound by a user, based on an enhanced audio signal from or derived from a signal processor DSP. Alternatively or additionally, depending on the specific application, the enhanced audio signal from the signal processor DSP may be further processed and / or transmitted to another device.

[0165] exist Figure 8 In one embodiment of the hearing device, the ITE portion includes an output unit in the form of a speaker (sometimes called a receiver) SPK for converting electrical signals into acoustic signals. Figure 8 The ITE portion of the embodiment also includes an input converter M for picking up sound from the environment. ITE (e.g., a microphone). Based on the acoustic environment, the input converter M... ITE It can pick up some sound from the output converter SPK (unintentional acoustic feedback). The ITE section also includes guides such as dome pieces or earmolds or miniature earmolds (DO) to guide and position the ITE section in the user's ear canal.

[0166] exist Figure 8 In this scenario, the (far-field) (target) sound source S (mixed with other ambient sounds) is propagated to the BTE microphone M in the BTE section. BTE The sound field at the location, the ITE microphone M in the ITE section ITE The sound field S at the location ITE and the sound field S at the eardrum ED .

[0167] Figure 8 The hearing device HD illustrated herein represents a portable device and also includes a battery BAT, such as a rechargeable battery, for powering the electronic components of the BTE and ITE sections. In several different embodiments, Figure 8 Hearing devices can implement the self-voice detector OVD according to the present invention (see, for example, see...) Figure 9 A self-voice detector can be used, for example, in conjunction with a telephone mode, and / or with a voice control interface, see, for example, see Figure 10 , 11 .

[0168] In embodiments, hearing devices HD such as hearing aids (e.g., processors DSPs) are adapted to provide frequency-varying gain and / or level-varying compression and / or frequency shifting (with or without frequency compression) of one or more frequency ranges to one or more other frequency ranges, for example, to compensate for a user's hearing loss.

[0169] Figure 8 The hearing device contains two input converters (M BTE and M ITE For example, a microphone, when the hearing device is installed on the user's head, a (M) ITE The ITE section is located in or within the user's ear canal, while the other (M) BTE In the BTE section, it is located elsewhere on the user's ear (such as behind the user's ear (auricle)). Figure 8 In some embodiments, the hearing device may be configured such that two input converters (M BTE and M ITE When the hearing device is installed in the user's ear and is in normal working condition, it is positioned along a substantially horizontal line OL (see, for example, see...). Figure 8 Input converter M in BTE M ITE (and double-headed dashed line OL). This has the advantage of helping the electrical input signal from the input converter to beam in the appropriate (horizontal) direction, such as in the user's "line of sight" (e.g., toward the target sound source). Alternatively, the microphones can be positioned such that their axes point toward the user's mouth. Alternatively, another microphone can be included to provide the microphone axis along with one of the other microphones, thereby improving the pickup of the wearer's voice.

[0170] Figure 9 An embodiment of the input stage of a sound capture device, such as a hearing device, including a self-voice detector (OVD) according to the present invention is shown. The self-voice detector (OVD) is configured to provide a self-voice control signal (OV) indicating whether or with what probability a given electrical input signal (X1, X2) or its processed version originates from the voice of a user wearing the device including the self-voice detector (such as a sound capture device or a hearing device such as a hearing aid). The self-voice detector is configured to receive a plurality of (M) electrical input signals (X1, X2) represented in time-frequency (k, l). m ,m=1,…,M, where M=2,X1,X2), where k and l are the frequency and time frame exponents, respectively. The self-voice detector (OVD) includes multiple input transducers (ITs) connected during operation. mA beamforming unit F-BF comprises at least two fixed beamformers, each including a target-holding beamformer (OMNI-REF, referred to as the reference beamformer) configured to maintain signal components from a fixed target direction with minimal or no attenuation relative to signal components from other directions, and providing a current reference signal ref. The beamforming unit F-BF also includes a target-cancelling beamformer TC-BF configured to attenuate signal components from the target direction while attenuating signal components from other directions with less attenuation relative to signal components from the target direction, and providing a current target-cancelling signal TC. The fixed target direction is, for example, the direction from a hearing aid (e.g., a hearing aid microphone) facing the user's mouth, and the target signal is the user's own speech. The fixed beamformers (ref, TC) are, for example, based on a plurality of frequency-varying beamformer weights (w) stored, for example, in memory. 11 ,w 12 ,w 21 ,w 22 ) combination Figure 6 and 7 The fixed beamformer discussed. The self-voice detector (OVD) also includes a controller OVD-PRO for determining the self-voice control signal OV based on the current reference signal ref and the current target cancellation signal TC. The controller OVD-PRO includes corresponding signal paths for the reference beamformer signal ref and the target voice cancellation beamformer signal TC. Each signal path includes modules abs, LP, and log to provide signals log(<|ref|>) and log(<|TC|>) respectively, and includes a summation unit ("+") for providing the difference between the two signals (in sub-band representation) (log(<|ref|>) - log(<|TC|>)), as combined with Figure 3 As described in the embodiment for the mode detector MODE-DET. (As follows) Figure 3 As in [the previous section], the smoothing preference provided by the low-pass filter LP is only performed when user voice is detected (optional features are indicated by the VAD dashed box and the VAD control signal to the LP unit). Differences found in separate channels (see [reference]). Figure 9 The SUM unit "+" (as in combination) Figure 3 As described above, the frequencies are combined in essentially the same way for joint decision-making (see COMB-F and the decision module) (large difference => high probability of self-voice presence, small difference => low probability of self-voice presence). Again, COMB-F and / or the decision module can be implemented as a logic module or a trained neural network.

[0171] Figure 10A voice control interface (VCI) for a sound capture device such as a microphone unit or a hearing device such as a hearing aid is shown. The voice control interface (VCI) is connected to a self-voice detector (OVD) according to the present invention (e.g., as shown in the diagram). Figure 9 (as shown in the image). Figure 10 The Voice Control Interface (VCI) includes a keyword detection system configured to present the current audio stream (here, signal Y, for example, from...) to the keyword detection system. Figure 6 In a self-voice beamformer (or 7), detect whether or with what probability a specific keyword KWx (x = 1, ..., Q) exists. Figure 10 In one embodiment, the keyword detection system includes a keyword detector KWD, which is divided into first and second parts (KWDa, KWDb). The first part KWDa of the keyword detector includes a wake-up word detector WWD, denoted as KWDa(WWD), used to detect a specific wake-up word KW1 of the device involved, such as the voice control interface (VCI) of a hearing device (thus saving energy). The second part KWDb of the keyword detector is configured to detect the remaining keywords (KWx, x = 2, ..., Q) from a finite number of keywords. The voice interface of the hearing device is configured to be activated by a specific wake-up word spoken by a user wearing the hearing device. Figure 10 In this embodiment, based on the detection of the wake word KW1 by the first part of the keyword detector KWDa (wake word detector) and the electrical input signals X1, X2, the activation of the second part of the keyword detector KWDb is made dependent on the self-voice indication signal OV from the self-voice detector OVD. The voice control interface VCI includes a memory MEM for storing the current time period of the input audio stream Y, thereby enabling the detection of time periods in the self-voice indication signal OV where self-voice is absent before the keyword detector detects the wake word (or other keywords). The first and / or second parts of the keyword detector may be implemented as corresponding (trained) neural networks, whose weights are determined before use (or during training, while using the device involved, such as a hearing device) and applied to the corresponding network. The voice control interface may be configured to control the functions of a device, forming, for example, part of a hearing device. Keywords detectable by the keyword detector may include command words configured to control the functions of the device, such as mode switching, volume control, program switching, telephone call control, directionality, etc. The Voice Control Interface (VCI) includes a Voice Control Interface Controller (VC-PRO), which converts the keyword KWx identified by the keyword detector KWDb into the corresponding control signal HA. ctr Used to control the formation, for example, as shown here. Figure 11 The function of a portion of the hearing aid device shown.

[0172] Figure 11A block diagram is shown of a hearing device HD, such as a hearing aid, configured to be worn by a user and to compensate for the user's hearing loss without being required. The hearing aid HD includes a self-voice detector OVD according to the invention, for example, combined with... Figure 9 The self-voice detector OVD provides a self-voice control signal OV, which indicates whether or with what probability a given electrical input signal (X1, X2) or its processed version originates from the user's voice. The hearing aid includes an input unit IU comprising first and second microphones (M1, M2) adapted to provide (time-domain, e.g., digitized) electrical input signals (x1, x2). The hearing device includes a corresponding analysis filter bank FB-A for providing the first and second electrical input signals (x1, x2) in a time-frequency representation. The (time-frequency domain) first and second electrical input signals (X1, X2) are fed to a self-voice beamformer OV-BF, which provides an estimate Y of the user's self-voice, as combined with... Figure 6 , 7 As stated above. Figure 11 In this embodiment, the self-voice detector OVD is segmented to share the provision of beamformer signals (ref and TC) with the self-voice beamformer OV-BF. Reference (target hold) and target cancellation beamformer signals (ref and TC, respectively) are fed to the (self-voice detection) controller OVD-PRO for determining the self-voice control signal OV based on the current reference signal ref and the current target cancellation signal TC, as combined with... Figure 9 The estimated user self-voice Y from the self-voice beamformer OV-BF and the corresponding self-voice identification signal from the self-voice detector (in this case, OVD-PRO) are fed to the voice interface VCI, as described above. Figure 10 The control signal HA is used to provide the function of controlling the hearing aid. ctr The hearing aid includes a forward (signal) path from the input unit IU to the output unit OU. This forward path includes a corresponding analysis filter bank FB-A, which, as described above, provides corresponding electrical input signals (X1, X2) in a time-frequency representation. The electrical input signals (X1, X2) are fed to the (far-field) beamformer unit FF-BF, which provides a beamforming signal Y representing (spatially filtered) sound from the environment (e.g., a voice from a communication partner). BF The forward path also includes a signal processor HA-PRO, used to apply one or more processing algorithms to the beamforming signal Y. BF One or more processing algorithms may include, for example, compression amplification algorithms for (by applying a gain that varies with frequency and level to a signal in the forward path, such as a beamforming signal Y). BF This compensates for the user's hearing loss. The signal processor HA-PRO, such as one or more processing algorithms, can receive control signals from the voice control interface (VCI).ctr Control. The signal processor HA-PRO provides the processed signal OUT to the synthesis filter bank FB-S, which converts the time-frequency domain signal OUT into a time-domain signal out, which is then fed to the output unit OU. The output unit may include appropriate digital-to-analog converter functionality and an output converter, such as in the form of a speaker for air conduction hearing aids and / or a vibrator for bone conduction hearing aids. The output unit may also, or alternatively, include an electrode array for cochlear implant hearing aids for electrical stimulation of the cochlear nerve, in which case the synthesis filter bank can be omitted.

[0173] Figure 12 A sound capture device SCD, such as a microphone unit, is shown, which, in a first use case, is adapted to be worn by a person and to pick up the voice of that person (wearer), and, optionally, in a second use case, is adapted to be placed on a surface such as a table and to pick up sound from the environment (such as from the speaker) in this mode. The sound capture device SCD includes a pattern detector MODE-DET according to the invention, such as in combination with... Figure 3 , 4 The mode detector MODE-DET provides a mode control signal MCTR based on the corresponding reference beamformer signal ref and target cancellation beamformer signal TC at a given time point. The input stage of the sound capture device SCD includes an input unit comprising first and second microphones (M1, M2) adapted to provide (time-domain, e.g., digitized) electrical input signals (x1, x2) respectively; and a corresponding analysis filter bank FB-A for providing the first and second electrical input signals (x1, x2) in a time-frequency representation (X1, X2). The first and second electrical input signals (X1, X2) in the time-frequency domain are fed to a configurable noise reduction system CONF-BF for providing a configurable output signal Y based on the mode control signal M-CTR. x In the first usage scenario, the sound capture device SCD is worn by the user, and the noise reduction system CONF-BF is configured to provide an estimate Y of the user's own voice. x For example, combining Figure 6 , 7 As stated above, when the mode control signal M-CTR indicates that the direction of the microphone of the input unit is well matched with the direction of the wearer's mouth (in... Figure 1A , 2A In 2D, these are M-DIR and OV-DIR respectively. In the first usage scenario, when the mode control signal M-CTR indicates a poor match between the microphone direction M-DIR of the input unit and the direction OV-DIR of the wearer's mouth (see...). Figure 1B , 2BThe noise reduction system CONF-BF is configured to provide an omnidirectional signal (e.g., from one of the microphones, such as M1 (or from the target-holding beamformer (signal ref))). In a second use case, the sound capture device SCD is located on a carrier such as a table, and the same function of the directional noise reduction system CONF-BF is provided according to the mode control signal M-CTR. However, in the second use case, the "directional mode" is satisfied only for a person along the microphone axis (M-DIR) of the sound capture device SCD. In cases where it is expected that only one person will be heard, the sound capture device SCD is preferably positioned such that the microphone axis points towards that person. Otherwise, the directional noise reduction system CONF-BF will be in omnidirectional mode, transmitting the signal Y x Provided as an omnidirectional signal. The sound capture device SCD also includes a synthesis filter bank FB-S for converting the time-frequency signal Y... x (k,l) is converted into a time-domain signal Y. x (n), where k is the frequency exponent and l and n are the time exponents. The sound capture device SCD also includes a transmitter Tx, used to transmit the signal Y representing the sound picked up by the sound capture device SCD. x (n) (e.g., wireless) transmits to another device such as a telephone, PC, hearing aid or other communication device (see “transmit to other devices”).

[0174] Freefall detection

[0175] Because the MICU (Mutable Inspire Capture Unit) may include motion sensors such as accelerometers, it can detect the onset of a free fall, which could be caused by a user accidentally dropping the device. Due to the risk that the MICU may fall onto a hard surface, there is an additional risk of impact noise from colliding with a hard surface such as a floor and potentially bouncing on that surface. This risk of loud noise needs to be mitigated, as it could cause interference noise in the hearing aid output converter. When the MICU detects a free fall, there are several options for mitigating potential impact noise. One option is to mute the input signal, i.e., stop recording the input signal from the microphone and then transmit a signal with no further sound information to the hearing aid, or interrupt the transmission of the signal to the hearing aid. Another option is to transmit the signal from the MICU to the hearing aid, indicating that a free fall of the MICU has been detected, and that the sound from the processor to the output converter will be muted or at least attenuated, or even special noise cancellation processing will be initiated.

[0176] To restore normal operation of the sound from the MICU (Mound Induction Unit), a timer function can be implemented. The timer can be triggered in the MICU and / or the hearing aid, after which the sound can be restored to its previous level before the freefall began. The restoration may include a gradual increase, where the volume increases from zero to a working level or a predetermined level over a predetermined time period or in fixed steps. This allows the user of the MICU to relocate the device using the sound signal and to regain a sense of sound in the surrounding environment. The resumption of sound transmission can also be canceled by the signal from the accelerometer where the MICU has first collided with the ground. In this case, some sound caused by the MICU's bounce may be transmitted to the hearing aid, but at a lower level than usual, thus causing less inconvenience to the user.

[0177] Since not all impact sounds are likely to annoy users, the start of a free fall can trigger a reduction in the output level for the first time period. If the fall continues beyond this first time period, the output volume can be reduced to zero, i.e., complete silence. This prevents all sounds from being silenced when the sound capture device has only fallen a short distance and when the sound transmitted from the sound capture device quickly returns to normal levels.

[0178] In addition to free fall, one can imagine the sound capturing device colliding with something (without free fall before the collision). Due to the small transmission delay, we could also have a few milliseconds to mute the hearing aid or stop the sound transmission from the sound device after the high acceleration (caused by the collision) has been detected.

[0179] When appropriately replaced by a corresponding process, the structural features of the apparatus described above, in detail in the "Detailed Description" section, and as defined in the claims can be combined with the steps of the method of the present invention.

[0180] Unless explicitly stated otherwise, the singular forms “a” and “the” used herein include the plural forms (i.e., meaning “at least one”). It should be further understood that the terms “having,” “comprising,” and / or “including” as used in the specification indicate the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof. It should be understood that, unless explicitly stated otherwise, when an element is referred to as “connected” or “coupled” to another element, it may be a direct connection or coupling to the other element, or there may be intermediate inserting elements. The term “and / or” as used herein includes any and all combinations of one or more of the listed related items. Unless explicitly stated otherwise, the steps of any method disclosed herein do not necessarily have to be performed in the exact order disclosed.

[0181] It should be understood that references to "an embodiment," "an embodiment," "an aspect," or "may" in this specification mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the invention. Furthermore, particular features, structures, or characteristics may be suitably combined in one or more embodiments of the invention. The foregoing description is provided to enable those skilled in the art to implement the various aspects described herein. Various modifications will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects.

[0182] The claims are not limited to the aspects shown herein, but encompass the full scope consistent with the language of the claims, wherein, unless expressly stated, an element referred to in the singular does not mean "one and only one," but rather "one or more." Unless expressly stated, the term "some" means one or more.

[0183] Therefore, the scope of this invention should be determined based on the claims.

[0184] References

[0185] ·[Elko&Pong; 1995] Gary W Elko, Anh-Tho Nguyen Pong, "A Simple AdaptiveFirst-Order Differential Microphone". Published in: Proceedings of 1995Workshop on Applications of Signal Processing to Audio and Acoustics, IEEE, Print ISBN: 0-7803-3064-1.

[0186] ·EP3588981A1(Oticon)01.01.2020.

[0187] ·EP3253075A1(Oticon)06.12.2017.

Claims

1. A sound capturing device configured to be worn by a person and / or located on a surface, the sound capturing device being configured to pick up target sound from a target sound source and comprising: - Contains multiple input converters IT m Input units of m=1, 2, …, M, where M is greater than or equal to 2, each input converter is configured to pick up sound from the environment of the sound capture device and provide a corresponding electrical input signal, each electrical input signal IN m ,m=1, …, M includes the target signal component and the noise signal component; - A housing in which the plurality of input converters are located, the housing having a longitudinal orientation; - A directional noise reduction system for providing an estimate of a target sound s, the directional noise reduction system comprising components connected to the plurality of input converters IT during operation. m Beamformer units of m=1, …, M, the beamformer unit includes -- A target holding, reference beamformer configured to maintain signal components from a fixed target direction with minimal or no attenuation relative to signal components from other directions, and to provide a current reference signal, wherein the fixed target direction is consistent with the longitudinal direction; and -- Target cancellation beamformer, configured to attenuate signal components from the target direction, while signal components from other directions are attenuated less than those from the target direction, and to provide a current target cancellation signal; The directional noise reduction system is configured to operate in at least two modes based on a mode control signal: -- Directional mode, where the estimate of the target sound s is based on the target signal component from a fixed target direction; and -- Non-directional, omnidirectional mode, where the estimate of the target sound s is based on the target signal components from all directions; - Antenna and transceiver circuitry for establishing an audio link to another device, wherein the sound capture device is configured to transmit an estimate of the target sound s to the other device; and - A mode controller, used to determine the mode control signal based on the current reference signal and the current target elimination signal; Wherein, when the mode control signal indicates that the cross-frequency difference between the current reference signal and the current target cancellation signal or its processed version is greater than a second threshold, the directional noise reduction system is adapted to be in directional mode; and when the mode control signal indicates that the cross-frequency difference between the current reference signal and the current target cancellation signal or its processed version is less than a first threshold, the directional noise reduction system is adapted to be in omnidirectional mode.

2. The sound capture device according to claim 1, wherein at least one of the input converters is a microphone.

3. The sound capture device according to claim 1, comprising a filter bank.

4. The sound capture apparatus of claim 3, wherein the magnitudes or processed versions of the corresponding current reference signal and the current target cancellation signal are averaged over time to provide corresponding smooth reference and target cancellation metrics.

5. The sound capture apparatus of claim 4, comprising a voice activity detector, and wherein the sound capture apparatus is configured such that only time-frame averaging occurs when the voice activity detector detects user voice.

6. The sound capture apparatus of claim 3, comprising a combined processor configured to compare a current reference signal and a current target cancellation signal or a processed version thereof in different sub-bands and provide a corresponding sub-band comparison signal.

7. The sound capture device of claim 3, further comprising a decision controller configured to provide a mode control signal indicating an appropriate operating mode of the directional noise reduction system based on a sub-band comparison signal.

8. The sound capture apparatus of claim 7, wherein the decision controller is configured to provide a mode control signal based on a weighted sum of the comparison signals of each sub-band.

9. The sound capturing device according to claim 1, wherein the sound capturing device is a microphone device.

10. A hearing system comprising a sound capture device according to claim 1 and another device, wherein the sound capture device and the other device are configured to establish a communication link therebetween, thereby enabling the exchange of data including audio data therebetween or the transmission of data from the sound capture device to the other device.

11. The hearing system according to claim 10, wherein the other device is a hearing device.

12. The hearing system of claim 11, wherein the hearing device comprises an air conduction hearing aid, a bone conduction hearing aid, a cochlear implant hearing aid, or a combination thereof.

13. The hearing system of claim 10, adapted to cause the sound capturing device to transmit an estimate of the target sound s to the other device.

14. A method of operating a sound capturing device configured to be worn by a person and / or located on a surface, the sound capturing device being configured to pick up target sound from a target sound source s, the method comprising: - Provides M electrical input signals, each electrical input signal IN m , m=1, …, M includes the target signal component and the noise signal component; - A directional noise reduction system for providing an estimate of a target sound s, the directional noise reduction system comprising: -- Target holding, reference beamformer, configured to attenuate signal components from other directions relative to a fixed target direction, while maintaining signal components from the fixed target direction with no or minimal attenuation relative to signal components from other directions, and providing a reference signal based on M electrical input signals, wherein the fixed target direction is consistent with the longitudinal direction formed by the housing of the sound capture device; -- Target cancellation beamformer, configured to attenuate signal components from the target direction, while signal components from other directions are attenuated less than those from the target direction, and to provide target cancellation signals based on M electrical input signals; - Provide at least two modes based on the mode control signal; -- Directional mode, where the estimate of the target sound s is based on the target signal component from a fixed target direction; and -- Non-directional, omnidirectional mode, where the estimate of the target sound s is based on the target signal components from all directions; - Establish an audio link to another device; - Transmit the estimate of the target sound s to the other device; and - Determine the mode control signal based on the reference signal and the target elimination signal; Wherein, when the mode control signal indicates that the cross-frequency difference between the current reference signal and the current target cancellation signal or its processed version is greater than a second threshold, the directional noise reduction system is adapted to be in directional mode; and when the mode control signal indicates that the cross-frequency difference between the current reference signal and the current target cancellation signal or its processed version is less than a first threshold, the directional noise reduction system is adapted to be in omnidirectional mode.

Citation Information

Patent Citations

  • A hearing aid comprising a beam former filtering unit comprising a smoothing unit

    EP3253075A1

  • A hearing device comprising an acoustic event detector

    EP3588981A1

  • Microphone device with an orientation sensor and corresponding method for operating the microphone device

    US7912237B2

  • Method and system for wireless hearing assistance

    US8391522B2

  • Hearing device comprising acoustic event detector

    CN110636429A