Microphone arrangement
The microphone device, which combines a main microphone array with adaptive and fixed beamformers, solves the problem of the lack of directional sensitivity in traditional microphones, and achieves high-quality voice capture and noise suppression in different wearing positions and environments.
Patent Information
- Application Number
- CN202310232386.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-03-21
- Filing Date
- 2023-03-10
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2043-03-10
AI Technical Summary
Traditional microphone devices lack directional sensitivity, which allows ambient noise to get in and affects the quality of voice pickup, especially noticeable in headsets and hands-free phones.
A combination of a main microphone array, an adaptive beamformer, a fixed beamformer, and an analyzer is used. The adaptive beamformer provides a first directional audio signal with optimized directional sensitivity, the fixed beamformer provides a predetermined second directional audio signal, and the analyzer calculates information on directional sensitivity misalignment to correct the incorrect positioning of the microphone device.
It improves voice quality, reduces environmental noise interference, and ensures effective positioning and high-quality audio signal capture of the microphone device in different wearing positions and environments.
Smart Images

Figure CN116801166B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to a microphone apparatus and related computer-implemented method. BACKGROUND
[0002] In conventional microphone apparatuses, it is not uncommon for the microphone array to include microphones that capture sound in a way that has no directional sensitivity, i.e. sound from all directions is captured equally. However, for many purposes, not having directional sensitivity is less suitable. For example, in a telephony process where the microphone should pick up the user’s speech, the lack of directional sensitivity causes sound produced by the environment around the user to be picked up by the microphone as well. The environment produced sound can be confusing or otherwise disturbing, making the user’s speech difficult to understand. This is especially a problem when the microphone is far away from the mouth, such as in earbuds, headsets and speakerphones.
[0003] To overcome the problems associated with having no directional sensitivity, beamforming is a commonly applied technique. Beamforming is a technique of further processing the audio signals picked up by the microphone array. Beamforming relies on the fact that sound waves produced by a sound source in the space around the microphone array will have different times of incidence for different microphones of the microphone array, and thus the phases of the sound waves picked up by different microphones will differ from each other. By filtering the audio signals and combining them, a new audio signal with directional sensitivity can thus be achieved. Beamforming can thus be used to focus the audio signal in the direction of the sound source. Furthermore, beamforming can help alleviate problems caused by poor placement of the microphone apparatus by compensating for incorrect positioning of the microphones. However, even with beamforming, correct placement of the microphones relative to the sound source is an important parameter for obtaining a high quality audio signal, whether it is to compensate for the distance and signal-to-noise ratio, the microphones are calibrated to a specific position, or due to the geometry of the microphone array.
[0004] US 7,346,176 B1 discloses an example of compensation for incorrect positioning, which discloses a system and method of detecting whether a microphone apparatus is incorrectly positioned relative to a sound source and automatically compensating for such incorrect positioning. A position estimation circuit determines whether the microphone apparatus is incorrectly positioned. A controller facilitates automatic compensation for the incorrect positioning.
[0005] Another example is disclosed in EP 3007170 A1 which discloses a method for optimizing noise cancellation in a headset comprising a headset receiver and a microphone unit comprising at least a first microphone and a second microphone, the method comprising: generating at least a first audio signal from the at least first microphone, wherein the first audio signal comprises a speech portion from a user of the headset and a noise portion from a surrounding environment; generating at least a second audio signal from the at least second microphone, wherein the second audio signal comprises a speech portion from the user of the headset and a noise portion from the surrounding environment; producing a noise cancellation output by filtering and summing at least a portion of the first audio signal and at least a portion of the second audio signal, and wherein the filtering is adaptively configured to continuously minimize a power of the noise cancellation output, and wherein the filtering is adaptively configured to continuously provide at least an amplitude spectrum of a speech portion of the noise cancellation output, which corresponds to a speech portion of a reference audio signal generated from the at least one microphone.
[0006] US 2013 / 297305 A1 discloses a non-spatial voice detection system comprising a plurality of microphones, the outputs of which are provided to a fixed beamformer. An adaptive beamformer is used to receive the outputs of the plurality of microphones, and one or more processors are used to process the outputs from the fixed beamformer and to identify speech from noise using an algorithm that utilizes a covariance matrix.
[0007] US 2010 / 177908 A1 describes an audio signal processing technique in which an adaptive beamformer processes input signals from microphones according to estimates received from a pre-filter. The adaptive beamformer can compute its parameters (e.g., weights) for each frame according to the estimates, either through a magnitude domain objective function or a log-magnitude domain objective function. The pre-filter can include a time-invariant beamformer and / or a non-linear spatial filter, and / or can include a spectral filter. The computed parameters can be adjusted according to a constraint condition, which can be selectively applied only at desired times.
[0008] WO 2018 / 127447 A1 discloses an apparatus for capturing audio, the apparatus comprising a first beamformer coupled with an array of microphones and arranged to generate a first beamformed audio output.
[0009] However, the correct positioning of the microphones and how to achieve this, or how to compensate for an incorrect positioning, is still a key issue, and there is still room for improvement. SUMMARY
[0010] It is an object of the present disclosure to provide an improved microphone arrangement which overcomes or at least mitigates problems of the prior art. These and other objects of the present disclosure are realized by the disclosure defined in the independent claims and explained in the following. Further objects of the present disclosure are realized by the embodiments and specific implementations of the disclosure defined in the dependent claims.
[0011] According to a first aspect of the present disclosure, a microphone arrangement is provided, comprising: a main microphone array, an adaptive beamformer, a fixed beamformer, and an analyzer, wherein the main microphone array comprises: a first microphone adapted to provide a first input audio signal representing sound at a first microphone entrance, a second microphone adapted to provide a second input audio signal representing sound at a second microphone entrance, wherein the first microphone entrance is spatially separated from the second microphone entrance, and wherein the main microphone array is configured to:
[0012] • provide a main input vector comprising the first input audio signal and the second input audio signal as components,
[0013] wherein the adaptive beamformer is configured to:
[0014] • provide a first directional audio signal based on the main input vector, wherein a directional sensitivity of the first directional audio signal is selected to optimize speech quality,
[0015] wherein the fixed beamformer is configured to:
[0016] • provide a second directional audio signal based on the main input vector, wherein a directional sensitivity of the second directional audio signal is predetermined,
[0017] and wherein the analyzer is configured to:
[0018] • determine a first relative score indicative of a difference between the first directional audio signal and the second directional audio signal based on the first directional audio signal and the second directional audio signal, wherein the first relative score gives information about a directional sensitivity misalignment between the adaptive beamformer and the fixed beamformer, and
[0019] • output the first relative score for controlling a further processing of the first input audio signal and the second input audio signal, or for determining a mispositioning of the microphone arrangement.
[0020] Thus, the first relative score gives information about the misalignment of the beam sensitivity between the adaptive beamformer and the fixed beamformer. The information about the misalignment can be used to control the further processing of the audio signal or can be used to determine a wrong positioning of the microphone arrangement. Thus, by having the relative score, a wrong positioning of the microphone arrangement can be compensated by processing or can be corrected by positioning the microphone arrangement correctly.
[0021] The microphone arrangement can be configured to be worn by a user. The microphone arrangement can be arranged at, on, above, in, in the ear canal of, behind and / or in the pinna of a user's ear, that is, the microphone arrangement is configured to be worn at the user's ear.
[0022] The microphone arrangement can be configured to be worn by a user on each ear, for example, a pair of in-ear earphones or a headphone with two earcups. In implementations where the microphone arrangement is to be worn on both ears, the components to be worn on each ear can be connected, for example, wirelessly and / or by wires and / or by a band. The components for being worn on each ear can be substantially identical or different from each other.
[0023] The microphone arrangement can be audible, like a headphone, a headset, an earpiece, an in-ear earphone, a hearing aid, an over-the-counter (OTC) hearing device, a hearing protection device, a one-size-fits-all microphone arrangement, a custom microphone arrangement, or other head-wearable microphone arrangement. The microphone arrangement can be a hands-free telephone, or other device not configured to be worn by a user.
[0024] The microphone arrangement can be embodied in various housing styles or form factors. Some of these form factors are in-ear earphones, on-ear headset, or over-ear headset. A person skilled in the art is aware of various types of microphone arrangements and the different options to arrange the microphone arrangement in the ear and / or at the ear of the wearer of the microphone arrangement.
[0025] The microphone arrangement comprises a plurality of input transducers. The plurality of input transducers can comprise a plurality of microphones. The plurality of input transducers can be configured for converting an acoustic signal into an electrical input signal. The electrical input signal can be an analog signal. The electrical input signal can be a digital signal. The plurality of input transducers can be coupled with one or more analog-to-digital converters configured for converting the analog input signal into a digital input signal.
[0026] The microphone device can include one or more antennas configured for wireless communication. The one or more antennas can include an electric antenna. The electric antenna is configured for wireless communication at a first frequency. The first frequency can be above 800 MHz, preferably between 900 MHz and 6 GHz. The first frequency can be 902 MHz to 928 MHz. The first frequency can be 2.4 GHz to 2.5 GHz. The first frequency can be 5.725 GHz to 5.875 GHz. The one or more antennas can include a magnetic antenna. The magnetic antenna can include a magnetic core. The magnetic antenna includes a coil. The coil can be wound around the magnetic core. The magnetic antenna is configured for wireless communication at a second frequency. The second frequency can be below 100 MHz. The second frequency can be between 9 MHz and 15 MHz.
[0027] The microphone device can include one or more wireless communication units. The one or more wireless communication units can include one or more wireless receivers, one or more wireless transmitters, one or more transmitter-receiver pairs, and / or one or more transceivers. At least one of the one or more wireless communication units can be coupled with the one or more antennas. The wireless communication unit can be configured to convert a wireless signal received by at least one of the one or more antennas into an electrical input signal. The microphone device can be configured for wired / wireless audio communication, e.g., to enable a user to listen to media, such as music or a radio broadcast, and / or to enable a user to make a phone call.
[0028] The wireless signal can originate from an external source, such as a mate microphone device, a wireless audio transmitter, a smart computer, and / or a distributed microphone array associated with a wireless transmitter.
[0029] The microphone device can be configured for wireless communication with one or more external devices, e.g., one or more accessory devices, such as a smart phone and / or a smart watch.
[0030] The microphone arrangement can comprise one or more processing units. The processing units can be configured for processing one or more input signals. The processing can comprise compensating for a hearing loss of a user, i.e. applying a frequency dependent gain to the input signal according to a frequency dependent hearing impairment of the user. The processing can comprise performing feedback cancellation, beamforming, tinnitus reduction / masking, noise reduction, noise cancellation, speech recognition, bass adjustment, treble adjustment, face equalization, and / or processing of a user input. The processing units can be processors, integrated circuits, applications, functional modules, etc. The processing units can be implemented in a signal processing chip or a printed circuit board (PCB). The processing units are configured to provide an electrical output signal based on processing the one or more input signals. The processing units can be configured to provide one or more further electrical output signals. The one or more further electrical output signals can be based on processing the one or more input signals. The processing units can comprise receivers, transmitters, and / or transceivers for receiving and transmitting wireless signals. The processing units can control one or more playback functions of the microphone arrangement.
[0031] The microphone arrangement can comprise an output transducer. The output transducer can be coupled with the processing units. The output transducer can be a loudspeaker, or any other device configured for converting an electrical signal into an acoustic signal. The receiver can be configured for converting the electrical output signal into an acoustic output signal.
[0032] The wireless communication unit can be configured for converting the electrical output signal into a wireless output signal. The wireless output signal can comprise synchronization data. The wireless communication unit can be configured for transmitting the wireless output signal through at least one of the one or more antennas.
[0033] The microphone arrangement can comprise a digital-to-analog converter configured to convert the electrical output signal or the wireless output signal into an analog signal.
[0034] The microphone arrangement can comprise a power supply. The power supply can comprise a battery providing a first voltage. The battery can be a rechargeable battery. The battery can be a replaceable battery. The power supply can comprise a power management unit. The power management unit can be configured to convert the first voltage into a second voltage. The power supply can comprise a charging coil. The charging coil can be provided by a magnetic antenna.
[0035] The microphone arrangement can comprise a memory, including volatile and non-volatile forms of memory.
[0036] The main microphone array can comprise two or more microphones. The main microphone array can comprise one or more directional microphones and / or one or more omnidirectional microphones. The main microphone array can comprise a uniform linear array. The main microphone array can comprise an end-fire array. The main microphone array can comprise a broadside array. The main microphone array comprises a first microphone adapted to provide a first input audio signal representing sound at an entrance of the first microphone. The main microphone array comprises a second microphone adapted to provide a second input audio signal representing sound at an entrance of the second microphone. The entrance of the first microphone and the entrance of the second microphone can be arranged in an end-fire array or a broadside array. The entrance of the first microphone is spatially separated from the entrance of the second microphone. The main microphone array is configured to provide a main input vector comprising the first input audio signal and the second input audio signal as components. The main input vector can be provided as an electrical signal. The main input vector can be provided as an analog or digital signal. The main microphone array can be communicatively connected with a processing unit of the microphone device, wired or wirelessly, and configured to transmit the main input vector to the processing unit of the microphone device. The main microphone array can comprise an analog-to-digital converter to convert analog signals to digital signals, e.g. to convert analog signals generated from the first microphone and the second microphone to digital signals.
[0037] In the context of the present disclosure, voice quality can be determined by a wide range of parameters. Voice quality can be determined as direct to reverb ratio, wherein a higher direct to reverb ratio indicates a higher voice quality. Voice quality can be determined as signal to noise ratio, wherein a higher signal to noise ratio indicates a higher voice quality. Voice quality can be determined as predicted MOS (Mean Opinion Score), wherein a higher MOS indicates a higher voice quality. Other audio parameters can also be used to define voice quality.
[0038] In the context of the present disclosure, the term beamformer or beamforming can be interpreted broadly as any processing or means for providing an audio signal with directional sensitivity.
[0039] In the context of the present disclosure, an audio signal with directional sensitivity can be understood as an audio signal, wherein sound emitted from a specific direction or a specific range of directions is focused on, e.g. sound from a specific direction or a specific range of directions is kept unchanged or amplified, while sound from other directions is suppressed or removed. When referring to a beamformer providing an audio signal with directional sensitivity, it can be understood that the audio signal provided by the beamformer is focused on sound emitted from a direction corresponding to the directional sensitivity, wherein sound emitted from a direction not corresponding to the directional sensitivity is completely or at least partially filtered out from the provided audio signal.
[0040] The adaptive beamformer can be an analog adaptive beamformer or a digital adaptive beamformer. The adaptive beamformer can be configured to receive analog or digital signals of the primary input vector from the primary microphone array. The adaptive beamformer is configured to provide the first directional audio signal from the primary input vector. The directional sensitivity of the first directional audio signal is selected to optimize speech quality. The adaptive beamformer can be set to optimize speech quality by optimizing the directional sensitivity of the first directional audio signal based on specific audio parameters, such as optimizing signal-to-noise ratio.
[0041] The adaptive beamformer can improve speech quality by applying one or more beamforming weights to the primary input vector. The adaptive beamformer can improve speech quality by applying a set of beamforming filters / weights to the primary input vector. The beamforming weights can be represented as a beamforming weight vector. The beamforming weights required for the computation can be applied using different adaptive algorithms, such as minimum variance distortion response, generalized eigenvalue, simple matrix inversion, least mean, conjugate gradient method, etc. The adaptive beamformer can be configured to process the primary input vector in the time domain. The adaptive beamformer can be configured to process the primary input vector in the frequency domain, for example, by determining the Fourier transform of the primary input vector before beamforming is performed.
[0042] In one embodiment, the adaptive beamformer can include a machine learning model. Model coefficients of the machine learning model can be stored in a memory of the microphone device. In one embodiment, the machine learning model can be a neural network trained offline. In one embodiment, the neural network can include one or more input layers, one or more intermediate layers, and / or one or more output layers. The one or more input layers of the neural network can receive the primary input vector as input. The one or more output layers of the neural network can provide the first directional audio signal as output. The one or more output layers of the neural network can provide one or more beamforming weights as output.
[0043] In one embodiment, the machine learning model of the adaptive beamformer can be a deep neural network. In one embodiment, the deep neural network can be a convolutional neural network. In one embodiment, the deep neural network can be a region-based convolutional neural network. In one embodiment, the deep neural network can be a wavenet neural network. In one embodiment, the deep neural network can be a Gaussian mixture model. In one embodiment, the deep neural network can be a regression model. In one embodiment, the deep neural network can be a linear factorization model. In one embodiment, the deep neural network can be a kernel regression model. In one embodiment, the deep neural network can be a non-negative matrix factorization model.
[0044] The fixed beamformer can be an analog fixed beamformer or a digital fixed beamformer. The fixed beamformer can be configured to receive the primary input vector from the primary microphone array as an analog or digital signal. The fixed beamformer is configured to provide a second directional audio signal from the primary input vector. The directional sensitivity of the second directional audio signal is predetermined. The directional sensitivity of the second directional audio signal can be predetermined during a tuning process of the microphone device. The tuning process can be performed by an audio expert in a laboratory environment. The tuning process can be performed by an end user of the microphone device. The predetermined directional sensitivity can be predetermined by a user of the microphone device based on user preferences or a setup procedure. The predetermined directional sensitivity can be adjusted by the user to adapt to a new environment of the microphone device or to adapt to a new user of the microphone device. The user can input a directional sensitivity to the fixed beamformer, e.g. input a desired direction or a desired direction range that the fixed beamformer should focus on. The fixed beamformer can be configured to process the primary input vector in the time domain. The fixed beamformer can be configured to process the primary input vector in the frequency domain, e.g. by determining a Fourier transform of the primary input vector before beamforming. The fixed beamformer can comprise one or more fixed audio filters for processing the primary input vector to provide the second directional audio signal.
[0045] The analyzer can be configured to receive the first directional audio signal and the second directional audio signal as analog or digital signals. The analyzer is configured to determine a first relative score based on the first directional audio signal and the second directional audio signal. The analyzer is configured to output the first relative score. The analyzer can determine the first relative score by determining one or more audio parameters of the first directional audio signal and one or more audio parameters of the second directional audio signal and comparing the one or more audio parameters of the first directional audio signal with the one or more audio parameters of the second directional audio signal.
[0046] The adaptive beamformer, the fixed beamformer and the analyzer can all be digital processing blocks composed of processing units, e.g. digital signal processors. The adaptive beamformer, the fixed beamformer and the analyzer can all be processing units composed of a plurality of processing units connected to each other, e.g. one processing unit comprising the adaptive beamformer and the fixed beamformer, another processing unit comprising the analyzer and connected to the processing unit comprising the adaptive beamformer and the fixed beamformer. Alternatively, the adaptive beamformer, the fixed beamformer can be provided as analog beamformers and the analyzer can be provided as a digital processing block within a processing unit, wherein an analog-to-digital converter is arranged between the beamformers and the analyzer.
[0047] When it is stated that the first microphone entrance and the second microphone entrance are spatially separated, it is to be understood that the entrances are located at different positions, i.e. are arranged at different positions on the microphone arrangement.
[0048] The first relative fraction can be determined as a difference in signal-to-noise ratio in the first directional audio signal and the second directional audio signal. The first relative fraction can be determined as a difference in speech quality in the first directional audio signal and the second directional audio signal. The first relative fraction can be determined as a difference in root mean square of the first directional audio signal and the second directional audio signal.
[0049] In an embodiment, the primary microphone array further comprises one or more microphones adapted to provide one or more further input audio signals representing sound at one or more further microphone entrances, and wherein the primary microphone array is configured to:
[0050] • provide a primary input vector comprising the first, second and one or more further input audio signals as components.
[0051] By providing additional input audio signals, it can improve the performance of the beamformer by providing additional data for processing.
[0052] In an embodiment, the first microphone and / or the second microphone is an omnidirectional microphone.
[0053] Thus, the microphone can be sensitive to sound from all directions and thus be able to deliver directional audio signals that can be focused in a multitude of different directions.
[0054] In an embodiment, the first microphone and / or the second microphone is a directional microphone.
[0055] A directional microphone is a microphone configured for picking up sound from one or more specific directions. The directional microphone can be a gradient microphone.
[0056] In an embodiment, the first microphone is an omnidirectional or directional microphone, the second microphone is an omnidirectional or directional microphone, and the primary microphone array further comprises one or more further directional or omnidirectional microphones adapted to provide one or more further audio signals representing sound at one or more further microphone entrances.
[0057] In an embodiment, the directional sensitivity of the second directional audio signal is predetermined in accordance with an expected position of the first microphone and / or the second microphone.
[0058] Thus, the directional sensitivity of the fixed beamformer can be optimized for a certain use case of the microphone arrangement, and thus, the first relative score can provide information on whether the microphone arrangement is used correctly. For example, the microphone arrangement can be a headset with a boom arm comprising the first microphone and / or the second microphone, wherein the directional sensitivity of the fixed beamformer is optimized for the boom arm being positioned at the end of a slide rail or in front of the mouth of a user. In this case, the relative score calculated by the analyzer gives information on whether the boom arm is correctly arranged at the end of the slide rail or in front of the mouth of the user.
[0059] The intended position is a position of the microphone arrangement relative to a sound source, wherein the directional sensitivity of the fixed beamformer is optimized for the sound source.
[0060] The intended position can be mechanically determined by a structure of the microphone arrangement, for example, if the first microphone and / or the second microphone are arranged on a boom arm that can be pivoted between two end positions, for example, a non-use position in which the boom arm is stowed and a use position in which the boom arm is used to pick up the sound of a user of the microphone arrangement. Then, the intended position can be the use position of the boom arm. Further, the boom arm can slide in a groove with built-in stops, for example, formed by notches or protrusions, and then one or more of the built-in stops can serve as the intended position of the boom arm.
[0061] The intended position can be determined by a tuning process of the microphone arrangement. The tuning process can be performed during the production process of the microphone arrangement. The tuning process can be performed by a user of the microphone arrangement. The tuning process can comprise determining a use position of the microphone arrangement relative to a sound source and changing the directivity of the fixed beamformer in dependence on the determined use position. The directivity of the fixed beamformer can be selected to optimize the speech quality at the use position.
[0062] The tuning process can comprise the user arranging the microphone arrangement in a desired use position, and then the user can provide a user input to a processing unit of the microphone arrangement indicating that the microphone arrangement is in the desired position, and then the user can provide a user audio signal to the microphone arrangement from the user position, for example, by speaking loudly, in response to receiving the user input and the user audio signal, the adaptive beamformer can determine the directivity to optimize the speech quality, and the processing unit can transfer parameters of the directional sensitivity of the adaptive beamformer to the fixed beamformer, for example, by transferring one or more beamforming weights from the adaptive beamformer to the fixed beamformer.
[0063] The expected position can be determined from a plurality of positions. For example, the expected position can be determined by a parametric function. The parametric function for the expected position can receive one or more input parameters, such as a room parameter, a room shape, a room size, a user position within the room, a user head shape, a user head size, and / or a number of users within the room, and then output the expected position based on the one or more input parameters, where the expected position is then the position of the microphone arrangement based on the one or more input parameters.
[0064] In one embodiment, the microphone arrangement further comprises:
[0065] a speech detector configured to:
[0066] • provide, based on the primary input vector, a speech probability signal indicative of a probability of speech in the first input audio signal and / or the second input audio signal,
[0067] and wherein the adaptive beamformer is further configured to:
[0068] • provide, based on the speech probability signal and the primary input vector, the first directional audio signal.
[0069] Thus, the adaptive beamformer can use information from the speech detector to further optimize the directional sensitivity based on the quality of speech.
[0070] The speech detector can be composed of a processing unit. The speech detector can be a digital processing block in a digital signal processor. The speech detector can be configured to receive the primary input vector from the primary microphone array as a digital signal. The speech detector can be configured to provide the speech probability signal indicative of a probability of speech in the first input audio signal and / or the second input audio signal. The speech detector can be configured to process the primary input vector in a frequency domain, for example, by determining a Fourier transform of the primary input vector. The speech probability signal can be generated based on the first input audio signal or the second input audio signal. The first microphone or the second microphone can be defined as a reference microphone, and the speech detector can be configured to determine the speech probability signal based on the input audio signal produced by the reference microphone.
[0071] In one embodiment, the speech detector can comprise a machine learning model. Model coefficients of the machine learning model can be stored in a memory of the microphone arrangement. In one embodiment, the machine learning model can be a neural network trained offline. In one embodiment, the neural network can comprise one or more input layers, one or more intermediate layers, and / or one or more output layers. The one or more input layers of the neural network can receive the primary input vector as input. The one or more output layers of the neural network can provide the speech probability signal as output.
[0072] In an embodiment, the machine learning model of the voice detector can be a deep neural network. In an embodiment, the deep neural network can be a convolutional neural network. In an embodiment, the deep neural network can be a Gaussian mixture model. In an embodiment, the deep neural network can be a regression model. In an embodiment, the deep neural network can be a linear factorization model. In an embodiment, the deep neural network can be a kernel regression model. In an embodiment, the deep neural network can be a representation learning model.
[0073] The voice probability signal can comprise a speech mask. The voice probability signal can be a dataset showing the voice probability as a function of time and frequency, wherein the voice probability is represented as a number from 0 to 1, wherein 1 represents the presence of speech and 0 represents the absence of speech. In other words, 0 and / or values in the range of 0 to 0.5 can define a voice inactive region and 1 and / or values in the range of 0.5 to 1 can define a voice active region.
[0074] The adaptive beamformer can be configured to determine the one or more beamforming weights based on the voice probability signal and the primary input vector. The adaptive beamformer can be configured to determine a covariance matrix from the voice probability signal, the noise probability signal, and the primary input vector. The adaptive beamformer can determine the one or more beamforming weights from the covariance matrix.
[0075] In an embodiment, the microphone arrangement comprises a signal path selector, and wherein the analyzer is further configured to:
[0076] • compare the first relative score to a first threshold, and
[0077] • provide a first pass signal if the first relative score exceeds the first threshold,
[0078] wherein the signal path selector is configured to, in response to providing the first pass signal,
[0079] • pass the first directional audio signal for further processing to provide the audio signal to be transmitted, and
[0080] • block the second directional audio signal from further processing.
[0081] Thus, the use of processing capacity can be allocated in an efficient manner without wasting processing capacity on further processing the second directional audio signal.
[0082] The further processing of the first directional audio signal can comprise encoding the first directional audio signal. The further processing of the first directional audio signal can comprise filtering of the first directional audio signal. The further processing of the first directional audio signal can comprise transmitting the first directional audio signal to a device external to the microphone arrangement. The further processing of the first directional audio signal can comprise outputting the first directional audio signal, e.g. by a loudspeaker or other output transducer.
[0083] The first threshold value can be determined during a tuning process of the microphone arrangement. The tuning process can be performed during development of the microphone arrangement. The tuning process can be performed during production of the microphone arrangement. The tuning process can be performed by a user of the microphone arrangement. The tuning process can be performed by a user listening test, wherein the user determines the first threshold value based on listening to the first directional audio signal and the second directional audio signal at different first relative scores.
[0084] The first threshold value can be set to a fixed value. The first threshold value can be set to 0.5 dB, 1 dB or 2 dB when the first relative score is determined as a signal-to-noise ratio difference between the fixed beamformer and the adaptive beamformer.
[0085] In one embodiment, the first threshold value comprises a plurality of threshold values, each of the plurality of threshold values being associated with a respective frequency band.
[0086] In one embodiment, the microphone arrangement comprises a signal path selector, and wherein the analyser is further configured to:
[0087] • compare the first relative score to the first threshold value, and
[0088] • provide a second pass signal if the first relative score does not exceed the first threshold value,
[0089] and wherein the signal path selector is configured to, in response to providing the second pass signal,
[0090] • pass the second directional audio signal for further processing to provide the audio signal to be transmitted, and
[0091] • prevent further processing of the first directional audio signal.
[0092] Thus, the signal path of processing the audio signal is simplified.
[0093] The further processing of the second directional audio signal can comprise encoding the second directional audio signal. The further processing of the second directional audio signal can comprise filtering the second directional audio signal. The further processing of the second directional audio signal can comprise transmitting the second directional audio signal to a device external to the microphone arrangement. The further processing of the second directional audio signal can comprise outputting the second directional audio signal, e.g. by a loudspeaker or other output transducer.
[0094] In an embodiment, the adaptive beamformer is further configured to:
[0095] • switch from the active mode to the passive mode in response to the analyzer providing the second pass signal.
[0096] Thus, processing capacity can be freed up for other purposes. Furthermore, the consumption of the battery can be reduced.
[0097] The active mode refers to an operational mode in which the adaptive beamformer provides the first directional audio signal in parallel with the fixed beamformer providing the second directional audio signal. Unless otherwise stated, the adaptive beamformer can be assumed to be in the active mode.
[0098] The passive mode refers to an operational mode in which the adaptive beamformer refrains from providing the first directional audio signal in full parallel with the fixed beamformer providing the second directional audio signal. In the passive mode, the adaptive beamformer can refrain from determining a covariance matrix for an incoming audio signal. In the passive mode, the adaptive beamformer can refrain from providing the first directional audio signal. In the passive mode, the adaptive beamformer can provide the first directional audio intermittently, e.g. every 0.5 seconds, 1 second, 2 seconds, 3 seconds, 4 seconds or 5 seconds. In response to the analyzer providing the first pass signal, the adaptive beamformer can switch from the passive mode to the active mode.
[0099] In an embodiment, the analyzer is further configured to:
[0100] • determine an initial speech quality parameter, wherein the initial speech quality parameter is associated with the first audio input, the second audio input or a combination of the first audio input and the second audio input,
[0101] • determine a first speech quality parameter, wherein the first speech quality parameter is associated with the first directional audio signal,
[0102] • determine a second speech quality parameter, wherein the second speech quality parameter is associated with the second directional audio signal,
[0103] • determine a first difference between the first speech quality parameter and the initial speech quality parameter,
[0104] • determining a second difference between the second speech quality parameter and the initial speech quality parameter,
[0105] • determining a first relative score from the first difference and the second difference.
[0106] Thus, a simple and effective way of comparing the first directional audio signal and the second directional audio signal is achieved. By observing the improvements achieved by the beamformers, a simple yardstick for comparing the beamformers to each other is obtained.
[0107] The first difference can be seen as a measure of the speech quality improvement achieved by the adaptive beamformer's processing of the primary input vector. The second difference can be seen as a measure of the speech quality improvement achieved by the fixed beamformer's processing of the primary input vector.
[0108] The initial speech quality parameter, the first speech quality parameter and the second speech quality parameter can be determined by a wide range of parameters. The speech quality parameter can be determined as a direct to reverb ratio. The speech quality parameter can be determined as a signal to noise ratio. The speech quality parameter can be determined as a MOS. The speech quality parameter can be determined as a noise suppression. The speech quality parameter can be determined as a signal to speech distortion. The speech quality parameter can be determined as a noise attenuation. Other audio parameters can also be used to define the speech quality parameter.
[0109] The initial speech quality parameter can be determined by defining the first microphone or the second microphone as a reference microphone and subsequently determining the initial speech quality parameter for the input audio signal associated with the reference microphone.
[0110] The analyzer can be configured to determine the initial speech quality parameter by receiving the first audio input, the second audio input, or a combination of the first audio input and the second audio input, and the speech probability signal. The analyzer can then determine a signal to noise ratio, where the signal is determined as the power of the first audio input, the second audio input, or the combination of the first audio input and the second audio input in a speech active region, and the noise is determined as the power of the first audio input, the second audio input, or the combination of the first audio input and the second audio input in a speech inactive region. The analyzer can be configured to determine the speech active region and the speech inactive region from the speech probability signal.
[0111] The analyzer can be configured to determine the first speech quality parameter by receiving the first directional audio signal. The analyzer can then determine a signal to noise ratio, where the signal is determined as the power of the first directional audio signal in a speech active region, and the noise is determined as the power of the first directional audio signal in a speech inactive region. The analyzer can be configured to determine the speech active region and the speech inactive region from the speech probability signal.
[0112] The analyser can be configured to determine the second speech quality parameter by receiving the second directional audio signal. The analyser can then determine a signal-to-noise ratio, where the signal is determined as the power of the second directional audio signal in the speech active region and the noise is determined as the power of the second directional audio signal in the speech inactive region. The analyser can be configured to determine the speech active region and the speech inactive region from the speech probability signal.
[0113] In an embodiment, the microphone arrangement is a headset comprising a movable arm, wherein the first microphone inlet and / or the second microphone inlet is arranged on the arm.
[0114] Thus, a user of the headset can correct any mispositioning by moving the arm.
[0115] In embodiments where the microphone arrangement is a hands-free telephone or a headset without an arm, mispositioning of the first microphone inlet and / or the second microphone inlet can be corrected by moving the entire microphone arrangement relative to one or more users of the microphone arrangement.
[0116] In an embodiment, the analyser is further configured to:
[0117] • compare the first relative score to a first threshold value, and
[0118] • provide a mispositioning signal if the first relative score exceeds the first threshold value,
[0119] • output the mispositioning signal as a user notification for notifying a user about mispositioning of the first microphone inlet and / or the second microphone inlet.
[0120] Thus, the user will be aware of any mispositioning of the microphone arrangement and can take action to correct the mispositioning.
[0121] The user notification can be provided by a loudspeaker comprised in the microphone arrangement or in an external device in communication connection with the microphone arrangement, e.g. giving a voice instruction to correct the position of the microphone arrangement. The user notification can be provided as a text message, displayed on a display screen of the microphone arrangement or on an external device in communication connection with the microphone arrangement.
[0122] In an embodiment, the microphone arrangement further comprises a mispositioning indicator, wherein the mispositioning indicator is configured to:
[0123] • receive the mispositioning signal, and
[0124] • in response to receiving the error positioning signal, providing a user stimulus for indicating an error positioning of the first microphone inlet and / or the second microphone inlet.
[0125] Thus, the user will be aware of any error positioning of the microphone arrangement and can take action to correct the error positioning.
[0126] The error positioning indicator can be an LED, a loudspeaker, a vibration module, or other means capable of providing a user stimulus.
[0127] The user stimulus can be an audible stimulus, a visual stimulus, a tactile stimulus, or other form of stimulus or multiple stimuli.
[0128] According to a second aspect of the present disclosure, a computer-implemented method is provided, comprising the steps of:
[0129] • receiving a primary input vector comprising as components a first input audio signal representing sound at a first microphone inlet and a second input audio signal representing sound at a second microphone inlet, wherein the first microphone inlet is spatially separated from the second microphone inlet,
[0130] • based on the primary input vector, providing a first directional audio signal, wherein a directional sensitivity of the first directional audio signal is selected to optimize speech quality,
[0131] • based on the primary input vector, providing a second directional audio signal, wherein a directional sensitivity of the second directional audio signal is predetermined,
[0132] • based on the first directional audio signal and the second directional audio signal, determining a first relative score indicating a difference between the first directional audio signal and the second directional audio signal, wherein the first relative score gives information about a directional sensitivity misalignment between an adaptive beamformer and a fixed beamformer, and
[0133] • outputting the first relative score for controlling a further processing of the first input audio signal and the second input audio signal or for determining an error positioning of the microphone arrangement.
[0134] It is readily understood that all steps described in the first aspect with respect to processing audio signals can be performed in the computer-implemented method.
[0135] In this document, the singular forms "a," "an," and "the" indicate the existence of one or more entities, such as features, operations, elements, or components, but do not preclude the existence of one or more entities or the addition of one or more entities. Also, the terms "have," "comprise," and "contain" are used to specify the presence of one or more entities but do not preclude the presence or addition of one or more entities. The term "and / or" specifies the presence of one or more associated entities. BRIEF DESCRIPTION OF DRAWINGS
[0136] The present disclosure will be explained in greater detail in connection with the preferred embodiments, and with reference to the drawings, in which:
[0137] Figure 1 A schematic block diagram showing one embodiment of a microphone device according to the present disclosure is shown.
[0138] Figure 2 A schematic block diagram showing another embodiment of a microphone device according to the present disclosure is shown.
[0139] Figure 3 An example of a voice mask output by a voice detector according to the present disclosure is shown.
[0140] Figure 4 A schematic block diagram showing yet another embodiment of a microphone device according to the present disclosure is shown.
[0141] Figure 5 A schematic block diagram showing one embodiment of a microphone device according to the present disclosure is shown.
[0142] Figure 6 A schematic block diagram showing one embodiment of a microphone device according to the present disclosure is shown.
[0143] The drawings are schematic and simplified for clarity and only show details which are essential in order to explain the present disclosure, while other details can have been omitted in order not to obscure the concept of the present disclosure. Wherever possible, reference has been made to similar numerals in the drawings and / or description to indicate like or corresponding parts. DETAILED DESCRIPTION
[0144] The detailed description set forth below and the recitation of specific examples given herein to represent the preferred embodiments of the present disclosure are intended for purposes of illustration only and are not intended to limit the scope of the present disclosure. Those skilled in the art will readily appreciate that the present disclosure is well adapted to carry out the objects and obtain the ends and advantages mentioned, as well as those inherent therein. The present disclosure is therefore to be considered as including any such changes and modifications to the described embodiments that come within the scope of the present disclosure. Any such changes or modifications are intended to be within the fair meaning of the terms used herein and the following claims. An aspect or an advantage described in connection with a particular embodiment is not necessarily limited to that embodiment and can be practiced in any other embodiments even if not so described or so explicitly claimed.
[0145] Reference is first made to Figure 1Fig. 1 illustrates a schematic block diagram of a microphone arrangement 1 according to an embodiment of the present disclosure. The microphone arrangement 1 comprises a main microphone array 10, an adaptive beamformer 20, a fixed beamformer 30, and an analyzer 40. The microphone arrangement 1 can be a headset with or without boom, a pair of in-ear earphones, a hands-free telephone, or other audio device. The main microphone array 10 comprises a first microphone 11 adapted to provide a first input audio signal representing sound at a first microphone entrance, a second microphone 12 adapted to provide a second input audio signal representing sound at a second microphone entrance. The first microphone entrance is spatially separated from the second microphone entrance. The main microphone array 10 is configured to capture sound produced by a sound source external to the microphone arrangement, e.g. the main microphone array 10 can capture a user signal 2 produced by a user 2 of the microphone arrangement 1, in turn the first input audio signal and the second input audio signal can be the user signal 2 captured by the first microphone 11 and the second microphone 12. The main microphone array 10 is configured to provide a main input vector comprising the first input audio signal and the second input audio signal as components. The main input vector is provided as a digital signal. The first microphone 11 and the second microphone 12 are omnidirectional microphones.
[0146] The main input vector is received by the adaptive beamformer 20 from the main microphone array 10. The adaptive beamformer 20 is configured to provide a first directional audio signal based on the main input vector. The directional sensitivity of the first directional audio signal is selected to optimize speech quality.
[0147] The main input vector is received by the fixed beamformer 30 from the main microphone array 10. The fixed beamformer 30 is configured to provide a second directional audio signal based on the main input vector. The directional sensitivity of the second directional audio signal is predetermined. The directional sensitivity of the second directional audio signal is predetermined according to an expected position of the first microphone 11 and / or the second microphone 12.
[0148] The analyzer 40 receives the first directional audio signal from the adaptive beamformer 20 and the second directional audio signal from the fixed beamformer 30. The analyzer is configured to determine, based on the first directional audio signal and the second directional audio signal, a first relative score indicative of a difference between the first directional audio signal and the second directional audio signal. The analyzer 40 then outputs the first relative score. The first relative score can be output for further processing of the microphone arrangement and / or to an external device of the microphone arrangement. The analyzer 40 can determine the first relative score by determining a first audio parameter associated with the first directional audio signal and a second audio parameter associated with the second directional audio signal and comparing the first audio parameter with the second audio parameter to determine the first relative score. The comparison between the first audio parameter and the second audio parameter can be a determination of a difference between the first audio parameter and the second audio parameter.
[0149] The main microphone array 10, the adaptive beamformer 20, the fixed beamformer 30 and the analyzer 40 can all constitute a part of a digital signal processor of the microphone arrangement 1. The main microphone array 10 can comprise an analog-to-digital converter configured to convert audio signals picked up by the first microphone 11 and the second microphone 12 into digital signals.
[0150] Reference is now made to Figure 2 and Figure 3 wherein Figure 2 depicts a schematic block diagram of one embodiment of a microphone arrangement 1 according to the present disclosure, Figure 3 depicts an example of a speech mask 51 output by a speech detector 50 according to the present disclosure. Figure 2 The microphone arrangement of Figure 1 differs from the microphone arrangement of Figure 3 in that it further comprises a speech detector 50. The speech detector 50 is configured to receive a main input vector from the main microphone array 10. The speech detector 50 is configured to provide, based on the main input vector, a speech probability signal indicative of a probability of speech in the first input audio signal and / or the second input audio signal. Preferably, the first microphone or the second microphone is selected as a reference microphone and the speech probability signal is provided based on the selected reference microphone. The speech probability signal can be provided as a speech mask 51 as shown in
[0151] The adaptive beamformer 20 receives the speech probability signal from the speech detector 50. The adaptive beamformer 20 is configured to provide the first directional audio signal based on the speech probability signal and the main input vector. The adaptive beamformer can determine a covariance matrix based on the speech probability signal and the main input vector and determine one or more beamforming weights based on the covariance matrix.
[0152] Reference is now made to Figure 4 ,Figure 4 A schematic block diagram depicting one embodiment of a microphone arrangement 1 according to the present disclosure is depicted. Figure 4 The embodiment of Figure 2 differs from the one described in
[0153] The analyzer 40 in the present embodiment is further configured to determine an initial speech quality parameter. The initial speech quality parameter can be associated with the first audio input, the second audio input, or a combination of the first audio input and the second audio input. However, in the present embodiment, the first microphone 11 is defined as a reference microphone, and the initial speech quality parameter is determined based on the first input audio signal provided by the first microphone 11. The initial speech quality parameter is determined as a signal-to-noise ratio in the first input audio signal. Wherein the signal is determined as the power of the first audio input in a speech active region, and the noise is determined as the power of the first audio input in a speech inactive region. The analyzer 40 is configured to determine the speech active region and the speech inactive region based on the speech probability signal, for example, the speech probability signal can be Figure 3 the speech mask 51 depicted above, wherein the speech active region and the speech inactive region can be determined directly based on the speech mask 51. The analyzer 40 is further configured to determine the first speech quality parameter by receiving the first directional audio signal. Then, the analyzer 40 determines the signal-to-noise ratio, wherein the signal is determined as the power of the first directional audio signal in the speech active region, and the noise is determined as the power of the first directional audio signal in the speech inactive region. The analyzer 40 is further configured to determine the second speech quality parameter by receiving the second directional audio signal. Then, the analyzer 40 determines the signal-to-noise ratio, wherein the signal is determined as the power of the second directional audio signal in the speech active region, and the noise is determined as the power of the second directional audio signal in the speech inactive region.
[0154] Then, the analyzer 40 determines a first difference between the first speech quality parameter and the initial speech quality parameter. The first difference can be seen as a measure of the improvement or degradation in speech quality after the fixed beamformer has processed the primary input vector. The analyzer 40 further determines a second difference between the second speech quality parameter and the initial speech quality parameter. The second difference can be seen as a measure of the improvement or degradation in speech quality after the adaptive beamformer has processed the primary input vector.
[0155] Based on the first difference and the second difference, the analyzer 40 is further configured to determine a first relative score. The first relative score is determined by determining the difference between the first difference and the second difference, and then the first difference can be expressed as the dB difference between the first difference and the second difference.
[0156] Reference is now made to Figure 5 , Figure 5 A schematic block diagram of one embodiment of a microphone arrangement 1 according to the present disclosure is described. Figure 5 The microphone arrangement of Figure 1 differs from the microphone arrangement of The microphone arrangement comprises a signal path selector 60. The signal path selector 60 is configured to pass a signal 61 for further processing. The signal path selector 60 is configured to receive the first directional audio signal from the adaptive beamformer 20. The signal path selector 60 is configured to receive the second directional audio signal from the fixed beamformer 30. The signal path selector 60 is configured to receive the first pass signal or the second pass signal from the analyzer 40. The analyzer 40 is configured to compare the first relative score to a first threshold value. The first threshold value is set to a fixed value of 1 dB. If the first relative score exceeds the first threshold value, the analyzer 40 is configured to provide the first pass signal. If the first relative score does not exceed the first threshold value, the analyzer 40 is configured to provide the second pass signal. In response to receiving the first pass signal, the signal path selector 60 is configured to pass the first directional audio signal for further processing and to block the second directional audio signal from further processing. In response to receiving the second pass signal, the signal path selector 60 is configured to pass the second directional audio signal for further processing and to block the first directional audio signal from further processing. The adaptive beamformer is further configured to receive the second pass signal and to switch from the active mode to the passive mode in response to receiving the second pass signal.
[0157] Finally, reference is made to Figure 6 , Figure 6 A schematic block diagram of one embodiment of a microphone arrangement 1 according to the present disclosure is described. Figure 6 The microphone arrangement 1 depicted in Figure 4 and Figure 5 The microphone arrangement 1 depicted in Figure 6 The microphone arrangement comprises both the voice detector 50 and the signal path selector 60. The microphone arrangement 1 further comprises an error localization indicator 70. The analyzer 40 is further configured to compare the first relative score to a first threshold value. The analyzer 40 is further configured to provide an error localization signal if the first relative score exceeds the first threshold value. The analyzer 40 is further configured to provide an error localization signal if the first relative score exceeds the first threshold value. The analyzer 40 is further configured to output the error localization signal as a user notification for notifying the user 2 about an error localization with respect to the first microphone inlet and / or the second microphone inlet. The error localization indicator 70 is configured to receive the error localization signal and to provide a user stimulus in response to receiving the error localization signal, indicating an error localization of the first microphone inlet and / or the second microphone inlet.
[0158] The present disclosure is not limited to the embodiments disclosed herein, but can be embodied in other ways within the subject matter defined in the following claims. As an example, features of the described embodiments can be combined in any combination, for example, to adapt the device according to the present disclosure to specific requirements.
[0159] Any reference signs in the claims and the specification shall not be construed as limiting the scope of the claims.
Claims
1. A microphone arrangement (1) comprising: a primary microphone array (10), an adaptive beamformer (20), a fixed beamformer (30) and an analyzer (40), wherein the primary microphone array (10) comprises: a first microphone (11) adapted to provide a first input audio signal representing sound at a first microphone entrance, a second microphone (12) adapted to provide a second input audio signal representing sound at a second microphone entrance, wherein the first microphone entrance and the second microphone entrance are spatially separated, and wherein the primary microphone array (10) is configured to: provide a primary input vector comprising the first input audio signal and the second input audio signal as components, wherein the adaptive beamformer (20) is configured to: provide a first directional audio signal based on the primary input vector, wherein a directional sensitivity of the first directional audio signal is selected to optimize speech quality, wherein the fixed beamformer (30) is configured to: provide a second directional audio signal based on the primary input vector, wherein a directional sensitivity of the second directional audio signal is predetermined, and wherein the analyzer is configured to: determine a first relative score indicative of a difference between the first directional audio signal and the second directional audio signal based on the first directional audio signal and the second directional audio signal, wherein the first relative score gives information about a misalignment of the directional sensitivity between the adaptive beamformer (20) and the fixed beamformer (30), and output the first relative score for controlling a further processing of the first input audio signal and the second input audio signal or for determining a mislocation of the microphone arrangement (1).
2. The microphone apparatus of claim 1, wherein, The first microphone (11) and the second microphone (12) are omnidirectional microphones.
3. The microphone apparatus of claim 1 or claim 2, wherein, The directional sensitivity of the second directional audio signal is predetermined based on an expected position of the first microphone (11) and / or the second microphone (12).
4. The microphone arrangement according to claim 1 or claim 2, further comprising: a speech detector (50) configured to: provide a speech probability signal indicative of a speech probability in the first input audio signal and / or the second input audio signal based on the primary input vector, and wherein the adaptive beamformer (20) is further configured to: provide the first directional audio signal based on the speech probability signal and the primary input vector.
5. The microphone apparatus of claim 1 or claim 2, further comprising: a signal path selector (60), and wherein the analyzer (40) is further configured to: compare the first relative score with a first threshold, and provide a first pass signal if the first relative score exceeds the first threshold, and wherein the signal path selector (60) is configured to, in response to the first pass signal being provided, pass the first directional audio signal for further processing to provide an audio signal to be transmitted, and further processing of the second directional audio signal is prevented.
6. The microphone apparatus of claim 1 or claim 2, further comprising: a signal path selector (60), and wherein the analyzer is further configured to: compare the first relative fraction to a first threshold, and if the first relative fraction does not exceed the first threshold, provide a second pass signal, and wherein the signal path selector (60) is configured to, in response to the second pass signal being provided, pass the second directional audio signal for further processing to provide an audio signal to be transmitted, and further processing of the first directional audio signal is prevented.
7. The microphone apparatus of claim 6, wherein, the adaptive beamformer is further configured to: switch from an active mode to a passive mode in response to the analyzer providing the second pass signal.
8. The microphone apparatus of claim 5, wherein, the analyzer is further configured to: determine an initial voice quality parameter, wherein the initial voice quality parameter is associated with the first input audio signal, the second input audio signal, or a combination of the first input audio signal and the second input audio signal, determine a first voice quality parameter, wherein the first voice quality parameter is associated with the first directional audio signal, determine a second voice quality parameter, wherein the second voice quality parameter is associated with the second directional audio signal, determine a first difference between the first voice quality parameter and the initial voice quality parameter, determine a second difference between the second voice quality parameter and the initial voice quality parameter, determine the first relative fraction based on the first difference and the second difference.
9. The microphone apparatus of claim 1 or claim 2, wherein, the microphone arrangement is a headset comprising: a movable arm, wherein the first microphone inlet and / or the second microphone inlet are arranged on the arm.
10. The microphone apparatus of claim 1 or claim 2, wherein, the analyzer is further configured to: compare the first relative fraction to a first threshold, and if the first relative fraction exceeds the first threshold, provide an error positioning signal, output the error positioning signal as a user notification for notifying a user about an error positioning of the first microphone inlet and / or the second microphone inlet.
11. The microphone apparatus of claim 10, wherein, the microphone arrangement further comprises an error positioning indicator (70), wherein the error positioning indicator is configured to: receive the error positioning signal, and in response to receiving the error positioning signal, provide a user stimulus for indicating an error positioning of the first microphone inlet and / or the second microphone inlet.
12. A computer-implemented method comprising the steps of: receiving a primary input vector comprising as components a first input audio signal representing sound at a first microphone inlet of a microphone arrangement and a second input audio signal representing sound at a second microphone inlet, wherein the first microphone inlet is spatially separated from the second microphone inlet, providing a first directional audio signal based on the primary input vector, wherein a directional sensitivity of the first directional audio signal is selected to optimize voice quality, providing a second directional audio signal based on the primary input vector, wherein a directional sensitivity of the second directional audio signal is predetermined, determining, based on the first directional audio signal and the second directional audio signal, a first relative score indicative of a difference between the first directional audio signal and the second directional audio signal, wherein the first relative score gives information about a misalignment between a direction sensitivity of an adaptive beamformer and a fixed beamformer, and outputting the first relative score for controlling a further processing of the first input audio signal and the second input audio signal or for determining a false localization of the microphone arrangement.
Citation Information
Patent Citations
Robust noise cancellation using uncalibrated microphones
EP3007170A1
Non-spatial speech detection system and method of using same
US20130297305A1
Auto-adjust noise canceling microphone with position sensor
US7346176B1
Method and apparatus for audio capture using beamforming
WO2018127447A1
Hearing device adapted for orientation
CN114125677A