Robust speaker positioning system and method in presence of strong noise interference
By estimating the RTF of the target audio using a multi-microphone array and generalized eigenvalue beamforming, the problem of target audio localization and enhancement in noisy environments is solved, achieving robust audio processing effects under strong noise conditions.
Patent Information
- Application Number
- CN202110655603.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-12
- Filing Date
- 2021-06-11
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2041-06-11
AI Technical Summary
In noisy environments, existing technologies struggle to effectively isolate target audio from noise, especially when the target audio comes from different directions and the noise interference is strong, resulting in poor target audio detection and processing performance.
By employing a multi-microphone array and generalized eigenvalue beamforming technology, the main noise sources are zeroed out by estimating the relative transfer function (RTF) of the target audio. Combined with an improved TDOA/DOA estimation method, the location of the target audio source is determined in real time, and the target audio is enhanced using a minimum variance distortionless response (MVDR) beamformer.
Robust localization and enhancement of target audio were achieved under strong noise interference, improving the detection accuracy and signal-to-noise ratio of target audio and enhancing the effectiveness of speech processing.
Smart Images

Figure CN113810825B_ABST
Abstract
Description
Technical Field
[0001] According to one or more embodiments, this disclosure generally relates to audio signal processing, and more particularly, for example, to systems and methods for robust speaker localization in the presence of strong noise interference. Background Technology
[0002] In recent years, smart speakers and other voice-controlled devices and appliances have become increasingly popular. Smart speakers typically include a microphone array for receiving audio input (e.g., a user's verbal commands) from the environment. When a target audio (e.g., a verbal command) is detected in the audio input, the smart speaker can translate the detected target audio into one or more commands and perform different tasks based on those commands. One challenge for these smart speakers is efficiently and effectively isolating the target audio (e.g., verbal commands) from noise in the operating environment. This challenge is exacerbated in noisy environments where the target audio could originate from any direction relative to the microphone.
[0003] In view of the foregoing, there is a need for improved systems and methods for processing audio signals received in noisy environments. Summary of the Invention
[0004] This disclosure provides systems and methods for improving audio signal processing in noisy environments. Various embodiments of the systems and methods are disclosed herein, and include: a plurality of audio input components configured to generate a plurality of audio input signals; and a logic device configured to receive the plurality of audio input signals, determine whether the plurality of audio signals include target audio associated with an audio source, estimate the relative position of the audio source relative to the plurality of audio input components based on the plurality of audio signals and the determination of whether the plurality of audio signals include the target audio, and process the plurality of audio signals by enhancing the target audio based on the estimated relative position to generate an audio output signal. The logic device is also configured to construct a cross-band aligned direction covariance matrix using a covariance based on relative transitivity, and find a direction that minimizes beam power under a distortion-free criterion.
[0005] The scope of this disclosure is defined by the claims, which are incorporated herein by reference. A more complete understanding of embodiments of this disclosure and the implementation of its additional advantages will be given to those skilled in the art by considering the detailed description of one or more of the following examples. Reference will be made to the accompanying drawings, which will first be briefly described. Attached Figure Description
[0006] A better understanding of the aspects and advantages of this disclosure can be achieved by referring to the following accompanying drawings and the following detailed description. It should be understood that similar reference numerals are used to identify similar elements illustrated in one or more of the drawings, wherein the illustrations in the drawings are for illustrative purposes and not for limiting the embodiments of this disclosure. The components in the drawings are not necessarily to scale, but rather the emphasis is on clearly illustrating the principles of this disclosure.
[0007] Figure 1 An example operating environment of an audio processing device according to one or more embodiments of the present disclosure is illustrated.
[0008] Figure 2 This is a block diagram of an example audio processing device according to one or more embodiments of the present disclosure.
[0009] Figure 3 This is a block diagram of an example audio signal processor according to one or more embodiments of the present disclosure.
[0010] Figure 4 The illustration depicts an example system architecture that provides robust speaker localization in the presence of strong noise interference, according to one or more embodiments.
[0011] Figure 5 This is a flowchart illustrating an example process for performing real-time audio signal processing according to one or more embodiments of the present disclosure. Detailed Implementation
[0012] This paper discloses systems and methods for detecting and enhancing target audio in noisy environments.
[0013] In various embodiments, a microphone array with multiple microphones senses target audio and noise in the operating environment and generates an audio signal for each microphone. Speaker localization using a microphone array in the form of Time Difference of Arrival (TDOA) or Direction of Arrival (DOA) is a well-known problem in far-field speech processing applications, including those where estimating the physical orientation of a speaker relative to an array is of interest in applications such as surveillance, human-computer interaction, camera manipulation, etc., and applications where estimating and tracking the positional information of one or more speakers results in one or more Voice Activity Detectors (VADs) that can supervise speaker enhancement and noise reduction tasks in methods such as beamforming or blind source separation (BSS).
[0014] This disclosure describes a system and method for robustly estimating the TDOA / DOA of one or more coexisting loudspeakers when a strong primary noise / interference source (e.g., loud television noise) is always present. In some embodiments, the system operates by employing some characteristics of a generalized eigenvalue (GEV) beamformer, which creates conditions for estimating the unique spatial fingerprint or relative transfer function (RTF) of the target loudspeaker. The target RTF is estimated by effectively zeroing out the primary noise source. By applying a modified TDOA / DOA estimation method using the RTF as input, the system described herein can obtain a robust localization estimate of the target loudspeaker. If multiple target loudspeakers are active in the presence of a stronger noise source (e.g., a noise source stronger than the target loudspeaker), the RTF of each source can be estimated intermittently and fed to a multi-source tracker using appropriate tuning, resulting in robust VAD for each source that can drive a multi-stream speech enhancement system.
[0015] This disclosure provides numerous advantages over conventional systems and methods. TDOA / DOA methods typically operate by taking the spatial correlation matrix of the raw input obtained from a microphone array and then scanning all possible directions / delays to form a pseudo-likelihood, where the peak(s) correspond to the TDOA / DOA of the source(s). These methods are suitable when a single source is present, or if multiple sources are present, their power is roughly at the same level. However, these methods fail when the target loudspeaker is masked in the presence of strong noise or interference sources, for example, when the signal-to-noise ratio (SNR) is negative, because the peak corresponding to the weaker target speech is not well distinguished or completely vanished relative to the peak corresponding to the stronger noise source. In various embodiments, the method proposed herein uses a modified TDOA / DOA estimation method that uses the estimated target RTF as input instead of the spatial correlation matrix of the raw microphone array signal. Since the RTF is estimated by effectively zeroing out the main noise sources, it contains less distorted spatial information of the target speech than the noisy raw microphone array correlation matrix, thus an improved localization estimate of the target loudspeaker can be obtained.
[0016] This disclosure can be used in conjunction with beamforming techniques incorporating generalized feature vector tracking (GEV) to enhance target audio in received audio signals. In one or more embodiments, a multi-channel audio input signal is received via an array of audio sensors (e.g., microphones). Each audio channel is analyzed to determine the presence of target audio, e.g., whether a target person is actively speaking. The system tracks the target and noise signals to determine the location of the target audio source (e.g., the target person) relative to the microphone array. The direction of the target audio can be determined in real time using a modified GEV process. The determined direction can then be used by a spatial filtering process (such as a minimum variance distortion-free response (MVDR) beamformer) to enhance the target audio. After processing the audio input signal, the enhanced audio output signal can be used, for example, as audio output transmitted to one or more speakers, as voice communication in a telephone or Voice over IP (VoIP) call, for timbre recognition or voice command processing, or other voice applications. The modified GEV system can be used to efficiently determine the direction of a target audio source in real time, regardless of whether the geometry of the microphone array or the audio environment is known.
[0017] Figure 1 An example operating environment 100 according to various embodiments of the present disclosure is illustrated, in which an audio processing system can operate. Operating environment 100 includes an audio processing device 105, a target audio source 110, and one or more noise sources 135-145. Figure 1 In the example illustrated, the operating environment 100 is depicted as a room; however, it is contemplated that the operating environment may include other areas, such as a vehicle interior, an office meeting room, a family room, an outdoor stadium, or an airport. According to various embodiments of this disclosure, the audio processing device 105 may include two or more audio sensing components 115a-115d (e.g., microphones), and optionally one or more audio output components 120a-120b, such as one or more loudspeakers.
[0018] Audio processing device 105 can be configured to sense sound via audio sensing components 115a-115d and generate multi-channel audio input signals, including two or more audio input signals. Audio processing device 105 can use audio processing techniques disclosed herein to process the audio input signals to enhance the audio signal received from target audio source 110. For example, the processed audio signal can be transmitted to other components within audio processing device 105 (such as a timbre recognition engine or voice command processor) or to an external device. Therefore, audio processing device 105 can be a standalone device for processing audio signals, or a device that transforms the processed audio signal into other signals (e.g., commands, instructions, etc.) for interacting with or controlling external devices. In other embodiments, audio processing device 105 can be a communication device, such as a mobile phone or a Voice over IP (VoIP) enabled device, and the processed audio signal can be transmitted over a network to another device for output to a remote user. The communication device can also receive processed audio signals from a remote device and output the processed audio signals via audio output components 120a-120b.
[0019] The target audio source 110 can be any source that generates target audio that can be detected by the audio processing device 105. The target audio can be defined based on criteria specified by a user or system requirement. For example, the target audio can be defined as a human voice, a sound produced by a specific animal, or a machine. In the illustrated example, the target audio is defined as a human voice, and the target audio source 110 is a person. In addition to the target audio source 110, the operating environment 100 may include one or more noise sources 135-145. In various embodiments, sounds that are not the target audio are processed as noise. In the illustrated example, noise sources 135-145 may include a loudspeaker 135 playing music, a television 140 playing a television program, movie, or sporting event, and background dialogue between non-target speakers 145. It will be understood that other noise sources may be present in various operating environments.
[0020] Note that the target audio and noise can reach the microphones 115a-115d of the audio processing device 105 from different directions. For example, noise sources 135-145 may generate noise at different locations within the room, and the target audio source 110 (e.g., a person) may be speaking while moving between locations within the room. Furthermore, the target audio and / or noise may be reflected from fixed objects (e.g., walls) within the room. For example, consider the path that the target audio can traverse from the target audio source 110 to reach each of the microphones 115a-115d. As indicated by arrows 125a-125d, the target audio can travel directly from the target audio source 110 to the microphones 115a-115d, respectively. Additionally, the target audio can be reflected from walls 150a and 150b and indirectly reach the microphones 115a-115d from the target audio source 110, as indicated by the arrows. According to various embodiments of the present disclosure, the audio processing device 105 may use the audio processing techniques disclosed herein to estimate the location of the target audio source 110 based on the audio input signal received by the microphones 115a-115d, and process the audio input signal to enhance the target audio and suppress noise based on the estimated location.
[0021] Figure 2 An example audio processing device 200 according to various embodiments of the present disclosure is illustrated. In some embodiments, the audio processing device 200 may be implemented as Figure 1 The audio processing device 105. The audio processing device 200 includes an audio sensor array 205, an audio signal processor 220, and a host system component 250.
[0022] The audio sensor array 205 includes two or more sensors, each of which can be implemented as a transducer that converts audio input in the form of sound waves into an audio signal. In the illustrated environment, the audio sensor array 205 includes a plurality of microphones 205a-205n, each microphone generating an audio input signal that is provided to the audio input circuitry 222 of the audio signal processor 220. In one embodiment, the sensor array 205 generates a multi-channel audio signal, wherein each channel corresponds to an audio input signal from one of the microphones 205a-n.
[0023] Audio signal processor 220 includes audio input circuitry 222, digital signal processor 224, and optional audio output circuitry 226. In various embodiments, audio signal processor 220 may be implemented as an integrated circuit including analog circuitry, digital circuitry, and digital signal processor 224, operable to execute program instructions stored in firmware. For example, audio input circuitry 222 may include an interface to audio sensor array 205, anti-aliasing filters, analog-to-digital converter circuitry, echo cancellation circuitry, and other audio processing circuitry and components as disclosed herein. Digital signal processor 224 is operable to process multi-channel digital audio signals to generate enhanced audio signals, which are output to one or more host system components 250. In various embodiments, digital signal processor 224 is operable to perform echo cancellation, noise cancellation, target signal enhancement, post-filtering, and other audio signal processing functions.
[0024] Optional audio output circuitry 226 processes the audio signals received from digital signal processor 224 for output to at least one speaker, such as speakers 210a and 210b. In various embodiments, audio output circuitry 226 may include a digital-to-analog converter that converts one or more digital audio signals into analog signals and one or more amplifiers for driving speakers 210a-210b.
[0025] The audio processing device 200 can be implemented as any device operable to receive and enhance target audio data, such as, for example, a mobile phone, smart speaker, tablet computer, laptop computer, desktop computer, voice-controlled appliance, or automobile. The host system component 250 may include various hardware and software components for operating the audio processing device 200. In the illustrated embodiment, the system component 250 includes a processor 252, a user interface component 254, a communication interface 256 for communicating with external devices and networks (such as network 280 (e.g., the Internet, cloud, local area network, or cellular network) and mobile device 284), and a memory 258.
[0026] Processor 252 and digital signal processor 224 may include one or more of the following: processor, microprocessor, single-core processor, multi-core processor, microcontroller, programmable logic device (PLD) (e.g., field-programmable gate array (FPGA)), digital signal processing (DSP) device, or other logic device that may be configured to perform the various operations discussed herein with respect to embodiments of this disclosure by means of hardwire, execution of software instructions, or a combination of both. Host system component 250 is configured to interface and communicate with audio signal processor 220 and other system components 250, such as via a bus or other electronic communication interface.
[0027] It will be understood that, although the audio signal processor 220 and the host system component 250 are shown as comprising a combination of hardware components, circuitry, and software, in some embodiments, at least some or all of the functionality that the hardware components and circuitry are operable to perform may be implemented as software modules executed by the processor 252 and / or the digital signal processor 224 in response to software instructions and / or configuration data stored in the firmware of the memory 258 or the digital signal processor 224.
[0028] Memory 258 may be implemented as one or more memory devices operable to store data and information, including audio data and program instructions. Memory 258 may include one or more different types of memory devices, including volatile and non-volatile memory devices, such as RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Read-Only Memory), flash memory, hard disk drive, and / or other types of memory.
[0029] Processor 252 may be operable to execute software instructions stored in memory 258. In various embodiments, voice recognition engine 260 is operable to process enhanced audio signals received from audio signal processor 220, including identifying and executing voice commands. Voice communication component 262 may be operable to facilitate voice communication with one or more external devices, such as mobile device 284 or user equipment 286, such as via voice calls over mobile or cellular telephone networks or VoIP calls over IP networks. In various embodiments, voice communication includes transmitting enhanced audio signals to external communication devices.
[0030] User interface component 254 may include a display, a touchpad display, a keypad, one or more buttons, and / or other input / output components operable to enable a user to interact directly with audio processing device 200.
[0031] Communication interface 256 facilitates communication between audio processing device 200 and external devices. For example, communication interface 256 may enable Wi-Fi (e.g., 802.11) or Bluetooth connectivity between audio processing device 200 and one or more local devices (such as mobile device 284) or a wireless router that provides network access to remote server 282 via network 280. In various embodiments, communication interface 256 may include other wired and wireless communication components that facilitate direct or indirect communication between audio processing device 200 and one or more other devices.
[0032] Figure 3Exemplary audio signal processors 300 according to various embodiments of the present disclosure are illustrated. In some embodiments, the audio input processor 300 is implemented as one or more integrated circuits, which include components derived from a digital signal processor (such as...) Figure 2 The digital signal processor 224 implements analog and digital circuitry and firmware logic. As illustrated, the audio signal processor 300 includes an audio input circuit 315, a sub-band frequency analyzer 320, a target activity detector 325, a target enhancement engine 330, and a synthesizer 335.
[0033] The audio signal processor 300 receives multi-channel audio input from multiple audio sensors, such as a sensor array 305 including at least two audio sensors 305a-n. The audio sensors 305a-305n may be integrated with an audio processing device (such as...) Figure 2 The audio processing device 200 may include an integrated microphone or an external component connected thereto. For the audio input processor 300 according to various embodiments of this disclosure, the arrangement of the audio sensors 305a-305n may be known or unknown.
[0034] The audio signal can initially be processed by audio input circuitry 315, which may include an anti-aliasing filter, an analog-to-digital converter, and / or other audio input circuitry. In various embodiments, audio input circuitry 315 outputs a digital, multi-channel, time-domain audio signal with N channels, where N is the number of sensor (e.g., microphone) inputs. The multi-channel audio signal is input to sub-band frequency analyzer 320, which divides the multi-channel audio signal into consecutive frames and decomposes each frame of each channel into multiple frequency sub-bands. In various embodiments, sub-band frequency analyzer 320 includes a Fourier transform process and outputs multiple frequency windows (bins). The decomposed audio signal is then provided to target activity detector 325 and target enhancement engine 330.
[0035] Target activity detector 325 is operable to analyze one or more frames in an audio channel and generate a signal indicating the presence of a target audio in the current frame. As discussed above, the target audio can be any audio to be identified by the audio system. When the target audio is a human voice, target activity detector 325 can be implemented as a speech activity detector. In various embodiments, a speech activity detector operable to receive audio data and make a determination about the presence or absence of a target audio can be used. In some embodiments, target activity detector 325 can apply a target audio classification rule to a sub-band frame to calculate a value. This value is then compared to a threshold to generate a target activity signal. In various embodiments, the signal generated by target activity detector 325 is a binary signal, such as an output of "1" indicating the presence of a target voice in a sub-band audio frame and a binary output of "0" indicating the absence of a target voice in a sub-band audio frame. The generated binary output is provided to target enhancement engine 330 for further processing of the multi-channel audio signal. In other embodiments, the target activity signal may include a probability of target presence, an indication that a determination of target presence cannot be made, or other target presence information as required by the system.
[0036] The target enhancement engine 330 receives subband frames from the subband frequency analyzer 320 and target activity signals from the target activity detector 325. According to various embodiments of this disclosure, the target enhancement engine 330 uses a modified generalized eigenvalue beamformer to process the subband frames based on the received activity signals, as will be described in more detail below. In some embodiments, processing the subband frames includes estimating the position of a target audio source (e.g., target audio source 110) relative to the sensor array 305. Based on the estimated position of the target audio source, the target enhancement engine 330 can enhance portions of the audio signal determined to originate from the target audio source and suppress other portions of the audio signal determined to be noise.
[0037] After enhancing the target audio signal, the target enhancement engine 330 can pass the processed audio signal to the synthesizer 335. In various embodiments, the synthesizer 335 reconstructs one or more of the multi-channel audio signals on a frame-by-frame basis by combining sub-bands to form an enhanced temporal audio signal. The enhanced audio signal can then be transformed back to the temporal domain and sent to system components or external devices for further processing.
[0038] Figure 4An example system architecture 400 providing robust speaker localization in the presence of strong noise interference is illustrated according to one or more embodiments. The system architecture 400 according to various embodiments of this disclosure can be implemented as a combination of digital circuitry and logic executed by a digital signal processor. System architecture 400 includes a sub-band analysis block 410, a covariance calculation block 420, an input speech activity detector (input VAD 450) making timbre / non-timbre determination, a target timbre relative transfer function estimate using a feature analysis module (RTF estimation module 430), and a modified covariance-based localization module 440.
[0039] The input VAD 450 drives the RTF estimation module 430 and operates to identify (with high confidence) the moments when non-speech-like noise is isolated. In other words, the input VAD 450 can be tuned to produce fewer false negatives (instances where the timbre is not present but the VAD incorrectly claims it is valid) than false positives (instances where the timbre is valid but the VAD incorrectly claims it is not). The reason for this relates to the process performed by the RTF estimation module 430, which will be discussed below.
[0040] During RTF estimation, VAD 450 is used to compute noise-only and noisy timbre covariance matrices. The noise covariance matrix is used to zero out the noise, but because the noise covariance is incorrectly updated using the actual timbre-like features during false negatives, it may ultimately cancel out the timbre, which reduces the accuracy of the target timbre's RTF and makes it less reliable for localization. Since we are dealing with a low SNR, the conventional power-based features used for VAD construction do not perform as expected, and instead rely on a spectrum-based classifier trained to distinguish between timbre and non-timbre audio.
[0041] This disclosure overcomes these limitations regarding conventional systems. Reference will now be made to... Figure 4 and Figure 5 Describe example process 500. In the following discussion, we use v(l) to indicate the state variable of VAD 450, which is defined as having a hypothetical value of “1” or “0” for the observed frame l if it is determined that the timbre is present or absent.
[0042] 1) RTF estimation using feature analysis
[0043] For a total of M microphones, we use x m (t), m = 1, ..., M indicates the sampled time-domain audio signal recorded at the m-th microphone. Through sub-band analysis 410, the signal is transformed to be represented as X mThe time-frequency domain of (k, l), where k indicates the frequency band index and l indicates the sub-band time index. We represent the vector of the entire array's time-frequency band as...
[0044] X(k,l)=[X1(k,l),...,X M (k, l)] T .
[0045] Next, in step 505, the covariance of the noise-only timbre and the noise-only timbre are calculated using the processing block. The covariance of the noise-only segment of the input VAD is calculated in block 420 as follows.
[0046]
[0047] In the online implementation, the covariance matrix is updated using frame l and can be estimated using first-order recursive smoothing.
[0048] P n (k, l) = α(l)P n (k,l-1)+(1-α(l))X(k,l)X(k,l) H
[0049] Where α(l) = max(γ, v(l)), and γ is the smoothing constant (<1). Similarly, the covariance of a noisy signal can be calculated as follows:
[0050] P x (k, l) = β(l)P x (k, l-1)+(1-β(l))X(k, l) H
[0051] Where β(l) = max(γ, 1 - v(l)). In step 510, the RTF of the target timbre source, represented by h(k, l), is found using a blind acoustic beamforming process based on general eigenvalue decomposition. The main feature vector is obtained (processing block 430).
[0052] The choice of parameter γ determines the rate at which this covariance matrix is updated. If multiple speakers are operating simultaneously in the presence of strong background noise sources, a faster rate may be required for P... x It can capture intermittent switching between target sources, which can later be fed into the tracker.
[0053] 2) Covariance-based localization
[0054] A microphone is selected as a reference microphone (e.g., a first microphone) and used to extract TDOA information of one or more sources relative to the reference microphone. The TDOA estimation method used here is based on a manipulated minimum variance (STMV) beamformer, and will now be described.
[0055] In step 520, the manipulation matrix is constructed for each frequency band as follows.
[0056]
[0057] Where τ m This refers to the TDOA of the m-th microphone relative to the first microphone for different scans (linear scans or scans corresponding to different azimuth and elevation angles), where τ = [τ2...τ]. M ] T , and f k It is the frequency at k.
[0058] When the timbre is valid (step 530), v(l) = 1, and instead of using the noisy timbre space covariance matrix (P... x (k,l) is used as the input covariance of the STMV algorithm (step 540). In step 550, we calculate h(k,l)h(k,l). H We use RTF-based covariance. We note that since the effect of noise is essentially nullified in the estimation of RTF h(k,l), we can examine this spatial covariance derived from the timbre-only component. Furthermore, when timbre is absent, v(l) = 0, and we can skip this calculation and use the previous estimate. Therefore, we make the timbre-only covariance matrix as...
[0059] P RTF (k, l) = (1 - v(l))P RTF (k,l-1)+v(l)h(k,l)h(k,l) H
[0060] Next, in step 560, the system constructs the directional covariance matrix coherently aligned across all frequency bands as follows:
[0061]
[0062] Next, in step 570, we find the direction that minimizes the beam power under the distortion-free criterion, where its equivalent pseudo-likelihood solution becomes
[0063]
[0064] Where 1 = [1, ..., 1] T .
[0065] In step 580, the TDOA that generates the maximum likelihood is then selected, denoted as
[0066]
[0067] The foregoing disclosure is not intended to limit the invention to the precise forms disclosed or to any particular field of use. Therefore, it is contemplated that various alternative embodiments and / or modifications to this disclosure are possible, whether expressly described or implied herein. Embodiments of this disclosure have been so described that those skilled in the art will recognize advantages over conventional methods, and changes in form and detail may be made without departing from the scope of this disclosure. Therefore, this disclosure is limited only by the claims.
Claims
1. A method for improved audio signal processing, comprising: receiving a multi-channel audio signal from a plurality of audio input components; determining whether the multi-channel audio signal includes a target audio associated with an audio source; estimating a relative position of the audio source with respect to the plurality of audio input components using a timbre-only covariance matrix based on the multi-channel audio signal and the determination of whether the multi-channel audio signal includes the target audio, wherein, when timbre is valid, calculating and using a covariance based on a target timbre relative transfer function as a timbre-only covariance matrix for a current frame, when timbre is not valid, using a timbre-only covariance matrix determined for a previous frame as the timbre-only covariance matrix for the current frame; and processing the multi-channel audio signal by enhancing the target audio in the multi-channel audio signal based on the estimated relative position to generate an audio output signal.
2. The method of claim 1, further comprising transforming the multi-channel audio signal into sub-band frames according to a plurality of frequency sub-bands, wherein the estimating the relative position of the audio source is further based on the sub-band frames.
3. The method of claim 1, further comprising calculating a noisy timbre covariance and a noise-only covariance.
4. The method of claim 1, further comprising estimating the target timbre relative transfer function using a feature analysis process.
5. The method of claim 1, further comprising calculating a covariance-based localization to identify a time difference of arrival.
6. The method of claim 1, further comprising determining whether an input audio frame is a timbre frame or a non-timbre frame.
7. The method of claim 1, further comprising constructing a steering matrix for each of a plurality of frequency sub-bands using one of the audio input components as a reference.
8. The method of claim 1, further comprising constructing a directional covariance matrix that is coherently aligned across all frequency sub-bands.
9. The method of claim 1, further comprising determining a direction that minimizes beam power under a distortionless norm condition; and picking a time difference of arrival that produces a maximum likelihood of the target audio associated with the audio source.
10. A system for improved audio signal processing, comprising: a plurality of audio input components configured to generate a plurality of audio input signals; a logic device configured to: receive the plurality of audio input signals; determine whether the plurality of audio input signals includes a target audio associated with an audio source; estimate a relative position of the audio source with respect to the plurality of audio input components using a timbre-only covariance matrix based on the plurality of audio input signals and the determination of whether the plurality of audio input signals includes the target audio, wherein, when timbre is valid, calculate and use a covariance based on a target timbre relative transfer function as a timbre-only covariance matrix for a current frame, when timbre is not valid, use a timbre-only covariance matrix determined for a previous frame as the timbre-only covariance matrix for the current frame; and process the plurality of audio input signals by enhancing the target audio based on the estimated relative position to generate an audio output signal. 11. The system of claim 10, wherein the logic device is further configured to transform the plurality of audio input signals into subband frames according to a plurality of frequency subbands, wherein the estimated relative positions of the audio sources are further based on the subband frames.
12. The system of claim 10, wherein the logic device is further configured to compute a noisy sound color covariance and a noise only covariance.
13. The system of claim 10, wherein the logic device is further configured to estimate the target sound color relative transfer function using a feature analysis process.
14. The system of claim 10, wherein the logic device is further configured to compute a covariance based localization to identify a time difference of arrival.
15. The system of claim 10, wherein the logic device is further configured to determine whether an input audio frame is a sound color frame or a non-sound color frame.
16. The system of claim 10, wherein the logic device is further configured to build a steering matrix for each of a plurality of frequency subbands using one of the audio input components as a reference.
17. The system of claim 10, wherein the logic device is further configured to build a directional covariance matrix that is coherently aligned across all frequency subbands.
18. The system of claim 10, wherein the logic device is further configured to determine a direction that minimizes beam power under a distortionless norm; and pick a time difference of arrival that produces a maximum likelihood of the target audio associated with the audio source.
Citation Information
Patent Citations
A microphone system and a hearing device comprising same
CN109040932A
Voice enhancement in audio signals through modified generalized eigenvalue beamformer
US20190172450A1