Generating restored spatial audio signal for occluded microphone
By detecting microphone occlusion events in multi-microphone devices and selecting appropriate spatial filters using occlusion information and DOA information, the problem of degradation in spatial audio recording quality caused by occlusion is solved, and effective recovery of the occluded channel and improvement of spatial audio recording quality is achieved.
Patent Information
- Application Number
- CN202380072809.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-20
- Filing Date
- 2023-10-04
- Publication Date
- 2025-05-27
AI Technical Summary
In a multi-microphone device, when one or more microphones are blocked, the quality of spatial audio recording will be degraded, and even damage or degraded spatial audio recording, thereby preventing the generation of spatial audio recordings.
By detecting microphone occlusion events, using the occlusion information and estimated direction of arrival (DOA) information of the sound source, an appropriate spatial filter is selected to restore the spatial audio output of the occluded channel.
The spatial audio signal associated with the obstructed microphone is effectively restored, the quality of spatial audio recording is improved, and an immersive spatial audio experience can still be generated under occlusion.
Smart Images

Figure CN120052005A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to audio signal processing. For example, aspects of the present disclosure relate to recovering spatial audio during one or more microphone occlusion events. Background Art
[0002] The spatialized audio rendering system outputs sounds that enable a user to perceive a three-dimensional (3D) audio space. Spatial audio (also referred to as three-dimensional or 3D audio) may refer to a variety of sound playback techniques that make it possible for a listener to perceive sounds around them without the need for multiple speaker setups. For example, spatial audio techniques may enable a listener to perceive three-dimensional sounds (e.g., spatial audio) based on simulating the acoustic interaction between real-world sound waves and the listener's ears. The interaction between sound waves and auditory anatomy (including the shape of the ears and head) may be used to provide spatial audio to a listener. For example, one or more human-related transfer functions (HRTFs) or other spatial sound filters may be used to enable a user to perceive a 3D audio space.
[0003] For example, a user may wear headphones, an augmented reality (AR) head mounted display (HMD), or a virtual reality (VR) HMD, and movement of at least a portion of the user (e.g., translational or rotational motion) may cause the perceived direction or distance of the sound to change. For example, a user may navigate from a first location in a visual (e.g., virtualized) environment to a second location in the visual environment. At the first location, the stream is in front of the user in the visual environment, and at the second location, the stream is to the right of the user in the visual environment. When the user navigates from the first location to the second location, the sound output by the spatialized audio rendering system may change so that the user perceives the sound of the stream as coming from the right side of the user rather than from the front of the user.
[0004] In order to render or provide an accurate and immersive spatial audio experience to a listener, high-quality and accurate spatial audio recordings are generally required. For example, a spatial audio recording may be captured using multiple microphones that allow spatial information to be captured with or otherwise determined from the original audio data. The spatial information may include the direction of arrival (DOA) of a particular sound, the time difference of arrival (ATD) of a given sound at different microphone positions, the difference in arrival level of a given sound at different microphone positions, etc. Summary of the invention
[0005] In some examples, systems and techniques for performing audio signal processing are described. For example, these systems and techniques may perform audio signal processing to generate restored or reconstructed spatial audio associated with a pair or group of microphones including at least one non-occluded microphone and at least one occluded microphone. According to at least one illustrative example, a device for generating a spatial audio recording is provided, which includes a memory (e.g., configured to store data, such as audio data, one or more audio frames, etc.) and one or more processors coupled to the memory (e.g., implemented in a circuit). The one or more processors are configured and may: detect occlusion of at least one of one or more audio frames associated with the spatial audio recording; and during the spatial audio recording, select between performing at least one of occluded spatial filtering of the one or more audio frames or non-occluded spatial filtering of the one or more audio frames based on detecting occlusion of the at least one audio frame.
[0006] In another illustrative example, a method of performing spatial audio recording is provided, the method comprising: detecting occlusion of at least one audio frame of one or more audio frames associated with the spatial audio recording; and during the spatial audio recording, selecting between performing at least one of occluded spatial filtering of the one or more audio frames or non-occluded spatial filtering of the one or more audio frames based on detecting the occlusion of the at least one audio frame.
[0007] In another example, a non-transitory computer-readable medium having instructions stored thereon is provided that, when executed by one or more processors, causes the one or more processors to: detect occlusion of at least one of one or more audio frames associated with a spatial audio recording; and during the spatial audio recording, select between performing at least one of occluded spatial filtering of the one or more audio frames or unoccluded spatial filtering of the one or more audio frames based on detecting occlusion of the at least one audio frame.
[0008] In another example, an apparatus for spatial audio recording is provided. The apparatus includes: means for detecting occlusion of at least one of one or more audio frames associated with the spatial audio recording; and means for selecting between performing at least one of occluded spatial filtering of the one or more audio frames or non-occluded spatial filtering of the one or more audio frames based on detecting occlusion of the at least one audio frame during the spatial audio recording.
[0009] In another example, an apparatus for generating a spatial audio recording is provided, comprising a memory (e.g., configured to store data, such as audio data, one or more audio frames, etc.) and one or more processors (e.g., implemented in a circuit) coupled to the memory. The one or more processors are configured and may: detect occlusion of one or more microphones associated with at least one of one or more audio frames associated with the spatial audio recording; and select at least one of an occluded spatial filter for the one or more audio frames or a non-occluded spatial filter for the one or more audio frames based on the detection of the occlusion.
[0010] In another illustrative example, a method of performing spatial audio recording is provided, the method comprising: detecting occlusion of one or more microphones associated with at least one audio frame of one or more audio frames associated with the spatial audio recording; and selecting at least one of an occluded spatial filter for the one or more audio frames or a non-occluded spatial filter for the one or more audio frames based on the detection of the occlusion.
[0011] In another example, a non-transitory computer-readable medium having instructions stored thereon is provided that, when executed by one or more processors, causes the one or more processors to: detect occlusion of one or more microphones associated with at least one of one or more audio frames associated with a spatial audio recording; and select at least one of an occluded spatial filter for the one or more audio frames or an unoccluded spatial filter for the one or more audio frames based on the detection of the occlusion.
[0012] In another example, an apparatus for spatial audio recording is provided. The apparatus includes: means for detecting occlusion of one or more microphones associated with at least one of one or more audio frames associated with the spatial audio recording; and selecting at least one of an occluded spatial filter for the one or more audio frames or a non-occluded spatial filter for the one or more audio frames based on the detection of the occlusion.
[0013] In some aspects, one or more of the devices described herein are, are part of, and / or include a mobile device or wireless communication device (e.g., a mobile phone or other mobile device), an extended reality (XR) device or system (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a wearable device (e.g., a network-connected watch or other wearable device), a camera, a personal computer, a laptop computer, a vehicle or a computing device or component of a vehicle, a server computer or server device, another device, or a combination thereof. In some aspects, the device includes one or more cameras for capturing one or more images. In some aspects, the device also includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the above-mentioned device may include one or more sensors (e.g., one or more inertial measurement units (IMUs), such as one or more gyroscopes, one or more gyrometers, one or more accelerometers, any combination thereof, and / or other sensors).
[0014] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used alone to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any or all drawings, and each claim.
[0015] The foregoing and other features and aspects will become more apparent upon reference to the following description, claims and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Illustrative aspects of the present application are described in detail below with reference to the following drawings:
[0017] Figure 1 is a diagram illustrating an example of a multi-microphone handset according to some examples;
[0018] Figure 2A is a diagram illustrating an example far-field model of plane wave propagation relative to a microphone pair according to some examples;
[0019] Figure 2B is a diagram illustrating examples of microphone placement of multiple microphone pairs in a linear array according to some examples;
[0020] Figure 3 is a diagram illustrating the use of a four-microphone array for omnidirectional and first-order capture for spatial coding according to some examples;
[0021] Figure 4A is a diagram illustrating an example spatial recording system that receives a multi-microphone input including unobstructed microphone signals according to some examples;
[0022] Figure 4B is a diagram illustrating an example spatial recording system that receives a multi-microphone input including at least one occluded microphone signal according to some examples;
[0023] Figure 5 is a diagram illustrating an example spatial recording system with occlusion handling according to some examples;
[0024] Figure 6 is a diagram of an example spatial record occlusion handling system according to some examples;
[0025] Figure 7 is a diagram of an example occlusion detection engine according to some examples;
[0026] Figure 8 is a diagram of an example spatial device selection engine according to some examples;
[0027] Fig. 9 is a diagram of an example spatial device selection engine receiving a dual channel input signal according to some examples;
[0028] Fig.10 is a diagram of an example spatial device selection engine receiving a three channel or greater input signal according to some examples;
[0029] Fig.11 is a diagram of an example selective spatial filtering engine according to some examples;
[0030] Fig.12 is a diagram of an example selective spatial filtering engine receiving a dual channel input signal according to some examples;
[0031] Fig.13 is a diagram of an example selective spatial filtering engine receiving three-channel or more input signals according to some examples;
[0032] Fig.14A is a diagram illustrating an example of selective spatial filtering based on three non-occluded microphone input signals according to some examples;
[0033] Fig. 14B is a diagram illustrating an example of selective spatial filtering based on restoration of two unoccluded microphone input signals and one occluded microphone input signal according to some examples;
[0034] Fig.15 is a diagram illustrating examples of spatial filter selection associated with a non-occluded microphone input signal and restored spatial filter selection associated with an occluded microphone input signal and a non-occluded microphone input signal according to some examples;
[0035] Fig.16 is a flowchart illustrating an example of a process for audio signal processing according to some examples; and
[0036] Fig.17 is a block diagram illustrating an example of a computing system according to some examples. DETAILED DESCRIPTION
[0037] Some aspects and aspects of the present disclosure are provided below. Some of these aspects and aspects can be applied independently, and some of them can be applied in combination, which is obvious to those skilled in the art. In the following description, specific details are set forth for explanation purposes to provide a thorough understanding of various aspects of the application. However, it will be apparent that various aspects can be implemented without these specific details. Each drawing and description are not intended to be restrictive.
[0038] The following description provides only exemplary aspects and is not intended to limit the scope, applicability or configuration of the present disclosure. On the contrary, the following description of the exemplary aspects will provide a description that can be used to implement the exemplary aspects to those skilled in the art. It should be understood that various changes can be made to the function and arrangement of elements without departing from the essence and scope of the present application as set forth in the appended claims.
[0039] Unless the context indicates otherwise, references to the "position" of a microphone of a multi-microphone audio sensing device indicate the location of the center of the acoustically sensitive surface of the microphone. Depending on the particular context, the term "channel" is sometimes used to indicate a signal path and at other times to indicate a signal carried by such a path. Unless otherwise indicated, the term "series" is used to indicate a sequence of two or more items. The term "logarithm" is used to indicate a logarithm with a base of ten, but extensions of this operation to other bases are within the scope of the present disclosure. The term "frequency component" is used to indicate one frequency or frequency band in a set of frequencies or frequency bands of a signal, such as a sample of a frequency domain representation of a signal (e.g., as produced by a fast Fourier transform) or a subband of a signal (e.g., a Bark scale or Mel scale subband).
[0040] Spatialized audio refers to capturing and reproducing audio signals in a manner that preserves or simulates the positional information of audio sources in an audio scene (e.g., a 3D audio space). For illustration, when listening to the playback of a spatial audio signal, the listener can perceive the relative positions of various audio sources relative to each other and relative to the listener in the audio scene. One format for creating and replaying spatial audio signals is channel-based. In channel-based audio, the loudspeaker feed is adjusted to create a reproduction of the audio scene. Another format for spatial audio signals is object-based audio. In object-based audio, audio objects are used to create spatial audio signals. Each audio object is associated with a 3D coordinate (and other metadata), and the audio object is simulated on the playback side to create the listener's perception that the sound is originating from a specific position of the audio object. An audio scene can be composed of several audio objects. Object-based audio is used in multiple systems (including video game systems). Higher order ambisonics (HOA) is another format for spatialized audio signals. HOA is used to capture, transmit, and render spatial audio signals. HOA represents the entire sound field in a compact and accurate manner and is intended to recreate the actual sound field of the capture location at the playback location (e.g., at the audio output device). The HOA signal enables the listener to experience the same audio spatialization as the listener would experience at the actual scene. In each of the above formats (e.g., channel-based audio, object-based audio, and HOA-based audio), multiple transducers (e.g., loudspeakers) are used for audio playback. If the audio playback is output by headphones, additional processing (e.g., binauralization) is performed to generate audio signals that "trick" the listener's brain into thinking that the sound is actually coming from different points in space rather than from the transducers in the headphones.
[0041] In one illustrative example, spatial audio (also referred to as "3D" or "360" audio) may refer to a variety of sound playback techniques that allow listeners to perceive sounds all around them. For example, unlike stereo and surround sound audio formats (e.g., such as 5.1 or 7.1 surround sound) that depict audio in two dimensions and are tied to a specific multi-speaker setup, spatial audio may be used to depict audio in three dimensions (e.g., may introduce a height dimension) without a multi-speaker setup dependency.
[0042] In some cases, spatial audio techniques may enable a listener to perceive three-dimensional sound (e.g., spatial audio) based on simulating the acoustic interaction between real-world sound waves and the user's ears. For example, the interaction between sound waves and auditory anatomy (including the shape of the ears and head) may be used to provide spatial audio to the listener. In some cases, the simulation may be based on one or more human-related transfer functions (HRTFs) and / or various other spatial sound filters (e.g., also referred to as "spatial filters").
[0043] In some examples, in order to render or provide an accurate and immersive spatial audio experience to a listener, a high-quality and accurate spatial audio recording may be required. For example, a spatial audio recording may be captured using multiple microphones that allow spatial information to be captured with or otherwise determined from the original audio data. The spatial information may include the direction of arrival (DOA) of a particular sound or sound signal, the time difference of arrival (ATD) of a given sound at different microphone positions, the difference in arrival level of a given sound at different microphone positions, etc.
[0044] Spatial audio recording may be performed using a mobile device equipped with multiple microphones (e.g., such as a smartphone or other multi-microphone device) or otherwise capable of providing a multi-microphone audio input to a spatial recording audio processor. In some cases, when a mobile device is used for spatial audio recording, one or more microphones may be occluded by a user's hand or finger. Microphone occlusion may reduce the quality of the resulting spatial audio recording, destroy or degrade the resulting spatial audio recording, and / or prevent the generation of a spatial audio recording. For example, a smartphone may include multiple microphones that are each recessed into the body of the smartphone (e.g., may include microphones that are seated in recessed openings included on an outer surface or body of the smartphone). When a user holds the smartphone and attempts to capture a spatial audio recording, the user's hand or finger may block, cover, or rub one or more of the microphone openings. Such occlusion may result in degradation of audio quality and / or loss of a spatial audio image. Systems and techniques are needed that can be used to generate a spatial audio recording from a multi-microphone (e.g., multi-channel) audio input that includes audio captured by one or more occluded (e.g., partially occluded or completely occluded) microphones. There is also a need for systems and techniques that can be used to recover spatial audio signals associated with one or more occluded microphones.
[0045] As described in more detail herein, systems, devices, methods, and computer-readable media (collectively, "systems and techniques") are described herein that can be used to generate spatial audio recordings (e.g., spatial audio outputs) from multi-microphone input signals including one or more channels with audio degradation caused by occlusion of corresponding microphones. For example, these systems and techniques can be used to recover spatial audio signals associated with one or more occluded microphones. In some aspects, these systems and techniques can perform occlusion detection to identify one or more microphones in a microphone array (e.g., or otherwise included in the same multi-microphone device) that are currently experiencing an occlusion event. An occlusion event can be associated with a user's hand or finger rubbing or blocking a microphone (e.g., rubbing or blocking an opening within which the microphone is located). In some aspects, friction occlusion can be associated with physical contact with a microphone or microphone opening that moves (e.g., changes) over time. Blockage occlusion can be associated with static physical contact with a microphone or microphone opening in which at least a portion of the microphone or microphone opening is blocked and remains blocked for a period of time.
[0046] In some examples, these systems and techniques can determine estimated directions of arrival (DOAs) of one or more sound sources associated with or represented in a multi-channel audio input. For example, these systems and techniques can determine one or more DOA estimates using non-occluded channels of the multi-channel audio input (e.g., using audio channel signal information generated by microphones that were not identified as occluded during occlusion detection as described above).
[0047] In some examples, the systems and techniques may use occlusion information (e.g., determined based on occlusion detection) and / or DOA information (e.g., determined based on DOA estimation) to select one or more spatial filters for restoring the spatial audio output of the occluded channel. For example, the systems and techniques may use occlusion information and DOA information to select one or more spatial filters for reconstructing some (or all) of the occluded channels. In some aspects, the systems and techniques may use occlusion information and DOA information to select one or more spatial filters for generating a replacement channel that is different from the occluded channel. For example, the occlusion information may indicate to a spatial recording engine (e.g., such as Figure 4A and Figure 4BThe illustrated spatial recording engine 430 may be used to estimate the microphones and corresponding sound field locations that were expected by the spatial recording engine 430 but become degraded or unavailable due to an occlusion event. The estimated DOA information may indicate the location of one or more sound sources and may be used in combination with the occlusion information to determine portions or locations of the sound field where unavailable or degraded audio signal input is associated with lost sound source measurements (e.g., portions or locations of the sound field where the occluded microphone is aligned with the DOA of the sound source). As will be described in more depth below, the spatial audio output may be restored by selecting one or more spatial filters for generating a beamformer that captures the portion of the sound field that is lost or degraded due to the occluded microphone.
[0048] Additional details regarding the systems and techniques will be described with respect to the accompanying figures.
[0049] Figure 1 Three different views (e.g., front view 120, rear view 130, and side view 140) of an example wireless communication device 120 are shown. For example, front view 120 may correspond to a first side of device 102 including display 106. The first side may include first microphone 104a, second microphone 104b, third microphone 104c, earpiece 108, first loudspeaker 110a, and second loudspeaker 110b. Rear view 130 may correspond to a second side of device 102 opposite to the first side. The second side may include camera 106, fourth microphone 104d, and fifth microphone 104e. Side view 140 may correspond to a third side of device 102 connecting the first side and the second side.
[0050] In one illustrative example, wireless communication device 120 may also be referred to as a multi-microphone handset (e.g., a multi-microphone device). Figure 1 In addition to the handheld specific implementation of the multi-microphone device 102 depicted in , various other examples of audio sensing devices may additionally or alternatively be utilized to capture audio data using a multi-microphone array. For example, the multi-microphone device 102 may be implemented using one or more of a portable computing device (e.g., a laptop computer, a notebook computer, an ultra-portable computer, a tablet computer, a mobile Internet device, a smart phone, a smart watch, a wearable device, etc.), an audio or video playback device (e.g., headphones, earbuds, speakers, etc.), an audio or video conferencing device, and / or a display screen (e.g., a computer monitor, a television), etc. In some examples, the multi-microphone device 102 may be implemented using and / or may be included in one or more of an XR device, a VR device, an AR device, a wearable device, an audible device, smart glasses, a robotic device, etc.
[0051] In some cases, microphones 104a to 104e may be included in one or more configurable microphone array geometries (e.g., based on or associated with different sound source directions). In some aspects, different combinations (e.g., different pairs) of microphones 104a to 104e included in a multi-microphone handset 120 may be selected or utilized to perform spatially selective audio recording in different sound source directions. For example, a first microphone pair may include a first microphone 104a and a second microphone 104b and may be associated with an axis extending in a left-right direction of the front of the device 102 (e.g., from the perspective of the front view 120). A given microphone may be included in multiple microphone pairs. For example, a second microphone pair may also include a first microphone 104a together with a fifth microphone 104e and may be associated with an axis extending in a front-to-back direction (e.g., orthogonal to the front of the device and / or the perspective of the front view 120). In some examples, a second microphone pair may be used to determine whether the sound source direction is along the front-to-back axis (e.g., whether the user is speaking in front of the device 102 or at the back of the device 102). In some cases, a front and back microphone pair (e.g., a second microphone pair of microphones 104a and 104e) may be used to resolve ambiguity between front and back directions that a left and right microphone pair (e.g., a first microphone pair of microphones 104a and 104b) may not resolve on its own.
[0052] In some examples, when the device 102 is used as a video camera (e.g., using the camera 106 and one or more of the microphones 104a-104e to capture audiovisual data), one or more front and back microphone pairs may be used to record audio in the front and back directions (e.g., by directing beams into and away from the camera 106, respectively). As previously mentioned, one front and back pair of microphones may include the first microphone 104a and the fifth microphone 104e. One or more different front and back pairs of microphones may also be utilized (e.g., the first microphone 104a and the fourth microphone 104d, the second microphone 104b and the fourth microphone 104d, the second microphone 104b and the fifth microphone 104e, the third microphone 104c and the fourth microphone 104d, the third microphone 104c and the fifth microphone 104e, etc.). In some aspects, one or more front and back pairs of microphones may be used to record audio in the front and back directions, where the left and right direction preference may be determined manually or automatically.
[0053] In some aspects, the configurable microphone 104a-104e array geometry can be used to compress and transmit three-dimensional (3D) audio (e.g., also referred to as spatial audio). For example, given a range of design methods (e.g., minimum variance distortionless response (MVDR), linear constrained minimum variance (LCMV), phased array, etc.), different beamformer data sets can be determined for various combinations of microphones included in the plurality of microphones 104a-104e of the device 102.
[0054] The multi-microphone device 102 can be used to determine the direction of arrival (DOA) of the source signal by measuring the difference (e.g., phase difference) between the microphone channels for each frequency bin to obtain an indication (or estimate) of the direction and averaging the direction indication over all frequency bins to determine whether the estimated direction is consistent over all bins. The range of frequency bins available for tracking can be constrained by the spatial aliasing frequency of the corresponding microphone pair (e.g., the first microphone 104a and the second microphone 104b). The upper limit of the range can be defined as the frequency at which the wavelength of the source signal is twice the distance d between the microphones included in the microphone pair (e.g., microphones 104a, 104b).
[0055] The source of the sound may be tracked using available frequency bins up to the Nyquist frequency and down to lower frequencies (e.g., by supporting the use of microphone pairs with larger inter-microphone distances). This approach may be implemented to select the best pair among all available pairs, rather than being limited to a single pair for tracking. This approach may be used to support source tracking and provide higher DOA resolution even in far-field scenarios (e.g., up to three to five meters or more). In some cases, an exact 2D representation of the active sound source may be obtained or otherwise determined.
[0056] The multi-microphone device 102 may be used to calculate a difference between a pair of channels of a multi-channel input signal (e.g., obtained using one or more of the microphones 104a to 104e). For example, each channel of the multi-channel signal may be based on a signal generated by a corresponding microphone (e.g., one of the microphones 104a to 104e). For each direction among the plurality of K candidate directions, a corresponding direction error may be determined based on the calculated difference. Based on the K direction errors, the multi-microphone device 102 may select a candidate direction.
[0057] A multi-channel input signal (e.g., an input audio signal obtained using multiple microphones of microphones 104a to 104e) may be processed as a series of segments or "frames". In some cases, the segment length may be in the range of from about five or ten milliseconds to about forty or fifty milliseconds, and the segments may be overlapping (e.g., where adjacent segments overlap by 25% or 50%) or non-overlapping. In one illustrative example, the multi-channel signal may be divided into a series of non-overlapping segments or frames, each of which has a length of ten milliseconds. In another example, each frame may have a length of twenty milliseconds. A segment as processed by the multi-microphone device 102 may also be a segment (e.g., a "subframe") of a larger segment as processed by a different operation, and vice versa.
[0058] Examples of differences between input channels captured by different ones of the microphones 104 a - 104 e may include gain differences or ratios, arrival time differences, and / or phase differences, etc. For example, the multi-microphone device 102 may calculate the difference between channels of a pair of input signals (e.g., a first input signal associated with the first microphone 104 a and a second input signal associated with the second microphone 104 b ) as a difference or ratio between corresponding gain values of the channels (e.g., a difference in amplitude or energy).
[0059] In some aspects, the multi-microphone device 102 may calculate a measure of gain of a segment of the multi-channel signal in the time domain (e.g., for each of a plurality of subbands of the signal) and / or in the frequency domain (e.g., for each of a plurality of frequency components of the signal in a transform domain such as a Fast Fourier Transform (FFT), a Discrete Cosine Transform (DCT), or a Modified DCT (MDCT) domain). Examples of such gain measures include, but are not limited to, one or more of: total amplitude (e.g., the sum of the absolute values of sample values), average amplitude (e.g., per sample), root mean square (RMS) amplitude, median amplitude, peak amplitude, peak energy, total energy (e.g., the sum of the squares of the sample values), and average energy (e.g., per sample).
[0060] In order to obtain accurate results with the gain difference technique, the responses of different microphone channels included in the multi-channel signal (e.g., a first input signal associated with the first microphone 104a and a second input signal associated with the second microphone 104b) may be calibrated relative to each other. The multi-microphone device 102 may apply a low-pass filter to the multi-channel signal so that the calculation of the gain measure is limited to the audio frequency components of the multi-channel signal.
[0061] In some aspects, the multi-microphone device 102 may calculate the difference between the gains as a difference between corresponding gain metric values (e.g., values in decibels) for each channel of the multi-channel signal in the logarithmic domain or equivalently as a ratio between the gain metric values in the linear domain. For a calibrated microphone pair (e.g., the first microphone 104a and the second microphone 104b), a gain difference of zero may be employed to indicate that the source is equidistant from each microphone (e.g., located in the broadside direction of the pair), a gain difference with a large positive value may be employed to indicate that the source is closer to one microphone (e.g., located in one endfire direction of the pair), and a gain difference with a large negative value may be employed to indicate that the source is closer to the other microphone (e.g., located in the other endfire direction of the pair).
[0062] In some examples, the multi-microphone device 102 may perform a cross-correlation on the input channels (e.g., a first input signal associated with the first microphone 104a and a second input signal associated with the second microphone 104b) to determine a difference, such as by calculating an arrival time difference based on a lag between the channels of the multi-channel signal.
[0063] In some cases, the multi-microphone device 102 may calculate the difference between the channels in a pair (e.g., a first input signal associated with the first microphone 104a and a second input signal associated with the second microphone 104b) as the difference between the phases of each channel (e.g., at a particular frequency component of the signal). In some aspects, such calculations may be performed for each frequency component among the plurality of frequency components.
[0064] For a signal received by a pair of microphones (e.g., microphones 104a and 104b) directly from a point source in a particular direction of arrival (DOA) relative to an axis of the microphone pair (e.g., microphones 104a, 104b), the phase delay may be different for each frequency component and may also depend on the spacing between microphones 104a and 104b. In some cases, multi-microphone device 102 may calculate the observed value of the phase delay at a particular frequency component (or "bin") as the inverse tangent (also referred to as the inverse tangent) of the ratio of the imaginary term of the complex FFT coefficients to the real term of the complex FFT coefficients.
[0065] refer to Figure 2A , a diagram of a far-field model of plane wave propagation relative to a microphone pair is shown and generally designated 200a. Figure 2B In FIG. 1 , a diagram of an example of microphone placement is shown and generally designated as 200 b. In some examples, microphone placement 200 b may correspond to Figure 1 The illustrated placement of the first microphone 104a, the second microphone 104b, the third microphone 104c and the fourth microphone 104d.
[0066] The multi-microphone device 102 may determine direction of arrival (DOA) information corresponding to respective input signals of the microphones 104a to 104c and 104d. For example, the far-field model 200a indicates a phase delay value of a source S01 of at least one microphone (e.g., microphones 104a to 104b) at a particular frequency f. can be related to the source DOA under the far-field (i.e., plane wave) assumption as follows:
[0067]
[0068] Here, d represents the distance between microphones 104a and 104b (e.g., in meters), θ represents the angle of arrival relative to a direction normal to the array axis (e.g., in radians), f represents the frequency (e.g., in Hertz (Hz)), and c represents the speed of sound (e.g., in m / s). The DOA estimation principles described herein can be extended to multiple microphone pairs in a linear array (e.g., such as Figure 2B For a single point source without reverberation, the ratio of phase delay to frequency is will have the same value at all frequencies:
[0069]
[0070] The DOA θ with respect to a microphone pair (eg, microphones 104a and 104b) is a one-dimensional measurement of the surface of a cone defined in space (eg, such that the axis of the cone is the axis of the array).
[0071] In some examples, the input audio signal (e.g., a speech signal) may be sparse in the time-frequency domain. If the sources of the input signal included in the multi-channel input (e.g., generated using multiple microphones 104a to 104e of device 102) do not intersect in the frequency domain, the multi-microphone device 102 may track two sources simultaneously. If the sources do not intersect in the time domain, the multi-microphone device 102 may track two sources at the same frequency. The microphone array of device 102 may include a number of microphones at least equal to the number of different source directions to be distinguished at any one time. The microphones (e.g., Figure 1 The illustrated microphones 104a through 104e) may be omnidirectional (eg, for a cellular telephone or dedicated conferencing equipment) or directional (eg, for a device such as a set-top box).
[0072] In some aspects, the multi-microphone device 102 may calculate a DOA estimate for a frame of a received multi-channel input signal (e.g., generated using one or more or all of the microphones 104a to 104e of the device 102). The multi-microphone device 102 may calculate the error of each candidate angle relative to the observed angle at each frequency bin, which is indicated by a phase delay. The target angle at that frequency bin may be the candidate with the smallest (or lowest) error. In one example, the errors may be summed across frequency bins to obtain a measure of the likelihood of the candidate. In another example, one or more of the target DOA candidates that appear most frequently across all frequency bins may be identified as the DOA estimate (or multiple DOA estimates) for a given frame.
[0073] The multi-microphone device 102 may obtain substantially instantaneous tracking results (e.g., with a delay of less than one frame). The delay may depend on the FFT size and the degree of overlap. For example, for a 512-point FFT with 50% overlap and a 16 kilohertz (kHz) sampling frequency, the resulting 256-sample delay may correspond to sixteen milliseconds. The multi-microphone device 102 may support differentiation of source directions for source-array distances of up to two to three meters, or up to five meters, etc.
[0074] The error may also be viewed as a variance (e.g., the degree to which an individual error deviates from an expected value). Converting a time-domain received signal into the frequency domain (e.g., by applying an FFT) may have the effect of averaging the frequency spectrum in each frequency bin. Such averaging may be more efficient if the multi-microphone device 102 uses a sub-band representation (e.g., a Mel scale or a Bark scale). In some aspects, the multi-microphone device 102 may perform time-domain smoothing on the DOA estimate (e.g., by applying a recursive smoother, such as a first-order infinite impulse response filter). In some examples, the multi-microphone device 102 may reduce the computational complexity of the error calculation operation (e.g., by using a search strategy such as a binary tree and / or applying known information such as DOA candidate selections from one or more previous frames).
[0075] Although the directional information may be measured in terms of phase delay, the multi-microphone device 102 may obtain a result indicative of the source DOA.The multi-microphone device 102 may calculate the directional error at frequency f for each DOA candidate in the inventory of K DOA candidates in terms of DOA rather than phase delay.
[0076] refer to Figure 3, an example arrangement of microphones is shown and generally designated as 300. In one illustrative example, the example arrangement of microphones 300 may be associated with omnidirectional and first-order capture for spatial coding using a four-microphone array, as will be described in more depth below. Using microphone arrangement 300, multi-microphone device 102 may generate output audio signals from input signals corresponding to microphones 104a to 104d. For example, multi-microphone device 102 may use microphone arrangement 300 to approximate first-order capture for spatial coding using a four-microphone setup (e.g., using microphones 104a to 104d). Examples of spatial audio coding methods that may be supported by a multi-microphone array as described herein may also include methods that may have been originally intended for use with specific microphones, such as Ambisonics B format or higher-order Ambisonics formats. For example, a processed multi-channel output of an Ambisonics encoding scheme may include a three-dimensional Taylor expansion at a measurement point, which may be approximated at least up to first order using a three-dimensionally positioned microphone array (e.g., corresponding to microphone arrangement 300). As a larger number of microphones are included in or otherwise associated with multi-microphone device 102 , the order of approximation may increase.
[0077] As illustrated, in some examples, the second microphone 104b may be separated from the first microphone 104a by a distance Δz in the z-direction, and the third microphone 104c may be separated from the first microphone 104a by a distance Δy in the y-direction. The fourth microphone 104d may be separated from the first microphone 104a by a distance Δx in the x-direction. In some aspects, the audio signals captured using the microphones 104a to 104e may be processed and / or filtered to obtain the DOA of the audio frame. In an illustrative example, one or more beamformers may be used to "shape" the audio signals captured using the microphones 104a to 104e. The shaped audio signals may be played back in a surround sound system, headphones, or other audio playback device to generate an immersive sound experience (e.g., a spatialized audio experience).
[0078] In some examples, a multi-microphone device (e.g., such as Figure 1 The illustrated multi-microphone device 102 may be used to capture or otherwise generate a multi-channel audio signal. The multi-channel audio signal may include multiple different audio signals (e.g., different audio channels captured by different microphones and / or microphone combinations) and may be used to generate a spatial audio output, as will be described in more depth below. Figure 4A4 is a diagram illustrating an example spatial audio recording system 400a that can generate spatial audio output based on receiving as input a multi-channel (e.g., multi-microphone) audio signal including non-occluded microphone signals. For example, the spatial audio recording system 400a can receive a multi-microphone input signal 410a captured by a multi-microphone device 402. In some aspects, the multi-microphone device 402 can be connected to Figure 1 The multi-microphone apparatus 102 illustrated and previously described above is the same or similar.
[0079] As illustrated, the example spatial audio recording system 400a may include a pre-processing engine 420 (e.g., which receives a multi-microphone input signal 410a as an input), a spatial recording engine 430 (e.g., which generates a spatial output signal 435a), and a post-processing engine 440. In some aspects, the multi-microphone input signal 410a provided to the pre-processing engine 420 may include one audio channel for each microphone of the device 402 used to capture the multi-microphone input signal 410a. For example, the multi-microphone input signal 410a may be a stereo signal that includes a left channel and a right channel captured by a respective microphone included in the device 402. In another example, the multi-microphone input signal 410a may include three or more channels (e.g., a left channel, a center channel, and a right channel) captured by a respective microphone included in the device 402.
[0080] In one illustrative example, one or more (or all) of pre-processing engine 420, spatial recording engine 430, and / or post-processing engine 440 may be included in or implemented by device 402 (e.g., spatial audio recording system 400a may be implemented locally on device 402). In some examples, one or more (or all) of pre-processing engine 420, spatial recording engine 430, and / or post-processing engine 440 may be implemented remotely from device 402 (e.g., using one or more remote servers, cloud computing platforms, etc.).
[0081] In some aspects, the pre-processing engine 420 may implement one or more audio pre-processing functions or operations. For example, the pre-processing engine 420 may perform one or more audio pre-processing functions or operations for some (or all) of the corresponding audio channels included in the multi-microphone input 410a. In some cases, the pre-processing engine 420 may perform audio pre-processing operations, which may include but are not limited to gain adjustment (e.g., gain increase or gain reduction), noise removal, noise suppression, etc. In some aspects, the one or more audio pre-processing operations performed by the pre-processing engine 420 may be linear processing operations (e.g., the pre-processing engine 420 may not perform non-linear processing operations). In some cases, one or more audio pre-processing operations may be performed (e.g., by the pre-processing engine 420) without modifying or changing the envelope of the corresponding audio signal associated with each channel of the multi-channel microphone input 410a.
[0082] After preprocessing is performed on the multi-microphone input 410a (e.g., in a preprocessing stage associated with the preprocessing engine 420), the preprocessed audio channels of the multi-microphone input 410a may be provided as input to the spatial recording engine 430. The spatial recording engine 430 may generate a spatial audio output 435a based on the preprocessed audio channel signals provided as output by the preprocessing engine 420. In some examples, preprocessing may not be performed on the multi-microphone input 410a (e.g., the input and output of the preprocessing engine 420 may be the same, or the preprocessing engine 420 may be removed from the example spatial audio recording system 400a), in which case the spatial recording engine 430 may use the original audio channel signals included in the multi-microphone input 410a to produce the spatial audio output 435a.
[0083] In some cases, the spatial audio output 435a may be provided to a post-processing engine 440, which generates a final spatial audio stream (e.g., also referred to as a spatial audio recording generated by the example spatial audio recording system 400a) as an output. The spatial audio stream may be generated based on applying or otherwise performing one or more post-processing operations on some (or all) of the corresponding audio channels included in the spatial audio output 435a. In some aspects, the number of channels included in the multi-microphone input 410a may be different from the number of channels included in the spatial audio output 435a. For example, when the device 402 includes multiple microphones and the multi-microphone input 410a includes corresponding multiple audio channels, the spatial audio output 435a may include a greater number of channels than the multi-microphone input 410a, or the spatial audio output 435a may include a lesser number of channels than the multi-microphone input 410a. In some aspects, the number of channels included in the spatial audio output 435a may be the same as the number of beamformers utilized by the spatial recording engine 430.
[0084] In some cases, spatial audio output 435a may be generated by spatial recording engine 430 using one or more predetermined spatial filters corresponding to an expected configuration of microphones used to capture original multi-microphone input 410a (e.g., an expected configuration of microphones included on multi-microphone device 402). For example, spatial recording engine 430 may include one or more predetermined two-channel spatial filters corresponding to a two-channel (e.g., stereo) configuration of a first microphone and a second microphone of device 402. The predetermined two-channel spatial filters may be used to generate spatial audio (e.g., spatial audio output 435a) from stereo input 410a captured using the first microphone and the second microphone of device 402.
[0085] The spatial recording engine 430 may additionally or alternatively include one or more predetermined multi-channel spatial filters (e.g., a three-channel spatial filter, a four-channel spatial filter, etc.) for generating spatial audio from a three-channel or more channel input 410a captured using corresponding three or more microphones of the device 402. In some aspects, when the device 402 includes three or more microphones and / or the multi-microphone input 410a includes three or more channels, the spatial recording engine 430 may utilize beamforming to generate the spatial audio output 435a. Each beamformer may be determined based on assigning predetermined beamformer weight values to audio channel information associated with a subset or selection of microphones included in the set of three or more microphones represented in the multi-channel input 410a.
[0086] For example, the spatial recording engine 430 may perform beamforming based on determining one or more direction of arrival (DOA) estimates. In an illustrative example, the spatial recording engine 430 may include a DOA estimator (not shown) that performs various operations on time-matched input data (e.g., time-matched audio channel data included in the multi-channel input 410a) to estimate the DOA of incident sounds in various frequency bands. The DOA estimator may utilize various techniques to estimate the DOA of sounds from one or more sound sources incident on the microphone array of the device 402 (e.g., incident on multiple microphones of the device 402). For example, the DOA estimator may estimate the spatial correlation matrix of input signals from a subset of microphones included in the microphone array of the device 402 and may perform eigenanalysis of the spatial correlation matrix to obtain a set of DOA estimates. These DOA estimates may then be used to assign weights to each microphone in a subset of microphones used in the DOA estimates. For example, various beamforming algorithms may be used to assign weights to each microphone in a subset of microphones used in the DOA estimates based on the DOA estimates. In some aspects, the beamforming algorithm selected or otherwise utilized by the spatial recording engine 430 can be based on the geometry of the microphone array of the device 402. Once the weights are assigned, a corresponding beamformer can be determined based at least in part on a weighted sum of each of the sound signals used in the DOA estimation.
[0087] In some aspects, the spatial audio output 435a may be generated by the spatial recording engine 430 using a predetermined spatial filter and / or a predetermined beamformer that depends on the multi-microphone input signal 410a that includes an expected number of audio channels that are each associated with a different microphone of the device 402 (e.g., based on a spatial filter and / or beamformer predetermined for the expected number of microphones and the geometry or arrangement of the microphones as provided on the device 402).
[0088] In some cases, the spatial recording engine 430 may be unable to generate the spatial audio output 435a if one or more of the expected audio channels are not included in the multi-microphone input 410a (e.g., based on blocking occlusion of the corresponding microphone or microphone opening on the device 402) and / or if one or more of the expected audio channels are included in the multi-microphone input 410a but exhibit audio quality degradation (e.g., based on friction or partial occlusion of the corresponding microphone or microphone opening on the device 402).
[0089] For example, Figure 4BAs illustrated, if one or more intended audio channels associated with the multi-channel input 410b are occluded such that the intended audio channels are absent or degraded, the spatial recording engine 430 may generate a corrupted spatial output 435b. In some cases, the corrupted spatial output 435b may be associated with a loss of spatial imaging (e.g., where the corrupted spatial output 435b is corrupted to the extent that a listener does not perceive the corrupted spatial output 435b as providing a spatial audio experience).
[0090] As previously mentioned, when a mobile device is used for spatial audio recording, one or more microphones may be blocked by a user's hand or finger. Microphone occlusion may reduce the quality of the resulting spatial audio recording, destroy or degrade the resulting spatial audio recording, and / or prevent the generation of a spatial audio recording. For example, a smartphone may include multiple microphones that are each sunken into the body of the smartphone (e.g., may include microphones that are seated in openings included on an outer surface or body of the smartphone). A user's hand or finger when holding the smartphone and attempting to capture a spatial audio recording may block or rub one or more of the microphone openings. In some aspects, both blocking of a microphone and rubbing of a microphone may be referred to as "occlusion" of a microphone. In some cases, the occlusion of a microphone may be partial or complete. Systems and techniques are needed to mitigate audio quality degradation and / or spatial audio image loss associated with such microphone occlusions.
[0091] As described in more detail below, systems and techniques are described herein that can be used to generate a spatial audio recording (e.g., a spatial audio output) from a multi-microphone input signal including one or more channels with audio degradation caused by occlusion of the corresponding microphone. As used herein, the terms "recording" and "recording" can be used interchangeably (e.g., "spatial audio record" and "spatial audio recording" can be used interchangeably). In some examples, these systems and techniques can be used to recover a spatial audio signal associated with one or more occluded microphones. In some aspects, these systems and techniques can perform occlusion detection to identify one or more microphones in a microphone array (e.g., or otherwise included in the same multi-microphone device) that are currently experiencing an occlusion event. An occlusion event can be associated with a user's hand or finger rubbing or blocking a microphone (e.g., rubbing or blocking an opening in which the microphone is located). In some aspects, friction occlusion can be associated with physical contact with a microphone or a microphone opening that moves (e.g., changes) over time. Blockage occlusion may be associated with static physical contact with the microphone or microphone opening wherein at least a portion of the microphone or microphone opening is blocked and remains blocked for a certain period of time.
[0092] In some examples, these systems and techniques can determine estimated directions of arrival (DOAs) of one or more sound sources associated with or represented in a multi-channel audio input. For example, these systems and techniques can determine one or more DOA estimates using non-occluded channels of the multi-channel audio input (e.g., using audio channel signal information generated by microphones that were not identified as occluded during occlusion detection as described above).
[0093] In some examples, the systems and techniques may use occlusion information (e.g., determined based on occlusion detection) and / or DOA information (e.g., determined based on DOA estimation) to select one or more spatial filters for restoring the spatial audio output of the occluded channel. For example, the systems and techniques may use occlusion information and DOA information to select one or more spatial filters for reconstructing some (or all) of the occluded channels. In some aspects, the systems and techniques may use occlusion information and DOA information to select one or more spatial filters for generating a replacement channel that is different from the occluded channel. For example, the occlusion information may indicate to a spatial recording engine (e.g., such as Figure 4A and Figure 4B The illustrated spatial recording engine 430 may be used to estimate the microphones and corresponding sound field locations that were expected by the spatial recording engine 430 but become degraded or unavailable due to an occlusion event. The estimated DOA information may indicate the location of one or more sound sources and may be used in combination with the occlusion information to determine portions or locations of the sound field where unavailable or degraded audio signal input is associated with lost sound source measurements (e.g., portions or locations of the sound field where the occluded microphone is aligned with the DOA of the sound source). As will be described in more depth below, the spatial audio output may be restored by selecting one or more spatial filters for generating a beamformer that captures the portion of the sound field that is lost or degraded due to the occluded microphone.
[0094] In some aspects, spatial filter selection (e.g., spatial construction selection) for recovering a spatial audio output of a multichannel input signal including at least one occluded channel may be performed based at least in part on the number of microphones used to capture the multichannel input signal (e.g., the number of microphones associated with the expected non-occluded input signal and / or the number of channels expected in the multichannel input signal). In an illustrative example, these systems and techniques may recover a spatial audio output of a stereo (e.g., two-channel) input including one occluded channel and one non-occluded channel. In some aspects, these systems and techniques may determine and use arrival time difference (ATD) information and arrival level difference (ALD) information to reconstruct a spatial (e.g., stereo) image over the duration of the occlusion event. For example, the ATD and ALD information may be used to reconstruct the occluded channel audio signal using the non-occluded channel audio signal.
[0095] In another illustrative example, these systems and techniques can recover a spatial audio output of a multi-channel input signal including three or more channels and where at least one channel is occluded. For example, based on performing occlusion detection to identify one or more microphones (e.g., input channels) that are occluded, these systems and techniques can recover the occluded microphone audio channel information by constructing one or more beamformers using non-occluded microphone audio channel information (e.g., where the one or more beamformers reconstruct the occluded channels).
[0096] For example, Figure 5 is a diagram illustrating an example spatial audio recording system 500 that may perform occlusion handling for spatial audio recording. In some aspects, the example spatial audio recording system 500 may include Figure 4A and Figure 4B The same multi-microphone input device 402 as illustrated and / or Figure 4B The illustrated same multi-microphone input 410b includes at least one occluded channel. In some aspects, the example spatial audio recording system 500 may include Figure 4A and Figure 4B The pre-processing engine 520 illustrated is the same as or similar to the pre-processing engine 420 and / or may include Figure 4A and Figure 4B The illustrated post-processing engine 440 is the same as or similar to the post-processing engine 540 .
[0097] In one illustrative example, spatial recording may be performed using an occlusion processing engine 532 and a spatial recording engine 534. In some aspects, the occlusion processing engine 532 may also be referred to as an occlusion handling system (OHS). As will be described in more depth below, the OHS 532 may be used to perform occlusion detection to identify one or more microphones in a microphone array (e.g., or otherwise included in the same multi-microphone device) that are occluded for a current or given frame of input audio data. The OHS 532 may additionally be used to perform DOA estimation and spatial filter selection, as will also be described in more depth below. The spatial recording engine 534 may receive occlusion information corresponding to one or more occluded microphone channels included in the multi-microphone input 410b from the OHS 532 and may utilize the occlusion information to generate a recovered spatial audio output 535.
[0098] Figure 6 is a diagram illustrating an example of an example spatial recording occlusion handling system 600. In some examples, the spatial recording occlusion handling system 600 may be used to implement Figure 5The illustrated OHS 532 and spatial recording engine 534. The spatial recording occlusion handling system 600 may receive as input a multi-channel audio input 610, which may include one or more occluded audio channels. The spatial recording occlusion handling system 600 may generate a restored spatial output 695 based on non-occluded audio channels included in the multi-channel audio input 610. In some cases where the multi-channel audio input 610 does not include any occluded audio channels, the spatial recording occlusion handling system 600 may generate a spatial output in the same or similar manner as described above with respect to the example spatial recording system 400.
[0099] As illustrated, the spatially recorded occlusion handling system 600 may include an occlusion detection engine 620, a delayed audio database 630, a spatial device selection engine 640 (e.g., also referred to as a spatial filter selection engine), and a selective spatial filtering engine 650, each of which is described in more detail below.
[0100] In some examples, the user interface (UI) / audio framework 625 may be external to the spatial recording occlusion handling system 600 (e.g., not included in the spatial recording occlusion handling system 600). The UI / audio framework 625 may receive an occlusion flag generated by the occlusion detection engine 620 for a given frame (e.g., the current frame) of the multi-channel audio input 610. For example, in some aspects, an occlusion warning or other UI element may be generated and displayed to a user of a corresponding multi-microphone device (e.g., a multi-microphone device whose occlusion flag is generated and / or received by the UI / audio framework 625). In some cases, an occlusion warning generated based on the UI / audio framework 625 receiving the occlusion flag may be a visual notification or warning displayed on a display of the corresponding multi-microphone device. For example, an occlusion warning may be used to inform a user of a corresponding multi-microphone device that one or more microphones of the device are currently occluded. Based on receiving an occlusion warning generated using the UI / audio framework 625, the user may be prompted to reposition his or her finger to avoid continued microphone occlusion.
[0101] In one illustrative example, multi-channel audio input 610 may include a frame of audio signal data for each respective audio channel (e.g., where each respective audio channel of multi-channel audio input 610 is captured by a corresponding microphone associated with the same device and / or spatial recording). An audio frame may also be referred to as an audio sample and may be the smallest discrete unit of data captured by a microphone. For example, if a given microphone has a sampling rate of 48,000 samples / second (e.g., 48kHz), one second of captured audio will include 48,000 audio frames. An audio frame may include amplitude (e.g., loudness) information at a particular point in time, where audio playback is performed by playing consecutive audio frames in sequential order. Multi-channel audio (such as multi-channel audio input 610) may include one audio frame per channel for each given point in time at which sampling was performed. For example, one second of mono audio sampled at 48 kHz may include 48,000 audio frames, one second of stereo audio sampled at 48 kHz may include 96,000 audio samples, one second of three-channel audio sampled at 48 kHz may include 144,000 audio frames, and so on.
[0102] In some aspects, a current audio frame (e.g., a most recently received frame) of the multi-channel audio input 610 may be provided to the occlusion detection engine 620. The occlusion detection engine 620 may analyze the current audio frame to determine whether occlusion (e.g., friction or blocking) has occurred for one or more of the microphones / channels represented in the multi-channel audio input 610. For example, the occlusion detection engine 620 may generate or otherwise set an occlusion flag indicating whether occlusion has been detected in the current audio frame for each channel included in the multi-channel audio input 610, as described below with respect to Figure 7 Will be described in more depth.
[0103] The current frame of the multi-channel audio input 610 may additionally be provided to a delay filter 615 coupled to a delay audio database 630. The delay audio database 630 may store one or more delayed audio signals based on the multi-channel audio input 610. For example, the delay audio database 630 may generate and store a delayed version of each audio channel included in the multi-channel audio input 610. In some aspects, the delay audio database 630 may use the delay filter 615 (e.g., shown as a delay filter z -m 615) to generate a delayed version of each audio channel. For example, in the delay filter z -mThe multi-channel audio input 610 received at 615 may be delayed by m samples (e.g., m frames) before being written to the delayed audio database 630. In one illustrative example, occlusion detection is performed by the occlusion detection engine 620 using a current frame of the audio input 610, while spatial filter selection is performed by the spatial device selection engine 640 using one or more delayed frames of the audio input 610 (e.g., where the delayed frames are based on the delayed filter z). -m 615 to be delayed relative to the current frame). In some aspects, the selective spatial filtering engine 650 can receive delay information from the delay filter 615. For example, the selective spatial filtering engine can receive a value for delay m (e.g., the number of samples m by which the multi-channel audio input 610 received at the delay filter 615 is delayed before being written to the delay audio database 630). In one illustrative example, the delay audio database 630 can be used to provide a time buffer between detecting the onset of an occlusion event and the occlusion event being represented in or associated with the spatial output 695. For example, if at time t 0 If an occlusion event is detected for a given set of audio frames, the same given set of audio frames is not used to generate the spatial output 695 until the delay time t d =t 0 +m (e.g., a given set of audio frames may be obtained and at t 0 Occlusion is detected, but these audio frames are not processed by the selective spatial filter engine 650 or read from the delayed audio database 620 until the delay time t d , the delay time is at t 0 In some cases, the selective spatial filtering engine 650 may be configured to filter the selected spatial filtering engine 650 at time t 0 Start the cross-fading from the non-occluded spatial filter to the occluded spatial filter so that at the delay time t d (For example, t 0 +m) to complete the cross-fade, at which delay time the selective spatial filtering engine 650 receives a delayed version of the first audio frame in which the occlusion is detected.
[0104] Figure 7 is a diagram illustrating an example architecture of an occlusion detection engine 620. In one illustrative example, a current frame of a multi-channel audio input 610 may be provided to a feature extraction engine 730, which may generate one or more audio features for performing occlusion detection. In some aspects, the feature extraction engine 730 may receive a current frame for each channel included in the multi-channel audio input 610. For example, audio features may be extracted or generated for each audio channel / microphone, and an occlusion flag 775 may be determined for each audio channel / microphone based on analyzing the extracted audio features.
[0105] In some examples, the feature extraction engine 730 may include an envelope tracker 732 for tracking envelope values associated with each audio channel. The tracked envelope values may be used as audio features for subsequent downstream processes of at least the occlusion detection engine 620.
[0106] In addition to the envelope tracker 732, the current frame of each audio channel included in the input 610 may also be provided to a series of filter banks for tracking power. For example, the feature extraction engine 730 may include a low pass filter (LPF) 734a, a first high pass filter (HPF1) 734b, and a second high pass filter (HPF2) 734c, each of which receives the current frame of each audio channel as an input. In some aspects, HPF1 734b may perform high pass filtering using a different cutoff frequency than HPF2 734c. For example, the HPF1 cutoff frequency may be higher than the HPF2 cutoff frequency, or the HPF1 cutoff frequency may be lower than the HPF2 cutoff frequency. In an illustrative example, based on systems and techniques for using HPF2 734c to detect windy conditions and / or mitigate false alarms caused by wind noise or wind turbulence captured at one or more microphones, the HPF1 cutoff frequency may be lower than the HPF2 cutoff frequency (e.g., HPF2 cutoff frequency>HPF1 cutoff frequency) (e.g., as will be described in more depth below).
[0107] Each of the filters 734a to 734c may be associated with a corresponding power tracker 736a to 736c. For example, the first power tracker 736a may generate an LPF power value for each audio channel included in the low pass filter output of the LPF 734a. The second power tracker 736b may generate an HPF1 power value for each audio channel included in the first high pass filter output of the HPF1 734b, and the third power tracker 736c may generate an HPF2 power value for each audio channel included in the second high pass filter output of the HPF2 734c.
[0108] As illustrated, for each audio channel included in the multi-channel input 610, the feature extraction engine 730 may generate an envelope value feature (e.g., using the envelope tracker 732) and may generate three power values (e.g., using the filters 734a to 734c and the corresponding power trackers 736a to 736c). For example, if the multi-channel input 610 is a stereo signal (e.g., including two channels), the feature extraction engine 730 may generate a total of eight features, four features for each of the two channels. Similarly, if the multi-channel input 610 is a three-channel signal, the feature extraction engine 730 may generate 12 features, four features for each of the three channels. In some aspects, the generated features may be instantaneous values (e.g., values determined for the current audio frame of each channel), differential values (e.g., for each channel, comparing the current frame value to one or more previous frame values), short-term averages, long-term averages, and / or maximum or minimum values, etc.
[0109] One or more (or all) of the extracted features for each channel may be provided as input to the friction detection engine 740 and the occlusion detection engine 750. The description provided below regarding friction and occlusion detection may be applied to each respective channel included in the multi-channel audio input 610.
[0110] As previously mentioned, friction detection may be performed to determine whether friction occlusion has occurred for a given audio frame (e.g., whether the corresponding microphone has experienced friction occlusion). In some aspects, friction occlusion may be associated with physical contact with a microphone or microphone opening, where the physical contact moves (e.g., changes) over time. In an illustrative example, the friction detection engine 740 may include a friction detector 742 and a friction state machine 744. The friction detector 742 may perform friction detection for a current frame of audio for each channel included in the multi-channel input 610. For example, friction detection may be performed based on identifying a sudden change in an envelope value of one or more channels (e.g., one of the features generated using the feature extraction engine 730 for each audio channel) using the friction detector 742.
[0111] In some aspects, friction detection may be performed based on evaluating one or more conditions using the friction detector 742. For example, the first condition 761 may be based on the envelope value of each audio channel. In some cases, the first condition 761 may compare the instantaneous envelope value (e.g., determined by the feature extraction engine 730 for the current frame) with the average envelope value (e.g., a long-term or short-term average) of the same channel to identify sudden changes in the envelope. In some aspects, a sudden change in the envelope of the audio input associated with a given channel may indicate friction (e.g., friction may cause a sudden increase in amplitude measured by a microphone being rubbed). In some examples, the friction detector 742 may perform friction detection based only on changes in the envelope of a given channel.
[0112] In one illustrative example, the friction detector 742 may also confirm a potential friction event for a given channel (e.g., potentially detected based on an envelope change) by evaluating one or more power features generated by the feature extraction engine 730 for the given channel. For example, the friction detector 742 may also evaluate the first condition 761 by analyzing one or more power features determined for the HPF2 filter 734c. For example, an instantaneous HPF2 power value (e.g., determined by the feature extraction engine 730 for the current frame) may be compared to one or more predetermined thresholds, where an instantaneous power greater than or equal to the predetermined threshold indicates a friction event. In some examples, the instantaneous HPF2 power value may be compared to a long-term and / or short-term average HPF2 power value to identify a sudden change in the HPF2 power of a given channel. In some examples, the cutoff frequency of the HPF2 734c may be selected to filter out wind noise artifacts and other noise artifacts that may be included in one or more channels of the input 610. For example, the cutoff frequency of HPF2 734c may be greater than the cutoff frequency of HPF1 734b in order to filter out wind noise artifacts and other noise artifacts that are not friction occlusions but that might otherwise cause the friction detector 742 to generate a false alarm.
[0113] Based on the first condition 761 of evaluating the change in the envelope value and / or the change in the HPF2 power value, the friction detector 742 may output an indication of a detected friction event or an indication of a non-detected friction event (e.g., a binary friction state may be determined for the current frame of each channel). The detected friction state of each channel may be provided to a friction state machine 744, which may track the detected friction state of each channel over time. In an illustrative example, the friction state machine 744 may use the friction state timing information to determine a friction flag 745 for each channel. For example, the friction state machine 744 may set the friction flag 745 to true in response to the friction detector 742 determining a friction event for the current frame. In some aspects, the friction state machine 744 may set the friction flag 745 to true based on the friction detector 742 determining a friction event for a predetermined percentage or number of frames within a confidence interval number of frames. In some cases, after entering the friction state (e.g., by setting the friction flag 745 to true), the friction state machine 744 may remain in the friction state for a dwell period before exiting the friction state. For example, if after entering the friction state, friction detector 742 subsequently determines at some future frame that friction is no longer detected, friction state machine 744 may exit the friction state after waiting a predetermined number of frames (eg, a dwell period).
[0114] The blockage detection engine 750 can perform blockage detection to determine whether a given audio frame has experienced blockage occlusion (e.g., whether the corresponding microphone of the current audio frame has experienced blockage occlusion). Blockage occlusion can be a complete blockage of a microphone / microphone opening or can be a partial blockage of a microphone / microphone opening. In some aspects, the blockage detection engine 750 may include a blockage eligibility detector 752 that evaluates a second condition 762 to determine whether a given frame qualifies for blockage detection based on features generated for a given audio frame in the current audio frame.
[0115] For example, the blockage qualification detector 752 can be used to mitigate or reduce false alarms associated with wind vectors (e.g., false blockage occlusions that might otherwise be erroneously detected due to wind noise at one of the microphones / channels included in the input 610). In some aspects, the blockage qualification test can be performed based on evaluating the second condition 762 to determine whether the HPF2 power is greater than or equal to a predetermined threshold. In some examples, the HPF2 power threshold associated with evaluating the second condition 762 during the blockage qualification test phase can be the same or similar to the HPF2 power threshold described above as being associated with evaluating the first condition 761 during the friction detection phase.
[0116] In some examples, the blockage eligibility detector 752 may also evaluate the second condition 762 by analyzing the HPF2 power value differences across each pair of channels included in the multi-channel input 610. For example, a relatively large HPF2 channel power difference may indicate a blockage occlusion, while a relatively small (or approximately zero) HPF2 channel power difference may indicate wind noise rather than a blockage occlusion. For example, since the microphones may be provided on the same device and may experience approximately equal wind interactions, wind noise may be detected at multiple (or all) of the different microphone channels included in the input 610. Blockage occlusion may be more localized than wind interactions, and thus may be associated with the larger HPF2 channel power differences mentioned above (e.g., wind noise may be experienced at most or all of the microphone channels, while blockage occlusion may be experienced at one microphone channel).
[0117] Based on evaluating the second condition 762 with respect to the HPF2 power and / or the HPF2 power difference between the channel pairs, the blockage eligibility detector 752 may output an indication of the eligibility of each channel for further blockage detection processing. For example, if the blockage eligibility detector 752 outputs an indication that the current frame associated with the given channel is eligible for further blockage detection processing, the current frame associated with the given channel may be advanced to the blockage detector 754. If the blockage eligibility detector 752 outputs an indication that the current frame associated with the given channel is not eligible for further blockage detection processing (e.g., because a wind vector is detected), the current frame associated with the given channel may instead be advanced to the final output stage of the blockage detection engine 750 (e.g., if the given channel is not eligible for blockage detection, the final blockage flag 755 output by the blockage detection engine 750 may be set to indicate that the given channel is not blocked).
[0118] When the blockage qualification detector 752 indicates that the current frame associated with a given channel is eligible, the blockage detector 754 may perform blockage occlusion detection based on evaluating the third condition 763. For example, the third condition 763 may cause the blockage detector 754 to evaluate the HPF1 channel power difference between some (or all) of the channel pairs that may be formed between the channels eligible for blockage detection. In some aspects, a relatively large HPF1 channel power difference may indicate blockage, while a relatively small HPF1 channel power difference may indicate no blockage. For example, if the HPF1 channel power difference is evaluated as channel 1 HPF1 power-channel 2 HPF1 power and channel 2 is blocked, the channel 2 HPF1 power value will be zero or close to zero and the HPF1 power difference will be evaluated to be approximately equal to the channel 1 HPF1 power value. However, if neither channel 2 nor channel 1 is blocked, the HPF1 power difference will be evaluated to be a relatively low value (e.g., approximately zero if channel 1 and channel 2 are close to each other and both measure approximately the same sound amplitude). In some aspects, the third condition 763 may additionally or alternatively cause the occlusion detector 754 to perform occlusion detection based on evaluating a ratio between the LPF power and the HPF1 power determined by the feature extraction engine 730 for each channel.
[0119] In some examples, the friction flag 745 determined by the friction detection engine 750 for each channel may be provided as an input to the blockage detector 754 of the blockage detection engine 740. In some aspects, if the friction flag 745 is set for a given channel (e.g., if the friction flag 745 is set to true), the blockage detector 754 may be configured to not update some (or all) of the running power averages tracked for the given channel. For example, for LPF, HPF1, and / or HPF2 power values determined for a given channel, long-term and / or short-term power averages may be tracked. Since in many cases, friction blocking may occur shortly before blocking occlusion (e.g., a user's finger rubs on or near the edge of the microphone opening before later moving to block the microphone opening), the power value determined for the frame in which the friction flag 745 is set to true should not be used to update the short-term and / or long-term power averages maintained by the blockage detector 754. In some aspects, when the friction flag 745 is set to true for the current frame of a given channel, the blockage detector 754 may exclude the current frame power value from any average or other update of the power estimate for the given channel.
[0120] Based on evaluating the third condition 763, the blockage detector 754 may output an indication that a blockage event is detected or an indication that a blockage event is not detected (e.g., a binary blockage state may be determined for the current frame of each channel that is eligible for blockage detection). The detected blockage state of each channel may be provided to a blockage state machine 756, which may track the detected blockage state of each channel over time. In an illustrative example, the blockage state machine 756 may use the blockage state timing information to determine a blockage flag 755 for each channel. For example, the blockage state machine 756 may set the blockage flag 755 to true in response to the blockage detector 754 determining a blockage event for the current frame.
[0121] In some aspects, the blocking state machine 756 may set the blocking flag 755 to true based on the blocking detector 754 determining a blocking event for a plurality of consecutive frames or a predetermined percentage or number of frames within a confidence interval number of frames. In some cases, after entering the blocking state (e.g., by setting the blocking flag 755 to true), the blocking state machine 756 may remain in the blocking state for a dwell period before exiting the blocking state. For example, if after entering the blocking state, the blocking detector 754 subsequently determines at some future frame that a blockage is no longer detected, the blocking state machine 756 may exit the blocking state after waiting for a predetermined number of frames (e.g., a dwell period). In some aspects, the blocking state machine 756 and the friction state machine 744 may use the same dwell period or may use different dwell periods.
[0122] The friction flag 745 set by the friction detection engine 740 for each channel included in the multi-channel audio input 610 and the blockage flag 755 set by the blockage detection engine 750 for each channel may be provided to the combiner 770. The combiner 770 may determine the blockage flag 775 for each channel based on the corresponding friction flag 745 and the corresponding blockage flag 755 determined for each channel. In an illustrative example, the combiner 770 may implement an “or” operation such that if at least one (or both) of the friction flag 745 and the blockage flag 755 are set to true for a given channel included in the multi-channel audio input 610, the blockage flag 775 is set to true.
[0123] Figure 8 is an example Figure 6Schematic diagram of an example architecture of an illustrated spatial device selection engine 640. In one illustrative example, the spatial device selection engine 640 may receive as input an occlusion flag 775 determined for a current frame for each channel included in the multi-channel audio input 610, wherein the spatial device selection engine 640 calculates spatial information and selects a spatial filter in response to the occlusion flag 775 being set to true (e.g., the spatial device selection engine 640 may be triggered when the occlusion flag 775 is true, but may not be triggered when the occlusion flag 775 is false). For example, these systems and techniques may reduce the use of computational power by only triggering the spatial device selection engine 640 to calculate spatial information for occluded frames (e.g., frames for which the occlusion flag 775 is set to true) rather than calculating spatial information for every frame.
[0124] As illustrated, the spatial device selection engine 640 may receive as input an occlusion flag 775 determined by the occlusion detection engine 620 for each channel included in the multi-channel input 610 and may also receive as input one or more frames of delayed audio from the delayed audio database 630. For example, the occlusion flag 775 and the delayed audio frames from the delayed audio database 630 may be provided to a load audio data validation engine 810 included in the direction of arrival (DOA) estimation framework 800. The load audio data validation engine 810 may obtain (e.g., load) delayed audio for a given channel from the delayed audio database 630 in response to evaluating the occlusion flag 775 as true for the given channel. In some cases, the audio validation engine 810 may load a predetermined number of delayed audio frames for each channel if the occlusion flag 775 is set to true. For example, the audio validation engine 810 may obtain the most recent 128 delayed audio frames available for the corresponding channel in the delayed audio database 630.
[0125] The audio validation engine 810 may validate the loaded delayed audio frame based on evaluating the first condition 861 for the loaded delayed audio frame. In some aspects, the first condition 861 may include one or more sub-conditions for evaluating the loaded delayed audio frame. For example, the audio validation engine 810 may compare one or more of the L1 norm, L2 norm, average power, average amplitude, peak power, peak amplitude, VAD, etc. with one or more corresponding thresholds (e.g., one or more predetermined thresholds for each sub-condition included in the first condition 861). In some aspects, the first condition 861 may be evaluated for the loaded delayed audio frame. Figure 7 Some (or all) of the different power values determined at the illustrated feature extraction engine 730 evaluate (eg, may be determined for LPF power, HPF1 power, and / or HPF2 power) a power-based sub-condition.
[0126] As follows about Fig. 9As will be described in more depth, the audio validation engine 810 may use the first condition 861 to evaluate the suitability of the loaded delayed audio frame (e.g., obtained for each occluded channel based on the occlusion flag 775) in the DOA estimation performed by the DOA estimation engine 850. For example, the audio validation engine 810 may use the first condition 861 to evaluate whether the loaded delayed audio frame is delayed to an earlier time point at which an unoccluded audio frame is obtained for the current occluded channel. The audio validation engine 810 may additionally use the first condition 861 to evaluate whether the loaded delayed audio frame includes non-zero audio data (e.g., to verify that the loaded delayed audio frame does not capture silence or low-amplitude sounds).
[0127] In one illustrative example, the loaded delayed audio frame is evaluated by the audio validation engine 810 using the first condition 861 to verify whether the loaded delayed audio frame can be used for subsequent processing within the DOA estimation framework 800. For example, the loaded delayed audio frame can be verified for use in subsequent processing operations (such as DOA estimation, power estimation, delay estimation, etc.).
[0128] The spatial device selection engine 640 may operate on delayed audio frames obtained from the delayed audio database 630 when the corresponding microphone is not occluded in order to extract valid spatial information from the audio recently captured for a given channel. For example, an occlusion event may occur at any time, including during silence or low sound. It may be difficult or impossible to determine valid spatial information from audio frames corresponding to silence or low amplitude sound. By verifying that the delayed audio frames loaded from the delayed audio database 630 into the DOA estimation framework 800 do not correspond to periods of silence or low amplitude sound, these systems and techniques can improve computational efficiency by only performing spatial information extraction (e.g., determination) after verifying that spatial information can be obtained from the delayed audio frames loaded from the delayed audio database 630. For example, if the loaded delayed audio frame is not verified by the audio confirmation engine 810, the delayed database load index 820 may be moved. For example, the most recent 128 frames of delayed audio frames may initially be loaded into the audio validation engine 810; if the loaded frames are not verified, operation 820 may move the delay database load index and cause the audio validation engine 810 to load the second most recent set of 128 frames of delayed audio from the delayed audio database 630. The process of loading delayed audio frames and moving the load index with the delay database load index 820 may be repeated until delayed audio frames that pass the verification check of the audio validation engine 810 are loaded.
[0129] After passing the validation check of the audio validation engine 810, the loaded delayed audio frame may be provided as input to a DOA estimator 852 included in the DOA estimation engine 850. In some aspects, the DOA estimation technique may be selected based on evaluating the second condition 862. For example, the second condition 862 may cause the DOA estimator 852 to analyze channel map information associated with the multi-channel input signal 610 (e.g., the number of input / output channels, location information of the corresponding microphone for each channel, etc.). In some aspects, the second condition 862 may additionally or alternatively cause the DOA estimator 852 to analyze information (such as available beamformer types, channels available for DOA estimation, etc.). In some examples, the DOA estimator 852 may use the occlusion information determined from the occlusion flag 775 and the information obtained based on evaluating the second condition 862 to determine one or more DOA estimates associated with the occluded channel.
[0130] In some examples, one or more DOA estimates determined by the DOA estimator 852 may be provided as input to a DOA check engine 854, which evaluates the DOA estimates based on a third condition 863. In some aspects, the third condition 863 may be used to perform recursive DOA estimation or otherwise improve the DOA estimation performance of the DOA estimation engine 850. For example, the third condition 863 may cause the DOA check engine 854 to perform recursive DOA estimation based on information that may include, but is not limited to, time difference of arrival (ATD), level difference of arrival (ALD) or loudness difference, beam energy, estimated angle, cross correlation, peak correlation, delay difference, microphone geometry and / or placement information, etc. The recursive DOA estimation performed by the DOA check engine 854 may also be based on the occlusion flag information 775 (e.g., in the same or similar manner as the initial DOA estimation performed by the DOA estimator 852 may be based on the occlusion flag information 775).
[0131] The recursive DOA estimate output by the DOA check engine 854 may be the same as the output of the DOA estimation engine 850 and may be provided as an input to the spatial filter selection engine 880. The spatial filter selection engine 880 may also receive as input the same validated delayed audio frame loaded by the audio validation engine 810 (e.g., the same validated delayed audio frame provided from the audio validation engine 810 to the DOA estimation engine 850). In one illustrative example, the spatial filter selection engine 880 may select a spatial filter for recovering the spatial audio output of the occluded channel identified by the occlusion flag 775. In some aspects, the spatial filter selection may be performed based on evaluating the fourth condition 864. The fourth condition 864 may cause the spatial filter selection engine 880 to select a spatial filter based on information such as channel map information, occluded channel information, available beamformer types, etc. The selected spatial filter may be used by these systems and techniques to recover the spatial audio image by reconstructing the occluded audio channel using the non-occluded audio channel frame.
[0132] Fig. 9 is a diagram illustrating an example dual channel (e.g., stereo) spatial device selection engine 940. In some cases, the dual channel spatial device selection engine 940 may be used to implement the spatial device selection engine 640. In order to validate the loaded delayed audio frames from the delayed audio database 630, the delayed audio frames may be loaded into the DOA estimation framework 900 for two stereo channels (e.g., L and R channels). For example, the delayed stereo audio frames may be loaded from the delayed audio database 630 into the amplitude check engine 912. The amplitude check engine 912 may perform an amplitude check on the delayed audio frames obtained for the occluded stereo channels indicated by the occlusion flag 775. The amplitude check engine 912 may be used to validate that the delayed audio frames are non-zero values (e.g., have non-zero amplitudes and do not correspond to periods of silence or near silence), as described above with respect to Figure 8 The illustrated audio validation engine 810 is described.
[0133] The delayed stereo audio frame may then be provided to an amplitude check engine 914, which in some aspects may evaluate the delayed stereo audio frame based on the first condition 861. For example, the amplitude check engine 914 may determine the peak amplitude or peak energy of the delayed audio frame (e.g., the maximum amplitude of the delayed audio frame within the corresponding time window of loading the delayed audio frame). In some aspects, the amplitude check engine 914 may use the first condition 861 to evaluate whether the peak amplitude or peak energy included within the corresponding time window of the delayed audio frame is greater than or equal to a predetermined threshold (e.g., the predetermined threshold given by the first condition 861). If the peak amplitude is greater than or equal to the threshold of the first condition 861, the loaded delayed audio frame may proceed to the DOA estimation performed by the dual-channel DOA estimation engine 950. If the peak amplitude is less than the threshold of the first condition 861, the analysis time window may be moved further into the past (e.g., to an earlier time in the delayed audio database 630), and the delayed audio frame loading and validation may be based on moving the loading index to the delayed database loading index 820 (e.g., as described above with respect to Figure 8 The process of iteratively repeating the above process ) is repeated until a delayed audio frame having a peak amplitude greater than a threshold given by the first condition 861 is loaded.
[0134] In one illustrative example, a dual-channel (e.g., stereo) DOA estimation engine 950 may be used to perform delay estimation between two stereo channels. For example, the stereo DOA estimation engine 950 may transform the delayed audio frame (e.g., previously verified by the peak amplitude check 914) into the frequency domain for further analysis. In some cases, the delayed audio frame may be provided to a fast Fourier transform (FFT) 951 and a bandpass filter 953 to obtain a band-limited frequency domain representation of the delayed audio frame. A cross-correlation 955 may be determined based on the band-limited frequency domain representation. In some cases, by using the bandpass filter 953 to obtain a band-limited frequency domain representation of the delayed audio frame, the DC bias effect may be eliminated by ignoring bands below a certain frequency.
[0135] The cross correlation 955 may be transformed back to the time domain using an inverse FFT (IFFT) 957. The time domain cross correlation may be provided as an input to a peak detector 958, which may detect the peak position of the cross correlation (e.g., the peak position in the time domain). In one illustrative example, a delay difference between two stereo channels may be determined based on detecting the peak position of the cross correlation in the time domain. In some aspects, a recursive delay estimation may be performed to improve the estimated delay difference between the two stereo channels. For example, if the peak position (e.g., determined by the peak detector 958) is greater than half the FFT size, the estimated delay difference will be biased.
[0136] In some cases, the DOA check engine 854 may be used to validate the estimated delay difference by a correlation value and / or by a maximum delay difference given by the geometric microphone placement (e.g., the maximum delay difference possible given the distance between the two stereo microphones). For example, using the third condition 863, the DOA check engine 854 may validate that the peak correlation is greater than a predetermined threshold and / or the delay difference is less than or equal to the maximum delay difference possible. If the DOA estimate is validated by the DOA check engine 854, the DOA estimate (and / or associated DOA estimate information) may be provided as input to the spatial filter selection engine 880. For example, in response to the DOA estimate being validated by the DOA check engine 854, the spatial filter selection engine 880 may receive corresponding delay information and / or gain information associated with the DOA estimate.
[0137] If the DOA estimate is not validated by the DOA check engine 854 (e.g., if the peak correlation is less than a predetermined threshold of the third condition 863 and / or if the delay difference is greater than the maximum possible delay difference), the delay database loading index 820 may be moved and the spatial device selection process repeated. For example, the peak correlation determined between the delayed audio frames of the two stereo channels when using the first loading index may not be high enough (e.g., less than a threshold) to accurately determine the delay difference between the two channels, and the first loading index will not be validated by the DOA check engine 854. By returning to the operation performed by the delay database loading index 820, a different set of delayed audio frames of the two stereo channels may be obtained using the second loading index, and the cross correlation may be determined again using the same process as described above. The process described above may be iteratively repeated until a delay estimate is obtained with a peak correlation greater than or equal to a threshold given by the third condition 863. In some aspects, a termination condition may be included in the third condition 863. For example, the termination condition may indicate a maximum number of iterations to be performed. In some aspects, if 60 different delayed audio database load indexes have produced valid (e.g., non-zero) data but have not resulted in a peak correlation above a predetermined threshold, the DOA check engine 854 may provide the maximum correlation index found in a different iteration to the spatial filter selection engine 880.
[0138] Fig.10 1000 and a DOA estimation engine 1050.
[0139] The delayed audio frames may be loaded from the delayed audio database 630 for three or more channels and may be validated based on checking that the magnitude of the peak amplitude included in the delayed audio frame is greater than or equal to a predetermined threshold given by the first condition 1061. In some examples, the first condition 1061 may be associated with Figure 8 and Fig. 9 The illustrated condition 861 is the same or similar. In one illustrative example, delayed audio frames may be loaded for each of three or more channels. Amplitude determination 1012 may be obtained for each delayed audio frame and provided to peak amplitude check 1014, which determines the maximum (e.g., peak) amplitude included in the delayed audio frame loaded for each channel. In some examples, amplitude determination 1012 may be combined with Fig. 9 The amplitude determination 912 illustrated and described above may be the same or similar, and / or the peak amplitude check 1014 may be the same as that also described in Fig. 9 The peak amplitude check 914 illustrated and also described above is the same or similar.
[0140] In some aspects, if a delayed audio frame loaded using a first load index (e.g., where the first load index corresponds to a most recent set of delayed audio frames written to the delayed audio database 630) does not pass the peak amplitude check 1014 (e.g., has a peak amplitude less than a threshold given by the first condition 1061), the load index can be moved and the peak amplitude check 1014 can be iteratively repeated until a delayed audio frame including a peak amplitude greater than the threshold is loaded (e.g., in the same manner as described above with respect to Figure 8 and Fig. 9 The illustrated lazy database load index 820 operates in the same or similar manner as described above).
[0141] After the delayed audio frames loaded for each channel (e.g., of three or more channels) have been found through the delayed database loading index of the peak amplitude check 1014, the loaded delayed audio frames can be provided as input to the DOA estimation engine 1050. In one illustrative example, the DOA estimation engine 1050 can be configured to perform DOA estimation for the three or more channels of the delayed audio frames based on performing multi-microphone beamforming in the frequency domain. For example, the loaded and verified delayed audio frames for each channel can be transformed into the frequency domain using a frequency domain transform 1051. In some aspects, the frequency domain transform can be an FFT 1051, which can be used with Fig. 9 The illustrated FFT 951 is the same or similar (the corresponding IFFT 1055 can be the same as Fig. 9 The illustrated IFFT 957 is the same or similar).
[0142] In an illustrative example, the loaded and verified delayed audio frames may be used by the DOA estimation engine 1050 to perform multi-microphone beamforming in the FFT domain. For example, the DOA estimation engine 1050 may implement a plurality of beamformers 1053 that divide a 360-degree sound field (e.g., centered around a device including microphones for obtaining three or more channels of audio frames) into different parts or directions. A beamformer may be determined for each of the divided parts or directions of the 360-degree sound field. For example, the 360-degree sound field may be divided into 8 different directions, and each direction may be set as a beamformer.
[0143] Each beamformer included in the plurality of beamformers 1053 may be associated with a plurality of beamformer coefficients. For example, the beamformer coefficients associated with a given beamformer in the beamformers 1053 may be predetermined (e.g., offline and / or generated in advance) based on the direction of the beamformer and the geometric information of the microphone positions. For example, the geometric information may include geometric information of the microphone positions on the same audio capture device and / or geometric information of the positions of the microphones relative to each other. In an illustrative example, the plurality of beamformers 1053 may be implemented based on a set of beamformer coefficients designed for different directions of arrival (DOA) (e.g., different directions used to divide a 360 degree sound field around a multi-microphone device). In some aspects, the second condition 1062 may include one or more (or both) of the beamformer information (e.g., the beamformer direction and / or the set of beamformer coefficients) and the microphone position geometric information.
[0144] The FFT domain delayed audio for each channel may be provided as input to the set of beamformer coefficients 1053. Based on the set of beamformer coefficients, a beamformer output for each given beamformer / direction may be generated based on combining the FFT domain delayed audio for each channel as weighted by the corresponding beamformer coefficients. The FFT beamformer output for each direction may be transformed back to the time domain using an IFFT 1055. The time domain beamformer output for each direction may be provided as input to an energy calculation engine 1057, which determines the energy of each beamformer output. The DOA selection engine 1054 may be based on evaluating the energy of the beamformer output, which may be similar in some aspects to the energy calculation engine 1057. Figure 8 and Fig. 9 The illustrated third condition 863 is executed by a third condition 1063 that is the same as or similar to the third condition 863 .
[0145] In one illustrative example, evaluating the third condition 1063 may cause the DOA selection engine 1054 to select the direction with the highest energy in the corresponding beamformer output. For example, the estimated DOA determined by the DOA estimation engine 1050 may be the direction with the highest energy of the beamformer output. The estimated DOA may be provided as an input to the spatial filter selection engine 880, which may select a spatial filter to use to restore spatial audio processing when one or more microphones are occluded. For example, the spatial filter selection engine 880 may be associated with the spatial filter selection engine 880. Figure 8 and / or Fig. 9 Spatial filter selection is performed in the same or similar manner as described for spatial filter selection performed by spatial filter selection engine 880 using the estimated DOA from DOA estimation engine 1050 and based on evaluating fourth condition 1064 .
[0146] Fig.11 is an example that can be used to implement Figure 6 650. As illustrated, the selective spatial filtering engine 650 may obtain delayed audio frames (e.g., for Figure 6 The first spatial filtering path may include a first switch 1122 and a first spatial filter 1124 for non-occluded (e.g., normal) spatial filtering. The first spatial filtering path may also be referred to as a non-occluded spatial filtering path. The second spatial filtering path may include a second switch 1132 and a second spatial filter 1134 for occluded spatial filtering. The second spatial filtering path may also be referred to as an occluded spatial filtering path.
[0147] In some examples, the selective spatial filtering engine 650 may be integrated into or implemented using one or more processors. For example, the selective spatial filtering engine 650 may use Fig.17 The illustrated processor 1710 is implemented. In some cases, the selective spatial filter engine 650 may use Figure 1 The processor included in the multi-microphone device 102 can be implemented with Fig.17 The processor 1710 illustrated is the same or similar. In some examples, the selective spatial filter engine 650 may be used on a processor (e.g., such as Fig.17 The embodiment is implemented by firmware running on the illustrated processor 1710).
[0148] In some aspects, the first spatial filter 1124 may be a non-occluded spatial filter, and the second spatial filter 1134 may be an occluded spatial filter. In some examples, the first spatial filter 1124 may include one or more filter groups for performing non-occluded spatial filtering. For example, one or more filter groups may each be a different filter group that generates a spatial audio output using different combinations of microphone channel inputs. In some examples, the second spatial filter 1134 may include one or more filter groups for performing occluded spatial filtering. For example, one or more filter groups included in the second spatial filter 1134 may each be a different filter group that generates a spatial audio output using different combinations of microphone channel inputs. In some aspects, the spatial audio output generated using one or more filter groups included in the second spatial filter 1134 may be a restored spatial audio output generated for an occluded microphone using one or more audio frames captured using one or more non-occluded microphones as input. In some cases, some (or all) of the one or more filter groups included in the first spatial filter 1124 (e.g., a non-occluded spatial filter) may be different from the one or more filter groups included in the second spatial filter 1134 (e.g., an occluded spatial filter). In some aspects, the number of filter groups included in the first spatial filter 1124 may be greater than or equal to the number of filter groups included in the second spatial filter 1134 .
[0149] In some examples, the first switch 1122 and the second switch 1132 may be implemented using a single switch (e.g., the single switch may include the first switch 1122 and the second switch 1132). For example, the single switch implementing the first switch 1122 and the second switch 1132 may receive as input one or more delayed audio frames from the delayed audio database 630, the first condition 1171, and the second condition 1172. Based at least in part on the first condition 1171 and the second condition 1172, the single switch may control (e.g., open and close) the first switch 1122 and may control (e.g., open and close) the second switch 1132.
[0150] In one illustrative example, the first switch 1122 may be controlled based on evaluating the first condition 1171 and the second switch 1132 may be controlled based on evaluating the second condition 1172. For example, the first condition 1171 and the second condition 1172 may be evaluated based on an occlusion flag (e.g., such as the occlusion flag 775), an occlusion start / end transition, etc.
[0151] In one illustrative example, when there is no occlusion in the currently processed frame of the delayed audio, the first switch 1122 can be closed (e.g., the non-occluded spatial filtering path can be utilized) and the second switch 1132 can be opened (e.g., the occluded spatial filtering path is disabled). Since the occluded spatial filtering path is disabled when there is no occlusion in the currently processed frame (e.g., by opening the second switch 1132), the output of the spatial filter 1124 can be the same as the final spatial audio output of the selective spatial filtering engine 650 (e.g., the input and output of the smart merging engine 1160 are the same).
[0152] In another illustrative example, when one or more of the currently processed frames of the delayed audio are occluded, the first switch 1122 may be opened (e.g., the non-occluded spatial filtering path is not utilized) and the second switch 1132 may be closed (e.g., the occluded spatial filtering path is utilized). Since the non-occluded spatial filtering path is disabled when the currently processed frame is occluded (e.g., based on opening the first switch 1122), the output of the spatial filter 1124 may be the same or similar to the final spatial audio output of the selective spatial filtering engine 650 (e.g., the smart merging engine 1160 does not merge the outputs of the non-occluded spatial filtering path and the occluded spatial filtering path).
[0153] For example, the first switch 1122 may be controlled based on evaluating the first condition 1171, such that the first switch 1122 is closed when the occlusion flag 775 is false (e.g., there is no occlusion) and is opened when the occlusion flag 775 is true (e.g., there is occlusion). In some aspects, the second switch 1132 may be controlled based on evaluating the second condition 1172, such that the second switch 1132 is closed when the occlusion flag 775 is true (e.g., there is occlusion) and is opened when the occlusion flag 775 is false (e.g., there is no occlusion).
[0154] Occluded spatial filtering 1134 may be performed based on estimated DOA information determined using spatial device selection engine 640 (and / or spatial device selection engine 940 and / or spatial device selection engine 1040). In some cases, occluded spatial filtering may be performed using Figures 8 to 10 The spatial filter selection engine 880 illustrated in one or more of the above examples may be used to determine the spatial filter to implement the occluded spatial filtering 1134. Fig.12 and Fig.13 Aspects of occluded spatial filtering are described in more depth.
[0155] refer to Fig.11The illustrated architecture of the selective spatial filter engine 650 may perform occluded spatial filtering 1134 based at least in part on the third condition 1173. For example, the third condition 1173 may include information such as the number of channels, non-occluded channel maps, occluded channel information, channel map information, etc. The occluded spatial filtering 1134 may additionally or alternatively receive as inputs the occlusion flag 775 and DOA information 1185. For example, the DOA information 1185 may include delay differences between channels (e.g., for stereo input), estimated DOA (e.g., for 3+ channel inputs where beamforming is used), etc.
[0156] The smart merging engine 1160 may be used to merge the non-occluded spatial filtering 1124 output and the occluded spatial filtering 1134 output. For example, during the occlusion start and / or occlusion end period, both the first switch 1122 and the second switch 1132 may be closed (e.g., and both the non-occluded spatial processing path and the occluded spatial processing path may be activated). During the occlusion start and / or occlusion end transition, the smart merging engine 1160 may perform a cross-fade and merge between the occluded spatial output and the non-occluded spatial output to provide a smoother transition that is less audible or not audible at all to the listener. In an illustrative example, the occlusion start and / or occlusion end transition period may be shorter than a delay period associated with a delayed audio database (e.g., the delayed audio database 630). For example, the occlusion start and / or occlusion end transition period associated with the smart merging engine 1160 may be shorter than a delay period associated with a delayed audio database (e.g., the delayed audio database 630). Figure 6 The delay filter z -m 615 (eg, which delays the multi-channel audio input 610 by m samples before being written to the delayed audio database 630). Figure 6 As illustrated, the selective spatial filtering engine 650 (eg, which may include a smart merging engine such as the smart merging engine 1160) may receive and delay filter z -m 615 associated information as input. In some aspects, the occlusion start transition period and the occlusion end transition period may be the same. In an illustrative example, the occlusion start and occlusion end transition periods may be between 30 milliseconds (ms) and 40 milliseconds.
[0157] In some cases, the occlusion start and / or occlusion end periods can be used to provide a smooth transition between the non-occluded spatial filter processing path output (e.g., associated with the first switch 1122) and the occluded spatial filter processing path output (e.g., associated with the second switch 1132). In some aspects, the occlusion start period associated with the smart merge engine 1160 can be used to provide latitude or tolerance for slow onset detection of occlusion events. For example, when no occlusion event is detected (e.g., not detected by the first switch 1122), the occlusion end period can be used to provide a smooth transition between the non-occluded spatial filter processing path output (e.g., associated with the second switch 1132). Figure 6 and Figure 7Slow detection of an occlusion event may occur when the illustrated occlusion detection engine 620 detects one or more initial audio frames corresponding to an occlusion event. In some cases, slow detection of an occlusion event may be more likely to occur if the occlusion occurs at a frame boundary. For example, based on the occlusion start period being less than a delay period associated with the delayed audio database 630, the smart merge engine 1160 may trigger and perform a merge even if the first frame of the occlusion event is missed or otherwise not detected.
[0158] In some aspects, the smart merging engine 1160 may perform smart merging based on the fourth condition 1174. For example, the fourth condition 1174 may be based on or include channel map information, occluded channel information, etc. In some cases, the smart merging engine 1160 may additionally or alternatively receive as input the occlusion flag 775 and DOA information 1185 (e.g., the same or similar to the occlusion flag 775 and DOA information provided as input to the occluded spatial filtering 1134).
[0159] Fig.12 1 is a diagram illustrating an example architecture of a dual channel (eg, stereo) selective spatial filter engine 1250 that may be used to implement the selective spatial filter engine 650. The stereo selective spatial filter engine 1250 may include Fig.11 The first switch 1222 illustrated is the same as or similar to the first switch 1122 and may include Fig.11 The illustrated second switch 1132 may be the same as or similar to the second switch 1232. In some examples, the first condition 1271 associated with the first switch 1122 may be the same as Fig.11 The illustrated first condition 1171 associated with the first switch 1222 is the same or similar. In some examples, the second condition associated with the second switch 1232 may be the same as Fig.11 The illustrated second condition 1172 associated with the second switch 1132 is the same or similar.
[0160] The first switch 1222 can be used to activate and deactivate the non-occlusion spatial processing path (e.g., as described above with respect to Fig.11 For example, when the first switch 1222 is closed, the non-occlusion spatial processing path is activated, and the delayed audio frame from the delayed audio database 630 can be provided as input to the non-occlusion spatial filtering 1224. In some examples, the non-occlusion spatial filtering 1224 can be performed based on bypassing the delayed signal associated with the output of the delayed audio database 630 (e.g., when there is no occlusion, the delayed audio frame from the delayed audio database 630 can be provided as input to the non-occlusion spatial processing path).
[0161] The stereo selective spatial filtering engine 1250 may receive as input (e.g., from the delayed audio database 630) delayed audio frames corresponding to a first channel and a second channel (e.g., L and R channels), one of which is occluded and one of which is not occluded. For example, occlusion may be detected in a current frame of the L or R channel, while the remaining channel is not occluded in its current frame. However, the selective spatial filtering engine 1250 obtains the delayed audio frames as input from the delayed audio database 630, and does not provide the occluded audio frames (e.g., associated with occlusion detected from the current frame) to the selective spatial filtering engine 1250 within a time period equal to the delay associated with implementing the delayed audio database 630.
[0162] In an illustrative example, during an occlusion onset period (e.g., the time period between detecting occlusion for a current audio frame and subsequently writing the current audio frame to and loading the current audio frame from the delayed audio database 630), the selective spatial filtering engine 1250 may cross-fade between a spatial output generated using non-occluded spatial filtering 1224 and a restored spatial output generated using occluded spatial filtering 1234 and merge the two spatial outputs.
[0163] The occluded spatial filtering 1234 may use the non-occluded stereo channels to generate a restored signal for the occluded stereo channels. In some examples, the occluded spatial filtering 1234 may be performed based on a third condition 1273, which may be related to the condition in Fig.11 1134. For example, if channel 1 is occluded and channel 2 is not occluded, the occluded spatial filtering 1234 may use the non-occluded channel 2 to generate the restored channel 1 signal. If channel 1 is not occluded and channel 2 is occluded, the occluded spatial filtering 1234 may use the non-occluded channel 1 to generate the restored channel 2 signal. In some aspects, the restored channel signal may be generated based on the audio frame of the non-occluded channel and the information associated with the third condition 1273 (e.g., occluded channel information, occlusion flag, channel map information, etc.). In an illustrative example, the occluded channel recovery engines 1235a, 1235b (e.g., respectively associated with generating the restored channel 1 signal and the restored channel 2 signal) may receive as input the delayed audio frames loaded from the delayed audio database 630. The occluded channel and the non-occluded channel may be identified based on the occlusion flag 775. The delayed audio frames associated with the non-occluded channel and the delay difference 1287 may be used to generate a restored signal for the occluded channel.
[0164] For example, the delay difference 1287 may be included in Fig.11In the illustrated DOA information 1185, the delay difference can be the same or similar to the estimated delay difference determined using the spatial device selection engine 640 and / or the spatial device selection engine 940. In some aspects, the delay difference 1287 can indicate an estimated delay between the same sound source being detected at a microphone associated with the occluded channel and being detected at a microphone associated with the non-occluded channel. In some aspects, the audio frames of the occluded channel can be adjusted by the estimated delay difference to generate a restored channel signal for the occluded channel. In one illustrative example, the estimated delay difference can be used based at least in part on the use of Figure 8 The illustrated spatial device selection engine 640 and Fig. 9 The spatial filter selected by the spatial filter selection engine 880 included in the illustrated spatial device selection engine 940 implements the occluded spatial filtering 1234 .
[0165] The recovered channel signals and the non-occluded channel signals may be provided to the smart merging engine 1260 as a recovered stereo signal 1263. In an illustrative example, the smart merging engine 1260 may apply one or more of a gain (e.g., loudness) adjustment 1267 and / or a coloration adjustment 1269. For example, the gain adjustment 1267 may be implemented as an increase or decrease in the gain (e.g., loudness) of the recovered stereo signal 1263. In some aspects, the coloration adjustment 1269 (e.g., also referred to as a coloration shift) may be implemented to change the coloration (e.g., pitch, timbre, etc.) of the recovered stereo signal 1263. For example, the coloration adjustment 1269 may be a coloration shift determined based on geometric information associated with one or more microphones used to capture the audio frames obtained from the delayed audio database 630.
[0166] In some aspects, the smart merging engine 1260 may be configured using a fourth condition 1274. For example, the fourth condition 1274 may include or otherwise be based on information, which may include, but is not limited to, occluded channel information, occlusion state information, and the like. In some examples, the smart merging engine 1260 may apply a gain adjustment 1267 and / or a coloration adjustment 1269 to the restored stereo signal 1263. In some aspects, the gain adjustment 1267 may be estimated or determined based on geometric information of two microphones associated with two stereo channels. For example, the estimated gain difference may be determined based on the distance between the occluded microphone and the non-occluded microphone. If the occluded microphone is closer to the sound source associated with the DOA estimate than the non-occluded microphone, the gain adjustment 1267 may be selected to increase the loudness of the restored channel signal of the occluded channel. If the occluded microphone is farther away from the sound source than the non-occluded microphone, the gain adjustment 1267 may be selected to reduce the loudness of the restored channel signal of the occluded channel.
[0167] Coloration adjustments 1269 may be estimated or determined based on geometric information of two microphones associated with two stereo channels. For example, these systems and techniques may determine that a sound source signal captured by an occluded microphone is associated with a coloration that is different from the same sound source signal captured by an unobstructed microphone. Since the recovered occluded channel signal is generated based on applying a delay adjustment and a gain adjustment to the unobstructed channel signal, the coloration of the recovered occluded channel signal and the unobstructed channel signal will be the same or similar. Based on applying the coloration adjustments 1269 to the recovered occluded channel signal, the intelligent merging engine 1260 may recover the coloration difference that existed between the two stereo microphones when both stereo microphones captured an unobstructed sound signal.
[0168] The recovered stereo signal 1263 and the occluded stereo signal 1262 may be provided as inputs to a crossfade and merge engine 1268 included in the smart merge engine 1260. As previously mentioned, the smart merge engine 1260 may crossfade between the two signals 1262 and 1263 during a transition window between the time when occlusion is detected for the current frame and a future time when the current frame is loaded from the delay database 630 and loaded into the selective spatial filter engine 1250. The output of the crossfade and merge engine 1268 (e.g., as well as the output of the selective spatial filter engine 1250) is a recovered stereo audio signal 1295 that is crossfaded between normal (e.g., non-occluded) spatial filtering 1224 at the time when occlusion is detected and occluded spatial filtering 1234 at a subsequent time when the occluded audio frame is loaded from the delay audio database 630 and processed by the selective spatial filter engine 1250.
[0169] Fig.13 1 is a diagram illustrating an example architecture of a selective spatial filter engine 1350 that can be used to perform selective spatial filtering on a multi-channel input signal including three or more audio channels. In some aspects, the selective spatial filter engine 1350 can be used to implement the selective spatial filter engine 650. The selective spatial filter engine 1350 can include an example architecture that can be used with Fig.11 The illustrated first switch 1122 is the same as or similar to and / or Fig.12 The selective spatial filter engine 1350 may additionally include a first switch 1322 that is the same as or similar to the first switch 1222 illustrated. Fig.11 The illustrated second switch 1132 and / or Fig.12 The second switch 1332 is the same as or similar to the second switch 1232 shown in the example. The first switch 1322 can be based on Fig.11 The first condition 1171 and / or Fig.12The second switch 1332 may be opened and closed (eg, controlled) by a first condition 1371 that is the same as or similar to the first condition 1271 illustrated. Fig.11 The second condition 1172 and / or Fig.12 The illustrated second condition 1272 is the same as or similar to the second condition 1372 for opening and closing (eg, controlling).
[0170] In one illustrative example, the selective spatial filtering engine 1350 may generate a reconstructed channel signal for one or more occluded channels by using beamforming. For example, the selective spatial filtering engine 1350 may perform spatial filtering by using a Fig.10 The illustrated spatial device selects 1040 the spatial filters selected to generate one or more beamforming outputs to generate reconstructed channel signals of one or more occluded channels to selectively restore the spatial audio output. In some aspects, the spatial filter may be selected based on Fig.10 The illustrated DOA selection engine 1054 and / or spatial filter selection engine 880 are used to generate reconstructed channel signals of one or more occluded channels as beamformer outputs.
[0171] In some aspects, a non-occlusion spatial filtering processing path may be implemented based on generating one or more beamformer outputs 1326. For example, Fig.14A 1 is a diagram illustrating an example of four beamformer outputs that may be generated using three audio frames captured by three non-obstructed microphones of device 102. As illustrated, a left beamformer output, a right beamformer output, a front beamformer output, and a rear beamformer output may be generated based on audio frames captured by three microphones (e.g., microphone 1, microphone 2, and microphone 3). For example, a first microphone may capture a first audio frame X. 1 , the second microphone can capture the second audio frame X 2 , and the third microphone may capture the third audio frame X 3 -. In some aspects, each beamformer output (e.g., of a total of four beamformer outputs) may be generated as W 1 X 1 (n,K)+W 2 X 2 (n,K)+W 3 X 3 (n,K), where W 1 -W 3 is the beamformer weight for a given beam direction (e.g., each of the four beamformers may have a different W 1 -W 3 weight value), and n and K are time and frequency indices respectively.
[0172] about Fig.13 , the non-occluded spatial processing path may use P microphone inputs to generate multiple beamformer outputs 1326. For example, if P = 3 (e.g., as in Fig.14A , which includes P=3 microphones), then each of the plurality of beamformer outputs 1326 may be generated based on applying different beamformer weights or coefficients to an audio frame captured using the corresponding P=3 microphones. When one or more microphones are occluded (e.g., when one or more of the P microphones / channels are occluded), the audio frame X 1 , X 2 , X 3 The corresponding audio frames in may include amplitude information of a degraded or zero value. Since the beamformer outputs 1326 are each generated as W 1 X 1 (n,K)+W 2 X 2 (n,K)+W 3 X 3 (n,K), thus including 1 , X 2 and / or X 3 The non-zero beamformer coefficients or weighted beamformer outputs of one or more occluded audio frames in the image may lose spatial perception.
[0173] For example, in Fig.14A In the example, the three microphones (eg, microphone 1, microphone 2, microphone 3) may not be blocked. Fig. 14B As depicted, one of the microphones (e.g., microphone 3) may be occluded, while the two remaining microphones (e.g., microphone 1 and microphone 2) are not occluded. When microphone 3 is occluded, the audio frame X captured by the occluded microphone 3 3 A zero or near zero amplitude may be measured (e.g., based on microphone 3 being completely blocked by a user's hand or finger). In some aspects, when microphone 3 is blocked, the audio frame X captured by the blocked microphone 3 is 3 The amplitude of the distortion or degradation may be measured (eg, the amplitude may increase based on noise from a user's finger causing frictional occlusion, may decrease to a non-zero value based on partial blocking of the occlusion, etc.).
[0174] In some cases, non-occluded audio frames X are used 1 and X 2 and the occluded audio frame X 3 One or more or all of the beamformer outputs generated by the method may be corrupted or otherwise lose spatial sense. For example, as previously mentioned, for coefficients W including non-zero values 3Any beamformer output of , generated using a set of non-occluded beamformer coefficients (e.g., based on W 1 X 1 (n,K)+W 2 X 2 (n,K)+W 3 X 3 The beamformer output 1326 generated by (n,K) may be corrupted or lose spatial sense.
[0175] In one illustrative example, the systems and techniques described herein can generate a plurality of spatially recovered beamformer outputs 1336, where the recovered beamformer outputs are generated using PK non-occluded microphone channels (e.g., where P represents the total number of microphone channels and K represents the number of non-occluded microphone channels) as inputs for generating the spatially recovered beamformer outputs. For example, when Fig. 14B When microphone 3 is occluded as illustrated (eg, where microphone 1 and microphone 2 remain unobstructed), the spatially recovered beamformer outputs may each be generated as W′ 1 X 1 (n,K)+W′ 2 X 2 (n, K). In some aspects, each of the four spatially recovered beamformer outputs (eg, left, right, front, back) may use different corresponding beamformer coefficients W' 1 and W' 2 to generate.
[0176] In one illustrative example, the Fig.13 The selective spatial filtering engine 1350 generates Fig. 14B The spatially recovered beamformer output is depicted in . For example, the beamformer coefficients W' 1 and W′ 2 Available Fig.10 The spatial filter selection engine 880 included in the illustrated 3+ channel spatial device selection engine 1040 determines and / or may use Figure 6 The spatial filter selection engine 880 included in the illustrated spatial device selection engine 640 determines.
[0177] In some aspects, spatially recovered beamforming 1336 may be performed using audio frames associated with a set of PK non-occluded microphones determined based on the occlusion flag 775 and / or the DOA information 1185. In some examples, spatially recovered beamforming 1336 may be performed based at least in part on the third condition 1373. For example, the third condition 1373 may be associated with Fig.12 The third condition 1273 and / or Fig.11The illustrated third condition 1173 is the same or similar. The spatially restored beamforming 1336 may output a set of beamformers that are the same or similar to the beamformers output by the non-occluded beamforming 1326 (e.g., both the spatially restored beamforming 1336 and the non-occluded beamforming 1326 may output the same number of beamformers each having the same corresponding direction, etc.). The spatially restored beamforming 1336 may output a beamformer generated using PK non-occluded outputs, while the non-occluded beamforming 1326 may output a beamformer generated using all P microphone inputs.
[0178] The restored beamforming output of the spatially restored beamforming 1336 and the non-occluded beamforming output of the non-occluded beamforming 1326 may be provided as input to a smart merging engine 1360, which may cross-fade and merge the restored stereo signal and the occluded stereo signal with the smart merging engine 1220 (e.g., as described above with respect to Fig.12 The two beamformed outputs are cross-faded and combined in the same or similar manner as described above.
[0179] In one illustrative example, the smart merging engine 1360 may include a coloration correction engine 1364 that receives the spatially restored beamforming 1336 output and performs coloration correction based on the occlusion flag 775 and the DOA information 1185. For example, the coloration correction engine 1364 may be used in conjunction with the above description of the coloration correction engine 1364. Fig.12 The smart merge engine 1260 may perform coloration correction in the same or similar manner as described above for coloration correction performed by the illustrated smart merge engine 1260. In some examples, the smart merge engine 1360 may be implemented or otherwise controlled based on the fourth condition 1374. For example, the fourth condition 1374 may be related to Fig.12 The fourth condition 1274 and / or Fig.11 The illustrated fourth condition 1174 is the same or similar. In some aspects, the fourth condition 1374 may include or indicate occluded channel information, occlusion status, etc.
[0180] Coloration correction 1364 may be performed to provide coloration adjustments (e.g., color shifts) to one or more (or all) of the spatially restored beamformed 1336 outputs. For example, the spatially restored beamformer 1336 and the non-occluded beamformer 1326 may perform coloration correction 1364 to provide coloration adjustments (e.g., color shifts) to one or more (or all) of the spatially restored beamformed 1336 outputs. Fig.14A and Fig. 14B The right beamformer, left beamformer, front beamformer, and rear beamformer are shown as examples, but may be generated based on the use of (eg, based on W' 1 X 1 (n,K)+W' 2 X2 (n, K)) spatially restored beamformer 1336 different microphone audio frame combinations (e.g., W 1 X 1 (n,K)+W 2 X 2 (n,K)+W 3 X 3 (n, K)) to generate the non-occluded beamformer 1326 and associated with different sound colorations. In some aspects, the coloration correction 1364 can determine the difference between the coloration information associated with each respective one of the non-occluded beamformed outputs 1326 and the corresponding spatially restored beamformed output 1336. Based on the coloration difference of each given pair of the non-occluded beamformed output 1326 and the corresponding spatially restored beamformed output 1336, the coloration correction 1364 can adjust the spatially restored beamformed output 1336. For example, the coloration of the spatially restored beamformed output 1336 can be adjusted (e.g., by the coloration correction 1364) to match the expected coloration of the corresponding non-occluded beamformed output 1326.
[0181] Fig.15 1 is a diagram illustrating an example of spatial filter selection associated with a set of non-occluded microphone audio frames and an example of spatial filter selection associated with a combination of occluded microphone audio frames and non-occluded microphone audio frames. For example, non-occluded spatial filter selection 1510 may be performed to generate a left channel beamformer 1512 and a right channel beamformer 1514 using audio frames captured by four non-occluded microphones (e.g., microphone 1, microphone 2, microphone 3, microphone 4) as inputs. In some aspects, when all four microphone inputs are available for selecting spatial filters to generate left beamformer output 1512 and right beamformer output 1514, left beamformer output 1512 may be generated based on input audio frames captured by microphone 1 and input audio frames captured by microphone 3. Right beamformer output 1514 may be generated based on input audio frames captured by microphone 2 and microphone 4.
[0182] If one of the four microphone inputs becomes occluded, a restored spatial filter selection 1520 may be performed to generate a restored left channel beamformer 1522 and a restored right channel beamformer 1524. For example, if an occlusion is detected for microphone 2, the right channel beamformer 1514 may no longer be accurately generated (e.g., because the right channel beamformer 1514 is generated based on using non-occluded audio frames captured by microphone 2 and microphone 4). The left channel 1512 is not generated based on the audio frames captured by microphone 2, so in some examples, the restored left channel beamformer 1522 and the non-occluded left channel beamformer 1512 may be the same.
[0183] A recovered right channel beamformer 1524 may be generated to recreate the right channel beamformer 1514 previously generated using the now occluded microphone 2 audio frame. Figure 6 and Figure 8 The illustrated spatial device selection engine 640, Fig. 9 The illustrated spatial device selection engine 940 and / or Fig.10 The illustrated spatial device selection engine 1040 may be used to select the best available or optimal spatial filter (e.g., beamformer) given a set of available non-occluded microphones (e.g., microphone 1, microphone 3, and microphone 4). In one illustrative example, the spatial selection device may select or otherwise determine a spatial filter or beamformer for generating a reconstructed right channel 1524 using non-occluded audio frames captured by microphone 3 and microphone 4. In some aspects, the spatial selection device may provide a set of beamformer coefficients for generating the reconstructed right channel beamformer 1524 to a selective spatial filtering engine, such as Figure 6 and Fig.11 The illustrated selective spatial filtering engine 650, Fig.12 The illustrated spatial filter engine 1250 and / or Fig.13 The illustrated selective spatial filtering engine 1350. Based on the spatial filters and / or beamformers determined by the spatial device selection engine described herein, the selective spatial filtering engine described herein can generate reconstructed channel signals for one or more occluded channels or beamformers. Based on the reconstructed channel signals, these systems and techniques can generate a restored spatial audio output that maintains a sense of space when one or more microphones or input audio channels are occluded.
[0184] Fig.1616 is a flow chart illustrating an example of a process 1600 for audio signal processing. At block 1602, the process 1600 includes detecting an occlusion of one or more microphones associated with at least one audio frame of one or more audio frames associated with a spatial audio recording. For example, the occlusion may be detected in connection with using a multi-microphone device such as Figure 1 The occlusion of at least one audio frame in one or more audio frames associated with the spatial audio recording generated by the illustrated multi-microphone device 102. In some cases, detecting the occlusion of at least one audio frame may be based on a channel map indicating the number of microphones that are blocked or rubbed. For example, the microphones may include Figure 1 The illustrated (also in FIG. 2 and Figure 3 One or more microphones 104a to 104e included on the multi-microphone device 102 (illustrated in FIG. 1 ).
[0185] In some examples, detecting occlusion of at least one audio frame includes detecting occlusion of a microphone associated with capturing the at least one audio frame. Figures 1 to 3 Occlusion of one or more (or all) of the illustrated microphones 104a to 104e. In some examples, a different one of the microphones may be used to obtain each audio frame included in the one or more audio frames. Each of the different microphones may be included in the same device (e.g., such as Figure 1 4, etc.). In some cases, the spatial audio recording can be a stereo audio recording including a first audio channel associated with at least a first microphone and a second audio channel associated with at least a second microphone. In some examples, the first microphone can be included on a first earpiece, and the second microphone can be included on a second earpiece associated with the first earpiece.
[0186] In some examples, the occlusion of the microphone can be a friction occlusion or a blocking occlusion. Figure 6 and Figure 7 An occlusion detection engine similar to or identical to the occlusion detection engine 620 illustrated in the example can be used to detect occlusion of the microphone. Figure 7 The friction detection engine 740 illustrated is the same as or similar to the friction detection engine 740 to detect friction occlusion. Figure 7 The occlusion detection engine 750 illustrated may be the same or similar to the occlusion detection engine to detect occlusion. In some cases, detecting occlusion may include generating Figures 7 to 13 The occlusion mark 775 shown is the same as or similar to the occlusion mark. In some cases, the occlusion mark 775 may be the same as or similar to the occlusion mark 775. Figure 6 and Figure 7Occlusion detection is performed on each audio channel of the illustrated multi-channel audio input 610).
[0187] In some examples, one or more audio frames may be sent to a delayed audio database. Figure 6 and Figures 8 to 13 The illustrated delayed audio database 630 sends one or more audio frames. The one or more audio frames may be written to the delayed audio database after a predetermined delay has elapsed. The predetermined delay may be the same as the Figure 6 The delay filter z - m 615 The associated delay periods are the same or similar.
[0188] At block 1604, process 1600 includes selecting at least one of an occluded spatial filter for one or more audio frames or a non-occluded spatial filter for one or more audio frames based on the detection of occlusion. For example, in some cases, process 1600 may select an occluded spatial filter for filtering the one or more audio frames and may select a non-occluded spatial filter for filtering the one or more audio frames. In some examples, process 1600 may select an occluded spatial filter for performing occluded spatial filtering on the one or more audio frames and may select a non-occluded spatial filter for performing non-occluded spatial filtering on the one or more audio frames. For example, the process 1600 may use the occluded spatial filter to filter the one or more audio frames. Figure 6 and Fig.11 The illustrated selective spatial filtering engine 650, Fig.12 The illustrated 2-channel selective spatial filtering engine 1250 and / or Fig.13 The illustrated 3+ channel selective spatial filter engine 1350 may be the same as or similar to a selective spatial filter engine to perform selective switching.
[0189] In some examples, performing occluded spatial filtering of one or more audio frames includes determining an estimated direction of arrival (DOA) associated with an occluded microphone used to obtain at least one audio frame associated with the occlusion. The occluded microphone may be associated with an occlusion marker (such as using Figure 6 and Figure 7 The occlusion detection engine 620 may be associated with or identified based on the occlusion flag 775 determined by the illustrated occlusion detection engine 620. In some cases, a spatial filter may be determined for generating a reconstructed signal of the occluded microphone, and the determined spatial filter may be used to perform occluded spatial filtering.
[0190] For example, you can use Figure 6 and Figure 8 The illustrated spatial device selection engine 640, Fig. 9 The illustrated 2-channel spatial device selection engine 940 and / or Fig.10 A spatial device selection engine similar to or identical to the illustrated 3+ channel spatial device selection engine 1040 determines a spatial filter for generating a reconstructed signal for the occluded microphone. In some examples, the determined spatial filter may be used (e.g., with Figure 6 and Fig.11 The illustrated selective spatial filtering engine 650, Fig.12 The illustrated 2-channel selective spatial filtering engine 1250 and / or Fig.13 The occluded spatial filtering may be performed by using an occluded spatial audio processing path included in a selective spatial filtering engine (same or similar to the illustrated 3+ channel selective spatial filtering engine 1350).
[0191] In some examples, performing occluded spatial filtering using a spatial filter may be based on one or more non-occluded audio frames included in the one or more audio frames, where the one or more non-occluded audio frames are not detected to be occluded. For example, the one or more non-occluded audio frames may be associated with different values of the occlusion flag 775 than the one or more occluded audio frames. In some cases, the one or more non-occluded audio frames may not be associated with the occlusion flag 775, and the one or more occluded audio frames may each be associated with the occlusion flag 775.
[0192] In some cases, selectively switching between performing occluded spatial filtering on one or more audio frames and performing non-occluded spatial filtering on one or more audio frames during spatial recording may include selecting, after a predetermined delay has elapsed, from a delayed audio database (e.g., such as Figures 6 to 13 The example delayed audio database 630) obtains one or more audio frames and corresponding one or more delayed audio frames. In some cases, selectively switching may include cross-fading between performing occluded spatial filtering on one or more delayed audio frames and performing non-occluded spatial filtering on one or more audio frames.
[0193] In some examples, the output signal generated based on performing occluded spatial filtering can be combined with the output signal generated based on performing non-occluded spatial filtering. Fig.11 A smart merging engine similar to or identical to the illustrated smart merging engine 1160 may be used to merge an occluded spatial filtering output signal (e.g., generated by occluded spatial filtering 1134) and a non-occluded spatial filtering output signal (e.g., generated by non-occluded spatial filtering 1124). In some examples, the merging may be performed based at least in part on a channel map associated with obtaining one or more audio frames. For example, the channel map information may be obtained or otherwise included in the Fig.11 The illustrated fourth condition 1174 information is provided as input to the smart merge engine 1160 .
[0194] In some examples, process 1600 may also include removing friction effects detected in one or more audio frames associated with the spatial audio recording. In some examples, process 1600 may also include removing scratching effects detected in one or more audio frames associated with the spatial audio recording. In some cases, the spatial audio recording may be a stereo audio recording including a first audio channel associated with at least a first microphone and a second audio channel associated with at least a second microphone. In some examples, the first microphone may be included on a first earpiece, and the second microphone may be included on a second earpiece associated with the first earpiece. In some examples, selective switching between at least a first microphone included on a first earpiece and a second microphone included on a second earpiece may be performed based on detecting an occlusion in one of the two microphones or earpieces.
[0195] In some examples, a stereo audio recording may be reconstructed by performing a time shift between a first microphone and a second microphone based on detecting an occlusion. In some examples, a stereo audio recording may be reconstructed by performing an acoustic shift between a first mono audio signal captured by a first microphone and a second mono audio signal captured by a second microphone based on detecting an occlusion. In some cases, the acoustic shift may be determined based on geometric information associated with the first microphone and the second microphone. In some examples, the first microphone and the second microphone may be included on a camera associated with the spatial audio recording, and selective switching may be performed between the first microphone and the second microphone based on detecting an occlusion.
[0196] In some examples, the processes described herein (e.g., process 1600 and / or other processes described herein) can be performed by a computing device or apparatus. In some examples, process 1600 can be performed by a wireless communication device. In one example, process 1600 can be performed by a multi-microphone device (e.g., such as Figure 1 102) and / or other audio playback devices. In another example, process 1600 may be performed by a Fig.17 The computing device of the computing system architecture 1700 shown in FIG. Fig.17 A wireless communication device of the illustrated computing architecture may include components of the multi-microphone device 102 and / or other audio playback devices and may implement the operations of process 1600 .
[0197] In some cases, a computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, a computing device may include a display, one or more network interfaces configured to communicate and / or receive data, any combination thereof, and / or other components. The one or more network interfaces may be configured to communicate and / or receive wired and / or wireless data, including data in accordance with 3G, 4G, 5G, and / or other cellular standards, data in accordance with WiFi (802.11x), data in accordance with Bluetooth TM Standard data, data according to the Internet Protocol (IP) standard and / or other types of data.
[0198] The components of the computing device may be implemented in circuits. For example, the components may include and / or be implemented using electronic circuits or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits), and / or may include and / or be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein.
[0199] Process 1600 is illustrated as a logic flow diagram, the operations of which represent a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, each operation represents a computer-executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the described operation. Generally speaking, computer-executable instructions include routines, programs, objects, components, and data structures, etc. that perform specific functions or implement specific data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations may be combined in any order and / or in parallel to implement the process.
[0200] Additionally, process 1600 and / or other processes described herein may be performed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed together on one or more processors, implemented by hardware, or implemented by a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program including a plurality of instructions that can be executed by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.
[0201] Fig.17 is a diagram illustrating an example of a system for implementing certain aspects of the present technology. Specifically, Fig.17 An example of a computing system 1700 is illustrated, which may be any computing device, for example, constituting an internal computing system, a remote computing system, a camera, or any components thereof, wherein the components of the system communicate with each other using a connection 1705. The connection 1705 may be a physical connection using a bus, or a direct connection into the processor 1710, such as in a chipset architecture. The connection 1705 may also be a virtual connection, a networked connection, or a logical connection.
[0202] In some aspects, computing system 1700 is a distributed system in which the functionality described in the present disclosure may be distributed within a data center, multiple data centers, a peer-to-peer network, etc. In some aspects, one or more of the system components represent a number of such components that each perform some or all of the functions for which the component is described. In some aspects, a component may be a physical or virtual device.
[0203] The example system 1700 includes at least one processing unit (CPU or processor) 1710 and connections 1705 that communicatively couple various system components including system memory 1715, such as read only memory (ROM) 1720 and random access memory (RAM) 1725, to the processor 1710. The computing system 1700 may include a cache 1714 of high-speed memory directly connected to, in close proximity to, or integrated as part of the processor 1710.
[0204] Processor 1710 may include any general purpose processor and hardware or software services, such as services 1732, 1734, and 1736 stored in storage device 1730, that are configured to control processor 1710 as well as a dedicated processor where software instructions are incorporated into the actual processor design. Processor 1710 may essentially be a completely independent computing system containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.
[0205] To enable user interaction, the computing system 1700 includes an input device 1745 that can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice, etc. The computing system 1700 may also include an output device 1735, which may be one or more of a plurality of output mechanisms. In some cases, a multimodal system may enable a user to provide multiple types of input / output to communicate with the computing system 1700.
[0206] The computing system 1700 may include a communication interface 1740, which may generally govern and manage user input and system output. The communication interface may perform or facilitate receiving and / or sending wired or wireless communications using wired and / or wireless transceivers, including using an audio jack / plug, a microphone jack / plug, a Universal Serial Bus (USB) port / plug, an Apple TM Lightning TM Ports / plugs, Ethernet ports / plugs, Fiber optic ports / plugs, Dedicated wired ports / plugs, 3G, 4G, 5G and / or other cellular data network wireless signal delivery, Bluetooth TM Wireless signal transmission, Bluetooth TM Low energy (BLE) wireless signal transmission, IBEACON TMThe communication interface 1740 may also include one or more global navigation satellite system (GNSS) receivers or transceivers for determining the location of the computing system 1700 based on receiving one or more signals from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the United States' Global Positioning System (GPS), Russia's Global Navigation Satellite System (GLONASS), China's BeiDou Navigation Satellite System (BDS), and Europe's Galileo GNSS. There is no restriction to operating on any particular hardware arrangement, and thus the base features herein may be easily substituted for improved hardware or firmware arrangements as they are developed.
[0207] The storage device 1730 may be a non-volatile and / or non-transitory and / or computer-readable memory device and may be a hard disk or other type of computer-readable medium that can store data accessible by a computer, such as a magnetic cassette, a flash memory card, a solid-state memory device, a digital versatile disk, a cassette, a floppy disk, a floppy disk, a hard disk, a magnetic tape, a magnetic stripe / magnetic stripe, any other magnetic storage medium, flash memory, a memristor memory, any other solid-state memory, a compact disk-read only memory (CD-ROM) optical disk, a rewritable compact disk (CD) optical disk, a digital video disk (DVD) optical disk, a Blu-ray disc (BDD) optical disk, a holographic optical disk, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, a Memory Card, or a memory card. card, a smart card chip, an EMV chip, a subscriber identity module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, a random access memory (RAM), a static RAM (SRAM), a dynamic RAM (DRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash EPROM (FLASH EPROM), a cache memory (e.g., a layer 1 (L1) cache, a layer 2 (L2) cache, a layer 3 (L3) cache, a layer 4 (L4) cache, a layer 5 (L5) cache, other (L#) cache), a resistive random access memory (RRAM / ReRAM), a phase change memory (PCM), a spin transfer torque RAM (STT-RAM), another memory chip or box, and / or a combination thereof.
[0208] Storage device 1730 may include software services, servers, services, etc., which, when the code defining such software is executed by processor 1710, causes the system to perform functions. In some aspects, hardware services that perform specific functions may include software components stored in a computer-readable medium connected to necessary hardware components (such as processor 1710, connection 1705, output device 1735, etc.) to perform functions. The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing or carrying instructions and / or data. Computer-readable media may include non-transient media in which data may be stored and does not include carrier waves and / or transient electronic signals propagated wirelessly or through wired connections. Examples of non-transient media may include, but are not limited to, disks or tapes, optical storage media (such as compact discs (CDs) or digital versatile discs (DVDs)), flash memory, memory, or memory devices. Computer readable media may store thereon code and / or machine executable instructions, which may represent a process, function, subprogram, program, routine, subroutine, module, software package, category, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by transmitting and / or receiving information, data, independent variables, parameters, or memory contents. Information, independent variables, parameters, data, etc. may be transmitted, forwarded, or sent via any suitable means, including memory sharing, message passing, token passing, network sending, etc.
[0209] Specific details are provided in the above description to provide a thorough understanding of the various aspects and examples provided herein, but those skilled in the art will recognize that the application is not limited thereto. Thus, although the exemplary aspects of the present application have been described in detail herein, it is to be understood that each inventive concept can be implemented and adopted in various other ways, and the appended claims are not intended to be interpreted as including these variations, unless limited by the prior art. The various features and aspects of the above-mentioned applications can be used individually or in combination. In addition, without departing from the broader scope of the specification, the various aspects can be utilized in any number of environments and applications beyond those described herein. Therefore, the specification and the accompanying drawings should be considered as illustrative rather than restrictive. For the purpose of illustration, each method is described in a specific order. It should be understood that, in alternative aspects, each method can be performed in a different order than described.
[0210] For the sake of explanation, in some cases, the present technology can be presented as including separate functional blocks, which include devices, device components, steps or routines in the method embodied in software or a combination of hardware and software. Additional components other than those components shown in the drawings and / or described herein can be used. For example, circuits, systems, networks, processes and other components can be shown as components in block diagram form to avoid confusing these aspects in unnecessary details. In other cases, well-known circuits, processes, algorithms, structures and techniques can be shown without unnecessary details to avoid confusing various aspects.
[0211] In addition, it will be appreciated by those skilled in the art that the various exemplary logic blocks, modules, circuits and algorithmic steps described in conjunction with the various aspects disclosed herein can be implemented as electronic hardware, computer software or a combination of the two. In order to clearly illustrate this interchangeability of hardware and software, various exemplary components, frames, modules, circuits and steps have been generally described in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints proposed to the entire system. Those skilled in the art can implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.
[0212] Various aspects may be described above as a process or method, which is depicted as a flow chart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flow chart may describe an operation as a sequential process, many operations in the operation may be performed in parallel or concurrently. In addition, the order of the operations may be rearranged. The process is terminated when the operation of the process is completed, but the process may have additional steps not included in the accompanying drawings. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, the termination of the process may correspond to the function returning to a calling function or a main function.
[0213] The processes and methods according to the above examples can be implemented using stored computer executable instructions or otherwise available computer executable instructions from a computer readable medium. Such instructions may include, for example, instructions and data that configure a general-purpose computer, a special-purpose computer, or a processing device to perform a certain function or group of functions. Portions of the computer resources used may be accessed via a network. Computer executable instructions may be, for example, binary, intermediate format instructions such as assembly language, firmware, source code. Examples of computer readable media that can be used to store instructions, information used, and / or information created during the methods according to the described examples include disks or optical disks, flash memory, USB devices with non-volatile memory, networked storage devices, etc.
[0214] In some aspects, computer readable storage devices, media, and memories may include wired or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer readable storage media expressly excludes media such as energy, carrier signals, electromagnetic waves, and signals themselves.
[0215] Those skilled in the art will appreciate that information and signals may be represented using any of a variety of different techniques and methods. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be mentioned in the above description may be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or optical particles, or any combination thereof in some cases, depending in part on the specific application, in part on the desired design, in part on the corresponding technology, etc.
[0216] The various illustrative logic blocks, modules, and circuits described in conjunction with the various aspects disclosed herein may be implemented or executed using hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof, and may be implemented in any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing the necessary tasks may be stored in a computer-readable or machine-readable medium. The processor may perform the necessary tasks. Examples of form factors include: laptops, smart phones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functions described herein may also be embodied in peripheral devices or add-in cards. By way of further example, such functions may also be implemented on circuit boards in different chips or different processes executed on a single device.
[0217] The instructions, the media for conveying such instructions, the computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functionality described in this disclosure.
[0218] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as general-purpose computers, wireless communication device handsets, or integrated circuit devices with multiple uses, including applications in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be implemented at least in part by a computer-readable data storage medium including a program code, which includes instructions for executing one or more of the above methods, algorithms, and / or operations when executed. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include a memory or data storage medium, such as a random access memory (RAM) (such as a synchronous dynamic random access memory (SDRAM)), a read-only memory (ROM), a non-volatile random access memory (NVRAM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, and a magnetic or optical data storage medium, etc. Additionally or alternatively, the techniques may be implemented at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.
[0219] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such processors may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in an alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Therefore, the term "processor" as used herein may refer to any of the aforementioned structures, any combination of the aforementioned structures, or any other structure or device suitable for implementing the techniques described herein.
[0220] It should be understood by those of ordinary skill in the art that the less than ("<") symbol and greater than (">") symbol or terminology used herein may be replaced by the less than or equal to ("≤") symbol and greater than or equal to ("≥") symbol, respectively, without departing from the scope of the present specification.
[0221] Where a component is described as being “configured to” perform certain operations, such configuration may be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., a microprocessor or other suitable electronic circuits) to perform the operations, or any combination thereof.
[0222] The phrases “coupled to” or “communicatively coupled to” refer to any component being physically connected directly or indirectly to another component, and / or any component being in communication directly or indirectly with another component (e.g., connected to the other component via a wired or wireless connection and / or other suitable communication interface).
[0223] Claim language or other language stating "at least one of" a set and / or "one or more of" a set indicates that one member of the set or multiple members of the set (in any combination) satisfies the claim. For example, claim language stating "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, claim language stating "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, or any repetition is information or data (e.g., A and A, B and B, C and C, A and A and B, etc.), or any other ordering, repetition, or combination of A, B, and C. The language "at least one of" a set and / or "one or more of" a set does not limit the set to the items listed in the set. For example, claim language stating "at least one of A and B" or "at least one of A or B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.
[0224] Exemplary embodiments of the present disclosure include:
[0225] Aspect 1. A device for generating a spatial audio recording, the device comprising: a memory; and one or more processors, the one or more processors coupled to the memory, the one or more processors configured to: detect occlusion of one or more microphones associated with at least one audio frame of one or more audio frames associated with the spatial audio recording; and selectively switch between an occluded spatial filter for the one or more audio frames and a non-occluded spatial filter for the one or more audio frames based on the detection of the occlusion.
[0226] Aspect 2. The apparatus according to aspect 1, wherein the one or more processors are configured to detect the occlusion of the one or more microphones based on a channel map indicating a number of microphones that are blocked or rubbed.
[0227] Aspect 3. An apparatus according to any one of Aspects 1 to 2, wherein the one or more processors are further configured to: send the one or more audio frames to a delayed audio database at a first time, wherein the one or more audio frames are written to the delayed audio database after a predetermined delay has passed.
[0228] Aspect 4. The apparatus according to aspect 3, wherein the one or more processors are further configured to: obtain the one or more audio frames from the delayed audio database at a second time, the second time being later than the first time; and select at least one of the occluded spatial filter or the non-occluded spatial filter based on obtaining the one or more audio frames at the second time.
[0229] Aspect 5. An apparatus according to Aspect 4, wherein in order to select at least one of the occluded spatial filter or the non-occluded spatial filter, the one or more processors are configured to: perform cross-fading between the occluded spatial filter and the non-occluded spatial filter, wherein the occluded spatial filter operates on the one or more audio frames obtained from the delayed audio database at the second time, and wherein the non-occluded spatial filter operates on the one or more audio frames obtained from the delayed audio database at the second time.
[0230] Aspect 6. An apparatus according to any one of Aspects 1 to 5, wherein the one or more processors are configured to: determine an estimated direction of arrival (DOA) associated with an occluded microphone used to obtain the at least one audio frame associated with the occlusion; determine a spatial filter to generate a reconstructed signal for the occluded microphone; and use the determined spatial filter as the occluded spatial filter.
[0231] Aspect 7. An apparatus according to Aspect 6, wherein the one or more processors are configured to use the determined spatial filter as the occluded spatial filter for one or more non-occluded audio frames included in the one or more audio frames, wherein no occlusion is detected for the one or more non-occluded audio frames.
[0232] Aspect 8. An apparatus according to any one of aspects 1 to 7, wherein the occluded spatial filter comprises one or more filter banks.
[0233] Aspect 9. An apparatus according to Aspect 8, wherein the one or more filter groups are configured to receive as input one or more non-occluded audio frames included in the one or more audio frames, and wherein no occlusion is detected for the one or more non-occluded audio frames.
[0234] Aspect 10. An apparatus according to any one of Aspects 8 to 9, wherein the one or more processors are configured to: select at least one filter group from the one or more filter groups included in the occluded spatial filter, wherein the selected at least one filter group is associated with one or more non-occluded microphones.
[0235] Aspect 11. An apparatus according to any one of Aspects to 10, wherein the non-occlusion spatial filter comprises one or more filter banks.
[0236] Aspect 12. An apparatus according to any one of Aspects 1 to 11, wherein to detect occlusion of at least one audio frame, the one or more processors are configured to detect occlusion of a microphone associated with capturing the at least one audio frame.
[0237] Aspect 13. The apparatus according to aspect 12, wherein the occlusion of the microphone comprises friction occlusion or blocking occlusion.
[0238] Aspect 14. An apparatus according to any one of aspects 1 to 13, wherein the one or more processors are configured to obtain each audio frame included in the one or more audio frames using a different microphone.
[0239] Clause 15. The apparatus of clause 14, wherein each different microphone is included on the same device.
[0240] Aspect 16. An apparatus according to any one of Aspects 1 to 15, wherein the one or more processors are further configured to merge, based on a channel map associated with obtaining the one or more audio frames, an output signal generated based on performing the occluded spatial filtering and an output signal generated based on performing the non-occluded spatial filtering.
[0241] Aspect 17. An apparatus according to any one of aspects 1 to 16, wherein the one or more processors are further configured to remove friction effects detected in the one or more audio frames associated with the spatial audio recording.
[0242] Aspect 18. The apparatus of any one of aspects 1 to 17, wherein the one or more processors are further configured to remove scratching effects detected in the one or more audio frames associated with the spatial audio recording.
[0243] Aspect 19. The apparatus of any one of Aspects 1 to 18, wherein the spatial audio recording is a stereo audio recording comprising a first audio channel associated with at least a first microphone and a second audio channel associated with at least a second microphone.
[0244] Aspect 20. The apparatus of aspect 19, wherein the first microphone is included on a first earpiece and the second microphone is included on a second earpiece associated with the first earpiece.
[0245] Aspect 21. The apparatus according to aspect 20, wherein the one or more processors are further configured to selectively switch between at least the first microphone included on the first earpiece or the second microphone included on the second earpiece based on detecting an occlusion.
[0246] Aspect 22. An apparatus according to any one of Aspects 19 to 21, wherein the one or more processors are further configured to reconstruct the stereo audio recording by performing a time shift between the first microphone and the second microphone based on detecting the occlusion.
[0247] Aspect 23. An apparatus according to any one of Aspects 19 to 22, wherein the one or more processors are further configured to reconstruct the stereo audio recording by performing a coloration shift between a first mono audio signal captured by the first microphone and a second mono audio signal captured by the second microphone based on detecting the occlusion.
[0248] Aspect 24. The apparatus of aspect 23, wherein the one or more processors are configured to determine the coloration shift based on geometric information associated with the first microphone and the second microphone.
[0249] Aspect 25. An apparatus according to any one of Aspects 19 to 24, wherein: the first microphone and the second microphone are included on a camera associated with the spatial audio recording; and the selective switching is performed between the first microphone and the second microphone based on detecting the occlusion.
[0250] Aspect 26. A method for performing spatial audio recording, comprising: detecting occlusion of at least one audio frame among one or more audio frames associated with the spatial audio recording; and during the spatial audio recording, selectively switching between performing occluded spatial filtering of the one or more audio frames and non-occluded spatial filtering of the one or more audio frames based on detecting the occlusion of the at least one audio frame.
[0251] Clause 27. The method according to clause 26, wherein detecting the occlusion of the at least one audio frame is based on a channel map indicating a number of microphones that are blocked or rubbed.
[0252] Aspect 28. The method according to any one of Aspects 26 to 27 further includes: sending the one or more audio frames to a delayed audio database, wherein the one or more audio frames are written to the delayed audio database after a predetermined delay has passed.
[0253] Aspect 29. The method according to Aspect 28 further includes: obtaining the one or more audio frames and the corresponding one or more delayed audio frames from the delayed audio database after the predetermined delay has passed, wherein the selective switching is performed based on obtaining the one or more audio frames and the corresponding one or more delayed audio frames from the delayed audio database.
[0254] Aspect 30. The method of aspect 29, wherein the selectively switching comprises cross-fading between performing occluded spatial filtering on the one or more delayed audio frames and performing non-occluded spatial filtering on the one or more delayed audio frames.
[0255] Aspect 31. A method according to any one of Aspects 26 to 30, wherein performing occluded spatial filtering on the one or more audio frames includes: determining an estimated direction of arrival (DOA) associated with an occluded microphone used to obtain the at least one audio frame associated with the occlusion; determining a spatial filter to generate a reconstructed signal for the occluded microphone; and using the spatial filter to perform occluded spatial filtering.
[0256] Aspect 32. A method according to aspect 31, wherein using the spatial filter to perform occluded spatial filtering is based on one or more non-occluded audio frames included in the one or more audio frames, wherein no occlusion is detected for the one or more non-occluded audio frames.
[0257] Aspect 33. The method according to any one of aspects 26 to 32, wherein detecting occlusion of at least one audio frame comprises detecting occlusion of a microphone associated with capturing the at least one audio frame.
[0258] Aspect 34. The method according to aspect 33, wherein the shielding of the microphone comprises friction shielding or blocking shielding.
[0259] Aspect 35. The method according to any one of aspects 26 to 34, wherein each audio frame included in the one or more audio frames is obtained using a different microphone.
[0260] Clause 36. The method of clause 35, wherein each different microphone is included on the same device.
[0261] Aspect 37. The method according to any one of Aspects 26 to 36 further includes merging an output signal generated based on performing the occluded spatial filtering and an output signal generated based on performing the non-occluded spatial filtering based on a channel map associated with obtaining the one or more audio frames.
[0262] Aspect 38. The method according to any one of aspects 26 to 37, further comprising removing friction effects detected in the one or more audio frames associated with the spatial audio recording.
[0263] Aspect 39. The method according to any one of aspects 26 to 37, further comprising removing scratching effects detected in the one or more audio frames associated with the spatial audio recording.
[0264] Aspect 40. The method according to any one of aspects 26 to 39, wherein the spatial audio recording is a stereo audio recording comprising a first audio channel associated with at least a first microphone and a second audio channel associated with at least a second microphone.
[0265] Aspect 41. The method of aspect 40, wherein the first microphone is included on a first earpiece and the second microphone is included on a second earpiece associated with the first earpiece.
[0266] Clause 42. The method according to clause 41, further comprising selectively switching between at least the first microphone included on the first earpiece and the second microphone included on the second earpiece based on detecting an occlusion.
[0267] Aspect 43. The method according to any one of aspects 40 to 42, further comprising reconstructing the stereo audio recording by performing a time shift between the first microphone and the second microphone based on detecting the occlusion.
[0268] Aspect 44. The method according to any one of Aspects 40 to 43, further comprising reconstructing the stereo audio recording by performing a coloration shift between a first mono audio signal captured by the first microphone and a second mono audio signal captured by the second microphone based on detecting the occlusion.
[0269] Clause 45. The method according to clause 44, wherein the coloration shift is determined based on geometric information associated with the first microphone and the second microphone.
[0270] Aspect 46. A method according to any one of Aspects 40 to 45, wherein: the first microphone and the second microphone are included on a camera associated with the spatial audio recording; and the selective switching is performed between the first microphone and the second microphone based on detecting the occlusion.
[0271] Aspect 47. A non-transitory computer-readable medium having instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to: detect occlusion of at least one audio frame of one or more audio frames associated with a spatial audio recording; and during the spatial audio recording, selectively switch between performing occluded spatial filtering of the one or more audio frames and unoccluded spatial filtering of the one or more audio frames based on detecting the occlusion of the at least one audio frame.
[0272] Aspect 48. The non-transitory computer-readable medium of aspect 47, wherein the instructions cause the one or more processors to detect the occlusion of the at least one audio frame based on a channel map indicating a number of microphones that are blocked or rubbed.
[0273] Aspect 49. A non-transitory computer-readable medium according to any one of Aspects 47 to 48, wherein the instructions further cause the one or more processors to: send the one or more audio frames to a delayed audio database, wherein the one or more audio frames are written to the delayed audio database after a predetermined delay has occurred.
[0274] Aspect 50. A non-transitory computer-readable medium according to Aspect 49, wherein the instructions further cause the one or more processors to: obtain the one or more audio frames and the corresponding one or more delayed audio frames from the delayed audio database after the predetermined delay has passed, wherein the selective switching is performed based on obtaining the one or more audio frames and the corresponding one or more delayed audio frames from the delayed audio database.
[0275] Aspect 51. A non-transitory computer-readable medium according to Aspect 50, wherein, in order to selectively switch, the instructions cause the one or more processors to: cross-fade between performing occluded spatial filtering on the one or more delayed audio frames and performing non-occluded spatial filtering on the one or more delayed audio frames.
[0276] Aspect 52. A non-transitory computer-readable medium according to any one of Aspects 47 to 51, wherein, in order to perform occluded spatial filtering of the one or more audio frames, the instructions cause the one or more processors to: determine an estimated direction of arrival (DOA) associated with an occluded microphone used to obtain the at least one audio frame associated with the occlusion; determine a spatial filter to generate a reconstructed signal for the occluded microphone; and use the spatial filter to perform occluded spatial filtering.
[0277] Aspect 53. A non-transitory computer-readable medium according to Aspect 52, wherein the instructions cause the one or more processors to perform occluded spatial filtering using the spatial filter based on one or more non-occluded audio frames included in the one or more audio frames, wherein no occlusion is detected for the one or more non-occluded audio frames.
[0278] Aspect 54. A non-transitory computer-readable medium according to any one of Aspects 47 to 53, wherein to detect occlusion of at least one audio frame, the instructions cause the one or more processors to detect occlusion of a microphone associated with capturing the at least one audio frame.
[0279] Aspect 55. The non-transitory computer-readable medium of aspect 54, wherein the occlusion of the microphone comprises a friction occlusion or a blocking occlusion.
[0280] Aspect 56. A non-transitory computer-readable medium according to any one of aspects 47 to 55, wherein the instructions cause the one or more processors to obtain each audio frame included in the one or more audio frames using a different microphone.
[0281] Clause 57. The non-transitory computer-readable medium of Clause 56, wherein each different microphone is included on the same device.
[0282] Aspect 58. A non-transitory computer-readable medium according to any one of Aspects 47 to 57, wherein the instructions further cause the one or more processors to merge an output signal generated based on performing the occluded spatial filtering and an output signal generated based on performing the non-occluded spatial filtering based on a channel map associated with obtaining the one or more audio frames.
[0283] Aspect 59. A non-transitory computer-readable medium according to any one of aspects 47 to 58, wherein the instructions further cause the one or more processors to remove friction effects detected in the one or more audio frames associated with the spatial audio recording.
[0284] Aspect 60. A non-transitory computer-readable medium according to any one of aspects 47 to 59, wherein the instructions further cause the one or more processors to remove a scratching effect detected in the one or more audio frames associated with the spatial audio recording.
[0285] Aspect 61. The non-transitory computer-readable medium of any one of Aspects 47 to 60, wherein the spatial audio recording is a stereo audio recording comprising a first audio channel associated with at least a first microphone and a second audio channel associated with at least a second microphone.
[0286] Aspect 62. The non-transitory computer-readable medium of aspect 61, wherein the first microphone is included on a first earpiece and the second microphone is included on a second earpiece associated with the first earpiece.
[0287] Aspect 63. A non-transitory computer-readable medium according to Aspect 62, wherein the instructions further cause the one or more processors to selectively switch between at least the first microphone included on the first earpiece or the second microphone included on the second earpiece based on detecting an occlusion.
[0288] Aspect 64. A non-transitory computer-readable medium according to any one of Aspects 61 to 63, wherein the instructions further cause the one or more processors to reconstruct the stereo audio recording by performing a time shift between the first microphone and the second microphone based on detecting the occlusion.
[0289] Aspect 65. A non-transitory computer-readable medium according to any one of Aspects 61 to 64, wherein the instructions further cause the one or more processors to reconstruct the stereo audio recording by performing an audible shift between a first mono audio signal captured by the first microphone and a second mono audio signal captured by the second microphone based on detecting the occlusion.
[0290] Aspect 66. The non-transitory computer-readable medium of aspect 65, wherein the instructions cause the one or more processors to determine the coloration shift based on geometric information associated with the first microphone and the second microphone.
[0291] Aspect 67. A non-transitory computer-readable medium according to any one of Aspects 61 to 66, wherein: the first microphone and the second microphone are included on a camera associated with the spatial audio recording; and the selective switching is performed between the first microphone and the second microphone based on detecting the occlusion.
[0292] Aspect 68. An apparatus comprising means for performing any of the operations according to aspects 1 to 25.
[0293] Aspect 69. An apparatus comprising means for performing any of the operations according to aspects 26 to 46.
[0294] Aspect 70. An apparatus comprising means for performing any of the operations according to aspects 47 to 67.
[0295] Aspect 71. A non-transitory computer-readable storage medium having stored thereon instructions, which, when executed by one or more processors, cause the one or more processors to perform any of the operations described in aspects 1 to 25.
[0296] Aspect 72. A non-transitory computer-readable storage medium having stored thereon instructions, which when executed by one or more processors cause the one or more processors to perform any of the operations described in aspects 26 to 46.
[0297] Aspect 73. A non-transitory computer-readable storage medium having stored thereon instructions, which, when executed by one or more processors, cause the one or more processors to perform any of the operations described in aspects 47 to 67.
[0298] Aspect 74. A device for generating a spatial audio output, the device comprising: a memory; and one or more processors, the one or more processors coupled to the memory, the one or more processors configured to: receive a signal comprising one or more audio frames from a plurality of microphones; generate the spatial audio output from the signal from the one or more microphones using a first spatial filter; detect occlusion of one or more of the plurality of microphones; and selectively switch between using the first spatial filter for the one or more audio frames and using a second spatial filter for the one or more audio frames based on the detection of the occlusion.
Claims
1. An apparatus for generating a spatial audio recording, the apparatus comprising: a memory; and one or more processors coupled to the memory, the one or more processors being configured to: detect occlusion of one or more microphones associated with at least one audio frame among one or more audio frames associated with the spatial audio recording; and select at least one of an occluded spatial filter for the one or more audio frames or a non-occluded spatial filter for the one or more audio frames based on the detection of the occlusion.
2. The apparatus according to claim 1, wherein the one or more processors are configured to detect the occlusion of the one or more microphones based on a channel map indicating the number of microphones blocked or rubbed.
3. The apparatus according to claim 1, wherein the one or more processors are further configured to: send the one or more audio frames to a delayed audio database at a first time, wherein the one or more audio frames are written to the delayed audio database after a predetermined delay.
4. The apparatus according to claim 3, wherein the one or more processors are further configured to: obtain the one or more audio frames from the delayed audio database at a second time, the second time being later than the first time; and select at least one of the occluded spatial filter or the non-occluded spatial filter based on obtaining the one or more audio frames at the second time.
5. The apparatus according to claim 4, wherein, in order to select at least one of the occluded spatial filter or the non-occluded spatial filter, the one or more processors are configured to: perform a cross-fade between the occluded spatial filter and the non-occluded spatial filter, wherein the occluded spatial filter operates on the one or more audio frames obtained from the delayed audio database at the second time, and wherein the non-occluded spatial filter operates on the one or more audio frames obtained from the delayed audio database at the second time.
6. The apparatus according to claim 1, wherein the one or more processors are configured to: determine an estimated direction of arrival (DOA) associated with an occluded microphone used to obtain the at least one audio frame associated with the occlusion; determine a spatial filter to generate a reconstructed signal for the occluded microphone; and use the determined spatial filter as the occluded spatial filter.
7. The apparatus according to claim 6, wherein the one or more processors are configured to use the determined spatial filter as the occluded spatial filter for one or more non-occluded audio frames included in the one or more audio frames, wherein no occlusion is detected for the one or more non-occluded audio frames.
8. The apparatus according to claim 1, wherein the occluded spatial filter comprises one or more filter banks.
9. The apparatus according to claim 8, wherein the one or more filter banks are configured to receive as input one or more non-occluded audio frames included in the one or more audio frames, and wherein for the one or more non-occluded audio frames, no occlusion is detected.
10. The apparatus according to claim 8, wherein the one or more processors are configured to: select at least one filter bank from the one or more filter banks included in the occluded spatial filter, wherein the at least one selected filter bank is associated with one or more non-occluded microphones.
11. The apparatus according to claim 1, wherein the non-occluded spatial filter includes one or more filter banks.
12. The apparatus according to claim 1, wherein in order to detect occlusion of the one or more microphones associated with the at least one audio frame, the one or more processors are configured to detect occlusion of the microphones used to capture the at least one audio frame.
13. The apparatus according to claim 12, wherein the occlusion of the microphone includes frictional occlusion or blocking occlusion.
14. The apparatus according to claim 1, wherein the one or more processors are configured to obtain each audio frame included in the one or more audio frames using different microphones.
15. The apparatus according to claim 14, wherein each different microphone is included on the same device.
16. The apparatus according to claim 1, wherein the one or more processors are further configured to merge an output signal generated based on the occluded spatial filter for the one or more audio frames and an output signal generated based on the non-occluded spatial filter for the one or more audio frames based on a channel map associated with the one or more microphones.
17. The apparatus according to claim 1, wherein the one or more processors are further configured to remove frictional effects detected in the one or more audio frames associated with the spatial audio recording.
18. The apparatus according to claim 1, wherein the one or more processors are further configured to remove scratching effects detected in the one or more audio frames associated with the spatial audio recording.
19. The apparatus according to claim 1, wherein the spatial audio recording is a stereo audio recording, which includes a first audio channel associated with at least a first microphone and a second audio channel associated with at least a second microphone.
20. The apparatus according to claim 19, wherein the first microphone is included on a first earpiece and the second microphone is included on a second earpiece associated with the first earpiece.
21. The apparatus according to claim 20, wherein the one or more processors are further configured to select at least one of the first microphone included on the first earpiece or the second microphone included on the second earpiece based on the detection of the occlusion.
22. The apparatus according to claim 19, wherein the one or more processors are further configured to reconstruct the stereophonic audio recording based on a time shift between the first microphone and the second microphone, wherein the time shift is determined based on the detection of the occlusion.
23. The apparatus according to claim 19, wherein the one or more processors are further configured to reconstruct the stereophonic audio recording based on a coloring shift between a first monophonic audio signal captured by the first microphone and a second monophonic audio signal captured by the second microphone, wherein the coloring shift is determined based on the detection of the occlusion.
24. The apparatus according to claim 23, wherein the one or more processors are configured to determine the coloring shift based on geometric information associated with the first microphone and the second microphone.
25. The apparatus according to claim 19, wherein: the first microphone and the second microphone are included on a camera associated with the spatial audio recording; and the one or more processors are configured to select at least one of the first microphone or the second microphone based on the detection of the occlusion.
26. A method of performing a spatial audio recording, comprising: detecting an occlusion of at least one audio frame among one or more audio frames associated with a spatial audio recording; and during the spatial audio recording, selecting between performing occluded spatial filtering of the one or more audio frames or non-occluded spatial filtering of the one or more audio frames based on the detection of the occlusion of the at least one audio frame.
27. The method according to claim 26, wherein detecting the occlusion of the at least one audio frame is based on a channel map indicating the number of microphones that are blocked or rubbed.
28. The method according to claim 26, further comprising: sending the one or more audio frames to a delayed audio database at a first time, wherein the one or more audio frames are written to the delayed audio database after a predetermined delay.
29. The method according to claim 28, further comprising: obtaining the one or more audio frames from the delayed audio database at a second time, the second time being later than the first time, wherein the selection between performing occluded spatial filtering or non-occluded spatial filtering is based on obtaining the one or more audio frames from the delayed audio database at the second time.
30. The method according to claim 29, wherein the selection comprises: performing a cross-fade between performing occluded spatial filtering of the one or more audio frames obtained from the delayed audio database at the second time and performing non-occluded spatial filtering of the one or more audio frames obtained from the delayed audio database at the second time.
31. The method according to claim 26, wherein performing occluded spatial filtering of the one or more audio frames comprises: Determine an estimated direction of arrival (DOA) associated with an occluded microphone that is used to obtain the at least one audio frame associated with the occlusion; Determine a spatial filter to generate a reconstructed signal for the occluded microphone; and Perform occluded spatial filtering using the spatial filter.
32. The method according to claim 31, wherein performing occluded spatial filtering using the spatial filter is based on one or more non-occluded audio frames included in the one or more audio frames, wherein for the one or more non-occluded audio frames, no occlusion is detected.
33. The method according to claim 26, wherein detecting the occlusion includes detecting an occlusion of a microphone associated with capturing the at least one audio frame.
34. The method according to claim 33, wherein the occlusion of the microphone includes a frictional occlusion or a blocking occlusion.
35. The method according to claim 26, wherein each audio frame included in the one or more audio frames is obtained using a different microphone.
36. The method according to claim 35, wherein each different microphone is included on the same device.
37. The method according to claim 26, further comprising merging an output signal generated based on performing the occluded spatial filtering and an output signal generated based on performing the non-occluded spatial filtering based on a channel map associated with obtaining the one or more audio frames.
38. The method according to claim 26, further comprising removing frictional effects detected in the one or more audio frames associated with the spatial audio recording.
39. The method according to claim 26, further comprising removing scratching effects detected in the one or more audio frames associated with the spatial audio recording.
40. The method according to claim 26, wherein the spatial audio recording is a stereo audio recording that includes a first audio channel associated with at least a first microphone and a second audio channel associated with at least a second microphone.
41. The method according to claim 40, wherein the first microphone is included on a first earpiece and the second microphone is included on a second earpiece associated with the first earpiece.
42. The method according to claim 41, further comprising selecting at least one of the first microphone included on the first earpiece or the second microphone included on the second earpiece based on detecting an occlusion.
43. The method according to claim 40, further comprising reconstructing the stereo audio recording by performing a time shift between the first microphone and the second microphone based on detecting the occlusion.
44. The method according to claim 40, further comprising reconstructing the stereo audio recording by performing a timbre shift between a first mono audio signal captured by the first microphone and a second mono audio signal captured by the second microphone based on detecting the occlusion.
45. The method according to claim 44, wherein the timbre shift is determined based on geometric information associated with the first microphone and the second microphone.
46. The method according to claim 40, wherein: the first microphone and the second microphone are included on a camera associated with the spatial audio recording; and selecting at least one of the first microphone or the second microphone is based on detecting the occlusion.
47. A non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to: detect an occlusion of at least one audio frame among one or more audio frames associated with a spatial audio recording; and during the spatial audio recording, based on detecting the occlusion of the at least one audio frame, select between performing occluded spatial filtering on the one or more audio frames or performing non-occluded spatial filtering on the one or more audio frames.
48. The non-transitory computer-readable medium according to claim 47, wherein the instructions cause the one or more processors to detect the occlusion of the at least one audio frame based on a channel map indicating the number of microphones that are blocked or rubbed.
49. The non-transitory computer-readable medium according to claim 47, wherein the instructions further cause the one or more processors to: send the one or more audio frames to a delayed audio database at a first time, wherein the one or more audio frames are written to the delayed audio database after a predetermined delay.
50. The non-transitory computer-readable medium according to claim 49, wherein the instructions further cause the one or more processors to: obtain the one or more audio frames from the delayed audio database at a second time, the second time being later than the first time; and based on obtaining the one or more audio frames from the delayed audio database at the second time, select between performing the occluded spatial filtering or the non-occluded spatial filtering.
51. The non-transitory computer-readable medium according to claim 50, wherein, in order to select between performing the occluded spatial filtering or the non-occluded spatial filtering, the instructions cause the one or more processors to: perform a cross-fade between performing occluded spatial filtering on the one or more audio frames obtained from the delayed audio database at the second time and performing non-occluded spatial filtering on the one or more audio frames obtained from the delayed audio database at the second time.
52. The non-transitory computer-readable medium according to claim 47, wherein, in order to perform occluded spatial filtering on the one or more audio frames, the instructions cause the one or more processors to: determine an estimated direction of arrival (DOA) associated with an occluded microphone used to obtain the at least one audio frame associated with the occlusion; determine a spatial filter to generate a reconstructed signal for the occluded microphone; and use the spatial filter to perform occluded spatial filtering.
53. The non-transitory computer-readable medium according to claim 52, wherein the instructions cause the one or more processors to perform occluded spatial filtering using the spatial filter based on one or more non-occluded audio frames included in the one or more audio frames, wherein no occlusion is detected for the one or more non-occluded audio frames.
54. The non-transitory computer-readable medium according to claim 47, wherein to detect occlusion of at least one audio frame, the instructions cause the one or more processors to detect occlusion of a microphone associated with capturing the at least one audio frame.
55. The non-transitory computer-readable medium according to claim 54, wherein the occlusion of the microphone includes frictional occlusion or blocking occlusion.
56. The non-transitory computer-readable medium according to claim 47, wherein the instructions cause the one or more processors to obtain each audio frame included in the one or more audio frames using different microphones.
57. The non-transitory computer-readable medium according to claim 56, wherein each different microphone is included on the same device.
58. The non-transitory computer-readable medium according to claim 47, wherein the instructions further cause the one or more processors to merge an output signal generated based on performing the occluded spatial filtering and an output signal generated based on performing the non-occluded spatial filtering based on a channel map associated with obtaining the one or more audio frames.
59. The non-transitory computer-readable medium according to claim 47, wherein the instructions further cause the one or more processors to remove frictional effects detected in the one or more audio frames associated with the spatial audio recording.
60. The non-transitory computer-readable medium according to claim 47, wherein the instructions further cause the one or more processors to remove scratching effects detected in the one or more audio frames associated with the spatial audio recording.
61. The non-transitory computer-readable medium according to claim 47, wherein the spatial audio recording is a stereo audio recording that includes a first audio channel associated with at least a first microphone and a second audio channel associated with at least a second microphone.
62. The non-transitory computer-readable medium according to claim 61, wherein the first microphone is included on a first earpiece and the second microphone is included on a second earpiece associated with the first earpiece.
63. The non-transitory computer-readable medium according to claim 62, wherein the instructions further cause the one or more processors to select at least one of the first microphone included on the first earpiece or the second microphone included on the second earpiece based on detecting occlusion.
64. The non-transitory computer-readable medium according to claim 61, wherein the instructions further cause the one or more processors to reconstruct the stereo audio recording by performing a time shift between the first microphone and the second microphone based on detecting the occlusion.
65. The non-transitory computer-readable medium according to claim 61, wherein the instructions further cause the one or more processors to reconstruct the stereophonic audio recording by performing a timbre shift between a first mono audio signal captured by the first microphone and a second mono audio signal captured by the second microphone based on detecting the occlusion.
66. The non-transitory computer-readable medium according to claim 65, wherein the instructions cause the one or more processors to determine the timbre shift based on geometric information associated with the first microphone and the second microphone.
67. The non-transitory computer-readable medium according to claim 61, wherein: the first microphone and the second microphone are included on a camera associated with the spatial audio recording; and the instructions cause the one or more processors to select at least one of the first microphone or the second microphone based on detecting the occlusion.