Stereo recording with closely spaced microphones

By generating directional virtual microphones from closely spaced omnidirectional microphones and applying inter-channel level differences, the method addresses the challenge of obtaining high-quality stereo recordings, enhancing the spatial characteristics and signal quality.

WO2025137504A1PCT designated stage expired Publication Date: 2025-06-26DOLBY LABORATORIES LICENSING CORP +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/061370
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-06
Filing Date
2024-12-20
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing technologies face challenges in obtaining high-quality stereo recordings using closely spaced omnidirectional microphones, as they often result in a narrow stereo image due to small time delays and intensity differences between the microphone signals.

Method used

The method involves generating directional virtual microphones by combining signals from closely spaced omnidirectional microphones, applying inter-channel level differences to correct for frequency imbalances, and combining spectral portions of these signals across different frequency ranges to construct high-quality stereo channels.

Benefits of technology

This approach effectively enhances the spatial characteristics of audio capture, improving the quality of stereo recordings by widening the stereo image and maintaining a good signal-to-noise ratio across various frequency bands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024061370_26062025_PF_FP_ABST
    Figure US2024061370_26062025_PF_FP_ABST
Patent Text Reader

Abstract

A method and apparatus for obtaining a stereo recording with closely spaced physical microphones, e.g., of a consumer-grade electronic device. Various examples construct signals for two stereo channels by combining different spectral portions of variously generated signals. In one example, a stereo channel of the stereo recording is constructed using a spectral portion of a directional virtual- microphone signal (332) derived from the audio signals captured by the physical microphones (302), a modified spectral portion of one of the audio signals (352), and an unmodified spectral portion of that audio signal (328). The directional virtual-microphone signal is generated using a pressure-gradient method (318). The modification of the audio signal (350) includes imposing an interchannel level difference (ILD) that matches the ILD of the directional virtual- microphone signals (332) in the adjacent frequency range (IMF). In various examples, the corresponding algorithm can be run at the electronic device that houses the closely spaced microphones or at a remote server.
Need to check novelty before this filing date? Find Prior Art

Description

STEREO RECORDING WITH CLOSELY SPACED MICROPHONES 1. Cross-Reference to Related Applications

[0001] This application claims the benefit of priority from Spanish Patent Application No. P202331063, filed on 21 December 2023, European Patent Application No.24174170.1, filed on 3 May 2024, and US Provisional Patent Application No.63 / 642,985, filed on 6 May 2024, each of which is incorporated by reference herein in its entirety. 2. Field of the Disclosure

[0002] Various example embodiments relate to audio equipment and, more specifically but not exclusively, to equipment for capturing and reproducing stereophonic sound. 3. Background

[0003] Stereophonic sound, sometimes referred to as “stereo,” is a method of sound reproduction that recreates a spatial audible perspective. A stereo effect is typically achieved with the use of two independent audio channels through a configuration of two loudspeakers (or stereo headphones) to create the impression of sound heard from various directions, as in natural hearing.

[0004] During typical two-channel stereo recording, two microphones are placed in appropriately chosen locations relative to the sound source, with both microphones recording simultaneously. The two recorded channels are similar but have distinct time-of-arrival and / or sound-pressure-level information. During playback, the listener’s brain uses those subtle differences in timing and sound level to sense the positions of the recorded objects, thereby detecting the stereo effect. Three most widely used stereo recording setups that capture time and / or level differences to create a stereo effect are the AB, XY, and ORTF setups. BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS

[0005] Example embodiments provide a method and apparatus for obtaining a stereo recording with closely spaced omnidirectional microphones, e.g., found in a consumer-grade electronic device. Various embodiments construct signals for two stereo channels of the stereo recording by combining different spectral portions of variously generated signals. In one example, a stereo channel of the stereo recording is constructed using a spectral portion of a directional virtual-microphone signal derived from the audio signals captured by theomnidirectional microphones, a modified spectral portion of one of the audio signals, and an unmodified spectral portion of that audio signal. The directional virtual-microphone signal is generated using a pressure-gradient method. The modification of the audio signal includes imposing an inter-channel level difference (ILD) that matches the ILD of the directional virtual- microphone signals in the adjacent frequency range. In different examples, the corresponding algorithm can be run locally at the electronic device that houses the closely spaced omnidirectional microphones or remotely, e.g., at a server that is network-connected to the electronic device.

[0006] According to an example embodiment, provided is a method of generating a stereo signal, the method comprising: obtaining a first directional virtual-microphone mono signal and a second directional virtual-microphone mono signal based on a first audio signal captured with a first microphone and further based on a second audio signal captured with a second microphone; determining a level difference between the first and second directional virtual- microphone signals in a first frequency range; applying a gain to the second audio signal in a second frequency range to generate a third mono audio signal, the gain being selected to cause a level difference between the third audio signal and the first audio signal in the second frequency range to match the determined level difference between the first and second directional virtual- microphone signals in the first frequency range; and generating the stereo signal using the first directional virtual-microphone signal, the second directional virtual-microphone signal, and the third audio signal.

[0007] According to another example embodiment, provided is a non-transitory computer- readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising the above method.

[0008] According to yet another example embodiment, provided is an apparatus for generating a stereo signal, the apparatus comprising: at least one processor; and at least one memory including program code; and wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: obtain a first directional virtual-microphone signal and a second directional virtual-microphone signal based on a first audio signal captured with a first microphone and further based on a second audio signal captured with a second microphone; determine a level difference between the first and second directional virtual-microphone signals in a first frequency range; apply a gain to the second audio signal in a second frequency range to generate a third audio signal, the gain being selected to cause a level difference between the third audio signal and the first audio signal in thesecond frequency range to match the determined level difference between the first and second directional virtual-microphone signals in the first frequency range; and generate the stereo signal using the first directional virtual-microphone signal, the second directional virtual-microphone signal, and the third audio signal. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Other aspects, features, and benefits of various disclosed embodiments will become more fully apparent, by way of example, from the following detailed description and the accompanying drawings, in which:

[0010] FIG.1 is a schematic diagram illustrating a 3D perspective view of a consumer-grade electronic device with which some embodiments can be practiced.

[0011] FIG.2 is a flowchart illustrating a method of generating a stereo recording according to an embodiment.

[0012] FIG.3 is a block diagram illustrating a machine-implemented algorithm for generating a stereo recording according to an embodiment.

[0013] FIGS.4A-4B graphically illustrate certain performance characteristics of the algorithm of FIG.3 according to an additional embodiment.

[0014] FIG.5 is a block diagram of an example computing device configured to perform at least some operations of the method of FIG.2 and / or the machine-implemented algorithm of FIG.3 according to various embodiments. DETAILED DESCRIPTION

[0015] As more content (e.g., audio and video) is created with smartphones, webcams, and other consumer-grade electronic devices, content creators and device manufacturers are looking for improvements in the quality of audio capture. One aspect contributing to the perceived audio quality is related to spatial characteristics of the audio capture. However, the form factor of many devices is typically such that it is difficult to accommodate directional pressure-gradient microphones for stereo signal capture therein. Instead, omnidirectional microphones are typically used. An example omnidirectional microphone includes a pressure transducer that is relatively simple to integrate into a device body by placing it behind a small opening in the outer surface of the device’s housing.

[0016] Stereo is a common recording and playback format. However, a high-quality stereo recording may be challenging to obtain from internal microphones of a consumer-grade electronic device. For example, two omnidirectional microphones of a smartphone may typically be closely spaced due to a relatively small size of the smartphone. When the pair of signals obtained from such microphones is directly used as a stereo signal, time delays between the channels are small. When the microphone placement is such that a part of the device body is located between the microphones, there may also be some intensity differences between the channels, e.g., at higher frequencies. Since the time delays and intensity differences are relatively small, the stereo image corresponding to the directly recorded signal may be disadvantageously or unacceptably narrow.

[0017] In some examples, a directional microphone can be obtained by combining the signals of two closely spaced omnidirectional microphones. According to one approach, a pressure gradient is approximated by the finite difference between two pressure signals taken at points in space whose distance is smaller than the acoustic wavelength of interest. The pressure gradient estimated in this manner represents the variation of pressure along the line connecting the two corresponding omnidirectional microphones and approximates the signal that a bidirectional microphone would pick up in that location. By adding a suitable time delay before subtracting the pressure signals, one can obtain a virtual directional microphone with a specific first order pickup pattern (e.g., a cardioid polar pattern). Two of such virtual directional microphones, e.g., pointing in opposite directions, can be obtained by suitably selecting the time delay value for the subtraction operation. The resulting virtual-microphone signals can then be used as two channels for the corresponding stereo recording.

[0018] Although useful, the above-indicated approach has certain limitations. For example, one limitation is that an approximation of the pressure gradient by a finite difference is sufficiently accurate only when the acoustic wavelength is significantly larger than the spacing between the microphones. Thus, for a microphone spacing of several centimeters, this approximation holds up to a frequency of approximately 1 kHz. Another limitation is that, as the spacing between the microphones decreases, the signal-to-noise ratio (SNR) of the corresponding pressure-gradient signal decreases as well, e.g., because the similarity between the two pressure signals increases and their difference is eventually overwhelmed by the transducer noise.

[0019] Another challenge is that, under a free-field (plane wave) condition, a pressure- gradient signal has a frequency response that rolls-off by about 6 dB per octave as the frequencydecreases. Therefore, to restore a timbre comparable to the pressure signal, one needs to apply a boost of 6 dB per octave to the pressure gradient. In at least some examples, such a boost may result in a poor SNR at relatively low frequencies. Furthermore, when a sound source is in proximity to the microphone, the free-field condition may not hold at all. In such cases, the curvature of the acoustic field results in different signal amplitudes at the two pressure transducers, which in turn produces a boost of low frequencies. This boost is referred to as the proximity effect, and the amount of boost depends on the frequency of sound, the distance between the sound source and the microphones, and the direction of incidence with respect to the direction along which the pressure gradient is measured. Directional microphones are typically designed to provide a flat low-frequency response on-axis at a specific distance chosen by the manufacturer and are subject to a bass boost at smaller distances or bass roll off at larger distances. Another aspect is that the phase response is also affected by the proximity effect and is frequency-, distance-, and direction-dependent. The latter characteristics pose an additional challenge, e.g., because one needs a controlled, similar phase response of the two microphones across all directions of interest (typically across the front semicircle) for a good-quality stereo capture.

[0020] Various embodiments disclosed herein beneficially address at least some of the above-indicated problems in the state of the art. For example, one embodiment provides a method of generating a high-quality stereo recording based on audio signals captured by two or more closely spaced omnidirectional microphones, such as the typical microphones found in mobile phones. The method includes building a pair of directional virtual microphones, ^^^^and ^^^^, pointing left and right, respectively. In some examples, the directional virtual microphones ^^^^and ^^^^are built using the above-mentioned pressure-gradient method, which works sufficiently well in the intermediate frequency range but may not be optimal for lower and higher frequency bands. The method corrects the uneven accuracy of the pressure-gradient estimate by first obtaining stereo cues in the lower to mid frequency range from the analysis of the ^^^^and ^^^^signals. The obtained stereo cues are then applied to the original microphone signals at lower frequencies, thereby providing modified low frequency signals. Finally, a left and right (L, R) stereo signal is obtained by appropriately combining, in the frequency domain, the modified low frequency signals, the ^^^^and ^^^^signals of the intermediate frequency range, and the original microphone signals of the high frequency band.

[0021] FIG.1 is a schematic diagram illustrating a 3D perspective view of a consumer-grade electronic device 100 with which some embodiments can be practiced. For illustration purposesand without any implied limitations, example embodiments are described below in reference to the shown device 100 being a smartphone. From the provided description, a person of ordinary skill in the pertinent art will readily understand how to practice various additional embodiments using other types of devices without any undue experimentation.

[0022] The device 100 has a length l, a width w, and a thickness d. In one example, the dimensions are l=15 cm, w=7.5 cm, and d=7 mm. In other examples, the device 100 may have other dimensions. The smallest external dimension (in this example, d) maybe referred to as the transverse dimension. The two larger external dimensions (in this example, l and w) maybe referred to as the lateral dimensions. In the example shown, the device 100 has three omnidirectional microphones 110, 120, 130. In other examples, other types of microphones can also be used to implement the microphones 110, 120, 130. The separation between the microphones 110 and 120 is l along the X-coordinate axis and approximately 0.5w along the Z- coordinate axis. The separation between the microphones 120 and 130 is approximately 0.1l along the X-coordinate axis and approximately 0.3w long the Z-coordinate axis.

[0023] FIG.2 is a flowchart illustrating a method 200 of generating a stereo recording according to an embodiment. The method 200 can be used to generate a high-quality stereo recording based on audio signals captured by two or more closely spaced omnidirectional microphones. For illustration purposes and without any implied limitations, the method 200 is described below in reference to the device 100. From the provided description, a person of ordinary skill in the pertinent art will readily understand how to practice additional embodiments of the method 200 using other types of devices without any undue experimentation.

[0024] The method 200 includes selecting a pair of microphones of the device 100 (in a block 202). In general, any two of the microphones 110, 120, 130 can be selected in the block 202 as source microphones M1 and M2. In various examples, the microphone selection is performed based on a user input or automatically, e.g., with an electronic controller of the device 100 running a programmatic microphone-selection script. In some examples, the microphones 110, 120 may be selected in the block 202 as the source microphones M1 and M2 for an audio capture to be performed with the device 100 being in the landscape orientation illustrated in FIG. 1. In this orientation, the above-described spacing of the microphones 110, 120 will typically lead to a suitable inter-channel phase difference (IPD), which helps spatial localization. In addition, at relatively high frequencies, the body of the device 100 will act as an acoustical obstacle, leading to some inter-channel level difference (ILD), which also contributes to better spatial localization.

[0025] The method 200 also includes obtaining two directional virtual microphones (in a block 204). In some examples, the two directional virtual microphones obtained in the block 204 are directional virtual microphones ^^^^and ^^^^pointing to the left and right, respectively. In one example, operations of the block 204 include determining a relative time delay corresponding to a desired polar sensitivity pattern based on the distance between the source microphones M1, M2 selected in the block 202. In some examples, the desired polar sensitivity pattern is a cardioid polar sensitivity pattern. Operations of the block 204 also include computing audio signals corresponding to the directional virtual microphones ^^^^and ^^^^using the audio signals captured by the source microphones M1 and M2, which are subjected to the determined relative time delay. More specifically, the ^^^^signal is computed by applying the relative time delay to the M2 signal and subtracting the resulting delayed signal from the M1 signal. The ^^^^signal is similarly computed by applying the relative time delay to the M1 signal and subtracting the resulting delayed signal from the M2 signal. In some examples, the applied relative time delay is a fractional delay, as it typically does not correspond to an integer number of signal samples. Some embodiments of the block 204 may benefit form certain features disclosed in the publication by Christof Faller, “Conversion of Two Closely Spaced Omnidirectional Microphone Signals to an XY Stereo Signal” presented at the 129thConvention of the Audio Engineering Society in November 2010 as Convention Paper 8188, which is incorporated herein by reference in its entirety.

[0026] In some examples, operations of the block 204 optionally include applying additional processing to the ^^^^and ^^^^signals obtained as described above to improve the audio quality thereof. Such additional processing may include: (i) replacing, in the frequency domain, high frequency portions of the ^^^^and ^^^^signals by the corresponding frequency portions of the M1 and M2 signals, respectively, and / or (ii) applying frequency-dependent equalization to low frequency portions of the ^^^^and ^^^^signals. The “replacing” operation is typically beneficial because the finite difference approximation of the pressure gradient only holds for acoustic wavelengths that are larger than the spacing between the source microphones M1 and M2. In the example described above in reference to FIG.1 the corresponding validity threshold frequency is approximately 1 kHz. The frequency portions of the ^^^^and ^^^^signals that are above this validity threshold frequency can therefore be discarded and replaced, e.g., by the corresponding frequency portions of the M1 and M2 signals, respectively. The “frequency-dependent equalization” is beneficial, e.g., because it can be used to compensate a typically present 6 dB / octave bass roll off in the ^^^^and ^^^^signals. In some examples, the frequency-dependentequalization starts at the validity threshold frequency (e.g., 1 kHz) and proceeds toward lower frequencies down to another threshold frequency at which a preselected fixed maximum boost value (e.g., 20 dB) is reached. The boost gain is then kept constant (i.e., frequency independent) at said preselected fixed maximum boost value in the frequency range below said another threshold frequency.

[0027] The method 200 also includes obtaining an inter-channel level difference (ILD) in an intermediate frequency range (in a block 206). As indicated above, the directional virtual microphones ^^^^and ^^^^serve as a good stereo pair in the intermediate frequency range, but their quality may be nonoptimal in a lower frequency range, e.g., due to a poor SNR and the above-mentioned proximity effect for sound sources that are relatively close to the microphones M1, M2. In the intermediate frequency range, spatial characteristics of the audio scene that is being recorded are well represented by the relative level difference between the ^^^^and ^^^^signals. This difference is the ILD, which is expressed in dB and can be obtained as the difference in dB between the energy of the ^^^^and ^^^^signals over a period of time.

[0028] Accordingly, in some examples, the block 206 includes the following operations. First, the signals of interest, in the time domain, are partitioned into overlapping frames. In some examples, the frame size is 1024 signal samples, and the frame overlap is 50%. Next, each of the frames is transformed to the frequency domain via a Fourier transform (e.g., a fast Fourier transform, FFT). The frequency bins of the resulting frequency-domain signals may subsequently be grouped into a smaller number of frequency bands. In some examples, the selection of edge frequencies for such frequency bands may be perceptually motivated. For an n-th frame, the corresponding ILD, denoted as ILD(n), is computed as the ratio in dB of the energy of the frequency-domain ^^^^signal across the frequency bands of interest and the energy of the frequency-domain ^^^^signal across the same frequency bands. The corresponding mathematical expression for this computation is as follows: ^^^^^^^^^^ ൌ 10 ൈ ^^^^^^ ∑ಳ ^మ^^ ^^,^^^(1) where the index b refers toone example, the frequency range B is located between 200 Hz and 500 Hz. Negative values of ILD(n) signify that the ^^^^signal is louder than the ^^^^signal. Positive values of ILD(n) signify that the ^^^^signal is louder than the ^^^^signal.

[0029] In some examples, operations of the block 206 also include averaging the ILD(n) values computed in accordance with Eq. (1) over a plurality of frames. In some examples, the averaging is performed using a sliding window. In some other examples, the averaging is performed using recursive smoothing. The corresponding mathematical expression used to implement the recursive smoothing is as follows: ^^^^^^_^^^^^^^^^^ℎ^^^^ ൌ ^1 െ ^^^^^^^^^_^^^^^^^^^^ℎ^^^ െ 1^ ^ ^^^^^^^^^^^^ (2)where 0 ^ ^^ ^ 1 is a constant that determines the effective time span of the smoothing timewindow.

[0030] The method 200 also includes applying an ILD to the M1, M2 signals in a lower frequency range (in a block 208). In some examples, the block 208 includes the following operations. First, the M1 and M2 signals are used to compute their average ILD in the lower frequency bands, e.g., located below 200 Hz. For an n-th frame, the corresponding average ILD, denoted as ^^^^^^^^^ଶ^^^^, is computed using the same approach as that expressed by Eq. (1). Second, a corresponding cross-channel gain value, denoted as ^^ௗ^^^^^, is determined as follows: ^^ௗ^^^^^ ൌ ^^^^^^^^^^ – ^^^^^^^^^ଶ^^^^ (3)Finally, the computed ^^ௗ^^^^^ or െ^^ௗ^^^^^, whichever is negative, is applied to the corresponding one of the After the cross-channel gain computed in thismanner is applied, there is no ILD discontinuity between the ^^^^, ^^^^signals and the M1, M2 signals at the junction between the intermediate frequency range and the lower frequency range. As indicated above, in one example, this junction is located at 200 Hz. In effect, these operations of the block 208 extrapolate the ^^^^^^^^^^ determined in the block 206 from the intermediate frequency range into the lower frequency range.

[0031] The method 200 also includes constructing left and right (L, R) stereo signals (in a block 210). In some examples, the block 210 includes the following operations. First, the L, R stereo signals are constructed in the frequency domain by using different portions of the various signals from the blocks 202, 204, and 208 in different frequency ranges. More specifically, for the lower frequency range, the L, R stereo signals receive the corresponding ILD-corrected portions of the M1, M2 signals, respectively, computed in the block 208. For the intermediate frequency range, the L, R stereo signals receive the corresponding portions of the ^^^^, ^^^^signals computed in the block 204. For the upper frequency range, the L, R stereo signals receive the corresponding portions of the M1, M2 signals of the block 202. After the lower, intermediate, and upper frequency ranges of the L, R stereo signals are populated in this manner, an inverse Fourier transform is applied to the resulting frequency domain signals to generate thecorresponding time-domain version of the L, R stereo signals. In some examples, the junction frequency between the lower frequency range and the intermediate frequency range is 200 Hz. The junction frequency between the intermediate frequency range and the upper frequency range is 1 kHz.

[0032] FIG.3 is a block diagram illustrating a machine-implemented algorithm 300 according to an embodiment. In some examples, the algorithm 300 is implemented in the device 100. In some other examples, the algorithm 300 is implemented in a computing device having access to the microphone signals captured by the device 100. In some cases, such a computing device can be a server that is network-connectable to the device 100. An input to the algorithm 300 includes a plurality of microphone signals 302, e.g., audio signals captured by the omnidirectional microphones 110, 120, and 130 of the device 100 illustrated in FIG.1. An output of the algorithm 300 includes left and right (L, R) stereo signals 380. At least some parts of the algorithm 300 can be implemented in accordance with the corresponding blocks of the method 200, e.g., as described in more detail below.

[0033] A block 310 of the algorithm 300 performs a microphone-signal selection in which two signals 312 from the plurality of microphone signals 302 are selected to serve as the M1, M2 signals. The microphone-signal selection is performed based on applicable metadata 304. In some examples, the metadata 304 include the orientation of the device 100 during the capture of the microphone signals 302, microphone locations on the device 100, pertinent sizes (e.g., the length l, width w, and thickness d) of the device 100, a desired first order pickup pattern, etc.

[0034] A block 314 of the algorithm 300 partitions each of the signals 312 into a corresponding sequence of partially overlapping audio frames M1(n) and M2(n). In various examples, the frame size is between 100 and 10000 signal samples, and the frame overlap is between 10% and 50%.

[0035] A block 318 of the algorithm 300 computes a sequence of partially overlapping frames Lpg(n) and Rpg(n) based on the sequence of partially overlapping frames M1(n) and M2(n) generated in the block 314 and further based on a pertinent portion of the metadata 304. In some examples, the frames Lpg(n) and Rpg(n) are computed in the block 318 using the corresponding operations of the block 204 of the method 200.

[0036] A block 322 of the algorithm 300 applies an FFT to each of the frames M1(n), M2(n), Lpg(n), and Rpg(n) to generate the corresponding discrete frequency spectra M1(f, n), M2(f, n), Lpg(f, n), and Rpg(f, n). Each of the discrete frequency spectra M1(f, n), M2(f, n), Lpg(f, n), andRpg(f, n) is then partitioned into respective spectral portions corresponding to three frequency ranges referred to as the lower, intermediate, and upper frequency (LF, IMF, and UF) ranges, respectively. The junction frequencies for the LF / IMF and IMF / UF ranges are determined based on the metadata 304. For one specific example of the device 100, which is described above in reference to FIGS.1-2, the LF / IMF junction frequency is 200 Hz, and the IMF / UF junction frequency is 1 kHz.

[0037] A block 330 of the algorithm 300 applies equalization (EQ) filtering to IMF portions 324 of the discrete frequency spectra Lpg(f, n), Rpg(f, n). In some examples, the EQ filtering implemented in the block 330 uses the corresponding frequency-dependent equalization operations of the block 204 of the method 200. The resulting equalized IMF portions are labeled in FIG.3 using the reference numeral 332.

[0038] A block 334 of the algorithm 300 computes ILD values 336 using the equalized IMF portions 332 computed in the block 330. In some examples, the ILD values 336 are computed in accordance with Eq. (1) or Eq. (2) using the corresponding operations of the block 206 of the method 200.

[0039] A block 350 of the algorithm 300 applies cross-channel gains ^^ௗ^^^^^ to LF portions 326 of the discrete frequency spectra M1(f, n), M2(f, n). In some examples, the cross-channel gains ^^ௗ^^^^^ are determined in accordance with Eq. (3) using the corresponding operations of the block 208 of the method 200 based on the ILD values 336 computed in the block 334. The resulting cross-channel-gain processed LF portions are labeled in FIG.3 using the reference numeral 352.

[0040] A block 360 of the algorithm 300 operates to construct discrete frequency spectra 362 using: (i) UF portions 328 of the of the discrete frequency spectra M1(f, n), M2(f, n); (ii) the equalized IMF portions 332 computed in the block 330; and (iii) the cross-channel-gain processed LF portions 352 computed in the block 350.

[0041] A block 370 of the algorithm 300 applies an inverse FFT to each of the discrete frequency spectra 362 constructed in the block 360. An output of the block 370 has a plurality of time-domain segments 372 corresponding to the discrete frequency spectra 362.

[0042] A block 376 of the algorithm 300 operates to appropriately combine the time-domain segments 372 computed in the block 370 to generate the L, R stereo signals 380. Since adjacent time-domain segments 372 are partially overlapped, operations of the block 376 includeappropriately truncating the overlapped portions of the adjacent time-domain segments 372 and concatenating the resulting truncated time-domain segments for each of the L and R channels of the stereo signal.

[0043] Various additional embodiments of the algorithm 300 can be implemented by incorporating one or more of the features described in more detail below. For illustration purposes, these features are described in continued reference to FIG.3.

[0044] In one additional embodiment, three microphone signals 312 from the plurality of microphone signals 302 are selected in the block 310 to serve as M1, M2, and M3 signals. A first pair of signals from the plurality of microphone signals 302 is selected as described above to serve as the M1 and M2 signals, which are then used in the block 318 to obtain pressure gradient stereo signals ^^^^^and ^^^^^in a first intermediate frequency (IMF1) range (e.g., below 1 kHz). A third microphone signal 312 is selected from the plurality of microphone signals 302 to serve as an M3 signal. The selected M3 signal is paired up with a selected one of the M1 and M2 signals to create a second signal pair representing a smaller distance between the originating omnidirectional microphones (pressure transducers) than that of the M1, M2 signal pair. This second signal pair is then used to compute pressure gradient stereo signals ^^^^ଶand ^^^^ଶin a second IMF (IMF2) range (e.g., between 1 kHz and 2 kHz). This use of the second signal pair enables an extension of the validity range of the pressure gradient approximation to a higher validity threshold frequency (e.g., 2 kHz). The discrete frequency spectra 362 are thereafter constructed in the modified block 360 by using: (i) UF portions of the discrete frequency spectra M1(f, n), M2(f, n) spectrally located above the IMF2 range, e.g., above 2 kHz; (ii) equalized portions computed based on the signals ^^^^ଶand ^^^^ଶ; (iii) equalized IMF1 portions computed based on the signals ^^^^^and ^^^^^, where IMF1 is the frequency range between the LF and IMF2 ranges; and (iv) the cross-channel-gain processed LF portions 352 computed in the block 350. The blocks 370 and 376 are then used to generate the L, R stereo signals 380 as described above, but based on the discrete frequency spectra 362 constructed over the four frequency ranges LF, IMF1, IMF2, and UF instead of the above-described three frequency ranges LF, IMF, and UF.

[0045] For the M1, M2 microphone distance of 15 cm and the M1, M3 microphone distance of 4 cm, the following example numerical ranges can be used in the corresponding embodiment of algorithm 300:(i) an example frequency range used to analyze and compute the ILD is 200 Hz – 500 Hz; (ii) an example frequency range where the computed ILD is imposed by extrapolation is 50 Hz – 200 Hz; (iii) an example IMF1 range is 200 Hz – 1 kHz; (iv) an example IMF2 range is 1 kHz – 2 kHz; and (v) an example UF range is 2 kHz – 20 kHz.

[0046] In another additional embodiment, once a pressure gradient stereo signal is obtained in a certain band ^^^, the ILD computed from ^^^^and ^^^^is applied to M1, M2 in the same band ^^^, or in a band ^^ଶthat is contained within ^^^, or that is partially overlapping with ^^^. This approach can be used, e.g., to avoid any audible artifacts caused by the filtering and signal processing used to obtain the pressure gradient, which in this case is used only as a guidance signal to steer the enhancement of M1 and M2, but not as an audio component for the L, R stereo signals 380.

[0047] In various examples, the M1, M2, and M3 signals can be raw signals from the microphones 110, 120, 130 or pre-processed versions of those signals. The pre-processing may include one or more of the following: (i) noise reduction directed at attenuating the electrical self-noise of the microphones; (ii) application of a broadband gain configured to compensate for different respective sensitivities of the microphones; and (iii) application of a frequency dependent gain configured to adjust the microphone frequency response to a desired frequency response, e.g., to a flat frequency response.

[0048] In yet another additional embodiment, the EQ filters to flatten the response of the pressure gradient signals described above in reference to the block 330 of the algorithm 300 and the block 204 of the method 200 are obtained empirically in a dynamic way, instead of being based on theoretical grounds. One purpose of those filters is to equalize the pressure gradient signal to the pressure signal, compensating for an inherent low-frequency loss caused by taking a finite difference. When a fixed filter is used, a tradeoff is realized, e.g., because the response of the pressure gradient changes depending on the source distance and the angle of incidence. In some examples, the tradeoff is directed to compensating for the far-field (plane wave) conditions and accepting the proximity effect substantially without compensation. However, in some cases, ground-truth pressure signals are available from the M1 and M2 microphones and can be used as an accurate reference regardless of the source distance and direction. For this additionalembodiment, in the block 206 of the method 200, one also computes the level per band of M1 and M2, ^^ଶ^^^^, ^^^ and ^^ଶଶ^^^, ^^^, and use these levels to compute equalization gains as follows:^^ ^^^, ^^^ ൌ 1ெభమ^^,^^^0 ∗ ^^^^^^^మ^^^^,^^െ ^^^^^^ೞ^^^^^^ (4)మ^ ^ where ^^ ^^^^ the overall ILDthe pressure In some examples, the offset gain values are determined as follows: ^^ ^^^^ ൌெభమ^^,^బ^ ^^^^ೞ^^ 10 ∗ ^^^^^^^మ^^^^,^బ^ (6) where ^^ is a pre-holds, for example, a

[0049] The gains ^^^and ^^ଶ, e.g., determined in accordance with Eqs. (4)-(7), are applied to the ^^^^and ^^^^signals, respectively, and typically ensure that their frequency characteristics correspond to those of the pressure transducers, regardless of the position of the sound source, effectively taking care of the proximity effect. In some examples, the gains ^^^and ^^ଶare smoothed over time and frequency and are clipped to not exceed a selected fixed maximum boost (e.g., 20 dB). In some examples, instead of computing the equalization gains for the left and right stereo channels based on the M1 and M2 signals, respectively, the equalization gains are computed based on an average of the M1 and M2 signals for each audio frame and / or frequency band or range.

[0050] FIGS.4A-4B graphically illustrate certain performance characteristics of the algorithm 300 according to an additional embodiment. Traces 401 and 402 illustrate frequency characteristics of the M1 and M2 signals, respectively, according to one example. A level difference, ^1, between the traces 401, 402 represents the ILD corresponding to the microphones M1 and M2. Traces 403 and 404 (FIG.4A) illustrate frequency characteristics of the pressure gradient signals computed based on the M1, M2, and M3 signals. The above-described bass roll off is evident from the illustrated frequency characteristics of the traces 403 and 404. Kinks in the traces 403 and 404 located at approximately 500 Hz are due to the use of two different pairs of microphones when obtaining the gradient signals. More specifically, the M1 and M2 signals are used below 500 Hz whereas the M1 and M3 are used above 500 Hz. A level difference, ^2, between the traces 403, 404 represents the ILD corresponding to the virtual directionalmicrophones. Traces 405 and 406 (FIG.4B) illustrate frequency characteristics of equalized pressure gradient signals obtained by applying frequency dependent boost to the initial pressure gradient signals illustrated by the traces 403 and 404 (FIG.4A). The level difference between the traces 405 and 406 represents the ILD of the equalized pressure gradient signals. This ILD is maintained at ^2 due to the above-described characteristics of the frequency dependent boost.

[0051] FIG.5 is a block diagram of an example computing device 500 configured to perform at least some operations of the method 200 and / or algorithm 300 according to various embodiments. In some embodiments, the computing device 500 can be a part of the device 100. In some other embodiments, a single computing device 500 or multiple computing devices 500 can be network connected to the device 100 to receive therefrom the microphone signals 302 and the metadata 304.

[0052] The computing device 500 of FIG.5 is illustrated as having a number of components, but any one or more of these components may be omitted or duplicated, as suitable for the application and setting. In some embodiments, some or all of the components included in the computing device 500 may be attached to one or more motherboards and enclosed in a housing. In some embodiments, some of those components may be fabricated onto a single system-on-a- chip (SoC) (e.g., the SoC may include one or more electronic processing devices 502 and one or more storage devices 504). Additionally, in various embodiments, the computing device 500 may not include one or more of the components illustrated in FIG.5, but may include interface circuitry for coupling to the one or more components using any suitable interface (e.g., a Universal Serial Bus (USB) interface, a High-Definition Multimedia Interface (HDMI) interface, a Controller Area Network (CAN) interface, a Serial Peripheral Interface (SPI) interface, an Ethernet interface, a wireless interface, or any other appropriate interface). For example, the computing device 500 may not include a display device 510, but may include display device interface circuitry (e.g., a connector and driver circuitry) to which an external display device 510 may be coupled.

[0053] The computing device 500 includes a processing device 502 (e.g., one or more processing devices). As used herein, the terms “electronic processor device” and “processing device” interchangeably refer to any device or portion of a device that processes electronic data from registers and / or memory to transform that electronic data into other electronic data that may be stored in registers and / or memory. In various embodiments, the processing device 502 may include one or more digital signal processors (DSPs), application-specific integrated circuits(ASICs), central processing units (CPUs), graphics processing units (GPUs), server processors, or any other suitable processing devices.

[0054] The computing device 500 also includes a storage device 504 (e.g., one or more storage devices). In various embodiments, the storage device 504 may include one or more memory devices, such as random-access memory (RAM) devices (e.g., static RAM (SRAM) devices, magnetic RAM (MRAM) devices, dynamic RAM (DRAM) devices, resistive RAM (RRAM) devices, or conductive-bridging RAM (CBRAM) devices), hard drive-based memory devices, solid-state memory devices, networked drives, cloud drives, or any combination of memory devices. In some embodiments, the storage device 504 may include memory that shares a die with the processing device 502. In such an embodiment, the memory may be used as cache memory and include embedded dynamic random-access memory (eDRAM) or spin transfer torque magnetic random-access memory (STT-MRAM), for example. In some embodiments, the storage device 504 may include non-transitory computer readable media having instructions thereon that, when executed by one or more processing devices (e.g., the processing device 502), cause the computing device 500 to perform any appropriate ones of the methods disclosed herein below or portions of such methods.

[0055] The computing device 500 further includes an interface device 506 (e.g., one or more interface devices 506). In various embodiments, the interface device 506 may include one or more communication chips, connectors, and / or other hardware and software to govern communications between the computing device 500 and other computing devices. For example, the interface device 506 may include circuitry for managing wireless communications for the transfer of data to and from the computing device 500. The term “wireless” and its derivatives may be used to describe circuits, devices, systems, methods, techniques, communications channels, etc., that may communicate data via modulated electromagnetic radiation through a nonsolid medium. The term does not imply that the associated devices do not contain any wires, although in some embodiments they might not. Circuitry included in the interface device 506 for managing wireless communications may implement any of a number of wireless standards or protocols, including but not limited to Institute for Electrical and Electronic Engineers (IEEE) standards including Wi-Fi (IEEE 802.11 family), IEEE 802.16 standards, Long-Term Evolution (LTE) project along with any amendments, updates, and / or revisions (e.g., advanced LTE project, ultramobile broadband (UMB) project (also referred to as “3GPP2”), etc.). In some embodiments, circuitry included in the interface device 506 for managing wireless communications may operate in accordance with a Global System for Mobile Communication(GSM), General Packet Radio Service (GPRS), Universal Mobile Telecommunications System (UMTS), High Speed Packet Access (HSPA), Evolved HSPA (E-HSPA), or LTE network. In some embodiments, circuitry included in the interface device 506 for managing wireless communications may operate in accordance with Enhanced Data for GSM Evolution (EDGE), GSM EDGE Radio Access Network (GERAN), Universal Terrestrial Radio Access Network (UTRAN), or Evolved UTRAN (E-UTRAN). In some embodiments, circuitry included in the interface device 506 for managing wireless communications may operate in accordance with Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Digital Enhanced Cordless Telecommunications (DECT), Evolution-Data Optimized (EV-DO), and derivatives thereof, as well as any other wireless protocols that are designated as 3G, 4G, 5G, and beyond. In some embodiments, the interface device 506 may include one or more antennas (e.g., one or more antenna arrays) configured to receive and / or transmit wireless signals.

[0056] In some embodiments, the interface device 506 may include circuitry for managing wired communications, such as electrical, optical, or any other suitable communication protocols. For example, the interface device 506 may include circuitry to support communications in accordance with Ethernet technologies. In some embodiments, the interface device 506 may support both wireless and wired communication, and / or may support multiple wired communication protocols and / or multiple wireless communication protocols. For example, a first set of circuitry of the interface device 506 may be dedicated to shorter-range wireless communications such as Wi-Fi or Bluetooth, and a second set of circuitry of the interface device 506 may be dedicated to longer-range wireless communications such as global positioning system (GPS), EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, or others. In some other embodiments, a first set of circuitry of the interface device 506 may be dedicated to wireless communications, and a second set of circuitry of the interface device 506 may be dedicated to wired communications.

[0057] The computing device 500 also includes battery / power circuitry 508. In various embodiments, the battery / power circuitry 508 may include one or more energy storage devices (e.g., batteries or capacitors) and / or circuitry for coupling components of the computing device 500 to an energy source separate from the computing device 500 (e.g., to AC line power).

[0058] The computing device 500 also includes a display device 510 (e.g., one or multiple individual display devices). In various embodiments, the display device 510 may include any visual indicators, such as a heads-up display, a computer monitor, a projector, a touchscreen display, a liquid crystal display (LCD), a light-emitting diode display, or a flat panel display.

[0059] The computing device 500 also includes additional input / output (I / O) devices 512. In various embodiments, the I / O devices 512 may include one or more data / signal transfer interfaces, audio I / O devices (e.g., microphones or microphone arrays, speakers, headsets, earbuds, alarms, etc.), audio codecs, video codecs, printers, sensors (e.g., thermocouples or other temperature sensors, humidity sensors, pressure sensors, vibration sensors, etc.), image capture devices (e.g., one or more cameras), human interface devices (e.g., keyboards, cursor control devices, such as a mouse, a stylus, a trackball, or a touchpad), etc.

[0060] Depending on the specific embodiment, various components of the interface devices 506 and / or I / O devices 512 can be configured to output suitable control signals, receive suitable control / telemetry signals, and receive and transmit data streams. In some examples, the interface devices 506 and / or I / O devices 512 include one or more analog-to-digital converters (ADCs) for transforming received analog signals into a digital form suitable for operations performed by the processing device 502 and / or the storage device 504. In some additional examples, the interface devices 506 and / or I / O devices 512 include one or more digital-to-analog converters (DACs) for transforming digital signals provided by the processing device 502 and / or the storage device 504 into an analog form suitable for being transmitted through a communication channel.

[0061] According to an example embodiment disclosed above, e.g., in the summary section and / or in reference to any one or any combination of some or all of FIGS.1-5, provided is a method of generating a stereo signal comprising: obtaining a first directional virtual-microphone signal and a second directional virtual-microphone signal based on a first audio signal captured with a first omnidirectional microphone and further based on a second audio signal captured with a second omnidirectional microphone; determining a level difference between the first and second directional virtual-microphone signals in a first frequency range; applying a gain to the second audio signal in a second frequency range to generate a third audio signal, the gain being selected to cause a level difference between the third audio signal and the first audio signal in the second frequency range to match the determined level difference between the first and second directional virtual-microphone signals in the first frequency range; and generating the stereo signal using the first directional virtual-microphone signal, the second directional virtual- microphone signal, and the third audio signal.

[0062] In some embodiments of the above method, a same electronic device houses the first omnidirectional microphone and the second omnidirectional microphone, at least one external lateral dimension of said electronic device being smaller than 20 cm.

[0063] In some embodiments of any of the above methods, the electronic device is a smartphone or a camera.

[0064] In some embodiments of any of the above methods, the second frequency range has lower frequencies than the first frequency range.

[0065] In some embodiments of any of the above methods, the generating comprises: combining a spectral portion of the first directional virtual-microphone signal corresponding to the first frequency range and a first spectral portion of the first audio signal corresponding to the second frequency range to generate a first spectral portion of a first stereo channel of the stereo signal; and combining a spectral portion of the second directional virtual-microphone signal corresponding to the first frequency range and a first spectral portion of the third audio signal corresponding to the second frequency range to generate a first spectral portion of a second stereo channel of the stereo signal.

[0066] In some embodiments of any of the above methods, the generating further comprises: combining the first spectral portion of the first stereo channel and a second spectral portion of the first audio signal corresponding to a third frequency range, the third frequency range having higher frequencies than the first frequency range; and combining the first spectral portion of the second stereo channel and a second spectral portion of the second audio signal corresponding to the third frequency range.

[0067] In some embodiments of any of the above methods, the method further comprises obtaining a third directional virtual-microphone signal and a fourth directional virtual- microphone signal based on a selected one of the first and second audio signals and further based on a fourth audio signal captured with a third omnidirectional microphone.

[0068] In some embodiments of any of the above methods, the generating further comprises: combining the first spectral portion of the first stereo channel and a spectral portion of the third directional virtual-microphone signal corresponding to a third frequency range to generate a larger spectral portion of the first stereo channel, the third frequency range having higher frequencies than the first frequency range; and combining the first spectral portion of the second stereo channel and a spectral portion of the fourth directional virtual-microphone signal corresponding to the third frequency range to generate a larger spectral portion of the second stereo channel.

[0069] In some embodiments of any of the above methods, the generating further comprises: combining the larger spectral portion of the first stereo channel and a second spectral portion of the first audio signal corresponding to a fourth frequency range, the fourth frequency range having higher frequencies than the third frequency range; and combining the larger spectral portion of the second stereo channel and a second spectral portion of the second audio signal corresponding to the fourth frequency range.

[0070] In some embodiments of any of the above methods, a same electronic device houses the first omnidirectional microphone, the second omnidirectional microphone, and the third omnidirectional microphone.

[0071] In some embodiments of any of the above methods, the obtaining comprises: applying a time delay to delay one of the first and second audio signals relative to the other one of the first and second audio signals; and computing a difference between the relatively delayed first and second audio signals.

[0072] In some embodiments of any of the above methods, the method further comprises selecting the time delay such that the first directional virtual-microphone signal or the second directional virtual-microphone signal represents a selected polar pickup pattern.

[0073] In some embodiments of any of the above methods, the obtaining comprises equalizing the difference to reduce a frequency dependent roll off.

[0074] In some embodiments of any of the above methods, the method further comprises: partitioning each of the first and second audio signals into a respective sequence of partially overlapping audio frames; and applying a Fourier transform to each of the audio frames to obtain respective signal portions corresponding to different frequency ranges.

[0075] In some embodiments of any of the above methods, the generating further comprises: applying an inverse Fourier transform to each of a plurality of discrete frequency spectra representing the stereo signal in a frequency domain to generate first and second sequences of audio-signal segments; combining the audio-signal segments of the first sequence to generate a time-domain signal for a first stereo channel of the stereo signal; and combining the audio-signal segments of the second sequence to generate a time-domain signal for a second stereo channel of the stereo signal.

[0076] In some embodiments of any of the above methods, the generating further comprises: combining frequency-domain segments of the signal portions corresponding to different frequency ranges to generate a frequency-domain signal for a first stereo channel of the stereo signal; combining frequency-domain segments of the signal portions corresponding to different frequency ranges to generate a frequency-domain signal for a second stereo channel of the stereo signal; and applying an inverse Fourier transform to each of the first and second frequency- domain stereo channels to generate the first and second stereo channels of the stereo signal in a time domain.

[0077] In some embodiments of any of the above methods, the first audio signal has a higher level than the second audio signal; and wherein the applied gain has a negative value.

[0078] According to another example embodiment disclosed above, e.g., in the summary section and / or in reference to any one or any combination of some or all of FIGS.1-5, provided is a non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising any one of the above methods.

[0079] According to yet another an example embodiment disclosed above, e.g., in the summary section and / or in reference to any one or any combination of some or all of FIGS.1-5, provided is an apparatus comprising: at least one processor; and at least one memory including program code; and wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: obtain a first directional virtual- microphone signal and a second directional virtual-microphone signal based on a first audio signal captured with a first omnidirectional microphone and further based on a second audio signal captured with a second omnidirectional microphone; determine a level difference between the first and second directional virtual-microphone signals in a first frequency range; apply a gain to the second audio signal in a second frequency range to generate a third audio signal, the gain being selected to cause a level difference between the third audio signal and the first audio signal in the second frequency range to match the determined level difference between the first and second directional virtual-microphone signals in the first frequency range; and generate the stereo signal using the first directional virtual-microphone signal, the second directional virtual- microphone signal, and the third audio signal.

[0080] In some embodiments of the above apparatus, a same electronic device houses the first omnidirectional microphone, the second omnidirectional microphone, the at least one processor, and the at least one memory.

[0081] In some embodiments of any of the above apparatus, the apparatus includes an electronic device and a computing device; wherein the electronic device includes the first omnidirectional microphone and the second omnidirectional microphone; and wherein the computing device includes the at least one processor and the at least one memory.

[0082] In some embodiments of any of the above apparatus, the at least one memory and the program code are further configured to, with the at least one processor, cause the apparatus to: obtain a third directional virtual-microphone signal and a fourth directional virtual-microphone signal based on a selected one of the first and second audio signals and further based on a fourth audio signal captured with a third omnidirectional microphone; and generate the stereo signal additionally using the third directional virtual-microphone signal and the fourth directional virtual-microphone signal.

[0083] With regard to the processes, systems, methods, heuristics, etc. described herein, it should be understood that, although the steps of such processes, etc. have been described as occurring according to a certain ordered sequence, such processes could be practiced with the described steps performed in an order other than the order described herein. It further should be understood that certain steps could be performed simultaneously, that other steps could be added, or that certain steps described herein could be omitted. In other words, the descriptions of processes herein are provided for the purpose of illustrating certain embodiments and should in no way be construed so as to limit the claims.

[0084] Accordingly, it is to be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided would be apparent upon reading the above description. The scope should be determined, not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technologies discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In sum, it should be understood that the application is capable of modification and variation.

[0085] All terms used in the claims are intended to be given their broadest reasonable constructions and their ordinary meanings as understood by those knowledgeable in the technologies described herein unless an explicit indication to the contrary is made herein. In particular, use of the singular articles such as “a,” “the,” “said,” etc. should be read to recite one or more of the indicated elements unless a claim recites an explicit limitation to the contrary.

[0086] The Abstract of the Disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in fewer than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separately claimed subject matter.

[0087] While this disclosure includes references to illustrative embodiments, this specification is not intended to be construed in a limiting sense. Various modifications of the described embodiments, as well as other embodiments within the scope of the disclosure, which are apparent to persons skilled in the art to which the disclosure pertains are deemed to lie within the principle and scope of the disclosure, e.g., as expressed in the following claims.

[0088] Some embodiments may be implemented as circuit-based processes, including possible implementation on a single integrated circuit.

[0089] Some embodiments can be embodied in the form of methods and apparatuses for practicing those methods. Some embodiments can also be embodied in the form of program code recorded in tangible media, such as magnetic recording media, optical recording media, solid state memory, floppy diskettes, CD-ROMs, hard drives, or any other non-transitory machine-readable storage medium, wherein, when the program code is loaded into and executed by a machine, such as a computer, the machine becomes an apparatus for practicing the patented invention(s). Some embodiments can also be embodied in the form of program code, for example, stored in a non-transitory machine-readable storage medium including being loaded into and / or executed by a machine, wherein, when the program code is loaded into and executed by a machine, such as a computer or a processor, the machine becomes an apparatus forpracticing the patented invention(s). When implemented on a general-purpose processor, the program code segments combine with the processor to provide a unique device that operates analogously to specific logic circuits.

[0090] Unless explicitly stated otherwise, each numerical value and range should be interpreted as being approximate as if the word “about” or “approximately” preceded the value or range.

[0091] The use of figure numbers and / or figure reference labels in the claims is intended to identify one or more possible embodiments of the claimed subject matter in order to facilitate the interpretation of the claims. Such use is not to be construed as necessarily limiting the scope of those claims to the embodiments shown in the corresponding figures.

[0092] Although the elements in the following method claims, if any, are recited in a particular sequence with corresponding labeling, unless the claim recitations otherwise imply a particular sequence for implementing some or all of those elements, those elements are not necessarily intended to be limited to being implemented in that particular sequence.

[0093] Reference herein to “one embodiment” or “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the disclosure. The appearances of the phrase “in one embodiment” in various places in the specification are not necessarily all referring to the same embodiment, nor are separate or alternative embodiments necessarily mutually exclusive of other embodiments. The same applies to the term “implementation.”

[0094] Unless otherwise specified herein, the use of the ordinal adjectives “first,” “second,” “third,” etc., to refer to an object of a plurality of like objects merely indicates that different instances of such like objects are being referred to, and is not intended to imply that the like objects so referred-to have to be in a corresponding order or sequence, either temporally, spatially, in ranking, or in any other manner.

[0095] Unless otherwise specified herein, in addition to its plain meaning, the conjunction “if” may also or alternatively be construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” which construal may depend on the corresponding specific context. For example, the phrase “if it is determined” or “if [a stated condition] is detected” may be construed to mean “upon determining” or “in response to determining” or“upon detecting [the stated condition or event]” or “in response to detecting [the stated condition or event].”

[0096] Also for purposes of this description, the terms “couple,” “coupling,” “coupled,” “connect,” “connecting,” or “connected” refer to any manner known in the art or later developed in which energy is allowed to be transferred between two or more elements, and the interposition of one or more additional elements is contemplated, although not required. Conversely, the terms “directly coupled,” “directly connected,” etc., imply the absence of such additional elements.

[0097] The functions of the various elements shown in the figures, including any functional blocks labeled as “processors” and / or “controllers,” may be provided through the use of dedicated hardware as well as hardware capable of executing software in association with appropriate software. When provided by a processor, the functions may be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors, some of which may be shared. Moreover, explicit use of the term “processor” or “controller” should not be construed to refer exclusively to hardware capable of executing software, and may implicitly include, without limitation, digital signal processor (DSP) hardware, network processor, application specific integrated circuit (ASIC), field programmable gate array (FPGA), read only memory (ROM) for storing software, random access memory (RAM), and non volatile storage. Other hardware, conventional and / or custom, may also be included. Similarly, any switches shown in the figures are conceptual only. Their function may be carried out through the operation of program logic, through dedicated logic, through the interaction of program control and dedicated logic, or even manually, the particular technique being selectable by the implementer as more specifically understood from the context.

[0098] As used in this application, the terms “circuit,” “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry); (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions); and (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.” This definition of circuitry applies to all uses of this term in this application,including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.

[0099] It should be appreciated by those of ordinary skill in the art that any block diagrams herein represent conceptual views of illustrative circuitry embodying the principles of the disclosure. Similarly, it will be appreciated that any flow charts, flow diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in computer readable medium and so executed by a computer or processor, whether or not such computer or processor is explicitly shown.

[0100] “BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS” in this specification is intended to introduce some example embodiments, with additional embodiments being described in “DETAILED DESCRIPTION” and / or in reference to one or more drawings. “BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS” is not intended to identify essential elements or features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.

Claims

CLAIMS What is claimed is:

1. A method of generating a stereo signal, the method comprising: obtaining a first directional virtual-microphone signal and a second directional virtual- microphone signal based on a first audio signal captured with a first microphone and further based on a second audio signal captured with a second microphone; determining a level difference between the first and second directional virtual- microphone signals in a first frequency range; applying a gain to the second audio signal in a second frequency range to generate a third audio signal, the gain being selected to cause a level difference between the third audio signal and the first audio signal in the second frequency range to match the determined level difference between the first and second directional virtual-microphone signals in the first frequency range; and generating the stereo signal using the first directional virtual-microphone signal, the second directional virtual-microphone signal, and the third audio signal; wherein generating the stereo signal comprises: combining a spectral portion of the first directional virtual-microphone signal corresponding to the first frequency range and a first spectral portion of the first audio signal corresponding to the second frequency range to generate a first spectral portion of a first stereo channel of the stereo signal; and combining a spectral portion of the second directional virtual-microphone signal corresponding to the first frequency range and a first spectral portion of the third audio signal corresponding to the second frequency range to generate a first spectral portion of a second stereo channel of the stereo signal.

2. The method of claim 1, wherein a same electronic device houses the first microphone and the second microphone, at least one external lateral dimension of said electronic device being smaller than 20 cm.

3. The method of claim 1 or claim 2, wherein the second frequency range has lower frequencies than the first frequency range.

4. The method of any one of claims 1-3, wherein generating the stereo signal further comprises: combining the first spectral portion of the first stereo channel and a second spectral portion of the first audio signal corresponding to a third frequency range, the third frequency range having higher frequencies than the first frequency range; and combining the first spectral portion of the second stereo channel and a second spectral portion of the second audio signal corresponding to the third frequency range.

5. The method of claim 4, further comprising obtaining a third directional virtual- microphone signal and a fourth directional virtual-microphone signal based on a selected one of the first and second audio signals and further based on a fourth audio signal captured with a third microphone.

6. The method of claim 5, wherein generating the stereo signal further comprises: combining the first spectral portion of the first stereo channel and a spectral portion of the third directional virtual-microphone signal corresponding to a third frequency range to generate a larger spectral portion of the first stereo channel, the third frequency range having higher frequencies than the first frequency range; and combining the first spectral portion of the second stereo channel and a spectral portion of the fourth directional virtual-microphone signal corresponding to the third frequency range to generate a larger spectral portion of the second stereo channel.

7. The method of claim 6, wherein generating the stereo signal further comprises: combining the larger spectral portion of the first stereo channel and a second spectral portion of the first audio signal corresponding to a fourth frequency range, the fourth frequency range having higher frequencies than the third frequency range; and combining the larger spectral portion of the second stereo channel and a second spectral portion of the second audio signal corresponding to the fourth frequency range.

8. The method of any one of claims 5-7, wherein a same electronic device houses the first microphone, the second microphone, and the third microphone.

9. The method of any one of claims 1-8, wherein the obtaining comprises: applying a time delay to delay one of the first and second audio signals relative to the other one of the first and second audio signals; andcomputing a difference between the relatively delayed first and second audio signals.

10. The method of claim 9, further comprising selecting the time delay such that the first directional virtual-microphone signal or the second directional virtual-microphone signal represents a selected polar pickup pattern.

11. The method of claim 9 or claim 10, wherein the obtaining comprises equalizing the difference to reduce a frequency dependent roll off.

12. The method of any one of claims 1-11, further comprising: partitioning each of the first and second audio signals into a respective sequence of partially overlapping audio frames; and applying a Fourier transform to each of the audio frames to obtain respective signal portions corresponding to different frequency ranges.

13. The method of claim 12, wherein generating the stereo signal further comprises: combining frequency-domain segments of the signal portions corresponding to different frequency ranges to generate a frequency-domain signal for a first stereo channel of the stereo signal; combining frequency-domain segments of the signal portions corresponding to different frequency ranges to generate a frequency-domain signal for a second stereo channel of the stereo signal; and applying an inverse Fourier transform to each of the first and second frequency-domain stereo channels to generate the first and second stereo channels of the stereo signal in a time domain.

14. The method of any one of claims 1-13, wherein the first audio signal has a higher level than the second audio signal; and wherein the applied gain has a negative value.

15. A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising the method of any one of claims 1-14.

16. An apparatus for generating a stereo signal, the apparatus comprising:at least one processor; and at least one memory including program code; and wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: obtain a first directional virtual-microphone signal and a second directional virtual- microphone signal based on a first audio signal captured with a first microphone and further based on a second audio signal captured with a second microphone; determine a level difference between the first and second directional virtual-microphone signals in a first frequency range; apply a gain to the second audio signal in a second frequency range to generate a third audio signal, the gain being selected to cause a level difference between the third audio signal and the first audio signal in the second frequency range to match the determined level difference between the first and second directional virtual-microphone signals in the first frequency range; and generate the stereo signal using the first directional virtual-microphone signal, the second directional virtual-microphone signal, and the third audio signal; wherein, to generate the stereo signal, the at least one memory and the program code are further configured to, with the at least one processor, cause the apparatus to: combine a spectral portion of the first directional virtual-microphone signal corresponding to the first frequency range and a first spectral portion of the first audio signal corresponding to the second frequency range to generate a first spectral portion of a first stereo channel of the stereo signal; and combine a spectral portion of the second directional virtual-microphone signal corresponding to the first frequency range and a first spectral portion of the third audio signal corresponding to the second frequency range to generate a first spectral portion of a second stereo channel of the stereo signal.

17. The apparatus of claim 16, wherein a same electronic device houses the first microphone, the second microphone, the at least one processor, and the at least one memory.

18. The apparatus of claim 16, wherein the apparatus includes an electronic device and a computing device; wherein the electronic device includes the first microphone and the second microphone; andwherein the computing device includes the at least one processor and the at least one memory.

19. The apparatus of any one of claims 16-18, wherein the at least one memory and the program code are further configured to, with the at least one processor, cause the apparatus to: obtain a third directional virtual-microphone signal and a fourth directional virtual- microphone signal based on a selected one of the first and second audio signals and further based on a fourth audio signal captured with a third microphone; and generate the stereo signal additionally using the third directional virtual-microphone signal and the fourth directional virtual-microphone signal.

Citation Information

Patent Citations

  • Magnified binaural CUES in a binaural hearing system

    WO2022112879A1