Stereo recording using close-spaced microphones

By combining closely spaced omnidirectional microphone signals and performing frequency domain processing, high-quality stereo recordings are generated, solving the problem of poor stereo recording quality in consumer electronic devices, especially significantly improving stereo performance in the mid-frequency range.

CN122460098APending Publication Date: 2026-07-24DOLBY LABORATORIES LICENSING CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DOLBY LABORATORIES LICENSING CORP
Filing Date
2024-12-20
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing technologies struggle to generate high-quality stereo recordings in consumer electronics devices, especially due to the small time delay and intensity differences caused by the close spacing of microphones, resulting in an unfavorable narrow stereo image and poor performance in the low and high frequency ranges.

Method used

By combining closely spaced omnidirectional microphone signals, a directional virtual microphone signal is generated. The non-uniformity of the pressure gradient estimation is corrected through frequency domain processing, and a high-quality stereo signal is generated using relative time delay and frequency domain equalization techniques.

Benefits of technology

It generates high-quality stereo signals in the mid-frequency range, corrects the inhomogeneity in the low-frequency band, improves the stereo effect, expands the frequency response range, and improves the signal-to-noise ratio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122460098A_ABST
    Figure CN122460098A_ABST
Patent Text Reader

Abstract

A method and apparatus for obtaining a stereo recording using closely spaced physical microphones, such as the microphones of a consumer electronic device. Various examples construct the signals of two stereo channels by combining different spectral portions of different generated signals. In one example, one stereo channel of a stereo recording is constructed using spectral portions (332) of a directional virtual microphone signal derived from audio signals (302) captured from the physical microphones, a modified spectral portion (352) of one of the audio signals, and an unmodified spectral portion (328) of the audio signal. The directional virtual microphone signal is generated using a pressure gradient method (318). The modification (350) of the audio signal includes applying an inter-channel level difference (ILD) that matches an inter-channel level difference (ILD) of the directional virtual microphone signal (332) in an adjacent frequency range (IMF). In various examples, corresponding algorithms can run at the electronic device housing the closely spaced microphones or at a remote server.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] 1. Cross-references to related applications This application claims priority to the following patent applications: Spanish Patent Application No. P202331063, filed December 21, 2023; European Patent Application No. 24174170.1, filed May 3, 2024; and U.S. Provisional Patent Application No. 63 / 642,985, filed May 6, 2024, each of which is incorporated herein by reference in its entirety. 2. Technical Field Various example embodiments relate to audio devices, and more specifically, but not exclusively, to apparatus for capturing and reproducing stereo sound. 3. Background Technology Stereo sound (sometimes called "stereo") is a method of sound reproduction that reconstructs the auditory perspective of space. Stereo effects are typically achieved using two independent audio channels through a configuration of two speakers (or stereo headphones) to create the impression of hearing sounds from different directions, much like natural hearing.

[0004] During typical two-channel stereo recording, two microphones are placed in appropriately chosen positions relative to the sound source, recording simultaneously using both microphones. The two recorded channels are similar but have different arrival times and / or sound pressure levels. During playback, the listener's brain uses these subtle differences in time and sound levels to perceive the location of the recorded object, thus detecting the stereo effect. The three most widely used stereo recording setups (AB setup, XY setup, and ORTF setup) create a stereo effect by capturing time and / or level differences. Summary of the Invention

[0005] Example embodiments provide a method and apparatus for obtaining stereo recordings using closely spaced omnidirectional microphones (such as microphones commonly found in consumer electronics devices). Various embodiments construct signals for two stereo channels of the stereo recording by combining different spectral portions of different generated signals. In one example, the stereo channels of the stereo recording are constructed using: a spectral portion of a directional virtual microphone signal derived from an audio signal captured by the omnidirectional microphone, a modified spectral portion of one audio signal from the audio signal, and an unmodified spectral portion of that audio signal. The directional virtual microphone signal is generated using a pressure gradient method. The modification of the audio signal includes applying an inter-channel level difference (ILD) that matches the inter-channel level difference (ILD) of the directional virtual microphone signal in adjacent frequency ranges. In various examples, the corresponding algorithms can run locally at the electronic device housing the closely spaced omnidirectional microphone, or remotely at a server networked to the electronic device, for example.

[0006] According to an example embodiment, a method for generating a stereo signal is provided, the method comprising: obtaining a first directional virtual microphone mono signal and a second directional virtual microphone mono signal based on a first audio signal captured by a first microphone and further based on a second audio signal captured by a second microphone; determining a level difference between the first directional virtual microphone signal and the second directional virtual microphone signal in a first frequency range; applying a gain to the second audio signal in a second frequency range to generate a third mono audio signal, the gain being selected such that the level difference between the third audio signal and the first audio signal in the second frequency range matches the determined level difference between the first directional virtual microphone signal and the second directional virtual microphone signal in the first frequency range; and generating a stereo signal using the first directional virtual microphone signal, the second directional virtual microphone signal, and the third audio signal.

[0007] According to another example embodiment, a non-transitory computer-readable medium is provided that stores instructions that, when executed by an electronic processor, cause the electronic processor to perform operations including the methods described above.

[0008] According to yet another example embodiment, an apparatus for generating a stereo signal is provided, the apparatus comprising: at least one processor; and at least one memory including program code; wherein the at least one memory and the program code are configured, together with the at least one processor, to enable the apparatus to at least: obtain a first directional virtual microphone signal and a second directional virtual microphone signal based on a first audio signal captured by a first microphone and further based on a second audio signal captured by a second microphone; determine a level difference between the first directional virtual microphone signal and the second directional virtual microphone signal in a first frequency range; apply a gain to the second audio signal in a second frequency range to generate a third audio signal, the gain being selected such that the level difference between the third audio signal and the first audio signal in the second frequency range matches the determined level difference between the first directional virtual microphone signal and the second directional virtual microphone signal in the first frequency range; and generate a stereo signal using the first directional virtual microphone signal, the second directional virtual microphone signal, and the third audio signal. Attached Figure Description

[0009] Other aspects, features, and advantages of the various disclosed embodiments will become more fully apparent from the following detailed description and accompanying drawings, wherein: Figure 1 This is a schematic diagram showing a 3D perspective view of a consumer electronic device, some embodiments of which can be implemented in the device.

[0010] Figure 2 This is a flowchart illustrating a method for generating stereo recordings according to an embodiment.

[0011] Figure 3 This is a block diagram illustrating a machine-implemented algorithm for generating stereo recordings according to an embodiment.

[0012] Figures 4A to 4B The illustration shows an embodiment according to the additional embodiment. Figure 3 The specific performance characteristics of the algorithm.

[0013] Figure 5 It is configured to execute according to various embodiments. Figure 2 Methods and / or Figure 3 A block diagram of an example computing device that implements at least some operations of the algorithm on a machine. Detailed Implementation

[0014] As more and more content, such as audio and video, is created via smartphones, webcams, and other consumer electronics devices, content creators and device manufacturers are seeking to improve the quality of audio capture. One aspect affecting perceived audio quality relates to the spatial characteristics of audio capture. However, the size specifications of many devices often make it difficult to house directional pressure gradient microphones used for stereo signal capture. Instead, omnidirectional microphones are typically used. Example omnidirectional microphones include pressure transducers, which are relatively easy to integrate into the device body by placing them behind small openings on the outer surface of the device housing.

[0015] Stereo is a common recording and playback format. However, obtaining high-quality stereo recordings from the internal microphones of consumer electronics devices can be challenging. For example, the two omnidirectional microphones of a smartphone are often close together due to the relatively small size of the phone. When the signal pairs obtained from such microphones are used directly as stereo signals, the time delay between the channels is small. When the microphones are positioned such that part of the device body is located between them, there may also be some intensity differences between the channels (e.g., at higher frequencies). Due to the relatively small time delay and intensity differences, the stereo image corresponding to the directly recorded signal may be unfavorably or unacceptably narrow.

[0016] In some examples, directional microphones can be obtained by combining the signals from two closely spaced omnidirectional microphones. According to one approach, the pressure gradient is approximated by a finite difference between two pressure signals acquired at two points in space, the distance between which is less than the acoustic wavelength of interest. The pressure gradient estimated in this way represents the pressure change along a straight line connecting the two corresponding omnidirectional microphones and approximates the signal that a bidirectional microphone would pick up at that location. By adding an appropriate time delay before subtracting the pressure signal, a virtual directional microphone with a specific first-order pickup pattern (e.g., cardioid polarity) can be obtained. By appropriately selecting the time delay value for the subtraction operation, two such virtual directional microphones can be obtained, for example, microphones pointing in opposite directions. The resulting virtual microphone signals can then be used as two channels for a corresponding stereo recording.

[0017] While the above method is useful, it has certain limitations. For example, one limitation is that the approximation of the pressure gradient using finite differences is only sufficiently accurate when the acoustic wavelength is significantly larger than the distance between the microphones. Therefore, for microphone spacing of a few centimeters, this approximation still holds true at frequencies as high as approximately 1 kHz. Another limitation is that as the distance between the microphones decreases, the signal-to-noise ratio (SNR) of the corresponding pressure gradient signal also decreases; for example, because the similarity between two pressure signals increases, their difference eventually becomes overwhelmed by transducer noise.

[0018] Another challenge is that, under free-field (plane wave) conditions, pressure gradient signals have a frequency response that rolls off by approximately 6 dB per octave as the frequency decreases. Therefore, to recover a timbre comparable to the pressure signal, a 6 dB boost per octave is needed to the pressure gradient. In at least some examples, this boost can result in poor SNR at relatively low frequencies. Furthermore, when the sound source is close to the microphone, the free-field condition may not hold at all. In this case, the curvature of the sound field causes different signal amplitudes at the two pressure transducers, resulting in a low-frequency boost. This boost is known as the proximity effect, and its magnitude depends on the sound frequency, the distance between the sound source and the microphone, and the direction of incidence relative to the direction in which the pressure gradient is measured. Directional microphones are typically designed to provide a flat on-axis low-frequency response at a manufacturer-selected distance, with a bass boost at smaller distances and a bass roll-off at larger distances. Another aspect is that the phase response is also affected by the proximity effect and is related to frequency, distance, and direction. The latter's characteristics present additional challenges; for example, to achieve high-quality stereo acquisition, two microphones are required to have controllable and similar phase responses in all directions of interest (typically across the front semicircle).

[0019] The various embodiments disclosed herein advantageously address at least some of the aforementioned problems in the prior art. For example, one embodiment provides a method for generating high-quality stereo recordings based on audio signals captured by two or more closely spaced omnidirectional microphones (such as microphones commonly found in mobile phones). This method includes constructing directional virtual microphone pairs. and They point to the left and right respectively. In some examples, directional virtual microphones... and Constructed using the pressure gradient method described above, this method works well enough in the mid-frequency range, but may not be optimal for lower and higher frequency bands. This method works by first... and The analysis of the signal yields stereo cues in the lower to mid-frequency range to correct for inhomogeneities in the pressure gradient estimation. These cues are then applied to the original microphone signal at lower frequencies, providing a modified low-frequency signal. Finally, the modified low-frequency signal is appropriately combined with the mid-frequency range in the frequency domain. and The original microphone signal and high-frequency band signal are used to obtain left and right (L, R) stereo signals.

[0020] Figure 1 This is a schematic diagram showing a 3D perspective view of a consumer electronic device 100, some embodiments of which may be practiced. For illustrative purposes and without any implied limitations, exemplary embodiments are described below with reference to the illustrated device 100 as a smartphone. Based on the provided description, those skilled in the art will readily understand how to practice various additional embodiments using other types of devices without excessive experimentation.

[0021] Device 100 has a length l ,width w and thickness d In one example, the size is l =15 cm w =7.5 cm d =7mm. In other examples, device 100 may have other dimensions. Minimum external dimension (in this example is 7mm). d This can be referred to as the lateral dimension. The two larger external dimensions (in this example, ) l and w This can be referred to as the lateral dimension. In the example shown, device 100 has three omnidirectional microphones 110, 120, and 130. In other examples, other types of microphones may also be used to implement microphones 110, 120, and 130. The spacing between microphones 110 and 120 is along the X-axis. l Approximately 0.5 along the Z-axis. wThe spacing between microphones 120 and 130 is approximately 0.1 mm along the X-axis. l The value is approximately 0.3 along the Z-axis. w .

[0022] Figure 2 This is a flowchart illustrating a method 200 for generating a stereo recording according to an embodiment. Method 200 can be used to generate a high-quality stereo recording based on audio signals captured by two or more closely spaced omnidirectional microphones. For illustrative purposes and without any implied limitation, method 200 is described below with reference to device 100. Based on the provided description, those skilled in the art will readily understand how to practice additional embodiments of method 200 using other types of devices without excessive experimentation.

[0023] Method 200 includes (in box 202) selecting a microphone pair for device 100. Typically, any two of microphones 110, 120, and 130 can be selected as source microphones M1 and M2 in box 202. In various examples, microphone selection is performed based on user input or automatically (e.g., by running a programmable microphone selection script via the electronic controller of device 100). In some examples, microphones 110 and 120 can be selected as source microphones M1 and M2 in box 202 for use when device 100 is in a specific state. Figure 1 The audio capture to be performed is shown in the horizontal orientation. In this orientation, the spacing of the microphones 110 and 120 will typically produce a suitable inter-channel phase difference (IPD), which aids in spatial localization. Furthermore, at relatively high frequencies, the body of device 100 will act as an acoustic barrier, resulting in a certain inter-channel level difference (ILD), which also contributes to better spatial localization.

[0024] Method 200 also includes (in box 204) obtaining two directional virtual microphones. In some examples, the two directional virtual microphones obtained in box 204 are directional virtual microphones pointing to the left and right, respectively. and In one example, the operation of box 204 includes determining a relative time delay corresponding to a desired polarity sensitivity pattern based on the distance between the source microphones M1 and M2 selected in box 202. In some examples, the desired polarity sensitivity pattern is a cardioid polarity sensitivity pattern. The operation of box 204 also includes using audio signals captured by the source microphones M1 and M2 (which are affected by the determined relative time delay) to calculate the corresponding directional virtual microphone. and The audio signal. More specifically, it is calculated by applying a relative time delay to the M2 signal and subtracting the resulting delayed signal from the M1 signal. The signal is calculated by applying a relative time delay to the M1 signal and subtracting the resulting delayed signal from the M2 signal. In some examples, the applied relative time delay is a fractional delay, as it typically does not correspond to an integer number of signal samples. Some embodiments of box 204 may benefit from certain features disclosed in Christof Faller’s publication entitled “Conversion of Two Closely Spaced Omnidirectional Microphone Signals to an XY Stereo Signal”, presented as conference paper 8188 at the 129th General Meeting of the Audio Engineering Society in November 2010, which is incorporated herein by reference in its entirety.

[0025] In some examples, the operation of box 204 may optionally include the operation of the method obtained as described above. and Additional processing is applied to the signal to improve its audio quality. This additional processing may include: (i) replacing the corresponding frequency portions of the M1 and M2 signals, respectively, in the frequency domain. and The high-frequency portion of the signal, and / or (ii) the pair and Frequency-dependent equalization is applied to the low-frequency portion of the signal. This "replacement" operation is generally beneficial because the finite-difference approximation of the pressure gradient only applies to acoustic wavelengths larger than the distance between source microphones M1 and M2. (Refer to the above...) Figure 1 In the described example, the corresponding validity threshold frequency is approximately 1 kHz. Therefore, and Frequency components in the signal above the validity threshold can be discarded and replaced, for example, by the corresponding frequency components of signals M1 and M2, respectively. Frequency-dependent equalization is beneficial, for example, because it can be used to compensate for… and A 6 dB / octave bass roll-off is typically present in the signal. In some examples, frequency-dependent equalization starts at an effectiveness threshold frequency (e.g., 1 kHz) and proceeds down to lower frequencies until another threshold frequency is reached, at which a preselected fixed maximum boost value (e.g., 20 dB) is achieved. Then, in the frequency range below the other threshold frequency, the boost gain remains constant at that preselected fixed maximum boost value (i.e., frequency-independent).

[0026] Method 200 also includes (in box 206) obtaining the inter-channel level difference (ILD) in the mid-frequency range. As described above, a directional virtual microphone and They function as a good stereo pair in the mid-frequency range, but their quality may not be optimal in the lower frequency range, for example, due to poor SNR and the proximity effect mentioned above for sound sources relatively close to microphones M1 and M2. In the mid-frequency range, the spatial characteristics of the audio scene being recorded are determined by… and The relative level difference between signals is well represented. This difference is called ILD, expressed in decibels, and can be measured over a period of time. and The decibel difference between the energies of the signals is obtained.

[0027] Therefore, in some examples, box 206 includes the following operations. First, the signal of interest in the time domain is divided into overlapping frames. In some examples, the frame size is 1024 signal samples, and the frame overlap rate is 50%. Next, each frame is transformed to the frequency domain by a Fourier transform (e.g., Fast Fourier Transform, FFT). The frequency bins of the resulting frequency domain signal can then be grouped into a smaller number of frequency bands. In some examples, the edge frequency selection of such frequency bands can be sense-driven. For the first... n Frame, corresponding ILD (represented as) ILD(n) The frequency domain is calculated as the frequency band of interest. Signal energy and frequency domain in the same frequency band The signal energy in decibels. The corresponding mathematical expression for this calculation is as follows: (1) Where index b This refers to the frequency range. B The frequency band. In one example, the frequency range B It is located between 200 Hz and 500 Hz. ILD(n) The negative value represents Signal-to-weight ratio The signal is louder. ILD(n) Positive values ​​represent Signal-to-weight ratio The signal is louder.

[0028] In some examples, the operation of block 206 also includes processing multiple frames calculated according to formula (1). ILD(n) The values ​​are averaged. In some examples, a sliding window is used to perform the averaging. In other examples, recursive smoothing is used to perform the averaging. The corresponding mathematical expression for implementing recursive smoothing is as follows: (2) in It is a constant used to determine the effective time span of the smoothing time window.

[0029] Method 200 also includes (in block 208) applying the ILD to the M1 and M2 signals in a lower frequency range. In some examples, block 208 includes the following operations. First, the average ILD of the M1 and M2 signals in the lower frequency band (e.g., below 200 Hz) is calculated. For the... n The frame, using the same method as expressed in formula (1), calculates the corresponding average ILD, denoted as: Secondly, determine the corresponding crosschannel gain value, expressed as... As shown below: (3) Finally, the calculated or (Take the negative value) and apply it to the corresponding one of the M1 and M2 signals. The crosschannel gain calculated in this way, after application, will be applied at the boundary between the mid-frequency range and the lower-frequency range. , There is no ILD discontinuity between the signal and signals M1 and M2. As mentioned above, in one example, this boundary is located at 200 Hz. In fact, these operations in block 208 will determine the ILD discontinuity defined in block 206. Extrapolate from the mid-frequency range to the lower-frequency range.

[0030] Method 200 further includes (in block 210) constructing left and right (L, R) stereo signals. In some examples, block 210 includes the following operations. First, L, R stereo signals are constructed in the frequency domain by using different portions of various signals from blocks 202, 204, and 208 in different frequency ranges. More specifically, for the lower frequency range, the L, R stereo signals receive the corresponding ILD correction portions of the M1 and M2 signals calculated in block 208, respectively. For the mid-frequency range, the L, R stereo signals receive the ILD correction portions of the M1 and M2 signals calculated in block 204. , The corresponding portions of the signal. For the higher frequency range, the corresponding portions of the M1 and M2 signals of the L / R stereo signal receiving block 202. After filling the lower, middle, and higher frequency ranges of the L / R stereo signal in this way, an inverse Fourier transform is applied to the resulting frequency domain signal to generate the corresponding time domain version of the L / R stereo signal. In some examples, the boundary frequency between the lower and middle frequency ranges is 200 Hz. The boundary frequency between the middle and higher frequency ranges is 1 kHz.

[0031] Figure 3This is a block diagram illustrating a machine-implemented algorithm 300 according to an embodiment. In some examples, algorithm 300 is implemented in device 100. In other examples, algorithm 300 is implemented in a computing device capable of accessing microphone signals captured by device 100. In some cases, such computing device may be a server that can be connected to device 100 via a network. The inputs to algorithm 300 include multiple microphone signals 302, for example, from... Figure 1 The omnidirectional microphones 110, 120, and 130 of the illustrated device 100 capture audio signals. The output of algorithm 300 includes left and right (L, R) stereo signals 380. At least some parts of algorithm 300 can be implemented according to corresponding blocks of method 200, for example, as described in more detail below.

[0032] Block 310 of algorithm 300 performs microphone signal selection, wherein two signals 312 are selected from a plurality of microphone signals 302 as signals M1 and M2. The microphone signal selection is performed based on applicable metadata 304. In some examples, metadata 304 includes the orientation of device 100 during the capture of microphone signals 302, the microphone position on device 100, and the relevant size of device 100 (e.g., length). l ,width w and thickness d ), and the desired first-order pickup pattern, etc.

[0033] Block 314 of Algorithm 300 divides each signal 312 into a corresponding partially overlapping sequence of audio frames. M 1( n )and M 2( n In various examples, the frame size ranges from 100 to 10,000 signal samples, and the frame overlap ranges from 10% to 50%.

[0034] Block 318 of Algorithm 300 is based on the partially overlapping frame sequence generated in Block 314. M 1( n )and M 2( n Furthermore, based on the relevant parts of metadata 304, the partially overlapping frame sequences are calculated. L pg ( n )and R pg ( n In some examples, frames L pg ( n )and R pg ( n The corresponding operation of method 200 in block 204 is used in block 318 to perform the calculation.

[0035] Algorithm 300's block 322 for each frame M 1( n ), M 2( n ), L pg ( n )and R pg ( n Apply FFT to generate the corresponding discrete spectrum. M 1( f, n ), M 2( f, n ), L pg ( f, n )and R pg ( f, n Then, each discrete spectrum M 1( f, n ), M 2( f, n ), L pg ( f, n )and R pg ( f, n The spectrum is divided into sections corresponding to three frequency ranges: lower frequency, middle frequency, and higher frequency (LF, IMF, and UF) ranges. The boundary frequencies between the LF / IMF and IMF / UF ranges are determined based on metadata 304. (See reference above.) Figures 1 to 2 A specific example of the described device 100 has an LF / IMF boundary frequency of 200 Hz and an IMF / UF boundary frequency of 1 kHz.

[0036] Block 330 of Algorithm 300 on discrete spectrum L pg ( f, n ), R pg ( f, n The IMF portion 324 of the method applies equalization (EQ) filtering. In some examples, the EQ filtering implemented in block 330 uses the corresponding frequency-related equalization operation of block 204 of method 200. The resulting equalized IMF portion is... Figure 3 The figure is labeled with reference numeral 332.

[0037] Block 334 of Algorithm 300 uses the balanced IMF portion 332 calculated in block 330 to calculate the ILD value 336. In some examples, the ILD value 336 is calculated using the corresponding operation of block 206 of Method 200 according to equation (1) or equation (2).

[0038] Block 350 of Algorithm 300 will increase the cross-channel gain. Applied to discrete spectrum M 1( f, n ), M 2( f, n The LF section 326 of the () is mentioned. In some examples, the cross-channel gain is also mentioned. It is determined according to formula (3), using the corresponding operation of block 208 of method 200, and based on the ILD value 336 calculated in block 334. The resulting LF portion after cross-channel gain processing is... Figure 3 It is indicated by reference numeral 352 in the attached figure.

[0039] Block 360 of Algorithm 300 is used to construct the discrete spectrum 362 using the following: (i) discrete spectrum M 1( f, n ), M 2( f, n (i) the UF portion 328 in block 330; (ii) the equalized IMF portion 332 calculated in block 330; and (iii) the LF portion 352 calculated in block 350 after cross-channel gain processing.

[0040] Block 370 of Algorithm 300 applies an inverse FFT to each discrete spectrum 362 constructed in block 360. The output of block 370 has multiple time-domain segments 372 corresponding to the discrete spectrum 362.

[0041] The block 376 operation of algorithm 300 is used to appropriately combine the time-domain segments 372 calculated in block 370 to generate L, R stereo signals 380. Since adjacent time-domain segments 372 partially overlap, the operation of block 376 includes: appropriately truncating the overlapping portions of adjacent time-domain segments 372, and splicing the resulting truncated time-domain segments together for each L and R channel of the stereo signal.

[0042] Various additional embodiments of Algorithm 300 can be implemented by combining one or more features described in more detail below. For illustrative purposes, these features will continue to be referenced. Figure 3 Describe it.

[0043] In an additional embodiment, in block 310, three microphone signals 312 are selected from a plurality of microphone signals 302 as signals M1, M2, and M3. As described above, a first signal pair is selected from the plurality of microphone signals 302 as signals M1 and M2, and these signals are then used in block 318 to obtain a pressure gradient stereo signal in a first intermediate frequency (IMF1) range (e.g., below 1 kHz). and A third microphone signal 312 is selected as the M3 signal from multiple microphone signals 302. The selected M3 signal is paired with one of the selected M1 and M2 signals to create a second signal pair, which represents a distance between the original omnidirectional microphones (pressure transducers) that is less than the distance between the M1 and M2 signal pairs. This second signal pair is then used to calculate the pressure gradient stereo signal in a second IMF (IMF2) range (e.g., between 1 kHz and 2 kHz). and This use of the second signal pair allows the effective range of the pressure gradient approximation to be extended to higher effective threshold frequencies (e.g., 2 kHz). Subsequently, in the modified block 360, the discrete spectrum 362 is constructed by using: (i) a discrete spectrum whose spectrum lies above the IMF2 range (e.g., above 2 kHz). M 1( f, n ), M 2( f, n (ii) based on the UF part; and (iii) Calculated equalization IMF2 component; and The calculated equalization IMF1 portion, where IMF1 is the frequency range between the LF and IMF2 ranges; and (iv) the LF portion 352 calculated in block 350 after cross-channel gain processing. Then, as described above, blocks 370 and 376 are used to generate L, R stereo signals 380, but based on a discrete spectrum 362 constructed over four frequency ranges LF, IMF1, IMF2, and UF (instead of the aforementioned three frequency ranges LF, IMF, and UF).

[0044] For the cases where the distance between microphones M1 and M2 is 15 cm and the distance between microphones M1 and M3 is 4 cm, the following example numerical ranges can be used in the corresponding embodiments of algorithm 300: (i) The example frequency range used for analyzing and calculating ILD is 200 Hz–500 Hz; (ii) The example frequency range of the ILD calculated by applying extrapolation is 50 Hz–200 Hz; (iii) Example IMF1 range is 200 Hz–1 kHz; (iv) Example IMF2 range is 1 kHz–2 kHz; and (v) Example UF range is 2 kHz–20 kHz.

[0045] In another additional embodiment, once in a specific frequency band To obtain the pressure gradient stereo signal, it will be based on and The calculated ILD is applied to the same frequency band M1, M2, or those included in Within the frequency band, or applied to... In partially overlapping frequency bands. This method can, for example, be used to avoid any audible artifacts caused by filtering and signal processing used to obtain the pressure gradient, in which case the pressure gradient is used only as a guide signal for the M1 and M2 enhancements, and not as an audio component of the L, R stereo signal 380.

[0046] In various examples, the signals M1, M2, and M3 can be the raw signals from microphones 110, 120, and 130, or preprocessed versions of these signals. Preprocessing may include one or more of the following: (i) noise reduction processing designed to attenuate microphone electrical self-generated noise; (ii) applying a broadband gain configured to compensate for the varying sensitivities of the individual microphones; and (iii) applying a frequency-dependent gain configured to adjust the microphone frequency response to a desired frequency response, such as a flat frequency response.

[0047] In yet another additional embodiment, the EQ filters described in block 330 of reference algorithm 300 and block 204 of method 200 for flattening the pressure gradient signal response are obtained empirically in a dynamic manner, rather than based on a theoretical benchmark. One purpose of these filters is to equalize the pressure gradient signal with the pressure signal, compensating for the inherent low-frequency loss caused by taking finite differences. When using fixed filters, a trade-off is made, for example, because the response of the pressure gradient varies with the distance from the sound source and the angle of incidence. In some examples, this trade-off aims to compensate for far-field (plane wave) conditions and accept near-field effects essentially uncompensated. However, in some cases, a reference real pressure signal can be obtained from microphones M1 and M2 and can be used as an accurate reference regardless of the distance and direction of the sound source. For this additional embodiment, in block 206 of method 200, the levels of each frequency band of M1 and M2 are also calculated. and And using these levels, the equalization gain is calculated as follows: (4) (5) in and This is the offset gain value, used to ensure that the overall ILD of the pressure gradient is preserved while the frequency response is flattened. In some examples, the offset gain value is determined as follows: (6) (7) in It is a predetermined frequency band in which the pressure gradient approximately holds, such as a frequency band centered at 800 Hz.

[0048] For example, the gain determined according to formulas (4) to (7) and Applied to respectively and Signals, typically ensuring their frequency characteristics correspond to those of the pressure transducer, effectively handle proximity effects regardless of the source's location. In some examples, gain... and The signal is smoothed over time and frequency and is clipped to not exceed a selected fixed maximum boost amount (e.g., 20 dB). In some examples, instead of calculating the equalization gain of the left and right stereo channels based on the M1 and M2 signals separately, the equalization gain is calculated based on the average value of the M1 and M2 signals for each audio frame and / or frequency band or range.

[0049] Figures 4A to 4B The figure illustrates some performance characteristics of algorithm 300 according to an additional embodiment. Traces 401 and 402 respectively illustrate the frequency characteristics of signals M1 and M2 according to an example. The level difference between traces 401 and 402... This indicates the ILD corresponding to microphones M1 and M2. Figure 4A Traces 403 and 404 show the frequency characteristics of the pressure gradient signal calculated based on signals M1, M2, and M3. The aforementioned bass roll-off is clearly visible in the frequency characteristics shown by traces 403 and 404. The inflection point at approximately 500 Hz in traces 403 and 404 is due to the use of two different microphone pairs when obtaining the gradient signal. More specifically, signals M1 and M2 are used below 500 Hz, while signals M1 and M3 are used above 500 Hz. The level difference between traces 403 and 404... This refers to the ILD corresponding to a virtual directional microphone. Figure 4B Traces 405 and 406 show the results obtained by (… Figure 4AThe frequency characteristics of the equalized pressure gradient signal obtained by applying frequency-dependent boosting to the original pressure gradient signal shown by traces 403 and 404 are illustrated. The level difference between traces 405 and 406 represents the ILD of the equalized pressure gradient signal. Due to the characteristics of the aforementioned frequency-dependent boosting, this ILD remains... .

[0050] Figure 5 This is a block diagram of an example computing device 500 configured according to various embodiments, which is configured to perform at least some operations of method 200 and / or algorithm 300. In some embodiments, computing device 500 may be part of device 100. In other embodiments, a single computing device 500 or multiple computing devices 500 may be network-connected to device 100 to receive microphone signal 302 and metadata 304 therefrom.

[0051] Figure 5 The computing device 500 is illustrated as having multiple components, but any one or more of these components may be omitted or repeated depending on the application and setup requirements. In some embodiments, some or all of the components included in the computing device 500 may be attached to one or more motherboards and packaged in a housing. In some embodiments, some of these components may be manufactured onto a single system-on-a-chip (SoC) (e.g., the SoC may include one or more electronic processing devices 502 and one or more storage devices 504). Furthermore, in various embodiments, the computing device 500 may not include... Figure 5 The device may include one or more components, but may also include interface circuitry for coupling to one or more components using any suitable interface, such as a Universal Serial Bus (USB) interface, an High Definition Multimedia Interface (HDMI) interface, a Controller Area Network (CAN) interface, a Serial Peripheral Interface (SPI) interface, an Ethernet interface, a wireless interface, or any other suitable interface. For example, computing device 500 may not include display device 510, but may include display device interface circuitry (e.g., connectors and driver circuitry) to which external display device 510 may be coupled.

[0052] Computing device 500 includes processing device 502 (e.g., one or more processing devices). As used herein, the terms "electronic processor device" and "processing device" are interchangeable to refer to any device or part of a device that processes electronic data from registers and / or memory to convert that electronic data into other electronic data that can be stored in registers and / or memory. In various embodiments, processing device 502 may include one or more digital signal processors (DSPs), application-specific integrated circuits (ASICs), central processing units (CPUs), graphics processing units (GPUs), server processors, or any other suitable processing devices.

[0053] The computing device 500 also includes a storage device 504 (e.g., one or more storage devices). In various embodiments, the storage device 504 may include one or more memory devices, such as random access memory (RAM) devices (e.g., static RAM (SRAM), magnetic RAM (MRAM), dynamic RAM (DRAM), resistive RAM (RRAM), or conductive bridged RAM (CBRAM) devices), hard disk-based memory devices, solid-state memory devices, network drives, cloud drives, or any combination of memory devices. In some embodiments, the storage device 504 may include memory sharing a wafer with the processing device 502. In such embodiments, the memory may be used as a cache memory and may include, for example, embedded dynamic random access memory (eDRAM) or spin-transfer torque magnetic random access memory (STT-MRAM). In some embodiments, the storage device 504 may include a non-transitory computer-readable medium having instructions thereon that, when executed by one or more processing devices (e.g., processing device 502), cause the computing device 500 to perform any suitable method or portion of such methods disclosed below.

[0054] The computing device 500 also includes an interface device 506 (e.g., one or more interface devices 506). In various embodiments, the interface device 506 may include one or more communication chips, connectors, and / or other hardware and software for managing communication between the computing device 500 and other computing devices. For example, the interface device 506 may include circuitry for managing wireless communication to enable the transfer of data to and from the computing device 500. The term "wireless" and its derivatives can be used to describe circuits, devices, systems, methods, technologies, communication channels, etc., that transmit data via modulated electromagnetic radiation in a non-solid medium. This term does not imply that the associated devices do not contain any wires, although they may not in some embodiments. The circuitry included in interface device 506 for managing wireless communications can implement any of a variety of wireless standards or protocols, including but not limited to Institute of Electrical and Electronics Engineers (IEEE) standards, including Wi-Fi (IEEE 802.11 series), IEEE 802.16 standards, Long Term Evolution (LTE) projects, and any amendments, updates, and / or revisions (e.g., Advanced LTE projects, Ultra Mobile Broadband (UMB) projects (also known as “3GPP2”), etc.). In some embodiments, the circuitry included in interface device 506 for managing wireless communications can operate according to Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Universal Mobile Telecommunications System (UMTS), High Speed ​​Packet Access (HSPA), Evolved HSPA (E-HSPA), or LTE networks. In some embodiments, the circuitry included in interface device 506 for managing wireless communications can operate according to Enhanced GSM Evolved Data (EDGE), GSM EDGE Radio Access Network (GERAN), Universal Terrestrial Radio Access Network (UTRAN), or Evolved UTRAN (E-UTRAN). In some embodiments, the circuitry included in the interface device 506 for managing wireless communications may operate according to Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Digital Enhanced Cordless Communication (DECT), Evolved Data Optimization (EV-DO) and its derivatives, as well as any other wireless protocol designated as 3G, 4G, 5G and higher. In some embodiments, the interface device 506 may include one or more antennas (e.g., one or more antenna arrays) configured to receive and / or transmit wireless signals.

[0055] In some embodiments, interface device 506 may include circuitry for managing wired communications, such as electrical, optical, or any other suitable communication protocol. For example, interface device 506 may include circuitry for supporting communications based on Ethernet technology. In some embodiments, interface device 506 may support both wireless and wired communications, and / or may support multiple wired communication protocols and / or multiple wireless communication protocols. For example, a first set of circuitry for interface device 506 may be dedicated to short-range wireless communications, such as Wi-Fi or Bluetooth, and a second set of circuitry for interface device 506 may be dedicated to long-range wireless communications, such as Global Positioning System (GPS), EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, or others. In some other embodiments, a first set of circuitry for interface device 506 may be dedicated to wireless communications, and a second set of circuitry for interface device 506 may be dedicated to wired communications.

[0056] The computing device 500 also includes a battery / power circuit 508. In various embodiments, the battery / power circuit 508 may include one or more energy storage devices (e.g., batteries or capacitors) and / or circuitry for coupling components of the computing device 500 to a power source (e.g., AC line power) that is separate from the computing device 500.

[0057] The computing device 500 also includes a display device 510 (e.g., one or more separate display devices). In various embodiments, the display device 510 may include any visual indicator, such as a head-up display, computer monitor, projector, touch screen display, liquid crystal display (LCD), light-emitting diode display, or flat panel display.

[0058] The computing device 500 also includes additional input / output (I / O) devices 512. In various embodiments, I / O devices 512 may include one or more data / signal transmission interfaces, audio I / O devices (e.g., microphones or microphone arrays, speakers, headphones, earphones, alarms, etc.), audio codecs, video codecs, printers, sensors (e.g., thermocouples or other temperature sensors, humidity sensors, pressure sensors, vibration sensors, etc.), image capture devices (e.g., one or more cameras), human-machine interface devices (e.g., keyboards, cursor control devices such as mice, styluses, trackballs, or touchpads), etc.

[0059] According to specific embodiments, various components of interface device 506 and / or I / O device 512 can be configured to output suitable control signals, receive suitable control / telemetry signals, and receive and transmit data streams. In some examples, interface device 506 and / or I / O device 512 includes one or more analog-to-digital converters (ADCs) for converting received analog signals into a digital form suitable for operation by processing device 502 and / or storage device 504. In some additional examples, interface device 506 and / or I / O device 512 includes one or more digital-to-analog converters (DACs) for converting digital signals provided by processing device 502 and / or storage device 504 into an analog form suitable for transmission over a communication channel.

[0060] According to the exemplary embodiments disclosed above, for example in the summary of the invention and / or references Figures 1 to 5 A method for generating a stereo signal is provided, comprising: obtaining a first directional virtual microphone signal and a second directional virtual microphone signal based on a first audio signal captured by a first omnidirectional microphone and further based on a second audio signal captured by a second omnidirectional microphone; determining a level difference between the first directional virtual microphone signal and the second directional virtual microphone signal in a first frequency range; applying a gain to the second audio signal in a second frequency range to generate a third audio signal, the gain being selected such that the level difference between the third audio signal and the first audio signal in the second frequency range matches the determined level difference between the first directional virtual microphone signal and the second directional virtual microphone signal in the first frequency range; and generating a stereo signal using the first directional virtual microphone signal, the second directional virtual microphone signal, and the third audio signal.

[0061] In some embodiments of the above method, the same electronic device houses a first omnidirectional microphone and a second omnidirectional microphone, and at least one external lateral dimension of the electronic device is less than 20 cm.

[0062] In some embodiments of any of the above methods, the electronic device is a smartphone or a camera.

[0063] In some embodiments of any of the above methods, the second frequency range has a lower frequency than the first frequency range.

[0064] In some embodiments of any of the above methods, the generation includes: combining a spectral portion of a first directional virtual microphone signal corresponding to a first frequency range with a first spectral portion of a first audio signal corresponding to a second frequency range to generate a first spectral portion of a first stereo channel of a stereo signal; and combining a spectral portion of a second directional virtual microphone signal corresponding to a first frequency range with a first spectral portion of a third audio signal corresponding to a second frequency range to generate a first spectral portion of a second stereo channel of a stereo signal.

[0065] In some embodiments of any of the above methods, the generation further includes: combining a first spectral portion of the first stereo channel with a second spectral portion of the first audio signal corresponding to a third frequency range, wherein the third frequency range has a higher frequency than the first frequency range; and combining a first spectral portion of the second stereo channel with a second spectral portion of the second audio signal corresponding to the third frequency range.

[0066] In some embodiments of any of the above methods, the method further includes obtaining a third directional virtual microphone signal and a fourth directional virtual microphone signal based on a first audio signal and a second audio signal, and further based on a fourth audio signal captured by a third omnidirectional microphone.

[0067] In some embodiments of any of the above methods, the generation further includes: combining a first spectral portion of the first stereo channel with a spectral portion of the third directional virtual microphone signal corresponding to a third frequency range to generate a larger spectral portion of the first stereo channel, the third frequency range having frequencies higher than the first frequency range; and combining a first spectral portion of the second stereo channel with a spectral portion of the fourth directional virtual microphone signal corresponding to the third frequency range to generate a larger spectral portion of the second stereo channel.

[0068] In some embodiments of any of the above methods, the generation further includes: combining a larger spectral portion of the first stereo channel with a second spectral portion of the first audio signal corresponding to a fourth frequency range, the fourth frequency range having a higher frequency than the third frequency range; and combining a larger spectral portion of the second stereo channel with a second spectral portion of the second audio signal corresponding to the fourth frequency range.

[0069] In some embodiments of any of the above methods, the same electronic device accommodates a first omnidirectional microphone, a second omnidirectional microphone, and a third omnidirectional microphone.

[0070] In some embodiments of any of the above methods, obtaining includes: applying a time delay such that one of the first audio signal and the second audio signal is delayed relative to the other of the first audio signal and the second audio signal; and calculating the difference between the relatively delayed first audio signal and the second audio signal.

[0071] In some embodiments of any of the above methods, the method further includes selecting a time delay such that the first directional virtual microphone signal or the second directional virtual microphone signal represents the selected polarity pickup mode.

[0072] In some embodiments of any of the above methods, the difference is equalized to reduce frequency-dependent roll-off.

[0073] In some embodiments of any of the above methods, the method further includes: dividing each of the first audio signal and the second audio signal into a sequence of partially overlapping audio frames; and applying a Fourier transform to each of the audio frames to obtain a signal portion corresponding to a different frequency range.

[0074] In some embodiments of any of the above methods, the generation further includes: applying an inverse Fourier transform to each of a plurality of discrete spectra representing a stereo signal in the frequency domain to generate a first sequence and a second sequence of audio signal segments; combining the audio signal segments of the first sequence to generate a time-domain signal of a first stereo channel of the stereo signal; and combining the audio signal segments of the second sequence to generate a time-domain signal of a second stereo channel of the stereo signal.

[0075] In some embodiments of any of the above methods, the generation further includes: combining frequency domain segments corresponding to signal portions of different frequency ranges to generate a frequency domain signal of a first stereo channel of a stereo signal; combining frequency domain segments corresponding to signal portions of different frequency ranges to generate a frequency domain signal of a second stereo channel of a stereo signal; and applying an inverse Fourier transform to each of the first and second frequency domain stereo channels to generate the first and second stereo channels of the stereo signal in the time domain.

[0076] In some embodiments of any of the above methods, the first audio signal has a higher level than the second audio signal; and the applied gain has a negative value.

[0077] According to another example embodiment disclosed above, such as in the summary of the invention and / or references Figures 1 to 5 Any one or any combination of some or all of the above provides a non-transitory computer-readable medium for storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations including any of the methods described above.

[0078] According to yet another example embodiment disclosed above, such as in the summary of the invention and / or references Figures 1 to 5An apparatus is provided, comprising: at least one processor; and at least one memory including program code; wherein the at least one memory and the program code are configured, together with the at least one processor, to cause the apparatus to perform at least the following operations: obtaining a first directional virtual microphone signal and a second directional virtual microphone signal based on a first audio signal captured by a first omnidirectional microphone and further based on a second audio signal captured by a second omnidirectional microphone; determining a level difference between the first directional virtual microphone signal and the second directional virtual microphone signal in a first frequency range; applying a gain to the second audio signal in a second frequency range to generate a third audio signal, the gain being selected such that the level difference between the third audio signal and the first audio signal in the second frequency range matches the determined level difference between the first directional virtual microphone signal and the second directional virtual microphone signal in the first frequency range; and generating a stereo signal using the first directional virtual microphone signal, the second directional virtual microphone signal, and the third audio signal.

[0079] In some embodiments of the above-described device, the same electronic device includes a first omnidirectional microphone, a second omnidirectional microphone, at least one processor, and at least one memory.

[0080] In some embodiments of any of the above-described devices, the device includes electronic devices and computing devices; wherein the electronic devices include a first omnidirectional microphone and a second omnidirectional microphone; and wherein the computing devices include at least one processor and at least one memory.

[0081] In some embodiments of any of the above-described devices, at least one memory and program code are further configured, together with at least one processor, to cause the device to: obtain a third directional virtual microphone signal and a fourth directional virtual microphone signal based on one of a first audio signal and a second audio signal, and further based on a fourth audio signal captured by a third omnidirectional microphone; and additionally generate a stereo signal using the third directional virtual microphone signal and the fourth directional virtual microphone signal.

[0082] Regarding the processes, systems, methods, heuristics, etc., described herein, it should be understood that although the steps of such processes, etc., are described as occurring in a specific ordered sequence, such processes may also execute the described steps in an order different from that described herein. It should also be understood that some steps may be performed simultaneously, other steps may be added, or some steps described herein may be omitted. In other words, the process descriptions herein are for illustrative purposes and should in no way be construed as limiting the claims.

[0083] Therefore, it should be understood that the above description is intended to be illustrative rather than limiting. Many other embodiments and applications will become apparent from the above description, in addition to the examples provided. Scope should not be determined by reference to the above description, but rather by the appended claims and the full scope of their equivalents. Future developments in the techniques discussed herein are foreseeable and anticipated, and the disclosed systems and methods will be incorporated into such future embodiments. In conclusion, it should be understood that modifications and variations are possible with this application.

[0084] All terms used in the claims, unless otherwise expressly stated herein, shall be given the broadest reasonable interpretation and the ordinary meaning as understood by one of those skilled in the art. In particular, the use of singular articles such as “a,” “the,” and “the” shall be understood to refer to one or more of the said elements, unless the claims expressly define the opposite meaning.

[0085] This abstract is intended to give the reader a quick understanding of the nature of the technical disclosure. It should be understood that this abstract is not intended to interpret or limit the scope or meaning of the claims. Furthermore, as can be seen from the detailed description above, various features are combined in different embodiments to make the disclosure more concise. This manner of disclosure should not be construed as indicating that the claimed embodiments include more features than those expressly recited in each claim. Rather, as the following claims show, the inventive subject matter lies in not all the features of a single disclosed embodiment. Therefore, the following claims are hereby incorporated into the detailed description, each claim being a separate claim.

[0086] While this disclosure contains references to illustrative embodiments, this specification is not intended to be construed as limiting. Various modifications to the described embodiments, as well as other embodiments within the scope of this disclosure, will be apparent to those skilled in the art to which this disclosure pertains, and such modifications and embodiments are considered to be within the principles and scope of this disclosure, for example, as set forth in the following claims.

[0087] Some embodiments can be implemented as a circuit-based process, including possibly on a single integrated circuit.

[0088] Some embodiments may be embodied in the form of methods and apparatus for practicing these methods. Some embodiments may also be embodied in the form of program code recorded in a tangible medium, such as a magnetic recording medium, an optical recording medium, a solid-state memory, a floppy disk, a CD-ROM, a hard disk drive, or any other non-transitory machine-readable storage medium, wherein when the program code is loaded into and executed by a machine (e.g., a computer), the machine becomes an apparatus for practicing the patented invention. Some embodiments may also be embodied in the form of program code, for example, stored in a non-transitory machine-readable storage medium that includes being loaded into and / or executed by a machine, wherein when the program code is loaded into and executed by a machine (e.g., a computer or processor), the machine becomes an apparatus for practicing the patented invention. When implemented on a general-purpose processor, the program code segments are combined with the processor to provide a unique device that operates similarly to a specific logic circuit.

[0089] Unless otherwise explicitly stated, every value and range should be interpreted as an approximation, as if the words “approximately” or “about” were used before the value or range.

[0090] The use of drawing numbers and / or drawing reference numerals in the claims is intended to identify one or more possible embodiments of the claimed subject matter to facilitate the interpretation of the claims. Such use should not be construed as necessarily limiting the scope of these claims to the embodiments shown in the corresponding drawings.

[0091] Although the elements are described in a particular order and marked accordingly in the following method claims (if any), these elements are not necessarily intended to be limited to being implemented in that particular order unless the description of the claims otherwise implies that some or all of these elements are to be implemented.

[0092] The terms "an embodiment" or "a particular embodiment" as used herein refer to an embodiment in which a specific feature, structure, or characteristic described in association may be included in at least one embodiment of this disclosure. The phrase "in one embodiment," appearing multiple times in the specification, does not necessarily refer to the same embodiment, and individual or alternative embodiments are not necessarily mutually exclusive with other embodiments. The same applies to the use of the term "implementation."

[0093] Unless otherwise expressly stated in this document, the use of ordinal adjectives such as “first,” “second,” “third,” etc., to refer to one of a plurality of similar objects indicates only that the objects referred to are different instances of these similar objects, and is not intended to imply that these similar objects must be in a corresponding order or sequence, whether in time, space, ranking, or any other way.

[0094] Unless otherwise specified herein, the conjunction “if” may also be interpreted, or alternatively, as “when,” “upon,” “in response to determining,” or “in response to detecting,” depending on the specific context. For example, the phrases “if it is determined” or “if [a stated condition] is detected” can be interpreted as “upon determining,” “in response to determining,” “upon detecting [the stated condition or event],” or “in response to detecting [the stated condition or event].”

[0095] Similarly, for the purposes of this description, the terms “couple,” “coupling,” “coupled,” “connect,” “connecting,” or “connected” refer to any manner known in the art or developed hereafter, in which energy is allowed to be transferred between two or more elements and one or more additional elements may be inserted, though not required. Conversely, the terms “direct coupling,” “direct connection,” etc., imply the absence of such additional elements.

[0096] The functions of the components shown in the diagram, including any functional blocks labeled “processor” and / or “controller”, can be provided using dedicated hardware and hardware capable of executing software in conjunction with appropriate software. When provided by a processor, these functions can be provided by a single dedicated processor, a single shared processor, or multiple independent processors, some of which may be shared. Furthermore, the explicit use of the terms “processor” or “controller” should not be construed as referring only to hardware capable of executing software, but may implicitly include (but is not limited to) digital signal processor (DSP) hardware, network processors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), read-only memory (ROM) for storing software, random access memory (RAM), and non-volatile memory. Other conventional and / or custom hardware may also be included. Similarly, any switches shown in the diagram are conceptual only. Their functions can be implemented through the operation of program logic, dedicated logic, the interaction between program control and dedicated logic, or even manually, with the specific techniques chosen by the implementer based on a more specific understanding of the context.

[0097] In this application, the terms "circuit" and "circuit system" may refer to one or more or all of the following: (a) a purely hardware circuit implementation (e.g., an implementation using only analog and / or digital circuitry); (b) a combination of hardware circuitry and software, such as (if applicable): (i) a combination of analog and / or digital hardware circuitry and software / firmware, and (ii) any portion of a hardware processor (including a digital signal processor), software, and memory that work together to enable a device (such as a mobile phone or server) to perform various functions; and (c) a hardware circuit and / or processor, such as a microprocessor or a portion thereof, whose operation requires software (e.g., firmware), but which may be absent when operation is not required. The definition of "circuit system" applies to all uses of the term in this application, including in any claim. As a further example, in this application, the term "circuit system" also covers an implementation of only hardware circuitry or a processor (or multiple processors) or a portion thereof and its accompanying software and / or firmware. The term "circuit system" also covers (e.g., and as applicable to certain claim elements) baseband integrated circuits or processor integrated circuits in mobile devices, or similar integrated circuits in servers, cellular network devices, or other computing or network devices.

[0098] Those skilled in the art will understand that any block diagram herein represents a conceptual view of an illustrative circuit embodying the principles of this disclosure. Similarly, it should be understood that any flowchart, diagrammatic representation, state transition diagram, pseudocode, etc., represents various processes that can be substantially represented in a computer-readable medium and thus executed by a computer or processor, whether or not such a computer or processor is explicitly shown.

[0099] The "Brief Overview of Some Specific Embodiments" in this specification is intended to introduce some exemplary embodiments, while other embodiments will be described in the "Detailed Description" and / or with reference to one or more accompanying drawings. The "Brief Overview of Some Specific Embodiments" is not intended to identify essential elements or features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.

Claims

1. A method for generating a stereo signal, the method comprising: Based on the first audio signal captured by the first microphone and further based on the second audio signal captured by the second microphone, a first directional virtual microphone signal and a second directional virtual microphone signal are obtained; Determine the level difference between the first directional virtual microphone signal and the second directional virtual microphone signal within a first frequency range; In a second frequency range, a gain is applied to a second audio signal to generate a third audio signal, the gain being selected such that the level difference between the third audio signal and the first audio signal in the second frequency range matches the level difference between the first directional virtual microphone signal and the second directional virtual microphone signal in the first frequency range as determined. as well as The stereo signal is generated using the first directional virtual microphone signal, the second directional virtual microphone signal, and the third audio signal; The generation of the stereo signal includes: The first spectral portion of the first directional virtual microphone signal corresponding to the first frequency range is combined with the first spectral portion of the first audio signal corresponding to the second frequency range to generate the first spectral portion of the first stereo channel of the stereo signal; as well as The second directional virtual microphone signal, corresponding to the spectral portion of the first frequency range, is combined with the third audio signal, corresponding to the first spectral portion of the second frequency range, to generate the first spectral portion of the second stereo channel of the stereo signal.

2. The method of claim 1, wherein the same electronic device houses the first microphone and the second microphone, and at least one external lateral dimension of the electronic device is less than 20 cm.

3. The method according to claim 1 or 2, wherein the second frequency range has a lower frequency than the first frequency range.

4. The method according to any one of claims 1 to 3, wherein generating the stereo signal further comprises: The first spectral portion of the first stereo channel is combined with the second spectral portion of the first audio signal corresponding to a third frequency range, wherein the third frequency range has a higher frequency than the first frequency range. as well as The first spectral portion of the second stereo channel is combined with the second spectral portion of the second audio signal corresponding to the third frequency range.

5. The method of claim 4, further comprising obtaining a third directional virtual microphone signal and a fourth directional virtual microphone signal based on one of the first audio signal and the second audio signal and further based on a fourth audio signal captured by the third microphone.

6. The method of claim 5, wherein generating the stereo signal further comprises: The first spectral portion of the first stereo channel is combined with the spectral portion of the third directional virtual microphone signal corresponding to the third frequency range to generate a larger spectral portion of the first stereo channel, the third frequency range having a higher frequency than the first frequency range. as well as The first spectral portion of the second stereo channel is combined with the spectral portion of the fourth directional virtual microphone signal corresponding to the third frequency range to generate a larger spectral portion of the second stereo channel.

7. The method of claim 6, wherein generating the stereo signal further comprises: The larger spectral portion of the first stereo channel is combined with the second spectral portion of the first audio signal corresponding to a fourth frequency range, the fourth frequency range having a higher frequency than the third frequency range; as well as The larger spectral portion of the second stereo channel and the second spectral portion of the second audio signal correspond to the fourth frequency range.

8. The method according to any one of claims 5 to 7, wherein the same electronic device houses the first microphone, the second microphone, and the third microphone.

9. The method according to any one of claims 1 to 8, wherein the obtaining comprises: A time delay is applied to delay one of the first audio signal and the second audio signal relative to the other of the first audio signal and the second audio signal; as well as Calculate the difference between the first and second audio signals with relative delay.

10. The method of claim 9, further comprising selecting the time delay such that the first directional virtual microphone signal or the second directional virtual microphone signal represents the selected polarity pickup mode.

11. The method of claim 9 or 10, wherein obtaining includes equalizing the difference to reduce frequency-dependent roll-off.

12. The method according to any one of claims 1 to 11, further comprising: Each of the first audio signal and the second audio signal is divided into a sequence of partially overlapping audio frames; as well as A Fourier transform is applied to each of the audio frames to obtain the respective signal portions corresponding to different frequency ranges.

13. The method of claim 12, wherein generating the stereo signal further comprises: The frequency domain segments corresponding to different frequency ranges of the signal are combined to generate a frequency domain signal for the first stereo channel of the stereo signal. The frequency domain segments corresponding to different frequency ranges of the signal are combined to generate a frequency domain signal for the second stereo channel of the stereo signal; as well as An inverse Fourier transform is applied to each of the first frequency domain stereo channel and the second frequency domain stereo channel to generate the first stereo channel and the second stereo channel of the stereo signal in the time domain.

14. The method according to any one of claims 1 to 13, Wherein the first audio signal has a higher level than the second audio signal; and The gain applied has a negative value.

15. A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations including the method according to any one of claims 1 to 14.

16. An apparatus for generating a stereo signal, the apparatus comprising: At least one processor; as well as At least one memory containing program code; as well as The at least one memory and the program code are configured, together with the at least one processor, to cause the device to at least: Based on the first audio signal captured by the first microphone and further based on the second audio signal captured by the second microphone, a first directional virtual microphone signal and a second directional virtual microphone signal are obtained; Determine the level difference between the first directional virtual microphone signal and the second directional virtual microphone signal within a first frequency range; A gain is applied to the second audio signal in a second frequency range to generate a third audio signal, the gain being selected such that the level difference between the third audio signal and the first audio signal in the second frequency range matches the level difference between the first directional virtual microphone signal and the second directional virtual microphone signal in the first frequency range as determined. as well as The stereo signal is generated using the first directional virtual microphone signal, the second directional virtual microphone signal, and the third audio signal; In order to generate the stereo signal, the at least one memory and the program code are further configured to, together with the at least one processor, enable the device to: The first spectral portion of the first directional virtual microphone signal corresponding to the first frequency range is combined with the first spectral portion of the first audio signal corresponding to the second frequency range to generate the first spectral portion of the first stereo channel of the stereo signal; as well as The second directional virtual microphone signal, corresponding to the spectral portion of the first frequency range, is combined with the third audio signal, corresponding to the first spectral portion of the second frequency range, to generate the first spectral portion of the second stereo channel of the stereo signal.

17. The apparatus according to claim 16, wherein, The same electronic device contains the first microphone, the second microphone, the at least one processor, and the at least one memory.

18. The apparatus according to claim 16, in, The device includes electronic equipment and computing devices; The electronic device includes the first microphone and the second microphone; and The computing device includes at least one processor and at least one memory.

19. The apparatus according to any one of claims 16 to 18, wherein, The at least one memory and the program code are also configured, together with the at least one processor, to enable the device to: Based on one of the first audio signal and the second audio signal, and further based on the fourth audio signal captured by the third microphone, a third directional virtual microphone signal and a fourth directional virtual microphone signal are obtained. as well as The stereo signal is generated by additionally using the third directional virtual microphone signal and the fourth directional virtual microphone signal.