Spatial audio capture
By processing audio signals with a multi-microphone system, estimating and removing the influence of the first sound source, and utilizing microphone shadows and angular differences to improve sound source estimation, the problem of unstable sound source direction in multi-sound-source environments is solved, and more accurate spatial audio synthesis is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-29
- Publication Date
- 2026-03-31
AI Technical Summary
In multi-source or background noise environments, the estimation of sound source direction and energy ratio is unstable, leading to artifacts and inaccuracies in spatial audio synthesis, which is particularly evident on mobile devices.
A multi-microphone system is used to process audio signals to estimate the direction and energy ratio of at least two sound sources. First, the influence of the first sound source is removed, and then the estimation of the second sound source is improved at low frequencies by utilizing the physical shadow and angular difference of the microphones, ensuring the accuracy of the estimation.
It improves the stability and accuracy of sound source estimation, generates more spatially stable and accurate sound source synthesis, and enhances the user's auditory experience.
Smart Images

Figure CN115942168B_ABST
Abstract
Description
Technical Field
[0001] This application relates to apparatus and methods for spatial audio capture, and more particularly, to apparatus and methods for determining the direction of arrival and energy-based ratio of two or more identified sound sources within a sound field captured by spatial audio capture. Background Technology
[0002] Spatial audio capture is performed using microphone arrays in many modern digital devices, such as mobile devices and cameras, and is often used in conjunction with video capture. The spatial audio capture can be played back through headphones or speakers to provide the user with an experience of the audio scene captured by the microphone array.
[0003] Parametric spatial audio capture methods enable spatial audio capture using different microphone configurations and arrangements, and therefore can be used in consumer devices such as mobile phones. These methods are based on signal processing solutions that analyze the spatial audio field around the device using available information from multiple microphones. Typically, these methods perceptually analyze the microphone audio signals to determine relevant information within the frequency band. This information includes, for example, the direction of the primary sound source (or audio source or audio object) and the relationship between the source energy and the total frequency band energy. Based on this determined information, spatial audio can be reproduced, for example, using headphones or speakers. Ultimately, the user or listener can thus experience ambient audio as if they were present in the audio scene being recorded by the capture device.
[0004] The better the audio analysis and synthesis performance, the more realistic the result will be for the user or listener. Summary of the Invention
[0005] According to a first aspect, an apparatus is provided comprising components configured to perform the following operations: acquiring corresponding two or more audio signals from two or more microphones; determining a first sound source direction parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals, wherein the processing of the two or more audio signals is further configured to provide one or more modified audio signals based on the two or more audio signals; and determining at least a second sound source direction parameter in one or more frequency bands of the two or more audio signals, at least in part based on the one or more modified audio signals.
[0006] The component configured to provide one or more modified audio signals based on two or more audio signals may be further configured to: generate two or more modified audio signals by modifying the two or more audio signals based on the projection of a first sound source defined by a first sound source direction parameter; and the component configured to determine at least a second sound source direction parameter in one or more frequency bands of the two or more audio signals based at least in part on the one or more modified audio signals is configured to: determine at least a second sound source direction parameter in one or more frequency bands of the two or more audio signals by processing the modified two or more audio signals.
[0007] The component may be further configured to: determine a first sound source energy parameter in one or more frequency bands of the two or more audio signals based on the processing of the two or more audio signals; and determine at least a second sound source energy parameter based at least in part on one or more modified audio signals and the first sound source energy parameter.
[0008] The first and second sound source energy parameters can be a direct ratio to the total energy, and wherein the component configured to determine at least the second sound source energy parameter based at least in part on one or more modified audio signals is configured to: determine a temporary second sound source energy parameter direct ratio to the total energy based on analysis of one or more modified audio signals; and generate the second sound source energy parameter direct ratio to the total energy based on one of the following: selecting the minimum of the temporary second sound source energy parameter direct ratio to the total energy or the value of the first sound source energy parameter direct ratio to the total energy minus the value of 1; or multiplying the temporary second sound source energy parameter direct ratio to the total energy by the value of the first sound source energy parameter direct ratio to the total energy minus the value of 1.
[0009] The component configured to determine at least a second sound source energy parameter based at least in part on one or more modified audio signals and a first sound source energy parameter may be further configured to: further determine at least a second sound source energy parameter based on a first sound source direction parameter, such that the second sound source energy parameter is scaled relative to the difference between the first and second sound source direction parameters.
[0010] A component configured to determine a first sound source direction parameter in one or more frequency bands of two or more audio signals based on the processing of two or more audio signals can be configured to: select a first pair of two or more microphones; select a first pair of corresponding audio signals from the selected pair of two or more microphones; determine a delay that maximizes the correlation between the first pair of corresponding audio signals from the selected pair of two or more microphones; and determine a direction pair associated with the delay that maximizes the correlation between the first pair of corresponding audio signals from the selected pair of two or more microphones, wherein the first sound source direction parameter is selected from the determined direction pair.
[0011] A component configured to determine a first sound source direction parameter in one or more frequency bands of two or more audio signals based on the processing of two or more audio signals can be configured to: select the first sound source direction parameter from the determined direction pair based on further determining another delay that maximizes another correlation between another pair of corresponding audio signals from another selected pair of two or more microphones.
[0012] A component configured to determine a first sound source energy parameter in one or more frequency bands of two or more audio signals based on the processing of two or more audio signals can be configured to: determine a first sound source energy ratio corresponding to a first sound source direction parameter by normalizing the correlation of the energy of a first pair of corresponding audio signals for the frequency bands to maximize the correlation.
[0013] A component configured to provide one or more modified audio signals based on two or more audio signals can be configured to: determine a delay between a first pair of corresponding audio signals based on a determined first sound source direction parameter; align the first pair of corresponding audio signals by applying the determined delay to one of the audio signals in the first pair of corresponding audio signals; identify a common component from each of the audio signals in the first pair of corresponding audio signals; subtract the common component from each of the audio signals in the first pair of corresponding audio signals; and recover the delay to the subtracted audio signal in the corresponding audio signal to generate one or more modified audio signals.
[0014] A component configured to provide one or more modified audio signals based on two or more audio signals can be configured to: determine a delay between a first pair of corresponding audio signals based on a determined first sound source direction parameter; align the first pair of corresponding audio signals by applying the determined delay to one of the audio signals in the first pair of corresponding audio signals; identify a common component from each of the audio signals in the first pair of corresponding audio signals; subtract a modified common component from each of the audio signals in the first pair of corresponding audio signals, the modified common component being the common component multiplied by a gain value associated with the microphones of the microphone pair; and restore the delay to the audio signal in the corresponding audio signal after subtracting the component multiplied by the gain, to generate two or more modified audio signals.
[0015] A component configured to provide one or more modified audio signals based on two or more audio signals can be configured to: determine a delay between a first pair of corresponding audio signals, the corresponding audio signals being from a selected first pair of two or more microphones, based on a determined first sound source direction parameter; align the first pair of corresponding audio signals by applying the determined delay to one of the audio signals in the first pair; select additional pairs of corresponding audio signals from the selected additional pair of two or more microphones; determine an additional delay between the additional pairs of corresponding audio signals based on a determined additional sound source direction parameter; align the additional pairs of corresponding audio signals by applying the determined additional delay to one of the audio signals in the additional pairs; identify a common component from the first pair and the second pair of corresponding audio signals; subtract the common component or a modified common component from each of the audio signals in the first pair of corresponding audio signals, the modified common component being the common component multiplied by a gain value associated with the microphones associated with the first microphone pair; and restore the delay to the audio signal in the corresponding audio signal after subtracting the component multiplied by the gain, to generate two or more modified audio signals.
[0016] A component configured to obtain two or more corresponding audio signals from two or more microphones may be further configured to: select a first pair of two or more microphones to obtain two or more audio signals, and select a second pair of two or more microphones to obtain a second pair of two or more audio signals, wherein the second pair of two or more microphones is in audio shadow relative to a first sound source direction parameter, and wherein a component configured to provide one or more modified audio signals based on the two or more audio signals is configured to: provide a second pair of two or more audio signals, wherein, based on the second pair of two or more audio signals, the component is configured to determine at least a second sound source direction parameter in one or more frequency bands of the two or more audio signals, at least in part based on the one or more modified audio signals.
[0017] One or more frequency bands may be below the threshold frequency.
[0018] According to a second aspect, a method for an apparatus is provided, the method comprising: acquiring corresponding two or more audio signals from two or more microphones; determining a first sound source direction parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals, wherein the processing of the two or more audio signals is further configured to provide one or more modified audio signals based on the two or more audio signals; and determining at least a second sound source direction parameter in one or more frequency bands of the two or more audio signals, at least in part based on the one or more modified audio signals.
[0019] Providing one or more modified audio signals based on two or more audio signals may further include: generating two or more modified audio signals by modifying the two or more audio signals based on the projection of a first sound source defined by a first sound source direction parameter; and determining at least a second sound source direction parameter in one or more frequency bands of the two or more audio signals, at least in part based on the one or more modified audio signals, may include: determining at least a second sound source direction parameter in one or more frequency bands of the two or more audio signals by processing the modified two or more audio signals.
[0020] The method may further include: determining a first sound source energy parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals; and determining at least a second sound source energy parameter based at least in part on one or more modified audio signals and the first sound source energy parameter.
[0021] The first and second sound source energy parameters can be a direct ratio to the total energy, and wherein determining at least the second sound source energy parameter based at least in part on one or more modified audio signals can include: determining a temporary second sound source energy parameter direct ratio to the total energy based on analysis of one or more modified audio signals; and generating the second sound source energy parameter direct ratio to the total energy based on one of the following: selecting the minimum of the temporary second sound source energy parameter direct ratio to the total energy or the value of the first sound source energy parameter direct ratio to the total energy minus the value of 1; or multiplying the temporary second sound source energy parameter direct ratio to the total energy by the value of the first sound source energy parameter direct ratio to the total energy minus the value of 1.
[0022] Determining at least a second sound source energy parameter based at least in part on one or more modified audio signals and a first sound source energy parameter may further include: further determining at least a second sound source energy parameter based on a first sound source direction parameter such that the second sound source energy parameter is scaled relative to the difference between the first and second sound source direction parameters.
[0023] Determining a first sound source direction parameter in one or more frequency bands of two or more audio signals based on the processing of two or more audio signals may include: selecting a first pair of two or more microphones; selecting corresponding audio signals from the first pair of the selected pair of two or more microphones; determining a delay that maximizes the correlation between the corresponding audio signals from the first pair of the selected pair of two or more microphones; and determining a direction pair associated with the delay that maximizes the correlation between the corresponding audio signals from the first pair of the selected pair of two or more microphones, wherein the first sound source direction parameter is selected from the determined direction pair.
[0024] Determining a first sound source direction parameter in one or more frequency bands of two or more audio signals based on the processing of two or more audio signals may include: selecting the first sound source direction parameter from the determined direction pair based on further determining another delay that maximizes another correlation between another pair of corresponding audio signals from another selected pair of two or more microphones.
[0025] Determining a first sound source energy parameter in one or more frequency bands of two or more audio signals based on the processing of two or more audio signals may include: determining a first sound source energy ratio corresponding to a first sound source direction parameter by normalizing the correlation of the maximum energy of a first pair of corresponding audio signals for the frequency band.
[0026] Providing one or more modified audio signals based on two or more audio signals may include: determining a delay between a first pair of corresponding audio signals based on a determined first sound source direction parameter; aligning the first pair of corresponding audio signals by applying the determined delay to one of the audio signals in the first pair of corresponding audio signals; identifying a common component from each of the audio signals in the first pair of corresponding audio signals; subtracting the common component from each of the audio signals in the first pair of corresponding audio signals; and restoring the delay to the subtracted audio signal in the corresponding audio signal to generate one or more modified audio signals.
[0027] Providing one or more modified audio signals based on two or more audio signals may include: determining a delay between a first pair of corresponding audio signals based on a determined first sound source direction parameter; aligning the first pair of corresponding audio signals by applying the determined delay to one of the audio signals in the first pair of corresponding audio signals; identifying a common component from each of the audio signals in the first pair of corresponding audio signals; subtracting a modified common component from each of the audio signals in the first pair of corresponding audio signals, the modified common component being the common component multiplied by a gain value associated with the microphones of the microphone pair; and restoring the delay to the audio signal in the corresponding audio signal after subtracting the component multiplied by the gain, to generate two or more modified audio signals.
[0028] Providing one or more modified audio signals based on two or more audio signals may include: determining a delay between a first pair of corresponding audio signals, the corresponding audio signals being from a selected first pair of two or more microphones, based on a determined first source direction parameter; aligning the first pair of corresponding audio signals by applying the determined delay to one of the audio signals in the first pair; selecting additional pairs of corresponding audio signals from the selected additional pair of two or more microphones; determining an additional delay between the additional pairs of corresponding audio signals based on a determined additional source direction parameter; aligning the additional pairs of corresponding audio signals by applying the determined additional delay to one of the audio signals in the additional pairs; identifying a common component from the first pair and the second pair of corresponding audio signals; subtracting the common component or a modified common component from each of the audio signals in the first pair of corresponding audio signals, the modified common component being the common component multiplied by a gain value associated with the microphones associated with the first microphone pair; and restoring the delay to the audio signal in the corresponding audio signal after subtracting the component multiplied by the gain, to generate two or more modified audio signals.
[0029] Obtaining two or more audio signals from two or more microphones includes: selecting a first pair of two or more microphones to obtain two or more audio signals, and selecting a second pair of two or more microphones to obtain a second pair of two or more audio signals, wherein the second pair of two or more microphones is in an audio shadow relative to a first sound source direction parameter, and wherein providing one or more modified audio signals based on the two or more audio signals includes: providing a second pair of two or more audio signals, and determining at least a second sound source direction parameter in one or more frequency bands of the two or more audio signals, based at least in part on the one or more modified audio signals, according to the second pair of two or more audio signals.
[0030] One or more frequency bands may be below the threshold frequency.
[0031] According to a third aspect, an apparatus is provided comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured, together with the at least one processor, to cause the apparatus to at least: acquire corresponding two or more audio signals from two or more microphones; determine a first sound source direction parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals, wherein the processing of the two or more audio signals is further configured to provide one or more modified audio signals based on the two or more audio signals; and determine at least a second sound source direction parameter in one or more frequency bands of the two or more audio signals, at least in part based on the one or more modified audio signals.
[0032] The means of providing one or more modified audio signals based on two or more audio signals can be further made to: modify the two or more audio signals based on the projection of a first sound source defined by a first sound source direction parameter to generate two or more modified audio signals; and the means of determining at least a second sound source direction parameter in one or more frequency bands of the two or more audio signals based at least in part on the one or more modified audio signals can be made to: determine at least a second sound source direction parameter in one or more frequency bands of the two or more audio signals by processing the two or more modified audio signals.
[0033] The apparatus can be further configured to: determine a first sound source energy parameter in one or more frequency bands of the two or more audio signals based on the processing of the two or more audio signals; and determine at least a second sound source energy parameter based at least in part on one or more modified audio signals and the first sound source energy parameter.
[0034] The first and second sound source energy parameters can be a direct ratio to the total energy, and wherein the means for determining at least the second sound source energy parameter based at least in part on one or more modified audio signals can be made to: determine a temporary second sound source energy parameter direct ratio to the total energy based on analysis of one or more modified audio signals; and generate the second sound source energy parameter direct ratio to the total energy based on one of the following: selecting the minimum of the temporary second sound source energy parameter direct ratio to the total energy or the value of the first sound source energy parameter direct ratio to the total energy minus the value of 1; or multiplying the temporary second sound source energy parameter direct ratio to the total energy by the value of the first sound source energy parameter direct ratio to the total energy minus the value of 1.
[0035] The means for determining at least a second sound source energy parameter based at least in part on one or more modified audio signals and a first sound source energy parameter may be further configured to: further determine at least a second sound source energy parameter based on a first sound source direction parameter such that the second sound source energy parameter is scaled relative to the difference between the first and second sound source direction parameters.
[0036] An apparatus for determining a first sound source direction parameter in one or more frequency bands of two or more audio signals based on the processing of two or more audio signals may be configured to: select a first pair of two or more microphones; select corresponding audio signals of the first pair from the selected pair of two or more microphones; determine a delay that maximizes the correlation between the corresponding audio signals of the first pair from the selected pair of two or more microphones; and determine a direction pair associated with the delay that maximizes the correlation between the corresponding audio signals of the first pair from the selected pair of two or more microphones, wherein the first sound source direction parameter is selected from the determined direction pair.
[0037] An apparatus that enables the determination of a first sound source direction parameter in one or more frequency bands of two or more audio signals based on the processing of two or more audio signals can be made to: select the first sound source direction parameter from the determined direction pair based on further determining another delay that maximizes another correlation between another pair of corresponding audio signals from another selected pair of two or more microphones.
[0038] An apparatus that enables the determination of a first sound source energy parameter in one or more frequency bands of two or more audio signals based on the processing of two or more audio signals can be made to: determine a first sound source energy ratio corresponding to a first sound source direction parameter by normalizing the correlation of the energy of a first pair of corresponding audio signals for the frequency bands to maximize the energy.
[0039] An apparatus that enables the provision of one or more modified audio signals based on two or more audio signals may be configured to: determine a delay between a first pair of corresponding audio signals based on a determined first sound source direction parameter; align the first pair of corresponding audio signals by applying the determined delay to one of the audio signals in the first pair of corresponding audio signals; identify a common component from each of the audio signals in the first pair of corresponding audio signals; subtract the common component from each of the audio signals in the first pair of corresponding audio signals; and recover the delay to the subtracted audio signal in the corresponding audio signal to generate one or more modified audio signals.
[0040] An apparatus that enables the provision of one or more modified audio signals based on two or more audio signals may be configured to: determine a delay between a first pair of corresponding audio signals based on a determined first sound source direction parameter; align the first pair of corresponding audio signals by applying the determined delay to one of the audio signals in the first pair of corresponding audio signals; identify a common component from each of the audio signals in the first pair of corresponding audio signals; subtract a modified common component from each of the audio signals in the first pair of corresponding audio signals, the modified common component being the common component multiplied by a gain value associated with a microphone in the microphone pair; and restore the delay to the audio signal in the corresponding audio signal after subtracting the component multiplied by the gain, to generate two or more modified audio signals.
[0041] An apparatus that enables the provision of one or more modified audio signals based on two or more audio signals may be configured to: determine a delay between a first pair of corresponding audio signals, the corresponding audio signals being from a selected first pair of two or more microphones, based on a determined first sound source direction parameter; align the first pair of corresponding audio signals by applying the determined delay to one of the audio signals in the first pair; select additional pairs of corresponding audio signals from the selected additional pair of two or more microphones; determine an additional delay between the additional pairs of corresponding audio signals based on a determined additional sound source direction parameter; align the additional pairs of corresponding audio signals by applying the determined additional delay to one of the audio signals in the additional pairs; identify a common component from the first pair and the second pair of corresponding audio signals; subtract the common component or a modified common component from each of the audio signals in the first pair of corresponding audio signals, the modified common component being the common component multiplied by a gain value associated with the microphones associated with the first microphone pair; and restore the delay to the audio signal in the corresponding audio signal after subtracting the component multiplied by the gain, to generate two or more modified audio signals.
[0042] The means of obtaining two or more audio signals from two or more microphones may be further made to: select a first pair of two or more microphones to obtain two or more audio signals, and select a second pair of two or more microphones to obtain a second pair of two or more audio signals, wherein the second pair of two or more microphones is in audio shadow relative to a first sound source direction parameter, and wherein the means of providing one or more modified audio signals based on the two or more audio signals is made to: provide a second pair of two or more audio signals, according to which the means is made to determine at least a second sound source direction parameter in one or more frequency bands of the two or more audio signals, at least in part based on the one or more modified audio signals.
[0043] One or more frequency bands may be below the threshold frequency.
[0044] According to a fourth aspect, an apparatus is provided, comprising components for performing the following operations: acquiring corresponding two or more audio signals from two or more microphones; determining a first sound source direction parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals, wherein the processing of the two or more audio signals is further configured to provide one or more modified audio signals based on the two or more audio signals; and determining at least a second sound source direction parameter in one or more frequency bands of the two or more audio signals, at least in part based on the one or more modified audio signals.
[0045] According to a fifth aspect, a computer program [or a computer-readable medium including program instructions] is provided, the instructions being configured to cause a device to perform at least the following operations: acquiring corresponding two or more audio signals from two or more microphones; determining a first sound source direction parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals, wherein the processing of the two or more audio signals is further configured to provide one or more modified audio signals based on the two or more audio signals; and determining at least a second sound source direction parameter in one or more frequency bands of the two or more audio signals, at least in part based on the one or more modified audio signals.
[0046] According to a sixth aspect, a non-transitory computer-readable medium is provided, comprising program instructions for causing a device to perform at least the following operations: acquiring corresponding two or more audio signals from two or more microphones; determining a first sound source direction parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals, wherein the processing of the two or more audio signals is further configured to provide one or more modified audio signals based on the two or more audio signals; and determining at least a second sound source direction parameter in one or more frequency bands of the two or more audio signals, at least in part based on the one or more modified audio signals.
[0047] According to a seventh aspect, an apparatus is provided, comprising: an acquisition circuit configured to acquire corresponding two or more audio signals from two or more microphones; a determination circuit configured to determine a first sound source direction parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals, wherein the processing of the two or more audio signals is further configured to provide one or more modified audio signals based on the two or more audio signals; and components for determining at least a second sound source direction parameter in one or more frequency bands of the two or more audio signals, at least in part based on the one or more modified audio signals.
[0048] According to an eighth aspect, a computer-readable medium is provided, comprising program instructions for causing a device to perform at least the following operations: acquiring corresponding two or more audio signals from two or more microphones; determining a first sound source direction parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals, wherein the processing of the two or more audio signals is further configured to provide one or more modified audio signals based on the two or more audio signals; and determining at least a second sound source direction parameter in one or more frequency bands of the two or more audio signals, at least in part based on the one or more modified audio signals.
[0049] An apparatus comprising components for performing the actions described above.
[0050] An apparatus configured to perform the actions described above.
[0051] A computer program comprising program instructions for causing a computer to perform the methods described above.
[0052] A computer program product stored on a medium can enable a device to perform the methods described herein.
[0053] An electronic device may include the apparatus described herein.
[0054] A chipset may include devices as described herein.
[0055] The embodiments of this application are intended to solve problems associated with the prior art. Attached Figure Description
[0056] To better understand this application, reference will now be made to the accompanying drawings by way of example, wherein:
[0057] Figure 1 This example illustrates sound source direction estimation when two equally loud sound sources are present.
[0058] Figure 2Example apparatus suitable for implementing some embodiments is illustrated schematically;
[0059] Figure 3 The following are examples of some embodiments. Figure 2 A flowchart illustrating the operation of the apparatus shown;
[0060] Figure 4 Another example apparatus suitable for implementing some embodiments is shown schematically;
[0061] Figure 5 The following are examples of some embodiments. Figure 4 A flowchart illustrating the operation of the apparatus shown;
[0062] Figure 6 Schematic illustration of some embodiments as shown in Figure 2 or Figure 4 The example spatial analyzer shown;
[0063] Figure 7 The following are examples of some embodiments. Figure 6 A flowchart illustrating the operation of the example spatial analyzer is shown below;
[0064] Figure 8 This illustrates an example where three microphones are used to estimate the direction of arrival of a sound source;
[0065] Figure 9 An example is shown for a set of estimated directions for simultaneous noise input from two directions for one frequency band;
[0066] Figure 10 Examples of sound source direction estimation based on estimation when two equally loud sound sources are present are shown according to some embodiments;
[0067] Figure 11 This illustrates the example microphone arrangement or configuration within the example device when operating in landscape mode;
[0068] Figure 12 Schematic illustration of some embodiments as shown in Figure 2 or Figure 4 The example spatial synthesizer shown;
[0069] Figure 13 Example apparatus suitable for implementing some embodiments is schematically shown; and
[0070] Figure 14 An example device suitable for implementing the illustrated apparatus is shown schematically. Detailed Implementation
[0071] The concepts discussed in further detail in this article with respect to the following embodiments relate to the capture of audio scenes.
[0072] In the following description, the term "sound source" is used to describe an element (artificial or real) defined within a sound field (or audio scene). The term "sound source" can also be defined as an audio object or an audio source, and these terms are interchangeable in understanding the implementation of the examples described herein.
[0073] The embodiments described herein relate to parametric audio capture apparatus and methods, such as spatial audio capture (SPAC) techniques. For each time-frequency tile, the apparatus is configured to estimate the direction of the primary sound source and the relative energy of the direct component of the sound source and the ambient component, which is expressed as a direct-to-total energy ratio.
[0074] The following examples are suitable for devices with challenging microphone arrangements or configurations (such as those found in typical mobile devices), where the size of the mobile device typically includes at least one short (or thin) dimension relative to other dimensions. In the examples shown herein, the captured spatial audio signal is a suitable input to a spatial synthesizer to generate spatial audio signals, such as binaural format audio signals for headphone listening or multichannel signal format audio signals for speaker listening.
[0075] In some embodiments, these examples can be implemented as part of a spatial capture front-end for IVAS standard codecs by generating audio signals and metadata compatible with Immersive Speech and Audio Services (IVAS).
[0076] Typical spatial analysis involves estimating the primary sound source directions and the direct-to-total energy ratio for each time-frequency tile. These parameters are driven by the human auditory system and are, in principle, based on similar characteristics. However, in certain given conditions, it is known that such a model does not provide optimal sound quality.
[0077] Typically, parameter estimation can be problematic when multiple simultaneous sound sources are present, or alternatively, when these sources are almost completely masked by background noise. In the first case, the direction of the analyzed primary sound source can jump between actual source directions, or depend on how the sounds from the sources superimpose, and the analysis may even end up as an average of the source directions. In the second case, a primary sound source may be found sometimes and not sometimes, depending on the instantaneous levels of the sound sources and the environment. Besides variations in direction values, the estimated energy ratios can be unstable in both of these cases.
[0078] In this context, direction and energy ratio analysis can lead to artifacts in the synthesized audio signal. For example, the direction of the sound source may sound unstable or inaccurate, and background audio may become reverberant.
[0079] As an example, such as Figure 1 As shown, an example direction estimation of the primary sound source is illustrated in a case where two equally loud sound sources are located at azimuth angles of 30 degrees and -20 degrees around the acquisition device. Figure 1 As shown, depending on the point in time, either of them can be identified as the primary sound source; therefore, both sound sources are synthesized by the spatial synthesizer into the estimated direction. Because the estimated direction jumps continuously between the two values, the result will be ambiguous, and the user or listener will have difficulty detecting which direction the two sound sources are coming from. Furthermore, this continuous jump in the estimated direction from one direction to another produces a synthesized sound field that sounds unsettling and unnatural.
[0080] Several techniques have been proposed to improve upon the aforementioned issues, including increasing the amount of available information. For example, it has been proposed to estimate parameters for each time-frequency tile in the two primary directions. For instance, the currently developing 3GPP IV AS standard plans to support two simultaneous directions.
[0081] However, for parametric audio coding with typical mobile device microphone settings, there is no reliable method for estimating the directions of the two primary sound sources. Furthermore, in cases of unreliable estimation, sound sources may be synthesized into directions where there are no actual sound sources, and / or the sound source locations may jump / move continuously from one location to another in an unstable manner. In other words, estimating more than one direction in cases of unreliable estimation offers no benefit and may worsen the spatial audio signal generated by the spatial synthesizer.
[0082] Therefore, in summary, the embodiments described herein relate to parametric spatial audio capture using two or more microphones. Furthermore, based on the audio signals from the two or more microphones, at least two direction and energy ratio parameters are estimated in each time-frequency tile.
[0083] In these embodiments, the influence of the first estimated direction is taken into account when estimating the second direction in order to improve the accuracy of multi-source direction detection. In some embodiments, this can lead to an improvement in the perceptual quality of the synthesized spatial audio.
[0084] In practice, the embodiments described herein produce estimates of sound sources that are perceived as more spatially stable and accurate (relative to their correct or actual locations).
[0085] In some embodiments, the first direction and energy ratio are estimated (and can be estimated) using any suitable estimation method. Furthermore, when estimating the second direction, the influence of the first direction is first removed from the microphone signal. In some embodiments, this can be achieved by first removing any delay between the signals based on the first direction and then subtracting the common component from the two signals. Finally, the original delay is recovered. The second direction parameters can then be estimated using a method similar to that used to estimate the first direction.
[0086] In some embodiments, different microphone pairs are used to estimate two different directions at low frequencies. This emphasizes the natural shadowing of sounds derived from the physical shape of the device and increases the likelihood of finding sound sources on different sides of the device.
[0087] In some embodiments, the energy ratio in the second direction is first analyzed using a method similar to that used to estimate the energy ratio in the first direction. Furthermore, in some embodiments, the second energy ratio is further modified based on the energy ratio in the first direction and the angular difference between the first and second estimated sound source directions.
[0088] about Figure 2 The diagram illustrates a device suitable for implementing the embodiments described herein.
[0089] In this example, the apparatus shown includes a microphone array 201. The microphone array 201 includes multiple (two or more) microphones configured to capture audio signals. The microphones within the microphone array can be of any suitable microphone type, arrangement, or configuration. The microphone audio signal 202 generated by the microphone array 201 can be passed to a spatial analyzer 203.
[0090] The device may include a spatial analyzer 203, which is configured to receive or otherwise acquire a microphone audio signal 202 and is configured to perform spatial analysis on the microphone audio signal to identify at least two primary sound or audio sources for each time-frequency block.
[0091] In some embodiments, the spatial analyzer may be a mobile device or the CPU of a computer. The spatial analyzer 203 is configured to generate a data stream 204 that includes audio signals and metadata of the spatial information being analyzed.
[0092] Depending on the use case, the data stream can be stored or compressed and transmitted to another location.
[0093] Furthermore, the device includes a spatial synthesizer 205. The spatial synthesizer 205 is configured to acquire a data stream including audio signals and metadata. In some embodiments, the spatial synthesizer 205 is implemented within the same device as the spatial analyzer 203 (e.g., Figure 2 (as shown), but in some embodiments it can also be implemented in different devices or equipment.
[0094] Spatial synthesizer 205 can be implemented within a CPU or similar processor. Spatial synthesizer 205 is configured to generate an output audio signal 206 based on the audio signal from data stream 204 and associated metadata.
[0095] Furthermore, depending on the usage, the output signal 206 can be any suitable output format. For example, in some embodiments, the output format is a binaural headphone signal (where the output device presenting the output audio signal is a set of headphones / earbuds or the like) or a multi-channel speaker audio signal (where the output device is a set of speakers). The output device 207 (as described above, which can be, for example, headphones or speakers) can be configured to receive the output audio signal 206 and present the output to a listener or user.
[0096] Figure 2 These operations of the example device shown can be performed by Figure 3 The flowchart shown is illustrated. Therefore, the operation of the example device can be summarized as follows.
[0097] Obtain the microphone audio signal, such as Figure 3 As shown in step 301.
[0098] Spatial analysis is performed on the microphone audio signal to generate spatial audio signals and metadata, which includes the direction and energy ratio for the first and second audio sources for each time-frequency tile, such as... Figure 3 As shown in step 303.
[0099] Spatial synthesis is applied to spatial audio signals to generate suitable output audio signals, such as... Figure 3 As shown in step 305.
[0100] Output the audio signal to the output device, such as Figure 3 As shown in step 307.
[0101] In some embodiments, spatial analysis can be used in conjunction with an IVAS codec. In this example, the spatial analysis output is an IVAS-compatible MASA (Metadata-Assisted Spatial Audio) format, which can be directly fed to the IVAS encoder. The IVAS encoder generates an IVAS data stream. At the receiving end, the IVAS decoder is able to directly produce the desired output audio format. In other words, in this embodiment, there is no separate spatial synthesis block.
[0102] This is for example about Figure 4 The device shown and Figure 5 The operation of the device is illustrated in the flowchart shown in the figure.
[0103] exist Figure 4 In the example shown, the device also includes a microphone array 201. The microphone array 201 is configured to generate a microphone audio signal 202, which is transmitted to a spatial analyzer 203.
[0104] The spatial analyzer 203 is configured to receive or otherwise acquire the microphone audio signal 202 and identify at least two primary sound or audio sources for each time-frequency block. The data stream generated by the spatial analyzer 203, namely the MASA format data stream (which includes the audio signal and metadata of the analyzed spatial information) 404, is then passed to the IVAS encoder 405.
[0105] The device may further include an IVAS encoder 405 configured to accept a MASA format data stream 404 and generate an IVAS data stream 406 that can be transmitted or stored, as shown by dashed line 416.
[0106] In addition, the device includes an IVAS decoder 407 (spatial synthesizer). The IVAS decoder 407 is configured to decode the IVAS data stream and also spatially synthesize the determined audio signal to generate an output audio signal 206 for the appropriate output device 207.
[0107] Output device 207 (as described above, which may be, for example, headphones or speakers) can be configured to receive output audio signal 206 and present the output to the listener or user.
[0108] Figure 4 These operations of the example device shown can be performed by Figure 5 The flowchart shown is used to illustrate this. Therefore, the operation of the example device can be summarized as follows.
[0109] Obtain the microphone audio signal, such as Figure 5 As shown in step 301.
[0110] Spatial analysis of the microphone audio signal is performed to generate MASA format output (spatial audio signal and metadata, which includes the direction and energy ratio for the first and second audio sources for each time-frequency tile), such as... Figure 5 As shown in step 503.
[0111] IVAS encodes the generated data stream, such as Figure 5 As shown in step 505.
[0112] The encoded IVAS data stream is decoded (and spatial synthesis is applied to the decoded spatial audio signal) to generate a suitable output audio signal, such as... Figure 5 Step 507 is shown in the diagram.
[0113] Output the audio signal to the output device, such as Figure 5 As shown in step 307.
[0114] In some embodiments, instead, the output audio signal is an ambisonic signal. In this embodiment, there may be no immediate direct output device.
[0115] refer to Figure 6 More details are shown in Figure 2 and Figure 4 The spatial analyzer is shown by reference numeral 203 in the attached figure.
[0116] In some embodiments, the spatial analyzer 203 includes a streaming audio signal generator 607. The streaming audio signal generator 607 is configured to receive microphone audio signals 202 and generate one or more streaming audio signals 608 to be passed to the multiplexer 609. The audio streaming signal is generated from the input microphone audio signals based on any suitable method. For example, in some embodiments, one or two microphone signals may be selected from the microphone audio signals 202. Alternatively, in some embodiments, the microphone audio signals 202 may be downsampled and / or compressed to generate the streaming audio signal 608.
[0117] In the following examples, spatial analysis is performed in the frequency domain; however, it should be understood that in some embodiments, analysis may also be performed in the time domain using a temporal sampled version of the microphone audio signal.
[0118] In some embodiments, the spatial analyzer 203 includes a time-frequency converter 601. The time-frequency converter 601 is configured to receive microphone audio signals 202 and convert them to the frequency domain. In some embodiments, the time-domain microphone audio signals may be represented as s before the conversion. i(t), where t is the time index and i is the microphone channel index. The transformation to the frequency domain can be achieved using any suitable time-frequency transform, such as STFT (Short Time Fourier Transform) or QMF (Quadrature Mirror Filter). The resulting time-frequency domain microphone signal 602 is represented as S i (b,n), where i is the microphone channel index, b is the frequency bin index, and n is the time frame index. The value of b is in the range of 0,...,B–1, where B is the number of bin indices at each time index n.
[0119] Frequency modules can be further combined into subbands k = 0, ..., K–1. Each subband consists of one or more frequency modules. Each subband k has a minimum module b. k,low and the highest warehouse b k,high The width of the subband is usually chosen based on the characteristics of human hearing; for example, equivalent rectangular bandwidth (ERB) or Bark scale can be used.
[0120] In some embodiments, the spatial analyzer 1203 includes a first direction analyzer 3603. The first direction analyzer 3603 is configured to receive a time-frequency domain microphone audio signal 602 and generate estimates of a first direction 614 and a first ratio 616 for each time-frequency tile for the first sound source.
[0121] The first direction analyzer 603 is configured to generate an estimate of the first direction based on any suitable method, such as SPAC (as described in more detail in US9313599).
[0122] In some embodiments, for example, by searching for a time shift τ that maximizes the correlation between the two (microphone audio signal) channels for subband k. k To estimate the most dominant direction used for time frame indexing. i (b,n) can be shifted by τ samples, as follows:
[0123]
[0124] Then, find the delay τ for each subband k. k It maximizes the correlation between the two microphone channels:
[0125]
[0126] In the formula above, the "optimal" delay is searched between microphones 1 and 2. Re indicates the real part of the result, and * is the complex conjugate of the signal. The delay search range parameter D is defined based on the distance between the microphones. max In other words, considering the distance between microphones and the speed of sound, τ is searched only within the physically possible range.k The value of .
[0127] The angle in the first direction can be defined as...
[0128]
[0129] As shown in the figure, the sign of the angle remains uncertain.
[0130] The directional analysis between microphones 1 and 2 has been defined above. A similar process can then be repeated between other microphone pairs to resolve ambiguities (and / or obtain the direction of another axis). In other words, information from other analysis pairs can be used to eliminate ambiguities. The ambiguity of symbols in [the text].
[0131] For example, Figure 8 The diagram illustrates a configuration where a microphone array comprises three microphones, with a first microphone 801, a second microphone 805, and a third microphone 803 arranged such that the first microphone pair (first microphone 801 and third microphone 803) is separated by a distance on a first axis, and the second microphone pair (first microphone 801 and second microphone 805) is separated by a distance on a second axis (in this example, the first axis is perpendicular to the second axis). Furthermore, in this example, the three microphones may be located on the same third axis, defined as perpendicular to both the first and second axes (and perpendicular to the plane of the paper on which the diagram is printed). Analysis of the delay between the first microphone pair 801 and 803 yields two alternative angles α807 and -α809. Analysis of the delay between the second microphone pair 801 and 805 can then be used to determine which alternative angle is correct. In some embodiments, the information required for this analysis is whether sound arrives first at microphone 801 or microphone 805. If sound arrives at microphone 805, angle α is correct. If not, -α is chosen.
[0132] Furthermore, based on inferences between several microphone pairs, the first spatial analyzer can determine or estimate the correct orientation angle.
[0133] In some embodiments with limited microphone configurations or arrangements, such as only two microphones, directional ambiguity cannot be resolved. In such embodiments, the spatial analyzer is configured to define that all sources are always in front of the device. The same is true when there are more than two microphones; however, their positions do not allow for, for example, front-to-back analysis.
[0134] Although not disclosed in this paper, multiple pairs of microphones on the vertical axis can determine elevation and azimuth estimates.
[0135] The first-direction analyzer 603 can also use, for example, a normalized correlation value c(k,n) to determine or estimate the energy ratio r1(k,n) corresponding to the angle θ1(k,n), for example:
[0136]
[0137] The value of r1(k,n) is between -1 and 1, and is usually further restricted to between 0 and 1.
[0138] In some embodiments, the first direction analyzer 603 is configured to generate a modified time-frequency microphone audio signal 604. The modified time-frequency microphone audio signal 604 is a signal from which a first sound source component has been removed.
[0139] Therefore, for example, regarding the first microphone pair (such as...) Figure 8 (Microphones 801 and 803 are shown in the example microphone configuration). For subband k, the delay that provides the highest relevance is τ. k For each sub-band k, the second microphone signal is shifted by τ. k One sample was used to obtain the shifted second microphone signal.
[0140] The source components can be estimated as the average of these time-aligned signals:
[0141]
[0142] In some embodiments, any other suitable method may be used to determine the sound source components.
[0143] The estimated C(b,n) of the sound source components has already been determined (for example, in the example formula above), and can then be removed from the microphone audio signal. On the other hand, other simultaneous sound sources are not in phase, causing them to be attenuated in C(b,n). Now, C(b,n) can be subtracted from the (shifted and unshifted) microphone signals:
[0144]
[0145]
[0146] In addition, the microphone audio signal after shifting modification Shifted backward by τ k 1 sample to obtain:
[0147]
[0148] Then, these modified signals and It can be passed to the second direction analyzer 305.
[0149] In some embodiments, the spatial analyzer 203 includes a second direction analyzer 605. The second direction analyzer 605 is configured to receive a time-frequency microphone audio signal 602 estimate, a modified time-frequency microphone audio signal 604 estimate, a first direction 614 estimate, and a first ratio 616 estimate, and generate a second direction 624 estimate and a second ratio 626 estimate.
[0150] The estimation of the second direction parameter value can adopt the same sub-band structure as the first direction estimation and follow similar operations as previously described for the first direction estimation.
[0151] Therefore, the second direction parameters θ2(k,n) and r′2(k,n) can be estimated. In this embodiment, a modified time-frequency microphone audio signal 604 is used. and Instead of using the time-frequency microphone audio signals 602 S1(b,n) and S2(b,n) to determine the direction estimate.
[0152] Furthermore, in some embodiments, the energy ratio r′2(k,n) is limited because the sum of the first ratio and the second ratio should not exceed 1.
[0153] In some embodiments, the second ratio is subject to the following limitations:
[0154] r2(k,n)=(1-r1(k,n))r′2(k,n)
[0155] or
[0156] r2(k,n)=min(r′2(k,n),1-r1(k,n))
[0157] The `min` function selects the smaller of the provided alternatives. Both alternatives have been found to provide good quality ratio values.
[0158] Note that in the example above, because there are several microphone pairs, the modified signal must be calculated separately for each pair; that is, when considering microphone pairs 801 and 805 or microphone pairs 801 and 803, They are not the same signal.
[0159] The first direction estimate 614, the first ratio estimate 616, the second direction estimate 624, and the second ratio estimate 626 are passed to a multiplexer (mux) 609, which is configured to generate a data stream 204 / 404 by combining these estimates with the streaming audio signal 608.
[0160] about Figure 7 This shows a summary Figure 6 The flowchart shows an example operation of the spatial analyzer.
[0161] Obtain the microphone audio signal, such as Figure 7 Step 701 is shown in the diagram.
[0162] Then, a streaming audio signal is generated from the microphone audio signal, such as... Figure 7 Step 702 is shown in the diagram.
[0163] It can also perform time-frequency domain transformation on microphone audio signals, such as... Figure 7 Step 703 is shown in the diagram.
[0164] Then, the first direction parameter estimate and the first ratio parameter estimate can be determined, such as Figure 7 Step 705 is shown in the diagram.
[0165] Then, the time-frequency domain microphone audio signal can be modified (to remove the first sound source component), such as... Figure 7 Step 707 is shown in the diagram.
[0166] Then, the modified time-frequency domain microphone audio signal is analyzed to determine the second direction parameter estimate and the second ratio parameter estimate, such as... Figure 7 Step 709 is shown in the diagram.
[0167] Then, the first direction parameter estimate, the first ratio parameter estimate, the second direction parameter estimate, the second ratio parameter estimate, and the streaming audio signal are multiplexed to generate a data stream (which can be a MASA format data stream), such as... Figure 7 Step 711 is shown in the diagram.
[0168] therefore, Figure 9 This example illustrates the results of directional analysis for a single subband. The input consists of two uncorrelated noise signals arriving simultaneously from two directions, where the signal arriving from the first direction is 1 dB louder than the signal arriving from the second direction. In most cases, the stronger source is found to be from the first direction, but occasionally the second source is also found to be from the first direction. If only one direction is estimated, the direction estimate will therefore jump between the two values, which could potentially lead to quality problems. In the case of two-directional analysis, both sound sources are included in either the first or second direction, and the quality of the synthesized signal remains consistently good.
[0169] Figure 10 For example, it is shown in the context of... Figure 1 (Where only one direction estimate is estimated for each time-frequency tile) The direction estimation results for the same case are shown. For comparison, the same case with two direction estimates better maintains the sound source in its position.
[0170] In some embodiments, other methods may be used to determine the common component C(b,n) (the first source component). For example, in some embodiments, principal component analysis (PCA) or other related methods may be used. In some embodiments, individual gains for different channels are applied when generating or subtracting the common component. Thus, for example, in some embodiments...
[0171]
[0172] and
[0173]
[0174]
[0175] In this embodiment, common components can be removed from the microphone signal while taking into account, for example, the different levels of the audio signal in the microphone.
[0176] Furthermore, although in the example above two microphone signals are used to generate the common component (combined signal) C(b,n), in some embodiments more microphones can be used. For example, with three microphones available, the “optimal” delay between microphone pairs 801 and 803 and 801 and 805 can be estimated. This delay is denoted as τ. k (1,2) and τ k (1,3). In this embodiment, the combined signal can be obtained as follows:
[0177]
[0178] As described above, the combined signal can then be removed from all three microphone signals before analyzing the second direction.
[0179] In the examples above, the methods used to estimate the two directions generally yielded good results. However, microphone positions in typical mobile device microphone configurations can be used to further improve the estimation and, in some examples, enhance the reliability of the second-direction analysis, especially at the lowest frequencies.
[0180] For example, Figure 11This illustrates a typical microphone configuration in a modern mobile device. The device has a display 1109 and a camera housing 1107. Microphones 1101 and 1105 are very close to each other, while microphone 1103 is located further away. The physical shape of the device affects the audio signals captured by the microphones. Microphone 1105 is located on the side of the device's main camera. Sound arriving from the display side of the device must travel around the edge of the device to reach microphone 1105. Due to this longer path, the signal is attenuated, and this attenuation depends on frequencies up to 6–10 dB. On the other hand, microphone 1101 is located at the edge of the device, so sound from the left side of the device can reach the microphone directly, while sound from the right side must travel around only one corner. Therefore, even though microphones 1101 and 1105 are close to each other, the signals they capture may be completely different.
[0181] The difference between the two microphone signals can be utilized in directional analysis. Using the equations given above, the optimal delay τ between the microphones in microphone pairs 1-2 (microphone labels 1101 and 1103) and 3-2 (microphone labels 1105 and 1103) can be estimated. k (1,2) and τ k (3,2), and the corresponding angle can be estimated. and Because the distance between the microphone pairs is different, this must be taken into account when calculating the angle.
[0182] In particular, if and If they clearly point in different directions, that is, they have found different main sound sources, then these two directions can be directly used as two direction estimates.
[0183]
[0184]
[0185] The energy ratio can be calculated similarly to that described above, and the value of r2(k,n) needs to be constrained again based on the value of r1(k,n). The ambiguity of the sign in the value can be resolved in a similar way to the above; in other words, the directional ambiguity can be resolved by using the microphone to pair 1-3.
[0186] These embodiments have been found to be particularly useful in the lowest frequency bands, where estimation in both directions is most challenging for typical microphone configurations.
[0187] In the above embodiments, the limitation of the energy ratio r2(k,n) in the second direction based on the value of the first energy ratio r1(k,n) has been discussed. In some embodiments, the angle difference between the first direction estimate and the second direction estimate is used to modify one or more ratios.
[0188] Therefore, in some embodiments, if θ1(k,n) and θ2(k,n) point in the same direction, the energy ratio parameter in the first direction already contains sufficient energy, and no energy needs to be allocated to the given second direction; that is, r2(k,n) can be set to zero. Conversely, when θ1(k,n) and θ2(k,n) point in opposite directions, the effect of the ratio r2(k,n) is most significant, and the value of r2(k,n) should be maintained to the maximum extent possible.
[0189] This can be implemented in some embodiments, where β(k,n) is the absolute angle between θ1(k,n) and θ2(k,n):
[0190] β(k,n)=θ1(k,n)-θ2(k,n)
[0191] Furthermore, the value of β(k,n) is contained between -π and π:
[0192] If β(k,n)>π, then β(k,n)=β(k,n)-2π
[0193] If β(k,n)<-π, then β(k,n)=β(k,n)+2π
[0194] The total effect of the energy ratio of the first direction to the second direction can then be calculated as follows:
[0195]
[0196] or
[0197]
[0198] Where r′2(k,n) is the initial ratio, and r2(k,n) is the modified ratio. In this example, the angle difference has a linear effect on the scaling of r2(k,n). In some embodiments, other weighting options exist, such as sine weighting.
[0199] about Figure 12 The following are examples: Figure 2 and Figure 4 The example shown is spatial synthesizer 205 or IVAS decoder 407.
[0200] In some embodiments, the spatial synthesizer 205 / IVAS decoder 407 includes a demultiplexer 1201. In some embodiments, the demultiplexer 1201 receives a data stream 204 / 404 and splits the data stream into a streaming audio signal 1208 and spatial parameter estimates, such as a first direction estimate 1214, a first ratio estimate 1216, a second direction estimate 1224, and a second ratio estimate 1226. In some embodiments where the data stream is encoded (e.g., using an IVAS encoder), the data stream can be decoded therein.
[0201] These are then passed to the space processor / synthesizer 1203.
[0202] The spatial synthesizer 205 / IVAS decoder 407 includes a spatial processor / synthesizer 1203 and is configured to receive these estimated and streamed audio signals and render the output audio signal. The spatial processing / synthesis can be any suitable two-way based synthesis, such as that described in EP3791605.
[0203] Figure 13 A schematic diagram illustrating an example implementation according to some embodiments is shown. The apparatus is a capture / playback device 1301, which includes a microphone array 201 assembly, a spatial analyzer 203 assembly, and a spatial synthesizer 205 assembly. Furthermore, device 1301 includes a storage device (memory) 1201 configured to store audio signals and metadata (data stream) 204.
[0204] In some embodiments, the capture / playback device 1301 may be a mobile device.
[0205] about Figure 14 This illustrates an example electronic device that can be used as a computer, encoder processor, decoder processor, or any functional block described herein. The device can be any suitable electronic device or apparatus. For example, in some embodiments, device 1600 is a mobile device, user equipment, tablet computer, computer, audio playback device, etc.
[0206] In some embodiments, device 1600 includes at least one processor or central processing unit 1607. Processor 1607 may be configured to execute various program codes, such as the methods described herein.
[0207] In some embodiments, device 1600 includes memory 1611. In some embodiments, at least one processor 1607 is coupled to memory 1611. Memory 1611 can be any suitable storage device. In some embodiments, memory 1611 includes a program code portion for storing program code that can be implemented on processor 1607. Furthermore, in some embodiments, memory 1611 may further include a storage data portion for storing data, such as data that has been processed or will be processed according to the embodiments described herein. The implemented program code stored in the program code portion and the data stored in the storage data portion can be retrieved by processor 1607 via memory-processor coupling when needed.
[0208] In some embodiments, device 1600 includes a user interface 1605. In some embodiments, user interface 1605 may be coupled to processor 1607. In some embodiments, processor 1607 may control the operation of user interface 1605 and receive input from user interface 1605. In some embodiments, user interface 1605 enables a user to input commands to device 1600, for example, via a keypad. In some embodiments, user interface 1605 enables a user to obtain information from device 1600. For example, user interface 1605 may include a display configured to show information from device 1600 to a user. In some embodiments, user interface 1605 may include a touchscreen or touch interface capable of allowing information to be input to device 1600 and further displaying the information to a user of device 1600.
[0209] In some embodiments, device 1600 includes an input / output port 1609. In some embodiments, input / output port 1609 includes a transceiver. In this embodiment, the transceiver may be coupled to processor 1607 and configured to communicate with other devices or electronic devices, for example, via a wireless communication network. In some embodiments, the transceiver or any suitable transceiver or transmitter and / or receiver device may be configured to communicate with other electronic devices or devices via wired or wired coupling.
[0210] The transceiver can communicate with other devices using any suitable known communication protocol. For example, in some embodiments, the transceiver may use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a Wireless Local Area Network (WLAN) protocol (such as IEEE 802.X), a suitable short-range radio frequency communication protocol (such as Bluetooth), or an Infrared Data Communication Path (IRDA).
[0211] The transceiver input / output port 1609 can be configured to send / receive audio signals, bit streams, and in some embodiments, the operations and methods described above are performed by using a processor 1607 that executes appropriate code.
[0212] Generally, various embodiments of the present invention can be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while others may be implemented in firmware or software executable by a controller, microprocessor, or other computing device, although the invention is not limited thereto. Although various aspects of the invention may be illustrated and described as block diagrams, flowcharts, or represented using certain other graphical representations, it is well understood that such blocks, apparatuses, systems, techniques, or methods described herein may be implemented (by way of non-limiting example) in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0213] Embodiments of the present invention can be implemented by computer software executable by a data processor of a mobile device (such as in a processor entity), or by hardware, or by a combination of software and hardware. Further, in this respect, it should be noted that any block in the logical flow of the figures may represent a program step, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on a physical medium such as a memory chip, or a memory block implemented within a processor, magnetic media, or optical media.
[0214] The memory can be of any type suitable for the local technical environment and can be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic storage devices and systems, optical storage devices and systems, fixed memory, and removable memory. The data processor can be of any type suitable for the local technical environment and can include one or more of the following as non-limiting examples: general-purpose computers, special-purpose computers, microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), gate-level circuits, and processors based on multi-core processor architectures.
[0215] Embodiments of the present invention can be practiced in a variety of elements, such as integrated circuit modules. The design of integrated circuits is largely a highly automated process. Sophisticated and powerful software tools can be used to transform logic-level designs into semiconductor circuit designs ready to be etched and formed on semiconductor substrates.
[0216] Programs (such as those provided by Synopsys in Mountain View, California, and Cadence Design in San Jose, California) use well-established design rules and pre-stored libraries of design modules to automatically route conductors and position components on semiconductor chips. Once the semiconductor circuit design is complete, the final design in a standardized electronic format (e.g., Opus, GDSII, etc.) can be sent to a semiconductor manufacturing facility or “fab” for fabrication.
[0217] The foregoing description has provided a complete and informative description of exemplary embodiments of the invention by way of exemplary and non-limiting examples. However, various modifications and adaptations will become apparent to those skilled in the art when read in conjunction with the accompanying drawings and appended claims, given the foregoing description. Nevertheless, all such modifications and similar alterations to the teachings of the invention will still fall within the scope of the invention as defined by the appended claims.
Claims
1. An apparatus for spatial audio capture, comprising: at least one processor; and at least one memory including computer program code; the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to: obtain respective two or more audio signals from two or more microphones; determine, based on processing of the two or more audio signals, a first sound source direction parameter in one or more frequency bands of the two or more audio signals, wherein the processing of the two or more audio signals is further configured to provide one or more modified audio signals based on the two or more audio signals, wherein a first sound source associated with the determined first sound source direction parameter is removed from the provided one or more modified audio signals; determine, based at least in part on the provided one or more modified audio signals, at least a second sound source direction parameter in the one or more frequency bands of the two or more audio signals; and generate a data stream based at least on the first sound source direction parameter, the second sound source direction parameter, and the two or more audio signals.
2. The apparatus of claim 1, wherein, providing the one or more modified audio signals causes the apparatus to: modify the two or more audio signals based on a projection with the first sound source defined by the first sound source direction parameter to generate modified two or more audio signals; and determining, based at least in part on the provided one or more modified audio signals, at least a second sound source direction parameter in the one or more frequency bands of the two or more audio signals causes the apparatus to determine at least the second sound source direction parameter in the one or more frequency bands of the two or more audio signals by processing the modified two or more audio signals.
3. The apparatus according to claim 1, further caused to: determine, based on the processing of the two or more audio signals, a first sound source energy parameter in the one or more frequency bands of the two or more audio signals; and determine, based at least in part on the one or more modified audio signals and the first sound source energy parameter, at least a second sound source energy parameter.
4. The apparatus of claim 3, wherein, the first sound source energy parameter and the second sound source energy parameter are a direct-to-total energy ratio, and wherein determining at least the second sound source energy parameter causes the apparatus to: determine, based on an analysis of the one or more modified audio signals, a temporary second sound source energy parameter direct-to-total energy ratio; and generate the second sound source energy parameter based on one of: selecting a minimum of the temporary second sound source energy parameter direct-to-total energy ratio or a value of the first sound source energy parameter subtracted from a value of 1; or multiplying the temporary second sound source energy parameter direct-to-total energy ratio by a value of the first sound source energy parameter subtracted from a value of 1.
5. The apparatus of claim 3, wherein, Determining at least a second sound source energy parameter causes the apparatus to determine at least the second sound source energy parameter further based on the first sound source direction parameter such that the second sound source energy parameter is scaled relative to a difference between the first sound source direction parameter and the second sound source direction parameter.
6. The apparatus of claim 1, wherein, Determining a first sound source direction parameter causes the apparatus to: select a first pair of the two or more microphones; select a first pair of respective audio signals from the selected pair of the two or more microphones; determine a delay that maximizes a correlation between the first pair of respective audio signals from the selected pair of the two or more microphones; and select the first sound source direction parameter from a determined pair of directions associated with the delay that maximizes the correlation between the first pair of respective audio signals from the selected pair of the two or more microphones.
7. The apparatus of claim 6, wherein, Determining a first sound source direction parameter based on processing of the two or more audio signals causes the apparatus to select the first sound source direction parameter from the determined pair of directions based on further determining a further delay that maximizes a further correlation between a further pair of respective audio signals from a further selected pair of the two or more microphones.
8. The apparatus of claim 6, wherein, Determining a first sound source energy parameter based on the processing of the two or more audio signals causes the apparatus to determine a first sound source energy ratio corresponding to the first sound source direction parameter by normalizing the correlation relative to a maximization of an energy of the first pair of respective audio signals for the frequency band.
9. The apparatus of claim 1, wherein, Providing one or more modified audio signals causes the apparatus to: determine a delay between the first pair of respective audio signals based on the determined first sound source direction parameter; align the first pair of respective audio signals based on applying the determined delay to one of the first pair of respective audio signals; identify a common component from each of the first pair of respective audio signals; subtract the common component from each of the first pair of respective audio signals; and restore the delay to the subtracted component audio signal of the first pair of respective audio signals to generate one or more modified audio signals.
10. The apparatus of claim 1, wherein, Providing one or more modified audio signals causes the apparatus to: determine a delay between the first pair of respective audio signals based on the determined first sound source direction parameter; align the first pair of respective audio signals based on applying the determined delay to one of the first pair of respective audio signals; identify a common component from each of the first pair of respective audio signals; subtract a modified common component from each of the first pair of respective audio signals, the modified common component being the common component multiplied by a gain value associated with a microphone associated with the microphone pair; and restore the delay to the subtracted component audio signal of the first pair of respective audio signals to generate the modified two or more audio signals.
11. The apparatus of claim 1, wherein, Providing one or more modified audio signals causes the apparatus to: determining, based on the determined first sound source direction parameter, a delay between a first pair of respective audio signals from a selected first pair of the two or more microphones; aligning the first pair of respective audio signals based on applying the determined delay to one of the first pair of respective audio signals; selecting an additional pair of respective audio signals from a selected additional pair of the two or more microphones; determining, based on the determined additional sound source direction parameter, an additional delay between the additional pair of respective audio signals; aligning the additional pair of respective audio signals based on applying the determined additional delay to one of the additional pair of respective audio signals; identifying a common component from the first pair of respective audio signals and the additional pair of respective audio signals; subtracting the common component or a modified common component from each of the first pair of respective audio signals, the modified common component being the common component multiplied by a gain value associated with a microphone associated with the first pair; and restoring the delay to the audio signals of the first pair of respective audio signals from which the component multiplied by the gain was subtracted to generate the modified two or more audio signals.
12. The apparatus of claim 1, wherein, obtaining two or more audio signals causes the apparatus to: select a first pair of the two or more microphones to obtain the two or more audio signals and select a second pair of the two or more microphones to obtain a second pair of two or more audio signals, wherein the second pair of the two or more microphones is in audio shadow with respect to the first sound source direction parameter, and wherein providing one or more modified audio signals causes the apparatus to provide the second pair of two or more audio signals based on a determined at least second sound source direction parameter associated at least in part with the one or more modified audio signals.
13. The apparatus of claim 12, wherein, the one or more frequency bands are below a threshold frequency.
14. A method of an apparatus for spatial audio capture, the method comprising: obtaining respective two or more audio signals from two or more microphones; determining, based on processing of the two or more audio signals, a first sound source direction parameter in one or more frequency bands of the two or more audio signals, wherein the processing of the two or more audio signals is further configured to provide one or more modified audio signals based on the two or more audio signals, wherein a first sound source associated with the determined first sound source direction parameter is removed from the provided one or more modified audio signals; determining, based at least in part on the provided one or more modified audio signals, at least a second sound source direction parameter in the one or more frequency bands of the two or more audio signals; and generating a data stream based at least on the first sound source direction parameter, the second sound source direction parameter, and the two or more audio signals.
15. The method of claim 14, wherein, providing one or more modified audio signals based on the two or more audio signals further comprises: modify the two or more audio signals based on a projection with the first sound source defined by the first sound source direction parameter, generating modified two or more audio signals; and determining at least a second sound source direction parameter in the one or more frequency bands of the two or more audio signals based at least in part on the one or more modified audio signals comprises determining at least the second sound source direction parameter in the one or more frequency bands of the two or more audio signals by processing the modified two or more audio signals.
16. The method of claim 14, wherein, The method further comprises: determining a first sound source energy parameter in the one or more frequency bands of the two or more audio signals based on the processing of the two or more audio signals; and determining at least a second sound source energy parameter based at least in part on the one or more modified audio signals and the first sound source energy parameter.
17. The method of claim 16, wherein, The first sound source energy parameter and the second sound source energy parameter are direct-to-total energy ratios, and wherein determining at least a second sound source energy parameter based at least in part on the one or more modified audio signals comprises: determining a provisional second sound source energy parameter direct-to-total energy ratio based on an analysis of the one or more modified audio signals; and generating the second sound source energy parameter based on one of: selecting a minimum of: the provisional second sound source energy parameter direct-to-total energy ratio, or a value of one minus the first sound source energy parameter; or multiplying the provisional second sound source energy parameter direct-to-total energy ratio by a value of one minus the first sound source energy parameter.
18. The method of claim 16, wherein, determining at least the second sound source energy parameter based at least in part on the one or more modified audio signals and the first sound source energy parameter further comprises determining at least the second sound source energy parameter further based on the first sound source direction parameter such that the second sound source energy parameter is scaled relative to a difference between the first sound source direction parameter and the second sound source direction parameter.
19. The method of claim 14, wherein, determining a first sound source direction parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals comprises: selecting a first pair of the two or more microphones; selecting a first pair of respective audio signals from the selected first pair of the two or more microphones; determining a delay that maximizes a correlation between the first pair of respective audio signals from the selected first pair of the two or more microphones; and determining a direction pair associated with the delay that maximizes the correlation between the first pair of respective audio signals from the selected first pair of the two or more microphones, the first sound source direction parameter being selected from the determined direction pair.
20. The method of claim 19, wherein, determining a first sound source direction parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals comprises selecting the first sound source direction parameter from the determined pair of directions based on further determining a further delay that maximizes a further correlation between another pair of respective audio signals from a selected another pair of the two or more microphones.
Citation Information
Patent Citations
An apparatus, method and computer program for audio signal processing
EP3791605A1
Apparatus and method for multi-channel signal playback
US9313599B2
An Apparatus, Method and Computer Program for Audio Signal Processing
US20210076130A1