Spatial Audio Capture

The method improves spatial audio capture by processing audio signals from multiple microphones to estimate and modify sound source directions and energy ratios, addressing inaccuracies in existing technologies and enhancing audio synthesis stability and accuracy.

JP7708730B2Active Publication Date: 2025-07-15NOKIA TECHNOLOGIES OY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022159375
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-10-04
Filing Date
2022-10-03
Publication Date
2025-07-15
Estimated Expiration
2042-10-03

AI Technical Summary

Technical Problem

Existing spatial audio capture technologies face challenges in accurately estimating multiple sound source directions and energy ratios, particularly when sources are masked by noise or when multiple sources are present, leading to unstable and unnatural synthesized audio.

Method used

A method and apparatus that utilizes a microphone array to process audio signals from multiple microphones, estimating a first sound source direction and energy ratio, and then modifies these signals to estimate a second direction and ratio, considering the influence of the first direction, to improve accuracy and stability.

Benefits of technology

Enhances the detection accuracy of multiple sound source directions, resulting in more stable and accurate spatial audio synthesis, reducing distortions and maintaining audio quality even in challenging conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007708730000027
    Figure 0007708730000027
  • Figure 0007708730000028
    Figure 0007708730000028
  • Figure 0007708730000029
    Figure 0007708730000029
Patent Text Reader

Abstract

To solve a problem related with prior arts.SOLUTION: An apparatus comprises means configured to: obtain two or more audio signals from respective two or more microphones; determine, in one or more frequency bands of the two or more audio signals, a first sound source direction parameter based on processing of the two or more audio signals, in which the processing of the two or more audio signals is further configured to provide one or more modified audio signal based on the two or more audio signals; and determine, in the one or more frequency bands of the two or more audio signals, at least a second sound source direction parameter at least based on at least in part the one or more modified audio signal.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to an apparatus and method for spatial audio capture, and more particularly to an apparatus and method for determining a ratio based on the arrival directions and energies of two or more specified sources in a sound field captured by spatial audio capture.

Background Art

[0002] Spatial audio capture using a microphone array is used in many up-to-date digital devices such as mobile terminals and cameras, and is often used in combination with video capture. Spatial audio can enable a user to experience an audio scene captured by a microphone array by playing it using headphones or a loudspeaker.

[0003] Parametric spatial audio capture methods can be adopted for consumer devices such as mobile terminals in order to enable spatial audio capture with various microphone configurations and arrangements. Parametric spatial audio capture methods are based on a signal processing solution for analyzing a spatial audio field around a device using information available from a plurality of microphones. Generally, these methods perceptually analyze the audio signals of the microphones and determine relevant information in the frequency band. This information includes, for example, the direction of a dominant sound source (or audio source or audio object), and the relationship between the sound source energy and the overall band energy. Based on this determined information, spatial audio can be played, for example, using headphones or a loudspeaker. Ultimately, a user or listener can experience environmental audio as if they were present in the audio scene recorded by the capture device.

[0004] The higher the performance of audio analysis and synthesis, the more realistic the results experienced by the user or listener.

Summary of the Invention

Problems to be Solved by the Invention

[0005] Embodiments of the present invention aim to solve problems related to the prior art.

Means for Solving the Problems

[0006] According to a first aspect, obtaining two or more audio signals from each of two or more microphones, and determining a first sound source direction parameter based on the processing of the two or more audio signals in one or more frequency bands of the two or more audio signals, wherein the processing of the two or more audio signals is further configured to provide one or more modified audio signals based on the two or more audio signals, determining, and in one or more frequency bands of the two or more audio signals, determining at least a second sound source direction parameter based at least in part on the one or more modified audio signals, and there is provided an apparatus including means configured to perform the above.

[0007] The means configured to provide one or more modified audio signals based on the two or more audio signals further includes generating the modified two or more audio signals based on modifying the two or more audio signals using the projection of the first sound source defined by the first sound source direction parameter, and in one or more frequency bands of the two or more audio signals, the means configured to determine at least the second sound source direction parameter based at least in part on the one or more modified audio signals may be configured to determine at least the second sound source direction parameter in one or more frequency bands of the two or more audio signals by processing the modified two or more audio signals.

[0008] This means may be further configured to determine a first sound source energy parameter based on processing of two or more audio signals in one or more frequency bands of the two or more audio signals, and to determine at least a second sound source energy parameter based at least in part on one or more modified audio signals and the first sound source energy parameter.

[0009] The first and second sound source energy parameters may be a direct-to-total energy ratio. The means for determining at least the second sound source energy parameter based at least in part on one or more modified audio signals is configured to determine an intermediate second sound source energy parameter direct-to-total energy ratio based on analysis of the one or more modified audio signals, and to generate the second sound source energy parameter direct-to-total energy ratio based on either selecting the minimum of the intermediate second sound source energy parameter direct-to-total energy ratio or the value obtained by subtracting the first sound source energy parameter direct-to-total energy ratio from a value of 1, or multiplying the intermediate second sound source energy parameter direct-to-total energy ratio by the value obtained by subtracting the first sound source energy parameter direct-to-total energy ratio from a value of 1.

[0010] The means configured to determine at least the second sound source energy parameter based at least in part on one or more modified audio signals and the first sound source energy parameter may be further configured to determine at least the second sound source energy parameter further based on a first sound source direction parameter such that the second sound source energy parameter is scaled relative to the difference between the first sound source direction parameter and a second sound source direction parameter.

[0011] Means configured to determine a first sound source direction parameter based on processing of two or more audio signals in one or more frequency bands of the two or more audio signals may be configured to select a first pair of two or more microphones, select a first pair of respective audio signals from the selected pair of two or more microphones, determine a delay that maximizes the correlation between the first pairs of respective audio signals from the selected pair of two or more microphones, and determine a pair of directions related to the delay that maximizes the correlation between the first pairs of respective audio signals from the selected pair of two or more microphones, wherein the first sound source direction parameter is determined to be selected from the determined pair of directions.

[0012] Means configured to determine a first sound source direction parameter based on processing of two or more audio signals in one or more frequency bands of the two or more audio signals may be configured to select from the pair of directions determined for the first sound source direction parameter based on determination of a further delay that maximizes a further correlation between further pairs of respective audio signals from a further selected pair of two or more microphones.

[0013] Means configured to determine a first sound source energy parameter based on processing of two or more audio signals in one or more frequency bands of the two or more audio signals may be configured to determine a first sound source energy ratio corresponding to the first sound source direction parameter by normalizing a maximized correlation with respect to the energy of each audio signal of the first pair for the frequency band.

[0014] Means configured to provide one or more modified audio signals based on two or more audio signals may determine a delay between a first pair of each audio signal based on a determined first sound source direction parameter, align the first pair of each audio signal based on application of the determined delay to one of the first pair of each audio signal, identify a common component from each of the first pair of each audio signal, subtract the common component from each of the first pair of each audio signal, and restore the delay to one of the subtracted components of each audio signal to generate one or more modified audio signals.

[0015] Means configured to provide one or more modified audio signals based on two or more audio signals may determine a delay between a first pair of each audio signal based on a determined first sound source direction parameter, align the first pair of each audio signal based on applying the determined delay to one of the first pair of each audio signal, identify a common component from each of the first pair of each audio signal, subtract from each of the first pair of each audio signal a modified common component, where the modified common component is associated with a pair of microphones and is the common component multiplied by a gain value associated with the microphones, and restore the delay to one of the subtracted gain-multiplied components of each audio signal to generate two or more modified audio signals.

[0016] Means configured to provide one or more modified audio signals based on two or more audio signals determine a first source direction parameter, based on each audio signal from a selected first pair of two or more microphones, determine a delay between the first pair of each audio signal; align the first pair of each audio signal based on applying the determined delay to one of the first pair of each audio signal; select an additional pair of each audio signal from a selected additional pair of two or more microphones; determine an additional delay between the additional pair of each audio signal based on the determined additional source direction parameter; align the additional pair of each audio signal based on applying the determined additional delay to one of the additional pair of each audio signal; identify a common component from the first and second pairs of each audio signal; subtract the common component or a modified common component from each of the first pair of each audio signal, where the modified common component is the common component multiplied by a gain value associated with the microphone associated with the first pair of microphones; restore the delay to one of the subtracted gain-multiplied components of each audio signal and generate two or more modified audio signals.

[0017] Means configured to obtain two or more audio signals from each of two or more microphones are further configured to select a first pair of two or more microphones to obtain two or more audio signals, and select a second pair of two or more microphones to obtain a second pair of two or more audio signals, wherein the second pair of two or more microphones is in an audio shadow with respect to a first sound source direction parameter, and means configured to provide one or more modified audio signals based on the two or more audio signals is configured to provide a second pair of the two or more audio signals based at least in part on at least one of the one or more modified audio signals in one or more frequency bands of the two or more audio signals, and means configured to determine at least a second sound source direction parameter based at least in part on at least one of the one or more modified audio signals in one or more frequency bands of the two or more audio signals.

[0018] One or more frequency bands may be lower than a threshold frequency.

[0019] According to a second aspect, a method for an apparatus is provided, the method comprising obtaining two or more audio signals from each of two or more microphones, and determining a first sound source direction parameter based on processing of the two or more audio signals in one or more frequency bands of the two or more audio signals, wherein processing the two or more audio signals is further configured to provide one or more modified audio signals based on the two or more audio signals, and determining at least a second sound source direction parameter based at least in part on at least one of the one or more modified audio signals in one or more frequency bands of the two or more audio signals.

[0020] Providing one or more modified audio signals based on two or more audio signals further includes generating the modified two or more audio signals based on modifying the two or more audio signals by the projection of a first sound source defined by a first sound source direction parameter, and determining at least a second sound source direction parameter at least partially based on the one or more modified audio signals in one or more frequency bands of the two or more audio signals. Determining at least the second sound source direction parameter may include processing the two or more modified audio signals in one or more frequency bands of the two or more audio signals.

[0021] The method may further include determining a first sound source energy parameter based on processing the two or more audio signals in one or more frequency bands of the two or more audio signals, and determining at least a second sound source energy parameter at least partially based on the one or more modified audio signals and the first sound source energy parameter.

[0022] The first and second sound source energy parameters may be a direct-to-total energy ratio. Determining at least the second sound source energy parameter at least partially based on the one or more modified audio signals includes determining an intermediate second sound source energy parameter direct-to-total energy ratio based on an analysis of the one or more modified audio signals, and selecting the minimum of the value obtained by subtracting the intermediate second sound source energy parameter direct-to-total energy ratio, or the value obtained by subtracting the first sound source energy parameter direct-to-total energy ratio from a value of 1, or multiplying the intermediate second sound source energy parameter direct-to-total energy ratio by the value obtained by subtracting the first sound source energy parameter direct-to-total energy ratio from a value of 1, and generating the second sound source energy parameter direct-to-total energy ratio based on one of the above. Generating the second sound source energy parameter direct-to-total energy ratio may include generating the second sound source energy parameter direct-to-total energy ratio based on one of the above.

[0023] Determining at least a second sound source energy parameter based at least in part on one or more modified audio signals and a first sound source energy parameter may further include determining at least the second sound source energy parameter based on the first sound source direction parameter such that the second sound source energy parameter is scaled relative to the difference between the first sound source direction parameter and the second sound source direction parameter.

[0024] Determining a first sound source direction parameter based on processing of two or more audio signals in one or more frequency bands of the two or more audio signals may include selecting a first pair of two or more microphones, selecting a first pair of respective audio signals from the selected pair of two or more microphones, determining a delay that maximizes the correlation between the first pair of respective audio signals from the selected pair of two or more microphones, and determining a pair of directions related to the delay that maximizes the correlation between the first pair of respective audio signals from the selected pair of two or more microphones, wherein the first sound source direction parameter is selected from the pair of directions in which the first sound source direction parameter is determined.

[0025] Determining a first sound source direction parameter based on processing of two or more audio signals in one or more frequency bands of the two or more audio signals may include selecting the first sound source direction parameter from the pair of directions in which the first sound source direction parameter is determined based on a further determination of a further delay that maximizes a further correlation between a further pair of respective audio signals from a selected further pair of two or more microphones.

[0026] Determining a first sound source energy parameter based on processing of two or more audio signals in one or more frequency bands of the two or more audio signals may include determining a first sound source energy ratio corresponding to the first sound source direction parameter by normalizing a maximized correlation with respect to the energy of the first pair of respective audio signals for the frequency band.

[0027] Providing one or more modified audio signals based on two or more audio signals may include determining a delay between a first pair of each audio signal based on a determined first sound source direction parameter; aligning the first pair of each audio signal based on the application of the determined delay to one of the first pair of each audio signal; identifying a common component from each of the first pair of each audio signal; subtracting the common component from each of the first pair of each audio signal; and restoring a delay to one of the subtracted components of each audio signal to generate one or more modified audio signals.

[0028] Providing one or more modified audio signals based on two or more audio signals may include determining a delay between a first pair of each audio signal based on a determined first sound source direction parameter; aligning the first pair of each audio signal based on the application of the determined delay to one of the first pair of each audio signal; identifying a common component from each of the first pair of each audio signal; subtracting a modified common component from each of the first pair of each audio signal, wherein the modified common component is the common component multiplied by a gain value associated with a microphone associated with the pair of microphones; and restoring a delay to one of the subtracted gain-multiplied components of each audio signal to generate two or more modified audio signals.

[0029] Providing one or more modified audio signals based on two or more audio signals includes determining a delay between a first pair of each audio signal based on a determined first sound source direction parameter, where each audio signal is from a selected first pair of two or more microphones; aligning the first pair of each audio signal based on the application of the determined delay to one of the first pair of each audio signal; selecting an additional pair of each audio signal from a selected additional pair of two or more microphones; determining an additional delay between the additional pair of each audio signal based on a determined additional sound source direction parameter; aligning the additional pair of each audio signal based on the application of the determined additional delay to one of the additional pair of each audio signal; identifying a common component from the first and second pairs of each audio signal; subtracting the common component or a modified common component from each of the first pair of each audio signal, where the modified common component is the common component multiplied by a gain value associated with a microphone associated with the first pair of microphones; restoring the delay to one subtracted gain multiplied component in each audio signal to generate two or more modified audio signals.

[0030] Obtaining two or more audio signals from two or more microphones includes selecting a first pair of two or more microphones to obtain two or more audio signals and selecting a second pair of two or more microphones to obtain a second pair of two or more audio signals, where the second pair of two or more microphones is in an audio shadow with respect to the first sound source direction parameter. Providing one or more modified audio signals based on two or more audio signals includes determining at least a second sound source direction parameter based at least in part on one or more modified audio signals in one or more frequency bands of the two or more audio signals, including providing the second pair of two or more audio signals.

[0031] One or more frequency bands may be lower than the threshold frequency.

[0032] According to a third aspect, there is provided an apparatus comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured to cause the at least one processor to perform at least: obtaining two or more audio signals from two or more microphones respectively; determining a first sound source direction parameter based on processing of the two or more audio signals in one or more frequency bands of the two or more audio signals, wherein the processing of the two or more audio signals is further configured to provide one or more modified audio signals based on the two or more audio signals; determining at least a second sound source direction parameter at least partially based on the one or more modified audio signals in one or more frequency bands of the two or more audio signals.

[0033] An apparatus configured to provide one or more modified audio signals based on two or more audio signals may further be configured to generate two or more modified audio signals based on modifying the two or more audio signals by the projection of a first sound source defined by the first sound source direction parameter, and an apparatus configured to determine at least a second sound source direction parameter at least partially based on the one or more modified audio signals in one or more frequency bands of the two or more audio signals may be configured to determine at least the second sound source direction parameter in one or more frequency bands of the two or more audio signals by processing the two or more modified audio signals.

[0034] The apparatus may further be configured to determine a first sound source energy parameter based on processing of two or more audio signals in one or more frequency bands of the two or more audio signals, and to determine at least a second sound source energy parameter based at least in part on one or more modified audio signals and the first sound source energy parameter.

[0035] The first and second sound source energy parameters may be direct-to-total energy ratios. The apparatus configured to determine at least the second sound source energy parameter based at least in part on one or more modified audio signals may be configured to determine an intermediate second sound source energy parameter direct-to-total energy ratio based on analysis of the one or more modified audio signals, and to generate a second sound source energy parameter direct-to-total energy ratio based on any of: selecting the minimum of the intermediate second sound source energy parameter direct-to-total energy ratio, or a value obtained by subtracting the first sound source energy parameter direct-to-total energy ratio from a value of 1; or multiplying the intermediate second sound source energy parameter direct-to-total energy ratio by a value obtained by subtracting the first sound source energy parameter direct-to-total energy ratio from a value of 1.

[0036] The apparatus configured to determine at least the second sound source energy parameter based at least in part on one or more modified audio signals and the first sound source energy parameter may be further configured to determine at least the second sound source energy parameter based further on a first sound source direction parameter such that the second sound source energy parameter is scaled relative to a difference between the first sound source direction parameter and a second sound source direction parameter.

[0037] In one or more frequency bands of two or more audio signals, an apparatus configured to determine a first sound source direction parameter based on processing of the two or more audio signals may perform: selecting a first pair of two or more microphones; selecting a first pair of respective audio signals from the selected pair of two or more microphones; determining a delay that maximizes the correlation between the first pair of respective audio signals from the selected pair of two or more microphones; determining a pair of directions related to the delay that maximizes the correlation between the first pair of respective audio signals from the selected pair of two or more microphones, wherein the first sound source direction parameter is selected from the determined pair of directions.

[0038] In one or more frequency bands of two or more audio signals, an apparatus configured to determine a first sound source direction parameter based on processing of the two or more audio signals may be configured to select the first sound source direction parameter from a pair of directions determined based on a further determination of a further delay that maximizes a further correlation between a further pair of respective audio signals from a further selected pair of two or more microphones.

[0039] In one or more frequency bands of two or more audio signals, an apparatus configured to determine a first sound source energy parameter based on processing of the two or more audio signals may be configured to determine a first sound source energy ratio corresponding to the first sound source direction parameter by normalizing the maximized correlation with respect to the energy of the first pair of respective audio signals for the frequency band.

[0040] An apparatus adapted to provide one or more modified audio signals based on two or more audio signals may perform: determining a delay between a first pair of each audio signal based on a determined first sound source direction parameter; aligning the first pair of each audio signal based on application of the determined delay to one of the first pair of each audio signal; identifying a common component from each of the first pair of each audio signal; subtracting the common component from each of the first pair of each audio signal; and restoring a delay to one of the subtracted components of each audio signal to generate one or more modified audio signals.

[0041] An apparatus adapted to provide one or more modified audio signals based on two or more audio signals may perform: determining a delay between a first pair of each audio signal based on a determined first sound source direction parameter; aligning the first pair of each audio signal based on application of the determined delay to one of the first pair of each audio signal; identifying a common component from each of the first pair of each audio signal; subtracting a modified common component from each of the first pair of each audio signal, the modified common component being the common component multiplied by a gain value associated with a microphone associated with a pair of microphones; and restoring a delay to one of the components of each audio signal multiplied by the subtracted gain to generate two or more modified audio signals.

[0042] An apparatus adapted to provide one or more modified audio signals based on two or more audio signals determines a delay between a first pair of each audio signal based on a determined first sound source direction parameter, wherein each audio signal is from a selected first pair of two or more microphones; aligns the first pair of each audio signal based on applying the determined delay to one of the first pair of each audio signal; selects an additional pair of each audio signal from a selected additional pair of two or more microphones; determines an additional delay between the additional pair of each audio signal based on a determined additional sound source direction parameter; aligns the additional pair of each audio signal based on applying the determined additional delay to one of the additional pair of each audio signal; identifies a common component from the first and second pairs of each audio signal; subtracts the common component or a modified common component from each of the first pair of each audio signal, wherein the modified common component is the common component multiplied by a gain value associated with a microphone associated with the first pair of microphones; restores the delay to one subtracted gain-multiplied component of each audio signal to generate two or more modified audio signals.

[0043] An apparatus configured to obtain two or more audio signals from each of two or more microphones is further configured to select a first pair of two or more microphones to obtain two or more audio signals, and to select a second pair of two or more microphones to obtain a second pair of two or more audio signals, wherein the second pair of two or more microphones is in an audio shadow with respect to a first sound source direction parameter, and an apparatus configured to provide one or more modified audio signals based on the two or more audio signals is configured to determine at least a second sound source direction parameter based at least in part on the one or more modified audio signals in one or more frequency bands of the two or more audio signals, and the second pair of two or more audio signals may be provided from the apparatus.

[0044] One or more frequency bands may be lower than a threshold frequency.

[0045] According to a fourth aspect, there is provided an apparatus comprising means for obtaining two or more audio signals from each of two or more microphones, means for determining a first sound source direction parameter based on processing of the two or more audio signals in one or more frequency bands of the two or more audio signals, wherein the processing of the two or more audio signals is further configured to provide one or more modified audio signals based on the two or more audio signals, means for determining at least a second sound source direction parameter based at least in part on the one or more modified audio signals in one or more frequency bands of the two or more audio signals.

[0046] According to a fifth aspect, the apparatus is caused to at least obtain two or more audio signals from respective two or more microphones, determine a first sound source direction parameter based on processing of the two or more audio signals in one or more frequency bands of the two or more audio signals, where the processing of the two or more audio signals is further configured to provide one or more modified audio signals based on the two or more audio signals, and determine at least a second sound source direction parameter based at least in part on the one or more modified audio signals in one or more frequency bands of the two or more audio signals, and a computer program including instructions [or a computer-readable medium including program instructions] for causing the above is provided.

[0047] According to a sixth aspect, a non-transitory computer-readable medium including program instructions for causing an apparatus to at least obtain two or more audio signals from respective two or more microphones, determine a first sound source direction parameter based on processing of the two or more audio signals in one or more frequency bands of the two or more audio signals, where the processing of the two or more audio signals is further configured to provide one or more modified audio signals based on the two or more audio signals, and determine at least a second sound source direction parameter based at least in part on the one or more modified audio signals in one or more frequency bands of the two or more audio signals is provided.

[0048] According to a seventh aspect, there is provided an apparatus comprising: an acquisition circuit configured to acquire two or more audio signals from each of two or more microphones; a determination circuit configured to determine a first sound source direction parameter based on processing of the two or more audio signals in one or more frequency bands of the two or more audio signals, wherein the processing of the two or more audio signals is further configured to provide one or more modified audio signals based on the two or more audio signals; and means for determining at least a second sound source direction parameter based at least in part on the one or more modified audio signals in one or more frequency bands of the two or more audio signals.

[0049] According to an eighth aspect, there is provided a computer-readable medium including program instructions for causing a device to perform at least: acquiring two or more audio signals from each of two or more microphones; determining a first sound source direction parameter based on processing of the two or more audio signals in one or more frequency bands of the two or more audio signals, wherein the processing of the two or more audio signals is further configured to provide one or more modified audio signals based on the two or more audio signals; and determining at least a second sound source direction parameter based at least in part on the one or more modified audio signals in one or more frequency bands of the two or more audio signals.

[0050] An apparatus including means for performing the operations of the above method.

[0051] An apparatus configured to perform the operations of the method described above.

[0052] A computer program including program instructions for causing a computer to execute the above method.

[0053] A computer program product stored in a medium can cause a device to execute the method described herein.

[0054] The electronic device may include an apparatus as described in this specification.

[0055] The chipset may be composed of an apparatus as described in this specification.

Brief Description of the Drawings

[0056] For a better understanding of the present application, reference is next made, by way of example, to the accompanying drawings.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

[0057] Regarding the following embodiments, the concepts described in more detail herein relate to the capture of an audio scene.

[0058] In the following description, the term sound source is used to describe a defined (artificial or real) element within a sound field (or audio scene). Also, the term sound source can be defined as an audio object or audio source, and these terms are interchangeable for the understanding of the examples described herein.

[0059] Embodiments herein relate to parametric audio capture devices and methods such as spatial audio capture (SPAC) technology. For each time-frequency tile, the device is configured to estimate the direction of the dominant sound source, and the relative energy of the direct and ambient components of the sound source is expressed as a direct-to-total energy ratio.

[0060] The following examples are suitable for devices having a challenging microphone arrangement or configuration as found within a typical portable terminal where the dimensions of the portable terminal include at least one short (or thin) dimension relative to other dimensions. In the examples shown herein, the captured spatial audio signal is a suitable input for a spatial synthesizer to generate a spatial audio signal such as a binaural format audio signal for headphone listening or a multi-channel signal format audio signal for loudspeaker listening.

[0061] In some embodiments, these examples can be implemented as part of a spatial capture front end of an Immersive Voice and Audio Services (IVAS) standard codec by generating IVAS-compatible audio signals and metadata.

[0062] General spatial analysis involves estimating the direction of the dominant sound source and the direct-to-total energy ratio for each time-frequency tile. These parameters are motivated by the human auditory system based on similar characteristics in principle. However, it is known that in some situations, such models may not be able to obtain optimal sound quality.

[0063] Generally, when multiple sound sources are present simultaneously, or when the sound source is mostly masked by background noise, problems may occur in parameter estimation. In the first case, the direction of the analyzed dominant sound source may deviate from the actual direction of the sound source, or depending on the total sound from the sound source, the analysis may result in the average value of the direction of the sound source. In the second case, depending on the instantaneous level and atmosphere of the sound source, the dominant sound source may or may not be found. In both of the above cases, in addition to the variation in the direction values, the estimated energy ratio may become unstable.

[0064] In such situations, distortion may occur in the synthesized audio signal due to the analysis of the direction and energy ratio. For example, the direction of the sound source may become unstable, sound inaccurately, or the background audio may become reverberant.

[0065] As an example, as shown in FIG. 1, an example of the main sound source direction estimation is shown when there are two sound sources of the same magnitude at azimuth angles of 30 degrees and -20 degrees around a capture device. As shown in FIG. 1, as time passes, it is determined that one of the sound sources is dominant, and both sound sources are synthesized in the estimated directions by the spatial synthesizer. At this time, since the estimated direction jumps continuously between two values, the result is ambiguous, and it is difficult for the user or listener to detect from which directions the two sound sources are emitted. Also, since this estimated direction changes continuously, the synthesized sound field becomes an unstable and unnatural sound.

[0066] When the amount of available information increases, techniques for improving the above problems have been proposed. For example, it has been proposed to estimate parameters for the two most dominant directions for each time-frequency tile. For example, in the currently being developed 3GPP (registered trademark) IVAS standard, it is planned to support two directions simultaneously.

[0067] However, in parametric audio coding using a microphone of a general mobile terminal, there is no highly reliable method for estimating the directions of two dominant sound sources. Furthermore, when the reliability of the estimation is low, there is a possibility that sound sources are synthesized in directions where there are actually no sound sources, the sound source position continuously moves from one position to another position, or it becomes unstable. That is, when the reliability of the estimation is low, there is no merit in estimating multiple directions, and the quality of the spatial audio signal generated by the spatial synthesizer may deteriorate.

[0068] Therefore, in short, the embodiments described in this specification are related to parametric spatial audio capture using two or more microphones. Furthermore, at least two directions and energy ratio parameters are estimated in all time-frequency tiles based on audio signals from two or more microphones.

[0069] In these embodiments, to achieve an improvement in the detection accuracy of multiple sound source directions, the influence of the first estimated direction is considered when estimating the second direction. This can result in an improvement in the perceived quality of the synthesized spatial audio in some embodiments.

[0070] In fact, the embodiments described herein generate an estimated value of the sound source that is recognized as being more spatially stable and more accurate (with respect to the correct or actual position).

[0071] In some embodiments, the first direction and the energy ratio are estimated (can be estimated) using any suitable estimation method. Further, when estimating the second direction, the influence of the first direction is first removed from the microphone signals. In some embodiments, this can be implemented by first removing any delay between the signals based on the first direction and then subtracting the common component from both signals. Finally, the original delay is restored. Next, the second direction parameter can be estimated using a method similar to the estimation of the first direction.

[0072] In some embodiments, different pairs of microphones are used to estimate two different directions at low frequencies. This emphasizes the natural shadowing of sound due to the physical shape of the device and improves the possibility of detecting sound sources on different sides of the device.

[0073] In some embodiments, the energy ratio of the second direction is first analyzed using a method similar to the estimation of the energy ratio of the first direction. In some further embodiments, the second energy ratio is further corrected based on the energy ratio of the first direction and based on the angular difference between the first estimated sound source direction and the second estimated sound source direction.

[0074] Regarding FIG. 2, it is a schematic diagram of an apparatus suitable for implementing the embodiments described herein.

[0075] In this example, an apparatus including a microphone array 201 is shown. The microphone array 201 is composed of a plurality (two or more) of microphones configured to capture audio signals. The microphones within the microphone array can be of any suitable microphone type, arrangement, or configuration. The microphone audio signal 202 generated by the microphone array 201 can be passed to a spatial analyzer 203.

[0076] The apparatus can include a spatial analyzer 203 configured to receive or otherwise obtain the microphone audio signal 202 and is configured to spatially analyze the microphone audio signal to determine at least two dominant sounds or audio sources for each time-frequency block.

[0077] In some embodiments, the spatial analyzer can be the CPU of a mobile terminal or a computer. The spatial analyzer 203 is configured to generate a data stream that includes not only the audio signal but also metadata of the analyzed spatial information 204.

[0078] Depending on the use case, the data stream can be saved, compressed, or transmitted to another location.

[0079] The apparatus further has a spatial synthesizer 205. The spatial synthesizer 205 is configured to obtain a data stream including the audio signal and metadata. In some embodiments, the spatial synthesizer 205 is implemented within the same apparatus as the spatial analyzer 203 (as shown in FIG. 2 here), but in some embodiments, it can also be implemented in a different apparatus or device.

[0080] The spatial synthesizer 205 can be implemented within a CPU or a similar processor. The spatial synthesizer 205 is configured to generate an output audio signal 206 based on the audio signal and the associated metadata from the data stream 204.

[0081] Furthermore, depending on the use case, the output signal 206 can be in any suitable output format. For example, in some embodiments, the output format is a binaural headphone signal (where the output device presenting the output audio signal is a set such as headphones / earphones, etc.), or a multi-channel loudspeaker audio signal (where the output device is a set of loudspeakers). The output device 207 (which can be, for example, headphones or loudspeakers as described above) can be configured to receive the output audio signal 206 and present the output to a listener or user.

[0082] These operations of the exemplary apparatus shown in FIG. 2 can be illustrated by the flowchart shown in FIG. 3. Thus, summarizing the operations of the exemplary apparatus, it is as follows.

[0083] As shown in FIG. 3, in step 301, a microphone audio signal is acquired.

[0084] As shown in FIG. 3, in step 303, the microphone audio signal is spatially analyzed to generate a spatial audio signal and metadata including the directions and energy ratios of the first and second audio sources for each time-frequency tile.

[0085] As shown in FIG. 3, in step 305, spatial synthesis is applied to the spatial audio signal to generate a suitable output audio signal.

[0086] As shown in FIG. 3, in step 307, the output audio signal is output to an output device.

[0087] In one embodiment, spatial analysis can be used in connection with an IVAS codec. In this example, the spatial analysis output is in an IVAS-compatible MASA (metadata-assisted spatial audio) format and can be fed directly to an IVAS encoder. The IVAS encoder generates an IVAS data stream. On the receiving side, the IVAS decoder can directly generate the desired output audio format. That is, in such an embodiment, there is no separate spatial synthesis block.

[0088] This is shown, for example, for the apparatus shown in FIG. 4 and the operation of the apparatus shown by the flow diagram of FIG. 5.

[0089] In this example shown in FIG. 4, the apparatus also includes a microphone array 201. It is configured to generate a microphone audio signal 202 that is passed to a spatial analyzer 203.

[0090] The spatial analyzer 203 is configured to receive or otherwise obtain the microphone audio signal 202 and determine at least two dominant sound sources or audio sources for each time-frequency block. The data stream generated by the spatial analyzer 203, the MASA format data stream 404 (including not only the audio signal but also the metadata of the analyzed spatial information), can then be passed to an IVAS encoder 405.

[0091] The apparatus can further include an IVAS encoder 405 configured to receive the MASA format data stream 404 and generate an IVAS data stream 406 that can be transmitted or stored as shown by the dashed line 416.

[0092] The apparatus further has an IVAS decoder 407 (spatial synthesizer). The IVAS decoder 407 is configured to decode the IVAS data stream and further spatially synthesize the determined audio signals to generate an output audio signal 206 to an appropriate output device 207.

[0093] The output device 207 (which can be, for example, headphones or a loudspeaker as described above) can be configured to receive the output audio signal 206 and present the output to a listener or user.

[0094] The operation of the device of the embodiment shown in FIG. 4 can be shown by the flowchart shown in FIG. 5. Accordingly, summarizing the operation of the device of this embodiment, it is as follows.

[0095] As shown in FIG. 5, at step 301, a microphone audio signal is acquired.

[0096] As shown in FIG. 5, at step 503, the microphone audio signal is spatially analyzed to generate an output in MASA format (metadata including the spatial audio signal and the direction and energy ratio of the first and second audio sources for each time-frequency tile).

[0097] As shown in FIG. 5, at step 505, the generated data stream is encoded in IVAS.

[0098] As shown in FIG. 5, at step 507, the encoded IVAS data stream is decoded (and spatial synthesis is performed on the decoded spatial audio signal) to generate an appropriate output audio signal.

[0099] As shown in FIG. 5, at step 307, the output audio signal is output to the output device.

[0100] In some embodiments, instead, the output audio signal is an ambisonic signal. In such embodiments, there may not be a directly available output device.

[0101] The spatial analyzer indicated by reference numeral 203 in FIGS. 2 and 4 is shown in further detail with reference to FIG. 5.

[0102] In some embodiments, the spatial analyzer 203 has a stream (transport) audio signal generator 607. The stream audio signal generator 607 is configured to receive the microphone audio signal 202 and generate a stream audio signal(s) 608 that is passed to the multiplexer 609. The audio stream signal is generated from the input microphone audio signal based on any suitable method. For example, in some embodiments, one or two microphone signals may be selected from the microphone audio signal 202. Alternatively, in some embodiments, the microphone audio signal 202 may be downsampled and / or compressed to generate the stream audio signal 608.

[0103] In the following example, the spatial analysis is performed in the frequency domain, but it will be understood that in some embodiments, the analysis can also be performed in the time domain using a time-domain sampled version of the microphone audio signal.

[0104] In some embodiments, the spatial analyzer 203 has a time-frequency converter 601. The time-frequency converter 601 is configured to receive the microphone audio signal 202 and convert it to the frequency domain. In some embodiments, prior to the conversion, the time-domain microphone audio signal can be represented as s i (t), where t is the time index and i is the microphone channel index. The conversion to the frequency domain can be performed by any suitable time-frequency conversion such as STFT (Short-Time Fourier Transform) or (complex modulation) QMF (Quadrature Mirror Filter Bank). The resulting time-frequency domain microphone signal 602 is denoted as S i (b,n), where i is the microphone channel index, b is the frequency bin index, and n is the time frame index. The value of b ranges from 0, ···, B - 1, and B is the number of bin indices for each time index n.

[0105] The frequency bins can further be combined with sub-bands k = 0, ···, K−1. Each sub-band is composed of one or more frequency bins. Each sub-band k has a lowest bin b k,low and a highest bin b k,high . The width of the sub-band is usually selected based on human auditory characteristics. For example, the equivalent rectangular bandwidth (ERB) or the Bark scale can be used.

[0106] In some embodiments, the spatial analyzer 203 includes a first direction analyzer 603. The first direction analyzer 603 is configured to receive the time-frequency domain microphone audio signal 602 and generate an estimated value of the first sound source for each time-frequency tile of a (first) first direction 614 and a (first) first ratio 616.

[0107] The first direction analyzer 603 is configured to generate an estimated value of the first direction based on any suitable method such as SPAC (as described in more detail in US9313599).

[0108] In some embodiments, for example, the most dominant direction for a time frame index is estimated by searching for a time shift τ k that maximizes the correlation between two (microphone audio signal) channels for sub-band k. S i (b,n) can be shifted by τ samples as follows.

Equation

[0109] Then, the delay τ k of each sub-band k that maximizes the correlation between the two microphone channels is determined.

Equation

[0110] In the above equation, an "optimal" delay is searched for between microphone 1 and microphone 2. Re represents the real part of the result, and * indicates the complex conjugate of the signal. Delay search range parameter D max is defined based on the distance between the microphones. That is, τ k is searched for only within the physically possible range considering the distance between the microphones and the speed of sound.

[0111] At this time, the angle in the first direction is defined as follows.

Equation

[0112] Thus, there is still uncertainty remaining in the sign of the angle.

[0113] In the above, the direction analysis between microphone 1 and microphone 2 was defined. By repeating the same procedure for other microphone pairs, the ambiguity can be resolved (and / or the direction with respect to other axes can be obtained). That is, To resolve the sign ambiguity of JPEG0007708730000004.jpg1084, information from other analysis pairs can be utilized.

[0114] For example, FIG. 8 shows an example where the microphone array includes three microphones, a first microphone 801, a second microphone 803, and a third microphone 805, and there are a first pair (the first microphone 801 and the third microphone 803) separated by a distance on a first axis, and a second pair (the first microphone 801 and the second microphone 805) separated by a distance on a second axis (in this example, the first axis is perpendicular to the second axis). Further, in this example, the three microphones can be located on the same third axis defined as perpendicular to the first axis and the second axis (and perpendicular to the plane of the paper on which the figure is printed). Analysis of the delay between the first pair 801 and 803 of microphones yields two alternative angles, α807 and -α809. Next, analysis of the delay between the second pair 801 and 805 of microphones can be used to determine which of the alternative angles is correct. In some embodiments, the information required from this analysis is which of the microphones 801 or 805 the sound arrives at first. If the sound arrives at microphone 805, the angle α is correct. Otherwise, -α is selected.

[0115] Furthermore, based on the estimation between multiple microphone pairs, the first spatial analyzer can determine or estimate the correct direction angle JPEG0007708730000005.jpg989.

[0116] In some limited microphone configurations or arrangements, for example, in some embodiments where there are only two microphones, the ambiguity in direction cannot be resolved. In such examples, the spatial analyzer may be configured to define that all sound sources are always in front of the device. This situation is the same even when there are two or more microphones, and depending on their positions, for example, analysis in the front - back direction may not be possible.

[0117] Although not disclosed herein, elevation and azimuth can be estimated with a plurality of pairs of microphones on the vertical axis.

[0118] The first-direction analyzer 603 can further determine or estimate the energy ratio r1(k,n) corresponding to the angle θ1(k,n) by using, for example, the correlation value c(k,n) after normalization as follows.

Number

[0119] The value of r1(k,n) is between -1 and 1, and usually further limited between 0 and 1.

[0120] In some embodiments, the first-direction analyzer 603 is configured to generate a modified time-frequency microphone audio signal 604. The modified time-frequency microphone audio signal 604 has the first sound source component removed from the microphone signal.

[0121] Therefore, for example, regarding the first microphone pair (microphones 801 and 803 shown in the microphone configuration example of FIG. 8), for sub-band k, the delay that gives the highest correlation is τ k . For each sub-band k, shift the second microphone signal by τ k samples only to obtain the shifted second microphone signal S 2,τk (b,n).

[0122] An estimated value of the sound source component can be obtained as the average value of these signals with aligned time axes.

Number

[0123] In some embodiments, any other suitable method for determining the sound source component can be used.

[0124] Once the estimated value of the sound source component C(b,n) is determined (for example, in the above mathematical formula example), it can be removed from the microphone audio signal. On the other hand, since the phases of other simultaneous sound sources are shifted, C(b,n) is attenuated. Here, C(b,n) can be reduced from the microphone signal (both the shifted and unshifted ones).

Number

[0125] Furthermore, the shifted and modified microphone audio signal JPEG0007708730000009.jpg1074 returns to τ k and returns.

Number

[0126] These modified signals JPEG0007708730000011.jpg10108 can then be passed to the second direction analyzer 605.

[0127] In some embodiments, the spatial analyzer 203 includes the second direction analyzer 605. The second direction analyzer 605 is configured to receive the time-frequency microphone audio signal 602, the modified time-frequency microphone audio signal 604, the first direction 614, and the estimated value of the first ratio 616, and generate the estimated values of the second direction 624 and the second ratio 626.

[0128] The estimation of the parameter values in the second direction adopts the same sub-band structure as the estimation in the first direction and can follow the same operations as described above for the estimation in the first direction.

[0129] Therefore, the second direction parameters θ2(k,n) and r2´(k,n) can be estimated. In such embodiments, in order to determine the direction estimation, instead of the time-frequency microphone audio signals 602 S1(b,n) and S2(b,n), the modified time-frequency microphone audio signals JPEG0007708730000012.jpg899 is used.

[0130] Furthermore, in some embodiments, the energy ratio r2´(k,n) is limited because the sum of the first and second ratios must not exceed 1.

[0131] In some embodiments, the second ratio is limited as follows.

Equation

Equation

[0132] Here, the function min selects the smaller of the given options. It has been found that both alternatives provide good quality ratio values.

[0133] In the above example, since there are multiple microphone pairs, the correction signal needs to be calculated separately for each pair. That is, considering microphone pairs 801 and 805, or pairs 801 and 803, Note that JPEG0007708730000015.jpg888 is not the same signal.

[0134] The first direction estimate 614, the first ratio estimate 616, the second direction estimate 624, and the second ratio estimate 626 are passed to a multiplexer (mux) 609 configured to generate the data stream 204 / 404 from the combination of the estimates and the stream audio signal 608.

[0135] Regarding FIG. 7, a flowchart summarizing the operation example of the spatial analyzer shown in FIG. 6 is shown.

[0136] As shown in FIG. 7, at step 701, the microphone audio signal is acquired.

[0137] Then, as shown in FIG. 7, a stream audio signal is generated from the microphone audio signal by step 702.

[0138] Furthermore, as shown in FIG. 7, the microphone audio signal can be transformed into the time-frequency domain by step 703.

[0139] Thereafter, as shown in FIG. 7, an estimated value of the parameter of the first direction and the first ratio can be determined by step 705.

[0140] Next, as shown in FIG. 7, the microphone audio signal in the time-frequency domain can be modified (in order to remove the first source component) by step 707.

[0141] Next, as shown in FIG. 7, the modified microphone audio signal in the time-frequency domain is analyzed to determine the estimated values of the second direction and the second ratio parameters by step 709.

[0142] Then, as shown in FIG. 7, the estimated values of the parameters of the first direction, the first ratio, the second direction, and the second ratio and the stream audio signal are multiplexed by step 711 to generate a data stream (which may be a data stream in the MASA format).

[0143] Therefore, as shown in FIG. 9, an example of the direction analysis result of one subband is shown. The input is an uncorrelated noise signal arriving simultaneously from two directions, and the signal arriving from the first direction is 1 dB stronger than that from the second direction. In many cases, the stronger sound source is detected as the first direction, but sometimes the sound source in the second direction may be detected as the first direction. If only one direction is estimated, the direction estimated value will jump between the two values, which may potentially cause quality problems. In the case of two-direction analysis, since both sound sources are included in the first or second direction, the quality of the synthesized signal is always kept good.

[0144] For example, FIG. 10 shows the direction estimation results in the same situation as FIG. 1 (direction estimation is performed only once for each time-frequency tile). For comparison, it can be seen that when two direction estimations are performed in the same situation, the position of the sound source is maintained.

[0145] In some embodiments, other methods may be employed to determine the common component C(b,n) (the first source component). For example, in some embodiments, principal component analysis (PCA) or other related methods can be employed. In some embodiments, when generating or subtracting the common component, individual gains for different channels are applied. Thus, for example, in some embodiments, the following occurs.

Number

Number

[0146] In such embodiments, for example, the common component can be removed from the microphone signals while taking into account the different levels of the audio signals in the microphones.

[0147] Furthermore, in the above example, the common component (combined signal) C(b,n) is generated using two microphone signals, but in some embodiments, more microphones can be employed. For example, if three available microphones are present, the "optimal" delays between microphone pairs 801 and 803, and 801 and 805 can be estimated. These are denoted as τ k (1,2) and τ k (1,3) respectively. In such embodiments, the combined signal is obtained as follows.

Number

[0148] Similar to the above, before analyzing the second direction, the combined signal can be removed from all three microphone signals.

[0149] In the above example, the method for estimating two directions generally provides good results. However, the microphone positions in a typical mobile terminal microphone configuration can be used to further improve the estimated values and, in some examples, to improve the reliability of the second direction analysis, particularly at the lowest frequencies.

[0150] For example, FIG. 11 shows the typical configured positions of microphones in a recent mobile terminal. This terminal has a display 1109 and a camera housing 1107. Microphones 1101 and 1105 are placed quite close to each other, while microphone 1103 is placed at a more distant position. The physical shape of the terminal affects the audio signals captured by the microphones. Microphone 1105 is on the main camera side of the terminal. Sound arriving from the display side of the terminal must circumnavigate the edge of the terminal to reach microphone 1105. Due to this long path, the signal is attenuated, and depending on the frequency, it can be attenuated by 6 to 10 dB. On the other hand, microphone 1101 is at the end of the device, and sound arriving from the left side of the device reaches the microphone directly, while sound arriving from the right side needs to go around a corner. Thus, even though microphones 1101 and 1105 are close, the signals they capture can be quite different.

[0151] The difference between these two microphone signals can be utilized for azimuth analysis. Using the equations shown above, the optimal delays τ k (1,2) and τ k (3,2) between the microphones in microphone pairs 1 - 2 (microphone numbers 1101 and 1103), 3 - 2 (microphone numbers 1105 and 1103) can be estimated, and the corresponding angles for JPEG0007708730000019.jpg10127 can also be estimated. Since the distances between microphone pairs are different, this needs to be considered when calculating the angles.

[0152] In particular, When JPEG0007708730000020.jpg12118 clearly points in different directions, that is, when different dominant sound sources are found, these two directions can be directly utilized as two - direction estimation.

Number

[0153] The energy ratio can be calculated in the same way as shown previously, and the value of r2(k,n) needs to be restricted again based on the value of r1(k,n). The ambiguity of the sign of the value of JPEG0007708730000022.jpg886 can be resolved in the same way as above. In other words, the microphone pairs 1 - 3 can be used to resolve the ambiguity of directionality.

[0154] These embodiments have revealed that they are particularly useful in the lowest - frequency band where two - direction estimation is most difficult with a general microphone configuration.

[0155] In the above embodiments, it has been discussed that the energy ratio r2(k,n) of the second direction is restricted based on the value of the first energy ratio r1(k,n). In some embodiments, the angular difference between the first and second direction estimations is used to correct the ratio(s).

[0156] Therefore, in some embodiments, when θ1(k,n) and θ2(k,n) are facing the same direction, the energy ratio parameter of the first direction already contains a sufficient amount of energy, and there is no need to allocate more energy to the given second direction. That is, r2(k,n) can be set to zero. Conversely, when θ1(k,n) and θ2(k,n) are facing opposite directions, the influence of the ratio r2(k,n) is the greatest, and it is necessary to maintain the value of r2(k,n) at the maximum.

[0157] This can be implemented in some embodiments where β(k,n) is the absolute angular difference between θ1(k,n) and θ2(k,n).

Number

Number

[0158] Then, the overall effect of the first direction on the energy ratio in the second direction can be calculated as follows.

Number

Number

[0159] Here, r2´(k,n) is the original ratio and r2(k,n) is the modified ratio. In this example, the angle difference has a linear effect on the scaling of r2(k,n). In some embodiments, there are other weighting options, such as sine wave weighting.

[0160] Referring to FIG. 12, examples of a spatial synthesizer 205 or an IVAS decoder 407 as shown in FIGS. 2 and 4 respectively are shown.

[0161] The spatial synthesizer 205 / IVAS decoder 407 in some embodiments has a demultiplexer 1201. The demultiplexer (Demux) 1201 in some embodiments receives the data stream 204 / 404 and separates the data stream into a stream audio signal 1208 and spatial parameter estimates such as an estimate of the first direction 1214, an estimate of the first ratio 1216, an estimate of the second direction 1224, and an estimate of the second ratio 1226. In some embodiments where the data stream is encoded (e.g., using an IVAS encoder), the data stream can be decoded here.

[0162] These are passed to the spatial processor / synthesizer 1203.

[0163] The spatial synthesizer 205 / IVAS decoder 407 includes the spatial processor / synthesizer 1203 and is configured to receive the estimated values and the stream audio signals and render the output audio signals. The spatial processing / synthesis can be any suitable two-way based synthesis as described in EP3791605.

[0164] FIG. 13 is a schematic diagram showing an example according to some embodiments. The apparatus is a capture / playback device 1301 including components of a microphone array 201, a spatial analyzer 203, and a spatial synthesizer 205. Further, the apparatus 1301 has a storage (memory) 1201 configured to store the audio signals and the metadata (data stream) 204.

[0165] In some embodiments, the capture / playback device 1301 can be a portable terminal.

[0166] With respect to FIG. 14, an exemplary electronic device that can be used as a computer, an encoder processor, a decoder processor, or any of the functional blocks described herein is shown. The device can be any suitable electronic device or apparatus. For example, in some embodiments, the device 1600 is a portable terminal, a user device, a tablet computer, a computer, an audio playback device, etc.

[0167] In some embodiments, the device 1600 has at least one processor or a central processing unit 1607. The processor 1607 can be configured to execute various program codes such as the methods described herein.

[0168] In some embodiments, device 1600 has a memory 1611. In some embodiments, at least one processor 1607 is connected to the memory 1611. The memory 1611 can be any suitable storage means. In some embodiments, the memory 1611 includes a program code section for storing program code implementable on the processor 1607. Further, in some embodiments, the memory 1611 can further include a stored data section for storing data, such as data processed according to, or to be processed according to, the embodiments described herein. The implemented program code stored within the program code section, and the data stored within the stored data section, can be retrieved by the processor 1607 via the memory-processor connection when needed.

[0169] In some embodiments, apparatus 1600 includes a user interface 1605. The user interface 1605 can be connected to the processor 1607 in some embodiments. In some embodiments, the processor 1607 can control the operation of the user interface 1605 and receive input from the user interface 1605. In some embodiments, the user interface 1605 can enable a user to input commands to the device 1600, for example, via a keypad. In some embodiments, the user interface 1605 can enable a user to obtain information from the apparatus 1600. For example, the user interface 1605 can include a display configured to display information from the apparatus 1600 to the user. The user interface 1605 can, in some embodiments, comprise a touch screen or touch interface capable of both enabling information to be input into the apparatus 1600 and displaying information to a user of the apparatus 1600.

[0170] In some embodiments, apparatus 1600 has an input / output port 1609. The input / output port 1609 in some embodiments has a transceiver. The transceiver in such embodiments is connected to processor 1607 and may be configured to enable communication with other devices or electronic equipment, for example, via a wireless communication network. The transceiver, or any suitable transceiver, or transmitting and / or receiving means may be configured to communicate with other electronic equipment or devices via a wired or wireless connection in some embodiments.

[0171] The transceiver can communicate with additional devices according to any suitable known communication protocol. For example, in some embodiments, the transceiver can use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a wireless local area network (WLAN) protocol such as IEEE802.X, a suitable short-range radio frequency communication protocol such as Bluetooth®, or an infrared data communication path (IRDA).

[0172] The transceiver input / output port 1609 may be configured to transmit / receive audio signals, bitstreams by using processor 1607 that executes suitable code, and in some embodiments, to perform operations and methods as described above.

[0173] Generally, various embodiments of the present invention may be implemented in hardware or special-purpose circuitry, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, a microprocessor, or other computing device, but the present invention is not limited thereto. Various aspects of the present invention may be illustrated and described as block diagrams, flowcharts, or in some other graphical representation, but these blocks, devices, systems, techniques, or methods described herein are, by way of non-limiting example, hardware, software, firmware, special-purpose circuitry or logic, general-purpose hardware or a controller or other computing device, or any combination thereof.

[0174] Embodiments of the present invention may be implemented by computer software executable by a data processor of a mobile terminal such as a processor entity, or by hardware, or by a combination of software and hardware. Further, in this regard, it should be noted that any block of the logic flow as illustrated in the figures can represent a program step, or interconnected logical circuits, blocks and functions, or a combination of program steps and logical circuits, blocks and functions. The software may be stored in a physical medium such as a memory chip, or a memory block implemented within a processor, a magnetic medium, and an optical medium.

[0175] The memory can be of any type suitable for the local technical environment and can be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor can be of any type suitable for the local technical environment and may include, by way of non-limiting example, one or more of a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a gate-level circuit, and a processor based on a multi-core processor architecture.

[0176] Embodiments of the present invention can be implemented in various components such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available to convert the logic-level design into a semiconductor circuit design suitable for etching and forming on a semiconductor substrate.

[0177] Programs such as Synopsys in Mountain View, California, and Cadence Design in San Jose, California, use established design rules and a library of pre-stored design modules to automatically route conductors and place components on a semiconductor chip. Once the design of the semiconductor circuit is complete, the design results can be sent in a standardized electronic format (Opus, GDSII, etc.) to a semiconductor manufacturing facility or "fab" for outsourcing the manufacturing.

[0178] The foregoing description has provided a complete and reference description of exemplary embodiments of the present invention by way of illustrative and non-limiting examples. However, upon reading the foregoing description in conjunction with the accompanying drawings and the appended claims, various modifications and applications will become apparent to those skilled in the relevant art. However, all such and similar modifications of the teachings of this invention will still fall within the scope of the invention as defined by the appended claims.

Claims

Claim 1 An apparatus comprising at least one processor and at least one memory including computer program code, wherein the at least one memory and the computer program code are configured to cause the at least one processor to cause the apparatus to at least obtain two or more audio signals from each of two or more microphones; determine a first sound source direction parameter based on processing of the two or more audio signals in one or more frequency bands of the two or more audio signals, wherein processing the two or more audio signals is further configured to generate one or more modified audio signals based on the two or more audio signals, and the first sound source associated with the determined first sound source direction parameter is removed from the generated one or more modified audio signals; determine at least a second sound source direction parameter associated with a second sound source based at least in part on the one or more generated modified audio signals in the one or more frequency bands of the two or more audio signals; generate a data stream based at least on the first sound source direction parameter, the second sound source direction parameter, and the two or more audio signals. An apparatus for performing the above. Claim 2 The apparatus for generating the one or more modified audio signals is further configured to generate the modified two or more audio signals based on modifying the two or more audio signals by projection of a first sound source defined by the first sound source direction parameter; The means configured to determine at least the second sound source direction parameter based at least in part on the one or more generated modified audio signals in the one or more frequency bands of the two or more audio signals is configured to determine at least the second sound source direction parameter by processing the modified two or more audio signals in the one or more frequency bands of the two or more audio signals. The apparatus according to claim 1. Claim 3 Determining a first sound source energy parameter based on the processing of the two or more audio signals in one or more frequency bands of the two or more audio signals; Determining at least a second sound source energy parameter based at least in part on the one or more modified audio signals and the first sound source energy parameter; The apparatus according to claim 1, further configured to perform the above.

4. The first and second sound source energy parameters are direct-to-total energy ratios, and the apparatus is configured to determine at least the second sound source energy parameter, Determining an intermediate second sound source energy parameter direct-to-total energy ratio based on an analysis of the one or more modified audio signals; Selecting the smallest of the intermediate second sound source energy parameter direct-to-total energy ratio or the value obtained by subtracting the value of the first sound source energy parameter direct-to-total energy ratio from 1, and multiplying the intermediate second sound source energy parameter direct-to-total energy ratio by the value obtained by subtracting the value of the first sound source energy parameter direct-to-total energy ratio from 1, and generating the second sound source energy parameter direct-to-total energy ratio based on one of the above; The apparatus according to claim 3, further configured to perform the above.

5. The determined at least second sound source energy parameter causes the apparatus to further determine the second sound source energy parameter based on the first sound source direction parameter so as to be scaled with respect to the difference between the first sound source direction parameter and the second sound source direction parameter. The apparatus according to claim 3.

6. The determined first sound source direction parameter causes the apparatus to Selecting a first pair of the two or more microphones; Selecting a first pair of respective audio signals from the selected pair of the two or more microphones; Determining a delay that maximizes the correlation between the first pair of the respective audio signals from the selected pair of the two or more microphones; Determining a pair of directions related to a delay that maximizes the correlation between the first pair of respective audio signals from the selected pair of the two or more microphones, wherein the first sound source direction parameter is selected from the determined pair of directions; The apparatus according to claim 1, which causes the above to be performed. **Claim 7** Based on the processing of the two or more audio signals, the determined first sound source direction parameter is configured to select the first sound source direction parameter from the determined pair of directions based on a further determination of a further delay that maximizes a further correlation between a further pair of respective audio signals from a further selected pair of the two or more microphones. The apparatus according to claim 6. **Claim 8** Based on the processing of the two or more audio signals, the determined first sound source energy parameter causes the apparatus to determine a first sound source energy ratio corresponding to the first sound source direction parameter by normalizing the maximized correlation with respect to the energy of each audio signal of the first pair for the frequency band. The apparatus according to claim 6. **Claim 9** The generated one or more modified audio signals cause the apparatus to determine a delay between the first pair of respective audio signals based on the determined first sound source direction parameter; align the first pair of respective audio signals based on applying the determined delay to one of the first pair of respective audio signals; identify a common component from each of the first pair of respective audio signals; subtract the common component from each of the first pair of respective audio signals; restore the delay to the subtracted components of the respective audio signals and generate the one or more modified audio signals; The apparatus according to claim 1, which causes the above to be performed. **Claim 10** The generated one or more modified audio signals cause the apparatus to determine a delay between the first pair of respective audio signals based on the determined first sound source direction parameter; align the first pair of respective audio signals based on applying the determined delay to one of the first pair of respective audio signals; Identifying a common component from each of the first pairs of the respective audio signals; Subtracting the modified common component from each of the first pairs of the respective audio signals, wherein the modified common component is the common component multiplied by a gain value associated with a microphone associated with the pair of microphones; Restoring the delay to the subtracted gain-multiplied component of each of the respective audio signals to generate the one or more modified audio signals; The apparatus according to claim 1, causing the above to be performed.

11. The one or more generated modified audio signals cause the apparatus to Determining a delay between a first pair of the respective audio signals based on the determined first sound source direction parameter, wherein each of the respective audio signals is from a selected first pair of the two or more microphones; Aligning the first pair of the respective audio signals based on applying the determined delay to one of the first pair of the respective audio signals; Selecting an additional pair of the respective audio signals from a selected additional pair of the two or more microphones; Determining an additional delay between the additional pair of the respective audio signals based on the determined additional sound source direction parameter; Aligning the additional pair of the respective audio signals based on applying the determined additional delay to one of the additional pair of the respective audio signals; Identifying a common component from the first and second pairs of the respective audio signals; Subtracting the common component or the modified common component from each of the first pairs of the respective audio signals, wherein the modified common component is the common component multiplied by a gain value associated with a microphone associated with the first pair of microphones; Restoring the delay to the subtracted gain-multiplied component of each of the respective audio signals to generate the modified one or more audio signals; The apparatus according to claim 1, causing the above to be performed.

12. The two or more acquired audio signals cause the apparatus to selecting a first pair of the two or more microphones to obtain the two or more audio signals, and selecting a second pair of the two or more microphones to obtain a second pair of the two or more audio signals, wherein the second pair of the two or more microphones is in an audio shadow with respect to the first sound source direction parameter, and the generated one or more modified audio signals cause the device to provide the second pair of the two or more audio signals based on at least a second sound source direction parameter determined at least partially in relation to the generated one or more modified audio signals. The device according to claim 1.

13. The device according to claim 12, wherein the one or more frequency bands are lower than a threshold frequency.

14. A method for a device, comprising: obtaining two or more audio signals from respective two or more microphones; determining a first sound source direction parameter based on processing of the two or more audio signals in one or more frequency bands of the two or more audio signals, wherein the processing of the two or more audio signals is further configured to generate one or more modified audio signals based on the two or more audio signals, and a first sound source associated with the determined first sound source direction parameter is removed from the generated one or more modified audio signals; determining at least a second sound source direction parameter associated with a second sound source based at least in part on the generated one or more modified audio signals in the one or more frequency bands of the two or more audio signals; generating a data stream based on at least the first sound source direction parameter, the second sound source direction parameter, and the two or more audio signals; A method comprising the above steps.

15. Generating one or more modified audio signals based on the two or more audio signals further comprises: generating the modified two or more audio signals based on modifying the two or more audio signals by a projection of a first sound source defined by the first sound source direction parameter; Determining at least the second sound source direction parameter, at least in part, based on at least the one or more modified audio signals in the one or more frequency bands of the two or more audio signals, including processing the two or more modified audio signals in the one or more frequency bands of the two or more audio signals to determine the at least second sound source direction parameter, and The method according to claim 14, including the above.

16. Determining a first sound source energy parameter based on the processing of the two or more audio signals in one or more frequency bands of the two or more audio signals; Determining at least a second sound source energy parameter, at least in part, based on at least the one or more modified audio signals and the first sound source energy parameter; The method according to claim 14, further including the above.

17. The first and second sound source energy parameters are direct-to-total energy ratios, and determining at least the second sound source energy parameter, at least in part, based on at least the one or more modified audio signals includes: Determining an intermediate second sound source energy parameter direct-to-total energy ratio based on the analysis of the one or more modified audio signals; Selecting the smallest of the intermediate second sound source energy parameter direct-to-total energy ratio or the value obtained by subtracting the value of the first sound source energy parameter direct-to-total energy ratio from 1, or Multiplying the intermediate second sound source energy parameter direct-to-total energy ratio by the value obtained by subtracting the value of the first sound source energy parameter direct-to-total energy ratio from 1; Generating the second sound source energy parameter direct-to-total energy ratio based on one of the above; The method according to claim 16, including the above.

18. Determining the at least second sound source energy parameter based at least in part on the one or more modified audio signals and the first sound source energy parameter includes determining the at least second sound source energy parameter further based on the first sound source direction parameter such that the second sound source energy parameter is scaled relative to a difference between the first sound source direction parameter and the second sound source direction parameter, the method according to claim 16.

19. Determining a first sound source direction parameter based on processing of the two or more audio signals in one or more frequency bands of the two or more audio signals comprises: selecting a first pair of the two or more microphones; selecting a first pair of respective audio signals from the selected pair of the selected two or more microphones; determining a delay that maximizes a correlation between the first pair of respective audio signals from the selected pair of the two or more microphones; determining a pair of directions related to the delay that maximizes the correlation between the first pair of respective audio signals from the selected pair of the two or more microphones, wherein the first sound source direction parameter is selected from the determined pair of directions, the determining; comprising the method according to claim 14.

20. Determining a first sound source direction parameter based on processing of the two or more audio signals in one or more frequency bands of the two or more audio signals includes selecting the first sound source direction parameter from the determined pair of directions based on a further determination of a further delay that maximizes a further correlation between a further pair of respective audio signals from a selected further pair of the two or more microphones, the method according to claim 19.

Citation Information

Patent Citations

  • Noise rejection device, noise rejection program, and noise rejection method

    JP2017097101A

  • Sound source survey device, sound source survey method, and program therefor

    JP2017151076A

  • An Apparatus, Method and Computer Program for Audio Signal Processing

    US20210076130A1

  • Wave-source-direction estimation device, wave-source-direction estimation method, and program storage medium

    WO2020003342A1