Spatial Audio Representation and Rendering
The system addresses the issue of excessive decorrelation artifacts in spatial audio rendering by controlling decorrelated energy based on spatial metadata, enhancing audio quality and spatial perception in complex sound scenes.
Patent Information
- Application Number
- JP2022572609
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-05-27
- Filing Date
- 2021-05-07
- Publication Date
- 2025-11-05
- Estimated Expiration
- 2041-05-07
AI Technical Summary
Existing spatial audio rendering techniques often result in overly reverberant or lacking spatiality, degrading audio quality, particularly for speech, due to excessive decorrelation artifacts.
A system that controls the amount of decorrelated audio signal based on spatial metadata and target characteristics to balance decorrelation artifacts with spatial perception, using mixing matrices and covariance matrices to generate output audio signals that maintain spatial coherence and reduce unnecessary decorrelation.
Improves audio quality by minimizing decorrelation artifacts while preserving spatiality, especially in complex sound scenes with multiple sources, by dynamically adjusting decorrelated energy based on spatial metadata.
Smart Images

Figure 0007764402000013 
Figure 0007764402000014 
Figure 0007764402000015
Abstract
Description
[Technical Field]
[0001] This application relates to an apparatus and method for spatial audio representation and rendering, but is not limited to audio representation for audio decoders. [Background technology]
[0002] Immersive audio codecs are implemented to support a variety of operating points, from low bitrate operation to transparency. An example of such a codec is the Immersive Voice and Audio Services (IVAS) codec, which is suitable for use in communication networks such as 3GPP® 4G / 5G networks and is also intended for use in immersive services such as virtual reality (VR) immersive voice and audio. This audio codec is expected to support the encoding, decoding, and rendering of speech, music, and general audio. Furthermore, it is expected to support channel-based audio and scene-based audio input, including spatial information about the sound field and sound source. It is also expected to operate with low latency to enable conversational services and support high error robustness under various transmission conditions.
[0003] Input signals can be presented to the IVAS encoder in one of several supported formats (and several permissible combinations of formats). For example, a mono audio signal (without metadata) can be encoded using the Enhanced Audio Services (EVS) encoder. Other input formats can take advantage of new encoding tools in IVAS. One input format proposed for IVAS is the Metadata-Assisted Spatial Audio (MASA) format, which allows the encoder to combine mono and stereo encoding tools with metadata encoding tools for efficient transmission of the format. MASA is a parametric spatial audio format suitable for spatial audio processing. Parametric spatial audio processing is a branch of audio signal processing that describes the spatial aspects of a sound (or sound field) using a set of parameters. For example, in parametric spatial audio capture from a microphone array, it is a typical and useful choice to estimate a set of parameters from the microphone array signal, such as the direction of the sound in a frequency band, expressed as a direct-to-whole ratio or ambient-to-whole energy ratio in a frequency band, and the relative energy of the directional and non-directional parts of the captured sound in a frequency band. These parameters are known to well describe the perceptual spatial characteristics of the sound captured at the microphone array location, and can be used for spatial sound synthesis according to binaural headphones, loudspeakers, or other formats such as Ambisonics.
[0004] For example, there may be a two-channel (stereo) audio signal and spatial metadata. The spatial metadata may further define parameters such as a direction index describing the direction of sound arrival in the time-frequency parameter interval, a level / phase difference, a direct-to-total energy ratio describing the energy ratio with respect to the direction index, coherence such as diffuseness and diffuse coherence describing the spread of energy with respect to the direction index, a diffuse-to-total energy ratio describing the energy ratio of omnidirectional sound with respect to the surrounding direction, a surround coherence describing the coherence of omnidirectional sound with respect to the surrounding direction, a reverberant-to-total energy ratio describing the energy ratio of reverberation (such as microphone noise) to meet the requirement that the sum of the energy ratios be 1, a distance describing the distance of the sound originating from the direction index in meters on a logarithmic scale, a covariance matrix for multi-channel loudspeaker signals, or data related to these covariance matrices, such as central prediction coefficients, other parameters guiding specific decoders for 1-to-2 decoding coefficients (used in MPEG Surround, etc.). Any of these parameters may be determined in frequency bands.
[0005] The rendering of parametric spatial audio (i.e., audio signal(s) and associated spatial metadata, e.g., a MASA stream) into a binaural (or other) output is known. A typical situation is one in which there are two audio channel signals in a stream along with metadata. There may be one or two (or more) directions per time-frequency interval in the metadata.
[0006] Vilkamo, J., Backstrom, T. and Kuntz, A., 2013. Optimized covariance domain framework for time-frequency processing of spatial audio. Journal of the Audio Engineering Society, 61(6), pp. 403-411, presented a method particularly suited to spatial audio rendering, in which the covariance matrix of the input signal is estimated in frequency bands and the target covariance matrix of the output signal is determined based on spatial metadata. Based on these matrices, a least-squares optimized mixing matrix is determined in the frequency bands that, when applied to the audio signal, produces an output signal with the desired target covariance matrix characteristics. Furthermore, if the target covariance matrix requires more incoherent signal components than are available in the input signal, the input signal can be further decorrelated to obtain a "residual signal," which, when mixed with the output signal, provides the required incoherence at the output. Summary of the Invention
[0007] According to a first aspect, there is provided an apparatus having means configured to receive a spatial audio signal, the spatial audio signal comprising at least one audio signal and spatial metadata associated with the at least one audio signal, generate at least one decorrelated audio signal based on the at least one audio signal, and determine at least one control parameter configured to control an amount of the at least one decorrelated audio signal in at least two output audio signals for spatial audio reproduction, the at least one control parameter being based at least on at least one further target characteristic of the at least two output audio signals and the spatial metadata and the at least one characteristic determined based on the at least one audio signal, and generate at least two output audio signals for spatial audio reproduction based on the spatial audio signal and the at least one decorrelated audio signal, wherein an amount of the at least one decorrelated audio signal in the at least two output audio signals is controlled based on the at least one control parameter.
[0008] The at least one control parameter may include at least one of: at least one processing gain applied to at least one of the at least one decorrelated audio signal or the at least one audio signal that has been decorrelated; at least one mixing matrix configured to control mixing of the at least one decorrelated audio signal and the at least one audio signal; at least one mixing matrix and at least one residual mixing matrix configured to control mixing of the at least one decorrelated audio signal and the at least one audio signal; and at least one covariance matrix configured to control generation of the at least one mixing matrix and / or the at least one residual mixing matrix, wherein the at least one mixing matrix and / or the at least one residual mixing matrix is configured to control mixing of the at least one decorrelated audio signal and / or the at least one audio signal.
[0009] The means configured to determine at least one control parameter configured to control the amount of at least one decorrelated audio signal in the at least two output audio signals for spatial audio reproduction may be further configured to: determine at least one further characteristic based on the at least one audio signal; determine at least one further target characteristic of the at least two output audio signals; determine at least one first control parameter based on the at least one further characteristic based on the at least one audio signal and the at least one further target characteristic of the at least two output audio signals; determine at least one second control parameter or modify the at least one first control parameter based on the spatial metadata and at least one of the at least one characteristic determined based on the at least one audio signal.
[0010] The means configured to generate at least two output audio signals for spatial audio reproduction may be further configured to mix the at least one audio signal and the at least one uncorrelated audio signal based on the at least one first control parameter and the at least one second control parameter or the at least one modified first control parameter.
[0011] The means may be further configured to output at least two output audio signals for spatial audio reproduction.
[0012] The means may be configured to determine the at least one second control parameter or the modified at least one first control parameter based on the at least one direct-to-total energy ratio parameter in the spatial metadata.
[0013] The at least one further characteristic based on the at least one audio signal may be a covariance, and the at least one further target characteristic of the at least two output audio signals may be a target covariance of the at least two output audio signals.
[0014] The means configured to determine the at least one second control parameter or modify the at least one first control parameter may be configured to determine a residual covariance characteristic based on the covariance characteristics of the at least two output audio signals and the target covariance characteristic, and to process the residual covariance characteristic based on spatial metadata associated with the at least one audio signal.
[0015] The means configured to process the residual covariance characteristic based on spatial metadata associated with the at least one audio signal may be configured to attenuate the residual covariance characteristic if the spatial metadata indicates that the at least one audio signal is highly directional, and to pass the residual covariance characteristic unprocessed if the spatial metadata indicates that the at least one audio signal is entirely ambient.
[0016] The means configured to determine a target covariance of the at least two output audio signals may be further configured to generate a total energy estimate based on the covariance characteristics, determine head-related transfer function data based on directional parameters from metadata associated with the at least one audio signal, and determine the target covariance characteristics of the at least two output audio signals further based on the head-related transfer function data and the total energy estimate.
[0017] The means may be configured to determine at least one characteristic based on the at least one audio signal, the at least one characteristic being an audio type, and the means configured to determine at least one control parameter configured to control the amount of at least one uncorrelated audio signal in the at least two output audio signals for spatial audio reproduction may be further configured to determine whether the audio type is a determined audio type, and to determine the at least one control parameter based on the audio type being the determined audio type.
[0018] The determined audio type may be speech.
[0019] The at least one audio signal may comprise a transport audio signal generated by an encoder.
[0020] According to a second aspect, a method may include receiving a spatial audio signal, the spatial audio signal including at least one audio signal and spatial metadata associated with the at least one audio signal; generating at least one decorrelated audio signal based on the at least one audio signal; determining at least one control parameter configured to control an amount of the at least one decorrelated audio signal in at least two output audio signals for spatial audio reproduction, the at least one control parameter being based at least on at least one of further target characteristics of the at least two output audio signals, the spatial metadata, and the at least one property determined based on the at least one audio signal; and generating at least two output audio signals for spatial audio reproduction based on the spatial audio signal and the at least one decorrelated audio signal, the amount of the at least one decorrelated audio signal in the at least two output audio signals being controlled based on the at least one control parameter.
[0021] The at least one control parameter may include at least one of: at least one processing gain applied to at least one of the at least one decorrelated audio signal or the at least one audio signal that has been decorrelated; at least one mixing matrix configured to control mixing of the at least one decorrelated audio signal and the at least one audio signal; at least one mixing matrix and at least one residual mixing matrix, where the at least one mixing matrix and the at least one residual mixing matrix are configured to control mixing of the at least one decorrelated audio signal and the at least one audio signal; and at least one covariance matrix configured to control generation of the at least one mixing matrix and / or the at least one residual mixing matrix, where the at least one mixing matrix and / or the at least one residual mixing matrix is configured to control mixing of the at least one decorrelated associated audio signal and / or the at least one audio signal.
[0022] Determining at least one control parameter configured to control the amount of at least one decorrelated audio signal in the at least two output audio signals for spatial audio reproduction may further include: determining at least one further property based on the at least one audio signal; determining at least one further target property of the at least two output audio signals; determining at least one first control parameter based on the at least one further property based on the at least one audio signal and the at least one further target property of the at least two output audio signals; and determining at least one second control parameter or modifying the at least one first control parameter based on at least one of the spatial metadata and the at least one property determined based on the at least one audio signal.
[0023] Generating at least two output audio signals for spatial audio reproduction may further include mixing the at least one audio signal and the at least one uncorrelated audio signal based on the at least one first control parameter and the at least one second control parameter, or the at least one modified first control parameter.
[0024] The method may further include outputting at least two output audio signals for spatial audio reproduction.
[0025] The method may further include determining at least one second control parameter or the modified at least one first control parameter based on the at least one direct-to-total energy ratio parameter in the spatial metadata.
[0026] The at least one further characteristic based on the at least one audio signal may be a covariance characteristic, and the at least one further target characteristic of the at least two output audio signals may be a target covariance characteristic of the at least two output audio signals.
[0027] Determining the at least one second control parameter or modifying the at least one first control parameter may include determining a residual covariance characteristic based on covariance characteristics and a target covariance characteristic of the at least two output audio signals, and processing the residual covariance characteristic based on spatial metadata associated with the at least one audio signal.
[0028] Processing the residual covariance characteristic based on spatial metadata associated with the at least one audio signal may include attenuating the residual covariance characteristic if the spatial metadata indicates that the at least one audio signal is highly directional, and passing the residual covariance characteristic unprocessed if the spatial metadata indicates that the at least one audio signal is entirely ambient.
[0029] Determining target covariance properties of the at least two output audio signals may further include generating a total energy estimate based on the covariance properties; determining head-related transfer function data based on directional parameters from metadata associated with the at least one audio signal; and further determining target covariance properties of the at least two output audio signals based on the head-related transfer function data and the total energy estimate.
[0030] The method may further include determining at least one characteristic based on the at least one audio signal, the at least one characteristic being an audio type, and determining at least one control parameter configured to control an amount of the at least one decorrelated audio signal in the at least two output audio signals for spatial audio reproduction may further comprise: determining whether the audio type is the determined audio type; and determining the at least one control parameter based on the audio type being the determined audio type.
[0031] The determined type of audio may be voice.
[0032] The at least one audio signal may include a transport audio signal generated by an encoder.
[0033] According to a third aspect, there is provided an apparatus comprising at least one processor and at least one memory including computer program code, wherein the at least one memory and the computer program code are configured by the at least one processor to cause the apparatus to perform at least: receiving a spatial audio signal, the spatial audio signal including at least one audio signal and spatial metadata associated with the at least one audio signal; generating at least one decorrelated audio signal based on the at least one audio signal; and determining at least one control parameter configured to control an amount of the at least one decorrelated audio signal in at least two output audio signals for spatial audio reproduction, the at least one control parameter being based on at least one of at least one further target characteristic of the at least two output audio signals and the at least one property determined based on the spatial metadata and the at least one audio signal; and generating at least two output audio signals for spatial audio reproduction based on the spatial audio signal and the at least one decorrelated audio signal, wherein an amount of the at least one decorrelated audio signal in the at least two output audio signals is controlled based on the at least one control parameter.
[0034] The at least one control parameter may comprise at least one of: at least one processing gain applied to the at least one decorrelated audio signal or at least one of the at least one audio signal that has been decorrelated; at least one mixing matrix configured to control mixing of the at least one decorrelated audio signal and the at least one audio signal; at least one mixing matrix and at least one residual mixing matrix configured to control mixing of the at least one decorrelated audio signal and the at least one audio signal; and at least one covariance matrix configured to control generation of the at least one mixing matrix and / or the at least one residual mixing matrix, wherein the at least one mixing matrix and / or the at least one residual mixing matrix is configured to control mixing of the at least one decorrelated audio signal and / or the at least one audio signal.
[0035] The apparatus for determining at least one control parameter configured to control the amount of at least one decorrelated audio signal in at least two output audio signals for spatial audio reproduction may further perform the following: determining at least one further property based on the at least one audio signal; determining at least one further target property of the at least two output audio signals; determining at least one first control parameter based on the at least one further property based on the at least one audio signal and the at least one further target property of the at least two output audio signals; and determining at least one second control parameter or modifying the at least one first control parameter based on at least one of the spatial metadata and the at least one property determined based on the at least one audio signal.
[0036] The device configured to generate at least two output audio signals for spatial audio reproduction may further be configured to mix the at least one audio signal and the at least one uncorrelated audio signal based on the at least one first control parameter and the at least one second control parameter or the at least one modified first control parameter.
[0037] The device may further be adapted to output at least two output audio signals for spatial audio reproduction.
[0038] The apparatus may further be adapted to determine the at least one second control parameter or the modified at least one first control parameter based on the at least one direct-to-total energy ratio parameter in the spatial metadata.
[0039] The at least one further characteristic based on the at least one audio signal may be a covariance characteristic, and the at least one further target characteristic of the at least two output audio signals may be a target covariance characteristic of the at least two output audio signals.
[0040] The device adapted to determine the at least one second control parameter or modify the at least one first control parameter may be adapted to determine a residual covariance characteristic based on the at least one first control parameter and a target covariance characteristic for the at least two output audio signals, and to process the residual covariance characteristic based on spatial metadata associated with the at least one audio signal.
[0041] An apparatus adapted to process residual covariance characteristics based on spatial metadata associated with at least one audio signal may be adapted to attenuate the residual covariance characteristics if the spatial metadata indicates that the at least one audio signal is highly directional, and to pass the residual covariance characteristics through unprocessed if the spatial metadata indicates that the at least one audio signal is entirely ambient.
[0042] The apparatus configured to determine target covariance characteristics of at least two output audio signals may further be configured to generate a total energy estimate based on the covariance characteristics, determine head-related transfer function data based on directional parameters from metadata associated with the at least one audio signal, and further determine the target covariance characteristics of the at least two output audio signals based on the head-related transfer function data and the total energy estimate.
[0043] The apparatus may be further configured to determine at least one characteristic based on the at least one audio signal, the at least one characteristic being an audio type, and the apparatus configured to determine at least one control parameter configured to control an amount of at least one uncorrelated audio signal in the at least two output audio signals for spatial audio reproduction may further be configured to determine whether the audio type is a determined audio type, and to determine the at least one control parameter based on the audio type being the determined audio type.
[0044] The determined audio type may be speech.
[0045] The at least one audio signal may include a transport audio signal generated by an encoder.
[0046] According to a fourth aspect, there is provided an apparatus comprising: a receiving circuit configured to receive a spatial audio signal, the spatial audio signal including at least one audio signal and spatial metadata associated with the at least one audio signal; a generating circuit configured to generate at least one decorrelated audio signal based on the at least one audio signal; a determining circuit configured to determine at least one control parameter configured to control an amount of the at least one decorrelated audio signal in at least two output audio signals for spatial audio reproduction, the at least one control parameter being based at least on at least one of: a further target characteristic of the at least two output audio signals and the at least one characteristic based on the spatial metadata and the at least one audio signal; and a generating circuit configured to generate at least two output audio signals for spatial audio reproduction based on the spatial audio signal and the at least one decorrelated audio signal, the amount of the at least one decorrelated audio signal in the at least two output audio signals being controlled based on the at least one control parameter.
[0047] According to a fifth aspect, there is provided a computer program comprising instructions (or a computer-readable medium comprising program instructions) to cause an apparatus to at least: receive a spatial audio signal, the spatial audio signal comprising at least one audio signal and spatial metadata associated with the at least one audio signal; generate at least one decorrelated audio signal based on the at least one audio signal; determine at least one control parameter configured to control an amount of the at least one decorrelated audio signal in at least two output audio signals for spatial audio reproduction, the at least one control parameter being based at least on at least one further target characteristic of the at least two output audio signals and the spatial metadata and the at least one characteristic determined based on the at least one audio signal; and generate at least two output audio signals for spatial audio reproduction based on the spatial audio signal and the at least one decorrelated audio signal, the amount of the at least one decorrelated audio signal in the at least two output audio signals being controlled based on the at least one control parameter.
[0048] According to a sixth aspect, there is provided a non-transitory computer-readable medium comprising program instructions to cause an apparatus to at least: receive a spatial audio signal, the spatial audio signal including at least one audio signal and spatial metadata associated with the at least one audio signal; generate at least one decorrelated audio signal based on the at least one audio signal; determine at least one control parameter configured to control an amount of the at least one decorrelated audio signal in at least two output audio signals for spatial audio reproduction, the at least one control parameter being based on at least one of at least one target further characteristic of the at least two output audio signals; generate at least two output audio signals for spatial audio reproduction based on the spatial metadata, the at least one property determined based on the at least one audio signal, the spatial audio signal, and the at least one decorrelated audio signal, the amount of the at least one decorrelated audio signal in the at least two output audio signals being controlled based on the at least one control parameter.
[0049] According to a seventh aspect, there is provided an apparatus comprising: means for receiving a spatial audio signal, the spatial audio signal including at least one audio signal and spatial metadata associated with the at least one audio signal; means for generating at least one decorrelated audio signal based on the at least one audio signal; means for determining at least one control parameter configured to control an amount of the at least one decorrelated audio signal in at least two output audio signals for spatial audio reproduction, the at least one control parameter being based at least on at least one further target characteristic of the at least two output audio signals and the spatial metadata and the at least one characteristic determined based on the at least one audio signal; and means for generating at least two output audio signals for spatial audio reproduction based on the spatial audio signal and the at least one decorrelated audio signal, the amount of the at least one decorrelated audio signal in the at least two output audio signals being controlled based on the at least one control parameter.
[0050] According to an eighth aspect, there is provided a computer-readable medium comprising program instructions to cause an apparatus to at least receive a spatial audio signal, the spatial audio signal including at least one audio signal and spatial metadata associated with the at least one audio signal; generate at least one decorrelated audio signal based on the at least one audio signal; and determine at least one control parameter configured to control an amount of the at least one decorrelated audio signal in at least two output audio signals for spatial audio reproduction, the at least one control parameter being based at least on at least one further target characteristic of the at least two output audio signals and the spatial metadata and the at least one characteristic determined based on the at least one audio signal; and generate at least two output audio signals for spatial audio reproduction based on the spatial audio signal and the at least one decorrelated audio signal, the amount of the at least one decorrelated audio signal in the at least two output audio signals being controlled based on the at least one control parameter.
[0051] An apparatus comprising means for carrying out the actions of the above methods.
[0052] An apparatus configured to perform the actions of the method described above.
[0053] A computer program comprising program instructions for causing a computer to carry out the above method.
[0054] A computer program product stored on the medium may cause an apparatus to perform the methods described herein.
[0055] The electronic device may include a device as described herein.
[0056] The chipset may include devices as described herein.
[0057] SUMMARY OF THE INVENTION Embodiments of the present invention aim to solve problems associated with the prior art. [Brief explanation of the drawings]
[0058] For a better understanding of the present application, reference will now be made, by way of example, to the accompanying drawings in which: [Figure 1] FIG. 1 shows a schematic diagram of a system of apparatus suitable for implementing some embodiments. [Figure 2] FIG. 2 is a flow diagram of the operation of an exemplary device according to some embodiments. [Figure 3] FIG. 3 illustrates a schematic diagram of an exemplary synthesis processor such as that shown in FIG. 1 in accordance with some embodiments. [Figure 4] FIG. 4 is a flow diagram of the operation of an exemplary synthesis processor such as that shown in FIG. 3 according to some embodiments. [Figure 5] FIG. 5 is a diagram that schematically illustrates an exemplary spatial synthesis processor such as that shown in FIG. 3 in accordance with some embodiments. [Figure 6] FIG. 6 is a flow diagram of the operation of an exemplary spatial synthesizer such as that shown in FIG. 5 according to some embodiments. [Figure 7] FIG. 7 illustrates an exemplary apparatus suitable for implementing the apparatus shown in the previous figures. DETAILED DESCRIPTION OF THE INVENTION
[0059] Rendering an audio signal as described above may produce a signal with a covariance matrix that matches the target covariance matrix, and therefore may produce good quality audio output because the spatial perception matches the target. Furthermore, decorrelated energy may be added when it is needed (i.e., when the mixing of the input signals does not provide the required incoherence). Thus, artifacts due to decorrelation (such as the perception of added reverberation) are minimized.
[0060] As used herein, the term audio signal can refer to a single audio channel or to an audio signal with two or more channels.
[0061] In many situations, for example, when the rendered audio signal primarily contains reverberation / ambience, the negative impact of (minimized) decorrelation may be negligible. However, even when decorrelation is minimized, the amount of decorrelation can degrade sound quality. That is, decorrelation is known to affect the perception of certain sounds, particularly speech, making them sound overly reverberant. Therefore, in situations where there are two sound sources in different directions, the incoherence to be synthesized may not be solely related to reverberation / ambience, but rather to generate incoherence for rendering multiple sound sources. In such cases, even with least-squares optimization, decorrelation artifacts may become audible. It may be possible to avoid excessive use of decorrelated energy by disabling its use. However, disabling the use of decorrelated energy may result in a perception of significantly reduced space and surrounds, as the output signal will not be mutually coherent and faithfully represent the ambient or reverberant sound scene.
[0062] The concepts discussed within the embodiments herein may be able to overcome any problems of complex sound scenes being rendered as being too reverberant or lacking in spaciousness and envelopment, thus degrading the audio quality.
[0063] Therefore, embodiments relate to parametric spatial sound rendering. Spatial parameter estimation may be based on microphone array signals. One example of determining spatial metadata including direction and ratio parameters is directional audio coding (DirAC), as discussed in Pulkki, V., 2007, "Spatial sound reproduction with directional audio coding." Journal of the Audio Engineering Society, 55(6), pp. 503-516, which uses the primary capture signal as input. A variant of DirAC is higher-order DirAC, which can simultaneously estimate many directions. Politis, A., Vilkamo, J. and Pulkki, V., 2015, "Sector-based parametric sound field reproduction in the spherical harmonic domain," IEEE Journal of Selected Topics in Signal Processing, 9(5), pp. 852-866, provides multiple direction estimates. Many additional parameter estimation methods exist, any of which may be implemented in some embodiments, for example, UK published patent application GB1619573.7 describes a suitable means for obtaining 360 / 3D spatial metadata from horizontally flat devices such as mobile phones. Any of the known spatial metadata determination techniques may be applied in some embodiments.
[0064] The embodiments discussed herein relate, for example, to rendering of parametric audio signals (including one or more audio signals and spatial metadata) in a spatial audio decoder. The embodiments may be configured to improve upon conventional rendering techniques that use measurements of input signal characteristics to control rendering and optimize the amount of decorrelation necessary to achieve a desired spatial output. The embodiments further provide a means to control the amount of decorrelated sound applied so as to attenuate decorrelated sounds when rendering those sound scenes where remaining decorrelation is expected to have a detrimental effect on perceived audio quality, while otherwise maintaining decorrelation to maintain adequate spatiality. The reduction in decorrelation may, in some embodiments, be based on monitoring spatial metadata, with the degree to which decorrelated sound energy is attenuated being determined based on a direct-to-total energy ratio parameter.
[0065] The concepts discussed in the embodiments herein relate to spatial audio reproduction of audio signals and associated spatial metadata containing information on how to spatially render the audio signal, and embodiments are provided that can render direct sound sources (even multiple simultaneous direct sound sources) without distracting decorrelation artifacts (such as added reverberation) while maintaining the correct spaciousness and ambience for reverberant / ambient sounds. Further, these embodiments may be configured to determine input covariance characteristics of an input signal and target covariance characteristics of an output signal, determine the amount of decorrelated energy required to reach the target covariance characteristics, determine a limit on the amount of decorrelated energy based on the spatial metadata, decorrelate the input audio signal, and render a spatial output signal based on the input audio signal, the decorrelated input audio signal, the determined limit on decorrelation, and the covariance characteristics.
[0066] In some embodiments, the determined covariance characteristic is a covariance matrix of the input signal, and the target covariance characteristic is a target covariance matrix (derived based on the audio signal and associated spatial metadata). A mixing matrix may be derived based on the determined covariance characteristic. Further, some embodiments may be configured to determine the amount of uncorrelated energy necessary to achieve the non-coherence characteristic of the target covariance matrix. Next, some embodiments may be configured to limit the amount of uncorrelated energy based on the spatial metadata. For example, if the spatial metadata includes a direct-to-total energy ratio, the maximum amount of uncorrelated energy may be limited using a factor 1-sum (direct-to-total energy ratio). Finally, in some embodiments, a spatial audio signal (e.g., a binaural audio signal) is rendered using the input audio signal, the uncorrelated input audio signal, the constraint information, and the mixing matrix.
[0067] In some embodiments, the direct sound components can be mostly rendered using blending and / or (complex-valued) gain processing without significant decorrelation, thus avoiding decorrelation artifacts. Furthermore, in some embodiments, the ambient / reverberant components are decorrelated when necessary, thus preserving the sense of space and surroundings. As a result, embodiments can be configured to provide good audio quality by avoiding decorrelation artifacts and still maintaining the sense of space and surroundings, even in the presence of multiple direct sound sources and reverberant / ambient.
[0068] The embodiments discussed herein are designed with the knowledge that the perception of reverberant spaciousness is related to the interaural correlation provided to the listener. For example, Borss, C. and Martin, R., 2009, February, "An improved parametric model for perception-based design of virtual acoustics," in Audio Engineering Society 35th International Conference, identified that when generating binaural reverberation (which is generally an example of ambience), interaural cross-correlation must be low or zero at mid-to-high frequencies to naturally generate a spacious perception in the listener. In other words, the left and right ear signals must be appropriately decoherent. In parametric spatial audio playback, input signals may not have such decoherence, so decorrelation processing is performed to generate decoherence, resulting in an appropriate sense of spaciousness.
[0069] Furthermore, embodiments are designed with the knowledge that decorrelation affects different sounds differently. For example, Vilkamo, J. and Pulkki, V., 201, “Minimization of decorrelator artifacts in directional audio coding by covariance domain rendering,” Journal of the Audio Engineering Society, 61(9), pp. 637-646, presents a listening test that includes two methods for rendering spatial sound. The first method is the one previously defined in Vilkamo, J., Backstrom, T. and Kuntz, A., 2013, “Optimized covariance domain framework for time-frequency processing of spatial audio,” Journal of the Audio Engineering Society, 61(6), pp. 403-411, and the second method is a conventional method that does not optimize the amount of decorrelated sound energy applied. The effective difference between these methods is primarily the relative difference in the amount of decorrelated sound energy, with the former method more effectively utilizing existing independent signals in the input. Listening tests provided perceptual quality results for the two methods for different sound scenes. The results show that speech quality degrades significantly with increasing decorrelation. On the other hand, reverberation (or more generally complex background ambience) is known to be unaffected by a well-designed decorrelation procedure, since such signals are already naturally decorrelated and further decorrelation has little adverse effect on the perceptual quality of such sounds.
[0070] Thus, embodiments may be configured to introduce a beneficial balance between decorrelation (artifacts) and the perception of spaciousness (or lack thereof).
[0071] In particular, sound fields where decorrelation is expected to degrade quality are processed with a reduced amount of decorrelation. An example of such a situation is when two talkers overlap (or when a talker and another sound source overlap). In such situations, the present invention provides the significant benefit of avoiding decorrelation artifacts, even though the perception of width may be temporarily reduced.
[0072] Sounds for which the deterioration of sound quality due to decorrelation is not expected are processed with an appropriate amount of decorrelation, such as reverberation, to achieve an appropriate sense of spaciousness in such situations.
[0073] Thus, the embodiments discussed herein are configured to provide an improved balance of combining good audio quality and preserving spaciousness, where prior art techniques may only achieve one of these goals.
[0074] In some embodiments, as described in further detail below, the audio processing device is configured to receive a spatial audio signal. The spatial audio signal may include at least one audio signal and spatial metadata associated with the at least one audio signal. The audio processing device may then, in some embodiments, be configured to determine at least one covariance characteristic associated with the at least one audio signal.
[0075] A target covariance characteristic (which is a target characteristic associated with the output spatial audio signal) may be determined based on at least the spatial metadata. In some embodiments, the audio processing device may then be further configured to determine a mixing matrix (or other suitable control) based on at least one of the covariance characteristic and the target covariance characteristic.
[0076] In some embodiments, the audio processing device may further be configured to generate at least one decorrelated audio signal based on the at least one audio signal. The residual covariance characteristics may further be determined by the audio processing device based on at least one covariance characteristic, the target covariance characteristic, and the mixing matrix.
[0077] The audio processing device may then attenuate the residual covariance characteristic to attenuate the uncorrelated energy based on the spatial metadata (and may generate a processed residual covariance characteristic).
[0078] In some embodiments, a residual mixing matrix is determined by the audio processing device using the processed residual covariance characteristics and the at least one covariance characteristic.
[0079] The audio processing device may be further configured to generate at least two output signals for spatial audio reproduction by applying a mixing matrix to at least one audio signal and applying a residual mixing matrix to at least one decorrelated audio signal.
[0080] In other words, in some embodiments, the spatial audio signal may include at least one audio signal and spatial metadata associated with the at least one audio signal. At least one decorrelated audio signal based on the at least one audio signal is also generated. At least one control parameter may then be determined, the at least one control parameter being configured to control the amount of the at least one decorrelated audio signal in the at least two output audio signals for spatial audio reproduction. In some embodiments, the at least one control parameter may be determined based on at least one further target characteristic of the at least two output audio signals (e.g., a target covariance characteristic of the at least two output audio signals) and at least one characteristic (e.g., an audio type) determined based on the spatial metadata and the at least one audio signal.
[0081] Then, at least two output signals for spatial audio reproduction may be generated based on the spatial audio signal and the at least one uncorrelated audio signal, and the amount of the at least one uncorrelated audio signal in the at least two output audio signals is controlled based on the at least one control parameter.
[0082] First, an embodiment will be described with respect to an example capture (or encoder / analyzer) and playback (or decoder / synthesizer) device or system as shown in FIG.
[0083] The system 199 is shown to include a capture section (encoder / analyzer) 101 and a playback section (decoder / synthesizer) 105 .
[0084] The capture unit 101 in some embodiments comprises an audio signal input configured to receive an input audio signal 110. The input audio signal may be from any suitable source, for example, two or more microphones attached to a mobile phone, other microphone arrays such as a B-format microphone or Eigenmike, Ambisonic signals such as First Order Ambisonic (FOA), Higher Order Ambisonic (HOA), loudspeaker surround mix, and / or object. The input audio signal 110 may be provided to an analysis processor 111 and a transport signal generator 113.
[0085] The capture unit 101 may include an analysis processor 111. The analysis processor 111 is configured to perform spatial analysis on the input audio signal, resulting in appropriate metadata 112. The purpose of the analysis processor 111 is therefore to estimate spatial metadata for frequency bands. For all of the aforementioned input types, there are known methods for generating appropriate spatial metadata, such as direction and direct-to-total energy ratios in frequency bands (or similar parameters such as diffuseness, i.e., ambient-to-total ratios). These methods will not be described in detail herein, but some examples may include performing an appropriate time-frequency transform on the input signal, estimating delay values between microphone pairs that maximize inter-microphone correlation in frequency bands if the input is a mobile phone microphone array, formulating direction values corresponding to those delays (as described in UK Patent Application No. 1619573.7 and PCT Patent Application No. PCT / FI2017 / 050778), and formulating ratio parameters based on the correlation values.
[0086] The metadata can take various forms and can include spatial and other metadata. A typical parameterization of spatial metadata is one directional parameter DOA(k,n) in each frequency band and an associated direct-to-total energy ratio r(k,n) in each frequency band, where k is the frequency band index and n is the time frame index. Determining or estimating the direction and ratio depends on the device or implementation in which the audio signal is acquired. For example, the metadata can be acquired or estimated using spatial audio capture (SPAC) using the methods described in UK Patent Application No. 1619573.7 and PCT Patent Application No. PCT / FI2017 / 050778. In other words, in this particular context, spatial audio parameters include parameters intended to characterize the sound field.
[0087] The spatial metadata in some embodiments may include information for rendering the audio signal into a spatial output, such as a binaural output, a surround loudspeaker output, a crosstalk-canceling stereo output, or an ambisonic output. For example, in some embodiments, the spatial metadata may include: loudspeaker level information, Inter-loudspeaker correlation information, Information about the amount of diffuse coherent sound, Information about the amount of surround coherence, (and / or any other suitable metadata).
[0088] In some embodiments, the generated parameters may differ for each frequency band. Thus, for example, all parameters are generated and transmitted in band X, but only one of the parameters is generated and transmitted in band Y, and no parameters are generated or transmitted in band Z. As a practical example, in some frequency bands, such as the highest frequency band, some of the parameters may not be needed for perceptual reasons.
[0089] If the input is a FOA signal or a B-format microphone, the analysis processor 111 can be configured to determine parameters such as intensity vectors from which directional parameters are derived, and to compare the length of the intensity vector with an estimate of the overall sound field energy to determine a ratio parameter, a method known in the literature as directional audio coding (DirAC).
[0090] If the input is an HOA signal, the analysis processor 111 can either take an FOA subset of the signal and apply the above method, or split the HOA signal into multiple sectors and apply the above method to each of them. This sector-based method is known in the literature as High-Order DirAC (HO-DirAC). In this case, multiple directional parameters exist simultaneously for each frequency band.
[0091] If the input is a loudspeaker surround mix and / or objects, the analysis processor 111 may be configured to convert the signal into FOA signal(s) (by use of spherical harmonic encoding gain) and analyze the direction and proportion parameters as described above.
[0092] Thus, the output of the analysis processor 111 is spatial metadata determined in frequency bands. The spatial metadata may include direction and proportion in frequency bands, but may also have any of the metadata types listed above. The spatial metadata may vary over time and with frequency.
[0093] In some embodiments, the spatial analysis may be performed outside of system 199. For example, in some embodiments, spatial metadata associated with the audio signal may be provided to the encoder as a separate bitstream. In some embodiments, the spatial metadata may be provided as a set of spatial (directional) index values.
[0094] The capture unit 101 may include a transport signal generator 113. The transport signal generator 113 is configured to receive an input signal and generate an appropriate transport audio signal 114. The transport audio signal may be a stereo or mono audio signal. The generation of the transport audio signal 114 may be performed using known methods, as summarized below.
[0095] If the input is a mobile phone microphone array audio signal, the transport signal generator 113 may be configured to select a left and right microphone pair and apply appropriate processing to the signal pair, such as automatic gain control, microphone noise cancellation, wind noise cancellation, equalization, etc.
[0096] It should be noted that when the input is an FOA / HOA signal or a B-format microphone, the transport signal generator 113 may be configured to form directional beam signals directed in the left and right directions, such as two opposing cardioid signals.
[0097] If the input is a loudspeaker surround mix and / or objects, the transport signal generator 113 may be configured to generate a downmix signal that combines the left channel into a left downmix channel, and similarly combines the right channel, and adds the center channel with an appropriate gain to both transport channels.
[0098] In some embodiments, the input audio signal bypasses the transport signal generator 113, for example in some situations where analysis and synthesis are performed in the same device in a single processing step without intermediate encoding. The number of transport channels can also be any suitable number (rather than one or two channels as discussed in the examples).
[0099] In some embodiments, the capture unit 101 may include an encoder / multiplexer 115. The encoder / multiplexer 115 may be configured to receive the transport audio signal 114 and the metadata 112. The encoder / multiplexer 115 may be further configured to generate an encoded or compressed form of the metadata information and the transport audio signal. In some embodiments, the encoder / multiplexer 115 may also interleave, multiplex, or embed the metadata within the encoded audio signal before transmission or storage into a single data stream 116. The multiplexing may be implemented using any suitable scheme.
[0100] For example, the encoder / multiplexer 115 may be implemented as an IVAS encoder or any other suitable encoder, and is thus configured to encode the audio signal and metadata to form the bitstream 116 (e.g., an IVAS bitstream).
[0101] This bitstream 116 may then be transmitted / stored 103, as indicated by the dashed line. In some embodiments, the encoder / multiplexer 115 is not present (and therefore the decoder / demultiplexer 121, as discussed below, is not present).
[0102] The system 199 may further include a playback (decoder / synthesizer) unit 105 configured to receive, acquire, or otherwise obtain the bitstream 116 and generate from the bitstream an appropriate audio signal that is presented to a listener / listener playback device.
[0103] The playback unit 105 may include a decoder / demultiplexer 121 configured to receive the bitstream, demultiplex the encoded stream, and then decode the audio signal to obtain a transport signal 124 and metadata 122.
[0104] Furthermore, in some embodiments, as described above, there may be no demultiplexer / decoder 121 (e.g., when both the capture unit 101 and the playback unit 105 are in the same device and therefore there is no associated encoder / multiplexer 115).
[0105] The playback unit 105 may include a synthesis processor 123 configured to receive a transport audio signal 124, spatial metadata 122, and generate a spatial output signal 128, such as a binaural audio signal playable over headphones.
[0106] The operation of this system can be summarized in the flow diagram shown in Figure 2.
[0107] FIG. 2, for example, illustrates the receipt of an input audio signal, as shown in step 201.
[0108] Next, the flow diagram shows in FIG. 2, by step 203, the analysis (spatial) of the input audio signal to generate spatial metadata.
[0109] A transport audio signal is then generated from the input audio signal, as shown in step 204 in FIG.
[0110] The generated transport audio signal and metadata may then be encoded and / or multiplexed, as shown in Figure 2 at step 205. This is shown in Figure 2 as an optional dashed box.
[0111] Further, in Figure 2, the encoded and / or multiplexed signal may be demultiplexed and / or decoded to generate a transport audio signal and spatial metadata, as indicated by step 207. This is also shown as an optional dashed box.
[0112] A spatial audio signal may then be synthesized based on the transport audio signal and the spatial metadata, as shown in step 209 in FIG.
[0113] The synthesized spatial audio signal may then be output to a suitable output device, such as a set of headphones, as shown in step 211 in FIG.
[0114] With reference to FIG. 3, the synthesis processor 123 is shown in further detail.
[0115] In some embodiments, the synthesis processor 123 comprises a forward filter bank (time-frequency transformer) 311 configured to receive the (time-domain) transport audio signal 124 and transform it into the time-frequency domain. Suitable forward filters or transforms include, for example, a short-time Fourier transform (STF) and a complex-modulated quadrature mirror filter bank (QMF). The resulting signal is i It can be expressed as (b,n), where i is the channel index, b is the frequency bin index of the time-frequency transform, and n is the time index. The time-frequency signal is expressed here, for example, in vector form (e.g., for two channels, the vector form is as follows:
number
[0116] The following processing operations may then be performed in the time-frequency domain across frequency bands. A frequency band may be one or more frequency bins (individual frequency components) of the applied time-frequency transformer (filter bank). In some embodiments, the frequency band may approximate a perceptually relevant resolution such as a Bark frequency band, which has spectrally higher selectivity at low frequencies than at high frequencies. Alternatively, in some embodiments, the frequency band may correspond to a frequency bin. The frequency band may be one for which spatial metadata has been determined (or approximated) by the analysis processor. Each frequency band k is a frequency band that is a subset of the lowest frequency bin b low (k) and the highest frequency bin b high It may be defined in terms of (k).
[0117] The time-frequency transport signal 302 in some embodiments may be provided to a spatial synthesizer 313 .
[0118] In some embodiments, the synthesis processor 123 includes a spatial synthesizer 313 configured to receive the time-frequency domain transport signal 302 and the spatial metadata 122 and to generate the spatial-time-frequency audio signal 304 by processing the time-frequency transport signal 302 based on the spatial metadata 122.
[0119] The synthesis processor 123 in some embodiments comprises an inverse filterbank 315 configured to receive the spatial-temporal frequency domain audio signal 304 and apply an inverse transform corresponding to the transform applied by the forward filterbank 311 to generate the time-domain spatial output signal 128. The output of the inverse filterbank 315 may therefore be a spatial output signal, e.g., a binaural audio signal for headphone listening.
[0120] The operation of this synthesis processor 123 can be summarized in a flow chart as shown in FIG.
[0121] FIG. 4, for example, illustrates the reception of an audio signal and spatial metadata, as shown in step 401.
[0122] Then, as shown in step 403 in FIG. 4, the audio signal is transformed into a time-frequency domain to generate an audio signal in the time-frequency domain.
[0123] Next, as shown in step 405 in FIG. 4, the time-frequency domain audio signal is processed based on the spatial metadata to generate a spatial-time-frequency domain audio signal.
[0124] Next, as shown in FIG. 4 at step 407, the spatial-time-frequency domain audio signal may be inverse transformed to generate a spatial (time-domain) audio signal.
[0125] The synthesized spatial audio signal may then be output, as shown in step 409 in FIG.
[0126] An example of the spatial synthesizer 313 of Figure 3 is shown in more detail in Figure 5. In the following example, the audio signal contains two channels, one "left" and one "right" channel. However, it will be understood that there are embodiments in which the same method can be implemented for any number of channels by a person skilled in the art without further inventive input.
[0127] 5, the time-frequency audio signal 302 may be provided to a mixer 531, a decorrelator 521, and a covariance matrix estimator 501. The spatial metadata 122 is provided to a target covariance matrix determiner 503 and a decorrelated (residual) energy attenuator 509.
[0128] In some embodiments, the spatial synthesizer 313 includes a covariance matrix estimator 501. The covariance matrix estimator 501 is configured to receive the time-frequency audio signal 302 and to estimate the covariance matrix of the time-frequency audio signal and its total energy estimate (in frequency bands). The covariance matrix may, for example, in some embodiments, be estimated as follows:
[0129]
number
[0130] where the superscript H denotes the complex conjugate, and b low (k) and b high (k) are the lowest and highest bin indices of frequency band k. The frequency bins may, in some embodiments, be the bins of the applied time-frequency transform, and frequency bands are typically configured to include a larger number of bins toward higher frequencies. The frequency bands may be such that spatial metadata has been determined. In some embodiments, C x (k,n) is averaged over time using an FIR or IIR (or any) window. The estimated covariance matrix 502 may, in some embodiments, be output to a target covariance matrix determiner 503, a residual covariance matrix determiner 505, a mixing matrix determiner 507, and a residual mixing matrix determiner 511.
[0131] In some embodiments, the spatial synthesizer 313 includes a target covariance matrix estimator 503. The target covariance matrix estimator 503 is configured to receive the estimated covariance matrix 502 and the spatial metadata 122. In this example, the spatial metadata includes one or more directional parameters DOA(k,n,p) for each frequency index k and time index n, where p=1···P, and P is the number of directional parameters (for a given time and frequency). In some embodiments, P may vary as a function of frequency and / or time, and in some embodiments, P may be constant, e.g., 1 or 2. In this example, the spatial metadata further comprises a direct overall ratio parameter r(k,n,p) that indicates the amount of energy associated with the direction DOA(k,n,p) compared to the overall sound energy. With this definition, JPEG0007764402000003.jpg12153 is established.
[0132] The target covariance matrix determiner 503 in some embodiments first calculates C x (k,n),k). In some embodiments, this value may be determined in or obtained from the covariance matrix estimator 501. In some embodiments where the output of the processing is a binaural audio signal, the target covariance matrix determiner 503 is configured to form, for each DOA(k,n,p),k, a head-related transferred function (HRTF) 2x1 column vector h(DOA(k,n,p),k) that includes the complex responses (amplitude and phase) of the left and right ear for the given DOA(k,n,p) and corresponds to the frequency (e.g., center frequency) of band k. In some embodiments, the diffuse-field binaural covariance matrix is determined by the directional DOA(k,n,p). d It may be obtained by choosing a uniform spatial distribution of (d=1···D) and by the following method:
[0133]
number
[0134] The target covariance matrix determiner in some embodiments is then configured to determine the target covariance matrix as follows:
[0135]
number
[0136] The target covariance matrix may then be output to a residual covariance matrix determiner 505 and a mixing matrix determiner 507 in some embodiments.
[0137] In some embodiments, the spatial synthesizer 313 includes a mixing matrix determiner 507. The mixing matrix determiner 507 is configured to receive the target covariance matrix 504 and the estimated covariance matrix 502. In some embodiments, the mixing matrix determiner 507 is configured to determine a mixing matrix. In some embodiments, this determination may employ the method described in Vilkamo, J., Backstrom, T. and Kuntz, A., 2013, “Optimized covariance domain framework for time-frequency processing of spatial audio”, Journal of the Audio Engineering Society, 61(6), pp. 403-411. This method utilizes a prototype matrix, e.g., in the case of binaural reproduction, JPEG0007764402000006.jpg19147. Also, when tracking the user's head direction, if the user is facing backwards (more than 90 degrees left or right), the prototype matrix should be set to JPEG0007764402000007.jpg17148. In summary, the embodiment uses the covariance matrix C x When applied to an input signal with (k,n), the target covariance matrix C yThe residual signal is then formulated as follows: the mixing matrix determiner 507 is configured to provide a mixing matrix M(k,n) that provides an output signal with a covariance matrix similar to the prototype signal Qx(b,n). This mixing solution may be least-squares optimized with respect to the prototype signal Qx(b,n). The formulation of the mixing matrix may, in some embodiments, be regularized to avoid arbitrarily large amplification of small independent signal components; thus, in practice, in many situations the target covariance matrix is not fully achieved. For this reason, the residual signal is formulated as described below. The mixing matrix determiner 507 is configured to output the mixing matrix M(k,n) 508 to the mixer 531 and the residual covariance matrix determiner 505.
[0138] In some embodiments, the spatial synthesizer 313 includes a residual covariance matrix determiner 505. The residual covariance matrix determiner 505 determines the estimated covariance matrix C x (k,n)502, the target covariance matrix C y (k,n) 504, and a mixing matrix M(k,n) 508. The residual covariance matrix determiner 505 is configured to determine a residual covariance matrix, which is formulated as follows:
[0139]
number
[0140] In other words, the residual covariance matrix is the target covariance matrix C y The residual covariance matrix determiner 505 determines the residual covariance matrix C r (k,n) 506 to a correlation (residual) energy attenuator 509 .
[0141] In some embodiments, the spatial synthesizer 313 comprises a decorrelation (residual) energy attenuator 509. The decorrelation (residual) energy attenuator 509 is configured to attenuate the residual mixing matrix C r(k,n) 506 and spatial metadata 122. The decorrelated (residual) energy attenuator 509 is configured to generate a processed residual covariance matrix 510. The residual signal is generated based on a decorrelated version of the input signal (described further below) because, if the target covariance matrix indicates so, a new independent signal is needed to achieve incoherence. However, the need for incoherent synthesis of the output signal can stem from a number of reasons. One possible reason is the presence of actual ambience or reverberation, and another possible reason is the active presence of multiple simultaneous sources. If the residual signal is not synthesized, the ambience will sound less spatial. Also, in some situations, if the residual signal is fully synthesized, the sound quality degradation caused by decorrelation will occur for more directional sounds. Therefore, the decorrelated (residual) energy attenuator 509 is configured to process or modify the residual covariance matrix based on the spatial metadata. For example, the modification in some embodiments may be as follows:
[0142]
number
[0143] In this example, the covariance matrix is determined at the same temporal resolution as the metadata (e.g., ratio) parameters. In some embodiments, the metadata may be determined at a different temporal resolution, e.g., multiple time indexes of the metadata contribute to one time index of the covariance matrix. In such cases, it is, for example, optional to take a time average (or energy-weighted time average) of the ratio parameters prior to this illustrated formula to modify the residual covariance matrix.
[0144] So, for example, if the sound is entirely ambient, the residual covariance matrix is raw, and if the sound is only directional, the residual covariance matrix is zero. Therefore, the decorrelation (residual) energy attenuator reduces the residual covariance matrix C' to the processed residual covariance matrix C'. r(k,n) 510 to provide the residual mixing matrix determiner 511 .
[0145] In some embodiments, the spatial synthesizer 313 includes a residual mixing matrix determiner 511. The residual mixing matrix determiner 511 determines the processed residual covariance matrix C' r (k,n) 510 and the estimated covariance matrix C x (k,n) 502. The residual mixing matrix determiner 511 operates in a similar manner to the mixing matrix determiner 507, but is configured to receive the covariance C x Instead of the (k,n) matrix 502, we use a diagonalized version of the input covariance matrix. In other words, this matrix has the covariance matrix C on its diagonal. x (k,n) 502 entries, but zeros elsewhere. This is because the residual mixing matrix is formulated to process a decorrelated version of the input signal. Furthermore, the target covariance matrix in this case is the processed residual covariance matrix C' r (k, n) 510. The remaining processing is the same as that of the mixing matrix determiner 507. The residual mixing matrix determiner 511 determines the residual mixing matrix 512 (M r (k, n) to the mixer 531.
[0146] In some embodiments, the spatial synthesizer 313 comprises a decorrelator 521 configured to receive the time-frequency audio signal x(b,n) 302 and generate a decorrelated d(b,n) version of it 522. The decorrelated audio signal d(b,n) 522 is then passed to a mixer 531.
[0147] In some embodiments, the spatial synthesizer 313 includes a mixer 531. The mixer 531 receives the time-frequency audio signal 302 and the decorrelated audio signal d(b,n) 522 and generates a mixing matrix 508M(k,n) and a residual mixing matrix M r (k,n) 512. The mixer 531 may generate an output, for example, as follows:
[0148]
number
[0149] The operation of spatial synthesizer 313 can be summarized in a flow diagram as shown in FIG.
[0150] In FIG. 6, as shown in step 601, inputs such as audio signals and spatial metadata are received.
[0151] In FIG. 6, the next operation is to estimate the covariance matrix, as shown in step 603.
[0152] Then, as shown in FIG. 6 at step 605, a target covariance matrix is generated based on the spatial metadata and the estimated covariance matrix.
[0153] Then, as shown in FIG. 6 at step 607, a mixing matrix is determined based on the estimated covariance matrix and the target covariance matrix.
[0154] Next, in FIG. 6, as shown in step 609, a residual covariance matrix is determined based on the covariance matrix, the target covariance matrix, and the mixing matrix.
[0155] In FIG. 6, as shown in step 611, after determining the residual covariance matrix, a processed residual covariance matrix is determined based on the residual covariance matrix and the spatial metadata.
[0156] Next, as shown in FIG. 6 at step 613, a residual mixing matrix is determined based on the processed residual covariance matrix and covariance matrix.
[0157] This produces a decorrelated audio signal, as shown in step 604 in FIG.
[0158] Then, as shown in FIG. 6 at step 615, a space-time-frequency audio signal is determined based on the time-frequency audio signal, the decorrelated audio signal, the mixing matrix, and the residual mixing matrix.
[0159] The space-time-frequency audio signal is then output as shown in step 617 in FIG.
[0160] The above describes processing audio signals in frequency bands. In some embodiments, all processing is performed in frequency bins. In such embodiments, all matrices, HRTFs, and other values are determined for each frequency bin. For example, spatial metadata is defined in frequency band k, so when selecting a DOA value (or any other metadata) for bin b, the DOA value for band k in which bin b resides is selected.
[0161] In some embodiments, the above procedure may be adapted for spatial outputs other than binaural audio signals. For example, the target covariance matrix may be determined based on a vector containing loudspeaker amplitude panning gains instead of HRTFs. Furthermore, for loudspeaker outputs, the diffuse sound field covariance matrix is diagonal.
[0162] In the above formulation, for simplicity of expression, we assumed that the temporal resolution of the time-frequency signal is the same as the temporal resolution of the spatial metadata. This may be true if the time-frequency transform has many bins, for example, when using a 2048-point short-time Fourier transform (STFT). In other embodiments, the filter bank may be, for example, a 60-bin complex-modulated quadrature mirror filter (QMF) bank, which results in much higher temporal resolution. In such embodiments, the metadata is not present at every temporal index n, but the indices associated with the metadata are more spaced apart (in time).
[0163] In some embodiments, the amount of uncorrelated energy can be limited using the following formula:
[0164]
number
[0165] where tr() is the trace of the matrix. In a practical implementation of such an embodiment, the total energy We limit the amount of decorrelation energy to be JPEG0007764402000012.jpg10149. As explained earlier, other formulas for limiting decorrelation can be used.
[0166] In embodiments as discussed herein, limiting the amount of decorrelated audio signals (in the decorrelated (residual) energy attenuator 509) is based on metadata. However, in some embodiments, limiting the amount of decorrelated audio signals to be present in the spatial output signal (or, in other words, attenuating the decorrelated audio signals) can be based on signal analysis. For example, the audio signal may be analyzed to determine whether it consists of a substantial speech component or other signal type known to cause a particular degradation in perceived audio quality. Thus, some embodiments include an audio type analyzer configured to determine the type of audio signal (e.g., speech), which can be used as an input to the decorrelated (residual) energy attenuator 509 to enable attenuation of the decorrelated (residual) signals. For example, if speech is detected, the amount of decorrelation can be attenuated by half. In such cases, it is also possible to perform suppression of decorrelated sounds based further on spatial metadata or without considering spatial metadata.
[0167] In the above embodiment, the suppression of uncorrelated sounds is implemented as a separate uncorrelated (residual) energy attenuator 509. This block has been described as performing the suppression by suppressing the residual covariance matrix, which subsequently reduces uncorrelated sounds in the spatial output signal. It will be apparent that the attenuation can also be performed in ways other than attenuating the residual covariance matrix, for example by attenuating the input signal to the decorrelator 521, by attenuating the output signal of the decorrelator 521, or by attenuating the residual mixing matrix 512.
[0168] 7 illustrates an exemplary electronic device that may be used as any of the device portions of the system as described above. The device may be any suitable electronic device or device. For example, in some embodiments, device 1700 is a mobile device, user equipment, tablet computer, computer, audio playback device, etc. The device may be configured to implement, for example, encoder / analyzer unit 101 and / or decoder / synthesizer unit 105 as shown in FIG. 1, or any of the functional blocks as described above.
[0169] In some embodiments, device 1700 has at least one processor or central processing unit 1707. Processor 1707 may be configured to execute various program code, such as the methods described herein.
[0170] In some embodiments, device 1700 comprises memory 1711. In some embodiments, at least one processor 1707 is coupled to memory 1711. Memory 1711 may be any suitable storage means. In some embodiments, memory 1711 has a program code section for storing program code implementable by processor 1707. Additionally, in some embodiments, memory 1711 may further comprise a storage data section for storing data, e.g., data that has been processed or to be processed in accordance with embodiments as described herein. The implemented program code stored in the program code section and the data stored in the storage data section may be retrieved by processor 1707 whenever needed via the memory-processor coupling.
[0171] In some embodiments, device 1700 comprises a user interface 1705. User interface 1705, in some embodiments, may be coupled to processor 1707. In some embodiments, processor 1707 may control the operation of user interface 1705 and receive input from user interface 1705. In some embodiments, user interface 1705 may allow a user to input instructions to device 1700, for example, via a keypad. In some embodiments, user interface 1705 may allow a user to obtain information from device 1700. For example, user interface 1705 may include a display configured to display information from device 1700 to the user. User interface 1705, in some embodiments, may be configured with a touchscreen or touch interface that can both allow information to be input into device 1700 and further display information to the user of device 1700. In some embodiments, user interface 1705 may be a user interface for communication.
[0172] In some embodiments, device 1700 has an input / output port 1709. The input / output port 1709 in some embodiments comprises a transceiver. The transceiver in such embodiments may be coupled to processor 1707 and configured to enable communication with other apparatuses or electronic devices, for example, via a wireless communication network. The transceiver or any suitable transceiver or transmitter and / or receiver means may, in some embodiments, be configured to communicate with other electronic devices or apparatuses via a wire or wired coupling.
[0173] The transceiver can communicate with the additional device via any suitable known communication protocol. For example, in some embodiments, the transceiver can use a suitable radio access architecture based on Long Term Evolution Advanced (LTE Advanced, LTE-A) or New Radio (NR) (also referred to as 5G), Universal Mobile Telecommunications System (UMTS) Radio Access Network (UTRAN or E-UTRAN), Long Term Evolution (e.g., LTE, E-UTRA), 2G networks (legacy network technologies), wireless local area networks (WLAN or WiFi), worldwide interoperability for microwave access (WiMAX), Bluetooth, Personal Communications Services (PCS), ZigBee, Wideband Code Division Multiple Access (WCDMA), Ultra-Wideband (UWB) technology systems, sensor networks, mobile ad hoc networks (MANETs), Cellular Internet of Things (IoT) RAN, and Internet Protocol Multimedia Subsystem (IMS), any other suitable alternatives, and / or any combination thereof.
[0174] The transceiver input / output port 1709 may be configured to receive a signal.
[0175] In some embodiments, device 1700 may be employed as at least a portion of a synthesis device. Input / output port 1709 may be coupled to headphones (which may be headphones with or without a head track) or the like.
[0176] In general, various embodiments of the present invention may be implemented in hardware or special purpose circuits, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Although various aspects of the present invention may be illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it will be appreciated that these blocks, apparatus, systems, techniques, or methods described herein may be implemented in, by way of non-limiting example, hardware, software, firmware, special purpose circuits or logic, general purpose hardware, or a controller or other computing device, or any combination thereof.
[0177] Embodiments of the invention may be implemented by computer software executable by a data processor of a mobile device, such as a processor entity, or by hardware, or by a combination of software and hardware. Further, in this regard, it should be noted that any blocks of the logic flow as illustrated may represent program steps, or interconnected logic circuits, blocks and functions, or combinations of program steps and logic circuits, blocks and functions. Software may also be stored on physical media, such as memory chips or memory blocks implemented within a processor, magnetic media, such as hard disks or floppy disks, and optical media, such as DVDs and their data variants, CDs.
[0178] The memory may be of any type suitable for the local technology environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed and removable memory, etc. The data processor may be of any type suitable for the local technology environment and may include, by way of non-limiting examples, one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), gate level circuits, and processors based on multi-core processor architectures.
[0179] Embodiments of the present invention can be implemented in a variety of components, such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available for converting logic-level designs into semiconductor circuit designs suitable for etching onto semiconductor substrates.
[0180] Programs such as those from Synopsys, Inc. of Mountain View, California, and Cadence Design, Inc. of San Jose, California, use established design rules and pre-stored libraries of design modules to automate the routing of conductors and placement of components on semiconductor chips. Once the design of a semiconductor circuit is complete, the resulting design may be sent in a standardized electronic format (such as Opus or GDSII) to a semiconductor manufacturing facility (fab) for fabrication.
[0181] The foregoing description has provided a full and informative description of exemplary embodiments of the present invention, by way of illustrative and non-limiting examples. However, various modifications and adaptations will become apparent to those skilled in the relevant art in light of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still fall within the scope of the present invention, as defined by the appended claims.
Claims
1. 1. An apparatus for spatial audio rendering, comprising at least one processor and at least one memory containing computer program code, the at least one memory and the computer program code configured to cause the apparatus, using the at least one processor, to perform at least: receiving a spatial audio signal, the spatial audio signal including at least one audio signal and spatial metadata associated with the at least one audio signal; generating at least one uncorrelated audio signal based on the at least one audio signal; determining at least one control parameter configured to control an amount of at least one uncorrelated audio signal in at least two output audio signals for spatial audio reproduction, wherein the at least one control parameter is based at least on at least one further target characteristic of the at least two output audio signals and on the spatial metadata; generating the at least two output audio signals for the spatial audio reproduction based on the spatial audio signal and the at least one uncorrelated audio signal, wherein the amount of the at least one uncorrelated audio signal in the at least two output audio signals is controlled based on the at least one control parameter; Execute the at least one further target characteristic is at least one target covariance matrix; Device.
2. The at least one control parameter is: at least one processing gain applied to at least one of the at least one decorrelated audio signal or the at least one audio signal to be decorrelated; at least one mixing matrix configured to control mixing of the at least one decorrelated audio signal and the at least one audio signal; at least one mixing matrix and at least one residual mixing matrix, the at least one mixing matrix and the at least one residual mixing matrix being configured to control mixing of the at least one decorrelated audio signal and the at least one audio signal; at least one covariance matrix configured to control generation of at least one mixing matrix and / or at least one residual mixing matrix, the at least one mixing matrix and / or the at least one residual mixing matrix configured to control mixing of the at least one decorrelated audio signal and / or the at least one audio signal; The apparatus of claim 1 , comprising at least one of:
3. The means configured to determine at least one control parameter comprises: determining at least one further characteristic based on the at least one audio signal; determining said at least one further target characteristic of said at least two output audio signals; determining at least one first control parameter based on the at least one further characteristic based on the at least one audio signal and the at least one further target characteristic of the at least two output audio signals; determining at least one second control parameter or modifying the at least one first control parameter based on at least one of the spatial metadata and an audio type determined based on the at least one audio signal; It is configured as follows: the at least one further characteristic is an input covariance matrix; 10. The apparatus of claim 1.
4. The means configured to generate the at least two output audio signals for spatial audio reproduction comprises: mixing the at least one audio signal and the at least one uncorrelated audio signal based on the at least one first control parameter and the at least one second control parameter or the at least one modified first control parameter; outputting the at least two output audio signals for spatial audio reproduction; The device of claim 3 , configured to:
5. The apparatus of claim 3 , wherein the determined at least one second control parameter or the modified at least one first control parameter is based on at least one direct-to-total energy ratio parameter in the spatial metadata.
6. 4. The apparatus of claim 3, wherein the at least one further characteristic based on the at least one audio signal is a covariance characteristic, and the at least one further target characteristic of the at least two output audio signals is a target covariance characteristic of the at least two output audio signals.
7. The means configured to determine at least one second control parameter or modify said at least one first control parameter comprises: determining a residual covariance characteristic based on the covariance characteristic and the target covariance characteristic of the at least two output audio signals; processing the residual covariance characteristics based on the spatial metadata associated with the at least one audio signal; The apparatus of claim 6 , configured to:
8. The means configured to process the residual covariance characteristics comprises: attenuating the residual covariance characteristic if the spatial metadata indicates that the at least one audio signal is highly directional; if the spatial metadata indicates that the at least one audio signal is entirely ambient, passing the residual covariance characteristic through unprocessed; The apparatus of claim 7 , configured to:
9. The means configured to determine the target covariance characteristic comprises: generating a total energy estimate based on the covariance property; determining head-related transfer function data based on directional parameters from the spatial metadata associated with the at least one audio signal; determining the target covariance characteristics of the at least two output audio signals based on the head-related transfer function data and the total energy estimate; The apparatus of claim 6 , configured to:
10. The means configured to determine at least one control parameter configured to control the amount of the at least one decorrelated audio signal in the at least two output audio signals for spatial audio reproduction comprises: Determine whether the audio type is the determined audio type; determining the at least one control parameter based on the determined audio type; The device of claim 1 configured to:
11. The apparatus of claim 10 , wherein the determined audio type is speech.
12. The apparatus of claim 1 , wherein the at least one audio signal comprises a transport audio signal generated by an encoder.
13. 1. A method for an apparatus for spatial audio rendering, the method comprising: receiving a spatial audio signal, the spatial audio signal including at least one audio signal and spatial metadata associated with the at least one audio signal; generating at least one uncorrelated audio signal based on the at least one audio signal; determining at least one control parameter configured to control an amount of at least one uncorrelated audio signal in at least two output audio signals for spatial audio reproduction, wherein the at least one control parameter is based at least on at least one further target characteristic of the at least two output audio signals and on the spatial metadata; generating the at least two output audio signals for spatial audio reproduction based on the spatial audio signal and at least one uncorrelated audio signal, wherein the amount of the at least one uncorrelated audio signal in the at least two output audio signals is controlled based on the at least one control parameter; Including, the at least one further target characteristic is a target covariance matrix; method.
14. The at least one control parameter is: at least one processing gain applied to at least one of the at least one decorrelated audio signal or the at least one audio signal to be decorrelated; the at least one decorrelated audio signal and at least one mixing matrix configured to control mixing of the at least one audio signal; at least one mixing matrix and at least one residual mixing matrix, the at least one mixing matrix and the at least one residual mixing matrix being configured to control mixing of the at least one decorrelated audio signal and the at least one audio signal; at least one covariance matrix configured to control generation of at least one mixing matrix and / or at least one residual mixing matrix, the at least one mixing matrix and / or at least one residual mixing matrix configured to control the at least one decorrelated audio signal and / or mixing of the at least one audio signal; The method of claim 13 , comprising at least one of:
15. Determining the at least one control parameter comprises: determining at least one further characteristic based on the at least one audio signal; determining at least one further target characteristic of the at least two output audio signals; determining at least one first control parameter based on the at least one further characteristic based on the at least one audio signal and the at least one further target characteristic of the at least two output audio signals; determining at least one second control parameter or modifying the at least one first control parameter based on at least one of the spatial metadata and a type of audio signal determined based on the at least one audio signal; Including, the at least one further characteristic is an input covariance matrix; The method of claim 13.
16. Generating the at least two output audio signals for spatial audio reproduction includes: mixing the at least one audio signal and at least one uncorrelated audio signal based on the at least one first control parameter and the at least one second control parameter or the at least one modified first control parameter; outputting the at least two output audio signals for spatial audio reproduction; 16. The method of claim 15, comprising:
17. 16. The method of claim 15, wherein determining the at least one second control parameter or the modified at least one first control parameter is based on at least one direct-to-total energy ratio parameter in the spatial metadata.
18. 16. The method of claim 15, wherein the at least one further characteristic based on the at least one audio signal is a covariance characteristic, and the at least one further target characteristic of the at least two output audio signals is a target covariance characteristic of the at least two output audio signals.
19. Determining the at least one second control parameter or modifying the at least one first control parameter may include: determining a residual covariance characteristic based on the covariance characteristic and the target covariance characteristic of the at least two output audio signals; processing the residual covariance characteristics based on the spatial metadata associated with the at least one audio signal; and Processing the residual covariance characteristics based on the spatial metadata associated with the at least one audio signal comprises: attenuating the residual covariance characteristic if the spatial metadata indicates that the at least one audio signal is highly directional; if the spatial metadata indicates that the at least one audio signal is entirely ambient, passing the residual covariance characteristic unprocessed; Including, 20. The method of claim 18.
20. Determining the target covariance characteristics of the at least two output audio signals comprises: generating a total energy estimate based on the covariance property; determining head-related transfer function data from the spatial metadata associated with the at least one audio signal based on directional parameters; determining the target covariance characteristics of the at least two output audio signals further based on the head-related transfer function data and the total energy estimate; 20. The method of claim 19, comprising:
Citation Information
Patent Citations
Output signal synthesis apparatus and synthesis method
JP2010525403A
A device for determining spatial output multichannel audio signals.
JP2011530913A
Processing spatially dispersed or large audio objects
JP2016530803A
Signaling spatial audio parameters
JP2021525392A
Binaural Audio Reproduction
US20160373877A1