Multi-channel audio encoding and decoding using direction metadata
By analyzing spatial audio signals for direction and energy information, the method generates a compressed representation that efficiently stores and transmits spatial audio scenes, ensuring high-quality reconstruction.
Patent Information
- Application Number
- JP2022524622
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-10-01
- Filing Date
- 2020-10-29
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2040-10-29
AI Technical Summary
Existing audio recording and playback technologies face challenges in efficiently capturing, storing, and transmitting spatial audio scenes due to the high bandwidth requirements of multi-channel or object-based signals, which often result in large data sizes.
A method for processing spatial audio signals to generate a compressed representation by analyzing the direction of arrival and energy information of audio elements, generating a channel-based audio signal with metadata, and using an inverse mixing matrix to reconstruct the original spatial audio signal.
This approach allows for a compact representation of spatial audio scenes with reduced data requirements while maintaining a high-quality reconstruction of the original audio experience.
Smart Images

Figure 0007711053000031 
Figure 0007711053000032 
Figure 0007711053000033
Abstract
Description
Technical Field
[0001] [Cross - Reference to Related Applications] This application claims priority to U.S. Provisional Patent Application No. 62 / 927,790, filed October 30, 2019, and U.S. Provisional Patent Application No. 63 / 086,465, filed October 1, 2020, the entire contents of each of which are incorporated herein by reference.
[0002] [Technical Field] The present disclosure generally relates to audio signal processing. In particular, the present disclosure relates to methods for processing a spatial audio signal (spatial audio scene) to generate a compressed representation of the spatial audio signal, and methods for processing a compressed representation of the spatial audio signal to generate a reconstructed representation of the spatial audio signal.
Background Art
[0003] By human hearing, listeners can perceive their environment in the form of a spatial audio scene. Here, the term "spatial audio scene" is used to refer to the acoustic environment around the listener or the acoustic environment perceived in the listener's mind.
[0004] Human experience is associated with the spatial audio scene, but audio recording and playback technologies include the capture, manipulation, transmission, and playback of audio signals or audio channels. The term "audio stream" is used to refer to a collection of one or more audio signals, particularly when the audio stream is intended to represent a spatial audio scene.
[0005] An audio stream can be reproduced to a listener via an electroacoustic transducer or by other means, providing a listening experience to one or more listeners in the form of a spatial audio scene. The goal of audio recording practitioners and audio artists is generally to create an audio stream intended to provide listeners with an experience of a particular spatial audio scene.
[0006] An audio stream may be accompanied by associated data called metadata that aids the playback process. The accompanying metadata may include information that changes over time. This information can be used to affect changes in the processing applied during the playback process.
[0007] Hereinafter, the term "captured audio experience" may be used to refer to metadata associated with an audio stream.
[0008] In some applications, the metadata consists only of data indicating the intended loudspeaker arrangement for playback. Assuming that the arrangement of the playback speakers is standardized, this metadata is often omitted. In this case, the captured audio experience consists only of the audio stream. One example of such a captured audio experience is a two-channel audio stream recorded on a compact disc. At this time, the intended playback system is assumed to be in the form of two loudspeakers placed in front of the listener.
[0009] Alternatively, a captured audio experience in the form of a scene-based multi-channel audio signal can be intended for presentation to a listener by processing the audio signal with a mixing matrix to generate a set of speaker signals. Each speaker signal is then reproduced at each respective loudspeaker. At this time, the loudspeakers can optionally be spatially arranged around the listener. In this example, the mixing matrix can be generated based on prior knowledge regarding the scene-based format and the arrangement of the playback speakers.
[0010] An example of a scene-based format is Higher Order Ambisonics (HOA), and an example of a method for calculating an appropriate mixing matrix is given in "Ambisonics", Franz Zotter and Matthias Frank, ISBN: 978-3-030-17206-0, Chapter 3, which is incorporated herein by reference.
[0011] Typically, such scene-based formats include a large number of channels or audio objects, so when transmitting or storing spatial audio signals in these formats, the bandwidth or storage requirements are relatively high.
[0012] Therefore, a compact representation of the spatial audio signal representing the spatial audio scene is required. This applies to both channel-based and object-based spatial audio signals. SUMMARY OF THE INVENTION
[0013] The present disclosure proposes a method for processing a spatial audio signal to generate a compressed representation of the spatial audio signal, a method for processing the compressed representation of the spatial audio signal to generate a reconstructed representation of the spatial audio signal, and corresponding devices, programs, and computer-readable storage media.
[0014] One aspect of the present disclosure relates to a method of processing a spatial audio signal to generate a compressed representation of the spatial audio signal. The spatial audio signal may be, for example, a multi-channel signal or an object-based signal. The compressed representation may be a compact or size-reduced representation. The method may include analyzing the spatial audio signal to determine the direction of arrival of one or more audio elements in an audio scene (spatial audio scene) represented by the spatial audio signal. The audio element may be a dominant audio element. The (dominant) audio element may be related to, for example, a (dominant) acoustic object, a (dominant) sound source, or a (dominant) acoustic component in the audio scene. The one or more audio elements may include, for example, from 1 to 10 audio elements, such as 4 audio elements. The direction of arrival may correspond to a position on a unit sphere indicating the perceived position of the audio element. The method may further include determining, for at least one frequency subband of the spatial audio signal (e.g., for all frequency subbands), an indication of each of the signal powers associated with the determined direction of arrival. The method may further include generating metadata including direction information and energy information, wherein the direction information includes an indication of the determined direction of arrival of one or more audio elements and the energy information includes an indication of each of the signal powers associated with the determined direction of arrival. The method may further include generating a channel-based audio signal having a predefined number of channels based on the spatial audio signal. The channel-based audio signal may sometimes be referred to as an audio mix signal or an audio mix stream. It is understood that the number of channels of the channel-based audio signal may be less than the number of channels or the number of objects of the spatial audio signal. The method may also further include outputting the channel-based audio signal and the metadata as a compressed representation of the spatial audio signal. The metadata may be related to a metadata stream.
[0015] Thereby, a compressed representation of a spatial audio signal can be generated to include a limited number of channels. Nevertheless, by appropriate use of the direction information and the energy information, the decoder can generate a reconstructed version of the original spatial audio signal that is a very good approximation of the original spatial audio signal as far as the representation of the original spatial audio signal is concerned.
[0016] In some embodiments, analyzing a spatial audio signal may be based on a plurality of frequency sub-bands of the spatial audio signal. For example, the analysis may be based on the full frequency range (i.e., the full signal) of the spatial audio signal. That is, the analysis may be based on all frequency sub-bands.
[0017] In some embodiments, analyzing a spatial audio signal may include applying scene analysis to the spatial audio signal. Thereby, the (direction of) dominant audio elements in the audio scene can be determined in a reliable and efficient manner.
[0018] In some embodiments, the spatial audio signal may be a multi-channel audio signal. Alternatively, the spatial audio signal may be an object-based audio signal. In this case, the method may further include converting the object-based audio signal to a multi-channel audio signal before applying scene analysis. This makes it possible to meaningfully apply scene analysis tools to the audio signal.
[0019] In some embodiments, an indication of the signal power associated with a given arrival direction may be related to the ratio of the signal power in a frequency sub-band for the given arrival direction to the total signal power in the frequency sub-band.
[0020] In some embodiments, an indication of signal power may be determined for each of a plurality of frequency sub-bands. In this case, they may be related to the ratio of the signal power in a given frequency sub-band for a given direction of arrival to the total signal power in the given frequency sub-band for the given direction of arrival and the given frequency sub-band. In particular, an indication of signal power may be determined for each sub-band, while the determination of the (dominant) direction of arrival may be performed for the total signal (i.e., based on all frequency sub-bands).
[0021] In some embodiments, analyzing the spatial audio signal, determining each indication of signal power, and generating the channel-based audio signal may be performed for each time segment. Accordingly, the compressed representation may be generated and output for each of a plurality of time segments by the downmixed audio signal and metadata (metadata block) of each time segment. Alternatively, or additionally, analyzing the spatial audio signal, determining each indication of signal power, and generating the channel-based audio signal may be performed based on the time-frequency representation of the spatial audio signal. For example, the above steps may be performed based on the discrete Fourier transform (e.g., STFT) of the spatial audio signal. That is, for each time segment (time block), the above steps may be performed based on the time-frequency bins (FFT bins) of the spatial audio signal, i.e., based on the Fourier coefficients of the spatial audio signal.
[0022] In some embodiments, the spatial audio signal may be an object-based audio signal that includes a plurality of audio objects and associated direction vectors. In that case, the method may further include generating a multi-channel audio signal by panning the audio objects into a predefined set of audio channels. Among them, each audio object may be panned into a predefined set of audio channels according to its direction vector. Further, the channel-based audio signal may be a downmix signal generated by applying a downmix operation to the multi-channel audio signal. The multi-channel audio signal may be, for example, a higher-order ambisonics signal.
[0023] In some embodiments, the spatial audio signal may be a multi-channel audio signal. In that case, the channel-based audio signal may be a downmix signal generated by applying a downmix operation to the multi-channel audio signal.
[0024] Another aspect of the present disclosure relates to a method of processing a compressed representation of a spatial audio signal to generate a reconstructed representation of the spatial audio signal. The compressed representation may include a channel-based audio signal having a predefined number of channels and metadata. The metadata may include direction information and energy information. The direction information may include an indication of the arrival direction of one or more audio elements in an audio scene (spatial audio scene). The energy information may include an indication of each of the signal powers associated with the arrival direction for at least one frequency subband. The method may include generating an audio signal of one or more audio elements based on the channel-based audio signal, direction information, and energy information. The method may further include generating a residual audio signal in which one or more audio elements are substantially absent based on the channel-based audio signal, direction information, and energy information. The residual signal may be represented in the same audio format as the channel-based audio signal and may have, for example, the same number of channels.
[0025] In some embodiments, an indication of signal power associated with a given direction of arrival may relate to a ratio of the signal power in a frequency sub - band for a given direction of arrival to the total signal power in the frequency sub - band.
[0026] In some embodiments, the energy information may include an indication of signal power for each of a plurality of frequency sub - bands. In that case, the indication of signal power may relate to a ratio of the signal power in a given frequency sub - band for a given direction of arrival to the total signal power in the given frequency sub - band for the given direction of arrival.
[0027] In some embodiments, the method may further include panning the audio signals of one or more audio elements to a set of channels of an output audio format. The method may also further include generating a reconstructed multi - channel audio signal in the output audio format based on the panned one or more audio elements and a residual audio signal. The output audio format may relate to an output representation such as, for example, HOA or any other suitable multi - channel format. Generating the reconstructed multi - channel audio signal may include up - mixing the residual signal to a set of channels of the output audio format. Generating the reconstructed multi - channel audio signal may further include adding together the panned one or more audio elements and the up - mixed residual signal.
[0028] In some embodiments, generating the audio signals of one or more audio elements may include determining the coefficients of an inverse mixing matrix M for mapping a channel - based audio signal to an intermediate representation that includes a residual audio signal and the audio signals of the one or more audio elements, based on direction information and energy information. The intermediate representation may sometimes be referred to as a separated or separable representation, or a hybrid representation.
[0029] In some embodiments, determining the coefficients of the unmixing matrix M includes, for each of one or more audio elements, determining a panning vector Pan for panning the audio element to a channel of the channel-based audio signal based on a direction of arrival dir of the audio element. down Determining coefficients of the unmixing matrix M may further include determining a mixing matrix E used to map the residual audio signal and the audio signals of the one or more audio elements to channels of the channel-based audio signal based on the determined panning vector. Determining coefficients of the unmixing matrix M may further include determining a covariance matrix S of the intermediate representation based on the energy information. Determining the covariance matrix S may include determining a covariance matrix S based on the determined panning vector Pan down Determining the coefficients of the unmixing matrix M above may also further include determining the coefficients of the unmixing matrix M based on the mixing matrix E and the covariance matrix S.
[0030] In some embodiments, the mixing matrix E is E=(I N |Pan down (dir1)|···|Pan down (dir P ) ) where I N may be an N×N identity matrix, where N denotes the number of channels of the channel-based audio signal, and Pan down (dir p ) is the associated direction of arrival dir for panning (mapping) the p-th audio element onto the N channels of the channel-based audio signal. p p may be a panning vector of the p-th audio element having p=1,...,P, where p=1,...,P denotes each one of the one or more audio elements and P denotes a total number of the one or more audio elements. Thus, the matrix E may be an NxP matrix. The matrix E may be determined for each of a plurality of time segments k. In that case, the matrix E and the direction of arrival dirp will have an index k indicating a time segment. For example, E k =(I N |Pan down (dir k,1 )|···|Pan down (dir k,P ))). Even if the proposed method can operate in band units, the matrix E will be the same for all frequency subbands.
[0031] In some embodiments, for 1 ≦ n ≦ N,
Number
Number
[0032] In some embodiments, determining the coefficients of the inverse mixing matrix M based on the mixing matrix E and the covariance matrix S may include determining a pseudo-inverse matrix based on the mixing matrix E and the covariance matrix S.
[0033] In some embodiments, the inverse mixing matrix M is M = S × E * × (E × S × E * ) -1 and can be determined according to this. Here, "×" represents the matrix product, and "*" represents the conjugate transpose of the matrix. The inverse mixing matrix M can be determined for each of a plurality of time segments k and / or for each of a plurality of frequency sub-bands b. In that case, the matrices M and S will have an index k indicating the time segment and / or an index b indicating the frequency sub-band, and the matrix E will have an index k indicating the time segment. For example, M k,b = S k,b × E * k × (E k × S k,b × E * k ) -1 is.
[0034] In some embodiments, the channel-based audio signal may be a first-order ambisonics signal.
[0035] Another aspect relates to an apparatus including a processor and a memory coupled to the processor, wherein the processor is configured to execute all steps of a method according to any one of the above aspects and embodiments.
[0036] Another aspect of the present disclosure relates to a program including instructions that, when executed by a processor, cause the processor to execute all steps of the above method.
[0037] Yet another aspect of the present disclosure relates to a computer-readable storage medium storing the above program.
[0038] Further embodiments of the present disclosure include an efficient way of representing a spatial audio scene in the form of an audio mixed stream and a direction metadata stream, the direction metadata stream including data indicating the positions of the directional acoustic elements in the spatial audio scene and, for a number of sub-bands, data indicating the power of each directional acoustic element relative to the total power of the spatial audio scene in that sub-band. Yet other embodiments relate to a method of determining a direction metadata stream from an input spatial audio scene and a method of generating a reconstructed audio scene from the direction metadata stream and an associated audio mixed stream.
[0039] In some embodiments, the method is used to represent the spatial audio scene in a more compact form as a compact spatial audio scene including an audio mixed stream and a direction metadata stream. At this time, the above audio mixed stream consists of one or more audio signals, and the above direction metadata stream consists of time-series direction metadata blocks, each of the direction metadata blocks being associated with a corresponding time segment of the audio signal. The spatial audio scene includes one or more directional acoustic elements, each associated with a respective direction of arrival. Each of the direction metadata blocks: ● Direction information indicating the direction of arrival for each of the directional acoustic elements, and ● For each of the directional acoustic elements and for each of the sets of two or more sub-bands, energy band fraction information indicating the energy of each directional acoustic element relative to the energy in the corresponding time segment of the audio signal. including.
[0040] In some embodiments, the method is used to process a compact spatial audio scene comprising an audio mixed stream and a direction metadata stream to generate a separated spatial audio stream and a residual stream comprising a set of one or more audio object signals. At this time, the above audio mixed stream consists of one or more audio signals, the above direction metadata stream consists of time-series direction metadata blocks, and each of the direction metadata blocks is related to the corresponding time segment of the audio signal. For each of a plurality of sub-bands, the method: ● determining coefficients of a demixing matrix (inverse mixing matrix) from direction information and energy band ratio information included in the direction metadata stream, and ● using the above demixing matrix to mix the audio signals to generate the above separated spatial audio stream is included.
[0041] In some embodiments, the method is used to process a spatial audio scene to generate a compact spatial audio scene comprising an audio mixed stream and a direction metadata stream. At this time, the above spatial audio scene includes one or more directional acoustic elements each associated with a respective arrival direction, the above direction metadata stream consists of time-series direction metadata blocks, and each of the direction metadata blocks is related to the corresponding time segment of the audio signal. The method: ● determining the arrival direction for one or more of the directional acoustic elements from an analysis of the spatial audio scene, ● determining which portion of the total energy in the spatial scene is contributed by the energy at each of the directional acoustic elements, and ● processing the spatial audio scene to generate an audio mixed stream is included.
[0042] It is understood that the above steps may be implemented by appropriate means or units, i.e., for example, they may be implemented by one or more computer processors.
[0043] It will also be understood that the mechanisms of the apparatus and the steps of the method may be interchanged in many ways. In particular, as will be understood by those skilled in the art, the details of the disclosed method are realizable by the corresponding apparatus, and vice versa. Furthermore, it is understood that any of the above descriptions made with respect to the method is equally applicable to the corresponding apparatus, and vice versa.
[0044] Exemplary embodiments of the present disclosure are illustrated by way of example in the accompanying drawings. In the drawings, the same reference numerals indicate the same or similar elements.
Brief Description of the Drawings
[0045]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
DETAILED DESCRIPTION OF THE INVENTION
[0046] Generally, the present disclosure relates to enabling storage and / or transmission of a spatial audio scene using a reduced amount of data.
[0047] Concepts of audio processing that may be used in the context of the present disclosure are described next.
[0048] [Panning Function] A multi-channel audio signal (or audio stream) may be formed by panning individual acoustic elements (or audio elements, audio objects) according to the law of linear mixing. For example, when a set of R audio objects is represented by R signals {o r (t): 1 ≤ r ≤ R}, the multi-channel pan mixture {z n (t): 1 ≤ n ≤ N} is: [Equation] It may be formed by.
[0049] The panning function Pan(θ r ) represents a column vector containing N scale factors (panning gains) that indicate the gains used to mix the object signal o r (t) to form the multi-channel output, where θ r indicates the position of each object.
[0050] One possible panning function is a first-order Ambisonics (FOA) panner. An example of an FOA panning function is: [Equation] Given by.
[0051] An alternative panning function is a third-order Ambisonics panner (3OA). An example of a 3OA panning function is: [Number] is given by.
[0052] As will be understood by those skilled in the art, it is understood that the present disclosure is not limited to FOA or HOA panning functions, and the use of other panning functions may be contemplated.
[0053] [Short-Time Fourier Transform] An audio stream consisting of one or more audio signals can be converted, for example, into the form of a short-term Fourier transform (STFT). For this purpose, a discrete Fourier transform can be applied to (optionally windowed) time segments of the audio signals of the audio stream (e.g., channels, audio object signals). This process applied to an audio signal x(t) can be expressed as follows: X c,k (f) = STFT{x c (t)} (4) The STFT is an example of a time-frequency transform, and it is understood that the present disclosure should not be limited to the STFT.
[0054] In Equation (4), the variable X c,k (f) is the short-time Fourier transform of the audio time segment k (Outer 1) TIFF0007711053000006.tif6170 shows the short-time Fourier transform of channel c (1 ≤ c ≤ NumChans). Here, F represents the number of frequency bins generated by the discrete Fourier transform. The terms used here are examples, and it is understood that the specific implementation details of various STFT methods (including various window functions) may be known in the art. The audio time segment k is defined, for example, as the range of audio samples centered at t = k × stride + constant so that the time segments are evenly spaced in time at intervals equal to the stride.
[0055] The numerical values of the STFT (e.g., X c,k (1), X c,k (2), ···, X c,k (F)) may be referred to as FFT bins.
[0056] Furthermore, the STFT format can be converted into an audio stream. The resulting audio stream can be an approximation to the original input:
Number
[0057] [Frequency-banded analysis] Characteristic data can be formed from the audio stream. The characteristic data is related to the number of frequency bands (frequency sub-bands), and the bands (sub-bands) are defined by regions of the frequency range.
[0058] As an example, the signal power in channel c of the stream in frequency band b (where the number of bands is B and 1 ≤ b ≤ B) when band b spans FFT bins f min ≤ f ≤ f max is given by:
Number
[0059] In a more general example, frequency band b may be defined by a weighting vector FR b (f) that assigns weights to each frequency bin, such that an alternative calculation of the power in a band is:
Number
[0060] In a further generalization of Equation (7), the STFT of a stream consisting of C audio signals can be processed to generate a covariance in a plurality of bands. At this time, the covariance R b,k is a C×C matrix, and the element {R b,k} i,j is:
Equation
[0061] In other examples, a band-pass filter may be used to form a filtered signal representing the original audio stream in a frequency band according to the band-pass filter response. For example, the audio signal x c (t) may be filtered to generate an x' c (t) representing a signal having energy mainly obtained from band b of x c,b (t). Therefore, an alternative method for calculating the covariance of the stream in band b for time block k (corresponding to time samples t min ≤t≤t max ) is:
Equation
[0062] [Frequency Banded Mixing] An audio stream consisting of N channels can be processed to generate an audio stream consisting of M channels according to an M×N linear mixing matrix Q such that:
Equation
Equation
[0063] Furthermore, an alternative mixing process may be implemented in the STFT domain, and the matrix Q can take different values at each time block t and each frequency band b. In this case, the process is: [Equation] by, or in the form of a matrix, [Equation] can be regarded as being approximately given by.
[0064] It is understood that an alternative method can be used to produce behavior equivalent to the process shown in Equation (13).
[0065] [Example implementation] Next, an example implementation of a method and apparatus according to an embodiment of the present disclosure will be described in more detail.
[0066] Broadly speaking, the method according to an embodiment of the present disclosure represents a spatial audio scene in the form of an audio mixing stream and a direction metadata stream, and the direction metadata stream includes data indicating the positions of the directional acoustic elements in the spatial audio scene and, among a number of sub-bands, data indicating the power of each directional acoustic element relative to the total power of the spatial audio scene in that sub-band. Further methods according to embodiments of the present disclosure relate to determining a direction metadata stream from an input spatial audio scene and generating a reconstructed (e.g., recovered) audio scene from the direction metadata stream and a related audio mixing stream.
[0067] Examples of methods according to embodiments of the present disclosure are efficient in representing a spatial audio scene (e.g., with respect to reduction of data for storage or transmission). The spatial audio scene can be represented by a spatial audio signal. The above method can be implemented by defining a storage or transmission format (e.g., Compact Spatial Audio Stream) consisting of an audio mixing stream and a metadata stream (e.g., a direction metadata stream).
[0068] The audio mixing stream has a number of audio signals carrying a reduced representation of the spatial audio scene. As such, the audio mixing stream can be related to a channel-based audio signal having a predefined number of channels. It is understood that the number of channels of the channel-based audio signal is less than the number of channels of the spatial audio signal or the number of audio objects. For example, the channel-based audio signal may be a first-order ambisonics audio signal. In other words, the Compact Spatial Audio Stream may include an audio mixing stream in the form of a first-order ambisonics representation of the sound field.
[0069] (Direction) The metadata stream has metadata defining the spatial characteristics of the spatial audio scene. The direction metadata can be composed of a sequence of direction metadata blocks. Each direction metadata block includes metadata indicating the characteristics of the spatial audio scene in the corresponding time segment within the audio mixing stream.
[0070] Generally, metadata includes direction information and energy information. The direction information includes an indication of the direction of arrival of one or more (dominant) audio elements in an audio scene. The energy information includes, for each direction of arrival, an indication of the signal power associated with the determined direction of arrival. In some implementations, the indication of signal power may be provided for one, some, or each of a plurality of bands (frequency sub-bands). Further, the metadata may be provided for each of a plurality of consecutive time segments, for example, in the form of metadata blocks.
[0071] In one example, the metadata (direction metadata) includes metadata indicating the characteristics of a spatial audio scene over a number of frequency bands, the metadata comprising: ● one or more directions (e.g., directions of arrival) indicating the position of audio objects (audio elements) in the spatial audio scene, and ● the proportion of energy (or spatial power) in each frequency band by each audio object (e.g., by each direction). is included.
[0072] Details regarding the determination of the direction information and the energy information are given below.
[0073] FIG. 1 schematically shows an example of an arrangement using an embodiment of the present disclosure. Specifically, the figure shows an arrangement 100 in which a spatial audio scene 10 is input to a scene encoder 200, and the scene encoder 200 generates an audio mixed stream 30 and a direction metadata stream 20. The spatial audio scene 10 may be represented by a spatial audio signal or a spatial audio stream input to the scene encoder 200. The audio mixed stream 30 and the direction metadata stream 20 together form an example of a compact spatial audio scene, i.e., a compressed representation of the spatial audio scene 10 (or of the spatial audio signal).
[0074] The compressed representations, namely, the mixed audio stream 30 and the direction metadata stream 20, are input into the scene decoder 300, and the scene decoder 300 generates the reconstructed audio scene 50. The audio elements existing within the spatial audio scene 10 are represented within the audio mix stream 30 according to the mixing panning function.
[0075] FIG. 2 schematically shows another example of an arrangement using an embodiment of the present disclosure. Specifically, the figure shows an alternative arrangement 110 in which a compact spatial audio scene consisting of an audio mix stream 30 and a direction metadata stream 20 is further encoded by supplying the audio mix stream 30 to an audio encoder 35 so as to generate a bitrate-reduced encoded audio stream 37, and by supplying the direction metadata stream 20 to a metadata encoder 25 so as to generate an encoded metadata stream 27. The bitrate-reduced encoded audio stream 37 and the encoded metadata stream 27 together form an encoded (bitrate-reduced) spatial audio scene.
[0076] The encoded spatial audio scene can be recovered by first applying the bitrate-reduced encoded audio stream 37 and the encoded metadata stream 27 to respective decoders 36 and 26 so as to generate a reproduced audio mix stream 38 and a reproduced direction metadata stream 28. The reproduced streams 38, 28 are the same as, or approximately equal to, the respective streams 30, 20. The reproduced audio mix stream 38 and the reproduced direction metadata stream 28 can be decoded by the decoder 300 so as to generate the reconstructed audio scene 50.
[0077] FIG. 3 schematically represents an example of an arrangement for generating a bitrate-reduced encoded audio stream and an encoded metadata stream from an input spatial audio scene. Specifically, the figure shows an arrangement 150 of a scene encoder 200 that supplies a direction metadata stream 20 and an audio mix stream 30 to respective encoders 25, 35 to generate an encoded spatial audio scene 40 including a bitrate-reduced encoded audio stream 37 and an encoded metadata stream 27. The encoded spatial audio stream 40 is preferably arranged to be suitable for storage and / or transmission with reduced data requirements for the data required for storage / transmission of the original spatial audio scene.
[0078] FIG. 4 schematically represents an example of an arrangement for generating a reconstructed spatial audio scene from a bitrate-reduced encoded audio stream and an encoded metadata stream. Specifically, the figure shows that an encoded spatial audio stream 40 consisting of a bitrate-reduced encoded audio stream 37 and an encoded metadata stream 27 is supplied as an input to respective decoders 36, 26 to generate an audio mix stream 38 and a direction metadata stream 28. Streams 38, 28 are then processed by a scene decoder 300 to generate a reconstructed audio scene 50.
[0079] Details for generating a compact spatial audio scene, i.e., a compressed representation of a spatial audio scene (or of a spatial audio signal / spatial audio stream), are described next.
[0080] FIG. 5 is a flowchart of an example of a method 500 for processing a spatial audio signal to generate a compressed representation of the spatial audio signal. Method 500 has steps S510 to S550.
[0081] In step S510, the spatial audio signal is analyzed to determine the direction of arrival of one or more audio elements (e.g., dominant audio elements) in the audio scene (spatial audio scene) represented by the spatial audio signal. (Dominant) audio elements may be related to, for example, (dominant) acoustic objects, (dominant) sound sources, or (dominant) acoustic components in the audio scene. Analyzing the spatial audio signal may include or be related to applying scene analysis to the spatial audio signal. It is understood that the scope of suitable scene analysis tools is known to those skilled in the art. The direction of arrival determined in this step may correspond to a position on the unit sphere indicating the (perceived) position of the audio element.
[0082] Consistent with the above description of the frequency-banded analysis, the analysis of the spatial audio signal in step S510 can be based on a plurality of frequency sub-bands of the spatial audio signal. For example, the analysis may be based on the entire frequency range of the spatial audio signal (i.e., the entire signal). That is, the analysis may be based on all frequency sub-bands.
[0083] In step S520, an indication of each of the signal powers associated with the determined direction of arrival is determined for at least one frequency sub-band of the spatial audio signal.
[0084] In step S530, metadata including direction information and energy information is generated. The direction information includes an indication of the determined direction of arrival of one or more audio elements. The energy information includes an indication of each of the signal powers associated with the determined direction of arrival. The metadata generated in this step may be related to a metadata stream.
[0085] In step S540, a channel-based audio signal having a predefined number of channels is generated based on the spatial audio signal.
[0086] Finally, in step S550, the channel-based audio signal and the metadata are output as a compressed representation of the spatial audio signal.
[0087] It is understood that the above steps may be executed in any order or in parallel with each other, as long as the necessary inputs for each step are guaranteed to be available according to the order of the steps.
[0088] Typically, a spatial scene (or spatial audio signal) can be considered to be composed of the sum of acoustic signals incident on the listener from a series of directions with respect to the listening position. Therefore, a spatial audio scene can be modeled as a set of R acoustic objects. Object r (1 ≤ r ≤ R) is associated with an audio signal o r incident on the listening position from an arrival direction defined by a direction vector θ r (t). The direction vector may also be a vector θ r (t) that changes with time.
[0089] Therefore, according to some implementations, a spatial audio signal (spatial audio stream) may be defined as an object-based spatial audio signal (object-based spatial audio scene) in the form of a set of an audio signal and an associated direction vector: Spatial audio scene (object-based) ={o r (t), θ r (t): 1 ≤ r ≤ R} (14) Furthermore, according to some implementations, a spatial audio signal (spatial audio stream) may be defined with respect to a short-time Fourier transform signal O r,k (f) according to Equation (4), and the direction vector may be specified according to a block index k, whereby: Spatial audio scene (object-based) ={O r,k (f), θ r(t): 1 ≦ r ≦ R} (15) is as follows.
[0090] Alternatively, the spatial audio signal (spatial audio stream) may be represented with respect to a channel-based spatial audio signal (channel-based spatial audio scene). The channel-based stream consists of a set of audio signals, and each acoustic object from the spatial audio scene is mixed into channels by a panning function (Pan(θ)) according to Equation (1). As an example, a Q-channel channel-based spatial audio scene {C q,k (f): 1 ≦ q ≦ Q} is
Equation
[0091] Many characteristics of the channel-based spatial audio scene are determined by the choice of the panning function. In particular, it will be understood that the length (Q) of the column vector returned by the panning function determines the number of audio channels included in the channel-based spatial audio scene. Generally speaking, a higher-quality representation of the spatial audio scene can be achieved by a channel-based spatial audio scene that includes a larger number of channels.
[0092] As an example, in step S540 of method 500, the spatial audio signal (spatial audio scene) may be processed to generate a channel-based audio signal (channel-based stream) according to Equation (16). The panning function may be selected to provide a relatively low-resolution representation of the spatial audio scene. For example, the panning function may be selected to be a first-order ambisonics (FOA) function as defined in Equation (2). As such, the compressed representation may be a compact or size-reduced representation.
[0093] FIG. 6 is a flowchart providing another formulation of a method 600 for generating a compact representation of a spatial audio scene. The method 600 is supplied with an input stream in the form of a spatial audio scene or a scene-based stream and generates a compact spatial audio scene as a compact representation. To this end, the method 600 has steps S610 to S660. Among them, step S610 may be regarded as corresponding to step S510, step 620 may be regarded as corresponding to step S520, step S630 may be regarded as corresponding to step S540, step S650 may be regarded as corresponding to step S530, and step S660 may be regarded as corresponding to step S550.
[0094] In step S610, the input stream is analyzed to determine a dominant arrival direction.
[0095] In step S620, for each band (frequency sub-band), the ratio of the energy assigned to each direction is determined with respect to the total energy in the stream in that band.
[0096] In step S630, a downmix stream including a plurality of audio channels representing the spatial audio scene is formed.
[0097] In step S640, the downmix stream is encoded to form a compressed representation of the stream.
[0098] In step S650, the direction information and the energy ratio information are encoded to form encoded metadata.
[0099] Finally, in step S660, the encoded downmix stream is combined with the encoded metadata to form a compact spatial audio scene.
[0100] It is understood that the above steps may be performed in any order or in parallel with each other, as long as the necessary inputs for each step are guaranteed to be available depending on the order of the steps.
[0101] Figures 7 to 11 schematically represent detailed examples of generating a compressed representation of a spatial audio scene according to embodiments of the present disclosure. Details such as, for example, the analysis of a spatial audio signal to determine the direction of arrival, the determination of an indication of signal power related to the determined direction of arrival, the generation of metadata including direction information and energy information, and / or the generation of a channel-based audio signal including a predefined number of channels, which will be described later, can be independent of a specific system arrangement and may be applied to any of the arrangements shown in Figures 7 to 11 or any suitable alternative arrangement, it is understood.
[0102] Figure 7 schematically represents a first detailed example of generating a compressed representation of a spatial audio scene. Specifically, Figure 7 shows a scene encoder 200 in which a spatial audio scene 10 is processed by a downmixing function 203 to generate an N-channel audio mixed stream 30, for example, according to steps S540 and S630. In some embodiments, the downmixing function 203 may include a panning process according to Equation (1) or Equation (16), and a downmix panning function is selected. That is,
Number
Number
[0103] For each audio time segment, the scene analysis 202 takes a spatial audio scene as input and determines, for example, according to steps S510 and S610, the arrival directions of up to the maximum P dominant acoustic components within the spatial audio scene. A typical value of P is between 1 and 10, and a preferred value of P is P≒4. Therefore, the one or more audio elements determined in step S510 may have between 1 and 10 audio elements, such as, for example, 4 audio elements.
[0104] The analysis 202 generates metadata 20 consisting of direction information 21 and energy band ratio information 22 (energy information). Optionally, the scene analysis 202 may also supply coefficients 207 to the downmix function 203 so as to enable the downmix to be changed.
[0105] Without intended limitation, analyzing the spatial audio signal (e.g., in step S510), determining each indication of signal power (e.g., in step S520), and generating the channel-based audio signal (e.g., in step S540) may be performed in units of time segments, for example, in accordance with the above description of STFT. This implies that the compressed representation has a downmix audio signal and metadata (metadata block) for each time segment, and is generated and output for each of the plurality of time segments.
[0106] For each time segment k, the direction information 21 (e.g., embodied by the arrival directions of one or more audio elements) can take the form of P direction vectors {dir k,p : 1≦p≦P}. The direction vector p indicates the direction associated with the dominant object index p, and in terms of unit vectors:
Number
Number
[0107] In some embodiments, each indication of the signal power determined in step S520 takes the form of a ratio of signal powers. That is, the indication of the signal power associated with a given direction of arrival in a frequency sub - band relates to the ratio of the signal power in the frequency sub - band for the given direction of arrival to the total signal power in the frequency sub - band.
[0108] Furthermore, in some embodiments, the indication of the signal power is determined for each of a plurality of frequency sub - bands (i.e., on a sub - band - by - sub - band basis). In that case, they relate to the ratio of the signal power in a given frequency sub - band for a given direction of arrival to the total signal power in the given frequency sub - band for the given direction of arrival and given frequency sub - band. In particular, even though the indication of the signal power can be determined sub - band by sub - band, the determination of the (dominant) direction of arrival can still be performed for the entire signal (i.e., based on all frequency sub - bands).
[0109] Still further, in some embodiments, analyzing the spatial audio signal (e.g., in step S510), determining each indication of the signal power (e.g., in step S520), and generating the channel - based audio signal (e.g., in step S540) are performed based on the time - frequency representation of the spatial audio signal. For example, the above steps and other appropriate steps can be performed based on the discrete Fourier transform (e.g., STFT) of the spatial audio signal. For example, for each time segment (time block), the above steps can be performed based on the time - frequency bins (FFT bins) of the spatial audio signal, i.e., the Fourier coefficients of the spatial audio signal.
[0110] In view of the anomaly, for each time segment k and for each dominant object index p (1 ≤ p ≤ P), the energy band ratio information 22 can include a fractional value e k,p,b for each band b (1 ≤ b ≤ B) of the set of bands. The fractional value e k,p,b is: [Number] It is determined according to time segment k.
[0111] Fractional value e k,p,b represents the energy of a plurality of acoustic objects in the original spatial audio scene, and the direction dir k,p is combined so as to represent a single dominant acoustic component to which the energy is assigned, and the direction dir k,p may represent a portion of the energy within the spatial region around. In some embodiments, the energy of all acoustic objects in the scene is weighted using an angular difference weighting function w(θ) that represents a greater weighting for directions θ closer to dir k,p and a smaller weighting for directions θ farther from dir k,p . The difference in direction may be considered close for an angular difference less than, for example, 10 degrees and far for an angular difference greater than, for example, 45 degrees. In alternative embodiments, the weighting function may be selected based on an alternative choice of near / far angular differences.
[0112] Generally, the input spatial audio signal for which the compressed representation is generated may be, for example, a multi-channel audio signal or an object-based audio signal. In the latter case, the method for generating the compressed representation of the spatial audio signal will further include a step of converting the object-based audio signal into a multi-channel audio signal before applying scene analysis (e.g., before step S510).
[0113] In the example of FIG. 7, the input spatial audio signal may be a multi-channel audio signal. In that case, the channel-based audio signal generated in step S540 will be a downmix signal generated by applying a downmix operation to the multi-channel audio signal.
[0114] Figure 8 schematically represents another example of details for generating a compressed representation of a spatial audio scene. The input spatial audio signal may, in this case, be an object-based audio signal including a plurality of audio objects and associated direction vectors. In this case, the method for generating a compressed representation of the spatial audio signal has generating a multi-channel audio signal as an intermediate representation or intermediate scene by panning the audio objects to a predefined set of audio channels. At this time, each audio object is panned to a predefined set of audio channels according to its direction vector. Thus, Figure 8 shows an alternative embodiment of a scene encoder 200 in which a spatial audio scene 10 is input to a converter 201 and the converter 201 generates an intermediate scene 11 (e.g., embodied by a multi-channel signal). The intermediate scene 11 may be generated according to Equation (1). At this time, the panning function is selected such that the inner product of the panning gain vectors Pan(θ1) and Pan(θ2) approximately represents the above angle difference weighting function.
[0115] In some embodiments, the panning function used in converter 201 is a third-order ambisonic panning function represented by Equation (3) (Outer 4) TIFF0007711053000024.tif8170. Thus, the multi-channel audio signal may be, for example, a higher-order ambisonic signal.
[0116] The intermediate scene 11 is then input to a scene analysis 202. The scene analysis 202 may determine the direction dir k,p of the dominant acoustic object in the spatial audio scene from the analysis of the intermediate scene 11. The determination of the dominant direction may be performed by estimating the energy in a set of directions, and the maximum estimated energy represents the dominant direction.
[0117] The energy band ratio information 22 of time segment k is a fractional value e for each band b derived from the energy in band b of the intermediate scene 11 in each direction with respect to the total energy in band b of the intermediate scene 11 within time segment k. k,p,b may include.
[0118] In this case, the audio mixing stream 30 (e.g., channel-based audio signal) of the compact spatial audio scene (e.g., compact representation) is a downmix signal generated by applying the downmix function 203 (downmix operation) to the spatial audio scene.
[0119] FIG. 10 shows an alternative arrangement of a scene encoder including a converter 201 that converts the spatial audio scene 10 into a scene-based intermediate format 11. The intermediate format 11 is input to the scene analysis 202 and the downmix function 203. In some embodiments, the downmix function 203 may include a matrix mixer having coefficients adapted to convert the intermediate format 11 into the audio mixing stream 30. That is, in this case, the audio mixing stream 30 (e.g., channel-based audio signal) of the compact spatial audio scene (e.g., compact representation) can be a downmix signal generated by applying the downmix function 203 (downmix operation) to the intermediate scene (e.g., multi-channel audio signal).
[0120] In an alternative embodiment shown in FIG. 11, the spatial encoder 200 can take the input in the form of a scene-based input 11. The acoustic objects are represented according to the panning rule Pan(θ). In some embodiments, the panning function may be a higher-order ambisonics panning function. In one exemplary embodiment, the panning function is a third-order ambisonics panning function.
[0121] In another alternative embodiment shown in FIG. 9, the spatial audio scene 10 is converted by the converter 201 in the spatial encoder 200 to generate an intermediate scene 11 that is input to the downmix function 203. Scene analysis 202 is supplied with input from the spatial audio scene 10.
[0122] FIG. 12 shows the direction information 21 and the energy band ratio information 22 that are input to the demixing matrix calculator 301 that determines the demixing matrix (inverse mixing matrix) used by the demixer 302.
[0123] Details for processing a compact spatial audio scene (e.g., a compressed representation of a spatial audio signal) to generate a reconstructed representation of the spatial audio signal are described next.
[0124] FIG. 13 is a flowchart of an example of a method 1300 for processing a compressed representation of a spatial audio signal to generate a reconstructed representation of the spatial audio signal. The compressed representation includes a channel-based audio signal (e.g., embodied by the audio mix stream 30) having a predefined number of channels and metadata, the metadata including direction information (e.g., embodied by the direction information 21) and energy information (e.g., embodied by the energy band ratio information 22), the direction information including an indication of the direction of arrival of one or more audio elements in the audio scene, and the energy information including an indication of each of the signal powers associated with the direction of arrival for at least one frequency subband. The channel-based audio signal may be, for example, a first-order ambisonics signal. The method 1300 has steps S1310 to S1320 and optionally has steps S1330 and S1340. It is understood that these steps may be performed, for example, by the scene decoder 300 of FIG. 12.
[0125] In step S1310, an audio signal of one or more audio elements is generated based on the channel-based audio signal, the direction information, and the energy information.
[0126] In step S1320, a residual audio signal with substantially no one or more audio elements is generated based on the channel-based audio signal, direction information, and energy information. Here, the residual signal may be represented in the same audio format as the channel-based audio signal, for example, may have the same number of channels as the channel-based audio signal.
[0127] In any step S1330, the audio signals of one or more audio elements are panned to a set of channels of the output audio format. Here, the output audio format may be related to the output representation, for example, like HOA or any other suitable multi-channel format.
[0128] In any step S1340, a reconstructed multi-channel audio signal in the output audio format is generated based on the panned one or more audio elements and the residual signal. Generating the reconstructed multi-channel audio signal may include upmixing the residual signal to a set of channels of the output audio format. Generating the reconstructed multi-channel audio signal may further include adding together the panned one or more audio elements and the upmixed residual signal.
[0129] It is understood that the above steps may be executed in any order or in parallel with each other, as long as the required inputs for each step are guaranteed to be available according to the order of the steps.
[0130] Consistent with the above description of the method of processing a spatial audio signal to generate a compressed representation of the spatial audio signal, an indication of the signal power related to a given arrival direction may be related to the ratio of the signal power in a frequency sub-band for the given arrival direction to the total signal power in the frequency sub-band.
[0131] Furthermore, in some embodiments, the energy information may include an indication of the signal power for each of a plurality of frequency sub-bands. In that case, the indication of the signal power may relate to the ratio of the signal power in a given frequency sub-band in a given direction of arrival to the total signal power in the given frequency sub-band for the given direction of arrival and the given frequency sub-band.
[0132] Generating the audio signals of the one or more audio elements in step S1310 may include determining the coefficients of the inverse mixing matrix M for mapping the channel-based audio signal to an intermediate representation that includes the residual audio signal and the audio signals of the one or more audio elements, based on the direction information and the energy information. The intermediate representation may also be referred to as a separated or separable representation, or a hybrid representation.
[0133] The details of the above determination of the coefficients of the inverse mixing matrix M are then described with reference to the flowchart of FIG. 14. The method 1400 represented by this flowchart has steps S1410 to S1440.
[0134] In step S1410, for each of the one or more audio elements, a panning vector Pan down (dir) for panning the audio element to the channels of the channel-based audio signal is determined based on the direction of arrival dir of the audio element.
[0135] In step S1420, a mixing matrix E used to map the residual audio signal and the audio signals of the one or more audio elements to the channels of the channel-based audio signal is determined based on the determined panning vectors.
[0136] In step S1430, the covariance matrix S of the intermediate representation is determined based on the energy information. The determination of the covariance matrix S may further be based on the determined panning vector Pan down and may be further based thereon.
[0137] Finally, in step S1440, the coefficients of the inverse mixing matrix M are determined based on the mixing matrix E and the covariance matrix S.
[0138] It is understood that the above steps may be executed in any order or in parallel with each other, as long as the required inputs for each step are guaranteed to be available according to the order of the steps.
[0139] Returning to FIG. 12, the demixing matrix calculator 301 calculates the demixing matrix 60 (inverse mixing matrix) M k,b according to a process including the following steps: 1. For each time segment k, the direction information dir k,p (1 ≤ p ≤ P) and the energy band ratio information e k,p,k (1 ≤ p ≤ P and 1 ≤ b ≤ B) are input to the demixing matrix calculator 301. P represents the number of dominant acoustic components, and B represents the number of frequency bands. 2. For each band b, the demixing matrix Mk,b is: M = S × E * × (E × S × E * ) -1 (20) is calculated according to. Here, "×" indicates the matrix product, and "*" indicates the conjugate transpose of the matrix. The calculation according to equation (20) may correspond to, for example, step S1440.
[0140] The demixing matrix M can be determined for each of a plurality of time segments k and / or for each of a plurality of frequency sub-bands b. In that case, the matrices M and S will have an index k indicating the time segment and / or an index b indicating the frequency sub-band, and the matrix E will have an index k indicating the time segment. For example, M k,b = S k,b × E * k × (E k × S k,b × E *k ) -1 (20a) is as follows.
[0141] In general, determining the coefficients of the inverse mixing matrix M based on the mixing matrix E and the covariance matrix S may include determining a pseudo-inverse matrix based on the mixing matrix E and the covariance matrix S. An example of such a pseudo-inverse matrix is given by equations (20) and (20a).
[0142] In equation (20), the matrix E k (mixing matrix) is formed by stacking the product of the N×N identity matrix (I N ) and P columns formed by panning functions applied to the directions of each of the P dominant acoustic components: E = (I N | Pan down (dir1) | ··· | Pan down (dir P ) ) (21) In equation (21), I N is the N×N identity matrix, N indicates the number of channels of the channel-based audio signal, Pan down (dir p ) is the panning vector of the p-th audio element having the associated direction of arrival dir p for panning the p-th audio element across the N channels of the channel-based audio signal, where p = 1, ···, P indicates each one of one or more audio elements, and P indicates the total number of one or more audio elements. The vertical bar in equation (21) indicates a matrix augmentation operation. Thus, the matrix E is an N×P matrix.
[0143] Furthermore, the matrix E may be determined for each of a plurality of time segments k. In that case, the matrix E and the direction of arrival dir p will have the index k indicating the time segment. For example: E k=(I N |Pan down (dir k,1 )|···|Pan down (dir k,P )) (21a) When the proposed method operates in band units, the matrix E is the same for all frequency subbands.
[0144] According to step S1420, the matrix E k is used to map the residual audio signal and the audio signals of one or more audio elements to the channels of the channel-based audio signal. As can be seen from equations (21) and (21a), the matrix E k is based on the panning vector Pan down (dir) determined in step S1410.
[0145] In equation (20), the matrix S is an (N + P) × (N + P) diagonal matrix. It can be regarded as the covariance matrix of the intermediate representation. Its coefficients can be calculated based on the energy information according to step S1430. The first N diagonal elements are for 1 ≤ n ≤ N:
Equation
[0146] The covariance matrix S can be determined for each of a plurality of time segments k and / or for each of a plurality of frequency subbands b. In that case, the covariance matrix S and the signal power e pwill have an index k indicating a time segment and / or an index b indicating a frequency sub-band. The first N diagonal elements are: [Number] are given by, and the remaining P diagonal elements are: {S k,b} N+p,N+p = e k , p,b (1 ≤ p ≤ P) (23a) are given by.
[0147] In a preferred embodiment, the demixing matrix M k,b is applied by the demixer 302 to generate the separated spatial audio stream 70 (as an example of an intermediate representation). According to the above implementation of step S1310, the first N channels are the residual stream 80, and the remaining P channels represent the dominant acoustic components.
[0148] The separated spatial stream 70 Y of N + P channels k (f), the P-channel dominant object signal 90 (as an example of the audio signal of one or more audio elements generated in step S1310) O k (f), and the N-channel residual stream 80 (as an example of the residual audio signal generated in step S1320) R k (f) are: [Number] are calculated from the N-channel audio mix 30 X k (f). The signals are represented in STFT format, and the representation with {Y k (f)} 1..N indicates the N-channel signal formed from channels 1..N of Y k (f), and {Y k (f)} N+1..N+P is Y kShows the P-channel signal formed from channels N+1..N+P of (f). Matrix M k,b It will be understood by those skilled in the art that the application of k,b can be achieved according to an alternative method known in the art that provides an approximation function equivalent to that of Equation (24).
[0149] In addition to the above, in some embodiments, the number P of dominant acoustic components can be adapted to take different values for each time segment. Thereby, P k can depend on time segment k. For example, the scene analysis 202 of the scene encoder 200 can determine the value of P k for each time segment. In general, the number of dominant acoustic components P can depend on time. The selection of P (or P k ) may include a trade-off between the data rate of the metadata and the quality of the reconstructed audio scene.
[0150] Returning to FIG. 12, the spatial decoder 300 generates the M-channel reconstructed audio scene 50. The M-channel stream is associated with the output panner (External 5) associated with TIFF0007711053000028.tif8170. This can be done according to step S1340 above. Examples of output panners include stereo panning functions, vector-based amplitude panning functions known in the art, and higher-order ambisonics panning functions known in the art.
[0151] For example, the object panner 91 in FIG. 12 is:[[]]
Number
[0152] FIG. 15 is a flowchart providing an alternative formulation of method 1500 for decoding a compact space audio scene to generate a reconstructed audio scene. Method 1500 includes steps S1510 through S1580.
[0153] In step S1510, a compact space audio scene is received and an encoded downmix stream and an encoded metadata stream are extracted.
[0154] In step S1520, the encoded downmix stream is decoded to form a downmix stream.
[0155] In step S1530, the encoded metadata stream is decoded to form direction information and energy ratio information.
[0156] In step S1540, a per-band demixing matrix is formed from the direction information and the energy ratio information.
[0157] In step S1550, the downmix stream is processed according to the demixing matrix to form separated streams.
[0158] In step S1560, object signals are extracted from the separated streams and panned according to the direction information and the desired output format to generate panned object signals.
[0159] In step S1570, residual signals are extracted from the separated streams and processed according to the desired output format to generate decoded residual signals.
[0160] Finally, in step S1580, the panned object signals and the decoded residual signals are combined to form a reconstructed audio scene.
[0161] It is understood that the above steps may be performed in any order or in parallel with each other, as long as the necessary inputs for each step are guaranteed to be available, depending on the order of the steps.
[0162] Methods for processing a spatial audio signal to generate a compressed representation of the spatial audio signal and for processing the compressed representation of the spatial audio signal to generate a reconstructed representation of the spatial audio signal have been described above. Further, the present disclosure relates also to an apparatus for performing these methods. An example of such an apparatus 1600 is schematically represented in FIG. 16. The apparatus 1600 may have a processor 1610 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), one or more application specific integrated circuits (ASICs), one or more radio frequency integrated circuits (RFICs), or any combination thereof), and a memory 1620 coupled to the processor 1610. The processor may be configured to perform some or all of the steps of the methods described throughout the present disclosure. When the apparatus 1600 operates as an encoder (e.g., a scene encoder), it may receive, as input 1630, for example, a spatial audio signal (i.e., a spatial audio scene). The apparatus 1600 may then generate, as output 1640, a compressed representation of the spatial audio signal. When the apparatus 1600 operates as a decoder (e.g., a scene decoder), it may receive, as input 1630, the compressed representation. The apparatus may then generate, as output 1640, a reconstructed audio scene.
[0163] Device 1600 may be a server computer, a client computer, a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular phone, a smartphone, a web appliance, a network router, a switch or bridge, or any machine capable of executing instructions that specify an operation performed by such device. Further, although only one device 1600 is shown in FIG. 16, the present disclosure is of course related to any collection of devices that individually or collectively execute instructions to perform any one or more of the method logics discussed herein.
[0164] The present disclosure further relates to a program (e.g., a computer program) having instructions that, when executed by a processor, cause the processor to perform some or all of the steps of the methods described herein.
[0165] Still further, the present disclosure relates to a computer-readable (or machine-readable) storage medium storing the above program. Here, the term "computer-readable storage medium" includes, but is not limited to, data repositories in the form of, for example, solid state memory, optical media, and magnetic media.
[0166] [Considerations Regarding Additional Configurations] Unless specifically stated otherwise, as will be apparent from the following discussion, throughout the present disclosure, discussions using terms such as "processing," "computing," "calculating," "determining," "analyzing," etc. refer to the operation and / or processing of a computer or computing system, or similar electronic computing device, that manipulates and / or transforms data represented as a physical quantity (such as electronic) into other data similarly represented as a physical quantity.
[0167] Similarly, the term "processor" can refer to any device or portion of a device that processes electronic data from, for example, registers and / or memory and converts that electronic data into other electronic data that can be stored, for example, in registers and / or memory. A "computer" or "computing machine" or "computing platform" may include one or more processors.
[0168] The method logic described herein, in one example embodiment, is executable by one or more processors that accept computer-readable (also called machine-readable) code that includes a set of instructions to perform at least one of the methods described herein when executed by one or more of the processors. Any processor capable of executing a set of instructions (sequential or otherwise) that specify the operations to be performed is included. Thus, one example is a typical processing system that includes one or more processors. Each processor may include one or more CPUs, a graphics processing unit, and a programmable DSP unit. The processing system may further include a memory subsystem that includes main RAM and / or static RAM, and / or ROM. A bus subsystem may be included for communication between components. The processing system may also be a distributed processing system having processors coupled by a network. If the processing system requires a display, such a display, for example, a liquid crystal display (LCD) or a cathode ray tube (CRT) display, may be included. If manual data input is required, the processing system also includes one or more input devices such as an alphanumeric input unit such as a keyboard, a pointing control device such as a mouse, etc. The processing system may also include a storage system such as a disk drive unit. The processing system may, in some configurations, include an audio output device and a network interface device. Thus, the memory subsystem includes a computer-readable carrier medium carrying computer-readable code (e.g., software) that includes a set of instructions to cause one or more of the methods described herein to be performed when executed by one or more processors. Note that if a method includes several elements, e.g., several steps, the order of such elements is not implied unless specifically stated. The software may reside permanently on the hard disk or, during its execution by the computer system, may reside completely or at least partially in RAM and / or in the processor.Accordingly, the memory and the processor also constitute a computer-readable carrier medium carrying computer-readable code. Further, the computer-readable carrier medium may form, or be included in, a computer program product.
[0169] In an alternative exemplary embodiment, one or more processors may operate as a stand-alone device or may be networked and connected, for example, to other processors in a networked deployment. One or more processors may operate as a server or user machine in a server-user network environment or as a peer machine in a peer-to-peer or distributed network environment. One or more processors may form any machine capable of executing a set (sequential or otherwise) of instructions that specify actions to be taken by a personal computer (PC), tablet PC, personal digital assistant (PDA), cellular telephone, web appliance, network router, switch or bridge, or any of those machines.
[0170] Note that the term "machine" is interpreted to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the method logics discussed herein.
[0171] Accordingly, one exemplary embodiment of each method described herein takes the form of a computer-readable carrier medium carrying a computer program that is executed by a set of instructions, e.g., by one or more processors, e.g., one or more processors that are part of a web server arrangement. Thus, as will be understood by those skilled in the art, the exemplary embodiments of the present disclosure may be embodied as a method, an apparatus such as a special purpose apparatus, an apparatus such as a data processing system, or a computer-readable carrier medium, e.g., a computer program product. The computer-readable carrier medium carries computer-readable code that includes a set of instructions that, when executed by one or more processors, cause the one or more processors to implement the method. Accordingly, aspects of the present disclosure can take the form of a method, an exemplary embodiment that is entirely hardware, an exemplary embodiment that is entirely software, or an exemplary embodiment that combines software and hardware aspects. Further, the present disclosure can take the form of a carrier medium (e.g., a computer program product on a computer-readable storage medium) carrying computer-readable program code embodied in the medium.
[0172] The software may also be transmitted or received over a network via a network interface device. The carrier medium is a single medium in the exemplary embodiments, but the term "carrier medium" should be interpreted to include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated cache and servers) that store one or more sets of instructions. The term "carrier medium" should also be interpreted to include any medium that can store, encode, or carry a set of instructions for execution by one or more processors and that can cause any one or more of the methods of the present disclosure to be executed by one or more processors. The carrier medium can take many forms including, but not limited to, non-volatile media, volatile media, and transmission media. Non-volatile media includes, for example, optical disks, magnetic disks, and magneto-optical disks. Volatile media includes dynamic memory such as main memory. Transmission media includes coaxial cables, copper wire, and fiber optics, including the wires that make up a bus subsystem. Transmission media can also take the form of acoustic or light waves such as those generated during radio frequency and infrared data communications. For example, the term "carrier medium" should be understood to include solid state memory, computer products embodied in optical and magnetic media, a medium having a propagated signal that is detectable by at least one processor or one or more processors and that represents a set of instructions that implement a method when executed, and a transmission medium in a network having a propagated signal that is detectable by at least one of one or more processors and that represents a set of instructions, but is not limited thereto.
[0173] The steps of the method being discussed are, in one exemplary embodiment, understood to be performed by a suitable processor (or processors) of a processing (e.g., computer) system that executes instructions (computer-readable code) stored in storage. Also, the present disclosure is not limited to any particular implementation or programming technique, and it will be understood that the present disclosure may be implemented by any suitable technique for implementing the functions described herein. The present disclosure is not limited to any particular programming language or operating system.
[0174] Throughout the present disclosure, references to "one exemplary embodiment", "some exemplary embodiments", or "an exemplary embodiment" mean that a particular feature, structure, or characteristic described in connection with the exemplary embodiment is included in at least one exemplary embodiment of the present disclosure. Thus, the appearances of the phrases "in one exemplary embodiment", "in some exemplary embodiments", or "in an exemplary embodiment" in various places throughout this specification are not necessarily all referring to the same exemplary embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more exemplary embodiments, as will be apparent to those skilled in the art from the present disclosure.
[0175] As used herein, the use of ordinal adjectives "first", "second", "third", etc. to describe common objects merely indicates that different instances of similar objects are being referred to, and is not intended to imply that the objects so described must be in a particular order, whether in time, space, ranking, or otherwise.
[0176] In the following claims and the description of this specification, any one of the terms "comprising", "comprised of", or "which comprises" is an open term meaning that it includes at least the following elements / features but does not exclude others. Thus, the term "comprising" should not be construed as limiting the means, elements, or steps listed thereafter when used in a claim. For example, the scope of the expression "a device having A and B" should not be limited to "a device including only elements A and B". Any one of the terms "including", "which includes", or "that includes" used in this specification also means that it includes at least the element / function following the term but does not exclude others. Thus, "including" is synonymous with and means "comprising".
[0177] In the above description of the exemplary embodiments of the present disclosure, it should be understood that the various features of the present disclosure are sometimes grouped together in a single exemplary embodiment, figure, or description thereof for the purpose of simplifying the disclosure and facilitating the understanding of one or more aspects of the invention. However, this method of disclosure should not be construed as reflecting an intention that the claims require more features than are explicitly recited in each claim. Rather, as reflected in the following claims, the aspects of the invention are in fewer features than all the features of the single disclosed exemplary embodiment described above. Accordingly, the claims following the description are expressly incorporated herein, and each claim stands independently as a separate exemplary embodiment of the present disclosure.
[0178] Furthermore, some of the exemplary embodiments described herein include some features that are included in other exemplary embodiments but not others, while combinations of features of different exemplary embodiments are intended to be within the scope of the present disclosure and, as will be understood by those skilled in the art, form another exemplary embodiment. For example, in the following claims, any of the claimed exemplary embodiments may be used in any combination.
[0179] In the description provided herein, many specific details are set forth. However, it is understood that the exemplary embodiments of the present disclosure may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this description.
[0180] Accordingly, while what is considered to be the best mode of the present disclosure is described, those skilled in the art will recognize that additional modifications can be made without departing from the spirit of the present disclosure, and it is intended to claim all such changes and modifications that are within the scope of the present disclosure. For example, the above formulas are merely representative of procedures that may be used. It is also possible to add or remove functions in the block diagrams, or exchange operations between functional blocks. Steps can also be added or removed from the methods described within the scope of the present disclosure.
[0181] Further aspects, embodiments, and examples of the present disclosure will become apparent from the exemplary embodiments (EEE) enumerated below.
[0182] EEE1 relates to a method of displaying a spatial audio scene as a compact spatial audio scene including an audio mixing stream and a direction metadata stream, the audio mixing stream consisting of one or more audio elements, the direction metadata stream consisting of time-series direction metadata blocks, each of the direction metadata blocks being related to a corresponding time segment in the audio signal, the spatial audio scene including one or more directional acoustic elements respectively related to each arrival direction, each of the direction metadata blocks including (a) direction information indicating the arrival direction for each of the directional acoustic elements and (b) energy band ratio information indicating the energy in each of the directional acoustic elements relative to the energy in the audio signal in the corresponding time segment for each of the directional acoustic elements and for each of a plurality of pairs of sub-bands.
[0183] EEE2 relates to the method according to EEE1, wherein (a) the energy band ratio information indicates the characteristics of the spatial audio scene in each of the plurality of sub-bands, and (b) for at least one arrival direction, the data included in the direction information indicates the characteristics of the spatial audio scene in two or more clusters of the sub-bands.
[0184] EEE3 relates to a method for processing a compact spatial audio scene including an audio mixed stream and a direction metadata stream to generate a separated spatial audio stream and a residual stream including a set of one or more audio object signals, wherein the audio mixed stream consists of one or more audio signals, the direction metadata stream consists of time-series direction metadata blocks, each of the direction metadata blocks is related to a corresponding time segment in the audio signal, and for each of a plurality of sub-bands, the method includes: (a) determining coefficients of a demixing matrix from direction information and energy band ratio information included in the direction metadata stream; and (b) using the demixing matrix to mix the audio mixed stream to generate the separated spatial audio stream.
[0185] EEE4 relates to the method described in EEE3, wherein each of the direction metadata blocks includes: (a) direction information indicating the direction of arrival for each of the directional acoustic elements; and (b) energy band ratio information indicating the energy of each of the directional acoustic elements with respect to the energy in the corresponding time segment in the audio signal for each of the directional acoustic elements and for each set of two or more sub-bands.
[0186] EEE5 relates to the method described in EEE3, wherein (a) for each of the direction metadata blocks, the direction information and the energy band ratio information are used to form a matrix S representing an approximate covariance of the separated spatial audio stream, and (a) the energy band ratio information is used to form an E representing a remixing matrix that defines the transformation of the separated spatial audio stream to the audio mixed stream, and (b) the demixing matrix E is calculated according to U = S × E * ×(E × S × E * ) -1 and is calculated according to.
[0187] EEE6 relates to the method described in EEE6, and the matrix S is a diagonal matrix.
[0188] EEE7 relates to the method described in EEE3, wherein (a) the residual stream is processed to generate a reconstructed residual stream, (b) each of the audio object signals is processed to generate a corresponding reconstructed object stream, and (c) each of the reconstructed residual stream and the reconstructed object stream is combined to form a reconstructed audio signal, and the reconstructed audio signal includes a directional acoustic element according to the compact spatial audio scene.
[0189] EEE8 relates to the method described in EEE7, wherein the reconstructed audio signal includes two signals for presentation to a listener by transducers at or near each ear to provide a binaural experience of a spatial audio scene including a directional acoustic element according to the compact spatial audio scene.
[0190] EEE9 relates to the method described in EEE7, wherein the reconstructed audio signal includes a plurality of signals representing a spatial audio scene in the form of spherical-harmonic panning functions.
[0191] EEE10 relates to a method of processing a spatial audio scene to generate a compact spatial audio scene including an audio mixed stream and a direction metadata stream, wherein the spatial audio scene includes one or more directional acoustic elements each associated with a respective direction of arrival, and the direction metadata stream consists of a time series of direction metadata blocks, each of which is related to a corresponding time segment in the audio signal, and the method includes (a) means for determining a direction of arrival for one or more of the directional acoustic elements from an analysis of the spatial audio scene, (b) means for determining what portion of the total energy in the spatial scene is contributed by the energy in each of the directional acoustic elements, and (c) means for processing the spatial audio scene to generate the audio mixed stream.
Claims
1. A method for processing a compressed representation of a spatial audio signal to generate a reconstructed representation of the spatial audio signal, wherein the compressed representation includes a channel-based audio signal having a predefined number of channels and metadata, the metadata includes direction information and energy information, the direction information includes an indication of the arrival direction of one or more audio elements in an audio scene, and the energy information includes an indication of each of the signal powers associated with the arrival direction for at least one frequency subband, in the method, generating an audio signal of the one or more audio elements based on the channel-based audio signal, the direction information, and the energy information; generating a residual audio signal in which the one or more audio elements are substantially absent based on the channel-based audio signal, the direction information, and the energy information and having generating the audio signal of the one or more audio elements includes determining coefficients of an inverse mixing matrix M for mapping the channel-based audio signal to an intermediate representation including the residual audio signal and the audio signal based on the direction information and the energy information, determining the coefficients of the inverse mixing matrix M includes for each of the one or more audio elements, determining a panning vector Pan_down(dir) for panning the audio element to the channels of the channel-based audio signal based on the arrival direction dir of the audio element; determining a mixing matrix E used to map the residual audio signal and the audio signal of the one or more audio elements to the channels of the channel-based audio signal based on the determined panning vectors; determining a covariance matrix S of the intermediate representation based on the energy information; determining the coefficients of the inverse mixing matrix M based on the mixing matrix E and the covariance matrix S and including a method.
2. The indication of signal power relates to the ratio of the signal power in a given frequency subband for a given arrival direction to the total signal power in the given frequency subband for the given arrival direction and the given frequency subband. The method according to claim 1.
3. The energy information includes an indication of signal power for each of a plurality of frequency sub-bands. The method according to claim 1 or 2.
4. Panning the audio signals of the one or more audio elements to a set of channels of an output audio format; Generating a reconstructed multi-channel audio signal in the output audio format based on the panned one or more audio elements and the residual audio signal. The method according to any one of claims 1 to 3, further comprising:
5. The mixing matrix E is E = (I N | Pan down (dir 1 ) |... | Pan down (dir P )) determined according to I N is an N×N identity matrix, N indicates the number of channels of the channel-based audio signal, Pan down (dir p ) is the panning vector of the p-th audio element having the associated direction of arrival dir for panning the p-th audio element over the N channels of the channel-based audio signal p where p = 1, ···, P indicates each one of the one or more audio elements, and P indicates the total number of the one or more audio elements The method according to any one of claims 1 to 4.
6. The covariance matrix S is, for 1 ≤ n ≤ N, 【Number 25】 According to, for 1 ≤ p ≤ P, {S} N+p,N+p = e p According to, it is determined as a diagonal matrix, and e p is the signal power related to the arrival direction of the p-th audio element, The method according to claim 5.
7. Determining the coefficients of the inverse mixing matrix based on the mixing matrix and the covariance matrix includes determining a pseudo-inverse matrix based on the mixing matrix and the covariance matrix. The method according to any one of claims 1 to 6.
8. The inverse mixing matrix M is M = S × E * × (E × S × E * ) -1 Determined according to × indicates a matrix product, and * indicates a conjugate transpose of a matrix. The method according to any one of claims 1 to 7.
9. The channel-based audio signal is a first-order ambisonics signal. The method according to any one of claims 1 to 8.
10. A program having instructions that, when executed by a processor, cause the processor to execute all steps of the method according to any one of claims 1 to 9.
11. A computer-readable storage medium storing the program according to claim 10.
12. Having a processor and a memory coupled to the processor, The processor is configured to execute all steps of the method according to any one of claims 1 to 9. An apparatus.
Citation Information
Patent Citations
Temporal spatial audio parameter smoothing
GB2571949A
Spatial audio rendering and encoding
JP2015509212A
Audio Object Extraction Supported by Video Content
JP2018511974A
Spatial audio coding based on universal spatial cues
US20070269063A1
Determination of targeted spatial audio parameters and associated spatial audio playback
WO2019086757A1