Multi-channel audio encoding and decoding using directional metadata

By analyzing spatial audio signals for directions and power, and using metadata to generate a channel-based signal, the method addresses the high bandwidth challenge of spatial audio representation, enabling efficient storage and transmission with high-quality reconstruction.

JP2025135018APending Publication Date: 2025-09-17DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025115521
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-10-01
Filing Date
2025-07-09
Publication Date
2025-09-17

AI Technical Summary

Technical Problem

Existing audio recording and playback technologies face challenges in efficiently representing spatial audio scenes due to high bandwidth and storage requirements, particularly in channel-based and object-based formats like Higher Order Ambisonics (HOA), necessitating a more compact representation.

Method used

A method for processing spatial audio signals to generate a compressed representation by analyzing directions of arrival and signal power, generating metadata with directional and energy information, and creating a channel-based audio signal, which can be reconstructed with a decoder using unmixing matrices.

Benefits of technology

This approach allows for a compact representation of spatial audio scenes that maintains a high-quality reconstruction, reducing data requirements for storage and transmission while preserving the spatial audio experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025135018000001_ABST
    Figure 2025135018000001_ABST
Patent Text Reader

Abstract

To provide a method for processing a spatial audio signal and generating compressed representation of the spatial audio signal.SOLUTION: The method includes: analyzing a spatial audio signal so as to determine an arriving direction of one or more audio elements; determining each indication of signal power associated with the arriving direction for at least one frequency sub-band; generating metadata including direction information including an indication of the arriving direction of the audio element and energy information including each indication of signal power; generating a channel base audio signal having the predefined number of channels based on the spatial audio signal; and outputting the channel base audio signal and the metadata as compressed representation. The present disclosure further relates to a method for processing compressed representation of the spatial audio signal and generating reconstructed representation of the spatial audio signal, and to a corresponding device, a program and a storage medium.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to U.S. Provisional Patent Application No. 62 / 927,790, filed October 30, 2019, and U.S. Provisional Patent Application No. 63 / 086,465, filed October 1, 2020, each of which is incorporated herein by reference in its entirety.

[0002] [Technical field] The present disclosure relates generally to audio signal processing, and more particularly to methods of processing a spatial audio signal (a spatial audio scene) to generate a compressed representation of the spatial audio signal, and to methods of processing the compressed representation of the spatial audio signal to generate a reconstructed representation of the spatial audio signal. [Background technology]

[0003] Human hearing allows listeners to perceive their environment in the form of a spatial audio scene, where the term "spatial audio scene" is used to refer to the acoustic environment around the listener or the acoustic environment perceived in the listener's mind.

[0004] While human experience is associated with a spatial audio scene, the art of audio recording and playback involves the capture, manipulation, transmission, and playback of audio signals or audio channels. The term "audio stream" is used to refer to a collection of one or more audio signals, especially when the audio stream is intended to represent a spatial audio scene.

[0005] The audio streams can be played back to listeners via electro-acoustic transducers or by other means to provide one or more listeners with a listening experience in the form of a spatial audio scene. The goal of audio recordists and audio artists is generally to create audio streams that are intended to provide listeners with the experience of a particular spatial audio scene.

[0006] Audio streams may be accompanied by associated data, called metadata, that aids in the playback process. The accompanying metadata may contain information that changes over time. This information may be used to affect changes in processing applied during the playback process.

[0007] In the following, the term "captured audio experience" may be used to refer to an audio stream and associated metadata.

[0008] In some applications, the metadata consists solely of data indicating the intended loudspeaker placement for playback. Often, this metadata is omitted, assuming that the placement of playback speakers is standardized. In this case, the captured audio experience consists of only the audio stream. One example of such a captured audio experience is a two-channel audio stream recorded on a compact disc, where the intended playback system is assumed to be in the form of two loudspeakers placed in front of the listener.

[0009] Alternatively, a captured audio experience in the form of a scene-based multi-channel audio signal may be intended for presentation to a listener by processing the audio signal with a mixing matrix to generate a set of speaker signals, each of which is then reproduced on a respective loudspeaker, which may optionally be spatially positioned around the listener. In this example, the mixing matrix may be generated based on a scene-based format and prior knowledge of the placement of the reproduction speakers.

[0010] An example of a scene-based format is Higher Order Ambisonics (HOA), and an example of how to calculate an appropriate mixing matrix is ​​given in "Ambisonics", Franz Zotter and Matthias Frank, ISBN: 978-3-030-17206-0, Chapter 3, which is incorporated herein by reference.

[0011] Such scene-based formats typically contain a large number of channels or audio objects, resulting in relatively high bandwidth or storage requirements when transmitting or storing spatial audio signals in these formats.

[0012] Therefore, there is a need for a compact representation of spatial audio signals that represent a spatial audio scene. This applies to both channel-based and object-based spatial audio signals. Summary of the Invention

[0013] The present disclosure proposes a method for processing a spatial audio signal to generate a compressed representation of the spatial audio signal, a method for processing the compressed representation of the spatial audio signal to generate a reconstructed representation of the spatial audio signal, and corresponding devices, programs, and computer-readable storage media.

[0014] One aspect of the present disclosure relates to a method for processing a spatial audio signal to generate a compressed representation of the spatial audio signal. The spatial audio signal may be, for example, a multi-channel signal or an object-based signal. The compressed representation may be a compact or reduced-size representation. The method may include analyzing the spatial audio signal to determine directions of arrival of one or more audio elements in an audio scene represented by the spatial audio signal (spatial audio scene). The audio elements may be dominant audio elements. The (dominant) audio elements may relate, for example, to (dominant) acoustic objects, (dominant) sound sources, or (dominant) acoustic components in the audio scene. The one or more audio elements may include one to ten audio elements, such as four audio elements. The directions of arrival may correspond to positions on a unit sphere indicating the perceived location of the audio elements. The method may further include determining, for at least one frequency subband of the spatial audio signal (e.g., for all frequency subbands), each indication of signal power associated with the determined directions of arrival. The method may further include generating metadata including directional information and energy information, where the directional information includes an indication of a determined direction of arrival of one or more audio elements and the energy information includes an indication of each of the signal powers associated with the determined direction of arrival. The method may further include generating a channel-based audio signal having a predefined number of channels based on the spatial audio signal. The channel-based audio signal may be referred to as an audio mixture signal or an audio mixture stream. It will be understood that the number of channels of the channel-based audio signal may be less than the number of channels or the number of objects of the spatial audio signal. The method may also further include outputting the channel-based audio signal and the metadata as a compressed representation of the spatial audio signal. The metadata may be associated with a metadata stream.

[0015] Thereby, a compressed representation of the spatial audio signal can be generated that includes a limited number of channels, but by appropriate use of the directional and energy information, the decoder can still generate a reconstructed version of the original spatial audio signal that is a very good approximation of the original spatial audio signal as far as the representation of the original spatial audio signal is concerned.

[0016] In some embodiments, analyzing the spatial audio signal may be based on multiple frequency subbands of the spatial audio signal. For example, the analysis may be based on the entire frequency range of the spatial audio signal (i.e., the entire signal), i.e., the analysis may be based on all frequency subbands.

[0017] In some embodiments, analyzing the spatial audio signal may include applying a scene analysis to the spatial audio signal, whereby (the direction of) dominant audio elements in the audio scene can be determined in a reliable and efficient manner.

[0018] In some embodiments, the spatial audio signal may be a multi-channel audio signal. Alternatively, the spatial audio signal may be an object-based audio signal. In this case, the method may further include converting the object-based audio signal into a multi-channel audio signal before applying the scene analysis. This allows for meaningful application of scene analysis tools to the audio signal.

[0019] In some embodiments, the indication of signal power associated with a given direction of arrival may relate to a ratio of the signal power in a frequency subband for the given direction of arrival to the total signal power in the frequency subband.

[0020] In some embodiments, an indication of signal power may be determined for each of a plurality of frequency subbands. In this case, they may relate to the ratio of the signal power in a given frequency subband for a given direction of arrival to the total signal power in the given frequency subband for a given direction of arrival and for a given frequency subband. In particular, the indication of signal power may be determined for each subband, while the determination of the (dominant) direction of arrival may be performed for the entire signal (i.e., based on all frequency subbands).

[0021] In some embodiments, analyzing the spatial audio signal, determining the respective indications of signal power, and generating the channel-based audio signal may be performed for each time segment. Thus, a compressed representation may be generated and output for each of a plurality of time segments by the downmix audio signal and metadata (metadata block) for each time segment. Alternatively or additionally, analyzing the spatial audio signal, determining the respective indications of signal power, and generating the channel-based audio signal may be performed based on a time-frequency representation of the spatial audio signal. For example, the above steps may be performed based on a discrete Fourier transform (e.g., STFT) of the spatial audio signal. That is, for each time segment (time block), the above steps may be performed based on the time-frequency bins (FFT bins) of the spatial audio signal, i.e., based on the Fourier coefficients of the spatial audio signal.

[0022] In some embodiments, the spatial audio signal may be an object-based audio signal including a plurality of audio objects and associated direction vectors. In that case, the method may further include generating a multi-channel audio signal by panning the audio objects to a predefined set of audio channels, wherein each audio object may be panned to the predefined set of audio channels according to its direction vector. Furthermore, the channel-based audio signal may be a downmix signal generated by applying a downmix operation to the multi-channel audio signal. The multi-channel audio signal may be, for example, a high-order Ambisonics signal.

[0023] In some embodiments, the spatial audio signal may be a multi-channel audio signal, in which case the channel-based audio signal may be a downmix signal generated by applying a downmix operation to the multi-channel audio signal.

[0024] Another aspect of the present disclosure relates to a method for processing a compressed representation of a spatial audio signal to generate a reconstructed representation of the spatial audio signal. The compressed representation may include a channel-based audio signal having a predefined number of channels and metadata. The metadata may include directional information and energy information. The directional information may include an indication of a direction of arrival of one or more audio elements in an audio scene (spatial audio scene). The energy information may include, for at least one frequency subband, an indication of a signal power associated with the direction of arrival. The method may include generating an audio signal of the one or more audio elements based on the channel-based audio signal, the directional information, and the energy information. The method may further include generating a residual audio signal substantially absent of the one or more audio elements based on the channel-based audio signal, the directional information, and the energy information. The residual signal may be represented in the same audio format as the channel-based audio signal, e.g., have the same number of channels.

[0025] In some embodiments, the indication of signal power associated with a given direction of arrival may relate to a ratio of signal power in a frequency subband for the given direction of arrival to the total signal power in the frequency subband.

[0026] In some embodiments, the energy information may include an indication of signal power for each of a plurality of frequency subbands, where the indication of signal power may relate to a ratio of signal power in a given frequency subband for a given direction of arrival to total signal power in the given frequency subband for a given direction of arrival and for the given frequency subband.

[0027] In some embodiments, the method may further include panning audio signals of one or more audio elements to a set of channels of the output audio format. The method may also include generating a reconstructed multi-channel audio signal in the output audio format based on the panned one or more audio elements and the residual audio signal. The output audio format may relate to an output representation such as, for example, HOA or any other suitable multi-channel format. Generating the reconstructed multi-channel audio signal may include upmixing the residual signal to the set of channels of the output audio format. Generating the reconstructed multi-channel audio signal may further include summing the panned one or more audio elements and the upmixed residual signal.

[0028] In some embodiments, generating the audio signals of the one or more audio elements may include determining, based on the directional information and the energy information, coefficients of an unmixing matrix M for mapping the channel-based audio signal to an intermediate representation including the residual audio signal and the audio signals of the one or more audio elements. The intermediate representation may also be referred to as a separated or separable representation, or a hybrid representation.

[0029] In some embodiments, determining the coefficients of the unmixing matrix M includes, for each of one or more audio elements, determining a panning vector Pan for panning the audio element to a channel of the channel-based audio signal based on a direction of arrival dir of the audio element. down (dir). Determining the coefficients of the unmixing matrix M may further include determining a mixing matrix E used to map the residual audio signal and the audio signals of the one or more audio elements to channels of the channel-based audio signal based on the determined panning vector. Determining the coefficients of the unmixing matrix M may further include determining a covariance matrix S of the intermediate representation based on the energy information. Determining the covariance matrix S may include determining a covariance matrix S based on the determined panning vector Pan down Determining the coefficients of the unmixing matrix M above may also further include determining the coefficients of the unmixing matrix M based on the mixing matrix E and the covariance matrix S.

[0030] In some embodiments, the mixing matrix E is E=(I N |Pan down (dir1)|···|Pan down (dir P )) where I N can be an N×N identity matrix, where N denotes the number of channels in the channel-based audio signal, and Pan down (dir p ) is the associated direction of arrival dir that pans (maps) the p-th audio element onto the N channels of the channel-based audio signal. p where p=1, . . . , P indicates each one of the one or more audio elements and P indicates the total number of the one or more audio elements. Thus, the matrix E can be an N×P matrix. The matrix E may be determined for each of a plurality of time segments k. In this case, the matrix E and the direction of arrival dir pwill have an index k indicating the time segment. For example, E k =(I N |Pan down (dir k,1 )|···|Pan down (dir k,P )). Even though the proposed method may operate band-wise, the matrix E will be the same for all frequency subbands.

[0031] In some embodiments, the covariance matrix S is, for 1≦n≦N,

number

number

[0032] In some embodiments, determining the coefficients of the unmixing matrix M based on the mixing matrix E and the covariance matrix S may include determining a pseudo-inverse matrix based on the mixing matrix E and the covariance matrix S.

[0033] In some embodiments, the unmixing matrix M is M=S×E * ×(E×S×E * ) -1 where "x" denotes matrix multiplication and "*" denotes the conjugate transpose of a matrix. An unmixing matrix M may be determined for each of a plurality of time segments k and / or for each of a plurality of frequency subbands b. In that case, matrices M and S will have an index k indicating the time segment and / or an index b indicating the frequency subband, and matrix E will have an index k indicating the time segment. For example, M k,b =S k,b ×E * k ×(E k ×S k,b ×E * k ) -1 is.

[0034] In some embodiments, the channel-based audio signal may be a first-order Ambisonics signal.

[0035] Another aspect relates to an apparatus comprising a processor and a memory coupled to the processor, the processor being configured to perform all the steps of the method according to any one of the above aspects and embodiments.

[0036] Another aspect of the present disclosure relates to a program comprising instructions that, when executed by a processor, cause the processor to perform all of the steps of the above method.

[0037] Yet another aspect of the present disclosure relates to a computer-readable storage medium storing the above program.

[0038] Further embodiments of the present disclosure include an efficient method for representing a spatial audio scene in the form of an audio mixing stream and a directional metadata stream, the directional metadata stream including data indicative of the position of directional acoustic elements in the spatial audio scene and data indicative of the power of each directional acoustic element in a number of sub-bands relative to the total power of the spatial audio scene in that sub-band. Still other embodiments relate to methods for determining a directional metadata stream from an input spatial audio scene and methods for generating a reconstructed audio scene from the directional metadata stream and the associated audio mixing stream.

[0039] In some embodiments, the method is used to represent a spatial audio scene in a more compact form as a compact spatial audio scene comprising an audio mixing stream and a directional metadata stream, where said audio mixing stream comprises one or more audio signals and said directional metadata stream comprises a time sequence of directional metadata blocks, each directional metadata block relating to a corresponding time segment of an audio signal. The spatial audio scene comprises one or more directional acoustic elements, each associated with a respective direction of arrival. Each directional metadata block comprises: Directional information indicating the direction of arrival for each of the directional acoustic elements, and For each directional acoustic element and for each set of two or more subbands, energy band fraction information indicating the energy in each directional acoustic element relative to the energy in a corresponding time segment of the audio signal. Includes.

[0040] In some embodiments, a method is used to process a compact spatial audio scene comprising an audio mixing stream and a directional metadata stream to generate a separated spatial audio stream and a residual stream comprising one or more sets of audio object signals, where said audio mixing stream comprises one or more audio signals and said directional metadata stream comprises a time sequence of directional metadata blocks, each directional metadata block associated with a corresponding time segment of an audio signal. For each of a plurality of sub-bands, the method comprises: Determining the coefficients of a demixing matrix from the directional information and energy band ratio information contained in the directional metadata stream; and Mixing the audio signals using the demixing matrix to generate the separated spatial audio streams. Includes.

[0041] In some embodiments, a method is used to process a spatial audio scene to generate a compact spatial audio scene comprising an audio mixing stream and a directional metadata stream, where said spatial audio scene comprises one or more directional sound elements each associated with a respective direction of arrival, and where said directional metadata stream consists of a time sequence of directional metadata blocks, each directional metadata block relating to a corresponding time segment of an audio signal. The method comprises: - determining a direction of arrival for one or more of the directional acoustic elements from an analysis of the spatial audio scene; determining what portion of the total energy in the spatial scene is contributed by the energy in each of the directional acoustic elements; and Processing the spatial audio scene to generate an audio mixing stream Includes.

[0042] It will be appreciated that the above steps may be implemented by any suitable means or unit, i.e., for example, by one or more computer processors.

[0043] It will also be understood that apparatus features and method steps may be interchanged in many ways. In particular, details of the disclosed methods can be implemented by corresponding apparatus, and vice versa, as will be understood by those skilled in the art. Furthermore, any statements made above with respect to a method will be understood to apply equally to the corresponding apparatus, and vice versa.

[0044] Exemplary embodiments of the present disclosure are illustrated by way of example in the accompanying drawings, in which like reference numbers indicate the same or similar elements and in which: [Brief explanation of the drawings]

[0045] [Figure 1] 1 schematically illustrates an example arrangement of an encoder for generating a compressed representation of a spatial audio scene and a corresponding decoder for generating a reconstructed audio scene from the compressed representation, according to an embodiment of the present disclosure; [Figure 2] 3A and 3B schematically represent another example arrangement of an encoder for generating a compressed representation of a spatial audio scene and a corresponding decoder for generating a reconstructed audio scene from the compressed representation, according to an embodiment of the present disclosure; [Figure 3] 1 schematically illustrates an example of generating a compressed representation of a spatial audio scene according to an embodiment of the present disclosure. [Figure 4] 1 schematically illustrates an example of decoding a compressed representation of a spatial audio scene to form a reconstructed audio scene according to an embodiment of the present disclosure. [Figure 5] 1 is a flowchart illustrating an example method for processing a spatial audio scene to generate a compressed representation of the spatial audio scene, according to an embodiment of the present disclosure. [Figure 6]1 is a flowchart illustrating an example method for processing a spatial audio scene to generate a compressed representation of the spatial audio scene, according to an embodiment of the present disclosure. [Figure 7] 3A-3C schematically illustrate example details of generating a compressed representation of a spatial audio scene according to an embodiment of the present disclosure. [Figure 8] 3A-3C schematically illustrate example details of generating a compressed representation of a spatial audio scene according to an embodiment of the present disclosure. [Figure 9] 3A-3C schematically illustrate example details of generating a compressed representation of a spatial audio scene according to an embodiment of the present disclosure. [Figure 10] 3A-3C schematically illustrate example details of generating a compressed representation of a spatial audio scene according to an embodiment of the present disclosure. [Figure 11] 3A-3C schematically illustrate example details of generating a compressed representation of a spatial audio scene according to an embodiment of the present disclosure. [Figure 12] 3A-3C schematically illustrate example details of decoding a compressed representation of a spatial audio scene to form a reconstructed audio scene according to an embodiment of the present disclosure. [Figure 13] 1 is a flowchart illustrating an example method for decoding a compressed representation of a spatial audio scene to form a reconstructed audio scene, according to an embodiment of the present disclosure. [Figure 14] 14 is a flowchart showing details of the method of FIG. 13. [Figure 15] 10 is a flowchart illustrating another example method for decoding a compressed representation of a spatial audio scene to form a reconstructed audio scene, according to an embodiment of the present disclosure. [Figure 16] 1 schematically illustrates an apparatus for generating a compressed representation of a spatial audio scene and / or decoding the compressed representation of a spatial audio scene to form a reconstructed audio scene, according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0046] Generally, the present disclosure relates to enabling storage and / or transmission of spatial audio scenes using a reduced amount of data.

[0047] Audio processing concepts that may be used within the context of the present disclosure are now described.

[0048] [Panning Function] A multi-channel audio signal (or audio stream) can be formed by panning individual acoustic elements (or audio objects) according to a linear mixing law. For example, a set of R audio objects can be panned over R signals {o r (t): 1 ≦ r ≦ R}, where n (t): 1≦n≦N} is:

number

[0049] Panning function Pan(θ r ) is used to generate the object signal o r represents a column vector containing N scale factors (panning gains) that indicate the gains used to blend (t), where θ r indicates the position of each object.

[0050] One possible panning function is the first-order Ambisonics (FOA) panner. An example of an FOA panning function is:

number

[0051] An alternative panning function is the third-order Ambisonics panner (3OA). An example of a 3OA panning function is:

number

[0052] As one skilled in the art will appreciate, it is understood that the present disclosure is not limited to FOA or HOA panning functions, and that other panning functions may be contemplated for use.

[0053] [Short-time Fourier transform] An audio stream consisting of one or more audio signals may be transformed, for example, into a short-term Fourier transform (STFT) form. For this purpose, a discrete Fourier transform may be applied to (optionally windowed) time segments of the audio signals (e.g., channels, audio object signals) of the audio stream. This processing applied to an audio signal x(t) may be expressed as follows: X c,k (f)=STFT{x c (t)} (4) It is understood that the STFT is an example of a time-frequency transform, and the present disclosure should not be limited to the STFT.

[0054] In equation (4), the variable X c,k (f) is the audio time segment k in frequency bin f (1≦f≦F). (outside 1) Figure 1 shows the short-time Fourier transform of channel c (1 < c < NumChans) for TIFF2025135018000007.tif6170, where F denotes the number of frequency bins generated by the Discrete Fourier Transform. It is understood that the terminology used here is exemplary, and specific implementation details of various STFT methods (including various window functions) may be known in the art. An audio time segment k is defined, for example, as a range of audio samples centered at t = k × stride + constant, such that the time segments are evenly spaced in time, at intervals equal to the stride.

[0055] STFT values ​​(e.g., X c,k (1), X c,k (2),···,X c,k (F)) is sometimes called an FFT bin.

[0056] Furthermore, the STFT format can be converted into an audio stream. The resulting audio stream can be an approximation to the original input:

number

[0057] Frequency Banded Analysis Characterization data may be formed from the audio stream, the characterization data relating to a number of frequency bands (frequency sub-bands), where a band (sub-band) is defined by a region of the frequency range.

[0058] For example, the signal power in channel c of a stream in frequency band b (where the number of bands is B and 1≦b≦B) is calculated by dividing band b by FFT bin f min ≦f≦f max When it comes to:

number

[0059] According to a more general example, a frequency band b is defined by a weighting vector FR that assigns a weight to each frequency bin. b (f), whereby an alternative calculation of power in a band may be defined by:

number

[0060] In a further generalization of equation (7), the STFT of a stream of C audio signals can be processed to generate covariances in multiple bands, where the covariance R b,k is a C × C matrix with elements {R b,k} i,j teeth:

number

[0061] In another example, a bandpass filter may be used to form a filtered signal representing the original audio stream in a frequency band according to the bandpass filter response. For example, c (t) is x c x' represents a signal whose energy is mainly derived from band b of (t). c,b (t), and thus time block k (time sample t min ≦t≦t max An alternative way to calculate the covariance of the stream in band b of

number

[0062] [Frequency Banded Mixture] An audio stream with N channels:

number

number

[0063] Furthermore, an alternative mixing process may be implemented in the STFT domain, where the matrix Q can take different values ​​for each time block t and for each frequency band b. In this case, the process is:

number

number

[0064] It will be appreciated that alternative methods may be used to produce behavior equivalent to the process shown in equation (13).

[0065] [Example Practice] Next, exemplary implementations of methods and apparatus according to embodiments of the present disclosure will be described in more detail.

[0066] Broadly speaking, a method according to an embodiment of the present disclosure represents a spatial audio scene in the form of an audio mixing stream and a directional metadata stream, the directional metadata stream including data indicative of the position of directional acoustic elements in the spatial audio scene and data indicative of the power of each directional acoustic element in a number of sub-bands relative to the total power of the spatial audio scene in that sub-band. A further method according to an embodiment of the present disclosure relates to determining a directional metadata stream from an input spatial audio scene and generating a reconstructed (e.g., restored) audio scene from the directional metadata stream and the associated audio mixing stream.

[0067]

[0006] Exemplary methods according to embodiments of the present disclosure are efficient (e.g., in terms of reducing storage or transmission data) in representing spatial audio scenes. The spatial audio scenes may be represented by spatial audio signals. The above methods may be implemented by defining a storage or transmission format (e.g., Compact Spatial Audio Stream) consisting of an audio mixing stream and a metadata stream (e.g., a directional metadata stream).

[0068] An audio mixture stream comprises a number of audio signals carrying a compact representation of a spatial audio scene. As such, the audio mixture stream may be associated with a channel-based audio signal having a predefined number of channels. It is understood that the number of channels of the channel-based audio signal is less than the number of channels of the spatial audio signal or the number of audio objects. For example, the channel-based audio signal may be a first-order Ambisonics audio signal. In other words, the compact spatial audio stream may include an audio mixture stream in the form of a first-order Ambisonics representation of a sound field.

[0069] The (directional) metadata stream contains metadata that defines the spatial characteristics of a spatial audio scene. The directional metadata may consist of a sequence of directional metadata blocks. Each directional metadata block contains metadata that indicates the characteristics of the spatial audio scene for a corresponding time segment in the audio mixing stream.

[0070] Typically, the metadata comprises directional information and energy information. The directional information comprises an indication of the direction of arrival of one or more (dominant) audio elements in the audio scene. The energy information comprises, for each direction of arrival, an indication of the signal power associated with the determined direction of arrival. In some implementations, an indication of the signal power may be provided for one, several, or each of a number of bands (frequency subbands). Furthermore, the metadata may be provided for each of a number of consecutive time segments, e.g., in the form of a metadata block.

[0071] In one example, the metadata (directional metadata) includes metadata that describes the characteristics of the spatial acoustic scene across multiple frequency bands, the metadata being: one or more directions (e.g., direction of arrival) indicating the location of an audio object (audio element) in the spatial audio scene, and The proportion of energy (or spatial power) in each frequency band attributed to each audio object (e.g., from each direction) Includes.

[0072] Details regarding the determination of directional and energy information are given below.

[0073] 1 illustrates schematically an example of an arrangement employing embodiments of the present disclosure. Specifically, the diagram shows an arrangement 100 in which a spatial audio scene 10 is input to a scene encoder 200, which generates an audio mixing stream 30 and a directional metadata stream 20. The spatial audio scene 10 may be represented by a spatial audio signal or a spatial audio stream that is input to the scene encoder 200. The audio mixing stream 30 and the directional metadata stream 20 together form an example of a compact spatial audio scene, i.e., a compressed representation of the spatial audio scene 10 (or of the spatial audio signal).

[0074] The compressed representations, i.e. the mixed audio stream 30 and the directional metadata stream 20, are input to a scene decoder 300, which generates a reconstructed audio scene 50. The audio elements present in the spatial audio scene 10 are represented in the audio mixed stream 30 according to a mixing panning function.

[0075] 2 schematically illustrates another example of an arrangement using embodiments of the present disclosure. Specifically, the diagram illustrates an alternative arrangement 110 in which a compact spatial audio scene consisting of an audio mixing stream 30 and a directional metadata stream 20 is further encoded by providing the audio mixing stream 30 to an audio encoder 35 to generate a bit-rate reduced encoded audio stream 37, and by providing the directional metadata stream 20 to a metadata encoder 25 to generate an encoded metadata stream 27. The bit-rate reduced encoded audio stream 37 and the encoded metadata stream 27 together form the encoded (bit-rate reduced encoded) spatial audio scene.

[0076] The encoded spatial audio scene can be recovered by first applying the bitrate reduced encoded audio stream 37 and the encoded metadata stream 27 to respective decoders 36 and 26 to generate a playback audio mixing stream 38 and a playback direction metadata stream 28. The playback streams 38, 28 will be the same or approximately equal to the respective streams 30, 20. The playback audio mixing stream 38 and the playback direction metadata stream 28 can be decoded by a decoder 300 to generate a reconstructed audio scene 50.

[0077] 3 schematically illustrates an example of an arrangement for generating a bit-rate reduced encoded audio stream and an encoded metadata stream from an input spatial audio scene. In particular, the diagram shows an arrangement 150 of a scene encoder 200 supplying a directional metadata stream 20 and an audio mixing stream 30 to respective encoders 25, 35 to generate an encoded spatial audio scene 40 comprising a bit-rate reduced encoded audio stream 37 and an encoded metadata stream 27. The encoded spatial audio stream 40 is preferably arranged to be suitable for storage and / or transmission with reduced data requirements relative to the data required for storage / transmission of the original spatial audio scene.

[0078] 4 shows a schematic representation of an example of an arrangement for generating a reconstructed spatial audio scene from a bit-rate reduced coded audio stream and a coded metadata stream. In particular, the diagram shows that a coded spatial audio stream 40, consisting of a bit-rate reduced coded audio stream 37 and a coded metadata stream 27, is provided as input to decoders 36, 26, respectively, to generate an audio mixing stream 38 and a directional metadata stream 28. The streams 38, 28 are then processed by a scene decoder 300 to generate a reconstructed audio scene 50.

[0079] The details of generating a compact spatial audio scene, ie a compressed representation of a spatial audio scene (or of a spatial audio signal / spatial audio stream) are now described.

[0080] 5 is a flow chart of an example method 500 for processing a spatial audio signal to generate a compressed representation of the spatial audio signal. The method 500 comprises steps S510 to S550.

[0081] In step S510, the spatial audio signal is analyzed to determine the directions of arrival of one or more audio elements (e.g., dominant audio elements) in the audio scene represented by the spatial audio signal (spatial audio scene). A (dominant) audio element may relate, for example, to a (dominant) acoustic object, a (dominant) sound source, or a (dominant) acoustic component in the audio scene. Analyzing the spatial audio signal may include or relate to applying a scene analysis to the spatial audio signal. It will be understood that a range of suitable scene analysis tools is known to those skilled in the art. The directions of arrival determined in this step may correspond to positions on a unit sphere indicating the (perceived) location of the audio element.

[0082] Consistent with the above description of frequency-banded analysis, the analysis of the spatial audio signal in step S510 may be based on multiple frequency subbands of the spatial audio signal. For example, the analysis may be based on the entire frequency range (i.e., the entire signal) of the spatial audio signal. That is, the analysis may be based on all frequency subbands.

[0083] In step S520, an indication of each of the signal powers associated with the determined directions of arrival is determined for at least one frequency subband of the spatial audio signal.

[0084] In step S530, metadata is generated that includes directional information and energy information. The directional information includes an indication of the determined direction of arrival of one or more audio elements. The energy information includes an indication of each of the signal powers associated with the determined direction of arrival. The metadata generated in this step may relate to a metadata stream.

[0085] In step S540, a channel-based audio signal having a predefined number of channels is generated based on the spatial audio signal.

[0086] Finally, in step S550, the channel-based audio signal and the metadata are output as a compressed representation of the spatial audio signal.

[0087] It will be appreciated that the above steps may be performed in any order, or in parallel with one another, so long as the order of the steps ensures that the necessary inputs for each step are available.

[0088] In general, a spatial scene (or spatial audio signal) can be considered to consist of the sum of acoustic signals incident on the listener from a set of directions relative to the listening position. A spatial audio scene can therefore be modeled as a set of R acoustic objects. Object r (1≦r≦R) is a set of objects with a direction vector θ r An audio signal o incident on the listening position from a direction of arrival defined by r (t). The direction vector is also related to the time-varying vector θ r (t) may also be used.

[0089] Thus, according to some implementations, a spatial audio signal (spatial audio stream) may be defined as an object-based spatial audio signal (object-based spatial audio scene) in the form of a set of audio signals and associated direction vectors: Spatial Audio Scene (Object-Based) ={o r (t),θ r (t):1≦r≦R} (14) Furthermore, according to some implementations, the spatial audio signal (spatial audio stream) is obtained by converting a short-time Fourier transform signal O r,k (f), and the direction vector may be specified according to the block index k, such that: Spatial Audio Scene (Object-Based) ={O r,k (f),θ r(t):1≦r≦R} (15) is.

[0090] Alternatively, a spatial audio signal (spatial audio stream) may be represented in terms of a channel-based spatial audio signal (channel-based spatial audio scene). A channel-based stream consists of a set of audio signals, where each sound object from a spatial audio scene is mixed into a channel by a panning function (Pan(θ)) according to equation (1). As an example, a channel-based spatial audio scene {C q,k (f): 1≦q≦Q} is

number

[0091] It will be appreciated that many characteristics of the channel-based spatial audio scene are determined by the choice of panning function, and in particular, the length (Q) of the column vector returned by the panning function determines the number of audio channels included in the channel-based spatial audio scene. Generally speaking, a higher quality representation of the spatial audio scene can be achieved by a channel-based spatial audio scene that includes a larger number of channels.

[0092] As an example, in step S540 of method 500, the spatial audio signal (spatial audio scene) may be processed to generate a channel-based audio signal (channel-based stream) according to equation (16). The panning function may be selected to result in a relatively low-resolution representation of the spatial audio scene. For example, the panning function may be selected to be a first-order Ambisonics (FOA) function as defined in equation (2). As such, the compressed representation may be a compact or reduced-size representation.

[0093] 6 is a flowchart providing another formulation of a method 600 for generating a compact representation of a spatial audio scene. The method 600 is provided with an input stream in the form of a spatial audio scene or a scene-based stream and generates a compact spatial audio scene as a compact representation. To this end, the method 600 comprises steps S610 to S660, among which step S610 may be considered to correspond to step S510, step S620 may be considered to correspond to step S520, step S630 may be considered to correspond to step S540, step S650 may be considered to correspond to step S530, and step S660 may be considered to correspond to step S550.

[0094] In step S610, the input stream is analyzed to determine the dominant directions of arrival.

[0095] In step S620, for each band (frequency subband), the proportion of energy allocated to each direction relative to the total energy in the stream in that band is determined.

[0096] In step S630, a downmix stream is formed that includes multiple audio channels representing a spatial audio scene.

[0097] In step S640, the downmix stream is encoded to form a compressed representation of the stream.

[0098] In step S650, the direction information and energy ratio information are encoded to form encoded metadata.

[0099] Finally, in step S660, the encoded downmix stream is combined with the encoded metadata to form a compact spatial audio scene.

[0100] It will be appreciated that the above steps may be performed in any order, or in parallel with one another, so long as the order of the steps ensures that the necessary inputs for each step are available.

[0101] 7 to 11 schematically represent example details for generating a compressed representation of a spatial audio scene according to an embodiment of the present disclosure. It will be understood that details described below, e.g., analyzing spatial audio signals to determine a direction of arrival, determining an indication of signal power associated with the determined direction of arrival, generating metadata including directional information and energy information, and / or generating a channel-based audio signal including a pre-defined number of channels, can be independent of the specific system configuration and may, for example, be applied to any of the configurations shown in Figures 7 to 11 or any suitable alternative configuration.

[0102] Fig. 7 schematically illustrates a first example of details for generating a compressed representation of a spatial audio scene. Specifically, Fig. 7 illustrates a scene encoder 200 in which a spatial audio scene 10 is processed by a downmix function 203 to generate an N-channel audio mixed stream 30, e.g., according to steps S540 and S630. In some embodiments, the downmix function 203 may include a panning operation according to equation (1) or equation (16), where a downmix panning function is selected, i.e.,

number

number

[0103] For each audio time segment, the scene analysis 202 takes as input the spatial audio scene and determines the directions of arrival of up to P dominant sound components in the spatial audio scene, for example according to steps S510 and S610. Typical values ​​of P are between 1 and 10, with a preferred value of P being P≈4. Thus, the one or more audio elements determined in step S510 may comprise between 1 and 10 audio elements, such as, for example, four audio elements.

[0104] The analysis 202 generates metadata 20 consisting of directional information 21 and energy band ratio information 22 (energy information). Optionally, the scene analysis 202 may also provide coefficients 207 to a downmix function 203 to allow the downmix to be modified.

[0105] Without any intended limitation, analyzing the spatial audio signal (e.g., in step S510), determining each indication of signal power (e.g., in step S520), and generating the channel-based audio signal (e.g., in step S540) may be performed on a time-segment basis, for example, consistent with the above description of the STFT. This implies that a compressed representation is generated and output for each of a plurality of time segments, with a downmix audio signal and metadata (metadata block) for each time segment.

[0106] For each time segment k, the direction information 21 (e.g., embodied by the direction of arrival of one or more audio elements) is calculated using P direction vectors {dir k,p :1≦p≦P}. The direction vector p indicates the direction relative to the dominant object index p and in terms of unit vectors:

number

number

[0107] In some embodiments, each indication of signal power determined in step S520 takes the form of a signal power ratio, i.e., the indication of signal power associated with a given direction of arrival in a frequency subband relates to the ratio of signal power in the frequency subband for the given direction of arrival to the total signal power in the frequency subband.

[0108] Furthermore, in some embodiments, the indications of signal power are determined for each of a plurality of frequency subbands (i.e., on a subband-by-subband basis). In that case, they relate to the ratio of the signal power in a given frequency subband for a given direction of arrival to the total signal power in the given frequency subband for a given direction of arrival and for a given frequency subband. In particular, even if the indications of signal power may be determined per subband, the determination of the (dominant) direction of arrival may still be performed for the entire signal (i.e., based on all frequency subbands).

[0109] Furthermore, in some embodiments, analyzing the spatial audio signal (e.g., in step S510), determining the respective indications of signal power (e.g., in step S520), and generating the channel-based audio signals (e.g., in step S540) are performed based on a time-frequency representation of the spatial audio signal. For example, the above steps and other appropriate steps may be performed based on a discrete Fourier transform (e.g., STFT) of the spatial audio signal. For example, for each time segment (time block), the above steps may be performed on time-frequency bins (FFT bins) of the spatial audio signal, i.e., based on the Fourier coefficients of the spatial audio signal.

[0110] In view of the anomaly, for each time segment k and for each dominant object index p (1≦p≦P), the energy band ratio information 22 is a fraction value e for each band b (1≦b≦B) of the set of bands: k,p,b It can contain fractional values ​​e k,p,b teeth:

number

[0111] Fractional value e k,p,b is the energy distribution of multiple acoustic objects in the original spatial audio scene in the direction dir k,p The direction dir is combined to represent a single dominant acoustic component assigned to k,p In some embodiments, the energy of all acoustic objects in the scene may represent the portion of the energy in the spatial region around dir. k,p A larger weight is assigned to the direction θ closer to dir k,p The directions may be weighted using an angular difference weighting function w(θ) that represents a smaller weighting for directions θ farther from θ. The direction difference may be considered close for angular differences less than, for example, 10 degrees, and far for angular differences greater than, for example, 45 degrees. In alternative embodiments, the weighting function may be selected based on an alternative selection of near / far angular differences.

[0112] In general, the input spatial audio signal for which a compressed representation is generated may, for example, be a multi-channel audio signal or an object-based audio signal, in which case the method for generating a compressed representation of a spatial audio signal will further comprise the step of converting the object-based audio signal into a multi-channel audio signal before applying the scene analysis (e.g. before step S510).

[0113] 7, the input spatial audio signal may be a multi-channel audio signal, in which case the channel-based audio signal generated in step S540 is a downmix signal generated by applying a downmix operation to the multi-channel audio signal.

[0114] FIG. 8 schematically illustrates another example of details for generating a compressed representation of a spatial audio scene. The input spatial audio signal may in this case be an object-based audio signal including a plurality of audio objects and associated direction vectors. In this case, the method for generating a compressed representation of the spatial audio signal comprises generating a multi-channel audio signal as an intermediate representation or intermediate scene by panning the audio objects to a predefined set of audio channels, where each audio object is panned to a predefined set of audio channels according to its direction vector. Thus, FIG. 8 illustrates an alternative embodiment of a scene encoder 200 in which a spatial audio scene 10 is input to a converter 201, which generates an intermediate scene 11 (e.g., embodied by a multi-channel signal). The intermediate scene 11 may be generated according to equation (1). Here, the panning function is selected such that the dot product of the panning gain vectors Pan(θ1) and Pan(θ2) approximately represents the above-mentioned angle difference weighting function.

[0115] In some embodiments, the panning function used in converter 201 is the third-order Ambisonic spanning function shown in equation (3): (outside 4) TIFF2025135018000025.tif8170. The multi-channel audio signal may therefore be, for example, a higher-order Ambisonics signal.

[0116] The intermediate scene 11 is then input to a scene analysis 202. From the analysis of the intermediate scene 11, the scene analysis 202 determines the directions dir of the dominant acoustic objects in the spatial audio scene. k,p The determination of the dominant direction may be performed by estimating the energy in the set of directions, with the maximum estimated energy representing the dominant direction.

[0117] The energy band ratio information 22 for time segment k is a fractional value e for each band b derived from the energy in band b of the intermediate scene 11 in each direction relative to the total energy in band b of the intermediate scene 11 in time segment k. k,p,b may include:

[0118] The audio mixing stream 30 (e.g., channel-based audio signal) of the compact spatial audio scene (e.g., compact representation) in this case is a downmix signal generated by applying a downmix function 203 (downmix operation) to the spatial audio scene.

[0119] 10 shows an alternative arrangement of a scene encoder including a converter 201 that converts a spatial audio scene 10 into a scene-based intermediate format 11. The intermediate format 11 is input to a scene analysis 202 and to a downmix function 203. In some embodiments, the downmix function 203 may include a matrix mixer having coefficients adapted to convert the intermediate format 11 into an audio mixing stream 30. That is, the audio mixing stream 30 (e.g., a channel-based audio signal) of the compact spatial audio scene (e.g., a compact representation) in this case may be a downmix signal generated by applying the downmix function 203 (downmix operation) to the intermediate scene (e.g., a multi-channel audio signal).

[0120] In an alternative embodiment shown in Figure 11, the spatial encoder 200 can take input in the form of a scene-based input 11. The audio objects are represented according to a panning rule Pan(θ). In some embodiments, the panning function may be a higher-order Ambisonic spanning function. In one example embodiment, the panning function is a third-order Ambisonic spanning function.

[0121] In another alternative embodiment, depicted in Figure 9, the spatial audio scene 10 is converted by a converter 201 within a spatial encoder 200 to generate an intermediate scene 11 which is input to a downmix function 203. A scene analysis 202 is provided with input from the spatial audio scene 10.

[0122] FIG. 12 shows the direction information 21 and energy band ratio information 22 input to the demixing matrix calculator 301 which determines the demixing matrix (unmixing matrix) used by the demixer 302.

[0123] Details of processing a compact spatial audio scene (eg, a compressed representation of a spatial audio signal) to generate a reconstructed representation of the spatial audio signal are now described.

[0124] 13 is a flowchart of an example method 1300 for processing a compressed representation of a spatial audio signal to generate a reconstructed representation of the spatial audio signal. The compressed representation includes a channel-based audio signal (e.g., embodied by an audio mixing stream 30) having a predefined number of channels and metadata, where the metadata includes directional information (e.g., embodied by direction information 21) and energy information (e.g., embodied by energy band ratio information 22). The directional information includes an indication of the direction of arrival of one or more audio elements in an audio scene, and the energy information includes, for at least one frequency subband, an indication of each of the signal powers associated with the direction of arrival. The channel-based audio signal may be, for example, a first-order Ambisonics signal. The method 1300 includes steps S1310 to S1320, and optionally includes steps S1330 and S1340. It will be understood that these steps may be performed, for example, by the scene decoder 300 of FIG. 12.

[0125] In step S1310, audio signals of one or more audio elements are generated based on the channel-based audio signals, the direction information, and the energy information.

[0126] In step S1320, a residual audio signal substantially absent of one or more audio elements is generated based on the channel-based audio signal, the directional information, and the energy information, where the residual signal may be represented in the same audio format as the channel-based audio signal, e.g., have the same number of channels as the channel-based audio signal.

[0127] In optional step S1330, the audio signals of one or more audio elements are panned to a set of channels of an output audio format, where the output audio format may relate to an output representation such as, for example, HOA or any other suitable multi-channel format.

[0128] In optional step S1340, a reconstructed multi-channel audio signal in the output audio format is generated based on the one or more panned audio elements and the residual signal. Generating the reconstructed multi-channel audio signal may include upmixing the residual signal to a set of channels of the output audio format. Generating the reconstructed multi-channel audio signal may further include summing the one or more panned audio elements and the upmixed residual signal.

[0129] It will be appreciated that the above steps may be performed in any order, or in parallel with one another, so long as the order of the steps ensures that the necessary inputs for each step are available.

[0130] Consistent with the above description of the method of processing spatial audio signals to generate a compressed representation of the spatial audio signals, the indication of the signal power associated with a given direction of arrival may relate to the ratio of the signal power in a frequency subband for the given direction of arrival to the total signal power in the frequency subband.

[0131] Further, in some embodiments, the energy information may include an indication of signal power for each of a plurality of frequency subbands, where the indication of signal power may relate to a ratio of signal power in a given frequency subband for a given direction of arrival to total signal power in the given frequency subband for a given direction of arrival and for the given frequency subband.

[0132] Generating audio signals of one or more audio elements in step S1310 may include determining coefficients of an unmixing matrix M for mapping the channel-based audio signal to an intermediate representation including the residual audio signal and the audio signals of the one or more audio elements based on the directional information and the energy information. The intermediate representation may also be referred to as a separated or separable representation, or a hybrid representation.

[0133] Details of the above determination of the coefficients of the unmixing matrix M will now be described with reference to the flowchart of Figure 14. The method 1400 represented by this flowchart has steps S1410 to S1440.

[0134] In step S1410, for each of one or more audio elements, a panning vector Pan for panning the audio element to a channel of the channel-based audio signal is calculated. down (dir) is determined based on the direction of arrival dir of the audio element.

[0135] In step S1420, a mixing matrix E used to map the residual audio signal and the audio signals of the one or more audio elements to the channels of the channel-based audio signal is determined based on the determined panning vector.

[0136] In step S1430, a covariance matrix S of the intermediate representation is determined based on the energy information. The covariance matrix S is determined by the determined panning vector Pan down It may further be based on.

[0137] Finally, in step S1440, the coefficients of the unmixing matrix M are determined based on the mixing matrix E and the covariance matrix S.

[0138] It will be appreciated that the above steps may be performed in any order, or in parallel with one another, so long as the order of the steps ensures that the necessary inputs for each step are available.

[0139] Returning to FIG. 12, the demixing matrix calculator 301 generates the demixing matrix 60 (unmixing matrix) M according to a process that includes the following steps: k,b Calculate: 1. For each time segment k, the demixing matrix calculator 301 receives directional information dir k,p (1≦p≦P) and energy band ratio information e k,p,k (1≦p≦P and 1≦b≦B) are input, where P represents the number of dominant acoustic components and B represents the number of frequency bands. 2. For each band b, the demixing matrix Mk,b is: M=S×E * ×(E×S×E * ) -1 (20) Here, "x" indicates a matrix multiplication, and "*" indicates a conjugate transpose of a matrix. The calculation according to equation (20) may correspond to step S1440, for example.

[0140] A demixing matrix M may be determined for each of a plurality of time segments k and / or for each of a plurality of frequency subbands b, where matrices M and S will have an index k indicating the time segment and / or an index b indicating the frequency subband, and matrix E will have an index k indicating the time segment. For example, M k,b =S k,b ×E * k ×(E k ×S k,b ×E *k ) -1 (20a) is.

[0141] In general, determining the coefficients of the unmixing matrix M based on the mixing matrix E and the covariance matrix S may include determining a pseudo-inverse matrix based on the mixing matrix E and the covariance matrix S. An example of such a pseudo-inverse matrix is ​​given in equations (20) and (20a).

[0142] In equation (20), matrix E k (mixing matrix) is an N × N identity matrix (I N ) and P columns formed by panning functions applied in the direction of each of the P dominant acoustic components: E=(I N |Pan down (dir1)|···|Pan down (dir P )) (twenty one) In formula (21), I N is an NxN identity matrix, where N denotes the number of channels in the channel-based audio signal, and Pan down (dir p ) is the associated direction of arrival dir for panning the p-th audio element to the N channels of the channel-based audio signal. p where p=1, , P denotes each one of the one or more audio elements, and P denotes the total number of the one or more audio elements. The vertical bars in equation (21) indicate matrix augmentation operations. Thus, matrix E is an N×P matrix.

[0143] Furthermore, the matrix E may be determined for each of a plurality of time segments k. In this case, the matrix E and the direction of arrival dir p will have an index k indicating the time segment. For example: E k =(I N|Pan down (dir k,1 )|···|Pan down (dir k,P )) (21a) If the proposed method operates band-wise, the matrix E will be the same for all frequency subbands.

[0144] According to step S1420, the matrix E k is used to map the residual audio signal and the audio signals of one or more audio elements to the channels of the channel-based audio signal. As can be seen from equations (21) and (21a), the matrix E k is the panning vector Pan determined in step S1410. down Based on (dir).

[0145] In equation (20), matrix S is a (N+P)×(N+P) diagonal matrix. It can be considered as the covariance matrix of the intermediate representation. Its coefficients can be calculated based on the energy information according to step S1430. The first N diagonal elements are, for 1≦n≦N:

number

[0146] A covariance matrix S may be determined for each of a plurality of time segments k and / or for each of a plurality of frequency subbands b. In that case, the covariance matrix S and the signal power e p will have index k indicating a time segment and / or index b indicating a frequency subband. The first N diagonal elements are:

number

[0147] In a preferred embodiment, the demixing matrix M k,b are applied by the demixer 302 to generate the separated spatial audio stream 70 (as an example of an intermediate representation). According to the above implementation of step S1310, the first N channels are the residual stream 80 and the remaining P channels represent the dominant acoustic components.

[0148] N+P channel separated spatial streams 70 Y k (f) P-channel dominant object signal 90 (as an example of the audio signal of one or more audio elements generated in step S1310) O k (f), and an N-channel residual stream 80 (as an example of the residual audio signal generated in step S1320) R k (f) is:

number

[0149] In addition to the above, in some embodiments, the number of dominant acoustic components P can be adapted to take different values ​​for each time segment, so that P k may depend on the time segment k. For example, the scene analysis 202 of the scene encoder 200 may calculate P k In general, the number of dominant acoustic components P may depend on time. k The choice of ) may involve a trade-off between the data rate of the metadata and the quality of the reconstructed audio scene.

[0150] Returning to Figure 12, the spatial decoder 300 generates an M-channel reconstructed audio scene 50. The M-channel stream is then fed to the output panner (outside 5) TIFF2025135018000029.tif8170. This may be done in accordance with step S1340 above. Examples of output panners include stereo panning functions, vector-based amplitude panning functions known in the art, and higher-order Ambisonic spanning functions known in the art.

[0151] For example, the object panner 91 in Figure 12:

number

[0152] 15 is a flowchart providing an alternative formulation of a method 1500 for decoding a compact spatial audio scene to generate a reconstructed audio scene. The method 1500 includes steps S1510 to S1580.

[0153] In step S1510, a compact spatial audio scene is received and the encoded downmix stream and the encoded metadata stream are extracted.

[0154] In step S1520, the encoded downmix stream is decoded to form a downmix stream.

[0155] In step S1530, the encoded metadata stream is decoded to form directional information and energy ratio information.

[0156] In step S1540, a demixing matrix for each band is formed from the direction information and the energy ratio information.

[0157] In step S1550, the downmix stream is processed according to a demixing matrix to form separated streams.

[0158] In step S1560, the object signal is extracted from the separated stream and panned according to the direction information and the desired output format to generate a panned object signal.

[0159] In step S1570, the residual signal is extracted from the separated stream and processed to generate a decoded residual signal according to the desired output format.

[0160] Finally, in step S1580, the panned object signal and the decoded residual signal are combined to form a reconstructed audio scene.

[0161] It will be appreciated that the above steps may be performed in any order, or in parallel with one another, so long as the order of the steps ensures that the necessary inputs for each step are available.

[0162] Methods for processing a spatial audio signal to generate a compressed representation of the spatial audio signal, and methods for processing the compressed representation of the spatial audio signal to generate a reconstructed representation of the spatial audio signal, have been described above. Furthermore, the present disclosure also relates to apparatuses for performing these methods. An example of such an apparatus 1600 is schematically represented in FIG. 16. The apparatus 1600 may include a processor 1610 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), one or more application specific integrated circuits (ASICs), one or more radio frequency integrated circuits (RFICs), or any combination thereof) and a memory 1620 coupled to the processor 1610. The processor may be configured to perform some or all of the steps of the methods described throughout this disclosure. When the apparatus 1600 operates as an encoder (e.g., a scene encoder), it may receive, for example, a spatial audio signal (i.e., a spatial audio scene) as input 1630. The apparatus 1600 may then generate a compressed representation of the spatial audio signal as output 1640. When the device 1600 operates as a decoder (e.g., a scene decoder), it may receive the compressed representation as input 1630. The device may then produce a reconstructed audio scene as output 1640.

[0163] Device 1600 may be a server computer, a client computer, a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a smartphone, a web appliance, a network router, a switch, or a bridge, or any machine capable of executing instructions that specify operations to be performed by the device. Additionally, although only one device 1600 is depicted in Figure 16, it should be understood that the present disclosure relates to any collection of devices that individually or collectively execute instructions to perform any one or more of the methodologies discussed herein.

[0164] The present disclosure further relates to a program (eg, a computer program) having instructions that, when executed by a processor, cause the processor to perform some or all of the steps of the methods described herein.

[0165] Still further, the present disclosure relates to a computer-readable (or machine-readable) storage medium having stored thereon the above-described program, where the term "computer-readable storage medium" includes, for example but not limited to, data repositories in the form of solid-state memory, optical media, and magnetic media.

[0166] [Additional configuration considerations] Unless specifically stated otherwise, as will be apparent from the discussion that follows, throughout this disclosure, discussions utilizing terms such as "processing," "computing," "calculating," "determining," "analyzing," and the like will be understood to refer to the operations and / or processing of a computer or computing system or similar electronic computing device that manipulates and / or converts data represented as physical quantities, such as electrons, into other data similarly represented as physical quantities.

[0167] Similarly, the term "processor" may refer to any device or part of a device that processes electronic data, e.g., from registers and / or memory, and converts the electronic data into other electronic data that may be stored, e.g., in registers and / or memory. A "computer" or "computing machine" or "computing platform" may include one or more processors.

[0168] The methodologies described herein, in one exemplary embodiment, are executable by one or more processors that accept computer-readable (also called machine-readable) code, including a set of instructions that, when executed by one or more of the processors, perform at least one of the methods described herein. Any processor capable of executing a set of instructions (sequential or otherwise) that specify operations to be performed is included. Thus, one example is a typical processing system including one or more processors. Each processor may include one or more CPUs, graphics processing units, and programmable DSP units. The processing system may further include a memory subsystem including main RAM and / or static RAM and / or ROM. A bus subsystem may also be included for communication between components. The processing system may also be a distributed processing system with processors coupled by a network. If the processing system requires a display, such a display may be included, for example, a liquid crystal display (LCD) or a cathode ray tube (CRT) display. If manual data entry is required, the processing system may also include one or more input devices, such as an alphanumeric input unit such as a keyboard, a pointing control device such as a mouse, etc. The processing system may also include a storage system, such as a disk drive unit. The processing system, in some configurations, may also include an audio output device and a network interface device. Thus, the memory subsystem includes a computer-readable carrier medium that carries computer-readable code (e.g., software) that includes a set of instructions that, when executed by one or more processors, cause the execution of one or more of the methods described herein. It should be noted that, when a method includes several elements, e.g., several steps, no order of such elements is implied unless specifically stated. The software may reside on a hard disk, or may reside completely or at least partially within RAM and / or within the processor during its execution by the computer system.Thus, the memory and processor also constitute a computer-readable carrier medium carrying computer-readable code. Furthermore, the computer-readable carrier medium may form, or be included in, a computer program product.

[0169] In alternative exemplary embodiments, one or more processors may operate as standalone devices or may be networked to other processors in a networked deployment, e.g., one or more processors may operate as server or user machines in a server-user network environment, or as peer machines in a peer-to-peer or distributed network environment. One or more processors may form a personal computer (PC), tablet PC, personal digital assistant (PDA), mobile phone, web appliance, network router, switch, or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify operations to be performed by the machine.

[0170] It should be noted that the term "machine" is intended to include any collection of machines that individually or collectively execute a set (or sets) of instructions to perform any one or more of the methodologies discussed herein.

[0171] Accordingly, one exemplary embodiment of each method described herein takes the form of a computer-readable carrier medium carrying a set of instructions, e.g., a computer program, executed on one or more processors, e.g., one or more processors that are part of a web server configuration. Accordingly, as will be appreciated by those skilled in the art, exemplary embodiments of the present disclosure may be embodied as a method, an apparatus such as a special purpose device, an apparatus such as a data processing system, or a computer-readable carrier medium, e.g., a computer program product. The computer-readable carrier medium carries computer-readable code, which, when executed on one or more processors, causes the one or more processors to implement the method. Accordingly, aspects of the present disclosure may take the form of a method, an entirely hardware exemplary embodiment, an entirely software exemplary embodiment, or an exemplary embodiment combining software and hardware aspects. Furthermore, the present disclosure may take the form of a carrier medium (e.g., a computer program product on a computer-readable storage medium) carrying computer-readable program code embodied in the medium.

[0172] Software may also be transmitted or received over a network via a network interface device. While the carrier medium is a single medium in the exemplary embodiment, the term "carrier medium" should be interpreted to include a single medium or multiple media (e.g., a centralized or distributed database and / or associated caches and servers) that store one or more sets of instructions. The term "carrier medium" should also be interpreted to include any medium that can store, encode, or carry a set of instructions for execution by one or more processors, causing the one or more processors to perform any one or more of the methodologies disclosed herein. Carrier media can take many forms, including, but not limited to, non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical disks, magnetic disks, and magneto-optical disks. Volatile media include dynamic memory, such as main memory. Transmission media include coaxial cable, copper wire, and fiber optics, including the wires that comprise a bus subsystem. Transmission media can also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications. For example, the term "carrier medium" should be interpreted to include, but not be limited to, solid-state memory, computer products embodied in optical and magnetic media, media having propagated signals detectable by at least one processor or one or more processors and representing a set of instructions that, when executed, implement a method, and transmission media within a network having propagated signals detectable by at least one processor of one or more processors and representing a set of instructions.

[0173] It will be understood that the steps of the methods discussed are, in one exemplary embodiment, performed by a suitable processor (or processors) of a processing (e.g., computer) system executing instructions (computer-readable code) stored in storage. It will also be understood that the present disclosure is not limited to any particular implementation or programming technique, and that the present disclosure may be implemented by any suitable technique for implementing the functions described herein. The present disclosure is not limited to any particular programming language or operating system.

[0174] References throughout this disclosure to "one exemplary embodiment," "some exemplary embodiments," or "exemplary embodiments" mean that a particular feature, structure, or characteristic described in connection with an exemplary embodiment is included in at least one exemplary embodiment of this disclosure. Thus, the appearances of the phrases "in one exemplary embodiment," "some exemplary embodiments," or "in exemplary embodiments" in various places throughout this disclosure do not necessarily all refer to the same exemplary embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner, as would be apparent to one of ordinary skill in the art from this disclosure, in one or more exemplary embodiments.

[0175] As used herein, the use of ordinal adjectives "first," "second," "third," etc. to describe a common object merely indicates that different instances of a similar object are being referenced, unless otherwise specified, and is not intended as to imply that the objects so described need be in a particular order, temporally, spatially, ranked, or otherwise.

[0176] In the following claims and in the description of this specification, any one of the terms "comprising," "comprised of," or "which comprises" is an open term meaning the inclusion of at least the following elements / features but not the exclusion of others. Thus, when used in a claim, the term "comprising" should not be interpreted as limiting the means or elements or steps listed thereafter. For example, the scope of the expression "a device having A and B" should not be limited to "a device having only elements A and B." As used herein, any one of the terms "including," "which includes," or "that includes" also means the inclusion of at least the following elements / features but not the exclusion of others. Thus, "including" is synonymous with and means "comprising."

[0177] In the foregoing description of exemplary embodiments of the present disclosure, it should be understood that various features of the present disclosure are sometimes grouped together in a single exemplary embodiment, figure, or description thereof for the purpose of streamlining the disclosure and aiding in understanding one or more of the various inventive aspects. This method of disclosure, however, should not be interpreted as reflecting an intention that the claims require more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in fewer than all features of a single foregoing disclosed exemplary embodiment. Accordingly, the claims following the description are expressly incorporated herein, with each claim standing on its own as a separate exemplary embodiment of the present disclosure.

[0178] Furthermore, while some exemplary embodiments described herein include some features but not others included in other exemplary embodiments, combinations of features from different exemplary embodiments are intended to be within the scope of this disclosure and will form separate exemplary embodiments, as would be understood by one of ordinary skill in the art. For example, in the following claims, any of the claimed exemplary embodiments may be used in any combination.

[0179] In the description provided herein, numerous specific details are set forth. However, it will be understood that exemplary embodiments of the present disclosure may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail in order not to obscure an understanding of this description.

[0180] Thus, while what is believed to be the best mode of the disclosure has been described, those skilled in the art will recognize that other and further modifications may be made without departing from the spirit of the disclosure, and it is intended to claim all such changes and modifications as fall within the scope of the disclosure. For example, the above formulas are merely representative of procedures that may be used. Functions may be added or deleted from the block diagrams, and operations may be interchanged between functional blocks. Steps may also be added or deleted to methods described within the scope of the disclosure.

[0181] Further aspects, embodiments, and examples of the present disclosure will become apparent from the enumerated example embodiments (EEE) below.

[0182] EEE1 relates to a method for representing a spatial audio scene as a compact spatial audio scene comprising an audio mixing stream and a directional metadata stream, wherein the audio mixing stream consists of one or more audio elements, the directional metadata stream consisting of a time series of directional metadata blocks, each of the directional metadata blocks being associated with a corresponding time segment in the audio signal, the spatial audio scene comprising one or more directional sound elements respectively associated with a respective direction of arrival, each of the directional metadata blocks comprising: (a) directional information indicating the direction of arrival for each of the directional sound elements; and (b) energy band ratio information, for each directional sound element and for each of a set of two or more sub-bands, indicating the energy in each of the directional sound elements relative to the energy in the corresponding time segment in the audio signal.

[0183] EEE2 relates to a method according to EEE1, wherein (a) the energy band ratio information indicates characteristics of the spatial audio scene in each of a plurality of the sub-bands, and (b) for at least one direction of arrival, data included in the direction information indicates characteristics of the spatial audio scene in two or more clusters of the sub-bands.

[0184] EEE3 relates to a method for processing a compact spatial audio scene comprising an audio mixing stream and a directional metadata stream to generate a separated spatial audio stream and a residual stream comprising a set of one or more audio object signals, wherein the audio mixing stream comprises one or more audio signals, the directional metadata stream comprises a time sequence of directional metadata blocks, each directional metadata block associated with a corresponding time segment in the audio signal, for each of a plurality of sub-bands, the method comprising: (a) determining coefficients of a demixing matrix from directional information and energy band ratio information comprised in the directional metadata stream; and (b) mixing the audio mixing stream using the demixing matrix to generate the separated spatial audio stream.

[0185] EEE4 relates to the method according to EEE3, wherein each of the directional metadata blocks includes (a) directional information indicating a direction of arrival for each of the directional acoustic elements, and (b) energy band ratio information indicating, for each of the directional acoustic elements and for each of a set of two or more sub-bands, the energy in each of the directional acoustic elements relative to the energy in the corresponding time segment in the audio signal.

[0186] EEE5 relates to the method according to EEE3, wherein (a) for each of the blocks of the directional metadata, directional information and energy band ratio information are used to form a matrix S representing an approximate covariance of the separated spatial audio streams, (a) the energy band ratio information is used to form E representing a remixing matrix defining a transformation of the separated spatial audio streams into the audio mixing stream, and (b) the demixing matrix E is calculated as follows: U=S×E * ×(E×S×E * ) -1 is calculated according to

[0187] EEE6 relates to the method described in EEE6, wherein the matrix S is a diagonal matrix.

[0188] EEE7 relates to a method according to EEE3, wherein (a) the residual stream is processed to generate a reconstructed residual stream; (b) each of the audio object signals is processed to generate a corresponding reconstructed object stream; and (c) the reconstructed residual stream and each of the reconstructed object streams are combined to form a reconstructed audio signal, the reconstructed audio signal comprising directional sound elements according to the compact spatial audio scene.

[0189] EEE8 relates to the method according to EEE7, wherein the reconstructed audio signal comprises two signals for presentation to a listener by transducers at or near each ear to provide a binaural experience of a spatial audio scene including directional acoustic elements according to the compact spatial audio scene.

[0190] EEE9 relates to a method as described in EEE7, wherein the reconstructed audio signal comprises a plurality of signals representing a spatial audio scene in the form of spherical-harmonic panning functions.

[0191] EEE10 relates to a method for processing a spatial audio scene to generate a compact spatial audio scene comprising an audio mixing stream and a directional metadata stream, the spatial audio scene comprising one or more directional sound elements each associated with a respective direction of arrival, the directional metadata stream consisting of a time series of directional metadata blocks, each directional metadata block relating to a corresponding time segment in an audio signal, the method comprising: (a) means for determining a direction of arrival for one or more of the directional sound elements from an analysis of the spatial audio scene; (b) means for determining what portion of the total energy in the spatial scene is contributed by the energy in each of the directional sound elements; and (c) means for processing the spatial audio scene to generate the audio mixing stream.

Claims

1. 1. A method of processing a spatial audio signal to generate a compressed representation of the spatial audio signal, comprising: analyzing the spatial audio signal to determine directions of arrival of one or more audio elements in an audio scene represented by the spatial audio signal; determining, for at least one frequency subband of the spatial audio signal, respective indications of signal power associated with the determined directions of arrival; generating metadata including directional information and energy information, the directional information including an indication of the determined directions of arrival of the one or more audio elements, and the energy information including an indication of each of the signal powers associated with the determined directions of arrival; generating a channel-based audio signal based on the spatial audio signal, the channel-based audio signal having a predefined number of channels; outputting the channel-based audio signal and the metadata as the compressed representation of the spatial audio signal; and the spatial audio signal is an object-based audio signal comprising a plurality of audio objects and associated direction vectors; The method further comprises generating a multi-channel audio signal by panning the plurality of audio objects onto a predefined set of audio channels, each audio object being panned onto the predefined set of audio channels according to its direction vector; the channel-based audio signal is a downmix signal generated by applying a downmix operation to the multi-channel audio signal; method.

2. analyzing the spatial audio signal based on a plurality of frequency subbands of the spatial audio signal; The method of claim 1.

3. analyzing the spatial audio signal includes applying a scene analysis to the spatial audio signal.

3. The method according to claim 1 or 2.

4. the indication of signal power associated with a given direction of arrival relates to a ratio of signal power in the frequency subband for the given direction of arrival to total signal power in the frequency subband; 4. The method according to any one of claims 1 to 3.

5. the indication of signal power is determined for each of a plurality of frequency subbands, and for a given direction of arrival and a given frequency subband, relates to a ratio of signal power in the given frequency subband for the given direction of arrival to a total signal power in the given frequency subband; 5. The method according to any one of claims 1 to 4.

6. analyzing the spatial audio signal, determining each indication of signal power, and generating the channel-based audio signal is performed for each time segment.

6. The method according to any one of claims 1 to 5.

7. analyzing the spatial audio signal, determining the respective indications of signal power, and generating the channel-based audio signal are performed based on a time-frequency representation of the spatial audio signal.

7. The method according to any one of claims 1 to 6.

8. the spatial audio signal is a multi-channel audio signal; the channel-based audio signal is a downmix signal generated by applying a downmix operation to the multi-channel audio signal; 8. The method according to any one of claims 1 to 7.

9. the channel-based audio signal is a first-order Ambisonics signal; 9. The method according to any one of claims 1 to 8.

10. A program comprising instructions which, when executed by a processor, cause the processor to carry out all the steps of the method according to any one of claims 1 to 9.

11. A computer-readable storage medium storing the program according to claim 10.

12. a processor and a memory coupled to the processor; The processor is configured to perform all the steps of the method according to any one of claims 1 to 9. Device.