Incorporation of spatial audio parameters

By combining spatial audio parameters, especially converting spherical direction vectors into Cartesian vectors and performing weighted normalization processing, the problem of high bit rate of spatial audio parameter encoding in the existing technology is solved, and more efficient encoding and transmission are achieved.

CN114846541BActive Publication Date: 2025-09-05NOKIA TECHNOLOGIES OY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080089375.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-23
Filing Date
2020-11-13
Publication Date
2025-09-05
Estimated Expiration
2040-11-13

AI Technical Summary

Technical Problem

Existing technologies have difficulty effectively compressing and encoding spatial metadata of multiple input types, such as speaker signals, audio object signals, and ambisonic signals, when encoding spatial audio parameters, especially signals captured by microphone arrays, resulting in excessively high bit requirements.

Method used

The number of bits required for encoding is reduced by incorporating spatial audio parameters, including converting spherical direction vectors into Cartesian vectors, and by weighting and normalizing the energy ratio, extension coherence, and surround coherence parameters of multiple samples.

Benefits of technology

It effectively reduces the bit rate requirement for each TF tile, improves the encoding efficiency of spatial audio parameters, is suitable for encoders of various input types, and supports more efficient storage and transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114846541B_ABST
    Figure CN114846541B_ABST
Patent Text Reader

Abstract

In particular, a device for spatial audio encoding is disclosed, which includes: a component for determining at least two spatial audio parameters of a type of spatial audio parameters of one or more audio signals, wherein a first spatial audio parameter of the type is associated with a first group of samples in a domain of the one or more audio signals, and a second spatial audio parameter of the type is associated with a second group of samples in the domain of the one or more audio signals; and a component for merging the first spatial audio parameter of the type and the second spatial audio parameter of the type into a merged spatial audio parameter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to an apparatus and method for encoding parameters related to a sound field, but is not limited to encoding parameters related to the time-frequency domain of an audio encoder and decoder. Background Art

[0002] Parametric spatial audio processing is an area of ​​audio signal processing that uses a set of parameters to describe the spatial aspects of sound. For example, when performing parametric spatial audio capture from a microphone array, estimating a set of parameters from the microphone array signal, such as the direction of the sound in a frequency band and the ratio of the directional to non-directional portion of the captured sound in the frequency band, is a typical and effective choice. It is well known that these parameters describe the perceived spatial characteristics of the captured sound at the position of the microphone array. These parameters can be used accordingly in the synthesis of spatial sound for use with binaural headphones, speakers, or other formats such as Ambisonics.

[0003] Therefore, the directionality and direct-to-total energy ratio in a frequency band are particularly effective parameterizations for spatial audio capture.

[0004] A parameter set including directional parameters in a frequency band and energy ratio parameters in a frequency band (indicating the directionality of the sound) can also be used as spatial metadata for an audio codec (which can also include other parameters such as surround coherence, spread coherence, number of directions, distance, etc.). For example, these parameters can be estimated from an audio signal captured by a microphone array, and, for example, a stereo or mono signal can be generated from the microphone array signal to be transmitted together with the spatial metadata. The stereo signal can, for example, be encoded with an AAC encoder, while the mono signal can be encoded with an EVS encoder. The decoder can decode the audio signal into a PCM signal and process the sound in the frequency band (using the spatial metadata) to obtain a spatial output, for example, a binaural output.

[0005] The aforementioned solution is particularly suitable for encoding spatial sound captured from a microphone array (e.g., in a mobile phone, a VR camera, or a standalone microphone array). However, it may be desirable for such an encoder to have other input types besides the signal captured by the microphone array, such as loudspeaker signals, audio object signals, or ambisonic signals.

[0006] The analysis of first-order ambisonics (FOA) input for spatial metadata extraction has been well documented in the scientific literature related to Directional Audio Coding (DirAC) and Harmonic Plane Wave Expansion (Harpex). This is due to the existence of microphone arrays that directly provide FOA signals (more precisely: its variant, B-format signals), and therefore the analysis of such input has become a research focus in this field. In addition, the analysis of higher-order ambisonics (HOA) input for multi-directional spatial metadata extraction has also been documented in the scientific literature related to Higher-Order Directional Audio Coding (HO DirAC).

[0007] Another input for the encoder may also be a multi-channel speaker input, such as a 5.1 or 7.1 channel surround sound input and an audio object.

[0008] However, regarding the component of spatial metadata, compression and encoding of the spatial audio parameters is quite important in order to minimize the total number of bits required to represent the spatial audio parameters. Summary of the Invention

[0009] According to a first aspect, a device for spatial audio encoding is provided, comprising: a component for determining at least two spatial audio parameters of a type of spatial audio parameters of one or more audio signals, wherein a first spatial audio parameter of the type is associated with a first group of samples in a domain of the one or more audio signals, and a second spatial audio parameter of the type is associated with a second group of samples in the domain of the one or more audio signals; and a component for merging the first spatial audio parameter of the type and the second spatial audio parameter of the type into merged spatial audio parameters.

[0010] The apparatus may further comprise means for determining whether the merged spatial audio parameter is encoded for storage and / or transmission, or whether at least two spatial audio parameters of the type are encoded for storage and / or transmission.

[0011] The apparatus may further include: a component for determining a metric for the first group of samples and the second group of samples; a component for comparing the metric with a threshold value, wherein the apparatus further includes a component for determining whether the merged spatial audio parameter is encoded for storage and / or transmission, or whether at least two of the spatial audio parameters of the type are encoded for storage and / or transmission, includes: a component for determining that when the metric is above the threshold value, at least two of the spatial audio parameters of the type are encoded for storage and / or transmission; and a component for determining that when the metric is below or equal to the threshold value, the merged spatial audio parameter is encoded for storage and / or transmission.

[0012] Alternatively, the device may further include: a component for determining a metric for a first group of samples and a second group of samples; a component for determining at least two other spatial audio parameters of a type of spatial audio parameters of one or more audio signals, wherein another first spatial audio parameter of the type is associated with another first group of samples in the domain of the one or more audio signals, and another second spatial audio parameter of the type is associated with another second group of samples in the domain of the one or more audio signals; a component for merging another first spatial audio parameter of the type of spatial audio parameters and another second spatial audio parameter of the type of spatial audio parameters into another merged spatial audio parameter; a component for determining a metric for another first group of samples and another second group of samples; and a component for determining that when the metric for another first group of samples and another second group of samples is higher than the metric for the first group of samples and the second group of samples, another first spatial audio parameter of the type of spatial audio parameters and another second spatial audio parameter of the type of spatial audio parameters are encoded for storage and / or transmission and the merged spatial audio parameters are encoded for storage and / or transmission.

[0013] The apparatus may further comprise means for determining an energy of a first set of samples of the one or more audio signals and an energy of a second set of samples of the one or more audio signals, wherein the value of the combined spatial audio parameter depends on the energy of the first set of samples of the one or more audio signals and the energy of the second set of samples of the one or more audio signals.

[0014] The spatial audio parameters of this type may include spherical direction vectors, and wherein the merged spatial audio parameters include merged spherical direction vectors, and wherein the component for merging the first spatial audio parameter of this type and the second spatial audio parameter of this type into the merged spatial audio parameter may include: a component for converting the first spherical direction vector into a first Cartesian vector and converting the second spherical direction vector into a second Cartesian vector, wherein the first Cartesian direction vector and the second Cartesian direction vector each include an x-axis component, a y-axis component and a z-axis component, and wherein, for each component one by one, the device includes: a component for converting the first Cartesian vector into a first Cartesian vector and a second Cartesian direction vector into a second Cartesian vector means for weighting the energies of a first set of samples of the one or more audio signals and a direct-to-total energy ratio calculated for the first set of samples of the one or more audio signals by an amount; means for weighting the components of a second Cartesian vector by the energies of a second set of samples of the one or more audio signals and a direct-to-total energy ratio calculated for the second set of samples of the one or more audio signals; and means for summing the weighted components of the first Cartesian vector and the corresponding weighted components of the second Cartesian vector to provide a corresponding combined Cartesian component vector; and means for converting the combined Cartesian x-axis component value, the combined Cartesian y-axis component value, and the combined Cartesian z-axis component value into a combined spherical direction vector.

[0015] The apparatus may further include means for combining the direct-to-total energy ratio of a first group of samples of the one or more audio signals and the direct-to-total energy ratio of a second group of samples of the one or more audio signals into a combined direct-to-total energy ratio by determining a length of the combined Cartesian vector, and normalizing the length of the combined Cartesian vector by the sum of an energy of the first group of samples of the one or more audio signals and an energy of the second group of samples of the one or more audio signals.

[0016] The apparatus may further include: means for determining a first extended coherence parameter associated with a first set of samples in a domain of one or more audio signals and a second extended coherence parameter associated with a second set of samples in the domain of the one or more audio signals; and means for merging the first extended coherence parameter and the second extended coherence parameter into a merged extended coherence parameter.

[0017] The component for combining the first extended coherence parameter and the second extended coherence parameter into a combined extended coherence parameter may include: a component for weighting the first extended coherence value by the energy of a first group of samples of one or more audio signals; a component for weighting the second extended coherence value by the energy of a second group of samples of the one or more audio signals; a component for summing the weighted first extended coherence value and the weighted second extended coherence value to give a combined extended coherence value; and a component for normalizing the combined extended coherence value by the sum of the energy of the first group of samples of the one or more audio signals and the energy of the second group of samples of the one or more audio signals.

[0018] The apparatus may further include: means for determining a first surround coherence parameter associated with a first set of samples in a domain of one or more audio signals and a second surround coherence parameter associated with a second set of samples in the domain of the one or more audio signals; and means for merging the first surround coherence parameter and the second surround coherence parameter into a merged surround coherence parameter.

[0019] The component for combining the first surround coherence parameter and the second surround coherence parameter into a combined surround coherence parameter may include: a component for weighting the first surround coherence value by the energy of a first group of samples of one or more audio signals; a component for weighting the second surround coherence value by the energy of a second group of samples of the one or more audio signals; a component for summing the weighted first surround coherence value and the weighted second surround coherence value to give a combined extended coherence value; and a component for normalizing the combined surround coherence value by the sum of the energy of the first group of samples of the one or more audio signals and the energy of the second group of samples of the one or more audio signals.

[0020] The means for determining the metric may include means for determining the sum of the length of the first Cartesian vector and the length of the second Cartesian vector; and means for determining the difference between the length of the combined Cartesian vector and the sum.

[0021] The first set of samples may be a first subframe in the time domain, and the second set of samples may be a second subframe in the time domain.

[0022] Alternatively, the first set of samples may be a first sub-band in the frequency domain and the second set of samples may be a second sub-band in the frequency domain.

[0023] According to a second aspect, there is a method for spatial audio encoding, comprising: determining at least two spatial audio parameters of a type of spatial audio parameters of one or more audio signals, wherein a first spatial audio parameter of the type is associated with a first group of samples in a domain of the one or more audio signals, and a second spatial audio parameter of the type is associated with a second group of samples in the domain of the one or more audio signals; and merging the first spatial audio parameter of the type and the second spatial audio parameter of the type into merged spatial audio parameters.

[0024] The method may further comprise determining whether the merged spatial audio parameter is encoded for storage and / or transmission, or whether at least two spatial audio parameters of the type are encoded for storage and / or transmission.

[0025] The method may further include: determining a metric for the first group of samples and the second group of samples; comparing the metric with a threshold, wherein the device further includes a component for determining whether the merged spatial audio parameter is encoded for storage and / or transmission, or whether at least two spatial audio parameters of the type are encoded for storage and / or transmission, includes: determining that when the metric is above the threshold, at least two spatial audio parameters of the type are encoded for storage and / or transmission; and determining that when the metric is below or equal to the threshold, the merged spatial audio parameter is encoded for storage and / or transmission.

[0026] Alternatively, the method may further include: determining a metric for a first group of samples and a second group of samples; determining at least two other spatial audio parameters of a type of spatial audio parameters of one or more audio signals, wherein another first spatial audio parameter of the type is associated with another first group of samples in the domain of the one or more audio signals, and another second spatial audio parameter of the type is associated with another second group of samples in the domain of the one or more audio signals; merging another first spatial audio parameter of the type of spatial audio parameters and another second spatial audio parameter of the type of spatial audio parameters into another merged spatial audio parameter; determining a metric for another first group of samples and another second group of samples; and determining that when the metric for another first group of samples and another second group of samples is higher than the metric for the first group of samples and the second group of samples, the another first spatial audio parameter of the type of spatial audio parameters and the another second spatial audio parameter of the type of spatial audio parameters are encoded for storage and / or transmission and the merged spatial audio parameter is encoded for storage and / or transmission.

[0027] The method may further include determining an energy of a first group of samples of the one or more audio signals and an energy of a second group of samples of the one or more audio signals, wherein the value of the combined spatial audio parameter depends on the energy of the first group of samples of the one or more audio signals and the energy of the second group of samples of the one or more audio signals.

[0028] The spatial audio parameters of the type may include spherical direction vectors, and wherein the merged spatial audio parameters include merged spherical direction vectors, and wherein merging the first spatial audio parameter of the type and the second spatial audio parameter of the type into the merged spatial audio parameter may include: converting the first spherical direction vector into a first Cartesian vector, converting the second spherical direction vector into a second Cartesian vector, wherein the first Cartesian direction vector and the second Cartesian direction vector each include an x-axis component, a y-axis component, and a z-axis component, wherein, for each component one by one, the apparatus includes: converting the first Cartesian direction vector into a first Cartesian vector, converting the second spherical direction vector into a second Cartesian vector, weighting the components of the vector by the energy of a first set of samples of the one or more audio signals and a direct-to-total energy ratio calculated for the first set of samples of the one or more audio signals; weighting the components of the second Cartesian vector by the energy of a second set of samples of the one or more audio signals and a direct-to-total energy ratio calculated for the second set of samples of the one or more audio signals; and summing the weighted components of the first Cartesian vector and the corresponding weighted components of the second Cartesian vector to provide a corresponding combined Cartesian component vector; and converting the combined Cartesian x-axis component value, the combined Cartesian y-axis component value, and the combined Cartesian z-axis component value into a combined spherical direction vector.

[0029] The method may further include combining the direct-to-total energy ratio of a first group of samples of the one or more audio signals and the direct-to-total energy ratio of a second group of samples of the one or more audio signals into a combined direct-to-total energy ratio by determining a length of the combined Cartesian vector, and normalizing the length of the combined Cartesian vector by the sum of the energy of the first group of samples of the one or more audio signals and the energy of the second group of samples of the one or more audio signals.

[0030] The method may further include determining a first extended coherence parameter associated with a first set of samples in a domain of one or more audio signals and a second extended coherence parameter associated with a second set of samples in the domain of the one or more audio signals; and merging the first extended coherence parameter and the second extended coherence parameter into a merged extended coherence parameter.

[0031] Combining the first extended coherence parameter and the second extended coherence parameter into a merged extended coherence parameter may include: weighting the first extended coherence value by the energy of a first group of samples of one or more audio signals; weighting the second extended coherence value by the energy of a second group of samples of the one or more audio signals; summing the weighted first extended coherence value and the weighted second extended coherence value to give a merged extended coherence value; and normalizing the merged extended coherence value by the sum of the energy of the first group of samples of the one or more audio signals and the energy of the second group of samples of the one or more audio signals.

[0032] The method may further include: determining a first surround coherence parameter associated with a first set of samples in a domain of one or more audio signals and a second surround coherence parameter associated with a second set of samples in the domain of the one or more audio signals; and merging the first surround coherence parameter and the second surround coherence parameter into a merged surround coherence parameter.

[0033] Combining the first surround coherence parameter and the second surround coherence parameter into a combined surround coherence parameter may include: weighting the first surround coherence value by the energy of a first group of samples of one or more audio signals; weighting the second surround coherence value by the energy of a second group of samples of the one or more audio signals; summing the weighted first surround coherence value and the weighted second surround coherence value to give a combined extended coherence value; and normalizing the combined surround coherence value by the sum of the energy of the first group of samples of the one or more audio signals and the energy of the second group of samples of the one or more audio signals.

[0034] Determining the metric may include determining a sum of a length of the first Cartesian vector and a length of the second Cartesian vector; and determining a difference between the length of the combined Cartesian vector and the sum.

[0035] The first set of samples may be a first subframe in the time domain, and the second set of samples may be a second subframe in the time domain.

[0036] The first set of samples may be a first sub-band in the frequency domain, and the second set of samples may be a second sub-band in the frequency domain.

[0037] According to a third aspect, there is an apparatus for spatial audio encoding, comprising at least one processor and at least one memory comprising computer program code, the at least one memory and the computer program code being configured to, together with the at least one processor, cause the apparatus to at least: determine at least two spatial audio parameters of a type of spatial audio parameter of one or more audio signals, wherein a first spatial audio parameter of the type is associated with a first set of samples in a domain of the one or more audio signals, and a second spatial audio parameter of the type is associated with a second set of samples in the domain of the one or more audio signals; and merge the first spatial audio parameter of the type and the second spatial audio parameter of the type into merged spatial audio parameters.

[0038] A computer program product stored on a medium may cause an apparatus to perform the method described herein.

[0039] An electronic device may include an apparatus as described herein.

[0040] A chipset may include the apparatus as described herein.

[0041] The embodiments of the present application are intended to solve the problems associated with the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings, in which:

[0043] Figure 1 Schematically illustrates an apparatus system suitable for implementing some embodiments;

[0044] Figure 2 schematically illustrates a metadata encoder according to some embodiments;

[0045] Figure 3 Showing how according to some embodiments Figure 2 A flowchart of the operation of the metadata encoder shown in ;

[0046] Figure 4 An example device suitable for implementing the illustrated means is schematically illustrated. DETAILED DESCRIPTION

[0047] Suitable devices and possible mechanisms for providing efficient spatial analysis-derived metadata parameters are described in more detail below. In the following discussion, multi-channel systems are discussed with respect to multi-channel microphone implementations. However, as discussed above, the input format can be any suitable input format, such as multi-channel loudspeakers, Ambisonics (FOA / HOA), etc. It will be appreciated that in some embodiments, the channel positions are based on the positions of the microphones or are virtual positions or directions. Furthermore, the output of the example system is a multi-channel loudspeaker arrangement. However, it will be appreciated that the output can be rendered to the user via means other than loudspeakers. Furthermore, the multi-channel loudspeaker signal can be summarized as two or more playback audio signals. Currently, the 3GPP standardization body is standardizing such a system as Immersive Voice and Audio Services (IVAS). IVAS is intended to be an extension of the existing 3GPP Enhanced Voice Services (EVS) codec to facilitate immersive voice and audio services on existing and future mobile (cellular) and fixed-line networks. IVAS applications can provide immersive voice and audio services over 3GPP fourth-generation (4G) and fifth-generation (5G) networks. In addition, the IVAS codec, which is an extension of EVS, can be used for store-and-forward applications, where audio and voice content is encoded and stored in a file for playback. It is understood that IVAS can be used in conjunction with other audio and voice coding techniques that encode samples of audio and voice signals.

[0048] For each considered time-frequency (TF) block or tile, in other words, for each time / frequency subband, the metadata includes at least the spherical direction (elevation, azimuth), at least one energy ratio of the resulting direction, the extended coherence, and the direction-independent surround coherence. In general, IVAS can have multiple different types of metadata parameters for each time-frequency (TF) tile. The types of spatial audio parameters that can constitute metadata for IVAS are shown in Table 1 below.

[0049] This data may be encoded by an encoder and sent (or stored) so that the spatial signal can be reconstructed at a decoder.

[0050] Furthermore, in some cases, metadata-assisted spatial audio (MASA) can support up to two directions per TF tile, which requires encoding and sending the above parameters for each direction on a per-TF tile basis, thus potentially doubling the required bitrate according to Table 1.

[0051]

[0052]

[0053] This data may be encoded by an encoder and sent (or stored) so that the spatial signal can be reconstructed at a decoder.

[0054] The bitrate allocated for metadata in practical immersive audio communication codecs may vary widely. A typical overall operating bitrate of a codec may leave only 2 to 10 kbps for transmission / storage of spatial metadata. However, some further implementations may allow up to 30 kbps or more for transmission / storage of spatial metadata. The encoding of directional parameters and energy ratio components, as well as the encoding of coherence data, have been examined previously. However, regardless of the transmission / storage bitrate allocated for spatial metadata, it will always be necessary to represent these parameters using as few bits as possible, especially when the TF tiles can support multiple directions corresponding to different sound sources in the spatial audio scene.

[0055] The concept as discussed below is to encode metadata spatial audio parameters for each TF tile by merging the spatial parameters across multiple frequency bands of a temporal subframe / frame and / or merging the spatial parameters across multiple temporal subframes / frames for a specific frequency band.

[0056] Therefore, the present invention is based on the consideration that the bit rate on a per-TF tile basis can be reduced by merging the spatial audio parameters associated with each TF tile across multiple frequency bands and / or multiple temporal subframes / frames.

[0057] in this regard, Figure 1 Depicted are example devices and systems for implementing embodiments of the present application. System 100 is shown as having an 'analysis' portion 121 and a 'synthesis' portion 131. The 'analysis' portion 121 is the portion that encodes the metadata and downmix signals from receiving the multi-channel loudspeaker signals, while the 'synthesis' portion 131 is the portion that decodes the encoded metadata and downmix signals to present the regenerated signals (e.g., in the form of multi-channel loudspeakers).

[0058] The input to the system 100 and the 'analysis' portion 121 is a multi-channel signal 102. In the following example, a microphone channel signal input is described; however, in other embodiments, any suitable input (or synthesized multi-channel) format may be implemented. For example, in some embodiments, the spatial analyzer and spatial analysis may be implemented externally to the encoder. For example, in some embodiments, spatial metadata associated with the audio signal may be provided to the encoder as a separate bitstream. In some embodiments, the spatial metadata may be provided as a set of spatial (directional) index values. These are examples of metadata-based audio input formats.

[0059] The multi-channel signal is passed to the transmission signal generator 103 and the analysis processor 105 .

[0060] In some embodiments, the transmission signal generator 103 is configured to receive a multi-channel signal, generate a suitable transmission signal comprising a predetermined number of channels, and output a transmission signal 104. For example, the transmission signal generator 103 may be configured to generate a 2-channel downmix of the multi-channel signal. The predetermined number of channels may be any suitable number of channels. In some embodiments, the transmission signal generator is configured to select or combine the input audio signal into the predetermined number of channels using beamforming techniques and output the selected channels as the transmission signal.

[0061] In some embodiments, the transmission signal generator 103 is optional and the multi-channel signal is passed unprocessed to the encoder 107 in the same manner as the transmission signal in this example.

[0062] In some embodiments, the analysis processor 105 is also configured to receive the multichannel signals and analyze these signals to generate metadata 106 associated with the multichannel signals and, therefore, the transmission signal 104. The analysis processor 105 can be configured to generate metadata that may include a direction parameter 108, an energy ratio parameter 110, and a coherence parameter 112 (and, in some embodiments, a diffuseness parameter) for each time-frequency analysis interval. In some embodiments, the direction, energy ratio, and coherence parameters may be considered spatial audio parameters. In other words, the spatial audio parameters include parameters intended to characterize the sound field created / captured by the multichannel signals (or, in general, two or more audio signals).

[0063] In some embodiments, the parameters generated may vary from frequency band to frequency band. Thus, for example, in frequency band X, all parameters are generated and transmitted, while in frequency band Y, only one parameter is generated and transmitted. Furthermore, in frequency band Z, no parameters are generated or transmitted. A practical example of this could be that for some frequency bands, such as the highest frequency band, certain parameters are not required for perceptual reasons. The transmission signal 104 and metadata 106 may be passed to the encoder 107.

[0064] The encoder 107 may include an audio encoder core 109 configured to receive the transport (e.g., downmix) signals 104 and generate a suitable encoding of these audio signals. In some embodiments, the encoder 107 may be a computer (running suitable software stored on memory and at least one processor), or alternatively may be a specific device using, for example, an FPGA or ASIC. The encoding may be implemented using any suitable scheme. The encoder 107 may also include a metadata encoder / quantizer 111 configured to receive metadata and output an encoded or compressed form of that information. In some embodiments, Figure 1 The encoder 107 may further interleave, multiplex into a single data stream, or embed metadata into the encoded downmix signal before transmission or storage as shown by the dashed line in FIG. Multiplexing may be achieved using any suitable scheme.

[0065] On the decoder side, the received or retrieved data (stream) can be received by a decoder / demultiplexer 133. The decoder / demultiplexer 133 can demultiplex the encoded stream and pass the audio encoded stream to a transport extractor 135, which is configured to decode the audio signal to obtain a transport signal. Similarly, the decoder / demultiplexer 133 can include a metadata extractor 137, which is configured to receive the encoded metadata and generate metadata. In some embodiments, the decoder / demultiplexer 133 can be a computer (running appropriate software stored in memory and on at least one processor), or alternatively can be a dedicated device, such as using an FPGA or ASIC.

[0066] The decoded metadata and transport audio signal may be passed to a synthesis processor 139 .

[0067] The 'synthesis' portion 131 of the system 100 further illustrates a synthesis processor 139 configured to receive the transmission and metadata and, based on the transmission signal and the metadata, recreate synthesized spatial audio in the form of the multi-channel signal 110 in any suitable format (which may be a multi-channel loudspeaker format, or in some embodiments any suitable output format such as a binaural or Ambisonics signal, depending on the use case).

[0068] Thus, in summary, firstly, the system (analysis part) is configured to receive a multi-channel audio signal.

[0069] Furthermore, the system (analysis part) is configured to generate a suitable transmission audio signal (eg by selecting or downmixing some audio signal channels) and spatial audio parameters as metadata.

[0070] The system is further configured to encode the transmission signal and metadata for storage / transmission.

[0071] After this, the system can store / send the encoded transmission and metadata.

[0072] The system can retrieve / receive the encoded transmission and metadata.

[0073] Furthermore, the system is configured to extract the transport and metadata from the encoded transport and metadata parameters, eg, demultiplex and decode the encoded transport and metadata parameters.

[0074] The system (synthesis part) is configured to synthesize and output a multi-channel audio signal based on the extracted transmission audio signal and metadata.

[0075] about Figure 2 , describes in more detail an example analysis processor 105 and metadata encoder / quantizer 111 (e.g., Figure 1 ).

[0076] Figure 1 and Figure 2 The metadata encoder / quantizer 111 and the analysis processor 105 are depicted as being coupled together. However, it should be understood that some embodiments may not so tightly couple the two respective processing entities, such that the analysis processor 105 may reside on a separate device from the metadata encoder / quantizer 111. Thus, there may be a device including the metadata encoder / quantizer 111 where the processing and encoding of the transmission signal and metadata stream are independent of the capture and analysis process. In this case, the energy estimator 205 may be configured as part of the metadata encoder / quantizer 111.

[0077] In some embodiments, the analysis processor 105 includes a time-frequency domain converter 201 .

[0078] In some embodiments, the time-frequency domain converter 201 is configured to receive the multi-channel signal 102 and apply a suitable time-frequency domain transform, such as a short-time Fourier transform (STFT), to convert the input time-domain signal into suitable time-frequency signals. These time-frequency signals can be passed to the spatial analyzer 203.

[0079] Thus, for example, the time-frequency signal 202 may be represented in the time-frequency domain as:

[0080] s i (b, n)

[0081] Where b is the frequency bin index, n is the time-frequency block (frame) index, and i is the channel index. In another expression, n can be considered as a time index with a sampling rate lower than the sampling rate of the original time domain signal. These frequency bins can be grouped into multiple subbands, which group one or more bins into subbands with frequency band index k=0,...,K-1. Each subband k has a minimum bin b k,low and the highest position b k,high , and the subband contains k,low to b k,high The widths of the sub-bands can be approximated by any suitable distribution, for example, the Equivalent Rectangular Bandwidth (ERB) scale or the Bark scale.

[0082] Thus, a time-frequency (TF) tile (or block) is a specific subband within a subframe of the frame.

[0083] It will be appreciated that the number of bits required to represent the spatial audio parameters may depend at least in part on the TF (time-frequency) tile resolution (i.e., the number of TF subframes or tiles). For example, a 20ms audio frame may be divided into 4 time domain subframes of 5ms each, and each time domain subframe may have a maximum of 24 frequency subbands divided in the frequency domain according to the Bark scale, its approximation, or any other suitable division. In this particular example, the audio frame may be divided into 96 TF subframes / tiles, in other words, 4 time domain subframes with 24 frequency subbands. Therefore, the number of bits required to represent the spatial audio parameters of an audio frame may depend on the TF tile resolution. For example, if each TF tile is to be encoded according to the distribution in Table 1 above, each TF tile will require 64 bits (for one sound source direction per TF tile).

[0084] Embodiments aim to reduce the number of bits on a per-frame basis by combining TF tiles in the time domain or the frequency domain.

[0085] Return to Figure 2 , the time-frequency signal 202 may be passed to the energy estimator 205, whereby the energy of each frequency sub-band k may be determined for all channels i of the time-frequency signal 202. In an embodiment, this operation may be expressed according to the following formula:

[0086]

[0087] The time-frequency audio signal is denoted as S(i, b, n), where i is the channel index, b is the frequency bin index, n is the time subframe index, and b is the time subframe index. k,low is the lowest bin of band k, b k,high It is the highest warehouse.

[0088] Furthermore, the energy of each subband k within the time subframe n may be passed to the spatial parameter merger 207 .

[0089] In an embodiment, the analysis processor 105 may comprise a spatial analyser 203. The spatial analyser 203 may be configured to receive the time-frequency signals 202 and based on these signals estimate the directional parameters 108. The directional parameters may be determined based on any audio based 'directional' determination.

[0090] For example, in some embodiments, the spatial analyzer 203 is configured to estimate the direction of a sound source using two or more signal inputs.

[0091] Therefore, the spatial analyzer 203 may be configured to provide at least one azimuth and elevation angle for each frequency band and temporal time-frequency block within a frame of the audio signal, denoted as azimuth and elevation angle, respectively. and the elevation angle θ(k,n). The directional parameters 108 for the time subframe may also be passed to the spatial parameter merger 207.

[0092] The spatial analyzer 203 may also be configured to determine an energy ratio parameter 110. The energy ratio may be considered as determining the energy of the audio signal that may be considered to arrive from one direction. The direct to total energy ratio r(k,n) may be estimated, for example, using a stability measure of the directional estimate, or using any correlation measure, or any other suitable method for obtaining a ratio parameter. Each direct to total energy ratio corresponds to a particular spatial direction and describes how much energy is coming from that particular spatial direction compared to the total energy. This value may also be expressed separately for each time-frequency tile. The spatial direction parameter and the direct to total energy ratio describe, for each time-frequency tile, how much of the total energy is coming from that particular direction. In general, the spatial direction parameter may also be considered as a direction of arrival (DOA).

[0093] In an embodiment, the direct-to-total energy ratio parameter may be estimated based on a normalized cross-correlation parameter cor′(k,n) between a pair of microphones in frequency band k, the value of which is between −1 and 1. By multiplying the normalized cross-correlation parameter with the diffuse field normalized cross-correlation parameter cor′ D (k, n), the total energy ratio parameter r(k, n) can be directly determined as:

[0094]

[0095] The total energy ratio is further described in PCT publication WO 2017 / 005978 (which is incorporated herein by reference). The energy ratio can be passed to the spatial parameter merger 207.

[0096] The spatial analyzer 203 may also be configured to determine a plurality of coherence parameters 112 , which may include surround coherence (γ(k,n)) and extended coherence (ζ(k,n)), both of which are analyzed in the time-frequency domain.

[0097] Each of the aforementioned coherence parameters is discussed below. All processing is performed in the time-frequency domain, so for the sake of brevity, the time-frequency indices k and n are removed where necessary.

[0098] Let us first consider the case where two spaced apart speakers (eg, front left and front right) are used instead of a single speaker to reproduce sound coherently. The coherence analyzer can be configured to detect that this approach has been applied to surround mixing.

[0099] It should be understood that the following section describes the analysis of extension and surround coherence in terms of a multi-channel loudspeaker signal input. However, similar practices can be applied when the input includes a microphone array as an input.

[0100] Therefore, in some embodiments, the spatial analyzer 203 may be configured to calculate a covariance matrix C for a given analysis interval comprising one or more time indices n and frequency bins b. The matrix size is N L x N L , and its elements are denoted as c ij , where N L is the number of speaker channels, i and j are the speaker channel indices.

[0101] Next, the spatial analyzer 203 may be configured to determine the speaker channel i closest to the estimated direction (in this example, the azimuth angle θ). c :

[0102] i c =arg(min(|θ-α i |))

[0103] Among them, α i is the angle of speaker i.

[0104] Furthermore, in such an embodiment, the spatial analyzer 203 is configured to determine the spatial c On the left side l and right i r Closest speaker.

[0105] The normalized coherence between loudspeakers i and j is expressed as:

[0106]

[0107] Using this equation, spatial analyzer 203 may be configured to calculate i l with i r The normalized coherence c′ between lr In other words, calculate:

[0108]

[0109] Furthermore, the spatial analyzer 203 may be configured to use the diagonal elements of the covariance matrix to determine the energy of loudspeaker channel i:

[0110] E i =c ii

[0111] And determine the speaker i l and i r With speaker i l 、i r and i c The energy ratio between them is:

[0112]

[0113] Furthermore, the spatial analyzer 203 may use these determined variables to generate a 'stereoness' parameter:

[0114] μ=c′ lr ξ lr / lrc

[0115] This 'stereoscopic' parameter has a value between 0 and 1. A value of 1 means that the speaker i l and i r A coherent sound is present in the sector and dominates the energy of the sector. This can be caused, for example, by the speaker mix using amplitude panning techniques to create an "airy" perception of the sound. A value of 0 means that no such technique has been applied, and the sound can simply be localized to the nearest speaker, for example.

[0116] Furthermore, the spatial analyzer 203 may be configured to detect or at least identify situations where three (or more) speakers are used to coherently reproduce sound to create a "close" perception (e.g., using left front, right front, and center instead of just center). This may be because the mixing engineer created this situation when performing a surround mix on a multi-channel speaker mix.

[0117] In this embodiment, the coherence analyzer uses the same previously identified loudspeaker i l 、i r and i c The normalized coherence value c′ is determined using the normalized coherence determination discussed previously. cl and c′ cr . In other words, calculate the following:

[0118]

[0119] Furthermore, the spatial analyzer 203 may determine a normalized coherence value c′ describing the coherence between the loudspeakers using the following equation: clr :

[0120] c′ clr =min(c′cl , c′ cr )

[0121] Additionally, the spatial analyzer 203 may be configured to determine a spatial distribution that describes how (ie, how uniform) the energy is in channel i. l 、i r and i c Parameters evenly distributed between:

[0122]

[0123] Using these variables, spatial analyzer 203 can determine a new coherence shift parameter κ as:

[0124] κ=c′ clr ξ clr

[0125] The coherence pan parameter κ has a value between 0 and 1. A value of 1 means that at all loudspeakers i l 、i r and i c coherent sound is present in the mix, and its energy is evenly distributed across the speakers. This can be because, for example, the speaker mix was generated using recording mixing techniques designed to create the perception of closer sound sources. A value of 0 means that no such techniques have been applied; for example, the sound is simply localized to the closest speaker.

[0126] Determine the metric i l and i r in (but not in i c The "stereo" parameter μ of the coherent sound volume and the measure of all i l 、i r and i c The spatial analyzer 203 is configured to use the coherence panning parameters κ of the coherent sound volume in φ(t) to determine a coherence parameter to be output as metadata.

[0127] Therefore, the spatial analyzer 203 is configured to combine the "stereoscopic" parameter μ and the coherence shift parameter κ to form an extended coherence ζ parameter, which has a value from 0 to 1. An extended coherence ζ value of 0 indicates a point source, in other words, the source should be detected with as few loudspeakers as possible (e.g., only loudspeaker i c ) to reproduce the sound. As the value of the extended coherence ζ increases, more energy is extended to the speaker i c surrounding speakers; until the value is 0.5, the energy is in speaker i l 、i r and i c When the value of the extended coherence ζ exceeds 0.5, the loudspeaker i cThe energy in the speaker decreases until the value is 1, c There is no energy in the speaker, and all the energy is in the speaker l and i r Place.

[0128] In some embodiments, using the above parameters μ and κ, the spatial analyzer 203 is configured to determine the extended coherence parameter ζ using the following expression:

[0129]

[0130] The above expression is merely an example, and it should be noted that the spatial analyzer 203 may adopt any other manner to estimate the extended coherence parameter ζ as long as it conforms to the above parameter definition.

[0131] In addition to being configured to detect the previous situations, the spatial analyzer 203 may also be configured to detect or at least identify situations in which sound is reproduced coherently from all (or nearly all) speakers to create an "inside-the-head" or "above" perception.

[0132] In some embodiments, the spatial analyzer 203 may be configured to determine the energy E having the maximum value. i and speaker channel i e Classify / sort.

[0133] Furthermore, the spatial analyzer 203 may be configured to determine whether this channel is L Normalized coherence c′ between the other loudest channels ij . Then, you can monitor the channel and M L These normalized coherences c′ between the other loudest channels ij In some embodiments, M L Can be N L -1, which would mean monitoring the coherence between the loudest channel and all other speaker channels. However, in some embodiments, M L It can be a smaller number, for example, N L Using these normalized coherence values, the coherence analyzer can be configured to determine the surround coherence parameter γ using the following expression:

[0134]

[0135] Among them, c′ iej is the loudest channel and M L Normalized coherence between the next loudest channels.

[0136] The surround coherence parameter γ has a value from 0 to 1. A value of 1 means that there is coherence between all (or almost all) of the loudspeaker channels. A value of 0 means that there is no coherence between all (or even almost all) of the loudspeaker channels.

[0137] The above expression is merely an example of estimating the surround coherence parameter γ, and any other way may be used as long as it satisfies the above parameter definition.

[0138] The spatial analyzer 203 may be configured to output the determined extended coherence parameter ζ and surround coherence parameter γ to the spatial parameter merger 207 .

[0139] Thus, for each subband k, there will be a set of spatial audio parameters associated with that subband. In this case, each subband k may have the following spatial parameters associated with it: at least one azimuth and elevation angle (denoted as azimuth φ(k,n) and elevation θ(k,n)), surround coherence (γ(k,n)), extended coherence (ζ(k,n)), and a direct-to-total energy ratio parameter r(k,n).

[0140] In an embodiment, the spatial parameter merger 207 may be configured to combine (or merge) each of the plurality of aforementioned parameters into a smaller number of frequency bands. For example, taking a TF tile having 24 frequency bands (i.e., k spans from 0 to 23), the spatial parameter values ​​for each of the 24 frequency bands are merged into values ​​associated with a smaller number of frequency bands, wherein each of the smaller number of frequency bands spans a continuous number of the original 24 frequency bands.

[0141] in this regard, Figure 3 Depicted are some processing steps that the spatial parameter merger 207 may be configured to perform in some embodiments.

[0142] The spatial parameter merger 207 can perform the above-mentioned merger by initially obtaining the azimuth angle φ(k,n) and elevation angle θ(k,n) spherical direction components for each of the K subbands and converting each direction component into its corresponding Cartesian coordinate vector. In turn, each Cartesian coordinate vector for subband k can be weighted by the corresponding energy E(k,n) for subband k (from the energy estimator 205) and the total energy ratio parameter r(k,n).

[0143] The conversion operation of the azimuth angle φ(k,n) and elevation angle θ(k,n) direction components of subband k gives the X-axis direction component as:

[0144] x(k,n)=E(k,n)r(k,n)coSφ(k,n)cosθ(k,n) (1)

[0145] The Y-axis component is:

[0146] y(k,n)=E(k,n)r(k,n)sinφ(k,n)cosθ(k,n) (2)

[0147] The Z-axis component is:

[0148] z(k,n)=E(k,n)r(k,n)sinθ(k,n) (3)

[0149] The above operation may be performed on all sub-bands k=0 to K-1.

[0150] The steps of converting the spherical direction component of each subband k of subframe n into its equivalent Cartesian coordinates x, y, z are as follows: Figure 3 As shown in processing step 301.

[0151] The steps of weighting each Cartesian coordinate x, y, z with respect to the energy of subband k and directly to the total energy parameter are as follows Figure 3 As shown in step 303.

[0152] in this regard, Figure 3 Also depicted is the step of receiving the energy for each sub-band from the energy estimator 205. This is shown as processing step 315. The respective energy for each sub-band is shown as being used in step 303.

[0153] Furthermore, the spatial parameter merger 207 may be configured to merge the above Cartesian coordinates for multiple subbands 0 to K-1 into a single "merged" frequency band. This merging process may be repeated for multiple consecutive subband groups so that all subbands 0 to K-1 have been merged into fewer merged frequency bands p = 0 to P-1, where P < K.

[0154] For example, the merging process for the first merged frequency band p=0 may include grouping of Cartesian coordinates for the first k1 (0 to k1-1) frequency bands in subbands 0 to K-1, the second merged frequency band p=1 may include grouping of Cartesian coordinates for the second k1 (k1 to 2*k1-1) frequency bands in subbands 0 to K-1, the third merged frequency band p=2 may include grouping of Cartesian coordinates for the third k1 (2*k1 to 3*k1-1) frequency bands in subbands 0 to K-1, and so on, until the final merged frequency band p=P-1 includes the Cartesian coordinates of the last subband in the K subbands.

[0155] It should be noted that the number of grouped subbands does not have to be fixed to k1, but can vary from one combined frequency band to another. In other words, the first combined frequency band p=0 can include the Cartesian coordinates of the first k1 subbands, while the second combined frequency band p=1 can include the Cartesian coordinates of the next k2 subbands, where k1 is a different number from k2.

[0156] In an embodiment, the grouping (or combining) mechanism may comprise a summing step, wherein the Cartesian coordinates are summed for the set of sub-bands assigned to a particular combined frequency band.

[0157] Returning to the above example of a subframe n having 24 subbands, the spatial parameter merger 207 can be configured to merge the Cartesian coordinates of the 24 subbands into four merged frequency bands, where each merged frequency band includes the merged Cartesian coordinates of six subbands. In this example, the x Cartesian coordinate merging process as performed by the spatial parameter merger 207 for the first merged frequency band can be expressed as:

[0158]

[0159] The second merged frequency band in this example can be given as:

[0160]

[0161] The third combined frequency band in this example can be given as:

[0162]

[0163] The fourth combined frequency band in this example may be given as:

[0164]

[0165] The above algorithm steps can be repeated for the y and z Cartesian coordinates to give y MF (p, n) and z MF (p, n) (p=0 to 3). Note that in the above expression, n is the time subframe index. In general, for the combined frequency band p, the above example can be expressed as:

[0166]

[0167] Among them, k p.low is the low-frequency subband of the combined frequency band p, k p.high is the high frequency sub-band of the merged frequency band p.

[0168] The step of merging the set of Cartesian coordinates into a plurality of merged frequency bands is performed in Figure 3This is shown as processing step 305 , where each merged frequency band comprises the Cartesian coordinates of a plurality of consecutive sub-bands k.

[0169] Once the Cartesian coordinates x, y, z for subbands k=0 to K-1 have been merged into Cartesian coordinates x for the merged frequency band p=0 to P-1 (where P<K), MF 、y MF and z MF (According to the above process steps), the combined Cartesian coordinate x MF 、y MF and z MF can be converted into its equivalent combined azimuth φ MF (p, n) and elevation angle θ MF (p, n) spherical direction components. In an embodiment, the P combined Cartesian coordinates x can be expressed as MF 、Y MF and z MF Each of these performs this transformation:

[0170]

[0171]

[0172] The atan function automatically detects the arc tangent of the correct quadrant of the angle.

[0173] The step of converting the combined Cartesian coordinates into their equivalent combined spherical coordinates for each combined frequency band is as follows Figure 3 As shown in step 307.

[0174] Following from the above, the corresponding combined direct-to-total energy ratio r can be determined for each combined frequency band p by taking the length of the vector formed from the Cartesian coordinates for the combined frequency band p as described above and normalizing the length of this vector by the energy of the combined frequency band p. MF (p, n). In an embodiment, the combined direct-to-total energy ratio r for the combined frequency band p is MF (p, n) can be expressed as:

[0175]

[0176] Among them, as mentioned above, E(k,n) is the original frequency band k for the p-th combined frequency band p.low to k p.high The energy of the signal contained in .

[0177] Determine the combined direct-to-total energy ratio r for each combined frequency band MF The step of (using input from process step 315 ) is shown as process step 309 .

[0178] Additionally, some embodiments may derive a combined extended coherence for each combined frequency band p by using the extended coherence value ζ(k,n) calculated for each sub-band k. Combined extended coherence ζ for combined frequency band p MF (p, n) may be calculated as the energy-weighted average of the extended coherence values ​​of the frequency sub-bands constituting the combined frequency band p. In an embodiment, the combined extended coherence for the combined frequency band p may be expressed as:

[0179]

[0180] Determine the combined extended coherence value ζ for each combined frequency band MF The step of is shown as processing step 311 (with input from processing step 315).

[0181] Similarly, some embodiments may derive the combined surround coherence for each combined frequency band p by using the surround coherence value γ(k,n) calculated for each sub-band k. Combined surround coherence γ for combined frequency band p MF (k, n) may be calculated as the energy-weighted average of the ambient coherence values ​​of the frequency sub-bands constituting the merged frequency band p. In an embodiment, the merged ambient coherence for the merged frequency band p may be expressed as:

[0182]

[0183] Determine the ambient coherence value γ for each combined frequency band MF The step of is shown as processing step 313 (with input from processing step 315).

[0184] In a further embodiment, the spatial parameter merger 207 may also be configured to combine spatial parameters such as the azimuth angle φ(k,n) and the elevation angle θ(k,n), the surround coherence (γ(k,n)) and the extended coherence (ζ(k,n)), and the direct-to-total energy ratio parameter r(k,n) across multiple time subframes n. For example, the spatial parameters for frequency band k may be combined (or merged) across multiple subframes n0 to N-1. In this case, the spatial parameter values ​​for multiple time subframes may be merged into a merged value associated with a smaller number of consecutive time subframes.

[0185] In a corollary of step 305, the spatial parameter merger 207 may be configured to merge the azimuth angle φ(k,n) and elevation angle θ(k,n) values ​​for a particular frequency subband k across multiple subframes n for a plurality of consecutive groups. In a manner similar to step 301, the spatial parameter merger may convert the azimuth angle φ(k,n) and elevation angle θ(k,n) values ​​for the particular subband k for n=0 to N-1 subframes into their corresponding Cartesian coordinate vectors for subframe n. In turn, each Cartesian coordinate for subframe n may be weighted by the corresponding energy E(k,n) for the particular subframe n (as generated by the energy estimator 205) and directly to the total energy parameter r(k,n).

[0186] The Cartesian coordinates x(k,n), y(k,n) and z(k,n) can be determined by calculating equations (1), (2) and (3) for subband k over time subframes (or frames) indexed n=0 to N-1.

[0187] Furthermore, the spatial parameter merger 207 may be configured to merge the Cartesian coordinates for a plurality of subframes into a single merged time frame q. In a manner similar to the frequency merger process embodiment described above, this merger process may be repeated for a plurality of consecutive subframe groups so that all subframes 0 to N-1 have been merged into fewer merged frames q = 0 to Q-1, where Q < N.

[0188] For example, the merging process for the first merged time frame q=0 may include grouping of Cartesian coordinates for the first n1 (0 to n1-1) time subframes in subframes 0 to N-1, the second merged time frame q=1 may include grouping of Cartesian coordinates for the second n1 (n1 to 2*n1-1) subframes in subframes 0 to N-1, the third merged time frame q=2 may include grouping of Cartesian coordinates for the third n1 (2*n1 to 3*n1-1) subframes in subframes 0 to N-1, and so on, until the final merged time frame q=Q-1 includes the Cartesian coordinates of the last subframe in N subframes.

[0189] It should be noted that the number n of subframes to be merged need not be fixed to n1, but may vary from one merged frame to another. In other words, the first merged frame q=0 may include the Cartesian coordinates of the first n1 subframes, and the second merged frame q=1 may include the Cartesian coordinates of the next n2 subframes, where n1 is a different number from n2.

[0190] Similarly, in these embodiments, the grouping mechanism may further comprise a summing step, wherein the Cartesian coordinates of a particular merged time frame are summed for the set of subframes assigned to the particular merged time frame.

[0191] Therefore, the x, y, and z coordinates x of the merged time frame q are MT 、y MT 、z MT can be expressed as:

[0192]

[0193]

[0194]

[0195] Among them, n q.low is the lower numbered subframe of the merged frame q, n q.high are the higher numbered subframes of the merged frame q.

[0196] In the corollary of processing step 307, the time subframe Cartesian coordinates x for the merged time frames q=0 to Q-1 (where Q<N) are MT 、y MT and z MT can also be converted into its equivalent combined azimuth φ MT (k, q) and elevation angle θ MT (k, q) spherical direction components. In an embodiment, the Q combined Cartesian coordinates x can be expressed as MT 、y MT and z MT Each of these performs this transformation:

[0197]

[0198]

[0199] As mentioned previously, the function atan automatically detects the correct quadrant for the arc tangent of an angle.

[0200] In a similar manner to the above-described embodiment where the merging process is across frequency subbands, the corresponding direct-to-total energy ratio r for the merged time frame q is MT (k,q) can be given as:

[0201]

[0202] in, E(k,n) is the original subframe for subband k for the qth merged subframe n q.low to n q.high The energy of the signal contained in .

[0203] Furthermore, the merged extended coherence for each merged time frame q for subband k can be derived by using the extended coherence value γ(k,n) calculated across the subframes of the merged time frame q:

[0204]

[0205] Similarly, the merged surround coherence for each merged time frame a for subband k may be derived by using the surround coherence values ​​γ(k,n) calculated across the subframes of the merged time frame q.

[0206]

[0207] In turn, the output of the spatial parameter merger 207 may comprise merged spatial audio parameters, which may be arranged to be passed to the metadata encoder / quantizer 111 for encoding and quantization.

[0208] In some embodiments, the merged spatial parameters may include a merged frequency band parameter θ for each of the merged frequency bands on a per-subframe basis. MF 、φ MF 、r MF , γ MF ,ζ MF .

[0209] In other embodiments, the combined spatial parameters may include a combined temporal frame parameter θ for each subband k. MT 、φ MT 、r MT , γ MT ,ζ MT .

[0210] In a further embodiment, the spatial parameter merger 207 may be configured such that the merger process is performed in a cascaded manner, whereby the spatial parameters may be merged first according to the above-described frequency band-based merger process followed by the above-described time frame-based merger process. Alternatively, the cascaded merger process as performed by the spatial parameter merger 207 may be reversed such that the above-described time frame-based merger process is followed by the above-described frequency band-based merger process.

[0211] In yet further embodiments, the spatial parameter merger 207 may be configured such that a merger process is performed such that parameters may be merged according to the above-described frequency band-based merger process together with a time frame-based merger process. This may be accomplished using the above-described merger formula according to n q.low and n q.high 、k p.low and k p.high restrictions to be enforced.

[0212] In an embodiment, the spatial parameter merger 207 may have an additional functional unit (actually an importance estimator) that provides an estimate (or measure) of the importance (actually an importance estimator) of the total number of spatial parameter sets (or directions) per TF tile relative to the reduced number of merged spatial parameter sets (and therefore based on a reduced number of directions per frame). In addition, the importance estimator may be used to determine whether a particular subband and / or temporal subframe should include merged or non-merged spatial audio parameters.

[0213] The importance estimates may be fed to a decision function within the spatial parameter merger 207 which decides whether the output (which is subsequently to be encoded) may include spatial audio parameters for each TF tile, or whether the output includes merged spatial audio parameters, or indeed whether specific subbands and / or groups of subframes in a time frame should have merged or non-merged spatial audio parameters.

[0214] Using the above example where spatial parameter sets are merged across frequency bands and / or across subframes in time, the role of the importance estimator may be to estimate the importance of the perceived audio quality of using the (unmerged) spatial audio parameter set for each TF tile rather than using the merged spatial audio parameter set across multiple frequency bands and / or multiple temporal subframes.

[0215] To this end, the importance measure may be estimated by comparing the length of the calculated merged Cartesian coordinate vector (derived as above) with the sum of the vector lengths of the (unmerged) Cartesian coordinates (summed over the merged subband and / or merged subframe).

[0216] Returning to the band-based merging example above, the sum of the vector lengths of the (unmerged) Cartesian coordinates (summed over the subbands merged into band p) can be expressed as:

[0217]

[0218] The length of the calculated merged Cartesian coordinate vector for the merged frequency band p can be written as:

[0219]

[0220] Furthermore, the importance estimate (or metric) λ(p,n) for the p-th merged frequency band can be expressed as:

[0221]

[0222] In this case, the choice of whether to encode and transmit the merged or unmerged spatial audio parameter set may be based on whether the importance measure λ(p,n) exceeds a threshold λth comparison.

[0223] Thus, if λ(p,n)>λ th , a decision can be made to encode and send the unmerged spatial audio parameters as metadata.

[0224] If λ(p,n)≤λ th , a decision can be made to encode and send the combined spatial audio parameters as metadata.

[0225] In case it is decided to send the unmerged spatial audio parameters as metadata, the spatial parameter merger 207 may be configured to output the original set of spatial audio parameters. For example, if the above comparison shows that it is advantageous to output the unmerged spatial audio parameters instead of the merged spatial audio parameters for the p-th merged frequency band, then the spatial parameter merger 207 may be configured to output the original set of spatial audio parameters for the subband k. p,low to k p,high The following spatial audio parameters φ(k,n), θ(k,n), γ(k,n), ζ(k,n) and r(k,n) of may form the output for the p-th merged frequency band.

[0226] In case it is decided to send the merged spatial audio parameters as metadata, in other words, in case it is decided to send a set of spatial audio parameters for a merged set of subbands and / or a merged set of subframes, the spatial parameter merger 207 may be configured to output the merged spatial audio parameters, and in case of a merged frequency band p, the output parameters may include the parameter set θ MF 、φ MF 、r MF , γ MF ,ζ MF .

[0227] In other embodiments, an average importance value may be determined for multiple subframes and / or subbands. This may be achieved by taking the average of the importance metrics over a set of importance metrics, such as:

[0228]

[0229] Where N is the number of subframes in frame m in this case, but may instead be averaged over multiple subbands, or in other embodiments, the importance measures may be averaged across a combination of frequency bands and time frames. An advantage of using an average of the importance measures is that signaling bits are only required for a group of merged frames and / or frequency bands, rather than for each merged time frame and / or frequency band.

[0230] It will be appreciated that in the above case, it may be necessary to include signaling bits in the metadata in order to indicate whether the spatial audio parameters are merged or not.

[0231] The importance metric may have the property that when all directions (across the merged subframes and / or subbands) point approximately in the same direction, the importance metric will tend to have a low value (close to a value of "zero"). However, in contrast, if all directions tend to point in opposite directions, and the direct-to-total energy ratio associated with each direction is approximately the same, then the importance metric may tend to have a value of 1. A further property exhibited by the importance metric may be that if one of the subbands / subframes has a significantly higher direct-to-total energy ratio than any other subband / subframe, then the importance metric will also tend to have a low value.

[0232] In an embodiment, the threshold λ is selected th The value of can be fixed, and experiments have found that a value of 0.3 gives favorable results.

[0233] In other embodiments, the importance threshold λ may be determined for a frame by the following operations: th : sorting a plurality of importance metrics λ(k,n) for a plurality of merged subbands and / or subframes in ascending order and determining a threshold as a value of the importance metric that gives a certain number of importance metrics (and therefore merged subbands and / or subframes) above the threshold, e.g., the threshold metric may be selected based on the presence of I merged subframes and / or subbands in a frame whose importance metrics are above the selected threshold.

[0234] In a further embodiment, the importance threshold λ th It can be adapted to the running median of the importance metric over the last N time subframes (e.g., the last 20 subframes). med (n) may indicate the median of the importance metrics for subframe n over the last N subframes in all frequency bands. th (n) can be expressed as λ th (n) = c th λ med (n), where c th is a coefficient that controls the value of the significance threshold, e.g., c th Can be assigned a value of 0.5.

[0235] Additionally, some embodiments may not employ a threshold. In these embodiments, a plurality of the most important TF tiles in a frame / subframe may be set to use the uncombined direction, while the remaining number of TF tiles in the frame / subframe are set to use the combined direction.

[0236] The metadata encoder / quantizer 111 may include a direction encoder. The direction encoder 205 is configured to receive the combined direction parameter (such as the azimuth angle φ MFor φ MT and elevation angle θ MF or θ MT ) (and, in some embodiments, the expected bit allocation), and generates a suitable encoded output therefrom. In some embodiments, the encoding is based on an arrangement of spheres (forming a spherical grid arranged in a ring on the 'surface' sphere), which are defined by a lookup table defined by the determined quantization resolution. In other words, the spherical grid uses the idea of ​​covering a sphere with smaller spheres, and considering the centers of the smaller spheres as points defining a grid of almost equidistant directions. The smaller spheres thus define cones or solid angles about the center point, which can be indexed according to any suitable indexing algorithm. Although spherical quantization is described herein, any other suitable quantization (linear or non-linear) may also be used.

[0237] The metadata encoder / quantizer 111 may include an energy ratio encoder. The energy ratio encoder may be configured to receive the combined energy ratio r MF or r MT , and determining suitable coding for compressing these energy ratios for the merged sub-bands and / or merged time-frequency blocks.

[0238] Similarly, the metadata encoder / quantizer 111 may also include a coherence encoder that may be configured to receive the combined ambient coherence value γ MF or γ MT and the extended coherence value ζ M or MT , and determining a suitable coding for compressing these surrounding and extended coherence values ​​for the merged subband and / or merged time-frequency block.

[0239] The encoded combining direction, energy ratio, and coherence value may be passed to a combiner 211. The combiner is configured to receive the encoded (or quantized / compressed) combining direction parameter, energy ratio parameter, and coherence parameter and combine these parameters to generate a suitable output (e.g., a metadata bitstream, which may be combined with the transmission signal or sent or stored separately from the transmission signal).

[0240] In some embodiments, the encoded data stream is passed to the decoder / multiplexer 133. The decoder / demultiplexer 133 demultiplexes the encoded merge direction index, merge energy ratio index, and merge coherence index and passes them to the metadata extractor 137. In addition, in some embodiments, the decoder / demultiplexer 133 can extract the transmission audio signal and pass it to the transmission extractor 135 for decoding and extraction.

[0241] In an embodiment, the decoder / demultiplexer 133 may be configured to receive and decode signaling bits indicating that the received encoded spatial audio parameters are encoded merged spatial audio parameters for a merged group of subbands and / or subframes, or that the received encoded spatial audio parameters are multiple sets of encoded spatial audio parameters, each parameter set corresponding to a subband or subframe.

[0242] The combined energy ratio index, direction index, and coherence index can be decoded by their respective decoders to generate the combined energy ratio, direction, and coherence for a subframe (when the merging is over the frequency band of the subframe, or for a specific subband when the merging is over consecutive time subframes). This can be performed by applying the inverse of the various encoding processes used at the encoder.

[0243] In case the signaling bits indicate that the spatial audio parameters are not merged, the received spatial audio parameter sets may be directly passed to various decoders for decoding.

[0244] The merged spatial parameters may be passed to a spatial parameter expander (which in some embodiments may form part of the metadata extractor 137) which is configured to expand the merged spatial parameters so as to reproduce the time and frequency resolution of the original spatial parameters at the decoder for subsequent processing and synthesis.

[0245] The combined spatial parameters are given by the combined frequency band parameters θ MF 、φ MF , γ MF ,ζ MF In the case of composition, the extension process may include copying the merged spatial parameters across the original frequency band k over which the spatial parameters were merged.

[0246] For example, in the combined elevation component θ MF In the case of (p, n), the expansion process may consist of simply copying the original frequency subband k p.low to k p.high The value θ for the pth combined frequency band MF (p, n).

[0247] In other words, with respect to the p-th merged frequency band, the extended spatial values ​​θ(k,n) associated with the subbands spanning the p-th merged frequency band can be expressed as:

[0248] θ(k,n) for k p.low to k p.high =θ MF (p, n)

[0249] Obviously, this may be repeated for each merged frequency band p=0 to P-1 to provide values ​​for all sub-bands k=0 to K-1.

[0250] The parameters θ can be calculated for all combined frequency bands. MF 、φ MF , γ MF ,ζ MF This above-described expansion process is performed so as to provide spatial parameters θ(k,n), φ(k,n), γ(k,n), ζ(k,n) for each subband k=0 to K-1.

[0251] The combined spatial parameters are given by the combined time frame parameters θ MT 、φ MT , γ MT ,ζ MT In the case of a composition, the extension process may comprise copying the merged spatial parameters of the original subframe n onto which the cross spatial parameters are merged. Thus, in the merged elevation component θ MT (k, q) case, the extension process may consist of simply copying the original subframe n q.low to n q.high The value θ for the qth merged time frame MT (k, q).

[0252] In other words, with respect to the qth merged time frame, the extended spatial value θ(k,n) associated with the subframes spanning the qth merged time frame can be expressed as:

[0253] θ(k,n) for n q.low to n q.high =θ MF (k, q)

[0254] Obviously, this may be repeated for each merged time frame q=0 to Q-1 to provide values ​​for all subframes n=0 to N-1.

[0255] In inference, we can calculate the parameter θ for all combined time frames. MT 、φ MT , γ MT ,ζ MT The above-described spreading process is performed to provide spatial parameters θ(k,n), φ(k,n), γ(k,n), ζ(k,n) for each subframe n=0 to N-1 (for a particular frequency band k).

[0256] In turn, the decoded and expanded spatial parameters may form decoded metadata output from the metadata extractor 137 , which is passed to a synthesis processor 139 in order to form the multi-channel signal 110 .

[0257] about Figure 4 , illustrates an example electronic device that can be used as an analysis or synthesis device. The device can be any suitable electronic device or apparatus. For example, in some embodiments, the device 1400 is a mobile device, a user device, a tablet computer, a computer, an audio playback device, etc.

[0258] In some embodiments, device 1400 includes at least one processor or central processing unit 1407. Processor 1407 may be configured to execute various program codes, such as the methods described herein.

[0259] In some embodiments, the device 1400 includes a memory 1411. In some embodiments, at least one processor 1407 is coupled to the memory 1411. The memory 1411 can be any suitable storage component. In some embodiments, the memory 1411 includes a program code portion for storing program codes that can be implemented on the processor 1407. In addition, in some embodiments, the memory 1411 can also include a storage data portion for storing data (e.g., data that has been processed or is to be processed according to the embodiments described herein). Whenever needed, the processor 1407 can obtain the implementation program code stored in the program code portion and the data stored in the storage data portion via the memory-processor coupling.

[0260] In some embodiments, device 1400 includes a user interface 1405. In some embodiments, user interface 1405 can be coupled to processor 1407. In some embodiments, processor 1407 can control the operation of user interface 1405 and receive input from user interface 1405. In some embodiments, user interface 1405 can enable a user to enter commands to device 1400, for example, via a keyboard. In some embodiments, user interface 1405 can enable a user to obtain information from device 1400. For example, user interface 1405 can include a display configured to display information from device 1400 to a user. In some embodiments, user interface 1405 can include a touch screen or touch interface that enables information to be entered into device 1400 and also displays information to a user of device 1400. In some embodiments, user interface 1405 can be a user interface for communicating with a location determiner as described herein.

[0261] In some embodiments, device 1400 includes input / output port 1409. In some embodiments, input / output port 1409 includes a transceiver. In such embodiments, the transceiver can be coupled to processor 1407 and configured to enable communication with other devices or electronic devices, for example, via a wireless communication network. In some embodiments, the transceiver or any suitable transceiver or transmitter and / or receiver device can be configured to communicate with other electronic devices or devices via a wired or wired coupling.

[0262] The transceiver can communicate with other devices via any suitable known communication protocol. For example, in some embodiments, the transceiver can use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a wireless local area network (WLAN) protocol such as, for example, IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth, or an infrared data communication path (IRDA).

[0263] The transceiver input / output port 1409 may be configured to receive signals and, in some embodiments, determine parameters as described herein using the processor 1407 executing appropriate code. Additionally, the device may generate appropriate downmix signals and parameter outputs to send to a synthesis device.

[0264] In some embodiments, the device 1400 may be implemented as at least part of a synthesis device. Thus, the input / output port 1409 may be configured to receive a downmix signal and, in some embodiments, parameters determined at a capture device or processing device as described herein, and to generate a suitable audio signal format output using the processor 1407 executing appropriate code. The input / output port 1409 may be coupled to any suitable audio output, such as a multi-channel speaker system and / or headphones or the like.

[0265] In general, various embodiments of the present invention may be implemented using hardware or dedicated circuitry, software, logic, or any combination thereof. For example, some aspects may be implemented using hardware, while other aspects may be implemented using firmware or software that may be executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Although various aspects of the present invention may be illustrated and described as block diagrams, flow charts, or using some other graphical representation, it is well known that the blocks, devices, systems, techniques, or methods described herein may be implemented using hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof, as non-limiting examples.

[0266] Embodiments of the present invention may be implemented by computer software executable by a data processor of a mobile device (such as in a processor entity), or by hardware, or by a combination of software and hardware. Furthermore, in this regard, it should be noted that any block of the logic flow as in the accompanying drawings may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on a physical medium such as a memory chip or a memory block implemented within a processor, on a magnetic medium such as a hard disk or floppy disk, and on an optical medium such as a DVD and its data variant CD.

[0267] The memory may be of any type suitable for the local technical environment and may be implemented using any appropriate data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor may be of any type suitable for the local technical environment and may include, by way of non-limiting example, one or more of a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a gate-level circuit based on a multi-core processor architecture, and a processor.

[0268] Embodiments of the present invention may be practiced in various components such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available to convert a logic-level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.

[0269] The program can use well-established design rules and a library of pre-stored design blocks to route conductors and position components on a semiconductor chip. Once the design of a semiconductor circuit is complete, the resulting design in a standardized electronic format can be transmitted to a semiconductor fabrication facility, or "fab," for fabrication.

[0270] The foregoing description has provided by way of exemplary and non-limiting examples a complete and informative description of the exemplary embodiments of the present invention. However, various modifications and adaptations will become apparent to those skilled in the relevant arts in view of the foregoing description when read in conjunction with the accompanying drawings and the appended claims. Nevertheless, all such and similar modifications of the teachings of this invention will still fall within the scope of the invention as defined by the appended claims.

Claims

1. An apparatus for spatial audio coding, comprising: means for determining or receiving at least two spherical direction vectors, wherein a first spherical direction vector is associated with a first frequency band and a second spherical direction vector is associated with a second frequency band; means for determining or receiving energy in the first frequency band and energy in the second frequency band; means for combining the first spherical direction vector and the second spherical direction vector into a combined spherical direction vector, wherein the apparatus comprises: means for converting the first spherical direction vector into a first Cartesian vector and converting the second spherical direction vector into a second Cartesian vector, wherein the first Cartesian direction vector and the second Cartesian direction vector each include an x-axis component, a y-axis component, and a z-axis component; means for weighting each of the x-axis component, the y-axis component, and the z-axis component of the first Cartesian vector by the energy of the first frequency band and a direct-to-total energy ratio calculated for the first frequency band to provide a weighted x-axis component, a weighted y-axis component, and a weighted z-axis component of the first Cartesian vector; means for weighting each of the x-axis component, the y-axis component, and the z-axis component of the second Cartesian vector by the energy of the second frequency band and a direct-to-total energy ratio calculated for the second frequency band to provide a weighted x-axis component, a weighted y-axis component, and a weighted z-axis component of the second Cartesian vector; means for summing the weighted x-axis component of the first Cartesian vector and the weighted x-axis component of the second Cartesian vector to provide a combined x-axis component of a combined Cartesian vector; means for summing the weighted y-axis component of the first Cartesian vector and the weighted y-axis component of the second Cartesian vector to provide a combined y-axis component of a combined Cartesian vector; means for summing the weighted z-axis component of the first Cartesian vector and the weighted z-axis component of the second Cartesian vector to provide a combined z-axis component of a combined Cartesian vector; means for converting the combined x-axis component, the combined y-axis component and the combined z-axis component into said combined spherical direction vector; and A component for combining the direct-to-total energy ratio of the first frequency band and the direct-to-total energy ratio of the second frequency band into a combined direct-to-total energy ratio by the following operations: determining the length of the combined Cartesian vector, and normalizing the length of the combined Cartesian vector by the sum of the energy of the first frequency band and the energy of the second frequency band.

2. The device according to claim 1, wherein The apparatus further comprises means for determining whether the combined spherical direction vector is encoded for storage and / or transmission, or whether the first spherical direction vector and the second spherical direction vector are encoded for storage and / or transmission.

3. The device according to claim 2, wherein The device further comprises: means for determining a metric for said first frequency band and said second frequency band; means for comparing the metric to a threshold, wherein the means further comprising determining whether the combined spherical direction vector is encoded for storage and / or transmission, or whether the first spherical direction vector and the second spherical direction vector are encoded for storage and / or transmission comprises: means for determining that the first spherical direction vector and the second spherical direction vector are encoded for storage and / or transmission when the metric is above the threshold; and means for determining that the combined spherical direction vector is encoded for storage and / or transmission when the metric is less than or equal to the threshold.

4. The device according to any one of claims 1 to 3, wherein The device further comprises: means for determining a first extended coherence associated with the first frequency band and a second extended coherence associated with the second frequency band; and Means for merging the first extended coherence and the second extended coherence into a merged extended coherence.

5. A device according to claim 4, wherein The means for merging the first extended coherence and the second extended coherence into a merged extended coherence comprises: means for weighting said first extended coherence by said energy of said first frequency band; means for weighting said second extended coherence by said energy of said second frequency band; means for summing the weighted first extended coherence and the weighted second extended coherence to give said combined extended coherence; and means for normalizing the combined extended coherence by a sum of the energy of the first frequency band and the energy of the second frequency band.

6. The device according to claim 1, wherein The device further comprises: means for determining a first surround coherence associated with the first frequency band and a second surround coherence associated with the second frequency band; and Means for combining the first ambient coherence and the second ambient coherence into a combined ambient coherence.

7. A device according to claim 6, wherein The means for merging the first ambient coherence and the second ambient coherence into a merged ambient coherence comprises: means for weighting said first surround coherence by said energy of said first frequency band; means for weighting said second surround coherence by said energy of said second frequency band; means for summing the weighted first surround coherence and the weighted second surround coherence to give said combined extended coherence; and means for normalizing the combined surround coherence by a sum of the energy of the first frequency band and the energy of the second frequency band.

8. The device according to claim 3, wherein Said means for determining a metric comprise: means for determining a sum of the length of the first Cartesian vector and the length of the second Cartesian vector; and Means for determining a difference between a length of the combined Cartesian vector and a sum of a length of the first Cartesian vector and a length of the second Cartesian vector.

9. A method for spatial audio coding, comprising: determining at least two spherical direction vectors, wherein a first spherical direction vector is associated with a first frequency band and a second spherical direction vector is associated with a second frequency band; determining or receiving energy in the first frequency band and energy in the second frequency band; Combining the first spherical direction vector and the second spherical direction vector into a combined spherical direction vector, the method comprising: converting the first spherical direction vector into a first Cartesian vector and converting the second spherical direction vector into a second Cartesian vector, wherein the first Cartesian direction vector and the second Cartesian direction vector each include an x-axis component, a y-axis component, and a z-axis component; weighting each of the x-axis component, the y-axis component, and the z-axis component of the first Cartesian vector by the energy of the first frequency band and a direct-to-total energy ratio calculated for the first frequency band to provide a weighted x-axis component, a weighted y-axis component, and a weighted z-axis component of the first Cartesian vector; weighting each of the x-axis component, the y-axis component, and the z-axis component of the second Cartesian vector by the energy of the second frequency band and the direct-to-total energy ratio calculated for the second frequency band to provide a weighted x-axis component, a weighted y-axis component, and a weighted z-axis component of the second Cartesian vector; summing the weighted x-axis component of the first Cartesian vector and the weighted x-axis component of the second Cartesian vector to provide a combined x-axis component of a combined Cartesian vector; summing the weighted y-axis component of the first Cartesian vector and the weighted y-axis component of the second Cartesian vector to provide a combined y-axis component of a combined Cartesian vector; summing the weighted z-axis component of the first Cartesian vector and the weighted z-axis component of the second Cartesian vector to provide a combined z-axis component of a combined Cartesian vector; converting the combined x-axis component, the combined y-axis component, and the combined z-axis component into the combined spherical direction vector; and The direct-to-total energy ratio of the first frequency band and the direct-to-total energy ratio of the second frequency band are combined into a combined direct-to-total energy ratio by determining a length of the combined Cartesian vector, and normalizing the length of the combined Cartesian vector by the sum of the energy of the first frequency band and the energy of the second frequency band.

10. The method according to claim 9, wherein: The method further comprises determining whether the combined spherical direction vector is encoded for storage and / or transmission, or whether the first spherical direction vector and the second spherical direction vector are encoded for storage and / or transmission.

11. The method according to claim 10, wherein: The method further comprises: determining a metric for the first frequency band and the second frequency band; comparing the metric to a threshold, wherein the method comprises determining whether the combined spherical direction vector is encoded for storage and / or transmission, or whether the first spherical direction vector and the second spherical direction vector are encoded for storage and / or transmission, comprises: determining that the first spherical direction vector and the second spherical direction vector are encoded for storage and / or transmission when the metric is above the threshold; and Determining that the combined spherical direction vector is encoded for storage and / or transmission is determined when the metric is below or equal to the threshold.

12. The method according to any one of claims 9 to 11, wherein The method further comprises: determining a first extended coherence associated with the first frequency band and a second extended coherence associated with the second frequency band; and The first extended coherence and the second extended coherence are combined into a combined extended coherence.

13. A method according to claim 12, wherein: Combining the first extended coherence and the second extended coherence into a combined extended coherence includes: weighting the first extended coherence by the energy of the first frequency band; weighting the second extended coherence by the energy of the second frequency band; summing the weighted first extended coherence and the weighted second extended coherence to give the combined extended coherence; and The combined extended coherence is normalized by the sum of the energy of the first frequency band and the energy of the second frequency band.

14. The method according to claim 9, wherein The method further comprises: determining a first surround coherence associated with the first frequency band and a second surround coherence associated with the second frequency band; and The first ambient coherence and the second ambient coherence are combined into a combined ambient coherence.

15. A method according to claim 14, wherein Combining the first ambient coherence and the second ambient coherence into a combined ambient coherence includes: weighting the first surround coherence by the energy of the first frequency band; weighting the second surround coherence by the energy of the second frequency band; summing the weighted first surround coherence and the weighted second surround coherence to give the combined extended coherence; and The combined surround coherence is normalized by the sum of the energy of the first frequency band and the energy of the second frequency band.

16. The method according to claim 11, wherein Determining metrics includes: determining a sum of a length of the first Cartesian vector and a length of the second Cartesian vector; and The difference between the length of the combined Cartesian vector and the sum of the length of the first Cartesian vector and the length of the second Cartesian vector is determined.

Citation Information

Patent Citations

  • Spatial audio processing apparatus

    WO2017005978A1

  • Object clustering for rendering object-based audio content based on perceptual criteria

    WO2014099285A1