Combining spatial audio streams

By determining an audio scene separation metric and using it to quantize spatial audio parameters, the method addresses the inefficiencies in current spatial audio coding systems, achieving improved encoding efficiency and spatial audio quality.

JP7689196B2Active Publication Date: 2025-06-05NOKIA TECHNOLOGIES OY
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2023558512
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-03-22
Publication Date
2025-06-05
Estimated Expiration
2041-03-22

Smart Images

  • Figure 0007689196000008
    Figure 0007689196000008
  • Figure 0007689196000009
    Figure 0007689196000009
  • Figure 0007689196000010
    Figure 0007689196000010
Patent Text Reader

Abstract

In particular, an apparatus for spatial audio coding is disclosed, the apparatus being configured to determine an audio scene separation metric between an input audio signal and an additional input audio signal, and to quantize at least one spatial audio parameter of the input audio signal using the audio scene separation metric.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present application relates to an apparatus and method for sound field related parameter coding, including but not limited to time-frequency domain coding of direction related parameters for speech coders and decoders. [Background technology]

[0002] Parametric spatial audio processing is a field of audio signal processing in which the spatial aspects of sound are described using a set of parameters. For example, in parametric spatial audio capture from a microphone array, it is a typical and effective option to estimate a set of parameters from the microphone array signal, such as the direction of the sound in a frequency band and the ratio of the directional and omnidirectional parts of the captured sound in the frequency band. These parameters are known to adequately describe the perceptual spatial characteristics of the captured sound at the position of the microphone array. Therefore, these parameters can be used to synthesize spatial sound, binaurally for headphones, for loudspeakers, or for other formats such as Ambisonics.

[0003] Therefore, the direction in a frequency band and the direct-to-total energy ratio (or energy ratio parameters) are particularly useful parameterizations for spatial audio capture.

[0004] A parameter set consisting of directional parameters in frequency bands (indicating sound directionality) and energy ratio parameters in frequency bands can also be used as spatial metadata for an audio codec (which may also include other parameters such as surround coherence, spread coherence, number of directions, distance, etc.). For example, these parameters can be estimated from an audio signal captured by a microphone array, from which, for example, a stereo or mono signal can be generated that is conveyed with the spatial metadata. The stereo signal can be encoded, for example, using an AAC encoder, and the mono signal can be encoded, for example, using an EVS encoder. A decoder can decode the audio signal into a PCM signal and process the sound in the frequency bands (using the spatial metadata) to obtain a spatial output, for example a binaural output.

[0005] The above solutions are particularly suitable for encoding captured spatial sound from microphone arrays (e.g. mobile phone microphone arrays, VR camera microphone arrays, stand-alone microphone arrays). However, it may be desirable for such encoders to also have other input types than the signals captured by the microphone array, e.g. loudspeaker signals, audio object signals or Ambisonic signals.

[0006] Analyzing first-order Ambisonics (FOA) inputs for spatial metadata extraction has been extensively studied in the scientific literature on Directional Audio Coding (DirAC) and Harmonic planewave expansion (Harpex). This is because there are microphone arrays that directly provide FOA signals (or more precisely its variant, B-format signals), and therefore analyzing such inputs has been the focus of research in this field. Moreover, the analysis of higher-order Ambisonics (HOA) inputs for multi-directional spatial metadata extraction has also been studied in the scientific literature on higher-order directional audio coding (HO-DirAC).

[0007] Furthermore, additional inputs to the encoder are multi-channel loudspeaker inputs such as 5.1 or 7.1 channel surround inputs and audio objects.

[0008] The above process may involve obtaining directional parameters such as azimuth and altitude as well as energy ratios as spatial metadata through multi-channel analysis in the time-frequency domain. On the other hand, directional metadata for individual audio objects may be processed in a separate processing chain. However, possible synergies in the processing of these two types of metadata are not efficiently exploited if they are processed separately. Summary of the Invention

[0009] According to a first aspect, there is provided a method for spatial audio coding, the method comprising determining an audio scene separation metric between an input audio signal and a further input audio signal, and quantizing at least one spatial audio parameter of the input audio signal using the audio scene separation metric.

[0010] The method may further include quantizing at least one spatial audio parameter of the further input audio signal using the audio scene separation metric.

[0011] Quantizing at least one spatial audio parameter of the input audio signal using the audio scene separation metric may include multiplying the audio scene separation metric by an energy ratio parameter calculated for a time-frequency tile of the input audio signal, quantizing a product of the audio scene separation metric and the energy ratio parameter to generate a quantization index, and using the quantization index to select a bit allocation for quantizing the at least one spatial audio parameter of the input audio signal.

[0012] Alternatively, quantizing at least one spatial audio parameter of the input audio signal using the audio scene separation metric may include selecting a quantizer from among a plurality of quantizers for quantizing an energy ratio parameter calculated for a time-frequency tile of the input audio signal, the selection being dependent on the audio scene separation metric; quantizing the energy ratio parameter using the selected quantizer to generate a quantization index; and using the quantization index to select a bit allocation for quantizing the energy ratio parameter together with the at least one spatial audio parameter of the input signal.

[0013] The at least one spatial audio parameter may be a directional parameter for a time-frequency tile of the input audio signal, and the energy ratio parameter may be a directional to global energy ratio.

[0014] Quantizing at least one spatial audio parameter of the further input audio signal using the audio scene separation metric may include selecting a quantizer from among a plurality of quantizers for quantizing the at least one spatial audio parameter, the selected quantizer being dependent on the audio scene separation metric, and quantizing the at least one spatial audio parameter using the selected quantizer.

[0015] The at least one spatial audio parameter of the further input audio signal may be an audio object energy ratio parameter for a time-frequency tile of a first audio object signal of the further input audio signal.

[0016] The audio object energy ratio parameter for the time-frequency tile of a first audio object signal of the further input audio signal may be determined by determining an energy of a first audio object signal of the plurality of audio object signals for the time-frequency tile of the further input audio signal, determining an energy of each remaining audio object signal of the plurality of audio object signals, and determining a ratio of the energy of the first audio object signal to a sum of the energies of the first audio object signal and the remaining audio object signals.

[0017] An audio scene separation metric may be determined between a time-frequency tile of the input audio signal and a time-frequency tile of the additional input audio signal, and determining a quantization of at least one spatial audio parameter of the additional input audio signal using the audio scene separation metric may include determining an additional audio scene separation metric between the additional time-frequency tile of the input audio signal and the additional time-frequency tile of the additional input audio signal, determining a factor for expressing the audio scene separation metric and the additional audio scene separation metric, selecting a quantizer from among a plurality of quantizers according to the factor, and quantizing the at least one additional spatial audio parameter of the additional input audio signal using the selected quantizer.

[0018] The at least one additional spatial audio parameter may be an audio object direction parameter for an audio frame of the additional input audio signal.

[0019] The factor for expressing the sound scene separation metric and the additional sound scene separation metric may be one of an average of the sound scene separation metric and the additional sound scene separation metric, or a minimum of the sound scene separation metric and the additional sound scene separation metric.

[0020] The stream separation index may provide a measure of the relative contribution of each of the input audio signal and the additional input audio signal to an audio scene comprising the input audio signal and the additional input audio signal.

[0021] Determining the audio scene separation metric may include converting the input audio signal into a plurality of time-frequency tiles, converting the additional input audio signal into a plurality of additional time-frequency tiles, determining an energy value of at least one time-frequency tile, determining an energy value of the at least one additional time-frequency tile, and determining the audio scene separation metric as a ratio of the energy value of the at least one time-frequency tile to a sum of the at least one time-frequency tile and the at least one additional time-frequency tile.

[0022] The input audio signal may comprise two or more audio channel signals and the further input audio signal may comprise a plurality of audio object signals.

[0023] According to a second aspect, there is provided a method for spatial audio decoding, the method comprising decoding a quantized audio scene separation metric and determining at least one quantized spatial audio parameter associated with a first audio signal using the quantized audio scene separation metric.

[0024] The method may further include determining at least one quantized spatial audio parameter associated with the second audio signal using the quantized audio scene separation metric.

[0025] Determining at least one quantized spatial audio parameter associated with the first audio signal using the quantized audio scene separation metric may include selecting a quantizer from among a plurality of quantizers to use for quantizing an energy ratio parameter calculated for a time-frequency tile of the first audio signal, the selection being dependent on the decoded quantized audio scene separation metric; determining a quantized energy ratio parameter from the selected quantizer; and decoding the at least one spatial audio parameter of the first audio signal using a quantization index of the quantized energy ratio parameter.

[0026] The at least one spatial audio parameter may be a directional parameter for a time-frequency tile of the first audio signal, and the energy ratio parameter may be a directional to global energy ratio.

[0027] Determining the quantized at least one spatial audio parameter representing the second audio signal using the quantized audio scene separation metric may include selecting a quantizer from among a plurality of quantizers to use for quantizing the at least one spatial audio parameter for the second audio signal, the selection being dependent on the decoded quantized audio scene separation metric, and determining the quantized at least one spatial audio parameter for the second audio signal from the selected quantizer to use for quantizing the at least one spatial audio parameter for the second audio signal.

[0028] The at least one spatial audio parameter of the second input audio signal may be an audio object energy ratio parameter for a time-frequency tile of the first audio object signal of the second input audio signal.

[0029] The stream separation index may provide a measure of the relative contribution of each of the first and second audio signals to an audio scene comprising the first and second audio signals.

[0030] The first audio signal may include two or more audio channel signals and the second input audio signal may include a plurality of audio object signals.

[0031] According to a third aspect, there is provided an apparatus for spatial audio coding, comprising: means for determining an audio scene separation metric between an input audio signal and a further input audio signal; and means for quantizing at least one spatial audio parameter of the input audio signal using the audio scene separation metric.

[0032] The apparatus may further comprise means for quantizing at least one spatial audio parameter of the further input audio signal using the audio scene separation metric.

[0033] The means for quantizing at least one spatial audio parameter of the input audio signal using the audio scene separation metric may comprise means for multiplying the audio scene separation metric by an energy ratio parameter calculated for a time-frequency tile of the input audio signal, means for quantizing the product of the audio scene separation metric and the energy ratio parameter to generate a quantization index, and means for using the quantization index to select a bit allocation for quantizing the at least one spatial audio parameter of the input audio signal.

[0034] Alternatively, the means for quantizing at least one spatial audio parameter of the input audio signal using the audio scene separation metric may comprise means for selecting a quantizer from among a plurality of quantizers for quantizing an energy ratio parameter calculated for a time-frequency tile of the input audio signal, the selection being dependent on the audio scene separation metric; means for quantizing the energy ratio parameter using the selected quantizer to generate a quantization index; and means for selecting a bit allocation for quantizing the energy ratio parameter together with the at least one spatial audio parameter of the input signal using the quantization index.

[0035] The at least one spatial audio parameter may be a directional parameter for a time-frequency tile of the input audio signal, and the energy ratio parameter may be a directional to global energy ratio.

[0036] The means for quantizing at least one spatial audio parameter of the further input audio signal using the audio scene separation metric may comprise means for selecting a quantizer from a plurality of quantizers for quantizing the at least one spatial audio parameter, the selected quantizer depending on the audio scene separation metric, and means for quantizing the at least one spatial audio parameter using the selected quantizer.

[0037] The at least one spatial audio parameter of the further input audio signal may be an audio object energy ratio parameter for a time-frequency tile of a first audio object signal of the further input audio signal.

[0038] The audio object energy ratio parameter for the time-frequency tile of a first audio object signal of the further input audio signal may be determined by means for determining an energy of a first audio object signal of the plurality of audio object signals for the time-frequency tile of the further input audio signal, means for determining an energy of each remaining audio object signal of the plurality of audio object signals, and means for determining a ratio of the energy of the first audio object signal to a sum of the energies of the first audio object signal and the remaining audio object signals.

[0039] The sound scene separation metric may be determined between a time frequency tile of the input audio signal and a time frequency tile of the additional input audio signal, and the means for determining the quantization of at least one spatial audio parameter of the additional input audio signal using the sound scene separation metric may comprise means for determining an additional sound scene separation metric between the additional time frequency tile of the input audio signal and the additional time frequency tile of the additional input audio signal, means for determining a factor for expressing the sound scene separation metric and the additional sound scene separation metric, means for selecting a quantizer from a plurality of quantizers in response to the factor, and means for quantizing the at least one additional spatial audio parameter of the additional input audio signal using the selected quantizer.

[0040] The at least one additional spatial audio parameter may be an audio object direction parameter for an audio frame of the additional input audio signal.

[0041] The factor for expressing the sound scene separation metric and the additional sound scene separation metric may be one of an average of the sound scene separation metric and the additional sound scene separation metric, or a minimum of the sound scene separation metric and the additional sound scene separation metric.

[0042] The stream separation index may provide a measure of the relative contribution of each of the input audio signal and the additional input audio signal to an audio scene comprising the input audio signal and the additional input audio signal.

[0043] The means for determining an audio scene separation metric may comprise means for converting the input audio signal into a plurality of time-frequency tiles, means for converting the additional input audio signal into a plurality of additional time-frequency tiles, means for determining an energy value of at least one time-frequency tile, means for determining an energy value of the at least one additional time-frequency tile, and means for determining the audio scene separation metric as a ratio of the energy value of the at least one time-frequency tile to a sum of the at least one time-frequency tile and the at least one additional time-frequency tile.

[0044] The input audio signal may comprise two or more audio channel signals and the further input audio signal may comprise a plurality of audio object signals.

[0045] According to a fourth aspect, there is provided an apparatus for spatial audio decoding, comprising: means for decoding a quantized audio scene separation metric; and means for determining at least one quantized spatial audio parameter associated with a first audio signal using the quantized audio scene separation metric.

[0046] The apparatus may further comprise means for determining at least one quantized spatial audio parameter associated with the second audio signal using the quantized audio scene separation metric.

[0047] The means for determining at least one quantized spatial audio parameter associated with the first audio signal using the quantized audio scene separation metric may comprise means for selecting a quantizer from a plurality of quantizers to use for quantizing an energy ratio parameter calculated for a time-frequency tile of the first audio signal, the selection being dependent on the decoded quantized audio scene separation metric; means for determining the quantized energy ratio parameter from the selected quantizer; and means for decoding the at least one spatial audio parameter of the first audio signal using a quantization index of the quantized energy ratio parameter.

[0048] The at least one spatial audio parameter may be a directional parameter for a time-frequency tile of the first audio signal, and the energy ratio parameter may be a directional to global energy ratio.

[0049] The means for determining at least one quantized spatial audio parameter representing the second audio signal using the quantized audio scene separation metric may comprise means for selecting a quantizer from a plurality of quantizers to use for quantizing the at least one spatial audio parameter for the second audio signal, the selection being dependent on the decoded quantized audio scene separation metric, and means for determining the quantized at least one spatial audio parameter for the second audio signal from the selected quantizer to use for quantizing the at least one spatial audio parameter for the second audio signal.

[0050] The at least one spatial audio parameter of the second input audio signal may be an audio object energy ratio parameter for a time-frequency tile of the first audio object signal of the second input audio signal.

[0051] The stream separation index may provide a measure of the relative contribution of each of the first and second audio signals to an audio scene comprising the first and second audio signals.

[0052] The first audio signal may include two or more audio channel signals, and the second input audio signal includes a plurality of audio object signals.

[0053] According to a fifth aspect, there is provided an apparatus for spatial audio encoding, comprising at least one processor and at least one memory containing computer program code, the at least one memory and the computer program code configured to determine an audio scene separation metric between an input audio signal and a further input audio signal, and to quantize at least one spatial audio parameter of the input audio signal using the audio scene separation metric.

[0054] According to a sixth aspect, there is provided an apparatus for spatial audio decoding, comprising at least one processor and at least one memory containing computer program code, the at least one memory and the computer program code configured to decode a quantized audio scene separation metric and use the quantized audio scene separation metric to determine at least one quantized spatial audio parameter associated with a first audio signal.

[0055] A computer program product stored on the medium can cause an apparatus to perform the methods described herein.

[0056] The electronic device may comprise the apparatus described herein.

[0057] A chipset may include the devices described herein.

[0058] The embodiments of the present application aim to solve problems associated with the current state of the art.

[0059] For a more complete understanding of the present application, reference will now be made, by way of example only, to the accompanying drawings in which: [Brief description of the drawings]

[0060] [Figure 1] FIG. 1 illustrates a schematic diagram of a system of apparatus suitable for implementing some embodiments. [Diagram 2] FIG. 2 illustrates a schematic diagram of a metadata encoder according to some embodiments; [Diagram 3] FIG. 1 illustrates a schematic diagram of a system of apparatus suitable for implementing some embodiments. [Figure 4] FIG. 2 illustrates a schematic diagram of an exemplary device suitable for implementing the depicted apparatus; DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0061] In the following, suitable devices and possible mechanisms for providing metadata parameters derived by effective spatial analysis are described in more detail. In the following discussion, the multi-channel system is discussed in terms of a multi-channel microphone implementation. However, as discussed above, the input format can be any suitable input format, such as multi-channel loudspeaker, Ambisonic (FOA / HOA), etc. It is understood that in some embodiments, the channel positions are based on the microphone positions, or the channel positions are virtual positions or directions. Furthermore, the output of the exemplary system is a multi-channel loudspeaker device. However, it is understood that the output may be provided to the user by means other than a loudspeaker. Furthermore, the multi-channel loudspeaker signal can be generalized as being two or more reproduced audio signals. Such a system is currently being standardized by the 3GPP standardization body as Immersive Voice and Audio Service (IVAS). IVAS is intended to be an extension to the existing 3GPP Enhanced Voice Service (EVS) codec to facilitate immersive voice and audio services over existing and future mobile (cellular) and fixed-line networks. An application of IVAS may be to provide immersive voice and audio services over 3GPP fourth generation (4G) and fifth generation (5G) networks. Furthermore, the IVAS codec as an extension to EVS may be used in store-and-forward applications to encode and store audio and speech content in a file for playback. It should be understood that IVAS may be used with other audio and speech encoding techniques that have the capability of encoding samples of audio and speech signals.

[0062] Metadata-assisted spatial audio (MASA) is one input format proposed for IVAS. The MASA input format may include several (e.g., one or two) audio signals along with corresponding spatial metadata. The MASA input stream can be captured using spatial audio capture with a microphone array, e.g., a microphone array that may be mounted in a mobile device. Spatial audio parameters can then be estimated from the captured microphone signals.

[0063] The MASA spatial metadata may consist at least of the spherical direction (altitude, azimuth), at least one energy ratio of the resulting directions, the spread coherence, and the direction-independent surround coherence for each considered time-frequency (TF) block or tile, in other words the time / frequency subband. Overall, the IVAS may have several metadata parameters of different types for each time-frequency (TF) tile. The types of spatial audio parameters constituting the spatial metadata for the MASA are shown in Table 1 below.

[0064] [Table 1]

[0065] This data can be encoded and transmitted (or stored) by an encoder so that the spatial signal can be reconstructed at a decoder.

[0066] Furthermore, in some examples, Metadata Assisted Spatial Audio (MASA) may support up to two directions per TF tile, which would require the above parameters to be encoded and transmitted for each direction per TF tile, thereby nearly doubling the required bitrate according to Table 1. Furthermore, it is easy to foresee that other MASA systems may support more than two directions per TF tile.

[0067] The bitrate allocated for metadata in a practical immersive audio communication codec can vary significantly. A typical overall operating bitrate of the codec may leave only 2-10 kbps for the transmission / storage of spatial metadata. However, some additional implementations may allow up to 30 kbps or more for the transmission / storage of spatial metadata. The coding of directional parameters and energy ratio components has been previously considered along with the coding of coherence data. However, whatever the transmission / storage bitrate allocated to the spatial metadata, it is always desirable to represent these parameters using as few bits as possible, especially when the TF tiles may support a large number of directions corresponding to different sound sources in the spatial audio scene.

[0068] In addition to the multi-channel input signal that is subsequently encoded as a MASA audio signal, the encoding system may also need to encode audio objects representing various sound sources. Each audio object may be accompanied by directional data in the form of bearing and altitude values ​​that indicate the location of the audio object in physical space, whether this is in the form of metadata or some other mechanism. Typically, an audio object may have one directional parameter value per audio frame.

[0069] The idea discussed below is to improve the coding of multiple inputs to a spatial audio coding system, such as an IVAS system, where such a system is presented with a multi-channel audio signal stream as discussed above and separate input streams of audio objects. Efficiency in coding can be achieved by exploiting synergies between these separate input streams.

[0070] In this regard, Figure 1 illustrates an exemplary apparatus and system for implementing embodiments of the present application, which is shown as having an "analysis" portion 121, which is the portion from receiving the multi-channel signal to encoding metadata and a downmix signal.

[0071] The input to the "analysis" portion 121 of the system is a multi-channel signal 102. In the examples below, microphone channel signal inputs are described, but in other embodiments, any suitable input (or synthetic multi-channel) format may be implemented. For example, in some embodiments, the spatial analyzer and spatial analysis may be implemented outside the encoder. For example, in some embodiments, spatial (MASA) metadata associated with the audio signal may be provided to the encoder as a separate bitstream. In some embodiments, the spatial (MASA) metadata may be provided as a set of spatial (directional) index values.

[0072] In addition, Fig. 1 further shows a number of audio objects 128 as additional inputs to the analysis portion 121. As mentioned above, these multiple audio objects (or audio object streams) 128 may represent various sound sources in the physical space. Each audio object can be characterized by an audio (object) signal and associated metadata including directional data (in the form of azimuth and altitude values) indicating the location of the audio object in the physical space on an audio frame basis;

[0073] The multi-channel signal 102 is passed to a transport signal generator 103 and an analysis processor 105 .

[0074] In some embodiments, the transport signal generator 103 is configured to receive the multi-channel signal, generate a suitable transport signal including a determined number of channels, and output the transport signal 104 (MASA transport audio signal). For example, the transport signal generator 103 may be configured to generate a two audio channel downmix of the multi-channel signal. The determined number of channels may be any suitable number of channels. In some embodiments, the transport signal generator is configured to select or combine the input audio signals to the determined number of channels in another manner, for example by beamforming techniques, and output these signals as a transport signal.

[0075] In some embodiments, the transport signal generator 103 is optional and the multi-channel signal is passed to the encoder 107 without processing, just like the transport signal in this example.

[0076] In some embodiments, the analysis processor 105 is also configured to receive the multi-channel signals and analyze them to generate metadata 106 related to the multi-channel signals and thus to the transport signal 104. The analysis processor 105 may be configured to generate metadata that may include direction parameters 108 and energy ratio parameters 110 as well as coherence parameters 112 (and in some embodiments diffusion parameters) for each time-frequency analysis interval. In some embodiments, these direction, energy ratio and coherence parameters may be considered to be MASA spatial audio parameters (or MASA metadata). In other words, spatial audio parameters include parameters that aim to characterize the sound field generated / captured by the multi-channel signal (or two or more audio signals in general).

[0077] In some embodiments, the generated parameters may be different for each frequency band. Thus, for example, in band X, all of the parameters are generated and transmitted, while in band Y, only one of the parameters is generated and transmitted, and in band Z, no parameters are generated or transmitted. A practical example of this may be that for some frequency bands, such as the highest band, some of the parameters are not needed for perceptual reasons. The MASA transport signal 104 and the MASA metadata 106 may be passed to an encoder 107.

[0078] The audio objects 128 may be passed to an audio object analyzer 122 for processing. In other embodiments, the audio object analyzer 122 may be located within the functionality of the encoder 107.

[0079] In some embodiments, the audio object analyzer 122 analyzes the object audio input stream 128 to generate an appropriate audio object transport signal 124 and audio object metadata 126. For example, the audio object analyzer 122 can be configured to generate the audio object transport signal 124 by downmixing the audio signals of the audio objects to stereo channels with amplitude panning based on the associated audio object direction. In addition, the audio object analyzer 122 can be configured to generate the audio object metadata 126 associated with the audio object input stream 128. The audio object metadata 126 may include at least a direction parameter and an energy ratio parameter for each time-frequency analysis interval.

[0080] The encoder 107 may comprise an audio encoder core 109 configured to receive the MASA transport audio (e.g. downmix) signal 104 and the audio object transport signal 124 in order to generate appropriate encodings of these audio signals. The encoder 107 may further comprise a MASA spatial parameter set encoder 111 configured to receive the MASA metadata 106 and to output the information in encoded or compressed form as encoded MASA metadata. The encoder 107 may further comprise an audio object metadata encoder 121 similarly configured to receive the audio object metadata 126 and to output the input information in encoded or compressed form as encoded audio object metadata.

[0081] In addition, the encoder 107 may further comprise a stream separation metadata determiner and encoder 123, which may be configured to determine the relative contribution percentage of the multi-channel signal 102 (MASA audio signal) and the audio objects 128 to the overall audio scene. This percentage measure generated by the stream separation metadata determiner and encoder 123 may be used to determine the percentage of quantization and encoding "effort" spent on the input multi-channel signal 102 and the audio objects 128. In other words, the stream separation metadata determiner and encoder 123 may generate a metric that quantifies the percentage of the encoding effort spent on the MASA audio signal 102 compared to the encoding effort spent on the audio objects 128. This metric may be used to drive the encoding of the audio object metadata 126 and the MASA metadata 106. Moreover, the metric determined by the separation metadata determiner and encoder 123 may also be used as an influencing factor in the encoding process of the MASA transport audio signal 104 and the audio object transport audio signal 124 performed by the audio encoder core 109. The output metrics from the stream separation metadata determiner and encoder 123 are represented as encoded stream separation metadata, and the output metrics may be combined with the encoded metadata stream from encoder 107 .

[0082] In some embodiments, the encoder 107 may be a computer or a mobile device (executing suitable software stored in memory and on at least one processor), or alternatively, the encoder 107 may be a specific device, for example utilizing an FPGA or an ASIC. This encoding may be performed using any suitable scheme. In some embodiments, the encoder 107 may further interleave, multiplex into a single data stream, or embed into the encoded (downmixed) transport audio signal, the encoded MASA metadata, audio object metadata, and stream separation metadata, prior to transmission or storage, as indicated by the dashed lines in FIG. 1. This multiplexing may be performed using any suitable scheme. So, in summary, the system (the analysis portion) is first arranged to receive a multi-channel audio signal.

[0083] The system (analysis part) is then configured to generate a suitable transport audio signal (eg by selecting or downmixing some of the audio signal channels) and to generate spatial audio parameters as metadata.

[0084] The system is then configured to encode the transport signal and the metadata for storage / transmission.

[0085] The system can then store / transmit the encoded transport and metadata.

[0086] With respect to FIG. 2, an exemplary analysis processor 105 and metadata encoder / quantizer 111 (shown in FIG. 1) according to some embodiments will be described in more detail.

[0087] 1 and 2 show the metadata encoder / quantizer 111 and the analysis processor 105 as being coupled together. However, it should be understood that some embodiments may not very tightly couple these two corresponding respective processing entities, such that the analysis processor 105 may reside on a different device than the metadata encoder / quantizer 111. As a result, the transport signal and metadata stream may be provided to a device that includes the metadata encoder / quantizer 111 for processing and encoding independent of the capture and analysis processes.

[0088] In some embodiments, the analysis processor 105 comprises a time-to-frequency domain transformer 201 .

[0089] In some embodiments, a time-to-frequency domain transformer 201 is configured to receive the multi-channel signal 102 and apply an appropriate time-to-frequency domain transform, such as a Short Time Fourier Transform (STFT), to transform the input time-domain signal into appropriate time-frequency signals that can be passed to a spatial analyzer 203.

[0090] Thus, for example, the time-frequency signal 202 may be S MASA (b,n,i) The time-frequency domain representation can be expressed by: where b is the frequency bin index, n is the time-frequency block (frame) index, and i is the channel index. In another expression, n can be considered as a time index with a lower sampling rate than the sampling rate of the original time domain signal. These frequency bins can be grouped into subbands, which group one or more of the bins into subbands with band index k=0,....,K-1. Each subband k is divided into the lowest bin b k,low and the highest bin b k,highand the subbands are b k,low From b k,high The widths of the subbands can approximate any suitable distribution, for example the Equivalent Rectangular Bandwidth (ERB) scale or the Bark scale.

[0091] Thus, a time-frequency (TF) tile (n, k) (or block) is a particular subband k within a subframe of frame n.

[0092] It should be noted that when appended to parameters, the subscript “MASA” means that the parameters are derived from the multi-channel input signal 102, and the subscript “Obj” means that the parameters are derived from the audio object input stream 128.

[0093] It may be appreciated that the number of bits required to represent spatial audio parameters may depend, at least in part, on the TF (time-frequency) tile resolution (i.e., the number of TF subframes or tiles). For example, for a "MASA" input multi-channel audio signal, a 20 ms audio frame may be divided into four time domain subframes of 5 ms each, each of which may have up to 24 frequency subbands divided in the frequency domain according to the Bark scale, an approximation thereof, or other suitable division. In this particular example, the audio frame may be divided into 96 TF subframes / tiles, or in other words, into four time domain subframes with 24 frequency subbands. Thus, the number of bits required to represent spatial audio parameters for an audio frame may depend on the TF tile resolution. For example, if each TF tile is coded according to the distribution in Table 1 above, each TF tile would require 64 bits per sound source direction. For two sound source directions per TF tile, 2×64 bits would be required for full coding of both directions. It should be noted that the use of the term sound source may refer to the dominant direction of propagating sound within the TF tile.

[0094] In an embodiment, the analysis processor 105 may comprise a spatial analyzer 203. The spatial analyzer 203 may be configured to receive the time-frequency signals 202 and estimate the direction parameters 108 based on these signals. The direction parameters may be determined based on any audio-based "direction" determination.

[0095] For example, in some embodiments, the spatial analyzer 203 is configured to estimate the direction of a sound source using two or more signal inputs.

[0096] Thus, the spatial analyzer 203 determines an orientation Φ for each frequency band and for a transient time-frequency block within a frame of the audio signal. MASA (k,n) and altitude θ MASAThe orientation parameters 108 for the temporal subframes may be passed to a MASA spatial parameter set (metadata) set encoder 111 for encoding and quantization.

[0097] The spatial analyzer 203 may further be configured to determine an energy ratio parameter 110. This energy ratio may be considered as a determination of the energy of the audio signal that may be considered as coming from one direction. The direction to global energy ratio r MASA (k,n) (in other words, the energy ratio parameters) can be estimated, for example, using a stability measure of the direction estimation, or using any correlation measure, or using any other suitable method of obtaining the ratio parameters. Each direction-to-total energy ratio corresponds to a particular spatial direction and describes how much energy comes from a particular spatial direction compared to the total energy. This value can also be expressed separately for each time-frequency tile. The spatial direction parameters and the direction-to-total energy ratios describe how much of the total energy comes from a particular direction for each time-frequency tile. In general, the spatial direction parameters can also be thought of as directions of arrival (DOA).

[0098] In general, the directional-to-global energy ratio parameter for a multi-channel captured microphone array signal can be estimated based on the normalized cross-correlation parameter cor'(k,n) between a pair of microphones in band k, where the value of the cross-correlation parameter is between -1 and 1. The directional-to-global energy ratio parameter r(k,n) can be calculated by multiplying the normalized cross-correlation parameter by the normalized diffuse field cross-correlation parameter cor' D By comparing with (k,n),

number

[0099] For this case of a multi-channel input audio signal, the directional-to-global energy ratio parameter r MASA The (k,n) ratios can be passed to a MASA spatial parameter set (metadata) set encoder 111 for encoding and quantization.

[0100] The spatial analyzer 203 may further be configured to determine a number of coherence parameters 112 (for the multi-channel signal 102), including the surrounding coherence (γ MASA (k,n)) and spread coherence (ζ MASA (k,n)), both of which are analyzed in the time-frequency domain.

[0101] The spatial analyzer 203 calculates the spread coherence parameter ζ MASA and the surrounding coherence parameter γ MASA to a MASA spatial parameter set (metadata) set encoder 111 for encoding and quantization.

[0102] Therefore, for each TF tile there will be a set of MASA spatial audio parameters associated with each sound source direction. In this example, each TF tile may have the following audio spatial parameters associated with it for each sound source direction: orientation Φ MASA (k,n) and altitude θ MASA The direction and altitude denoted by (k,n), the spread coherence (γ MASA (k,n)), and the directional-to-global energy ratio parameter (r MASAIn addition, each TF tile also estimates the surround coherence (ζ MASA (k,n)).

[0103] In a manner similar to the processing performed by the analysis processor 105, the audio object analyzer 122 analyzes the input audio object stream to determine: S obj (b,n,i) It is possible to generate an audio object time-frequency domain signal which can be denoted as:

[0104] where b is the frequency bin index, n is the time-frequency block (TF tile) (frame) index, and i is the channel index, as previously described. The resolution of the audio object time-frequency domain signal can be the same as the corresponding MASA time-frequency domain signal, so that both signal sets are aligned in terms of time and frequency resolution. For example, the audio object time-frequency domain signal S obj (b,n,i) can have the same time resolution based on TF tile n, and the frequency bins b can be grouped into the same subband k pattern as that deployed for the MASA time-frequency domain signal. In other words, each subband k of the audio object time-frequency domain signal can also be grouped into the lowest bin b k,low and the highest bin b k,high and subband k can have b k,low From b k,highIn some embodiments, the processing of the audio object stream may not necessarily follow the same level of granularity as the processing of the MASA audio signal. For example, the MASA processing may have a different time-frequency resolution than that of the time-frequency resolution for the audio object stream. In these examples, various techniques such as parameter interpolation may be deployed to align the audio object stream processing and the MASA audio signal processing, or one parameter set may be deployed as a superset of the other.

[0105] Therefore, the resulting resolution of the time-frequency (TF) tiles for the audio object time-frequency domain signal can be the same as the resolution of the time-frequency (TF) tiles for the MASA time-frequency domain signal.

[0106] It should be noted that in FIG. 1, the audio object time-frequency domain signal may be referred to as the object transport audio signal, and the MASA time-frequency domain signal may be referred to as the MASA transport audio signal.

[0107] The audio object analyzer 122 can determine directional parameters for each audio object on an audio frame basis. The audio object directional parameters may include a heading and an altitude for each audio frame. The directional parameters include an orientation Φ obj and altitude θ obj It can be shown as:

[0108] The audio object analyzer 122 further calculates an audio object-to-total energy ratio r for each audio object signal i. obj (k,n,i) (in other words, the audio object ratio parameter). In an embodiment, the audio object to global energy ratio r obj(k,n,i) can be estimated as the ratio of the energy of object i to the energy of all audio objects.

[0109]

number

[0110] In the above formula,

number

[0111] Spatial audio parameters (metadata) associated with the audio object signal, i.e., the audio object-to-global energy ratio r for each TF tile of the audio frame for audio object i obj (k,n,i) and the direction component Φ obj and altitude θ obj To generate the audio object signal, the audio object analyzer 122 may essentially comprise similar functional processing blocks as the analysis processor 105. In other words, the audio object analyzer 122 may comprise similar processing blocks as the time domain transformer and spatial analyzer present in the analysis processor 105. The spatial audio parameters (or metadata) associated with the audio object signal may then be passed to an audio object spatial parameter set (metadata) set encoder 121 for encoding and quantization.

[0112] Audio object to total energy ratio r obj It should be understood that the (k,n,i) processing steps can be performed for each TF tile. In other words, the processing required for the directional to global energy ratio is performed for each subband k and subframe n of the speech frame, but with respect to the directional component, the orientation Φobj,i and altitude θ obj,i is obtained on an audio frame basis for audio object i.

[0113] As mentioned above, the stream separation metadata determiner and encoder 123 may be arranged to accept the MASA transport audio signal 104 and the object transport audio signal 124. The stream separation metadata determiner and encoder 123 may then use these signals to determine stream separation metrics / metadata.

[0114] In an embodiment, the stream separation metric can be found by first determining the energy of each of the MASA transport audio signal 104 and the object transport audio signal 124. This is done for each TF tile by:

number

[0115] In an embodiment, the stream separation metadata determiner and encoder 123 may then be arranged to determine a stream separation metric by calculating the ratio of MASA energy to total speech energy on a TF tile basis (total speech energy being the MASA energy combined with the speech object energy), which may be expressed as the ratio of the MASA energy in each of the MASA transport audio signals to the total energy in each of the MASA and object transport audio signals.

[0116] Therefore, this stream separation metric (or audio stream separation metric) is given by TF tile-based (k,n):

number

[0117] The stream separation metric μ(k,n) may then be quantized by the stream separation metadata determiner and encoder 123 to facilitate subsequent transmission or storage of the parameters. The stream separation metric μ(k,n) may also be referred to as the MASA-to-total energy ratio.

[0118] An example procedure for quantizing the stream separation metric μ(k,n) (for each TF tile) may include: - Arrange all MASA to total energy ratios in the speech frame as an (MxN) matrix, where M is the number of subframes in the speech frame and N is the number of subbands in the speech frame. - Transform this matrix using a 2D DCT (Discrete Cosine Transform). The optimized codebook can then be used to quantize the zeroth order DCT coefficients. The remaining DCT coefficients can be scalar quantized using the same resolution. The indices of the scalar quantized DCT coefficients can then be coded using Golomb Rice codes. - The quantized MASA to total energy ratio within the audio frame can then be formed into a suitable bitstream format by having the index of the zeroth order coefficient (at fixed rate), followed by as many GR coded indices as are allowed according to the number of bits allocated to quantize the MASA to total energy ratio. These indices can then be placed in the bitstream in a zigzag fashion according to a second diagonal direction, starting from the top left corner. The number of indices added to the bitstream is limited by the amount of available bits for coding the MASA-to-overall ratio.

[0119] The output from the stream separation metadata determiner and encoder 123 is the quantized stream separation metric μ q (k,n), which is sometimes called the quantized MASA-to-total energy ratio, which can be passed to the MASA spatial parameter set encoder 111 to drive or influence the encoding and quantization of the MASA spatial audio parameters (in other words, MASA metadata).

[0120] For a spatial audio coding system that codes the MASA audio signal alone, the quantization of the MASA spatial audio direction parameters for each TF tile is determined by the (quantized) direction-to-global energy ratio r MASA In such a system, the directional-to-global energy ratio r for that TF tile can then be determined first. MASA (k,n) can be quantized using a scalar quantizer. Then, the directional-to-global energy ratio r for that TF tile is MASA Using the index assigned to quantize (k,n), (the directional-to-total energy ratio r MASA It is possible to determine the number of bits to allocate for quantization of all MASA spatial audio parameters for that TF tile (including (k,n)).

[0121] However, the inventive spatial audio coding system is configured to code both a multi-channel audio signal (MASA audio signal) and audio objects. In such a system, the entire audio scene may be constructed as a contribution from the multi-channel audio signal and as a contribution from the audio object. As a result, the quantization of the MASA spatial audio direction parameters for a particular TF tile of interest is determined by the MASA direct-to-total energy ratio r MASA (k,n) alone, but instead depends on the MASA direction-to-total energy ratio r MASA It may depend on a combination of (k,n) and the stream separation metric μ(k,n).

[0122] In an embodiment, this coupling of dependencies is first calculated using the quantized MASA direction-to-total energy ratio r MASA Let (k,n) be the quantized stream separation metric μ for that TF tile. q (k,n) (or MASA to total energy ratio) to obtain the weighted MASA direction to total energy ratio wr MASA This can be expressed by giving (k,n). wr MASA (k,n)=μ q (k,n)*r MASA (k,n)

[0123] Then, to determine the number of bits to allocate for TF tile-based quantization of the set of MASA spatial audio parameters being transmitted to the decoder, the weighted MASA direction-to-total energy ratio wr MASA (k,n) can be quantized using a scalar quantizer, e.g., a 3-bit quantizer. For clarity, the set of MASA spatial audio parameters includes at least the directional parameter Φ MASA (k,n) and altitude θ MASA (k,n), and the directional to total energy ratio r MASA Includes (k,n).

[0124] For example, the weighted MASA direction vs. total energy wr MASA The index from the 3-bit quantizer used to quantize (k,n) can give the bit allocation from the following array: [11,11,10,9,7,6,5,3].

[0125] Then, by using some of the exemplary processes detailed in patent application publications WO 2020 / 089510, WO 2020 / 070377, WO 2020 / 008105, WO 2020 / 193865 and WO 2021 / 048468, a directional parameter Φ using bit allocations from arrays such as those described above is obtained. MASA (k,n), θ MASA (k,n), and can then proceed to encode the spread coherence and surround coherence (in other words the remaining spatial audio parameters for that TF tile).

[0126] In another embodiment, the resolution of the quantization step is determined by the MASA direction-to-total energy ratio r MASA (k,n) can be variable. For example, the MASA to total energy ratio μ q When (k,n) is low (e.g., smaller than 0.25), a low-resolution quantizer, e.g., a 1-bit quantizer, is used to obtain the MASA direction-to-total energy ratio r MASA (k,n) can be quantized. However, the MASA to total energy ratio μ q If (k,n) is higher (e.g., between 0.25 and 0.5), a higher resolution quantizer, e.g., a 2-bit quantizer, can be used. However, the MASA-to-total energy ratio μ q If (k,n) is greater than 0.5 (or some other threshold higher than the threshold for the next lower resolution quantizer), then a higher resolution quantizer, for example a 3-bit quantizer, can be used.

[0127] The output from the MASA spatial parameter set encoder 121 may then be quantization indices representing the quantized MASA directional-to-total energy ratio, the quantized MASA directional parameters, the quantized spread and the surround coherence parameters. In Figure 1, this is shown as encoded MASA metadata.

[0128] For a similar purpose, i.e. to drive or influence the coding and quantization of audio object spatial audio parameters (in other words audio object metadata), the quantized MASA-to-total energy ratio μ q It is also possible to pass (k,n) to the audio object spatial parameter set encoder 121 .

[0129] As mentioned above, the MASA to total energy ratio μ q Using (k,n), we obtain the audio object-to-global energy ratio r for audio object i. obj It is possible to influence the quantization of (k,n,i). For example, if the MASA to global energy ratio is low, a low-resolution quantizer, e.g. a 1-bit quantizer, can be used to reduce the audio object to global energy ratio r obj (k,n,i) can be quantized. However, if the MASA to total energy ratio is higher, a higher resolution quantizer can be used, e.g., a 2-bit quantizer. However, if the MASA to total energy ratio is greater than 0.5 (or some other threshold higher than the threshold for the next lower resolution quantizer), a higher resolution quantizer can be used, e.g., a 3-bit quantizer.

[0130] Furthermore, the MASA to total energy ratio μ q (k,n) can also be used to affect the quantization of the audio object direction parameters for the audio frame. Typically, this is done by first determining the MASA to global energy ratio μ FThis can be accomplished by finding an overall factor that represents F , the MASA to total energy ratio μ q Another embodiment is to use the MASA to total energy ratio μ q μ so that it becomes the average value of (k,n) F Then, the MASA to global energy ratio μ F can be used to guide the quantization of the audio object direction parameters for that frame. For example, the MASA to global energy ratio μ F If ,is high, a low-resolution quantizer can be used to quantize the audio object direction parameters, and the MASA-to-global energy ratio μ F When θ is low, a high-resolution quantizer can be used to quantize the audio object direction parameters.

[0131] The output from the speech object parameter set encoder 121 is then the quantized speech object-to-global energy ratio r obj (k,n,i), and a quantization index representing the quantized audio object direction parameters for each audio object i. In Figure 1, this is shown as the coded audio object metadata.

[0132] With regard to the audio encoder core 109, this processing block may be arranged to receive the MASA transport audio (e.g. downmix) signal 104 and the audio object transport signal 124 and combine them into a single combined audio transport signal. The combined audio transport signal may then be encoded using a suitable audio encoder. Examples of suitable audio encoders may include the 3GPP Enhanced Voice Services codec or the MPEG Advanced Audio codec.

[0133] A bitstream for storage or transmission can then be formed by multiplexing the encoded MASA metadata, the encoded stream separation metadata, the encoded audio object metadata and the encoded combined transport audio signal.

[0134] The system is capable of retrieving / receiving encoded transport and metadata.

[0135] The system is then configured to extract the transport and metadata from the encoded transport and metadata parameters, eg, demultiplex and decode the encoded transport and metadata parameters.

[0136] The system (the synthesis portion) is configured to synthesize an output multi-channel audio signal based on the extracted transport audio signal and the metadata.

[0137] In this regard, Fig. 3 illustrates an exemplary apparatus and system for implementing embodiments of the present application, which is shown as having a "synthesis" portion 331 illustrating the decoding of the encoded metadata and downmix signal for presentation of a regenerated spatial audio signal (e.g., in a multi-channel loudspeaker format).

[0138] 3, the received or extracted data (streams) may be received by a demultiplexer, which may demultiplex the encoded streams (encoded MASA metadata, encoded stream separation metadata, encoded audio object metadata and encoded transport audio signal) and pass the encoded streams to a decoder 307.

[0139] The encoded audio stream may be passed to an audio decoding core 304 configured to decode the encoded transport audio signal to obtain a decoded transport audio signal.

[0140] Similarly, the demultiplexer may be arranged to pass the encoded stream separation metadata to the stream separation metadata decoder 302. The stream separation metadata decoder 302 may then be arranged to decode the encoded stream separation metadata by doing the following: - Deindexing the zeroth order DCT coefficients. - Golomb Rice decoding the remaining DCT coefficients, provided that the number of decoded bits is within the range of the number of allowed bits. - Setting the remaining coefficients to zero. - the decoded quantized MASA to total energy ratio μ for a TF tile of an audio frame q Apply the inverse 2D DCT transform to obtain (k,n).

[0141] As shown in Fig. 3, the MASA to total energy ratio μ q (k,n) may be passed to the MASA metadata decoder 301 and the audio object metadata decoder 303 to facilitate decoding of their corresponding respective spatial audio (metadata) parameters.

[0142] The MASA metadata decoder 301 receives the encoded MASA metadata and outputs the MASA-to-total energy ratio μ q (k,n) to provide the decoded MASA spatial audio parameters. In an embodiment, this may take the following form for each audio frame:

[0143] First, we use the inverse steps of those used by the encoder to find the MASA direction-to-total energy ratio r MASA Deindex (k,n). The result of this step is the directional-to-global energy ratio r MASA (k,n).

[0144] Then, the weighted directional-to-total energy ratio wr MASA To provide the directional to global energy ratio r MASA (k,n) corresponds to the MASA to total energy ratio μ q (k,n), which is repeated for all TF tiles in the audio frame.

[0145] Then, the same optimized scalar quantizer as used in the encoder, e.g., an optimized 3-bit scalar quantizer, is used to obtain the weighted directional-to-global energy ratio wr MASA (k,n) can be scalar quantized.

[0146] As in the case of the encoder, an index from the scalar quantizer can be used to determine the number of allocated bits to be used to code the remaining MASA spatial audio parameters. For example, in the example given for the encoder, an optimized 3-bit scalar quantizer was used to determine the bit allocation for the quantization of the MASA spatial audio parameters. After the bit allocation has been determined, the remaining quantized MASA spatial audio parameters can be determined. This can be done according to at least one of the methods described in the following patent application publications: WO 2020 / 089510, WO 2020 / 070377, WO 2020 / 008105, WO 2020 / 193865 and WO 2021 / 048468.

[0147] The above steps in the MASA metadata decoder 301 are performed for every TF tile in an audio frame.

[0148] The audio object metadata decoder 301 receives the encoded audio object metadata and calculates a quantized MASA-to-total energy ratio μ q (k,n) to provide the decoded audio object spatial audio parameters. In an embodiment, this may take the following form for each audio frame:

[0149] In some embodiments, for each audio object i and TF tile (k,n) of the audio frame, an audio object-to-global energy ratio r obj (k,n,i) is the received audio object to total energy ratio r obj (k,n,i) can be de-indexed with the help of a quantizer with accurate resolution from among multiple quantizers that can be used for the purpose of decoding. As mentioned above, the audio object to global energy ratio r objThe (k,n,i) can be quantized using one of several quantizers of various resolutions. The used object-to-global energy ratio r obj The specific quantizer that quantizes (k,n,i) is the quantized MASA-to-total energy ratio for the TF tile μ q As a result, the audio object to global energy ratio r obj To select the corresponding dequantizer for (k,n,i), we use the quantized MASA-to-total energy ratio μ q (k,n) is used. In other words, the MASA to total energy ratio μ q There may be a mapping between the range of (k,n) values ​​and different inverse quantizers.

[0150] Or the whole speech frame μ F To give an overall factor that represents the MASA-to-global energy ratio for q (k,n) can also be transformed. According to a particular embodiment implemented in the encoder, μ F The derivation of is the minimum quantized MASA to total energy ratio μ q The form of selecting (k,n), or the MASA to total energy ratio μ q It can take the form of determining the average value for all of (k,n). F The value of can be used to select a particular inverse quantizer (among multiple inverse quantizers) for inverse quantizing the audio object direction parameters for an audio frame.

[0151] The output from the audio object metadata decoder 301 is then combined with the decoded quantized audio object direction parameters for the audio frame for each audio object, and the decoded quantized audio object to global energy ratio r for the TF tile of the audio frame.obj (k,n,i). In Figure 3, these parameters are shown as the decoded audio object metadata.

[0152] In some embodiments, the decoder 307 may be a computer or mobile device (executing appropriate software stored in memory and on at least one processor), or alternatively, the decoder 307 may be a specialized device, such as one that utilizes an FPGA or an ASIC.

[0153] The decoded metadata and transport audio signal may be passed to a spatial synthesis processor 305 .

[0154] A spatial synthesis processor 305 configured to receive the transport and metadata and to regenerate, based on the transport signal and the metadata, a synthesized spatial audio signal in the form of a multi-channel signal in any suitable format (which may be a multi-channel loudspeaker format or, in some embodiments, any suitable output format such as a binaural or Ambisonics signal or indeed a MASA format, depending on the use case). An example of a suitable spatial synthesis processor 305 is given in patent application publication WO 2019 / 086757.

[0155] In other embodiments, the spatial synthesis processor 305 may take different approaches to generate the multi-channel output signal. In these embodiments, the rendering may be performed in the metadata domain by combining the MASA metadata and the audio object metadata in the metadata domain. The combined metadata spatial parameters may be referred to as rendering metadata spatial parameters, and the combined metadata spatial parameters may be matched on a spatial audio direction basis. For example, having a multi-channel input signal to the encoder with one spatial audio direction identified, the rendered MASA spatial audio parameters may be set as follows: θ render (k,n,i)=θ MASA (k,n) Φ render (k,n,i)=Φ MASA (k,n) ζ render (k,n,i)=ζ MASA (k,n) r render (k,n,i)=r MASA (k,n)μ(k,n) In the above formula, i denotes the direction number. For example, for one spatial audio direction related to the input multi-channel input signal, i can take the value 1 to indicate this one MASA spatial audio direction. Furthermore, the MASA-to-global energy ratio can be used to calculate the "rendered" direction-to-global energy ratio r render (k,n,i) can be changed on a TF tile basis.

[0156] The audio object spatial audio parameters can be added to the combined metadata spatial parameters as follows: θ render (k,n,i obj +1)=θ obj (n,i obj ) Φ render (k,n,i obj +1)=Φ obj (n,i obj ) ζ render (k,n,i obj +1)=0 r render (k,n,i obj +1)=r obj (1-μ(k,n)) In the above formula, i obj is the audio object number. In this example, the audio objects are determined to have no spread coherence ζ. Finally, the MASA to global energy ratio (μ) is used to modify the spread to global energy ratio (ψ), and the surround coherence (γ) is set directly. ψ render (k,n)=ψ MASA (k,n)μ(k,n) Gamma render (k,n)=γ MASA (k,n)

[0157] 4, an exemplary electronic device that can be used as an analysis or synthesis device is shown. The device can be any suitable electronic device or apparatus. For example, in some embodiments, device 1400 is a mobile device, user equipment, tablet computer, computer, audio playback device, etc.

[0158] In some embodiments, device 1400 includes at least one processor or central processing unit 1407. Processor 1407 may be configured to execute various program code, such as, for example, the methods described herein.

[0159] In some embodiments, the device 1400 comprises a memory 1411. In some embodiments, the memory 1411 is coupled to at least one processor 1407. The memory 1411 can be any suitable storage means. In some embodiments, the memory 1411 comprises a program code section for storing program code executable on the processor 1407. Moreover, in some embodiments, the memory 1411 can further comprise a storage data section for storing data, for example data processed or to be processed according to the embodiments described herein. The executed program code stored in the program code section and the data stored in the storage data section can be retrieved by the processor 1407 via the memory-processor coupling whenever required.

[0160] In some embodiments, the device 1400 comprises a user interface 1405. In some embodiments, the user interface 1405 can be coupled to a processor 1407. In some embodiments, the processor 1407 can control the operation of the user interface 1405 and receive input from the user interface 1405. In some embodiments, the user interface 1405 can enable a user to input commands into the device 1400, for example, via a keypad. In some embodiments, the user interface 1405 can enable a user to obtain information from the device 1400. For example, the user interface 1405 can comprise a display configured to display information from the device 1400 to the user. In some embodiments, the user interface 1405 can comprise a touch screen or touch interface that can both enable information to be input into the device 1400 and also display information to the user of the device 1400. In some embodiments, the user interface 1405 can be a user interface for communicating with a position determiner as described herein.

[0161] In some embodiments, the device 1400 comprises an input / output port 1409. In some embodiments, the input / output port 1409 comprises a transceiver. In such embodiments, the transceiver may be coupled to the processor 1407 and may be configured to allow communication with other apparatuses or electronic devices, for example via a wireless communication network. In some embodiments, the transceiver, or any suitable transceiver or transmitting and / or receiving means, may be configured to communicate with other electronic devices or apparatuses via a conductor or wired coupling.

[0162] The transceiver may communicate with the additional device by any suitable known communication protocol, for example, in some embodiments, the transceiver may use a suitable universal mobile telecommunications system (UMTS) protocol, a wireless local area network (WLAN) protocol such as IEEE 802.X, or a suitable short-range radio frequency communication protocol such as Bluetooth or an infrared data communication pathway (IRDA).

[0163] The transceiver input / output port 1409 can be configured to receive signals and, in some embodiments, determine the parameters described herein by using the processor 1407 executing appropriate code, which can further generate an appropriate downmix signal and parameter output for transmission to a combining device.

[0164] In some embodiments, device 1400 can be used as at least a portion of a synthesis device. As such, input / output port 1409 can be configured to receive a downmix signal and, in some embodiments, parameters determined by a capture device or processing device described herein, and generate an appropriate audio signal format output by using processor 1407 executing appropriate code. Input / output port 1409 can be coupled to any appropriate audio output, such as a multi-channel speaker system and / or headphones, or similar device.

[0165] In general, various embodiments of the present invention can be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. For example, some aspects can be implemented in hardware, and other aspects can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. However, the present invention is not limited thereto. Although various aspects of the present invention may be illustrated or described as block diagrams or flow diagrams, or using some other pictorial representation, it is fully understood that these blocks, devices, systems, techniques, or methods described herein can be implemented, as non-limiting examples, in hardware, software, firmware, dedicated circuits or logic, general-purpose hardware or controller, or other computing device, or some combination thereof.

[0166] The embodiments of the present invention may be implemented by computer software executable by a data processor of a mobile device, for example within that processor entity, or by hardware, or by a combination of software and hardware. Moreover, in this regard, it should be noted that any blocks of the logical flow of the diagrams may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on a physical medium, such as a memory chip, or a memory block embodied within a processor, a magnetic medium, such as a hard disk or floppy disk, and an optical medium, such as, for example, DVDs and their data variants, CDs, etc.

[0167] The memory may be any type of memory suitable for the local technology environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed and removable memories, etc. The data processor may be any type of data processor suitable for the local technology environment and may include, by way of non-limiting examples, one or more of a general purpose computer, a special purpose computer, a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), gate level circuits and processors based on multi-core processor architectures.

[0168] Embodiments of the present invention can be implemented in a variety of components, such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available for converting logic level designs into semiconductor circuit designs ready to be etched and formed on a semiconductor substrate.

[0169] The program can use well-established design rules and a library of pre-stored design modules to route conductors and place components on a semiconductor chip. After the design of a semiconductor circuit is completed, the resulting design can be sent in a standardized electronic format to a semiconductor manufacturing facility or "fab" for manufacturing.

[0170] The above description provides an informative and thorough description of exemplary embodiments of the present invention as illustrative and non-limiting examples. However, various modifications and adaptations in view of the above description may become apparent to those skilled in the art when read in conjunction with the accompanying drawings and the appended claims. Nevertheless, all such modifications and similar modifications of the teachings of the present invention are within the scope of the present invention as defined in the appended claims.

Claims

1. 1. A method for spatial audio signal coding, comprising: determining an audio scene separation metric between an input audio signal comprising two or more audio channel signals and a further input audio signal comprising a plurality of audio object signals, Transforming the input audio signal into a number of time-frequency tiles; transforming the further input audio signal into a plurality of further time-frequency tiles; determining an energy value for at least one time-frequency tile; determining an energy value of at least one additional time-frequency tile; and determining the audio scene separation metric as a ratio of the energy value of the at least one time-frequency tile to a sum of the at least one time-frequency tile and the at least one additional time-frequency tile; determining said audio scene separation metric comprising: quantizing at least one spatial audio parameter of the input audio signal using the audio scene separation metric; The method includes:

2. quantizing at least one spatial audio parameter of the further input audio signal using the audio scene separation metric. The method of claim 1 further comprising:

3. quantizing the at least one spatial audio parameter of the input audio signal using the audio scene separation metric; multiplying said audio scene separation metric by an energy ratio parameter calculated for a time-frequency tile of said input audio signal; quantizing a product of the audio scene separation metric and the energy ratio parameter to generate a quantization index; and using the quantization index to select a bit allocation for quantizing the at least one spatial audio parameter of the input audio signal. The method of claim 1 or 2, comprising:

4. quantizing the at least one spatial audio parameter of the input audio signal using the audio scene separation metric; selecting a quantizer from among a plurality of quantizers for quantizing an energy ratio parameter calculated for a time-frequency tile of the input audio signal, said selection being dependent on the audio scene separation metric; quantizing the energy ratio parameter using the selected quantizer to generate a quantization index; and using said quantization index to select a bit allocation for quantizing said energy ratio parameter together with said at least one spatial audio parameter of said input audio signal. The method of claim 1 or 2, comprising:

5. The method of claim 3 or 4, wherein the at least one spatial audio parameter is a directional parameter for the time-frequency tile of the input audio signal and the energy ratio parameter is a directional to global energy ratio.

6. quantizing the at least one spatial audio parameter of the further input audio signal using the audio scene separation metric; selecting a quantizer for quantizing the at least one spatial audio parameter from among a plurality of quantizers, the quantizer selected being dependent on the audio scene separation metric; and quantizing the at least one spatial audio parameter using the selected quantizer. The method according to any one of claims 2 to 5, comprising:

7. The method of claim 6 , wherein the at least one spatial audio parameter of the further input audio signal is an audio object energy ratio parameter for a time-frequency tile of a first audio object signal of the further input audio signal.

8. The audio object energy ratio parameter for the time-frequency tile of the first audio object signal of the further input audio signal is determining an energy of the first audio object signal of a plurality of audio object signals for the time-frequency tile of the further input audio signal; determining an energy of each remaining one of the plurality of audio object signals; and determining a ratio of the energy of the first audio object signal to a sum of the energies of the first audio object signal and the remaining audio object signals; The method of claim 7, wherein the value is determined by

9. the audio scene separation metric is determined between time-frequency tiles of the input audio signal and time-frequency tiles of the further input audio signal, and the audio scene separation metric is used to determine the quantization of at least one spatial audio parameter of the further input audio signal; determining additional audio scene separation metrics between additional time-frequency tiles of the input audio signal and additional time-frequency tiles of the additional input audio signal; determining factors for representing said audio scene separation metric and said further audio scene separation metric; selecting a quantizer from among a plurality of quantizers in response to said factor; and quantizing at least one additional spatial audio parameter of the additional input audio signal using the selected quantizer; The method according to any one of claims 2 to 8, comprising:

10. The method of claim 9 , wherein the at least one additional spatial audio parameter is an audio object direction parameter for an audio frame of the additional input audio signal.

11. The factors for representing the audio scene separation metric and the additional audio scene separation metric are: an average of the audio scene separation metric and the additional audio scene separation metric; or the minimum of the audio scene separation metric and the additional audio scene separation metric The method according to claim 9 or 10, wherein the

12. 12. The method of claim 1, wherein the audio scene separation metric provides a measure of the relative contribution of each of the input audio signal and the further input audio signal to an audio scene comprising the input audio signal and the further input audio signal.

13. An apparatus for spatial audio signal coding, comprising:

1. A method for determining an audio scene separation metric between an input audio signal comprising two or more audio channel signals and a further input audio signal comprising a plurality of audio object signals, the method comprising: means for converting the input audio signal into a number of time-frequency tiles; means for converting the additional input audio signal into a plurality of additional time-frequency tiles; means for determining an energy value of at least one time-frequency tile; means for determining an energy value of at least one additional time-frequency tile; means for determining the audio scene separation metric as a ratio of the energy value of the at least one time-frequency tile to a sum of the at least one time-frequency tile and the at least one additional time-frequency tile; means for determining said audio scene separation metric, said means comprising: means for quantizing at least one spatial audio parameter of the input audio signal using the audio scene separation metric; An apparatus comprising:

14. means for quantizing at least one spatial audio parameter of said further input audio signal using said audio scene separation metric; The apparatus of claim 13 further comprising:

15. said means for quantizing said at least one spatial audio parameter of said input audio signal using said audio scene separation metric, means for multiplying said audio scene separation metric by an energy ratio parameter calculated for a time-frequency tile of said input audio signal; means for quantizing a product of said audio scene separation metric and said energy ratio parameter to generate a quantization index; means for selecting a bit allocation for quantizing the at least one spatial audio parameter of the input audio signal using the quantization index; 15. The apparatus according to claim 13 or 14, comprising:

16. said means for quantizing said at least one spatial audio parameter of said input audio signal using said audio scene separation metric, means for selecting a quantizer from among a plurality of quantizers for quantizing an energy ratio parameter calculated for a time-frequency tile of the input audio signal, said selection being dependent on the audio scene separation metric; means for quantizing the energy ratio parameter using the selected quantizer to generate a quantization index; means for selecting a bit allocation for quantizing the energy ratio parameter together with the at least one spatial audio parameter of the input audio signal using the quantization index; 15. The apparatus according to claim 13 or 14, comprising:

17. 17. Apparatus according to claim 15 or 16, wherein said at least one spatial audio parameter is a directional parameter for said time-frequency tile of said input audio signal and said energy ratio parameter is a directional to global energy ratio.

18. said means for quantizing said at least one spatial audio parameter of said further input audio signal using said audio scene separation metric, means for selecting a quantizer from among a plurality of quantizers for quantizing the at least one spatial audio parameter, the quantizer selected being dependent on the audio scene separation metric; means for quantizing the at least one spatial audio parameter using the selected quantizer; The apparatus according to any one of claims 14 to 17, comprising:

19. The apparatus of claim 18 , wherein the at least one spatial audio parameter of the further input audio signal is an audio object energy ratio parameter for a time-frequency tile of a first audio object signal of the further input audio signal.

20. The audio object energy ratio parameter for the time-frequency tile of the first audio object signal of the further input audio signal is means for determining an energy of the first audio object signal of a plurality of audio object signals for the time-frequency tile of the further input audio signal; means for determining an energy of each remaining one of the plurality of audio object signals; means for determining a ratio of the energy of the first audio object signal to a sum of the energies of the first audio object signal and the remaining audio object signals; 20. The apparatus of claim 19, wherein the determination is made by:

21. The audio scene separation metric is determined between a time-frequency tile of the input audio signal and a time-frequency tile of the further input audio signal, and the means for determining the quantization of at least one spatial audio parameter of the further input audio signal using the audio scene separation metric comprises: means for determining additional audio scene separation metrics between additional time-frequency tiles of the input audio signal and additional time-frequency tiles of the additional input audio signal; means for determining factors for representing said audio scene separation metric and said further audio scene separation metric; means for selecting a quantizer from among a plurality of quantizers in response to said factor; means for quantizing at least one additional spatial audio parameter of the additional input audio signal using the selected quantizer; The apparatus according to any one of claims 14 to 20, comprising:

22. The apparatus of claim 21 , wherein the at least one additional spatial audio parameter is an audio object direction parameter for an audio frame of the additional input audio signal.

23. The factors for representing the audio scene separation metric and the additional audio scene separation metric are: an average of the audio scene separation metric and the additional audio scene separation metric; or the minimum of the audio scene separation metric and the additional audio scene separation metric 23. The device according to claim 21 or 22, wherein

24. 24. The apparatus of claim 13, wherein a stream separation index provides a measure of a relative contribution of each of the input audio signal and the further input audio signal to an audio scene comprising the input audio signal and the further input audio signal.

Citation Information

Patent Citations

  • Apparatus and method for encoding or decoding directional audio coding parameters using different time / frequency resolutions

    JP2021503627A

  • Segmentation-based feature extraction for acoustic scene classification

    US20200265864A1

  • Audio scene encoder, audio scene decoder and related methods using hybrid encoder-decoder spatial analysis

    US20200357421A1

  • Method and system for semantically segmenting an audio sequence

    WO2005093712A1

  • Audio coding

    WO2019170955A1