Combination of spatial audio parameters

By combining the spatial audio parameters of frequency subbands and employing Cartesian vector transformation technology, the problem of high bit rate in the encoding of multi-channel loudspeaker and audio object signals is solved, achieving more efficient spatial audio parameter encoding and adapting to multi-source directional scenarios.

CN114846542BActive Publication Date: 2025-10-31NOKIA TECHNOLOGIES OY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080089404.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-23
Filing Date
2020-11-13
Publication Date
2025-10-31
Estimated Expiration
2040-11-13

AI Technical Summary

Technical Problem

Existing technologies suffer from high bit rate requirements when encoding spatial audio parameters, especially for multi-channel loudspeaker and audio object signals, making it difficult to effectively compress and encode spatial audio parameters to reduce the total number of bits required for representation.

Method used

By combining the first and second spatial audio parameters of the frequency sub-bands, combined spatial audio parameters are generated, and the encoding method is determined by comparing the metric with the threshold. This reduces unnecessary parameter transmission, and Cartesian vector transformation and combination techniques are used to reduce bit rate requirements.

Benefits of technology

It effectively reduces the bit rate per TF block, improves the coding efficiency of spatial audio parameters, adapts to multi-source directional scenarios, and reduces the number of bits required for transmission and storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114846542B_ABST
    Figure CN114846542B_ABST
Patent Text Reader

Abstract

In particular, an apparatus for spatial audio coding is disclosed, comprising: means for determining a first spatial audio parameter of a frequency sub-band of one or more audio signals and a second spatial audio parameter of the frequency sub-band of the one or more audio signals; and means for combining the first spatial audio parameter and the second spatial audio parameter to provide combined spatial audio parameters for the frequency sub-band.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to apparatus and methods for sound field-related parameter encoding, but is not limited to time-frequency domain orientation-related parameter encoding of audio encoders and decoders. Background Technology

[0002] Parametric spatial audio processing is a field of audio signal processing that uses a set of parameters to describe the spatial aspects of sound. For example, in parametric spatial audio capture from a microphone array, estimating a set of parameters from the microphone array signal is a typical and effective choice. These parameters include, for instance, the direction of sound in a frequency band and the ratio of directional to non-directional portions of the captured sound within that band. These parameters are well-known for describing the perceived spatial characteristics of the captured sound at the location of the microphone array. These parameters can then be used accordingly in the synthesis of spatial sound for use with binaural headphones, loudspeakers, or other formats such as Ambisonics.

[0003] Therefore, the direction and direct-to-total energy ratio in the frequency band are particularly effective parameters for spatial audio capture.

[0004] A set of parameters, including directional parameters and energy ratio parameters (indicating the directionality of sound) within a frequency band, can also be used as spatial metadata for audio codecs (which may also include other parameters such as surround coherence, spread coherence, number of directions, distance, etc.). For example, these parameters can be estimated from audio signals captured by a microphone array, and stereo or mono signals can be generated from the microphone array signal, for example, to be transmitted along with the spatial metadata. Stereo signals can be encoded, for example, using an AAC encoder, while mono signals can be encoded using an EVS encoder. The decoder can decode the audio signal into a PCM signal and (using the spatial metadata) process the sound within the frequency band to obtain a spatial output, such as a binaural output.

[0005] The aforementioned solution is particularly well-suited for encoding spatial sound captured from microphone arrays (e.g., mobile phones, VR cameras, individual microphone arrays). However, it is expected that such an encoder will accept other input types besides the signals captured by the microphone array, such as speaker signals, audio object signals, or ambisonic signals.

[0006] The analysis of first-order Ambisonics (FOA) inputs for spatial metadata extraction has been extensively documented in the scientific literature related to Directional Audio Coding (DirAC) and Harpex plane wave expansion. This is because microphone arrays directly provide FOA signals (more precisely, their variants, B-format signals), and therefore, analyzing such inputs has become a focus of research in this field. Furthermore, analyses of higher-order Ambisonics (HOA) inputs for multi-directional spatial metadata extraction have also been documented in the scientific literature related to Higher-order Directional Audio Coding (HO DirAC).

[0007] Another input for the encoder can also be a multi-channel speaker input, such as a 5.1 or 7.1 channel surround sound input and an audio object.

[0008] However, regarding the components of spatial metadata, the compression and encoding of spatial audio parameters are quite important in order to minimize the total number of bits required to represent the spatial audio parameters. Summary of the Invention

[0009] According to a first aspect, an apparatus for spatial audio coding is provided, comprising: means for determining first spatial audio parameters of a frequency sub-band of one or more audio signals and second spatial audio parameters of the frequency sub-band of the one or more audio signals; and means for combining the first spatial audio parameters and the second spatial audio parameters to provide combined spatial audio parameters for the frequency sub-band.

[0010] The apparatus may further include: a component for determining whether combined spatial audio parameters for a frequency subband are encoded for storage and / or transmission, or whether a first spatial audio parameter of the frequency subband and a second spatial audio parameter of the frequency subband are encoded for storage and / or transmission.

[0011] The apparatus may further include: means for determining a metric for a frequency sub-band of one or more audio signals; means for comparing the metric with a threshold, wherein the apparatus further includes means for determining whether combined spatial audio parameters for the frequency sub-band are encoded for storage and / or transmission, or whether a first spatial audio parameter and a second spatial audio parameter of the frequency sub-band are encoded for storage and / or transmission. The apparatus may also include: means for determining that, when the metric is above a threshold, the first spatial audio parameter and the second spatial audio parameter of the frequency sub-band are encoded for storage and / or transmission; and means for determining that, when the metric is below or equal to a threshold, the combined spatial audio parameters for the frequency sub-band are encoded for storage and / or transmission.

[0012] The apparatus may further include: components for determining a metric for a frequency sub-band of one or more audio signals; components for determining a first spatial audio parameter and a second spatial audio parameter for at least one other frequency sub-band of the one or more audio signals; components for combining the first spatial audio parameter and the second spatial audio parameter of the at least one other frequency sub-band of the one or more audio signals to provide combined spatial audio parameters for the other frequency sub-band of the one or more audio signals; components for determining another metric for the at least one other frequency sub-band; and components for determining that when the metric is higher than the other metric, the first spatial audio parameter and the second spatial audio parameter of the frequency sub-band of the one or more audio signals are encoded for storage and / or transmission, and the combined spatial audio parameters for the at least one other frequency sub-band of the one or more audio signals are encoded for storage and / or transmission.

[0013] The first spatial audio parameter can be a first spherical direction vector calculated for a frequency sub-band, the first spherical direction vector including azimuth and elevation components, wherein the second spatial audio parameter can be a second spherical direction vector calculated for the same frequency sub-band, the second spherical direction vector including azimuth and elevation components, and wherein the combined spatial audio parameter can be a combined spherical direction vector.

[0014] The components for combining the first and second spatial audio parameters may include: components for converting a first spherical direction vector into a first Cartesian vector, and components for converting a second spherical direction vector into a second Cartesian vector, wherein the first and second Cartesian vectors each include an x-axis component, a y-axis component, and a z-axis component, wherein, for each corresponding component, the apparatus may include: components for weighting the corresponding component of the first Cartesian vector to calculate a first direct-to-total energy ratio for a frequency sub-band; components for weighting the corresponding component of the second Cartesian vector to calculate a second direct-to-total energy ratio for the frequency sub-band; and components for summing the corresponding weighted components of the first and second Cartesian vectors to give a corresponding combined Cartesian component, wherein the combined x-axis Cartesian component, the combined y-axis Cartesian component, and the combined z-axis Cartesian component form the components of the combined Cartesian vector; and components for converting the combined x-axis Cartesian component, the combined y-axis Cartesian component, and the combined z-axis Cartesian component into a combined spherical direction vector.

[0015] The device may further include: a component for determining an ambient energy value for a frequency sub-band by subtracting a first direct-to-total energy ratio calculated for the frequency sub-band and a second direct-to-total energy ratio calculated for the frequency sub-band from 1.

[0016] The device may further include a component for combining a first direct-to-total power ratio calculated for a frequency sub-band and a second direct-to-total power ratio calculated for the frequency sub-band to provide a combined direct-to-total power ratio for the frequency sub-band.

[0017] A component for combining a first direct-to-total energy ratio calculated for a frequency sub-band and a second direct-to-total energy ratio calculated for that frequency sub-band to provide a combined direct-to-total energy ratio for that frequency sub-band may include a component for determining the combined direct-to-total energy ratio based on the ratio of the vector length of the combined Cartesian vector to the sum of the first direct-to-total energy ratio calculated for that frequency sub-band and the second direct-to-total energy ratio calculated for that frequency sub-band, and an ambient energy value.

[0018] The device may further include a component for combining a first extended coherence value calculated for a frequency subband and a second extended coherence value calculated for the frequency subband to provide a combined extended coherence parameter for the frequency subband.

[0019] The components for combining a first extended coherence value calculated for a frequency sub-band and a second extended coherence value calculated for the frequency sub-band to provide a combined extended coherence parameter for the frequency sub-band may include: components for determining a first sum, wherein the first sum includes the product of the first extended coherence value calculated for the frequency sub-band and a first direct-to-total power ratio calculated for the frequency sub-band, and the product of the second extended coherence value calculated for the frequency sub-band and a second direct-to-total power ratio calculated for the frequency sub-band; components for determining a second sum, wherein the second sum includes the first direct-to-total power ratio calculated for the frequency sub-band and a second direct-to-total power ratio calculated for the frequency sub-band; and components for determining a ratio of the first sum to the second sum to provide the combined extended coherence parameter.

[0020] The apparatus for spatial audio coding may further include: means for calculating a surround coherence value for a frequency sub-band; means for determining another ambient energy value for the frequency sub-band by subtracting a combined direct-to-total energy ratio from 1; means for determining a surround coherence energy by determining a product of a combined extended coherence parameter and the difference between the other ambient energy value for the frequency sub-band and the ambient energy value for the frequency sub-band; and means for adding the surround coherence energy to the product of the ambient energy for the frequency sub-band and the surround coherence value for the frequency sub-band, and normalizing it to another ambient energy value for the frequency sub-band to provide a combined surround coherence value.

[0021] The apparatus, including components for determining the metric, may include components for determining the difference between the sum of a first direct-to-total power ratio calculated for a frequency subband and a second direct-to-total power ratio calculated for the frequency subband and the length of the combined Cartesian vector.

[0022] The first spatial audio parameter can be associated with the direction of the first sound source in the frequency sub-band, and the second spatial audio parameter can be associated with the direction of the second sound source in the same frequency sub-band.

[0023] According to the second aspect, there is a method for spatial audio coding, comprising: determining a first spatial audio parameter of a frequency sub-band of one or more audio signals and a second spatial audio parameter of the frequency sub-band of the one or more audio signals; and combining the first spatial audio parameter and the second spatial audio parameter to provide combined spatial audio parameters for the frequency sub-band.

[0024] The method may further include: determining whether the combined spatial audio parameters for a frequency subband are encoded for storage and / or transmission, or whether a first spatial audio parameter of the frequency subband and a second spatial audio parameter of the frequency subband are encoded for storage and / or transmission.

[0025] The method may further include: determining a metric for a frequency sub-band of one or more audio signals; comparing the metric with a threshold, wherein the apparatus further includes components for determining whether combined spatial audio parameters for the frequency sub-band are encoded for storage and / or transmission, or whether a first spatial audio parameter and a second spatial audio parameter of the frequency sub-band are encoded for storage and / or transmission, and may include: determining that when the metric is higher than the threshold, the first spatial audio parameter and the second spatial audio parameter of the frequency sub-band are encoded for storage and / or transmission; and determining that when the metric is lower than or equal to the threshold, the combined spatial audio parameters for the frequency sub-band are encoded for storage and / or transmission.

[0026] The method may further include: determining a metric for a frequency sub-band of one or more audio signals; determining a first spatial audio parameter and a second spatial audio parameter for at least one other frequency sub-band of the one or more audio signals; combining the first spatial audio parameter and the second spatial audio parameter of the at least one other frequency sub-band of the one or more audio signals to provide a combined spatial audio parameter for the other frequency sub-band of the one or more audio signals; determining another metric for the at least one other frequency sub-band; and determining that when the metric is higher than the other metric, the first spatial audio parameter and the second spatial audio parameter of the frequency sub-band of the one or more audio signals are encoded for storage and / or transmission, and the combined spatial audio parameter for the at least one other frequency sub-band of the one or more audio signals is encoded for storage and / or transmission.

[0027] The first spatial audio parameter can be a first spherical direction vector calculated for a frequency sub-band, the first spherical direction vector including azimuth and elevation components, wherein the second spatial audio parameter can be a second spherical direction vector calculated for the same frequency sub-band, the second spherical direction vector including azimuth and elevation components, and wherein the combined spatial audio parameter can be a combined spherical direction vector.

[0028] Combining the first spatial audio parameters and the second spatial audio parameters may include: converting a first spherical direction vector into a first Cartesian vector, and converting a second spherical direction vector into a second Cartesian vector, wherein the first Cartesian vector and the second Cartesian vector each include an x-axis component, a y-axis component, and a z-axis component, wherein, for each corresponding component, the method may include: weighting the corresponding component of the first Cartesian vector to calculate a first direct-to-total energy ratio for a frequency sub-band; weighting the corresponding component of the second Cartesian vector to calculate a second direct-to-total energy ratio for that frequency sub-band; and summing the corresponding weighted components of the first Cartesian vector and the corresponding weighted components of the second Cartesian vector to give a corresponding combined Cartesian component, wherein the combined x-axis Cartesian component, the combined y-axis Cartesian component, and the combined z-axis Cartesian component form the components of the combined Cartesian vector; and converting the combined x-axis Cartesian component, the combined y-axis Cartesian component, and the combined z-axis Cartesian component into a combined spherical direction vector.

[0029] The method may further include: determining the ambient energy value for a frequency sub-band by subtracting a first direct-to-total energy ratio calculated for the frequency sub-band and a second direct-to-total energy ratio calculated for the frequency sub-band from 1.

[0030] The method may further include: combining a first direct-to-total power ratio calculated for a frequency sub-band and a second direct-to-total power ratio calculated for the frequency sub-band to provide a combined direct-to-total power ratio for the frequency sub-band.

[0031] Combining a first direct-to-total energy ratio calculated for a frequency sub-band and a second direct-to-total energy ratio calculated for that frequency sub-band to provide a combined direct-to-total energy ratio for that frequency sub-band may include: determining the combined direct-to-total energy ratio based on the ratio of the vector length of the combined Cartesian vector to the sum of the first direct-to-total energy ratio calculated for that frequency sub-band and the second direct-to-total energy ratio calculated for that frequency sub-band, and an ambient energy value.

[0032] The method may further include: combining a first extended coherence value calculated for a frequency sub-band and a second extended coherence value calculated for the frequency sub-band to provide a combined extended coherence parameter for the frequency sub-band.

[0033] Combining a first extended coherence value calculated for a frequency sub-band and a second extended coherence value calculated for the same frequency sub-band to provide a combined extended coherence parameter for that frequency sub-band may include: determining a first sum, the first sum including the product of the first extended coherence value calculated for the frequency sub-band and a first direct-to-total power ratio calculated for the same frequency sub-band, and the product of the second extended coherence value calculated for the frequency sub-band and a second direct-to-total power ratio calculated for the same frequency sub-band; determining a second sum, the second sum including the first direct-to-total power ratio calculated for the same frequency sub-band and the second direct-to-total power ratio calculated for the same frequency sub-band; and determining a ratio of the first sum to the second sum to provide the combined extended coherence parameter.

[0034] The method for spatial audio coding may further include: calculating a surround coherence value for a frequency sub-band; determining another ambient energy value for the frequency sub-band by subtracting the combined direct-to-total energy ratio from 1; determining a surround coherence energy by determining the product of a combined extended coherence parameter and the difference between the other ambient energy value for the frequency sub-band and the ambient energy value for the frequency sub-band; and adding the surround coherence energy to the product of the ambient energy for the frequency sub-band and the surround coherence value for the frequency sub-band, and normalizing it to the other ambient energy value for the frequency sub-band to provide the combined surround coherence value.

[0035] The determination of the metric may include: determining the difference between the sum of a first direct-to-total power ratio calculated for a frequency subband and a second direct-to-total power ratio calculated for that frequency subband and the length of the combined Cartesian vector.

[0036] The first spatial audio parameter can be associated with the direction of the first sound source in the frequency sub-band, and the second spatial audio parameter can be associated with the direction of the second sound source in the same frequency sub-band.

[0037] According to a third aspect, there is an apparatus for spatial audio coding, comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured, together with the at least one processor, to cause the apparatus to at least: determine a first spatial audio parameter of a frequency sub-band of one or more audio signals and a second spatial audio parameter of the frequency sub-band of the one or more audio signals; and combine the first spatial audio parameter and the second spatial audio parameter to provide combined spatial audio parameters for the frequency sub-band.

[0038] A computer program product stored on a medium can enable a device to perform the methods described herein.

[0039] An electronic device may include the apparatus described herein.

[0040] A chipset may include the devices described herein.

[0041] The embodiments of this application are intended to solve problems associated with the prior art. Attached Figure Description

[0042] To better understand this application, reference will now be made to the accompanying drawings by way of example, wherein:

[0043] Figure 1 An apparatus system suitable for implementing some embodiments is schematically shown;

[0044] Figure 2 A metadata encoder according to some embodiments is illustrated schematically;

[0045] Figure 3 The following are examples illustrating some embodiments: Figure 2 The flowchart shown illustrates the operation of the metadata encoder;

[0046] Figure 4 An example device suitable for implementing the illustrated apparatus is shown schematically. Detailed Implementation

[0047] The following describes in more detail suitable apparatuses and possible mechanisms for providing metadata parameters derived from effective spatial analysis. In the following discussion, multi-channel systems are discussed in relation to multi-channel microphone implementations. However, as discussed above, the input format can be any suitable input format, such as multi-channel speakers, Ambisonic (FOA / HOA), etc. It is understood that in some embodiments, channel positions are based on microphone positions or virtual positions or orientations. Furthermore, the output of the example system is a multi-channel speaker arrangement. However, it is understood that the output can be rendered to the user via means other than speakers. Furthermore, the multi-channel speaker signal can be summarized as two or more playback audio signals. Currently, the 3GPP standardization body is standardizing such systems as Immersive Voice and Audio Services (IVAS). IVAS is designed to be an extension of the existing 3GPP Enhanced Voice Service (EVS) codec to facilitate immersive voice and audio services on existing and future mobile (cellular) and fixed-line networks. IVAS applications can provide immersive voice and audio services through 3GPP fourth-generation (4G) and fifth-generation (5G) networks. Furthermore, the IVAS codec, as an extension of EVS, can be used in storage and forwarding applications, where audio and speech content is encoded and stored in files for playback. It is understood that IVAS can be used in conjunction with other audio and speech coding techniques that have the capability to encode samples of audio and speech signals.

[0048] For each considered time-frequency (TF) block or tile—in other words, for a time / frequency sub-band—the metadata includes at least the spherical orientation (elevation, azimuth), at least one energy ratio of the obtained orientation, extended coherence, and orientation-independent surround coherence. In general, IVAS can have multiple different types of metadata parameters for each time-frequency (TF) tile. The spatial audio parameter types that constitute the metadata used for IVAS are shown in Table 1 below.

[0049]

[0050]

[0051] This data can be encoded and transmitted (or stored) by the encoder so that the spatial signal can be reconstructed at the decoder.

[0052] Furthermore, in some cases, Metadata-Assisted Spatial Audio (MASA) can support up to two directions per TF tile, which requires encoding and transmitting the aforementioned parameters for each direction on a per-TF tile basis. Therefore, according to Table 1, the required bit rate is potentially doubled. Additionally, it is easy to foresee other MASA systems that can support more than two directions per TF tile.

[0053] In practical immersive audio communication codecs, the bit rate allocated for metadata can vary considerably. A typical overall operating bit rate for a codec might leave only 2 to 10 kbps for spatial metadata transmission / storage. However, some further implementations can allow up to 30 kbps or higher for spatial metadata transmission / storage. The encoding of directional parameters and energy fractions, as well as coherence data, have been previously examined. However, regardless of the bit rate allocated for spatial metadata transmission / storage, it will always be necessary to use as few bits as possible to represent these parameters, especially when TF tiles can support multiple directions corresponding to different sound sources in a spatial audio scene.

[0054] The concept discussed below is to combine each spatial audio parameter associated with each direction into one or more combined spatial audio parameters on a per TF tile basis.

[0055] Therefore, the present invention is based on the consideration that the bit rate per TF tile can be reduced by combining spatial audio parameters associated with each direction.

[0056] in this regard, Figure 1 Example apparatus and systems for implementing embodiments of this application are depicted. System 100 is shown having an 'analysis' section 121 and a 'synthesis' section 131. The 'analysis' section 121 is the portion that encodes from received multi-channel speaker signals to metadata and a downmixed signal, while the 'synthesis' section 131 is the portion that decodes the encoded metadata and the downmixed signal to present the regenerated signal (e.g., in the form of a multi-channel speaker).

[0057] The input to system 100 and the 'analysis' section 121 is a multi-channel signal 102. The microphone channel signal input is described in the following example; however, in other embodiments, any suitable input (or synthesized multi-channel) format can be implemented. For example, in some embodiments, the spatial analyzer and spatial analysis can be implemented externally to the encoder. For example, in some embodiments, spatial metadata associated with the audio signal can be provided to the encoder as a separate bitstream. In some embodiments, spatial metadata can be provided as a set of spatial (direction) index values. These are examples of metadata-based audio input formats.

[0058] The multi-channel signal is transmitted to the transmission signal generator 103 and the analysis processor 105.

[0059] In some embodiments, the transmission signal generator 103 is configured to receive a multi-channel signal, generate a suitable transmission signal comprising a defined number of channels, and output a transmission signal 104. For example, the transmission signal generator 103 may be configured to generate a 2-channel audio downmix of the multi-channel signal. The defined number of channels can be any suitable number of channels. In some embodiments, the transmission signal generator is configured to otherwise select or combine, for example, by beamforming the input audio signal into a defined number of channels and outputting it as a transmission signal.

[0060] In some embodiments, the transmission signal generator 103 is optional, and the multi-channel signal is passed to the encoder 107 unprocessed in the same manner as the transmission signal in this example.

[0061] In some embodiments, the analysis processor 105 is also configured to receive multichannel signals and analyze these signals to generate metadata 106 associated with the multichannel signals and therefore with the transmission signal 104. The analysis processor 105 may be configured to generate metadata that may include a direction parameter 108, an energy ratio parameter 110, and a coherence parameter 112 (and in some embodiments, a diffusion parameter) for each time-frequency analysis interval. In some embodiments, the direction, energy ratio, and coherence parameters may be considered spatial audio parameters. In other words, spatial audio parameters include parameters designed to characterize the sound field created / captured by the multichannel signals (or generally, two or more audio signals).

[0062] In some embodiments, the generated parameters may differ between frequency bands. Thus, for example, in frequency band X, all parameters are generated and transmitted, while in frequency band Y, only one parameter is generated and transmitted; furthermore, in frequency band Z, no parameters are generated or transmitted. A practical example of this could be that for some frequency bands, such as the highest frequency band, certain parameters are not needed for sensing purposes. Transmitted signal 104 and metadata 106 can be passed to encoder 107.

[0063] Encoder 107 may include an audio encoder core 109 configured to receive transmitted (e.g., downmixed) signals 104 and generate suitable encodings of these audio signals. In some embodiments, encoder 107 may be a computer (running suitable software stored on memory and at least one processor), or alternatively, a specific device such as using an FPGA or ASIC. Encoding can be implemented using any suitable scheme. Encoder 107 may also include a metadata encoder / quantizer 111 configured to receive metadata and output the information in an encoded or compressed form. In some embodiments, in Figure 1Prior to transmission or storage, as shown by the dashed lines, encoder 107 can further interleave, multiplex into a single data stream, or embed metadata within the encoded submixed signal. Multiplexing can be achieved using any suitable scheme.

[0064] On the decoder side, the received or retrieved data (stream) can be received by decoder / demultiplexer 133. Decoder / demultiplexer 133 can demultiplex the encoded stream and pass the audio encoded stream to transport extractor 135, which is configured to decode the audio signal to obtain the transport signal. Similarly, decoder / demultiplexer 133 may include metadata extractor 137, which is configured to receive encoded metadata and generate metadata. In some embodiments, decoder / demultiplexer 133 may be a computer (with suitable software running on memory and stored on at least one processor), or alternatively, may be a specific device, such as using an FPGA or ASIC.

[0065] The decoded metadata and transmitted audio signals can be passed to the synthesis processor 139.

[0066] The 'synthesis' section 131 of system 100 further illustrates a synthesis processor 139, which is configured to receive transmissions and metadata and, based on the transmission signals and metadata, recreate synthesized spatial audio in the form of a multi-channel signal 110 in any suitable format (which, depending on the use case, may be a multi-channel speaker format, or in some embodiments, any suitable output format such as a binaural or ambisonics signal).

[0067] Therefore, in general, firstly, the system (analysis section) is configured to receive multi-channel audio signals.

[0068] Furthermore, the system (analysis section) is configured to generate appropriate transmitted audio signals (e.g., by selecting or downmixing some audio signal channels) and spatial audio parameters as metadata.

[0069] The system is then configured to encode transmitted signals and metadata for storage / transmission.

[0070] After this, the system can store / send encoded transmissions and metadata.

[0071] The system can retrieve / receive encoded transmissions and metadata.

[0072] Furthermore, the system is configured to extract transport and metadata from encoded transport and metadata parameters, for example, to demultiplex and decode encoded transport and metadata parameters.

[0073] The system (synthesis section) is configured to synthesize and output multi-channel audio signals based on the extracted transmitted audio signals and metadata.

[0074] about Figure 2 The analysis processor 105 and metadata encoder / quantizer 111 (e.g., according to some embodiments) are described in more detail. Figure 1 (as shown in the image).

[0075] Figure 1 and Figure 2 The metadata encoder / quantizer 111 and the analysis processor 105 are depicted as being coupled together. However, it should be understood that some embodiments may not couple these two respective processing entities so tightly that the analysis processor 105 may reside on a different device than the metadata encoder / quantizer 111. Thus, a device including the metadata encoder / quantizer 111 may exist where the processing and encoding of the transmitted signals and metadata stream are independent of the capture and analysis processes. In this case, the energy estimator 205 may be configured as part of the metadata encoder / quantizer 111.

[0076] In some embodiments, the analysis processor 105 includes a time-frequency domain converter 201.

[0077] In some embodiments, the time-frequency domain transformer 201 is configured to receive the multi-channel signal 102 and apply a suitable time-frequency domain transform, such as the short-time Fourier transform (STFT), to convert the input time-domain signal into a suitable time-frequency signal. These time-frequency signals can then be passed to the spatial analyzer 203.

[0078] Therefore, for example, the time-frequency signal 202 can be represented in the time-frequency domain as:

[0079] s i (b,n)

[0080] Where b is the frequency bin index, n is the time-frequency block (frame) index, and i is the channel index. In another expression, n can be considered as the time index with a sampling rate lower than the sampling rate of the original time-domain signal. These frequency bins can be divided into multiple sub-bands, which divide one or more bins into sub-bands k = 0, ..., K-1 of the frequency band index. Each sub-band k has the lowest bin b. k,low and the highest warehouse b k,high And the subband contains from b k,low to b k,high All bins. The width of the subband can approximate any suitable distribution. For example, the Equivalent Rectangular Bandwidth (ERB) scale or the Barkscale.

[0081] Therefore, a time-frequency (TF) patch (or block) is a specific sub-band within a subframe of that frame.

[0082] It is understandable that the number of bits required to represent spatial audio parameters can depend at least in part on the TF (time-frequency) tile resolution (i.e., the number of TF subframes or tiles). For example, a 20ms audio frame can be divided into four time-domain subframes of 5ms each, and each time-domain subframe can have up to 24 frequency subbands in the frequency domain, divided according to the Bark scale, its approximation, or any other suitable division. In this particular example, the audio frame can be divided into 96 TF subframes / tiles, in other words, four time-domain subframes with 24 frequency subbands. Therefore, the number of bits required to represent the spatial audio parameters of the audio frame can depend on the TF tile resolution. For example, if each TF tile is to be encoded according to the distribution in Table 1 above, each TF tile will require 64 bits per sound source direction. For the two sound source directions of each TF tile, 2 x 64 bits will be needed to fully encode both directions. It should be noted that the term "sound source" can be used to indicate the dominant direction of propagating sound in a TF tile.

[0083] The implementation aims to reduce the number of bits when each TF block has more than one sound source direction.

[0084] In one embodiment, the analysis processor 105 may include a spatial analyzer 203. The spatial analyzer 203 may be configured to receive time-frequency signals 202 and estimate direction parameters 108 based on these signals. The direction parameters may be determined based on any audio-based 'direction' determination.

[0085] For example, in some embodiments, the spatial analyzer 203 is configured to estimate the direction of a sound source using two or more signal inputs.

[0086] Therefore, the spatial analyzer 203 can be configured to provide at least one azimuth and elevation angle for each frequency band and temporal time-frequency block within a frame of the audio signal, denoted as azimuth angle, respectively. And the elevation angle θ(k,n). The direction parameter 108 for the time subframe can also be passed to the spatial parameter merger 207.

[0087] The spatial analyzer 203 can also be configured to determine the energy ratio parameter 110. The energy ratio can be considered as determining the energy of an audio signal that can be considered to arrive from one direction. The direct-to-total energy ratio r(k,n) can be estimated, for example, using a stability metric for direction estimation, or any correlation metric, or any other suitable method for obtaining the ratio parameter. Each direct-to-total energy ratio corresponds to a specific spatial direction and describes how much energy, compared to the total energy, comes from that specific spatial direction. This value can also be represented individually for each time-frequency patch. The spatial direction parameter and the direct-to-total energy ratio describe how much of the total energy for each time-frequency patch comes from that specific direction. In general, the spatial direction parameter can also be considered as the direction of arrival (DOA).

[0088] In an embodiment, the direct pair total energy ratio parameter can be estimated based on the normalized cross-correlation parameter cor′(k,n) between microphone pairs in frequency band k, where the value of the cross-correlation parameter is between -1 and 1. This is achieved by interpolating the normalized cross-correlation parameter with the diffuse field normalized cross-correlation parameter cor′. D Compared to (k,n), the total energy ratio parameter r(k,n) can be determined directly as:

[0089]

[0090] The total energy ratio is further explained directly in PCT Publication WO2017 / 005978 (which is incorporated herein by reference). The energy ratio can be passed to the spatial parameter merger 207.

[0091] In embodiments, high-order directional audio coding with HOA input or methods with mobile device input, as proposed in PCT Publication WO2019 / 215391, can be used to analyze parameters related to the second direction (for TF tiles). Detailed information on high-order directional audio coding can be found in "Sector-Based Parametric Sound Field Reproduction in the Spherical Harmonic Domain" (IEEE Journal of Signal Processing, Vol. 9, No. 5).

[0092] The spatial analyzer 203 can also be configured to determine a plurality of coherence parameters 112, which may include surround coherence (γ(k,n)) and extended coherence (ζ(k,n)), both of which are analyzed in the time-frequency domain.

[0093] Next, we will discuss each of the aforementioned coherence parameters. All processing is performed in the time-frequency domain; therefore, for the sake of simplicity, the time-frequency indices k and n are omitted where necessary.

[0094] Let's first consider the case where two spaced-out speakers (e.g., front left and front right) are used instead of a single speaker to coherently reproduce the sound. A coherence analyzer can be configured to detect that this method has been applied to the surround mix.

[0095] It should be understood that the following sections illustrate the analysis of extended and surround coherence based on multi-channel speaker signal input. However, similar practices can be applied when the input includes a microphone array as input.

[0096] Therefore, in some embodiments, the spatial analyzer 203 can be configured to compute a covariance matrix C for a given analysis interval comprising one or more time indices n and frequency bins b. This matrix has a size of N. L x N L And its elements are represented as c ij , where N L This represents the number of speaker channels, and i and j are the speaker channel indices.

[0097] Next, the spatial analyzer 203 can be configured to determine the speaker channel i that is closest to the estimated direction (azimuth angle θ in this example). c :

[0098] i c =arg(min(|θ-α) i |))

[0099] Where, α i It is the angle of speaker i.

[0100] Furthermore, in this embodiment, the spatial analyzer 203 is configured to determine the location of the speaker i on the left side i. l and the right i r The closest speaker.

[0101] The normalized coherence between loudspeakers i and j is expressed as:

[0102]

[0103] Using this equation, the spatial analyzer 203 can be configured to calculate i l with i r Normalized coherence c′ between lr In other words, calculate:

[0104]

[0105] Furthermore, the spatial analyzer 203 can be configured to use the diagonal elements of the covariance matrix to determine the energy of speaker channel i:

[0106] E i =c ii

[0107] And determine speaker i l and i r With speaker i l i r and i c The energy ratio between them is:

[0108]

[0109] Furthermore, the spatial analyzer 203 can use these determined variables to generate the 'stereoness' parameter:

[0110] μ=c′ lr ξ lr / lrc

[0111] This 'stereometry' parameter has a value between 0 and 1. A value of 1 means that there is coherent sound in the speaker's il and ir, and that sound dominates the energy of that sector. This could be because, for example, the speaker uses amplitude shifting techniques to create an "airy" perception of sound. A value of 0 means that no such techniques have been applied, and, for example, the sound can simply be localized to the nearest speaker.

[0112] Furthermore, the spatial analyzer 203 can be configured to detect or at least identify situations where three (or more) speakers coherently reproduce sound to create a "close" perception (e.g., using front left, front right, and center instead of just center). This could be because a mixing engineer has created such a situation when performing surround mixing on a multi-channel speaker mix.

[0113] In this embodiment, the coherence analyzer uses the same previously identified speaker i l i r and i c The normalized coherence value c′ is determined using the previously discussed normalized coherence determination method. cl and c′ cr In other words, calculate the following values:

[0114]

[0115] Furthermore, the spatial analyzer 203 can use the following formula to determine the normalized coherence value c′ describing the coherence between these loudspeakers. clr :

[0116] c′ clr =min(c′) cl ,c′ cr )

[0117] Additionally, the spatial analyzer 203 can be configured to determine how energy (i.e., uniformity) is described in channel i. l i r and i c Parameters that are uniformly distributed between them:

[0118]

[0119] Using these variables, the spatial analyzer 203 can determine the new coherent translation parameter κ as:

[0120] κ=c′ clr ξ clr

[0121] This coherent translation parameter κ has a value between 0 and 1. A value of 1 means that in all speakers il, i r and i c There is coherent sound present, and the energy of this sound is evenly distributed among these speakers. This could be because, for example, speaker mixing is generated using a recording mixing technique that uses perception to create a closer sound source. A value of 0 means that no such technique has been applied; for example, the sound can simply be localized to the nearest speaker.

[0122] Determine metric i l and i r In (but not in i) c The "stereo" parameter μ of the coherent sound quantity in (in the middle) and the measure of all i l i r and i c The spatial analyzer 203, which analyzes the coherence translation parameter κ of the coherent sound quantity, is configured to use these parameters to determine the coherence parameters to be output as metadata.

[0123] Therefore, the spatial analyzer 203 is configured to combine the "stereometry" parameter μ and the coherence translation parameter κ to form an extended coherence ζ parameter, which has a value from 0 to 1. An extended coherence ζ value of 0 indicates a point source; in other words, as few speakers as possible should be used (e.g., only speaker i). c To reproduce sound. As the value of extended coherence ζ increases, more energy is extended to the speaker i. c Surrounding speakers; until the value is 0.5, the energy in speaker i l i r and i cThey are evenly distributed between them. When the value of the extended coherence ζ exceeds 0.5, the speaker i c The energy in the speaker decreases until it reaches a value of 1. c There is no energy in it, and all the energy is in the speaker. l and i r Place.

[0124] In some embodiments, using the parameters μ and κ described above, the spatial analyzer 203 is configured to determine the extended coherence parameter ζ using the following expression:

[0125]

[0126] The above expression is merely an example, and it should be noted that the spatial analyzer 203 may use any other method to estimate the extended coherence parameter ζ, as long as it conforms to the parameter definition above.

[0127] In addition to being configured to detect previous situations, the spatial analyzer 203 can also be configured to detect or at least identify situations in which sound is coherently reproduced from all (or almost all) speakers to create a perception of "inside-the-head" or "above".

[0128] In some embodiments, the spatial analyzer 203 can be configured to analyze the determined energy E with the maximum value. i and speaker channel i e Perform classification / sorting.

[0129] Furthermore, the spatial analyzer 203 can be configured to determine the relationship between this channel and M. L Normalized coherence c′ among the other loudest channels ij Furthermore, this channel and M can be monitored. L These normalized coherences c′ between the other loudest channels ij Value. In some embodiments, M L It can be N L -1, which would mean monitoring the coherence between the loudest channel and all other speaker channels. However, in some embodiments, M L It can be a smaller number, for example, N L -2. Using these normalized coherence values, the coherence analyzer can be configured to determine the surrounding coherence parameter γ using the following expression:

[0130]

[0131] in, The loudest channel and M L Normalized coherence among the loudest channels.

[0132] The surround coherence parameter γ has values ​​from 0 to 1. A value of 1 means that there is coherence between all (or almost all) speaker channels. A value of 0 means that there is no coherence between all (or even almost all) speaker channels.

[0133] The above expression is merely an example of estimating the surrounding coherence parameter γ, and any other method can be used as long as it conforms to the parameter definition above.

[0134] The spatial analyzer 203 can be configured to output the determined extended coherence parameter ζ and the surrounding coherence parameter γ to the spatial parameter combiner 207.

[0135] Therefore, for each subband k, there will be a set of spatial audio parameters associated with that subband. In this case, each subband k may have the following associated spatial parameters: at least one azimuth and elevation angle (denoted as azimuth φ(k,n) and elevation θ(k,n)), surround coherence (γ(k,n)), extended coherence (ζ(k,n)), and a direct parameter r(k,n) to the total energy ratio.

[0136] In an embodiment, the spatial parameter combiner 207 can be configured to combine each of the aforementioned parameters for each sound source direction into a combined set of parameters for a smaller number of directions. For example, a typical example may exist where a TF patch may have been assigned two spatial audio parameter sets, one spatial audio parameter set for one direction. In this case, the spatial parameter combiner can be configured to combine the two spatial audio parameter sets into a single combined spatial audio parameter set on a per-TF patch basis.

[0137] Therefore, in general, the spatial parameter combiner 207 can be configured to combine N spatial parameter sets (one spatial parameter set per direction) into Q combined spatial parameter sets on a per TF tile basis, where Q < N. For example, in the case of three directions per TF tile, the corresponding spatial parameter sets can be combined into a single combined spatial audio parameter set. Another example could include four directions per TF tile. In this case, the spatial audio parameter sets associated with each direction (four in total) can be combined into two combined spatial parameter sets.

[0138] For clarity, the following description is provided with the assumption that each TF tile has two source directions. However, it should be understood that combinations can be made on a set of spatial audio parameters associated with a greater number of source directions.

[0139] in this regard, Figure 3Some processing steps that the spatial parameter combiner 207 can be configured to perform in some embodiments are described.

[0140] It should be understood that the following processing steps are performed on a per-TF tile basis. In other words, processing is performed for each subband k in subframe n.

[0141] The spatial parameter combiner 207 can perform the combination by initially acquiring the spherical direction components of the azimuth angle φ1(k,n) and elevation angle θ1(k,n) for the first direction and the spherical direction components of the azimuth angle φ2(k,n) and elevation angle θ2(k,n) for the second direction, and converting each direction component into their corresponding Cartesian coordinates.

[0142] Furthermore, each Cartesian coordinate can be weighted with respect to the corresponding direct total energy ratio parameter r(k,n) for the corresponding direction.

[0143] The conversion operation of the azimuth angle φ1(k,n) and elevation angle θ1(k,n) directional components in the first direction gives the X-axis directional component of the first direction as follows:

[0144] x1(k,n)=r1(k,n)cosθ1(k,n)cosφ1(k,n)

[0145] The Y-axis component is:

[0146] y1(k,n)=r1(k,n)cosθ1(k,n)sinφ1(k,n)

[0147] The Z-axis component is:

[0148] z1(k,n)=r1(k,n)sinθ1(k,n)

[0149] The same steps can be performed on the second direction to give the X-axis component of the second direction as follows:

[0150] x2(k,n)=r2(k,n)cosθ2(k,n)cosφ2(k,n)

[0151] The second direction Y-axis component is:

[0152] y2(k,n)=r2(k,n)cosθ2(k,n)sinφ2(k,n)

[0153] The second direction Z-axis component is:

[0154] z2(k,n)=r2(k,n)sinθ2(k,n)

[0155] The steps to convert the spherical direction components for each direction into their equivalent Cartesian coordinates x, y, z are as follows: Figure 3 The processing step 301 is shown in the middle.

[0156] The steps for weighting each Cartesian coordinate x, y, z with their corresponding direct relation to the total energy parameter are as follows: Figure 3 The processing step 303 is shown in the middle.

[0157] Furthermore, the spatial parameter combiner 207 can be configured to sequentially combine each corresponding Cartesian coordinate for each direction to give a combined Cartesian coordinate. This combination step for each Cartesian coordinate can be expressed as:

[0158] x c (k,n) = x1(k,n) + x2(k,n)

[0159] y c (k,n)=y1(k,n)+y2(k,n)

[0160] z c (k,n)=z1(k,n)+z2(k,n)

[0161] The steps for combining Cartesian coordinates for each direction are as follows: Figure 3 The processing step 305 is shown in the middle.

[0162] Once the Cartesian coordinates x, y, z for all directions have been combined into Cartesian coordinates x c y c and z c The combined Cartesian coordinates can be converted into their equivalent combined azimuth angle φ. c (k,n) and elevation angle θ c (k,n) spherical direction components. In an embodiment, the combined Cartesian coordinates x can be expressed using the following expression. c y c and z c Perform this transformation on each of the following:

[0163]

[0164]

[0165] The function atan is the arc tangent of the angle in the correct quadrant for automatic detection.

[0166] The steps to transform the merged Cartesian coordinates into their equivalent merged spherical coordinates for each merged frequency band are as follows: Figure 3 The processing step 307 is shown in the middle.

[0167] In an embodiment, the combined Cartesian coordinates calculated as part of step 305 can be used in conjunction with the direct-to-total energy ratio for each direction to determine the combined direct-to-total energy ratio for both directions. The combined direct-to-total energy ratio r can be determined from the following expression: c (k,n):

[0168]

[0169] It can be seen that the molecule is the length of a combined Cartesian coordinate vector, which is directly related to the sum of the total energy ratios (r1(k,n)+r2(k,n)) and the additional factor ca according to the first and second directions. 12 (k,n) is normalized.

[0170] Item a 12 (k,n) is the value of the ambient energy, that is, the energy remaining in the TF tile after the energy in these two directions has been removed. In the embodiment, the ambient energy can be expressed as:

[0171] a 12 (k,n)=1-(r1(k,n)+r2(k,n))

[0172] Factor c is an adjustable factor whose value can be between 0 and 1 (e.g., c = 0.5), which controls the balance between direct flow and environmental flow.

[0173] Determine the combined direct-to-total energy ratio r c The steps are shown as processing step 309.

[0174] Additionally, some embodiments can derive combined extended coherence ζ for the two directions ζ(k,n) and ζ2(k,n). c (k,n) can be calculated as a ratio-weighted average of the extended coherence for each direction by using the direct-to-total energy ratio (r1(k,n),r2(k,n)) for both directions.

[0175] In an embodiment, this can be expressed as:

[0176]

[0177] Determine the combined extended coherence value ζ for the first and second directions. c The steps are shown as processing step 311.

[0178] The spatial parameter combiner 307 can also calculate the combined circumferential coherence γ for the first and second directions in the TF patch. c The value of (k,n).

[0179] In an embodiment with two-directional TF tiles, a single circumferential coherence value γ may exist. 12 (k,n), as mentioned earlier, is a measure of the coherence of non-directional sound. In this case, the amount of non-directional sound can be obtained as a 12 (k,n)=1-(r1(k,n)+r2(k,n)), which is the amount of energy after the contribution of the two directional components has been removed according to the corresponding direct ratio to the total energy.

[0180] Combined circumferential coherence γ c (k,n) can be derived based on the premise that the increase in quantized non-directional sound is coherent or incoherent. In the embodiment, the combined surround coherence γ c (k,n) can be written as:

[0181]

[0182] Where a(k,n)=1-r c (k,n) is the energy of the non-directional sound, i.e., the combined ambient sound from the first and second directions. The increase in coherent energy around the captured sound field can be calculated as (a(k,n) - a) 12 (k,n))ζ c (k,n), and the energy of the non-directional coherent sound in the captured sound field can be given as a 12 (k,n)γ 12 (k,n). In this example of the derivation of combined encircling coherence, it is assumed that if the extended coherence of the original direction is large, the increase in non-directional energy will be coherent, while if the extended coherence of the original direction is small, the increase in non-directional energy will be incoherent.

[0183] Determine the combined circumferential coherence value γ for the first and second directions. c The steps are shown as processing step 313.

[0184] In an embodiment, the spatial parameter combiner 207 may have an additional functional unit that provides an estimate (or metric) of the importance (effectively an importance estimator) of the total number of spatial parameter sets (or orientations) for each TF tile, compared to a reduced number of combined spatial parameter sets (and thus a reduced number of orientations). This estimate can then be fed into a decision function unit within the spatial parameter combiner 207, which determines whether the output for the TF tile can have spatial parameters for each orientation, or whether the output for the TF tile can include a combined set of spatial audio parameters. Furthermore, in embodiments with three or more orientations, the decision function unit may decide whether to combine spatial parameters associated with some orientations while leaving spatial parameters for other orientations uncombined.

[0185] Starting with the example above where each TF patch has two directions, the role of the importance estimator can be to estimate the importance of perceived audio quality with a set of spatial audio parameters for both directions, rather than with a single combined set of spatial audio parameters.

[0186] Therefore, the importance metric can be estimated (or derived) by comparing the sum of the total energies of the direct pairs for each direction with the length of the combined Cartesian coordinate vector derived above.

[0187] Therefore, for TF tiles, the importance estimate (or metric) λ(k,n) can be expressed as:

[0188]

[0189] In this case, the choice between sending two (original) sets of spatial parameters for both directions or a combined set of spatial parameters for one direction can be based on whether the importance metric λ(k,n) exceeds a threshold λ. th A comparison.

[0190] Therefore, if λ(k,n)>λ th Then, a decision can be made to encode and send the raw spatial audio parameters for both directions as metadata.

[0191] If λ(k,n)≤λ th Then, a decision can be made to encode and send the combined spatial audio parameters as metadata.

[0192] When it is decided to send in these two directions—in other words, when it is decided to send these two raw spatial audio parameter sets as metadata for TF tiles—the spatial parameter combiner 207 can be configured to output the raw (uncombined) spatial audio parameter sets θ1(k,n), θ2(k,n), φ1(k,n), φ2(k,n), r1(k,n), r2(k,n), ζ1(k,n), ζ2(k,n), and γ for the first and second directions. 12 (k,n).

[0193] If it is decided to send in one direction, in other words, if it is decided to send a combined set of spatial audio parameters for a TF tile, the spatial parameter combiner 207 will be configured to output the combined set of spatial audio parameters θ. c (k,n), φ c (k,n), r c (k,n), ζ c (k,n) and γ c (k,n).

[0194] It should be understood that, in the above situation, it will be necessary to determine whether signaling transmission is carried out in two directions or in one direction per TF block.

[0195] It should be understood that, in the above cases, it may be necessary to include signaling bits in the metadata to indicate whether the spatial audio parameters are for one direction (i.e., the combined set of spatial audio parameters) or for both directions (i.e., the original / uncombined set of spatial audio parameters).

[0196] In other embodiments, the selection performed by the spatial parameter combiner 207 can be performed at a higher granularity than for each TF tile. For example, signaling transmission for a set of TF tiles can be advantageous. This can be achieved by averaging the importance metric over a set of N subframes, such that the importance metric can be given by:

[0197]

[0198] Where N is the number of subframes in frame m. The advantage of using the average of the importance metric is that only signaling bits are needed for a set of merged frames and / or frequency bands, rather than for each merged time frame and / or frequency band.

[0199] Importance measures can have the property that if two directions point roughly in the same direction, the importance measure λ(k,n) will tend to have a lower value (in other words, tend to be "zero"). This can be achieved by comparing (r1(k,n)+r2(k,n)) with... Similar values ​​can be used to interpret this. If one of the direct pair total energy ratios is significantly greater than the other, the importance metric λ(k,n) will also tend to have a lower value. However, in contrast, if the two directions tend to point in opposite directions and the direct pair total energy ratios associated with each direction are approximately the same, the importance metric λ(k,n) will tend to have a value of 1.

[0200] In the embodiment, the threshold λ is selected. th The value can be fixed, and experiments have shown that a value of 0.3 gives favorable results.

[0201] In other embodiments, the importance threshold λ for a frame can be determined by the following operations. th The N importance metrics λ(k,n) in a frame are sorted in ascending order, and a threshold is determined to be the value of the importance metric that gives a specific number of importance metrics in the frame that are above the threshold. For example, the threshold metric can be adjusted so that I subframes in the frame have an importance metric higher than the adjusted threshold. In this case, I subframes will use 2 directions per TF tile, while NI subframes (those below the importance threshold) will use 1 combined direction per TF tile.

[0202] Additionally, some embodiments may not deploy a threshold. In these embodiments, the most important TF tiles in a frame / subframe may be set to use an uncombined orientation, while the remaining number of TF tiles in the frame / subframe are set to use a combined orientation.

[0203] Furthermore, additional embodiments can determine whether a particular TF tile should be set to be encoded on an average basis with either a combined or uncombined orientation. This can include setting an average number of TF tiles to be encoded with a combined orientation, and setting an average number of TF tiles to be encoded with an uncombined orientation.

[0204] In a further embodiment, the importance threshold λ th It can be adapted to the runtime value of the importance metric on the last N time subframes (e.g., the last 20 subframes). Therefore, λ med (n) can be represented as the median of the importance metric for subframe n across all frequency bands in the last N subframes. Furthermore, the importance threshold λ for subframe n... th (n) can be expressed as λ th (n)=c th λ med (n), where c th It is a coefficient that controls the value of the importance threshold, for example, c th It can be assigned the value 0.5.

[0205] Metadata encoder / quantizer 111 may include a direction encoder. This direction encoder can be configured to receive combined direction parameters (such as azimuth angle φ). c and elevation angle θ c (In some embodiments, a predetermined bit allocation is also received), thereby generating a suitable encoded output. In some embodiments, the encoding is based on an arrangement of spheres (forming a spherical grid arranged in a ring on a 'surface' sphere), which are defined by a lookup table defined by a determined quantization resolution. In other words, the spherical grid uses the concept of covering a sphere with smaller spheres and considering the center of the smaller sphere as a point defining a grid of nearly equidistant directions. Thus, the smaller sphere defines a cone or solid angle about the center point, which can be indexed according to any suitable indexing algorithm. Although spherical quantization has been described herein, any other suitable quantization (linear or nonlinear) may also be used.

[0206] The metadata encoder / quantizer 111 may include an energy ratio encoder. The energy ratio encoder 207 can be configured to receive the combined energy ratio r for each TF patch. c And determine the appropriate encoding for compressing these energy ratios.

[0207] Similarly, the metadata encoder / quantizer 111 may also include a coherent encoder, which can be configured to receive combined surround coherent values ​​γ. c and extended coherence value ζ c And determine the appropriate encoding for compressing these wraparound and extended coherence values ​​for TF tiles.

[0208] Encoded combination direction, energy ratio, and coherence value can be passed to combiner 211. The combiner is configured to receive encoded (or quantized / compressed) combination direction parameters, energy ratio parameters, and coherence parameters, and combine these parameters to generate a suitable output (e.g., a metadata bitstream, which can be combined with the transmitted signal, or sent or stored separately from the transmitted signal).

[0209] It should be noted that in embodiments deploying the aforementioned importance estimator, the metadata encoder / quantizer 111 may receive either combined spatial audio parameters per TF patch or an uncombined raw set of spatial audio parameters per TF patch for each direction, as described above. In the latter case, the uncombined set of spatial parameters for each direction, rather than the combined set, is passed to the various encoders. In this case, the metadata for each patch may be accompanied by signaling bits indicating whether the spatial parameter data is combined or uncombined.

[0210] An implementation may deploy a method for entropy coding to indicate whether a TF patch is encoded with bits in one or more directions. This can be useful when a fixed number of subbands are allocated in a frame with multiple directions.

[0211] In some embodiments, the encoded data stream can be passed to decoder / demultiplexer 133. Decoder / demultiplexer 133 demultiplexes / extracts the encoded combined direction index, combined energy ratio index, and combined coherence index for each TF tile and passes them to metadata extractor 137. In addition, in some embodiments, decoder / demultiplexer 133 can extract the transmitted audio signal and pass it to transmission extractor 135 for decoding and extraction.

[0212] In an embodiment, the decoder / demultiplexer 133 may be configured to receive and decode signaling bits that indicate whether the received coded spatial audio parameters are combined or uncombined for a specific TF tile.

[0213] The encoded combined energy ratio index, orientation index, and coherence index can be decoded by their respective decoders to generate the combined energy ratio, orientation, and coherence for the TF patch. This can be performed by applying the inverse of the various encoding processes used at the encoder.

[0214] When the signaling bits indicate that the spatial audio parameters are not combined, the received set of spatial audio parameters (for each direction of the TF tile) can be directly passed to various decoders for decoding.

[0215] Furthermore, the decoded spatial audio parameters can form decoded metadata output from metadata extractor 137, which is passed to synthesis processor 139 to form multichannel signal 110.

[0216] about Figure 4 Example electronic devices that can be used as analysis or synthesis devices are shown. The device can be any suitable electronic device or apparatus. For example, in some embodiments, device 1400 is a mobile device, user equipment, tablet computer, computer, audio playback device, etc.

[0217] In some embodiments, device 1400 includes at least one processor or central processing unit 1407. Processor 1407 may be configured to execute various program codes such as the methods described herein.

[0218] In some embodiments, device 1400 includes memory 1411. In some embodiments, at least one processor 1407 is coupled to memory 1411. Memory 1411 can be any suitable storage component. In some embodiments, memory 1411 includes a program code portion for storing program code that can be implemented on processor 1407. Furthermore, in some embodiments, memory 1411 may also include a stored data portion for storing data (e.g., data that has been processed or will be processed according to the embodiments described herein). Whenever needed, processor 1407 can access the implementation program code stored in the program code portion and the data stored in the stored data portion via memory-processor coupling.

[0219] In some embodiments, device 1400 includes a user interface 1405. In some embodiments, user interface 1405 may be coupled to processor 1407. In some embodiments, processor 1407 may control the operation of user interface 1405 and receive input from user interface 1405. In some embodiments, user interface 1405 may enable a user to input commands to device 1400, for example, via a keyboard. In some embodiments, user interface 1405 may enable a user to obtain information from device 1400. For example, user interface 1405 may include a display configured to display information from device 1400 to a user. In some embodiments, user interface 1405 may include a touchscreen or touch interface that enables information to be input to device 1400 and also displays information to a user of device 1400. In some embodiments, user interface 1405 may be a user interface for communicating with a location determiner as described herein.

[0220] In some embodiments, device 1400 includes an input / output port 1409. In some embodiments, input / output port 1409 includes a transceiver. In this embodiment, the transceiver may be coupled to processor 1407 and configured to enable communication with other devices or electronic devices, for example, via a wireless communication network. In some embodiments, the transceiver or any suitable transceiver or transmitter and / or receiver device may be configured to communicate with other electronic devices or devices via wired or wired coupling.

[0221] The transceiver can communicate with other devices using any suitable known communication protocol. For example, in some embodiments, the transceiver can use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a wireless local area network (WLAN) protocol such as IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth, or an infrared data communication path (IRDA).

[0222] The transceiver input / output port 1409 can be configured to receive signals, and in some embodiments, a processor 1407 executing appropriate code determines parameters as described herein. Furthermore, the device can generate appropriate downmixed signals and parameter outputs for transmission to a synthesis device.

[0223] In some embodiments, device 1400 may be at least part of a synthesis device. Thus, input / output port 1409 may be configured to receive a submixed signal and, in some embodiments, parameters determined at a capture or processing device as described herein, and to generate a suitable audio signal format output using processor 1407 executing appropriate code. Input / output port 1409 may be coupled to any suitable audio output, such as to a multi-channel speaker system and / or headphones or the like.

[0224] Generally, various embodiments of the present invention can be implemented using hardware or dedicated circuitry, software, logic, or any combination thereof. For example, some aspects can be implemented in hardware, while others can be implemented using firmware or software executable by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Although various aspects of the invention may be illustrated and described as block diagrams, flowcharts, or other graphical representations, it is well known that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples using hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.

[0225] Embodiments of the present invention can be implemented by computer software executable by a data processor of a mobile device (such as in a processor entity), or by hardware, or by a combination of software and hardware. Furthermore, in this regard, it should be noted that any block of the logical flow as shown in the figures may represent a program step, or interconnected logical circuits, blocks and functions, or a combination of program steps and logical circuits, blocks and functions. The software may be stored on a physical medium such as a memory chip or a memory block implemented within a processor, on a magnetic medium such as a hard disk or floppy disk, and on an optical medium such as a DVD and its data variant, the CD.

[0226] The memory can be of any type suitable for the local technical environment and can be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor can be of any type suitable for the local technical environment and, by way of non-limiting example, can include one or more of a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a gate-level circuit based on a multi-core processor architecture, and a processor.

[0227] Embodiments of the invention can be practiced in various components such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools can be used to transform logic-level designs into semiconductor circuit designs ready for etching and formation on semiconductor substrates.

[0228] The program can use well-established design rules and a pre-stored library of design modules to route conductors and position components on semiconductor chips. Once the semiconductor circuit design is complete, the resulting design in a standardized electronic format can be transferred to a semiconductor manufacturing facility or "fab" for fabrication.

[0229] The foregoing description has provided a complete and beneficial description of exemplary embodiments of the invention through exemplary and non-limiting examples. However, various modifications and adaptations will become apparent to those skilled in the art when read in conjunction with the accompanying drawings and appended claims, given the foregoing description. Nevertheless, all such modifications and similar alterations to the teachings of the invention will still fall within the scope of the invention as defined by the appended claims.

Claims

1. An apparatus for spatial audio coding, comprising: A component for determining or receiving a first spherical direction vector of a time-frequency patch of one or more audio signals and a second spherical direction vector of the time-frequency patch of the one or more audio signals, wherein the first spherical direction vector includes an azimuth component and an elevation component, and the second spherical direction vector includes an azimuth component and an elevation component; and A component for combining the first spherical direction vector and the second spherical direction vector to provide a combined spherical direction vector for the time-frequency patch, wherein the component for combining includes: A component for converting the first spherical direction vector into a first Cartesian vector, and A component for converting the second spherical direction vector into a second Cartesian vector, wherein the first Cartesian vector and the second Cartesian vector each include an x-axis component, a y-axis component, and a z-axis component, wherein, for each corresponding component, the device includes: A component for calculating the first direct-to-total energy ratio by weighting the corresponding components of the first Cartesian vector over the time-frequency patch; A component for calculating the second direct-to-total energy ratio by weighting the corresponding components of the second Cartesian vector over the time-frequency patch; and A component for summing the corresponding weighted components of the first Cartesian vector and the corresponding weighted components of the second Cartesian vector to give the corresponding combined Cartesian components, wherein the combined x-axis Cartesian component, the combined y-axis Cartesian component, and the combined z-axis Cartesian component form the components of the combined Cartesian vector; and A component for converting the combined x-axis Cartesian component, the combined y-axis Cartesian component, and the combined z-axis Cartesian component into the combined spherical direction vector.

2. The apparatus according to claim 1, wherein, The apparatus further includes: a component for determining whether the combined spherical orientation vector for the time-frequency patch is encoded for storage and / or transmission, or whether the first spherical orientation vector of the time-frequency patch and the second spherical orientation vector of the time-frequency patch are encoded for storage and / or transmission.

3. The apparatus according to claim 2, wherein, The device further includes: A component for determining a first metric for the time-frequency block of the one or more audio signals; The component for comparing the first metric with a threshold, further comprising means for determining whether the combined spherical orientation vector for the time-frequency patch is encoded for storage and / or transmission, or whether the first spherical orientation vector of the time-frequency patch and the second spherical orientation vector of the time-frequency patch are encoded for storage and / or transmission, includes: Components for determining that when the first metric is higher than the threshold, the first spherical orientation vector of the time-frequency patch and the second spherical orientation vector of the time-frequency patch are encoded for storage and / or transmission; and A component for determining that when the first metric is less than or equal to the threshold, the combined spherical orientation vector for the time-frequency patch is encoded for storage and / or transmission.

4. The apparatus according to claim 1, wherein, The device further includes: A component for determining a first metric for the time-frequency block of the one or more audio signals; Components for determining a first spherical orientation vector of at least one other time-frequency patch of the one or more audio signals and a second spherical orientation vector of the at least one other time-frequency patch of the one or more audio signals; A component for combining the first spherical direction vector of the one or more audio signals and the second spherical direction vector of the one or more audio signals and the at least one other time-frequency patch to provide a combined spherical direction vector for the other time-frequency patch of the one or more audio signals; Components for determining a second metric for the at least one other time-frequency patch; and A component for determining that, when the first metric is higher than the second metric, the first spherical direction vector of the time-frequency patch of the one or more audio signals and the second spherical direction vector of the time-frequency patch of the one or more audio signals are encoded for storage and / or transmission, and the combined spherical direction vector for the at least one other time-frequency patch of the one or more audio signals is encoded for storage and / or transmission.

5. The apparatus according to claim 1, wherein, The apparatus further includes: a component for determining a first ambient energy value for the time-frequency patch by subtracting from 1 a first direct-to-total energy ratio calculated for the time-frequency patch and a second direct-to-total energy ratio calculated for the time-frequency patch.

6. The apparatus according to claim 5, wherein, The apparatus further includes a component for combining a first direct-to-total power ratio calculated for the time-frequency patch and a second direct-to-total power ratio calculated for the time-frequency patch to provide a combined direct-to-total power ratio for the time-frequency patch.

7. The apparatus according to claim 6, wherein, The component for combining the first direct-to-total energy ratio calculated for the time-frequency patch and the second direct-to-total energy ratio calculated for the time-frequency patch to provide a combined direct-to-total energy ratio for the time-frequency patch includes: A component for determining the combined direct-to-total energy ratio based on the ratio of the vector length of the combined Cartesian vector to the sum of the first direct-to-total energy ratio calculated for the time-frequency patch and the second direct-to-total energy ratio calculated for the time-frequency patch, and the first ambient energy value.

8. The apparatus according to claim 5, wherein, The apparatus further includes: a component for combining a first extended coherence value calculated for the time-frequency patch and a second extended coherence value calculated for the time-frequency patch to provide a combined extended coherence value for the time-frequency patch.

9. The apparatus according to claim 8, wherein, The component for combining the first extended coherence value calculated for the time-frequency patch and the second extended coherence value calculated for the time-frequency patch to provide a combined extended coherence value for the time-frequency patch includes: The component for determining a first sum, wherein the first sum includes the product of a first extended coherence value calculated for the time-frequency patch and a first direct-to-total power ratio calculated for the time-frequency patch, and the product of a second extended coherence value calculated for the time-frequency patch and a second direct-to-total power ratio calculated for the time-frequency patch; Components for determining a second sum, wherein the second sum includes a first direct-to-total energy ratio calculated for the time-frequency patch and a second direct-to-total energy ratio calculated for the time-frequency patch; and Components used to determine the ratio of the first sum to the second sum in order to provide the combined extended coherence value.

10. The apparatus according to claim 8, wherein, The apparatus for spatial audio coding further includes: Components used to calculate the surround coherence value for the time-frequency plot; A component for determining a second ambient energy value for the time-frequency patch by subtracting the combined direct-to-total energy ratio from 1; A component for determining the surrounding coherence energy by determining the product of the combined extended coherence value and the difference between the second ambient energy value for the time-frequency patch and the first ambient energy value for the time-frequency patch; and Components for adding the surrounding coherent energy to the product of the ambient energy for the time-frequency patch and the surrounding coherent value for the time-frequency patch, and normalizing it to the second ambient energy value for the time-frequency patch to provide a combined surrounding coherent value.

11. The apparatus according to claim 3, wherein, The apparatus including the component for determining the first metric comprises: A component for determining the difference between the sum of the first direct-to-total energy ratio calculated for the time-frequency patch and the second direct-to-total energy ratio calculated for the time-frequency patch, and the length of the combined Cartesian vector.

12. The apparatus according to claim 1, wherein, The first spherical direction vector is associated with the direction of the first sound source in the time-frequency block, and the second spherical direction vector is associated with the direction of the second sound source in the time-frequency block.

13. A method for spatial audio coding, comprising: Determine or receive a first spherical direction vector of a time-frequency patch of one or more audio signals and a second spherical direction vector of the time-frequency patch of the one or more audio signals, wherein the first spherical direction vector includes an azimuth component and an elevation component, and the second spherical direction vector includes an azimuth component and an elevation component; and The first spherical direction vector and the second spherical direction vector are combined to provide a combined spherical direction vector for the time-frequency patch, wherein the combination includes: The method involves converting the first spherical direction vector into a first Cartesian vector and the second spherical direction vector into a second Cartesian vector, wherein the first Cartesian vector and the second Cartesian vector each include an x-axis component, a y-axis component, and a z-axis component, and for each corresponding component, the method includes: The first direct pair total energy ratio is calculated by weighting the corresponding components of the first Cartesian vector for the time-frequency patch; The second direct-pair total energy ratio is calculated by weighting the corresponding components of the second Cartesian vector over the time-frequency patch; and The corresponding weighted components of the first Cartesian vector and the corresponding weighted components of the second Cartesian vector are summed to give the corresponding combined Cartesian components, wherein the combined x-axis Cartesian component, the combined y-axis Cartesian component, and the combined z-axis Cartesian component form the components of the combined Cartesian vector; and The combined x-axis Cartesian component, the combined y-axis Cartesian component, and the combined z-axis Cartesian component are converted into the combined spherical direction vector.

14. The method according to claim 13, wherein, The method further includes: determining whether the combined spherical direction vector for the time-frequency patch is encoded for storage and / or transmission, or whether the first spherical direction vector of the time-frequency patch and the second spherical direction vector of the time-frequency patch are encoded for storage and / or transmission.

15. The method according to claim 14, wherein, The method further includes: Determine a first metric for the time-frequency block used for the one or more audio signals; The method of comparing the first metric with a threshold, further comprising determining whether the combined spherical orientation vector for the time-frequency patch is encoded for storage and / or transmission, or whether the first spherical orientation vector of the time-frequency patch and the second spherical orientation vector of the time-frequency patch are encoded for storage and / or transmission, includes: When the first metric is higher than the threshold, the first spherical orientation vector of the time-frequency patch and the second spherical orientation vector of the time-frequency patch are determined to be encoded for storage and / or transmission; and When the first metric is less than or equal to the threshold, the combined spherical orientation vector for the time-frequency patch is determined to be encoded for storage and / or transmission.

16. The method according to claim 13, wherein, The method further includes: Determine a first metric for the time-frequency block used for the one or more audio signals; Determine a first spherical direction vector of at least one other time-frequency patch of the one or more audio signals and a second spherical direction vector of the at least one other time-frequency patch of the one or more audio signals; The first spherical orientation vector of the at least one other time-frequency patch of the one or more audio signals and the second spherical orientation vector of the at least one other time-frequency patch of the one or more audio signals are combined to provide a combined spherical orientation vector for the other time-frequency patch of the one or more audio signals; Determine a second metric for the at least one other time-frequency patch; and When the first metric is higher than the second metric, the first spherical direction vector of the time-frequency patch of the one or more audio signals and the second spherical direction vector of the time-frequency patch of the one or more audio signals are encoded for storage and / or transmission, and the combined spherical direction vector for the at least one other time-frequency patch of the one or more audio signals is encoded for storage and / or transmission.

17. The method according to claim 13, wherein, The method further includes: determining a first ambient energy value for the time-frequency patch by subtracting the first direct-to-total energy ratio calculated for the time-frequency patch and the second direct-to-total energy ratio calculated for the time-frequency patch from 1.

18. The method according to claim 17, wherein, The method further includes: combining a first direct-to-total energy ratio calculated for the time-frequency patch and a second direct-to-total energy ratio calculated for the time-frequency patch to provide a combined direct-to-total energy ratio for the time-frequency patch.

19. The method according to claim 18, wherein, Combining the first direct-to-total energy ratio calculated for the time-frequency patch and the second direct-to-total energy ratio calculated for the time-frequency patch to provide a combined direct-to-total energy ratio for the time-frequency patch includes: The combined direct-to-total energy ratio is determined based on the ratio of the vector length of the combined Cartesian vector to the sum of the first direct-to-total energy ratio calculated for the time-frequency patch and the second direct-to-total energy ratio calculated for the time-frequency patch, and the first ambient energy value.

20. The method of claim 17, wherein, The method further includes: combining a first extended coherence value calculated for the time-frequency patch and a second extended coherence value calculated for the time-frequency patch to provide a combined extended coherence value for the time-frequency patch.

21. The method according to claim 20, wherein, Combining the first extended coherence value calculated for the time-frequency patch and the second extended coherence value calculated for the time-frequency patch to provide a combined extended coherence value for the time-frequency patch includes: A first sum is determined, the first sum comprising the product of the first extended coherence value calculated for the time-frequency patch and the first direct-to-total power ratio calculated for the time-frequency patch, and the product of the second extended coherence value calculated for the time-frequency patch and the second direct-to-total power ratio calculated for the time-frequency patch; Determine a second sum, the second sum comprising the first direct-to-total energy ratio calculated for the time-frequency patch and the second direct-to-total energy ratio calculated for the time-frequency patch; and The ratio of the first sum to the second sum is determined to provide the combined extended coherence value.

22. The method according to claim 20, wherein, The method for spatial audio coding further includes: Calculate the surround coherence value for the time-frequency patch; The second ambient energy value for the time-frequency patch is determined by subtracting the combined direct-to-total energy ratio from 1. The surrounding coherence energy is determined by multiplying the combined extended coherence value with the product of the difference between the second ambient energy value for the time-frequency patch and the first ambient energy value for the time-frequency patch; and The surrounding coherence energy is added to the product of the ambient energy for the time-frequency patch and the surrounding coherence value for the time-frequency patch, and then normalized to the second ambient energy value for the time-frequency patch to provide a combined surrounding coherence value.

23. The method according to claim 15, wherein, Determining the first metric includes: The difference between the sum of the first direct-pair total energy ratio calculated for the time-frequency patch and the second direct-pair total energy ratio calculated for the time-frequency patch, and the length of the combined Cartesian vector is determined.

24. The method according to claim 13, wherein, The first spherical direction vector is associated with the direction of the first sound source in the time-frequency block, and the second spherical direction vector is associated with the direction of the second sound source in the time-frequency block.

Citation Information

Patent Citations

  • Spatial audio processing apparatus

    WO2017005978A1

  • An apparatus, method and computer program for audio signal processing

    WO2019215391A1

  • Object clustering for rendering object-based audio content based on perceptual criteria

    WO2014099285A1