Spatial metadata direction coordination

By optimizing the sorting of directional metadata parameters in the time-frequency block grid, the problem of low coding efficiency and decreased perceptual quality caused by incorrect sorting of directional metadata parameters in immersive audio codecs is solved, achieving high-quality spatial audio output at low bit rates.

CN120937074APending Publication Date: 2025-11-11NOKIA TECHNOLOGIES OY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202480024034.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-03-31
Filing Date
2024-02-29
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing immersive audio codecs suffer from problems such as directional metadata parameter sorting errors and shuffling leading to low encoding efficiency and reduced perceptual quality when processing spatial sound captured by microphone arrays.

Method used

By identifying sorting errors in the time-frequency block grid, the directional metadata parameters are reordered to meet predetermined conditions regarding their similarity or difference metrics with adjacent time-frequency blocks, thereby optimizing the sorting of directional metadata.

Benefits of technology

It improves coding efficiency and perceptual quality, ensuring high-quality spatial audio output even at low bit rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120937074A_ABST
    Figure CN120937074A_ABST
Patent Text Reader

Abstract

An apparatus comprises means for obtaining, with respect to at least two sources within an audio scene, ordered directivity metadata parameters associated with the at least two sources, the directivity metadata parameters identifying directions of arrival with respect to the at least two sources and being arranged in a frame, the frame is arranged as a time-frequency block grid with respect to a time axis and a frequency axis; determining a sorting error about at least one time-frequency block directivity metadata parameter, the ordering error is configured to identify that at least one adjacent time-frequency block directivity metadata parameter associated with another order index is more similar to and / or less different from the at least one time-frequency block directivity metadata parameter than the at least one adjacent time-frequency block directivity metadata parameter associated with the same order index; and reordering the determined at least one time-frequency block directivity metadata parameter to another order index.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to apparatus and methods for directional coordination of spatial metadata. Background Technology

[0002] Capturing spatial audio from inputs (such as microphone arrays and other sources) is a typical and effective choice for estimating a set of parameters based on the input (microphone array signal), such as the direction of sound in a frequency band and the ratio of the directional to the non-directional portion of the captured sound in that band. These parameters are known to well describe the perceptual spatial characteristics of the captured sound at the location of the microphone array. These parameters can then be used accordingly for the synthesis of spatial sound for binaural headphones, speakers, or other formats such as stereo.

[0003] Therefore, the direction in the frequency band, as well as the direct-to-total energy ratio and the diffuse energy ratio, is a particularly effective parameterization for spatial audio capture.

[0004] A set of parameters consisting of directional parameters and energy ratio parameters (indicating the directionality of sound) in a frequency band can also be used as spatial metadata for an audio codec (which may also include other parameters such as surround coherence, spread coherence, number of directions, distance, etc.). For example, these parameters can be estimated based on audio signals captured by a microphone array, and, for example, stereo or single-channel transmitted audio signals can be generated based on microphone array signals to be transmitted along with the spatial metadata.

[0005] Immersive audio codecs are being implemented to support a variety of operating points, from low bitrate operation to transparency. One example of such a codec is the Immersive Voice and Audio Services (IVAS) codec, designed for use on communication networks such as 3GPP 4G / 5G networks, including immersive services like immersive voice and audio for virtual reality (VR). This audio codec is expected to handle the encoding, decoding, and rendering of speech, music, and general audio. Furthermore, it is expected to support channel-based audio, object-based audio, and scene-based audio input, including spatial information about the sound field and sound sources. The codec is also expected to operate with low latency for session services and support high error robustness under various transport conditions.

[0006] For example, the transmitted audio signal can be encoded using the IVAS audio core codec, or the AAC (Advanced Audio Coding) or EVS (Enhanced Voice Services) encoder. The decoder can decode the audio signal into a PCM (Pulse Code Modulation) signal and process the sound in the frequency band (using spatial metadata) to obtain a spatial output, such as a binaural output.

[0007] The aforementioned immersive audio codec is particularly well-suited for encoding captured spatial sound from microphone arrays (e.g., in mobile phones, VR cameras, and standalone microphone arrays). However, such an encoder can have other input types, such as speaker signals, audio object signals, or stereo signals. Summary of the Invention

[0008] According to a first aspect, an apparatus is provided, comprising components for: acquiring ordered directional metadata parameters associated with at least two sources within an audio scene, the ordered directional metadata parameters identifying directions of arrival with respect to the at least two sources and being arranged in a frame arranged as a time-frequency tile grid about a time-axis and a frequency axis; determining an ordering error for at least one time-frequency tile directional metadata parameter, the ordering error being configured to identify that at least one neighboring time-frequency tile directional metadata parameter associated with another order index is more similar to and / or less different from at least one neighboring time-frequency tile directional metadata parameter associated with the same order index; and reordering the determined at least one time-frequency tile directional metadata parameter to another order index.

[0009] The component used to determine an ordering error for at least one time-frequency block directional metadata parameter can be used to: determine that the similarity measure between at least one adjacent time-frequency block directional metadata parameter associated with another sequential index and at least one time-frequency block directional metadata is greater than, equal to, or greater than the similarity measure between at least one adjacent time-frequency block directional metadata parameter associated with the same sequential index and at least one time-frequency block directional metadata.

[0010] The component used to determine an ordering error for at least one time-frequency block directional metadata parameter can be used to: determine that the difference measure between at least one adjacent time-frequency block directional metadata parameter associated with another sequential index and at least one time-frequency block directional metadata is less than, equal to, or less than the similarity measure between at least one adjacent time-frequency block directional metadata parameter associated with the same sequential index and at least one time-frequency block directional metadata.

[0011] At least one adjacent time-frequency block directional metadata parameter can be at least one of the following: the preceding time time directional metadata parameter; the succeeding time time directional metadata parameter; the preceding frequency time-frequency block directional metadata parameter; the succeeding frequency time-frequency block directional metadata parameter; the preceding time and frequency time-frequency block directional metadata parameter; the preceding time and succeeding frequency time-frequency block directional metadata parameter; and the succeeding time and preceding frequency time-frequency block directional metadata parameter.

[0012] The component for reordering at least one determined time-frequency block directional metadata parameter to another sequential index can be used to: reassign the at least one determined time-frequency block directional metadata parameter to another sequential index.

[0013] The component for reassigning another order index to at least one determined time-frequency block directional metadata parameter can be used to: determine which other order index, at least one adjacent time-frequency block directional metadata parameter, and at least one time-frequency block directional metadata parameter associated with the same order index are more similar and / or less different; and reassign at least one sub-frame metadata parameter to the determined other order.

[0014] According to a second aspect, a method is provided, comprising: acquiring ordered directional metadata parameters associated with at least two sources within an audio scene, the ordered directional metadata parameters identifying the direction of arrival with respect to the at least two sources and being arranged in a frame, the frame being arranged as a time-frequency block grid with respect to a time axis and a frequency axis; determining an ordering error with respect to at least one time-frequency block directional metadata parameter, the ordering error being configured to indicate that: at least one adjacent time-frequency block directional metadata parameter associated with another order index is more similar to and / or less different from at least one adjacent time-frequency block directional metadata parameter associated with the same order index; and reordering the determined at least one time-frequency block directional metadata parameter to another order index.

[0015] Determining an ordering error for at least one time-frequency block directional metadata parameter may include: determining that the similarity measure between at least one adjacent time-frequency block directional metadata parameter associated with another sequential index and at least one time-frequency block directional metadata is greater than, equal to, or greater than the similarity measure between at least one adjacent time-frequency block directional metadata parameter associated with the same sequential index and at least one time-frequency block directional metadata.

[0016] Determining an ordering error for at least one time-frequency block directional metadata parameter may include: determining that the difference measure between at least one adjacent time-frequency block directional metadata parameter associated with another sequential index and at least one time-frequency block directional metadata is less than, equal to, or less than the similarity measure between at least one adjacent time-frequency block directional metadata parameter associated with the same sequential index and at least one time-frequency block directional metadata.

[0017] At least one adjacent time-frequency block directional metadata parameter can be at least one of the following: the previous time-frequency block directional metadata parameter; the next time-frequency block directional metadata parameter; the previous frequency-frequency block directional metadata parameter; the next frequency-frequency block directional metadata parameter; the previous time and frequency-frequency block directional metadata parameter; the previous time and next frequency-frequency block directional metadata parameter; and the next time and previous frequency-frequency block directional metadata parameter.

[0018] Reordering at least one determined time-frequency block directional metadata parameter to another sequential index may include: reassigning another sequential index to the at least one determined time-frequency block directional metadata parameter.

[0019] Reassigning another order index to at least one determined time-frequency block directional metadata parameter may include: determining which other order index, at least one adjacent time-frequency block directional metadata parameter, and at least one time-frequency block directional metadata parameter associated with the same order index are more similar and / or less different; and reassigning at least one subframe metadata parameter to the determined other order.

[0020] According to a third aspect, an apparatus is provided, comprising at least one processor and at least one memory storing instructions, which, when executed by the at least one processor, cause the system to perform at least: with respect to at least two sources within an audio scene, acquiring ordered directional metadata parameters associated with at least two sources, the ordered directional metadata parameters identifying directions of arrival with respect to the at least two sources and being arranged in a frame, the frame being arranged as a time-frequency block grid with respect to a time axis and a frequency axis; determining an ordering error with respect to at least one time-frequency block directional metadata parameter, the ordering error being configured to indicate that: at least one adjacent time-frequency block directional metadata parameter associated with another order index is more similar to and / or less different from at least one adjacent time-frequency block directional metadata parameter associated with the same order index; and reordering the determined at least one time-frequency block directional metadata parameter to another order index.

[0021] The means for determining an ordering error with respect to at least one time-frequency block directional metadata parameter can be made to: determine that the similarity measure between at least one adjacent time-frequency block directional metadata parameter associated with another order index and at least one time-frequency block directional metadata is greater than, equal to, or greater than the similarity measure between at least one adjacent time-frequency block directional metadata parameter associated with the same order index and at least one time-frequency block directional metadata.

[0022] The means for determining an ordering error with respect to at least one time-frequency block directional metadata parameter can be made to: determine that the difference measure between at least one adjacent time-frequency block directional metadata parameter associated with another sequential index and at least one time-frequency block directional metadata is less than, equal to, or less than the similarity measure between at least one adjacent time-frequency block directional metadata parameter associated with the same sequential index and at least one time-frequency block directional metadata.

[0023] At least one adjacent time-frequency block directional metadata parameter can be at least one of the following: the previous time-frequency block directional metadata parameter; the next time-frequency block directional metadata parameter; the previous frequency-frequency block directional metadata parameter; the next frequency-frequency block directional metadata parameter; the previous time and frequency-frequency block directional metadata parameter; the previous time and next frequency-frequency block directional metadata parameter; and the next time and previous frequency-frequency block directional metadata parameter.

[0024] The means by which an action is made to reorder at least one determined time-frequency block directional metadata parameter to another sequential index can be made to: reassign another sequential index to at least one determined time-frequency block directional metadata parameter.

[0025] The means of causing to perform the reassignment of another order index for at least one determined time-frequency block directional metadata parameter can be made to: determine which other order index, at least one adjacent time-frequency block directional metadata parameter, and at least one time-frequency block directional metadata parameter associated with the same order index are more similar and / or less different; and reassign at least one subframe metadata parameter to the determined other order.

[0026] According to a fourth aspect, an apparatus is provided, comprising: an acquisition circuit system configured to acquire ordered directional metadata parameters with respect to at least two sources within an audio scene, the ordered directional metadata parameters being associated with at least two sources, the directional metadata parameters identifying directions of arrival with respect to the at least two sources and being arranged in a frame, the frame being arranged as a time-frequency block grid with respect to a time axis and a frequency axis; a determination circuit system configured to determine an ordering error with respect to at least one time-frequency block directional metadata parameter, the ordering error being configured to identify that: at least one adjacent time-frequency block directional metadata parameter associated with another order index is more similar to and / or less different from at least one adjacent time-frequency block directional metadata parameter associated with the same order index; and a reordering circuit system configured to reorder the determined at least one time-frequency block directional metadata parameter to another order index.

[0027] According to a fifth aspect, a computer program [or a computer-readable medium including program instructions] is provided, the instructions being configured to cause a device to perform at least the following: with respect to at least two sources within an audio scene, acquiring ordered directional metadata parameters, the ordered directional metadata parameters being associated with at least two sources, the directional metadata parameters identifying directions of arrival with respect to the at least two sources and being arranged in a frame, the frame being arranged as a time-frequency block grid with respect to a time axis and a frequency axis; determining an ordering error with respect to at least one time-frequency block directional metadata parameter, the ordering error being configured to indicate that: at least one adjacent time-frequency block directional metadata parameter associated with another order index is more similar to and / or less different from at least one adjacent time-frequency block directional metadata parameter associated with the same order index; and reordering the determined at least one time-frequency block directional metadata parameter to another order index.

[0028] According to a sixth aspect, a non-transitory computer-readable medium is provided, the non-transitory computer-readable medium comprising program instructions for causing a device to perform at least the following: with respect to at least two sources in an audio scene, acquiring ordered directional metadata parameters, the ordered directional metadata parameters being associated with at least two sources, the directional metadata parameters identifying directions of arrival with respect to the at least two sources and being arranged in a frame, the frame being arranged as a time-frequency block grid with respect to a time axis and a frequency axis; determining an ordering error with respect to at least one time-frequency block directional metadata parameter, the ordering error being configured to indicate that: at least one adjacent time-frequency block directional metadata parameter associated with another order index is more similar to and / or less different from at least one adjacent time-frequency block directional metadata parameter associated with the same order index; and reordering the determined at least one time-frequency block directional metadata parameter to another order index.

[0029] According to a seventh aspect, an apparatus is provided, comprising: means for acquiring ordered directional metadata parameters with respect to at least two sources within an audio scene, the ordered directional metadata parameters being associated with at least two sources, the directional metadata parameters identifying directions of arrival with respect to the at least two sources and being arranged in a frame, the frame being arranged as a time-frequency block grid with respect to a time axis and a frequency axis; means for determining an ordering error with respect to at least one time-frequency block directional metadata parameter, the ordering error being configured to identify that: at least one adjacent time-frequency block directional metadata parameter associated with another order index is more similar to and / or less different from at least one adjacent time-frequency block directional metadata parameter associated with the same order index; and means for reordering the determined at least one time-frequency block directional metadata parameter to another order index.

[0030] According to an eighth aspect, a computer-readable medium is provided, the computer-readable medium comprising program instructions for causing a device to perform at least the following: with respect to at least two sources in an audio scene, acquiring ordered directional metadata parameters, the ordered directional metadata parameters being associated with at least two sources, the directional metadata parameters identifying directions of arrival with respect to the at least two sources and being arranged in a frame, the frame being arranged as a time-frequency block grid with respect to a time axis and a frequency axis; determining an ordering error with respect to at least one time-frequency block directional metadata parameter, the ordering error being configured to indicate that: at least one adjacent time-frequency block directional metadata parameter associated with another order index is more similar to and / or less different from at least one adjacent time-frequency block directional metadata parameter associated with the same order index; and reordering the determined at least one time-frequency block directional metadata parameter to another order index.

[0031] An apparatus includes components for performing the actions described above.

[0032] An apparatus is configured to perform the actions described above.

[0033] A computer program includes program instructions for causing a computer to perform the methods described above.

[0034] A computer program product stored on a medium can cause a device to perform the methods described herein.

[0035] An electronic device may include the means as described herein.

[0036] A chipset may include devices as described herein.

[0037] The embodiments of this application are intended to solve problems related to the prior art. Attached Figure Description

[0038] To better understand this application, reference will now be made to the accompanying drawings by way of example, in which:

[0039] Figure 1 The apparatus for extracting MASA metadata is illustrated schematically;

[0040] Figure 2 The example MASA metadata frame subframe structure is illustrated schematically;

[0041] Figure 3 The time-frequency structure of a sample MASA metadata frame is illustrated schematically;

[0042] Figure 4 An example system suitable for implementing some embodiments is illustrated schematically;

[0043] Figure 5 The diagram illustrates a known metadata analyzer, metadata, and audio encoder.

[0044] Figure 6 Example input data with two direction fields is shown, each direction field corresponding to a physical direction;

[0045] Figure 7 Example input data with two direction fields is shown, each direction field corresponding to a physical direction, where the direction fields in the sub-band are shuffled;

[0046] Figure 8 An example of low spatial resolution mode coercion is shown, where orientation blending is caused by inconsistency in temporally consecutive subframes;

[0047] Figure 9 and Figure 10 An example metadata analyzer, metadata, and audio encoder are illustrated schematically according to some embodiments;

[0048] Figure 11 and Figure 12 Various embodiments are shown respectively. Figure 9 and Figure 10 The flowcharts shown below illustrate the operation of the example metadata analyzer, metadata and audio encoder; and

[0049] Figure 13 An example device suitable for implementing the apparatus shown in the previous figures is illustrated. Detailed Implementation

[0050] The following describes in more detail suitable apparatuses and possible mechanisms for encoding parametric spatial audio signals, including transmitted audio signals and spatial metadata. As mentioned above, immersive audio codecs (such as 3GPPIVAS) are being planned, supporting multiple operating points from low bitrate operation to transparency.

[0051] Metadata-Assisted Spatial Audio (MASA) is an example of a parametric spatial audio format and representation suitable as an input format for IVAS.

[0052] It can be viewed as an audio representation consisting of "N channels + spatial metadata". It is a scene-based audio format, particularly well-suited for spatial audio capture on practical devices such as smartphones. The idea is to describe a sound scene based on time-frequency variations in source direction and, for example, energy ratio. Sound energy in the scene that is not defined (described) by direction is described as diffuse (from all directions).

[0053] As described above, spatial metadata associated with an audio signal can include multiple parameters for each time-frequency block (such as multiple directions, and direct-to-total energy ratio, spread coherence, distance, etc. associated with each direction (or direction value)). Spatial metadata can also include other parameters, or can be associated with other parameters considered non-directional (such as surround coherence, spread-to-total energy ratio, residual-to-total energy ratio), but when combined with directional parameters, it can be used to define the characteristics of the audio scene. For example, a reasonable design choice to produce high-quality output is to determine that the spatial metadata includes one or more directions for each time-frequency subframe (and direct-to-total ratio, spread coherence, distance values, etc. associated with each direction).

[0054] Figure 1 An example MASA analyzer 101 is shown. The MASA analyzer 101 is configured to receive one or more input audio signals 100 and analyze the input audio signals to generate one or more transmission audio signals 102 and spatial metadata 104.

[0055] Examples of MASA spatial metadata are shown in the table below. These values ​​can be used for each time-frequency tile (TF tile). In other words, the metadata is arranged as frames comprising multiple TF tiles or time-frequency elements, which can be arranged in a 'grid' of TF tiles or TF elements, arranged along the time and frequency axes. In some implementations, the frame is subdivided into 24 frequency bands and 4 time subframes. In other implementations, other divisions of frequency and time may be used. Furthermore, in some implementations, the frame size (e.g., as implemented in IVAS) is 20 ms (and therefore the time subframe is 5 ms). However, similarly, other frame lengths may be used in other embodiments. In some embodiments, the MASA analyzer is configured to determine one or two directions for each time-frequency tile (i.e., each time-frequency tile has one or two direction indices, direct energy to total energy ratio, and extended coherence parameters). However, in some embodiments, the analyzer is configured to generate more than two directions for a time-frequency tile.

[0056] MASA streams can be rendered as various outputs, such as multichannel speaker signals (e.g., 5.1) or binaural signals.

[0057] A direction index is an encoded form of the azimuth and elevation angles (or other direction values, such as vectors based on Cartesian 2D or 3D or polar coordinates) of a sound or sound source.

[0058] As mentioned above, the frame size in IVAS is 20ms. An example of the (IVAS) frame structure is shown below. Figure 2 As shown, metadata frame 201 includes four 5ms-long time subframes. For example, Figure 2 The previous metadata subframe 4 200 is shown, followed by the current metadata frame 201, which includes metadata subframes 1 202, 2 204, 3 206, and 4 208. This is followed by the subsequent or next-next metadata subframe 1 210.

[0059] In addition, such as Figure 3 As shown, an example raw high-resolution metadata frame 300 with both high frequency resolution and high temporal resolution is illustrated. Frame 300 is shown with respect to TF blocks arranged along the time axis 303 and the frequency axis 301. Therefore, the TF blocks are arranged according to the time axis 303, which has four subframes or divisions: metadata subframe 1 302, metadata subframe 2 304, metadata subframe 3 306, and metadata subframe 4 308. Furthermore, a series of frequency bands or divisions are shown on the frequency axis 301 (although these frequency bands or divisions are...). Figure 3 (Not separately marked in the text). Therefore, for a specific TF block 350, there may be adjacent time TF blocks 360 and 370 or adjacent frequency TF blocks 353 and 354.

[0060] Furthermore, the IVAS codec is expected to operate at a wide range of bit rates, from very low bit rates (e.g., 13.2 kbps) to relatively high bit rates (e.g., 512 kbps or even 768 kbps). Since the raw bit rate of MASA metadata is approximately 300-500 kbps (depending on whether one or two directions are encoded simultaneously), the metadata is significantly compressed (especially at the lowest bit rates).

[0061] One aspect of compression can be methods for reducing the temporal and / or frequency resolution of metadata (which can be used in conjunction with other methods for compressing data).

[0062] For example, the original high-resolution metadata frame may include 24 frequency bands on the frequency axis and 4 time subframes (subframes 1 to 4) on the time axis, which represents the total time-frequency block (also known as a TF block). Known methods for reducing the number of time-frequency blocks to be transmitted, and thus significantly reducing the required bit rate, can be based on the methods described in UKIPO patent applications 1919130.3 and 1919131.1 and WO2021 / 130405, which propose methods for combining metadata from multiple frequency bands and / or time subframes into fewer frequency bands and / or time subframes.

[0063] For example, depending on the bit rate, 5-24 frequency bands and 1-4 subframes can be transmitted.

[0064] Therefore, this method includes a metadata resolution selector configured to select and generate at least one of the following: a 1sf high-frequency resolution (low-temporal resolution) metadata frame and a 4sf (low-frequency resolution) high-temporal resolution metadata frame, which can then be encoded and output.

[0065] Because MASA streams can be generated from a variety of devices, such as microphone arrays on mobile devices and dedicated stereo microphone arrays like Eigenmike, the methods used to determine spatial metadata can vary considerably between implementations. Some methods may have high temporal resolution but low frequency resolution, while others may have low temporal resolution but high frequency resolution.

[0066] To improve the coding efficiency of these two time-frequency resolutions, it has been proposed that MASA metadata be encoded in two different modes, as shown in PCT application WO2021250312. The first metadata frame resolution is a low time resolution (1sf) mode, which has only one time subframe mode but has high frequency resolution, and the other metadata frame resolution is a high time resolution (4sf) mode that maintains 4 time subframes but has low frequency resolution.

[0067] In this example, the low temporal resolution mode (1sf) is selected when the encoder receives spatial metadata that is determined or detected to be the same (or substantially the same or similar) across all subframes of the frame.

[0068] If the spatial metadata is not the same (or substantially different or dissimilar) across all subframes, a high temporal resolution (4sf) mode is used.

[0069] For example, at a specific bit rate, the low temporal resolution mode (1sf) can transmit 18 bands and 1 subframe (in other words, a total of 18 TF blocks), and the high temporal resolution mode (4sf) can transmit 5 bands and 4 subframes (in other words, a total of 20 TF blocks), which is roughly equivalent to the similar size of data transmitted at the same total bit rate.

[0070] In PCT application WO2019105575, the use of variable input metadata time-frequency resolution has been proposed. This achieves a similar trade-off to the approach in PCT application WO2021250312, but the decision is implemented outside the codec and can be based on the specific capture algorithm of the microphone array used.

[0071] Therefore, the above method illustrates a method for allowing the encoding quality to be maintained, wherein the time and frequency resolution are tuned or adjusted with respect to the audio input.

[0072] Figure 4 An example system in which some embodiments can be implemented is shown. The inputs are a transmitted audio signal 102 and spatial metadata 104. The transmitted audio signal 102 and spatial metadata 104 are passed to an encoder 401, which generates an encoded bitstream 402. The encoded bitstream 402 is received by a decoder 403, which is configured to generate a spatial audio output 404.

[0073] As described above, the system's input, transmitted audio signal 102, and spatial metadata 104 can be acquired in the form of a MASA stream. For example, the MASA stream can originate from a mobile device (containing a microphone array), or, as an alternative example, it can be created by an audio server that can potentially process MASA streams in some way.

[0074] Furthermore, in some embodiments, encoder 401 may be an IVAS encoder.

[0075] In some embodiments, decoder 403 may be configured to directly output spatial audio output 404 for rendering by an external renderer or editing / processing by an audio server. In some embodiments, decoder 403 includes a suitable renderer configured to render the output in a suitable format, such as binaural audio signals or multi-channel speaker signals (such as 5.1 or 7.1+4 channel formats), which are also examples of spatial audio output 404.

[0076] Encoder 401 in Figure 5 This is shown in more detail below. The encoder 401 in this example includes a spatial metadata encoder that is configured to operate such that when it sees four subframes with different metadata, the encoding uses a high temporal resolution of 4sf, but uses a low frequency resolution.

[0077] The spatial metadata encoder is configured to receive spatial metadata 104. The spatial metadata 104 is passed to a subframe analyzer 501, which is configured to analyze the subframes in the spatial metadata 104 to detect whether all four subframes are similar and whether the 1sf encoding mode is usable.

[0078] In some embodiments, the example similarity test can be implemented by comparing spatial metadata fields element-by-element, and if the difference between values ​​in a field is greater than a given threshold, the spatial metadata is considered different. If the metadata is not different, they are considered similar.

[0079] For example, the following can be implemented as a similarity check:

[0080] Check the populated orientation space metadata fields (1 or 2 orientations are active).

[0081] Examine the spatial metadata parameters in each time-frequency block.

[0082] If the difference in the azimuth parameter is greater than a given threshold, such as 0.5 degrees, then the metadata is different.

[0083] If the difference in the elevation parameter is greater than a given threshold, such as 0.5 degrees, then the metadata is different.

[0084] If the difference between the direct energy and total energy ratio parameter is greater than a given threshold, such as 0.1, then the metadata is different.

[0085] If the difference in the extended coherence parameter is greater than a given threshold, such as 0.1, then the metadata is different.

[0086] If the difference in the surrounding coherence parameter is greater than a given threshold, such as 0.1, then the metadata is different.

[0087] However, any suitable similarity test can be implemented. For example, an importance metric such as that proposed in UKIPO patent applications 1919130.3 and 1919131.1 can be used to compare direction and reach-total ratio, i.e., to compare direction vectors with lengths having reach-total ratios.

[0088] The analysis results 502 and spatial metadata 104 can be passed to a coherent detector and a 2dir analyzer 503, which are configured to examine the input and determine the presence of meaningful coherent metadata. The coherent detector and 2dir analyzer 503 can also be configured to analyze the spatial metadata and determine, on a per-band basis, whether one or two directions should be used.

[0089] Then, the analysis results 504 and spatial metadata 104 can be passed to the metadata codec configurator 505, which is used to generate configuration information 506, which can be passed to the metadata reducer (metadata encoder) 507.

[0090] The encoder is also configured to receive the transmitted audio signal 102 and pass it to the audio encoder 511 and also to the metadata reducer (metadata encoder) 507.

[0091] Then, the configuration information 506, the transmitted audio signal 102, and the spatial metadata 104 can be passed to the metadata reducer 507. The metadata reducer is configured to reduce the amount of metadata and generate encoded metadata 508.

[0092] In addition, the encoder includes an audio and metadata combiner (multiplexer) 513, which is configured to receive encoded transmitted audio signals 512 and encoded metadata 508, and generate a bitstream 514 that can be output based on them.

[0093] In some embodiments, two or more orientation fields may exist in the (MASA) spatial metadata.

[0094] Some (MASA) capture and analysis systems do not explicitly assign a direction between the (physical sound) source orientation and the metadata orientation field. Within each parameter TF block, capture might simply analyze the orientation of the source with the highest energy (in the TF block) and assign it to the first orientation field, then analyze the dominant orientation of the remaining sound field (e.g., the orientation of another sound source) and assign it to the second orientation field. In other cases, capture and analysis systems might divide the space into non-overlapping regions, analyze the orientation of sources within each region, and assign each region to a dedicated orientation field.

[0095] When the relative level of a sound source varies with time (and frequency), the spatial parameters associated with each physical sound source can be distributed across two (or more) other directional fields. Example methods for performing this analysis are given in EP application EP3791605 and UK patent application GB2114186.6.

[0096] For example, in a room, two sound sources can speak simultaneously from different directions. In practice, the direction associated with each source can change rapidly with time and frequency, whether they are in the first direction field or the second direction field.

[0097] Furthermore, while a source can be a physical source—in other words, a physical source of an audio signal, such as a speaker or musical instrument—it is understood that a source is not a physical source, but rather the result of a capture or capture analysis. A capture or capture analysis assigns or sorts some metadata (with a low energy ratio) to a direction, and in some cases, can represent a set of related audio sources with low energy ratios. This might occur, for example, when a capture analysis can be specified to generate two directions without a significant second physical source.

[0098] Furthermore, even though the (MASA) capture and analysis system can assign physical orientations to specific metadata orientation fields, this arrangement can be disrupted by data processing. For example, the (IVAS) encoder can reorder metadata orientation fields so that fields with a higher (or highest) direct energy to total energy ratio are assigned to the first position or orientation field (e.g., via the encoding function ivas_qmetadata_reorder_2dir_bands() in the IVAS encoder). This is not a problem in the first round of encoding. However, if the decoder output transmits audio signals and spatial metadata (in the so-called external output), and it is used as input to a second encoder (in the so-called concatenated encoding), the original orientation in the parameter TF block can be reassigned or shuffled to another orientation from the first round of encoding.

[0099] For several reasons, this shuffling of directional data (regardless of where the shuffling occurs) may not be optimal for the encoding algorithm. A prominent reason is that if the bit rate does not allow for the transmission of spatial metadata at full TF resolution, then, as mentioned above, the TF grid is coarser (containing fewer TF blocks) by combining TF blocks across time and / or frequency. This combination can be performed using energy-weighted averaging or some other method. When combining or averaging TF blocks containing spatial metadata from different directions, the metadata is effectively obfuscated, and the perceptual quality of any resulting decoded output will be significantly lower than the perceptual quality achievable at the same bit rate with directional alignment.

[0100] Another encoding-related drawback is that some metadata encoding systems use differential encoding to further reduce the bit rate of encoded metadata. In this case, the first value is encoded as is, but subsequent values ​​are encoded based on the difference with respect to the previous value. When the changes between values ​​are small or slow, this can be achieved as an efficient encoding scheme by changing the distribution of the data to be encoded. However, if the spatial metadata changes significantly, because the shuffling of metadata indicates that fields are associated with different elements (e.g., different physical orientations), differential encoding may perform poorly.

[0101] For example, an (MASA) audio scene may include two direct sound sources with approximately constant spatial locations.

[0102] It can be analyzed and determined that the example scenario has two direction fields, where each physical direction is assigned to a metadata direction field. (See below) Figure 6 and Figure 7 In the diagram, solid lines or boxes correspond to the first direct sound source, and dashed lines or boxes correspond to the second direct sound source.

[0103] For example, Figure 6 A series of metadata frames are shown, namely, frame #1 600, frame #2 602, and frame #3 604; and direction 1 parameters for the first direction 610 and direction 2 parameters for the second direction 620. The direction parameters are represented as follows: ,in θ It is the source azimuth. φ It is the source elevation angle, r It is the ratio of direct energy to total energy. srcIdx It is an index representing the physical source index (similar to a solid / dashed line visual representation), and sfIdx This is the index of the parameter subframe within the parameter frame. In the following text, it is assumed that the sources have roughly the same physical location (or at least move slowly). Therefore, this assumption can be expressed as... .

[0104] In addition, such as Figure 7 As shown, after reordering the shuffling operation on frames or subframes, the data in the spatial metadata direction field can be sorted so that at least one subframe is shuffled between two directions. Therefore, for example, Figure 7 A series of metadata frames are shown, namely frame #1 700, frame #2 702, and frame #3 704. These metadata frames are related to... Figure 6 The difference in the metadata frames shown is that some subframes of the direction 1 parameter of the first source or direction 610 are located in the direction 2 parameter field of the second direction 720, and some subframes of the direction 2 parameter of the second source or direction 620 are located in the direction 1 parameter field of the first direction 720.

[0105] Therefore, for example, for frame #1 700: Direction 1 parameter 710 is: First subframe 701 in the first direction; First direction, second subframe 703; Second direction, third subframe 721; and First direction, fourth subframe 705. Direction 2 parameter 720 is: Second direction, first subframe 751; Second direction, second subframe 753; First direction, third subframe 723; and Second direction, fourth subframe 755.

[0106] Regarding frame #2 702: Direction 1 parameter 710 is: Second direction, fifth subframe 761; First direction, sixth subframe 731; Second direction, seventh subframe 763; and Second direction, eighth subframe 765. Direction 2 parameter 720 is: First direction, fifth subframe 771; Second direction, sixth subframe 733; First direction, seventh subframe 773; and First direction, eighth subframe 775.

[0107] Regarding frame #3 704: Direction 1 parameter 710 is: First direction, ninth subframe 741; Second direction, tenth subframe 781; Second direction, eleventh subframe 783; and Second direction, twelfth subframe 785. Direction 2 parameter 720 is: Second direction, ninth subframe 743; First direction, tenth subframe 791; First direction, eleventh subframe 793; and First direction, twelfth subframe, 795.

[0108] For example, reordering or shuffling processes can be implemented, resulting in a higher direct energy to total energy ratio for the encoding system. r The direction is assigned to direction 1. Furthermore, as mentioned above, some capture algorithms may not assign a direction to the direction field based on the physical direction, and the resulting metadata may look different from... Figure 7 The direct similarity shown.

[0109] For the example (MASA) metadata input to the encoder and the directional shuffled subframes, where, for example, the bit rate is relatively limited, the encoder can use a low temporal resolution coding mode (1sf mode) to combine every 4 consecutive subframes. The result is a highly ambiguous representation of the directional parameters, as follows: Figure 8 As shown.

[0110] For example, Figure 8The diagram illustrates a 1sf mode combination of direction 1 parameter 810 and metadata frame #1 800, which is a functional combination f(.) of first direction first subframe 701, first direction second subframe 703, second direction third subframe 721, and first direction fourth subframe 705. Similarly, direction 2 parameter 820 and metadata frame #1 800 are functional combinations of second direction first subframe 751, second direction second subframe 753, first direction third subframe 723, and second direction fourth subframe 755.

[0111] Direction 1 parameter 810 and metadata frame #2 802 are functional combinations of the fifth subframe 761 of the second direction, the sixth subframe 731 of the first direction, the seventh subframe 763 of the second direction, and the eighth subframe 765 of the second direction. Similarly, direction 2 parameter 820 and metadata frame #2 802 are functional combinations of the fifth subframe 771 of the first direction, the sixth subframe 733 of the second direction, the seventh subframe 773 of the first direction, and the eighth subframe 775 of the first direction.

[0112] Furthermore, direction 1 parameter 810 and metadata frame #3 804 are a functional combination of the ninth subframe 741 of the first direction, the tenth subframe 781 of the second direction, the eleventh subframe 783 of the second direction, and the twelfth subframe 785 of the second direction. Similarly, direction 2 parameter 820 and metadata frame #3 804 are a functional combination of the ninth subframe 743 of the second direction, the tenth subframe 791 of the first direction, the eleventh subframe 793 of the first direction, and the twelfth subframe 795 of the first direction. The function f(.) can be any suitable combination function.

[0113] As shown in the figure, this will produce a 1sf frame in which the resulting direction 1 parameter and direction 2 parameter are blurred because the resulting direction 1 parameter contains some of the original direction 2 value, and vice versa.

[0114] This example illustrates a problem that the embodiments discussed herein attempt to overcome when applied along the time axis of data. Similar problems may arise when looking at spatial parameters aggregated across a frequency axis, and some embodiments can be applied to these parameters in a similar manner to those discussed in the following examples. This is because the number of frequency bands is typically reduced when spatial metadata is encoded at a finite bit rate (in the case of MASA, from 24 to 5 at the coarsest resolution). If the frequency bands combined in the encoding correspond to significantly different spatial orientations, the operations used to combine them may distort the resulting spatial representation.

[0115] The concepts discussed in further detail with respect to the following embodiments and examples relate to the encoding of parametric spatial audio (i.e., (one or more) audio signals and spatial metadata), wherein the spatial metadata is encoded in frames and subframes containing two (or more) direction fields.

[0116] In these embodiments, the apparatus and method are configured to preprocess spatial metadata with respect to orientation field assignment (sorting), such that any subsequent encoding operation maintains orientation accuracy better than without preprocessing.

[0117] In some embodiments, this can be achieved in the following ways: Two consecutive (sub)frames to acquire spatial metadata; The values ​​of the comparison direction field; and Determine the order of metadata direction fields to minimize the difference between two consecutive (sub)frames.

[0118] In some embodiments, in addition to the above, this can also be achieved by: obtaining spatial metadata from two spatial metadata direction fields and two adjacent frequency bands, comparing the values ​​of the direction fields, and determining the order of the metadata direction fields such that the difference between the two adjacent frequency bands is minimized.

[0119] As described in the following embodiments, a processing step is provided that attempts to align (or coordinate) the (MASA) metadata orientation field such that the total difference in orientation between adjacent parameter blocks is minimized. In some embodiments, coordination can be implemented on the time dimension as described below (and thus coordination can be used effectively when combining multiple subframe data, such as when operating in a low temporal resolution 1sf encoding mode).

[0120] In some embodiments, coordination can also be achieved in the frequency dimension (useful when reducing frequency resolution). Furthermore, in some embodiments, coordination can be jointly achieved in both the time and frequency dimensions, depending on the TF block grouping used in the encoding.

[0121] In some embodiments, aligning the spatial metadata orientation within grouped TF blocks aims to provide an advantage: any average will result in less distortion of the underlying data compared to cases where the orientation of the underlying data is shuffled. Furthermore, aligning the spatial metadata orientation across groups (or typically across TF domains) has the advantage of making data encoding more efficient due to reduced variations in the data.

[0122] about Figure 9 This shows that based on Figure 5 The encoder shown is an example encoder, but it includes metadata orientation alignment processing for the input spatial metadata and provides the aligned spatial metadata as a result for further processing.

[0123] In some embodiments, encoder 991 includes audio encoder 511, which is configured to encode an audio signal and generate an encoded transmit audio signal 512.

[0124] In this example, encoder 991 is configured to receive spatial metadata 104. Spatial metadata 104 is passed to metadata orientation aligner 901, which generates aligned spatial metadata 904.

[0125] The encoder 991 also includes a subframe analyzer 501, which is configured to analyze subframes in the aligned spatial metadata 904 to detect whether all four subframes are similar and whether the 1sf encoding mode is available.

[0126] The analysis result 502 and the aligned spatial metadata 904 can be passed to a coherent detector and a 2dir analyzer 503, which are configured to examine the input and determine whether meaningful coherent metadata exists. The coherent detector and 2dir analyzer 503 can also be configured to analyze the spatial metadata and determine whether one or both directions should be used on a per-band basis.

[0127] The analysis results 504 and the aligned spatial metadata 904 can then be passed to the metadata codec configurator 505, which generates configuration information 506, which can then be passed to the metadata reducer (metadata encoder) 507.

[0128] The encoder is also configured to receive the transmitted audio signal 102 and pass it to the audio encoder 511 and also to the metadata reducer (metadata encoder) 507.

[0129] Then, the configuration information 506, the transmitted audio signal 102, and the alignment spatial metadata 904 can be passed to the metadata reducer 507. The metadata reducer is configured to reduce the amount of metadata and generate encoded metadata 508.

[0130] In addition, the encoder includes an audio and metadata combiner (multiplexer) 513, which is configured to receive encoded transmitted audio signals 512 and encoded metadata 508, and generate a bitstream 514 that can be output based on them.

[0131] also, Figure 10 It shows from Figure 5 Another example encoder, 1091, is a modified version of the encoder shown. Figure 10In the example shown, the metadata orientation aligner can be located near or inside the metadata reducer or metadata encoder within the encoding chain or path. In this configuration, the metadata orientation aligner is configured to receive the metadata encoding configuration as additional information and use that information to determine the axis on which to operate (time axis operation across subframes and frames, frequency axis operation across parameter bands, or a combination of both), and perform metadata orientation field coordination on that axis.

[0132] In this example, encoder 1091 is configured to receive spatial metadata 104. Spatial metadata 104 is passed to subframe analyzer 501, which is configured to analyze the subframes in spatial metadata 104 to detect whether all four subframes are similar and whether the 1sf encoding mode is usable.

[0133] The analysis results 502 and spatial metadata 104 can be passed to a coherent detector and a 2dir analyzer 503, which are configured to examine the input and determine the presence of meaningful coherent metadata. The coherent detector and 2dir analyzer 503 can also be configured to analyze the spatial metadata and determine, on a per-band basis, whether one or two directions should be used.

[0134] The analysis results 504 and spatial metadata 104 can then be passed to the metadata codec configurator 505, which generates configuration information 506, which can then be passed to the metadata reducer (metadata encoder) 507 and the metadata orientation aligner 1001.

[0135] Metadata orientation aligner 1001 is configured to receive configuration 506 and spatial metadata 104, and generate aligned spatial metadata 1004 based thereon.

[0136] The encoder is also configured to receive the transmitted audio signal 102 and pass it to the audio encoder 511 and also to the metadata reducer (metadata encoder) 507.

[0137] In some embodiments, encoder 1091 includes audio encoder 511, which is configured to encode an audio signal and generate an encoded transmit audio signal 512.

[0138] Then, the configuration information 506, the transmitted audio signal 102, and the alignment spatial metadata 1004 can be passed to the metadata reducer 507. The metadata reducer is configured to reduce the amount of metadata and generate encoded metadata 508.

[0139] In addition, the encoder includes an audio and metadata combiner (multiplexer) 513, which is configured to receive encoded transmitted audio signals 512 and encoded metadata 508, and generate a bitstream 514 that can be output based on them.

[0140] Figure 11 An example flowchart is shown, illustrating... Figure 9 The operation of the encoder is shown.

[0141] Therefore, 1101 illustrates the operation for acquiring transmitted audio signals.

[0142] The encoding of the transmitted audio signal is shown in 1102.

[0143] In addition, 1103 shows the operations used to obtain spatial metadata.

[0144] Then there are the operations used to align spatial metadata, as shown in 1105.

[0145] Following the alignment of spatial metadata are operations used to analyze subframes, as shown in 1107.

[0146] Then there are operations for analyzing the coherence direction and the second direction, as shown in 1109.

[0147] Next is the configuration of the metadata codec, as shown in 1111.

[0148] Then, the metadata is encoded / reduced based on the configuration, as shown in 1113.

[0149] The encoded audio and metadata can then be combined to generate a bitstream, as shown in 1115.

[0150] Then comes the output of the bitstream, as shown in 1117.

[0151] Figure 12 An example flowchart is shown, illustrating... Figure 10 The operation of the encoder is shown.

[0152] Therefore, 1201 illustrates the operation for acquiring transmitted audio signals.

[0153] The encoding of the transmitted audio signal is shown in Figure 1202.

[0154] In addition, 1203 illustrates the operations used to obtain spatial metadata.

[0155] Then there are the operations used to analyze subframes, as shown in 1205.

[0156] Then there are operations for analyzing the coherence direction and the second direction, as shown in 1207.

[0157] Next is the configuration of the metadata codec, as shown in 1209.

[0158] Then there are the operations used to align spatial metadata, as shown in 1211.

[0159] Then, the metadata is encoded / reduced based on the configuration, as shown in 1213.

[0160] The encoded audio and metadata can then be combined to generate a bitstream, as shown in 1215.

[0161] Then comes the output of the bitstream, as shown in 1217.

[0162] In some embodiments, the operation of metadata orientation aligners 901 and 1001 can be implemented in the following ways: 0. Obtain or receive direction information from direction field 1. and the direction information in direction field 2 The given spatial metadata subframe. This is the initialization, and the output metadata subframe has a direction field. and . 1. Acquire or receive directional information and The next spatial metadata subframe. 2. When using the orientation field from the given data for allocation, determine a measure of the total difference in orientation between two subframes. Furthermore, if the allocation of the two directional fields is to be reversed, then the total difference measure is determined. 3. Compare the two difference measures and determine whether the orientation field assignment in the next subframe should be retained or reversed: ○ If Then the direction field assignment is maintained, where and ○ Otherwise, reverse the direction field assignment, where and 4. Assign using the current output direction field , For reference, repeat from step 1.

[0163] In this example, the metric is a "difference" metric, but in some embodiments, a "similarity" metric may be used instead. In these embodiments, the less-than comparison shown in step 3 above will be replaced with a greater-than comparison.

[0164] In addition, in some embodiments, less than or equal to comparisons can be used instead of less than comparisons (or similarly, for similarity measures, greater than or equal to comparisons can be used instead of greater than comparisons).

[0165] In some embodiments, difference measurement D (·) can be defined, for example, as angular distance (sensitive only to direction): Or Cartesian distance (also considering radius or distance):

[0166] Cartesian distance may be preferred in some cases because it more closely approximates what might happen in some embodiments of parameter aggregation in subframe grouping. In such embodiments, the parameter set may include parameters based on the energy of the transmitted audio signal. E Direct energy to total energy ratio r The determined weights are used instead of the ratio of direct energy to total energy. And the direction information is .

[0167] The examples above illustrate possible metrics, and in some embodiments, other difference metrics can be implemented. Furthermore, it is understood that a difference metric can also be referred to as a distance metric (e.g., the difference metric values ​​described above are determined based on a distance function).

[0168] The examples and embodiments above illustrate two orientation fields for each subframe / frame. In some embodiments, this can be extended to a higher number of simultaneous orientation fields. In such embodiments, instead of determining a difference / similarity metric for two candidate orientation fields, a metric for... N An evaluation of the difference / similarity metrics of all available candidate orders in each direction. For example, the following can be performed. 0. Use initialization N One direction field, . 1. Obtain directional information The next spatial metadata subframe. 2. Generate candidate order ord For example, by listing N All combinations of directions. 3. Assess the measure of difference

[0169] In these embodiments, in all N Summation is performed in each direction. From the original direction n Assigned to candidate directions dirIdx Directional metadata in the middle. 4. Choose the order that minimizes the difference measure: and assign output . It is a function that returns the measure of minimum difference. d ord order ord . 5. Repeat the above steps for the next subframe starting from step 1.

[0170] Furthermore, in some embodiments, the above examples focus on MASA spatial metadata with two direction fields. However, practical capture systems can also switch between analyzing a single direction and two directions, resulting in inconsistent codec inputs based on the number of directions. For example, a capture system might be configured to capture two directions, but due to the characteristics of spatial signals, it might only find one candidate direction and therefore either output only the single direction of the frame or set the second direction to zero (i.e., set the direct energy to total energy ratio of that direction to zero). The latter can even occur on individual TF blocks.

[0171] In some embodiments, energy-based averaging can correctly handle such data (e.g., zero-energy components do not cause bias in the average orientation data). In some embodiments, various implementations can track orientation data consistency across subframes and frames purely based on the orientation values ​​themselves, without energy weighting. Therefore, as part of the metadata orientation alignment, the zero-energy orientation should (in such implementations) be reset based on extrapolation (e.g., copying previous orientation data) or interpolation (e.g., averaging between previous and next orientation data). This also provides consistent 2dir MASA metadata in cases where the original input data switches between 1dir and 2dir. This step is also important in the case of concatenated coding operations, where prior quantization of the zero-energy orientation may result in a low-energy orientation after decoding.

[0172] Similar to the above, this example can be extended to more than two directions. (Since the IVAS MASA format specification is part of the IVAS design constraints (Tdoc S4-221619), and the corresponding IVAS codec implementation is expected to be used for IVAS candidate submissions, the focus is on two direction fields, switching between one and two.)

[0173] In the above embodiments, coordination is implemented in directions across the time dimension (on subframes and frames). This is useful when the following operations benefit from a time-consistent direction field. When encoding combines multiple frequency bands into a lower number of bands (in the case of MASA, the highest resolution is 24 bands and the lowest resolution is 5 bands), it is more beneficial to coordinate the spatial metadata in the bands that will be grouped together in the following processes. Coordination can be implemented in a manner similar to the embodiments described above for cross-time coordination data, but using... to replace .here, bandIdx It is an index of the frequency bands. Coordination can be performed on all 24 frequency bands of the spatial metadata (in the case of MASA) or within each subset of frequency bands grouped together at a lower frequency resolution.

[0174] In some embodiments, it may be advantageous not to use a fixed frequency band to determine the alignment starting point. Instead, embodiments may select a frequency band with the highest energy (e.g., determined based on the transmitted audio signal) and then apply the alignment method from that band to higher and lower frequencies.

[0175] The presented embodiment determines the orientation field order progressively within each subframe. This can be extended to consider multiple consecutive subframes simultaneously and determine the orientation field order for multiple consecutive subframes at the same time.

[0176] In addition, such as Figure 9 The illustrated embodiment describes the method as a preprocessing step, typically used before encoding and sending metadata. However, when the goal is to output metadata as part of the MASA format from the codec output, this method can also be used as a post-processing step after decoding the metadata from the bitstream. This ensures that any other possible codecs or renderers obtain the MASA format in an optimal form similar to preprocessing. Generally, it is beneficial for the proposed method to perform this process at least once for the metadata in any operation chain using the MASA format.

[0177] The presented embodiment uses orientation information (azimuth and elevation) from spatial metadata to determine alignment. This is only one possibility, and other embodiments may also consider other spatial metadata fields, such as extended coherence, when determining the total difference measure of the order of candidates for the orientation field.

[0178] Furthermore, in the example above, the three-dimensional orientation representation used in MASA's spatial metadata is considered. This is based on the azimuth (left and right angles on the horizontal plane) and elevation (angles with respect to the horizontal plane) of the orientation in a spherical coordinate system. This should be considered as an example implementation only. All operations can be implemented using other orientation parameters, such as azimuth and polar angles (angles with respect to the vertical plane), and in the case of two-dimensional orientation, only azimuth or elevation is considered.

[0179] also, Figure 9 and Figure 10 The encoder shown illustrates two possible locations for processing within the encoder. Processing (alignment) can also be applied at other locations within the processing chain, or even simultaneously at multiple locations. For example, an instance of processing could be placed near the input of the metadata encoder, similar to... Figure 9 It operates along the timeline. In addition, a second instance of the invention, similar to the one described above, may exist near the metadata encoder. Figure 10 It operates along the frequency axis. Other configurations are also possible.

[0180] about Figure 13 The example electronic device can be used as any component of the system described above. This device can be any suitable electronic device or apparatus. For example, in some embodiments, device 2200 is a mobile device, user equipment, tablet computer, computer, audio playback device, etc. For example, the device can be configured to implement the encoder and / or decoder or any functional block as described above.

[0181] In some embodiments, device 2200 includes at least one processor or central processing unit 2207. Processor 2207 may be configured to execute various program codes, such as the methods described herein.

[0182] In some embodiments, device 2200 includes at least one memory 2211. In some embodiments, at least one processor 2207 is coupled to memory 2211. Memory 2211 can be any suitable storage component. In some embodiments, memory 2211 includes a program code portion for storing program code that can be implemented on processor 2207. Furthermore, in some embodiments, memory 2211 may also include a data storage portion for storing data, such as data that has been processed or will be processed according to the embodiments described herein. The implemented program code stored in the program code portion and the data stored in the data storage portion can be retrieved by processor 2207 via memory processor coupling when needed.

[0183] In some embodiments, device 2200 includes a user interface 2205. In some embodiments, user interface 2205 may be coupled to processor 2207. In some embodiments, processor 2207 may control the operation of user interface 2205 and receive input from user interface 2205. In some embodiments, user interface 2205 may enable a user to input commands to device 2200, for example, via a keypad. In some embodiments, user interface 2205 may enable a user to obtain information from device 2200. For example, user interface 2205 may include a display configured to display information from device 2200 to a user. In some embodiments, user interface 2205 may include a touchscreen or touch interface that enables information to be input to device 2200 and further displayed to a user of device 2200. In some embodiments, user interface 2205 may be a user interface for communication.

[0184] In some embodiments, device 2200 includes an input / output port 2209. In some embodiments, input / output port 2209 includes a transceiver. In such embodiments, the transceiver may be coupled to processor 2207 and configured to communicate with other devices or electronic devices, such as via a wireless communication network. In some embodiments, the transceiver or any suitable transceiver or transmitter and / or receiver component may be configured to communicate with other electronic devices or devices via wired or wired coupling.

[0185] The transceiver can communicate with other devices via any suitable known communication protocol. For example, in some embodiments, the transceiver can use a suitable radio access architecture based on: Advanced Long Term Evolution (Advanced LTE, LTE-A) or New Radio (NR) (or may be referred to as 5G), Universal Mobile Telecommunications System (UMTS) Radio Access Network (UTRAN or E-UTRAN), Long Term Evolution (LTE, the same as E-UTRA), 2G networks (legacy network technologies), Wireless Local Area Network (WLAN or Wi-Fi), Global Microwave Access Interoperability (WiMAX), Bluetooth®, Personal Communication Services (PCS), ZigBee®, Wideband Code Division Multiple Access (WCDMA), systems using Ultra Wideband (UWB) technology, sensor networks, Mobile Ad Hoc Networks (MANET), Cellular Internet of Things (IoT) RAN and Internet Protocol Multimedia Subsystem (IMS), any other suitable options and / or any combination thereof.

[0186] The transceiver input / output port 1409 can be configured to receive signals.

[0187] In some embodiments, device 1400 may be used as at least part of a synthesis device. Input / output port 1409 may be coupled to headphones (which may be over-ear or non-over-ear headphones) and speakers.

[0188] Generally, various embodiments of the present invention can be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while others may be implemented in firmware or software, which may be executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. While various aspects of the invention may be shown and described using block diagrams, flowcharts, or some other graphical representation, it is well understood that, by way of non-limiting example, the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.

[0189] Embodiments of the present invention can be implemented by computer software executable by a mobile device's data processor, such as in a processor entity, or by hardware, or by a combination of software and hardware. Furthermore, it should be noted in this regard that any block of the logic flow shown can represent a program step, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software can be stored on physical media such as memory blocks implemented within a memory chip or processor, on magnetic media such as hard disks or floppy disks, and on optical media such as DVDs and their data variants, CDs, etc.

[0190] The memory can be of any type suitable for the local technical environment and can be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor can be of any type suitable for the local technical environment and can include one or more of the following as non-limiting examples: general-purpose computers, special-purpose computers, microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), gate-level circuits, and processors based on multi-core processor architectures.

[0191] Embodiments of the present invention can be practiced in various components such as integrated circuit modules. The design of integrated circuits is largely a highly automated process. Complex and powerful software tools can be used to transform logic-level designs into semiconductor circuit designs ready to be etched and formed on semiconductor substrates.

[0192] Programs such as those offered by Synopsys in Mountain View, California, and CadenceDesign in San Jose, California, use well-established design rules and pre-stored libraries of design modules to automatically route conductors and position components on semiconductor chips. Once the semiconductor circuit design is complete, the final design in a standardized electronic format (e.g., Opus, GDSII, etc.) can be sent to a semiconductor manufacturing facility or "manufacturing plant" for production.

[0193] As used in this application, the term "circuit system" may refer to one or more or all of the following:

[0194] (a) Hardware circuit implementation only (such as implementation in analog and / or digital circuit systems only) and

[0195] (b) A combination of hardware circuitry and software, such as (if applicable):

[0196] (i) A combination of analog and / or digital hardware circuitry and software / firmware, and

[0197] (ii) Any part of a hardware processor (including one or more digital signal processors), software, and memory (one or more) that works together to enable a device such as a mobile phone or server to perform various functions, and

[0198] (One or more) hardware circuits and / or (one or more) processors, such as (one or more) microprocessors or a portion thereof, which require software (e.g., firmware) to function, but may be absent when not needed to function.

[0199] This definition of circuit system applies to all uses of the term in this application, including in any claim. As another example, as used herein, the term circuit system also covers implementations of hardware circuitry or processors (or processors in general) and a portion thereof, along with their accompanying software and / or firmware. For instance, if applicable to a particular claim element, the term circuit system also covers baseband integrated circuits or processor integrated circuits for mobile devices, or similar integrated circuits in servers, cellular network devices, or other computing or network devices.

[0200] The term “non-transient” as used in this article refers to a limitation on the medium itself (i.e., tangible, not signaling), rather than a limitation on the persistence of data storage (e.g., RAM vs. ROM).

[0201] As used herein, “at least one of the following: ” and “at least one of ” and similar wording (where a list of two or more elements is connected by “and” or “or”) means at least any one of these elements, or at least any two or more of these elements, or at least all of these elements.

[0202] The foregoing description provides a complete and informative description of exemplary embodiments of the invention through exemplary and non-limiting examples. However, various modifications and adaptations will become apparent to those skilled in the art when read in conjunction with the accompanying drawings and appended claims, given the foregoing description. Nevertheless, all such and similar modifications to the teachings of the invention will still fall within the scope of the invention as defined in the appended claims.

Claims

1. An apparatus comprising components for: For at least two sources within an audio scene, obtain ordered directional metadata parameters associated with the at least two sources. These directional metadata parameters identify the direction of arrival of the at least two sources and are arranged in a frame. The frame is arranged as a time-frequency tile grid about the time axis and frequency axis. Identify an ordering error regarding at least one time-frequency tile directional metadata parameter, the ordering error being configured to identify that: the directional metadata parameter of at least one neighboring time-frequency tile associated with another order index is more similar to and / or less different from the directional metadata parameter of the at least one neighboring time-frequency tile associated with the same order index; and The determined at least one time-frequency block directional metadata parameters are reordered to the other sequential index.

2. The apparatus of claim 1, wherein the component for determining an ordering error with respect to at least one time-frequency block directional metadata parameter is configured to: determine that the similarity measure between the at least one adjacent time-frequency block directional metadata parameter associated with another order index and the at least one time-frequency block directional metadata is greater than, equal to, or greater than the similarity measure between the at least one adjacent time-frequency block directional metadata parameter associated with the same order index and the at least one time-frequency block directional metadata.

3. The apparatus of claim 1, wherein the component for determining an ordering error regarding at least one time-frequency block directional metadata parameter is configured to: determine that the difference measure between the at least one adjacent time-frequency block directional metadata parameter associated with another order index and the at least one time-frequency block directional metadata is less than, equal to, or less than the similarity measure between the at least one adjacent time-frequency block directional metadata parameter associated with the same order index and the at least one time-frequency block directional metadata.

4. The apparatus according to any one of claims 1 to 3, wherein the at least one adjacent time-frequency block directional metadata parameter is at least one of the following: Previous time-frequency block directional metadata parameters; The directional metadata parameters of the time-frequency block at the succeeding time. The preceding frequency is a frequency block directivity metadata parameter. The succeeding frequency is a time-frequency block directivity metadata parameter. Previous time and frequency block directional metadata parameters; The time and frequency block directivity metadata parameters of the next time and frequency; The previous time and next frequency time-frequency block directivity metadata parameters; and The time-frequency block directional metadata parameters of the next time and the previous frequency.

5. The apparatus according to any one of claims 1 to 4, wherein the component for reordering the determined at least one time-frequency block directional metadata parameter to the other sequential index is used to: reassign the other sequential index for the determined at least one time-frequency block directional metadata parameter.

6. The apparatus of claim 5, wherein the component for reassigning the other sequential index for the determined at least one time-frequency block directional metadata parameter is used to: Determine which of the following is more similar and / or less different regarding another sequential index, at least one adjacent time-frequency block directional metadata parameter, and the at least one time-frequency block directional metadata parameter associated with the same sequential index; At least one sub-frame metadata parameter is reassigned to the determined other order.

7. A method for use in an apparatus, comprising: For at least two sources within an audio scene, obtain ordered directional metadata parameters, which are associated with the at least two sources. The directional metadata parameters identify the direction of arrival with respect to the at least two sources and are arranged in a frame, which is arranged as a time-frequency block grid with respect to the time axis and frequency axis. Identify an ordering error for at least one time-frequency block directional metadata parameter, the ordering error being configured to identify that: the directional metadata parameter of at least one adjacent time-frequency block associated with another order index is more similar to and / or less different from the directional metadata parameter of the at least one adjacent time-frequency block associated with the same order index; as well as The determined at least one time-frequency block directional metadata parameter is reordered to the other sequential index.

8. The method of claim 7, wherein determining an ordering error regarding at least one time-frequency block directivity metadata parameter comprises: The similarity metric between the at least one adjacent time-frequency block directional metadata parameter associated with another sequential index and the at least one time-frequency block directional metadata is determined to be greater than, equal to, or greater than the similarity metric between the at least one adjacent time-frequency block directional metadata parameter associated with the same sequential index and the at least one time-frequency block directional metadata.

9. The method of claim 7, wherein determining an ordering error regarding at least one time-frequency block directional metadata parameter comprises: The difference metric between the at least one adjacent time-frequency block directional metadata parameter associated with another sequential index and the at least one time-frequency block directional metadata is determined to be less than, equal to, or less than the similarity metric between the at least one adjacent time-frequency block directional metadata parameter associated with the same sequential index and the at least one time-frequency block directional metadata.

10. The method according to any one of claims 7 to 9, wherein the at least one adjacent time-frequency block directional metadata parameter is at least one of the following: Previous time-frequency block directional metadata parameters; The directional metadata parameters of the next time-frequency block; The frequency block directivity metadata parameters of the previous frequency; The frequency block directivity metadata parameters of the next frequency; Previous time and frequency block directional metadata parameters; The time and frequency block directivity metadata parameters of the next time and frequency; The previous time and next frequency time-frequency block directivity metadata parameters; and The time-frequency block directional metadata parameters of the next time and the previous frequency.

11. The method according to any one of claims 7 to 10, wherein reordering the determined at least one time-frequency block directional metadata parameter to the other sequential index comprises: The other sequential index is reassigned to the determined at least one time-frequency block directional metadata parameter.

12. The method of claim 11, wherein reassigning the other sequential index to the determined at least one time-frequency block directional metadata parameter comprises: Determine which of the following is more similar and / or less different regarding another sequential index, at least one adjacent time-frequency block directional metadata parameter, and the at least one time-frequency block directional metadata parameter associated with the same sequential index; as well as At least one subframe metadata parameter is reassigned to the determined other order.

13. An apparatus comprising at least one processor and at least one memory storing instructions, the instructions, when executed by the at least one processor, causing the system to perform at least: For at least two sources within an audio scene, obtain ordered directional metadata parameters, which are associated with the at least two sources. The directional metadata parameters identify the direction of arrival with respect to the at least two sources and are arranged in a frame, which is arranged as a time-frequency block grid with respect to the time axis and frequency axis. Identify an ordering error for at least one time-frequency block directional metadata parameter, the ordering error being configured to identify that: the directional metadata parameter of at least one adjacent time-frequency block associated with another order index is more similar to and / or less different from the directional metadata parameter of the at least one adjacent time-frequency block associated with the same order index; as well as The determined at least one time-frequency block directional metadata parameter is reordered to the other sequential index.

14. The apparatus of claim 13, wherein determining an ordering error with respect to at least one time-frequency block directional metadata parameter is performed by: determining that the similarity metric between the at least one adjacent time-frequency block directional metadata parameter associated with another order index and the at least one time-frequency block directional metadata is greater than, equal to, or greater than the similarity metric between the at least one adjacent time-frequency block directional metadata parameter associated with the same order index and the at least one time-frequency block directional metadata.

15. The apparatus of claim 13, wherein performing the determination of an ordering error with respect to at least one time-frequency block directional metadata parameter, further comprises performing: determining that the difference measure between the at least one adjacent time-frequency block directional metadata parameter associated with another order index and the at least one time-frequency block directional metadata is less than, equal to, or less than the similarity measure between the at least one adjacent time-frequency block directional metadata parameter associated with the same order index and the at least one time-frequency block directional metadata.

16. The apparatus according to any one of claims 13 to 15, wherein the at least one adjacent time-frequency block directional metadata parameter is at least one of the following: Previous time-frequency block directional metadata parameters; The directional metadata parameters of the next time-frequency block; The frequency block directivity metadata parameters of the previous frequency; The frequency block directivity metadata parameters of the next frequency; Previous time and frequency block directional metadata parameters; The time and frequency block directivity metadata parameters of the next time and frequency; The previous time and next frequency time-frequency block directivity metadata parameters; and The time-frequency block directional metadata parameters of the next time and the previous frequency.

17. The apparatus according to any one of claims 13 to 16, wherein performing the reordering of the determined at least one time-frequency block directional metadata parameter to the other sequential index is performed by: reallocating the other sequential index for the determined at least one time-frequency block directional metadata parameter.

18. The apparatus of claim 17, wherein the reassignment of the other sequential index for the determined at least one time-frequency block directional metadata parameter is performed: Determine which of the following is more similar and / or less different regarding another sequential index, at least one adjacent time-frequency block directional metadata parameter, and the at least one time-frequency block directional metadata parameter associated with the same sequential index; and At least one subframe metadata parameter is reassigned to the determined other order.

Citation Information

Patent Citations

  • An apparatus, method and computer program for audio signal processing

    EP3791605A1

  • Determination of spatial audio parameter encoding and associated decoding

    WO2019105575A1

  • Combining of spatial audio parameters

    WO2021130405A1

  • The reduction of spatial audio parameters

    WO2021250312A1