Encoding of frame-level asynchronous metadata

By detecting and addressing framing asynchronousness in spatial metadata, the encoding method optimizes time-frequency resolution, enhancing encoding efficiency and reducing data loss in immersive audio codecs.

JP2026511174APending Publication Date: 2026-04-10NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
NOKIA TECHNOLOGIES OY
Filing Date
2024-02-13
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing immersive audio codecs face challenges in efficiently encoding spatial metadata due to asynchronous framing issues, leading to suboptimal encoding modes and potential data loss when metadata subframes are not synchronized with the encoding process.

Method used

The proposed solution involves detecting framing asynchronousness by analyzing spatial metadata across multiple subframes, determining framing offsets, and adjusting the encoding mode to optimize time-frequency resolution, allowing for efficient encoding even when metadata is not synchronized.

Benefits of technology

This approach improves encoding efficiency by accurately identifying situations where lower time-resolution encoding can be used, reducing data loss and maintaining encoding quality without requiring synchronization with the framing of the stream.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026511174000001_ABST
    Figure 2026511174000001_ABST
Patent Text Reader

Abstract

The apparatus comprises means for acquiring at least one spatial metadata parameter associated with at least one audio signal, wherein the at least one spatial metadata parameter is located in a frame comprising at least two subframes; means for acquiring asynchronous information, wherein the asynchronous information is based on the asynchronousness between a sequence of at least two time subframes comprising similar values ​​and a frame for further processing of the at least one spatial metadata parameter; and means for processing a frame comprising at least two time subframes based on the asynchronous information so that a processed frame comprising at least two time subframes of the at least one spatial metadata parameter can be further processed.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates to an apparatus and method for encoding frame-level out-of-sync metadata. [Background technology]

[0002] Parametric spatial audio capture from inputs such as microphone arrays and other sources is a typical and effective choice for estimating a set of parameters from the input (microphone array signal), including the direction of sound in the frequency band and the ratio between the directional and non-directional portions of the captured sound in the frequency band. These parameters are known to well describe the perceptual spatial characteristics of the captured sound at the location of the microphone array. These parameters can be used in the synthesis of spatial acoustics accordingly for binaural headphones, speakers, or other formats such as ambisonics.

[0003] Therefore, in the frequency band, direction, and direct energy to total energy ratios and diffuse-to-total energy ratios are particularly effective parameterizations for spatial audio capture.

[0004] A parameter set consisting of frequency band directional parameters (indicating sound directivity) and frequency band energy ratio parameters can also be used as spatial metadata for audio codecs (which may also include other parameters such as surround coherence, spread coherence, directionality, and distance). For example, these parameters can be estimated from audio signals captured by a microphone array, and a stereo or mono transport audio signal can be generated from the microphone array signal transmitted via spatial metadata.

[0005] Immersive audio codecs are implemented supporting a wide range of operating points, from low-bitrate operation to transparency. An example of such a codec is the Immersive Voice and Audio Service (IVAS) codec, designed for use over communication networks such as 3GPP 4G / 5G networks, including use in immersive services such as immersive voice and sound for virtual reality (VR). This audio codec is expected to handle the encoding, decoding, and rendering of speech, music, and general audio. It is further expected to support channel-based audio, object-based audio, and scene-based audio inputs, including spatial information about the sound field and sound source. The codec is also expected to operate with low latency to enable conversational services and to support high error robustness under various transmission conditions.

[0006] The transport audio signal may be encoded, for example, using the IVAS audio core codec, or using an AAC (Advanced Audio Coding) or EVS (Enhanced Voice Service) encoder. The decoder may decode the audio signal into a PCM (Pulse Code Modulation) signal and process the sound in the frequency band (using spatial metadata) to obtain a spatial output, such as a binaural output.

[0007] The aforementioned immersive audio codecs are particularly well-suited for encoding spatial audio captured from microphone arrays (e.g., mobile phones, VR cameras, standalone microphone arrays). However, such encoders may also have other input types, such as speaker signals, audio object signals, and ambisonic signals. [Overview of the Initiative]

[0008] According to a first embodiment, the provided apparatus includes means for acquiring at least one spatial metadata parameter associated with at least one audio signal, wherein the at least one spatial metadata parameter is located in a frame comprising at least two subframes; means for acquiring asynchrony information, wherein the asynchrony information is based on the asynchronous relationship between a sequence of at least two time subframes comprising similar values ​​and a frame for further processing of the at least one spatial metadata parameter; and means for processing a frame comprising at least two time subframes based on the asynchrony information such that a processed frame comprising at least two time subframes of the at least one spatial metadata parameter can be further processed.

[0009] A sequence of at least two time subframes containing similar values ​​may be located within the same frame or across two consecutive frames.

[0010] The means may be for further processing a processed frame that includes at least two time subframes of at least one spatial metadata parameter.

[0011] Means for further processing a processed frame containing at least two time subframes of at least one spatial metadata parameter may be for encoding a processed frame containing at least two time subframes of at least one spatial metadata parameter.

[0012] The means may further include a means for acquiring at least one audio signal and a means for encoding at least one audio signal.

[0013] The means for obtaining asynchronous information may be one of the following: one for analyzing at least one spatial metadata parameter to determine the asynchronous information; one for receiving the asynchronous information from at least one further device that has determined the asynchronous information; and one for receiving the asynchronous information from at least one further device that has analyzed at least one spatial metadata parameter to determine the asynchronous information.

[0014] The means for obtaining asynchronous information may be for obtaining an offset value that identifies the time difference between a sequence of at least two time subframes and the encoded frame, the time difference being obtained as the number of subframes.

[0015] A means for processing a sequence of at least two time subframes based on asynchronous information may be for processing at least one spatial metadata parameter based on an offset value.

[0016] A means for processing at least one spatial metadata parameter based on an offset value may be for generating a single spatial metadata parameter subframe to represent at least two time subframes based on processing a sequence of at least two time subframes.

[0017] Means for generating a single spatial metadata parameter subframe to represent at least two time subframes based on the processing of a sequence of at least two time subframes may be one of the following: selecting one of at least two subframes to be a single spatial metadata parameter subframe, and determining the aggregation of at least two subframes to be a single spatial metadata parameter subframe.

[0018] Means for determining the aggregation of at least two subframes so that they constitute a single spatial metadata parameter subframe may be one of the following: determining an aggregation function based on the vector mean of the directional components of at least one spatial metadata parameter; or determining an aggregation function based on the vector mean of the directional components of at least one spatial metadata parameter weighted by signal energy weighting.

[0019] A means for obtaining asynchronous information may be for obtaining an encoding mode that identifies an encoding mode for encoding at least one spatial metadata parameter based on the asynchronous information.

[0020] Means for analyzing at least one spatial metadata parameter to determine asynchronous information may include, for the current frame, analyzing at least one spatial metadata parameter associated with at least one audio signal to determine the number of mutually similar subframes from the start of the current frame and the number of mutually similar subframes from the end of the current frame, and determining asynchronous information based on the number of mutually similar subframes from the start of the current frame and / or the number of mutually similar subframes from the end of the current frame.

[0021] The means for analyzing at least one spatial metadata parameter to determine asynchronous information may further include means for analyzing whether the last subframe of the previous frame and the first subframe of the current frame are similar subframes, and means for obtaining the number of mutually similar subframes for the previous frame. Also, the means for determining asynchronous information based on the number of mutually similar subframes from the start of the current frame and / or the number of mutually similar subframes from the end of the current frame may further include determining asynchronous information based on whether the last subframe of the previous frame and the first subframe of the current frame are similar subframes, and the number of mutually similar subframes for the previous subframe.

[0022] According to a second aspect, a method for an apparatus includes obtaining at least one spatial metadata parameter associated with at least one audio signal, where the at least one spatial metadata parameter is arranged in a frame including at least two subframes; obtaining asynchronous information, where the asynchronous information is based on the asynchrony between a sequence of at least two time subframes including similar values and a frame for further processing the at least one spatial metadata parameter; and processing a frame including at least two time subframes based on the asynchronous information such that a processed frame including at least two time subframes of the at least one spatial metadata parameter can be further processed.

[0023] The sequence of at least two time subframes including similar values can be one of those located within the same frame and those located across two consecutive frames.

[0024] The method may further include processing a processed sequence of at least two time sub-frames of at least one spatial metadata parameter.

[0025] Further processing a frame processed with at least one spatial metadata parameter including at least two time sub-frames may include encoding a processed frame including at least two time sub-frames of at least one spatial metadata parameter.

[0026] The method may further include obtaining at least one audio signal and encoding at least one audio signal.

[0027] Obtaining asynchronous information may include one of analyzing at least one spatial metadata parameter to determine asynchronous information, receiving asynchronous information from at least one additional device that has determined the asynchronous information, and receiving asynchronous information from at least one additional device that has analyzed at least one spatial metadata parameter to determine the asynchronous information.

[0028] Obtaining asynchronous information may include obtaining an offset value that identifies a time difference between a sequence of at least two time sub-frames and an encoded frame, where the time difference is obtained as the number of sub-frames.

[0029] Processing a sequence of at least two time sub-frames based on asynchronous information may include processing at least one spatial metadata parameter based on the offset value.

[0030] Processing at least one spatial metadata parameter based on the offset value may include generating a single spatial metadata parameter sub-frame to represent at least two time sub-frames based on the processing of the sequence of at least two time sub-frames.

[0031] Generating a single spatial metadata parameter subframe to represent at least two time subframes based on the processing of a sequence of at least two time subframes may include one of selecting one of at least two subframes to be a single spatial metadata parameter subframe, or determining the aggregation of at least two subframes to be a single spatial metadata parameter subframe.

[0032] Determining the aggregation of at least two subframes so that they constitute a single spatial metadata parameter subframe may include one of the following: determining the aggregation function based on the vector mean of the directional components of at least one spatial metadata parameter; and determining the aggregation function based on the vector mean of the directional components of at least one spatial metadata parameter weighted by signal energy weighting.

[0033] Obtaining asynchronous information may include obtaining an encoding mode that identifies an encoding mode for encoding at least one spatial metadata parameter based on the asynchronous information.

[0034] Analyzing at least one spatial metadata parameter to determine asynchronous information may include, for the current frame, analyzing at least one spatial metadata parameter associated with at least one audio signal to determine the number of similar subframes from the start of the current frame and the number of similar subframes from the end of the current frame, and determining asynchronous information based on the number of similar subframes from the start of the current frame and / or the number of similar subframes from the end of the current frame.

[0035] Analyzing at least one spatial metadata parameter to determine asynchronous information may further include analyzing whether the last subframe of the previous frame and the first subframe of the current frame are similar subframes, and obtaining the number of similar subframes for the previous frame; determining asynchronous information based on the number of similar subframes from the start of the current frame and / or the number of similar subframes from the end of the current frame further includes determining asynchronous information based on whether the last subframe of the previous frame and the first subframe of the current frame are similar subframes, as well as the number of similar subframes for the previous subframe.

[0036] According to a third aspect, the device is provided, which includes at least one processor and at least one memory that stores instructions causing the system, when executed by the at least one processor, to acquire at least one spatial metadata parameter associated with at least one audio signal, wherein the at least one spatial metadata parameter is located in a frame comprising at least two subframes; to acquire asynchronous information, wherein the asynchronous information is based on asynchronousness between a sequence of at least two time subframes comprising similar values ​​and a frame for further processing of the at least one spatial metadata parameter; and to process a frame comprising at least two time subframes based on the asynchronous information so that a processed frame comprising at least two time subframes of the at least one spatial metadata parameter can be further processed.

[0037] A sequence of at least two time subframes containing similar values ​​is one of two things: one located within the same frame, and one located across two consecutive frames.

[0038] The device may be configured to further process the processed frame, which includes at least two time subframes of at least one spatial metadata parameter.

[0039] An apparatus configured to further process a processed frame containing at least two time subframes of at least one spatial metadata parameter may further encode the processed frame containing at least two time subframes of at least one spatial metadata parameter.

[0040] The device may further perform the tasks of acquiring at least one audio signal and encoding at least one audio signal.

[0041] A sequence of at least two subframes may be further arranged as a frame containing two or more subframes.

[0042] A device configured to acquire asynchronous information may perform one of the following: analyze at least one spatial metadata parameter to determine the asynchronous information; receive the asynchronous information from at least one further device that has determined the asynchronous information; or receive the asynchronous information from at least one further device that has analyzed at least one spatial metadata parameter to determine the asynchronous information.

[0043] A device configured to acquire asynchronous information may further acquire an offset value that identifies the time difference between a sequence of at least two time subframes and the encoded frame, the time difference being acquired as the number of subframes.

[0044] An instrument configured to process a sequence of at least two time subframes based on asynchronous information may further configure to process at least one spatial metadata parameter based on an offset value.

[0045] An apparatus configured to process at least one spatial metadata parameter based on an offset value may further configure to generate a single spatial metadata parameter subframe to represent at least two time subframes based on the processing of a sequence of at least two time subframes.

[0046] An apparatus configured to generate a single spatial metadata parameter subframe to represent at least two time subframes based on the processing of a sequence of at least two time subframes may perform one of the following: select one of the at least two subframes to be a single spatial metadata parameter subframe, or determine the aggregation of the at least two subframes to be a single spatial metadata parameter subframe.

[0047] An apparatus configured to determine the aggregation of at least two subframes so that they constitute a single spatial metadata parameter subframe may further configure to determine an aggregation function based on the vector mean of the directional components of at least one spatial metadata parameter, and to determine an aggregation function based on the vector mean of the directional components of at least one spatial metadata parameter weighted by signal energy weighting.

[0048] A device configured to acquire asynchronous information may further acquire an encoding mode that identifies an encoding mode for encoding at least one spatial metadata parameter based on the asynchronous information.

[0049] An apparatus configured to analyze at least one spatial metadata parameter to determine asynchronous information may further configure to analyze at least one spatial metadata parameter associated with at least one audio signal for the current frame to determine the number of mutually similar subframes from the start of the current frame and the number of mutually similar subframes from the end of the current frame, and to determine asynchronous information based on the number of mutually similar subframes from the start of the current frame and / or the number of mutually similar subframes from the end of the current frame.

[0050] A device configured to analyze at least one spatial metadata parameter to determine asynchronous information may further analyze whether the last subframe of the previous frame and the first subframe of the current frame are similar subframes, and obtain the number of similar subframes for the previous frame. A device configured to determine asynchronous information based on the number of similar subframes from the start of the current frame and / or the number of similar subframes from the end of the current frame may further determine asynchronous information based on whether the last subframe of the previous frame and the first subframe of the current frame are similar subframes, and the number of similar subframes for the previous subframe.

[0051] According to a fourth aspect, the apparatus is provided, comprising: an acquisition circuit configured to acquire at least one spatial metadata parameter associated with at least one audio signal, wherein the at least one spatial metadata parameter is located in a frame comprising at least two subframes; an acquisition circuit configured to acquire asynchronous information, wherein the asynchronous information is based on asynchronousness between a sequence of at least two time subframes comprising similar values ​​and a frame for further processing of the at least one spatial metadata parameter; and a processing circuit configured to process a frame comprising at least two time subframes based on the asynchronous information so that a processed frame comprising at least two time subframes of the at least one spatial metadata parameter can be further processed.

[0052] According to a fifth aspect, a computer program is provided [or a computer-readable medium including program instructions] that causes a device to perform at least the following: obtaining at least one spatial metadata parameter associated with one audio signal, wherein the at least one spatial metadata parameter is located in a frame comprising at least two subframes; obtaining asynchronous information, wherein the asynchronous information is based on asynchronousness between a sequence of at least two time subframes comprising similar values ​​and a frame for further processing of the at least one spatial metadata parameter; and processing a frame comprising at least two time subframes based on the asynchronous information so that a processed frame comprising at least two time subframes of the at least one spatial metadata parameter can be further processed.

[0053] According to the sixth aspect, a non-temporary computer-readable medium is provided, which includes program instructions for causing a device to perform at least the following: obtaining at least one spatial metadata parameter associated with at least one audio signal, wherein the at least one spatial metadata parameter is located in a frame comprising at least two subframes; obtaining asynchronous information, wherein the asynchronous information is based on asynchronousness between a sequence of at least two time subframes comprising similar values ​​and a frame for further processing of the at least one spatial metadata parameter; and processing a frame comprising at least two time subframes based on the asynchronous information so that a processed frame comprising at least two time subframes of the at least one spatial metadata parameter can be further processed.

[0054] According to a seventh aspect, the apparatus is provided, comprising: means for acquiring at least one spatial metadata parameter associated with at least one audio signal, wherein the at least one spatial metadata parameter is located in a frame comprising at least two subframes; means for acquiring asynchronous information, wherein the asynchronous information is based on asynchronousness between a sequence of at least two time subframes comprising similar values ​​and a frame for further processing of the at least one spatial metadata parameter; and means for processing a frame comprising at least two time subframes based on the asynchronous information such that a processed frame comprising at least two time subframes of the at least one spatial metadata parameter can be further processed.

[0055] According to the eighth aspect, a computer-readable medium is provided, which includes program instructions for causing a device to perform at least the following: obtaining at least one spatial metadata parameter associated with at least one audio signal, wherein the at least one spatial metadata parameter is located in a frame comprising at least two subframes; obtaining asynchronous information, wherein the asynchronous information is based on asynchronousness between a sequence of at least two time subframes comprising similar values ​​and a frame for further processing of the at least one spatial metadata parameter; and processing a frame comprising at least two time subframes based on the asynchronous information so that a processed frame comprising at least two time subframes of the at least one spatial metadata parameter can be further processed.

[0056] The apparatus includes means for performing actions in the manner described above.

[0057] The device is configured to perform actions in the manner described above.

[0058] A computer program contains program instructions that cause a computer to perform the methods described above.

[0059] A computer program product stored on a medium can cause the device to perform the methods described herein.

[0060] Electronic devices may include the apparatus described herein.

[0061] The chipset may include the devices described herein.

[0062] The embodiments of this application aim to address problems related to the state of the art.

[0063] For a better understanding of this application, references to the attached drawings are made here as an example. [Brief explanation of the drawing]

[0064] [Figure 1] A schematic diagram of the equipment used for extracting MASA metadata is shown. [Figure 2] A schematic diagram of an exemplary MASA metadata frame subframe structure is shown below. [Figure 3] A schematic representation of the time-frequency structure of an exemplary MASA metadata frame is shown. [Figure 4] An illustrative method for adjusting and reconstructing metadata resolution is outlined below. [Figure 5] This document presents exemplary application scenarios demonstrating tandem coding and multistream coupling. [Figure 6] This shows an exemplary delay implemented by the encoder / decoder regarding metadata. [Figure 7] This illustrates an exemplary asynchronous situation when encoding / decoding metadata. [Figure 8] A schematic diagram illustrates an exemplary system of an apparatus suitable for carrying out several embodiments. [Figure 9] Several embodiments schematically illustrate alternatives to timing in metadata framing. [Figure 10] The potential metadata contributions using different framing offsets between IVAS and MASA framing in several embodiments are schematically shown. [Figure 11] A schematic representation of known metadata analyzers and encoders is provided below. [Figure 12] This section provides a schematic overview of the metadata subframes available when constructing transport data in history mode and history-free or history-less mode. [Figure 13] Several encoders suitable for employing certain embodiments are schematically shown. [Figure 14] Figure 13 shows a flowchart illustrating the operation of an exemplary encoder in several embodiments. [Figure 15] Figure 13 shows a flowchart illustrating the operation of the asynchronous analyzer according to several embodiments. [Figure 16] A schematic diagram of a decoder suitable for employing several embodiments is shown below. [Figure 17] Figure 13 shows a flowchart illustrating the operation of an example of a further asynchronous analyzer, based on several embodiments. [Figure 18] Further exemplary systems of apparatus suitable for carrying out several embodiments are schematically shown. [Figure 19] An exemplary device suitable for implementing the apparatus shown in the previous figure is shown. [Modes for carrying out the invention]

[0065] The following sections further describe suitable devices and possible mechanisms for encoding parametric spatial audio signals, including transport audio signals and spatial metadata. As mentioned above, immersive audio codecs (such as 3GPP IVAS) are planned, supporting a large number of operating points, ranging from low bitrate operation to transparency.

[0066] Metadata-Assisted Spatial Audio (MASA) is an example of a parametric spatial audio format and representation suitable as an input format for IVAS.

[0067] Audio representation can be thought of as consisting of "N channels + spatial metadata." This is a scene-based audio format particularly well-suited for spatial audio capture on practical devices such as smartphones. The idea is to describe the acoustic scene in terms of time-varying and frequency-varying sound source direction, and, for example, energy ratio. Acoustic energy within a scene that is not defined (not described) by direction is described as diffusion (coming from all directions).

[0068] As described above, spatial metadata associated with an audio signal may include multiple parameters (multiple directions and, associated with each direction (or direction value), such as direct-to-total energy ratio, spread coherence, distance, etc.) on a time-frequency tile basis. Spatial metadata can also be considered non-directional (such as surround coherence, diffuse energy to total energy ratio, residual energy to total energy ratio), but when combined with directional parameters, it may include, or be associated with, other parameters that can be used to define the characteristics of an audio scene. For example, a reasonable design choice that can produce a good output is one in which the spatial metadata is determined to include one or more directions (and, associated with each direction, such as direct-to-total ratio, spread coherence, distance value, etc.) for each time-frequency subframe.

[0069] With respect to Figure 1, an exemplary MASA analyzer 101 is shown. The MASA analyzer 101 is configured to receive and analyze an input audio signal 100 in order to generate a transport audio signal 102 and spatial metadata 104.

[0070] Examples of MASA spatial metadata are presented in the table below. These values ​​are available for each time-frequency tile. In some embodiments, a frame is subdivided into 24 frequency bands and 4 time subframes. In other embodiments, other divisions of frequency and time may be employed. Furthermore, in some embodiments, the frame size is 20 ms (and therefore the time subframe is 5 ms), (as is done, for example, in IVAS). However, similarly, other frame lengths may be employed in other embodiments. In some embodiments, the MASA analyzer is configured to determine one or two directions for each time-frequency tile (i.e., for each time-frequency tile, there are one or two direction indices, a direct energy to total energy ratio, and a spread coherence parameter). However, in some embodiments, the analyzer is configured to generate more than two directions for a time-frequency tile.

[0071] [Table 1]

[0072] The MASA stream can be rendered to various outputs, such as multi-channel speaker signals (e.g., 5.1) or binaural signals.

[0073] As mentioned above, the frame size of IVAS is 20ms. Figure 2 shows an example of a representative frame structure in which metadata frame 201 contains four time subframes, each 5ms long. Figure 2 shows, for example, metadata subframe 4 200 of the previous frame, followed by metadata subframe 1 202, metadata subframe 2 204, metadata subframe 3 206, and metadata subframe 4 208 of the current metadata frame 201. This is followed by metadata subframe 1 210 of the subsequent or next frame.

[0074] Furthermore, the IVAS codec is expected to operate at a variety of bitrates, ranging from very low bitrates (e.g., 13.2kbps) to relatively high bitrates (e.g., 512kbps or even 768kbps). Since the raw bitrate of MASA metadata is approximately 300-500kbps (depending on whether there are one or two concurrently encoded directions), the metadata will be significantly compressed (especially at the lowest bitrates).

[0075] One form of compression may be a method of reducing the temporal resolution and / or frequency resolution of metadata (which may be employed in conjunction with other methods for compressing data).

[0076] As described above and as shown in Figure 3, a raw high-resolution example metadata frame 300 can contain 24 frequency bands on the frequency axis 301 and 4 time subframes (subframes 1-4 302, 304, 306, 308) on the time axis 303, totaling 96 time-frequency tiles (also called TF tiles).

[0077] With respect to Figure 4, a method is shown for reducing the number of time-frequency tiles transmitted, thereby significantly reducing the required bitrate. Such a method is described in UKIPO patent applications 1919130.3 and 1919131.1, which present a method for combining metadata to reduce multiple frequency bands and / or time subframes to fewer frequency bands and / or time subframes.

[0078] For example, as shown in Figure 4, depending on the bitrate, 5 to 24 frequency bands and 1 to 4 subframes may be transmitted.

[0079] Therefore, such a method includes a metadata resolution selector 401 and an adjuster 403 configured to generate a low-time-resolution metadata frame 404 with 1 sf, low-time-resolution (or high-frequency-resolution), and a high-time-minute-resolution metadata frame 406 with 4 sf, high-time-resolution (or low-frequency-resolution).

[0080] The metadata is unpacked by a metadata unpacker 405, which is configured in the decoder to unpack the common representation resolution metadata frame 408 (a frame with an unknown underlying TF resolution).

[0081] Since MASA streams can be generated from various types of devices (e.g., microphone arrays on mobile devices and dedicated ambisonic microphone arrays such as Eienmike), the methods used to determine spatial metadata can vary significantly between embodiments. Some methods may have high temporal resolution, while others may have low temporal resolution (or high frequency resolution).

[0082] To improve encoding efficiency for both types of time-frequency resolution, it has been suggested that MASA metadata can be encoded in two different modes, as shown in PCT application WO2021250312 and also shown in Figure 4 above. The first metadata frame resolution has one with a higher frequency resolution (1 sf) but only one time subframe per frame, while the other metadata frame resolution has a higher time resolution (or lower frequency resolution) (4 sf) but maintains four time subframes.

[0083] In this example, the former mode (1sf) is selected when the encoder receives spatial metadata that is determined or detected to be identical (or substantially identical or similar) across all subframes of a frame.

[0084] If the spatial metadata is not identical (or substantially identical or similar) across all subframes, the latter mode (4sf) is adopted.

[0085] For example, at a given bitrate, the former mode (1sf) may transmit 18 frequency bands and 1 subframe (in other words, a total of 18 TF tiles), while the latter mode (4sf) may transmit 5 frequency bands and 4 subframes (in other words, a total of 20 TF tiles), which is roughly equivalent to a similar size of data transmitted at the same overall bitrate.

[0086] PCT application WO2019105575 proposes using variable input metadata time-frequency resolution. This achieves a similar trade-off to the method in PCT application WO2021250312, but the decision is made outside the codec and may be based on a specific capture algorithm for the microphone array being used.

[0087] Therefore, the method described above demonstrates a way to maintain encoding quality, in which the temporal and frequency resolutions are tuned or adjusted with respect to the audio input.

[0088] As described with respect to these methods, the low-time-resolution (or high-frequency-resolution) mode is selected when all subframes of a frame have identical (or substantially identical, or at least sufficiently similar) data. A microphone front-end creating a MASA stream may create metadata to ensure this is the case, but it cannot guarantee that subframes will always be synchronized. In other words, metadata framing can exhibit asynchronous behavior.

[0089] An example of an application scenario in which asynchronous metadata framing may be introduced is shown with respect to Figure 5. For example, on the left side of Figure 5, a tandem encoder 501 is shown configured to receive an input transport signal 102 and metadata 104 as input, where the metadata (and transport audio) stream has already been encoded and then decoded, and is provided as input to a second encoder, which is configured to produce an output transport signal 502 and metadata 504. In this example, it cannot be guaranteed that the framing (or subframe grouping) of this second encoding performed by the tandem encoder 501 will match that of the previous encoder.

[0090] Furthermore, the right side of Figure 5 shows an MCU (Multipoint Control Unit) application, in which the multistream coupler 503 is configured to combine multiple transport streams 1021, 1022 and metadata streams 1041, 1042 into a single stream to generate output transport signals 512 and metadata 514. In this case as well, it cannot be guaranteed that the framing synchronization in the MCU is the same for all combined streams.

[0091] With respect to Figure 6, timing delays that may be experienced in both scenarios within the exemplary system are shown. The MASA stream (including the input transport audio signal 102 and metadata 104) may be acquired by system 601. System 601 can be thought of as including an encoder configured to generate a bitstream and a decoder configured to receive the bitstream and to provide an exemplary output including the transport audio signal and spatial metadata. These are shown coupled together in a single schematic system block. In some embodiments of system 601, the system can be thought of as including an audio core encoder 603 configured to implement an IVAS codec (e.g., having a delay of 32 ms in the MASA format) configured to generate the decoded transport audio signal 602. This exemplary 32 ms delay of the MASA includes a 20 ms framing delay and a 12 ms look-ahead delay in the audio core encoder. This look-ahead delay corresponds to a delay of (approximately) two subframes (2 × 5 ms = 10 ms ≈ 12 ms). The metadata delay 605 may be applied to the metadata to generate delayed spatial metadata 606 in an attempt to realign the decoded transport audio signal and spatial metadata by applying a look-ahead delay of two subframes to the metadata.

[0092] Therefore, for example, input condition 610 is an input condition in which the audio signal and spatial metadata are aligned.

[0093] Next, following the application of the encoder / decoder, the post-coded situation 620 shows that the decoded transport audio signal is delayed, for example, 12 ms relative to the spatial metadata 621.

[0094] Next, following the application of the metadata delay, the post-delay situation 630 shows that the application of the two subframe look-ahead delays 631 to the metadata is nearly identical to the decode delay 621, and therefore nearly resynchronizes the decoded transport audio signal and the delayed spatial metadata.

[0095] This delay solution for metadata subframes may cause synchronization issues, although this does not cause significant perceptual problems. However, the introduction of a framing offset relative to the “global” clock, in other words, the timing of delayed spatial metadata and any further encoding operations, cannot guarantee that, for further stages, the delayed spatial metadata in the bitstream is a “frame” synchronized with the frame clock of the further stages. For example, there is an encoding mode, as described herein, in which a low time resolution (or high frequency resolution) encoding mode is selected when the current frame has substantially similar values, and it may be determined that the encoded frame of metadata has identical values ​​for all subframes, even if the later encoded frame is not synchronized with the delayed metadata frame.

[0096] For example, a system may exist in which the output of the decoding stage is audio with a 12-millisecond delay and metadata with a 10-millisecond delay, and the second encoding stage operates with the same global absolute framing clock as the first encoding round. Due to the delay, the second encoding stage receives the delayed audio and metadata, even though framing begins immediately. In that case, the first 12ms of the audio will be zero, and the first two metadata subframes will be empty before the decoded real data becomes available.

[0097] Therefore, the first frame has a half-frame offset, with half being zeros and the other half being real data (the first half of the original first frame). The second frame of the second encoding has a half-frame offset, with the second half of the original first frame and the first half of the original second frame.

[0098] The potential asynchronous nature of this framing is illustrated in relation to Figure 7. Figure 7 shows, for example, the original metadata frame timing, and includes a first "original" metadata frame 1 701 containing metadata subframes 1 711, 2 713, 3 715, and 4 717, and a second "original" metadata frame 2 703 containing metadata subframes 1 721, 2 723, 3 725, and 4 727.

[0099] Next, all four potential offset values ​​that can occur for the four subframes used in MASA are shown.

[0100] For example, in a situation where offset = 0, the first re-encoded metadata frame, frame 1731, and the second re-encoded metadata frame, frame 2732, are perfectly aligned or synchronized with the first "original" metadata frame 1701 and the second "original" metadata frame 2703, respectively.

[0101] Situation 741 is also shown for a subframe with an offset of +1 (or -3) applied to frame 1 (or frame 2), in which case the re-encoded metadata frame is shifted by 1 subframe relative to the first "original" metadata frame 1 701, or by -3 subframes relative to the second "original" metadata frame 2 703.

[0102] Furthermore, scenario 751 is shown for a subframe with an offset of +2 (or -2) with respect to frame 1 (or frame 2), in which case the re-encoded metadata frame is shifted by 2 subframes relative to the first "original" metadata frame 1 701, or by -2 subframes relative to the second "original" metadata frame 2 703.

[0103] Next, scenario 761 is shown of a subframe with an offset of +3 (or -1) with respect to frame 1 (or frame 2), in which case the re-encoded metadata frame is shifted by 3 subframes relative to the first "original" metadata frame 1 701, or by -1 subframe relative to the second "original" metadata frame 2 703.

[0104] Furthermore, as described above, there may be tandem coding situations in which the server is configured to re-encode the MASA stream after editing and mixing multiple streams, either before storing the signal or before sending the signal to the user.

[0105] The encoder checks whether the input is identical (or similar) across all subframes, and if it determines that they are not identical (or similar) due to offsets, it is configured to encode spatial metadata using a "high time resolution mode" of 4sf, in which case four subframes are encoded instead of the "low time resolution (or high frequency resolution) mode" of 1sf. As a result, fewer frequency bands are encoded, for example, only 5 instead of 18, and frequency resolution is compromised.

[0106] Therefore, the objective of the following embodiments described herein is to improve the (MASA) metadata encoding method to detect and address framing asynchronousness (offset) from data. This may manifest, for example, in a way that better identifies or determines situations in which a lower time-resolution (or higher frequency-resolution) encoding mode may be employed rather than reverting to a higher time-resolution 4sf encoding mode, even when a signal is asynchronous.

[0107] Accordingly, the embodiments relate to the encoding of parametric spatial audio (in other words, audio signals and spatial metadata), in which the spatial metadata is encoded in frames containing multiple frequency bands and multiple subframes. In some embodiments, this is carried out by apparatus and methods that enable the encoding of spatial metadata with optimized time-frequency resolution even when the spatial metadata is not synchronized with the framing of the stream (i.e., similar / identical data may be found in subframes of different frames).

[0108] Furthermore, in some embodiments, this can be achieved by obtaining one or more frames of spatial metadata, comparing the values ​​of spatial metadata in subframes of one or more frames to determine whether the metadata framing has an offset, selecting an encoding mode based on the comparison, potentially determining new spatial metadata based on the metadata values ​​of the subframes, and encoding this spatial metadata.

[0109] In some embodiments, the apparatus, This involves analyzing spatial metadata, detecting framing (i.e., subframe grouping) asynchronousness, and determining framing offsets. Based on the detected offset and spatial metadata subframe, determine new spatial metadata for this frame, It is configured to implement a method that includes the following.

[0110] In some embodiments, the determination of framing asynchronousness may be performed in a device separate from the device that performs the determination of new spatial metadata. In other words, in some embodiments, the device or method is one in which new spatial metadata is determined for a frame, in which case the determination is made based on an input in which it has been determined that the frame is asynchronous with respect to subframes. This further device may determine the asynchronousness information, for example, based on (or by performing an analysis of the spatial metadata after processing) in the further device. In such embodiments, the asynchronousness information, such as an offset value, is determined as a "hardcoded" value and provided to the encoder.

[0111] In some embodiments, the encoder may be configured to analyze metadata subframes of the current and previous frames and utilize the presence of multiple (e.g., four) subframes with similar metadata to indicate a potential low-time resolution (or high-frequency resolution, 1sf mode coding) grouping.

[0112] In the following example, there are four subframes within one frame, but this is a specific example, and the embodiment can be generalized to N subframes within one frame.

[0113] In some embodiments, the apparatus and method analyze spatial metadata from the current and previous frames to detect any framing asynchronousness. In such embodiments, once the method detects a non-zero framing offset and determines that the framing mode should be low time resolution (or high frequency resolution), 1sf encoding mode, the encoder may be configured to apply an aggregation / interpolation function to calculate or determine new spatial metadata values ​​from the current values ​​of the spatial metadata. These new spatial metadata values ​​are then encoded in place of the original spatial metadata values.

[0114] In such embodiments, the decoder in question can be any suitable known decoder (in other words, the decoder is not modified).

[0115] In some embodiments, the asynchronous analyzer is configured to use spatial metadata from the current frame being encoded, without memory for analyzing previous metadata. In these embodiments, less working memory is used for analysis, but the reliability of asynchronous detection may be reduced.

[0116] In some embodiments, there may be combinations of these methods in which the method further selects whether or not the analysis uses memory of the previous frame, depending on any encoder complexity limitations. For example, memory (previous decisions) may be used in situations where the additional complexity of asynchronous detection using memory is not an issue, but history-independent analysis may be employed when encoder complexity is constrained.

[0117] With respect to Figure 8, an exemplary system is shown in which several embodiments may be implemented. The inputs are a transport audio signal 102 and spatial metadata 104. The transport audio signal 102 and spatial metadata 104 are passed to an encoder 801, which generates an encoded bitstream 802. The encoded bitstream 802 is received by a decoder 803, which is configured to generate a spatial audio output 804.

[0118] As described above, the system input, transport audio signal 102, and spatial metadata 104 may be acquired in the form of a MASA stream. The MASA stream may originate, for example, from a mobile device (including a microphone array), or, as an alternative example, be created by an audio server that is potentially processing the MASA stream in some way.

[0119] In some embodiments, the encoder 801 may further be an IVAS encoder.

[0120] In some embodiments, the decoder 803 may be configured to directly output a spatial audio output 804 which is rendered by an external renderer or edited / processed by an audio server. In some embodiments, the decoder 803 includes a suitable renderer, which is configured to render the output in a suitable format such as a binaural audio signal or a multi-channel speaker signal (5.1 or 7.1+4 channel format), and these formats are also examples of spatial audio output 804.

[0121] In some embodiments, the spatial metadata 104 arrives at the encoder 801 as a sequence of subframes without indication of the original framing (grouping of subframes). If the framing of the metadata-providing process is synchronized with the framing of the encoding process, there are no potential problems. However, if some asynchronousity exists, for example from the tandem encoding described above or other application scenarios, the encoding may be suboptimal in the sense that an incorrect encoding mode is selected, resulting in some data loss.

[0122] Figures 9 and 10 illustrate exemplary offset scenarios in framing asynchronous. For example, Figure 9 shows alternative and underlying MASA metadata framing for metadata framing asynchronous in the case of four subframes. It shows the previous metadata frame 900 (in re-encoding of already transmitted data), the current metadata frame 902 (in re-encoding), and the future data frame 904 (not available for modification). Each of these frames is further shown along with the subframe boundary 906.

[0123] Furthermore, Case 1: The correct offset of 910 is shown, where the framing of the input metadata is synchronized with the current frame.

[0124] Case 2: The input metadata framing is synchronized with the current frame by +3 subframes, offset +3 920, or the input metadata framing is synchronized with the current frame by -1 subframe, offset -1 is also shown.

[0125] Furthermore, Case 3 shows that the framing of the input metadata is synchronized with the current frame by +2 subframes, with an offset of +2 930, or that the framing of the input metadata is synchronized with the current frame by -2 subframes, with an offset of -2.

[0126] Additionally, Case 3 is shown as follows: the framing of the input metadata is synchronized with the current frame by +1 subframe, offset +1 940, or the framing of the input metadata is synchronized with the current frame by -3 subframes, offset -3.

[0127] Figure 10 shows the sequence of IVAS encoded frames: frame 1 1000, frame 2 1010, frame 3 1020, and frame 4 1030.

[0128] Additionally, the subframes are shown for 1001 with a 0 subframe offset, 1003 with a +1 (or -3) subframe offset, 1005 with a +2 (or -2) subframe offset, and 1007 with a +3 (or -1) subframe offset.

[0129] Because future metadata is inaccessible (causal processing), the current frame must be encoded using only data from the current frame and potentially previous frames.

[0130] An exemplary spatial metadata encoder is shown with respect to Figure 11. The spatial metadata encoder is configured to operate in such a way that when it finds that four subframes have different metadata (regardless of whether these actually come from frames that would benefit from high temporal resolution, or simply because the macroframing does not match the original framing), the encoding will use 4sf mode for high temporal resolution, but will use the highest possible temporal resolution (or low frequency resolution).

[0131] As mentioned above, this can result in suboptimal perceived quality.

[0132] The spatial metadata encoder is configured to receive spatial metadata 104. The spatial metadata 104 is passed to the subframe analyzer 1101, which is configured to analyze the subframes in the spatial metadata 104 to detect whether all four subframes are similar and whether the 1sf coding mode can be used.

[0133] The analysis results 1102 and spatial metadata 104 can be passed to a coherence detector and 2dir analyzer 1103, which is configured to examine the input and determine the presence of significant coherence metadata. The coherence detector and 2dir analyzer 1103 may further be configured to analyze the spatial metadata and determine, on a band-by-band basis, whether one-way or two-way should be used.

[0134] The analysis results 1104 and spatial metadata 104 can then be passed to the metadata codec configurator 1105, which is used to generate the configuration information 1106.

[0135] The configuration information 1106 and spatial metadata 104 can then be passed to a metadata reducer 1107 configured to generate encoded metadata 1108.

[0136] Furthermore, Figure 12 shows an example of a metadata subframe that is potentially available for analysis during encoding. Therefore, subframe 3 sf (-1,3) 1201 and subframe 4 sf (-1,4) History metadata 1200 having 1203 and subframe 1 sf (0,1) 1205, Subframe 2 sf (0,2) 1207, Subframe 3 sf (0,3) 1209, and subframe 4 sf (0,4) The current metadata frame 1202, which has 1211, is shown. In the example using subframe history as shown herein (in other words, one or more past subframes are stored and available), all subframes are available, but in history-independent analysis, historical metadata subframes such as subframe 3 1201 and subframe 4 1203 are not available.

[0137] With respect to Figure 13, exemplary encoders, such as those shown in Figure 8, are shown in more detail according to several embodiments. In some embodiments, the encoder is configured to acquire or receive transport audio signals 102 and pass them to an audio encoder 1301. The audio encoder 1301 is configured to generate encoded transport audio 1306 and pass it to an audio and metadata coupler (or multiplexer) 1309.

[0138] Furthermore, the encoder is configured to acquire or receive spatial metadata 104 and pass it to the asynchronous analyzer 1303. In some embodiments, the asynchronous analyzer has access to at least the subframes "sf(-1,4)" from the previous frame, the subframes "sf(0,1)", "sf(0,2)", "sf(0,3)", "sf(0,4)" from the current frame, and the number of similar subframes at the end of the previous frame (which may be defined as the value of NstopPrev).

[0139] The asynchronous analyzer 1303 is configured to analyze the spatial metadata 104, generate an arbitrary determined asynchronous 1300, and pass it to the metadata determiner 1305 for transmission. In some embodiments, the encoder does not feature the analyzer 1303, the analysis is performed elsewhere, and the encoder is configured to receive the determined asynchronous information along with the transport audio signal and spatial metadata.

[0140] In some embodiments, the encoder further includes a transmitting metadata determinator 1305 configured to receive spatial metadata 104 and determined asynchronous 1300, and output a spatial metadata frame 1302 to a spatial metadata encoder 1307. In some embodiments, the transmitting metadata determinator is configured to further receive a transport audio signal so that the determinator can generate weighted values ​​for generating spatial metadata based on the transport signal energy within each subframe.

[0141] In some embodiments, the encoder also includes a spatial metadata encoder 1307 configured to receive a spatial metadata frame 1302 and generate encoded spatial metadata 1304, which is passed to an audio and metadata (multiplexer) combiner 1309.

[0142] The encoder further includes an audio and metadata (multiplexer) combiner 1309 configured to receive or acquire encoded spatial metadata 1304 and encoded transport audio 1306 and generate a bitstream 804.

[0143] Regarding Figure 14, an illustrative flowchart of the operation of the encoder shown in Figure 13 is provided.

[0144] Therefore, it is shown that the transport audio signal is obtained, as shown in 1401.

[0145] Furthermore, 1403 demonstrates the acquisition of spatial metadata.

[0146] Next, line 1411 demonstrates encoding an audio signal.

[0147] 1405 presents an analysis of asynchronous behavior related to metadata.

[0148] After analyzing the asynchronous nature within the metadata (and optionally, after obtaining the transport audio signal), 1407 then indicates the decision of which metadata to encode for transmission (or storage).

[0149] Next, 1409 shows the encoding of the determined spatial metadata.

[0150] Once the transport audio signal and metadata are encoded, an operation is performed to combine the encoded transport audio signal and metadata to generate a bitstream, as shown in 1413.

[0151] Finally, the bitstream is output (either in the form of transmission or storage of the bitstream). As shown in 1415, the bitstream contains the encoded transport audio signal and metadata.

[0152] Regarding Figure 15, a further, enlarged flowchart is shown concerning the operation of the asynchronous analyzer 1303 (detection using memory).

[0153] In these embodiments, the asynchronous analyzer 1303 is configured to acquire an input frame containing spatial metadata, as shown in 1501.

[0154] Furthermore, in some embodiments, a delay is applied to the spatial metadata, as shown in 1503, and configured to delay the spatial metadata for subframes. The delay can effectively be seen as generating a history frame.

[0155] Next, 1505 demonstrates the behavior of determining a measure of similarity between the first subframe of the current frame and the last subframe of the history frame. This can effectively be seen as the generation of an SFlag indicator when similarity is determined.

[0156] Furthermore, 1507 indicates the determination of the number of mutually similar subframes from the start of the frame. This determination can generate an Nstart value.

[0157] Furthermore, 1509 indicates the determination of the number of similar subframes from the end of the frame. This determination can generate an Nstop value.

[0158] The number of mutually similar subframes from the end of a frame can be further delayed, as shown in 1511, to generate the NstopPrev value.

[0159] Next, as shown in 1513, a decision may be made to determine whether the Nstart value is 4 or whether the Nstop value is 4.

[0160] If the Nstart value is 4 or the Nstop value is 4, the determined asynchronous output is one that assigns or fixes a mode of 1sf, as shown in 1517.

[0161] Otherwise, as shown in 1515, further decisions may be made to determine if the Nstart value is 3 and if there is an SFlag.

[0162] If the Nstart value is 3 and there is an SFlag, the determined asynchronous output is one that assigns or fixes a mode of 1sf but has an offset value of -1 (subframe), as shown in 1521.

[0163] Otherwise, as shown in 1519, further decisions may be made to determine if the Nstart value is 2, the NstopPrev value is 2 or greater, and whether there is an SFlag.

[0164] If the Nstart value is 2, and the NstopPrev value is 2 or greater, and there is an SFlag, then the determined asynchronous output will either assign or fix a mode of 1sf, but with an offset value of +2 (subframes), as shown in 1525.

[0165] Otherwise, as shown in 1523, further decisions may be made to determine whether the Nstop value is 3.

[0166] If the Nstop value is 3, the determined asynchronous output will either be assigned a mode of 1sf or fixed, but with an offset value of +1 (subframe), as shown in 1529.

[0167] Otherwise, as shown in 1527, the determined asynchronous output is one that assigns or fixes the mode to 4sf.

[0168] In this example, the current spatial metadata frame is shown as the input frame, and the previous spatial metadata frame is shown as the history frame. For analysis, the history frame may be (or only a part of it, for example, just the last subframe of the previous frame).

[0169] Therefore, the analysis uses two parts of the historical information: the number of similar subframes at the end of the previous frame (NstopPrev), and an indication (Sflag) that determines whether the first subframe of the current frame is similar to the last subframe of the previous frame, which is determined using the historical frame and the input frame.

[0170] In these examples, the definition of similarity is not important and can be any appropriate measure of similarity. One embodiment of similarity may be when two subframes are interchangeable and the output signal is perceptually similar to the original signal.

[0171] Exemplary similarity tests can be performed in some embodiments by comparing spatial metadata fields element by element. If the difference in values ​​of a field is greater than a given threshold, the two metadata fields are different. If the metadata fields are not different, they are similar.

[0172] For example, the following may be performed as a similarity check:

[0173] Check the input directional spatial metadata field (one or two directions are active).

[0174] Check the spatial metadata parameters for each time-frequency tile.

[0175] If the difference in the azimuth parameter is greater than a given threshold, for example, 0.5 degrees, the metadata is considered different.

[0176] If the difference in elevation angle parameters is greater than a given threshold, for example, 0.5 degrees, the metadata is different.

[0177] If the difference in the parameter for the direct energy to total energy ratio is greater than a given threshold, for example, 0.1, the metadata is different.

[0178] If the difference in the spread coherence parameter is greater than a given threshold, for example, 0.1, the metadata is different.

[0179] If the difference in surround coherence parameters is greater than a given threshold, for example, 0.1, the metadata is considered different.

[0180] However, any appropriate similarity test can be performed. For example, direction and direct-to-total ratio can be compared using a measure of importance such as those presented in UKIPO patent applications 1919130.3 and 1919131.1, and PCT patent application WO2021 / 130405, i.e., direction vectors having a direct-to-total ratio length can be compared.

[0181] This historical information enables the detection of 1sf frames that span encoded frame boundaries. The result of the analysis process is that the determined asynchronous 1300 enables the determination of the encoding mode (1sf or 4sf) and the detected framing offset (0 or none, -1 or +3, +2 or -2, +1 or -3 subframes). These can be used by metadata determiners to send to construct spatial metadata frames.

[0182] In some embodiments, any decision that does not determine the offset may either maintain the offset detected in the previous frame or set the offset to 0. In some embodiments, a particular order of decisions may differ from that in Figure 15 without changing the final result as a function of the input.

[0183] As described above, the transmitting metadata determiner 1305 is configured to acquire or receive the spatial metadata 104 and the determined asynchronousness 1300, and to determine the metadata to be encoded based on the determined asynchronousness 1300. In some embodiments, the metadata is determined by metadata interpolation.

[0184] For example, when the encoding mode is high time-resolution mode (4sf), the metadata subframes sf(0,1), sf(0,2), sf(0,3), and sf(0,4) of the current frame are provided as they are for further analysis / processing / encoding.

[0185] When the encoding mode is low time resolution (or high frequency resolution) mode (1sf), a single archetypal metadata subframe sf(0,ξ) is determined that represents the four metadata subframes for transmitting / encoding sf(0,1), sf(0,2), sf(0,3), and sf(0,4).

[0186] In some embodiments, this single archetypal metadata subframe is a representative subframe for the entire frame, which can be achieved by selecting one of the subframes to use only the selected subframes. This can be expressed as follows: sf(0,ξ)=sf(0,1)

[0187] It should be understood that a representative subframe for the entire frame can be any other subframe.

[0188] This is a suitable solution, for example, for the 1sf mode when the offset is equal to 0 (because all subframes contain the same data).

[0189] However, in some embodiments, the selected subframe may not optimally represent all subframes in a frame with other offsets, and therefore, in some embodiments, the subframes to be encoded are determined in a different way.

[0190] For example, in some embodiments, the metadata decisioner to transmit is configured to compute an aggregate function across subframes to determine the archetypal subframe sf(0,ξ).

[0191] An example of such an aggregate function is the vector average of directional spatial metadata. This average can be weighted by signal energy, which can be weighted based on subframe positions within a frame or some other weighting function. The following example illustrates the use of transport signal energy for weighting.

[0192] The transport signal is x i It can be denoted by (t), where i is the transport channel index and t is the time index. The transport signal can be converted to the time-frequency domain, for example, using a complex-modulated low-latency filter bank (CLDFB). The signal in this domain can be denoted by Xi(k,n), where k is the frequency bin index and n is the slot index (time and frequency are in the CLDFB slots and bins here, which may differ from the MASA parameter tile definition). The transport signal energy E in the MASA parameter TF tile (b,s), where b is the bandwidth and s is the subframe, is:

[0193]

number

[0194] In this example, one or more CLDFB slots n are grouped into parameter subframes s, and one or more CLDFB bins k are grouped into parameter bands b. In a preferred embodiment, each parameter-time subframe s corresponds to one of the spatial subframes sf(0,1), sf(0,2), sf(0,3), and sf(0,4). In other words, sf(0,s).

[0195] The parameter bandwidth index b is not visible in the subframe structure and is often determined by the available bitrate. Alternatively, calculations can be performed across all 24 bandwidths of the MASA parameterization. Each spatial metadata subframe sf(0,s) has an azimuth angle θ.b,s , elevation angle φ b,s , and the direct energy to total energy ratio r b,s The aggregated azimuth angle θ b,ξ , elevation angle φ b,ξ , and the direct energy to total energy ratio r b,ξ parameters (i.e., they represent all sub - frames s of the frame ξ) are calculated by interpreting the parameter values for each sub - frame as spherical coordinates and converting them

[0196] [Number] to Cartesian representation x b,s , y b,s , z b,s using

[0197] These are then [Number] averaged / summed over the sub - frames s ∈ [1,4] using

[0198] [Number] and inverse - transformed to spherical - coordinate parameterization using where [Number] and tan -1 () is the inverse - tangent (or arctangent) variant that solves for the correct quadrant

[0199] Spread coherence parameter [Number] is the value for each sub - frame [Number] As an energy-weighted average

number

[0200] Similarly, surround coherence parameters

number

number

number

[0201] As mentioned above, the transport signal energy E b,s Alternatively, some other weighting could be used, but otherwise, the behavior would be the same.

[0202] The spatial metadata parameter sf(0,ξ) for bandwidth b represents the result of aggregation / interpolation across subframes. This subframe metadata can be replicated to all subframes of the current frame, replacing the original values.

[0203]

number

[0204] These subframes are then typically used in further analysis and encoding. The decision-maker can then swap the spatial metadata 104 or include them in the spatial metadata frame.

[0205] The results of these embodiments are metadata in 1sf (low temporal resolution (or high frequency resolution)) mode, which does not require further alignment in the decoder. In 4sf (high temporal resolution) mode, no adjustments are made, and the spatial metadata frame is handled without further modification.

[0206] With respect to Figure 16, an exemplary known decoder 803 is shown, as is shown in more detail in Figure 8.

[0207] The decoder is configured to receive the bitstream 804. In some embodiments, the decoder 803 includes a separator configured to unpack the bitstream 804 and separate and output the encoded transport audio signal 1600 and the encoded spatial metadata 1600.

[0208] The encoded transport audio signal 1600 is then provided to the audio decoder 1603, which decodes it to provide the transport audio signal 1602.

[0209] The encoded spatial metadata 1602 is further processed by the spatial metadata decoder 1605, which then provides the spatial metadata frame 1604 to the (MASA) renderer 1606.

[0210] The (MASA) renderer 1606 is configured to use the spatial metadata frame 1604 to determine the output audio signal 1606 from the transport audio signal 1606. In some embodiments, the (MASA) renderer 1606 may be implemented within the (IVAS) decoder embodiment (as described above) or it may be a so-called external renderer.

[0211] However, the decoder can be implemented using any suitable known method, given the operation of the above-described embodiment of the encoder.

[0212] In some embodiments, the asynchronous analyzer 1303 described above utilizes spatial metadata from the previous frame. However, in some situations, spatial metadata from the previous frame is unavailable, for example, due to limitations in working memory.

[0213] In some embodiments, as shown in the flowchart of Figure 17, asynchronous analysis can be performed by the asynchronous analyzer 1303 without requiring the previous frame (or historical frame).

[0214] In other words, analyzer 1303 determines the analysis based on metadata subframes sf(0,1), sf(0,2), sf(0,3), and sf(0,4) (as shown in Figure 12) in order to detect framing asynchronousness.

[0215] The analysis shown in the flowchart of Figure 17 demonstrates how to obtain the input frame as shown in 1701.

[0216] As shown in 1703, this involves determining the number of mutually similar subframes from the start of the frame. This determination can generate an Nstart value.

[0217] Furthermore, 1705 indicates the determination of the number of mutually similar subframes from the end of the frame. This determination can generate an Nstop value.

[0218] Next, as shown in 1707, a decision may be made to determine whether the Nstart value is 4 or whether the Nstop value is 4.

[0219] If the Nstart value is 4 or the Nstop value is 4, the determined asynchronous output is one that assigns or fixes a mode of 1sf, as shown in 1711.

[0220] Otherwise, as shown in 1709, further decisions may be made to determine whether the Nstart value is 3.

[0221] If the Nstart value is 3, the determined asynchronous output will either be assigned a mode of 1sf or fixed, but with an offset value of -1 (subframe), as shown in 1715.

[0222] Otherwise, as shown in 1713, further decisions may be made to determine whether the Nstart value is 2 and the Nstop value is 2.

[0223] If the Nstart value is 2 and the Nstop value is 2, then the determined asynchronous output is one that assigns or fixes a mode of 1sf but has an offset value of +2 (subframes), as shown in 1719.

[0224] Otherwise, as shown in 1717, further decisions may be made to determine whether the Nstop value is 3.

[0225] If the Nstop value is 3, the determined asynchronous output will either be assigned a mode of 1sf or fixed, but with an offset value of +1 (subframe), as shown in 1723.

[0226] Otherwise, as shown in 1721, the determined asynchronous output is one that assigns or fixes the mode to 4sf.

[0227] In other words, in some embodiments, the analyzer is configured to determine the number of similar subframes Nstart from the start of the current frame and the number of similar subframes Nstop from the end of the current frame. These values ​​are then used to determine the frame mode and offset. The remaining operation of the encoder and decoder may be carried out as previously presented.

[0228] The two exemplary embodiments presented above illustrate alternative embodiments aimed at improving audio quality.

[0229] In some embodiments, the choice between adopting a history-based or non-history-based embodiment may be made based on the constraints of the embodiment. For example, the device may have limited memory between frames.

[0230] For example, as shown in Figure 18, encoder 1801 is similar to encoder 801 shown in Figure 8, but has an additional input called operating mode 1880. The operating mode can be selected from any of the above embodiments.

[0231] Furthermore, the operating mode 1880 may be determined by the operating mode controller 1811. The operating mode controller 1811 is configured to receive operating parameters 1860, for example, the overall allowable complexity or memory, and to select the operating mode that best fits the given operating parameters based on these parameters.

[0232] For example, in some embodiments, the encoder may have constraints on memory usage. Constraints on metadata encoding may differ at different bit rates. Therefore, as an example, the operating parameters may also include the total bit rate for encoding the input signal.

[0233] In some embodiments, for example, the operating mode controller may be configured to select an asynchronous analysis method based on the bitrate used and predetermined memory constraints on that bitrate.

[0234] In the example where the complex-valued low-latency filter bank (CLDFB) is shown as the frequency domain representation, other methods of time-frequency domain representation, such as the Short-time Fourier Transform (STFT) or the Quadrature Mirrored Filterbank (QMF), may be used.

[0235] Although the parametric format described above was the MASA format, the embodiment can be extended to other parametric formats, such as parametric coding of ambisonics or multichannel mixes.

[0236] With respect to Figure 19, the exemplary electronic device may be used as any of the device components of the system described above. The device may be any suitable electronic device or apparatus. For example, in some embodiments, device 2000 is a mobile device, user equipment, tablet computer, computer, audio playback device, etc. The device may be configured to implement, for example, an encoder and / or decoder, or any of the functional blocks described above.

[0237] In some embodiments, the device 2000 includes at least one processor or central processing unit 2007. The processor 2007 may be configured to execute various program code, such as in the manner described herein.

[0238] In some embodiments, device 2000 includes at least one memory 2011. In some embodiments, at least one processor 2007 is coupled to memory 2011. Memory 2011 can be any suitable storage means. In some embodiments, memory 2011 includes a program code section for storing program code that can be implemented on processor 2007. Furthermore, in some embodiments, memory 2011 may further include a stored data section for storing data, for example, data that has been or will be processed according to embodiments described herein. The implementable program code stored in the program code section and the data stored in the stored data section can be retrieved by processor 2007 at any time as needed via memory-processor coupling.

[0239] In some embodiments, device 2000 includes a user interface 2005. In some embodiments, the user interface 2005 may be coupled to a processor 2007. In some embodiments, the processor 2007 may control the operation of the user interface 2005 and receive input from the user interface 2005. In some embodiments, the user interface 2005 may allow a user to input commands into device 2000, for example, via a keypad. In some embodiments, the user interface 2005 may allow a user to retrieve information from device 2000. For example, the user interface 2005 may include a display configured to display information from device 2000 to the user. In some embodiments, the user interface 2005 may include a touchscreen or touch interface capable of both allowing information to be input into device 2000 and further displaying information to the user of device 2000. In some embodiments, the user interface 2005 may be a user interface for communication.

[0240] In some embodiments, device 2000 includes an input / output port 2009. In some embodiments, the input / output port 2009 includes a transceiver. The transceiver in such embodiments may be coupled to a processor 2007 and may be configured to enable communication with other devices or electronic devices, for example, via a wireless communication network. The transceiver or any suitable transceiver or transmitter and / or receiver means may, in some embodiments, be configured to communicate with other electronic devices or devices via a wire or wired coupling.

[0241] The transceiver may communicate with further devices by any suitable known communication protocol. For example, in some embodiments, the transceiver may use a suitable radio access architecture based on Long-Term Evolution Advanced (LTE Advanced, LTE-A) or New Radio (NR) (or 5G), Universal Mobile Communications System (UMTS) Radio Access Network (UTRAN or E-UTRAN), Long-Term Evolution (LTE, same as E-UTRA), 2G Network (Legacy Network Technology), Wireless Local Area Network (WLAN or Wi-Fi), WiMAX (Worldwide Interoperability for Microwave Access), Bluetooth®, Personal Communication Services (PCS), ZigBee®, Wideband Code Division Multiple Access (WCDMA), Systems using Ultra-Wideband (UWB) Technology, Sensor Networks, Mobile Ad Hoc Networks (MANET), Cellular Internet of Things (IoT) RAN and Internet Protocol Multimedia Subsystem (IMS), any other suitable option and / or any combination thereof.

[0242] The transceiver input / output port 1409 can be configured to receive signals.

[0243] In some embodiments, device 1400 may be used as at least part of a synthesis device. The input / output port 1409 may be coupled to headphones (which may be head-tracking or non-head-tracking headphones) or similar devices, and speakers.

[0244] In general, various embodiments of the present invention may be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. For example, some embodiments may be implemented in hardware, while others may be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Various embodiments of the invention may be shown and described as block diagrams, flowcharts, or using any other graphic representation, but it should be understood that these blocks, devices, systems, techniques, or methods described herein may, in non-limiting examples, be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers, or other computing devices, or any combination thereof.

[0245] Embodiments of the present invention may be implemented by computer software executable by the data processor of a mobile device, such as in a processor entity, by hardware, or by a combination of software and hardware. Furthermore, it should be noted that any block of the logic flow as shown in the drawings may represent a program step or an interconnected logic circuit, block, and function, or a combination of a program step and a logic circuit, block, and function. The software may be stored on a physical medium such as a memory chip or memory block implemented in a processor, a magnetic medium such as a hard disk or floppy disk, and an optical medium such as a DVD and its data variants, or a CD.

[0246] Memory may be of any type suitable for the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processor may be of any type suitable for the local technical environment and may include, in non-limiting examples, one or more of the following: general-purpose computers, dedicated computers, microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), gate-level circuits and processors based on multicore processor architectures.

[0247] Embodiments of the invention can be put into practice in various components, such as integrated circuit modules. Designing integrated circuits is generally a highly automated process. Complex and high-performance software tools are available to translate logic-level designs into ready-to-etch and formed semiconductor circuit designs on semiconductor substrates.

[0248] Programs such as those offered by Synopsys, Inc. in Mountain View, California, and Cadence Design in San Jose, California, use established design rules and a library of pre-stored design modules to automatically route conductors and place components on semiconductor chips. Once the design for the semiconductor circuit is complete, the resulting design in a standard electronic format (e.g., Opus or GDSII) can be sent to a semiconductor manufacturing facility or "fab" for production.

[0249] The term "circuit" as used in this application refers to, (a) Hardware-only circuit embodiments (such as embodiments consisting only of analog and / or digital circuits), and (b) combination of hardware circuits and software (to the extent applicable): (i) combinations of analog and / or digital hardware circuits and software / firmware, (ii) Any part of a hardware processor, software, and memory having software (including a digital signal processor) that cooperates to cause a device such as a mobile phone or server to perform various functions, A processor, such as a microprocessor or a part of a microprocessor, which requires hardware circuitry and software (e.g., firmware) for operation, but the software may not be present when not required for operation. This may refer to one, more, or all of them. This definition of circuit applies to all uses of this term in this application, including in any claim. Further examples include, as used herein, embodiments of a mere hardware circuit or processor (or multiple processors), or a portion of a hardware circuit or processor, and its associated software and / or firmware. Where applicable to a particular claim, the term circuit also includes, for example, a baseband integrated circuit or processor integrated circuit for a mobile device, or a similar integrated circuit in a server, cellular network device, or other computing or network device.

[0250] As used herein, the term “non-transient” is a limitation of the medium itself (i.e., tangible and not signal-based), not a limitation of data storage persistence (e.g., RAM vs. ROM).

[0251] As used herein, “at least one of the following <list of two or more elements>” and “at least one of the <list of two or more elements>” and similar phrases mean at least one of the elements, or at least two or more of the elements, or at least all of the elements, when the lists of two or more elements are linked by “and” or “or”.

[0252] The foregoing description has provided a complete and informative description of exemplary embodiments of the invention, in illustrative and non-limiting terms. However, various modifications and adaptations may become apparent to those skilled in the art in light of the foregoing description, when read in conjunction with the accompanying drawings and the accompanying claims. However, all such modifications and similar modifications of the teachings of the invention still fall within the scope of the invention as defined in the accompanying claims.

Claims

1. It is a device, Means for obtaining at least one spatial metadata parameter associated with at least one audio signal, wherein the at least one spatial metadata parameter is located in a frame that includes at least two subframes, Means for acquiring asynchronous information, wherein the asynchronous information is based on the asynchronous relationship between a sequence of at least two time subframes containing similar values ​​and a frame for further processing at least one spatial metadata parameter, Means for processing the frame, which includes the at least two time subframes, based on the asynchronous information, such that the processed frame, which includes the at least two time subframes of at least one spatial metadata parameter, can be further processed. A device equipped with the following features.

2. A sequence of at least two time subframes containing the aforementioned similar values ​​is: Things located within the same frame, and Something that is located across two consecutive frames, The apparatus according to claim 1, which is one of the following.

3. The apparatus according to claim 1 or 2, wherein the means is for further processing the processed frame, which includes the at least two time subframes of at least one spatial metadata parameter.

4. The apparatus according to claim 3, wherein means for further processing the processed frame of at least one spatial metadata parameter including the at least two time subframes are for encoding the processed frame including the at least two time subframes of the at least one spatial metadata parameter.

5. The means further, A device for acquiring at least one audio signal, A device for encoding at least one audio signal, The apparatus according to any one of claims 1 to 4.

6. The means for acquiring the asynchronous information is, For analyzing the at least one spatial metadata parameter in order to determine the asynchronous information, A device for receiving the asynchronous information from at least one further device that determined the asynchronous information, A device for receiving the asynchronous information from at least one further device that has analyzed the at least one spatial metadata parameter in order to determine the asynchronous information, The apparatus according to any one of claims 1 to 5, which is one of the above.

7. The apparatus according to any one of claims 1 to 6, wherein the means for acquiring the asynchronous information is for acquiring an offset value that identifies the time difference between the sequence of at least two time subframes and the encoded frame, the time difference being acquired as the number of subframes.

8. The apparatus according to claim 7, wherein the means for processing the sequence of the at least two time subframes based on the asynchronous information is for processing the at least one spatial metadata parameter based on the offset value.

9. The apparatus according to claim 8, wherein the means for processing the at least one spatial metadata parameter based on the offset value is for generating a single spatial metadata parameter subframe to represent the at least two time subframes based on the processing of the sequence of the at least two time subframes.

10. The means for generating a single spatial metadata parameter subframe to represent the at least two time subframes based on the processing of the sequence of the at least two time subframes, A means for selecting one of the at least two subframes so that it is the single spatial metadata parameter subframe, A means for determining the aggregation of the at least two subframes so that they form a single spatial metadata parameter subframe, The apparatus according to claim 9, which is one of the present inventions.

11. The means for determining the aggregation of the at least two subframes so that they form a single spatial metadata parameter subframe is: For determining an aggregation function based on the vector mean of the directional component of at least one of the spatial metadata parameters, For determining an aggregation function based on the vector mean of the directional components of at least one spatial metadata parameter, weighted by signal energy weighting, The apparatus according to claim 10, which is one of the present inventions.

12. The apparatus according to any one of claims 1 to 11, wherein the means for acquiring asynchronous information is for acquiring an encoding mode that identifies an encoding mode for encoding the at least one spatial metadata parameter based on the asynchronous information.

13. The means for analyzing the at least one spatial metadata parameter in order to determine the asynchronous information is, For analyzing at least one spatial metadata parameter associated with at least one audio signal for the current frame in order to determine the number of mutually similar subframes from the start of the current frame and the number of mutually similar subframes from the end of the current frame, A means for determining the asynchronous information based on the number of similar subframes from the start of the current frame and / or the number of similar subframes from the end of the current frame, The apparatus according to claim 6.

14. The means for analyzing the at least one spatial metadata parameter in order to determine the asynchronous information further includes, This is for analyzing whether the last subframe of the previous frame and the first subframe of the current frame are similar to each other, This is for obtaining the number of subframes that are similar to each other in the aforementioned previous frame, The apparatus according to claim 13, wherein the means for determining the asynchronous information based on the number of similar subframes from the start of the current frame and / or the number of similar subframes from the end of the current frame further determines the asynchronous information based on whether the last subframe of the previous frame and the first subframe of the current frame are similar subframes, and the number of similar subframes for the previous subframe.

15. A method for an apparatus, Obtaining at least one spatial metadata parameter associated with at least one audio signal, wherein the at least one spatial metadata parameter is located in a frame that includes at least two subframes, Obtaining asynchronous information, wherein the asynchronous information is based on the asynchronous relationship between a sequence of at least two time subframes containing similar values ​​and a frame for further processing at least one spatial metadata parameter. Processing the frame, which includes the at least two time subframes, based on the asynchronous information, such that the processed frame, which includes the at least two time subframes of at least one spatial metadata parameter, can be further processed. Methods that include...

16. A sequence of at least two time subframes containing the aforementioned similar values ​​is: Things located within the same frame, and Something that is located across two consecutive frames, The method according to claim 15, which is one of the methods.

17. The method according to claim 15 or 16, further comprising processing the processed frame which includes the at least two time subframes of at least one spatial metadata parameter.

18. The method according to claim 17, wherein further processing of the processed frame of at least one spatial metadata parameter including the at least two time subframes comprises encoding the processed frame including the at least two time subframes of the at least one spatial metadata parameter.

19. Acquiring at least one audio signal, Encoding the aforementioned at least one audio signal, The method according to claim 18, further comprising:

20. Obtaining asynchronous information is To determine the asynchronous information, the at least one spatial metadata parameter is analyzed, Receiving the asynchronous information from at least one further device that determined the asynchronous information, Receiving the asynchronous information from at least one further device that has analyzed the at least one spatial metadata parameter in order to determine the asynchronous information, The method according to any one of claims 15 to 19, including one of the following.

21. The method according to any one of claims 15 to 20, wherein obtaining asynchronous information includes obtaining an offset value that identifies a time difference between the sequence of at least two time subframes and the encoded frame, the time difference being obtained as the number of subframes.

22. The method according to claim 21, wherein processing the sequence of the at least two time subframes based on the asynchronous information includes processing the at least one spatial metadata parameter based on the offset value.

23. The method according to claim 22, wherein processing the at least one spatial metadata parameter based on the offset value includes generating a single spatial metadata parameter subframe to represent the at least two time subframes based on the processing of the sequence of the at least two time subframes.

24. Based on the processing of the sequence of the at least two time subframes, generating a single spatial metadata parameter subframe to represent the at least two time subframes is: Selecting one of the at least two subframes so that it is the single spatial metadata parameter subframe, Determining the aggregation of the at least two subframes so that they form a single spatial metadata parameter subframe, The method according to claim 23, comprising one of the following.

25. Determining the aggregation of the at least two subframes so that they form a single spatial metadata parameter subframe means The aggregation function is determined based on the vector mean of the directional component of at least one of the spatial metadata parameters, Determining an aggregation function based on the vector mean of the directional components of at least one spatial metadata parameter weighted by signal energy weighting, The method according to claim 24, comprising one of the following.

26. The method according to any one of claims 15 to 25, wherein obtaining asynchronous information includes obtaining an encoding mode that identifies an encoding mode for encoding the at least one spatial metadata parameter based on the asynchronous information.

27. Analyzing the at least one spatial metadata parameter to determine the asynchronous information is To determine the number of mutually similar subframes from the start of the current frame and the number of mutually similar subframes from the end of the current frame, the at least one spatial metadata parameter associated with the at least one audio signal is analyzed for the current frame. The asynchronous information is determined based on the number of similar subframes from the start of the current frame and / or the number of similar subframes from the end of the current frame. The method according to claim 20, including the method described in claim 20.

28. Analyzing the at least one spatial metadata parameter to determine the asynchronous information is Analyze whether the last subframe of the previous frame and the first subframe of the current frame are similar to each other, Regarding the aforementioned previous frame, obtain the number of subframes that are similar to each other, It further includes, The method according to claim 27, wherein determining the asynchronous information based on the number of similar subframes from the start of the current frame and / or the number of similar subframes from the end of the current frame further comprises determining the asynchronous information based on whether the last subframe of the previous frame and the first subframe of the current frame are similar subframes, and the number of similar subframes for the previous subframe.