Decoding of frame-level out-of-sync metadata
By detecting and adjusting encoding modes for frame-level asynchrony in spatial metadata, the solution addresses coding efficiency and quality issues in immersive audio codecs, ensuring synchronized framing and optimal encoding even in asynchronous conditions.
Patent Information
- Application Number
- GB2023004334
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-03-24
- Publication Date
- 2025-06-25
AI Technical Summary
Existing immersive audio codecs face challenges in maintaining optimal coding efficiency and audio quality due to frame-level asynchrony between spatial metadata and audio signals, particularly in scenarios involving tandem coding and multi-stream combining, which can lead to sub-optimal encoding modes and data loss.
The proposed solution involves detecting frame-level asynchrony in spatial metadata by analyzing sub-frames for similarity and adjusting encoding modes to ensure synchronized framing, either through explicit signaling or implicit detection, thereby optimizing time-frequency resolution and maintaining audio quality.
This approach enhances coding efficiency and audio quality by aligning spatial metadata with audio signals, ensuring optimal encoding modes are used even in asynchronous conditions, reducing data loss and improving perceptual quality.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Field The present application relates to apparatus and methods for decoding frame-level out-of-sync metadata. Background Parametric spatial audio capture from inputs, such as microphone arrays and other sources, is a typical and an effective choice to estimate from the input (microphone array signals) a set of parameters such as directions of the sound in frequency bands, and the ratios between the directional and non-directional parts of the captured sound in frequency bands. These parameters are known to well describe the perceptual spatial properties of the captured sound at the position of the microphone array. These parameters can be utilized in synthesis of the spatial sound accordingly, for headphones binaurally, for loudspeakers, or to other formats, such as Ambisonics. The directions and direct-to-total and diffuse-to-total energy ratios in frequency bands are thus a parameterization that is particularly effective for spatial audio capture. A parameter set consisting of a direction parameter in frequency bands and an energy ratio parameter in frequency bands (indicating the directionality of the sound) can be also utilized as the spatial metadata (which may also include other parameters such as surround coherence, spread coherence, number of directions, distance etc) for an audio codec. For example, these parameters can be estimated from microphone-array captured audio signals, and, for example, a stereo or mono transport audio signal can be generated from the microphone array signals to be conveyed with the spatial metadata. Immersive audio codecs are being implemented supporting a multitude of operating points ranging from a low bit rate operation to transparency. An example of such a codec is the Immersive Voice and Audio Services (IVAS) codec which is being designed to be suitable for use over a communications network such as a 3GPP 4G / 5G network including use in such immersive services as, for example immersive voice and audio for virtual reality (VR). This audio codec is expected to handle the encoding, decoding and rendering of speech, music and generic audio. It is furthermore expected to support channel-based audio, object-based audio, and scene-based audio inputs including spatial information about the sound field and sound sources. The codec is also expected to operate with low latency to enable conversational services as well as support high error robustness under various transmission conditions. The transport audio signal could be encoded, for example, using an IVAS audio core codec, or with an AAC (Advanced Audio Coding) or EVS (Enhanced Voice Services) encoder. A decoder can decode the audio signals into PCM (Pulse code modulation) signals and process the sound in frequency bands (using the spatial metadata) to obtain the spatial output, for example, a binaural output. The aforementioned immersive audio codecs are particularly suitable for encoding captured spatial sound from microphone arrays (e.g., in mobile phones, VR cameras, stand-alone microphone arrays). However, such an encoder can have other input types, for example, loudspeaker signals, audio object signals, Ambisonic signals. Summary According to a first aspect there is provided an apparatus comprising means for: obtaining, at least one encoded spatial metadata associated with at least one encoded audio signal, the at least one spatial metadata parameter being arranged in frames comprising at least two sub-frames; obtaining asynchrony information, wherein the asynchrony information is based on asynchrony between a sequence of at least two time sub-frames containing similar values and a frame for any further processing the at least one spatial metadata parameter; decoding the at least one encoded spatial metadata parameter to generate a sequence of the at least two sub-frames of spatial metadata parameters; and processing the frame comprising the at least two time sub-frames based on the asynchrony information, such that the processed frame comprising the at least two time sub-frames of at least one spatial metadata parameter is able to be employed in a rendering of audio signals. The sequence of at least two time sub-frames containing similar values may be one of: located within the same frame; and located over two consecutive frames. The means may be further for: obtaining the at least one encoded audio signal; decoding the at least one encoded audio signal; and rendering an output audio signal from the decoded at least one encoded audio signal based on the processed frame comprising the at least two time sub-frames of at least one spatial metadata parameter. The means for obtaining asynchrony information may be one of: analysing the decoded at least one encoded spatial metadata parameter to determine the asynchrony information; receiving the asynchrony information from at least one further apparatus having obtained the asynchrony information; and receiving the asynchrony information from at least one further apparatus having analysed the at least one spatial metadata parameter to determine the asynchrony information. The means for obtaining the at least one encoded audio signal may be for receiving the at least one encoded audio signal from the at least one further apparatus. The means for obtaining asynchrony information may be for obtaining an offset value identifying a temporal difference between the frame comprising at least two time sub-frames and a rendering frame, the temporal difference obtained as a number of sub-frames. The means for processing the frame comprising the at least two time subframes based on the asynchrony information may be for processing the frame comprising the at least two time sub-frames based on the offset value. The means for processing the frame comprising the at least two time subframes based on the offset value may be for applying an alignment buffer configured based on the offset value to modify a time associated with the processed decoded at least one encoded spatial metadata parameter relative to the decoded at least one encoded audio signal. The means for obtaining asynchrony information may be for obtaining an encoding mode identifying an encoding mode for the at least one encoded spatial metadata parameter. The means for analysing the decoded at least one encoded spatial metadata parameter to determine the asynchrony information may be for: analysing for a current frame of the decoded at least one encoded spatial metadata parameter to determine: a number of mutually similar sub-frames from the current frame start and a number of mutually similar sub-frames from the current frame end; and determining the asynchrony information based on the number of mutually similar sub-frames from the current frame start and / or the number of mutually similar subframes from the current frame end. The means for analysing the at least one spatial metadata parameter to determine the asynchrony information may be further for: analysing for a last-sub-frame of a previous frame and a first-sub-frame of a current frame whether these are mutually similar sub-frames; and obtaining for the previous frame a number of mutually similar sub-frames, wherein the means for determining the asynchrony information based on the number of mutually similar sub-frames from the current frame start and / or the number of mutually similar sub-frame from the current frame end may be further for determining the asynchrony information based on whether the last-sub-frame of a previous frame and a first-sub-frame of a current frame are mutually similar sub-frames; and the number of mutually similar sub-frames for the previous sub-frame. The means may be further for obtaining an operation mode indicator; and wherein the means for processing the frame comprising the at least two time subframes based on the asynchrony information may be for processing the decoded at least one encoded spatial metadata parameter based on the asynchrony information and the operation mode indicator. According to a second aspect there is provided an apparatus comprising means for: obtaining at least one spatial metadata parameter associated with at least one audio signal, the at least one spatial metadata being arranged in frames comprising at least two sub-frames; obtaining asynchrony information, wherein the asynchrony information is based on asynchrony between a sequence of at least two time sub-frames containing similar values and a frame for encoding the at least one spatial metadata parameter; encoding the asynchrony information, and the frame comprising the at least two time sub-frames based on the asynchrony information. The sequence of at least two time sub-frames containing similar values may be one of: located within the same frame; and located over two consecutive frames. The means may be further for: obtaining the at least one audio signal; and encoding the at least one audio signal based on the frame for encoding the at least one spatial metadata parameter. The means for obtaining asynchrony information may be for one of: analysing the at least one spatial metadata parameter to determine the asynchrony information; receiving the asynchrony information from at least one further apparatus having determined the asynchrony information; and receiving the asynchrony information from at least one further apparatus having analysed the at least one spatial metadata parameter to determine the asynchrony information. The means for obtaining asynchrony information may be for obtaining an offset value identifying a number of sub-frames which is a difference between the frame comprising at least two time sub-frames and the encoding frame, the temporal difference obtained as a number of sub-frames. The means for encoding the frame comprising the at least two time subframes based on the asynchrony information may be for: processing the at least one spatial metadata parameter based on the offset value; and encoding the processed at least one spatial metadata parameter. The means processing the at least one spatial metadata parameter based on the offset value may be for determining an aggregation of the at least two subframes to be a single spatial metadata parameter sub-frame to represent the frame, The means for obtaining asynchrony information may be for obtaining an encoding mode identifying an encoding mode for encoding the at least one audio signal and the at least one spatial metadata parameter based on the asynchrony information. The means for analysing the at least one spatial metadata parameter to determine the asynchrony information may be for: analysing for a current frame at least one spatial metadata parameter associated with the at least one audio signal to determine: a number of mutually similar sub-frames from the current frame start and a number of mutually similar sub-frames from the current frame end; determining the asynchrony information based on the number of mutually similar sub-frames from the current frame start and / or the number of mutually similar subframes from the current frame end. The means for analysing the at least one spatial metadata parameter to determine the asynchrony information may be further for: analysing for a last-sub-frame of a previous frame and a first-sub-frame of a current frame whether these are mutually similar sub-frames; and obtaining for the previous frame a number of mutually similar sub-frames, wherein the means for determining the asynchrony information based on the number of mutually similar sub-frames from the current frame start and / or the number of mutually similar sub-frame from the current frame end is further for determining the asynchrony information based on whether the last-sub-frame of a previous frame and a first-sub-frame of a current frame are mutually similar sub-frames; and the number of mutually similar sub-frames for the previous sub-frame. The means for encoding the asynchrony information, and the at least one spatial metadata parameter may be for, based on the asynchrony information identifying asynchrony between the spatial metadata parameter frame comprising at least two time sub-frames and the encoding frame for encoding the at least one audio signal, encoding the spatial metadata parameters in a first encoding mode for a number of frames before switching to processing and encoding the spatial metadata parameters in a second encoding mode. According to a third aspect there is provided a method for an apparatus, the method comprising: obtaining, at least one encoded spatial metadata associated with at least one encoded audio signal, the at least one spatial metadata parameter being arranged in frames comprising at least two sub-frames; obtaining asynchrony information, wherein the asynchrony information is based on asynchrony between a sequence of at least two time sub-frames containing similar values and a frame for any further processing the at least one spatial metadata parameter; decoding the at least one encoded spatial metadata parameter to generate a sequence of the at least two sub-frames of spatial metadata parameters; and processing the frame comprising the at least two time sub-frames based on the asynchrony information, such that the processed frame comprising the at least two time subframes of at least one spatial metadata parameter is able to be employed in a rendering of audio signals. The sequence of at least two time sub-frames containing similar values may be one of: located within the same frame; and located over two consecutive frames. The method may further comprise: obtaining the at least one encoded audio signal; decoding the at least one encoded audio signal; and rendering an output audio signal from the decoded at least one encoded audio signal based on the processed frame comprising the at least two time sub-frames of at least one spatial metadata parameter. Obtaining asynchrony information may comprise one of: analysing the decoded at least one encoded spatial metadata parameter to determine the asynchrony information; receiving the asynchrony information from at least one further apparatus having obtained the asynchrony information; and receiving the asynchrony information from at least one further apparatus having analysed the at least one spatial metadata parameter to determine the asynchrony information. Obtaining the at least one encoded audio signal may comprise receiving the at least one encoded audio signal from the at least one further apparatus. Obtaining asynchrony information may comprise obtaining an offset value identifying a temporal difference between the frame comprising at least two time sub-frames and a rendering frame, the temporal difference obtained as a number of sub-frames. Processing the frame comprising the at least two time sub-frames based on the asynchrony information may comprise processing the frame comprising the at least two time sub-frames based on the offset value. Processing the frame comprising the at least two time sub-frames based on the offset value may comprise applying an alignment buffer configured based on the offset value to modify a time associated with the processed decoded at least one encoded spatial metadata parameter relative to the decoded at least one encoded audio signal. Obtaining asynchrony information may comprise obtaining an encoding mode identifying an encoding mode for the at least one encoded spatial metadata parameter. Analysing the decoded at least one encoded spatial metadata parameter to determine the asynchrony information may comprise: analysing for a current frame of the decoded at least one encoded spatial metadata parameter to determine: a number of mutually similar sub-frames from the current frame start and a number of mutually similar sub-frames from the current frame end; and determining the asynchrony information based on the number of mutually similar sub-frames from the current frame start and / or the number of mutually similar sub-frames from the current frame end. Analysing the at least one spatial metadata parameter to determine the asynchrony information may further comprise: analysing for a last-sub-frame of a previous frame and a first-sub-frame of a current frame whether these are mutually similar sub-frames; and obtaining for the previous frame a number of mutually similar sub-frames, wherein determining the asynchrony information based on the number of mutually similar sub-frames from the current frame start and / or the number of mutually similar sub-frame from the current frame end may further comprise determining the asynchrony information based on whether the last-sub-frame of a previous frame and a first-sub-frame of a current frame are mutually similar sub-frames; and the number of mutually similar sub-frames for the previous sub-frame. The method may further comprise obtaining an operation mode indicator; and wherein processing the frame comprising the at least two time sub-frames based on the asynchrony information may further comprise processing the decoded at least one encoded spatial metadata parameter based on the asynchrony information and the operation mode indicator. According to a fourth aspect there is provided a method for an apparatus, the method comprising: obtaining at least one spatial metadata parameter associated with at least one audio signal, the at least one spatial metadata being arranged in frames comprising at least two sub-frames; obtaining asynchrony information, wherein the asynchrony information is based on asynchrony between a sequence of at least two time sub-frames containing similar values and a frame for encoding the at least one spatial metadata parameter; encoding the asynchrony information, and the frame comprising the at least two time sub-frames based on the asynchrony information. The sequence of at least two time sub-frames containing similar values may be one of: located within the same frame; and located over two consecutive frames. The method may further comprise: obtaining the at least one audio signal; and encoding the at least one audio signal based on the frame for encoding the at least one spatial metadata parameter. Obtaining asynchrony information may comprise one of: analysing the at least one spatial metadata parameter to determine the asynchrony information; receiving the asynchrony information from at least one further apparatus having determined the asynchrony information; and receiving the asynchrony information from at least one further apparatus having analysed the at least one spatial metadata parameter to determine the asynchrony information. Obtaining asynchrony information may comprise obtaining an offset value identifying a number of sub-frames which is a difference between the frame comprising at least two time sub-frames and the encoding frame, the temporal difference obtained as a number of sub-frames. Encoding the frame comprising the at least two time sub-frames based on the asynchrony information may comprise: processing the at least one spatial metadata parameter based on the offset value; and encoding the processed at least one spatial metadata parameter. Processing the at least one spatial metadata parameter based on the offset value may comprise determining an aggregation of the at least two sub-frames to be a single spatial metadata parameter sub-frame to represent the frame, Obtaining asynchrony information may comprise obtaining an encoding mode identifying an encoding mode for encoding the at least one audio signal and the at least one spatial metadata parameter based on the asynchrony information. Analysing the at least one spatial metadata parameter to determine the asynchrony information may comprise: analysing for a current frame at least one spatial metadata parameter associated with the at least one audio signal to determine: a number of mutually similar sub-frames from the current frame start and a number of mutually similar sub-frames from the current frame end; determining the asynchrony information based on the number of mutually similar sub-frames from the current frame start and / or the number of mutually similar subframes from the current frame end. Analysing the at least one spatial metadata parameter to determine the asynchrony information may comprise: analysing for a last-sub-frame of a previous frame and a first-sub-frame of a current frame whether these are mutually similar sub-frames; and obtaining for the previous frame a number of mutually similar subframes, wherein determining the asynchrony information based on the number of mutually similar sub-frames from the current frame start and / or the number of mutually similar sub-frame from the current frame end may further comprise determining the asynchrony information based on whether the last-sub-frame of a previous frame and a first-sub-frame of a current frame are mutually similar subframes; and the number of mutually similar sub-frames for the previous sub-frame. Encoding the asynchrony information, and the at least one spatial metadata parameter may comprise, based on the asynchrony information identifying asynchrony between the spatial metadata parameter frame comprising at least two time sub-frames and the encoding frame for encoding the at least one audio signal, encoding the spatial metadata parameters in a first encoding mode for a number of frames before switching to processing and encoding the spatial metadata parameters in a second encoding mode. According to a fifth aspect there is provided an apparatus comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the system at least to perform: obtaining, at least one encoded spatial metadata associated with at least one encoded audio signal, the at least one spatial metadata parameter being arranged in frames comprising at least two sub-frames; obtaining asynchrony information, wherein the asynchrony information is based on asynchrony between a sequence of at least two time sub-frames containing similar values and a frame for any further processing the at least one spatial metadata parameter; decoding the at least one encoded spatial metadata parameter to generate a sequence of the at least two sub-frames of spatial metadata parameters; and processing the frame comprising the at least two time sub-frames based on the asynchrony information, such that the processed frame comprising the at least two time sub-frames of at least one spatial metadata parameter is able to be employed in a rendering of audio signals. The sequence of at least two time sub-frames containing similar values may be one of: located within the same frame; and located over two consecutive frames. The apparatus may further be caused to perform: obtaining the at least one encoded audio signal; decoding the at least one encoded audio signal; and rendering an output audio signal from the decoded at least one encoded audio signal based on the processed frame comprising the at least two time sub-frames of at least one spatial metadata parameter. The apparatus caused to perform obtaining asynchrony information may be caused to perform one of: analysing the decoded at least one encoded spatial metadata parameter to determine the asynchrony information; receiving the asynchrony information from at least one further apparatus having obtained the asynchrony information; and receiving the asynchrony information from at least one further apparatus having analysed the at least one spatial metadata parameter to determine the asynchrony information. The apparatus caused to perform obtaining the at least one encoded audio signal may be caused to perform receiving the at least one encoded audio signal from the at least one further apparatus. The apparatus caused to perform obtaining asynchrony information may be caused to perform obtaining an offset value identifying a temporal difference between the frame comprising at least two time sub-frames and a rendering frame, the temporal difference obtained as a number of sub-frames. The apparatus caused to perform processing the frame comprising the at least two time sub-frames based on the asynchrony information may be caused to perform processing the frame comprising the at least two time sub-frames based on the offset value. The apparatus caused to perform processing the frame comprising the at least two time sub-frames based on the offset value may be caused to perform applying an alignment buffer configured based on the offset value to modify a time associated with the processed decoded at least one encoded spatial metadata parameter relative to the decoded at least one encoded audio signal. The apparatus caused to perform obtaining asynchrony information may be caused to perform obtaining an encoding mode identifying an encoding mode for the at least one encoded spatial metadata parameter. The apparatus caused to perform analysing the decoded at least one encoded spatial metadata parameter to determine the asynchrony information may be caused to perform: analysing for a current frame of the decoded at least one encoded spatial metadata parameter to determine: a number of mutually similar sub-frames from the current frame start and a number of mutually similar subframes from the current frame end; and determining the asynchrony information based on the number of mutually similar sub-frames from the current frame start and / or the number of mutually similar sub-frames from the current frame end. The apparatus caused to perform analysing the at least one spatial metadata parameter to determine the asynchrony information may further be caused to perform: analysing for a last-sub-frame of a previous frame and a first-sub-frame of a current frame whether these are mutually similar sub-frames; and obtaining for the previous frame a number of mutually similar sub-frames, wherein determining the asynchrony information based on the number of mutually similar sub-frames from the current frame start and / or the number of mutually similar sub-frame from the current frame end may further comprise determining the asynchrony information based on whether the last-sub-frame of a previous frame and a first-sub-frame of a current frame are mutually similar sub-frames; and the number of mutually similar sub-frames for the previous sub-frame. The apparatus may be further caused to perform obtaining an operation mode indicator; and wherein the apparatus caused to perform processing the frame comprising the at least two time sub-frames based on the asynchrony information may further be caused to perform processing the decoded at least one encoded spatial metadata parameter based on the asynchrony information and the operation mode indicator. According to a sixth aspect there is provided an apparatus comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the system at least to perform: obtaining at least one spatial metadata parameter associated with at least one audio signal, the at least one spatial metadata being arranged in frames comprising at least two sub-frames; obtaining asynchrony information, wherein the asynchrony information is based on asynchrony between a sequence of at least two time subframes containing similar values and a frame for encoding the at least one spatial metadata parameter; encoding the asynchrony information, and the frame comprising the at least two time sub-frames based on the asynchrony information. The sequence of at least two time sub-frames containing similar values may be one of: located within the same frame; and located over two consecutive frames. The method may further comprise: obtaining the at least one audio signal; and encoding the at least one audio signal based on the frame for encoding the at least one spatial metadata parameter. The apparatus caused to perform obtaining asynchrony information may be caused to perform one of: analysing the at least one spatial metadata parameter to determine the asynchrony information; receiving the asynchrony information from at least one further apparatus having determined the asynchrony information; and receiving the asynchrony information from at least one further apparatus having analysed the at least one spatial metadata parameter to determine the asynchrony information. The apparatus caused to perform obtaining asynchrony information may be caused to perform obtaining an offset value identifying a number of sub-frames which is a difference between the frame comprising at least two time sub-frames and the encoding frame, the temporal difference obtained as a number of subframes. The apparatus caused to perform encoding the frame comprising the at least two time sub-frames based on the asynchrony information may be caused to perform: processing the at least one spatial metadata parameter based on the offset value; and encoding the processed at least one spatial metadata parameter. The apparatus caused to perform processing the at least one spatial metadata parameter based on the offset value may be caused to perform determining an aggregation of the at least two sub-frames to be a single spatial metadata parameter sub-frame to represent the frame. The apparatus caused to perform obtaining asynchrony information may be caused to perform obtaining an encoding mode identifying an encoding mode for encoding the at least one audio signal and the at least one spatial metadata parameter based on the asynchrony information. The apparatus caused to perform analysing the at least one spatial metadata parameter to determine the asynchrony information may be caused to perform: analysing for a current frame at least one spatial metadata parameter associated with the at least one audio signal to determine: a number of mutually similar subframes from the current frame start and a number of mutually similar sub-frames from the current frame end; determining the asynchrony information based on the number of mutually similar sub-frames from the current frame start and / or the number of mutually similar sub-frames from the current frame end. The apparatus caused to perform analysing the at least one spatial metadata parameter to determine the asynchrony information may be caused to perform: analysing for a last-sub-frame of a previous frame and a first-sub-frame of a current frame whether these are mutually similar sub-frames; and obtaining for the previous frame a number of mutually similar sub-frames, wherein the apparatus caused to perform determining the asynchrony information based on the number of mutually similar sub-frames from the current frame start and / or the number of mutually similar sub-frame from the current frame end may be caused to perform determining the asynchrony information based on whether the last-sub-frame of a previous frame and a first-sub-frame of a current frame are mutually similar subframes; and the number of mutually similar sub-frames for the previous sub-frame. The apparatus caused to perform encoding the asynchrony information, and the at least one spatial metadata parameter may be caused to perform, based on the asynchrony information identifying asynchrony between the spatial metadata parameter frame comprising at least two time sub-frames and the encoding frame for encoding the at least one audio signal, encoding the spatial metadata parameters in a first encoding mode for a number of frames before switching to processing and encoding the spatial metadata parameters in a second encoding mode. According to a seventh aspect there is provided an apparatus comprising: obtaining circuitry configured to obtain at least one encoded spatial metadata associated with at least one encoded audio signal, the at least one spatial metadata parameter being arranged in frames comprising at least two sub-frames; obtaining circuitry configured to obtain asynchrony information, wherein the asynchrony information is based on asynchrony between a sequence of at least two time subframes containing similar values and a frame for any further processing the at least one spatial metadata parameter; decoding circuitry configured to decode the at least one encoded spatial metadata parameter to generate a sequence of the at least two sub-frames of spatial metadata parameters; and processing circuitry configured to process the frame comprising the at least two time sub-frames based on the asynchrony information, such that the processed frame comprising the at least two time sub-frames of at least one spatial metadata parameter is able to be employed in a rendering of audio signals. According to an eighth aspect there is provided an apparatus comprising: obtaining circuitry configured to obtain at least one spatial metadata parameter associated with at least one audio signal, the at least one spatial metadata being arranged in frames comprising at least two sub-frames; obtaining circuitry configured to obtain asynchrony information, wherein the asynchrony information is based on asynchrony between a sequence of at least two time sub-frames containing similar values and a frame for encoding the at least one spatial metadata parameter; encoding circuitry configured to encode the asynchrony information, and the frame comprising the at least two time sub-frames based on the asynchrony information. According to a ninth aspect there is provided a computer program comprising instructions [or a computer readable medium comprising program instructions] for causing an apparatus to perform at least the following: obtaining, at least one encoded spatial metadata associated with at least one encoded audio signal, the at least one spatial metadata parameter being arranged in frames comprising at least two sub-frames; obtaining asynchrony information, wherein the asynchrony information is based on asynchrony between a sequence of at least two time sub-frames containing similar values and a frame for any further processing the at least one spatial metadata parameter; decoding the at least one encoded spatial metadata parameter to generate a sequence of the at least two sub-frames of spatial metadata parameters; and processing the frame comprising the at least two time sub-frames based on the asynchrony information, such that the processed frame comprising the at least two time sub-frames of at least one spatial metadata parameter is able to be employed in a rendering of audio signals. According to a tenth aspect there is provided a computer program comprising instructions [or a computer readable medium comprising program instructions] for causing an apparatus to perform at least the following: obtaining at least one spatial metadata parameter associated with at least one audio signal, the at least one spatial metadata being arranged in frames comprising at least two subframes; obtaining asynchrony information, wherein the asynchrony information is based on asynchrony between a sequence of at least two time sub-frames containing similar values and a frame for encoding the at least one spatial metadata parameter; encoding the asynchrony information, and the frame comprising the at least two time sub-frames based on the asynchrony information. According to an eleventh aspect there is provided a non-transitory computer readable medium comprising program instructions for causing an apparatus to perform at least the following: obtaining, at least one encoded spatial metadata associated with at least one encoded audio signal, the at least one spatial metadata parameter being arranged in frames comprising at least two sub-frames; obtaining asynchrony information, wherein the asynchrony information is based on asynchrony between a sequence of at least two time sub-frames containing similar values and a frame for any further processing the at least one spatial metadata parameter; decoding the at least one encoded spatial metadata parameter to generate a sequence of the at least two sub-frames of spatial metadata parameters; and processing the frame comprising the at least two time sub-frames based on the asynchrony information, such that the processed frame comprising the at least two time sub-frames of at least one spatial metadata parameter is able to be employed in a rendering of audio signals. According to a twelfth aspect there is provided a non-transitory computer readable medium comprising program instructions for causing an apparatus to perform at least the following: obtaining at least one spatial metadata parameter associated with at least one audio signal, the at least one spatial metadata being arranged in frames comprising at least two sub-frames; obtaining asynchrony information, wherein the asynchrony information is based on asynchrony between a sequence of at least two time sub-frames containing similar values and a frame for encoding the at least one spatial metadata parameter; encoding the asynchrony information, and the frame comprising the at least two time sub-frames based on the asynchrony information. According to a thirteenth aspect there is provided an apparatus comprising: means for obtaining at least one encoded spatial metadata associated with at least one encoded audio signal, the at least one spatial metadata parameter being arranged in frames comprising at least two sub-frames; means for obtaining asynchrony information, wherein the asynchrony information is based on asynchrony between a sequence of at least two time sub-frames containing similar values and a frame for any further processing the at least one spatial metadata parameter; means for decoding the at least one encoded spatial metadata parameter to generate a sequence of the at least two sub-frames of spatial metadata parameters; and means for processing the frame comprising the at least two time sub-frames based on the asynchrony information, such that the processed frame comprising the at least two time sub-frames of at least one spatial metadata parameter is able to be employed in a rendering of audio signals. According to a fourteenth aspect there is provided an apparatus comprising: means for obtaining at least one spatial metadata parameter associated with at least one audio signal, the at least one spatial metadata being arranged in frames comprising at least two sub-frames; means for obtaining asynchrony information, wherein the asynchrony information is based on asynchrony between a sequence of at least two time sub-frames containing similar values and a frame for encoding the at least one spatial metadata parameter; means for encoding the asynchrony information, and the frame comprising the at least two time sub-frames based on the asynchrony information. According to a fifteenth aspect there is provided a computer readable medium comprising program instructions for causing an apparatus to perform at least the following: obtaining, at least one encoded spatial metadata associated with at least one encoded audio signal, the at least one spatial metadata parameter being arranged in frames comprising at least two sub-frames; obtaining asynchrony information, wherein the asynchrony information is based on asynchrony between a sequence of at least two time sub-frames containing similar values and a frame for any further processing the at least one spatial metadata parameter; decoding the at least one encoded spatial metadata parameter to generate a sequence of the at least two sub-frames of spatial metadata parameters; and processing the frame comprising the at least two time sub-frames based on the asynchrony information, such that the processed frame comprising the at least two time subframes of at least one spatial metadata parameter is able to be employed in a rendering of audio signals. According to a sixteenth aspect there is provided a computer readable medium comprising program instructions for causing an apparatus to perform at least the following: obtaining at least one spatial metadata parameter associated with at least one audio signal, the at least one spatial metadata being arranged in frames comprising at least two sub-frames; obtaining asynchrony information, wherein the asynchrony information is based on asynchrony between a sequence of at least two time sub-frames containing similar values and a frame for encoding the at least one spatial metadata parameter; encoding the asynchrony information, and the frame comprising the at least two time sub-frames based on the asynchrony information. An apparatus comprising means for performing the actions of the method as described above. An apparatus configured to perform the actions of the method as described above. A computer program comprising program instructions for causing a computer to perform the method as described above. A computer program product stored on a medium may cause an apparatus to perform the method as described herein. An electronic device may comprise apparatus as described herein. A chipset may comprise apparatus as described herein. Embodiments of the present application aim to address problems associated with the state of the art. Summary of the Figures For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings in which: Figure 1 shows schematically an apparatus for MASA metadata extraction; Figure 2 shows schematically an example MASA metadata frame sub-frame structure; Figure 3 shows schematically an example MASA metadata frame time-frequency structure; Figure 4 shows schematically an example metadata resolution adjustment and reconstruction method; Figure 5 shows example application scenarios showing tandem coding and multi-stream combining; Figure 6 shows example delays implemented by the encoder / decoder with respect to the metadata; Figure 7 shows example asynchrony situations in encoding / decoding the metadata; Figure 8 shows schematically an example system of apparatus suitable for implementing some embodiments; Figure 9 shows schematically timing alternatives in metadata framing according to some embodiments; Figure 10 shows schematically potential metadata contributions with different framing offsets between IVAS and MASA framings according to some embodiments; Figure 11 shows schematically a known metadata analyser and encoder; Figure 12 shows schematically metadata sub-frames available when constructing transport data in history and history-free or history-less modes; Figure 13 shows schematically an encoder suitable for employing some embodiments; Figure 14 shows a flow diagram of the operation of the example encoder shown in Figure 13 according to some embodiments; Figure 15 shows schematically a decoder suitable for employing in some embodiments; Figure 16 shows a flow diagram of the decoder as shown in Figure 15 according to some embodiments; Figure 17 shows example metadata-to-audio alignment delay buffer with respect to the decoder as shown in Figure 15; Figure 18 shows an example metadata decoder configured to determine offsets without signalling according to some embodiments; Figure 19 shows a flow diagram of the operation of the example metadata decoder configured to determine offsets without signalling according to some embodiments; Figure 20 shows a flow diagram of the asynchrony detection and offset determination according to some embodiments; Figure 21 shows schematically a further example system of apparatus suitable for implementing some embodiments; and Figure 22 shows an example device suitable for implementing the apparatus shown in previous figures. Embodiments of the Application The following describes in further detail suitable apparatus and possible mechanisms for the encoding of parametric spatial audio signals comprising transport audio signals and spatial metadata. As indicated above immersive audio codecs (such as 3GPP IVAS) are being planned which support a multitude of operating points ranging from a low bit rate operation to transparency. Metadata-Assisted Spatial Audio (MASA) is an example of a parametric spatial audio format and representation suitable as an input format for IVAS. It can be considered an audio representation consisting of ‘N channels + spatial metadata’. It is a scene-based audio format particularly suited for spatial audio capture on practical devices, such as smartphones. The idea is to describe the sound scene in terms of time- and frequency-varying sound source directions and, e.g., energy ratios. Sound energy in the scene that is not defined (described) by the directions, is described as diffuse (coming from all directions). As discussed above spatial metadata associated with the audio signals may comprise multiple parameters (such as multiple directions and associated with each direction (or directional value) a direct-to-total energy ratio, spread coherence, distance, etc.) per time-frequency tile. The spatial metadata may also comprise other parameters or may be associated with other parameters which are considered to be non-directional (such as surround coherence, diffuse-to-total energy ratio, remainder-to-total energy ratio) but when combined with the directional parameters are able to be used to define the characteristics of the audio scene. For example, a reasonable design choice which is able to produce a good quality output is one where the spatial metadata comprises one or more directions for each time-frequency subframe (and associated with each direction direct-to-total ratios, spread coherence, distance values etc) are determined. With respect to Figure 1 is shown an example MASA analyser 101. The MASA analyser 101 is configured to receive the input audio signal(s) 100 and analyse the input audio signals to generate transport audio signal(s) 102 and spatial metadata 104. Examples of MASA spatial metadata is presented in the following table. These values are available for each time-frequency tile. In some implementations 5 a frame is subdivided into 24 frequency bands and 4 temporal sub-frames. In other implementations other divisions of frequency and time can be employed. Furthermore, in some implementations a frame size (for example, as implemented in IVAS) is 20 ms (and thus the temporal sub-frame is 5 ms). However, similarly, other frame lengths can be employed in other embodiments. In some embodiments 10 the MASA analyser is configured to determine 1 or 2 directions for each time-frequency tile (i.e., there are 1 or 2 direction index, direct-to-total energy ratio, and spread coherence parameters for each time-frequency tile). However, in some embodiments the analyser is configured to generate more than 2 directions for a time-frequency tile. Field bits Description Direction index 16 Direction of arrival of the sound at a time-frequency parameter interval. Spherical representation at about 1-degree accuracy. Range of values: “covers all directions at about 1° accuracy” Values stored as 16-bit unsigned integers. Direct-to-total energy ratio 8 Energy ratio for the direction index (i.e., time-frequency subframe). Calculated as energy in direction / total energy. Range of values: [0.0,1.0] Values stored as 8-bit unsigned integers with uniform spacing of mapped values. Spread coherence 8 Spread of energy for the direction index (i.e., timefrequency subframe). Defines the direction to be reproduced as a point source or coherently around the direction. Range of values: [0.0,1.0] Values stored as 8-bit unsigned integers with uniform spacing of mapped values. Diffuse-to-total energy ratio 8 Energy ratio of non-directional sound over surrounding directions. Calculated as energy of non-directional sound / total energy. Range of values: [0.0,1.0] (Parameter is independent of number of directions provided.) Values stored as 8-bit unsigned integers with uniform spacing of mapped values. Surround coherence 8 Coherence of the non-directional sound over the surrounding directions. Range of values: [0.0,1.0] (Parameter is independent of number of directions provided.) Values stored as 8-bit unsigned integers with uniform spacing of mapped values. Remainder-to-total energy ratio 8 Energy ratio of the remainder (such as microphone noise) sound energy to fulfil requirement that sum of energy ratios is 1. Calculated as energy of remainder sound / total energy. Range of values: [0.0,1.0] (Parameter is independent of number of directions provided.) Values stored as 8-bit unsigned integers with uniform spacing of mapped values. The MASA stream can be rendered to various outputs, such as multichannel loudspeaker signals (e.g., 5.1) or binaural signals. As discussed above the frame size in IVAS is 20 ms. An example of the 5 example frame structure is shown in Figure 2 where the metadata frame 201 comprises four temporal sub-frames which are 5 ms long. Figure 2 shows, for example, the previous frame metadata sub-frame 4 200, then the current metadata frame 201 comprising metadata sub-frame 1 202, metadata sub-frame 2 204, metadata sub-frame 3 206, and metadata sub-frame 4 208. Following this is the succeeding or next frame metadata sub-frame 1 210. Furthermore, the IVAS codec is expected to operate at various bit rates ranging from very low bit rates (for example 13.2 kbps) to relatively high bit rates (for example 512 kbps or even 768 kbps). As the raw bit rate of the MASA metadata is about 300-500 kbps (depending on whether there are encoded one or two simultaneous directions), the metadata is significantly compressed (especially at the lowest bit rates). One aspect of compression can be methods that reduce the temporal and / or frequency resolution of the metadata (which can be employed alongside other methods for compressing the data). As mentioned above and as shown in Figure 3, an example metadata frame 300 in raw high resolution 350 can comprise 24 frequency bands on the frequency axis 301 and 4 temporal subframes (sub-frames 1 to 4 302, 304, 306, 308) on the time axis 303, meaning in total 96 time-frequency tiles (also called TF-tiles). With respect to Figure 4 is shown a method of reducing the number of time-frequency tiles to be transmitted and therefore reduce the required bitrate significantly. Such a method is described in UKIPO patent applications 1919130.3 and 1919131.1 and WO2021 / 130405 presents methods that allow combining metadata from multiple frequency bands and / or temporal subframes to fewer frequency bands and / or temporal subframes. As an example, such as shown in Figure 4, depending on the bitrate, 5-24 frequency bands and 1-4 subframes may be transmitted. Such a method therefore comprises a metadata resolution selector 401 and adjuster 403 configured to generate a 1 sf, high frequency resolution low temporal resolution metadata frame, 404 and a 4sf, low frequency resolution high temporal resolution metadata frame, 406. The metadata is unpacked in a metadata unpacker 405 which is configured to unpack the common representation resolution metadata frame 408 at the decoder (the frame having an unknown underlying TF-resolution). As the MASA stream can be created from various types of devices (e.g., from microphone arrays on mobile devices as well as dedicated Ambisonics microphone arrays, such as the Eigenmike), the methods used for determining the spatial metadata may vary significantly between implementations. Some methods may have high temporal resolution but lower frequency resolution, whereas some methods may have low temporal resolution but higher frequency resolution. In order to improve coding efficiency for both kind of time-frequency resolutions, it has been suggested that the MASA metadata could be encoded in two different modes as shown in PCT application WO2021250312, and also shown above in Figure 4. The first metadata frame resolution is one having more frequency resolution (1sf) but only one temporal sub-frame per frame, the other metadata frame resolution having a lower frequency resolution (4sf) but keeping the 4 temporal subframes. In this example the former mode (1 sf) is selected when the encoder receives spatial metadata which is determined or detected to be identical (or substantially identical or similar) over all subframes of the frame. If the spatial metadata is not identical (or not substantially identical or not similar) over all subframes then the latter mode (4sf) is employed. As an example, at a certain bitrate, the former mode (1 sf) may transmit 18 frequency bands and 1 subframe (in other words a total of 18 TF-tiles), and the latter mode (4sf) may transmit 5 frequency bands and 4 subframes (in other words a total of 20 TF-tiles) which roughly equates to similar size of transmitted data at the same overall bit rate. In PCT application WO2019105575, it has been proposed to use variable input metadata time-frequency resolution. This achieves a similar trade-off as methods of PCT application WO2021250312, however the decision is implemented outside of the codec and can be based on the specific capture algorithm for the microphone array being used. The methods described above therefore show ways to permit an encoding quality to be maintained, where the temporal and the frequency resolution is tuned or adjusted with respect to the audio input. As discussed about these methods select the high frequency resolution mode when all the subframes of the frame have identical (or substantially identical or at least similar-enough) data. A microphone front-end which creates the MASA stream can create the metadata in a way that this is true, but cannot guarantee that the sub-frames are always synchronised. In other words, metadata framing may show asynchrony. Examples of application scenarios which may introduce metadata framing asynchrony are shown with respect to Figure 5. For example, on the left side of Figure 5 shows a tandem coder 501 which is configured to receive as inputs transport signals 102 and metadata 104 in which the metadata (and transport audio) stream was already encoded and then decoded, and it is provided as an input to a second encoder, the tandem coder 501 configured to generate an output transport signal(s) 502 and metadata 504. In this example it cannot be guaranteed that the framing (or sub-frame grouping) of this second encoding implemented by the tandem coder 501 matches the earlier encoder. Furthermore the right side of Figure 5 shows a MCU (Multipoint Control Unit) application where a multi-stream combiner 503 is configured to combine multiple transport 102i, 1022 and metadata 104i, 1042 streams into one stream to generate an output transport signal(s) 512 and metadata 514. Once again it cannot be guaranteed that the synchronization of the framing in the MCU is the same as in all combined streams. With respect to Figure 6 is shown the timing delays which can be experienced in both scenarios within an example system. The MASA stream (comprising the input transport audio signals 102 and metadata 104) can be obtained by the system 601. The system 601 can be effectively considered to comprise an encoder, which is configured to generate a bitstream and decoder which is configured to receive the bitstream and configured to provide an example output comprising a transport audio signal(s) and spatial metadata. These are shown combined into a single schematic system block. The system 601 in some embodiments the system can be considered to comprise an audio core coder 603 configured to implement an IVAS codec (which for example has 32 ms of delay for the MASA format) configured to generate the decoded transport audio signals 602. The metadata delay 605 can be applied to apply a 2 sub-frame look-ahead delay to the metadata to generate delayed spatial metadata 606 in an attempt to re-align the decoded transport audio signals and spatial metadata. Thus, for example, the input situation 610 is one in which the audio signals and spatial metadata are aligned. Then following the application of the encoder / decoder, the post-codec situation 620, shows that the decoded transport audio signals are delayed 621, for example by 12 ms, with respect to the spatial metadata. Then following the application of the metadata delay, the post-delay situation 630, shows that the application of the 2 sub-frame look ahead delay 631 to the metadata approximately matches the decoding delay 621 and therefore approximately re-synchronizes the decoded transport audio signals and the delayed spatial metadata. Although this delay resolution of the sub-frames of the metadata can cause synchronization issues between the transport audio signals and the spatial metadata even where they are accurately synchronised it cannot be guaranteed for a further stage that the delayed spatial metadata within a bitstream is synchronised with the further stage frame clock. In other words, even where the delayed bitstream has identical values for all subframes of a frame, this can no longer be guaranteed that a framing that a later apparatus employs will also have identical values for all subframes of a frame. For example there can be a system where an output of a decoding stage is audio with 12 ms delay and metadata with 10 ms delay and a second coding stage operates with the same global absolute framing clock as the first coding round. Because of the delays, the second encoding stage receives the audio and metadata delayed, even though the framing starts running immediately. In which case the first 12 ms of audio will be zeros and the first two metadata sub-frames will be empty before the decoded real data is available. Therefore the first frame has half zeros and half real data (the first half of the first original frame). The second frame of the second encoding has the second half of the first original frame and the first half of the second original frame, and so forth with half-a-frame offset. This potential asynchrony of the framing is shown with respect to Figure 7. Figure 7 for example shows the original metadata frame timing where there is shown a first ‘original’ metadata frame 1 701 comprising metadata sub-frame 1 711, metadata sub-frame 2 713, metadata sub-frame 3 715, metadata sub-frame 4 717, and a second ‘original’ metadata frame 2 703 comprising metadata sub-frame 1 721, metadata sub-frame 2 723, metadata sub-frame 3 725, metadata sub-frame 4 727. Then is shown all potential 4 offset values that can take place in the 4 subframe case used in MASA. For example, the offset = 0 situation where a first re-encoding metadata frame, frame, 1 731 and a second re-encoding metadata frame, frame 2, 732 are completely aligned or synchronised with respective first ‘original’ metadata frame 1 701 and second ‘original’ metadata frame 2 703. Also is shown the offset = +1 (or -3) sub-frames situation 741 applied with respect to the frame 1 (or frame 2), where the re-encoding metadata frame is shifted by 1 sub-frame with respect to the first ‘original’ metadata frame 1 701 or -3 subframes with respect to the second ‘original’ metadata frame 2 703. Furthermore is shown the offset = +2 (or -2) sub-frames situation 751 with respect to the frame 1 (or frame 2), where the re-encoding metadata frame is shifted by 2 sub-frames with respect to the first ‘original’ metadata frame 1 701 or -2 subframes with respect to the second ‘original’ metadata frame 2 703. Then is shown the offset = +3 (or -1) sub-frames situation 761 with respect to the frame 1 (or frame 2), where the re-encoding metadata frame is shifted by 3 sub-frame with respect to the first ‘original’ metadata frame 1 701 or -1 subframe with respect to the second ‘original’ metadata frame 2 703. Furthermore as indicated above there can be a tandem coding situation where after editing a MASA stream and mixing multiple streams, the server is configured to encode the MASA stream once again prior to storing the signal or transmitting the signal to a user. However there can be a problem with detecting framing asynchrony (offset) from the data and address the asynchrony with respect to metadata coding and decoding methods in that they implement a coding mode (i.e., lower temporal resolution / higher frequency resolution, 1 sf, or higher temporal resolution I lower frequency resolution, 4sf) only by checking the current frame. If the metadata has become out-of-sync with the framing coding efficiency may be compromised. For example, an ‘identical’ 4 subframes of metadata which is split, because of delays, between two different frames to be encoded will be determined as two frames of ‘non-identical’ metadata and therefore will be encoded with higher temporal resolution and lower frequency resolution and as a result, significantly fewer frequency bands may be transmitted than could have been transmitted. This in turn can cause significant deterioration of the perceived audio quality, as the human hearing is sensitive to frequency resolution (and there is no additional time resolution despite the higher temporal resolution 4sf mode being triggered by default). As such the concept relates to handling (encoding and decoding) of parametric spatial audio (in other words, audio signal(s) and spatial metadata), where the spatial metadata is coded in frames containing multiple frequency bands and multiple sub-frames. In some embodiments this is implemented by apparatus and methods that enable coding the spatial metadata with optimized time-frequency resolution also when the spatial metadata is not in synchronization with the framing of the stream (i.e., similar / identical data may be found in subframes of different frames). Additionally the apparatus and methods as described in the embodiments herein are configured to: obtain one or more frames of spatial metadata; compare the values of the spatial metadata in the subframes of the one or more frames to determine if the metadata framing has an offset; select the coding mode based on the comparison; determine spatial metadata to be encoded; encode the spatial metadata and the offset information; and decode the metadata in such a manner that the decoding enables modification of the decoded spatial metadata sub-frames depending on the offset information. In some embodiments the offset (side) information is not encoded and the information is not received, but rather the decoder is configured to detect the offset situation based only on the spatial metadata. In some embodiments the apparatus and methods are configured to analyze the spatial metadata, and detect or determine any framing (i.e., sub-frame grouping) asynchrony and determine any framing offset. (This determination can be implemented in some embodiments in a manner similar to UK application number GB2304318.5) In some embodiments the analysis is implemented in the encoder or may be implemented prior to the encoder. Additionally the apparatus and methods can be configured to determine any new spatial metadata for this frame based on the detected offset and spatial metadata sub-frames. Furthermore in some embodiments the apparatus and methods are configured to decode the received metadata, determine the metadata offset either based on the metadata structure or received additional side information, and modify the decoded and unpacked metadata before further processing in the decoder and / or renderer. The encoding and decoding modes employed in the embodiments can differ according to the implementations in which the embodiments are employed. Thus for example different embodiments may use different encoding and decoding modes. Considering IVAS encoding, high frequency resolution for the metadata is available when the sub-frames in one transport frame are so similar over time that it is determined to be more beneficial using the available transport resources for the frequency resolution instead of the temporal resolution. In some embodiments the encoder and / or decoder can be configured to analyse the metadata sub-frames in the current and previous frames and use the presence of multiple (e.g., 4) sub-frames with similar metadata as an indication of a potential low temporal resolution / high frequency resolution (or 1 sf mode coding) grouping. In the following examples there are 4 sub-frames in one frame, but this is a specific example and the embodiments can be generalized into N sub-frames in one frame. In some embodiments the apparatus and method analyse the spatial metadata using current and possibly one or more metadata sub-frames of the previous frame. In some embodiments when asynchrony is detected, new spatial metadata to be encoded is determined from the spatial metadata of the current frame and possibly one or more metadata sub-frames of the previous frame such that the framing is re-aligned. In some embodiments the coding mode based at least in part on this analysis is determined from the spatial metadata sub-frames, and the metadata is encoded. The determined offset can furthermore be signalled to the decoder (for example, by including the offset information into the spatial metadata, or using extra signalling outside of the spatial metadata.) In such embodiments the decoder uses the offset information and an alignment buffer for modifying the location of the decoded spatial metadata subframes relative to the audio data. The benefit is that these embodiments can be configured to correct the framing asynchrony in many use cases without additional data loss at the potential cost of additional decoder complexity and increased total size of the data to be transmitted. In some embodiments the offset information is not explicitly signalled from the encoder to the decoder. Instead, in such embodiments the decoder is configured to detect or determine this information based on the spatial metadata received from the encoder. This option where the offset information is not signalled can be implemented in some embodiments by the following: When the encoder detects an asynchrony situation, the encoder is configured to keep encoding and sending the metadata in the bitstream in the higher temporal resolution 4sf encoding mode for a number of frames before switching to determining the data to be sent similar to that as described above and sending it with the lower temporal resolution / higher frequency resolution 1 sf encoding mode. The relevance of the lower temporal resolution I higher frequency resolution 1 sf encoding mode is that asynchrony can only be detected when there are frames with all subframes with similar or same data. The decoder then is configured to detect that in the received frames the first sub-frames contain mutually similar spatial metadata, and the remaining last subframes contain mutually similar spatial metadata. It can use the detection or determination method and described and set a detection or determination flag. Once the decoder subsequently receives a frame in which all sub-frames are identical (i.e., it has been encoded with a lower temporal resolution I higher frequency resolution 1 sf encoding mode), it applies the offset on the data on based on a method similar to that described in the embodiments where the signaling is explicit. When receiving a frame with lower temporal resolution / higher frequency resolution 1sf encoding mode, the decoder can be configured to compare the decoded metadata of the current frame with the decoded metadata from the earlier frame (if that was using higher temporal resolution / lower frequency resolution 4sf encoding mode). If one or more earlier metadata sub-frames are similar enough to the current sub-frames taking the TF-tile groupings into account, the decoder can now be configured to detect that the current frame is a version of the metadata with higher frequency resolution. The offset is determined by the number of similar subframes at the end of the earlier frame. Once the offset has been detected, the processing is similar to that as described above. The benefit of these methods over the explicit signalling embodiments is that the additional signalling of the offset can be omitted reducing the total bitrate, but additional complexity on the decoder is required. In some embodiments either explicit or implicit signalling is selected or employed depending on bitrate and / or decoder complexity limitations. For example, in situations with tight bitrate constraints the implicit signalling embodiments (without additional signalling) are employed, or the metadata interpolation as described herein can be employed used. If additionally high packet loss is expected or otherwise random access to the encoded data is needed, or handling of any asynchrony offset is required and the bitrate allows it, an embodiment with explicitly signalled offset information can be employed for improved perceptual quality. With respect to Figure 8 is shown an example system within which some embodiments can be implemented. As an input are the transport audio signals 102 and the spatial metadata 104. The transport audio signals 102 and the spatial metadata 104 are passed to an encoder 801 which generates an encoded bitstream 802. The encoded bitstream 802 is received by the decoder 803 which is configured to generate a spatial audio output 804. As discussed above the input to the system, the transport audio signals 102 and the spatial metadata 104 can be obtained in the form of a MASA stream. The MASA stream can, for example, originate from a mobile device (containing a microphone array), or as an alternative example, it may have been created by an audio server that has potentially processed a MASA stream in some way. The encoder 801 can furthermore, in some embodiments, be an IVAS encoder. The decoder 803, in some embodiments, can be configured to directly output the spatial audio output 804 to be rendered by an external renderer, or edited / processed by an audio server. In some embodiments, the decoder 803 comprises a suitable renderer, which is configured to render the output in a suitable form, such as binaural audio signals or multichannel loudspeaker signals (such as 5.1 or 7.1+4 channel format), which are also examples of spatial audio output 804. In some embodiments the spatial metadata 104 arrives to the encoder 801 as a sequence of sub-frames with no indication of the original framing (grouping of sub-frames). If the framing of the process that provides the metadata is synchronized with the framing of the encoding process, there is no potential problem, however if there is some asynchrony present, for example, from tandem coding or other application scenario as discussed earlier, the encoding may be sub-optimal in the sense that wrong coding mode is selected, and some data will be lost consequently. With respect to Figures 9 and 10 are shown example offset scenarios in the framing asynchrony. Figure 9 for example shows a metadata framing asynchrony alternatives in the case of 4 sub-frames and the underlying MASA metadata framing. There is shown a previous metadata frame (in re-encoding data already sent) 900, a current metadata frame (in re-encoding) 902 and future data frame (unavailable for modification) 904. Each of these frames are further shown with the sub-frame borders 906. Furthermore is shown a case 1: correct offset 910 where the framing of the input metadata is synchronised with the current frame. Also is shown case 2: offset +3 920 where the framing of the input metadata is +3 sub-frame asynchronised with the current frame or offset -1 where the framing of the input metadata is -1 sub-frame asynchronised with the current frame. The two cases, the +3 and -1 sub-frame asynchronised cases, are equal. Further is shown case 3: offset +2 930 where the framing of the input metadata is +2 sub-frame asynchronised with the current frame or offset -2 where the framing of the input metadata is -2 sub-frame asynchronised with the current frame. These two cases, the +2 and -2 sub-frame asynchronised cases, are equal. Additionally is shown case 3: offset +1 940 where the framing of the input metadata is +1 sub-frame asynchronised with the current frame or offset -3 where the framing of the input metadata is -3 sub-frame asynchronised with the current frame. The two cases, the +1 and -3 sub-frame asynchronised cases, are equal. Figure 10 shows a sequence of IVAS encoding frames, frame 1 1000, frame 2 1010, frame 3 1020 and frame 4 1030. Additionally is shown subframes in the cases where 1001 there is a 0 subframe offset, 1003 there is a +1 (or -3) subframe offset, 1005 there is a +2 (or -2) subframe offset and 1007 there is a +3 (or -1) subframe offset. As there is no access to future metadata (causal processing) the current frame must be encoded with only the data from the current and potentially earlier frames. An example spatial metadata encoder is shown with respect to Figure 11. The spatial metadata encoder is configured to operate such that when it sees 4 sub-frames with different metadata (regardless, if these are really from a frame benefiting from high temporal resolution or only because the macro-framing is not matching the original framing), the encoding uses 4sf mode for high temporal resolution, but with lower frequency resolution. As mentioned above, this may produce a sub-optimal perceptual quality. The spatial metadata encoder is configured to receive the spatial metadata 104. The spatial metadata 104 is passed to a sub-frame analyser 1101 which is configured to analyse sub-frames in spatial metadata 104 to detect if all 4 subframes are similar and the 1 sf coding mode could be used. This analysis result 1102 and the spatial metadata 104 can be passed to a coherence detector and 2dir analyser 1103 which is configured to inspect the inputs and determine the presence of meaningful coherence metadata. The coherence detector and 2dir analyser 1103 furthermore can be configured to analyse the spatial metadata and determines on a per-band basis whether one or two directions should be used. The analysis result 1104 and the spatial metadata 104 can then be passed to the metadata codec configurer 1105 which is used to generate configuration information 1106. The configuration information 1106 and the spatial metadata 104 can then be passed to the metadata reducer 1107 configured to generate the encoded metadata 1108. Additionally is shown with respect to Figure 12 an example of the metadata sub-frames potentially available for analysis when encoding. Thus is shown the history metadata 1200 with sub-frame 3 sf(-i,3) 1201 and sub-frame 4 sf(-i,4) 1203, and the current metadata frame 1202, with sub-frame 1 sf(o,i) 1205, sub-frame 2 sf(o,2) 1207, sub-frame 3 sf(o,3) 1209 and sub-frame 4 sf(o,4) 1211. In the examples shown herein with subframe history (in other words past sub-frames are stored and are available) all of the sub-frames are available but in a history-free analysis then the history metadata subframes such as sub-frame 3 1201 and sub-frame 4 1203 are not available. As discussed above in some embodiments a decoder is configured to receive the spatial metadata containing an offset, and in addition, information identifying the amount of offset is received. First with respect to the system such as shown in Figure 8, the encoder producing this data is considered as shown in Figure 13, and then the decoder is discussed as shown in Figure 15. In these following embodiments, it is assumed that the audio has the aforementioned 12 ms delay due to encoding / decoding, and thus the metadata is delayed by two subframes (i.e., 10 ms) in the decoder. The values selected to transmit are selected with this assumption in mind. In other embodiments, there may be different amount of delay, or no delay at all. In those embodiments, the transmitted values may be selected based on some other criteria. Thus Figure 13 shows schematically an encoder block according to some embodiments. The encoder in some embodiments is configured to obtain or receive the transport audio signals 102 and pass these to the audio encoder 1301. The audio encoder 1301 is configured to generate encoded transport audio 1306 and pass this to the audio and metadata combiner (or multiplexer) 1309. Furthermore the encoder is configured to obtain or receive spatial metadata 104 and pass this to an asynchrony obtainer 1303. In some embodiments the asynchrony analyser has access to at least the sub-frame “sf(-1,4)” from the previous frame, sub-frames “sf(0,1)”, “sf(0,2)”, “sf(0,3)”, “sf(0,4)” from the current frame, and the number of mutually similar sub-frames in the end of the previous frame (which can be defined as a value of NstopPrev). The asynchrony obtainer 1303 is configured to obtain the asynchrony information, for example, to analyse the spatial metadata 104 and generate the asynchrony information 1300 and pass this to a metadata to transmit determiner 1305, a spatial metadata encoder 1307 and to the audio and metadata combiner or multiplexer 1309. In some embodiments the encoder does not feature an asynchrony obtainer 1303 and the analysis is performed elsewhere, with the encoder configured to receive the determined asynchrony information with the transport audio signals and the spatial metadata. In some embodiments the encoder further comprises a metadata to transmit determiner 1305 configured to receive the spatial metadata 104 and the determined asynchrony information 1300 and output spatial metadata frame 1302 to a spatial metadata encoder 1307. In some embodiments the metadata to transmit determiner is configured to further receive the transport audio signals in order that the determiner is able to generate weighting values for the generation of spatial metadata based on the transport signal energy in each sub-frame. The metadata to transmit determiner 1305, in some embodiments, can be configured to select the suitable metadata to be encoded (and stored or transmitted) from the data of the previous and the current frames, based on the received asynchrony information 1300. The available spatial metadata sub-frame values as described above are those shown in Figure 12. In some embodiments if the delay asynchrony compensation is limited to frames with a higher frequency resolution 1sf coding mode, the spatial metadata sub-frame sf(-1,3) is not needed. In some embodiments depending on the determined offset and coding mode as indicated by the asynchrony information then the determiner can be configured to determine the new spatial metadata frame in the following manner: If “offset == 0”, the metadata to be transmitted is determined based on the sub-frames of current frame: the metadata to transmit I encode is sf(0,1) 1205, sf(0,2) 1207, sf(0,3) 1209, and sf(0,4) 1211. If “offset == +1 (or -3)”, and the metadata to be transmitted is determined based on the last three sub-frames of the current frame, using an aggregation function f( ): f(sf(0,2), sf(0,3), sf(0,4)). This is implemented because this offset condition requires access to future data, which is not possible in causal processing, and the coding mode used is therefore a lower temporal resolution I higher frequency resolution 1 sf coding mode. If “offset == +2 (or -2)”, the metadata to be transmitted is determined based on two history sub-frames and the first two sub-frames of current frame: sf(-1,3) 1201, sf(-1, 4) 1203, sf(0,1) 1205, sf(0,2) 1207. This way the metadata framing in the encoding matches the original framing and the optimal encoding mode can be determined based on the metadata. The resulting delay of 2 sub-frames in the metadata is compensated for in the decoder by writing the data earlier in the output buffer. It is possible to limit the operation of the system into applying asynchrony correction only with lower temporal resolution / higher frequency resolution 1 sf coding mode frames, in which case no history values are needed here on the account of the 4 sub-frames containing similar data, and only sf(0,1) 1205, sf(0,2) 1027 or one of them is used. If “offset == +3 (or -1)”, the metadata to be transmitted is determined based on one history sub-frame and the first three sub-frames of current frame: sf(-1, 4) 1203, sf(0,1) 1205, sf(0,2) 1207, sf(0,3) 1209. This way the metadata framing in the encoding matches the original framing and the optimal encoding mode can be determined based on the metadata. The resulting delay of 1 sub-frame in the metadata is compensated for in the decoder by writing the data earlier in the output buffer. It is possible to limit the operation of the system into applying asynchrony correction only with lower temporal resolution I higher frequency resolution 1 sf coding mode frames, in which case no history values are needed here on the account of the 4 sub-frames containing similar data, and only sf(0,1), sf(0,2), sf(0,3), or one of them is used. An example of the aggregation function f( ) is the energy-weighted vector average of the directional spatial metadata over the sub-frames for determining the prototypal sub-frame sf. This average can, in some embodiments, be weighted with the signal energy, weighting based on the sub-frame location within the frame, or some other weighting function. The following focuses on the embodiment of using transport signal energy for the weighting. Therefore in some embodiments the metadata to transmit determiner is configured to further receive the transport audio signals. The following example describes on employing transport signal energy for the weighting. The transport signal can be denoted with (t), where i is the index of the transport channel and t is the time index. The transport signal can be converted into time-frequency domain, e.g., with the use of a complex modulated low delay filter bank (CLDFB). The signal in this domain can be denoted with X^k.n), where k is frequency bin index and n is the slot index (time and frequency are now in CLDFB slots and bins, which may be different from the MASA parameter tile definitions). The transport signal energy E in MASA parameterTF-tile (b,s) (where b band and s is sub-frame) can be estimated by Eb,s = xi(k>n^>\2 keb nes i In this example, one or more CLDFB slots n are grouped into parameter sub-frames s, and one or more CLDFB bins k are grouped into parameter bands b. In the preferred embodiment, each parameter time sub-frame s corresponds to one of the spatial sub-frames sf(0,1), sf(0,2), sf(0,3), and sf(0,4). In other words sf(O, s). The parameter band index b is not visible in the sub-frame structure and is often determined by the available bitrate. Alternatively, the computation can be performed at the full 24 bands of MASA parametrization. Each spatial metadata sub-frame sf(0, s) contains values for azimuth 0bs, elevation (pbs, and direct-to-total energy ratio rbs. The aggregated azimuth 0b^, elevation (pb^, and direct-to-total energy ratio rb^ parameters (i.e., they depict all the subframes s of the frame f) are computed by interpreting the per-sub-frame parameter values as spherical coordinates and transforming them into Cartesian representation xbs, yb s, zbs with xbs cos(0bs) rbsEbs Xb.s sin(6b,s) cos((f)brbsEbs %b.s ^in(0b,s) ^b,sEb,s These are then averaged / summed over the sub-frames s e [1,4] with r 1\ —! •^b.^ “ N / . * %b,s -Jsef 1\ —i yb.^ ~ N A > yb,s -^SE^ 1\ —1 ^b.^ ~ Nl * %b,s -•sE^ and converted back into spherical coordinate parameterization with tan 1---- (pbi = tan 1 Xbf+yb.^2 *bt2 + yb.^2 + z«2 rb,t = where b,s and the tan-1() is the arcus tangent (or inverse tangent) variant resolving the correct quadrant. The spread coherence parameter c^r can be determined as an energy-weighted average of the per-sub-frame values with spr 1 \ ' spr Cb£ ~~c / , Eb,scb,s ^b,^ ^—‘sE^ Similarly, the surround coherence parameter c^“r can be determined as an energy-weighted average of the per-sub-frame values c^r with 1 V ^sur ___ X p „sur cbp ~ c / cb,scb,s ^b,^ 'sef As mentioned earlier, some other weighting can be used in the place of the transport signal energy Eb,s, however the operations remain otherwise similar. The spatial metadata parameters for sf(O. for the band b represent the result from the aggregation / interpolation over sub-frames. This sub-frame metadata can be replicated to all sub-frames of the current frame replacing the original values: sf (0,1) = sf (0,2) = sf (0,3) = sf (0,4) = sf (0,0 and these sub-frames are then used in the further analysis and encoding normally. The determiner can then replace the spatial metadata 104 or include them in the spatial metadata frame. In some embodiments the determiner is alternatively configured to copy one of the sub-frames available, for example, to use sf(0,2) for all subframes. This approach is especially effective when the data in the different subframes within the aggregation range is identical. This spatial metadata frame 1302 determined by the metadata to transmit determiner 1305 can then be passed to the spatial metadata encoder 1307. In some embodiments the encoder also comprises a spatial metadata encoder 1307 which is configured to receive the spatial metadata frame 1302 and asynchrony information 1300 and generate encoded spatial metadata 1304 which is passed to the audio and metadata (multiplexer) combiner 1309. The encoder furthermore comprises an audio and metadata (multiplexer) combiner 1309 configured to receive or obtain the encoded spatial metadata 1304 and the encoded transport audio 1306 and the asynchrony information 1300 and generate a bitstream 804. The combiner (or multiplexer) is configured to obtain the offset value from the asynchrony information 1300 and multiplexes this information also to the stream 804. The resulting bitstream 804 can then be output (and stored or transmitted to the decoder). In some embodiments the asynchrony information may also be signalled in a different way, e.g., as part of the encoded spatial metadata (in other words the spatial metadata encoder 1307 is configured to receive the asynchrony information and encodes it as another ‘spatial metadata parameter’ prior to the audio and metadata combiner 1309) or the asynchrony information is signalled separately outside of the bitstream. With respect to Figure 14 is shown an example flow diagram of the operations of the encoder shown in Figure 13. Thus is shown obtaining transport audio signals as indicated by 1401. Further is shown obtaining spatial metadata by 1403. Then the encoding the audio signals is shown by 1411. The obtaining of the asynchrony information with respect to the metadata (for example, analysis of the metadata, or the receiving the asynchrony information from an external analyser) is shown by 1405. Having analysed the asynchrony within the metadata then is shown the determination of the metadata to encode for transmission (or storage) by 1407. Then is shown by 1409 is the encoding of the determined spatial metadata. Having encoded the transport audio signals and the metadata is the operation of combining the encoded transport audio signals, encoded metadata and offset from the asynchrony information to generate the bitstream as shown by 1413. Finally the bitstream is output (either in the form or transmission or storing of the bitstream). The bitstream comprising the encoded transport audio signals and metadata as shown by 1415. With respect to Figure 15 is shown a schematic view of the decoder according to some embodiments. The decoder in some embodiments comprises a separator (or audio and spatial metadata demultiplexer) 1501. The separator (or audio and spatial metadata demultiplexer) 1501 can be configured to receive or otherwise obtain the bitstream 804 and unpack and separate it into encoded transport audio signal 1500, encoded spatial metadata 1510 and asynchrony information 1506 parts. In some embodiments the asynchrony information 1506 can be output from the separator 1501 in the case of out-of-band signalling (or the asynchrony information can in some embodiments be obtained outside of the bitstream 804). In some embodiments the decoder comprises audio decoder 1503, configured to receive the encoded transport audio and is further configured to generate the decoded transport audio signals 1502 and pass these to the (MASA) renderer 1509. Furthermore in some embodiments the decoder comprises a spatial metadata decoder 1505 configured to obtain the encoded spatial metadata 1510 and from this provide the (intermediate) decoded spatial metadata 1512 (and in some embodiments the asynchrony information 1508 in the case of in-band signalling) to the spatial metadata modifier 1507. The decoder in some embodiments comprises a spatial metadata modifier 1507 which is configured to receive the decoded spatial metadata 1512 and further receives the asynchrony information (and which can be the asynchrony information 1506 from the separator or the asynchrony information 1508 from the spatial metadata decoder 1505). The spatial metadata modifier 1507 is configured, based on the asynchrony information 1506 / 1508 to determine the spatial metadata frame that is forwarded to the (MASA) renderer 1509. The (MASA) renderer 1509, in some embodiments, is configured to receive or otherwise obtain the modified spatial metadata frame 1514 and determine output audio signals 1516 from the decoded transport audio signals 1502. In some embodiments the (MASA) renderer 1509 may be inside the decoder or can be implemented outside of the decoder (a so-called external renderer). The renderer can be any suitable renderer method. In some embodiments the asynchrony information is provided as an output from the decoder. For example, the asynchrony information is used in an external renderer, or as a part of spatial metadata frame or as an additional signalling. With respect to Figure 16 is shown a flow diagram showing the operation of the example decoder as shown in Figure 15. Thus is shown by 1601 the obtaining or receiving of the bitstream. Then as shown by 1603 is the separating bitstream into encoded transport audio signals, encoded metadata (and asynchrony information). Furthermore is shown by 1605, decoding the encoded transport audio signals to generate decoded transport audio signals. Additionally is shown by 1607, decoding encoded spatial metadata to generate decoded spatial metadata and asynchrony information. Then is shown by 1609, modifying the decoded spatial metadata based on asynchrony information. Furthermore as shown by 1611 is rendering audio signals from decoded transport audio signals based on the modified spatial metadata. Then as shown by 1613 is outputting the audio signals. The modification of the spatial metadata is shown in further detail hereafter. In some embodiments, as described above, the spatial metadata modifier is configured to receive or otherwise obtain the intermediate (or decoded) spatial metadata frames and asynchrony information and modifies the spatial metadata such that it is better synchronized with the audio stream. This modification may be implemented utilizing a metadata-to-audio synchronization buffer such as illustrated in Figure 17. Figure 17 thus shows a metadata-to-audio synchronization buffer 1721 showing buffered sub-frames, sub-frame 1 1703, sub-frame 2 1707, sub-frame 3 1709, sub-frame 4 1711, sub-frame 5 1713, and sub-frame 6 1715. Some embodiments may contain an additional sub-frame buffer slot 1717. In some embodiments it is possible that such a buffer does not exist for the metadata-to-audio synchronization and in such a case a delay buffer functionality needs to be implemented. A dedicated delay buffer may operate on a temporal resolution higher than a sub-frame (which is used here for the description), but the synchronization can still be done similarly by varying the offset between writing and reading indices. In this example the unpacked metadata frame 1703 range is delayed by 2 sub-frames relative to the metadata applied on audio signal 1701. This delay operation allows modifying the metadata of the last 2 sub-frames of the previous frame still in the current frame enabling the synchronization processing as discussed herein. In some embodiments as the spatial metadata needs to be delayed by 12 ms with respect to the audio in order to attempt to synchronize them. In some embodiments this can be implemented by employing a delay buffer such as conceptually depicted in Figure 17. In some embodiments therefore the current spatial metadata frame 1703 is decoded / unpacked into nominal sub-frames with indices 3 to 6 (as shown by subframe 3 1709, sub-frame 4 1711, sub-frame 5 1713, and sub-frame 6 1715), while the spatial metadata applied on the audio signals 1701 will be read from nominal indices 1 to 4 (as shown by sub-frame 1 1705, sub-frame 2 1707, sub-frame 3 1709, and sub-frame 4 1711). After each write and read step, the buffer contents may be shifted left by 4 indices (or in some embodiments, the read and write indices are updated using a modulo operation for a ring buffer). For illustration purposes the temporal resolution of the delay buffer is set to one sub-frame (leading into concrete delay of 10 ms instead of the ideal 12 ms), but it may also be higher, e.g., the 1.25 ms resolution of the CLDFB. In some embodiments the decoder is configured to receive or decode the offset information from the encoder and after unpacking the 4 metadata sub-frames (when the transport I encoding mode is the lower frequency resolution / higher frequency resolution encoding mode, 1 sf, 4 identical metadata sub-frames are obtained from the decoded data), determines the location where it should write the data in the delay buffer. In some embodiments this can be implemented as follows (using the indices which refer to the sub-frames of Figure 17): If “offset == 0”, the metadata is written into locations 3-6 (as shown by subframe 3 1709, sub-frame 4 1711, sub-frame 5 1713, and sub-frame 6 1715) and read from indices 1-4 (as shown by sub-frame 1 1705, sub-frame 2 1707, subframe 3 1709, and sub-frame 4 1711). If “offset == +31 -1”, the metadata is written into locations 2-5 (as shown by sub-frame 2 1702, sub-frame 3 1709, sub-frame 4 1711 and sub-frame 5 1713), i.e., one sub-frame index earlier. Optionally, the sub-frame written to index 5 (subframe 5 1713) is written also to index 6 (sub-frame 6 1715). The metadata is read from indices 1-4 (as shown by sub-frame 1 1705, sub-frame 2 1707, sub-frame 3 1709, and sub-frame 4 1711). If “offset == +2 I -2”, the metadata is written into locations 1-4 (as shown by sub-frame 1 1705, sub-frame 2 1707, sub-frame 3 1709, and sub-frame 4 1711), i.e., two sub-frame indices earlier. Optionally, the indices 5 and 6 (sub-frame 5 1713 and sub-frame 6 1713) are filled with extrapolated values or by copying some of the current sub-frames. The metadata is read from indices 1-4 (sub-frame 1 1705, sub-frame 2 1707, sub-frame 3 1709, and sub-frame 4 1711). If “offset == -1 I -3”, the buffer is extended by one additional index (7: “subframe buffer”) and the unpacked metadata is written into locations 4-7 (sub-frame 4 1711, sub-frame 5 1713, sub-frame 6 1715, and sub-frame buffer slot 1717). If the previous frame had “offset != -3”, sub-frame index 3 (sub-frame 3 1709) needs to be filled with a new value. This value can be determined from the metadata in the earlier and following sub-frames, e.g., by using a “hold” or data interpolation. If the previous frame also had “offset == -3”, the buffer contains the sub-frame data from the previous frame. The metadata is read from indices 1-4 (sub-frame 1 1705, sub-frame 2 1707, sub-frame 3 1709, and sub-frame 4 1711). In some embodiments the system is implemented where the Encoder and Decoder are implemented with a signalled offset. In the example presented above, the asynchrony analysis used (partial) metadata from the previous frame for determining the asynchrony. In these embodiments, the asynchrony information is determined without any history. This can therefore be implemented with the following differences compared to the above embodiments. If the encoder does not use history metadata in analysing for asynchrony information (a method for which is described in UK application number 2304318.5), then it can furthermore be configured to implement a metadata to transmit determiner without the history information. This is possible in the following manner: If “offset == 0”: use the sub-frames of the current frame without any modification: “sf(0,1), sf(0,2), sf(0,3), sf(0,4)”. If “offset == +3 I -1”: the metadata to transmit is determined based on the first 3 sub-frames of the current frame. If “mode == 1sf”: Determine a prototypal sub-frame sf(0,O by applying an aggregation function “f(sf(0,1), sf(0,2), sf(0,3))” on the 3 first sub-frames. Examples of the aggregation function include the metadata interpolation described above applied on the three metadata sub-frames or selecting one of the three metadata subframes. If “mode == 4sf”: Use the first three metadata sub-frames of the current frame and create a new metadata sub-frame for the missing value using an aggregation function: “f(sf(0,1), sf(0,2), sf(0,3)), sf(0,1), sf(0,2), sf(0,3)”. Here, the aggregation function may be similar to the aggregation function above or it may be different, alternatively, the offset is ignored and the sub-frames “sf(0,1), sf(0,2), sf(0,3), sf(0,4)” are used as-is. If “offset == +2 I -2”: the metadata to transmit is determined based on the first 2 sub-frames of the current frame. Determine a prototypal sub-frame sf (0, 0 by applying an aggregation function “f(sf(0,1), sf(0,2)),”on the 2 first sub-frames. Examples of the aggregation function include the metadata interpolation described above applied on the two metadata sub-frames or selecting one of the two metadata sub-frames. If “mode == 4sf”, Use the first two metadata sub-frames of the current frame and create two new metadata sub-frames for the missing values using an aggregation function: f(sf(0,1), sf(0,2)), f(sf(O,1), sf(0,2)), sf(0,1), sf(0,2)”. Here, the aggregation function may be similar to the aggregation function above or it may be different, alternatively, the offset is ignored and the sub-frames “sf(0,1), sf(0,2), sf(0,3), sf(0,4)” are used as-is. If “offset ==+1 / -3”: the metadata to transmit is determined from the last 3 sub-frames of the current frame. If “mode == 1sf”: Determine a prototypal sub-frame sf (0,0 by applying an aggregation function “f(sf(0,2), sf(0,3), sf(0,4))” on the 3 last sub-frames. Examples of the aggregation function include the metadata interpolation described above applied on the three metadata sub-frames or selecting one of the three metadata sub-frames. If “mode == 4sf”, Use the last three metadata sub-frames of the current frame and create a new metadata sub-frame for the missing value using an aggregation function: “sf(0,2), sf(0,3), sf(0,4), f(sf(0,2), sf(0,3), sf(0,4))”. Here, the aggregation function may be similar to the aggregation function above or it may be different, alternatively, the offset is ignored and the sub-frames “sf(0,1), sf(0,2), sf(0,3), sf(0,4)” are used as-is. In all cases with non-zero offset and “mode == 1 sf”, the prototypal sub-frame value sf (0,0 can be replicated to all sub-frames for further use replacing the original values: sf (0,1) = sf (0,2) = sf (0,3) = sf (0,4) = sf (0,0 In some embodiments the decoder does not receive explicit side information signalling the asynchrony offset but detects the asynchrony condition and offset itself. This can be implemented in a manner with a decoder similar to that described above but the bitstream does not obtain the asynchrony information from the bitstream. With respect to Figure 18 is shown a schematic view of part of the decoder as shown in Figure 15. In particular the spatial metadata decoder 1801 can replace the spatial metadata decoder 1505, the spatial metadata modifier 1805 replacing the spatial metadata modifier 1507 and an asynchrony analyser 1803 located between the spatial metadata decoder 1801 and spatial metadata modifier 1805. In some embodiments the spatial metadata decoder 1801 is configured to receive the encoded spatial metadata 1510 and generates (intermediate) decoded spatial metadata 1800. The asynchrony analyser 1803 is configured to receive the (intermediate) decoded spatial metadata 1800 and generate asynchrony information 1508 and (intermediate) decoded spatial metadata 1804, which can be both passed to the spatial metadata modifier 1805. The asynchrony analyser 1803 in some embodiments is configured to analyse the spatial metadata in a manner similar to the asynchrony analyser implemented within the encoder (as described in detail in UK application number 2304318.5). When an offset is found, the encoder can be configured to raise an internal flag, while it keeps on sending the data as if no offset was found. In other words, the high temporal resolution coding mode, 4sf, used and the each 4 similar subframes of the original encoder input data are divided into two consecutive transport frames. For example, the boundary of IVAS encoding frame 3 1020 and IVAS encoding frame 4 1030 as shown in Figure 10. The encoder stays in this mode of sub-optimal transport for one or more frames. After the pre-determined number of frames, the metadata to be transmitted is selected and low temporal resolution I high frequency resolution, 1 sf, coding mode is used. The larger the number of frames to stay in the intermediate mode is, the more reliable the detection in the decoder will be (robustness against packetloss), but during this intermediate state sub-optimal coding is used. With respect to the asynchrony analyser 1803 the following occurs: In the initial state, no offset has been found; While no offset is found, the metadata received is analyzed for similar subframes. The analysis may be similar to the analysis with state history or analysis without state history as discussed in UK application number 2304318.5. When a non-zero offset has been found and the decoder receives 4 subframes of similar data (i.e. data was encoded using the low temporal resolution mode 1 sf), it knows that the encoder has switched into a mode compensating for the offset. In other words, when a non-zeros offset has been found and then the analysis returns “1sf” and “offset=0” as the result, the earlier non-zero offset is used in a modification method similar to that described above adjusting location of the metadata in the synchronization buffer. In some embodiments the asynchrony analyser is configured to rely on a part of the metadata being transmitted twice: once with low (using the high temporal resolution / low frequency resolution mode 4sf) and once with high (using low temporal resolution I high frequency resolution mode 1 sf) frequency resolution, as illustrated by the flow diagram in Figure 19. This requires that the encoder also is also configured to perform the asynchrony determination of detection, and adjusting the metadata transmitted I encoded based on the analysis result. - Framing offset 0 frames: o Everything is processed without additional adjustments. - Framing offset +2 (or -2) sub-frames: o Encoder: The transition frame from “4sf to “1 sf’, i.e., “IVAS encoding frame 3” in Figure 10, is transmitted as-is with 4 sub-frames, using high temporal resolution mode 4sf. The encoder detects the framing asynchrony and transmits the next frame “IVAS encoding frame 4” with low temporal resolution mode 1sf using “offset = 2”. In other words, the solid light grey sub-frames in “IVAS encoding frame 3” are transmitted twice: in “IVAS encoding frame 3” with high temporal resolution I low frequency resolution mode 4sf and then in IVAS encoding frame 4” with low temporal resolution I high frequency resolution mode 1sf. o Decoder: The decoder compares the unpacked content of the current and previous frame. It detects that the values in the previous frame (originally encoded using high temporal resolution mode 4sf) are in fact stemming from the same data when using the lower frequency resolution. The detection algorithm is described below. - Framing offset +3 (or -1) sub-frames: o Operation is similar to the case of “framing offset +2 sub-frames”. The detection or determination can be shown in Figure 19. For example, a first check as shown by 1901 is one of determining whether the previous frame was a low frequency resolution (high temporal resolution). If no then asynchrony offset value is kept as shown in Figure 19 by 1903. If yes then a further check whether the current frame is a high frequency resolution is made as shown by 1905. If previous frame was encoded using high temporal resolution mode 4sf and the current frame is encoded using high temporal resolution mode 4sf, nothing is done and the asynchrony offset value is kept as shown in Figure 19 by 1907. If previous frame was encoded using high temporal resolution mode 4sf and current frame is encoded using low temporal resolution mode 1 sf, check if asynchrony is present. For each sub-frame s in the current (“4sf”) frame, do: 1. Compute the per parameter tile transport signal energy Eb s with Eb,s = kEb nEs i o The transport signal is denoted with x^t), where i is the index of the transport channel and t is the time index. The transport signal is converted into time-frequency domain, e.g., with the use of CLDFB. This is denoted with X^k, ri), where k is frequency bin index and n is the time slot index (now in CLDFB bins and slots, which may be different from the MASA parameter bands and frames). 2. Compute the per parameter tile vector representation from the azimuth dbiS, elevation <pbs, and direct-to-total energy ratio rb,s with xb,s COS^0£, ^b^Eb,s Xb.s xb,sEb,s Zb,s xb,sEb,s 3. Compute the parameter vector representation using the parameter frequency resolution of the previous frame. This is done with (1) yb,s b where NB is the number of high-resolution parameter bands assigned to the low-resolution parameter band B. 4. Convert the parameters back to the azimuth + elevation representation with b,s = tan 1 — xB.s _ Z^ $ 0b,s = tan-1 -^=== VXB.s + yB.s2 y xB.s T yB.s T XB.s rB,s =---------7----------- V cB,s -] where EBs = —2b Eb,s and tan-1() is the arcus tangent variant resolving the correct quadrant. 5. Analogously, reduce the frequency resolution of the spread coherence c^‘ and surround coherence cBul parameters with ' bB,s^—‘b and „sur _ \1 p ^sur ^B,s p / ^b.s^b.s bB,s ^b This produces a set of metadata sub-frames with lower frequency resolution: “sf(01j, sf(02), sf(0;3), sf(04)" . The offset can be now detected using the algorithm depicted in the flow diagram of Figure 20. With respect to Figure 20 is shown the detection of the offset. The previous frame is received or obtained as shown by 2000. The current frame is received or obtained as shown by 2001. The current frame has a frequency resolution reduction as shown by 2003. Then a check whether there is a similarity between the current frame first sub-frame and the previous frame last sub-frame as shown by 2005. If they are not similar then keep the previous offset as shown by 2007. If they are similar then check whether there is a similarity between the current frame first sub-frame and the previous frame third sub-frame as shown by 2009. If they are similar then the offset is a +2 sub-frame offset as shown by 2011. If they are not similar then the offset is a -1 sub-frame offset as shown by 2013. In some embodiments the asynchrony information, in other words the offset, determined is provided to the spatial metadata modifier along with the original metadata before frequency resolution reduction. The processing and decoding can then be implemented in a manner as described earlier. In some embodiments, as shown in Figure 21, there is shown a system wherein an encoder 2101 which is similar to the encoder 801 as shown in Figure 8, but with an additional input referred as operation mode 2180. The operation mode can select which of the above embodiments to employ. Furthermore the operation mode 2180 can be determined by an operation mode controller 2111. The operation mode controller 2111 is configured to receive operation parameters 2160, for example, a total allowable complexity or memory and based on these select the operation mode that is most fitting to the given operation parameters. For example in some embodiments the encoder may have constraints on the usage of the memory. The constraints for the metadata encoding may be different at different bitrates. So, as an example, the operation parameters could also comprise a total bitrate for encoding the input signals. The operation mode controller could, for example, in some embodiments be configured to select the asynchrony analysis method based on the used bitrate and the pre-determined memory constraint for that bitrate. For example, the additional signalling needed by the earlier signalling of asynchrony information may not be feasible at extremely low bitrates, or the encoder embodiments with the additional history may not be feasible due to implementation constraints. In some embodiments the selection of the operation mode 2180 depending on the operation parameters include: If the bitrate budget is larger than a given threshold, select an “Operation mode” making use of additional signaling, If the bitrate budget is below a given threshold, and no additional decoder complexity is allowed, use an operation mode without history and without explicit signaling, If the bitrate budget is below a given threshold and additional decoder complexity is allowed, use an operation mode with no explicit signaling and no history. In the examples shown above Complex-Values Low-Delay Filter Bank (CLDFB) as the frequency-domain representation other methods of time-frequency domain representation can be employed such as Short-time Fourier Transform (STFT) or Quadrature Mirrored Filterbank (QMF). The parametric format described above has been MASA format but the embodiments can be extended to other parametric formats such as parametric coding of Ambisonics or multi-channel mixes. With respect to Figure 22 an example electronic device which may be used as any of the apparatus parts of the system as described above. The device may be any suitable electronics device or apparatus. For example, in some 52 embodiments the device 2200 is a mobile device, user equipment, tablet computer, computer, audio playback apparatus, etc. The device may for example be configured to implement the encoder and / or decoder or any functional block as described above. In some embodiments the device 2200 comprises at least one processor or central processing unit 2207. The processor 2207 can be configured to execute various program codes such as the methods such as described herein. In some embodiments the device 2200 comprises at least one memory 2211. In some embodiments the at least one processor 2207 is coupled to the memory 2211. The memory 2211 can be any suitable storage means. In some embodiments the memory 2211 comprises a program code section for storing program codes implementable upon the processor 2207. Furthermore, in some embodiments the memory 2211 can further comprise a stored data section for storing data, for example data that has been processed or to be processed in accordance with the embodiments as described herein. The implemented program code stored within the program code section and the data stored within the stored data section can be retrieved by the processor 2207 whenever needed via the memory-processor coupling. In some embodiments the device 2200 comprises a user interface 2205. The user interface 2205 can be coupled in some embodiments to the processor 2207. In some embodiments the processor 2207 can control the operation of the user interface 2205 and receive inputs from the user interface 2205. In some embodiments the user interface 2205 can enable a user to input commands to the device 2200, for example via a keypad. In some embodiments the user interface 2205 can enable the user to obtain information from the device 2200. For example, the user interface 2205 may comprise a display configured to display information from the device 2200 to the user. The user interface 2205 can in some embodiments comprise a touch screen or touch interface capable of both enabling information to be entered to the device 2200 and further displaying information to the user of the device 2200. In some embodiments the user interface 2205 may be the user interface for communicating. In some embodiments the device 2200 comprises an input / output port 2209. The input / output port 2209 in some embodiments comprises a transceiver. The transceiver in such embodiments can be coupled to the processor 2207 and configured to enable a communication with other apparatus or electronic devices, for example via a wireless communications network. The transceiver or any suitable transceiver or transmitter and / or receiver means can in some embodiments be configured to communicate with other electronic devices or apparatus via a wire or wired coupling. The transceiver can communicate with further apparatus by any suitable known communications protocol. For example, in some embodiments the transceiver can use a suitable radio access architecture based on long term evolution advanced (LTE Advanced, LTE-A) or new radio (NR) (or can be referred to as 5G), universal mobile telecommunications system (UMTS) radio access network (UTRAN or E-UTRAN), long term evolution (LTE, the same as E-UTRA), 2G networks (legacy network technology), wireless local area network (WLAN or Wi-Fi), worldwide interoperability for microwave access (WiMAX), Bluetooth®, personal communications services (PCS), ZigBee®, wideband code division multiple access (WCDMA), systems using ultra-wideband (UWB) technology, sensor networks, mobile ad-hoc networks (MANETs), cellular internet of things (loT) RAN and Internet Protocol multimedia subsystems (IMS), any other suitable option and / or any combination thereof. The transceiver input / output port 1409 may be configured to receive the signals. In some embodiments the device 1400 may be employed as at least part of the synthesis device. The input / output port 1409 may be coupled to headphones (which may be a headtracked or a non-tracked headphones) or similar and loudspeakers. In general, the various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto. While various aspects of the invention may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof. The embodiments of this invention may be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD. The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processors may be of any type suitable to the local technical environment, and may include one or more of general-purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), gate level circuits and processors based on multi-core processor architecture, as non-limiting examples. Embodiments of the inventions may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate. Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre-stored design modules. Once the design for a semiconductor circuit has been completed, the resultant design, in a standardized electronic format (e.g., Opus, GDSII, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication. As used in this application, the term “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry) and (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation. This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device. The term “non-transitory,” as used herein, is a limitation of the medium itself (i.e., tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM). As used herein, “at least one of the following: ” and “at least one of ” and similar wording, where the list of two or more elements are joined by “and” or “or”, mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements. The foregoing description has provided by way of exemplary and nonlimiting examples a full and informative description of the exemplary embodiment of this invention. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, 5 when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still fall within the scope of this invention as defined in the appended claims.
Claims
1. An apparatus comprising means for:obtaining, at least one encoded spatial metadata associated with at least one encoded audio signal, the at least one spatial metadata parameter being arranged in frames comprising at least two sub-frames;obtaining asynchrony information, wherein the asynchrony information is based on asynchrony between a sequence of at least two time sub-frames containing similar values and a frame for any further processing the at least one spatial metadata parameter;decoding the at least one encoded spatial metadata parameter to generate a sequence of the at least two sub-frames of spatial metadata parameters; andprocessing the frame comprising the at least two time sub-frames based on the asynchrony information, such that the processed frame comprising the at least two time sub-frames of at least one spatial metadata parameter is able to be employed in a rendering of audio signals.
2. The apparatus as claimed in claim 1, wherein the sequence of at least two time sub-frames containing similar values is one of:located within the same frame; andlocated over two consecutive frames.
3. The apparatus as claimed in any of claims 1 or 2, wherein the means is further for:obtaining the at least one encoded audio signal;decoding the at least one encoded audio signal; andrendering an output audio signal from the decoded at least one encoded audio signal based on the processed frame comprising the at least two time subframes of at least one spatial metadata parameter.
4. The apparatus as claimed in any of claims 1 to 3, wherein the means for obtaining asynchrony information is for one of:analysing the decoded at least one encoded spatial metadata parameter to determine the asynchrony information;receiving the asynchrony information from at least one further apparatus having obtained the asynchrony information; andreceiving the asynchrony information from at least one further apparatus having analysed the at least one spatial metadata parameter to determine the asynchrony information.
5. The apparatus as claimed in claim 4 when dependent on claim 2, wherein the means for obtaining the at least one encoded audio signal is for receiving the at least one encoded audio signal from the at least one further apparatus.
6. The apparatus as claimed in any of claims 1 to 5, wherein the means for obtaining asynchrony information is for obtaining an offset value identifying a temporal difference between the frame comprising at least two time sub-frames and a rendering frame, the temporal difference obtained as a number of subframes.
7. The apparatus as claimed in claim 6, wherein the means for processing the frame comprising the at least two time sub-frames based on the asynchrony information is for processing the frame comprising the at least two time sub-frames based on the offset value.
8. The apparatus as claimed in claim 7, wherein the means for processing the frame comprising the at least two time sub-frames based on the offset value is for applying an alignment buffer configured based on the offset value to modify a time associated with the processed decoded at least one encoded spatial metadata parameter relative to the decoded at least one encoded audio signal.
9. The apparatus as claimed in any of claims 1 to 8, wherein the means for obtaining asynchrony information is for obtaining an encoding mode identifying an encoding mode for the at least one encoded spatial metadata parameter.
10. The apparatus as claimed in claim 4, or any claim dependent on claim 4, wherein the means for analysing the decoded at least one encoded spatial metadata parameter to determine the asynchrony information is for:analysing for a current frame of the decoded at least one encoded spatial metadata parameter to determine: a number of mutually similar sub-frames from the current frame start and a number of mutually similar sub-frames from the current frame end; anddetermining the asynchrony information based on the number of mutually similar sub-frames from the current frame start and / or the number of mutually similar sub-frames from the current frame end.
11. The apparatus as claimed in claim 10, wherein the means for analysing the at least one spatial metadata parameter to determine the asynchrony information is further for:analysing for a last-sub-frame of a previous frame and a first-sub-frame of a current frame whether these are mutually similar sub-frames; andobtaining for the previous frame a number of mutually similar sub-frames;wherein the means for determining the asynchrony information based on the number of mutually similar sub-frames from the current frame start and / or the number of mutually similar sub-frame from the current frame end is further for determining the asynchrony information based on whether the last-sub-frame of a previous frame and a first-sub-frame of a current frame are mutually similar subframes; and the number of mutually similar sub-frames for the previous sub-frame.
12. The apparatus as claimed in any of claims 1 to 11, wherein the means is further for obtaining an operation mode indicator; and wherein the means for processing the frame comprising the at least two time sub-frames based on the asynchrony information is for processing the decoded at least one encoded spatial metadata parameter based on the asynchrony information and the operation mode indicator.
13. An apparatus comprising means for:obtaining at least one spatial metadata parameter associated with at least one audio signal, the at least one spatial metadata being arranged in frames comprising at least two sub-frames;obtaining asynchrony information, wherein the asynchrony information is based on asynchrony between a sequence of at least two time sub-frames containing similar values and a frame for encoding the at least one spatial metadata parameter; andencoding the asynchrony information, and the frame comprising the at least two time sub-frames based on the asynchrony information.
14. The apparatus as claimed in claim 13, wherein the sequence of at least two time sub-frames containing similar values is one of:located within the same frame; andlocated over two consecutive frames15. The apparatus as claimed in any of claims 13 or 14, wherein the means is further for:obtaining the at least one audio signal; andencoding the at least one audio signal based on the frame for encoding the at least one spatial metadata parameter.
16. The apparatus as claimed in any of claims 13 to 15, wherein the means for obtaining asynchrony information is for one of:analysing the at least one spatial metadata parameter to determine the asynchrony information;receiving the asynchrony information from at least one further apparatus having determined the asynchrony information; andreceiving the asynchrony information from at least one further apparatus having analysed the at least one spatial metadata parameter to determine the asynchrony information.
17. The apparatus as claimed in any of claims 13 to 16, wherein the means for obtaining asynchrony information is for obtaining an offset value identifying anumber of sub-frames which is a difference between the frame comprising at least two time sub-frames and the encoding frame, the temporal difference obtained as a number of sub-frames.
18. The apparatus as claimed in claim 17, wherein the means for encoding the frame comprising the at least two time sub-frames based on the asynchrony information is for:processing the at least one spatial metadata parameter based on the offset value; andencoding the processed at least one spatial metadata parameter.
19. The apparatus as claimed in claim 18, wherein the means processing the at least one spatial metadata parameter based on the offset value is for determining an aggregation of the at least two sub-frames to be a single spatial metadata parameter sub-frame to represent the frame,20. The apparatus as claimed in any of claims 13 to 19, wherein the means for obtaining asynchrony information is for obtaining an encoding mode identifying an encoding mode for encoding the at least one audio signal and the at least one spatial metadata parameter based on the asynchrony information.
21. The apparatus as claimed in claim 16, or any claim dependent on claim 16, wherein the means for analysing the at least one spatial metadata parameter to determine the asynchrony information is for:analysing for a current frame at least one spatial metadata parameter associated with the at least one audio signal to determine: a number of mutually similar sub-frames from the current frame start and a number of mutually similar sub-frames from the current frame end;determining the asynchrony information based on the number of mutually similar sub-frames from the current frame start and / or the number of mutually similar sub-frames from the current frame end.
22. The apparatus as claimed in claim 21, wherein the means for analysing the at least one spatial metadata parameter to determine the asynchrony information is further for:analysing for a last-sub-frame of a previous frame and a first-sub-frame of a current frame whether these are mutually similar sub-frames; andobtaining for the previous frame a number of mutually similar sub-frames;wherein the means for determining the asynchrony information based on the number of mutually similar sub-frames from the current frame start and / or the number of mutually similar sub-frame from the current frame end is further for determining the asynchrony information based on whether the last-sub-frame of a previous frame and a first-sub-frame of a current frame are mutually similar subframes; and the number of mutually similar sub-frames for the previous sub-frame.
23. The apparatus as claimed in any of claims 13 to 22, wherein the means for encoding the asynchrony information, and the at least one spatial metadata parameter is for, based on the asynchrony information identifying asynchrony between the spatial metadata parameter frame comprising at least two time subframes and the encoding frame for encoding the at least one audio signal, encoding the spatial metadata parameters in a first encoding mode for a number of frames before switching to processing and encoding the spatial metadata parameters in a second encoding mode.
24. A method comprising:obtaining, at least one encoded spatial metadata associated with at least one encoded audio signal, the at least one spatial metadata parameter being arranged in frames comprising at least two sub-frames;obtaining asynchrony information, wherein the asynchrony information is based on asynchrony between a sequence of at least two time sub-frames containing similar values and a frame for any further processing the at least one spatial metadata parameter;decoding the at least one encoded spatial metadata parameter to generate a sequence of the at least two sub-frames of spatial metadata parameters; andprocessing the frame comprising the at least two time sub-frames based on the asynchrony information, such that the processed frame comprising the at least two time sub-frames of at least one spatial metadata parameter is able to be employed in a rendering of audio signals.
525. A method comprising:obtaining at least one spatial metadata parameter associated with at least one audio signal, the at least one spatial metadata being arranged in frames comprising at least two sub-frames;10 obtaining asynchrony information, wherein the asynchrony information isbased on asynchrony between a sequence of at least two time sub-frames containing similar values and a frame for encoding the at least one spatial metadata parameter; andencoding the asynchrony information, and the frame comprising the at least 15 two time sub-frames based on the asynchrony information.
Citation Information
Patent Citations
Spatial audio parameter encoding and associated decoding
US20220366918A1
Method and apparatus for calculating down-mixed signal
WO2019227931A1
Silence descriptor using spatial parameters
WO2023031498A1