Spatial metadata post enhancement with machine learning during packet losses
Patent Information
- Application Number
- GB2025002048
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2026-09-16
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Field The present application relates to apparatus and methods for metadata processing for a spatial audio stream, but not exclusively for post enhancement spatial metadata processing during packet losses. Background Immersive audio codecs are being implemented supporting a multitude of operating points ranging from a low bit rate operation to transparency. An example of such a codec is the 3GPP Immersive Voice and Audio Services (IVAS) codec which is designed to be suitable for use over a communications network such as a 4G / 5G network including use in such immersive services as for example immersive voice and audio for virtual reality (VR). This audio codec handles the encoding, decoding and rendering of speech, music, and generic audio. It supports a variety of input formats, such as channel-based audio, object-based audio, and scenebased audio inputs including spatial information about the sound field and sound sources, as well as MASA (Metadata-assisted spatial audio) inputs. IVAS operates with low latency to enable conversational services as well as supports high error robustness under various transmission conditions. The IVAS codec operates on a wide range bit rates from very low (13.2 kbps) to relatively high bit rates (512 kbps). Additionally, the application of machine learning (ML), and more specifically the application of artificial or deep neural networks (ANNs, DNNs) to assist in processing operations is known. A neural network (NN) model is composed of a number of interconnected layers, each layer representing a set of operations (e.g., matrix multiplications, additions, convolutions, or non-linear operations) defining a graph of computational operations. These operations may have processing coefficients (or parameters, or weights) that are adapted during the model training phase based on the training data. A benefit of ML methods, especially those of DNNs, is that they are able to model highly complex relationships in the data without the need for manually describing those relationships. In many cases the relationships are so complex that it is currently not feasible or even possible to describe them manually. Instead, the ML methods are able to determine and model the complex relationships based on the training data examples with method inputs and expected model outputs. Summary There is provided according to a first aspect an apparatus for selectively enhancing at least one spatial metadata parameter, the apparatus comprising at least one processor and at least one memory including a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform: obtaining a spatial bitstream, the spatial bitstream defining spatial audio content and comprising: at least one metadata parameter associated with at least one audio signal, the at least one metadata parameter being an encoded version of an original metadata parameter; obtaining a first machine-learning model for enhancing the at least one metadata parameter; obtaining a second machine-learning model for determining the at least one metadata parameter when the at least one metadata parameter is missing; determining for the spatial bitstream whether content comprising spatial audio content has been obtained in a current frame or is missing; selecting one of the first machine-learning model or second machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content has been obtained in the current frame or is missing; and enhancing the at least one metadata parameter for the current frame which has been obtained or determining the at least one metadata parameter which was missing based on the selected one of the first machine-learning model or second machine-learning model. The apparatus may be further caused to perform determining at least one current frame input feature for the first machine-learning model based on the determination for spatial bitstream whether content comprising spatial audio content has been obtained for the current frame. The apparatus caused to perform determining at least one current frame input feature for the selected first machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content has been obtained for the current frame may be further caused to perform determining the at least one current frame input feature for the selected first machine-learning model based on the at least one metadata parameter associated with the at least one audio signal from the current frame of the spatial bitstream. The apparatus caused to perform determining at least one current frame input feature for the selected first machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content has been obtained for the current frame may be further caused to perform: determining the at least one current frame input feature for the selected first machine-learning model further based on the at least one audio signal from the current frame of the spatial bitstream. The apparatus may be further caused to perform determining at least one spatial metadata parameter without determining input features for the current frame based on the determination of the spatial bitstream whether content comprising spatial audio content is missing. The apparatus caused to perform selecting one of the first machine-learning model or second machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content is missing may be caused to perform selecting the second machine-learning model. The apparatus caused to perform determining at least one current frame input feature for the second machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content is missing may be further caused to perform: determining at least one of: an estimated audio signal based on at least one previous frame audio signal; and at least one estimated metadata parameter associated with the at least one audio signal based on at least one previous frame metadata parameter; and determining the at least one current frame input feature based on the at least one of: the estimated audio signal; and the at least one estimated metadata parameter. The apparatus may be further caused to perform determining for all from a determined number of previous frames whether content comprising spatial audio content has been obtained or any content comprising spatial audio content is missing. The apparatus caused to perform determining at least one previous frame input feature for the selected one of the first machine-learning model or second machine-learning model may be caused to perform determining the at least one previous frame input feature based on determining for all from a determined number of previous frames whether content comprising spatial audio content has been obtained or any content comprising spatial audio content is missing. The apparatus caused to perform determining the at least one previous frame input feature for the selected one of the first machine-learning model or second machine-learning model may be further caused to perform: when the at least one previous frame content comprising spatial audio content has been obtained determining the at least one input feature based on the at least one metadata parameter associated with the at least one audio signal from the at least one previous frame; and when at least one from the determined number of previous frames content comprising spatial audio content is missing: determining, for the previous frame with missing spatial audio content, at least one generated metadata parameter based on at least one enhanced metadata parameter and / or predicted metadata property of the previous frame with missing spatial audio content; and determining the at least one input feature based on the generated metadata parameter. The apparatus caused to perform determining the at least one previous frame input feature for the selected one of the first machine-learning model or second machine-learning model may be further caused to perform: when the at least one previous frame content comprising spatial audio content has been obtained determining the at least one input based on the at least one audio signal from the at least one previous frame of the spatial bitstream; and when at least one from the determined number of previous frames content comprising spatial audio content is missing: determining at least one generated audio signal, for the previous frame with missing spatial audio content, based on at least one prior frame audio signal, the prior frame being before the previous frame; and determining the at least one input feature based on the generated audio signal. The apparatus caused to perform obtaining the first machine-learning model may be further caused to perform, at least one of: obtaining the first machine learning model based on the spatial audio content; and receiving the first machine-learning model from at least one further apparatus. The apparatus caused to perform obtaining the second machine-learning model may be further caused to perform at least one of: obtaining the second machine-learning model based on the spatial audio content with a temporal offset first machine-learning model spatial audio content; and receiving the second machine-learning model from at least one further apparatus. The apparatus caused to perform enhancing the at least one metadata parameter for the current frame which has been obtained or determining the at least one metadata parameter for the current frame which was missing based on the selected one of the first machine-learning model or second machine-learning model may be further caused to perform: employing the selected one of the first machine-learning model or second machine-learning model, the selected one of the first machine-learning model or second machine-learning model configured to output at least one predicted metadata property; and generating at least one enhanced metadata parameter for the current frame which has been obtained or predicting the at least one metadata parameter for the current frame which was missing based on the at least one predicted metadata property. According to a second aspect there is provided an apparatus for selectively enhancing at least one spatial metadata parameter, the apparatus comprising means configured to: obtain a spatial bitstream, the spatial bitstream defining spatial audio content and comprising: at least one metadata parameter associated with at least one audio signal, the at least one metadata parameter being an encoded version of an original metadata parameter; obtain a first machine-learning model for enhancing the at least one metadata parameter; obtain a second machine-learning model for determining the at least one metadata parameter when the at least one metadata parameter is missing; determine for the spatial bitstream whether content comprising spatial audio content has been obtained in a current frame or is missing; select one of the first machine-learning model or second machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content has been obtained in the current frame or is missing; and enhance the at least one metadata parameter for the current frame which has been obtained or determining the at least one metadata parameter which was missing based on the selected one of the first machine-learning model or second machine-learning model. The means may be further configured to determine at least one current frame input feature for the first machine-learning model based on the determination for spatial bitstream whether content comprising spatial audio content has been obtained for the current frame. The means configured to determine at least one current frame input feature for the selected first machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content has been obtained for the current frame may further be configured to determine the at least one current frame input feature for the selected first machine-learning model based on the at least one metadata parameter associated with the at least one audio signal from the current frame of the spatial bitstream. The means configured to determine at least one current frame input feature for the selected first machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content has been obtained for the current frame may be configured to determine the at least one current frame input feature for the selected first machine-learning model further based on the at least one audio signal from the current frame of the spatial bitstream. The means may be configured to determine at least one spatial metadata parameter without determining input features for the current frame based on the determination of the spatial bitstream whether content comprising spatial audio content is missing. The means configured to select one of the first machine-learning model or second machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content is missing may be configured to select the second machine-learning model. The means configured to determine at least one current frame input feature for the second machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content is missing may be further configured to: determine at least one of: an estimated audio signal based on at least one previous frame audio signal; and at least one estimated metadata parameter associated with the at least one audio signal based on at least one previous frame metadata parameter; and determine the at least one current frame input feature based on the at least one of: the estimated audio signal; and the at least one estimated metadata parameter. The means may be configured to determine for all from a determined number of previous frames whether content comprising spatial audio content has been obtained or any content comprising spatial audio content is missing. The means configured to determine at least one previous frame input feature for the selected one of the first machine-learning model or second machine-learning model may be configured to determine the at least one previous frame input feature based on determining for all from a determined number of previous frames whether content comprising spatial audio content has been obtained or any content comprising spatial audio content is missing. The means configured to determine the at least one previous frame input feature for the selected one of the first machine-learning model or second machine-learning model may be further configured to: when the at least one previous frame content comprising spatial audio content has been obtained determine the at least one input feature based on the at least one metadata parameter associated with the at least one audio signal from the at least one previous frame; and when at least one from the determined number of previous frames content comprising spatial audio content is missing: determine, for the previous frame with missing spatial audio content, at least one generated metadata parameter based on at least one enhanced metadata parameter and / or predicted metadata property of the previous frame with missing spatial audio content; and determine the at least one input feature based on the generated metadata parameter. The means configured to determine the at least one previous frame input feature for the selected one of the first machine-learning model or second machine-learning model may be further configured to: when the at least one previous frame content comprising spatial audio content has been obtained determine the at least one input based on the at least one audio signal from the at least one previous frame of the spatial bitstream; and when at least one from the determined number of previous frames content comprising spatial audio content is missing: determine at least one generated audio signal, for the previous frame with missing spatial audio content, based on at least one prior frame audio signal, the prior frame being before the previous frame; and determine the at least one input feature based on the generated audio signal. The means configured to obtain the first machine-learning model may be configured to, at least one of: obtain the first machine-learning model based on the spatial audio content; and receive the first machine-learning model from at least one further apparatus. The means configured to obtain the second machine-learning model may be further configured to at least one of: obtain the second machine-learning model based on the spatial audio content with a temporal offset first machine-learning model spatial audio content; and receive the second machine-learning model from at least one further apparatus. The means configured to enhance the at least one metadata parameter for the current frame which has been obtained or determine the at least one metadata parameter for the current frame which was missing based on the selected one of the first machine-learning model or second machine-learning model may be further configured to: employ the selected one of the first machine-learning model or second machine-learning model, the selected one of the first machine-learning model or second machine-learning model configured to output at least one predicted metadata property; and generate at least one enhanced metadata parameter for the current frame which has been obtained or predicting the at least one metadata parameter for the current frame which was missing based on the at least one predicted metadata property. According to a third aspect there is provided a method for an apparatus for selectively enhancing at least one spatial metadata parameter, the method comprising: obtaining a spatial bitstream, the spatial bitstream defining spatial audio content and comprising: at least one metadata parameter associated with at least one audio signal, the at least one metadata parameter being an encoded version of an original metadata parameter; obtaining a first machine-learning model for enhancing the at least one metadata parameter; obtaining a second machine-learning model for determining the at least one metadata parameter when the at least one metadata parameter is missing; determining for the spatial bitstream whether content comprising spatial audio content has been obtained in a current frame or is missing; selecting one of the first machine-learning model or second machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content has been obtained in the current frame or is missing; and enhancing the at least one metadata parameter for the current frame which has been obtained or determining the at least one metadata parameter which was missing based on the selected one of the first machine-learning model or second machine-learning model. The method may further comprise determining at least one current frame input feature for the first machine-learning model based on the determination for spatial bitstream whether content comprising spatial audio content has been obtained for the current frame. Determining at least one current frame input feature for the selected first machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content has been obtained for the current frame may further comprise determining the at least one current frame input feature for the selected first machine-learning model based on the at least one metadata parameter associated with the at least one audio signal from the current frame of the spatial bitstream. Determining at least one current frame input feature for the selected first machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content has been obtained for the current frame may further comprise: determining the at least one current frame input feature for the selected first machine-learning model further based on the at least one audio signal from the current frame of the spatial bitstream. The method may further comprise determining at least one spatial metadata parameter without determining input features for the current frame based on the determination of the spatial bitstream whether content comprising spatial audio content is missing. Selecting one of the first machine-learning model or second machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content is missing method may further comprise selecting the second machine-learning model. Determining at least one current frame input feature for the second machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content is missing may further comprise: determining at least one of: an estimated audio signal based on at least one previous frame audio signal; and at least one estimated metadata parameter associated with the at least one audio signal based on at least one previous frame metadata parameter; and determining the at least one current frame input feature based on the at least one of: the estimated audio signal; and the at least one estimated metadata parameter. The method may further comprise determining for all from a determined number of previous frames whether content comprising spatial audio content has been obtained or any content comprising spatial audio content is missing. Determining at least one previous frame input feature for the selected one of the first machine-learning model or second machine-learning model may further comprise determining the at least one previous frame input feature based on determining for all from a determined number of previous frames whether content comprising spatial audio content has been obtained or any content comprising spatial audio content is missing. Determining the at least one previous frame input feature for the selected one of the first machine-learning model or second machine-learning model may further comprise: when the at least one previous frame content comprising spatial audio content has been obtained determining the at least one input feature based on the at least one metadata parameter associated with the at least one audio signal from the at least one previous frame; and when at least one from the determined number of previous frames content comprising spatial audio content is missing: determining, for the previous frame with missing spatial audio content, at least one generated metadata parameter based on at least one enhanced metadata parameter and / or predicted metadata property of the previous frame with missing spatial audio content; and determining the at least one input feature based on the generated metadata parameter. Determining the at least one previous frame input feature for the selected one of the first machine-learning model or second machine-learning model may further comprise: when the at least one previous frame content comprising spatial audio content has been obtained determining the at least one input based on the at least one audio signal from the at least one previous frame of the spatial bitstream; and when at least one from the determined number of previous frames content comprising spatial audio content is missing: determining at least one generated audio signal, for the previous frame with missing spatial audio content, based on at least one prior frame audio signal, the prior frame being before the previous frame; and determining the at least one input feature based on the generated audio signal. Obtaining the first machine-learning model may further comprise, at least one of: obtaining the first machine-learning model based on the spatial audio content; and receiving the first machine-learning model from at least one further apparatus. Obtaining the second machine-learning model may further comprise at least one of: obtaining the second machine-learning model based on the spatial audio content with a temporal offset first machine-learning model spatial audio content; and receiving the second machine-learning model from at least one further apparatus. Enhancing the at least one metadata parameter for the current frame which has been obtained or determining the at least one metadata parameter for the current frame which was missing based on the selected one of the first machine-learning model or second machine-learning model may further comprise: employing the selected one of the first machine-learning model or second machine-learning model, the selected one of the first machine-learning model or second machine-learning model configured to output at least one predicted metadata property; and generating at least one enhanced metadata parameter for the current frame which has been obtained or predicting the at least one metadata parameter for the current frame which was missing based on the at least one predicted metadata property. According to a fourth aspect there is provided an apparatus for selectively enhancing at least one spatial metadata parameter, the apparatus comprising: means for obtaining a spatial bitstream, the spatial bitstream defining spatial audio content and comprising: at least one metadata parameter associated with at least one audio signal, the at least one metadata parameter being an encoded version of an original metadata parameter; means for obtaining a first machine-learning model for enhancing the at least one metadata parameter; means for obtaining a second machine-learning model for determining the at least one metadata parameter when the at least one metadata parameter is missing; means for determining for the spatial bitstream whether content comprising spatial audio content has been obtained in a current frame or is missing; means for selecting one of the first machine-learning model or second machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content has been obtained in the current frame or is missing; and means for enhancing the at least one metadata parameter for the current frame which has been obtained or determining the at least one metadata parameter which was missing based on the selected one of the first machine-learning model or second machine-learning model. According to a fifth aspect there is provided a computer program comprising instructions [or a computer readable medium comprising program instructions] for selectively enhancing at least one spatial metadata parameter, to perform at least the following: obtaining a spatial bitstream, the spatial bitstream defining spatial audio content and comprising: at least one metadata parameter associated with at least one audio signal, the at least one metadata parameter being an encoded version of an original metadata parameter; obtaining a first machine-learning model for enhancing the at least one metadata parameter; obtaining a second machine-learning model for determining the at least one metadata parameter when the at least one metadata parameter is missing; determining for the spatial bitstream whether content comprising spatial audio content has been obtained in a current frame or is missing; selecting one of the first machine-learning model or second machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content has been obtained in the current frame or is missing; and enhancing the at least one metadata parameter for the current frame which has been obtained or determining the at least one metadata parameter which was missing based on the selected one of the first machine-learning model or second machine-learning model. According to a sixth aspect there is provided a non-transitory computer readable medium comprising program instructions for causing an apparatus for selectively enhancing at least one spatial metadata parameter, to perform at least the following: obtaining a spatial bitstream, the spatial bitstream defining spatial audio content and comprising: at least one metadata parameter associated with at least one audio signal, the at least one metadata parameter being an encoded version of an original metadata parameter; obtaining a first machine-learning model for enhancing the at least one metadata parameter; obtaining a second machine-learning model for determining the at least one metadata parameter when the at least one metadata parameter is missing; determining for the spatial bitstream whether content comprising spatial audio content has been obtained in a current frame or is missing; selecting one of the first machine-learning model or second machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content has been obtained in the current frame or is missing; and enhancing the at least one metadata parameter for the current frame which has been obtained or determining the at least one metadata parameter which was missing based on the selected one of the first machine-learning model or second machine-learning model.. According to a seventh aspect there is provided an apparatus for selectively enhancing at least one spatial metadata parameter, the apparatus comprising: obtaining circuitry configured to obtain a spatial bitstream, the spatial bitstream defining spatial audio content and comprising: at least one metadata parameter associated with at least one audio signal, the at least one metadata parameter being an encoded version of an original metadata parameter; obtaining circuitry configured to obtain a first machine-learning model for enhancing the at least one metadata parameter; obtaining circuitry configured to obtain a second machine-learning model for determining the at least one metadata parameter when the at least one metadata parameter is missing; determining circuitry configured to determine for the spatial bitstream whether content comprising spatial audio content has been obtained in a current frame or is missing; selecting circuitry configured to select one of the first machine-learning model or second machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content has been obtained in the current frame or is missing; and enhancing circuitry configured to enhance the at least one metadata parameter for the current frame which has been obtained or determine the at least one metadata parameter which was missing based on the selected one of the first machine-learning model or second machine-learning model. According to an eighth aspect there is provided a computer readable medium comprising program instructions for selectively enhancing at least one spatial metadata parameter, to perform at least the following: obtaining a spatial bitstream, the spatial bitstream defining spatial audio content and comprising: at least one metadata parameter associated with at least one audio signal, the at least one metadata parameter being an encoded version of an original metadata parameter; obtaining a first machine-learning model for enhancing the at least one metadata parameter; obtaining a second machine-learning model for determining the at least one metadata parameter when the at least one metadata parameter is missing; determining for the spatial bitstream whether content comprising spatial audio content has been obtained in a current frame or is missing; selecting one of the first machine-learning model or second machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content has been obtained in the current frame or is missing; and enhancing the at least one metadata parameter for the current frame which has been obtained or determining the at least one metadata parameter which was missing based on the selected one of the first machine-learning model or second machine-learning model. An apparatus comprising means for performing the actions of the method as described above. An apparatus configured to perform the actions of the method as described above. A computer program comprising program instructions for causing a computer to perform the method as described above. A computer program product stored on a medium may cause an apparatus to perform the method as described herein. An electronic device may comprise apparatus as described herein. A chipset may comprise apparatus as described herein. Embodiments of the present application aim to address problems associated with the state of the art. Summary of the Figures For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings in which: Fig.1 shows an example system of apparatus suitable for implementing some embodiments; Fig.2 shows a flow diagram of the operation of the apparatus shown in Fig. 1 according to some embodiments; Fig.3 shows schematically an example encoder as shown in Fig.1 in further detail according to some embodiments; Fig.4 shows a flow diagram of the operations of the example encoder shown in Fig.3 according to some embodiments; Fig.5 shows schematically an example decoder as shown in Fig.1 in further detail according to some embodiments; Fig.6 shows a flow diagram of the operations of the example decoder shown in Fig.5 according to some embodiments; Fig.7 shows schematically an example metadata enhancer as shown in Fig.5 in further detail according to some embodiments; Fig.8 shows a flow diagram of the operations of the example metadata enhancer shown in Fig.7 according to some embodiments; Fig.9 illustrates conceptually the temporal dependencies for a Dilated Residual Block block of the ReslINet which is a component of the WNet within an example embodiment; Figs. 10 and 11 show example temporal relationships of ML model input and output for an example enhancer ML model and Packet loss concealment ML model according to some embodiments; Fig. 12 shows schematically an example history metadata generator as shown in Fig.7 in further detail according to some embodiments; Fig. 13 shows a flow diagram of the operations of the example history metadata generator shown in Fig. 12 according to some embodiments; Fig. 14 shows an example trace showing the effect of the applications of some embodiments; Figs. 15 and 16 shows an example ML model structure suitable for implementing some embodiments; Fig. 17 shows an example structure of the ReslINet; and Fig. 18 shows an example Dilated ResBlock (residual block) structure according to some embodiments; and Fig. 19 shows schematically an example device suitable for implementing the apparatus shown herein; Embodiments of the Application The concept as discussed herein in further detail with respect to the following embodiments is related to encoding and decoding parametric spatial audio. In the following examples an IVAS codec is used to show practical implementations or examples of the concept. However, it would be appreciated that the embodiments presented herein may be extended to other codecs without inventive input. As discussed earlier the metadata-assisted spatial audio (MASA) is one of the input formats supported by IVAS. It uses audio signal(s) together with corresponding spatial metadata (containing, e.g., directions and direct-to-total energy ratios in frequency bands). The MASA stream can, e.g., be obtained by capturing spatial audio with microphones of, e.g., a mobile device, where the set of spatial metadata is estimated based on the microphone signals. The MASA stream can be obtained also from other sources, such as specific spatial audio microphones (such as Ambisonics), studio mixes (e.g., 5.1 multichannel mix), or other content by means of a suitable format conversion. MASA spatial metadata values are available for each time-frequency tile (TF-tile) (there can, for example, be 24 frequency bands and 4 temporal sub-frames in each frame). The frame size in IVAS is 20 ms (and thus the temporal sub-frame is 5 ms). In addition, MASA supports 1 or 2 directions for each time-frequency tile (i.e., there are 1 or 2 direction index values, and associated 1 or 2 direct-to-total energy ratios, and spread coherence parameters for each time-frequency tile. Other parameters and parameters can also be defined). Additionally, IVAS supports also audio objects (Independent streams with metadata, ISM) as an input. The audio objects contain for each object an audio signal and associated metadata (e.g., the direction of the object). IVAS supports not only metadata such as direction(s), energy ratio(s), etc. but can also comprise extended metadata. The extended metadata parameters, for example, can comprise parameters such as a yaw, a pitch, and a radius. In IVAS, the coding of the extended metadata is supported for higher bitrates, for example, for ISM format input encoding the extended metadata is supported for bratelVAS >64 kbps. Furthermore, IVAS supports a variety of bitrates, ranging from 13.2 kbps to 512 kbps. This bitrate is shared by various signaling bits, the audio signal coding, and the spatial metadata coding. With the MASA input, the bitrate allocated for the spatial metadata coding varies from about 2.25 kbps to about 65 kbps, depending on the total IVAS bitrate and the actual spatial metadata content. Spatial metadata encoding and decoding operates on defined or engineered rules. For example, in IVAS, the MASA metadata is compressed using various data reduction methods, such as representing the direction parameter with sparser resolution when the direct-to-total energy ratio is small; or applying non-uniform encoding codebooks to the parameters that emphasize precision of the parameters at some data ranges over others. As described in the following embodiments machine learning methods can be applied to assist the design of efficient compression. The machine learning methods as described herein do not directly address the compression but attempt to enhance the data degraded by compression. The concepts as discussed in the following embodiments focuses on controlling the utilization of the machine learning method based enhancement of data which has suffered packet loss. The (machine learning) model or (deep neural) network before being able to generate useful outputs is typically required to be trained. In each training step, a set of input data elements (from the training dataset) are presented to the network and the network performs the defined computational operations resulting into the model output. The model output is compared against a known reference (i.e., ground truth or target) with a loss function. Then the network coefficients are adjusted such that the model output is closer to the reference target. If the comparison uses an error function or loss, this adaptation aims at reducing the value of the error or loss. If the comparison uses a similarity or fitness function, the adaptation aims at increasing this. The network training is finished when the network is converged, i.e., the error no longer is reduced or the fitness is no longer improved, or the magnitude of the change is considered too small, or a given maximum number of training steps is reached. This decision may also be done using a separate validation dataset that is not used for adjusting the model parameters, but only to assess the performance. Common framework tools for defining a neural network model include PyTorch and TensorFlow. Using the trained network is referred to as inference. During this stage, the model parameters are usually no longer adjusted. The inference can be done using the same framework which was used for defining and training the model, or the model can be saved into another known network definition format, such as the Open Neural Network Exchange (ONNX) format or the TensorFlow Lite format. These formats may include a software library for the inference, which is capable of performing computational operations according to the format definitions. 3GPP TS 26.253 IVAS describes a set of techniques to convey the spatial metadata at a multitude of bitrates, and encoding and decoding signals with machine learning is known. For example, a machine-learning based audio encoder and decoder was described in Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., &Tagliasacchi, M. (2022). SoundStream: An end-to-end neural audio codec. IEEE / ACM Transactions on Audio, Speech, and Language Processing, 30, pp. 495-507 where it was possible to obtain high audio quality at very low bitrates. The state-of-the art parametric spatial metadata coding methods (such as the MASA coding methods in IVAS) control the level of compression depending on the bitrate at the cost of the reconstruction accuracy of the spatial metadata at the decoder. With lower bitrates, a higher compression factor is needed, and this is obtained by trading off some of the reconstruction accuracy. Especially at very low bitrates (such as 13.2 kbps with IVAS), the bitrate that can be allocated to metadata encoding is so low (such as 2.25 kbps with IVAS) that it affects the perceptual quality of the spatial audio output at the decoder in a negative manner. This includes instability and inaccuracies of the perceived sound sources. There has been proposed, for example, within UK patent application GB2411721, methods to improve the spatial metadata quality using a machine-learning based post-processing block. The post-processing block employs input features from a temporal context preceding (and in some embodiments also following) the current (sub)frame to be enhanced in addition to feature values in the current (sub)frame. Although it is possible to make the enhancement work without temporal context, the quality is significantly worse. It was found that using a history context of approximately 200 ms provides a good enhancement performance because there are typically some correlations in the metadata values over time. It is understood that the value of 200 ms is an example value that was found to work for a specific range of scenarios, and that other values may be employed based on the scenario or implementation. In real-world communication codecs, the transport capability of the physical layer of the network may vary over time, which may lead to packet losses. This means that some packets carrying one or more IVAS frames are not received. For these frames, some sensible spatial metadata has to be generated, in order to have continuous spatial audio playback. A simple solution to the problem of missing or lost data is to repeat the metadata of the previous received IVAS frame (or that of the last subframe of the last received frame), as is done, e.g., in the IVAS decoder (see 3GPP TS 26.253). However, typically, this does not produce an optimal solution, as the spatial metadata typically varies over time. Thus, spatial errors can be caused in the rendering because of the lost packets, which may disturb the spatial audio perception (for example, an unnatural sounding scene can be output to the listener where a sound is perceived from a direction other than the viewed object which causes a break in the user’s immersion of the scene). Moreover, the lost packets can also disturb the spatial metadata enhancement during the frames when the packets are received. For example, the method of GB2411721 uses input features from a temporal context preceding the current frame to be enhanced in addition to the feature values in the current frame. Furthermore, conventional ML methods produce degraded output results when the current frame is missing. As discussed above it has been found that using a history context of approximately 200 ms provides a good enhancement performance wellexceeding the performance of not using the context. However, lost packets disturb the history context that the ML spatial metadata enhancement uses, which may lead to inferior performance, or even noticeable spatial artefacts (e.g., sudden changes in direction and / or spaciousness). Hence, there is a known problem with current parametric audio decoders in that in lost packet situations they can provide poor quality spatial metadata. Firstly, they provide only very rough approximation of possible content of the spatial metadata for the lost packets. Secondly, ML-based spatial metadata enhancers utilizing history context are disturbed by the lost packets. As a result, the decoded spatial metadata typically has errors, which can lead to poor perceived spatial audio quality (e.g., sudden changes in direction and / or spaciousness, and unnatural spatiality). The concept as discussed in the following examples is apparatus or methods which attempts to implement a ML-based spatial metadata post-enhancement of a coded parametric spatial audio stream (containing transport audio signal(s) and associated spatial metadata) in a system comprising encoding and decoding devices. In other words, apparatus and methods for processing the decoded spatial metadata to make it more similar to the original spatial metadata than the decoded spatial metadata using a ML model, for example, such as disclosed in GB2411721. In these embodiments, as described in further detail later, the apparatus or methods for the spatial metadata enhancement are designed for packet loss operations (i.e., spatial metadata is not received for one or more frames), and aim to generate high-quality prediction of spatial metadata for the lost packets and provide enhancement to the decoded spatial metadata for the received packets after the lost packets, by generating suitable history context for the lost packets based on the predicted spatial metadata and by utilizing dedicated packet loss concealment ML model for the frames having the packet losses. In some embodiments these apparatus or methods can achieve such advantages by: determining if one or more of the / V previous frames have packet losses; predicting history context features for those frames which have packet losses by generating emulated spatial metadata associated with the lost packet or missing decoded spatial metadata and employing this emulated spatial metadata to compute history context features for a machine learning model; obtaining history context features for others of the / V previous frames from the decoded spatial metadata; determining or obtaining packet loss information, in other words, information identifying if the current frame or packet has been lost, and if not, decoding a bitstream to obtain decoded spatial metadata and decoded transport audio signal(s); selecting, based on the (current frame) packet loss information, one of at least two machine learning models: a first, or packet loss concealment, machine learning model configured to perform concealment for lost packets based on the history context features; and a second, or default or enhancement, machine learning model configured to perform enhancement for received packets based on decoded spatial metadata (and optionally decoded transport audio signal(s)) obtained by decoding a bitstream and the history context features; generating improved or enhanced spatial metadata for the current frame based on the output of the selected one of the machine learning models; and rendering spatial audio (e.g., binaural audio signals) using the decoded audio signal(s) and the enhanced spatial metadata. In some embodiments, instead of generating emulated spatial metadata for missing decoded spatial metadata for the history context and employing this emulated spatial metadata to compute input features (for example history context features or current frame features) for the selected ML model, the apparatus and methods are configured to directly generate emulated input features associated with the missing decoded spatial metadata which is used as the input for the selected ML model. In some further embodiments, in addition to generating the emulated spatial metadata for the history context or metadata-based model input features, features computed from the transport audio signals are generated for the history context. In other words, missing frames can also employ emulated or generated transport audio signals as inputs to the generation of the features. In some embodiments the current frame feature of the selected ML model when the current frame of the spatial bitstream is not received or is missing data is determined by determining at least one of: an estimated audio signal based on at least one previous frame audio signal; and at least one estimated metadata parameter associated with the at least one audio signal based on at least one previous frame metadata parameter and then determining the at least one input feature based on the at least one of: the estimated audio signal. Error! Reference source not found, presents an example system suitable for implementing some embodiments as described in further detail herein. In this example the transport audio signals (or more generally input audio signals) 100 are passed to an encoder 101, which is furthermore configured to receive or otherwise obtain spatial metadata 104. The encoder can then encode the spatial metadata 104 and transport audio signals to be incorporated into a bitstream 110 in some form. The bitstream 110 can then be obtained and processed by a decoder and renderer 111 to generate output audio signals 114. The output audio signals or spatial audio output can be any suitable output format, for example, binaural audio signals. In embodiments where there is no spatial metadata 104 input then the transport audio signals (or input audio signals) 100 can be analysed to determine the spatial metadata. In some other embodiments the spatial metadata 104 can be input with the transport audio signals 100 as a combined input audio signals input passed to the encoder 101. In other words, the input to the encoder 101 can be a spatial audio stream comprising transport audio signals 100 and spatial metadata 104. The spatial audio stream can, for example, be in a metadata-assisted spatial audio (MASA) format or any suitable input format. With respect to Fig.2 is shown a flow diagram showing example operations of the system shown in Fig. 1. For example, as shown in Fig.2 by 201 is the operation of obtaining transport audio signals and spatial metadata (which may be known as the spatial audio streams). Then as shown in Fig.2 by 203 is the operation of encoding the spatial audio streams (the transport audio signals and spatial metadata) into a bitstream. Then is the operation of transmitting / receiving (or storing / retrieving the bitstream comprising the encoded spatial audio streams as shown in Fig.2 by 205. The received / retrieved bitstream can then be parsed or demultiplexed into encoded transport audio signals and encoded spatial metadata as shown in Fig.2 by 207. The encoded transport audio signals and encoded spatial metadata can then be decoded to generate decoded transport audio signals and decoded spatial metadata as shown in Fig.2 by 209. Then based on the decoded transport audio signals and decoded spatial metadata there is as shown in Fig.2 by 211 a rendering of spatial or output audio signals. Finally, the rendered spatial or output audio signals are output as shown in Fig.2 by 213. Fig.3 furthermore shows a schematic example of the encoder 101 as shown in Fig.1 in further detail. The input to the encoder 101, is as shown in Fig.3 and Fig.1 are transport audio signals 100 and spatial metadata 104. As shown in Fig.3, the spatial metadata 104 is forwarded to a metadata encoder 305, which encodes the spatial metadata. The encoding can use any suitable method, such as the MASA encoding methods of the IVAS encoder, to generate encoded spatial metadata 306 to be passed to the multiplexer 307. Furthermore, as shown in Fig.3, the transport audio signals are forwarded to an audio encoder 303, which encodes the transport audio signals 100 using any suitable audio-signal encoder, such as the IVAS core coder, EVS, or AAC, to generate encoded transport audio signals 304 to be passed to the multiplexer (Mux) 307. The multiplexer 307 (or Mux) is configured to receive the encoded transport audio signals 304 and the encoded spatial metadata 306, and multiplexes them to a bitstream 110, which is the output of the encoder 101. With respect to Fig.4 is shown a flow diagram showing example operations of the encoder shown in Fig. 3. For example, as shown in Fig.4 by 401 is the operation of obtaining transport audio signals and spatial metadata. Additionally, as shown in Fig.4 by 403 is the operation of generating encoded spatial metadata based on the spatial metadata. Then as shown in Fig.4 by 405 is the operation of generating encoded transport audio signals based on the transport audio signals. Then is the operation, as shown in Fig.4 by 407, of multiplexing the encoded transport audio signals and the encoded spatial metadata to generate the bitstream. Finally, the bitstream is output as shown in Fig.4 by 409. Fig.5 furthermore shows a schematic example of the decoder as shown in Fig.1 in further detail. The input to the decoder 111 is, as shown in Fig.5 and Fig. 1, the bitstream 110. The bitstream 110 is forwarded to a demultiplexer 501 (or Demux), which demultiplexes the received or obtained bitstream and generates encoded transport audio signals 502, encoded spatial metadata 506, and a packet lost flag (or lost packed information) 504. The demultiplexer 501 can furthermore be configured to determine whether a packet (or data in general) has been lost. For example, in some embodiments the demultiplexer 501 is configured to determine when the encoded transport audio signals 502 (or in some embodiments the decoded transport audio signals 510) does not contain any data and / or similarly when the encoded spatial metadata 506 (or in some embodiments the decoded spatial metadata 508) does not contain any data for the corresponding frame(s). Having determined the lost packet (or more generally the lack of data from the encoded and / or decoded versions of the transport audio signals and / or spatial metadata) then the demultiplexer is configured to generate and signal this information in a suitable manner, for example, by generating and signaling a packet lost flag 504. The packet lost flag 504 or more generally the lost packet or data information can be forwarded to an audio generator 513 and the metadata enhancer 507. The encoded transport audio signals 502 are forwarded to an audio decoder 503, which decodes the encoded audio signals, using a decoder that is compatible with the encoder used for encoding the audio signals. The output of the audio decoder 503 are decoded transport audio signals 510. The encoded spatial metadata 506 are forwarded to a metadata decoder 505, which decodes the encoded spatial metadata 506 using a decoder that is compatible with the encoder used for encoding the spatial metadata. The output of the metadata decoder 505 is the decoded spatial metadata 508. In some embodiments an audio generator 513 is configured to generate suitable audio signals for any lost packets, based on the preceding decoded transport audio signals 510. In other words the audio generator 513 is configured to receive the decoded transport audio signals 510 and the packet lost flag (or lost packet or lost data information) 504, and when indicated by the packet lost flag 504 generate a generated transport audio signals 520 based on the decoded transport audio signals 510 from previous frames (and in some embodiments future frames or additional packet loss concealment side information). For example, in some embodiments the generated transport audio signals 520 comprise similar spectra as the decoded transport audio signals 510 of the previous received frames. In some embodiments the generated transport audio signals can be generated using methods similar to those found in the IVAS algorithmic description (3GPP TS 26.253). The generated transport audio signals 520 can then be forwarded to the spatial audio synthesizer 509 to be used instead of the decoded transport audio signals 510 for the frames where the packets were lost. In addition, the generated transport audio signals can be forwarded to the metadata enhancer 507. The decoded spatial metadata 508, the decoded transport audio signals 510, the packet lost flag (or lost packet or lost data information) 504, and generated transport audio signals 520 can be passed to the metadata enhancer 507. The metadata enhancer 507 can comprise a post-processing metadata enhancer which produces enhanced decoded spatial metadata 512 that attempts to be more similar to the original spatial metadata 104 than the decoded spatial metadata 508. In addition, the metadata enhancer 507 receives the packet lost flag 504, which identifies whether the packet has been lost or not (in other words whether there is content in the decoded transport audio signals or the generated transport audio signals and furthermore whether the decoded spatial metadata comprise any content or not). The metadata enhancer 507 is configured to process the decoded spatial metadata 508 using a machine learning model with optionally (as shown by the dashed lines) the aid of the decoded transport audio signals 510 and / or generated transport audio signals 520 (based on the packet lost flag 504). The details of the metadata enhancer 507 according to some embodiments is described further below with respect to Fig.7. The output of the metadata enhancer, the enhanced decoded spatial metadata 512 in some embodiments is employed as an input to the spatial synthesizer 509. The decoded transport audio signals 510, generated transport audio signals 520 and the enhanced decoded spatial metadata 512 can be forwarded to the spatial synthesizer 509, which is configured to render a spatial audio output 114 or output audio signals (for example, binaural audio signals) based at least in part on the decoded transport audio signals 510, generated transport audio signals 520 and the enhanced decoded spatial metadata 512. The spatial synthesizer 509 can be any suitable spatial synthesizer implementation. The spatial synthesizer 509 in some embodiments further is configured to receive the packet lost flag (lost packet information) 504 in order to determine whether there is has been a lost packet or in some embodiments is configured to determine whether there is a lost packet from any missing decoded transport audio signals (or the presence of generated transport audio signals). With respect to Fig.6 is shown a flow diagram showing example operations of the decoder shown in Fig.5. For example, as shown in Fig.6 by 601 is the operation of obtaining the bitstream. Additionally, as shown in Fig.6 by 603 is the operation of demultiplexing the bitstream to generate encoded spatial metadata and encoded transport audio signals, and obtaining (from the bitstream or otherwise the packet lost flag or lost packet or data information). Then as shown in Fig.6 by 605 is the operation of generating a decoded transport audio signal based on the encoded transport audio signal. Also as shown in Fig.6 by 607 is the operation of generating a decoded spatial metadata based on the encoded spatial metadata. Furthermore, as shown in Fig.6 by 608 is the operation of determining (or generating) a generated transport audio signals based on the decoded transport audio signals and packet lost flag (lost packet information), where the generated transport audio signals are generated where there is no data within the current decoded transport audio signals (or encoded transport audio signals). Then is the operation, as shown in Fig.6 by 609, of generating enhanced decoded spatial metadata based on the decoded spatial metadata and packet lost flag (lost packet information) and optionally on the decoded and / or generated transport audio signals. Furthermore, as shown in Fig.6 by 611 is the operation of spatially synthesizing (rendering) an output audio signal or spatial audio signal based on the enhanced decoded spatial metadata packet lost flag (lost packet information), generated transport audio signals and / or the decoded transport audio signals. Finally, the output audio signals, the spatial audio signals are output as shown in Fig.6 by 613. Fig.7 shows schematically an example metadata enhancer 507 according to some embodiments. In the following examples the metadata enhancer 507 is configured to enhance the decoded version of the encoded metadata. However, in some embodiments the enhancement is applied to the encoded metadata. In such embodiments the enhanced encoded spatial metadata can then be decoded. Furthermore, in some embodiments the metadata enhancer 507 can be configured to enhance the spatial metadata (encoded or decoded) based on either decoded transport audio signals 510, generated transport audio signals 520 or encoded transport audio signals 502. The input to the metadata enhancer 507 is the decoded spatial metadata 508 and decoded transport audio signals 510, generated transport audio signals 520, and packet lost flag (lost packet information) 504. In some embodiments the metadata enhancer 507 can be configured to obtain or otherwise receive the encoded spatial metadata 506 and / or encoded transport audio signals 502. Both the decoded spatial metadata 508 and decoded transport audio signals 510 have been encoded and decoded, so the data that the metadata enhancer 507 receives has been compressed in some way in information theoretic sense. Similarly, the encoded spatial metadata 506 and encoded transport audio signals 502 are compressed in an information theoretic manner and thus typically deviate from the original spatial metadata 104 and transport audio signals 100. Moreover, the decoded spatial metadata 508 may contain the information at a low time-frequency (TF) resolution due to, for example, being coded by a low-bitrate IVAS codec. In some embodiments the metadata enhancer 507 is based on, and configured to operate in a manner similar to, the metadata enhancer described in GB2411721. The embodiments described herein differ from the examples in GB2411721 in that the metadata enhancer 507 comprises an ability to select or switch between trained models. For some of the frames, the packets may have been lost, and for these frames the decoded spatial metadata 508 and the decoded transport audio signals 510 do not contain any data. This can, as discussed above, be signaled using the packet lost flag (lost packet information) 504. Moreover, for these frames, the generated transport audio signals 520 can be obtained, which contain the generated audio signals determined to replace the lost audio signals. In some embodiments the task of the metadata enhancer 507 is to enhance the decoded spatial metadata 508 for the frames in which the packets have not been lost and to generate suitable spatial metadata for the frames in which the packets have been lost. The result or output of the metadata enhancer is the enhanced decoded spatial metadata 512 that is output from the block. In some embodiments the metadata enhancer 507 is configured to operate in four different modes: Mode 1: when no packets have been lost in the current subframe and during the preceding N subframes (“Normal mode”); Mode 2: when a packet is lost in the current subframe, but no packets have been lost in the preceding N subframes (“Current subframe packet lost mode”); Mode 3: when a packet is not lost in the current subframe, but packets have been lost in one or more of the preceding N subframes (“Previous subframes packets lost mode”); Mode 4: when a packet is lost in the current subframe and in one or more of the preceding N subframes (“Current and previous subframes packets lost mode”); The value of N can be the length of the temporal context of the ML models. For example, it can be 40 subframes (i.e., 200 milliseconds) but could be any suitable value. In some embodiments the value N may be different in different modes. For example, in some embodiments the state or value of the current frame loss flag is configured to determine which ML is employed - enhancer ML model 705 for modes 1 &3 or PLC ML model 715 for modes 2 &4 and can further determine a value of N (and can be different for the enhancer ML model 705 for the PLC ML model 715. These modes can be implemented, for example, in some embodiments by the following functions. A feature computer 701, which is configured to receive or otherwise obtain the decoded spatial metadata 508, decoded transport audio signals 510, generated transport audio signals 520 and generated history metadata 706. The feature computer 701 is configured to generate at least one feature 702 which is passed to at least one machine learning model for the generation of predicted metadata properties based on the mode. The feature computer 701 or feature determiner is configured to determine at least one feature 702 which describes relevant properties of the sound scene described by the decoded spatial metadata 508 and decoded transport audio signals 510 or generated transport audio signals 520. The feature computer 701 can be implemented in some embodiments using expert-designed features, or it may be implemented as a machine learning algorithm trained alongside with the enhancer ML model 705 or PLC ML model 715. The at least one feature 702 can comprise features determined from the decoded spatial metadata 508 and also include features determined from the decoded transport audio signals 510 or generated transport audio signals 520. In some embodiments the at least one feature 702 may be represented as a sequence of features, where each sequence element corresponds to one temporal frame, sub-frame, slot, or some other temporal unit in the metadata and audio stream. In this example embodiment the selection of which machine learning model to employ (for example, one of the enhancer ML model 705 or the PLC ML model 715) to be active or in the data path is based on the packet lost flag (lost packet information) 504. In the example shown in Fig.7 the selection is performed by two pairs of switches controlled by the packet loss flag (lost packet information) 504 value, such that when the packet loss flag has a first or normal mode value, then the one of the enhancer ML model 705 is selected and is configured to receive the at least one feature 702 from the feature computer 701 and output predicted metadata properties 714 to the metadata determiner 707, and when the packet loss flag (lost packet information) 504 has a second or packet loss mode value, then the PLC ML model 715 is selected and is configured to receive the at least one feature 702 from the feature computer 701 and output predicted metadata properties 724 to the metadata determiner 707. However, in some embodiments any suitable selection or control method can be implemented, for example, both ML models receive the at least one feature 702, but only one outputs predicted metadata properties to the metadata determiner 707. The selected ML model can then output the at least one predicted metadata property 714 / 724 to the metadata determiner 707. The metadata determiner 707 is configured to employ the at least one predicted metadata property 714 / 724 for enhancing the decoded spatial metadata 508 (or generate spatial metadata) or when there is a current missing packet to determine an enhanced spatial metadata from the at least one predicted metadata property 714 / 724. In other words, the metadata determiner 707, when the current packet is present, can be configured to process the decoded spatial metadata 508 parameters based on the at least one predicted metadata property 714 / 724 to attempt to reduce any error between the output spatial metadata parameters (which are the enhanced decoded spatial metadata 512) and the original spatial metadata parameters 104 input to the encoder 101 as shown in Fig.1. Furthermore, the metadata determiner 707 is configured to generate spatial metadata 508 to replace the missing or lost packet spatial metadata when the current packet is lost or missing. In some embodiments this can result in an increase in the effective TF-resolution of the decoded spatial metadata, as well as an increase of the dequantization accuracy of the parameter values. The output of the metadata determiner 707 can be the enhanced decoded spatial metadata 512, which can be output from the metadata enhancer 507. The following then describes the operations of this example embodiment apparatus with respect to the example modes described above. When operating in mode 1, “Normal mode” or first or default mode, the processing employs a metadata enhancer method similar to that described in GB2411721, and employing or selecting the Enhancer ML model 705. Thus, in the normal mode, the feature computer 701 is configured to receive the decoded spatial metadata 508 and decoded transport audio signals 510 (as it has been determined that the decoded spatial metadata 508 and decoded transport audio signals 510 comprise relevant information). The feature computer 710 can then, as described above, determine at least one feature describing relevant properties of the sound scene as described by the decoded spatial metadata 508 and decoded transport audio signals 510. The at least one feature can comprise features determined from the decoded spatial metadata 508 and it may also include features determined from the decoded transport audio signals 510. The at least one feature can be represented as a sequence of features, where each sequence element corresponds to one temporal frame, sub-frame, slot, or some other temporal unit in the metadata and I or audio stream. In an example embodiment, for the decoded spatial metadata 508 the at least one feature can be obtained by transforming the directional MASA metadata into Cartesian XYZ-representation (3D-coordinates in X, Y, Z -axes). The relevant fields of the low-resolution decoded spatial metadata 508 can be denoted with azi(K,n), ele(K,n), andrQc.n) corresponding to the azimuth angle, elevation angle, and direct-to-total energy ratio in TF-tile K,n, where k is the parameter band index k = 1, ...,n_bands_meta, and n is the (sub-)frame index. This spherical representation is transformed into Cartesian vector representation with v( / c,n) = vx(k,h) vy(K, n) = r(K,n) cos(azi(K, n)) cos (eZe( / c, n)) sin(azi(K, n)) cos (eZe( / c, n)) sin (eZe( / c,n)) The metadata feature is the decoded spatial metadata 508 represented with such vectors for each TF-tile. It can be considered as a 3-dimensional tensor with the shape (n_bands_metaln_frames,n_features_meta), where n_features_meta = 3 is the number of input features per tile, i.e., corresponding to the 3 elements of the vector v( / c,n), nframes is the number of spatial metadata (sub-)frames that are processed at once, and n_bands_meta = 5 corresponding to the 5 parameter bands of the low-resolution metadata (in case of 13.2 kbps, potentially other number of bands at other bitrates). Optionally, a number of features can be determined from decoded transport audio signals 510 and / or generated transport audio signals 520. For example, these features could be generated in a manner similar to those generated in GB2411721. For example, these can be different audio feature sets “Cov” and “SPAC”, and Cov5 from the audio feature set “Cov” can be employed in the loss function. These include features that describe the inter-channel (of two channels of the transport audio signal) cross-correlation properties in frequency bands, features describing the inter-channel level differences in frequency bands, local signal energy evolution over time in bands, and band-energy ratios. The exact set of features can be implementation specific and can be other features in some embodiments. The features 702 are provided to the enhancer ML model 705, which is described in further detail later. The enhancer ML model 705 is a machine learning processor or method, employing the at least one feature 701 as an input and configured to determine at least one predicted metadata property 714. An example of the at least one predicted metadata property 714 is the MASA directional information (azimuth, elevation, and direct-to-total energy ratio) in a suitable representation, e.g., in Cartesian vector format. As indicated above, the enhancer ML model 705 can employ history context to perform the spatial metadata enhancement. In this example embodiment, the history context is approximately 200 ms long, though other values may be used in other embodiments. In the normal mode the original decoded spatial metadata 508 (with a low TF-resolution) and the at least one predicted metadata property 714 are input to the metadata determiner 707. This metadata determiner 707 in some embodiments is configured to use the at least one predicted metadata property 714 for enhancing the decoded spatial metadata 508. For example, in some embodiments the metadata determiner 707 is configured to modify the decoded spatial metadata 508 to bring it closer to the original spatial metadata obtained as the input to the encoder in Fig.1. For example, in some embodiments, metadata determiner 707 is configured to increase the effective TF-resolution of the decoded spatial metadata 508, as well as increasing the de-quantization accuracy of the parameter values. The output of the metadata determiner 707 in the mode 1 is the enhanced decoded spatial metadata 512. The mode 2, or current subframe packet lost mode, in some embodiments is activated when the packet lost flag 504 indicates that the packet has been lost for this subframe. In this case, there is no spatial metadata (or decoded transport audio signals) available for the current subframe. However, in this mode, no packets have been lost for the preceding N subframes. The value of N as indicated above can be any suitable value, for example the length of the temporal context of the ML models, such as 40 subframes (i.e., 200 milliseconds). In this mode, no new features are computed (or determined or output). Instead, only previous N subframe features computed from the previous N subframe decoded spatial metadata and decoded transport audio signals are used. These features 702 are passed to the packet loss concealment (PLC) ML model 715. The PLC ML model 715 is a machine learning method that determines predicted metadata properties 724 for the current subframe using the features 702 of the N previous subframes. The details of the PLC ML model 715 are presented further below. The predicted metadata properties 724 from the PLC ML model 715 are forwarded to the metadata determiner 707, which determines the enhanced decoded spatial metadata 512 in a manner similar as described with respect to the mode 1 or normal mode whereas presented above. For example, when the current packet is lost or more generally the data is not present for the current packet, there is no decoded spatial metadata 508 that could be enhanced. Instead, the metadata determiner 707 is configured to process or operate predicted metadata properties 724 from the PLC ML model 715. In some embodiments the metadata determiner 707 is further configured to input and generate spatial metadata 508 based on some signal-specific configuration information. The mode 3, or previous subframes packets lost mode, in some embodiments is activated when the packet lost flag indicates that for the current subframe no packet is lost (or in other words indicates that the current packet is present), but for one or more of the previous N subframes packets have been lost or are missing. This can be implemented in some embodiments by employing a mode 3, or previous subframe packet lost, flag which is set on determination of any lost packet (or data) and only expires after a counter reaches the determined N number of subframe packets, where the counter is reset when a further lost packet is determined. In some embodiments the mode 3 is active where the mode 3 or previous subframe packet lost flag is set but the packet lost flag is not active. In the mode 3, or previous subframes packets lost mode, the decoded spatial metadata 508 and decoded transport audio signals 510 for this subframe are forwarded to the feature computer 701 and the features 702 are computed as in the mode 1 or normal mode as described above. However, for some subframes in the N subframe history, packets were lost, and the decoded spatial metadata 508 and the decoded transport audio signals 510 are not available for the subframes corresponding to those packets. Thus, for those subframes, the at least one feature 702 cannot be computed using the decoded spatial metadata 508 and the decoded transport audio signals 510. In mode 3, history spatial metadata 706 is generated for the lost subframes using the history metadata generator 703. For the subframes for which the packets were lost the history metadata generator 703 receives the corresponding enhanced decoded spatial metadata 512 as an input. This contains a spatial metadata in the full time-frequency resolution. Then, the history metadata generator 703 is configured to generate the generated history spatial metadata 706 by emulating the effect of the encoding and decoding applied to the spatial metadata by the encoder and the decoder for making the generated history spatial metadata 706 compatible with the decoded spatial metadata 508. The details of the history metadata generator 703 are presented further below. The history metadata generator 703 generates the generated history spatial metadata 706 for the subframes for which the packet was lost (i.e., the decoded spatial metadata 508 does not have any content). The metadata related history features 702 are computed in the feature computer 701 using the generated history spatial metadata 706 for the subframes where the packet has been lost and using the decoded spatial metadata 508 for the subframes where the packet has not been lost. Similarly, the audio related history features 701 are computed using the generated transport audio signals 520 for the subframes where the packet has been lost and using the decoded transport audio signals 510 for the subframes where the packet has not been lost. The determined features 702 are forwarded to the enhancer ML model 705, which operates as described above for the mode 1 or normal mode. The resulting predicted metadata properties 714 are then passed to the metadata determiner 707, which is configured to determine the enhanced decoded spatial metadata 512 as presented herein. The mode 4, or current and previous subframes packets lost mode, in some embodiments is activated when the packet lost flag 504 indicates that for the current and for one or more of the previous N subframes packets have been lost (for example, as indicated by the mode 3 or previous subframe packet lost flag). This mode can be seen as a combination of the mode 3, current subframe packet lost mode, and the mode 2, previous subframes packets lost mode, and as such could be detected or determined in some embodiments when both the packet lost flag is set and the previous subframes packets lose flag is also set. In the mode 4, or current and previous subframes packets lost mode, no new features 702 are computed using the data from the current subframe. Instead, only the features determined from the history data are employed. The features 702 based on history data are computed using the decoded spatial metadata 508 and the decoded transport audio signals 510 for the subframes where the packets have not been lost and using the generated history spatial metadata 706 and the generated transport audio signals 520 for the subframes where the packets have not been lost, similar as in mode 3 or previous subframes packets lost mode. The determined features 702 are forwarded to the PLC ML model 715, which operates as described above, and the resulting predicted metadata properties 724 are forwarded to the metadata determiner 707, which determines the enhanced decoded spatial metadata 512. Thus, for example, in some embodiments the modes and operation of the modes can be summarized in the following table. Mode packet lost flag previous subframe Features are generated from ML Model employed packet lost flag Mode 1 0 0 decoded spatial metadata (and decoded transport audio signals) Enhancer ML model Mode 2 1 0 previous present subframes: PLC ML model decoded spatial metadata (and decoded transport audio signals) Mode 3 0 1 Current and anv present previous subframe: decoded spatial metadata (and decoded transport audio signals) previous lost sub-frames: generated spatial metadata (and generated transport audio signals) Enhancer ML model Mode 4 1 1 Previous lost subframes: generated spatial metadata (and generated transport audio signals) previous present sub-frames: decoded spatial metadata (and decoded transport audio signals) PLC ML model With respect to Fig.8 is shown a flow diagram showing examp e operations of the metadata enhancer shown in Fig.7. For example, as shown in Fig.8 by 801 is the operation of obtaining the inputs such as decoded spatial metadata, decoded transport audio signals, 5 generated transport audio signals, and packet lost flag (lost packet information). Additionally, as shown in Fig.8 by 802 is the operation of determining a mode of operation, for example, based on the packet lost flag (lost packet or data information) and optionally based on the previous subframe packet lost flag. Furthermore, for a situation where there has been at least one previous lost 10 packet (for example, indicated by the previous subframe packet lost flag) such as found in either of modes 3 or 4, there is the operation of generating or obtaining (for any lost sub-frame packet) history spatial metadata based on enhanced decoded spatial metadata from sub-frames prior to the missing or lost packet, as shown in Fig.8 by 803 Then, as shown in Fig.8 by 805 is the operation of generating or determining at least one feature based on a selection determined by the mode of at least one of: the generated history spatial metadata; decoded spatial metadata; decoded transport audio signals and generated transport audio signals. Also as shown in Fig.8 by 807 is the operation of selecting, based on the packet lost flag (lost packet information), between enhancer ML model and PLC ML model to generate predicted metadata properties from the generated or determined at least one feature. Then follows, as shown in Fig.8 by 809, is the operation of generating the predicted metadata properties based on the generated at least one feature passed to the selected ML model. Then is the operation, as shown in Fig.8 by 811, of generating enhanced decoded spatial metadata based on decoded spatial metadata (where available) and predicted metadata properties. Following this is the outputting of the enhanced decoded spatial metadata as shown in Fig.8 by 813. The enhancer ML model 705 in some embodiments can be a machine learning model similar in operation and design as described in detail in GB2411721. The at least one feature 702 input to the enhancer ML model 705 comprises three groups of features determined from both metadata and transport audio signals: metadata features and two sets of features from audio. However, these are example features employed in an example embodiment and other features can be employed in other embodiments. The features 702 input to the DNN that produces the predicted metadata properties 714, which can then be used to determine enhanced spatial metadata 512. The exact architecture of the DNN employed can be implementation specific and the following examples are examples only. The disclosure of GB2411721 employed a DNN consisting of separate U-Net -type sub-models for different feature inputs followed by one more U-Net -type model for producing the actual predicted metadata properties based on the outputs of the feature sub-models. The ll-Net -type model used in this example is similar to ResU-Net as discussed in Zhang, Z., Liu, Q. &Wang, Y. (2018). Road extraction by deep residual ll-Net. IEEE Geoscience and Remote Sensing Letters, vol. 15, issue 5, pp. 749-753. DOI: 10.1109 / LGRS.2018.2802944. Each macro layer consists of ResNet blocks similar to those disclosed in He, K., Zhang, X. Ren, S. &Sun, J. (2015). Deep residual learning for image recognition. arXiv: 1512.03385. Normal ResNet blocks in the literature are non-causal, when considering applying them of a TF-representation of a signal. This means that they utilize values from future time instants for determining the output for the current time instant. However, it is possible to modify the block to operate in a causal manner, and this is used in the current example embodiment. The temporal dependencies of the operations of a ResNet block is conceptually illustrated in Fig.9. This illustration is simplified to show only the data axis corresponding to time. The example embodiment can further also contain convolutional operations along the axis corresponding to frequency. However, these convolutions are not shown for clarity reasons. The illustrated ResNet block in Fig.9 is configured to operate in causal mode, i.e., it does not access future inputs. In other words, only the current and past (indices n - 6,n - 5, are used for determining the output for time instant n. The output in the current frame depends on the current input n and a number (here, 6) of earlier inputs. As there are multiple blocks of this kind stacked in the ResU-Net, the temporal dependencies accumulate and each output has a dependency to a large number (e.g., 20-40) earlier inputs in addition to the current input. This history context has proven to be useful in the enhancement process. With respect to Fig. 15 is shown an example ML model 705 configuration in further detail. Fig.9 exemplifies the temporal operations of a single Dilated Residual Block (DRB) 1756A-1756D, as shown in Fig. 18. The DRBs are building blocks that can be used in ResU-Net structure shown in Fig. 15, Fig. 16, and that can be used in an example embodiment of ML model 705 as illustrated in Fig.15. Different structures can be used for the machine learning models 705 in different examples. Fig. 15 shows an example structure for a machine learning model 705 that can be used in some examples. Other structures for the machine learning model 705 can be used in other examples. The input features are provided or obtained as an input to a DNN that produces predicted metadata properties, which can then be used to determine enhanced spatial metadata. The exact architecture of the DNN is not critical and any suitable arrangement or implementation can be employed. The disclosure of GB2411721 uses a DNN consisting of separate ll-Net -type sub-models for different feature inputs followed by one more ll-Net -type model for producing the actual predicted metadata properties based on the outputs of the feature submodels. The ll-Net -type model used in this example is similar to ResU-Net as disclosed in Zhang, Z., Liu, Q. &Wang, Y. (2018). Road extraction by deep residual U-Net. IEEE Geoscience and Remote Sensing Letters, vol. 15, issue 5, pp. 749-753. DOI: 10.1109 / LGRS.2018.2802944. In this implementation each macro layer consists of ResNet blocks as disclosed in He, K., Zhang, X. Ren, S. &Sun, J. (2015). Deep residual learning for image recognition. arXiv: 1512.03385. Normal ResNet blocks in the literature are non-causal, when considering applying them of a TF-representation of a signal. This means that they utilize values from future time instants for determining the output for the current time instant. However, it is possible to modify the block to operate in a causal manner, and this is used in the current example embodiment. In the example of Fig. 15 the machine learning model 705 comprises a W-Net structure. The structure is referred to a W-Net because the path from the input to the output goes through two U-Net structures. Fig. 16 shows an example U-Net structure that can be used in the machine learning model 705 or in the W-Net structure 1503 used in the machine learning model 705. The U-Net structure comprises a downsampling part 1510 followed by an upsampling part 1514. The downsampling part 1510 comprises a sequence of downsampling layers. The respective downsampling layers can comprise convolutional operations such as 2D convolutional layers applying convolutional operation over the spatial dimensions of the data. The downsampling layers of the downsampling part 1510 can reduce the dimensions of an input along at least some axes. The output of the downsampling part 1510 has a smaller number of data elements in at least one axis compared to the original input. The output of the downsampling part 1510 is provided as an input to the upsampling part 1514. The upsampling part 1514 is configured to produce output data. The upsampling part 1514 comprises a sequence of upsampling layers. The upsampling part 1514 can comprise X upsampling layers where X can be also the number of downsampling layers in the downsampling part 1510. The upsampling layers can comprise transposed convolutional operations and / or upsampling operations possibly followed or preceded by convolutional operations. The upsampling layers of the upsampling part 1514 can increase the dimensions of an input along at least some axes. The ll-Net structures in this example also comprise skip connections 1512. The skip connections 1512 are configured to relay skip connection signals from respective downsampling layers to corresponding upsampling layers. The skip connection signals can reintroduce features from the downsampling part 1510 back into corresponding layers of the upsampling part 1514. The upsampling layers can comprise operations such as concatenating operations to combine data from a skip connection signal with input data. The upsampling layers can also comprise operations to increase the dimensions of data in at least one axis compared to the input that is input to the upsampling part 1514. The upsampling part 1514 provides output data. The output data of the upsampling part 1514 typically has the same number of data elements in at least one dimension as the input data that is originally provided to the downsampling part 1510. In the example of Figs. 15 and 16 the output of the downsampling part 1510 is provided as an input to the upsampling part 1514. In other examples there can comprise one or more intervening components such as a bottleneck and / or any other suitable operations or combinations of operations. In the example of Fig. 15 the machine learning model 705 comprises three pre-networks 1520A, 1520B, 1520C. Each of the pre-networks 1520A, 1520B, 1520C comprises a ll-Net. The ll-Net can be as shown in Fig. 16 (some of the reference numbers are omitted in Fig. 15 for clarity) or can have any other suitable arrangement. The machine learning model 705 receives the features 702 as an input. In this example the features 702 can comprise the metadata features 1504 and the audio features. The audio features can comprise SPAC features 1502 and Cov features 1500. The respective feature inputs are provided to different pre-networks 1520. The metadata features 1504 are provided as an input to a first pre-network 1520A, the SPAC features 1502 are provided as an input to a second pre-network 1520B, and the Cov features 1500 are provided as an input to a third pre-network 1520C. The respective pre-networks 1520A, 1520B, 1520C are arranged to process the respective inputs to provide intermediate representations 1522A, 1522B, 1522C as an output. The settings and weights of the respective pre-networks 1520A, 1520B, 1520C are specific for each of the input features. The first pre-network 1520A processes the input metadata features 1504 to provide a metadata intermediate representation 1522A as an output, the second pre-network 1520B processes the input SPAC features 1502 to provide a SPAC intermediate representation 1522B as an output, and the third pre-network 1520C processes the input Cov features 1500 to provide a Cov intermediate representation 1522C as an output. The intermediate representations 1522A, 1522B, 1522C are provided to a concatenation block 1524. The concatenation block 1524 is configured to combine the intermediate representations 1522A, 1522B, 1522C. The concatenation block 1524 can concatenate the intermediate representations 1522A, 1522B, 1522C along the feature axis or perform any other suitable combination. The concatenation block 1524 provides a combined intermediate representation 1526 as an output. The combined intermediate representation 1526 is provided as an input to a combined prediction network 1528. The combined prediction network 1528 can comprise another ll-Net structure. The weights and settings of the ll-Net structure of each of the four networks, the combined prediction network 1528 and the pre-networks 1520A, 1520B, 1520C can be different. The combined prediction network 1528 provides pre-scale predicted metadata properties 1530 as an output. The pre-scale predicted metadata properties 1530 are provided as an input to an XYZ scale block 1532. The XYZ scale block 1532 is configured to apply appropriate scaling to the pre-scale predicted metadata properties 1530. The XYZ scale block 1532 provides predicted spatial metadata properties 714 as an output. In the following, the structure and settings of the ML model 705 are described. An example metadata features pre-network 1520A (for metadata from, for example, 13.2 kbps IVAS coding with 5 frequency bands) is shown with respect to Fig. 17 and described as follows. Corresponding structures can be used for the SPAC features pre-network 1520B, the Cov features pre-network 1520C, and the combined prediction network 1528 with the appropriate possible adjustments for the specific feature representation dimensions. The input in this case is the metadata features. This input has shape (n_bands_meta, n_fram.es, n_features_meta). The input is provided to a dimension adjustment 1750 which comprises a linear layer 1752. The linear layer 1752 can be fully connected. The linear layer 1752 can operate on the input dimension that corresponds to the frequency bands. In this case this is the first dimension of the input. In this description a single input is processed. In examples of the disclosure multiple inputs can be processed in parallel as a batch. In such examples the stacking of multiple inputs adds one dimension in front of the actual data dimensions. The operation of the linear layer 1752 provides an intermediate feature tensor Y 1754 as an output. The metadata features pre-network 1520A can further comprise multiple residual blocks (ResBlock) or dilated residual blocks (DRB). The ResBlocks and DRBs 1756 comprise a stack of layers that is arranged so that the output of a given layer is taken and added to a subsequent layer deeper within the ResBlock or DRB. An example DBR is shown in Fig. 18. Each DRB 1756 has stride and dilation settings (str, dil)=(x,y) given within the block, for example, (str, dil)=(2, 2). The first number corresponds to the stride along the frequency axis (first dimension) and the second number corresponds to the convolutional kernel dilation factor parameter along the temporal axis (second dimension). The number of output channels (third dimension) from each DRB is given as the last number in the triplet following the respective DRBs. The DRBs in the downsampling part may have the convolutional kernel size parameters of (3, 3). The metadata features pre-network 1520A can comprise a sequence of downwards DRBs 1756A-1756D and a sequence of upsampling or upwards DRBs 1756E-1756H possibly paired with a nearest neighbor upsampling layer. The sequence of downwards DRBs 1756A-1756D can provide a downsampling part of the metadata features pre-network. The sequence of upwards DRBs 1756E-1756H and upsampling blocks 1762E-1762G can provide an upsampling or upwards part of the metadata features pre-network 1520A. For example, the first DRB 1756A has (str, dil)=( 1, 1) with 8 output channels. The second DRB 1756B has (str, dil)=(2, 2) with 16 output channels. The second DRB 1756B is arranged to perform factor 2 sub-sampling in the spatial (time and frequency) dimensions. The third DRB 1756C has (str, dil)=(2,2) with 32 output channels. The third DRB 1756C is arranged to perform factor 2 sub-sampling in the spatial dimensions. The fourth DRB 1756D has (str, dil)=(2, 2) with 64 output channels. The fourth DRB 1756D is arranged to perform factor 2 sub-sampling in the spatial dimensions. All DRBs (first through fourth) of the downsampling path have the convolutional kernel size parameter of (3, 3). The output of the last downwards DRB 1756D is passed through a convolutional block (ConvBlock) 1758. The ConvBlock 1758 comprises a two-dimensional convolution with kernel size of (3, 1) for (frequency, time) with 128 output channels, followed by a Batch Normalization (BatchNorm) and a Rectified Linear Unit (ReLU) activation (not shown in Fig. for clarity). The output of the ConvBlock 1758 is passed through Transposed Convolution (TransConv) 1760 block. The TransConv 1760 block comprises a two-dimensional transposed convolution with kernel size of (3, 1) and 64 output channels. This effectively up-samples along the frequency axis. This is followed by a BatchNorm and ReLU activation (also not shown in Fig. for clarity). The output of the TransConv 1760 block (including the BatchNorm and activation) is concatenated with the output of the last downward DRB 1756D. The concatenation of the output of the TransConv 1760 with the output of the last downward DRB 1756D is performed along the channel dimension. The resulting tensor is provided as an input to the upsampling part of the metadata features pre-network 1520A. The upsampling part comprises a fifth DRB 1756E with (str, dil)=(1, 1) followed by an Upsampling (2, 1) 1762E block. The Upsampling 1762E block is arranged to upsample the layer input with the spatial size scaler factor given in the parenthesis (frequency, time) using, for example, the nearest neighbor upsampling method. Therefore Upsampling (2, 1) increases the size of the spatial dimension corresponding to frequency by a factor of two using nearest neighbor upsampling. The output of the Upsampling is the output of this layer of the upsampling part of the metadata features pre-network 1520A. The output of the Upsampling block 1762E is concatenated with a corresponding matching skip connection tensor. This comprises data from the matching level downsampling layer 1756C. The concatenation is performed along the feature dimension (third dimension in this example). The output of the concatenation is provided to a sixth DRB 1756F with (str, dil)=(1, 1) and following Upsampling (2, 1) block 1762F. The output of the Upsampling block 1762F is concatenated with a corresponding matching skip connection tensor from the downward DRB 1756B. The concatenation is performed along the feature dimension (third dimension in this example). The output of this concatenation is provided to a seventh DRB 1756G with (str, dil)=(1, 1) and following Upsampling (2, 1) block 1762G. The convolutions in the upsampling part of the metadata features pre-network 1520A have the number of output channels of 32, 16, and 8. The convolutions in the upsampling part of the metadata features pre-network 1520A have the kernel size parameters of (3, 1), corresponding to kernel size along the axes (frequency, time). The output of the Upsampling block 1762G is concatenated with a corresponding matching skip connection tensor from the downward DRB 1756A. The last layer of the upsampling part comprises an eighth DRB 1756H with (str, dil)=(1, 1). The last DRB 1756H has five output channels and no following Upsampling blocks. The shape of the last DRB 1756H provides the output of the metadata features pre-network. The output of the metadata features pre-network is now (24,n_frames, 5). In this example the output of the metadata features pre-network is the intermediate metadata features. Other U-Net structures can be arranged in a similar manner but would have different inputs and outputs. Fig. 18 shows an example structure for a dilated residual block (DRB). The DRB could be used in the ll-Net structures of the machine learning model. The example DRB could be used in a downsampling part. A DRB with the same internal structure could also be used in the upsampling part but different stride and kernel size settings could be used in the upsampling part. The internal structure of the DRB has a pre-activation ordering. This means that BatchNorm layers 1812, 1818, and the ReLU activation layers 1814, 1820 are before the convolution operations 1866, 1882. The DRB comprises two paths for an input 1800. The first path 1830 is shown on the left side of Fig. 18 and the second path 1840 is shown on the right side of Fig. 18. The path on the left side corresponds to the residual or by-pass path 1830. The path on the right side corresponds to the convolutional core path 1840. The residual or by-pass path 1830 comprises convolution operations 1862 and a BatchNorm layer 1804. The convolution operations 1862 comprises a Conv2D layer with the kernel size of (1, 1) and striding str_f along the first dimension corresponding to frequency axis matching the stride settings of the DRB. The convolution operations 1862 adjust the size of the feature dimension of the input to match the feature dimension of the last convolution in the convolution core. The convolution operations 1862 are followed by the BatchNorm layer 1804. The BatchNorm layer 1804 comprises a BatchNorm2D layer that operates on the feature dimension. The convolutional core path 1840 consists of a sequence of blocks. The sequence comprises BatchNorm layers 1812, 1818 followed by ReLU activation layers 1814, 1820 which are then followed by convolution operations 1866, 1882. In this example the first BatchNorm layer 1812 is a BatchNorm2D that operates on the feature dimension (third dimension). The first ReLU activation layer 1814 is applied on each element from the BatchNorm layer 1812. The output of the ReLU activation layer 1814 is passed to the first convolution operations 1866. The first convolutions operations 1866 comprise a Conv2D with the kernel size corresponding to the kernel size of the DRB, for example, (3, 3) and the stride setting str_f matching the frequency axis stride settings of the DRB. The output of the first convolution operations 1866 is passed through a second BatchNorm layer 1818 and a second ReLU activation layer 1820 before being passed to the second convolution operations 1882. The second convolution operations 1882 processes the input with a Conv2D with kernel size corresponding to the kernel size of the DRB, for example, (3, 3), and stride of 1 along the first dimension corresponding to the frequency axis, and dilation corresponding to the dil_t setting of the DRB, for example, 2. In this example, the padding corresponding to the convolutional shrinkage due to both convolutions 1866, 1882 along the first dimension, corresponding to the frequency axis, is applied as reflection padding along the frequency axis before the first convolution 1866. The output of the residual or by-pass path and the output of the convolutional core are provided to an addition block 1850. The addition block 1850 adds the output of the residual or by-pass path and the output of the convolutional core in an element-wise manner. The output of the addition block 1850 is the output 1852 of the DRB. The SPAC feature pre-network can have a similar structure to the metadata features pre-network described above however different dimensions and settings could be used. The input to the SPAC feature pre-network would be the SPAC features. This input has shape (60, n_frames, 33). This input can be provided to a dimension adjustment as described above. The dimension adjustment can also comprise a fully connected linear layer. The dimension adjustment provides an intermediate feature tensor Y with shape (48, n_frames, 33). In the example SPAC feature pre-network 1520B the downsampling part also comprises four DRBs. All DRBs in the downsampling path may have the kernel size parameter (3, 3). The first DRB with (str, dil)=(2, 1) has 64 output channels. The second DRB with (str, dil)=(2, 2) has 128 output channels. The third DRB with (str, dil)=(2, 2) has 256 output channels. The fourth DRB with (str, dil) = (2,2) has 512 output channels. In the example SPAC feature pre-network 1520B the output of the fourth DRB is passed through a convolutional block (ConvBlock). The ConvBlock comprises a two-dimensional convolution with kernel size of (3, 1) for (frequency, time) with 1024 output channels, followed by a TransConv block. The TransConv Block has a corresponding kernel size of (3, 1) and 256 output channels. In the example SPAC feature pre-network 1520B the upsampling part also comprises a further four DRBs and corresponding Upsampling blocks. The DRBs in the upsampling part may have kernel size parameters of (3, 1). In this example the fifth DRB with (str, dil)=(1, 1) has 256 output channels followed by an Upsampling (2, 1) block. The Upsampling block does not provide any Upsampling in the temporal axis. The sixth DRB with (str, dil)=(1, 1) has 128 output channels followed by an Upsampling (2, 1) block. The seventh DRB with (str, dil)=(1, 1) has 64 output channels followed by an Upsampling (2, 1) block. The eighth DRB with (str, dil)=(1, 1) has 20 output. The DRBs in the upsampling part may have the convolutional kernel size parameters of (3, 1). The output of the SPAC feature pre-network 1520B has shape (24, n_fram.es, 20). The Cov feature pre-network 1520C can have a similar structure to the metadata features pre-network 1520A and also the SPAC feature pre-network 1520B, however different dimensions and settings could be used. For the case of the Cov feature pre-network 1520C the dimension adjustment block can be omitted and the input to the U-Net structure would be an input of shape (2A,n_subfram.es, 5). In the example Cov feature pre-network 1520C the downsampling part also comprises four DRBs. The first DRB with (str, dil)=( 1, 1) has 8 output channels. The second DRB with (str, dil)=(2, 2) has 16 output channels. The third DRB with (str, dli)=(2, 2) has 32 output channels. The fourth DRB with (str, dil)=(2, 2) has 64 output channels. The DRBs in the downsampling part may have the convolutional kernel parameters of (3, 3). In the example Cov feature pre-network 1520C the output of the fourth DRB is passed through a ConvBlock. The ConvBlock comprises a two-dimensional convolution with kernel size of (3, 1) with 128 output channels, followed by a TransConv block. The TransConv Block has a corresponding kernel size of (3, 1) and 64 output channels. In the example Cov feature pre-network 1520C the upsampling portion also comprises a further four DRBs and corresponding Upsampling blocks. In this example the fifth DRB with (str, dil)=(1, 1) has 32 output channels followed by an Upsampling (2, 1) block. The sixth DRB with (str, dil)=(1, 1) has 16 output channels followed by an Upsampling (2, 1) block. The seventh DRB with (str, dil)=(1, 1) has 8 output channels followed by an Upsampling (2, 1) block. The eighth DRB with (str, dil)=(1, 1) has 8 output channels. There is no Upsampling block following the eighth DRB in the Cov feature pre-network 1520C. The DRBs in the upsampling part may have the convolutional kernel parameters of (3, 1). The output of the Cov feature pre-network has shape (24, n_frames, 8). The metadata feature pre-network 1520A can provide a metadata intermediate representation 1522A as an output, the SPAC feature pre-network 1520B can provide a SPAC intermediate representation 1522B as an output, and the Cov feature pre-network 1520C can provide a Cov intermediate representation 1522C as an output. Using the example U-Net structures as described above the metadata intermediate representation 1522A has a shape (24,n_frames, 5), the SPAC intermediate representation 1522B has shape (24,n_frames, 20), and the Cov intermediate representation 1522C has shape (24, n_frames, 8). The intermediate representations are provided to a concatenation block 1524. The concatenation block 1524 concatenates the intermediate representations 1522A, 1522B, 1522C along the feature axis. The concatenation block provides a combined intermediate representation 1526 as an output. The combined intermediate representation 1526 has a shape (24,n_frames, 33). The combined intermediate representation 1526 is provided as an input to a fourth U-Net structure 1528. The fourth U-Net structure 1528 is the combined prediction network. The combined prediction network 1528 has a similar structure to the metadata features pre-network 1520A with some differences. For the case of the combined prediction network the dimension adjustment block can be omitted and the input to the U-Net structure would be an input of shape (24, n_frames, 33). In the example combined prediction network the downsampling portion also comprises four DRBs. The first DRB with (str, dil)=(1, 1) has 32 output channels. The second DRB with (str, dil)=(2, 2) has 48 output channels. The third DRB with (str, dil)=(2, 2) has 64 output channels. The fourth DRB with (str, dil)=(2, 2) has 128 output channels. All four DRBs of the downsampling portion may have the convolution kernel size parameter of (3, 3). In the example combined prediction network the output of the fourth DRB is passed through a ConvBlock. The ConvBlock comprises a two-dimensional convolution with kernel size of (3, 1) and 512 output channels, followed by a TransConv block with a corresponding kernel size of (3, 1) and 256 output channels. In the example combined prediction network the upsampling part also comprises a further four DRBs and corresponding Upsampling blocks. In this example the fifth DRB with (str, dil)=(1, 1) has 64 output channels followed by an Upsampling (2, 1) block. The sixth DRB with (str, dil)=(1, 1) has 48 output channels followed by an Upsampling (2, 1) block. The seventh DRB with (str, dil)=(1, 1) has 12 output channels followed by an Upsampling (2, 1) block. The eighth DRB with (str, dil)=(1, 1) has 3 output channels. There is no Upsampling block following the eighth DRB in the combined prediction network 1528. The output 1530 of the combined prediction network 1528 has shape (24,n_frames, 3). The output 1530 of the combined prediction network provides pre-scale predicted metadata properties 1530. These represent the directional metadata in each (24, n_frames) TF-tiles in XYZ vector representation ,n). The pre-scale predicted metadata properties are provided as an input to an XYZ scale block 1532. The XYZ scale block 1532 is configured to apply appropriate scaling to the pre-scale predicted metadata properties 1530. The XYZ scale block 1532 can apply final constraints on the length of the vector. In some examples the scaling applied by the XYZ scale block 1532 can comprise determining the length of the input vectors. The length of the input vectors can be determined with rin(k,n) = ^rescalex(k,n) + escale>y(k,n) + v*rescalez^,n) This is passed through hyperbolic tangent and scaled with a constant, for example, cxyz = 1.1 for obtaining the scaled length ^scaled ^xyztanh (rin(fc, n)) Without the scaling, the output of the hyperbolic tangent would require infinite value of the input for the output to reach value of 1.0. When the model is used for inference the value of rscaied(k,n) is limited to the range of 0... 1.0 with rscaied^^ = max (o,min(l, rscaled(k,n))} and this value is used in place of rscaied(k,ri). A scaling value s(k,n) can be determined from these two lengths with s(k, n) ^scaled ^.k, rin(k, n) The output 714 of the XYZ scale block 1532 is the input 1530 multiplied by this scaling value Upred(k,n) s(k,n)vprescaie(k,n) The XYZ scale block 1532 provides predicted spatial metadata properties 714 as an output which is the output of the enhancer machine learning model 705. For the PLC ML model 715 the output 724 is for predicting a current subframe rather than output 714 which is for enhancing the current subframe. In some embodiments the PLC output 724 is the output of 1532 when the W-Net structure 1503 is operating in the configuration of PLC ML model. The machine learning model can be trained using any suitable process. In some examples the machine learning model can be trained using a set of spatial audio items in MASA format comprising an audio signal such as a transport audio signal and MASA spatial metadata with one directional field. The audio signal can comprise two channels. It is assumed that the original spatial metadata has high temporal resolution and all four sub-frames in each frame may contain unique values. The training data comprises 4530 items that are 4 seconds in length (that is, 200 frames or 800 subframes). The directional spatial metadata consisting of azimuth azi(k,n), elevation ele(k,n), and direct-to-total energy ratio r(k,n) are transformed into Cartesian xyz vector representation using target, x (k, ri)' V target ^.k, Tl) V target,y (k, n) Ttarget, z (k, n) = r(k,n) cos^azi^k, n)) cos (ele(k, ri)) sin(azi(fc,n)) cos (ele(k, ri)) sin (ele(k,n)) This is the reference or target data during the training. The validation data of 799 items is selected from this same pool of items. The description of metadata features pre-network 1520A above may be used for spatial metadata coded, for example, with 13.2 kbps IVAS codec. Other bitrates may result into different metadata feature representation 1504 being employed. For example, IVAS coding with 128 kbps may result in embodiments employing 8 distinct frequency bands. The ML model 705 for processing content encoded with, for example, 128 kbps, may be otherwise similar as the ML model 705 described above, but the metadata features pre-network 1520A may be adjusted. For example, the n_bands_meta may be 8, changing the size of the input 1704 in the first dimension. The dimension adjustment layer 1750 may still use a linear layer 1752 and adjust the size of the first dimension to, for example, 24 (or any suitable number). The rest of the metadata feature pre-network 1520A may be similar as described above, or it may feature other adjustments or changes. Similarly, the Cov feature pre-network 1520C, the SPAC feature pre-network 1520B, and combined prediction network 1528 may be configured differently for different bitrates, or as in the example embodiment, similar configurations can be used for different bitrates. For the input to ML model 705, the training items are passed through an IVAS codec (encoder and decoder) with a specific bitrate, for example, 13.2 kbps or 128 kbps, and the decoded MASA metadata is obtained using the external renderer (EXT) output mode of the IVAS decoder. The IVAS codec may reduce the frequency resolution of the spatial metadata into 5 bands (for the bitrate of 13.2 kbps, or 8 bands for 128 kbps) from the original 24 bands and applies quantization to the values. The IVAS codec may additionally reduce the temporal resolution of the spatial metadata by assigning the same parameter value to each of the four sub-frames. This metadata is transformed into the Cartesian XYZ vector representation in the same way as the reference data. The training uses batch size of 32, that is, the parameters of the model are adjusted after each 32 training examples. AdaDelta optimizer is used with learning rate of 1.0. The training is run for the maximum of 1000 epochs or until early stopping is triggered. The early stopping is triggered when the per-epoch validation loss is not lower than the best per-epoch validation loss in 50 consecutive epochs. The validation batch size is 8 items. The loss can be computed as follows. First, the direct-to-total energy ratios of the target and the predicted data are computed rtarget^n} = v^argetx(k, n) + v^argety(k,n) + vlargetz(k,n) rpred(fc,n) = ^p2redx(k,n) + vlredy(k,n) + v^,redz(k,n) Then, the absolute value of the energy ratio difference is computed ^absdiff^’^) \^predC^> ^target| Then, unit-length direction vectors are computed for the target and the predicted data v ■, (k n) - Vtar^k’n) ^tarqet.unitleny^f M fir \ rtarget^.n) x, _ Vpred(k,Yl) ^pred,unitlen ~ \ ^pred Then, direction error vector is computed by Verror (fc, 71) ^pred,unitlen $■> Tl) target,unitlen^-t^i n) and the length of the direction error vector is computed by error terror,x ^l) T terror,y ^l) T terror,z ^l) This length is weighted by the target direct-to-total energy ratio terror,wterror $■> ^)^"target^~t^i n) Using the determined absolute value of the energy ratio difference and the determined weighted length of the direction error vector, the combined error measure is determined by £>comb(k> ^absdiff Tt)lerror w^k, Tl) Then, an energy-weighting metric is determined. The energy-weighting metric is for weighting the loss based on the energies of the corresponding time-frequency tiles (that is, time-frequency tiles having larger energy should have a larger effect on the loss). The energy-weighting metric can be formulated in any suitable manner. In some examples, the local signal energy evolution over time in frequency bands with frequency dependent weighting (referred to as Cov5(nSf,k) which can be determined in a manner described in GB2411721) as computed by the feature computation block may be used as the weight. As Cov5(nSf,k) is computed in subframes nsf, and the loss is computed in frames n, the mean of the values Cov5(nSf,k) for the corresponding frame n are computed and set as the weight Eloss,w(.k>n) jy (^s / ' Sf nsfEn where n is the frame index, nsf the subframe index, and Nsf = 4 the number of subframes in a frame. Then, the energy-weighted combined error measure is determined by ^comb,w ^comb ^)^loss,w Using ^ombiW(k,n), the loss is determined by computing the mean over time and frequency over the entire training example ^comb,w ’ %comb,w(k’ k n which is the loss that is output from the loss function. With respect to Fig.9 is shown the temporal behavior of dilated residual block (DRB) described in Fig. 18 in the configuration for the enhancer ML model 705. In some embodiments the at least one features input 901 comprises three groups of features determined from both metadata and transport audio signals: metadata features and two sets of features from audio. However, these are example input features and in some other embodiments other features can be used or employed as inputs. The operation of a DRB block is conceptually illustrated in Fig.9 which shows an example simplified arrangement to show only the data axis corresponding to time. Thus, is shown the input features 901 a first convolution 963 to the intermediate features 967, and a second convolution 969 to the output features 971 and a by-pass 905 from the input features to output features 901. The example embodiment contains also convolutional operations along the axis corresponding to frequency which are not illustrated in the figure. The input features 901 boxes at the top row are the 7 consecutive input feature frames. These are processed with a first convolution 963 kernel of stride 1 and dilation 1 (or any suitable values) resulting in the intermediate features 967 on the middle row. These are then processed with a second convolution 969 stride 1 and dilation factor 2 (or any suitable values). This result is combined with the by-pass 905 (which may contain further convolutions) coming from the input features 901 to obtain the output features 971 on the bottom row. However, these convolutions are not directly affecting the feature generation and operation modes described above. The illustrated DRB block in Fig.9 is configured to operate in causal mode, in other words, the DRB block does not access future inputs. In other words, only the current and past (indices n - 6,n - 5, are used for determining the output for time instant n. The output in the current frame depends on the current input n and a number (here, 6) of earlier inputs. As there are multiple blocks of this kind stacked in the ll-Net, the temporal dependencies accumulate and each output has a dependency to a large number (e.g., 20-40) earlier inputs in addition to the current input. This history context has proven to be useful in the enhancement process. The example difference between the two models can be illustrated with respect to below in Fig. 10 and 11. In the enhancer ML model as shown in Fig. 10, the input features 702 are passed to the enhancer ML model 705 to generate the predicted metadata properties 714, whereas in Fig.11 the input features 702 are passed to the PLC ML model 715 to generate the predicted metadata properties 724. During the model training the target output for the enhancer ML model 705 is temporally synchronized with the model input: for producing the output for time instance n the inputs from time instances (n - N + 1) to n are used. This results that during inference time, the output at time instance n is determined based on the input features at time instances (n - N + 1) to n. In the training and use of PLC ML model, illustrated in Fig.11, the training target is selected one time index later. In other words, the input features from time indices (n - N) to (n - 1) are used to determine the predicted metadata properties 724 at time index n. This change means that it is possible to construct the PLC ML model 715 to have the same internal structure as enhancement ML model 705, and to use the same training procedure with training data in which the target has been temporally shifted by one index later. The difference in the temporal indices must be considered in features 702 when providing them to the enhancer ML model 705 and PLC ML model 715. With respect to Fig. 12 is shown schematically an example history metadata generator 703 as shown in Fig.7 in further detail according to some embodiments. The history metadata generator 703 is configured to emulate encoding and decoding the original spatial metadata into decoded spatial metadata. However, instead of employing the unavailable original spatial metadata the history metadata generator 703 is configured to employ the enhanced decoded spatial metadata 512 of the earlier frames as an input and produces a generated history spatial metadata 706. The history metadata generator 703 receives or obtains the enhanced decoded spatial metadata 512, and optionally the decoded transport audio signals as an input. The enhanced decoded spatial metadata 512 has the spatial metadata in the time-frequency resolution of the original MASA spatial metadata. In other words, in the current example there are 24 frequency bands and 4 temporal subframes. Then the history metadata generator 703 is configured to convert the obtained enhanced decoded spatial metadata 512 to the form that is compatible with the current ML model. In some embodiments the history metadata generator 703 can furthermore employ the optional decoded transport audio signals 504 input for determining transport audio signal energy in TF-tiles. This in turn can be used in a frequency band combiner 1211 and subframe combiner 1213 for applying additional weighting in the operations to emulate the operations of the (IVAS) encoder more closely. A frequency band limits determiner is configured to determine the number of frequency bands, and the limits or ranges 1202 of the frequency bands (in other words which original MASA frequency bands belong to each coded frequency band). This information and the frequency band limits are forwarded as frequency band limits 1202 to a frequency band combiner 1211. The frequency band combiner 1211 is configured to receive or obtain the enhanced decoded spatial metadata 512 and the frequency band limits 1202 as an input and combines the input 24 frequency bands to the target coding frequency bands (e.g., 5 frequency bands in our example), with optionally the aid of the decoded transport audio signals 504. The conversion can be implemented using the same or similar methods as is done in the (IVAS) encoder in order to have similar features for the spatial metadata as has been used in the training of the ML model (which has been trained using (IVAS) encoded spatial metadata). For example, the methods presented in UKIPO patent applications 1919130.3 and 1919131.1 may be used for determining the 5 combined frequency bands from the original 24 bands. The output of the block is frequency combined spatial metadata 1212. A subframe limit determiner is configured to determine the number of the subframes, and the limits or ranges of the subframes (in other words, which original MASA subframes belong to each coded subframe). For example, a bitrate employed can be 13.2 kbps, which means that the spatial metadata has 1 subframe. This information and the subframes limits are forwarded as the subframe limits 1204 to a subframe combiner 1213. The subframe combiner 1213 is configured to receive or obtain the frequency combined spatial metadata 1212 and the subframe limits 1204 as an input and combines the input 4 subframes to the target coding subframes (for example, 1 subframe in this example), with the optional aid of decoded transport audio signals 504. The conversion can be implemented employing the same or similar methods as performed in the (IVAS) encoder in order to have similar features for the spatial metadata as has been used in the training of the ML model (which has been trained using (IVAS) encoded spatial metadata). For example, the methods presented in UKIPO patent applications 1919130.3 and 1919131.1 can be used for determining combined subframes. If the original MASA spatial metadata and the target spatial metadata have both the same number of subframes, the combination can be skipped, and the frequency combined spatial metadata can be passed to the output. The output of the block is generated history spatial metadata 706. In addition to the frequency band combiner 1211 and the subframe combiner 1213, any other suitable processing can be performed or implemented as well in order to generate or determine data more resembling the decoded spatial metadata at the current bitrate. The output of the history metadata generator 703 is the generated history spatial metadata 706, that has the correct properties in the spatial metadata, for example, in this example 5 frequency bands and 1 subframe. The above example history metadata generator 703 is based on an IVAS codec. In some embodiments, when another codec or with different metadata resolution reduction method is employed, the history metadata generator is configured to emulate the operations performed by the codec on the metadata employed in the encoder. In some embodiments the obtained generated history spatial metadata can be obtained by employing a metadata encoding and decoding using the employed codec within the encoder. With respect to Fig. 13 is shown a flow diagram of the operations of the example history metadata generator as shown in Fig. 12. For example, as shown in Fig. 13 by 1301 is the operation of obtaining enhanced decoded spatial metadata and decoded transport audio signals. Additionally, as shown in Fig. 13 by 1303 is the operation of obtaining frequency band limits. Furthermore, as shown in Fig. 13 by 1305 is the operation of generating frequency combined spatial metadata based on the frequency band limits, enhanced decoded spatial metadata and decoded transport audio signals. Also, there is the operation of obtaining subframe as shown in Fig. 13 by 1307 Then, as shown in Fig. 13 by 1309 is the operation of generating history spatial metadata based on the frequency combined spatial metadata, subframe limits, and decoded transport audio signals. Then is the operation as shown in Fig. 13 by 1311 of outputting the history spatial metadata. The benefit of the current invention is demonstrated with an example embodiment. The two ML models are trained: enhancer ML model 705 and PLC ML model 715. The training of enhancer ML model 705 in this example was performed in a manner described in GB2411721. In the training of the PLC ML model, the model target is shifted by one temporal subframe earlier as described earlier. This means that given the input, the model’s task is to produce the next spatial metadata subframe as the output instead of only producing the enhanced output of the current subframe. Since also the loss function may utilize some signaldependent factors as weighting (e.g., factors computed from the audio signal energy) the same temporal shift is applied on these factors in order to keep them synchronized with the target output. With the exception of this temporal alignment change, the training of the PLC ML model 715 is performed similarly as the training of the enhancer ML model 705. Le., maximum of 1000 epochs, but early exit when the validation loss has not improved during the last 50 epochs. The optimizer used is Adadelta with the learning rate of 1.0. The training and validation batch size is 32 examples, each example being 4 seconds, i.e., 200 subframes, in length. The enhancer ML model 705 training converged after 638 epoch to training loss of 0.166 and validation loss of 0.185. The PLC ML model 715 converged after 702 epochs to training loss of 0.172 and validation loss of 0.196. The longer training, higher absolute loss values, and larger difference between the training and validation loss for the PLC ML model 715 compared to the enhancer ML model 705 can be interpreted as indications that the task of predicting the next spatial metadata frame is a more difficult task than the enhancement of current spatial metadata frame. The temporal receptive field of both models is 44 subframes. The simulation or experiment simulates packet losses so that with the interval of 25 frames, i.e., 100 subframes, one or two frames (alternating) of input is marked as lost. In a frame marked lost, the ML model input features 702 are set to zero for all four subframes (i.e., for 8 consecutive subframes for the loss of two frames). This is then given as the input to two systems: baseline 1401 and proposed 1403 systems. The baseline system 1401 is a system without implementing the embodiments described herein, i.e., operating only as described in GB2411721, for received and also for the lost frames. The proposed 1403 system operates as described in these embodiments progressing subframe-by-subframe over the input: When the current subframe has not been lost and no subframe has been lost in the receptive field length of temporal history, mode 1 or normal mode is used, i.e., the enhancer ML model is active with the real input features; When the current subframe has been lost, but no subframes have been lost in the receptive field length of temporal history, mode 2 or current subframe packet lost mode is used. The PLC ML model is used with the real input features for generating predicted metadata properties for the lost subframe, and the enhanced decoded spatial metadata from these. This result is given to history metadata generator and feature computer which then generate the metadata features for the lost subframe. These features are inserted in place of the lost subframe in the history context. The audio-based features are taken from the original features, assuming perfect audio packet loss concealment. However, as described in this document, these could be also generated based on the enhanced decoded spatial metadata if necessary; When the current subframe has been lost and subframes have been lost in the receptive field length of temporal history, the PLC ML model is used with mode 4 or current and previous subframes packets lost mode. For the subframes with packet loss in the history, the features from the generated history spatial metadata are used instead of the zero features from the loss. The resulting enhanced decoded spatial metadata are again fed back to history metadata generator similar to above; When the current subframe has not been lost, but subframes have been lost in the receptive field length temporal history, the enhancer ML model is used with (mode 3 or previous subframes packets lost mode). The ML model enhances the current spatial metadata subframe, but one or more history subframes contain features from generated history spatial metadata. The evaluation of the produced enhanced decoded spatial metadata is done by comparing against the target (spatial metadata at the encoder input). The produced enhanced decoded spatial metadata and spatial metadata are compared using an error function. The error function used has the original spatial metadata represented in a Cartesian vector representation vrefQc,n) for each spatial metadata sub-frame and band as the reference, and the enhanced spatial metadata represented in a Cartesian vector representation vprobeQc,n) for each spatial metadata sub-frame as the probe value, computes the length of the error vector for each TF-tile with e(K,n) = ||vref(K, n) — vprobe(K,n)|| computes the mean error over frequency bands as the time-varying error e(n) = _____1_____^<n_bands_meta n_bands_meta k = l e(K,n), and computes the overall error by averaging this over all subframes with _ 1 Yn-subframes z \ n_sub frames n=^ The error is computed for a reference case where no packet losses have been introduced and the spatial metadata enhancement is done using the enhancer ML model only (denoted no packet loss 1411, 1413 in Fig. 14). The second condition is when the packet loss has been introduced as described above. The baseline 1411 method uses this input and only the enhancer ML model for the processing. This corresponds to the situation without the current invention. The last version is the proposed method 1413, in which the current invention is used. To have a single value for comparison, the mean over per-item error values over all test items is computed. The resulting overall errors are From these results ^baseline, then the error is packet loss, as expected. ^reference 0.294226 ^baseline 0.309651 ^proposed 0.297571 it is possible to see that since ^reference <^proposed <the smallest in the reference situation 1405, without However, the packet losses increase the error in the baseline 1401 case. Furthermore, when the embodiments as described herein that the proposed 1403 method, is able to reduce the error compared to baseline 1401. This demonstrates the benefit and applicability of the invention. The spatial synthesizer 509 as shown in Fig.5 is configured to receive the enhanced decoded spatial metadata 512 and decoded transport audio signals 510, and employ any suitable spatial synthesis to generate the spatial audio signals or output audio signals. For example, the spatial synthesizer 509 may operate according to the principles described in PCT application WO2019086757A1. The cited reference describes also synthesis based on spread and surround coherence metadata parameters, but in some embodiments they can be assumed or set to be zero. Furthermore, even though the enhancement of the coherence parameters are not discussed herein, it would be understood that coherence parameters could be enhanced in a similar manner. In such implementations if the decoded spatial metadata has the coherence parameters, but they are not enhanced, then the decoded coherence parameters can be directly output (i.e., there would not be enhancement for those parameters). These values would then be used in the spatial synthesizer. Furthermore, in some embodiments the spatial synthesizer 509 can implement operations as described in UK application GB2218103.6, and the synthesis methods in 3GPP TS 26.253 IVAS specification. In some embodiments, the enhancer ML model and PLC ML model may use a different size of input along the temporal dimension, for example the PLC ML model can employ an input size of N - 1 so that the size of the history features to be stored into the past remains constant. In some embodiments, the audio generator is capable of producing generated transport audio signals of such high fidelity that they can be used directly in place of decoded transport audio signals as an input to the feature computer when determining the features for the current subframe. In some embodiments, the audio generator is not capable of producing generated transport audio signals of sufficiently high fidelity that they could be used directly in place of decoded transport audio signals in feature computer. In these cases, for example, only the energies are computed from the generated transport audio signals, and only the monaural features are computed using them alone, and the inter-channel features are computed using the energies and the generated history spatial metadata (e.g., inter-channel coherences can be estimated using the energies and the direct-to-total energy ratios). With respect to Fig. 19 an example electronic device which may be used as the computer, encoder processor, decoder processor, or any of the functional blocks described herein is shown. The device may be any suitable electronics device or apparatus. For example, in some embodiments the device 1900 is a mobile device, user equipment, tablet computer, computer, audio playback apparatus, a laptop, or a teleconferencing system. In some embodiments the device 1900 comprises at least one processor or central processing unit (CPU or processor) 1907. The processor 1907 can be configured to execute various program codes such as the methods such as described herein. The device 1900 furthermore comprises a transceiver 1909 which is configured to receive the bitstream and provide it to the processor 1907. Typically, the connection is wirelessly received data from a remote device or a server, however, in some embodiments the bitstream is received via a wired connection or read from a local memory of the device. The transceiver can communicate with further apparatus by any suitable known communications protocol. For example, in some embodiments the transceiver can use a suitable 5G New Radio (5G NR) protocol, a Wi-Fi protocol such as, for example, IEEE 802.11 be, a suitable short-range radio frequency communication protocol such as Bluetooth, or Li-Fi). The device may furthermore comprise a user interface (UI) 1905 which may display to the user an interface for interacting with the device. The device 1900 may further comprise memory (MEM) 1911 which is coupled to the processor 1907. In some embodiments the memory 1911 comprises the program code 1921 which is executed by the processor 1907. The program code may involve instructions to perform the operations of the spatial synthesizer described above. The processor 1907 can then be configured to output the spatial audio signals, which in this example was a binaural output, to a digital to analogue converter (DAC) / Bluetooth 1901 converter. The combination of the processor, CPU, 1907 and memory, MEM, can implement the IVAS decoder 1931 functionality described above. The DAC / Bluetooth 1901 is configured to convert the spatial audio signals to an analogue form if the headphones are conventional wired (analogue) headphones. For wireless connections, the DAC / Bluetooth 1901 may be a Bluetooth transceiver. The DAC / Bluetooth 1901 block provides (either wired or wirelessly) the spatial audio to be played back with the headphones 1903 to the user. In some embodiments, the headphones 1903 may have a head tracker which may provide orientation and / or position information of the user’s head to the processor 1907 of the rendering apparatus, so that user’s head orientation is accounted for at the spatial synthesizer. It should be understood that the apparatuses may comprise or be coupled to other units or modules used in or for transmission and / or reception. Although the apparatuses have been described as one entity, different modules and memory may be implemented in one or more physical or logical entities. Embodiments of the invention can be practised as a computer software product in the form of an App. The App can either reside on the electronic / computing device or in an App repository such as an “App Store” which is typically sited remotely from the electronic / computing device. When the App is sited in an “App Store,” the app is typically downloaded from the “App Store” to an electronic / computing device over a communication network, such as an IP based network. The downloaded App, comprising the invention, can execute as a computer software product on the electronic / computing device. It is also noted herein that while the above describes example embodiments, there are several variations and modifications which may be made to the disclosed solution without departing from the scope of the present invention. As used herein, “at least one of the following: ” and “at least one of ” and similar wording, where the list of two or more elements are joined by “and” or “or”, mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements. In general, the various embodiments may be implemented in hardware or special purpose circuitry, software, logic or any combination thereof. Some aspects of the disclosure may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the disclosure is not limited thereto. While various aspects of the disclosure may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof. As used in this application, the term “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry) and (b) combinations of hardware circuits and software, such as (as applicable): (c) a combination of analog and / or digital hardware circuit(s) with software / firmware and (i) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions); and (ii) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation. This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device. The embodiments of this disclosure may be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Computer software or program, also called program product, including software routines, applets and / or macros, may be stored in any apparatus-readable data storage medium and they comprise program instructions to perform particular tasks. A computer program product may comprise one or more computerexecutable components which, when the program is run, are configured to carry out embodiments. The one or more computer-executable components may be at least one software code or portions of it. Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as DVD and the data variants thereof, CD. The physical media is a non-transitory media. The term “non-transitory,” as used herein, is a limitation of the medium itself (i.e., tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM). The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processors may be of any type suitable to the local technical environment, and may comprise one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), FPGA, gate level circuits and processors based on multi core processor architecture, as non-limiting examples. Embodiments of the disclosure may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate. The scope of protection sought for various embodiments of the disclosure is set out by the independent claims. The embodiments and features, if any, described in this specification that do not fall under the scope of the independent claims are to be interpreted as examples useful for understanding various embodiments of the disclosure. The foregoing description has provided by way of non-limiting examples a full and informative description of the exemplary embodiment of this disclosure. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this disclosure will still fall within the scope of this invention as defined in the appended claims. Indeed, there is a further embodiment comprising a combination of one or more embodiments with any of the other embodiments previously discussed. List of abbreviations AAC - Advanced Audio Coding ANN - artificial neural network BN - Batch Normalization DNN - deep neural network EVS - 3GPP Enhanced Voice Services FSAC - Future Speech and Audio Codec IVAS - 3GPP Immersive Voice and Audio Services kbps - kilobits per second MASA - Metadata-Assisted Spatial Audio MDCT- modified discrete cosine transform ML - machine learning NN - neural network ONNX- Open Neural Network exchange PLC - Packet loss concealment STFT - Short-time Fourier transform TF - time / frequency VR - virtual reality
Claims
1. An apparatus for selectively enhancing at least one spatial metadata parameter, the apparatus comprising at least one processor and at least one memory including a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform:obtaining a spatial bitstream, the spatial bitstream defining spatial audio content and comprising: at least one metadata parameter associated with at least one audio signal, the at least one metadata parameter being an encoded version of an original metadata parameter;obtaining a first machine-learning model for enhancing the at least one metadata parameter;obtaining a second machine-learning model for determining the at least one metadata parameter when the at least one metadata parameter is missing;determining for the spatial bitstream whether content comprising spatial audio content has been obtained in a current frame or is missing;selecting one of the first machine-learning model or second machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content has been obtained in the current frame or is missing; andenhancing the at least one metadata parameter for the current frame which has been obtained or determining the at least one metadata parameter which was missing based on the selected one of the first machine-learning model or second machine-learning model.
2. The apparatus as claimed in claim 1, further caused to perform determining at least one current frame input feature for the first machine-learning model based on the determination for spatial bitstream whether content comprising spatial audio content has been obtained for the current frame.
3. The apparatus as claimed in claim 2, caused to perform determining at least one current frame input feature for the selected first machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content has been obtained for the current frame is further caused to perform determining the at least one current frame input feature for the selected first machine-learning model based on the at least one metadata parameter associated with the at least one audio signal from the current frame of the spatial bitstream.
4. The apparatus as claimed in claim 3, caused to perform determining at least one current frame input feature for the selected first machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content has been obtained for the current frame is further caused to perform: determining the at least one current frame input feature for the selected first machine-learning model further based on the at least one audio signal from the current frame of the spatial bitstream.
5. The apparatus as claimed in any of claims 1 to 4, caused to perform determining at least one spatial metadata parameter without determining input features for the current frame based on the determination of the spatial bitstream whether content comprising spatial audio content is missing.
6. The apparatus as claimed in claim 5, caused to perform selecting one of the first machine-learning model or second machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content is missing is caused to perform selecting the second machine-learning model.
7. The apparatus as claimed in any of claims 1 to 4, caused to perform determining at least one current frame input feature for the second machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content is missing is further caused to perform:determining at least one of:an estimated audio signal based on at least one previous frame audio signal; andat least one estimated metadata parameter associated with the at least one audio signal based on at least one previous frame metadata parameter;determining the at least one current frame input feature based on the at least one of: the estimated audio signal; and the at least one estimated metadata parameter.
8. The apparatus as claimed in any of claims 1 to 7, further caused to perform determining for all from a determined number of previous frames whether content comprising spatial audio content has been obtained or any content comprising spatial audio content is missing.
9. The apparatus as claimed in claim 8, caused to perform determining at least one previous frame input feature for the selected one of the first machine-learning model or second machine-learning model is caused to perform determining the at least one previous frame input feature based on determining for all from a determined number of previous frames whether content comprising spatial audio content has been obtained or any content comprising spatial audio content is missing.
10. The apparatus as claimed in claim 9, caused to perform determining the at least one previous frame input feature for the selected one of the first machine-learning model or second machine-learning model is further caused to perform:when the at least one previous frame content comprising spatial audio content has been obtained determining the at least one input feature based on the at least one metadata parameter associated with the at least one audio signal from the at least one previous frame; andwhen at least one from the determined number of previous frames content comprising spatial audio content is missing:determining, for the previous frame with missing spatial audio content, at least one generated metadata parameter based on at least oneenhanced metadata parameter and / or predicted metadata property of the previous frame with missing spatial audio content; anddetermining the at least one input feature based on the generated metadata parameter.
11. The apparatus as claimed in claim 10, caused to perform determining the at least one previous frame input feature for the selected one of the first machine-learning model or second machine-learning model is further caused to perform:when the at least one previous frame content comprising spatial audio content has been obtained determining the at least one input based on the at least one audio signal from the at least one previous frame of the spatial bitstream; andwhen at least one from the determined number of previous frames content comprising spatial audio content is missing:determining at least one generated audio signal, for the previous frame with missing spatial audio content, based on at least one prior frame audio signal, the prior frame being before the previous frame; anddetermining the at least one input feature based on the generated audio signal.
12. The apparatus as claimed in any of claims 1 to 11, caused to perform obtaining the first machine-learning model is caused to perform, at least one of:obtaining the first machine-learning model based on the spatial audio content; andreceiving the first machine-learning model from at least one further apparatus.
13. The apparatus as claimed in any of claims 1 to 12, caused to perform obtaining the second machine-learning model is further caused to perform at least one of:obtaining the second machine-learning model based on the spatial audio content with a temporal offset first machine-learning model spatial audio content; andreceiving the second machine-learning model from at least one further apparatus.
14. The apparatus as claimed in any of claims 1 to 13, caused to perform enhancing the at least one metadata parameter for the current frame which has been obtained or determining the at least one metadata parameter for the current frame which was missing based on the selected one of the first machine-learning model or second machine-learning model is further caused to perform:employing the selected one of the first machine-learning model or second machine-learning model, the selected one of the first machine-learning model or second machine-learning model configured to output at least one predicted metadata property; andgenerating at least one enhanced metadata parameter for the current frame which has been obtained or predicting the at least one metadata parameter for the current frame which was missing based on the at least one predicted metadata property.
15. A method for an apparatus for selectively enhancing at least one spatial metadata parameter, the method comprising:obtaining a spatial bitstream, the spatial bitstream defining spatial audio content and comprising: at least one metadata parameter associated with at least one audio signal, the at least one metadata parameter being an encoded version of an original metadata parameter;obtaining a first machine-learning model for enhancing the at least one metadata parameter;obtaining a second machine-learning model for determining the at least one metadata parameter when the at least one metadata parameter is missing;determining for the spatial bitstream whether content comprising spatial audio content has been obtained in a current frame or is missing;selecting one of the first machine-learning model or second machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content has been obtained in the current frame or is missing; andenhancing the at least one metadata parameter for the current frame which has been obtained or determining the at least one metadata parameter which was missing based on the selected one of the first machine-learning model or second machine-learning model.
16. The method as claimed in claim 15, further comprising determining at least one current frame input feature for the first machine-learning model based on the determination for spatial bitstream whether content comprising spatial audio content has been obtained for the current frame.
17. The method as claimed in claim 16, wherein determining at least one current frame input feature for the selected first machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content has been obtained for the current frame further comprises determining the at least one current frame input feature for the selected first machine-learning model based on the at least one metadata parameter associated with the at least one audio signal from the current frame of the spatial bitstream.
18. The method as claimed in claim 17, wherein determining at least one current frame input feature for the selected first machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content has been obtained for the current frame further comprises: determining the at least one current frame input feature for the selected first machine-learning model further based on the at least one audio signal from the current frame of the spatial bitstream.
19. The method as claimed in any of claims 15 to 18, further comprising determining at least one spatial metadata parameter without determining input features for the current frame based on the determination of the spatial bitstream whether content comprising spatial audio content is missing.
20. The method as claimed in claim 19, wherein selecting one of the first machine-learning model or second machine-learning model based on thedetermination for the spatial bitstream whether content comprising spatial audio content is missing comprises selecting the second machine-learning model.
21. The method as claimed in any of claims 15 to 18, wherein determining at least one current frame input feature for the second machine-learning model based on the determination for the spatial bitstream whether content comprising spatial audio content is missing further comprises:determining at least one of:an estimated audio signal based on at least one previous frame audio signal; andat least one estimated metadata parameter associated with the at least one audio signal based on at least one previous frame metadata parameter; anddetermining the at least one current frame input feature based on the at least one of: the estimated audio signal; and the at least one estimated metadata parameter.
22. The method as claimed in any of claims 15 to 21, further comprising determining for all from a determined number of previous frames whether content comprising spatial audio content has been obtained or any content comprising spatial audio content is missing.
23. The method as claimed in claim 22, wherein determining at least one previous frame input feature for the selected one of the first machine-learning model or second machine-learning model further comprises determining the at least one previous frame input feature based on determining for all from a determined number of previous frames whether content comprising spatial audio content has been obtained or any content comprising spatial audio content is missing.
24. The method as claimed in claim 23, wherein determining the at least one previous frame input feature for the selected one of the first machine-learning model or second machine-learning model further comprises:when the at least one previous frame content comprising spatial audio content has been obtained determining the at least one input feature based on the at least one metadata parameter associated with the at least one audio signal from the at least one previous frame; andwhen at least one from the determined number of previous frames content comprising spatial audio content is missing:determining, for the previous frame with missing spatial audio content, at least one generated metadata parameter based on at least one enhanced metadata parameter and / or predicted metadata property of the previous frame with missing spatial audio content; anddetermining the at least one input feature based on the generated metadata parameter.
25. The method as claimed in claim 24, wherein determining the at least one previous frame input feature for the selected one of the first machine-learning model or second machine-learning model further comprises:when the at least one previous frame content comprising spatial audio content has been obtained determining the at least one input based on the at least one audio signal from the at least one previous frame of the spatial bitstream; andwhen at least one from the determined number of previous frames content comprising spatial audio content is missing:determining at least one generated audio signal, for the previous frame with missing spatial audio content, based on at least one prior frame audio signal, the prior frame being before the previous frame; anddetermining the at least one input feature based on the generated audio signal.
26. The method as claimed in any of claims 15 to 25, wherein obtaining the first machine-learning model comprises, at least one of:obtaining the first machine-learning model based on the spatial audio content; andreceiving the first machine-learning model from at least one further apparatus.
27. The method as claimed in any of claims 15 to 26, wherein obtaining the second machine-learning model further comprises at least one of:obtaining the second machine-learning model based on the spatial audio content with a temporal offset first machine-learning model spatial audio content; andreceiving the second machine-learning model from at least one further apparatus.
28. The method as claimed in any of claims 15 to 27, wherein enhancing the at least one metadata parameter for the current frame which has been obtained or determining the at least one metadata parameter for the current frame which was missing based on the selected one of the first machine-learning model or second machine-learning model comprises:employing the selected one of the first machine-learning model or second machine-learning model, the selected one of the first machine-learning model or second machine-learning model configured to output at least one predicted metadata property; andgenerating at least one enhanced metadata parameter for the current frame which has been obtained or predicting the at least one metadata parameter for the current frame which was missing based on the at least one predicted metadata property.
Citation Information
Patent Citations
ViewGB2615323AonEspacenetopensinnewtab
ViewUS2004039464A1onEspacenetopensinnewtab
ViewWO2023141034A1onEspacenetopensinnewtab
ViewGB2617055AonEspacenetopensinnewtab