Spatial metadata for rendering of spatial audio

By processing encoded audio and spatial metadata with machine learning models, the challenges of maintaining audio quality at low bitrates in spatial audio rendering are addressed, achieving improved accuracy and resolution of sound source perception.

GB2643271APending Publication Date: 2026-02-11NOKIA TECHNOLOGIES OY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
GB2024011721
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-08
Publication Date
2026-02-11

AI Technical Summary

Technical Problem

Existing spatial audio rendering technologies face challenges in maintaining audio quality at low bitrates due to the adverse effects of compression on spatial metadata, leading to inaccuracies and instability in sound source perception.

Method used

Utilizing machine learning models, particularly deep neural networks, to process encoded audio and spatial metadata, enhancing spatial metadata by leveraging redundancies between audio signals and metadata to improve temporal and frequency resolution, and accuracy of sound source direction and energy ratios.

Benefits of technology

Enhances spatial audio rendering by improving the accuracy and resolution of spatial metadata, resulting in more stable and accurate sound source perception even at low bitrates, thus maintaining audio quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention uses spatial metadata to render spatial audio by using information from corresponding audio signals, such as in metadata assisted spatial audio (MASA). An encoded audio signal 204 and co
Need to check novelty before this filing date? Find Prior Art

Description

TECHNOLOGICAL FIELD Examples of the disclosure relate to spatial metadata for rendering spatial audio. Some relate to enhancing the spatial metadata for rendering spatial audio by using information from the corresponding audio signals. BACKGROUND Metadata assisted spatial audio (MASA) uses one or more audio signals together with corresponding spatial metadata to generate spatial audio. The spatial metadata can comprise direction information, direct-to-total energy ratios in frequency bands, or any other suitable information. A MASA stream can, in some examples, be obtained by capturing spatial audio with microphones and estimating the corresponding spatial metadata from the microphone signals. BRIEF SUMMARY According to various, but not necessarily all, examples of the disclosure there is provided an apparatus comprising means for: receiving an encoded audio signal and corresponding encoded spatial metadata; determining at least one input for processing wherein the input is based on the encoded audio signal and the encoded spatial metadata; processing the at least one input to use information from the encoded audio signal and the encoded spatial metadata to generate processed spatial metadata; and enabling spatial rendering to be performed on one or more audio signals using the processed spatial metadata. The means may be for: determining a first input for processing wherein the first input is based on the encoded audio signal; determining a second input for processing wherein the second input is based on the encoded spatial metadata; and processing the first input and the second input to use information from the encoded audio signal and the encoded spatial metadata to generate processed spatial metadata. Determining an input based on the encoded audio signal may comprise at least partially decoding the encoded audio signal. Determining an input based on the encoded spatial metadata may comprise at least partially decoding the encoded spatial metadata. The processed spatial metadata may comprise at least one of: improved temporal resolution compared to the received spatial metadata; improved frequency resolution compared to the received spatial metadata; or improved accuracy of restoration of information from spatial metadata from which the encoded spatial metadata was obtained compared to the received spatial metadata. The processed spatial metadata may comprise one or more energy-related parameters and one or more directional parameters. The energy-related parameters may comprise at least one of: a direct-to-total energy ratio; a diffuse-to-total energy ratio. The directional parameters may comprise at least one of: an azimuth angle; an elevation angle; a direction index. The means may be for obtaining one or more features from the received encoded audio signal and using the one or more obtained features as the input to the processing. Processing the at least one input may comprise obtaining one or more features from the received encoded audio signal. The means may be for obtaining one or more features from the received encoded spatial metadata and using the one or more obtained features as, at least part of, the input for processing. Processing the at least one input may comprise obtaining one or more features from the received encoded spatial metadata The processing the at least one input may comprise obtaining predicted spatial metadata properties and using the predicted spatial metadata properties and the received encoded spatial metadata to generate the processed spatial metadata. The spatial rendering may comprise at least one of: binaural rendering; multi-loudspeaker rendering; Ambisonics rendering; stereo rendering; or cross talk-cancelled stereo rendering. The processing may be performed, at least in part, using a machine learning model. The machine learning model may comprise a deep neural network model. The one or more audio signals may be obtained by decoding the received encoded audio signal. The one or more audio signals may comprise one or more channels. The one or more encoded audio signals may comprise one or more channels. According to various, but not necessarily all, examples of the disclosure there is provided a method comprising: receiving an encoded audio signal and corresponding encoded spatial metadata; determining at least one input for processing wherein the input is based on the encoded audio signal and the encoded spatial metadata; processing the at least one input to use information from the encoded audio signal and the encoded spatial metadata to generate processed spatial metadata; and enabling spatial rendering to be performed on one or more audio signals using the processed spatial metadata. According to various, but not necessarily all, examples of the disclosure there is provided a computer program comprising instruction which, when executed by a processor, cause the processor to perform: receiving an encoded audio signal and corresponding encoded spatial metadata; determining at least one input for processing wherein the input is based on the encoded audio signal and the encoded spatial metadata; processing the at least one input to use information from the encoded audio signal and the encoded spatial metadata to generate processed spatial metadata; and enabling spatial rendering to be performed on one or more audio signals using the processed spatial metadata. According to various, but not necessarily all, embodiments there is provided an apparatus comprising at least one processor; and at least one memory including computer program code; the at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to perform at least a part of one or more methods described herein. According to various, but not necessarily all, embodiments there is provided an apparatus comprising means for performing at least part of one or more methods described herein. The description of a function and / or action should additionally be 5 considered to also disclose any means suitable for performing that function and / or action. Functions and / or actions described herein can be performed in any suitable way using any suitable method. According to various, but not necessarily all, embodiments there is provided examples as claimed in the appended claims. While the above examples of the disclosure and optional features are described separately, it is to be understood that their provision in all possible combinations and permutations is contained within the disclosure. It is to be understood that various examples of the disclosure can comprise any or all the features described in respect of other examples of the disclosure, and vice versa. Also, it is to be appreciated that any one or more or all the features, in any combination, may be implemented by / comprised in / performable by an apparatus, a method, and / or computer program instructions as desired, and as appropriate. The description of a function should additionally be considered to also disclose any means suitable for performing that function BRIEF DESCRIPTION Some examples will now be described with reference to the accompanying drawings in which: FIG. 1 shows an example system; FIG. 2 shows an example encoder; FIG. 3 shows an example decoder; FIG. 4 shows an example method; FIG. 5 shows an example metadata enhancer; FIG. 6 shows another example metadata enhancer; FIGS. 7A and 7B show an example machine learning model; FIG. 8 shows an example pre-network; FIG. 9 shows an example ResBlock; FIG. 10 shows an example system; FIGS. 11A to 11D show example results; FIG. 12 shows an example device; and FIG. 13 shows an example apparatus. The figures are not necessarily to scale. Certain features and views of the figures can be shown schematically or exaggerated in scale in the interest of clarity and conciseness. For example, the dimensions of some elements in the figures can be exaggerated relative to other elements to aid explication. Corresponding reference numerals are used in the figures to designate corresponding features. For clarity, all reference numerals are not necessarily displayed in all figures. DETAILED DESCRIPTION Fig. 1 shows an example system 100 that can be used for spatial audio. The system 100 comprises an encoder 106 and a decoder 110. The encoder 106 and decoder 110 can be in different devices. For example, the system 100 can be part of a telecommunications system in which the encoder 106 can be provided within a first communications device and the decoder 110 can be provided within a different communications device. The encoder 106 receives a spatial audio stream as an input. The spatial audio stream comprises audio signals 102 and spatial metadata 104. The spatial audio stream can be provided in a metadata-assisted spatial audio (MASA) format or in any other suitable format. In some examples the audio signals 102 can comprise microphone signals or signals that are obtained by processing microphone signals. In some examples the audio signals 102 can originate from other inputs such as multi-channel audio signals (for example, 5.1 or 7.1+4) or audio objects. In such cases the audio signals 102 can comprise a downmix of the originally obtained inputs. The audio signals 102 can comprise transport audio signals. The transport audio signals are audio signals that have been processed into a format for transport or transmission to another communication device. The spatial metadata 104 comprises information that can be used to render the audio signals 102 to generate spatial audio. The spatial metadata 104 can comprise direction information, direct-to-total energy ratios in frequency bands, or any other suitable information. The spatial metadata 104 can be estimated from the microphone signals or from any other suitable signals. The encoder 106 is configured to encode the audio signals 102 and the spatial metadata 104 to form a bitstream 108. The bitstream 108 can be sent to the decoder 110. Encoder 106 and the decoder 110 can be in different devices and so the bitstream can be sent via any suitable communication network. The decoder 110 receives the bitstream 108 as an input. The decoder 110 is configured to decode the bitstream 108 to render spatial audio output 112. The spatial audio output can comprise binaural audio signals, stereo signals, Ambisonics signals or any other suitable types of signals. Fig. 2 shows an example encoder 106 in more detail. The encoder 106 is configured so that the audio signals 102 are provided to an audio encoder 200. The audio encoder 200 is configured to encode the audio signals 102. The audio encoder 200 can use any suitable audio-signal encoder to encode the audio signals 102. For example, the audio encoder 200 can use the Immersive Voice and Audio Services (IVAS) core coder, Enhanced Voice Services (EVS), or Advanced Audio Coding (AAC). The audio encoder 200 provides encoded audio signals 204 as an output. The encoded audio signals 410 can comprise one or more channels. The encoder 106 is also configured so that the spatial metadata 104 is provided to a spatial metadata encoder 202. The spatial metadata encoder 202 is configured to encode the spatial metadata 104. The spatial metadata encoder 202 can use any method to encode the spatial metadata 104. For example, the spatial metadata encoder 202 can use the MASA encoding methods of the IVAS encoder or any other suitable encoding methods. The spatial metadata encoder 202 provides encoded spatial metadata 206 as an output. The encoded audio signals 204 and encoded spatial metadata 206 are provided as an input to a multiplexer 208. The multiplexer 208 multiplexes the encoded audio signals 204 and encoded spatial metadata 206 to a bitstream 108. The bitstream 108 is the output of the encoder 106. The encoding methods that are used to encode the spatial metadata 104 control the level of compression depending on the bitrate. Using a lower bitrate can reduce the reconstruction accuracy of the spatial metadata 104 at the decoder 110. With a low bitrate a higher compression factor is needed and this is obtained by trading off some of the reconstruction accuracy. For example, at a very low bitrates (such as 13.2 kbps with IVAS), the bitrate that can be allocated to spatial metadata 104 encoding is so low (such as about 2.25 kbps with IVAS) that it adversely affects the perceptual quality of the spatial audio output 112 at the decoder 110. These adverse effects can include instability and inaccuracies of the perceived sound sources. Allocating more bits to the spatial metadata 104 does not resolve this issue because there are constraints on the total bitrate. Increasing the bits used for the spatial metadata 104 would reduce the bits available for the coding of the audio signals 102. This could result in a perceivable degradation of audio quality. There are some redundancies between the audio signals 102 and the spatial metadata 104. For example, when there is an onset in the audio (such as, when someone starts to speak), typically both the energy of the audio signal 102 and the direct-to-total energy ratio of the spatial metadata increase rapidly in the time-frequency tiles at the onset. Similarly, if there is a dominant audio source on the left side, the directional spatial metadata at those time-frequency tiles points at left. Examples of the disclosure make use of these redundancies to enable improved or enhanced spatial metadata to be obtained. The improved or enhanced spatial metadata can be obtained by using information from the audio signals to process the spatial metadata. This can help to account for any information that is lost in the encoding of the spatial metadata 102 or any other encoding and avoid perceivable degradations of audio quality even when a low bitrate is used. Fig. 3 shows an example decoder 110 that can be used in examples of the disclosure. The decoder 110 can be part of a system 100 as shown in Fig. 1. The decoder 110 receives the bitstream 108 as an input. The bitstream 108 can be generated by an encoder 106 as shown in Figs. 1 and 2. The bitstream 108 is provided to a demultiplexer 300. The demultiplexer 300 demultiplexes the bitstream 108 to provide encoded audio signals 204 and encoded spatial metadata 206. The encoded audio signals 204 can be encoded transport audio signals or any other type of audio signals. The encoded audio signals 204 are provided to an audio decoder 302. The audio decoder 302 is configured to decode the encoded audio signals. The audio decoder 302 can use any suitable type of decoder. The audio decoder 302 can use a decoder that is compatible with the audio encoder 200 that was used to encode the audio signals 102. The audio decoder 302 provides decoded audio signals 306 as an output. These can be decoded transport audio signals. The encoded spatial metadata 206 is provided to a metadata decoder 304. The metadata decoder 304 is configured to decode the encoded spatial metadata 206. The metadata decoder 304 can use any suitable type of decoder. The metadata decoder 304 can use a decoder that is compatible with the spatial metadata encoder 202 that was used to encode the spatial metadata 104. The metadata decoder 304 provides decoded spatial metadata 308 as an output. The decoded audio signals 306 and the decoded spatial metadata 308 are provided as an input to a metadata processor 310. The metadata processor 310 is configured to process the decoded spatial metadata 308 so that information from the decoded audio signals 306 can be used to improve or enhance the spatial metadata and provide improved spatial audio. The metadata processor 310 can comprise a machine learning model or any other suitable means. The metadata processor 310 provides processed spatial metadata 312 as an output. This can be processed decoded spatial metadata. The processed decoded spatial metadata 312 and the decoded audio signals 306 are provided as an input to a spatial synthesizer 314. The spatial synthesizer 314 is configured to process the processed decoded spatial metadata 312 and the decoded audio signals 306 to render spatial audio output 316. The spatial audio output can comprise binaural audio signals, stereo audio signals, multi-channel audio signals (e.g., 5.1 or 7.1+4), Ambisonics signals, or any other suitable type of spatial audio. Fig. 4 shows an example method that can be used in examples of the disclosure. The method could be implemented by a decoder 110 or an apparatus or means within a decoder 110. The method comprises, at block 400, receiving an encoded audio signal 204 and corresponding encoded spatial metadata 206. In some examples the encoded audio signal 204 and corresponding encoded spatial metadata 206 can be received in a combined bitstream 108 and then separated out by a demultiplexer 300 or any other suitable means. The received encoded audio signal 204 can comprise one or more channels. The encoded spatial metadata 206 corresponds to the encoded audio signal 204 in that the spatial metadata and the audio signals represent the same sound scene. The spatial metadata can be used to process the audio signals to render spatial audio. At block 402 the method comprises determining at least one input for processing. The input is based on the encoded audio signal 204 and the encoded spatial metadata 206. The processing can comprise processing that is performed on the spatial metadata so as to improve the spatial audio that can be achieved using the spatial metadata. Determining an input based on the encoded audio signal 204 can comprise at least partially decoding the encoded audio signal 204. Determining an input based on the encoded spatial metadata 206 can comprise at least partially decoding the encoded spatial metadata 206. In some examples determining an input based on the encoded audio signal 204 can comprise obtaining one or more features from the received encoded audio signal 204 and using the one or more obtained features as the input to the processing. The features can comprise information relating to the physical properties of the audio signal. The features can comprise peak energy levels, evolution of energy levels, repetitions of energy levels, or any other suitable information. Similarly, in some examples determining an input based on the encoded spatial metadata 206 can comprise obtaining one or more features from the received encoded spatial metadata 206 and using the one or more obtained features as the input to the processing. The features can comprise information relating to the properties of the spatial metadata. At block 404 the method comprises processing the at least one input to use information from the encoded audio signal 204 and the encoded spatial metadata 206 to generate processed spatial metadata 312. The information from the encoded audio signal 204 can comprise information that is inherent within the encoded audio signal 204, features that are extracted from the encoded audio signal 204 or any other suitable information. The processing the at least one input to use information from the encoded audio signal 204 and the encoded spatial metadata 206 to generate processed spatial metadata 312 can comprise obtaining predicted spatial metadata properties and using the predicted spatial metadata properties and the received encoded spatial metadata 206 to generate the processed spatial metadata 312. The processed spatial metadata 312 that is generated by the processing can comprise one or more energy-related parameters and one or more directional parameters and / or any other suitable parameters. The energy related parameters can comprise a direct-to-total energy ratio, a diffuse-to-total energy ratio and / or any other suitable parameters. The directional parameters can comprise any information that indicates a direction of arrival of sound. The directional parameters can comprise an azimuth angle, an elevation angle, a direction index and / or any other suitable information. The processed spatial metadata 312 can comprise enhanced or improved spatial metadata compared to the original spatial metadata. The original spatial metadata can be the metadata from which the encoded spatial metadata is obtained 206. In some examples the processed spatial metadata 312 can comprise improved temporal resolution compared to the received spatial metadata. In some examples the processed spatial metadata 312 can comprises improved frequency resolution compared to the received spatial metadata. In some examples the processed spatial metadata 312 can comprise improved accuracy of restoration of information from spatial metadata from which the encoded spatial metadata 206 was obtained compared to the received spatial metadata. The accuracy of restoration of information is the accuracy with which a value that is closer to the original spatial metadata is obtained compared to the encoded spatial metadata 206. In some examples the input for the processing can comprise one or more features. For instance, determining the input can comprise determining one or more features of the encoded audio signals 204 or one or more features of the encoded spatial metadata 206. In other examples the processing can comprise obtaining one or more features from the received encoded audio signals 204 or the received encoded spatial metadata 206. The processing can be performed, at least in part, using a machine learning model. The machine learning model can comprise a deep neural network model and / or any other suitable operations or combinations of operations. The method comprises, at block 406, enabling spatial rendering to be performed on one or more audio signals 306 using the processed spatial metadata. The spatial rendering can be performed on decoded audio signals. The decoded audio signals can be obtained by decoding the encoded audio signals 204 that were received at block 400. The spatial rendering can comprise binaural rendering, multi-channel loudspeaker rendering, Ambisonics rendering, stereo rendering, cross talk-cancelled stereo rendering, or any other suitable type of rendering. Variations to this method could be used. In some examples a single input can be provided for the processing. In other examples multiple inputs can be provided. If multiple inputs are used the method can comprise determining a first input for processing wherein the first input is based on the encoded audio signal 204 and determining a second input for processing wherein the second input is based on the encoded spatial metadata 206. In such examples the first input and the second input would be processed to use information from the encoded audio signal 204 and the encoded spatial metadata 206 to generate processed spatial metadata 312. Fig. 5 shows an example metadata processor 310 that can be used in some examples. The metadata processor 310 can be provided within a decoder 110 as shown in Fig. 3. The metadata processor 310 can be used to implement methods such as the method of Fig. 4 and / or any variations of this method. Decoded audio signals 306 and the decoded spatial metadata 308 are provided as an input to the metadata processor 310. The decoded audio signals 306 and the decoded spatial metadata 308 have both been encoded and decoded so the data that is received by the metadata processor 310 has been compressed in some way. In some use case scenarios the decoded spatial metadata 308 can comprise information at a low time-frequency resolution. This could be because the spatial metadata is coded using a low-bitrate IVAS codec. The values of the decoded spatial metadata 308 can also be quantized to a lower resolution than the values of the original spatial metadata. The feature computation block 500 receives the decoded audio signals 306 and the decoded spatial metadata 308 as inputs. The feature computation block 500 is configured to determine features from the received decoded audio signals 306 and the decoded spatial metadata 308 that can be used as inputs for processing. The features can comprise information that describes physical properties of the sound scene represented by the decoded audio signals 306 and the decoded spatial metadata 308. The features can comprise peak energy levels, evolution of energy levels, repetitions of energy levels, or any other suitable information. In some examples the features can comprise information that describes some properties or features of the audio signals and / or the spatial metadata. Such information provides indirect information relating to the physical properties of the sound scene. Such information could comprise interchannel features or other suitable information. As shown in Fig. 5 the feature computation block 500 receives the decoded audio signals 306 as an input. In other examples the feature computation block 500 could receive partially decoded audio signals or encoded audio signals as an input. Similarly, in Fig. 5 the feature computation block 500 receives the decoded spatial metadata 308 as an input. In other examples, the feature computation block 500 could receive partially decoded spatial metadata or encoded metadata as an input. The feature computation block 500 provides features 502 as an output. The features 502 can comprise features from the decoded audio signals 306 and features the decoded spatial metadata 308. The features 502 are provided as an input to a machine learning model 504. The machine learning model 504 can comprise multiple defined processing steps, and can be similar to the processing instructions related to conventional program code. The difference between conventional program code and the machine learning model 504 is that the instructions of the conventional program code are defined more explicitly at the programming time. The instructions of the machine learning model 504 are defined by combining a set of predefined processing blocks (such as convolutions, data normalizations, other operators), where the weights of the model are unknown at the model definition time. The weights of the machine learning model 504 are optimized by providing the machine learning model 504 with a large amount of input and reference data, and the model weights then converge so that the machine learning model 504 is trained to solve a given task. In this case the task is processing the inputs to generate processed spatial metadata that can provide for improved spatial audio. In examples of the disclosure, when the machine learning model 504 is used, the machine learning model 504 would be fixed and would correspond to a set of processing instructions. The machine learning model 504 is trained to provide predicted metadata properties 506 as an output. The predicted metadata properties 506 can comprise predicted values for parameters of the spatial metadata. The predicted metadata properties 506 can comprise predicted directional properties and / or predicted energy properties. In some examples the predicted metadata properties 506 can comprise information such as MASA directional information (azimuth, elevation, and direct-to-total energy ratio). The predicted metadata properties 506 can be provided in any suitable representation, such as Cartesian “xyz” vector format. The predicted metadata properties 506 are provided as an input to a metadata determination block 508. The metadata determination block 508 also receives the decoded spatial metadata 308 as an input. The metadata determination block 508 is configured to use the predicted metadata properties 506 to process or enhance the decoded spatial metadata 308. The processing of decoded spatial metadata 308 using the predicted metadata properties 506 uses information that was comprised in the decoded audio signals 306 to make the decoded spatial metadata 308 closer to the original spatial metadata that was obtained by the encoder. The metadata determination 508 provides the processed spatial metadata 312 as an output. The processed spatial metadata 312 can have increased time-frequency resolution compared to the decoded spatial metadata 308. The de-quantization accuracy of the parameter values of the processed spatial metadata 312 can be increased compared to the decoded spatial metadata. The increase in de-quantization accuracy provides for improved accuracy of restoration of information from the original spatial metadata that was obtained by the encoder. The processed spatial metadata 312 is provided as the output of the metadata processor 310. In the example of Fig. 5 the feature computation block 500 is shown as a separate block to the machine learning model 504. In some examples the feature computation could be performed by the machine learning model 504. In such cases there would be no separate feature computation block 500 and the machine learning model 504 could receive the decoded audio signals 306 and the decoded spatial metadata 308 as inputs, In some examples the feature computation block 500 can comprise a machine learning algorithm that is different to the machine learning model 504. The machine learning algorithm of the feature computation block 500 could be trained in conjunction with the machine learning model 504 so as to provide appropriate inputs. Other means could be used in other examples. Fig. 6 shows another example metadata processor 310. The example metadata processor 310 of Fig. 6 is similar to the metadata processor 310 of Fig. 5 and corresponding reference numerals are used for corresponding components. The metadata processor 310 of Fig. 6 differs from the metadata processor 310 of Fig. 5 in that it comprises a metadata feature computation block 600 and an audio feature computation block 602 rather than a single feature computation block 500. The metadata feature computation block 600 receives the decoded spatial metadata 308 as inputs. The metadata feature computation block 600 is configured to determine features from the received decoded spatial metadata 308 that can be used as inputs for processing. The metadata feature computation block 600 provides metadata features 604 as an output. The metadata features 604 can comprise features from the decoded spatial metadata 308. The audio feature computation block 602 receives the decoded audio signals 306 as inputs. The audio feature computation block 602 is configured to determine features from the received decoded audio signals 306 that can be used as inputs for processing. The audio feature computation block 602 provides audio features 606 as an output. The audio features 606 can comprise features from the decoded audio signals 306. The metadata features 604 and the audio features 606 can be provided as inputs to the machine learning model 504 which can be arranged to process the metadata features 604 and the audio features 606 to provide predicted metadata properties 506. In the example of Fig. 6 the metadata feature computation block 600 and the audio feature computation block 602 are shown as separate blocks to the machine learning model 504. In some examples the feature computation could be performed by the machine learning model 504. In such cases there would be no separate metadata feature computation block 600 or audio feature computation block 602 and the machine learning model 504 could receive the decoded audio signals 306 and the decoded spatial metadata 308 as inputs. In some examples the metadata feature computation block 600 and the audio feature computation block 602 can comprise a machine learning algorithm that is different to the machine learning model 504. The machine learning algorithms of the metadata feature computation block 600 and the audio feature computation block 602 could be trained in conjunction with the machine learning model 504 so as to provide appropriate inputs. Other means could be used in other examples. In the following example it is assumed that there is one direction active in the MASA metadata. The high-resolution spatial metadata, that is the spatial metadata as it is originally estimated before it is encoded and decoded, is assumed to have 24 bands (all MASA bands) with distinct values and all 4 sub-frames within a frame having the same values. This means that the temporal resolution of the “high-resolution” spatial metadata is lower than what could be supported by the MASA format. The resolution of the spatial metadata 104 is reduced by the metadata encoding as shown in Fig. 2. In this example, the spatial metadata 104 is encoded into 5 parameter bands and 1 sub-frame per frame. Thus, for example, a single parameter set is transmitted and the value is used for all 4 sub-frames. The encoding of the spatial metadata 104 therefore reduces the frequency resolution of the spatial metadata and the value resolution through quantization. This resolution corresponds to the resolution the IVAS codec is using when operating at the 13.2 kbps total bitrate for the MASA content having approximately 2.25 kbps for the MASA spatial metadata. It should be noted that this is only one example. Examples of the disclosure can be used with any number of original and coded frequency bands and temporal subframes, as well as any suitable quantization scheme, as well as with any suitable bit rate. The metadata feature computation block 600 can receive the low resolution decoded spatial metadata 308 as an input. The low resolution decoded spatial metadata 308 can be denoted azi^K, n), ele(K,n), and r(K,n) corresponding to the azimuth angle, elevation angle, and direct-to-total energy ratio in TF(time-frequency)-tile K,n, where k is the parameter band index k = 1, ...,n_bands_meta, and n is the frame index. This spherical representation can be transformed into Cartesian xyz vector representation with v(K,n) = vx(k, n) vy( / c, n) vz(k, n) = t(k, ri) cos(azi(K, n)) cos (ele(K, ri)) sm^azi^K, n)) cos (ele(K, ri)) sin (ele(K, n)) The metadata feature computation block 600 can provide metadata features 604 as an output. The metadata features 604 can be represented with such vectors for each TF-tile. The metadata features 604 can be considered as a 3-dimensional tensor with the shape (n_bands_meta,n_frames,n_features_meta), where n_features_meta = 3 is the number of input features per tile (for example three corresponding to the three elements of the vector v(k, n)), n_frames is the number of spatial metadata frames that are processed at once in implementations of the disclosure. In this example this is n_frames = 200 corresponding to 4 seconds a spatial audio stream with 50 frames per second. n_bands_meta = 5 corresponds to the 5 parameter bands of the low-resolution spatial metadata. Also in this example the decoded audio signals 306 can be assumed to comprise two audio channels. This assumption is valid for this example and audio feature representation but different audio feature representations can be used for the same number of audio channels in the decoded audio signals or different numbers of audio channels in the decoded audio signals 306. The channels of the decoded audio signals 306 can be denoted as xL(t) and xR(t) where xL(t) corresponds to the left channel and xR(t) corresponds to the right channel). The channels xL(t) and xR(t) can be transformed into the frequency-domain. Any suitable means can be used to transform the channels xL(t) and xR(t) to the frequency domains such as a complex-valued low-delay filter bank (CLDFB). In this example the transformation produces 60 complex-valued frequency-domain samples for XL(b,m) and XR(b,m) for each 60 time-domain samples t corresponding to the time range of the slot m and frequency bin b. Other transforms, such as Short-time Fourier transform (STFT), can be used in other examples. The audio feature computation block 602 determines audio features 606. The audio features 606 can comprise information that is potentially useful for enhancing the resolution of the spatial metadata. Different features or types of features can be used in different examples of the disclosure. In some examples the audio features 606 can comprise SPAC-based features that refer to spatial audio capture and Cov-based features that refer to covariance properties. The SPAC-based audio features can be determined with c(n, d, kSPAC) ^stopW bsh^^PAC) j2n.freq^dlyW _ ^-‘m=mstart(n)^,b=bSPAC^kspAC) ' 7 7 (uSPACfi \ \ / i^SPAC / i \ \ Ymstop(n) ^bhigh (kSPAC) . „ z, \|2| (^mstop(n) ^bMgh (.kSPAC) / -r™\|2] A™-™ ^SPAC„ AAL\°>mJ\ ^SPAC„ AAR\.a,mJ\ I m-mstart(n) b=b^c(kSPACy / \ m-MstartW b=b^c(kSPAC) ) where n is the feature output temporal index, for example, corresponding to the frame n, mstart(n) and mstop(n) are the starting and ending time slot indices of feature frame n, d = 1, ...,D with D = 33 is a delay index, kSPAC = 1, ...,KSPAC with KSPAC = 60 is a frequency band index, b^c(kSPAC) and bbk%\kSPAC) are the CLDFB bin limits for frequency band index kSPAC, freq(b) is the center frequency of CLDFB bin b, dly(d) is the delay value corresponding to the delay index d, XR is the complex conjugate transpose of and j = V-l is the imaginary unit. The feature temporal index n can have the same or different spacing as for the spatial metadata and other audio features. The set of delay-values dly(d) can be determined so that they span a reasonable range given the assumed or virtual microphone spacing. For example, for a typical smart phone at a landscape mode, the delays could be equispaced in the range between -0.43 and 0.43 milliseconds, corresponding to the microphone distance of 0.150 m with the speed of sound of 343 m / s. The number of bands k and the band borders b^c(kSPAC) and b^gb (.kSpAc) can be a design decision and may be different from the definitions in MASA format and between different metadata and audio features. In some examples these can be made to approximate the 24 MASA bands. In other examples, including the present example, the input CLDFB bins are not grouped into bands, but each band consists of exactly one CLDFB bin. The feature can be computed over the time range corresponding to one MASA frame. In other words, the number of CLDFB slots in each feature frame is mstop(n) - mstart(n) + 1 = 16. As a result, for the input spatial audio stream of length of 4 seconds, the SPAC feature is of shape (n_bands_SPAC,n_frames_SPAC,n_features_SPAC), where n_features_SPAC = D = 33, n_frames_SPAC = 200, and n_bands_SPAC = KSPAC = 60. In this example the audio features 606 also comprise a second set of features. These are parallel to the SPAC features and are referred to as the Cov features. The Cov features comprise number of related audio features Covi(n,k), where i is the feature index. The first (i = 1) feature is the normalized covariance value between the channels of the decoded audio signals 306 determined with Covr(nsf, k) XL(b,m)X^b,m) _ ™.=mstart(nsf)‘Ljb=bCov('k^ LV f^nstop(nsfj m)|21 \ m=mstart(nsf) ^b^^W L J \ m=mstart(nsf) ^b=bc™^ / N where nsf is the feature output temporal index, for example, corresponding to the subframe, mstart(nsf) and mstop(nsf) are the starting and ending time slot indices of feature (sub-)frame nsf, k = 1,..., KCov with KCov = 24 is a frequency band index, W and bhfghW are CLDFB bin limits for frequency band index k which may correspond to the MASA frequency bands. Both the granularity of the Cov features, that is the number of bands KCov and the temporal spacing mstart(ns / ) -mstart(nsf 1) and temporal range mstop(nsf) - mstart(ns / ) can be different from the SPAC features. This is the case in this example. In other examples different audio features 606 could have the same granularity. If the TF-settings are the same for both the SPAC features and the Cov features then, Cov^n-sf, k) is a subset of c(n, d, kSPAC) for dly(d) = 0. The second (i = 2) feature describes the level difference between the two channels of the decoded audio signals 306 and is computed with Cov2(nsf,k) 1 = -—max kild -7? / LD, min RILD, 101og10 zmstop(nsf) \XR(b,m)\2 rn=mstart(nsf)^b=bcov(ky ' where R[LD = 20 is a channel-level difference normalization factor. The third (i = 3) feature describes the signal energy evolution over time with z x 1 / / E(nsf,k) \\ Cov3[nsf,k) = 1+---max -J?£TO,min REV0,101oglo-^---------—— \ \ £ / N E (n , k) where REV0 = 24 is a normalization factor for the energy evolution, NEV0 = 15 is the length of the energy evolution context, and mstop(ns / ) E(nsf,k) = \XL(b,m)\2 + \XR(b,m)\2 m=mstart(nsf) b=bf™(k) The fourth (i = 4) feature describes the signal energy distribution over bands with / x 2 / / E(nsf,k) Cov4(nsf,k) = 1 + ----max 0,min Rber, 10log10 K ---— rber \ L^E^^k') where Rber = 40 is a normalization factor for the band energy ratio. The fifth (i = 5) feature describes the signal energy over time with frequencydependent normalization with E(nSf, k) / E(nsf,k) \°'3 H ~^sf~^EVO x K 1 Cov5(nsf,k) = All the presented divisions and logarithms may contain numerical regularization limiting the argument to be strictly positive. As a result, for the input spatial audio stream of length of 4 seconds, the Cov feature is of shape (n_bands_cov,n_frames_cov,n_features_cov), where n_features_cov = 5, n_frames_cov = 800, and n_bands_cov = KCov = 24. In this example the frequency bin grouping of the Cov feature is independent of the frequency bin grouping of the SPAC feature. The Cov feature can have a different frequency bin grouping to the SPAC feature. The length of the temporal feature frame for the Cov feature can also be independent of the length of the temporal feature frame for the SPAC feature. In this example the SPAC feature has frame spacing of 20 ms corresponding to one MASA or IVAS frame while the Cov features have frame length of 5 ms corresponding to sub-frame temporal resolution in MASA and IVAS. The machine learning model 504 receives the features as an input. In this example the machine learning model 504 receives the metadata features 604 and the audio features 606. The audio features 606 comprise SPAC features and the Cov features as described. Other types of features can be used in other examples. In this example the machine learning model 504 comprises a deep neural network (DNN). Other types of machine learning model 504 can be used in other examples. The machine learning model 504 receives the metadata features 604 and the audio features 606 and produces a predicted metadata properties 506 as an output. The output can be have a shape of (n_bands_MASA,n_frames,n_features_out'), where n_features_out = 3, n_bands_MASA = 24, and n_fram.es corresponds to the number of MASA frames from the input spatial audio stream. For the example signal of 4 seconds in length, n_frames = 200. The predicted metadata properties 506 that are provided as the output of the machine learning model 504 describe the directional spatial metadata in the Cartesian xyz vector format. The format for the predicted spatial metadata properties 506 can be similar to the format used for the metadata features 604 however the spatial metadata properties 506 can now have a higher frequency resolution and finer quantization values. Different structures can be used for the machine learning model 504 in different examples. Figs. 7A and 7B shows an example structure for a machine learning model 504 that can be used in some examples. Other structures for the machine learning model 504 can be used in other examples. In the example of Fig. 7A the machine learning model 504 comprises a W-Net structure. The structure is referred to a W-net because the path from the input to the output goes through two ll-Net structures. Fig. 7B shows an example ll-Net structure that can be used in the machine learning model 504. The ll-Net structure comprises a downsampling part 720 followed by an upsampling part 724. The downsampling part 720 comprises a sequence of downsampling layers. The respective downsampling layers can comprise convolutional operations such as convolutional neural networks (CNN). The downsampling part 720 can capture high level features from an input. The downsampling layers of the downsampling part 720 can reduce the dimensions of an input along at least some axes. The output of the downsampling part 720 has a smaller number of data elements in at least one axis compared to the original input. The output of the downsampling part 720 is provided as an input to the upsampling part 724. The upsampling part 724 is configured to produce output data. The upsampling part 724 comprises a sequence of upsampling layers. The upsampling part 724 can comprise X upsampling layers where X is also the number of downsampling layers in the downsampling part 720. The upsampling layers can comprise transposed convolutional operations such as transposed CNNs. The upsampling layers of the upsampling part 724 can increase the dimensions of an input along at least some axes. The ll-Net structures in this example also comprise skip connections 722. The skip connections 722 are configured to relay skip connection signals from respective downsampling layers to corresponding upsampling layers. The skip connection signals can reintroduce features from the downsampling part 720 back into corresponding layers of the upsampling part 724. The upsampling layers can comprise operations such as concatenating operations to combine data from a skip connection signal with input data. The upsampling layers can also comprise operations to increase the dimensions of data that is input to the upsampling part 724. The upsampling part 724 provides output data. The output data of the upsampling part 724 has the same number of data elements in at least one dimension as the input data that is originally provided to the downsampling part 720. In the example of Figs. 7 A and 7B the output of the downsampling part 720 is provided as an input to the upsampling part 724. In other examples there can comprise one or more intervening components such as a bottleneck and / or any other suitable operations or combinations of operations. In the example of Fig. 7A the machine learning model 504 comprises three prenetworks 704A, 704B, 704C. Each of the pre-networks 704A, 704B, 704C comprises a U-Net. The U-Net can be as shown in Fig. 7B (some of the reference numbers are omitted in Fig. 7A for clarity) or can have any other suitable arrangement. In this example the machine learning model receives the metadata features 604 and the audio features 606 as an input. The audio features 606 comprise SPAC features 700 and Cov features 702. The respective feature inputs are provided to different prenetworks 704. The metadata features 604 are provided as an input to a first prenetwork 704A, the SPAC features 700 are provided as an input to a second prenetwork 704B, and the Cov features 702 are provided as an input to a third pre-network 704C. The respective pre-networks 704A, 704B, 704C are arranged to process the respective inputs to provide intermediate representations 706A, 706B, 706C as an output. The settings and weights of the respective pre-networks 704A, 704B, 704C are specific for each of the input features. The first pre-network 704A processes the input metadata features 604 to provide a metadata intermediate representation 706A as an output, the second pre-network 704B processes the input SPAC features 700 to provide a SPAC intermediate representation 706B as an output, and the third pre-network 704C processes the input Cov features 702 to provide a Cov intermediate representation 706C as an output. The intermediate representations 706A, 706B, 706C are provided to a concatenation block 708. The concatenation block 708 is configured to combine the intermediate representations 706A, 706B, 706C. The concatenation block 708 can concatenate the intermediate representations 706A, 706B, 706C along the feature axis or perform any other suitable combination. The concatenation block 708 provides a combined intermediate representation 710 as an output. The combined intermediate representation 710 is provided as an input to a combined prediction network 712. The combined prediction network 712 can comprise another ll-Net structure. The weights and settings of the ll-Net structure of the combined prediction network 712 can be different to the weights and settings used for the U-Nets in the pre-networks 704A, 704B, 704C. The combined prediction network 712 provides pre-scale predicted metadata properties 714 as an output. The pre-scale predicted metadata properties 714 are provided as an input to an XYZ scale block 716. The XYZ scale block 716 is configured to apply appropriate scaling to the pre-scale predicted metadata properties 714. The XYZ scale block 716 provides predicted spatial metadata properties 506 as an output. Fig. 8 shows an example structure of a pre-network 704. The example of Fig. 8 shows a structure for a metadata features pre-network 704A. Corresponding structures can be used for the SPAC features pre-network 704B, the Cov features pre-network 704C, and the combined prediction network 712. The input in this case is the metadata features 604. This input has shape (n_bands_meta,n_frames,n_features_meta). The input is provided to a dimension adjustment 800. The dimension adjustment 800 comprises a linear layer 802. The linear layer 802 can be fully connected. The linear layer 802 can operate on the input dimension that corresponds to the frequency bands. In this case this is the first dimension of the input. In this description a single input is processed. In examples of the disclosure multiple inputs can be processed in parallel as a batch. In such examples the stacking of multiple inputs adds one dimension in front of the actual data dimensions. The operation of the linear layer 802 provides an intermediate feature tensor Y 804 as an output. The metadata features pre-network 704A comprises multiple residual blocks (ResBlock) 806. The ResBlocks 806 comprise a stack of layers that is arranged so that the output of a given layer is taken and added to a subsequent layer deeper within the ResBlock 806. An example ResBlock 806 is shown in Fig. 9. Each ResBlock 806 has the stride settings (x,y) given within the block for example (str=1, 1). The first number corresponds to the stride along the frequency axis (first dimension) and the second number corresponds to the stride along the temporal axis (second dimension). The number of output channels (third dimension) from each ResBlock 806 is given as the last number in the triplet following the respective ResBlocks 806 in Fig. 8. The metadata features pre-network 704A can comprise a sequence of downwards ResBlocks 806A-806D and a sequence of further ResBlocks 806E-806H. The sequence of downwards ResBlocks 806A-806D can provide a downsampling part of the metadata features pre-network 704A. The sequence of upwards ResBlocks 806AE-806H can provide a upsampling part of the metadata features pre-network 704A. In the example of Fig. 8 the first ResBlock 806A has (str=1, 1) with 8 output channels. The second ResBlock 806B has (str=2, 2) with 16 output channels. The second ResBlock 806B is arranged to perform factor 2 sub-sampling in the spatial (time and frequency) dimensions. The third ResBlock 806C has (str=2, 2) with 32 output channels. The third ResBlock 806C is arranged to perform factor 2 sub-sampling in the spatial dimensions. The fourth ResBlock 806D has (str=2, 2) with 64 output channels. The fourth ResBlock 806D is arranged to perform factor 2 sub-sampling in the spatial dimensions. The output of the last downwards ResBlock 806D is passed through a convolutional block (ConvBlock) 808. The ConvBlock 808 comprises a two-dimensional convolution with kernel size of (3, 1) for (frequency, time) with 128 output channels, followed by a Batch Normalization (BatchNorm) and a Rectified Linear Unit (ReLU) activation. The output of the ConvBlock 808 is passed through Transposed Convolution (TransConv) block 810. The TransConv block 810 comprises a two-dimensional transposed convolution with kernel size of (3, 1) and 64 output channels. This effectively up-samples along the frequency axis. This is followed by a BatchNorm and ReLU activation. The output of the TransConv block 810 is concatenated with the output of the last downward ResBlock 806D. The solid arrows in Fig. 8 represent concatenations. The concatenation of the output of the TransConv 810 with the output of the last downward ResBlock 806D is performed along the channel dimension. The concatenation produces a tensor with shape (3, N / 8,128). This tensor is provided as an input to the upsampling part of the metadata features pre-network 704A. The upsampling part comprises a fifth ResBlock 806E followed by an Upsampling (2, 2) block 812E. The Upsampling block 812E is arranged to upsample the two-dimensional layer with the spatial size scaler factor given in the parenthesis (frequency, time) using, for example, the nearest neighbor upsampling method. Therefore Upsampling (2,2) increases the size of both spatial dimensions by a factor of two using nearest neighbor upsampling. The output of the Upsampling is the output of this layer of the upsampling part of the metadata features pre-network 704A. The output of the Upsampling block 812E is concatenated with a corresponding matching skip connection tensor. This comprises data from the downsampling layer. The concatenation is performed along the feature dimension (third dimension in this example). The output of the concatenation is provided to a sixth ResBlock 806F and following Upsampling block 812F. The output of the Upsampling block 812F is concatenated with a corresponding matching skip connection tensor. The concatenation is performed along the feature dimension (third dimension in this example). The output of this concatenation is provided to a seventh ResBlock 806G and following Upsampling block 812G. The convolutions in the upsampling part of the metadata features pre-network 704A have the number of output channels of 32, 16, and 8. The last layer of the upsampling part comprises an eight ResBlock 806H. the last ResBlock 806H has five output channels and no following Upsampling blocks. The shape of the last ResBlock 806H provides the output 706A of the metadata features pre-network 704A. The output 706A of the metadata features pre-network 704A is now (24, n_frames, 5). In this example the output 706A of the metadata features prenetwork 704A is the intermediate metadata features 706A. Other U-Net structures can be arranged in a similar manner but would have different inputs and outputs. Fig. 9 shows an example structure for a ResBlock 806. The ResBlock 806 could be used in the U-Net structures of the machine learning model 504. The example ResBlock 806 could be used in a downsampling part. A ResBlock 806 with the same internal structure could also be used in the upsampling part but different stride and kernel size settings could be used in the upsampling part. The internal structure of the ResBlock 806 has a pre-activation ordering. This means that Batch Norm layers 904 and the ReLU activation layers 906 are before the convolution operations 908. The ResBlock 806 comprises two paths for an input. The first path is shown on the left side of Fig. 9 and the second path is shown on the right side of Fig. 9. The path on the left side corresponds to the residual or by-pass path. The path on the right side corresponds to the convolutional core. The residual or by-pass path comprises convolution operations 900 and a BatchNorm layer 902. The convolution operations 900 comprises a Conv2D layer with the kernel size of (1, 1) and striding (x, y) matching the stride settings of the ResBlock 806. The convolution operations 900 adjust the size of the feature dimension of the input to match the feature dimension of the last convolution in the convolution core. The convolution operations 900 are followed by the BatchNorm layer 902. The BatchNorm layer 902 comprises a BatchNorm2D layer that operates on the feature dimension. The convolutional core consists of a sequence of blocks. The sequence comprises a BatchNorm layer 904 followed by a ReLU activation layers 906 which is then followed by convolution operations 908. In this example the first BatchNorm layer 904A is a BatchNorm2D that operates on the feature dimension (third dimension). The first ReLU activation layer 906A is applied one each element from the BatchNorm layer 904A. The output of the ReLU activation layer 906A is passed to the first convolution operations 908A. The first convolutions operations 908A comprise a Conv2D with the kernel size of (3, 3) and the stride setting (x, y) matching the stride settings of the ResBlock 806. The output of the first convolution operations 908A is passed through a second BatchNorm layer 904B and a second ReLU activation layer 906B before being passed to the second convolution operations 908B. The second convolution operations 908B processes the input with a Conv2D with kernel size of (3, 3) and stride of 1. In this example both of the convolution operations 908 use reflection padding of size 1 along all four sides of the input. The output of the residual or by-pass path and the output of the convolutional core are provided to an addition block 910. The addition block 910 adds the output of the residual or by-pass path and the output of the convolutional core in an element wise manner. The output of the addition block 910 is the output of the ResBlock 806. The SPAC feature pre-network 704B can have a similar structure to the metadata features pre-network 704A as shown in Fig. 8 however different dimensions and settings could be used. The input to the SPAC feature pre-network 704B would be the SPAC features 702. This input has shape (60, n_fram.es, 33). This input can be provided to a dimension adjustment 800 as shown in Fig. 8. The dimension adjustment 800 can also comprise a fully connected linear layer 802. The dimension adjustment provides an intermediate feature tensor Y with shape (48, n_frames, 33) . In the example SPAC feature pre-network 704B the downsampling part also comprises four ResBlocks 806A-806D. The first ResBlock 806A has (str=2, 2) with 64 output channels. The second ResBlock 806B has (str=2, 2) with 128 output channels. The third ResBlock 806C has (str=2, 2) with 256 output channels. The fourth ResBlock 806D has (str=2, 2) with 512 output channels. In this example there is no downsampling along the temporal axis in the last downsampling layer. In the example SPAC feature pre-network 704B the output of the fourth ResBlock 806D is passed through a convolutional block (ConvBlock) 808. The ConvBlock 808 comprises a two-dimensional convolution with kernel size of (3,1) for (frequency, time) with 1024 output channels, followed by a TransConv block 810. The TransConv Block 810 has a corresponding kernel size of (3, 1) and 256 output channels. In the example SPAC feature pre-network 704B the upsampling part also comprises a further four ResBlocks 806E-806H and corresponding Upsampling blocks 812E-812H. In this example the fifth ResBlock 806E has (str=1, 1) with 256 output channels followed by an Upsampling (2, 1) block 812E. The Upsampling block 812E does not provide any Upsampling in the temporal axis. The sixth ResBlock 806F has (str=1, 1) with 128 output channels followed by an Upsampling (2, 2) block 812F. The seventh ResBlock 806G has (str=1, 1) with 64 output channels followed by an Upsampling (2, 2) block 812G. The eighth ResBlock 806H has (str=1, 1) with 20 output channels followed by an Upsampling (2, 1) block 812H. The Upsampling block 812H does not provide any Upsampling in the temporal axis. This differs from the upsampling part of the metadata features pre-network 704A because the metadata features pre-network 704A does not comprise an Upsampling block after the last ResBlock 806 in the upsampling part. The output of the SPAC feature pre-network 704B has shape (24, n_frames, 20). The Cov feature pre-network 704C can have a similar structure to the metadata features pre-network 704A as shown in Fig. 8 and also the SPAC feature pre-network 704B, however different dimensions and settings could be used. For the case of the Cov feature pre-network 704C the dimension adjustment block 800 can be omitted and the input to the u-net structure would be an input of shape (24, n_subframes, 5). In the example Cov feature pre-network 704C the downsampling part also comprises four ResBlocks 806A-806D. The first ResBlock 806A has (str=1, 2) with 8 output channels. In this case there would be no downsampling along the frequency axis. The second ResBlock 806B has (str=2, 2) with 16 output channels. The third ResBlock 806C has (str=2, 2) with 32 output channels. The fourth ResBlock 806D has (str=2, 2) with 64 output channels. In the example SPAC feature pre-network 704B the output of the fourth ResBlock 806D is passed through a ConvBlock 808. The ConvBlock 808 comprises a two-dimensional convolution with kernel size of (3, 1) with 128 output channels, followed by a TransConv block 810. The TransConv Block 810 has a corresponding kernel size of (3, 1) and 64 output channels. In the example Cov feature pre-network 704C the upsampling portion also comprises a further four ResBlocks 806E-806H and corresponding Upsampling blocks 812E-812G. In this example the fifth ResBlock 806E has (str=1, 1) with 32 output channels followed by an Upsampling (2, 2) block 812E. The sixth ResBlock 806F has (str=1, 1) with 16 output channels followed by an Upsampling (2, 2) block 812F. The seventh ResBlock 806G has (str=1, 1) with 8 output channels followed by an Upsampling (2, 2) block 812G. The eighth ResBlock 806H has (str= 1,2) with 8 output channels. There is no Upsampling block following the eighth ResBlock 806H in the Cov feature prenetwork 704C. This is similar to the metadata features pre-network 704A which also does not comprise an Upsampling block after the last ResBlock 806 in the upsampling part. The Cov feature pre-network 704C differs from the metadata features pre-network 704A (and also the SPAC features pre-network 704B) in that, in the Cov feature prenetwork 704C the last ResBlock 806H has downsampling along the temporal axis with the stride (1, 2) setting. The output of the Cov feature pre-network 704C has shape (24,n_subframes / 4,8), which is equal to (24, n_frames, 8). Referring to Fig. 7 metadata feature pre-network 704A provides a metadata intermediate representation 706A as an output, the SPAC feature pre-network 704B provides a SPAC intermediate representation 706B as an output, and the Cov feature pre-network 704C provides a Cov intermediate representation 706C as an output. Using the example U-Net structures of Fig. 8 the metadata intermediate representation 706A has a shape (24, n_frames, 5), the SPAC intermediate representation 706B has shape (24, n_frames, 20), and the Cov intermediate representation 706C has shape (24, n_frames, 8). The intermediate representations 706A, 706B, 706C are provided to a concatenation block 708. The concatenation block 708 concatenates the intermediate representations 706A, 706B, 706C along the feature axis (the third dimension in Fig. 7). The concatenation block 708 provides a combined intermediate representation 710 as an output. The combined intermediate representation 710 has a shape (24, n_frames, 33). The combined intermediate representation 710 is provided as an input to a fourth U-Net structure. The fourth U-Net structure is the combined prediction network 712. The combined prediction network 712 has a similar structure to the metadata features prenetwork 704A as shown in Fig. 8 with some differences. For the case of the combined prediction network 712 the dimension adjustment block 800 can be omitted and the input to the u-net structure would be an input of shape (24,n_frames, 33). In the example combined prediction network 712 the downsampling portion also comprises four ResBlocks 806A-806D. The first ResBlock 806A has (str=1, 1) with 32 output channels. The second ResBlock 806B has (str=2, 2) with 48 output channels. The third ResBlock 806C has (str=2, 2) with 64 output channels. The fourth ResBlock 806D has (str=2, 2) with 128 output channels. In the example combined prediction network 712 the output of the fourth ResBlock 806D is passed through a ConvBlock 808. The ConvBlock 808 comprises a two dimensional convolution with kernel size of (3, 1) and 512 output channels, followed by a TransConv block 810 with a corresponding kernel size of (3, 1) and 256 output channels. In the example combined prediction network 712 the upsampling part also comprises a further four ResBlocks 806E-806H and corresponding Upsampling blocks 812E-812G. In this example the fifth ResBlock 806E has (str=1, 1) with 64 output channels followed by an Upsampling (2, 2) block 812E. The sixth ResBlock 806F has (str=1, 1) with 48 output channels followed by an Upsampling (2, 2) block 812F. The seventh ResBlock 806G has (str=1, 1) with 12 output channels followed by an Upsampling (2, 2) block 812G. The eighth ResBlock 806H has(str=1,1) with 3 output channels. There is no Upsampling block following the eighth ResBlock 806H in the combined prediction network 712. This is similar to the metadata features pre-network 704A which also does not comprise an Upsampling block after the last ResBlock 806 in the upsampling part. The output of the combined prediction network 712 has shape (24, n_frames, 3). The output of the combined prediction network 712 provides pre-scale predicted metadata properties 714. These represent the directional metadata in each (24, n_frames) TF-tiles in xyz vector representation vprescaie(k,n). The pre-scale predicted metadata properties 714 are provided as an input to an XYZ scale block 716. The XYZ scale block 716 is configured to apply appropriate scaling to the pre-scale predicted metadata properties 714. The XYZ scale block 716 can apply final constraints on the length of the vector. In some examples the scaling applied by the XYZ scale block 716 can comprise determining the length of the input vectors. The length of the input vectors can be determined with Vprescale,n) + ^prescale,y T" ^pr escale, This is passed through hyperbolic tangent and scaled with a constant, for example, cxyz = 1-1 for obtaining the scaled length ^scaled (k,n) = cxyztanh(r^ Without the scaling, the output of the hyperbolic tangent would require infinite value of the input for the output to reach value of 1.0. When the model is used for inference the value of rscaied(k,n) is limited to the range of 0...1.0 with rscaiedtk.n) = max (o,min(l, rscaled(k,n)^ and this value is used in place of rscaied(k, n). A scaling value s(k,n) can be determined from these two lengths with s(k,n) ^scaled ^k, 7l) rin(k, n) The output of the XYZ scale block 716 is the input multiplied by this scaling value ^pred$(k> prescale(Kt The XYZ scale block 716 provides predicted spatial metadata properties 506 as an output which is the output of the machine learning model 504. The machine learning model 504 can be trained using any suitable process. In some examples the machine learning model 504 can be trained using a set of spatial audio items in MASA format comprising an audio signal such as a transport audio signal and MASA spatial metadata with one directional field. The audio signal can comprise two channels. It is assumed that the original spatial metadata has low temporal resolution and all four sub-frames in each frame contain the same values. The training data comprises 4530 items that are 4 seconds in length (that is, 200 frames). The directional spatial metadata consisting of azimuth azi(k,n), elevation ele(k,n), and direct-to-total energy ratio r(k,ri) are transformed into Cartesian xyz vector representation using ^target^> Vtarget,x(.^> ^-) ^target,y ytarget,z^k, n). = r(k, n) cos(azi(k,n)) cos (ele(k, n)) sin(azi(k, n)) cos (ele(k, n)) sin (ele(k, n)) This is the reference or target data during the training. The validation data of 799 items is selected from this same pool of items. For the low-resolution model input, the training items are passed through an IVAS codec (encoder and decoder) operating at 13.2 kbps total bitrate (having about 2.25 kbps for the MASA metadata), and the decoded MASA metadata is obtained using the external renderer (EXT) output mode of the IVAS decoder. The IVAS codec reduces the frequency resolution of the spatial metadata into 5 bands from the original 24 bands and applies quantization to the values. This metadata is transformed into the Cartesian xyz vector representation in the same way as the reference data. The training uses batch size of 32, that is, the parameters of the model are adjusted after each 32 training examples. AdaDelta optimizer is used with learning rate of 1.0. The training is run for the maximum of 1000 epochs or until early stopping is triggered. The early stopping is triggered when the per-epoch validation loss is not lower than the best per-epoch validation loss in 50 consecutive epochs. The validation batch size is 8 items. The loss can be computed as follows. First, the direct-to-total energy ratios of the target and the predicted data are computed ^"targetO^'^) ^target, x&' + ^target,y + ^target,z^’^) ^pred (k, ri) (Ji, n) + ^pred:y (k,n) + v^redz(k,n) Then, the absolute value of the energy ratio difference is computed rabsdiff(fc>n) pred (ri, n) ^target(^> w) | Then, unit-length direction vectors are computed for the target and the predicted data ^target,unitlen(^> Vpred,unitlen(k / ri) ^target(^> ^target (^> ri) ^predO^i ri) Tpred (k> ri) Then, direction error vector is computed by Verror (^, n) pred,unit ten (k, n) — v target.unitlen (k, n) and the length of the direction error vector is computed by error ^error,x (&’ ri) + Vgrrory (k, n) + Vgrrorz (k,n) This length is weighted by the target direct-to-total energy ratio I err or,w (^-> ri) (error (^>'^)^target((^>ri) Using the determined absolute value of the energy ratio difference and the determined weighted length of the direction error vector, the combined error measure is determined by (comb (ri, n) — l^absd iff ri) Ierror,w (ri, 7l) Then, an energy-weighting metric is determined. The energy-weighting metric is for weighting the loss based on the energies of the corresponding time-frequency tiles (that is, time-frequency tiles having larger energy should have a larger effect on the loss). The energy-weighting metric can be formulated in any suitable manner. In some examples, the signal energy evolution over time with frequency dependent weighting (i.e., Cov5(nsf,k)) as computed by the audio feature computation block 602 may be used as the weight. As Cov5(nspfc) is computed in subframes nsf, and the loss is computed in frames n, the mean of the values Cov5(nsf, k) for the corresponding frame n are computed and set as the weight where n is the frame index, nsf the subframe index, and Nsf = 4 the number of subframes in a frame. Then, the energy-weighted combined error measure is determined by %comb,w ^comb ^-) Eloss, Using ^comtliW(k,n), the loss is determined by computing the mean over time and frequency over the entire training example ^comb,w j ’ ^comb,w (k> k n which is the loss that is output from the loss function. In the example metadata processors 310 such as those shown in Figs. 5 and 6 the metadata determination block 508 receives the decoded spatial metadata 308 and predicted metadata properties 506 as an input. The decoded spatial metadata 308 can comprise all MASA metadata variables. The metadata determination block 508 is arranged to enhance the decoded spatial metadata 308 based on the predicted metadata properties 506 by updating the direction (azimuth and elevation) and the direct-to-total energy ratio based on the predicted metadata properties 506 that are generated by the machine learning model 504. In some examples the process of enhancing the decoded spatial metadata 308 can comprise transforming the Cartesian xyz vector representation of the model prediction vpred(k,n) into azimuth azidec(k,n), elevation eledec(k,n), and direct-to-total energy ratio rdec(k, n) representation with azidec(k,n) = atan2 (vprediy(k,n),vpred:X(k,n eledec(k,n) = atan2 [vpredz(k,n), v^redx(k,n) + v^redy(k,n) rdec(k,n) = vpredx(k,n) + v*redy(k,n) + v*redz(k,n) Here atan2Q) is the inverse tangent function resolving the correct quadrant. The perframe temporal resolution of predicted metadata properties 506 is adjusted to the per-sub-frame temporal resolution of MASA metadata by repeating the same value for all four sub-frames in a frame. The diffuse-to-total energy ratio rdiff(k,n) in the spatial metadata is adjusted to reflect the new direct-to-total energy ratio by rdiff(k,n) = 1 - rdec(k,n) In this example, the surround and spread coherence parameter values are not adjusted, but the values from the low-resolution decoded spatial metadata 308 are used instead, by replicating the 5 values to cover the all 24 bands. In some other examples the surround and / or spread coherence parameter values can be adjusted. In some examples all parameter values can be obtained from the predicted metadata properties 506. In these examples the metadata determination block 508 does not need to have the decoded spatial metadata 308 as an input because all spatial metadata values may be obtained from the predicted metadata properties 506. The architecture described above for the machine learning model 504 uses convolutions in time and temporal downsampling and upsampling operations. This means that the architecture is not causal and thus not suitable for low-latency applications. However, this example is only used for demonstrating that the assumptions of additional information in an audio signal are valid. It is possible to implement the machine learning model using other DNN architectures, some of which are causal (not utilizing information from the future) and low-latency. One way to define a causal architecture is to restrict the downsampling and upsampling operations along the time axis to consider only the current and past values instead of considering also future indices, and aligning the by-pass paths between the blocks of downsampling and upsampling parts along the time axis. This alignment can be done, for example, by applying the padding only in the direction of past samples on the time axis, and aligning the sample corresponding to the most recent time index after temporal upsampling with the most recent time index at the corresponding by-pass path. Alternatively, the strides in the convolution operations along the time axis may be replaced using dilated convolutions, or by implementing the temporal modelling using recurrent modelling. A further alternative is to omit temporal context modelling entirely by using convolution kernel size of 1, stride length of 1, and upsampling factor of 1 on the time axis instead of the values provided in the example embodiment. The examples of the disclosure provide processed spatial metadata as an output. The processed spatial metadata can be used for spatial synthesis. Any suitable methods can be used for the spatial synthesis. Fig. 10 shows another example system 100 that can be used in examples of the disclosure. In this example the metadata processor 310 is provided as a post processing block that is outside of, or separate to, the decoder 110. This arrangement can enable examples of the disclosure to be used with existing codecs without having to make any changes within the codec. The system 100 comprises an encoder 106 that receives audio signals 102 and spatial metadata 104 as an input. The encoder 106 provides a bitstream 108 that is sent from the encoder 106 to the decoder 110. The decoder 110 is arranged to decode the bitstream 108 to provide decoded audio signals 306 and decoded spatial metadata 308 as outputs. The decoded audio signals 306 and decoded spatial metadata 308 are provided as inputs to the metadata processor 310. The metadata processor 310 can comprise a machine learning model 504 as described herein or any other suitable means. The metadata processor 310 provides processed spatial metadata 312 as an output. The processed decoded spatial metadata 312 and the decoded audio signals 306 are provided as an input to a spatial synthesizer 314. The spatial synthesizer 314 is configured to process the processed decoded spatial metadata 312 and the decoded audio signals 306 to render spatial audio output 316. Other variations to the system 100 and components of the system 100 can be made in examples of the disclosure. For instance, in the examples the described the inputs to the encoder 106 comprise audio signals 102 and spatial metadata 104 but other formats for the inputs could be used. In some examples the audio signals 102 and spatial metadata 104 could be generated inside the encoder 106. Similarly, different types of inputs could be provided to the metadata processor 310. For example, any associated audio signals and spatial metadata could be used and not necessarily decoded signals. In the described examples the directional parameters of the spatial metadata are given as azimuth and elevation values. In other examples, the directional parameters can be received in other formats, such as spherical indices, and can be converted to and from azimuths and elevations and / or other suitable representations where appropriate. In the described examples, one set of audio-based features were used. Some of them may be more useful with spaced microphones (such as the SPAC feature), while some of them may be more useful with directional coincident microphones (such as the inter channel level difference feature). These are merely example features, and in other examples, other kinds of audio-based features can be used additionally or instead of the described features. Figs. 11A to 11D show example results that can be obtained using examples of the disclosure. The results validate the performance of the examples of the disclosure. The experiments that were used to obtain the data shown in Figs. 11A to 11D used spatial metadata as described above and the original audio signals was used for the audio-based features. That is, the spatial metadata was passed through an IVAS codec for reducing the resolution, but the audio signals from the IVAS codec were replaced with the ideal transport audio signals. The performance of the system 100 was measured by evaluating the loss function for the low-resolution model input and then for the processed spatial metadata 512 from the output of the metadata processor 310. Fig. 11A shows the loss function of the machine learning model 504 using metadata features 604 and audio features 606. The first horizontal line 1100 shows the loss function evaluated on the decoded spatial metadata 308 from which the machine learning model 504 input is determined. The second horizontal line 1102 shows the loss function evaluated on the machine learning model 504 for validation data. The third line 1104 shows the loss function for the machine learning model 504 output for the training data and the fourth line 1106 shows the loss function for the machine learning model 504 for the validation data. The data in Fig. 11A shows that even though the training loss keeps on decreasing, the validation loss has converged, and training has been ended with early exit. Both training and validation data exhibit a clearly lower loss value for the machine learning model 504 output compared to the machine learning model 504 input indicating that the processing has enhanced the spatial metadata. The data in Fig. 11A shows that the examples of the disclosure are able to enhance the low-resolution spatial metadata with the help of an audio signal such as the transport audio signal as additional information source. This means that, the processed spatial metadata 312 obtained using examples of the disclosure is closer to the original high-resolution spatial metadata when the distance is measured with the loss function. In addition to this computational evaluation that provided the results shown in Figs. 11A to 11D, an informal subjective evaluation was done. The informal subjective evaluation compared the binaural rendering based on the ideal transport audio signals using the original high-resolution metadata, the low-resolution metadata from the model input, and the metadata obtained by applying the example methods described herein. It was noticed that the enhanced version (that is using the processed spatial metadata obtained as described herein) was perceptually spatially closer to the reference than the low-resolution input version. The directions of the sound sources were more correct, the sound source directions were more stable, and the spaciousness was closer to the reference. The subjective evaluation indicates that when using the processed spatial metadata 312 in binaural rendering, the reproduced audio scene is perceptually closer to the original audio scene than when using the model input for the rendering. Figs. 11B and 11C show data obtained in a second experiment. In the second experiment the performance of the machine learning model 504 as described was compared with partial models. The results in Fig. 11B were obtained from a comparison using only spatial metadata features and the results in Fig.11C were obtained from a comparison using only audio signal-based features. In the comparison using only spatial metadata features the audio-based feature inputs (such as the SPAC features and the Cov features) are not processed. Only the metadata features pre-network 704A is active. In this case the combined intermediate representation 710 contains only the intermediate metadata features 706A and the number of input channels to the combined prediction network 712 is adjusted accordingly. In Fig. 11B The first horizontal line 1110 shows the loss function evaluated on the machine learning model 504 for training data. The second horizontal line 1112 shows the loss function evaluated on the machine learning model 504 for validation data. The third line 1114 shows the loss function for the machine learning model 504 output for the training data and the fourth line 1116 shows the loss function for the machine learning model 504 for the validation data. In this example, even though the training loss keeps on decreasing, the validation loss has converged, and training has been ended with early exit. Both the training and validation data exhibit clearly lower loss value at the machine learning model 504 output compared to the machine learning model 504 input meaning that the machine learning model 504 has been able to enhance the metadata. However, the value of the loss shown in Fig.11B remains above the values shown in Fig.11A indicating that the machine learning model 504 has been able to obtain useful information from the audio features. In the comparison using only audio-based features the metadata feature inputs. Only the SPAC features pre-network 704B and the Cov features pre-network 704C are active. In this case the combined intermediate representation 710 comprises the SPAC intermediate representation 706B and the Cov intermediate representation 706C. The number of input channels to the combined prediction network 712 is adjusted accordingly. This example that uses only audio-based features can be considered an example of a machine learning based metadata estimation from two-channel transport audio signals. In Fig. 11C The first horizontal line 1120 shows the loss function evaluated on the machine learning model 504 for training data. The second horizontal line 1122 shows the loss function evaluated on the machine learning model 504 for validation data. The third line 1124 shows the loss function for the machine learning model 504 output for the training data and the fourth line 1126 shows the loss function for the machine learning model 504 for the validation data. In this example, even though the training loss keeps on decreasing, the validation loss has converged, and training has been ended with early exit. Both the training and validation data exhibit clearly lower loss value at the machine learning model 504 output compared to the machine learning model 504 input meaning that the machine learning model 504 has been able to enhance the metadata. However, the value of the loss shown in Fig.11C remains above the values shown in Fig.11A indicating that the machine learning model 504 has been able to obtain useful information from having both the metadata features and the audio features. The results shown in Figs.HB and 11C indicate that using the metadata processor 310 as described but with just metadata features or with just audio features will still provide improved spatial metadata compared to the low-resolution input. However, the use of both metadata features 604 and audio feature 606 provides even better performance. This validates the assumption that there is useful information present in audio signals such as transport audio signals and this can be used for enhancing the spatial metadata. The same conclusion was done based on the informal subjective evaluation with binaural audio signals. A further experiment was performed in which the ideal transport audio signals were replaced with transport audio signals that have been encoded using IVAS codec in stereo mode at 64 kbps. This uses the modified discrete cosine transform (MDCT) stereo codec. The intention of this experiment was to use non-ideal transport audio signals so that the information from the transport audio signals is closer to the ones used in real application. In Fig. 11D The first horizontal line 1130 shows the loss function evaluated on the machine learning model 504 for training data. The second horizontal line 1132 shows the loss function evaluated on the machine learning model 504 for validation data. The third line 1134 shows the loss function for the machine learning model 504 output for the training data and the fourth line 1136 shows the loss function for the machine learning model 504 for the validation data. Fig. 11D shows that even though the training loss keeps on decreasing, the validation loss has converged, and training has been ended with early exit. Both the training and validation data exhibit clearly lower loss value at the machine learning model 504 output compared to the machine learning model 504 input indicating that the machine learning model 504 has been able to enhance the metadata even with the coded transport audio. The loss reaches lower values than when using only metadata in as shown in Fig. 11B. This indicates that there is useful information present even in the coded transport audio signals even though the performance of uncoded transport audio signals as shown in Fig. 11A is not reached. Fig. 12 shows an example device 1200 that can be used to implement examples of the disclosure. In this example the device 1200 is a mobile device with connected headphone 1218. The headphone 1218 can be connected to the device 1200 via a wired connection or a wireless connection 1214. Other types of devices 1200 and connected peripheral devices can be used in other examples of the disclosure. The device comprises a processor 1206 and a memory 1210. The memory 1210 can comprise program code 1212. The memory 1210 can comprise a machine learning model 504 which can be arranged as described herein to generate processed spatial metadata. The processor 1206 and memory 1210 can be arranged to provide a decoder 110 as described above. The user 1220 of the device 1200 can be using the device to participate in an immersive call or other communication that uses spatial audio. The device 1200 obtains a bitstream 108. The bitstream 108 can comprise the bitstream 108 that is transferred between an encoder 106 and a decoder 108 as shown in Fig.1. The bitstream 108 can be received via a transceiver 1204 or can be retrieved from storage 1202. In some examples the bitstream 108 can be received via the transceiver 1204 and then stored to the storage and then accessed later by the processor 1206. The processor 1206 is arranged to convert the bitstream 108 to a spatial audio output 1208. The processor 1206 uses the program code 1212 and the machine learning model 504 that are stored in the memory 1210 to convert the bitstream 108 to the spatial audio output 1208. In this example the spatial audio output 1208 is a binaural output. Other types of spatial audio output 1208 can be used in other examples. The spatial audio output 1208 is provided to a headphone connection 1214. The headphone connection 1214 connects the headphone 1218 to the device 1200 and enables signals from the device 1200 to be provided to the headphone 1218. The headphone connection 1214 can be a wired connection or a wireless connection. The headphone connection 1214 provides an audio signal 1216 to the headphones to enable the spatial audio to be played back to the user 1220. In the example of Fig. 12 the machine learning model 504 is stored in the memory 1210. The machine learning model 504 can be trained by an external device and then provided to the device 1200 so that it can be stored in the memory 1210. The external device that performs the training of the machine learning model 504 can have a higher processing capacity than the device 1200 shown in Fig. 12 or other devices that implement the spatial audio communications. For example, the external device that performs the training of the machine learning model 504 could comprise a workstation with multiple graphic processing units (GPUs) dedicated to the training of the machine learning model 504. Fig. 13 schematically illustrates an apparatus 1300 that can be used to implement examples of the disclosure. In this example the apparatus 1300 comprises a controller 1302. The controller 1302 can be a chip or a chip-set. In some examples the controller 1302 can be provided within a communications device such as telephone or teleconferencing device or any other suitable type of device that enables audio signals to be communicated. In the example of Fig. 13 the implementation of the controller 1302 can be as controller circuitry. In some examples the controller 1302 can be implemented in hardware alone, have certain aspects in software including firmware alone or can be a combination of hardware and software (including firmware). As illustrated in Fig. 13 the controller 1302 can be implemented using instructions that enable hardware functionality, for example, by using executable instructions of a computer program 1304 in a general-purpose or special-purpose processor 1206 that can be stored on a computer readable storage medium (disk, memory etc.) to be executed by such a processor 1206. The processor 1206 is configured to read from and write to the memory 1210. The processor 1206 can also comprise an output interface via which data and / or commands are output by the processor 1206 and an input interface via which data and / or commands are input to the processor 1206. The memory 1210 is configured to store a computer program 1304 comprising computer program instructions (computer program code 1212) that controls the operation of the controller 1302 when loaded into the processor 1206. The computer program instructions, of the computer program 1304, provide the logic and routines that enables the controller 1302 to perform the methods illustrated in the Figs. The processor 1206 by reading the memory 1210 is able to load and execute the computer program 1304. The apparatus 1300 therefore comprises: at least one processor 1206; and at least one memory 1210 including computer program code 1304, the at least one memory 1210 and the computer program code 1212 configured to, with the at least one processor 1206, cause the apparatus 1300 at least to perform: receiving 400 an encoded audio signal and corresponding encoded spatial metadata; determining 402 at least one input for processing wherein the input is based on the encoded audio signal and the encoded spatial metadata; processing 404 the at least one input to use information from the encoded audio signal and the encoded spatial metadata to generate processed spatial metadata; and enabling 406 spatial rendering to be performed on one or more audio signals using the processed spatial metadata. As illustrated in Fig. 13 the computer program 1304 can arrive at the controller 1300 via any suitable delivery mechanism 1306. The delivery mechanism 1306 can be, for example, a machine readable medium, a computer-readable medium, a non-transitory computer-readable storage medium, a computer program product, a memory device, a record medium such as a Compact Disc Read-Only Memory (CD-ROM) or a Digital Versatile Disc (DVD) or a solid state memory, an article of manufacture that comprises or tangibly embodies the computer program 1304. The delivery mechanism 1306 can be a signal configured to reliably transfer the computer program 1304. The controller 1302 can propagate or transmit the computer program 1304 as a computer data signal. In some examples the computer program 1304 can be transmitted to the controller 1302 using a wireless protocol such as Bluetooth, Bluetooth Low Energy, Bluetooth Smart, 6LoWPan (IPv6 over low power personal area networks) ZigBee, ANT+, near field communication (NFC), Radio frequency identification, wireless local area network (wireless LAN) or any other suitable protocol. The computer program 1304 comprises computer program instructions that when executed by an apparatus 1300 cause the apparatus 1300 to perform at least the following: receiving 400 an encoded audio signal and corresponding encoded spatial metadata; determining 402 at least one input for processing wherein the input is based on the encoded audio signal and the encoded spatial metadata; processing 404 the at least one input to use information from the encoded audio signal and the encoded spatial metadata to generate processed spatial metadata; and enabling 406 spatial rendering to be performed on one or more audio signals using the processed spatial metadata. The computer program instructions can be comprised in a computer program 1304, a non-transitory computer readable medium, a computer program product, a machine readable medium. In some but not necessarily all examples, the computer program instructions can be distributed over more than one computer program 1304. Although the memory 1210 is illustrated as a single component / circuitry it can be implemented as one or more separate components / circuitry some or all of which can be integrated / removable and / or can provide permanent / semi-permanent / dynamic / cached storage. Although the processor 1206 is illustrated as a single component / circuitry it can be implemented as one or more separate components / circuitry some or all of which can be integrated / removable. The processor 1206 can be a single core or multi-core processor. References to “computer-readable storage medium”, “computer program product”, “tangibly embodied computer program” etc. or a “controller”, “computer”, “processor” etc. should be understood to encompass not only computers having different architectures such as single / multi- processor architectures and sequential (Von Neumann) / parallel architectures but also specialized circuits such as field-programmable gate arrays (FPGA), application specific circuits (ASIC), signal processing devices and other processing circuitry. References to computer program, instructions, code etc. should be understood to encompass software for a programmable processor or firmware such as, for example, the programmable content of a hardware device whether instructions for a processor, or configuration settings for a fixed-function device, gate array or programmable logic device etc. As used in this application, the term “circuitry” can refer to one or more or all of the following: (a) hardware-only circuitry implementations (such as implementations in only analog and / or digital circuitry) and (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions and (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g. firmware) for operation, but the software cannot be present when it is not needed for operation. This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit for a mobile device or a similar integrated circuit in a server, a cellular network device, or other computing or network device. The apparatus 1300 as shown in Fig. 13 can be provided within any suitable device. In some examples the apparatus 1300 can be provided within an electronic device such as a mobile telephone, a teleconferencing device, a camera, a computing device, a server or any other suitable device. The blocks illustrated in the Figs, can represent steps in a method and / or sections of code in the computer program 1304. The illustration of a particular order to the blocks does not necessarily imply that there is a required or preferred order for the blocks and the order and arrangement of the blocks can be varied. Furthermore, it can be possible for some blocks to be omitted. The apparatus can be provided in an electronic device, for example, a mobile terminal, according to an example of the present disclosure. It should be understood, however, that a mobile terminal is merely illustrative of an electronic device that would benefit from examples of implementations of the present disclosure and, therefore, should not be taken to limit the scope of the present disclosure to the same. While in certain implementation examples, the apparatus can be provided in a mobile terminal, other types of electronic devices, such as, but not limited to: mobile communication devices, hand portable electronic devices, wearable computing devices, portable digital assistants (PDAs), pagers, mobile computers, desktop computers, televisions, gaming devices, laptop computers, cameras, video recorders, GPS devices and other types of electronic systems, can readily employ examples of the present disclosure. Furthermore, devices can readily employ examples of the present disclosure regardless of their intent to provide mobility. The term ‘comprise’ is used in this document with an inclusive not an exclusive meaning. That is any reference to X comprising Y indicates that X may comprise only one Y or may comprise more than one Y. If it is intended to use ‘comprise’ with an exclusive meaning then it will be made clear in the context by referring to ‘comprising only one...’ or by using ‘consisting.’ In this description, the wording ‘connect’, ‘couple’ and ‘communication’ and their derivatives mean operationally connected / coupled / in communication. It should be appreciated that any number or combination of intervening components can exist (including no intervening components), i.e., to provide direct or indirect connection / coupling / communication. Any such intervening components can include hardware and / or software components. As used herein, the term "determine / determining" (and grammatical variants thereof) can include, not least: calculating, computing, processing, deriving, measuring, investigating, identifying, looking up (for example, looking up in a table, a database, or another data structure), ascertaining and the like. Also, "determining" can include receiving (for example, receiving information), accessing (for example, accessing data in a memory), obtaining and the like. Also, "determine / determining" can include resolving, selecting, choosing, establishing, and the like. In this description, reference has been made to various examples. The description of features or functions in relation to an example indicates that those features or functions are present in that example. The use of the term ‘example’ or ‘for example’ or ‘can’ or ‘may’ in the text denotes, whether explicitly stated or not, that such features or functions are present in at least the described example, whether described as an example or not, and that they can be, but are not necessarily, present in some of or all other examples. Thus ‘example’, ‘for example’, ‘can’, or ‘may’ refers to a particular instance in a class of examples. A property of the instance can be a property of only that instance or a property of the class or a property of a sub-class of the class that includes some but not all the instances in the class. It is therefore implicitly disclosed that a feature described with reference to one example but not with reference to another example, can where possible be used in that other example as part of a working combination but does not necessarily have to be used in that other example. As used herein, “at least one of the following: ” and “at least one of ” and similar wording, where the list of two or more elements are joined by “and” or “or” mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements. Although examples have been described in the preceding paragraphs with reference to various examples, it should be appreciated that modifications to the examples given can be made without departing from the scope of the claims. Features described in the preceding description may be used in combinations other than the combinations explicitly described above. Although functions have been described with reference to certain features, those functions may be performable by other features whether described or not. The description of a feature, such as an apparatus or a component of an apparatus, configured to perform a function, or for performing a function, should additionally be considered to also disclose a method of performing that function. For example, description of an apparatus configured to perform one or more actions, or for performing one or more actions, should additionally be considered to disclose a method of performing those one or more actions with or without the apparatus. Although features have been described with reference to certain examples, those features may also be present in other examples whether described or not. The term ‘a’, ‘an’ or ‘the’ is used in this document with an inclusive not an exclusive meaning. That is any reference to X comprising a / an / the Y indicates that X may comprise only one Y or may comprise more than one Y unless the context clearly indicates the contrary. If it is intended to use ‘a’, ‘an’ or ‘the’ with an exclusive meaning then it will be made clear in the context. In some circumstances the use of ‘at least one’ or ‘one or more’ may be used to emphasis an inclusive meaning but the absence of these terms should not be taken to infer any exclusive meaning. The presence of a feature (or combination of features) in a claim is a reference to that feature or (combination of features) itself and to features that achieve substantially the same technical effect (equivalent features). The equivalent features include, for example, features that are variants and achieve substantially the same result in substantially the same way. The equivalent features include, for example, features that perform substantially the same function, in substantially the same way to achieve substantially the same result. In this description, reference has been made to various examples using adjectives or adjectival phrases to describe characteristics of the examples. Such a description of a characteristic in relation to an example indicates that the characteristic is present in some examples exactly as described and is present in other examples substantially as described. The above description describes some examples of the present disclosure however those of ordinary skill in the art will be aware of possible alternative structures and method features which offer equivalent functionality to the specific examples of such structures and features described herein above and which for the sake of brevity and clarity have been omitted from the above description. Nonetheless, the above description should be read as implicitly including reference to such alternative structures and method features which provide equivalent functionality unless such alternative structures or method features are explicitly excluded in the above description of the examples of the present disclosure. Whilst endeavoring in the foregoing specification to draw attention to those features believed to be of importance the Applicant may seek protection via the claims in respect of any patentable feature or combination of features hereinbefore referred to 5 and / or shown in the drawings whether or not emphasis has been placed thereon.

Claims

1. An apparatus comprising means for:receiving an encoded audio signal and corresponding encoded spatial metadata;determining at least one input for processing wherein the input is based on the encoded audio signal and the encoded spatial metadata;processing the at least one input to use information from the encoded audio signal and the encoded spatial metadata to generate processed spatial metadata; and enabling spatial rendering to be performed on one or more audio signals using the processed spatial metadata.

2. An apparatus as claimed in claim 1 wherein the means are for:determining a first input for processing wherein the first input is based on the encoded audio signal;determining a second input for processing wherein the second input is based on the encoded spatial metadata; andprocessing the first input and the second input to use information from the encoded audio signal and the encoded spatial metadata to generate processed spatial metadata.

3. An apparatus as claimed in any preceding claim wherein determining an input based on the encoded audio signal comprises at least partially decoding the encoded audio signal.

4. An apparatus as claimed in any preceding claim wherein determining an input based on the encoded spatial metadata comprises at least partially decoding the encoded spatial metadata.

5. An apparatus as claimed in any preceding claim wherein the processed spatial metadata comprises at least one of:improved temporal resolution compared to the received spatial metadata;improved frequency resolution compared to the received spatial metadata; orimproved accuracy of restoration of information from spatial metadata from which the encoded spatial metadata was obtained compared to the received spatial metadata.

6. An apparatus as claimed in any preceding claim wherein the processed spatial metadata comprises one or more energy-related parameters and one or more directional parameters.

7. An apparatus as claimed in claim 6 wherein the energy-related parameters comprise at least one of:a direct-to-total energy ratio;a diffuse-to-total energy ratio.

8. An apparatus as claimed in any of claims 6 to 7 wherein the directional parameters comprise at least one of:an azimuth angle;an elevation angle;a direction index.

9. An apparatus as claimed in any preceding claim wherein the means are for obtaining one or more features from the received encoded audio signal and using the one or more obtained features as the input to the processing.

10. An apparatus as claimed in any of claims 1 to 8 wherein processing the at least one input comprises obtaining one or more features from the received encoded audio signal.

11. An apparatus as claimed in any preceding claim wherein the means are for obtaining one or more features from the received encoded spatial metadata and using the one or more obtained features as, at least part of, the input for processing.

12. An apparatus as claimed in any of claims 1 to 10 wherein processing the at least one input comprises obtaining one or more features from the received encoded spatial metadata13. An apparatus as claimed in any preceding claim wherein the processing the at least one input comprises obtaining predicted spatial metadata properties and using the predicted spatial metadata properties and the received encoded spatial metadata to generate the processed spatial metadata.

14. An apparatus as claimed in any preceding claim wherein the spatial rendering comprises at least one of:binaural rendering;multi-loudspeaker rendering;Ambisonics rendering;stereo rendering; orcross talk-cancelled stereo rendering.

15. An apparatus as claimed in any preceding claim wherein the processing is performed, at least in part, using a machine learning model.

16. An apparatus as claimed in claim 15 wherein the machine learning model comprises a deep neural network model.

17. An apparatus as claimed in any preceding claim wherein the one or more audio signals are obtained by decoding the received encoded audio signal.

18. An apparatus as claimed in any preceding claim wherein the one or more audio signals comprise one or more channels.

19. An apparatus as claimed in any preceding claim wherein the one or more encoded audio signals comprise one or more channels.

20. A method comprising:receiving an encoded audio signal and corresponding encoded spatial metadata;determining at least one input for processing wherein the input is based on the encoded audio signal and the encoded spatial metadata;processing the at least one input to use information from the encoded audio signal and the encoded spatial metadata to generate processed spatial metadata; and enabling spatial rendering to be performed on one or more audio signals using the processed spatial metadata.

21. A computer program comprising instruction which, when executed by a processor, cause the processor to perform:receiving an encoded audio signal and corresponding encoded spatial metadata;determining at least one input for processing wherein the input is based on the encoded audio signal and the encoded spatial metadata;processing the at least one input to use information from the encoded audio signal and the encoded spatial metadata to generate processed spatial metadata; and enabling spatial rendering to be performed on one or more audio signals using the processed spatial metadata.58

Citation Information

Patent Citations

  • Apparatus, methods and computer programs for enabling rendering of spatial audio

    GB2615323A

  • Parametric spatial audio rendering

    GB2615607A

  • Apparatus, methods and computer programs for enabling rendering of spatial audio

    GB2617055A

  • Recording and Rendering Spatial Audio Signals

    US20200260206A1

  • Spatial audio Capture, Transmission and Reproduction

    US20210250717A1