Spatial metadata for rendering of spatial audio
By employing machine learning models to process audio and metadata features, the patent enhances spatial audio rendering at low bitrates, addressing quality degradations and ensuring accurate sound source stability and directionality in spatial audio formats.
Patent Information
- Application Number
- PCT/EP2025/070109
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-08
- Filing Date
- 2025-07-14
- Publication Date
- 2026-02-12
AI Technical Summary
Existing spatial audio rendering technologies face challenges in maintaining audio quality at low bitrates due to inadequate reconstruction accuracy of spatial metadata, leading to perceivable degradations in sound source stability and directionality.
Enhance spatial metadata by utilizing machine learning models, particularly deep neural networks, to process audio and metadata features, improving temporal and frequency resolution, and restoring directional information, thereby generating processed spatial metadata for accurate spatial rendering.
The enhanced spatial metadata improves audio quality by maintaining sound source stability and directionality even at low bitrates, ensuring accurate spatial rendering across various formats like binaural, multi-loudspeaker, and Ambisonics.
Smart Images

Figure EP2025070109_12022026_PF_FP_ABST
Abstract
Description
[0001]TITLE Spatial Metadata for Rendering of Spatial Audio TECHNOLOGICAL FIELDExamples of the disclosure relate to spatial metadata for rendering spatial audio. Some relate toenhancing the spatial metadata for rendering spatial audio by using information from thecorresponding audio signals. BACKGROUND Metadata assisted spatial audio (MASA) uses one or more audio signals together with corresponding spatial metadata to generate spatial audio. The spatial metadata can comprisedirection information, direct-to-total energy ratios in frequency bands, or any other suitableinformation. A MASA stream can, in some examples, be obtained by capturing spatial audiowith microphones and estimating the corresponding spatial metadata from the microphone signals. BRIEF SUMMARY According to various, but not necessarily all, examples of the disclosure there is provided anapparatus comprising means for:receiving an encoded audio signal and corresponding encoded spatial metadata; determining at least one input for processing wherein the input is based on the encoded audio signal and the encoded spatial metadata; processing the at least one input to use information from the encoded audio signal andthe encoded spatial metadata to generate processed spatial metadata; andenabling spatial rendering to be performed on one or more audio signals using theprocessed spatial metadata.The means may be for: determining a first input for processing wherein the first input is based on the encoded audio signal; determining a second input for processing wherein the second input is based on the encoded spatial metadata; andprocessing the first input and the second input to use information from the encoded audiosignal and the encoded spatial metadata to generate processed spatial metadata.Determining an input based on the encoded audio signal may comprise at least partially decoding the encoded audio signal. Determining an input based on the encoded spatial metadata may comprise at least partially decoding the encoded spatial metadata. The processed spatial metadata may comprise at least one of: improved temporal resolution compared to the received spatial metadata; improved frequency resolution compared to the received spatial metadata; or improved accuracy of restoration of information from spatial metadata from which the encoded spatial metadata was obtained compared to the received spatial metadata. The processed spatial metadata may comprise one or more energy-related parameters and one or more directional parameters. The energy-related parameters may comprise at least one of: a direct-to-total energy ratio; a diffuse-to-total energy ratio. The directional parameters may comprise at least one of: an azimuth angle; an elevation angle;a direction index. The means may be for obtaining one or more features from the received encoded audio signal and using the one or more obtained features as the input to the processing.Processing the at least one input may comprise obtaining one or more features from the receivedencoded audio signal. The means may be for obtaining one or more features from the received encoded spatial metadata and using the one or more obtained features as, at least part of, the input for processing.Processing the at least one input may comprise obtaining one or more features from the receivedencoded spatial metadataThe processing the at least one input may comprise obtaining predicted spatial metadataproperties and using the predicted spatial metadata properties and the received encoded spatialmetadata to generate the processed spatial metadata.The spatial rendering may comprise at least one of: binaural rendering; multi-loudspeaker rendering; Ambisonics rendering; stereo rendering; or cross talk-cancelled stereo rendering. The processing may be performed, at least in part, using a machine learning model. The machine learning model may comprise a deep neural network model. The one or more audio signals may be obtained by decoding the received encoded audio signal. The one or more audio signals may comprise one or more channels. The one or more encoded audio signals may comprise one or more channels. According to various, but not necessarily all, examples of the disclosure there is provided a method comprising: receiving an encoded audio signal and corresponding encoded spatial metadata; determining at least one input for processing wherein the input is based on the encoded audio signal and the encoded spatial metadata; processing the at least one input to use information from the encoded audio signal andthe encoded spatial metadata to generate processed spatial metadata; andenabling spatial rendering to be performed on one or more audio signals using theprocessed spatial metadata. According to various, but not necessarily all, examples of the disclosure there is provided a computer program comprising instruction which, when executed by a processor, cause the processor to perform: receiving an encoded audio signal and corresponding encoded spatial metadata; determining at least one input for processing wherein the input is based on the encoded audio signal and the encoded spatial metadata; processing the at least one input to use information from the encoded audio signal andthe encoded spatial metadata to generate processed spatial metadata; andenabling spatial rendering to be performed on one or more audio signals using theprocessed spatial metadata.According to various, but not necessarily all, embodiments there is provided an apparatus comprising at least one processor; and at least one memory including computer program code; the at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to perform at least a part of one or more methods described herein. According to various, but not necessarily all, embodiments there is provided an apparatus comprising means for performing at least part of one or more methods described herein. The description of a function and / or action should additionally be considered to also disclose any means suitable for performing that function and / or action. Functions and / or actions described herein can be performed in any suitable way using any suitable method.According to various, but not necessarily all, embodiments there is provided examples as claimedin the appended claims. While the above examples of the disclosure and optional features are described separately, it is to be understood that their provision in all possible combinations and permutations is containedwithin the disclosure. It is to be understood that various examples of the disclosure can compriseany or all the features described in respect of other examples of the disclosure, and vice versa.Also, it is to be appreciated that any one or more or all the features, in any combination, may beimplemented by / comprised in / performable by an apparatus, a method, and / or computer programinstructions as desired, and as appropriate. The description of a function should additionally beconsidered to also disclose any means suitable for performing that function BRIEF DESCRIPTIONSome examples will now be described with reference to the accompanying drawings in which:FIG.1 shows an example system; FIG.2 shows an example encoder; FIG.3 shows an example decoder; FIG.4 shows an example method; FIG.5 shows an example metadata enhancer; FIG.6 shows another example metadata enhancer;FIGS.7A and 7B show an example machine learning model;FIG.8 shows an example pre-network;FIG.9 shows an example ResBlock;FIG.10 shows an example system; FIGS.11A to 11D show example results; FIG.12 shows an example device; and FIG.13 shows an example apparatus.The figures are not necessarily to scale. Certain features and views of the figures can be shownschematically or exaggerated in scale in the interest of clarity and conciseness. For example, thedimensions of some elements in the figures can be exaggerated relative to other elements to aidexplication. Corresponding reference numerals are used in the figures to designate correspondingfeatures. For clarity, all reference numerals are not necessarily displayed in all figures.DETAILED DESCRIPTIONFig. 1 shows an example system 100 that can be used for spatial audio. The system 100comprises an encoder 106 and a decoder 110. The encoder 106 and decoder 110 can be in different devices. For example, the system 100 can be part of a telecommunications system inwhich the encoder 106 can be provided within a first communications device and the decoder 110can be provided within a different communications device.The encoder 106 receives a spatial audio stream as an input. The spatial audio streamcomprises audio signals 102 and spatial metadata 104. The spatial audio stream can beprovided in a metadata-assisted spatial audio (MASA) format or in any other suitable format. In some examples the audio signals 102 can comprise microphone signals or signals that are obtained by processing microphone signals. In some examples the audio signals 102 can originate from other inputs such as multi-channel audio signals (for example, 5.1 or 7.1+4) or audio objects. In such cases the audio signals 102 can comprise a downmix of the originally obtained inputs. The audio signals 102 can comprise transport audio signals. The transport audio signals are audio signals that have been processed into a format for transport or transmission to another communication device. The spatial metadata 104 comprises information that can be used to render the audio signals 102 to generate spatial audio. The spatial metadata 104 can comprise direction information, direct- to-total energy ratios in frequency bands, or any other suitable information. The spatial metadata 104 can be estimated from the microphone signals or from any other suitable signals. The encoder 106 is configured to encode the audio signals 102 and the spatial metadata 104 to form a bitstream 108. The bitstream 108 can be sent to the decoder 110. Encoder 106 and the decoder 110 can be in different devices and so the bitstream can be sent via any suitable communication network. The decoder 110 receives the bitstream 108 as an input. The decoder 110 is configured to decode the bitstream 108 to render spatial audio output 112. The spatial audio output can comprise binaural audio signals, stereo signals, Ambisonics signals or any other suitable types of signals.Fig.2 shows an example encoder 106 in more detail. The encoder 106 is configured so that theaudio signals 102 are provided to an audio encoder 200. The audio encoder 200 is configured to encode the audio signals 102. The audio encoder 200 can use any suitable audio-signalencoder to encode the audio signals 102. For example, the audio encoder 200 can use theImmersive Voice and Audio Services (IVAS) core coder, Enhanced Voice Services (EVS), orAdvanced Audio Coding (AAC). The audio encoder 200 provides encoded audio signals 204 asan output. The encoded audio signals 410 can comprise one or more channels.The encoder 106 is also configured so that the spatial metadata 104 is provided to a spatialmetadata encoder 202. The spatial metadata encoder 202 is configured to encode the spatialmetadata 104. The spatial metadata encoder 202 can use any method to encode the spatialmetadata 104. For example, the spatial metadata encoder 202 can use the MASA encodingmethods of the IVAS encoder or any other suitable encoding methods. The spatial metadataencoder 202 provides encoded spatial metadata 206 as an output.The encoded audio signals 204 and encoded spatial metadata 206 are provided as an input to amultiplexer 208. The multiplexer 208 multiplexes the encoded audio signals 204 and encodedspatial metadata 206 to a bitstream 108. The bitstream 108 is the output of the encoder 106. The encoding methods that are used to encode the spatial metadata 104 control the level of compression depending on the bitrate. Using a lower bitrate can reduce the reconstructionaccuracy of the spatial metadata 104 at the decoder 110. With a low bitrate a higher compressionfactor is needed and this is obtained by trading off some of the reconstruction accuracy. For example, at a very low bitrates (such as 13.2 kbps with IVAS), the bitrate that can be allocated tospatial metadata 104 encoding is so low (such as about 2.25 kbps with IVAS) that it adverselyaffects the perceptual quality of the spatial audio output 112 at the decoder 110. These adverse effects can include instability and inaccuracies of the perceived sound sources. Allocating more bits to the spatial metadata 104 does not resolve this issue because there are constraints on the total bitrate. Increasing the bits used for the spatial metadata 104 would reducethe bits available for the coding of the audio signals 102. This could result in a perceivabledegradation of audio quality.There are some redundancies between the audio signals 102 and the spatial metadata 104. Forexample, when there is an onset in the audio (such as, when someone starts to speak), typicallyboth the energy of the audio signal 102 and the direct-to-total energy ratio of the spatial metadataincrease rapidly in the time-frequency tiles at the onset. Similarly, if there is a dominant audiosource on the left side, the directional spatial metadata at those time-frequency tiles points at left. Examples of the disclosure make use of these redundancies to enable improved or enhanced spatial metadata to be obtained. The improved or enhanced spatial metadata can be obtained by using information from the audio signals to process the spatial metadata. This can help to account for any information that is lost in the encoding of the spatial metadata 102 or any other encoding and avoid perceivable degradations of audio quality even when a low bitrate is used.Fig. 3 shows an example decoder 110 that can be used in examples of the disclosure. Thedecoder 110 can be part of a system 100 as shown in Fig.1. The decoder 110 receives the bitstream 108 as an input. The bitstream 108 can be generated by an encoder 106 as shown in Figs.1 and 2. The bitstream 108 is provided to a demultiplexer 300. The demultiplexer 300 demultiplexes the bitstream 108 to provide encoded audio signals 204 and encoded spatial metadata 206. The encoded audio signals 204 can be encoded transport audio signals or any other type of audio signals.The encoded audio signals 204 are provided to an audio decoder 302. The audio decoder 302is configured to decode the encoded audio signals. The audio decoder 302 can use any suitabletype of decoder. The audio decoder 302 can use a decoder that is compatible with the audio encoder 200 that was used to encode the audio signals 102. The audio decoder 302 provides decoded audio signals 306 as an output. These can be decoded transport audio signals.The encoded spatial metadata 206 is provided to a metadata decoder 304. The metadatadecoder 304 is configured to decode the encoded spatial metadata 206. The metadata decoder 304 can use any suitable type of decoder. The metadata decoder 304 can use a decoder that is compatible with the spatial metadata encoder 202 that was used to encode the spatial metadata104. The metadata decoder 304 provides decoded spatial metadata 308 as an output.The decoded audio signals 306 and the decoded spatial metadata 308 are provided as an input to a metadata processor 310. The metadata processor 310 is configured to process the decoded spatial metadata 308 so that information from the decoded audio signals 306 can be used to improve or enhance the spatial metadata and provide improved spatial audio. The metadata processor 310 can comprise a machine learning model or any other suitable means. The metadata processor 310 provides processed spatial metadata 312 as an output. This can be processed decoded spatial metadata.The processed decoded spatial metadata 312 and the decoded audio signals 306 are providedas an input to a spatial synthesizer 314. The spatial synthesizer 314 is configured to processthe processed decoded spatial metadata 312 and the decoded audio signals 306 to render spatialaudio output 316. The spatial audio output can comprise binaural audio signals, stereo audiosignals, multi-channel audio signals (e.g., 5.1 or 7.1+4), Ambisonics signals, or any other suitabletype of spatial audio.Fig. 4 shows an example method that can be used in examples of the disclosure. The methodcould be implemented by a decoder 110 or an apparatus or means within a decoder 110. The method comprises, at block 400, receiving an encoded audio signal 204 and correspondingencoded spatial metadata 206. In some examples the encoded audio signal 204 andcorresponding encoded spatial metadata 206 can be received in a combined bitstream 108 and then separated out by a demultiplexer 300 or any other suitable means. The received encoded audio signal 204 can comprise one or more channels. The encoded spatial metadata 206 corresponds to the encoded audio signal 204 in that the spatial metadata and the audio signals represent the same sound scene. The spatial metadata canbe used to process the audio signals to render spatial audio.At block 402 the method comprises determining at least one input for processing. The input isbased on the encoded audio signal 204 and the encoded spatial metadata 206. The processingcan comprise processing that is performed on the spatial metadata so as to improve the spatial audio that can be achieved using the spatial metadata. Determining an input based on the encoded audio signal 204 can comprise at least partiallydecoding the encoded audio signal 204. Determining an input based on the encoded spatialmetadata 206 can comprise at least partially decoding the encoded spatial metadata 206. In some examples determining an input based on the encoded audio signal 204 can comprise obtaining one or more features from the received encoded audio signal 204 and using the one ormore obtained features as the input to the processing. The features can comprise informationrelating to the physical properties of the audio signal. The features can comprise peak energylevels, evolution of energy levels, repetitions of energy levels, or any other suitable information.Similarly, in some examples determining an input based on the encoded spatial metadata 206 can comprise obtaining one or more features from the received encoded spatial metadata 206 and using the one or more obtained features as the input to the processing. The features cancomprise information relating to the properties of the spatial metadata.At block 404 the method comprises processing the at least one input to use information from theencoded audio signal 204 and the encoded spatial metadata 206 to generate processed spatialmetadata 312. The information from the encoded audio signal 204 can comprise information thatis inherent within the encoded audio signal 204, features that are extracted from the encoded audio signal 204 or any other suitable information.The processing the at least one input to use information from the encoded audio signal 204 andthe encoded spatial metadata 206 to generate processed spatial metadata 312 can compriseobtaining predicted spatial metadata properties and using the predicted spatial metadataproperties and the received encoded spatial metadata 206 to generate the processed spatialmetadata 312.The processed spatial metadata 312 that is generated by the processing can comprise one ormore energy-related parameters and one or more directional parameters and / or any othersuitable parameters. The energy related parameters can comprise a direct-to-total energy ratio,a diffuse-to-total energy ratio and / or any other suitable parameters. The directional parameterscan comprise any information that indicates a direction of arrival of sound. The directionalparameters can comprise an azimuth angle, an elevation angle, a direction index and / or any othersuitable information. The processed spatial metadata 312 can comprise enhanced or improved spatial metadatacompared to the original spatial metadata. The original spatial metadata can be the metadatafrom which the encoded spatial metadata is obtained 206. In some examples the processed spatial metadata 312 can comprise improved temporal resolution compared to the received spatial metadata. In some examples the processed spatial metadata 312 can comprises improved frequency resolution compared to the received spatial metadata. In some examples the processed spatial metadata 312 can comprise improved accuracy of restoration of informationfrom spatial metadata from which the encoded spatial metadata 206 was obtained compared tothe received spatial metadata. The accuracy of restoration of information is the accuracy withwhich a value that is closer to the original spatial metadata is obtained compared to the encodedspatial metadata 206. In some examples the input for the processing can comprise one or more features. For instance, determining the input can comprise determining one or more features of the encoded audio signals 204 or one or more features of the encoded spatial metadata 206. In other examplesthe processing can comprise obtaining one or more features from the received encoded audiosignals 204 or the received encoded spatial metadata 206.The processing can be performed, at least in part, using a machine learning model. The machinelearning model can comprise a deep neural network model and / or any other suitable operations or combinations of operations. The method comprises, at block 406, enabling spatial rendering to be performed on one or moreaudio signals 306 using the processed spatial metadata. The spatial rendering can beperformed on decoded audio signals. The decoded audio signals can be obtained by decoding the encoded audio signals 204 that were received at block 400. The spatial rendering can comprise binaural rendering, multi-channel loudspeaker rendering, Ambisonics rendering, stereo rendering, cross talk-cancelled stereo rendering, or any other suitable type of rendering. Variations to this method could be used. In some examples a single input can be provided for the processing. In other examples multiple inputs can be provided. If multiple inputs are used the method can comprise determining a first input for processing wherein the first input is based on the encoded audio signal 204 and determining a second input for processing wherein the second input is based on the encoded spatial metadata 206. In such examples the first input and thesecond input would be processed to use information from the encoded audio signal 204 and theencoded spatial metadata 206 to generate processed spatial metadata 312.Fig. 5 shows an example metadata processor 310 that can be used in some examples. Themetadata processor 310 can be provided within a decoder 110 as shown in Fig.3. The metadataprocessor 310 can be used to implement methods such as the method of Fig. 4 and / or anyvariations of this method. Decoded audio signals 306 and the decoded spatial metadata 308 are provided as an input tothe metadata processor 310. The decoded audio signals 306 and the decoded spatial metadata308 have both been encoded and decoded so the data that is received by the metadata processor310 has been compressed in some way. In some use case scenarios the decoded spatialmetadata 308 can comprise information at a low time-frequency resolution. This could be because the spatial metadata is coded using a low-bitrate IVAS codec. The values of the decoded spatial metadata 308 can also be quantized to a lower resolution than the values of the original spatial metadata. The feature computation block 500 receives the decoded audio signals 306 and the decoded spatial metadata 308 as inputs. The feature computation block 500 is configured to determinefeatures from the received decoded audio signals 306 and the decoded spatial metadata 308 thatcan be used as inputs for processing. The features can comprise information that describes physical properties of the sound scene represented by the decoded audio signals 306 and thedecoded spatial metadata 308. The features can comprise peak energy levels, evolution of energylevels, repetitions of energy levels, or any other suitable information. In some examples thefeatures can comprise information that describes some properties or features of the audio signalsand / or the spatial metadata. Such information provides indirect information relating to thephysical properties of the sound scene. Such information could comprise inter- channel featuresor other suitable information. As shown in Fig.5 the feature computation block 500 receives the decoded audio signals 306 as an input. In other examples the feature computation block 500 could receive partially decoded audio signals or encoded audio signals as an input. Similarly, in Fig.5 the feature computationblock 500 receives the decoded spatial metadata 308 as an input. In other examples, the featurecomputation block 500 could receive partially decoded spatial metadata or encoded metadata as an input. The feature computation block 500 provides features 502 as an output. The features 502 can comprise features from the decoded audio signals 306 and features the decoded spatial metadata 308. The features 502 are provided as an input to a machine learning model 504. The machine learning model 504 can comprise multiple defined processing steps, and can be similar to the processing instructions related to conventional program code. The difference between conventional program code and the machine learning model 504 is that the instructions of the conventional program code are defined more explicitly at the programming time. Theinstructions of the machine learning model 504 are defined by combining a set of predefinedprocessing blocks (such as convolutions, data normalizations, other operators), where the weights of the model are unknown at the model definition time. The weights of the machine learning model 504 are optimized by providing the machine learning model 504 with a large amount of input and reference data, and the model weights then converge so that the machine learning model 504 is trained to solve a given task. In this case the task is processing the inputs to generate processed spatial metadata that can provide for improved spatial audio. In examples of the disclosure, when the machine learning model 504 is used, the machine learning model 504 would be fixed and would correspond to a set of processing instructions. The machine learning model 504 is trained to provide predicted metadata properties 506 as an output. The predicted metadata properties 506 can comprise predicted values for parameters of the spatial metadata. The predicted metadata properties 506 can comprise predicted directional properties and / or predicted energy properties. In some examples the predicted metadata properties 506 can comprise information such as MASA directional information (azimuth, elevation, and direct-to-total energy ratio). The predicted metadata properties 506 canbe provided in any suitable representation, such as Cartesian “xyz” vector format.The predicted metadata properties 506 are provided as an input to a metadata determinationblock 508. The metadata determination block 508 also receives the decoded spatial metadata 308 as an input. The metadata determination block 508 is configured to use the predictedmetadata properties 506 to process or enhance the decoded spatial metadata 308. Theprocessing of decoded spatial metadata 308 using the predicted metadata properties 506 usesinformation that was comprised in the decoded audio signals 306 to make the decoded spatial metadata 308 closer to the original spatial metadata that was obtained by the encoder. The metadata determination 508 provides the processed spatial metadata 312 as an output. The processed spatial metadata 312 can have increased time-frequency resolution compared to the decoded spatial metadata 308. The de-quantization accuracy of the parameter values of the processed spatial metadata 312 can be increased compared to the decoded spatial metadata.The increase in de-quantization accuracy provides for improved accuracy of restoration ofinformation from the original spatial metadata that was obtained by the encoder. The processed spatial metadata 312 is provided as the output of the metadata processor 310. In the example of Fig.5 the feature computation block 500 is shown as a separate block to the machine learning model 504. In some examples the feature computation could be performed by the machine learning model 504. In such cases there would be no separate feature computationblock 500 and the machine learning model 504 could receive the decoded audio signals 306 andthe decoded spatial metadata 308 as inputs, In some examples the feature computation block 500 can comprise a machine learning algorithm that is different to the machine learning model 504. The machine learning algorithm of the feature computation block 500 could be trained in conjunction with the machine learning model 504 so as to provide appropriate inputs. Other means could be used in other examples.Fig. 6 shows another example metadata processor 310. The example metadata processor 310of Fig.6 is similar to the metadata processor 310 of Fig.5 and corresponding reference numeralsare used for corresponding components.The metadata processor 310 of Fig.6 differs from the metadata processor 310 of Fig.5 in that itcomprises a metadata feature computation block 600 and an audio feature computation block 602 rather than a single feature computation block 500. The metadata feature computation block 600 receives the decoded spatial metadata 308 as inputs. The metadata feature computation block 600 is configured to determine features from the received decoded spatial metadata 308 that can be used as inputs for processing. The metadata feature computation block 600 provides metadata features 604 as an output. Themetadata features 604 can comprise features from the decoded spatial metadata 308.The audio feature computation block 602 receives the decoded audio signals 306 as inputs. The audio feature computation block 602 is configured to determine features from the received decoded audio signals 306 that can be used as inputs for processing. The audio feature computation block 602 provides audio features 606 as an output. The audio features 606 can comprise features from the decoded audio signals 306. The metadata features 604 and the audio features 606 can be provided as inputs to the machine learning model 504 which can be arranged to process the metadata features 604 and the audio features 606 to provide predicted metadata properties 506. In the example of Fig. 6 the metadata feature computation block 600 and the audio featurecomputation block 602 are shown as separate blocks to the machine learning model 504. In someexamples the feature computation could be performed by the machine learning model 504. In such cases there would be no separate metadata feature computation block 600 or audio featurecomputation block 602 and the machine learning model 504 could receive the decoded audiosignals 306 and the decoded spatial metadata 308 as inputs. In some examples the metadata feature computation block 600 and the audio feature computation block 602 can comprise a machine learning algorithm that is different to the machine learning model 504. The machine learning algorithms of the metadata feature computation block 600 and the audio feature computation block 602 could be trained in conjunction with the machine learning model 504 so as to provide appropriate inputs. Other means could be used in other examples. In the following example it is assumed that there is one direction active in the MASA metadata. The high-resolution spatial metadata, that is the spatial metadata as it is originally estimatedbefore it is encoded and decoded, is assumed to have 24 bands (all MASA bands) with distinctvalues and all 4 sub-frames within a frame having the same values. This means that the temporalresolution of the “high-resolution” spatial metadata is lower than what could be supported by theMASA format.The resolution of the spatial metadata 104 is reduced by the metadata encoding as shown in Fig.2. In this example, the spatial metadata 104 is encoded into 5 parameter bands and 1 sub-frameper frame. Thus, for example, a single parameter set is transmitted and the value is used for all 4sub-frames. The encoding of the spatial metadata 104 therefore reduces the frequency resolutionof the spatial metadata and the value resolution through quantization. This resolution correspondsto the resolution the IVAS codec is using when operating at the 13.2 kbps total bitrate for the MASA content having approximately 2.25 kbps for the MASA spatial metadata.It should be noted that this is only one example. Examples of the disclosure can be used with anynumber of original and coded frequency bands and temporal subframes, as well as any suitable quantization scheme, as well as with any suitable bit rate. The metadata feature computation block 600 can receive the low resolution decoded spatial metadata 308 as an input. The low resolution decoded spatial metadata 308 can be denoted^^^(^, ^), ^^^(^, ^), and ^(^, ^) corresponding to the azimuth angle, elevation angle, and direct-to-total energy ratio in TF(time-frequency)-tile ^, ^, where ^ is the parameter band index ^ =1, … , ^_^^^^^_^^^^, and ^ is the frame index. This spherical representation can be transformedinto Cartesian xyz vector representation with The metadata feature computation block 600 can provide metadata features 604 as an output. The metadata features 604 can be represented with such vectors for each TF-tile. The metadatafeatures 604 can be considered as a 3-dimensional tensor with the shape(^_^^^^^_^^^^, ^_^^^^^^, ^_^^^^^^^^_^^^^) , where ^_^^^^^^^^_^^^^ = 3 is the number ofinput features per tile (for example three corresponding to the three elements of the vector^(^, ^)), ^_^^^^^^ is the number of spatial metadata frames that are processed at once inimplementations of the disclosure. In this example this is ^_^^^^^^ = 200 corresponding to 4seconds a spatial audio stream with 50 frames per second. ^_^^^^^_^^^^ = 5 corresponds tothe 5 parameter bands of the low-resolution spatial metadata.Also in this example the decoded audio signals 306 can be assumed to comprise two audiochannels. This assumption is valid for this example and audio feature representation butdifferent audio feature representations can be used for the same number of audio channels in thedecoded audio signals or different numbers of audio channels in the decoded audio signals 306.The channels of the decoded audio signals 306 can be denoted as ^^(^) and ^^(^) where ^^(^)corresponds to the left channel and ^^(^) corresponds to the right channel). The channels ^^(^)and ^^(^) can be transformed into the frequency-domain. Any suitable means can be used totransform the channels ^^(^) and ^^(^) to the frequency domains such as a complex-valuedlow-delay filter bank (CLDFB). In this example the transformation produces 60 complex-valuedfrequency-domain samples for ^^(^, ^) and ^^(^, ^) for each 60 time-domain samples ^corresponding to the time range of the slot ^ and frequency bin ^. Other transforms, such asShort-time Fourier transform (STFT), can be used in other examples.The audio feature computation block 602 determines audio features 606. The audio features 606 can comprise information that is potentially useful for enhancing the resolution of the spatialmetadata. Different features or types of features can be used in different examples of thedisclosure. In some examples the audio features 606 can comprise SPAC-based features that refer to spatial audio capture and Cov-based features that refer to covariance properties. The SPAC-based audio features can be determined with^(^, ^, ^^^^^ ) where ^ is the feature output temporal index, for example, corresponding to the frame ^ ,^^^^^^(^) and ^^^^^(^) are the starting and ending time slot indices of feature frame ^, ^ =with ^ = 33 is a delay index, ^^^^^ = 1, … , ^^^^^ with ^^^^^ = 60 is a frequency bandindex, ^^^^^ ^^^^^^^ (^^^^^ ) and ^^^^^ (^^^^^ ) are the CLDFB bin limits for frequency band index ^^^^^ ,^^^^(^) is the center frequency of CLDFB bin ^, ^^^(^) is the delay value corresponding to thedelay index ^, ^∗^ is the complex conjugate transpose of ^^ , and ^ = √−1 is the imaginaryunit. The feature temporal index ^ can have the same or different spacing as for the spatialmetadata and other audio features.The set of delay-values ^^^(^) can be determined so that they span a reasonable range giventhe assumed or virtual microphone spacing. For example, for a typical smart phone at a landscape mode, the delays could be equispaced in the range between -0.43 and 0.43 milliseconds,corresponding to the microphone distance of 0.150 m with the speed of sound of 343 m / s. Thenumber of bands ^ and the band borders ^ ^^^^^^^ (^^^^^ ) and ^ ^^^^^^^^ (^^^^^) can be a designdecision and may be different from the definitions in MASA format and between different metadataand audio features. In some examples these can be made to approximate the 24 MASA bands.In other examples, including the present example, the input CLDFB bins are not grouped intobands, but each band consists of exactly one CLDFB bin. The feature can be computed over thetime range corresponding to one MASA frame. In other words, the number of CLDFB slots in eachfeature frame is ^^^^^(^) − ^^^^^^(^) + 1 = 16. As a result, for the input spatial audio stream oflength of 4 seconds, the SPAC feature is of shape(^_^^^^^_^^^^, ^_^^^^^^_^^^^, ^_^^^^^^^^_^^^^ ) , where ^_^^^^^^^^_^^^^ = ^ = 33 ,^_^^^^^^_^^^^ = 200, and ^_^^^^^_^^^^ = ^^^^^ = 60.In this example the audio features 606 also comprise a second set of features. These are parallel to the SPAC features and are referred to as the Cov features. The Cov featurescomprise number of related audio features ^^^^(^, ^), where ^ is the feature index. The first (^ =1) feature is the normalized covariance value between the channels of the decoded audio signals306 determined withwhere ^^^ is the feature output temporal index, for example, corresponding to the sub-frame,^ and ^^^^^^^^^^ are the starting and ending time slot indices of feature (sub-)^^^, ^ = 1, … , ^^^^ with ^^^^ = 24 is a frequency band index, ^ ^^^^^^ (^) and ^ ^^^^^^^ (^) are theCLDFB bin limits for frequency band index ^ which may correspond to the MASA frequencybands. Both the granularity of the Cov features, that is the number of bands ^^^^ and thetemporal spacing − 1^ and temporal range can be different from the SPAC features. This is the case in this example. In other examplesdifferent audio features 606 could have the same granularity.If the TF-settings are the same for both the SPAC features and the Cov features then,^^^^^^^^ , ^^ is a subset of ^(^, ^, ^^^^^ ) for ^^^(^) = 0.The second (^ = 2) feature describes the level difference between the two channels of thedecoded audio signals 306 and is computed with where ^^^^ = 20 is a channel-level difference normalization factor.The third (^ = 3) feature describes the signal energy evolution over time with where ^^^^ = 24 is a normalization factor for the energy evolution, ^^^^ = 15 is the length ofthe energy evolution context, and The fourth (^ = 4) feature describes the signal energy distribution over bands withwhere ^^^^ = 40 is a normalization factor for the band energy ratio.The fifth ( ^ = 5 ) feature describes the signal energy over time with frequency-dependentnormalization with All the presented divisions and logarithms may contain numerical regularization limiting the argument to be strictly positive. As a result, for the input spatial audio stream of length of 4 seconds, the Cov feature is of shape(^_^^^^^_^^^, ^_^^^^^^_^^^, ^_^^^^^^^^_^^^ ) , where ^_^^^^^^^^_^^^ = 5 , ^_^^^^^^_^^^ =800, and ^_^^^^^_^^^ = ^^^^ = 24.In this example the frequency bin grouping of the Cov feature is independent of the frequency bin grouping of the SPAC feature. The Cov feature can have a different frequency bin grouping to the SPAC feature. The length of the temporal feature frame for the Cov feature can also be independent of the length of the temporal feature frame for the SPAC feature. In this example theSPAC feature has frame spacing of 20 ms corresponding to one MASA or IVAS frame while theCov features have frame length of 5 ms corresponding to sub-frame temporal resolution in MASAand IVAS. The machine learning model 504 receives the features as an input. In this example the machine learning model 504 receives the metadata features 604 and the audio features 606. The audio features 606 comprise SPAC features and the Cov features as described. Other types of features can be used in other examples. In this example the machine learning model 504 comprises a deep neural network (DNN). Other types of machine learning model 504 can be used in other examples. The machine learning model 504 receives the metadata features 604 and the audio features 606 and produces a predicted metadata properties 506 as an output. The output can be have a shape of(^_^^^^^_^^^^, ^_^^^^^^, ^_^^^^^^^^_^^^ ), where ^_^^^^^^^^_^^^ = 3, ^_^^^^^_^^^^ = 24,and ^_^^^^^^ corresponds to the number of MASA frames from the input spatial audio stream.For the example signal of 4 seconds in length, ^_^^^^^^ = 200.The predicted metadata properties 506 that are provided as the output of the machine learningmodel 504 describe the directional spatial metadata in the Cartesian xyz vector format. Theformat for the predicted spatial metadata properties 506 can be similar to the format used for the metadata features 604 however the spatial metadata properties 506 can now have a higher frequency resolution and finer quantization values. Different structures can be used for the machine learning model 504 in different examples. Figs.7A and 7B shows an example structure for a machine learning model 504 that can be used insome examples. Other structures for the machine learning model 504 can be used in otherexamples.In the example of Fig. 7A the machine learning model 504 comprises a W-Net structure. Thestructure is referred to a W-net because the path from the input to the output goes through two U-Net structures. Fig.7B shows an example U-Net structure that can be used in the machine learning model 504.The U-Net structure comprises a downsampling part 720 followed by an upsampling part 724.The downsampling part 720 comprises a sequence of downsampling layers. The respectivedownsampling layers can comprise convolutional operations such as convolutional neuralnetworks (CNN). The downsampling part 720 can capture high level features from an input. Thedownsampling layers of the downsampling part 720 can reduce the dimensions of an input alongat least some axes.The output of the downsampling part 720 has a smaller number of data elements in at least oneaxis compared to the original input.The output of the downsampling part 720 is provided as an input to the upsampling part 724. Theupsampling part 724 is configured to produce output data. The upsampling part 724 comprises asequence of upsampling layers. The upsampling part 724 can comprise X upsampling layerswhere X is also the number of downsampling layers in the downsampling part 720. Theupsampling layers can comprise transposed convolutional operations such as transposed CNNs.The upsampling layers of the upsampling part 724 can increase the dimensions of an input alongat least some axes. The U-Net structures in this example also comprise skip connections 722. The skip connections722 are configured to relay skip connection signals from respective downsampling layers tocorresponding upsampling layers. The skip connection signals can reintroduce features fromthe downsampling part 720 back into corresponding layers of the upsampling part 724. Theupsampling layers can comprise operations such as concatenating operations to combine datafrom a skip connection signal with input data. The upsampling layers can also compriseoperations to increase the dimensions of data that is input to the upsampling part 724.The upsampling part 724 provides output data. The output data of the upsampling part 724 hasthe same number of data elements in at least one dimension as the input data that is originallyprovided to the downsampling part 720.In the example of Figs. 7A and 7B the output of the downsampling part 720 is provided as aninput to the upsampling part 724. In other examples there can comprise one or more interveningcomponents such as a bottleneck and / or any other suitable operations or combinations of operations. In the example of Fig.7A the machine learning model 504 comprises three pre-networks 704A,704B, 704C. Each of the pre-networks 704A, 704B, 704C comprises a U-Net. The U-Net canbe as shown in Fig.7B (some of the reference numbers are omitted in Fig. 7A for clarity) or canhave any other suitable arrangement.In this example the machine learning model receives the metadata features 604 and the audiofeatures 606 as an input. The audio features 606 comprise SPAC features 700 and Cov features702. The respective feature inputs are provided to different pre-networks 704. The metadata features 604 are provided as an input to a first pre-network 704A, the SPAC features 700 are provided as an input to a second pre-network 704B, and the Cov features 702 are provided as an input to a third pre-network 704C. The respective pre-networks 704A, 704B, 704C are arranged to process the respective inputs to provide intermediate representations 706A, 706B, 706C as an output. The settings and weights of the respective pre-networks 704A, 704B, 704C are specific for each of the input features. The first pre-network 704A processes the input metadata features 604 to provide a metadata intermediate representation 706A as an output, the second pre-network 704B processes the input SPAC features 700 to provide a SPAC intermediate representation 706B as an output, and the third pre-network 704C processes the input Cov features 702 to provide a Cov intermediate representation 706C as an output. The intermediate representations 706A, 706B, 706C are provided to a concatenation block 708. The concatenation block 708 is configured to combine the intermediate representations 706A,706B, 706C. The concatenation block 708 can concatenate the intermediate representations706A, 706B, 706C along the feature axis or perform any other suitable combination. The concatenation block 708 provides a combined intermediate representation 710 as an output. The combined intermediate representation 710 is provided as an input to a combined prediction network 712. The combined prediction network 712 can comprise another U-Net structure. The weights and settings of the U-Net structure of the combined prediction network 712 can be different to the weights and settings used for the U-Nets in the pre-networks 704A, 704B, 704C. The combined prediction network 712 provides pre-scale predicted metadata properties 714 as an output. The pre-scale predicted metadata properties 714 are provided as an input to an XYZ scale block 716. The XYZ scale block 716 is configured to apply appropriate scaling to the pre-scale predicted metadata properties 714. The XYZ scale block 716 provides predicted spatial metadata properties 506 as an output.Fig. 8 shows an example structure of a pre-network 704. The example of Fig. 8 shows astructure for a metadata features pre-network 704A. Corresponding structures can be used forthe SPAC features pre-network 704B, the Cov features pre-network 704C, and the combinedprediction network 712. The input in this case is the metadata features 604. This input has shape(^_^^^^^_^^^^, ^_^^^^^^, ^_^^^^^^^^_^^^^). The input is provided to a dimension adjustment800. The dimension adjustment 800 comprises a linear layer 802. The linear layer 802 can be fully connected. The linear layer 802 can operate on the input dimension that corresponds to the frequency bands. In this case this is the first dimension of the input. In this description a single input is processed. In examples of the disclosure multiple inputs can be processed inparallel as a batch. In such examples the stacking of multiple inputs adds one dimension in frontof the actual data dimensions. The operation of the linear layer 802 provides an intermediate feature tensor Y 804 as an output.The metadata features pre-network 704A comprises multiple residual blocks (ResBlock) 806.The ResBlocks 806 comprise a stack of layers that is arranged so that the output of a given layeris taken and added to a subsequent layer deeper within the ResBlock 806. An exampleResBlock 806 is shown in Fig.9.Each ResBlock 806 has the stride settings (x,y) given within the block for example (str=1, 1). Thefirst number corresponds to the stride along the frequency axis (first dimension) and the second number corresponds to the stride along the temporal axis (second dimension). The number ofoutput channels (third dimension) from each ResBlock 806 is given as the last number in thetriplet following the respective ResBlocks 806 in Fig.8.The metadata features pre-network 704A can comprise a sequence of downwards ResBlocks 806A-806D and a sequence of further ResBlocks 806E-806H. The sequence of downwardsResBlocks 806A-806D can provide a downsampling part of the metadata features pre-network704A. The sequence of upwards ResBlocks 806AE-806H can provide a upsampling part of themetadata features pre-network 704A.In the example of Fig. 8 the first ResBlock 806A has (str=1, 1) with 8 output channels. Thesecond ResBlock 806B has (str=2, 2) with 16 output channels. The second ResBlock 806B isarranged to perform factor 2 sub-sampling in the spatial (time and frequency) dimensions. Thethird ResBlock 806C has (str=2, 2) with 32 output channels. The third ResBlock 806C isarranged to perform factor 2 sub-sampling in the spatial dimensions. The fourth ResBlock 806Dhas (str=2, 2) with 64 output channels. The fourth ResBlock 806D is arranged to perform factor2 sub-sampling in the spatial dimensions.The output of the last downwards ResBlock 806D is passed through a convolutional block(ConvBlock) 808. The ConvBlock 808 comprises a two-dimensional convolution with kernel size of (3, 1) for (frequency, time) with 128 output channels, followed by a Batch Normalization(BatchNorm) and a Rectified Linear Unit (ReLU) activation.The output of the ConvBlock 808 is passed through Transposed Convolution (TransConv) block810. The TransConv block 810 comprises a two-dimensional transposed convolution with kernelsize of (3, 1) and 64 output channels. This effectively up-samples along the frequency axis. Thisis followed by a BatchNorm and ReLU activation.The output of the TransConv block 810 is concatenated with the output of the last downwardResBlock 806D. The solid arrows in Fig. 8 represent concatenations. The concatenation of the output of the TransConv 810 with the output of the last downward ResBlock 806D is performedalong the channel dimension. The concatenation produces a tensor with shape (3, N / 8, 128).This tensor is provided as an input to the upsampling part of the metadata features pre-network704A.The upsampling part comprises a fifth ResBlock 806E followed by an Upsampling (2, 2) block812E. The Upsampling block 812E is arranged to upsample the two-dimensional layer with thespatial size scaler factor given in the parenthesis (frequency, time) using, for example, the nearestneighbor upsampling method. Therefore Upsampling (2, 2) increases the size of both spatialdimensions by a factor of two using nearest neighbor upsampling. The output of the Upsamplingis the output of this layer of the upsampling part of the metadata features pre-network 704A.The output of the Upsampling block 812E is concatenated with a corresponding matching skipconnection tensor. This comprises data from the downsampling layer. The concatenation isperformed along the feature dimension (third dimension in this example). The output of the concatenation is provided to a sixth ResBlock 806F and following Upsampling block 812F.The output of the Upsampling block 812F is concatenated with a corresponding matching skipconnection tensor. The concatenation is performed along the feature dimension (thirddimension in this example). The output of this concatenation is provided to a seventh ResBlock806G and following Upsampling block 812G.The convolutions in the upsampling part of the metadata features pre-network 704A have thenumber of output channels of 32, 16, and 8.The last layer of the upsampling part comprises an eight ResBlock 806H. the last ResBlock806H has five output channels and no following Upsampling blocks. The shape of the lastResBlock 806H provides the output 706A of the metadata features pre-network 704A. Theoutput 706A of the metadata features pre-network 704A is now (24, ^_^^^^^^, 5). In this examplethe output 706A of the metadata features pre-network 704A is the intermediate metadata features706A. Other U-Net structures can be arranged in a similar manner but would have different inputsand outputs.Fig.9 shows an example structure for a ResBlock 806. The ResBlock 806 could be used in theU-Net structures of the machine learning model 504. The example ResBlock 806 could be usedin a downsampling part. A ResBlock 806 with the same internal structure could also be used inthe upsampling part but different stride and kernel size settings could be used in the upsamplingpart. The internal structure of the ResBlock 806 has a pre-activation ordering. This means that BatchNorm layers 904 and the ReLU activation layers 906 are before the convolution operations 908. The ResBlock 806 comprises two paths for an input. The first path is shown on the left side ofFig. 9 and the second path is shown on the right side of Fig. 9. The path on the left sidecorresponds to the residual or by-pass path. The path on the right side corresponds to theconvolutional core.The residual or by-pass path comprises convolution operations 900 and a BatchNorm layer 902.The convolution operations 900 comprises a Conv2D layer with the kernel size of (1, 1) andstriding (x, y) matching the stride settings of the ResBlock 806. The convolution operations 900adjust the size of the feature dimension of the input to match the feature dimension of the lastconvolution in the convolution core. The convolution operations 900 are followed by the BatchNorm layer 902. The BatchNorm layer 902 comprises a BatchNorm2D layer that operates on the feature dimension. The convolutional core consists of a sequence of blocks. The sequence comprises a BatchNorm layer 904 followed by a ReLU activation layers 906 which is then followed by convolution operations 908. In this example the first BatchNorm layer 904A is a BatchNorm2D that operates on the feature dimension (third dimension). The first ReLU activation layer 906A is applied one each element from the BatchNorm layer 904A. The output of the ReLU activation layer 906A is passed to the first convolution operations 908A. The first convolutions operations 908A comprise a Conv2D with the kernel size of (3, 3) and the stride setting (x, y) matching the stride settings of the ResBlock 806.The output of the first convolution operations 908A is passed through a second BatchNorm layer904B and a second ReLU activation layer 906B before being passed to the second convolutionoperations 908B. The second convolution operations 908B processes the input with a Conv2D with kernel size of (3, 3) and stride of 1. In this example both of the convolution operations 908 use reflection padding of size 1 along all four sides of the input.The output of the residual or by-pass path and the output of the convolutional core are providedto an addition block 910. The addition block 910 adds the output of the residual or by-pass pathand the output of the convolutional core in an element wise manner. The output of the additionblock 910 is the output of the ResBlock 806. The SPAC feature pre-network 704B can have a similar structure to the metadata features pre- network 704A as shown in Fig.8 however different dimensions and settings could be used. The input to the SPAC feature pre-network 704B would be the SPAC features 702. This inputhas shape (60, ^_^^^^^^, 33). This input can be provided to a dimension adjustment 800 asshown in Fig.8. The dimension adjustment 800 can also comprise a fully connected linear layer 802. The dimension adjustment provides an intermediate feature tensor Y with shape(48, ^_^^^^^^, 33) .In the example SPAC feature pre-network 704B the downsampling part also comprises fourResBlocks 806A-806D. The first ResBlock 806A has (str=2, 2) with 64 output channels. Thesecond ResBlock 806B has (str=2, 2) with 128 output channels. The third ResBlock 806C has(str=2, 2) with 256 output channels. The fourth ResBlock 806D has (str=2, 2) with 512 outputchannels. In this example there is no downsampling along the temporal axis in the lastdownsampling layer.In the example SPAC feature pre-network 704B the output of the fourth ResBlock 806D is passedthrough a convolutional block (ConvBlock) 808. The ConvBlock 808 comprises a two- dimensional convolution with kernel size of (3, 1) for (frequency, time) with 1024 output channels, followed by a TransConv block 810. The TransConv Block 810 has a corresponding kernel size of (3, 1) and 256 output channels.In the example SPAC feature pre-network 704B the upsampling part also comprises a further fourResBlocks 806E-806H and corresponding Upsampling blocks 812E-812H. In this example thefifth ResBlock 806E has (str=1, 1) with 256 output channels followed by an Upsampling (2, 1)block 812E. The Upsampling block 812E does not provide any Upsampling in the temporal axis.The sixth ResBlock 806F has (str=1, 1) with 128 output channels followed by an Upsampling (2,2) block 812F. The seventh ResBlock 806G has (str=1, 1) with 64 output channels followed byan Upsampling (2, 2) block 812G. The eighth ResBlock 806H has (str=1, 1) with 20 outputchannels followed by an Upsampling (2, 1) block 812H. The Upsampling block 812H does notprovide any Upsampling in the temporal axis. This differs from the upsampling part of themetadata features pre-network 704A because the metadata features pre-network 704A does notcomprise an Upsampling block after the last ResBlock 806 in the upsampling part.The output of the SPAC feature pre-network 704B has shape (24, ^_^^^^^^, 20).The Cov feature pre-network 704C can have a similar structure to the metadata features pre-network 704A as shown in Fig.8 and also the SPAC feature pre-network 704B, however differentdimensions and settings could be used. For the case of the Cov feature pre-network 704C the dimension adjustment block 800 can be omitted and the input to the u-net structure would be aninput of shape (24, ^_^^^^^^^^^, 5).In the example Cov feature pre-network 704C the downsampling part also comprises fourResBlocks 806A-806D. The first ResBlock 806A has (str=1, 2) with 8 output channels. In thiscase there would be no downsampling along the frequency axis. The second ResBlock 806Bhas (str=2, 2) with 16 output channels. The third ResBlock 806C has (str=2, 2) with 32 outputchannels. The fourth ResBlock 806D has (str=2, 2) with 64 output channels.In the example SPAC feature pre-network 704B the output of the fourth ResBlock 806D is passedthrough a ConvBlock 808. The ConvBlock 808 comprises a two-dimensional convolution withkernel size of (3, 1) with 128 output channels, followed by a TransConv block 810. TheTransConv Block 810 has a corresponding kernel size of (3, 1) and 64 output channels.In the example Cov feature pre-network 704C the upsampling portion also comprises a furtherfour ResBlocks 806E-806H and corresponding Upsampling blocks 812E-812G. In this examplethe fifth ResBlock 806E has (str=1, 1) with 32 output channels followed by an Upsampling (2, 2)block 812E. The sixth ResBlock 806F has (str=1, 1) with 16 output channels followed by anUpsampling (2, 2) block 812F. The seventh ResBlock 806G has (str=1, 1) with 8 output channelsfollowed by an Upsampling (2, 2) block 812G. The eighth ResBlock 806H has (str=1, 2) with 8output channels. There is no Upsampling block following the eighth ResBlock 806H in the Covfeature pre-network 704C. This is similar to the metadata features pre-network 704A which alsodoes not comprise an Upsampling block after the last ResBlock 806 in the upsampling part. The Cov feature pre-network 704C differs from the metadata features pre-network 704A (and also the SPAC features pre-network 704B) in that, in the Cov feature pre-network 704C the last ResBlock 806H has downsampling along the temporal axis with the stride (1, 2) setting.The output of the Cov feature pre-network 704C has shape (24, ^_^^^^^^^^^ / 4,8), which isequal to (24, ^_^^^^^^, 8).Referring to Fig. 7 metadata feature pre-network 704A provides a metadata intermediaterepresentation 706A as an output, the SPAC feature pre-network 704B provides a SPAC intermediate representation 706B as an output, and the Cov feature pre-network 704C provides a Cov intermediate representation 706C as an output. Using the example U-Net structures ofFig.8 the metadata intermediate representation 706A has a shape (24, ^_^^^^^^, 5), the SPACintermediate representation 706B has shape (24, ^_^^^^^^, 20) , and the Cov intermediaterepresentation 706C has shape (24, ^_^^^^^^, 8).The intermediate representations 706A, 706B, 706C are provided to a concatenation block 708.The concatenation block 708 concatenates the intermediate representations 706A, 706B, 706Calong the feature axis (the third dimension in Fig. 7). The concatenation block 708 provides a combined intermediate representation 710 as an output. The combined intermediaterepresentation 710 has a shape (24, ^_^^^^^^, 33).The combined intermediate representation 710 is provided as an input to a fourth U-Net structure.The fourth U-Net structure is the combined prediction network 712. The combined prediction network 712 has a similar structure to the metadata features pre-network 704A as shown in Fig. 8 with some differences. For the case of the combined prediction network 712 the dimension adjustment block 800 can be omitted and the input to the u-net structure would be an input ofshape (24, ^_^^^^^^, 33).In the example combined prediction network 712 the downsampling portion also comprises fourResBlocks 806A-806D. The first ResBlock 806A has (str=1, 1) with 32 output channels. Thesecond ResBlock 806B has (str=2, 2) with 48 output channels. The third ResBlock 806C has(str=2, 2) with 64 output channels. The fourth ResBlock 806D has (str=2, 2) with 128 outputchannels.In the example combined prediction network 712 the output of the fourth ResBlock 806D is passedthrough a ConvBlock 808. The ConvBlock 808 comprises a two-dimensional convolution withkernel size of (3, 1) and 512 output channels, followed by a TransConv block 810 with acorresponding kernel size of (3, 1) and 256 output channels.In the example combined prediction network 712 the upsampling part also comprises a furtherfour ResBlocks 806E-806H and corresponding Upsampling blocks 812E-812G. In this examplethe fifth ResBlock 806E has (str=1, 1) with 64 output channels followed by an Upsampling (2, 2)block 812E. The sixth ResBlock 806F has (str=1, 1) with 48 output channels followed by anUpsampling (2, 2) block 812F. The seventh ResBlock 806G has (str=1, 1) with 12 outputchannels followed by an Upsampling (2, 2) block 812G. The eighth ResBlock 806H has (str=1,1) with 3 output channels. There is no Upsampling block following the eighth ResBlock 806H inthe combined prediction network 712. This is similar to the metadata features pre-network 704Awhich also does not comprise an Upsampling block after the last ResBlock 806 in the upsamplingpart.The output of the combined prediction network 712 has shape (24, ^_^^^^^^, 3). The output ofthe combined prediction network 712 provides pre-scale predicted metadata properties 714.These represent the directional metadata in each (24, ^_^^^^^^) TF-tiles in xyz vectorrepresentation ^^^^^^^^^(^, ^).The pre-scale predicted metadata properties 714 are provided as an input to an XYZ scale block 716. The XYZ scale block 716 is configured to apply appropriate scaling to the pre-scalepredicted metadata properties 714. The XYZ scale block 716 can apply final constraints on thelength of the vector. In some examples the scaling applied by the XYZ scale block 716 can comprise determining the length of the input vectors. The length of the input vectors can be determined with This is passed through hyperbolic tangent and scaled with a constant, for example, ^^^^ = 1.1for obtaining the scaled length (^, ^)^ Without the scaling, the output of the hyperbolic tangent would require infinite value of the inputfor the output to reach value of 1.0. When the model is used for inference the value of ^^^^^^^(^, ^)is limited to the range of 0…1.0 with ^^^^^^^^ (^, ^) = ma^ ^0, min^1, ^^^^^^^(^, ^)^^and this value is used in place of ^^^^^^^(^, ^).A scaling value ^(^, ^) can be determined from these two lengths with The output of the XYZ scale block 716 is the input multiplied by this scaling value The XYZ scale block 716 provides predicted spatial metadata properties 506 as an output which is the output of the machine learning model 504. The machine learning model 504 can be trained using any suitable process. In some examples the machine learning model 504 can be trained using a set of spatial audio items in MASA format comprising an audio signal such as a transport audio signal and MASA spatial metadata with one directional field. The audio signal can comprise two channels. It is assumed that the original spatial metadata has low temporal resolution and all four sub-frames in each frame contain the same values. The training data comprises 4530 items that are4 seconds in length (that is, 200 frames). The directional spatial metadata consisting of azimuth^^^(^, ^) , elevation ^^^(^, ^) , and direct-to-total energy ratio ^(^, ^) are transformed intoCartesian xyz vector representation using This is the reference or target data during the training. The validation data of 799 items is selected from this same pool of items.For the low-resolution model input, the training items are passed through an IVAS codec (encoderand decoder) operating at 13.2 kbps total bitrate (having about 2.25 kbps for the MASA metadata), and the decoded MASA metadata is obtained using the external renderer (EXT) output mode of the IVAS decoder. The IVAS codec reduces the frequency resolution of the spatial metadata into 5 bands from the original 24 bands and applies quantization to the values. This metadata is transformed into the Cartesian xyz vector representation in the same way as thereference data. The training uses batch size of 32, that is, the parameters of the model areadjusted after each 32 training examples. AdaDelta optimizer is used with learning rate of 1.0. The training is run for the maximum of 1000 epochs or until early stopping is triggered. The early stopping is triggered when the per-epoch validation loss is not lower than the best per-epoch validation loss in 50 consecutive epochs. The validation batch size is 8 items.The loss can be computed as follows. First, the direct-to-total energy ratios of the target and thepredicted data are computed Then, the absolute value of the energy ratio difference is computed Then, unit-length direction vectors are computed for the target and the predicted data Then, direction error vector is computed by and the length of the direction error vector is computed by This length is weighted by the target direct-to-total energy ratio Using the determined absolute value of the energy ratio difference and the determined weighted length of the direction error vector, the combined error measure is determined by Then, an energy-weighting metric is determined. The energy-weighting metric is for weightingthe loss based on the energies of the corresponding time-frequency tiles (that is, time-frequencytiles having larger energy should have a larger effect on the loss). The energy-weighting metriccan be formulated in any suitable manner. In some examples, the signal energy evolution overtime with frequency dependent weighting (i.e., ^^^^^^^^ , ^^) as computed by the audio featurecomputation block 602 may be used as the weight. As ^^^^^^^^ , ^^ is computed in subframes^^^ , and the loss is computed in frames ^ , the mean of the values ^^^^^^^^ , ^^ for thecorresponding frame ^ are computed and set as the weight where ^ is the frame index, ^^^ the subframe index, and ^^^ = 4 the number of subframes ina frame. Then, the energy-weighted combined error measure is determined by ^^^^^,^(^, ^) = ^^^^^(^, ^)^^^^^,^(^, ^)Using ^^^^^,^(^, ^), the loss is determined by computing the mean over time and frequency overthe entire training example which is the loss that is output from the loss function.In the example metadata processors 310 such as those shown in Figs. 5 and 6 the metadatadetermination block 508 receives the decoded spatial metadata 308 and predicted metadata properties 506 as an input. The decoded spatial metadata 308 can comprise all MASA metadata variables. The metadata determination block 508 is arranged to enhance the decoded spatial metadata 308 based on the predicted metadata properties 506 by updating the direction (azimuth and elevation) and the direct-to-total energy ratio based on the predicted metadata properties 506 that are generated by the machine learning model 504. In some examples the process of enhancing the decoded spatial metadata 308 can comprisetransforming the Cartesian xyz vector representation of the model prediction ^^^^^(^, ^) intoazimuth ^^^^^^(^, ^) , elevation ^^^^^^(^, ^) , and direct-to-total energy ratio^^^^(^, ^) representation with Here ^^^^2(∙) is the inverse tangent function resolving the correct quadrant. The per-frametemporal resolution of predicted metadata properties 506 is adjusted to the per-sub-frame temporal resolution of MASA metadata by repeating the same value for all four sub-frames in aframe. The diffuse-to-total energy ratio ^^^^^(^, ^) in the spatial metadata is adjusted to reflectthe new direct-to-total energy ratio by In this example, the surround and spread coherence parameter values are not adjusted, but the values from the low-resolution decoded spatial metadata 308 are used instead, by replicating the5 values to cover the all 24 bands. In some other examples the surround and / or spread coherenceparameter values can be adjusted. In some examples all parameter values can be obtained fromthe predicted metadata properties 506. In these examples the metadata determination block 508 does not need to have the decoded spatial metadata 308 as an input because all spatialmetadata values may be obtained from the predicted metadata properties 506.The architecture described above for the machine learning model 504 uses convolutions in time and temporal downsampling and upsampling operations. This means that the architecture is not causal and thus not suitable for low-latency applications. However, this example is only used for demonstrating that the assumptions of additional information in an audio signal are valid. It is possible to implement the machine learning model using other DNN architectures, some of whichare causal (not utilizing information from the future) and low-latency. One way to define a causalarchitecture is to restrict the downsampling and upsampling operations along the time axis toconsider only the current and past values instead of considering also future indices, and aligningthe by-pass paths between the blocks of downsampling and upsampling parts along the time axis.This alignment can be done, for example, by applying the padding only in the direction of pastsamples on the time axis, and aligning the sample corresponding to the most recent time indexafter temporal upsampling with the most recent time index at the corresponding by-pass path.Alternatively, the strides in the convolution operations along the time axis may be replaced usingdilated convolutions, or by implementing the temporal modelling using recurrent modelling. Afurther alternative is to omit temporal context modelling entirely by using convolution kernel sizeof 1, stride length of 1, and upsampling factor of 1 on the time axis instead of the values providedin the example embodiment. The examples of the disclosure provide processed spatial metadata as an output. The processed spatial metadata can be used for spatial synthesis. Any suitable methods can be used for the spatial synthesis.Fig. 10 shows another example system 100 that can be used in examples of the disclosure. Inthis example the metadata processor 310 is provided as a post processing block that is outside of, or separate to, the decoder 110. This arrangement can enable examples of the disclosure to be used with existing codecs without having to make any changes within the codec. The system 100 comprises an encoder 106 that receives audio signals 102 and spatial metadata 104 as an input. The encoder 106 provides a bitstream 108 that is sent from the encoder 106 to the decoder 110. The decoder 110 is arranged to decode the bitstream 108 to provide decoded audio signals 306 and decoded spatial metadata 308 as outputs. The decoded audio signals 306 and decoded spatial metadata 308 are provided as inputs to the metadata processor 310. The metadata processor 310 can comprise a machine learning model 504 as described herein orany other suitable means. The metadata processor 310 provides processed spatial metadata312 as an output. The processed decoded spatial metadata 312 and the decoded audio signals306 are provided as an input to a spatial synthesizer 314. The spatial synthesizer 314 isconfigured to process the processed decoded spatial metadata 312 and the decoded audiosignals 306 to render spatial audio output 316.Other variations to the system 100 and components of the system 100 can be made in examplesof the disclosure. For instance, in the examples the described the inputs to the encoder 106comprise audio signals 102 and spatial metadata 104 but other formats for the inputs could beused. In some examples the audio signals 102 and spatial metadata 104 could be generated inside the encoder 106. Similarly, different types of inputs could be provided to the metadata processor 310. Forexample, any associated audio signals and spatial metadata could be used and not necessarilydecoded signals. In the described examples the directional parameters of the spatial metadata are given as azimuthand elevation values. In other examples, the directional parameters can be received in otherformats, such as spherical indices, and can be converted to and from azimuths and elevationsand / or other suitable representations where appropriate. In the described examples, one set of audio-based features were used. Some of them may bemore useful with spaced microphones (such as the SPAC feature), while some of them may bemore useful with directional coincident microphones (such as the inter channel level differencefeature). These are merely example features, and in other examples, other kinds of audio-basedfeatures can be used additionally or instead of the described features.Figs. 11A to 11D show example results that can be obtained using examples of the disclosure.The results validate the performance of the examples of the disclosure. The experiments that were used to obtain the data shown in Figs. 11A to 11D used spatial metadata as described above and the original audio signals was used for the audio-based features. That is, the spatial metadata was passed through an IVAS codec for reducing the resolution, but the audio signals from the IVAS codec were replaced with the ideal transport audiosignals. The performance of the system 100 was measured by evaluating the loss function for thelow-resolution model input and then for the processed spatial metadata 512 from the output ofthe metadata processor 310. Fig.11A shows the loss function of the machine learning model 504 using metadata features 604 and audio features 606. The first horizontal line 1100 shows the loss function evaluated on the decoded spatial metadata 308 from which the machine learning model 504 input is determined. The second horizontal line 1102 shows the loss function evaluated on the machine learning model 504 for validation data. The third line 1104 shows the loss function for the machine learning model 504 output for the training data and the fourth line 1106 shows the loss function for themachine learning model 504 for the validation data.The data in Fig.11A shows that even though the training loss keeps on decreasing, the validation loss has converged, and training has been ended with early exit. Both training and validation dataexhibit a clearly lower loss value for the machine learning model 504 output compared to themachine learning model 504 input indicating that the processing has enhanced the spatialmetadata. The data in Fig. 11A shows that the examples of the disclosure are able to enhance the low- resolution spatial metadata with the help of an audio signal such as the transport audio signal as additional information source. This means that, the processed spatial metadata 312 obtained using examples of the disclosure is closer to the original high-resolution spatial metadata when the distance is measured with the loss function.In addition to this computational evaluation that provided the results shown in Figs.11A to 11D,an informal subjective evaluation was done. The informal subjective evaluation compared thebinaural rendering based on the ideal transport audio signals using the original high-resolution metadata, the low-resolution metadata from the model input, and the metadata obtained byapplying the example methods described herein. It was noticed that the enhanced version (thatis using the processed spatial metadata obtained as described herein) was perceptually spatiallycloser to the reference than the low-resolution input version. The directions of the sound sources were more correct, the sound source directions were more stable, and the spaciousness was closer to the reference.The subjective evaluation indicates that when using the processed spatial metadata 312 inbinaural rendering, the reproduced audio scene is perceptually closer to the original audio scene than when using the model input for the rendering. Figs.11B and 11C show data obtained in a second experiment. In the second experiment the performance of the machine learning model 504 as described was compared with partial models. The results in Fig.11B were obtained from a comparison using only spatial metadata features and the results in Fig.11C were obtained from a comparison using only audio signal-based features. In the comparison using only spatial metadata features the audio-based feature inputs (such as the SPAC features and the Cov features) are not processed. Only the metadata features pre-network 704A is active. In this case the combined intermediate representation 710 contains onlythe intermediate metadata features 706A and the number of input channels to the combinedprediction network 712 is adjusted accordingly. In Fig. 11B The first horizontal line 1110 shows the loss function evaluated on the machine learning model 504 for training data. The second horizontal line 1112 shows the loss function evaluated on the machine learning model 504 for validation data. The third line 1114 shows the loss function for the machine learning model 504 output for the training data and the fourth line 1116 shows the loss function for the machine learning model 504 for the validation data. In this example, even though the training loss keeps on decreasing, the validation loss has converged, and training has been ended with early exit. Both the training and validation dataexhibit clearly lower loss value at the machine learning model 504 output compared to themachine learning model 504 input meaning that the machine learning model 504 has been ableto enhance the metadata. However, the value of the loss shown in Fig.11B remains above thevalues shown in Fig.11A indicating that the machine learning model 504 has been able to obtainuseful information from the audio features. In the comparison using only audio-based features the metadata feature inputs. Only the SPAC features pre-network 704B and the Cov features pre-network 704C are active. In this case the combined intermediate representation 710 comprises the SPAC intermediate representation706B and the Cov intermediate representation 706C. The number of input channels to thecombined prediction network 712 is adjusted accordingly. This example that uses only audio- based features can be considered an example of a machine learning based metadata estimation from two-channel transport audio signals. In Fig. 11C The first horizontal line 1120 shows the loss function evaluated on the machine learning model 504 for training data. The second horizontal line 1122 shows the loss functionevaluated on the machine learning model 504 for validation data. The third line 1124 shows theloss function for the machine learning model 504 output for the training data and the fourth line 1126 shows the loss function for the machine learning model 504 for the validation data. In this example, even though the training loss keeps on decreasing, the validation loss has converged, and training has been ended with early exit. Both the training and validation dataexhibit clearly lower loss value at the machine learning model 504 output compared to themachine learning model 504 input meaning that the machine learning model 504 has been ableto enhance the metadata. However, the value of the loss shown in Fig.11C remains above thevalues shown in Fig.11A indicating that the machine learning model 504 has been able to obtainuseful information from having both the metadata features and the audio features. The results shown in Figs.11B and 11C indicate that using the metadata processor 310 as described but with just metadata features or with just audio features will still provide improved spatial metadata compared to the low-resolution input. However, the use of both metadata features 604 and audio feature 606 provides even better performance. This validates the assumption that there is useful information present in audio signals such as transport audio signals and this can be used for enhancing the spatial metadata. The same conclusion was done based on the informal subjective evaluation with binaural audio signals. A further experiment was performed in which the ideal transport audio signals were replaced with transport audio signals that have been encoded using IVAS codec in stereo mode at 64 kbps.This uses the modified discrete cosine transform (MDCT) stereo codec. The intention of thisexperiment was to use non-ideal transport audio signals so that the information from the transportaudio signals is closer to the ones used in real application. In Fig. 11D The first horizontal line 1130 shows the loss function evaluated on the machine learning model 504 for training data. The second horizontal line 1132 shows the loss function evaluated on the machine learning model 504 for validation data. The third line 1134 shows the loss function for the machine learning model 504 output for the training data and the fourth line 1136 shows the loss function for the machine learning model 504 for the validation data. Fig.11D shows that even though the training loss keeps on decreasing, the validation loss has converged, and training has been ended with early exit. Both the training and validation data exhibit clearly lower loss value at the machine learning model 504 output compared to the machine learning model 504 input indicating that the machine learning model 504 has been able to enhance the metadata even with the coded transport audio. The loss reaches lower values than when using only metadata in as shown in Fig. 11B. This indicates that there is useful information present even in the coded transport audio signals even though the performance of uncoded transport audio signals as shown in Fig.11A is not reached.Fig.12 shows an example device 1200 that can be used to implement examples of the disclosure.In this example the device 1200 is a mobile device with connected headphone 1218. Theheadphone 1218 can be connected to the device 1200 via a wired connection or a wirelessconnection 1214. Other types of devices 1200 and connected peripheral devices can be usedin other examples of the disclosure.The device comprises a processor 1206 and a memory 1210. The memory 1210 can compriseprogram code 1212. The memory 1210 can comprise a machine learning model 504 which can be arranged as described herein to generate processed spatial metadata. The processor 1206 and memory 1210 can be arranged to provide a decoder 110 as described above. The user 1220 of the device 1200 can be using the device to participate in an immersive call or other communication that uses spatial audio. The device 1200 obtains a bitstream 108. The bitstream 108 can comprise the bitstream 108 that is transferred between an encoder 106 and a decoder 108 as shown in Fig.1. The bitstream 108 can be received via a transceiver 1204 or canbe retrieved from storage 1202. In some examples the bitstream 108 can be received via thetransceiver 1204 and then stored to the storage and then accessed later by the processor 1206.The processor 1206 is arranged to convert the bitstream 108 to a spatial audio output 1208. Theprocessor 1206 uses the program code 1212 and the machine learning model 504 that are storedin the memory 1210 to convert the bitstream 108 to the spatial audio output 1208. In thisexample the spatial audio output 1208 is a binaural output. Other types of spatial audio output1208 can be used in other examples.The spatial audio output 1208 is provided to a headphone connection 1214. The headphoneconnection 1214 connects the headphone 1218 to the device 1200 and enables signals from the device 1200 to be provided to the headphone 1218. The headphone connection 1214 can be a wired connection or a wireless connection. The headphone connection 1214 provides an audio signal 1216 to the headphones to enable the spatial audio to be played back to the user 1220. In the example of Fig.12 the machine learning model 504 is stored in the memory 1210. The machine learning model 504 can be trained by an external device and then provided to the device 1200 so that it can be stored in the memory 1210. The external device that performs the training of the machine learning model 504 can have a higher processing capacity than the device 1200 shown in Fig. 12 or other devices that implement the spatial audio communications. For example, the external device that performs the training of the machine learning model 504 could comprise a workstation with multiple graphic processing units (GPUs) dedicated to the training of the machine learning model 504.Fig. 13 schematically illustrates an apparatus 1300 that can be used to implement examples ofthe disclosure. In this example the apparatus 1300 comprises a controller 1302. The controller 1302 can be a chip or a chip-set. In some examples the controller 1302 can be provided within a communications device such as telephone or teleconferencing device or any other suitable type of device that enables audio signals to be communicated.In the example of Fig. 13 the implementation of the controller 1302 can be as controller circuitry.In some examples the controller 1302 can be implemented in hardware alone, have certain aspects in software including firmware alone or can be a combination of hardware and software (including firmware).As illustrated in Fig. 13 the controller 1302 can be implemented using instructions that enablehardware functionality, for example, by using executable instructions of a computer program 1304 in a general-purpose or special-purpose processor 1206 that can be stored on a computer readable storage medium (disk, memory etc.) to be executed by such a processor 1206. The processor 1206 is configured to read from and write to the memory 1210. The processor 1206 can also comprise an output interface via which data and / or commands are output by the processor 1206 and an input interface via which data and / or commands are input to the processor 1206.The memory 1210 is configured to store a computer program 1304 comprising computer programinstructions (computer program code 1212) that controls the operation of the controller 1302 when loaded into the processor 1206. The computer program instructions, of the computer program 1304, provide the logic and routines that enables the controller 1302 to perform the methodsillustrated in the Figs. The processor 1206 by reading the memory 1210 is able to load andexecute the computer program 1304. The apparatus 1300 therefore comprises: at least one processor 1206; and at least one memory1210 including computer program code 1304, the at least one memory 1210 and the computerprogram code 1212 configured to, with the at least one processor 1206, cause the apparatus1300 at least to perform: receiving 400 an encoded audio signal and corresponding encoded spatial metadata; determining 402 at least one input for processing wherein the input is based on the encoded audio signal and the encoded spatial metadata; processing 404 the at least one input to use information from the encoded audio signaland the encoded spatial metadata to generate processed spatial metadata; andenabling 406 spatial rendering to be performed on one or more audio signals using theprocessed spatial metadata.As illustrated in Fig. 13 the computer program 1304 can arrive at the controller 1300 via anysuitable delivery mechanism 1306. The delivery mechanism 1306 can be, for example, a machinereadable medium, a computer-readable medium, a non-transitory computer-readable storage medium, a computer program product, a memory device, a record medium such as a Compact Disc Read-Only Memory (CD-ROM) or a Digital Versatile Disc (DVD) or a solid state memory, an article of manufacture that comprises or tangibly embodies the computer program 1304. Thedelivery mechanism 1306 can be a signal configured to reliably transfer the computer program1304. The controller 1302 can propagate or transmit the computer program 1304 as a computerdata signal. In some examples the computer program 1304 can be transmitted to the controller1302 using a wireless protocol such as Bluetooth, Bluetooth Low Energy, Bluetooth Smart, 6LoWPan (IPv6 over low power personal area networks) ZigBee, ANT+, near field communication (NFC), Radio frequency identification, wireless local area network (wireless LAN) or any other suitable protocol.The computer program 1304 comprises computer program instructions that when executed by anapparatus 1300 cause the apparatus 1300 to perform at least the following: receiving 400 an encoded audio signal and corresponding encoded spatial metadata; determining 402 at least one input for processing wherein the input is based on the encoded audio signal and the encoded spatial metadata; processing 404 the at least one input to use information from the encoded audio signaland the encoded spatial metadata to generate processed spatial metadata; andenabling 406 spatial rendering to be performed on one or more audio signals using theprocessed spatial metadata.The computer program instructions can be comprised in a computer program 1304, a non- transitory computer readable medium, a computer program product, a machine readable medium. In some but not necessarily all examples, the computer program instructions can be distributed over more than one computer program 1304.Although the memory 1210 is illustrated as a single component / circuitry it can be implemented asone or more separate components / circuitry some or all of which can be integrated / removable and / or can provide permanent / semi-permanent / dynamic / cached storage. Although the processor 1206 is illustrated as a single component / circuitry it can be implemented as one or more separate components / circuitry some or all of which can be integrated / removable. The processor 1206 can be a single core or multi-core processor. References to “computer-readable storage medium”, “computer program product”, “tangibly embodied computer program” etc. or a “controller”, “computer”, “processor” etc. should be understood to encompass not only computers having different architectures such as single / multi- processor architectures and sequential (Von Neumann) / parallel architectures but also specialized circuits such as field-programmable gate arrays (FPGA), application specific circuits (ASIC), signal processing devices and other processing circuitry. References to computer program, instructions, code etc. should be understood to encompass software for a programmable processor or firmware such as, for example, the programmable content of a hardware device whether instructions for a processor, or configuration settings for a fixed-function device, gate array or programmable logic device etc. As used in this application, the term “circuitry” can refer to one or more or all of the following: (a) hardware-only circuitry implementations (such as implementations in only analog and / or digital circuitry) and (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions and (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g. firmware) for operation, but the software cannot be present when it is not needed for operation. This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit for a mobile device or a similar integrated circuit in a server, a cellular network device, or other computing or network device.The apparatus 1300 as shown in Fig. 13 can be provided within any suitable device. In someexamples the apparatus 1300 can be provided within an electronic device such as a mobile telephone, a teleconferencing device, a camera, a computing device, a server or any other suitable device. The blocks illustrated in the Figs. can represent steps in a method and / or sections of code in the computer program 1304. The illustration of a particular order to the blocks does not necessarily imply that there is a required or preferred order for the blocks and the order and arrangement of the blocks can be varied. Furthermore, it can be possible for some blocks to be omitted. The apparatus can be provided in an electronic device, for example, a mobile terminal, according to an example of the present disclosure. It should be understood, however, that a mobile terminal is merely illustrative of an electronic device that would benefit from examples of implementations of the present disclosure and, therefore, should not be taken to limit the scope of the presentdisclosure to the same. While in certain implementation examples, the apparatus can be providedin a mobile terminal, other types of electronic devices, such as, but not limited to: mobile communication devices, hand portable electronic devices, wearable computing devices, portable digital assistants (PDAs), pagers, mobile computers, desktop computers, televisions, gaming devices, laptop computers, cameras, video recorders, GPS devices and other types of electronicsystems, can readily employ examples of the present disclosure. Furthermore, devices canreadily employ examples of the present disclosure regardless of their intent to provide mobility. The term ‘comprise’ is used in this document with an inclusive not an exclusive meaning. That isany reference to X comprising Y indicates that X may comprise only one Y or may comprise morethan one Y. If it is intended to use ‘comprise’ with an exclusive meaning then it will be made clearin the context by referring to ‘comprising only one...’ or by using ‘consisting.’In this description, the wording ‘connect’, ‘couple’ and ‘communication’ and their derivatives meanoperationally connected / coupled / in communication. It should be appreciated that any number orcombination of intervening components can exist (including no intervening components), i.e., toprovide direct or indirect connection / coupling / communication. Any such intervening componentscan include hardware and / or software components. As used herein, the term "determine / determining" (and grammatical variants thereof) can include, not least: calculating, computing, processing, deriving, measuring, investigating, identifying,looking up (for example, looking up in a table, a database, or another data structure), ascertainingand the like. Also, "determining" can include receiving (for example, receiving information),accessing (for example, accessing data in a memory), obtaining and the like. Also, "determine / determining" can include resolving, selecting, choosing, establishing, and the like.In this description, reference has been made to various examples. The description of features orfunctions in relation to an example indicates that those features or functions are present in that example. The use of the term ‘example’ or ‘for example’ or ‘can’ or ‘may’ in the text denotes, whether explicitly stated or not, that such features or functions are present in at least the described example, whether described as an example or not, and that they can be, but are not necessarily,present in some of or all other examples. Thus ‘example’, ‘for example’, ‘can’, or ‘may’ refers to aparticular instance in a class of examples. A property of the instance can be a property of only that instance or a property of the class or a property of a sub-class of the class that includes somebut not all the instances in the class. It is therefore implicitly disclosed that a feature describedwith reference to one example but not with reference to another example, can where possible beused in that other example as part of a working combination but does not necessarily have to beused in that other example. As used herein, “at least one of the following: ” and “at least one of ” and similar wording, wherethe list of two or more elements are joined by “and” or “or” mean at least any one of the elements,or at least any two or more of the elements, or at least all the elements. Although examples have been described in the preceding paragraphs with reference to various examples, it should be appreciated that modifications to the examples given can be made without departing from the scope of the claims. Features described in the preceding description may be used in combinations other than thecombinations explicitly described above.Although functions have been described with reference to certain features, those functions may be performable by other features whether described or not. The description of a feature, such as an apparatus or a component of an apparatus, configured to perform a function, or for performing a function, should additionally be considered to alsodisclose a method of performing that function. For example, description of an apparatusconfigured to perform one or more actions, or for performing one or more actions, should additionally be considered to disclose a method of performing those one or more actions with or without the apparatus. Although features have been described with reference to certain examples, those features mayalso be present in other examples whether described or not.The term ‘a’, ‘an’ or ‘the’ is used in this document with an inclusive not an exclusive meaning.That is any reference to X comprising a / an / the Y indicates that X may comprise only one Y ormay comprise more than one Y unless the context clearly indicates the contrary. If it is intendedto use ‘a’, ‘an’ or ‘the’ with an exclusive meaning then it will be made clear in the context. In somecircumstances the use of ‘at least one’ or ‘one or more’ may be used to emphasis an inclusivemeaning but the absence of these terms should not be taken to infer any exclusive meaning.The presence of a feature (or combination of features) in a claim is a reference to that feature or(combination of features) itself and to features that achieve substantially the same technical effect(equivalent features). The equivalent features include, for example, features that are variants and achieve substantially the same result in substantially the same way. The equivalent features include, for example, features that perform substantially the same function, in substantially the same way to achieve substantially the same result. In this description, reference has been made to various examples using adjectives or adjectivalphrases to describe characteristics of the examples. Such a description of a characteristic inrelation to an example indicates that the characteristic is present in some examples exactly as described and is present in other examples substantially as described. The above description describes some examples of the present disclosure however those of ordinary skill in the art will be aware of possible alternative structures and method features which offer equivalent functionality to the specific examples of such structures and features described herein above and which for the sake of brevity and clarity have been omitted from the abovedescription. Nonetheless, the above description should be read as implicitly including referenceto such alternative structures and method features which provide equivalent functionality unless such alternative structures or method features are explicitly excluded in the above description of the examples of the present disclosure. Whilst endeavoring in the foregoing specification to draw attention to those features believed tobe of importance the Applicant may seek protection via the claims in respect of any patentablefeature or combination of features hereinbefore referred to and / or shown in the drawings whether or not emphasis has been placed thereon. I / we claim:
Claims
CLAIMS1. An apparatus comprising means for:receiving an encoded audio signal and corresponding encoded spatial metadata; determining at least one input for processing wherein the input is based on the encoded audio signal and the encoded spatial metadata; processing the at least one input to use information from the encoded audio signal andthe encoded spatial metadata to generate processed spatial metadata; andenabling spatial rendering to be performed on one or more audio signals using theprocessed spatial metadata.
2. An apparatus as claimed in claim 1, wherein the means are for:determining a first input for processing wherein the first input is based on the encoded audio signal; determining a second input for processing wherein the second input is based on the encoded spatial metadata; and processing the first input and the second input to use information from the encoded audiosignal and the encoded spatial metadata to generate processed spatial metadata.
3. An apparatus as claimed in any preceding claim, wherein determining an input based onthe encoded audio signal comprises at least partially decoding the encoded audio signal.
4. An apparatus as claimed in any preceding claim, wherein determining an input based onthe encoded spatial metadata comprises at least partially decoding the encoded spatial metadata.
5. An apparatus as claimed in any preceding claim, wherein the processed spatial metadatacomprises at least one of: improved temporal resolution compared to the received spatial metadata; improved frequency resolution compared to the received spatial metadata; or improved accuracy of restoration of information from spatial metadata from which the encoded spatial metadata was obtained compared to the received spatial metadata.
6. An apparatus as claimed in any preceding claim, wherein the processed spatial metadatacomprises one or more energy-related parameters and one or more directional parameters.
7. An apparatus as claimed in claim 6, wherein the energy-related parameters comprise atleast one of: adirect-to-total energy ratio; anda diffuse-to-total energy ratio.
8. An apparatus as claimed in any of claims 6 to 7, wherein the directional parameterscomprise at least one of: an azimuth angle; an elevation angle; anda direction index.
9. An apparatus as claimed in any preceding claim, wherein the means are for obtaining oneor more features from the received encoded audio signal and using the one or more obtained features as the input to the processing.
10. An apparatus as claimed in any of claims 1 to 8, wherein processing the at least one inputcomprises obtaining one or more features from the received encoded audio signal.
11. An apparatus as claimed in any preceding claim, wherein the means are for obtaining oneor more features from the received encoded spatial metadata and using the one or more obtained features as, at least part of, the input for processing.
12. An apparatus as claimed in any of claims 1 to 10, wherein processing the at least oneinput comprises obtaining one or more features from the received encoded spatial metadata13. An apparatus as claimed in any preceding claim, wherein the processing the at least oneinput comprises obtaining predicted spatial metadata properties and using the predicted spatialmetadata properties and the received encoded spatial metadata to generate the processed spatialmetadata.
14. An apparatus as claimed in any preceding claim, wherein the spatial rendering comprisesat least one of: binaural rendering; multi-loudspeaker rendering; Ambisonics rendering; stereo rendering; orcross talk-cancelled stereo rendering.
15. An apparatus as claimed in any preceding claim, wherein the processing is performed, atleast in part, using a machine learning model.
16. An apparatus as claimed in claim 15, wherein the machine learning model comprises adeep neural network model.
17. An apparatus as claimed in any preceding claim, wherein the one or more audio signalsare obtained by decoding the received encoded audio signal.
18. An apparatus as claimed in any preceding claim, wherein the one or more audio signalscomprise one or more channels.
19. An apparatus as claimed in any preceding claim, wherein the one or more encoded audiosignals comprise one or more channels.
20. A method comprising:receiving an encoded audio signal and corresponding encoded spatial metadata; determining at least one input for processing wherein the input is based on the encoded audio signal and the encoded spatial metadata; processing the at least one input to use information from the encoded audio signal andthe encoded spatial metadata to generate processed spatial metadata; andenabling spatial rendering to be performed on one or more audio signals using theprocessed spatial metadata.
21. A computer program comprising instruction which, when executed by a processor, causethe processor to perform: receiving an encoded audio signal and corresponding encoded spatial metadata; determining at least one input for processing wherein the input is based on the encoded audio signal and the encoded spatial metadata; processing the at least one input to use information from the encoded audio signal andthe encoded spatial metadata to generate processed spatial metadata; andenabling spatial rendering to be performed on one or more audio signals using theprocessed spatial metadata.
22. An apparatus comprising at least one processor; and at least one memory includingcomputer program code; the at least one memory storing instructions that, when executed by theat least one processor, cause the apparatus to: receive an encoded audio signal and corresponding encoded spatial metadata; determine at least one input for processing wherein the input is based on the encoded audio signal and the encoded spatial metadata; process the at least one input to use information from the encoded audio signal and theencoded spatial metadata to generate processed spatial metadata; andenable spatial rendering to be performed on one or more audio signals using theprocessed spatial metadata.
Citation Information
Patent Citations
Multi-audio object coding and decoding method applied to low bit rate
CN113096672A
Apparatus, methods and computer programs for enabling rendering of spatial audio
WO2023148426A1
Parametric spatial audio rendering
WO2023156176A1
Binaural audio rendering of spatial audio
WO2024115045A1