Spatial metadata for rendering at spatial audio

By leveraging redundancies and machine learning models to encode spatial metadata, the method addresses the inefficiencies in spatial audio encoding, ensuring accurate and high-quality spatial audio rendering at reduced bitrates.

GB2643269APending Publication Date: 2026-02-11NOKIA TECHNOLOGIES OY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
GB2024011718
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-08
Publication Date
2026-02-11

AI Technical Summary

Technical Problem

Existing spatial audio encoding methods face challenges in efficiently compressing spatial metadata while maintaining perceptual quality, as low bitrates lead to instability and inaccuracies in sound source perception, and allocating more bits to spatial metadata reduces audio quality due to bitrate constraints.

Method used

Utilizing redundancies between audio signals and spatial metadata to generate compact spatial metadata, encoded with machine learning models, allowing for a bitstream with fewer bits while maintaining accurate spatial audio rendering.

Benefits of technology

The method achieves efficient use of bitstream resources by reducing spatial metadata bits, improving spatial audio rendering accuracy and maintaining audio quality, even at low bitrates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A Metadata Assisted Spatial Audio (MASA) encoder (for eg. Immersive Voice Audio Services IVAS) encodes audio objects in an audiostream by generating processed spatial metadata formed into a stream hav
Need to check novelty before this filing date? Find Prior Art

Description

Metadata assisted spatial audio (MASA) uses one or more audio signals together with corresponding spatial metadata to generate spatial audio. The spatial metadata can comprise direction information, direct-to-total energy ratios in frequency bands, or any other suitable information. A MASA stream can be obtained by capturing spatial audio with microphones and estimating the corresponding spatial metadata from the microphone signals. BRIEF SUMMARY According to various, but not necessarily all, examples of the disclosure there is provided an apparatus for encoding spatial audio signals comprising means for: obtaining one or more audio signals and associated spatial metadata; determining at least one input for processing wherein the at least one input is based on the obtained one or more audio signals and on the obtained spatial metadata; processing the at least one input to generate processed spatial metadata wherein the processed spatial metadata is formed into a bitstream using a fewer number of bits compared to the required number of bits for the obtained spatial metadata; and forming the bitstream based on the processed spatial metadata and encoded one or more audio signals. Determining at least one input for processing may comprise: determining a first input for processing wherein the first input is based on the obtained one or more audio signals; and determining a second input for processing wherein the second input is based on the obtained spatial metadata. Forming the bitstream comprising the processed spatial metadata and encoded one or more audio signals may comprise encoding the obtained one or more audio signals and combining the encoded one or more audio signals with the processed spatial metadata generated by the processing. The processing may provide a bitstream comprising the processed spatial metadata and encoded one or more audio signals as an output. Determining the at least one input for the processing may comprise encoding the obtained one or more audio signals. Determining the at least one input for the processing may comprise at least partially decoding encoded one or more audio signals. Determining the at least one input for the processing may comprise determining one or more features by processing at least one of: the obtained one or more audio signals; one or more decoded audio signals; one or more partially decoded audio signals. The means may be for enabling transmission of the bitstream comprising the processed spatial metadata and encoded one or more audio signals. The processed spatial metadata may be generated so as to be decoded by decoding processing in an apparatus for decoding spatial audio signals. The encoded one or more audio signals may be obtained by encoding the obtained one or more audio signals. The processing may be performed, at least in part, using a machine learning model. The machine learning model may comprise a deep neural network model. The machine learning model may be trained in conjunction with a decoding machine learning model for an apparatus for decoding spatial audio signals. According to various, but not necessarily all, examples of the disclosure there is provided a method comprising: obtaining one or more audio signals and associated spatial metadata; determining at least one input for processing wherein the at least one input is based on the obtained one or more audio signals and on the obtained spatial metadata; processing the at least one input to generate processed spatial metadata wherein the processed spatial metadata is formed into a bitstream using a fewer number of bits compared to the required number of bits for the obtained spatial metadata; and forming the bitstream based on the processed spatial metadata and encoded one or more audio signals. According to various, but not necessarily all, examples of the disclosure there is provided a computer program comprising instructions which, when executed by a processor, cause the processor to perform: obtaining one or more audio signals and associated spatial metadata; determining at least one input for processing wherein the at least one input is based on the obtained one or more audio signals and on the obtained spatial metadata; processing the at least one input to generate processed spatial metadata wherein the processed spatial metadata is formed into a bitstream using a fewer number of bits compared to the required number of bits for the obtained spatial metadata; and forming the bitstream based on the processed spatial metadata and encoded one or more audio signals. According to various, but not necessarily all, examples of the disclosure there is provided an apparatus for decoding spatial audio signals comprising means for: receiving a bitstream comprising encoded one or more audio signals and processed spatial metadata wherein the processed spatial metadata has been formed into a bitstream using a fewer number of bits compared to the required number of bits for originally obtained spatial metadata; determining at least one input for decoding processing wherein the at least one input is based on the encoded one or more audio signals and on the processed spatial metadata; performing decoding processing on the at least one input to generate decoded spatial metadata; and using the decoded spatial metadata and decoded one or more audio signals to enable rendering of a spatial audio output. Determining at least one input for decoding processing may comprise: determining a first input for a decoding processing based on the encoded one or more audio signals; and determining a second input for the decoding processing based on the processed spatial metadata; Determining the at least one input for the decoding processing may comprise decoding the encoded one or more audio signals received in the bitstream. Determining the at least one input for the decoding processing comprises processing the encoded one or more audio signals to determine one or more features. The spatial audio output may comprise at least one of: binaural output; multi-loudspeaker output; Ambisonics output; stereo output; and cross talk-cancelled stereo output. The processed spatial metadata may be generated by processing in an apparatus for encoding spatial audio signals. The decoding processing may be performed, at least in part, using a machine learning model. The machine learning model may comprise a deep neural network model. The machine learning model may be trained in conjunction with a machine learning model for an apparatus for encoding spatial audio signals. According to various, but not necessarily all, examples of the disclosure there is provided a method comprising: receiving a bitstream comprising encoded one or more audio signals and processed spatial metadata wherein the processed spatial metadata has been formed into a bitstream using a fewer number of bits compared to the required number of bits for originally obtained spatial metadata; determining at least one input for decoding processing wherein the at least one input is based on the encoded one or more audio signals and on the processed spatial metadata; performing decoding processing on the at least one input to generate decoded spatial metadata; and using the decoded spatial metadata and decoded one or more audio signals to enable rendering of a spatial audio output. According to various, but not necessarily all, examples of the disclosure there is provided a computer program comprising instructions which, when executed by a processor, cause the processor to perform: receiving a bitstream comprising encoded one or more audio signals and processed spatial metadata wherein the processed spatial metadata has been formed into a bitstream using a fewer number of bits compared to the required number of bits for originally obtained spatial metadata; determining at least one input for decoding processing wherein the at least one input is based on the encoded one or more audio signals and on the processed spatial metadata; performing decoding processing on the at least one input to generate decoded spatial metadata; and using the decoded spatial metadata and decoded one or more audio signals to enable rendering of a spatial audio output. According to various, but not necessarily all, embodiments there is provided an apparatus comprising at least one processor; and at least one memory including computer program code; the at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to perform at least a part of one or more methods described herein. According to various, but not necessarily all, embodiments there is provided an apparatus comprising means for performing at least part of one or more methods described herein. The description of a function and / or action should additionally be considered to also disclose any means suitable for performing that function and / or action. Functions and / or actions described herein can be performed in any suitable way using any suitable method. According to various, but not necessarily all, embodiments there is provided examples as claimed in the appended claims. While the above examples of the disclosure and optional features are described separately, it is to be understood that their provision in all possible combinations and permutations is contained within the disclosure. It is to be understood that various examples of the disclosure can comprise any or all the features described in respect of other examples of the disclosure, and vice versa. Also, it is to be appreciated that any one or more or all the features, in any combination, may be implemented by / comprised in / performable by an apparatus, a method, and / or computer program instructions as desired, and as appropriate. The description of a function should additionally be considered to also disclose any means suitable for performing that function BRIEF DESCRIPTION Some examples will now be described with reference to the accompanying drawings in which: FIG. 1 shows an example system; FIG. 2 shows an example encoder device; FIG. 3 shows an example processor; FIG. 4 shows an example encoder; FIG. 5 shows an example method; FIG. 6 shows an example decoder device; FIG. 7 shows an example decoder; FIG. 8 shows an example method; FIG. 9 shows an example metadata encoder; FIG. 10 shows an example method; FIG. 11 shows an example metadata decoder; FIG. 12 shows an example method; FIG. 13 shows training of example machine learning models; FIGS. 14A and 14B show example combination blocks; FIG. 15 shows an example encoder machine learning model; FIG. 16 shows an example decoder machine learning model; FIG. 17 shows example results; FIG. 18 shows an example apparatus. The figures are not necessarily to scale. Certain features and views of the figures can be shown schematically or exaggerated in scale in the interest of clarity and conciseness. For example, the dimensions of some elements in the figures can be exaggerated relative to other elements to aid explication. Corresponidng reference numerals are used in the figures to designate corresponding features. For clarity, all reference numerals are not necessarily displayed in all figures. DETAILED DESCRIPTION Fig. 1 shows an example system 100 that can be used for spatial audio. The system 100 comprises an encoder 106 and a decoder 110. The encoder 106 and decoder 110 can be in different devices. For example, the system 100 can be part of a telecommunications system in which the encoder 106 can be provided within a first communications device and the decoder 110 can be provided within a different communications device. The encoder 106 receives a spatial audio stream as an input. The spatial audio stream comprises audio signals 102 and spatial metadata 104. The spatial audio stream can be provided in a metadata-assisted spatial audio (MASA) format or in any other suitable format. In some examples the audio signals 102 can comprise microphone signals or signals that are obtained by processing microphone signals. In some examples the audio signals 102 can originate from other inputs such as multi-channel audio signals (for example, 5.1 or 7.1+4) or audio objects. In such cases the audio signals 102 can comprise a downmix of the originally obtained inputs. The audio signals 102 can comprise transport audio signals. The transport audio signals are audio signals that have been processed into a format for transport or transmission to another communication device. The spatial metadata 104 comprises information that can be used to render the audio signals 102 to generate spatial audio. The spatial metadata 104 can comprise direction information, direct-to-total energy ratios in frequency bands, or any other suitable information. The spatial metadata 104 can be estimated from the microphone signals or from any other suitable signals. The encoder 106 is configured to encode the audio signals 102 and the spatial metadata 104 to form a bitstream 108. The bitstream 108 can be sent to the decoder 110. The encoder 106 and the decoder 110 can be in different devices and the bitstream 108 can be sent via any suitable communication network. The decoder 110 receives the bitstream 108 as an input. The decoder 110 is configured to decode the bitstream 108 to render spatial audio output 112. The spatial audio output can comprise binaural audio signals, stereo audio signals, multi-channel audio signals (e.g., 5.1 or 7.1+4), Ambisonics signals, or any other suitable types of signals. The encoding methods that are used to encode the spatial metadata 104 control the level of compression depending on the bitrate. Using a lower bitrate can reduce the reconstruction accuracy of the spatial metadata 104 at the decoder 110. With a low bitrate a higher compression factor is needed and this is obtained by trading off some of the reconstruction accuracy. For example, at a very low bitrates (such as 13.2 kbps with IVAS), the bitrate that can be allocated to spatial metadata 104 encoding is so low (such as about 2.25 kbps with IVAS) that it adversely affects the perceptual quality of the spatial audio output 112 at the decoder 110. These adverse effects can include instability and inaccuracies of the perceived sound sources. Allocating more bits to the spatial metadata 104 does not resolve this issue because there are constraints on the total bitrate. Increasing the bits used for the spatial metadata 104 would reduce the bits available for the coding of the audio signals 102. This could result in a perceivable degradation of audio quality. There are some redundancies between the audio signals 102 and the spatial metadata 104. For example, when there is an onset in the audio (such as, when someone starts to speak), typically both the energy of the audio signal 102 and the direct-to-total energy ratio of the spatial metadata increase rapidly in the time-frequency tiles at the onset. Similarly, if there is a dominant audio source on the left side, the directional spatial metadata at those time-frequency tiles points to the left and correspondingly the left audio signal is louder than the right audio signal Examples of the disclosure make use of these redundancies to enable compact spatial metadata to be obtained. The compact spatial metadata can use a fewer number of bits when formed into a bitstream compared to the original spatial metadata. This provides for a more efficient use of the bitstream 108. The compact spatial metadata can be decoded using the audio signals 102 (or signals based on the audio signals) to generate decoded spatial metadata that can be used for rendering of the spatial audio output 112. Fig. 2 shows an example encoder device 200 that can be used in examples of the disclosure. The encoder device 200 can be any communications device that comprises an encoder. In examples a communications device can be both an encoder device 200 and a decoder device however just the components relevant to encoding are shown in Fig. 2. In this example the encoder device 200 is a mobile device. Other types of devices can be used as an encoder device 200 in other examples of the disclosure. The example encoder device 200 comprises multiple microphones 202, a processor 206, a memory 208, storage 216 and a transceiver 214. The encoder device 200 can also comprise other components that are not shown in Fig. 2 such as a camera or any other suitable components. The microphones 202 can comprise any means that can detect acoustic signals and convert them to microphone audio signals 204. The microphones 202 can be used to capture a sound environment. In the example of Fig. 2 the encoder device 200 comprises three microphones 202. A first microphone 202A is provided at a first edge of the encoder device 200, a second microphone 202B is provided at a second edge of the encoder device 200, and a third microphone 202C is provided at the rear of the encoder device 200. The third microphone 202C could be located near to a camera of the encoder device 200. Other numbers and arrangements of the microphones 202 could be used in other examples. For example, an encoder device 200 could be arranged to receive the microphone audio signals 204 from an external microphone arrangement such as an Eigenmike. The encoder device 200 is arranged so that the microphones 202 provide microphone audio signals 204 to the processor 206. The microphone audio signals 204 can be provided in any suitable format, for example the microphone audio signals 204 can be provided in a digital format such as pulse code modulation (PCM) format. In some examples the microphones 202 can comprise analogue microphones. In such examples an analogue-to-digital converter is provided between the microphones 202 and the processor 206. The processor 206 and the memory 208 are arranged to convert the microphone audio signals 204 to audio signals 102 and corresponding spatial metadata 104. The processor 206 and memory 208 can provide an encoder 106 as shown in Fig. 1. The audio signals 102 can be transport audio signals that are in a suitable format for transmission to a decoder device or any other suitable device. The processor 206 and memory 208 can also perform other functions such as encoding of the audio signals 102 and spatial metadata 104. The processor 206 and memory 208 can be arranged to use examples of the disclosure to generate the spatial metadata 104. Any suitable information can be stored in the memory 208 and accessed by the processor 206. In examples of the disclosure the memory 208 can comprise a program code 210 and a machine learning model 212. The machine learning model 212 can be as described below and can be used to generate processed spatial metadata. The machine learning model 212 is stored in the memory 208. The machine learning model 212 can be trained by an external device and then provided to the encoder device 200 so that it can be stored in the memory 208. The external device that performs the training of the machine learning model 212 can have a higher processing capacity than the encoder device 200 shown in Fig. 2 or other devices that implement the spatial audio communications. For example, the external device that performs the training of the machine learning model 212 could comprise a workstation with multiple graphic processing units (GPUs) dedicated to the training of the machine learning model 212. The processor 206 runs the program code 210 and the machine learning model 212 and generates a bitstream 108. The bitstream 108 can be provided to storage 216 for later use and / or can be provided to the transceiver 214 for transmission to another communications device or to a server. In some examples the audio-based bitstream 108 that is generated by the processor 206 can be accompanied with a video bitstream. The video bitstream can be based on images captured by one or more cameras of the encoder device 200. Fig. 3 shows an example operation of a processor 206 of an encoder device 200. The processor 206 receives the microphone audio signals 204 as an input. The microphone audio signals 204 are provided to a spatial metadata analyzer 300 and also to a microphone signal preprocessor 302. The spatial metadata analyzer 300 can use any suitable procedure to determine spatial metadata 104. The spatial metadata analyzer 300 provides the spatial metadata 104 as an output. The spatial metadata 104 can be provided in any suitable format. In some examples the spatial metadata 104 can comprise directions and direct-to-total energy ratios in frequency bands. The microphone signal preprocessor 302 can use any suitable procedure to generate audio signals 102. The audio signals 102 can comprise transport audio signals that are suitable for being transmitted to another communication device, or any other suitable type of audio signals. In some examples the microphone signal preprocessor 302 can select two microphones 202 from the available microphones and perform processing operations on the microphone audio signals 204 from the selected microphones 202. In some examples the microphones 202A, 202B on opposing edges of the encoder device 200 could be selected as the microphones. Example operations that can be performed by the microphone signal preprocessor 302 can comprise equalization, automatic gain control, microphone noise removal, wind noise removal, speech enhancement, or any other suitable operation or combination of operations. The spatial metadata 104 and the audio signals 102 are provided to an encoder 106. The encoder 106 encodes the spatial metadata 104 and the audio signals 102 and forms a bitstream 108. The bitstream 108 can comprise the encoded spatial metadata and the encoded audio signals. Variations to the processor 206 shown in Fig. 3 could be used in examples of the disclosure. For instance, in some examples the procedures performed by the microphone signal preprocessor 302 can be performed on the microphone audio signals 204 before they are provided to the spatial metadata analyzer 300. For instance, significant noise removal or speech enhancing might alter the signal content significantly, and therefore it might be better to generate the spatial metadata 104 13 based on the audio signals 102 that are provided by the microphone signal preprocessor 302. Fig. 4 shows an example encoder 106 that can be used in some examples of the disclosure. The spatial metadata 104 and the audio signals 102 are provided as an input to the encoder 106. The audio signals 102 are provided to an audio encoder 400. The audio encoder 400 is configured to encode the audio signals 400. Any suitable process can be used to encode the audio encoder 400 such as the Immersive Voice and Audio Services (IVAS) core coder, Enhanced Voice Services (EVS) or Advanced Audio Coding (AAC) or any other suitable process. The audio encoder 400 provides encoded audio signals 402 as an output. The encoded audio signals 402 are provided to an audio decoder 404. The audio decoder 404 is configured to decode the encoded audio signals 402. The audio decoder 404 uses a decoding process that is compatible with the encoding process used by the audio encoder 400. The audio decoder 404 provides decoded audio signals 406 as an output. The decoded audio signals 406 and the spatial metadata 104 are provided to a metadata encoder 408. The metadata encoder 408 is configured to process the decoded audio signals 406 and the spatial metadata 104 to form a metadata bitstream 410 that represents the spatial metadata. The metadata bitstream 410 uses a fewer number of bits compared to the original spatial metadata 104. The metadata bitstream represents the spatial metadata in such a way that is intended to be later decoded with use of the decoded audio signals 406. The metadata encoder 408 can use any suitable processes. The metadata encoder 408 can use machine learning based processes to generate the metadata bitstream 410. The metadata encoder 408 can use the fact that the spatial metadata 104 and the decoded audio signals 406 (or other corresponding audio signals) represent different aspects of the same spatial audio signal, and so have at least some mutual correspondences to generate the compact representation of the metadata in the metadata bitstream 410. The encoded audio signals 402 and the metadata bitstream 410 are provided as an input to a multiplexer 412. The multiplexer 412 is configured to multiplex the encoded audio signals 402 and the metadata bitstream 410 into the bitstream 108. The bitstream 108 is the output of the encoder 106. Variations to the encoder 106 can be used in examples of the disclosure. For example instead of passing the audio signals through the audio encoder 400 and then the audio decoder 404 the audio signals 102 to could be provided directly to the metadata encoder 406. This variation could be used in low-delay cases. Fig. 5 shows an example method. The method could be implemented by an encoder 106 as shown in Fig. 4 or by any other suitable apparatus for encoding audio signals. At block 500 the method comprises obtaining one or more audio signals 102 and associated spatial metadata 104. The spatial metadata 104 is associated with the audio signals 102 in that the spatial metadata 104 can be estimated from the microphone signals or from any other suitable signals. The spatial metadata 104 can be used to process the audio signals 102 to render spatial audio. The one or more audio signals 102 and the associated spatial metadata 104 can be obtained from external sources such as a spatial metadata analyzer 300 and a microphone signal preprocessor 302 or they can be determined based on one or more received signals. The spatial metadata 104 can comprise spatial information that enables spatial rendering of the audio signals 102. The spatial metadata 104 can comprise one or more directional parameters and one or more energy parameters. For instance, the spatial metadata 104 can comprise direction information (which can be provided in frequency bands), direct-to-total energy ratios (which can be provided in frequency bands), or any other suitable information. At block 502 the method comprises determining at least one input for processing wherein the at least one input is based on the obtained one or more audio signals 102 and on the obtained spatial metadata 104. The processing can comprise processing that is performed on the spatial metadata 104 so as to improve the efficiency with which the spatial metadata 104 can be formed into a bitstream. In some examples determining at least one input for processing can comprise determining a first input for processing and determining a second input for processing. The first input can be based on the obtained one or more audio signals 102 and the second input can be based on the obtained spatial metadata 104. In some examples determining the at least one input for the processing can comprise encoding the obtained one or more audio signals 102. In some examples determining the at least one input for the processing can comprise at least partially decoding encoded one or more audio signals. For instance, as shown in Fig. 4 the audio signals 102 can be provided to an audio encoder 400 and, in some cases, also an audio decoder 404 before being provided to the metadata encoder 408 for processing. The encoding and decoding of the audio signals 102 means that the version of the audio signals that are used for processing are more similar to the version of the audio signals that can be used by a decoder 110 for decoding the spatial metadata. This can provide for improved spatial rendering. In other cases the encoding and decoding of the audio signals 102 before they are provided for processing can be omitted so as to improve latency. In some examples determining the at least one input for the processing comprises determining one or more features. The features can be determined by processing, the obtained one or more audio signals 102, one or more decoded audio signals, one or more partially decoded audio signals. The features can comprise information relating to the physical properties of the audio signal. The features can comprise peak energy levels, evolution of energy levels, repetitions of energy levels, or any other suitable information. Similarly, in some examples determining an input based on the spatial metadata 102 can comprise obtaining one or more features from the obtained spatial metadata 102 and using the one or more obtained features as the input to the processing. The features can comprise information relating to the properties of the spatial metadata. In some examples the features can comprise information that describes some properties or features of the audio signals and / or the spatial metadata. Such information provides indirect information relating to the physical properties of the sound scene. Such information could comprise inter-channel features or other suitable information. At block 504 the method comprises processing the at least one input to generate processed spatial metadata. The processed spatial metadata is formed into a bitstream 108 (such as the bitstream formed at block 506) using a fewer number of bits compared to the required number of bits for the obtained spatial metadata 104. That is, the processed spatial metadata provides a more compact representation of the spatial metadata but can be decoded by using the audio signals 102 so as to get closer to the original obtained spatial metadata 104. The processed spatial metadata can be encoded spatial metadata. The processed spatial metadata can comprise binary sequences or can be converted to binary sequences. The processed spatial metadata is generated so as to be decoded by decoding processing in an apparatus for decoding spatial audio signals. The decoding processing can be compatible with the processing used by the apparatus for encoding audio signals. For example, if the processing is performed using a machine learning model the decoding processing can be performed using a corresponding machine learning model. The corresponding machine learning models can be trained in conjunction with each other. The machine learning model 212 that is stored in the memory 208 could be an encoding machine learning model and / or a decoding machine learning model. The processing to generate the processed spatial metadata can be performed, at least in part, using a machine learning model 212. The machine learning model 212 can comprise a deep neural network (DNN) model or any other suitable type of model. At block 506 the method comprises forming the bitstream 108 based on the processed spatial metadata and encoded one or more audio signals. The encoded one or more audio signals that are used to form the bitstream 108 can be obtained by encoding the obtained one or more audio signals. In some examples forming the bitstream 108 comprising the processed spatial metadata and encoded one or more audio signals can comprise encoding the obtained one or more audio signals 102 and combining the encoded one or more audio signals with the processed spatial metadata generated by the processing. In some examples the encoding of the audio signals 102 and the forming of the bitstream 108 can be performed by the same entity such as a metadata encoder 408 in Fig. 4. In such examples the processing provides a bitstream 108 comprising the processed spatial metadata and encoded one or more audio signals as an output. In other examples the encoding of the audio signals 102 and the forming of the bitstream 108 can be performed by different entities. For instance, a metadata encoder 408 can generate the processed spatial metadata and a separate encoder can encode the audio signals 102. Any suitable processing can be performed on the processed spatial metadata and or the encoded audio signals before they are formed into a bitstream 108. The bitstream 108 comprising the processed spatial metadata and encoded one or more audio signals can be transmitted to a decoder device or to a media server or to any other suitable entity. Fig. 6 shows an example decoder device 600. The bitstream 108 that is generated by the encoder 106 or encoder device 200 can be transmitted to the decoder device 600. In this example the decoder device 600 is a mobile device with connected headphone 608. The headphone 608 can be connected to the decoder device 600 via a wired connection or a wireless connection. Other types of decoder devices 600 and connected peripheral devices can be used in other examples of the disclosure. The example decoder device 600 comprises a processor 206, a memory 208, storage 216, a transceiver 214 and a headphone connection 604. The decoder device 600 can also comprise other components that are not shown in Fig. 6 such as a camera or any other suitable components. The reference numerals used in Fig. 6 correspond to the reference numerals used in Fig. 2 for the encoder device 200. A mobile device can typically perform both encoding and decoding operations during a communication session. Only decoding operations are referred to in relation to Fig. 6 and the following Figs. The user 610 of the decoder device 600 can be using the decoder device 600 to participate in an immersive call or other communication that uses spatial audio. The decoder device 600 obtains a bitstream 108. The bitstream 108 can comprise the bitstream 108 that is transferred between an encoder 106 and a decoder 108 as shown in Fig. 1. The bitstream 108 can comprise processed spatial metadata that has been generated using the method of Fig. 5 or any other suitable method. The bitstream 108 can be received via a transceiver 214 or can be retrieved from storage 216. In some examples the bitstream 108 can be received via the transceiver 214 and then stored to the storage 216 and accessed later by the processor 206. The processor 206 is arranged to convert the bitstream 108 to a spatial audio output 602. The processor 206 and memory 208 can provide a decoder 110 as shown in Fig. 1. The processor 206 uses the program code 210 and the machine learning model 212 that are stored in the memory 208 to convert the bitstream 108 to the spatial audio output 602. In this example the spatial audio output 602 is a binaural output. Other types of spatial audio output 602 can be used in other examples. Any suitable information can be stored in the memory 208 and accessed by the processor 206. In examples of the disclosure the memory 208 can comprise a program code 210 and a machine learning model 212. The machine learning model 212 can be as described below and can be used to decode processed spatial metadata. The machine learning model 212 is stored in the memory 208. The machine learning model 212 can be trained by an external device and then provided to the decoder device 600 so that it can be stored in the memory 208. The external device that performs the training of the machine learning model 212 can have a higher processing capacity than the decoder device 600 shown in Fig. 6 or other devices that implement the spatial audio communications. For example, the external device that performs the training of the machine learning model 212 could comprise a workstation with multiple graphic processing units (GPUs) dedicated to the training of the machine learning model 212. The spatial audio output 602 is provided to the headphone connection 604. The headphone connection 604 connects the headphone 608 to the decoder device 600 and enables signals from the decoder device 600 to be provided to the headphone 608. The headphone connection 604 can be a wired connection or a wireless connection. The headphone connection 604 provides an audio signal 606 to the headphone 608 to enable the spatial audio to be played back to the user 610. Fig. 7 shows an example decoder 110. The processor 206 and memory 208 of the decoder device 600 can perform the operations of the decoder 110 as shown in Fig. 7. The bitstream 108 is provided to a demultiplexer 700. The demultiplexer 700 is arranged to demultiplex the bitstream 108 to provide encoded audio signals 402 and a metadata bitstream 410. The encoded audio signals 402 can be encoded transport audio signals or any other type of audio signals. The encoded audio signals 402 are provided to an audio decoder 702. The audio decoder 702 is configured to decode the encoded audio signals 402. The audio decoder 702 can use any suitable type of decoder. The audio decoder 702 can use a decoder that is compatible with the audio encoder 400 that was used to encode the audio signals 102. The audio decoder 702 provides decoded audio signals 704 as an output These can be decoded transport audio signals. The decoded audio signals 704 and the metadata bitstream 410 are provided as an input to a metadata decoder 706. The metadata decoder 706 is configured to process the metadata bitstream 410 so that information from the decoded audio signals 704 can be used to decode the spatial metadata. The metadata decoder 706 can comprise a machine learning model 212 or any other suitable means. The machine learning model 212 of the metadata decoder 706 can have been trained in conjunction with the machine learning model 212 of the metadata encoder 408. The metadata decoder 706 provides decoded spatial metadata 708 as an output. The decoded spatial metadata 708 and the decoded audio signals 704 are provided as an input to a spatial synthesizer 710. The spatial synthesizer 710 is configured to process the decoded spatial metadata 708 and the decoded audio signals 704 to render spatial audio output 112. The spatial audio output 112 can comprise binaural audio signals, stereo audio signals, multi-channel audio signals (e.g., 5.1 or 7.1+4), Ambisonics signals or any other suitable type of spatial audio. The spatial audio output 112 of the decoder 110 can be the spatial audio output 602 of the processor 206 as shown in Fig. 6 Fig. 8 shows an example method. The method could be implemented by a decoder 110 as shown in Fig. 7 or by any other suitable apparatus for decoding audio signals. The method comprises, at block 800, receiving a bitstream 108. The bitstream 108 comprises encoded one or more audio signals 402 and processed spatial metadata. The processed spatial metadata has been formed into a bitstream 410. This can be referred to as a spatial metadata bitstream or a metadata bitstream. The spatial metadata bitstream 410 uses a fewer number of bits compared to the required number of bits for originally obtained spatial metadata 104. The originally obtained spatial metadata 104 can be the spatial metadata 104 that is obtained by processing the microphone audio signals 204. This can be obtained by the encoder device 200 or by any other suitable device. The processed spatial metadata can be generated by processing an apparatus for encoding spatial audio signals such as the encoder device 200 in Fig. 2 or any other suitable device. The bitstream 108 can be received from an encoder device 200 or from any other suitable communications device. In some examples the bitstream 108 can be received by being retrieved from storage 216. At block 802 the method comprises determining at least one input for decoding processing. The at least one input is based on the encoded one or more audio signals and on the processed spatial metadata. In some examples determining at least one input for decoding processing can comprise determining a first input for decoding processing and determining a second input for decoding processing. The first input can be based on the encoded one or more audio signals and the second input can be based on the processed spatial metadata. In some examples determining the at least one input for the decoding processing can comprise decoding the encoded one or more audio signals 402 that were received in the bitstream 108. The decoding can be performed by an audio decoder 702 as shown in Fig. 7 and / or by any other suitable entity. This can provide an input based on the encoded one or more audio signals In some examples determining the at least one input for the decoding processing comprises processing the encoded one or more audio signals 402 (or decoded audio signals 704) to determine one or more features. The features can comprise information relating to the properties of the audio signal. The features can comprise peak energy levels, evolution of energy levels, repetitions of energy levels, or any other suitable information. Similarly, in some examples determining the at least one input for the decoding processing comprises using the processed spatial metadata, or a further processed version of the processed spatial metadata as part of the at least one input. For example, the processed spatial metadata could be codebook indices, and this further processing could comprise converting the codebook indices to data vectors. Similarly, in some examples determining an input based on the processed spatial metadata can comprise obtaining one or more features from the processed spatial metadata and using the one or more obtained features as the input to the decoding processing. The features can comprise information relating to the physical properties of the spatial metadata. This can provide an input based on the processed spatial metadata. This method could be used in examples that do not use vector-encoded data. At block 804 the method comprises performing decoding processing on the at least one input to generate decoded spatial metadata 708. The decoding processing can be performed by a metadata decoder 706 as shown in Fig. 7 or by any other suitable means. The decoding processing can be performed, at least in part, using a machine learning model 212. The machine learning model 212 can comprise a deep neural network (DNN) model or any other suitable type of model. A machine learning model 212 that is used for decoding processing can be training in conjunction with a corresponding machine learning model used for encoding the spatial metadata. At block 806 the method comprises using the decoded spatial metadata 708 and decoded one or more audio signals 704 to enable rendering of a spatial audio output 112. The spatial output can comprise a binaural output, multi-channel loudspeaker output, Ambisonics output, stereo output, cross talk-cancelled stereo output, or any other suitable type of output. Fig. 9 shows an example metadata encoder 408 that can be used in some examples of the disclosure. The metadata encoder 408 can be provided in an encoder 106 as shown in Figs. 3 and 4. The metadata encoder 408 can be arranged to perform the processing of inputs to generate processed spatial metadata as shown in Fig. 5. In this example the metadata encoder 408 is implemented using a machine learning model. Other types of processing could be used in other examples. Figs. 9 and 10 show the inference of the machine learning model. An example of training the machine learning model is shown in Fig. 13. The metadata encoder 408 receives the decoded audio signals 406 and the spatial metadata 104 as inputs. Different inputs can be used in different examples. For instance, in examples which need lower latency and / or computation requirements the first input to the metadata encoder 408 could be the audio signals 102. The spatial metadata 104 is provided to a spatial metadata preprocessor 900. The spatial metadata 104 can comprise one or more directional parameters and one or more energy parameters. In this example the spatial metadata 104 comprises an azimuth parameter azi(k,n), an elevation parameter ele(k,n) and a direct-to-total energy ratio parameter r( / c,n), where k = 1,..., 24 is a frequency band index and n is a subframe index. The frequency bands k correspond to frequency intervals that are wider in the higher frequencies, and narrower in the lower frequencies. One subframe corresponds to 5 milliseconds of audio. The spatial metadata preprocessor 900 is arranged to convert the spatial metadata 104 to a vector form by cos(azj( / c, n)) cos (ele(k,n)) v(k, ri) — r(k, n) sin(azi(k, n)J cos (ele(k, n)) sin (ele(k, n)) The data in v(k,n) is the preprocessed metadata 902. The preprocessed metadata 902 is provided as the output of the spatial metadata preprocessor 900. In some examples the metadata encoder 408 can be operating in 20 millisecond audio frames. In such examples, there are four 5 millisecond subframes for each frame of data; 24 frequency bands and three dimensions (x, y, z) of the vector v(k,n). Therefore, one frame of preprocessed metadata 902 in this example has a size (1 x 24 x 4 x 3), where the first dimension is the batch size of the network call (which is 1 in the inference stage); the second dimension is the frequency dimension; the third dimension is the time dimension; and the fourth dimension is the feature dimension. In this case the feature dimension is the x, y, z values of vector v(k, n). The decoded audio signals 406 are provided to an audio feature analysis block 904. The audio feature analysis block 904 is arranged to preprocess the information in the decoded audio signals 406 to a suitable form for input to the encoder machine learning model 908. In the example of Fig. 9, the processing performed by the audio feature analysis block 904 obtains a form of the audio features 906 that presents the spectral data of the decoded audio signals 406 in a time-frequency resolution of the spatial metadata 104. A first operation of the audio feature analysis block 904 is to convert the decoded audio signals 406 to the time-frequency representation. In the example of Fig. 9, a complex-modulated low-delay filter bank (CLDFB) is used that provides 60 uniform frequency bins of data, where every frequency bin has 400 Hz bandwidth (for 48 kHz sampling rate signals). Then, energy is formulated by tiW b^k) 2 E(k, n) = I Z t=t0(n) b=b0(k) i=l where S(b, k, i) is the CLDFB transformed version of the decoded audio signals 406, i is the transport audio channel index (in this example case there are two transport audio channels, but a different number is possible, such as one for mono-MASA), t and b are the time and frequency indices of the CLDFB transformed signal, b0(k) and b^lk) are the lowest and highest CLDFB bins of band k, and t0(n) and t^n) are the first and last CLDFB sample (or “slot”) of subframe n. The audio feature analysis block 904 the processes the energy data using the following procedure. For temporal steps where the history data would refer to negative indices of t, it is assumed that such data is zeros. The processing steps comprise: E(k,n) z / E(k,ri) \03 Eber(k,n)- [^kE(k n)) Eproc(.k,Tt) (^Eevo(kt n)Eber(k, ii)^) Where Eevo is the evolution of energy over time, Eber is the band to total energy ratio, and Eproc is the processed energy values. The divisions can be regularized, for example, the divisor can be bottom limited to Ie-12. The feature computation parameters, such as the context length for Eevo(k,ri) or the value powers for Eber(k,n) and Eproc(k,n) are provided as examples that are suitable for this application. Different feature computation parameters can be used in other examples. The values Eproc(k, ri) are the audio features 906 that are provided as the output of the audio feature analysis block 904. The audio features 906 for one 20 milliseconds frame could be provided in the shape (1 x 24 x 4 x 1), where the dimensions are (batch, frequency, time, feature). This audio feature has been aggregated over all input audio signal channels, and therefore the method is applicable also for the use case of having only one audio signal, or the use case of having more than two audio signals. The preprocessed metadata 902 and the audio features 906 are provided as an input to the encoder machine learning model 908. The encoder machine learning model 908 is configured to process the preprocessed metadata 902 and the audio features 906 to form the metadata bitstream 410. Any suitable means can be used to implement the encoder machine learning model 908. In some examples the implementation of the encoder machine learning model 908 can comprise two parts. The first part can comprise the definition of the machine learning model in a known format. The second part can comprise software capable of performing inference according to the known format. For example, the inference software could be the TensorFlow Ute interpreter, and the machine learning model could be stored in the TensorFlow Lite file format. A second example is that the model is stored in the ONNX format and the inference software may be ONNX Runtime. The machine learning model used for the encoder machine learning model 908, and any other machine learning models used in implementations of the disclosure can comprise a set of processing instructions, such as multiplications, additions, nonlinearities, data combining, data splitting, table lookups, and any other suitable operations. As such, the processing instructions of a machine learning model are not fundamentally different from the processing instructions of conventional program code. The key difference is that the coefficients applied at these processing instructions have been, at the offline training stage, adapted based on training data. It is possible to train one machine learning model with a specific architecture, then derive another machine learning model from that using processes such as compilation, pruning, quantization, or distillation. The term machine learning model also covers all these use cases and the outputs of them. The machine learning model can be executed using any suitable apparatus, for example central processing unit (CPU), graphics processing unit (GPU), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), compute-in-memory, analogue, or digital, or optical apparatus or any other suitable means. It is also possible to execute the machine learning model in apparatus that combine features from any number of these, for instance, digital-optical or analogue-digital hybrids. In some examples, the weights and required computations in these systems can be programmed to correspond to the machine learning model. In some examples, the apparatus can be designed and manufactured so as to perform the task defined by the machine learning model so that the apparatus is configured to perform the task when it is manufactured without the apparatus being programmable as such. The encoder machine learning model 908 of the metadata encoder 408 can be trained at an offline stage (training taking place temporally before the use of the model in inference) together with a decoder machine learning model. The decoder machine learning model can be as described herein. The encoder machine learning model 908 is trained to create a metadata bitstream 410 that the decoder machine learning model can decode. The overall system (comprising the encoder machine learning model 908 and the decoder machine learning model) can be optimized, or substantially optimized, in the training stage, to compare metadata predicted by the decoder machine learning model against the original spatial metadata 104. As such, the exact organization of the metadata bitstream 410 is not predetermined, but adapts to a certain form in the training process. In the training phase, the encoder machine learning model 908 and decoder machine learning model can be considered in practice as a unified machine learning model that has a low-bitrate bottleneck in the center between the encoder part and the decoder part. The two network portions are split for the inference stage as encoder and decoder parts. Fig. 10 shows an example method of encoding spatial metadata. The method can be implemented by a metadata encoder 408 as shown in Fig. 9 or by any other suitable apparatus or means. At block 1000 the metadata encoder 408 obtains audio signals. The audio signals can comprise decoded audio signals 406 or any other suitable version of the audio signals. At block 1002 audio features 906 are estimated based on the obtained audio signals. This block can be an optional block depending on the audio signals that are obtained at block 1000. For example, the signals obtained at block 1000 could comprise the audio features 906, in which case no further processing would be needed. At block 1004 the metadata encoder 408 obtains spatial metadata 104. At block 1006 the preprocessed metadata 902 is generated by processing the spatial metadata 104. This block can be an optional block depending on the spatial metadata 104 that is obtained at block 1004. For example, the spatial metadata 104 that is obtained at block 1004 could already be in a format which is suitable for input to the encoder machine learning model 908, in which case no preprocessing of the spatial metadata 104 would be needed. At block 1008 the preprocessed spatial metadata 902 and the audio features 906 are processed to form the metadata bitstream 410 and at block 1010 the metadata bitstream 410 is provided as an output. Fig. 11 shows an example metadata decoder 706 that can be used in some examples of the disclosure. The metadata decoder 706 can be provided in a decoder 110 as shown in Fig. 7. The metadata decoder 706 can be arranged to perform the decoding processing of the inputs to generate decoded spatial metadata 708. In this example the metadata decoder 706 is implemented using a machine learning model. Other types of processing could be used in other examples. Figs. 11 and 12 show the inference of the machine learning model. An example of training the machine learning model is shown in Fig. 13. The metadata decoder 706 receives the decoded audio signals 704 and the metadata bitstream 410 as inputs. The decoded audio signals 704 are provided to an audio feature analysis 1100 block. The audio feature analysis block 1100 can operate in the same manner as the audio feature analysis block 904 of the metadata encoder 408 as described herein. The audio feature analysis block 1100 provides audio features 1102 as an output. The format of the audio features 1102 can be as describe above. The audio features 1102 and the metadata bitstream 410 are provided as an input to the decoder machine learning model 1104. As described in context of the encoder machine learning model 908, the decoder machine learning model 1104 can be defined in a TensorFlow Lite format and run with the TensorFlow Lite interpreter or can be implemented in any other suitable manner. The decoder machine learning model 1104 provides the predicted format spatial metadata 1106. In this example the predicted format spatial metadata 1106 is in a form of vectors vpred( / c,n), which could for a 20-millisecond frame of audio data be in a form of (1 x 24 x 4 x 3) where the last dimension is the three dimensions of the vector. The predicted format spatial metadata 1106 is provided to the spatial metadata postprocessor 1108. The spatial metadata postprocessor 1108 processes the predicted format spatial metadata 1106 to provide decoded spatial metadata 708 as an output. The predicted format spatial metadata 1106 can be denoted vpred(k,n) = [vPred,xCk,n) vpredy(k,n) vpred>z(k,n)]T, and the spatial metadata postprocessor 1108 converts the predicted format spatial metadata 1106 to decoded spatial metadata 708 by ^^^dec^.k,Ti) atan2(vprediy(k,n),vpredxQk,n)^ eledec(kfn) = atan2 (vpred:Z(k,n), v^redx(k,n) + v^redy(k,n)] rdec(k,n) = v*redx(k,n) + v^redy(k,n) + v^redz(k,n) where atanK^T) is the inverse tangent function resolving the correct quadrature, and the rdec(k,n) may be upper limited to 1. This is the decoded spatial metadata 708 that is provided as the output of the spatial metadata postprocessor 1108 of Fig.11 and correspondingly of the metadata decoder 706 of Fig. 7. Fig. 12 shows an example method of decoding spatial metadata. The method can be implemented by a metadata decoder 706 as shown in Fig. 11 or by any other suitable apparatus or means. At block 1200 the metadata decoder 706 obtains audio signals. The audio signals can be decoded audio signals 704 as shown in Fig. 11. At block 1202 audio features 1102 are estimated based on the obtained audio signals. This block can be an optional block depending on the audio signals that are obtained at block 1200. For example, the signals obtained at block 1200 could comprise the audio features 1102, in which case no further processing would be needed. At block 1204 the metadata decoder 706 obtains the metadata bitstream 1204. At block 1206 the metadata bitstream 410 and the audio features 1102 are processed to obtain spatial metadata. The metadata bitstream 410 and the audio features 1102 can be processed using a decoder machine learning model 1104 or any other suitable decoding processing. The spatial metadata that is provided by the processing can comprise predicted format spatial metadata 1106. At block 1208 post processing of the spatial metadata can be performed. The post processing of the spatial metadata can process the spatial metadata into a suitable format for use in rendering the spatial audio output 112. The post processing of the spatial metadata can be an optional step. In some examples the decoder machine learning model 1104 can provide an output which is already in a suitable format for use in rendering the spatial audio output 112 and which does not need any postprocessing. At block 1210 spatial metadata is output. The output can comprise decoded spatial metadata 708 as shown in Fig. 11. The spatial metadata can be provided to a spatial synthesizer 710 as shown in Fig. 7 to enable rendering of the spatial audio output 112. In some examples the spatial synthesizer 710 can be separate to the metadata decoder 706. For example, a metadata decoder 706 can be provided in a decoding device 600 and a spatial synthesizer 710 can be provided on a peripheral device such as a headphone 608. In other examples the synthesizer 710 and the metadata decoder 706 can be part of the same software or other means. Any suitable process can be used for rendering the spatial audio output 112. Fig. 13 shows an example training process 1300 for training encoder machine learning models 908 and corresponding decoder machine learning models 1104. More precisely, the figure arrows depict the signal flow in the forward pass of the training procedure. Fig. 13 does not show all known aspects that can be used as part of the training process 1300 such as the back propagation of the error gradient from the loss function 1302. The back propagation is a process to determine an incremental update for each trainable weight of the respective machine learning models 908, 1104. The incremental updates are determined based on each training data batch. By repeated steps of providing the system of Fig. 13 new sets of training data and then updating the weights incrementally, the machine learning models 908, 1104 are trained to minimize the loss function 1302 for the training data set. The machine learning models 908, 1104 can comprise any suitable operations. The operations can be conventional signal processing operations. Some of the operations that are shown as separate to the machine learning models 908, 1104 in Fig. 13 could be part of the machine learning models 908, 1104 in other examples. For example, Fig. 13 shows the audio feature analysis block 904 as being outside of the machine learning models 908, 1104 but the audio feature analysis could be implemented as part of the machine learning models 908, 1104 in other examples. Similarly, some of the operations that are shown as within the machine learning models 908, 1104 in Fig. 13 could be outside of the machine learning models 908, 1104 in other examples. The input data for the training process 1300 comprises a set of audio signals 102 and corresponding spatial metadata 104. The audio signals 102 can comprise transport audio signals. In this specific example the audio signals 102 and corresponding spatial metadata 104 are based on 32-microphone Eigenmike recordings. The microphone audio signals are analyzed using the MASA analysis reference software that is publicly available. In the MASA analysis reference software the Eigenmike microphone audio signals are analyzed to determine the spatial metadata 104. The audio signals 102 are two microphone signals selected from the left and right edges of the Eigenmike. The spatial metadata 104 and audio signals 102 can be obtained from other types of devices in other examples. For example, they can be obtained from Ambisonic-capable microphone arrangements, mobile consumer devices such as phones, cameras, laptops, and any other suitable devices. Any suitable software and / or processes can be used to obtain the spatial metadata 104 and audio signals 102 from such devices. In some examples the spatial metadata 104 and audio signals 102 can be obtained from a multi-channel signal such as 5.1, 7.1+4 sounds, or object-based sound scenes, or mixtures of any of these formats. In the example training process of Fig. 13 only data from the Eigenmike was used as the input data. The same training process could be used for a variety of input data. For example, the input data could comprise data from multiple different types of data source. The audio signals 102 are encoded using an audio encoder 400. The audio encoder 400 provides encoded audio signals 402 as an output. The encoded audio signals 402 are provided to the audio decoder 702. The audio encoder 400 and the audio decoder 702 that are used during the training process 1300 can be the same as the audio decoder 702 that are used during the inference processes. The audio decoder 702 provides decoded audio signals 704 as an output. The decoded audio signals 704 are provided to the audio feature analysis block 904. The audio feature analysis block 904 that is used during the training process 1300 can operate in the same way as the audio feature analysis block 904 that is used in the inference, for example as shown in Fig. 9. The audio feature analysis block 904 provides the audio features 906 as an output. The audio features 906 are provided to the machine learning models 908, 1104. The encoder machine learning model 908 and the decoder machine learning model 1104 can be trained jointly using the training process 1300 of Fig. 13 or any other suitable process. The spatial metadata preprocessor 900 is arranged to perform processing of the spatial metadata 104. The processing of the spatial metadata 104 can be as described in relation to Fig. 9 or can comprise any other suitable processing. The spatial metadata preprocessor 900 provides preprocessed metadata 902 as an output. The preprocessed metadata 902 is provided as input training data for the joint training of the machine learning models 908, 1104. The preprocessed metadata 902 is provided to the encoder machine learning model 908 and also to the loss function 1302. The training data comprising the preprocessed metadata 902 and the audio features 906 can be provided in any suitable format for the joint training of the machine learning models 908, 1104. In the described example the training data comprising the preprocessed metadata 902 is provided in data sizes of (30 x 24 x 12000 x 3) and the training data comprising the audio features 906 is provided in data sizes of (30 x 24 x 12000 x 1) for “Audio features”. In this case the used batch size was 30, and the temporal length of the samples was 12000 subframes, which corresponds to 60 seconds of audio data. The machine learning models 908, 1104 receive the training data comprising the preprocessed metadata 902 and the audio features 906 and produce an output comprising the predicted format spatial metadata 1106. The predicted format spatial metadata 1106 can be provided in any suitable format. In the described example the predicted format spatial metadata 1106 has shape (30 x 24 x 12000 x 3). This format corresponds to the vectors vpred(fc,n) that can be converted to the decoded spatial metadata 708 as shown in Fig. 11. However, in the training stage, the vectors are not converted, but provided to the loss function 1302. In the example of Fig. 13 there is no additional processing of the predicted format spatial metadata 1106 before it is provided to the loss function 1302. The loss function 1302 receives the predicted format spatial metadata 1106 (for example, the network prediction vector data vpred(k,n)), preprocessed metadata 902 (for example, the reference vector data v(k,n)), and audio features 906 Eproc(k,n). The audio features 906 in this example are energy-based, and can be used as a weighting function for the loss. The loss function 1302 can apply any suitable procedure. In some examples the loss function can operate as follows. The vectors are first scaled in length by v'(k,n) = v(k, n) max ^||v( / c,n)||4, —) l|v(^7l)|| And, correspondingly for vpred(k,n) to obtain v'pred(k,n), where ||v(k,n)|| = JvT(k, n)v(k, n) denotes the vector length. The above function weights the vector length so that vector length errors when ||v(fc, n)|| and ||vpred(fc,n)|| are closer to 1 have a larger significance. This function causes the vector length to be mapped to a scale where a just noticeable difference of an error in the vector length is more similar throughout all vector lengths. The loss for the vectors is then ^vec । V pred (.k, n) 11 Eproc k n where the summation is done over the range of k,n in the training example. The overall loss function also has a quantizer loss, as explained below. The training may be performed with any suitable optimizer such as an AdamW optimizer with a learning rate 1e-3 and a weight decay factor of 1e-6. Figs. 14A and 14B show example combination blocks 1400,1420 that can be used in the machine learning models 908, 1104. The combination block 1400 shown in Fig. 14A is referred to as BRC which is an abbreviation of batch normalization, rectified linear unit, and convolution. The combination block 1420 shown in Fig. 14B is referred to as BRCT which is an abbreviation of batch normalization, rectified linear unit, and convolution where T denotes usage of transposed convolution. As shown in Fig. 14A the BRC block 1400 comprises a sequence of batch normalization 1402, rectified linear unit (ReLU) 1404, and convolution 1406. The BRC block 1400 also has a bypass convolution 1410 without a nonlinear operator. The convolution 1406 and the bypass convolution 1410 have mutually all the same configuration parameters, but have their own trainable weights. A sum block 1408 sums the output of the convolution 1406 and the bypass convolution 1410. The respective blocks of the BRC block 1400 can operate in a conventional manner. For example, the batch normalization 1402 can normalize the mean and standard deviation of each feature data in the training time and is converted to fixed bias and scale operators (per feature) in the inference time. The BRCT block 1420, as shown in Fig. 14B, is similar to the BRC block 1400. The BRCT block 1420 comprises a sequence of batch normalization 1422, rectified linear unit (ReLU) 1424, and transposed convolution 1426. The BRCT block 1420 also has a bypass transposed convolution 1430 without a nonlinear operator. The transposed convolution 1426 and the bypass transposed convolution 1430 have mutually all the same configuration parameters, but have their own trainable weights. A sum block 1428 sums the output of the transpose convolution 1426 and the bypass transpose convolution 143O.The respective blocks of the BRCT block 1420 can operate in a conventional manner. Fig. 15 shows an example encoder machine learning model 908 that can be used in some examples of the disclosure. The encoder machine learning model 908 receives the preprocessed metadata 902 and the audio features 906 as inputs. The first operation comprises a concatenation 1500. The concatenation 1500 concatenates the preprocessed metadata 902 and the audio features 906 along the feature axis (in this example all concatenation operators in the encoder machine learning model 908 are along the feature axis). The parameters of the BRC blocks 1504 to 1520 in Fig. 8 indicate the parameters of both convolution blocks 1406, 1410 within the respective BRC blocks, where “s” denotes the stride, “d” denotes the dilation, “f” denotes the number of output features, and “same” and “valid” refer to the padding mode. When not mentioned, the strides and dilations are (1,1). The order of data is (frequency, time), and therefore, for example, “BRC (1,3) valid f=64 d=(1,4)” would mean that the convolutions have a kernel size 3 along time axis, but with a dilation rate of 4 along the time axis, and 64 output features. Furthermore, the indicator “valid” would mean that no padding is implemented. Correspondingly, “same” would mean zero-padding that results in unchanged data size in case no strides are applied. After the first concatenation 1500 the data is in the shape of (batch_size, num_T, num_F, 4), where the last dimension is the feature dimension having the three elements from the metadata vectors of the preprocessed metadata 902, and one element from the audio features 906. A convolution operation 1502 is provided after the concatenation 1500. The convolution operation 1502 does not change the data size. A first BRC block 1504 is provided after the convolution operation 1502. The first BRC block 1504 converts the data which temporally corresponds to 5 millisecond subframes, to data that corresponds to 20 millisecond frames (because stride and kernel size is 4 at the time axis). The remainder of the encoder machine learning model 908 operates on this data that temporally corresponds to frames. The design of the encoder machine learning model 908 shown in Fig. 15 can converted to operate on a frame-by-frame basis. When the inference time system receives one frame (4 subframes) of data, it can be first converted to frame data by the operators described above. Then, when any of the BRC operators need to access the history data, that history data can be retrieved from memory storage, so that the encoder machine learning model 908 does not need to re-calculate anything. The BRC operators 1504 to 1520 that operate on the time axis can be considered to always receive one frame worth of data, but then to perform a specific padding along the time axis (towards the history direction), where this history padding is the data that that BRC operator had received in the prior calls. During the training of the encoder machine learning model 908 the encoder machine learning model 908 receives longer temporal sequences of data (and does notoperate frame by frame). Therefore, no memory storage of prior calls is needed during the training procedure. Instead, the temporal “valid” padding consumes the available data history at the input data examples. Therefore, during the training procedure, the encoder machine learning model 908 as shown in Fig. 15 provides output data that is temporally shorter than what it received as input. In the encoder machine learning model 908 shown in Fig. 15, the output data comprise 80 subframes (that is, 20 frames, 0.4 seconds shorter than the input). This corresponds to the receptive field (towards the history direction) of the encoder machine learning model 908. Therefore, at the subsequent loss function 1302, the reference data is cropped from the beginning by the same amount (80 subframes worth of data) as the encoder machine learning model 908 had consumed the temporal data as history data. In an alternative implementation of the encoder machine learning model 908, the input data examples can be zero-padded (or otherwise padded) from the beginning prior to calling the encoder machine learning model 908. This enables the output during the training process to have the same temporal length as the input examples. After the first BRC block 1504 that maps subframe data to frame data, four pairs of BRC blocks follow. The pairs of BRC blocks are indicated by the brackets 1530, 1532, 1534 and 1536 in Fig. 15. The first pair 1530 comprises BRC blocks 1506 and 1508. The second pair 1532 comprises BRC blocks 1510 and 1512. The third pair 1534 comprises BRC blocks 1514 and 1516. The fourth pair 1536 comprises BRC blocks 1518 and 1520. The first BRC block 1506, 1510, 1514, 1518 in each pair applies convolution along the frequency axis while also downsampling the data using strides. The second BRC block 1508, 1512, 1516, 1520 in each pair applies the convolution along time axis. The time-axis convolution applies dilation (not stride) to enable the longer receptive field when needed. Unlike using temporal strides, the dilation structure supports straightforwardly converting the encoder machine learning model 908 for inference time frame-by-frame processing, as described above. On the other hand, using strides on the frequency axis does not prevent the conversion process to frame-by-frame processing. The operation of the BRC blocks 1504 to 1520 reduces the frequency dimension to 1. The resulting data is processed with a fully-connected (FC) layer 1522 with 64 output features and tanh activation 1524. These 64 values per frame, which have values between -1 and 1, are the latent representation of the metadata that is being quantized with the vector quantizer. At block 1526 Gaussian noise with a standard deviation of 0.05 is added to this latent representation. The addition of Gaussian noise is only performed during the training procedure. The addition of Gaussian noise is performed for regularization purposes. During inference, no noise is added. The purpose of the added noise is to ensure that the encoder machine learning model 908 is trained to use the latent representation in a balanced manner. The latent data that is appended with the noise process is not strictly limited between values -1 and 1. The latent data with added noise is provided to a vector quantizer 1528. The vector quantizer 1528 applies a five-level residual vector quantizer with 8 bits (256 entries) codebook for each level. The vector quantizer can use the exponential moving average method for adapting the codebook vectors described in A. Razavi, A. van den Oord, and O. Vinyals, “Generating diverse high-fidelity images with VQ-VAE-2,” arXiv: 1906.00446, 2019, or any other suitable process. Considering only one level of the five-level residual vector quantizer, the vector quantizer has a codebook that has 256 vectors, each 64 elements long, and the codebook entries are adapted to represent the received data distribution. The vector quantizer with exponential moving averages operates for one update step based on a batch of data as follows. First, each received data vector (which for the first level of the vector quantizers is the data vectors from block 1526) is associated with the closest vector currently at the vector quantizer code book. Then each of the code book vectors are updated towards the average of the data points associated with it. In the exponential moving average method, this average is recursively updated in an HR fashion. By repeated such iterations with new data batches, the vector quantizer learns to encode the data distribution it receives. In a multi-level residual vector quantization, each level operates independently as above, but with different data. The first level receives data to be encoded and selects from the codebook the vector that is closest to the received vector (and performs the codebook updates when in the training stage). Then, the second level encodes the difference of the quantized vector and original vector, and so forth, for five levels. Each level then forms a codebook to represent the data it receives. The total bit rate in this example is 5*8=40 bits per frame, which corresponds to 2 kilobits per second. A quantized vector provided by the multi-level residual vector quantizer is then the sum of the five selected codebook vectors, one from each level of the residual vector quantizer. The vector quantizer is updated only at the training stage, and not at the inference stage. As described in the foregoing references, the vector quantizer that operates within a system of a trainable encoder and a decoder, which is the present situation, quantizes the data in the forward pass, but provides the gradient as pass-through in the backward pass. This is because the quantization process is not differentiable. The vector quantizer also provides a quantizer loss, that is the L2 norm of the non-quantized and quantized data. This loss is added to the loss function Lvec, without weighting. This loss enables the encoder to create latent representation values that are closer to the quantization data points. The residual vector quantizer 1528 expresses the 64 feature values (for each frame) with 5*8=40 bits of data. The vector quantizer 1528 provides quantized vectors 1538 as an output. During training of the encoder machine learning model 908 the quantized vectors 1538 comprise the 64 feature length vectors for each frame. During training the quantized vectors 1538 are of the same shape as the input to the vector quantizer 1528. That is, the quantized vectors 1538 are approximations of the inputs. During inference the output of the vector quantizer 1528 comprises the 40 bits data indicating the vector indices in the five-level codebook. These bit sequences can be converted to the quantized vectors 1538, by selecting the vector elements from the five code books and adding them together. These 40 bits per frame are the metadata bitstream 410 as shown in Fig. 4. Fig. 16 shows an example decoder machine learning model 1104 that can be used in some examples of the disclosure. The decoder machine learning model 1104 receives audio features 906 and quantized vectors 1538 as inputs. During inference the quantized vectors 1538 can be converted from the metadata bitstream 410. The decoder machine learning model 1104 is arranged so that the audio features 906 are initially processed by a first BRC block 1600. The first BRC block 1600 combines the temporal data from subframes to frames (for example, combining 4 steps of temporal data to 1 step). The result from the first BRC block 1600 is processed with a further four BRC blocks 1602, 1604, 1606, 1608. The further four BRC blocks 1602, 1604, 1606, 1608 operate to combine the frequency dimension from 24 down to 1. The output of the BRC block 1608 is provided to a concatenation block 1610. The concatenation block 1610 concatenates the output of the BRC block 1608 with the quantized vectors 1538. During training of the decoder machine learning model 1104 the quantized vectors 1538 can be received directly from the encoder machine learning model 908. In some examples there might not be any processing of the output of the encoder machine learning model 908 before it is provided to the decoder machine learning model. During inference using the decoder machine learning model 1104 the data can be provided in any suitable format such as the 5*8 bit vector table lookup indices, which are then used for retrieving the quantized vectors 1538. The concatenated data is provided to a set of BRCT blocks 1612, 1616, 1620, 1624. In this example the set of BRCT blocks 1612, 1616, 1620, 1624 comprises four BRCT blocks. The set of BRCT blocks 1612, 1616, 1620, 1624 is arranged to expand the frequency dimension step by step from 1 to 24. After each BRCT blocks in the set a concatenation block 1614, 1618, 1622, 1626 is provided. The respective concatenation blocks 1614, 1618, 1622, 1626 concatenate the data to be processed with the audio feature based data (energy data) of the same dimension. After the last concatenation block 1626 in the set of BRCT blocks 1612, 1616, 1620, 1624 the data is provided to a BRC block 1628 followed by a further BRCT block 1630. The further BRCT block 1630 processes the data from frame-resolution to the subframe resolution. The output of the further BRCT block 1630 is provided to a concatenation block 1632 which concatenates the data with the audio features 906. The output of this concatenation block 1632 is then provided to two further BRC blocks 1634, 1636. The operations applied by the two further BRC blocks 1634, 1636 do not modify the data size, apart from the feature dimension. The final BRC block 1636 provides 3 feature outputs (per subframe and frequency band), which correspond to the x, y, z vector dimensions. The output of the final BRC block 1636 is provided to a postprocess length block 1638. The postprocess length block 1638 applies a tanh-function to the vector length and scales the result by 1.1. If there were no scaling, the tanh function would have required infinite vector length as an input to be able to provide an output vector with length 1. During inference, the postprocess length block 1638 also upper limits the vector length to 1 after the scaling. The output of the postprocess length block 1638 is the predicted format spatial metadata 1106 which is the output of the decoder machine learning model 1104. Fig. 17 shows example results that were obtained using an example of the disclosure. In this example an encoder machine learning model 908 as shown in Fig. 15 and a decoder machine learning model 1104 as shown in Fig. 16 were used in a system as described above. The respective machine learning models 908, 1104 were implemented in TensorFlow. The example implementation differed from the foregoing description so that the audio encoder 400 and audio decoder 404 were simply passthroughs. The encoding of the spatial metadata was otherwise performed as described above. The approximation of using unencoded audio is representable, since the machine learning models 908, 1104 do not have access to the fine features of the audio signal 102, but only to the 24 bands and 5 milliseconds resolution spectrograms. To obtain the results shown in Fig. 17 two sets of training were performed, using 3899 training examples of 60 seconds length each. The training was run for 1200 epochs, in three stages. For the first 400 epochs the quantizer was turned off so that the encoder machine learning model 908 and decoder machine learning model 1104 were trained with unquantized bottleneck. However, the gaussian noise addition at the bottleneck was present. For the remaining 800 steps, the vector quantizer was also switched on and trained. In some embodiments, the network may be trained in three steps. For example, during the first 400 steps adapting only encoder and decoder without the quantizer at the bottleneck. Then, for 200 steps adapting only the quantizer. Then, for the remaining 600 epochs, not updating the quantizer (but using it) and resuming the training of the encoder machine learning model 908 and decoder machine learning model 1104 otherwise. The test implementation demonstrates two properties. Firstly, the described encoder machine learning model 908 and decoder machine learning model 1104 improves the encoding efficiency by itself. Secondly, using the audio features at the metadata encoding further improves the encoding efficiency of the encoder machine learning model 908 and decoder machine learning model 1104. To demonstrate these properties the training system was run in two different modes: In the first training mode, the processing was as described in the foregoing using the audio features. In the second training mode the audio features were not used and instead the audio spectrogram that the machine learning models 908, 1104 received was replaced with noise. Otherwise, the structure of the machine learning models 908, 110, the training data and the training parameters were identical. Fig. 17 shows the validation loss value Lvec in the two different training modes (with using the audio features, and without using them), compared against same loss value formulated based on spatial metadata that was encoded and decoded with IVAS, in an operating mode that allocated 43 bits per frame for metadata encoding. In comparison, the machine learning models 908, 1104 used only 40 bits per frame. In Fig. 17 the first plot 1700 shows the validation loss value Lvec from a non-machine learning model IVAS metadata encoder, the second plot 1702 shows the validation loss value Lvec from machine learning models 908, 1104 that did not use audio features, and the third plot 1704 shows the validation loss value Lvec from machine learning models 908,1104 that did use audio features. Fig. 17 shows that, in terms of the validation loss value Lvec the machine learning models 908, 1104 clearly exceeded the performance of the non-machine learning model IVAS metadata encoder. The jump in the loss values of Fig. 17 is when the vector quantizer is turned on at 400 steps. When it is turned on, its codebook entries are initially random, and it then converges to represent the data distribution. When the machine learning models 908, 1104 use the audio features, the validation loss value Lvec is further reduced, that is, the metadata is more accurately reconstructed. This indicates that for the given bitrate, if the encoder machine learning model 908 has the access to features of the audio signals 102, it learns to encode the metadata more efficiently. It allocates the available bits more efficiently than the manually-designed allocation method in the IVAS codec, as it does not need to convey information that is available in the audio signals 102. The results shown in Fig. 17 clearly show the advantages of the examples of the disclosure. The use of the machine learning models 908, 1104 is particularly suitable when the bit rates become low like in the present example, as it becomes difficult for non-machine learning means to use the few available bits effectively. Furthermore, the machine learning-based method can take into account audio signal features to further optimize the metadata encoding quality at low bitrates. The validity of the processing of the examples of the disclosure was also further confirmed by using an IVAS rendering system to render a binaural audio output based on the audio signals 102 and the decoded spatial metadata, and the successful encoding and decoding was noticed by listening to the result. Variation to the examples described herein can be used in implementations of the disclosure. For instance, in the examples described above the audio features 906 comprising on the audio spectrogram is used as the feature data. In other examples, the audio features 906 can comprise other features as well, for example, inter-channel information such as coherences and / or delays in frequency bands. Providing further audio features 906 for the machine learning models 908, 1104 can enable a higher coding efficiency to be achieved. In some examples the input to the encoder 106 can be in some other format than described above. For example, the system 100 could be arranged so that the encoder 106 receives multi-microphone signals 204, from which the audio signals 102 and spatial metadata 104 are generated inside the encoder 106. In another example, system 100 can be arranged so that the encoder 106 receives multi-loudspeaker or multi-object signals, or Ambisonic signals. Any of these formats or their mixtures can be converted to audio signals 102 and spatial metadata 104 for encoding. In the examples described above (not the training examples used for Fig. 17), the metadata encoding was based on the decoded audio signals 406. In other examples, the metadata encoding can be performed based on the original audio signals 102. The spectrogram information is very similar for the original audio signals 102 and the decoded audio signals 406, and therefore in some examples the encoder 106 can be arranged to encode the spatial metadata based on information from the original audio signals 102. The decoder may decode the metadata based on information from the decoded audio signals 406. In some examples, the metadata encoder 408 can obtain feature information about the audio signals 102 from the audio encoder 400 that encodes the audio signals 102. For example, the audio encoder 400 may have information of the audio spectral properties, or any other features of the audio signal 102. This information can be conveyed to be used by the metadata encoder 408. Similarly, the decoder machine learning model 1108 can retrieve feature information of the encoded audio signals 402 from the coded representation (for example, from the audio decoder 702). The metadata encoder 408 described above was for a single bitrate but can be adapted to be used in a multitude of bitrates, for example, by training the encoder machine learning model 908 and decoder machine learning model 1104 while varying the number of levels at the residual vector quantizer. Furthermore, the bitrate of the encoding of the audio signal can also be varied during the training of the respective machine learning models 908, 1104 to ensure that the metadata encoder 408 can account for the variations in the encoding of the audio signals 102. The above described examples show encoding of directions and ratio parameters. As an extension or an alternative, other metadata can be included to be encoded. For example, there could be additional input and output features in the encoder machine learning model 908 and decoder machine learning model 1104, which would express the spread and surround coherences or any other spatial metadata. The encoder machine learning model 908 and decoder machine learning model 1104 can be composed of trained signal processing operations that are fixed in the inference time. Therefore, there need not be a strict boundary which operations are considered to be part of the respective machine learning models 908, 1104 and which operations are considered to not be part of the respective machine learning models 908, 1104. As an illustrative example, the entire operation of both machine learning models 908,1104 could be implemented with conventional programming languages. Then there would be no machine learning models 908, 1104 apart from the conventional program code. In the examples described above, the encoder machine learning model 908 viewed several frames of history when generating the metadata bit stream, but the decoder machine learning model 1104 did not In other examples, the decoder machine learning model 1104 can also store history data at the inference time and have operators that account for also previous frames’ data. Similarly, in some examples the encoder machine learning model 908 might only use data that is present in the current frame and would not use history data beyond this. In some examples, the encoder machine learning model 908 and decoder machine learning model 1104 operate on different time and frequency resolutions than described above. For example, when using a different filterbank or different sampling frequency of the audio signals 102. Fig. 18 schematically illustrates an apparatus 1800 that can be used to implement examples of the disclosure. In this example the apparatus 1800 comprises a controller 1802. The controller 1802 can be a chip or a chip-set. In some examples the controller 1802 can be provided within a communications device such as telephone or teleconferencing device or any other suitable type of device that enables audio signals to be communicated. In the example of Fig. 18 the implementation of the controller 1802 can be as controller circuitry. In some examples the controller 1802 can be implemented in hardware alone, have certain aspects in software including firmware alone or can be a combination of hardware and software (including firmware). As illustrated in Fig. 18 the controller 1802 can be implemented using instructions that enable hardware functionality, for example, by using executable instructions of a computer program 1804 in a general-purpose or special-purpose processor 206 that can be stored on a computer readable storage medium (disk, memory etc.) to be executed by such a processor 206. The processor 206 is configured to read from and write to the memory 208. The processor 206 can also comprise an output interface via which data and / or commands are output by the processor 206 and an input interface via which data and / or commands are input to the processor 206. The memory 208 is configured to store a computer program 1804 comprising computer program instructions (computer program code 210) that controls the operation of the controller 1802 when loaded into the processor 206. The computer program instructions, of the computer program 1804, provide the logic and routines that enables the controller 1802 to perform the methods illustrated in the Figs. The processor 206 by reading the memory 208 is able to load and execute the computer program 1804. The apparatus 1800 therefore comprises: at least one processor 206; and at least one memory 208 including computer program code 1804, the at least one memory 208 and the computer program code 210 configured to, with the at least one processor 206, cause the apparatus 1800 at least to perform: obtaining 500 one or more audio signals and associated spatial metadata; determining 502 at least one input for processing wherein the at least one input is based on the obtained one or more audio signals and on the obtained spatial metadata; processing 504 the at least one input to generate processed spatial metadata wherein the processed spatial metadata is formed into a bitstream using a fewer number of bits compared to the required number of bits for the obtained spatial metadata; and forming 506 the bitstream based on the processed spatial metadata and encoded one or more audio signals. The apparatus 1800 therefore comprises: at least one processor 206; and at least one memory 208 including computer program code 1804, the at least one memory 208 and the computer program code 210 configured to, with the at least one processor 206, cause the apparatus 1800 at least to perform: receiving 800 a bitstream comprising encoded one or more audio signals and processed spatial metadata wherein the processed spatial metadata has been formed into a bitstream using a fewer number of bits compared to the required number of bits for originally obtained spatial metadata; determining 802 at least one input for decoding processing wherein the at least one input is based on the encoded one or more audio signals and on the processed spatial metadata; performing 804 decoding processing on the at least one input to generate decoded spatial metadata; and using the decoded spatial metadata and decoded one or more audio signals to enable 806 rendering of a spatial audio output. As illustrated in Fig. 18 the computer program 1804 can arrive at the controller 1800 via any suitable delivery mechanism 1806. The delivery mechanism 1806 can be, for example, a machine readable medium, a computer-readable medium, a non-transitory 47 computer-readable storage medium, a computer program product, a memory device, a record medium such as a Compact Disc Read-Only Memory (CD-ROM) or a Digital Versatile Disc (DVD) or a solid state memory, an article of manufacture that comprises or tangibly embodies the computer program 1804. The delivery mechanism 1806 can be a signal configured to reliably transfer the computer program 1804. The controller 1802 can propagate or transmit the computer program 1804 as a computer data signal. In some examples the computer program 1804 can be transmitted to the controller 1802 using a wireless protocol such as Bluetooth, Bluetooth Low Energy, Bluetooth Smart, 6LoWPan (IPv6 over low power personal area networks) ZigBee, ANT+, near field communication (NFC), Radio frequency identification, wireless local area network (wireless LAN) or any other suitable protocol. The computer program 1804 comprises computer program instructions that when executed by an apparatus 1800 cause the apparatus 1800 to perform at least the following: obtaining 500 one or more audio signals and associated spatial metadata; determining 502 at least one input for processing wherein the at least one input is based on the obtained one or more audio signals and on the obtained spatial metadata; processing 504 the at least one input to generate processed spatial metadata wherein the processed spatial metadata is formed into a bitstream using a fewer number of bits compared to the required number of bits for the obtained spatial metadata; and forming 506 the bitstream based on the processed spatial metadata and encoded one or more audio signals. The computer program 1804 comprises computer program instructions that when executed by an apparatus 1800 cause the apparatus 1800 to perform at least the following: receiving 800 a bitstream comprising encoded one or more audio signals and processed spatial metadata wherein the processed spatial metadata has been formed into a bitstream using a fewer number of bits compared to the required number of bits for originally obtained spatial metadata; determining 802 at least one input for decoding processing wherein the at least one input is based on the encoded one or more audio signals and on the processed spatial metadata; performing 804 decoding processing on the at least one input to generate decoded spatial metadata; and using the decoded spatial metadata and decoded one or more audio signals to enable 806 rendering of a spatial audio output. The computer program instructions can be comprised in a computer program 1804, a non-transitory computer readable medium, a computer program product, a machine readable medium. In some but not necessarily all examples, the computer program instructions can be distributed over more than one computer program 1804. Although the memory 208 is illustrated as a single component / circuitry it can be implemented as one or more separate components / circuitry some or all of which can be integrated / removable and / or can provide permanent / semi-permanent / dynamic / cached storage. Although the processor 206 is illustrated as a single component / circuitry it can be implemented as one or more separate components / circuitry some or all of which can be integrated / removable. The processor 206 can be a single core or multi-core processor. References to “computer-readable storage medium”, “computer program product”, “tangibly embodied computer program” etc. or a “controller”, “computer”, “processor” etc. should be understood to encompass not only computers having different architectures such as single / multi- processor architectures and sequential (Von Neumann) / parallel architectures but also specialized circuits such as field-programmable gate arrays (FPGA), application specific circuits (ASIC), signal processing devices and other processing circuitry. References to computer program, instructions, code etc. should be understood to encompass software for a programmable processor or firmware such as, for example, the programmable content of a hardware device whether instructions for a processor, or configuration settings for a fixed-function device, gate array or programmable logic device etc. As used in this application, the term “circuitry” can refer to one or more or all of the following: (a) hardware-only circuitry implementations (such as implementations in only analog and / or digital circuitry) and (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions and (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g. firmware) for operation, but the software can not be present when it is not needed for operation. This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit for a mobile device or a similar integrated circuit in a server, a cellular network device, or other computing or network device. The apparatus 1800 as shown in Fig. 18 can be provided within any suitable device. In some examples the apparatus 1800 can be provided within an electronic device such as a mobile telephone, a teleconferencing device, a camera, a computing device, a server or any other suitable device. The blocks illustrated in the Figs, can represent steps in a method and / or sections of code in the computer program 1804. The illustration of a particular order to the blocks does not necessarily imply that there is a required or preferred order for the blocks and the order and arrangement of the blocks can be varied. Furthermore, it can be possible for some blocks to be omitted. The apparatus can be provided in an electronic device, for example, a mobile terminal, according to an example of the present disclosure. It should be understood, however, that a mobile terminal is merely illustrative of an electronic device that would benefit from examples of implementations of the present disclosure and, therefore, should not be taken to limit the scope of the present disclosure to the same. While in certain implementation examples, the apparatus can be provided in a mobile terminal, other types of electronic devices, such as, but not limited to: mobile communication devices, hand portable electronic devices, wearable computing devices, portable digital assistants (PDAs), pagers, mobile computers, desktop computers, televisions, gaming devices, laptop computers, cameras, video recorders, GPS devices and other types of electronic systems, can readily employ examples of the present disclosure. Furthermore, devices can readily employ examples of the present disclosure regardless of their intent to provide mobility. The term ‘comprise’ is used in this document with an inclusive not an exclusive meaning. That is any reference to X comprising Y indicates that X may comprise only one Y or may comprise more than one Y. If it is intended to use ‘comprise’ with an exclusive meaning then it will be made clear in the context by referring to ‘comprising only one...’ or by using ‘consisting.’ In this description, the wording ‘connect’, ‘couple’ and ‘communication’ and their derivatives mean operationally connected / coupled / in communication. It should be appreciated that any number or combination of intervening components can exist (including no intervening components), i.e., to provide direct or indirect connection / coupling / communication. Any such intervening components can include hardware and / or software components. As used herein, the term "determine / determining" (and grammatical variants thereof) can include, not least: calculating, computing, processing, deriving, measuring, investigating, identifying, looking up (for example, looking up in a table, a database, or another data structure), ascertaining and the like. Also, "determining" can include receiving (for example, receiving information), accessing (for example, accessing data in a memory), obtaining and the like. Also, "determine / determining" can include resolving, selecting, choosing, establishing, and the like. In this description, reference has been made to various examples. The description of features or functions in relation to an example indicates that those features or functions are present in that example. The use of the term ‘example’ or ‘for example’ or ‘can’ or ‘may’ in the text denotes, whether explicitly stated or not, that such features or functions are present in at least the described example, whether described as an example or not, and that they can be, but are not necessarily, present in some of or all other examples. Thus ‘example’, ‘for example’, ‘can’, or ‘may’ refers to a particular instance in a class of examples. A property of the instance can be a property of only that instance or a property of the class or a property of a sub-class of the class that includes some but not all the instances in the class. It is therefore implicitly disclosed that a feature described with reference to one example but not with reference to another example, can where possible be used in that other example as part of a working combination but does not necessarily have to be used in that other example. As used herein, “at least one of the following: ” and “at least one of ” and similar wording, where the list of two or more elements are joined by “and” or “or” mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements. Although examples have been described in the preceding paragraphs with reference to various examples, it should be appreciated that modifications to the examples given can be made without departing from the scope of the claims. Features described in the preceding description may be used in combinations other than the combinations explicitly described above. Although functions have been described with reference to certain features, those functions may be performable by other features whether described or not. The description of a feature, such as an apparatus or a component of an apparatus, configured to perform a function, or for performing a function, should additionally be considered to also disclose a method of performing that function. For example, description of an apparatus configured to perform one or more actions, or for performing one or more actions, should additionally be considered to disclose a method of performing those one or more actions with or without the apparatus. Although features have been described with reference to certain examples, those features may also be present in other examples whether described or not. The term ‘a’, ‘an’ or ‘the’ is used in this document with an inclusive not an exclusive meaning. That is any reference to X comprising a / an / the Y indicates that X may comprise only one Y or may comprise more than one Y unless the context clearly indicates the contrary. If it is intended to use ‘a’, ‘an’ or ‘the’ with an exclusive meaning then it will be made clear in the context. In some circumstances the use of ‘at least one’ or ‘one or more’ may be used to emphasis an inclusive meaning but the absence of these terms should not be taken to infer any exclusive meaning. The presence of a feature (or combination of features) in a claim is a reference to that feature or (combination of features) itself and to features that achieve substantially the same technical effect (equivalent features). The equivalent features include, for example, features that are variants and achieve substantially the same result in substantially the same way. The equivalent features include, for example, features that perform substantially the same function, in substantially the same way to achieve substantially the same result. In this description, reference has been made to various examples using adjectives or adjectival phrases to describe characteristics of the examples. Such a description of a characteristic in relation to an example indicates that the characteristic is present in some examples exactly as described and is present in other examples substantially as described. The above description describes some examples of the present disclosure however those of ordinary skill in the art will be aware of possible alternative structures and method features which offer equivalent functionality to the specific examples of such structures and features described herein above and which for the sake of brevity and clarity have been omitted from the above description. Nonetheless, the above description should be read as implicitly including reference to such alternative structures and method features which provide equivalent functionality unless such alternative structures or method features are explicitly excluded in the above description of the examples of the present disclosure. 5 Whilst endeavoring in the foregoing specification to draw attention to those features believed to be of importance the Applicant may seek protection via the claims in respect of any patentable feature or combination of features hereinbefore referred to and / or shown in the drawings whether or not emphasis has been placed thereon. 10

Claims

1. An apparatus for encoding spatial audio signals comprising means for: obtaining one or more audio signals and associated spatial metadata; determining at least one input for processing wherein the at least one input is based on the obtained one or more audio signals and on the obtained spatial metadata;processing the at least one input to generate processed spatial metadata wherein the processed spatial metadata is formed into a bitstream using a fewer number of bits compared to the required number of bits for the obtained spatial metadata; andforming the bitstream based on the processed spatial metadata and encoded one or more audio signals.

2. An apparatus as claimed in claim 1 wherein determining at least one input for processing comprises:determining a first input for processing wherein the first input is based on the obtained one or more audio signals; anddetermining a second input for processing wherein the second input is based on the obtained spatial metadata.

3. An apparatus as claimed in any preceding claim wherein forming the bitstream comprising the processed spatial metadata and encoded one or more audio signals comprises encoding the obtained one or more audio signals and combining the encoded one or more audio signals with the processed spatial metadata generated by the processing.

4. An apparatus as claimed in any of claims 1 to 2 wherein the processing provides a bitstream comprising the processed spatial metadata and encoded one or more audio signals as an output.

5. An apparatus as claimed in any preceding claim wherein determining the at least one input for the processing comprises encoding the obtained one or more audio signals.

6. An apparatus as claimed in any preceding claim wherein determining the at least one input for the processing comprises at least partially decoding encoded one or more audio signals.

7. An apparatus as claimed in any preceding claim wherein determining the at least one input for the processing comprises determining one or more features by processing at least one of:the obtained one or more audio signals;one or more decoded audio signals;one or more partially decoded audio signals.

8. An apparatus as claimed in any preceding claim wherein the means are for enabling transmission of the bitstream comprising the processed spatial metadata and encoded one or more audio signals.

9. An apparatus as claimed in any preceding claim wherein the processed spatial metadata is generated so as to be decoded by decoding processing in an apparatus for decoding spatial audio signals.

10. An apparatus as claimed in any preceding claim wherein the encoded one or more audio signals are obtained by encoding the obtained one or more audio signals.

11. An apparatus as claimed in any preceding claim wherein the processing is performed, at least in part, using a machine learning model.

12. An apparatus as claimed in claim 11 wherein the machine learning model is trained in conjunction with a decoding machine learning model for an apparatus for decoding spatial audio signals.

13. An apparatus for decoding spatial audio signals comprising means for: receiving a bitstream comprising encoded one or more audio signals and processed spatial metadata wherein the processed spatial metadata has been formed into a bitstream using a fewer number of bits compared to the required number of bits for originally obtained spatial metadata;determining at least one input for decoding processing wherein the at least one input is based on the encoded one or more audio signals and on the processed spatial metadata;performing decoding processing on the at least one input to generate decoded spatial metadata; andusing the decoded spatial metadata and decoded one or more audio signals to enable rendering of a spatial audio output.

14. An apparatus as claimed in claim 13 wherein determining at least one input for decoding processing comprises:determining a first input for a decoding processing based on the encoded one or more audio signals; anddetermining a second input for the decoding processing based on the processed spatial metadata;15. An apparatus as claimed in any of claims 13 to 14 wherein determining the at least one input for the decoding processing comprises decoding the encoded one or more audio signals received in the bitstream.

16. An apparatus as claimed in any of claims 13 to 14 wherein determining the at least one input for the decoding processing comprises processing the encoded one or more audio signals to determine one or more features.

17. An apparatus as claimed in any of claims 13 to 16 wherein the spatial audio output comprises at least one of:binaural output;multi-loudspeaker output;Ambisonics output;stereo output; andcross talk-cancelled stereo output.

18. An apparatus as claimed in any of claims 13 to 19 wherein the processed spatial metadata is generated by processing in an apparatus for encoding spatial audio signals.

19. An apparatus as claimed in any of claims 13 to 18 wherein the decoding processing is performed, at least in part, using a machine learning model.

20. An apparatus as claimed in claim 19 wherein the machine learning model is 5 trained in conjunction with a machine learning model for an apparatus for encoding spatial audio signals.

Citation Information

Patent Citations

  • Metadata processing for first order ambisonics

    WO2023088560A1