Signalling of pass-through mode in spatial audio coding

By repurposing bits from MASA metadata and transport channels signal, the modified signaling mechanism addresses the challenge of encoding multiple audio objects with multichannel audio, ensuring accurate decoding and separation of audio objects at low bit rates, enhancing immersive audio experiences.

GB2640563APending Publication Date: 2025-10-29NOKIA TECHNOLOGIES OY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
GB2024005792
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-25
Publication Date
2025-10-29

AI Technical Summary

Technical Problem

Existing immersive audio codecs face challenges in efficiently encoding and decoding spatial audio signals at low bit rates, particularly when multiple audio objects are combined with multichannel audio, as there are insufficient bits to signal the number of audio objects, leading to inseparability of audio objects from the overall spatial audio mix.

Method used

A modified signaling mechanism repurposes bits from the MASA metadata and transport channels signal to encode the number of audio objects, allowing for accurate decoding and rendering of combined audio signals by allocating additional bits for encoding these signals, even at low bit rates.

Benefits of technology

Ensures that audio objects can be correctly identified and separated at the decoder, maintaining the integrity of the audio scene even in low-bitrate conditions, thereby improving the decoding process for immersive audio experiences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A Metadata Assisted Spatial Audio (MASA) encoder (for eg. Immersive Voice Audio Services IVAS) encodes audio objects using two bits to indicate the number of audio objects input to the encoder. If it determines that one or two objects are input, a further bit encodes whether it is one or two, otherwise the further bit is used to encode a combined signal comprising an analysed multichannel audio signal and an analysed audio objects signal.
Need to check novelty before this filing date? Find Prior Art

Description

Field The present application relates to apparatus and methods for signalling for spatial audio encoding and decoding Background Parametric spatial audio processing is a field of audio signal processing where the spatial aspect of the sound is described using a set of parameters. For example, in parametric spatial audio capture from microphone arrays, it is a typical and an effective choice to estimate from the microphone array signals a set of parameters such as directions of the sound in frequency bands, and the ratios between the directional and non-directional parts of the captured sound in frequency bands. These parameters are known to well describe the perceptual spatial properties of the captured sound at the position of the microphone array. These parameters can be utilized in synthesis of the spatial sound accordingly, for headphones binaurally, for loudspeakers, or to other formats, such as Ambisonics. The directions and direct-to-total energy ratios in frequency bands are thus a parameterization that is particularly effective for spatial audio capture. A parameter set consisting of a direction parameter in frequency bands and an energy ratio parameter in frequency bands (indicating the directionality of the sound) can be also utilized as the spatial metadata (which may also include other parameters such as surround coherence, spread coherence, number of directions, distance etc) for an audio codec. For example, these parameters can be estimated from microphone-array captured audio signals, and for example a stereo or mono signal can be generated from the microphone array signals to be conveyed with the spatial metadata. The stereo signal could be encoded, for example, with an AAC encoder and the mono signal could be encoded with an EVS encoder. A decoder can decode the audio signals into PCM signals and process the sound in frequency bands (using the spatial metadata) to obtain the spatial output, for example a binaural output. Immersive audio codecs are being implemented supporting a multitude of operating points ranging from a low bit rate operation to transparency. An example of such a codec is the Immersive Voice and Audio Services (IVAS) codec which is designed to be suitable for use over a communications network such as a 3GPP 4G / 5G network including use in such immersive services as for example immersive voice and audio for virtual reality (VR). This audio codec handles the encoding, decoding and rendering of speech, music and generic audio. It furthermore supports channelbased audio and scene-based audio inputs including spatial information about the sound field and sound sources. The codec operates with low latency to enable conversational services as well as support high error robustness under various transmission conditions. The aforementioned immersive audio codecs are particularly suitable for encoding captured spatial sound from microphone arrays (e.g., in mobile phones, VR cameras, stand-alone microphone arrays). Furthermore, such an encoder can have other input types, for example, loudspeaker signals, audio object signals, Ambisonic signals. Summary According to a first aspect there is an apparatus configured to: encode two bits with a value representing a number of audio objects at the input to an audio encoder; determine whether the number of audio objects at the input to the encoder is either a first number of audio objects or a second number of audio objects; when the number of audio objects at the input to the audio encoder is determined to be either the first number of audio objects or the second number of audio objects, encode a further bit with a value which distinguishes between the first number of audio objects and the second number of audio objects; and when the number of audio objects at the input to the audio encoder is determined to be neither the first number of audio objects nor the second number of audio objects, use the further bit for the encoding of a combined audio signal, wherein the combined audio signal comprises an analysed multichannel audio signal and an analysed audio objects signal. The analysed multichannel audio signal may comprise at least one transport audio signal and spatial audio parameter metadata, and the analysed audio objects signal may comprise at least one audio object transport audio signal and audio object spatial audio parameter metadata. The two bits maybe reserved from a group of bits for encoding the spatial audio parameter metadata and the audio object spatial audio parameter metadata, and wherein the further bit may be associated with the at least one transport audio signal. The apparatus configured to encode two bits with a value representing a number of audio objects at the input to an audio encoder may be configured to; encode the two bits with a single two-bit value for an instance of the number of audio objects being either the first number of audio objects or the second number of audio objects; and encode the two bits with a further two-bit value for an instance of the number of audio objects not being the first number of audio objects or the second number of audio objects. The first number of audio objects may be one and the second number of audio objects may be two. The audio encoder may be an audio encoder according to the Immersive Voice Audio Service (IVAS) standard, the spatial audio parameter metadata may be MASA metadata according to the IVAS standard and the further bit may be a MASA number of transport channel signal bit according to the IVAS standard. According to a second aspect there is an apparatus configured to: decode two bits to give a value representing a number of audio objects; determine whether the value representing the number of audio objects is either a first number of audio objects or a second number of audio objects; when the value representing the number of audio objects is determined to be either a first number of audio objects or a second number of audio objects decode a further bit, wherein the further bit is encoded with a value which distinguishes between the first number of audio objects and the second number of audio objects; and when the value representing the number of audio objects is determined to be neither the first number of audio objects nor the second number of audio objects use the further bit for the decoding of a combined audio signal, wherein the combined audio signal comprises an analysed multichannel audio signal and an analysed audio objects signal. The analysed multichannel audio signal may comprise at least one transport audio signal and spatial audio parameter metadata, and the analysed audio objects signal may comprise at least one audio object transport audio signal and audio object spatial audio parameter metadata. The two bits can be reserved from a group of bits for encoding the spatial audio parameter metadata and the audio object spatial audio parameter metadata, and wherein the further bit may be associated with the at least one transport audio signal. The apparatus configured to decode two bits to give a value representing a number of audio objects may be configured to; decode a single two-bit value of the two bits for an instance of the number of audio objects being either the first number of audio objects or the second number of audio objects; and decode a further two-bit value of the two bits for an instance of the number of audio objects not being the first number of audio objects or the second number of audio objects. The first number of audio objects may be one and the second number of audio objects may be two. The audio decoder may be an audio decoder according to the Immersive Voice Audio Service (IVAS) standard, the spatial audio parameter metadata may be MASA metadata according to the IVAS standard and the further bit may be a MASA number of transport channel signal bit according to the IVAS standard. According to a third aspect there is a method comprising: encoding two bits with a value representing a number of audio objects at the input to an audio encoder; determining whether the number of audio objects at the input to the encoder is either a first number of audio objects or a second number of audio objects; when the number of audio objects at the input to the audio encoder is determined to be either the first number of audio objects or the second number of audio objects, encoding a further bit with a value which distinguishes between the first number of audio objects and the second number of audio objects; and when the number of audio objects at the input to the audio encoder is determined to be neither the first number of audio objects nor the second number of audio objects, using the further bit for the encoding of a combined audio signal, wherein the combined audio signal comprises an analysed multichannel audio signal and an analysed audio objects signal. The analysed multichannel audio signal may comprise at least one transport audio signal and spatial audio parameter metadata, and the analysed audio objects signal may comprise at least one audio object transport audio signal and audio object spatial audio parameter metadata. The two bits maybe reserved from a group of bits for encoding the spatial audio parameter metadata and the audio object spatial audio parameter metadata, and wherein the further bit may be associated with the at least one transport audio signal. The method comprising encoding two bits with a value representing a number of audio objects at the input to an audio encoder may comprise; encoding the two bits with a single two-bit value for an instance of the number of audio objects being either the first number of audio objects or the second number of audio objects; and encoding the two bits with a further two-bit value for an instance of the number of audio objects not being the first number of audio objects or the second number of audio objects. The first number of audio objects may be one and the second number of audio objects may be two. The audio encoder may be an audio encoder according to the Immersive Voice Audio Service (IVAS) standard, the spatial audio parameter metadata may be MASA metadata according to the IVAS standard and the further bit may be a MASA number of transport channel signal bit according to the IVAS standard. According to a fourth aspect there is a method comprising decoding two bits to give a value representing a number of audio objects; determining whether the value representing the number of audio objects is either a first number of audio objects or a second number of audio objects; when the value representing the number of audio objects is determined to be either a first number of audio objects or a second number of audio objects decoding a further bit, wherein the further bit is encoded with a value which distinguishes between the first number of audio objects and the second number of audio objects; and when the value representing the number of audio objects is determined to be neither the first number of audio objects nor the second number of audio objects using the further bit for the decoding of a combined audio signal, wherein the combined audio signal comprises an analysed multichannel audio signal and an analysed audio objects signal. The analysed multichannel audio signal may comprise at least one transport audio signal and spatial audio parameter metadata, and the analysed audio objects signal may comprise at least one audio object transport audio signal and audio object spatial audio parameter metadata. The two bits can be reserved from a group of bits for encoding the spatial audio parameter metadata and the audio object spatial audio parameter metadata, and wherein the further bit may be associated with the at least one transport audio signal. The method comprising decoding two bits to give a value representing a number of audio objects may comprise; decoding a single two-bit value of the two bits for an instance of the number of audio objects being either the first number of audio objects or the second number of audio objects; and decoding a further two-bit value of the two bits for an instance of the number of audio objects not being the first number of audio objects or the second number of audio objects. The first number of audio objects may be one and the second number of audio objects may be two. The audio decoder may be an audio decoder according to the Immersive Voice Audio Service (IVAS) standard, the spatial audio parameter metadata may be MASA metadata according to the IVAS standard and the further bit may be a MASA number of transport channel signal bit according to the IVAS standard. According to a fifth aspect an apparatus comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to: encode two bits with a value representing a number of audio objects at the input to an audio encoder;_determine whether the number of audio objects at the input to the encoder is either a first number of audio objects or a second number of audio objects; when the number of audio objects at the input to the audio encoder is determined to be either the first number of audio objects or the second number of audio objects, encode a further bit with a value which distinguishes between the first number of audio objects and the second number of audio objects; and when the number of audio objects at the input to the audio encoder is determined to be neither the first number of audio objects nor the second number of audio objects, use the further bit for the encoding of a combined audio signal, wherein the combined audio signal comprises an analysed multichannel audio signal and an analysed audio objects signal. According to a sixth aspect an apparatus comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to: decode two bits to give a value representing a number of audio objects; determine whether the value representing the number of audio objects is either a first number of audio objects or a second number of audio objects; when the value representing the number of audio objects is determined to be either a first number of audio objects or a second number of audio objects decode a further bit, wherein the further bit is encoded with a value which distinguishes between the first number of audio objects and the second number of audio objects; and when the value representing the number of audio objects is determined to be neither the first number of audio objects nor the second number of audio objects use the further bit for the decoding of a combined audio signal, wherein the combined audio signal comprises an analysed multichannel audio signal and an analysed audio objects signal. An apparatus configured to perform the actions of the method as described above. A computer program comprising program instructions for causing a computer to perform the method as described above. A computer program product stored on a medium may cause an apparatus to perform the method as described herein. An electronic device may comprise apparatus as described herein. A chipset may comprise apparatus as described herein. Embodiments of the present application aim to address problems associated with the state of the art. Summary of the Figures For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings in which: Figure 1 shows schematically a system or apparatus suitable for implementing the lowest rate coding mode (mode A); Figure 2 shows schematically an example encoding mode selector as shown in the system of apparatus as shown in Figure 1 according to some embodiments; Figure 3 shows a flow diagram of the operation of the example encoding mode selector shown in Figure 2 according to some embodiments; Figure 4 shows a flow diagram of the operation of the example first, lowest rate coding mode (mode A), or only MASA bitrate encoding mode shown in Figure 3 according to some embodiments; Figure 5 shows a flow diagram of the operation of the audio object number encoder 110 shown in Figure 1; Figure 6 shows schematically an example of the audio decoder &renderer 129 shown in Figure 1 operating in a pass through mode with respect to an embodiment of the first encoding mode; Figure 7 shows a flow diagram of the operation of the audio object number determiner 1013 shown in Figure 6; Figure 8 shows an example device suitable for implementing the apparatus shown in previous figures. Embodiments of the Application The following describes in further detail suitable apparatus and possible mechanisms for the encoding of parametric spatial audio signals comprising transport audio signals and spatial metadata. As indicated above immersive audio codecs (such as 3GPP TS 26.253 Immersive Voice and Audio Services (IVAS)) support a multitude of operating points ranging from a low bit rate operation to transparency. It supports channel-based audio and scene-based audio inputs including spatial information about the sound field and sound sources. In the following the example codec is configured to be able to receive multiple input formats. In particular the codec is configured to obtain or receive a multi audio signal (for example a spatial audio metadata and transport audio signal(s) analysed from a microphone array, or as a multi-channel audio format input, an ambisonics format input) and one or more audio object signal (these can also be called an independent stream with metadata - ISM format). Furthermore, in some situations the codec is configured to handle more than one input format at a time. This combined (input) format mode can, for example, enable simultaneous encoding of two different audio input formats. An example of two different audio input formats which IVAS currently supports is the combination of the MASA format with audio object format. Metadata-Assisted Spatial Audio (MASA) is an example of a parametric spatial audio format and representation suitable as an input format for IVAS. It can be considered as an audio representation consisting of ‘N channels + spatial metadata’. It is a scene-based audio format particularly suited for spatial audio capture on practical devices, such as smartphones. The idea is to describe the sound scene in terms of time- and frequency-varying sound source directions and, e.g., energy ratios. Sound energy that is not defined (described) by the directions, is described as diffuse (coming from all directions). As discussed above spatial metadata associated with the audio signals may comprise multiple parameters (such as multiple directions and associated with each direction (or directional value) a direct-to-total energy ratio, spread coherence, distance, etc.) per time-frequency tile. The spatial metadata may also comprise other parameters or may be associated with other parameters which are considered to be non-directional (such as surround coherence, diffuse-to-total energy ratio, remainder-to-total energy ratio) but when combined with the directional parameters are able to be used to define the characteristics of the audio scene. For example, a reasonable design choice which is able to produce a good quality output is one where the spatial metadata comprises one or more directions for each time-frequency subframe (and associated with each direction direct-to-total ratios, spread coherence, distance values etc) are determined. As described above, parametric spatial metadata representation can use multiple concurrent spatial directions. With MASA, the proposed maximum number of concurrent directions is two. For each concurrent direction, there may be associated parameters such as: Direction index; Direct-to-total ratio; Spread coherence; and Distance. In some embodiments other parameters such as Diffuse-to-total energy ratio; Surround coherence; and Remainder-to-total energy ratio are defined. The IVAS codec is also configured to operate in what is known as a “pass through” mode where the format of the (spatial) audio output at the decoder / renderer is of the same form as the input to the IVAS encoder. For instance, taking the above example of encoding the audio input combination of the MASA format with the audio object format. Then an IVAS codec operating in “pass through” mode would produce an output at the decoder / renderer which adheres to the same format as the audio input, which in this case would be an output combination of MASA format and audio object format. Furthermore, the IVAS coding model also provides for encoding of a combined format at various bitrates. This coding model enables a parametric encoding of the audio object input that includes variable rate encoding for an object direction parameter. As will be seen below when the IVAS codec provides for the simultaneous encoding of audio object format and MASA format, this pairing of formats can be encoded using one of four separate modes depending on the available bit rate. For the two lowest bit rate coding modes the encoding of some or all audio objects are amalgamated with the encoding of the MASA format resulting in an audio output where audio objects are inseparable from the overall spatial audio mix. For the third coding mode the audio objects can be separated to some degree by using sent parametric information relating to the audio objects. In the fourth mode, which is the highest bit rate coding mode, all audio objects are separable from the overall spatial audio mix at the decoder / renderer. Therefore, when the IVAS codec is operating in a “pass through” mode the decoder / renderer is able to recreate the same format of output as the input to the codec irrespective of the choice of encoding mode. For instance when the audio objects and MASA input have been encoded using an encoding mode which amalgamates the encoding of the audio objects with the MASA input such that the individual audio objects cannot be recovered from the decoded audio streams, such as the first two encoding modes (mode A and mode B), the decoder is still able to produce an output with the same format as the input to the encoder. . However, when the IVAS codec is operating in a “pass through” mode at the lowest rate coding mode (known as mode A) there is an insufficient number of bits which would allow for the number of individual audio objects at the input to the encoder to be encoded into the bitstream using dedicated bits from the encoded bitstream. One potential solution to this problem is discussed in the patent application GB2315543.5 which describes the repurposing of two bits from the bit allocation for the encoding of the MASA metadata for the frame in order to convey an indication of the number of audio objects. In GB2315543.5 these “reserved” two bits can be configured to signal the number of audio objects according to the following table. - ‘00’ indicates only a multichannel audio signal / MASA input - ‘01’ indicates a multichannel audio signal / MASA input with 1 audio object input - ‘10’ indicates a multichannel audio signal / MASA input with 2 audio object input - ‘11’ indicates a multichannel audio signal / MASA input with either 3 or 4 audio objects. Furthermore GB2315543.5 describes the repurposing of the MASA number of transport channels signal bit from its original role of signalling the number of transport audio channels to a new role of performing the function of distinguishing between 3 or 4 audio objects when a “reserved” two-bit code of ‘11’ is used. As a note, when the IVAS encoder is operating solely as a MASA input format encoder (the MASA input format being the combination of the MASA spatial audio parameter metadata and the transport audio signals), the encoder can be configured to operate with either one or two transport audio channel(s) which is signalled using the MASA number of transport channels signal bit. However, when the IVAS encoder is operating in the mode which supports the simultaneous encoding of the MASA input format and audio object format the number of transport audio channels / signals is consistently set as two. Consequently, the MASA number of transport channels signal bit can be repurposed for use elsewhere, such as for signalling the number of audio objects when the encoder is operating in mode A When the IVAS codec is operating in mode A the input audio objects are analysed where the respective audio object transport audio signals are combined with the (MASA) transport audio signals to form a combined encoded transport audio signal and the audio object MASA metadata is combined with the multichannel MASA metadata to form a combined encoded MASA metadata stream There is no specific encoding directed to the actual individual audio object per se. Therefore, the contribution of the audio objects to the overall audio scene is solely contained within the audio mix of the combined transport audio signals and the spatial audio scene of the combined MASA metadata. Therefore, it would be advantageous to allocate additional bits for the encoding of these signals. Embodiments of the present application aim to address this problem associated with the state of the art. The concept as discussed in further detail herein proceeds on the basis of providing a signalling mechanism which carries information relating to the number of audio objects which have subjected to combined encoding with the MASA metadata. This information can then be used at the decoder / renderer as part of the IVAS decoding process to recreate a same number of audio objects as that presented at the encoder. In this regard Figure 1 depicts a general apparatus 100 and system for providing the joint coding of the MASA input format and the audio object format for the lowest rate coding mode (mode A). The system is shown with an analyser 111. The analyser 111 is the part from receiving the multi-channel audio signals 102 up to the encoding of the MASA metadata 114 (derived from the multichannel audio signals) and transport audio signals 106 (derived from downmixing the multichannel audio signals.) The input to the system analyser 111 is the multi-channel audio signals 102. In the following examples a microphone channel signal input is described, however any suitable input (or synthetic multi-channel) format may be implemented in other embodiments. For example, in some embodiments the spatial analyser and the spatial analysis may be implemented external to the encoder. For example, in some embodiments the (spatial) MASA metadata associated with the multichannel audio signals 102 may be provided to an encoder as a separate bit-stream. Additionally, Figure 1 also depicts multiple audio objects 104 (each audio object comprising an audio object signal and audio object metadata) as a further input to the analysis part. As mentioned above these multiple audio objects (or audio object stream) 104 may represent various sound sources within a physical space. Each audio object may be characterized by an audio (object) signal and accompanying metadata comprising directional data (in the form of azimuth and elevation values) which indicate the position or direction of the audio object within a physical space on an audio frame basis. The multi-channel audio signals 102 are passed to the analyser 111, and specifically a transport signal generator 105 and to a MASA metadata generator 103. In some embodiments the MASA metadata generator 103 is also configured to receive the multi-channel signals 102 and analyse the signals to produce multichannel MASA metadata 114 which are the spatial audio parameters associated with the multi-channel audio signals 102, and also the transport audio signals 106 which are the downmixed audio signals associated with the multichannel audio signals 102. The MASA metadata generator 103 may be configured to generate the multichannel MASA metadata 114 which can comprise, for each time-frequency analysis interval, a direction parameter and an energy ratio parameter and a coherence parameter (and in some embodiments a diffuseness parameter). The direction, energy ratio and coherence parameters may in some embodiments be considered to be MASA spatial audio parameters ( multichannel MASA metadata 114). In other words, the spatial audio parameters comprise parameters which aim to characterize the sound-field created / captured by the multi-channel signals (or two or more audio signals in general). In some embodiments the parameters generated can differ from frequency band to frequency band. Thus, for example in band X all of the parameters are generated and transmitted, whereas in band Y only one of the parameters is generated and transmitted, and furthermore in band Z no parameters are generated or transmitted. A practical example of this can be that for some frequency bands such as the highest band some of the parameters are not required for perceptual reasons. The transport audio signals 106 and the multichannel MASA metadata 114 may be passed to a combined encoder core 109. In some embodiments the transport signal generator 105 is configured to receive the multi-channel audio signals 102 and generate a suitable transport signal comprising a determined number of channels and output the transport audio signals 106 (MASA transport audio signals). For example, the transport signal generator 105 can be configured to generate a 2-audio channel downmix of the multi-channel audio signals 102. The determined number of channels may be any suitable number of channels. The transport signal generator 105 in some embodiments is configured to otherwise select or combine, for example, by beamforming techniques the input audio signals to the determined number of channels and output these as transport signals. In some embodiments the transport signal generator 105 is optional and the multichannel audio signals 102 are passed unprocessed to a combined encoder core 109 in the same manner as the transport signal are in this example. The audio objects 104 can be passed to the audio object analyser 107 for processing. For mode A, the audio object analyser 107 analyses the object audio input stream 104 in order to produce suitable audio object transport signals and audio object metadata. For example, the audio object analyser can be configured to produce the audio object transport audio signals 126 by downmixing the audio signals of the audio objects into a stereo channel together using amplitude panning based on the associated audio object directions. Additionally, the audio object analyser 107 can also be configured to produce the MASA format metadata based on the audio objects 104 (audio object MASA metadata 108). This metadata 108 may then be fed to the combined encoder core 109. For other modes of operation, the audio object analyser 107 can be configured to produce the audio object metadata (rather than MASA format metadata for mode A) associated with the audio object input stream 104. The audio object metadata (not shown) may comprise direction values which are applicable for all sub-bands. So, if there are 4 objects, there are 4 directions. In the examples described herein the direction values also apply across all of the subframes of the frame, but in some embodiments the temporal resolution of the direction values can differ and the directions values apply for one or more than one sub-frames of the frame. Furthermore, energy ratios (or ISM ratios) may be determined for each object. The energy ratio (ISM ratio) defines the contribution of the object within the object part of the total audio environment. In the following examples the energy ratios (or ISM ratios) are for each time-frequency tile for each object. The audio object analyser 107 in the other modes of operation can be additionally configured to determine an audio object for separation from the stream of audio objects 104 and create an separated object audio signal (not shown). In embodiments this can be a mono audio signal. The separated audio object audio signal may then be encoded by an audio signal encoder (such as an EVS encoder) to form an encoded separated object audio signal. This encoded separated object audio signal can then form part of the bitstream 118. The audio signal encoder may be a mono channel encoder of IVAS. Additionally, an identifier index of the separated audio object can form part of the bitstream 118. In some embodiments, the audio object analyser 107 may be sited elsewhere and the audio objects 104 input to the analyser and encoder 101 is audio objects transport signals 126 and audio object MASA metadata 108. The encoder 101 can comprise a combined encoder core 109 which is configured to receive the transport audio (for example downmix) signals 106 and audio objects transport signals 126 in order to combine and generate a suitable encoding of these audio signals into the encoded transport audio signals 138. Additionally, the combined encoder core 109 is arranged to receive the multichannel MASA metadata 114, and in mode A the combined encoder core 109 also receives the audio object MASA metadata 108. The combined encoder core 109 is also arranged to form the encoded MASA metadata stream 116, which for mode A will also include the combining of the audio object MASA metadata 108 and multichannel MASA metadata 114. For other encoding modes, the encoder 101 can also comprise an audio object metadata encoder (not shown) which is similarly configured to receive the audio object metadata and output an encoded or compressed form of the input information as encoded audio object metadata (not shown). For other encoding modes, the combined encoder core 109 is also configured to implement a stream separation metadata determiner and encoder which can be configured to determine the relative contributory proportions of the multi-channel audio signals 102 (MASA audio signals) and audio objects 104 to the overall audio scene. The encoder 101 is also depicted in Figure 1 as comprising a bitstream generator 113 configured to obtain the encoded MASA metadata 116, the encoded transport audio signals 138 and generate the bitstream 118 for potential transmission or storage. The encoder 101 can comprise an encoder controller 115. The encoder controller 115 can control the encoding implemented by the combined encoder core 109. The encoder controller 115 is configured to determine the bitrate for the bitstream 118 and based on the bitrate control the encoding. The encoder controller 115 is further configured to control at least one of the audio object analyser 107, transport signal generator 105 and MASA metadata generator 103 in generating parameters. The encoder 101 can be a computer or mobile device (running suitable software stored on memory and on at least one processor), or alternatively a specific device utilizing, for example, FPGAs or ASICs. The encoding may be implemented using any suitable scheme. In some embodiments the encoder may further interleave, multiplex to a single data stream or embed the encoded MASA metadata, audio object metadata and stream separation metadata within the encoded (downmixed) transport audio signals before transmission or storage shown in Figure 1 by the dashed line. The multiplexing may be implemented using any suitable scheme. Furthermore, with respect to Figure 1 is shown an associated decoder and renderer 129 operating in mode A, which is configured to obtain the bitstream 118 comprising encoded MASA metadata 116 and encoded transport audio signals 138 and from these generate suitable spatial audio output signals. The decoding and processing of such audio signals in relation to decoding in a “pass- through” mode of operation will be discussed later on. With respect to Figure 2 is shown in further detail the encoder controller 115 according to some embodiments. In this example the encoder controller 115 comprises a bitrate determiner / monitor 201 configured to determine and / or monitor the available bitrate for the bandwidth for the encoded audio and metadata. This could be determined based on a transmission path bandwidth estimation (and for example be based on an estimated signal strength) or a bandwidth storage determination to maintain the file for a determined time to be below a required size or by any suitable manner. The bitrate determiner / monitor 201 can furthermore be configured to control an encoding mode selector 203. The encoder controller 115 can comprise an encoding mode selector 203 configured to select an encoding mode, for example based on the determined bandwidth, bitrate and / or number of audio objects since the overall bit rate is distributed between the encoding of the MASA metadata &transport signals and the audio objects. Once the encoding mode has been determined the encoder controller 115 then controls the encoders, for example the combined encoder core 109 (and audio object metadata encoder for non-mode A coding modes), according to the encoding mode. Bandwidth and bitrate information can also be passed onto the audio object analyser 107 so that this information can be used [by the audio object analyser 107] to determine whether audio objects are to be separated in accordance with the second and third modes of encoding. As will be seen later these modes are also referred to as mode B and mode C respectively. With respect to Figure 3 is shown a flow diagram of an example operation of the encoder controller shown in Figure 2 In this example there is an initial operation of receiving or obtaining or otherwise determining the bitrate or bandwidth for encoded parameters and audio data as shown in Figure 3 by step 301. Having obtained the available bandwidth or bitrate then a check can be made to determine whether the bitrate is below a first (or lowest) threshold limit as shown in Figure 3 by step 303. Where the available bandwidth or bitrate is below the first (or lowest) threshold limit then the encoders can be controlled to encode the transport channels and MASA metadata only in mode A as shown in Figure 3 by step 304. Where the available bandwidth or bitrate is above the first (or lowest) threshold limit then a further check can be made to determine whether the bitrate is below a second (or lower or one audio object) threshold limit as shown in Figure 3 by step 305. Where the available bandwidth or bitrate is below the second (or lower or one object) threshold limit then the encoders can be controlled to encode the combined input of MASA and audio objects in a mode B. Therefore, in mode B the encoders can be controlled to encode in a combined manner the multichannel transport audio channels / signals, multichannel MASA metadata, remaining audio objects MASA metadata, remaining audio objects transport audio signals, one audio object data (separated audio object metadata and separated audio object signal) and some in embodiments the one audio object identifier as shown Figure 3 by step 306. Where the available bandwidth or bitrate is above the second (or lower or one object) threshold limit then a further check can be made to determine whether the bitrate is below a third, higher or full object threshold limit as shown in Figure 3 by step 307. Where the available bandwidth or bitrate is below the third, higher or full object threshold limit then the encoders can be controlled to encode transport channels, MASA metadata, ISM metadata (all objects), MASA to total ratios, ISM ratios, and 1 object audio data, with 1 object identifier (also shown as Mode C) as shown in Figure 3 by step 308. Where the available bandwidth or bitrate is above the third, higher or full object threshold limit then the encoders can be controlled to encode transport channels, MASA metadata, ISM metadata (all objects), All objects audio data (also shown as Mode D) as shown in Figure 3 by step 310. The encoding modes can for example be summarised by the following table: Mode Format Encoded parameters A MASA_FORMAT -transport channels - MASA metadata - input format - number of input objects B I SM_M AS A_M 0 D E_MAS A_O N E_O B J - Transport channels of mix of MASA and remaining audio objects - MASA metadata of mix of input MASA metadata and remaining audio objects equivalent MASA metadata - 1 separated object audio data 1 separated object metadata C I SM_MAS A_MO D E_PARAM_O N E_O B J - Transport channels of mix of MASA and remaining audio objects - MASA metadata - all objects metadata - MASA to total ratios - ISM ratios -1 object audio data -1 object identifier D ISM_MASA_MODE_DISC - MASA Transport channels - MASA metadata - ISM metadata - all objects audio data Mode Bitrate 1obj Bitrate 2 obj Bitrate 3obj Bitrate 4obj A <24kbps <32kbps <32kbps <32kbps B - - 32kbps 48kbps 32kbps 48kbps C 32kbps 64kbps 80kbps 64kbps 80kbps 96kbps D >=24kbps >=48kbps >=96kbps >=128kbps The bitrates shown herein are examples and it would be understood that they can be other specific values. With respect to Figure 4 there is shown flow a diagram showing a first (or lowest or combined) encoding mode as shown in Figure 3 by step 304 in further detail. Thus, for very low total bitrates (for example less or equal to 32kbps) all encoding is implemented using a MASA representation. Thus, for example there is an operation of receiving / obtaining the object based streams (independent streams with metadata (ISM)) and multichannel based (MASA stream) transport audio signals and metadata as shown in Figure 4 by step 401. Then, as shown in Figure 4 by step 403, there is an operation of generating an object based MASA stream from the object streams (independent streams with metadata). This object based MASA stream can in some embodiments be created from the object stream using, for example, the methods presented in WO2019086757A1. After this, as shown in Figure 4 by step 405, the object based MASA stream and multichannel based MASA stream are combined. In some embodiments the original MASA stream and the MASA stream created from the objects can be combined using the method presented in GB2574238. The decoder received the information about the audio objects and the MASA audio content in the MASA format. Then the combined stream is output as shown in Figure 4 by step 407. In such embodiments the object audio content (together with the MASA audio content) is present in the decoded audio scene, but the objects cannot be edited nor separated from the scene at the decoder based on the decoded parameters. For the sake of brevity, the other encoding modes are not described herein. Details of these other encoding modes can be found in patent applications PCT / EP2023 / 080897 and GB2315543. Returning to Figure 1, the decoder &renderer 129 is also arranged to receive a further input (140) which indicates whether the decoder is to operate in a “pass through” EXT mode. As explained previously the decoder requires information relating to the number of audio objects at the encoder so that the decoder can operated in a “pass though” EXT mode. In some of the above combined encoding modes the number of audio objects is explicitly included into the frame header as part of the encoding process. However, other combined coding modes, especially the lower bit rate modes, may not have the bit rate capacity to include the number of audio objects in each frame header. In these cases, the encoder is arranged to repurpose other bits in the bitstream in order to encode information relating to the number of audio objects 104. For instance, when the lowest coding rate mode (mode A) is selected for encoding, two bits may be reserved from the bit allocation for a frame of the encoded MASA metadata 116, and as mentioned earlier the MASA number of transport channels signal bit can also be repurposed for coding the number of audio objects 104. As mentioned above there already exists a mechanism (from GB2315543.5) for signalling the number of audio objects at the lowest encoding rate (mode A). However, as shall be seen below the current signalling mechanism can be modified such that an additional bit can be “freed up” for the purposes of encoding the (combined) transport audio signals and (combined) MASA metadata (i.e. the encoding of the transport audio signals 106 and the audio objects transport audio signals 126, and the encoding of the audio object MASA metadata and multichannel MASA metadata)when the input to the encoder comprises a higher number of audio objects (i.e for the cases of 3 and 4 audio objects). To that end, the tables below outline the modified signalling mechanism for A mode. Firstly, the two bits reserved from the encoded frame for the MASA data can be allocated according to Table 1 below. ‘00’ indicates only a multichannel audio signal / MASA input - ‘01’ indicates a multichannel audio signal / MASA input with 4 audio object input - ‘10’ indicates a multichannel audio signal / MASA input with 3 audio object input - ‘11 ’ indicates a multichannel audio signal / MASA input with either 1 or 2 audio objects. It is to be understood that the two reserved bits can be encoded differently. However, the allocations of two-bit codes should adhere to the principle that the signalling of a MASA input with either 1 or 2 audio objects is encoded as single two-bit code. For example, a different legitimate mapping of the above two-bit codes can entail: ‘01’ indicating a multichannel audio signal / MASA input with either 1 or 2 audio objects; ‘10’ indicates a multichannel audio signal / MASA input with 3 audio object input; ‘11’ indicates a multichannel audio signal / MASA input with 4 audio object input; and ‘00’ indicates only a multichannel audio signal / MASA input. In other words, any combination of two-bits codes is permissible as long as the encoding scheme does not encode either of 1 audio object and 2 audio objects with a unique two-bit code. Secondly, the MASA number of transport channels signal bit can be used to discern the difference between 1 and 2 audio objects by using Table 2 below. - ‘0’ indicates a multichannel audio signal / MASA input with 1 audio object input - ‘1’ indicates a multichannel audio signal / MASA input with 2 audio objects input. Consequently, when the number of input audio objects is either 3 or 4 the MASA number of transport channels signal bit is not required to distinguish between the aforementioned numbers of audio objects. Instead, the MASA number of transport channels bit can be repurposed for use in the combined encoding of the transport audio signals 106 and audio object transport audio signals 128 and the combined encoding of the audio object MASA metadata 108 and multichannel MASA metadata 114. That is the MASA number of transport channels signal bit can be added to (or allocated to) the common pool of bits for encoding the combined MASA metadata and combined transport audio signals. Therefore, when compared to the current signalling scheme as laid out in the current IVAS specification 3GPP 26.253 v2.0, the above modified scheme offers the advantage of repurposing the MASA number of transport channels signal bit for use in the encoding of the audio transport signals 106 and 128 and MASA metadata 108 and 114 when the overall audio mix comprises the higher number of audio objects. As a note, when modes B and C are selected as the method for combined encoding then information relating to the number of audio objects 104 is included in the frame header of the encoded bitstream. Thus, for these encoding modes, the combined encoder core 109 is not required to use the above reserved bits from the MASA metadata stream. With respect to Figure 5 there is shown a flow diagram show for implementing in the above modified mode A signalling scheme for an IVAS encoder such as that depicted in 101 in Figure 1. To that end the encoder can comprise an audio object number encoder 110 which can be arranged to receive a count (audio object count 117) signifying the number of audio objects 104 (step 501 in Figure 5). The audio object number encoder 110 is arranged to use the object count input (117) in order to encode the number of audio objects into the bitstream accordingly. As described above, this can be performed initially by using the two reserved bits from the bits allocated for the encoding of the audio frame’s MASA metadata (MASA metadata frame) according to Table 1 above. The process of encoding the number of audio objects into the reserved bits of the MASA metadata frame is shown as step 502 in Figure 5. Sequentially, or in parallel the audio object number encoder 110 is also arranged to check whether the encoding of step 502 results (or would result) in the case of whether the number of audio objects is either 1 or 2. This can be performed by determining whether the two reserved bits have been encoded as ’11or by simply inspecting the received audio object count 117. Figure 5 depicts this check as the processing step 503. When the number of audio objects at processing step 503 is determined to be 1 or 2 the process can proceed to step 504 where the MASA number of transport channels signal bit is encoded according to Table 2 above. When the number of audio objects is determined (at step 503) as being neither 1 nor 2, the process can proceed to step 505 where the MASA number of transport channels signal bit is allocated to the pool of bits for the encoding the combined transport audio signals (Multichannel and audio object) and combined MASA metadata (multichannel and audio object) Steps 503 and 505 result in a signal being sent to the combined encoding core 109 along the connection 130 (Figure 1) indicating that the MASA number of transport channels signal bit is to be used for the encoding of the combined transport audio signals and combined MASA metadata Steps 502 and 504 result in the encoded MASA number of transport channels signal bit and encoded reserved two bits (of the MASA data frame) being conveyed to the Bitstream generator 113 for incorporation into the bitstream 118. This is depicted in Figure 1 as the connections 131 and 132. With respect to Figure 6 there is shown a decoder for use in “pass through” mode (EXT mode) when combined encoding according to mode A has been used. Specifically for each frame, the demux 160 is arranged to separate the reserved two bits from the encoded MASA metadata frame 1005 together with the MASA number of transport channels signal bit 1006 from the bitstream 118 and feed them to the audio objects number determiner 1013. The audio objects number determiner 1013 is then arranged to determine the number of audio objects by using in part the above allocation Tables 1 and 2 to give the number of audio objects value 1007. Also, shown in Figure 10 is the connection 10060 which directs the MASA number of transport channels signal bit to the MASA decoder core 1003. The number of audio objects value 1007 is passed to the “null” audio object producer 1015. The “null” audio object producer is arranged to produce as an output the number of audio objects value of “null” audio objects. Where a “null” audio object is in effect a dummy audio object having a zero valued audio object signal and a zero valued audio object metadata. The “null” audio object serves as place holder for future development of the IVAS coding system. Additionally, for each IVAS frame, the demux 160 is arranged to extract the encoded MASA metadata 1002 and the encoded transport audio signals 1138 from the bitstream 118. These are then decoded by the MASA decoder 1003 to provide decoded MASA signal (comprising decoded MASA metadata 10041 and transport audio signals 10042) 1004. The decoded MASA metadata signal 10041 is the decoded resultant signal from combining at the encoder the multichannel MASA metadata and the audio object MASA metadata, and likewise the decoded transport audio signals 10042 is the resultant signal from combining at the encoder the transport audio signals and the audio object transport audio signals. The output frame 150 from the decoder &renderer 129 when operating in EXT mode is the combination of decoded MASA signal 1004 and the requisite number of “null” audio objects 1009. One advantage of having the same number of audio objects (whether they are either actual audio objects or “null” audio objects) at the output of the IVAS decoder as found at the input to the IVAS encoder is that same renderer can be used for all bitrates and encoder operating modes. Returning to the audio objects number determiner 1013. Figure 7 shows a flow diagram of the processing steps performed by the determiner 1013. As mentioned above the audio object number determiner 1013 can be arranged to receive the reserved two bits from the MASA (encoded) metadata frame 1005 together with the MASA number of transport channels signal bit 1006. These are shown as the processing steps 701 and 702 in Figure 7. The audio object number determiner 1013 is then arranged decode the two reserved bits from the MASA metadata frame using Table 1 in order to determine the number of audio objects. This is shown as processing step 703 in Figure 7. Upon decoding the two bits reserved from the MASA metadata frame, the audio object number determiner is able to determine whether either the number of audio objects is one of 3 or 4, that is the reserved two bits will hold either of the values of ‘01’ or ‘10’, or whether the reserved two bits decodes to the case corresponding to code ‘11’ from Table 1. In other words, the code (‘11’) which indicates that the number of audio objects is one of two possibilities (1 or 2) and therefore subsequently requires the decoding of the MASA number of transport channels signal bit in order to distinguish between the two possibilities. This determining step is shown as 704 in Figure 7. When at step 704 it is determined that the reserved two bits of the MASA metadata frame indicates that the number of audio objects is either 3 or 4 number of audio objects, the number of audio objects determiner 1013 proceeds to step 705. At step 705 the number of audio objects determiner 1013 is arranged to pass a signal to the MASA decoder core 1003 indicating that the MASA number of transport channels signal bit is to be used for the decoding of the encoded transport audio signals 1138 and encoded MASA metadata 1002. With respect to Figure 6, this signal is depicted by the link 1006. Furthermore, at step 705 the number of audio objects determiner 1013 can also be configured to pass the number of audio objects via the output 1007, which in this case will either take the value of 3 or 4 depending on the encoded value of the reserved two bits. When at step 704 it is determined from the reserved two bits that the number of audio objects is one of two possibilities (1 or 2 audio objects), the number of audio objects determiner 1013 proceeds to step 706. At step 706 the audio objects number determiner 1013 is arranged to decode the MASA number of transport channels signal bit in order to distinguish which of the two possibilities gives the number of audio objects. In embodiments step 706 can be performed by using Table 2. The output 1007 from step 706 is therefore the number of audio objects as given by the decoded MASA number of transport channels signal bit, which will be either 1 audio object or 2 audio objects. With respect to Figure 8 an example electronic device which may be used as any of the apparatus parts of the system as described above. The device may be any suitable electronics device or apparatus. For example, in some embodiments the device 1700 is a mobile device, user equipment, tablet computer, computer, audio playback apparatus, etc. The device may for example be configured to implement the encoder / analyser part and / or the decoder part as shown in Figures 1 and 8 or any functional block as described above. In some embodiments the device 1700 comprises at least one processor or central processing unit 1707. The processor 1707 can be configured to execute various program codes such as the methods such as described herein. In some embodiments the device 1700 comprises at least one memory 1711. In some embodiments the at least one processor 1707 is coupled to the memory 1711. The memory 1711 can be any suitable storage means. In some embodiments the memory 1711 comprises a program code section for storing program codes implementable upon the processor 1707. Furthermore, in some embodiments the memory 1711 can further comprise a stored data section for storing data, for example data that has been processed or to be processed in accordance with the embodiments as described herein. The implemented program code stored within the program code section and the data stored within the stored data section can be retrieved by the processor 1707 whenever needed via the memory-processor coupling. In some embodiments the device 1700 comprises a user interface 1705. The user interface 1705 can be coupled in some embodiments to the processor 1707. In some embodiments the processor 1707 can control the operation of the user interface 1705 and receive inputs from the user interface 1705. In some embodiments the user interface 1705 can enable a user to input commands to the device 1700, for example via a keypad. In some embodiments the user interface 1705 can enable the user to obtain information from the device 1700. For example the user interface 1705 may comprise a display configured to display information from the device 1700 to the user. The user interface 1705 can in some embodiments comprise a touch screen or touch interface capable of both enabling information to be entered to the device 1700 and further displaying information to the user of the device 1700. In some embodiments the user interface 1705 may be the user interface for communicating. In some embodiments the device 1700 comprises an input / output port 1709. The input / output port 1709 in some embodiments comprises a transceiver. The transceiver in such embodiments can be coupled to the processor 1707 and configured to enable a communication with other apparatus or electronic devices, for example via a wireless communications network. The transceiver or any suitable transceiver or transmitter and / or receiver means can in some embodiments be configured to communicate with other electronic devices or apparatus via a wire or wired coupling. The transceiver can communicate with further apparatus by any suitable known communications protocol. For example in some embodiments the transceiver can use a suitable radio access architecture based on long term evolution advanced (LTE Advanced, LTE-A) or new radio (NR) (or can be referred to as 5G), universal mobile telecommunications system (UMTS) radio access network (UTRAN or E-UTRAN), long term evolution (LTE, the same as E-UTRA), 2G networks (legacy network technology), wireless local area network (WLAN or Wi-Fi), worldwide interoperability for microwave access (WiMAX), Bluetooth®, personal communications services (PCS), ZigBee®, wideband code division multiple access (WCDMA), systems using ultra-wideband (UWB) technology, sensor networks, mobile ad-hoc networks (MANETs), cellular internet of things (loT) RAN and Internet Protocol multimedia subsystems (IMS), any other suitable option and / or any combination thereof. The transceiver input / output port 1709 may be configured to receive the signals. In some embodiments the device 1700 may be employed as at least part of the synthesis device. The input / output port 1709 may be coupled to headphones (which may be a headtracked or a non-tracked headphones) or similar and loudspeakers. In general, the various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto. While various aspects of the invention may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof. The embodiments of this invention may be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD. The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processors may be of any type suitable to the local technical environment, and may include one or more of general-purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), gate level circuits and processors based on multi-core processor architecture, as non-limiting examples. Embodiments of the inventions may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate. Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre-stored design modules. Once the design for a semiconductor circuit has been completed, the resultant design, in a standardized electronic format (e.g., Opus, GDSII, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication. As used in this application, the term “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry) and (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation. This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device. The term “non-transitory,” as used herein, is a limitation of the medium itself (i.e., tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM). As used herein, “at least one of the following: ” and “at least one of ” and similar wording, where the list of two or more elements are joined by “and” or “or”, mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements. 5 The foregoing description has provided by way of exemplary and non-limiting examples a full and informative description of the exemplary embodiment of this invention. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. 10 However, all such and similar modifications of the teachings of this invention will still fall within the scope of this invention as defined in the appended claims.

Claims

1. An apparatus configured to:encode two bits with a value representing a number of audio objects at the input to an audio encoder;determine whether the number of audio objects at the input to the encoder is either a first number of audio objects or a second number of audio objects;when the number of audio objects at the input to the audio encoder is determined to be either the first number of audio objects or the second number of audio objects, encode a further bit with a value which distinguishes between the first number of audio objects and the second number of audio objects; andwhen the number of audio objects at the input to the audio encoder is determined to be neither the first number of audio objects nor the second number of audio objects, use the further bit for the encoding of a combined audio signal, wherein the combined audio signal comprises an analysed multichannel audio signal and an analysed audio objects signal.

2. The apparatus as claimed in Claim 1, wherein the analysed multichannel audio signal comprises at least one transport audio signal and spatial audio parameter metadata, and wherein the analysed audio objects signal comprises at least one audio object transport audio signal and audio object spatial audio parameter metadata.

3. The apparatus as claimed in Claim 2, wherein the two bits are reserved from a group of bits for encoding the spatial audio parameter metadata and the audio object spatial audio parameter metadata, and wherein the further bit is associated with the at least one transport audio signal.

4. The apparatus as claimed in Claims 1 to 3, wherein the apparatus configured to encode two bits with a value representing a number of audio objects at the input to an audio encoder is configured to;encode the two bits with a single two-bit value for an instance of the number of audio objects being either the first number of audio objects or the second number of audio objects; andencode the two bits with a further two-bit value for an instance of the number of audio objects not being the first number of audio objects or the second number of audio objects.

5. The apparatus as claimed in Claims 1 to 4, wherein the first number of audio objects is one and the second number of audio objects is two.

6. The apparatus as claimed in Claims 1 to 5, wherein the audio encoder is an audio encoder according to the Immersive Voice Audio Service (IVAS) standard, wherein the spatial audio parameter metadata is MASA metadata according to the IVAS standard and wherein the further bit is a MASA number of transport channel signal bit according to the IVAS standard.

7. An apparatus configured to:decode two bits to give a value representing a number of audio objects;determine whether the value representing the number of audio objects is either a first number of audio objects or a second number of audio objects;when the value representing the number of audio objects is determined to be either a first number of audio objects or a second number of audio objects decode a further bit, wherein the further bit is encoded with a value which distinguishes between the first number of audio objects and the second number of audio objects; andwhen the value representing the number of audio objects is determined to be neither the first number of audio objects nor the second number of audio objects use the further bit for the decoding of a combined audio signal, wherein the combined audio signal comprises an analysed multichannel audio signal and an analysed audio objects signal.

8. The apparatus as claimed in Claim 7, wherein the analysed multichannel audio signal comprises at least one transport audio signal and spatial audio parametermetadata, and wherein the analysed audio objects signal comprises at least one audio object transport audio signal and audio object spatial audio parameter metadata.

9. The apparatus as claimed in Claim 8, wherein the two bits are reserved from a group of bits for encoding the spatial audio parameter metadata and the audio object spatial audio parameter metadata, and wherein the further bit is associated with the at least one transport audio signal.

10. The apparatus as claimed in Claims 7 to 9, wherein the apparatus configured to decode two bits to give a value representing a number of audio objects is configured to;decode a single two-bit value of the two bits for an instance of the number of audio objects being either the first number of audio objects or the second number of audio objects; anddecode a further two-bit value of the two bits for an instance of the number of audio objects not being the first number of audio objects or the second number of audio objects.

11. The apparatus as claimed in Claims 7 to 10, wherein the first number of audio objects is one and the second number of audio objects is two.

12. The apparatus as claimed in Claims 7 to 11, wherein the audio decoder is an audio decoder according to the Immersive Voice Audio Service (IVAS) standard, wherein the spatial audio parameter metadata is MASA metadata according to the IVAS standard and wherein the further bit is a MASA number of transport channel signal bit according to the IVAS standard.

13. A method comprising:encoding two bits with a value representing a number of audio objects at the input to an audio encoder;determining whether the number of audio objects at the input to the encoder is either a first number of audio objects or a second number of audio objects;when the number of audio objects at the input to the audio encoder is determined to be either the first number of audio objects or the second number of audio objects, encoding a further bit with a value which distinguishes between the first number of audio objects and the second number of audio objects; andwhen the number of audio objects at the input to the audio encoder is determined to be neither the first number of audio objects nor the second number of audio objects, using the further bit for the encoding of a combined audio signal, wherein the combined audio signal comprises an analysed multichannel audio signal and an analysed audio objects signal.

14. A method comprising:decoding two bits to give a value representing a number of audio objects;determining whether the value representing the number of audio objects is either a first number of audio objects or a second number of audio objects;when the value representing the number of audio objects is determined to be either a first number of audio objects or a second number of audio objects decoding a further bit, wherein the further bit is encoded with a value which distinguishes between the first number of audio objects and the second number of audio objects; andwhen the value representing the number of audio objects is determined to be neither the first number of audio objects nor the second number of audio objects using the further bit for the decoding of a combined audio signal, wherein the combined audio signal comprises an analysed multichannel audio signal and an analysed audio objects signal.

Citation Information

Patent Citations

  • Spatial audio parameter merging

    GB2574238A

  • Parametric spatial audio decoding with pass-through mode

    GB2634524A

  • Determination of targeted spatial audio parameters and associated spatial audio playback

    WO2019086757A1

  • Parametric spatial audio encoding

    WO2024115051A1