Apparatus and methods

The apparatus and methods for muting specific audio components in immersive codecs address inefficiencies by identifying and managing mute requests, enhancing audio control and resource optimization in dynamic environments.

GB2640667APending Publication Date: 2025-11-05NOKIA TECHNOLOGIES OY
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
GB2024006053
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-30
Publication Date
2025-11-05

AI Technical Summary

Technical Problem

Existing immersive audio codecs lack efficient mechanisms for muting specific components or streams within immersive audio services, particularly in dynamic and unpredictable environments like virtual reality, augmented reality, and spatial voice communication, which can lead to unwanted audio signals and resource inefficiencies.

Method used

Implementing apparatus and methods for identifying and muting specific components or streams within immersive audio bitstreams by reducing energy, applying silence descriptors, or adjusting bit rates, using IVAS encoder and decoder protocols to manage mute requests and unmute requests through RTP and RTCP packets.

Benefits of technology

Enables precise control over audio signals, reducing unwanted audio components, optimizing resource usage, and ensuring seamless audio experience in immersive environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A Metadata Assisted Spatial Audio (MASA) encoder (for eg. Immersive Voice Audio Services IVAS) encodes audio objects in a bitstream and implements discontinuous adaptation for muting. One component or stream of the immersive audio bitstream to be muted is identified, a mute request 10 is transmitted to another apparatus 120 , and a bitstream having reduced energy, reduced bit rate or a silence descriptor indicator is received in return. The mute request may be transmitted via eg. a Real Time Transport (RTP) packet or header.
Need to check novelty before this filing date? Find Prior Art

Description

Field The present application relates to apparatus and methods for implementing discontinuous adaptation for muting within an immersive voice and audio service environment. Background Immersive audio codecs are being implemented supporting a multitude of operating points ranging from a low bit rate operation to transparency. An example of such a codec is the immersive voice and audio services (IVAS) codec which is designed to be suitable for use over a communications network such as a 3GPP 4G / 5G network. Such immersive services include uses for example in immersive voice and audio for applications such as virtual reality (VR), augmented reality (AR) and mixed reality (MR) as well as spatial voice communication including teleconferencing. This audio codec handles the encoding, decoding and rendering of speech, music and generic audio. It provides backwards interoperable EVS mono operation and furthermore supports stereo, scene-based audio (SBA), metadata-assisted spatial audio (MASA), channel-based audio, object-based audio, and certain combinations of these input formats. The codec operates with low latency to enable conversational services as well as support high error robustness under various transmission conditions. The input signals are presented to the IVAS encoder in one of the supported formats (and in some allowed combinations of the formats). In addition, combinations of formats are supported such as: Objects with MASA (OMASA) and Objects with SBA (OSBA). The IVAS output formats include mono, stereo, multichannel (including custom loudspeaker layouts), FOA, HOA2, HOA3, and binaural. In addition, so-called pass-through operation is possible allowing, e.g., MASA output for MASA input. Similarly, the decoder can output the audio in several supported formats including a pass-through operation, where the audio can be provided in its original format after transmission (encoding / decoding). As a spatial audio codec supporting at least three degrees of rotation freedom (yaw, pitch, roll) for all spatial inputs, the IVAS codec is expected to be used in a variety of scenarios, all of which cannot be known beforehand. The IVAS codec algorithm is described in 3GPP TS 26.253 (Codec for Immersive Voice and Audio Services; Detailed Algorithmic Description incl. RTP payload format and SDP parameter definitions), currently at v2.0.0. Furthermore the IVAS codec floating-point C code is provided in 3GPP TS 26.258. Additionally RTP (Real-Time Transport Protocol) is intended for an end-to-end, real-time transfer of streaming media and provides facilities for jitter compensation and detection of packet loss and out-of-order delivery. RTP allows data transfer to multiple destinations through IP multicast or to a specific destination through IP unicast. The majority of the RTP implementations are built on top of the User Datagram Protocol (UDP). Other transport protocols may also be utilized. RTP is used in together with other protocols such as H.323 and Real Time Streaming Protocol (RTSP). The RTP specification describes two protocols: RTP and RTCP. RTP is used for the transfer of multimedia data, and its companion protocol (RTCP) is used to periodically send control information and QoS (Quality of Service) parameters. For a class of applications (e.g., audio, video), a RTP profile may be defined. For a media format (e.g., a specific video coding format), an associated RTP payload format may be defined. Every instantiation of RTP in a particular application may require a profile and payload format specifications. The profile defines the codecs used to encode the payload data and their mapping to payload format codes in the protocol field Payload Type (PT) of the RTP header. For example, the RTP profile for audio and video conferences with minimal control is defined in RFC 3551. The profile defines a set of static payload type assignments, and a dynamic mechanism for mapping between a payload format, and a PT value using Session Description Protocol (SDP). The latter mechanism is used for newer video codecs such as RTP payload format for H.264 Video defined in RFC 6184 or RTP Payload Format for High Efficiency Video Coding (HEVC) defined in RFC 7798. IVAS RTP payload format is currently being developed, and the latest state is described in TS 26.253 Annex A (v2.0.0, S4-240376). Summary There is provided according to a first aspect an apparatus comprising means configured to: obtain, from a further apparatus, at least one bitstream comprising at least one component or stream of an immersive audio bitstream; identify at least one of the at least one component or stream of the immersive audio bitstream to be muted; transmit, to a further apparatus, a mute request indicating the identified at least one component or stream to be muted; obtain, from the further apparatus, the at least one bitstream, wherein the at least one bitstream is modified with respect to the identified at least one component or stream of the at least one immersive audio bitstream to be muted, wherein the modification in the at least one bitstream, is at least one of: the at least one component or stream having reduced energy; the at least one component or stream not being received; the at least one component or stream comprises a silence descriptor indicator for indicating the at least one component or stream is to be muted; the at least one component or stream having a reduced bit rate; and at least one other component or stream having an increased bit rate. The means configured to obtain, from the further apparatus, the bitstream, may be configured to depacketize and decode the at least one immersive audio bitstream from the bitstream. The depacketized and decoded at least one immersive audio bitstream may comprise at least one decoded IVAS frame, the at least one decoded IVAS frame comprising one of: at least one component having been the identified at least one component to be muted such that the at least one component has reduced energy; or at least one stream having been the identified at least one stream to be muted such that the at least one stream has reduced energy. The silence descriptor indicator for indicating the at least one component or stream may comprise at least one silence descriptor frame within the at least one component or stream. The means may be further configured to render from the obtained immersive audio bitstream at least one audio signal for playback wherein the at least one component or stream of an immersive audio bitstream to be muted, wherein the means configured to render is configured to generate the at least one audio signal for playback comprising at least one of: the identified at least one component or stream of the immersive audio bitstream with a significantly decreased signal energy compared to a signal energy of the identified at least one component or stream of the immersive audio bitstream prior to the mute request; a selectively muted at least one component or stream; and comfort noise to selectively replace the at least one component or stream. The mute request may comprise metadata comprising an identifier field for indicating the identified at least one component or stream to be muted. The metadata may further comprise a validity field configured to identify a time over which the muting is to be applied. The means configured to obtain, from the further apparatus, the bitstream or modified bitstream may be configured to obtain, from the further apparatus, metadata for identifying a status of the at least one component or stream. The metadata for identifying a status of the at least one component or the stream may be configured to indicate which component or stream is muted. The metadata may be a processing information frame. The means configured to transmit to the further apparatus the mute request indicating the identified at least one component or stream to be muted may be configured to, at least one of: transmit the mute request within an IVAS RTP payload header; transmit the mute request within a RTP header extension; transmit the mute request within a RTCP payload; transmit the mute request using an IVAS E-byte; and transmit the mute request using IVAS PI data. The means may be configured to transmit, to the further apparatus, an unmute request indicating the identified at least one component or stream to be unmuted. The means configured to render may be further configured to generate the at least one audio signal for playback comprising at least one of: the identified at least one component or stream of the immersive audio bitstream with a significantly increased signal energy following the unmute request compared to the signal energy of the identified at least one component or stream of the immersive audio bitstream following the mute request; a selectively unmuted at least one component or stream; and the at least one component or stream to selectively replace the comfort noise. The means configured to transmit to the further apparatus the unmute request may be configured to, at least one of: transmit the unmute request within an IVAS RTP payload header; transmit the unmute request within a RTP header extension; transmit the unmute request within a RTCP payload; transmit the unmute request using an IVAS E-byte; and transmit the mute request using IVAS PI data. The means may be further configured to negotiate with the further apparatus support for handling the mute request indicating the identified at least one component or stream to be muted. The means configured to negotiate with the further apparatus support for handling the mute request indicating the identified at least one component or stream to be muted may be configured to negotiate employing a session description file. The means configured to identify at least one component or stream of the immersive audio bitstream is to be muted may be configured to obtain at least one mute input for identifying the at least one component or stream of the immersive audio bitstream. The at least one immersive audio bitstream may comprise at least two components. According to a second aspect there is provided an apparatus comprising means configured to: transmit to a further apparatus, at least one bitstream comprising at least one component or stream of an immersive audio bitstream; obtain, from a further apparatus, a mute request identifying at least one component or stream from the at least one component or stream to be muted; encode, based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream; transmit, to the further apparatus, the modified at least one bitstream, wherein the modified at least one bitstream comprising the identified at least one component or stream, wherein the modification in the at least one bitstream, is at least one of: the at least one component or stream having reduced energy; the at least one component or stream not being transmitted; the at least one component or stream comprises a silence descriptor indicator for indicating the at least one component or stream is to be muted; the at least one component or stream having a reduced bit rate; and at least one other component or stream having an increased bit rate. The means configured to encode, based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream may be configured to packetize and encode the identified at least one component or stream. The means configured to encode, based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream may be configured to employ an IVAS encoder to generate at least one encoded IVAS frame, the at least one encoded IVAS frame comprising one of: at least one component, identified by the mute request; or at least one stream, identified by the mute request. The means configured to encode, based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream may be configured to: apply a gain control processing to the identified at least one component or stream; encode at least the gain control processed at least one component or stream to generate the modified at least one bitstream. The means configured to encode at least the gain control processed at least one component or stream to generate the modified at least one bitstream may be configured to generate an encoded IVAS frame comprising at least one silence descriptor frame based on a signal energy of the gain controlled at least one component or stream being below the IVAS encoder signal level threshold. The means configured to encode at least the gain control processed at least one component or stream to generate a modified at least one bitstream may be configured to adaptively encode the at least one component or stream and / or the at least one other component or stream based on at least one of: a signal level of the gain controlled at least one component or stream; a relative signal level of the gain controlled at least one component or stream and at least one other component or stream. The mute request may comprise metadata comprising an identifier field for indicating the identified at least one component or stream to be muted. The metadata may further comprise a validity field configured to identify a time over which the muting is to be applied. The means configured to transmit, to the further apparatus, the modified at least one bitstream may be configured to transmit, to the further apparatus, metadata for identifying a status of the identified at least one component or stream. The metadata for identifying a status of the identified at least one component or stream may be configured to indicate the at least one component or at least one stream is muted. The metadata may be a processing information frame. The means configured to receive from the further apparatus the mute request indicating the identified at least one component or stream to be muted may be configured to, at least one of: receive the mute request within an IVAS RTP payload header; receive the mute request within a RTP header extension; receive the mute request within a RTCP payload; receive the mute request using an IVAS E-byte; and receive the mute request using IVAS PI data. The means may be configured to obtain, from the further apparatus, an unmute request indicating the identified at least one component or stream to be unmuted. The means may be further configured to de-apply the muting to the indicated at least one component or stream identified by the unmute request. The means configured to receive from the further apparatus the unmute request indicating the identified at least one component or stream to be unmuted may be configured to, at least one of: receive the unmute request within an IVAS RTP payload header; receive the unmute request within a RTP header extension; receive the unmute request within a RTCP payload; receive the unmute request using an IVAS E-byte; and receive the unmute request using IVAS PI data. The means may be further configured to negotiate with the further apparatus support for handling the mute request indicating the identified the at least one component or stream to be muted. The means configured to negotiate with the further apparatus support for handling the mute request indicating the identified at least one component or stream to be muted may be configured to negotiate employing a session description file. According to a third aspect there is provided a method for an apparatus, the method comprising: obtaining, from a further apparatus, at least one bitstream comprising at least one component or stream of an immersive audio bitstream; identifying at least one of the at least one component or stream of the immersive audio bitstream to be muted; transmitting, to a further apparatus, a mute request indicating the identified at least one component or stream to be muted; obtaining, from the further apparatus, the at least one bitstream, wherein the at least one bitstream is modified with respect to the identified at least one component or stream of the at least one immersive audio bitstream to be muted, wherein the modification in the at least one bitstream, is at least one of: the at least one component or stream having reduced energy; the at least one component or stream not being received; the at least one component or stream comprises a silence descriptor indicator for indicating the at least one component or stream is to be muted; the at least one component or stream having a reduced bit rate; and at least one other component or stream having an increased bit rate. Obtaining, from the further apparatus, the bitstream, may comprise depacketizing and decoding the at least one immersive audio bitstream from the bitstream. The depacketized and decoded at least one immersive audio bitstream may comprise at least one decoded IVAS frame, the at least one decoded IVAS frame comprising one of: at least one component having been the identified at least one component to be muted such that the at least one component has reduced energy; or at least one stream having been the identified at least one stream to be muted such that the at least one stream has reduced energy. The silence descriptor indicator for indicating the at least one component or stream may comprise at least one silence descriptor frame within the at least one component or stream. The method may further comprise rendering from the obtained immersive audio bitstream at least one audio signal for playback wherein the at least one component or stream of an immersive audio bitstream to be muted, wherein rendering may comprise generating the at least one audio signal for playback comprising at least one of: the identified at least one component or stream of the immersive audio bitstream with a significantly decreased signal energy compared to a signal energy of the identified at least one component or stream of the immersive audio bitstream prior to the mute request; a selectively muted at least one component or stream; and comfort noise to selectively replace the at least one component or stream. The mute request may comprise metadata comprising an identifier field for indicating the identified at least one component or stream to be muted. The metadata may further comprise a validity field configured to identify a time over which the muting is to be applied. Obtaining, from the further apparatus, the bitstream or modified bitstream may comprise obtaining, from the further apparatus, metadata for identifying a status of the at least one component or stream. The metadata for identifying a status of the at least one component or the stream may be configured to indicate which component or stream is muted. The metadata may be a processing information frame. Transmitting to the further apparatus the mute request indicating the identified at least one component or stream to be muted may comprise at least one of: transmitting the mute request within an IVAS RTP payload header; transmitting the mute request within a RTP header extension; transmitting the mute request within a RTCP payload; transmitting the mute request using an IVAS E-byte; and transmit the mute request using IVAS PI data. The method may further comprise transmitting, to the further apparatus, an unmute request indicating the identified at least one component or stream to be unmuted. Rendering may further comprise generating the at least one audio signal for playback comprising at least one of: the identified at least one component or stream of the immersive audio bitstream with a significantly increased signal energy following the unmute request compared to the signal energy of the identified at least one component or stream of the immersive audio bitstream following the mute request; a selectively unmuted at least one component or stream; and the at least one component or stream to selectively replace the comfort noise. Transmitting to the further apparatus the unmute request may comprise at least one of: transmitting the unmute request within an IVAS RTP payload header; transmitting the unmute request within a RTP header extension; transmitting the unmute request within a RTCP payload; transmitting the unmute request using an IVAS E-byte; and transmitting the mute request using IVAS PI data. The method may further comprise negotiating with the further apparatus support for handling the mute request indicating the identified at least one component or stream to be muted. Negotiating with the further apparatus support for handling the mute request indicating the identified at least one component or stream to be muted may comprise negotiating employing a session description file. Identifying at least one component or stream of the immersive audio bitstream is to be muted may comprise obtaining at least one mute input for identifying the at least one component or stream of the immersive audio bitstream. The at least one immersive audio bitstream may comprise at least two components. According to a fourth aspect there is provided a method for an apparatus, the method comprising: transmitting to a further apparatus, at least one bitstream comprising at least one component or stream of an immersive audio bitstream; obtaining, from a further apparatus, a mute request identifying at least one component or stream from the at least one component or stream to be muted; encoding, based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream; transmitting, to the further apparatus, the modified at least one bitstream, wherein the modified at least one bitstream comprising the identified at least one component or stream, wherein the modification in the at least one bitstream, is at least one of: the at least one component or stream having reduced energy; the at least one component or stream not being transmitted; the at least one component or stream comprises a silence descriptor indicator for indicating the at least one component or stream is to be muted; the at least one component or stream having a reduced bit rate; and at least one other component or stream having an increased bit rate. Encoding, based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream may comprise packetizing and encoding the identified at least one component or stream. Encoding, based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream may comprise employing an IVAS encoder to generate at least one encoded IVAS frame, the at least one encoded IVAS frame comprising one of: at least one component, identified by the mute request; or at least one stream, identified by the mute request. Encoding, based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream may comprising: applying a gain control processing to the identified at least one component or stream; encoding at least the gain control processed at least one component or stream to generate the modified at least one bitstream. Encoding at least the gain control processed at least one component or stream to generate the modified at least one bitstream may comprise generating an encoded IVAS frame comprising at least one silence descriptor frame based on a signal energy of the gain controlled at least one component or stream being below the IVAS encoder signal level threshold. Encoding at least the gain control processed at least one component or stream to generate a modified at least one bitstream may comprise adaptively encoding the at least one component or stream and / or the at least one other component or stream based on at least one of: a signal level of the gain controlled at least one component or stream; a relative signal level of the gain controlled at least one component or stream and at least one other component or stream. The mute request may comprise metadata comprising an identifier field for indicating the identified at least one component or stream to be muted. The metadata may further comprise a validity field configured to identify a time over which the muting is to be applied. Transmitting, to the further apparatus, the modified at least one bitstream may comprise transmitting, to the further apparatus, metadata for identifying a status of the identified at least one component or stream. The metadata for identifying a status of the identified at least one component or stream may be configured to indicate the at least one component or at least one stream is muted. The metadata may be a processing information frame. Receiving from the further apparatus the mute request indicating the identified at least one component or stream to be muted may comprise at least one of: receiving the mute request within an IVAS RTP payload header; receiving the mute request within a RTP header extension; receiving the mute request within a RTCP payload; receiving the mute request using an IVAS E-byte; and receiving the mute request using IVAS PI data. The method may further comprise obtaining, from the further apparatus, an unmute request indicating the identified at least one component or stream to be unmuted. The method may further comprise de-applying the muting to the indicated at least one component or stream identified by the unmute request. Receiving from the further apparatus the unmute request indicating the identified at least one component or stream to be unmuted may comprise at least one of: receiving the unmute request within an IVAS RTP payload header; receiving the unmute request within a RTP header extension; receiving the unmute request within a RTCP payload; receiving the unmute request using an IVAS E-byte; and receive the unmute request using IVAS PI data. The method may further comprise negotiating with the further apparatus support for handling the mute request indicating the identified the at least one component or stream to be muted. Negotiating with the further apparatus support for handling the mute request indicating the identified at least one component or stream to be muted may comprise negotiating employing a session description file. According to a fifth aspect there is provided an apparatus comprising at least one processor and at least one memory including a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to: obtain, from a further apparatus, at least one bitstream comprising at least one component or stream of an immersive audio bitstream; identify at least one of the at least one component or stream of the immersive audio bitstream to be muted; transmit, to a further apparatus, a mute request indicating the identified at least one component or stream to be muted; obtain, from the further apparatus, the at least one bitstream, wherein the at least one bitstream is modified with respect to the identified at least one component or stream of the at least one immersive audio bitstream to be muted, wherein the modification in the at least one bitstream, is at least one of: the at least one component or stream having reduced energy; the at least one component or stream not being received; the at least one component or stream comprises a silence descriptor indicator for indicating the at least one component or stream is to be muted; the at least one component or stream having a reduced bit rate; and at least one other component or stream having an increased bit rate. The apparatus caused to obtain, from the further apparatus, the bitstream, may be caused to depacketize and decode the at least one immersive audio bitstream from the bitstream. The depacketized and decoded at least one immersive audio bitstream may comprise at least one decoded IVAS frame, the at least one decoded IVAS frame comprising one of: at least one component having been the identified at least one component to be muted such that the at least one component has reduced energy; or at least one stream having been the identified at least one stream to be muted such that the at least one stream has reduced energy. The silence descriptor indicator for indicating the at least one component or stream may comprise at least one silence descriptor frame within the at least one component or stream. The apparatus may be further caused to render from the obtained immersive audio bitstream at least one audio signal for playback wherein the at least one component or stream of an immersive audio bitstream to be muted, wherein the apparatus caused to render is caused to generate the at least one audio signal for playback comprising at least one of: the identified at least one component or stream of the immersive audio bitstream with a significantly decreased signal energy compared to a signal energy of the identified at least one component or stream of the immersive audio bitstream prior to the mute request; a selectively muted at least one component or stream; and comfort noise to selectively replace the at least one component or stream. The mute request may comprise metadata comprising an identifier field for indicating the identified at least one component or stream to be muted. The metadata may further comprise a validity field configured to identify a time over which the muting is to be applied. The apparatus caused to obtain, from the further apparatus, the bitstream or modified bitstream may be caused to obtain, from the further apparatus, metadata for identifying a status of the at least one component or stream. The metadata for identifying a status of the at least one component or the stream may be configured to indicate which component or stream is muted. The metadata may be a processing information frame. The apparatus caused to transmit to the further apparatus the mute request indicating the identified at least one component or stream to be muted may be caused to, at least one of: transmit the mute request within an IVAS RTP payload header; transmit the mute request within a RTP header extension; transmit the mute request within a RTCP payload; transmit the mute request using an IVAS E-byte; and transmit the mute request using IVAS PI data. The apparatus may be caused to transmit, to the further apparatus, an unmute request indicating the identified at least one component or stream to be unmuted. The apparatus caused to render may be further cased to generate the at least one audio signal for playback comprising at least one of: the identified at least one component or stream of the immersive audio bitstream with a significantly increased signal energy following the unmute request compared to the signal energy of the identified at least one component or stream of the immersive audio bitstream following the mute request; a selectively unmuted at least one component or stream; and the at least one component or stream to selectively replace the comfort noise. The apparatus caused to transmit to the further apparatus the unmute request may be caused to at least one of: transmit the unmute request within an IVAS RTP payload header; transmit the unmute request within a RTP header extension; transmit the unmute request within a RTCP payload; transmit the unmute request using an IVAS E-byte; and transmit the mute request using IVAS PI data. The apparatus may be further caused to negotiate with the further apparatus support for handling the mute request indicating the identified at least one component or stream to be muted. The apparatus caused to negotiate with the further apparatus support for handling the mute request indicating the identified at least one component or stream to be muted may be caused to negotiate employing a session description file. The apparatus caused to identify at least one component or stream of the immersive audio bitstream is to be muted is caused to obtain at least one mute input for identifying the at least one component or stream of the immersive audio bitstream. The at least one immersive audio bitstream may comprise at least two components. According to a sixth aspect there is provided an apparatus comprising at least one processor and at least one memory including a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to: transmit to a further apparatus, at least one bitstream comprising at least one component or stream of an immersive audio bitstream; obtain, from a further apparatus, a mute request identifying at least one component or stream from the at least one component or stream to be muted; encode, based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream; transmit, to the further apparatus, the modified at least one bitstream, wherein the modified at least one bitstream comprising the identified at least one component or stream, wherein the modification in the at least one bitstream, is at least one of: the at least one component or stream having reduced energy; the at least one component or stream not being transmitted; the at least one component or stream comprises a silence descriptor indicator for indicating the at least one component or stream is to be muted; the at least one component or stream having a reduced bit rate; and at least one other component or stream having an increased bit rate. The apparatus caused to encode, based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream may be caused to packetize and encode the identified at least one component or stream. The apparatus caused to encode, based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream may be caused to employ an IVAS encoder to generate at least one encoded IVAS frame, the at least one encoded IVAS frame comprising one of: at least one component, identified by the mute request; or at least one stream, identified by the mute request. The apparatus caused to encode, based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream may be caused to: apply a gain control processing to the identified at least one component or stream; encode at least the gain control processed at least one component or stream to generate the modified at least one bitstream. The apparatus caused to encode at least the gain control processed at least one component or stream to generate the modified at least one bitstream may be caused to generate an encoded IVAS frame comprising at least one silence descriptor frame based on a signal energy of the gain controlled at least one component or stream being below the IVAS encoder signal level threshold. The apparatus caused to encode at least the gain control processed at least one component or stream to generate a modified at least one bitstream may be caused to adaptively encode the at least one component or stream and / or the at least one other component or stream based on at least one of: a signal level of the gain controlled at least one component or stream; a relative signal level of the gain controlled at least one component or stream and at least one other component or stream. The mute request may comprise metadata comprising an identifier field for indicating the identified at least one component or stream to be muted. The metadata may further comprise a validity field configured to identify a time over which the muting is to be applied. The apparatus caused to transmit, to the further apparatus, the modified at least one bitstream may be caused to transmit, to the further apparatus, metadata for identifying a status of the identified at least one component or stream. The metadata for identifying a status of the identified at least one component or stream may be configured to indicate the at least one component or at least one stream is muted. The metadata may be a processing information frame. The apparatus caused to receive from the further apparatus the mute request indicating the identified at least one component or stream to be muted may be caused to, receive at least one of: the mute request within an IVAS RTP payload header; the mute request within a RTP header extension; the mute request within a RTCP payload; receive the mute request using an IVAS E-byte; and receive the mute request using IVAS PI data. The apparatus may be further caused to obtain, from the further apparatus, an unmute request indicating the identified at least one component or stream to be unmuted. The apparatus may be further caused to de-apply the muting to the indicated at least one component or stream identified by the unmute request. The apparatus may be further caused to receive from the further apparatus the unmute request indicating the identified at least one component or stream to be unmuted may be caused to at least one of: receive the unmute request within an IVAS RTP payload header; receive the unmute request within a RTP header extension; receive the unmute request within a RTCP payload; receive the unmute request using an IVAS E-byte; and receive the unmute request using IVAS PI data. The apparatus may be further caused to negotiate with the further apparatus support for handling the mute request indicating the identified the at least one component or stream to be muted. The apparatus caused to negotiate with the further apparatus support for handling the mute request indicating the identified at least one component or stream to be muted may be caused to negotiate employing a session description file. According to a seventh aspect there is provided an apparatus comprising: obtaining circuitry configured to obtain, from a further apparatus, at least one bitstream comprising at least one component or stream of an immersive audio bitstream; identifying circuitry configured to identify at least one of the at least one component or stream of the immersive audio bitstream to be muted; transmitting circuitry configured to transmit, to a further apparatus, a mute request indicating the identified at least one component or stream to be muted; obtaining circuitry configured to obtain, from the further apparatus, the at least one bitstream, wherein the at least one bitstream is modified with respect to the identified at least one component or stream of the at least one immersive audio bitstream to be muted, wherein the modification in the at least one bitstream, is at least one of: the at least one component or stream having reduced energy; the at least one component or stream not being received; the at least one component or stream comprises a silence descriptor indicator for indicating the at least one component or stream is to be muted; the at least one component or stream having a reduced bit rate; and at least one other component or stream having an increased bit rate. According to an eighth aspect there is provided an apparatus comprising transmitting circuitry configured to transmit to a further apparatus, at least one bitstream comprising at least one component or stream of an immersive audio bitstream; obtaining circuitry configured to obtain, from a further apparatus, a mute request identifying at least one component or stream from the at least one component or stream to be muted; encoding circuitry configured to encode, based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream; transmitting circuitry configured to transmit, to the further apparatus, the modified at least one bitstream, wherein the modified at least one bitstream comprising the identified at least one component or stream, wherein the modification in the at least one bitstream, is at least one of: the at least one component or stream having reduced energy; the at least one component or stream not being transmitted; the at least one component or stream comprises a silence descriptor indicator for indicating the at least one component or stream is to be muted; the at least one component or stream having a reduced bit rate; and at least one other component or stream having an increased bit rate. According to a ninth aspect there is provided a computer program comprising instructions [or a computer readable medium comprising program instructions] for causing an apparatus to perform at least the following: obtain, from a further apparatus, at least one bitstream comprising at least one component or stream of an immersive audio bitstream; identify at least one of the at least one component or stream of the immersive audio bitstream to be muted; transmit, to a further apparatus, a mute request indicating the identified at least one component or stream to be muted; obtain, from the further apparatus, the at least one bitstream, wherein the at least one bitstream is modified with respect to the identified at least one component or stream of the at least one immersive audio bitstream to be muted, wherein the modification in the at least one bitstream, is at least one of: the at least one component or stream having reduced energy; the at least one component or stream not being received; the at least one component or stream comprises a silence descriptor indicator for indicating the at least one component or stream is to be muted; the at least one component or stream having a reduced bit rate; and at least one other component or stream having an increased bit rate. According to a tenth aspect there is provided a computer program comprising instructions or a computer readable medium comprising program instructions for causing an apparatus to perform at least the following: transmit to a further apparatus, at least one bitstream comprising at least one component or stream of an immersive audio bitstream; obtain, from a further apparatus, a mute request identifying at least one component or stream from the at least one component or stream to be muted; encode, based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream; transmit, to the further apparatus, the modified at least one bitstream, wherein the modified at least one bitstream comprising the identified at least one component or stream, wherein the modification in the at least one bitstream, is at least one of: the at least one component or stream having reduced energy; the at least one component or stream not being transmitted; the at least one component or stream comprises a silence descriptor indicator for indicating the at least one component or stream is to be muted; the at least one component or stream having a reduced bit rate; and at least one other component or stream having an increased bit rate. According to an eleventh aspect there is provided a non-transitory computer readable medium comprising program instructions for causing an apparatus to perform at least the following: obtain, from a further apparatus, at least one bitstream comprising at least one component or stream of an immersive audio bitstream; identify at least one of the at least one component or stream of the immersive audio bitstream to be muted; transmit, to a further apparatus, a mute request indicating the identified at least one component or stream to be muted; obtain, from the further apparatus, the at least one bitstream, wherein the at least one bitstream is modified with respect to the identified at least one component or stream of the at least one immersive audio bitstream to be muted, wherein the modification in the at least one bitstream, is at least one of: the at least one component or stream having reduced energy; the at least one component or stream not being received; the at least one component or stream comprises a silence descriptor indicator for indicating the at least one component or stream is to be muted; the at least one component or stream having a reduced bit rate; and at least one other component or stream having an increased bit rate. According to a twelfth aspect there is provided a non-transitory computer readable medium comprising program instructions for causing an apparatus to perform at least the following: transmit to a further apparatus, at least one bitstream comprising at least one component or stream of an immersive audio bitstream; obtain, from a further apparatus, a mute request identifying at least one component or stream from the at least one component or stream to be muted; encode, based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream; transmit, to the further apparatus, the modified at least one bitstream, wherein the modified at least one bitstream comprising the identified at least one component or stream, wherein the modification in the at least one bitstream, is at least one of: the at least one component or stream having reduced energy; the at least one component or stream not being transmitted; the at least one component or stream comprises a silence descriptor indicator for indicating the at least one component or stream is to be muted; the at least one component or stream having a reduced bit rate; and at least one other component or stream having an increased bit rate. According to a thirteenth aspect there is provided an apparatus comprising: means for obtaining, from a further apparatus, at least one bitstream comprising at least one component or stream of an immersive audio bitstream; means for identifying at least one of the at least one component or stream of the immersive audio bitstream to be muted; means for transmitting, to a further apparatus, a mute request indicating the identified at least one component or stream to be muted; means for obtaining, from the further apparatus, the at least one bitstream, wherein the at least one bitstream is modified with respect to the identified at least one component or stream of the at least one immersive audio bitstream to be muted, wherein the modification in the at least one bitstream, is at least one of: the at least one component or stream having reduced energy; the at least one component or stream not being received; the at least one component or stream comprises a silence descriptor indicator for indicating the at least one component or stream is to be muted; the at least one component or stream having a reduced bit rate; and at least one other component or stream having an increased bit rate. According to a fourteenth aspect there is provided an apparatus comprising: means for obtaining at least one immersive audio signal comprising at least two components or at least two streams; means for transmitting to a further apparatus, at least one bitstream comprising at least one component or stream of an immersive audio bitstream; means for obtaining, from a further apparatus, a mute request identifying at least one component or stream from the at least one component or stream to be muted; means for encoding, based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream; means for transmitting, to the further apparatus, the modified at least one bitstream, wherein the modified at least one bitstream comprising the identified at least one component or stream, wherein the modification in the at least one bitstream, is at least one of: the at least one component or stream having reduced energy; the at least one component or stream not being transmitted; the at least one component or stream comprises a silence descriptor indicator for indicating the at least one component or stream is to be muted; the at least one component or stream having a reduced bit rate; and at least one other component or stream having an increased bit rate. According to a fifteenth aspect there is provided a computer readable medium comprising program instructions for causing an apparatus to perform at least the following: obtain, from a further apparatus, at least one bitstream comprising at least one component or stream of an immersive audio bitstream; identify at least one of the at least one component or stream of the immersive audio bitstream to be muted; transmit, to a further apparatus, a mute request indicating the identified at least one component or stream to be muted; obtain, from the further apparatus, the at least one bitstream, wherein the at least one bitstream is modified with respect to the identified at least one component or stream of the at least one immersive audio bitstream to be muted, wherein the modification in the at least one bitstream, is at least one of: the at least one component or stream having reduced energy; the at least one component or stream not being received; the at least one component or stream comprises a silence descriptor indicator for indicating the at least one component or stream is to be muted; the at least one component or stream having a reduced bit rate; and at least one other component or stream having an increased bit rate. According to a sixteenth aspect there is provided a computer readable medium comprising program instructions for causing an apparatus to perform at least the following: transmit to a further apparatus, at least one bitstream comprising at least one component or stream of an immersive audio bitstream; obtain, from a further apparatus, a mute request identifying at least one component or stream from the at least one component or stream to be muted; encode, based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream; transmit, to the further apparatus, the modified at least one bitstream, wherein the modified at least one bitstream comprising the identified at least one component or stream, wherein the modification in the at least one bitstream, is at least one of: the at least one component or stream having reduced energy; the at least one component or stream not being transmitted; the at least one component or stream comprises a silence descriptor indicator for indicating the at least one component or stream is to be muted; the at least one component or stream having a reduced bit rate; and at least one other component or stream having an increased bit rate. An apparatus comprising means for performing the actions of the method as described above. An apparatus configured to perform the actions of the method as described above. A computer program comprising program instructions for causing a computer to perform the method as described above. A computer program product stored on a medium may cause an apparatus to perform the method as described herein. An electronic device may comprise apparatus as described herein. A chipset may comprise apparatus as described herein. Embodiments of the present application aim to address problems associated with the state of the art. Summary of the Figures For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings in which: Figure 1 shows schematically example server and peer-to-peer teleconferencing systems within which embodiments may be implemented; Figure 2 shows schematically an example Table of Content (ToC) byte structure for a Header-Full IVAS frame according to some embodiments; Figure 3 shows schematically an example IVAS Table of Content (ToC) byte structure according to some embodiments; Figure 4 shows schematically an example processing information (PI) frame in an IVAS RTP payload according to some embodiments; Figure 5 shows schematically an example IVAS extra (E) byte header structure according to some embodiments; Figure 6a shows schematically an example RTP packet structure with RTP Header with IVAS payload structure and frame data according to some embodiments; Figure 6b shows schematically a further example RTP packet structure with RTP Header with IVAS payload structure with separated IVAS payload header, IVAS frame data and PI data block sections according to some embodiments; Figure 6c shows schematically an example PI data block structure with PI header section and PI data frames according to some embodiments; Figure 6d shows schematically an example PI header byte structure according to some embodiments; Figure 7 shows schematically an example MUTE PI frame structure according to some embodiments; Figure 8 shows schematically an example implementation employing the MUTE request for a stream component according to some embodiments; Figure 9 shows schematically a further example implementation employing the MUTE request for a stream component according to some embodiments; Figure 10 shows schematically a multi user example implementation employing the MUTE request for a stream component according to some embodiments; Figures 11 to 13 show flow diagrams of an example operation of the MUTE request according to some embodiments; Figures 14 and 15 show flow diagrams of summary operations of the embodiments; and Figure 16 shows an example device suitable for implementing the apparatus shown. Embodiments of the Application The following describes in further detail suitable apparatus and possible mechanisms for the provision of efficient IVAS audio. An example system within which embodiments may be implemented is shown in Figure 1. Figure 1, for example, shows an example teleconferencing system within which some embodiments can be implemented. In this example there is shown two sites or rooms, Room A 100 and Room B 102. Room A 100 comprises a ‘talker’ or user, Talker 1 103. Room B 102 comprises one ‘talker’ or user, Talker RX 141. In the following example within room A is a suitable teleconference apparatus (or more generally telecommunications apparatus 110) configured to spatially capture and encode the audio environment and furthermore is configured to render a spatial audio signal to the room. The apparatus can in some embodiments be implemented by a user equipment (UE) operating within a cellular communications system or accessing any suitable access network. Within each of the other rooms may be a suitable teleconference apparatus (or more generally telecommunications apparatus such as apparatus 120 within room B) configured to render a spatial audio signal to the room and furthermore is configured to capture and encode at least a mono audio and optionally configured to spatially capture and encode the audio environment. In the following examples each room is provided with the means to spatially capture, encode spatial audio signals, receive spatial audio signals and render these to a suitable listener. It would be understood that there may be other embodiments where the system comprises some apparatus configured to only capture and encode audio signals (in other words the apparatus is a ‘transmit’ only apparatus), and other apparatus configured to only receive and render audio signals (in other words the apparatus is a ‘receive’ only apparatus). In such embodiments the system within which embodiments may be implemented may comprise apparatus with varying abilities to capture / render audio signals. The teleconference apparatus (for each site or room) 110, 120 can be configured to call into a teleconference controlled by and implemented over a server 111. In some embodiments the communications or teleconferencing system comprises a (peer-to-peer) communications system (rather than the server based system shown in Figure 1) within which some embodiments can be implemented. Thus, for example, two or more UEs can be configured to interact directly with each other (for example to implement an immersive audio phone call between users). In such a scenario one of the UEs can be configured to deliver spatial ambience as one stream and employ a close-up microphone (for example a Lavalier microphone) to capture the speech as an audio object or audio source. The sender UE can be configured to encode the spatial ambience audio signals in a MASA format stream and the close-up microphone audio signal as an object format stream. The two audio streams can then be delivered as separated IVAS streams. The sender UE, in addition, can be configured to encode processing information during the encoding to deliver the PI frames together with the IVAS frames to the receiver UE. The teleconference apparatus can be configured to spatially capture and encode the audio environment and furthermore can be configured to render a spatial audio signal to the room. In this example only the communications or signalling path from the Room A 100 to the Room B 102 is shown for simplicity but a duplex or multipoint communication system comprising multiple signalling paths can be implemented using the methods as described herein without significant inventive input. The teleconference apparatus (for each site or room) 110, 120 is further configured to communicate with each other to implement a teleconference function. As shown in Figure 1, the apparatus 110, 120 and server 111 can comprise suitable encoder and decoder functionality. For example the apparatus 110 is shown comprising an (IVAS) encoder / packetizer function 105 which can be divided into IVAS stream encoder (encoder 1) 1011 and IVAS stream encoder (encoder 2) 1012, the server 111 is shown comprising a (IVAS) decoder and encoder functions 121 and the apparatus 120 is shown comprising an (IVAS) depacketizer / decoder function 141 which can be divided into IVAS stream decoder (decoder 1) 1311 and IVAS stream decoder (decoder 2) 1312. In such a manner a first audio stream 104i (the audio signals representing the user or talker 1 103) can be encoded by the encoder 1 1011 and a second audio stream 1042 (for example representing a Finnish translation of the user or talker 1 103) can be encoded by the encoder 2 1012 which generates a single RTP payload 106 that comprises two IVAS bitstreams from the encoder 1 1011 and encoder 2 1012 to be passed to a server 111. The server 111 can then decode, (optionally then mix with other objects and otherwise process the audio signals) and encode then to generate the single RTP payload 108 with multiple IVAS bitstreams to be passed to the apparatus 120. In some embodiments the server 111 decoder and encoder 121 comprises multiple stream encoder and decoder instances. The apparatus 120 can then decode the audio signals and present them to the user or talker ‘Talker RX’ 141. The apparatus 120 thus can comprise a decoder function 131 configured to receive the RTP payload 108. In some embodiments the decoder function 131 comprises a first decoder (decoder 1) 1311 configured to decode the first stream and a second decoder (decoder 2) 1312 configured to decode the second stream. Although this example shows a teleconference application the encoder / decoder functionality can be applied to the streaming of any suitable media. The IVAS decoder / renderer for each of the teleconference apparatus 102 can be furthermore configured to handle multiple input streams that may each originate from a different encoder. Similarly although in the example shown in Figure 1 has a server 111 between the apparatus 110, 120 in some embodiments the communication can be direct without any intermediate (IVAS) decoder / encoder As discussed previously RTP is intended for an end-to-end, real-time transfer of streaming media and provides facilities for jitter compensation and detection of packet loss and out-of-order delivery. RTP is furthermore designed to carry a multitude of multimedia formats, which permit the transport of new formats without revising the RTP standard. To this end, the information required by a specific application of the protocol is not included in the generic RTP header. For a class of applications (e.g., audio, video), an RTP profile may be defined. For a media format (e.g., a specific video coding format), an associated RTP payload format may be defined. Every instantiation of RTP in a particular application may therefore require a profile and payload format specifications. The profile is configured to define the codec used to encode the payload data and the mapping to payload format codes in the protocol field Payload Type (PT) of the RTP header. For example, the RTP profile for audio and video conferences with minimal control is defined in RFC 3551. The profile defines a set of static payload type assignments, and a dynamic mechanism for mapping between a payload format, and a PT value using Session Description Protocol (SDP). The latter mechanism is used for newer video codec such as RTP payload format for H.264 Video defined in RFC 6184 or RTP Payload Format for High Efficiency Video Coding (HEVC) defined in RFC 7798. An RTP session can be established for each multimedia stream. Audio and video streams may be implemented which use separate RTP sessions, enabling a receiver to selectively receive components of a particular stream. The RTP specification can furthermore be configured to recommend port numbers for RTP, and furthermore to recommend the use of the next odd port number for the associated RTCP session. A single port can be used for RTP and RTCP in applications that multiplex the protocols. Each RTP stream can comprise RTP packets, and the RTP packet in turn can comprise a RTP header and payload pair. Enhanced Voice Services (EVS) is a mono voice codec standardized in 3GPP and described in the TS 26.445 specification document. The codec can have two operating modes: EVS Primary and EVS AMR-WB IO (Adaptive Multi Rate Wideband Inter-Operable). The IVAS codec can be considered to be an extension to the EVS codec and as such the IVAS and EVS codecs can have some similarities in terms of design and implementation. The RTP payload structure, while not having been specified for the IVAS codec, is envisioned to have similarities to the RTP payload structure in EVS. The RTP payload format of EVS is described in 3GPP TS 26.445 Annex A. In EVS, the RTP payload format is divided into two different embodiments: a Compact format and a Header-Full format. In the EVS Compact payload format, a RTP packet includes only a single EVS speech frame for EVS Primary mode. For EVS AMR-WB IO mode, the compact RTP packet also includes a 3-bit Codec Mode Request (CMR) field in front of the speech frame. In the EVS Compact format, the different modes and bitrates for the speech frames are identified by the size of the RTP payload. For example, an RTP packet of size 328 bits is assigned for EVS Primary mode with 16.4 kbps bitrate, as is shown in Table A.1 in TS 26.445 Annex A. In EVS Header-Full format, the RTP payload consists of the speech frame(s) accompanied by an optional CMR byte and Table of Content (ToC) bytes. The CMR byte is used to request a change in bitrate or coding mode that the receiver wants to receive. The request is sent as a CMR byte as part of the Header-Full EVS packet. In EVS AMR-WB IO Compact format, the CMR functionality is also present as a 3-bit signalling at the beginning of the packet. One aspect or technique of audio encoding and employed in various speech processing algorithms is Voice Activity Detection (VAD), also known as speech activity detection or more generally as signal activity detection. For example VAD can be employed in speech codecs, for detecting the presence or absence of human speech. It can be generalized to detection of active signal, i.e., a sound source other than background noise. Based on a VAD decision, it is possible to utilize, e.g., a certain encoding mode in a speech encoder, e.g., active signal encoding instead of background noise encoding. Discontinuous Transmission (DTX) is a technique utilizing VAD intended to temporarily shut off parts of active signal processing (such as speech coding according to certain modes) and the frame-by-frame transmission of encoded audio (e.g., transmission of RTP packets). Instead of normal encoded frames, e.g., it can be sent simplified update frames to drive a CNG at the decoder. Typically, this is done at slower update rate, e.g., only every / Vth frame is updated / transmitted. Otherwise, no data is transmitted. The use of DTX can help with reducing interference and / or preserving / reallocating capacity in a mobile network. It can also help with battery life of the device. Consequently, the transmission of simplified update frames (to drive a CNG at the decoder) can be used to detect DTX operation. Comfort Noise Generation (CNG) is a technique for creating a synthetic background noise to fill silence periods that would otherwise be observed, e.g., under the DTX operation. A complete silence can be confusing or annoying to a receiving user. For example, the listener could judge that the transmission may have been lost and then unnecessarily say “hello, are you still there?” to confirm or simply hang up. On the other hand, sudden changes in sound level (from total silence to active background and speech or vice versa) could also be very annoying. Thus, CNG is applied. Typically, this is based on a highly simplified transmission of noise parameters (e.g., spectral shape) derived from the real captured background noise. Silence Descriptor (SID) frames are sent during speech inactivity to keep the receiver CNG sufficiently aligned with the background noise level and spectral shape at the sender side. This is of particular importance at the onset of each new talk spurt. Thus, SID frames should not be too old, when a new talks burst starts. Commonly SID frames are sent regularly, e.g., every 8th frame, but a codec may allow also variable rate SID updates. SID frames are typically quite small, e.g., 2.4 kbps SID bitrate equals 48 bits per frame for the typical 20-ms frame size used in modern communications codecs (AMR, AMR-WB, ITU-T G.718, EVS, and IVAS). IVAS SID frame size is 5.2 kbps (104 bits) for a 20-ms frame, and the default SID frame interval in IVAS is 8 frames. Other update intervals are furthermore supported. IVAS includes EVS, and in EVS primary mode operation the 48-bit SID is kept. In an IVAS session, there can be one or more streams (e.g., MASA, ISM instances coded separately) or stream components (e.g., each individual object in ISM, OMASA, and OSBA operation) arriving at a receiver. It can be beneficial for the listener experience to be able to mute one or more such stream or stream component for playback. One example is the streaming of language tracks, where user is likely to listen to a single language at a time. Muting can then be applied to other languages than the one the user selects. Thus, a simpler or more focused playback experience is achieved with no overtalk. When muting is applied at the receiver only, there is no implication or effect for the transmission. In other words, even though muting is applied, communication resources are still used (spent) transmitting all streams (stream components) and a limited transmission channel is allocated for signal content that is not used by the receiver. This is inefficient and also decreases overall perceptual quality if adaptive bit budget allocation is not properly used. PI data frames, as described in UK application GB2317940.1, enable a request control by sending PI frames in the backwards (or reverse) direction from receiver to sender. A MUTE request can, for example, be sent this way. This provides the mechanism to signal a change in playback status of a stream (stream component). However, there is currently lacking an efficient mechanism for implementing this within an IVAS codec. Consequently, the concept as discussed in detail with respect to the following embodiments are apparatus and methods configured to handle MUTE requests (from the receiver UE) for audio inputs, IVAS processing, and transmission. Additionally as discussed in some embodiments there is discussed specific mechanisms with respect to rendering based on the MUTE requests. In some embodiments there can be provided apparatus and methods for requesting a muting of individual audio streams or stream components in an immersive audio session. In these embodiments the apparatus or method can be configured to control (or indirectly affect) how an audio input is processed at an audio encoder in order to enable muting of audio in the playback for a listener while simultaneously reducing the bitrate for encoded audio bitstream transmission. For example in some embodiments there can be a session between a sender apparatus UE1 110 and a receiver apparatus UE2 120. Furthermore the sender apparatus can be configured to send an OMASA stream (with MASA and audio object ISM components). In such an example the muting of a single audio object in OMASA format can be achieved by the following: Sender apparatus UE1 110: Receive a mute request from a receiver apparatus UE2 120; The mute request is configured to indicate one object 11 stream in OMASA format to mute; Apply muting to 11 stream (for example via gain control); The IVAS encoder allocates bitrate differently as the muted audio input 11 is not considered important due to its energy; Packetize the encoded bitstream into an IVAS RTP packet and transmit to receiver apparatus UE2. Receiver apparatus UE2: Receive IVAS RTP packet from sender apparatus UE1; De-packetize the packet and decode the IVAS bitstream; Provide the decoded IVAS frames to a renderer; Provide the rendered audio for playback. The muted audio input 11 is not heard by the listener, the overall quality for other audio inputs can be improved due to different bit rate allocation In a further example in a session between a sender apparatus UE1 and a receiver apparatus UE2, the sender apparatus is configured to send individual MASA and audio object ISM 11 streams that are encoded by separate IVAS encoders. In such an example the muting of a single audio object 11 stream can be achieved by the following: Sender apparatus UE1: Receive a mute request from a receiver apparatus UE2; The mute request indicates to mute audio object stream 11; Apply muting to 11 stream via gain control; The IVAS encoder for 11 enters DTX mode; The IVAS encoder for MASA remains un-modified; Packetize the encoded bitstreams from the separate encoders into IVAS RTP packet(s) and transmit to receiver UE2. Receiver apparatus UE2: Receive IVAS RTP packet(s) from sender apparatus UE1; De-packetize the packet(s) and decode the IVAS bitstreams for MASA and ISM 11; Provide the decoded IVAS frames to a renderer; Provide the rendered audio for playback; The muted audio input 11 is not heard by the listener The Header-Full IVAS frames (or IVAS frame as RTP payload with a payload header described as ToC) comprise of one or several Table-of-Content (ToC) indicators and their associated IVAS speech / audio frames. In some embodiments the bitrate could be moved from the encoder for the 11 stream to the MASA encoder. This is possible as the UE1 knows that the 11 stream is muted. With respect to Figure 2 is shown an example Table-of-Content (ToC) byte structure 201 for an IVAS frame. The (F) bit 203 indicates whether more frames follow this entry. The F (1 bit) 203: If set to 1, the bit indicates that the corresponding frame is followed by another speech or PI frame in this payload, implying that another ToC byte follows this entry. If set to 0, the bit indicates that this frame is the last frame in this payload and no further header entry follows this entry. The Bitrate Level Indication (BLI) 207 describes the content of the associated frame, e.g., the bitrate of an IVAS frame. The Supplemental Signaling Bits (SSB) 205 can be used to describe various things, for example IVAS input format specific signaling. The SSB 205 can in some embodiments be referred to as extra bits. The BLI (4-5 bits) 207 can, for example, be configured to indicate the bitrate or other frame content indication for the frame. From the content indication, the receiver can determine the size of the received frame either directly from the bitrate or from pre-determined frame sizes, e.g., for the SPEECH_LOST and NO_DATA frames. The BLI (or in some implementations the Frame Type or FT) bits can indicate, for example, the bitrate of an IVAS speech frame, a SPEECH_LOST, NO_DATA or comfort noise (SID) frames. An example for BLI bit values (for 5 bits) are presented in the following table. At this point of the codec development, it is not decided if the aforementioned three types (SPEECH_LOST, NO_DATA, SID) are supported in the final codec. The SPEECH_LOST, NO_DATA and SID frames are part of the EVS codec and at it is 5 likely that at least the SID and NO_DATA frame types are being incorporated into the IVAS specification. It is understood that the final IVAS codec specification, however, might have different frame types present than presented here. BLI bits Bitrate (kbps) or other frame content indication 00000 13.2 10 00001 16.4 00010 24.4 00011 32 00100 48 00101 64 00110 80 00111 96 01000 128 01001 160 01010 192 01011 256 01100 384 01101 512 01110 SPEECH_LOST 01111 N0_DATA 10000 SID (5.2) 10001 -11111 (reserved for future use) In some embodiments, 4 bits are reserved for the BLI part. The frame type bit values and content indications when using 4 bits are presented below. In these embodiments the BLI part identifies the bitrate where the BLI bits are bit values other than a defined value (for example 1111) and can be used to identify other 15 aspects, for example SPEECH_LOST, NO_DATA and SID frames with the combination of the BLI indicator OTHER (the defined value) and the extra bits. BLI bits Bitrate (kbps) or other frame content indication 0000 13.2 0001 16.4 0010 24.4 0011 32 0100 48 0101 64 0110 80 0111 96 1000 128 1001 160 1010 192 1011 256 1100 384 1101 512 1110 Reserved for future use 1111 OTHER Extra + BLI bits Frame content indication 000 + 1111 SPEECH_LOST 001+1111 NO_DATA 010 + 1111 SID (2.4) 011 + 1111 Reserved for future use 100 + 1111 Reserved for future use 101 +1111 Reserved for future use 110 + 1111 Reserved for future use 111 + 1111 Reserved for future use As shown above there are 15 available bit allocations for the BLI bits reserved for future use (bit allocations 10001 - 11111), when 5 bits are reserved 5 for the BLI indicator. In some embodiments when 4 bits are reserved for the BLI indicator, there are 5 available bit allocations for future use, when BLI value 1111 is used. BLI value 1110 provides an additional 8 bit allocations for future use when combined with the extra bits. These available bit allocations can be used in the future to indicate frame contents that will be defined, such as PI frames described hereafter. These bit allocation values are examples only and it would be appreciated that in some embodiments the bit allocation values can be otherwise configured. The number of bits in the SSB 205 can vary between 2-3 depending on how many bits are reserved for the BLI bits 207 in the IVAS ToC byte 201. The SSB (2-3 bits) 205 can be bits reserved for future use. If 4 bits are reserved for the BLI indicator, the extra 3 bits can be partly used to identify other frame content than bitrates / frame-size (e.g., SPEECH_LOST, NO_DATA, SID) as demonstrated in further detail in GB2313472.9, where the BLI bits are referred to as FT bits (frame type index) and SSB as extra bits. Annex A of 3GPP TS 26.253, v2.0.0 describes the most up-to-date design for the IVAS ToC byte. Furthermore, “Corrections to TS 26.253 Annex A” CR S4-240664 describes further updates to the IVAS RTP payload format. Figure 3 shows an example IVAS ToC byte structure according to the current design where the ToC byte from EVS is extended to cover IVAS bitrates. In summary, the ToC byte consists of an H-bit 301 (1 bit, set to 0 for ToC byte) which separates the ToC byte from the E-byte (which is further described below), a F-bit (1 bit) 303 similar to the F-bit mentioned above, and ToC data 305 comprising 2-bit mode field 307 differentiating between IVAS, EVS and AMR-WB IO modes and 4-bit field 309 indicating the bitrate. In order to enable rendering or consumption of the IVAS audio bitstream data, additional (non-audio) data can be added to streamed IVAS RTP packets. Specifically, there can be inserted data that requires maintaining a sufficient or exact alignment with the IVAS audio bitstream data. The alignment can be considered, for example, relative to an IVAS audio frame. For example, any external orientations from the sender UE side could be included in these (audio) processing information (PI) frames. A RTP payload with respect to these embodiments can comprise one or more of the following: immersive audio coded bitstream; the immersive audio coded bitstream payload header. The immersive audio coded bitstream payload header can be a ToC header byte, which can comprise supplemental signaling bits or frame type indication, bitrate level indication bits, etc; processing information header; and processing information frame. Packetization furthermore with respect to these embodiments can be a process of including the immersive audio coded bitstream header, and at least one of the immersive audio coded bitstream or processing information data block as RTP payload in order to deliver the immersive audio coded bitstream over real-time transport protocol. In some embodiments, the packetization can also be performed for inclusion and carriage of control information as payload of the real-time transport control protocol. An example PI frame structure is shown with respect to Figure 4. The example PI frame structure 401 comprises a “format” field 403. The format field 403 is configured to describe the type of the PI data, for example, orientation data. The example PI frame structure 401 can further comprise a “usage” field 405. The usage field 405 is configured to describe how the PI data should be used. For example, the usage field 405 is configured to define whether the data should be applied to the next IVAS frame, to all frames in the RTP packet or if the data is more general information to be sent to the receiver. In other words, the usage field can describe that the PI data should not be taken account in the rendering, i.e., that the data is more general information sent to the receiver. Additionally the example PI frame structure 401 can further comprise a “validity” field 407. The validity field 407 in some embodiments describes how long the PI data is valid at the receiver end. For example, the field could describe that for X amount of processing frames the receiver should apply the PI data described in the frame, and if no new PI data is received within X frames, stop applying the data. The validity could also be indicated in other time formats than number of processing frames, e.g., in milliseconds. Furthermore, some audio frame values may be “hold” whereas other values may be “instantaneous” only. This enables the renderer to perform rendering accordingly. Furthermore, the example PI frame structure 401 can comprise an optional “size” field 409. The size field 409 in some embodiments is configured to define the size of the PI data 411. In some embodiments to minimize the number of additional bits introduced by the PI frames, the maximum size of the PI frames can be restricted. For example, the frames can be restricted to have a maximum size of K bits that is common for all IVAS bitrates. Alternatively, the maximum size of the PI frames could be linked to the used bitrate of the IVAS speech frames, e.g., by having the maximum number of bits for PI frames be some percentage share of the bits used for the speech frames. In addition, the use of PI frames could be restricted to be used only in the highest IVAS bitrates to reduce the risk of adding too much load to the sent packets. In some embodiments the “usage” and “validity” fields could be also combined into a single field, for example, to a “scope” field. The “scope” field would indicate both how the data should be applied and how long the data is valid. In another embodiment, the PI data carried is metadata carried in addition to the audio data (e.g., IVAS speech or audio frames and SID frames) as part of the RTP stream via indication that this data supplements the IVAS frames. The PI data may also carry indication of whether it is applied in the forward direction (e.g., receiver of the RTP stream utilizes the PI data for rendering) or the reverse direction (e.g., receiver of the RTP stream utilizes the PI data for encoding). The processing information can also be referred to as rendering metadata or PI data or extension metadata or any suitable term which is addition to the IVAS audio data. Additionally, the PI frame structure comprises the PI data 411. Figure 5 shows an example Extra byte (E-byte) structure as defined by the IVAS RTP payload design (described in Annex A of TS 26.253, v2.0.0). The payload header bytes in an IVAS session are divided into two categories based on the first (H) bit 501. H=0 indicates that the header byte is a ToC byte describing the bitrate for an associated audio frame. H=1 indicates that the header byte is an extra information (E) byte which contains or identifies additional information related to the session (e.g., CMR or PI frames). Furthermore the extra byte structure comprises the E-data 505 or extra data or extra information related to the session. Figure 6a shows an example full structure for an IVAS RTP packet 600. An RTP header (with possible RTP header extension) 601 precedes the IVAS payload 602. The IVAS payload 602 consists of the IVAS payload header 603 and data frames (PI or IVAS) 605. The IVAS payload header 603 identifies the different data frames (IVAS or PI) in the payload (through ToC or E-bytes) and contains optional information, e.g., CMR through E-bytes. Figure 6b shows another example full structure for an IVAS RTP packet 611. In this approach and compared to Figure 6a, the PI data is separated into a separate data section at the end of the payload 612, which comprises the IVAS payload header 615, the IVAS frame data 617 and the PI data block 619. The presence of the PI data block 619 may be indicated explicitly in the IVAS payload header 615 (through, e.g., E-bytes) or the presence of the PI data block 619 can be determined implicitly by detecting non-zero data after the IVAS frame data section 617. Figure 6c shows an example structure for a PI data block 619 presented in Figure 6b. The PI data block 619 comprises a PI header section 621 and PI data frames section 623. The PI header section 621 can identify the different PI data frames in the payload (e.g., through PI header bytes) and also contains information to which audio frame(s) the PI data frames 623 are associated with. Figure 6d shows an example PI header section 621 byte structure. The PI header section 621 byte structure comprises a F-bit 631 which indicates whether another PI header byte follows this byte or not. In some embodiments a bit value of 1 indicate a following byte and a bit value of 0 indicates no further bytes. The example PI header section 621 byte structure further comprises M-bits 633 (marker or marking or association bits) which identify to which audio frame(s) the PI data frames are associated with. The example PI header section 621 byte structure further comprises PI type 635 bits indicate the type for the associated PI frame. The M-bits 633 can have values of, for example: - 00 for reserved value; - 01 for indicating that the associated PI frame is not the last PI frame for the current audio frame in sequence; - 10 for indicating that the associated PI frame is the last PI frame for the current audio frame in sequence (and that the counter for the current audio frame should be increased); and - 11 for indicating that the associated PI frame data should be applied for all the audio frames in the payload. In normal operation mode in an IVAS session, the PI frames are tied / linked / associated to the timestamp of the associated audio frames (IVAS or EVS in mono mode). That is, a PI frame does not increase the media time of the packet. The RTP timestamp defines the sampling instant (media time) of the first sample of the first audio frame in an RTP packet. In DTX (discontinuous transmission) operation mode, SID frames can be transmitted at a given SID update interval (for example, 8 frames interval or about six SID frames per second). The media timeline progresses with the SID frames, however, the update rate is lower, because typically the SID frames are transmitted at a lower frame rate to reduce bitrate consumption. In some situations, relevant PI data frames could still be available for transmission. For example, the updated orientation values of the capturing device could still be available for transmission. These values can then be received at suitable intervals from a tracker that does not depend on the input audio in anyway. As mentioned above, the PI frames are tied / linked / associated to the timestamp of associated audio frames, however in DTX, these audio frames are not available, rather only lower frame rate SID frames may be available. In some embodiments, the receiver apparatus (for example UE2) can request to mute some incoming streams. For example, in the translation example above, the receiver UE2 can request the sender UE1 to mute all language streams except the one the receiver is listening to. For example, there can be the translation scenario where the sender apparatus or transmitter apparatus (UE1) is transmitting multiple translations (e.g., English, Finnish, Swedish, etc.) of speech in separate streams. The receiver apparatus (UE2) can then be configured to choose only one language to listen to and prioritize that stream. For example, UE2 can choose Finnish as the preferred stream language and request a higher bitrate for that specific stream and lower bitrate for the other language streams. For example, in the translation example above, the receiver UE2 can request the sender UE1 to mute all language streams except the one the receiver is listening to. The muted streams would then be indicated as N0_DATA or with similar notation or they could be left out from the transmitted payload. In some embodiments, the mute request could be a PI frame type, where the PI frame data indicates the stream(s) to be muted, as presented in Figure 7. The Mute PI frame is shown in Figure 7, for example, comprises a MUTE type 701, a “usage” field 703, a “validity” field 705 and “size” field 707 as described above together with an affected_stream 709 indicator. In some embodiments, the affected_stream indicator could be replaced with, e.g., affected_input_formats indicator to target the mute request to specific IVAS input formats. With respect to Figure 8 is shown an example system of apparatus implementing some embodiments. For example in Figure 8 is shown the UE 1 110 comprising the IVAS encoder 807 and packetizer 809. The IVAS encoder 807 is configured to receive the OMASA encoder input 805 which comprises a MASA stream component 801 and an object ISM 11 stream component 803. The IVAS encoder 807 is configured to generate an IVAS output which is passed to the packetizer 809. The packetizer 809 is then configured to generate RTP packets which are then transmitted to the UE2 120. The de-packetizer 811 of UE2 120 is configured to depacketize the received RTP packets and pass IVAS encoded signals to the IVAS decoder 813. The IVAS decoder 813 is configured to receive the IVAS encoded signals and decode these to generate an audio output 815 comprising the MASA stream component 817 and ISM 11 stream component 819. Additionally in Figure 8 is shown the operations of the implementation of the muting of the object 11 stream component 819. Thus for example the user is shown by 851 in Figure 8 as wishing to mute the 11 stream component 819. This can be expressed by any suitable user input, over any suitable user interface, for example touch, keyboard, mouse, voice or otherwise. Then as shown, by 852 in Figure 8, is the generating and sending a MUTE request, such as shown in Figure 7 and described above wherein the indicator can identify the stream component 11 to be muted. The MUTE request is received by the UE1, and the UE1 110 is configured to apply the muting by applying a gain control for 11. In other words, generate or provide a silent object, gain-adjusted ISM 11 stream component 803’ in the OMASA input 805’. Thus the IVAS encoder 807 is then configured to implement an encoding of the OMASA encoder input 805’ comprising the MASA stream component 801 and the gain-adjusted ISM 11 stream component 803’. In this scenario, since the ISM1 stream component 803’ is muted, the encoder can be configured to allocate more bitrate budget for MASA encoding. In such a manner the IVAS encoder is configured to implement an improved quality encoding of the unaffected part or stream components, for example the non-silent MASA stream component 801 in the OMASA encoder input 805’ as shown by 854 in Figure 8. This encoded IVAS signal is packetized, transmitted, received at the UE2 120, where it is depacketized and decoded to generate an audio output 815’ comprising the MASA stream component 817’ and 11 muted stream component 819’ such that the muted part 11 (at least substantially muted) is not heard by the user in the audio output as shown by 855 in Figure 8. Where Figure 8 shows an embodiment wherein a stream component is muted, Figure 9 shows how a similar approach can be applied to separate streams. For example in Figure 9 is shown the UE 1 110 comprising IVAS encoders 907a and 907b for encoding separate streams and packetizer 809. The IVAS encoders 907a and 907b are respectively configured to receive a MASA stream 901 and an object ISM 11 stream 903 and generate IVAS stream outputs which are passed to the packetizer 809. The packetizer 809 is then configured to generate RTP packets which are then transmitted to the UE2 120. The de-packetizer 811 of UE2 120 is configured to depacketize the received RTP packets and pass IVAS encoded signals for each stream to IVAS decoders 913a and 913b. The IVAS decoders 913a and 913b are configured to receive the IVAS encoded signal streams and decode these to generate an audio output comprising the MASA stream 917 and ISM 11 stream 919. In this example although there are shown separate encoders and decoders for each stream it would be appreciated that a single encoder and decoder implementation configured to handle multiple IVAS streams could also be implemented. Additionally in Figure 9 is shown the operations of the implementation of the muting of the object 11 stream 919. Thus, for example the user is shown by 951 in Figure 9 as wishing to mute the 11 stream 919. This can be expressed by any suitable user input, over any suitable user interface, for example touch, keyboard, mouse, voice or otherwise. Then as shown, by 952 in Figure 9, is the generating and sending a MUTE request, such as shown in Figure 7 and described above wherein the indicator can identify the stream component 11 to be muted. The MUTE request is received by the UE1, and the UE1 110 is configured to apply the muting by applying a gain control for the 11 stream. In other words, generate or provide a silent object, gain-adjusted ISM 11 stream 903’ to the IVAS encoder 907b. Thus, the IVAS encoder 907b can be configured to implement an encoding of the gain-adjusted ISM 11 stream component 903’, for example signalling a SID frame. In such a manner the IVAS encoders implement a (more efficient) encoding of the streams. Furthermore, in some embodiments where bitrate is distributed between encoders then the encoder configured to encode the non-silent MASA stream 901 with additional bits. This encoded IVAS signals are packetized, transmitted, received at the UE2 120, where it is depacketized and decoded to generate an audio output comprising the MASA stream 917’ and muted 11 stream (signalled for example by SID frames) 919’ such that the muted part 11 (at least substantially muted) is not heard by the user in the audio output as shown by 955 in Figure 9. With respect to Figure 10 is shown a further example, where UE1 110, UE2 120 and UE3 1010 are operating in a peer-to-peer telecommunications network. In this example, the user UE2 120 is shown by 1001 in Figure 10 as wishing to mute the incoming 11 stream from UE3 1010. This can be expressed by any suitable user input, over any suitable user interface, for example touch, keyboard, mouse, voice or otherwise. The MUTE request is thus sent, as shown by 1002 in Figure 10 from receiving UE2 120 to UE3 1010 to mute one of two audio inputs that are encoded separately using two IVAS encoder instances on two separate UEs. The MUTE request is then handled by UE 3 1010 and the user of UE2 120 does not (at least substantially) hear the muted stream. In some embodiments, as described above the muting can be applied by at least modifying the corresponding audio input gain prior to ingest by IVAS encoder. For example, when a zero-gain is applied, an audio input with silence (e.g., digital zeros) is obtained. This modified audio input can then be input to the IVAS encoder. The IVAS encoder can then be configured to encode the audio input, the encoded bitstream and any auxiliary data are packed into RTP packets, and the packets transmitted over the radio interface. This achieves the muting. However, if the IVAS encoder does not have DTX operation activated (i.e., ‘DTX off’) or DTX capability, there is no efficiency gain in terms of transmission. It is noted that IVAS multi-channel, OMASA, and OSBA operation do not currently support DTX operation. It is furthermore noted that the metadata that may be associated with the audio (waveform) are not affected. In various examples, the gain that is applied can be, for example, a small gain that can be zero or non-zero. In the following, a zero-gain and zero-signal are used for simplicity of description. Considering the example as shown in Figure 8, an example flow diagram is shown in Figure 11 which aims to achieve stream component muting: Thus, as shown in Figure 11 by 1101 is the operation of obtaining the mute input (at the receiving UE2) and generating the MUTE request. The MUTE request is then sent from the receiving UE2 to the sending side (UE1 or server) as shown by 1103 in Figure 11. Following reception of the MUTE request, the application of the gain control for audio indicated via the MUTE request is shown by 1105 in Figure 11. Then is the provision of the audio input modification by the gain control to the encoder as shown by 1107 in Figure 11. There is then an operation of encoding and transmitting at least the audio input modified by gain control as shown by 1109 in Figure 11. This can result in the operation of receiving and decoding at least the encoded audio modified by gain control as shown by 1111 in Figure 11. In such a manner the audio input for which a muting is requested is seen as a zero-signal by the IVAS encoder. The IVAS encoder can then, depending on the operation (e.g., ISM, OMASA, OSBA), can allocate the available bitrate differently as the muted audio input is not considered important due to its energy. Then the IVAS frames can be transmitted normally. The IVAS decoder / renderer receives the IVAS frames and provides the rendered audio for playback. The muted audio input is not heard by the listener, the overall quality for other audio inputs can be improved due to different bit rate allocation. The behaviour of some operations (mono, stereo, MASA, SBA, ISM1) can furthermore depend on or be based on whether DTX operation is activated. For example, where DTX can be implemented the following operations for muting can be employed. In other words, the muting of the audio input can again be applied by directly modifying the corresponding audio input gain prior to IVAS encoder. In addition, the IVAS encoder operation is set to ‘DTX on’, i.e., DTX operation is activated. For example, during negotiation, DTX operation can be negotiated (allowed). In some embodiments the MUTE PI frame can furthermore implicitly or explicitly indicate or signal a request to also activate DTX operation in the send direction. For example, this activation can be limited to a MUTE activation time range (e.g., based on a “validity” parameter within the PI frame, which indicates how long a PI frame is valid). In other words though there can be implemented other methods for reducing a bitrate, such as, by pausing the RTP stream (e.g., in case of a single RTP stream carrying a single IVAS stream), the RTP pause mechanism does not enable continued SID frames which have advantages to maintain subjective experience compared to entirely stopping the RTP stream. The characteristics and benefits of employing Silence Descriptor (SID) frames are that SID frames can be sent during speech inactivity to keep the receiver CNG sufficiently aligned with the background noise level and spectral shape at the sender side. This is of particular importance at the onset of each new talk spurt. Thus, SID frames should not be too old, when a new talks burst starts. Commonly SID frames are sent regularly, e.g., every 8th frame, but a codec may allow also variable rate SID updates. SID frames are typically quite small, e.g., 2.4 kbps SID bitrate equals 48 bits per frame for the typical 20-ms frame size used in modern communications codecs (AMR, AMR-WB, ITU-T G.718, EVS, and IVAS). IVAS SID frame size is 5.2 kbps (104 bits) for a 20-ms frame, and the default SID frame interval in IVAS is 8 frames. Other update intervals are furthermore supported. IVAS includes EVS, and in EVS primary mode operation the 48-bit SID is kept. An example flow diagram is shown in Figure 12 in relation to the muting of IVAS streams. Thus as shown in Figure 12 by 1201 is the operation of obtaining the mute input (at the receiving UE2) and generating the MUTE request. The MUTE request is then sent from the receiving UE2 to the sending side (UE1 or server) as shown by 1203 in Figure 12. Following the reception of the MUTE request, the application of the gain control for audio indicated via the MUTE request is shown by 1205 in Figure 12. Then is the provision of the audio encoder DTX operation (for example switch from DTX off to DTX on) based on the MUTE request as shown by 1207 in Figure 12. There is then an operation of providing the audio input modified by gain control to the encoder as shown by 1209 in Figure 12. There is then an operation of encoding and transmitting at least the audio input modified by gain control as shown by 1211 in Figure 12. This can result in the operation of receiving and decoding at least the encoded audio modified by gain control as shown by 1213 in Figure 12. In such a manner any audio input for which a muting is requested is seen as a zero-signal by the IVAS encoder. The IVAS encoder encodes the zeroed audio input and enters DTX based on VAD. SID frames are then intermittently transmitted to receiver according to DTX operation settings. The IVAS decoder / renderer receives the SID frames for the affected frames and processes them accordingly using CNG. The muted audio input is not heard by the listener, in its place may be heard comfort noise according to SID / CNG. Any other audio playback (i.e., streams that are not muted) is unaffected. In some embodiments, the muting of the audio input can be implemented by directly modifying the corresponding audio input gain prior to IVAS encoder, and the IVAS encoder operation can be set to ‘DTX on’. In addition, a new processing is employed for handling of SID frames at the decoder / renderer based on receiver UE2 having locally the MUTE information it has transmitted to sending UE1 (and UE3) to mute the audio signal. In other words, this processing guarantees the muted audio is fully muted in playback. Thus in some embodiments, when the IVAS decoder / renderer receives and decodes a SID frame (or NO_DATA), or an external renderer receives a decoded signal based on SID frames (or NO_DATA) and this corresponds with a local MUTE instruction for the stream or stream component, said stream or stream component is muted or zeroed, for example by applying a zero gain to it. The zero-signal can then be provided for rendering and playback. As a result in this example, no audio at all (in other words also no comfort noise under any circumstances) is heard by the listener for the muted audio. Thus, the muting is not dependent on SID / CNG implementation (or PLC), while more efficient transmission is achieved by applying DTX. This example can, for example, be shown with respect to the flow diagram of Figure 13. Thus, as shown in Figure 13 by 1301 is the operation of obtaining the mute input (at the receiving UE2) and generating the MUTE request. The MUTE request is then sent from the receiving UE2 to the sending side (UE1 or server) as shown by 1303 in Figure 13. Following the reception of the MUTE request, the application of the gain control for audio indicated via the MUTE request as shown by 1305 in Figure 13. Then is the provision of the audio encoder DTX operation (for example switch from DTX off to DTX on) based on the MUTE request as shown by 1307 in Figure 13. There is then an operation of providing the audio input modified by gain control to the encoder as shown by 1309 in Figure 13. There is then an operation of encoding and transmitting at least the audio input modified by gain control as shown by 1311 in Figure 13. This can result in the operation of receiving and decoding at least the encoded audio modified by gain control as shown by 1313 in Figure 13. Then is shown the operation of upon obtaining SID (or NO_DATA) frame at decoder, apply gain control for audio indicated by corresponding MUTE request before playback as shown by 1315 in Figure 13. It is noted that the above processing ignores the local muting for a received active frame that is requested muted by the receiving UE2. This distinction is intentional in case, for example, a service considers a specific input important thus not permitting it being selectively muted by user. In other examples, the local muting can be applied directly. In some embodiments, a MUTE PI frame, in addition to activating DTX, can be configured to signal or indicate a SID update rate, for example set the SID update rate to its maximum possible value or to set the length of the MUTE period (for example by employing the validity parameter). In such a manner a further transmission efficiency is achieved since SID frames are not needed when even non-zero CNG output is zeroed for playback. Thus, there is no need to keep the receiver CNG aligned with the background noise level and spectral shape on the sender side, which would be achieved by the SID frame calculation and transmission. In some embodiments the receiver is configured to know if some streams (or stream components) are muted by the sender. For example, the receiver, in some embodiments is configured to identify if more streams (or stream components) are available, but are currently muted. In the language stream example, the receiver might want to switch to receive some other language that was muted earlier in the session. If the receiver is not aware of the other muted language streams (e.g., the receiver has not kept a track record of the muted streams), then these available muted streams could be explicitly signalled to the receiver by the sender. The available streams (or stream components) that are muted can in some embodiments be signalled through a suitable PI frame. For example, PI frames of type STREAM-STATUS or COMPONENT-STATUS could be used for the signalling. The STREAM-STATUS PI frame, for example, can be employed to indicate the status (muted or not muted) for each available stream in the session. The COMPONENT-STATUS PI frame can be employed to indicate the status (muted or not muted) of individual stream components. For example, the COMPONENT-STATUS PI frame could be used to indicate the status of the MASA part and the individual audio objects in OMASA format. In some embodiments to switch the received stream (e.g., from one language to another), the receiver can mute a currently received stream and unmute another stream that is currently muted. The un-muting can be signalled with a UN_MUTE (or similar) PI frame type. The UN_MUTE PI frame type can be defined similar to the MUTE PI frame, but instead of muting a stream (or a stream component) the targeted or identified stream (or stream component) is un-muted (or resumed to transmission). If the targeted stream (or stream component) is currently not muted, the sender can ignore the UN_MUTE request. In some examples, where an immersive audio encoder input comprises a combined audio input, e.g., MASA + 2 or more objects (e.g., OMASA input for IVAS is one such example), the combined audio input processing can select the most important object for encoding separately, while the at least one additional object can be downmixed with the MASA spatial audio. Such selection can be based, e.g., on energy at speech onset or similar method. In other words, the object can be separately encoded as long as it remains active. When an audio object that is selected as the most important object is now muted, e.g., based on an indicator to the encoder triggered by a MUTE request, the encoder algorithm may select the next important object as separated instead. In some embodiments, the mute request can be transmitted within the IVAS RTP payload header, e.g., with the IVAS E-bytes. For example, if the IVAS payload supports multiple E-bytes, some subsequent E-byte could be used to indicate a mute request. In some embodiments, the mute request can be transmitted in the RTP Header Extension or via RTCP. In some embodiments, the muting method can be negotiated for a session, e.g., via SDP (session description protocol) or SIP (session initiation protocol). The method could be negotiated, for example, with a mute-method SDP parameter, where the possible values for the parameter would include: - Gain editing of the muted component or stream; - Stopping the transmission of the muted component or stream; and - Replacing the muted component or stream with SID frames; The mute-method parameter could include only a single possible value to indicate which of the above methods is used for muting. In another example, the mutemethod parameter could include a list of the supported methods, in which case the negotiated list would indicate a variety of methods for muting. The mute-method parameter could be named differently in other examples. In some embodiments, the muting via gain editing could be negotiated further for a session, e.g., via SDP or SIP. For example, the session participants could negotiate how much the gain is edited for a muted component or stream. For example, an SDP parameter of mute-gain could be negotiated to indicate the gain applied to the muted component or stream by the sender. The mute-gain parameter could be named differently in other examples. In further examples, the gain editing parameter (e.g., mute-gain) could be combined with the mute-method parameter so that the mute-method parameter would directly indicate the amount of gain to be applied to the muted component or stream. In some examples, the mute-gain parameter could indicate how much gain is reduced for the muted component or stream (i.e., the mute-gain would indicate the reverse of the above example). Figure 14 shows a flow diagram of a summary of the operation of the handling of muted channels or streams. For example there can be an operation of obtaining, from a further apparatus, at least one bitstream comprising at least one component or stream of an immersive audio bitstream as shown in Figure 14 by 1401. Then is shown an operation of identifying at least one of the at least one component or stream of the immersive audio bitstream to be muted in Figure 14 by 1403. Furthermore there is an operation of transmitting, to the further apparatus, a mute request indicating the identified at least one component or stream to be muted as shown in Figure 14 by 1405. This can result in obtaining, from the further apparatus, the at least one bitstream, wherein the at least one bitstream is modified with respect to the identified at least one component or stream of the at least one immersive audio bitstream to be muted as shown in Figure 14 by 1407. In some embodiments herein the modification in the at least one bitstream is at least one of: the at least one component or stream having reduced energy; the at least one component or stream not being received; the at least one component or stream comprises a silence descriptor indicator for indicating the at least one component or stream is to be muted; the at least one component or stream having a reduced bit rate; and at least one other component or stream having an increased bit rate. In some embodiments as shown herein obtaining, from the further apparatus, the bitstream, may comprise depacketizing and decoding the at least one immersive audio bitstream from the bitstream. The depacketized and decoded at least one immersive audio bitstream may comprise at least one decoded IVAS frame, the at least one decoded IVAS frame comprising one of: at least one component having been the identified at least one component to be muted such that the at least one component has reduced energy; or at least one stream having been the identified at least one stream to be muted such that the at least one stream has reduced energy. The silence descriptor indicator for indicating the at least one component or stream may comprise at least one silence descriptor frame within the at least one component or stream. The embodiments may further comprise rendering from the obtained immersive audio bitstream at least one audio signal for playback wherein the at least one component or stream of an immersive audio bitstream to be muted, wherein rendering may comprise generating the at least one audio signal for playback comprising at least one of: the identified at least one component or stream of the immersive audio bitstream with a significantly decreased signal energy compared to a signal energy of the identified at least one component or stream of the immersive audio bitstream prior to the mute request; a selectively muted at least one component or stream; and comfort noise to selectively replace the at least one component or stream. The mute request may comprise metadata, such as described above, comprising an identifier field for indicating the identified at least one component or stream to be muted. The metadata may further comprise a validity field configured to identify a time over which the muting is to be applied. The obtaining, from the further apparatus, the bitstream or modified bitstream may further comprise obtaining, from the further apparatus, metadata for identifying a status of the at least one component or stream. The metadata for identifying a status of the at least one component or the stream may be configured to indicate which component or stream is muted. The metadata may be a processing information frame. In some embodiments the transmitting to the further apparatus the mute request indicating the identified at least one component or stream to be muted may comprise transmitting at least one of: the mute request within an IVAS RTP payload header; the mute request within a RTP header extension; the mute request within a RTCP payload; transmitting the mute request using an IVAS E-byte; and transmit the mute request using IVAS PI data. The method may further comprise transmitting, to the further apparatus, an unmute request indicating the identified at least one component or stream to be unmuted. The rendering may further comprise generating the at least one audio signal for playback comprising at least one of: the identified at least one component or stream of the immersive audio bitstream with a significantly increased signal energy following the unmute request compared to the signal energy of the identified at least one component or stream of the immersive audio bitstream following the mute request; a selectively unmuted at least one component or stream; and the at least one component or stream to selectively replace the comfort noise. Furthermore the transmitting to the further apparatus the unmute request may comprise transmitting at least one of: the unmute request within an IVAS RTP payload header; the unmute request within a RTP header extension; the unmute request within a RTCP payload; the unmute request using an IVAS E-byte; and the mute request using IVAS PI data. As described above the method may further comprise negotiating with the further apparatus support for handling the mute request indicating the identified at least one component or stream to be muted. This negotiating with the further apparatus support for handling the mute request indicating the identified at least one component or stream to be muted may comprise negotiating employing a session description file. Furthermore identifying at least one component or stream of the immersive audio bitstream is to be muted may comprise obtaining at least one mute input for identifying the at least one component or stream of the immersive audio bitstream. As shown in Figure 15 is a flow diagram of example operations summarizing the embodiments with respect to the generation of suitable streams or components (of which one is to be muted). For example there is shown in Figure 15 an operation of transmitting to a further apparatus, at least one bitstream comprising at least one component or stream of an immersive audio bitstream by 1501. Following this can be an operation of obtaining, from the further apparatus, a mute request identifying at least one component or stream from the at least one component or stream to be muted as shown in Figure 15 by 1503. After this can be an operation of encoding, based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream as shown in Figure 15 by 1505. After the encoding, there is an operation of transmitting, to the further apparatus, the modified at least one bitstream as shown in Figure 15 by 1507. The modified at least one bitstream as described above can comprise the identified at least one component or stream, wherein the modification in the at least one bitstream, is at least one of: the at least one component or stream having reduced energy; the at least one component or stream not being transmitted; the at least one component or stream comprises a silence descriptor indicator for indicating the at least one component or stream is to be muted; the at least one component or stream having a reduced bit rate; and at least one other component or stream having an increased bit rate. In some embodiments the encoding, based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream can comprise packetizing and encoding the identified at least one component or stream. Furthermore, encoding based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream can comprise employing an IVAS encoder to generate at least one encoded IVAS frame, the at least one encoded IVAS frame comprising one of: at least one component, identified by the mute request; or at least one stream, identified by the mute request. In some embodiments as discussed above the encoding, based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream may comprise: applying a gain control processing to the identified at least one component or stream; encoding at least the gain control processed at least one component or stream to generate the modified at least one bitstream. The encoding at least the gain control processed at least one component or stream to generate the modified at least one bitstream may further comprise generating an encoded IVAS frame comprising at least one silence descriptor frame based on a signal energy of the gain controlled at least one component or stream being below the IVAS encoder signal level threshold. Furthermore the encoding at least the gain control processed at least one component or stream to generate a modified at least one bitstream may comprise adaptively encoding the at least one component or stream and / or the at least one other component or stream based on at least one of: a signal level of the gain controlled at least one component or stream; a relative signal level of the gain controlled at least one component or stream and at least one other component or stream. The mute request can in some embodiments comprise metadata comprising an identifier field for indicating the identified at least one component or stream to be muted. The metadata may further comprise a validity field configured to identify a time over which the muting is to be applied. Transmitting, to the further apparatus, the modified at least one bitstream may comprise transmitting, to the further apparatus, metadata for identifying a status of the identified at least one component or stream. The metadata for identifying a status of the identified at least one component or stream may be configured to indicate the at least one component or at least one stream is muted. The metadata may be a processing information frame. Receiving from the further apparatus the mute request indicating the identified at least one component or stream to be muted may comprise receiving at least one of: the mute request within an IVAS RTP payload header; the mute request within a RTP header extension; the mute request within a RTCP payload; the mute request using an IVAS E-byte; and the mute request using IVAS PI data. The method may further comprise obtaining, from the further apparatus, an unmute request indicating the identified at least one component or stream to be unmuted. The method may further comprise de-applying the muting to the indicated at least one component or stream identified by the unmute request. Receiving from the further apparatus the unmute request indicating the identified at least one component or stream to be unmuted may comprise receiving at least one of: the unmute request within an IVAS RTP payload header; the unmute request within a RTP header extension; the unmute request within a RTCP payload; receiving the unmute request using an IVAS E-byte; and receive the unmute request using IVAS PI data. The method may further comprise negotiating with the further apparatus support for handling the mute request indicating the identified the at least one component or stream to be muted. Negotiating with the further apparatus support for handling the mute request indicating the identified at least one component or stream to be muted may comprise negotiating employing a session description file. With respect to Figure 16 an example electronic device is shown. The device may be any suitable electronics device or apparatus. For example in some embodiments the device 1900 is a mobile device, user equipment, tablet computer, computer, audio playback apparatus, etc. In some embodiments the device 1900 comprises at least one processor or central processing unit 1907. The processor 1907 can be configured to execute various program codes such as the methods such as described herein. In some embodiments the device 1900 comprises a memory 1911. In some embodiments the at least one processor 1907 is coupled to the memory 1911. The memory 1911 can be any suitable storage means. In some embodiments the memory 1911 comprises a program code section for storing program codes implementable upon the processor 1907. Furthermore, in some embodiments the memory 1911 can further comprise a stored data section for storing data, for example data that has been processed or to be processed in accordance with the embodiments as described herein. The implemented program code stored within the program code section and the data stored within the stored data section can be retrieved by the processor 1907 whenever needed via the memory-processor coupling. In some embodiments the device 1900 comprises a user interface 1905. The user interface 1905 can be coupled in some embodiments to the processor 1907. In some embodiments the processor 1907 can control the operation of the user interface 1905 and receive inputs from the user interface 1905. In some embodiments the user interface 1905 can enable a user to input commands to the device 1900, for example via a keypad. In some embodiments the user interface 1905 can enable the user to obtain information from the device 1900. For example, the user interface 1905 may comprise a display configured to display information from the device 1900 to the user. The user interface 1905 can in some embodiments comprise a touch screen or touch interface capable of both enabling information to be entered to the device 1900 and further displaying information to the user of the device 1900. In some embodiments the device 1900 comprises an input / output port 1909. The input / output port 1909 in some embodiments comprises a transceiver. The transceiver in such embodiments can be coupled to the processor 1907 and configured to enable a communication with other apparatus or electronic devices, for example via a wireless communications network. The transceiver or any suitable transceiver or transmitter and / or receiver means can in some embodiments be configured to communicate with other electronic devices or apparatus via a wire or wired coupling. The transceiver can communicate with further apparatus by any suitable known communications protocol. For example, in some embodiments the transceiver can use a suitable universal mobile telecommunications system (UMTS) protocol, a wireless local area network (WLAN) protocol such as for example IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth, or infrared data communication pathway (IRDA). The transceiver input / output port 1909 may be configured to receive the signals and in some embodiments obtain the focus parameters as described herein. In some embodiments the device 1900 may be employed to generate a suitable audio signal using the processor 1907 executing suitable code. The input / output port 1909 may be coupled to any suitable audio output for example to a multichannel speaker system and / or headphones (which may be a headtracked or a non-tracked headphones) or similar. In general, the various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto. While various aspects of the invention may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof. The embodiments of this invention may be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD. The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processors may be of any type suitable to the local technical environment, and may include one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), gate level circuits and processors based on multi-core processor architecture, as non-limiting examples. Embodiments of the inventions may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate. Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre-stored design modules. Once the design for a semiconductor circuit has been completed, the resultant design, in a standardized electronic format (e.g., Opus, GDSII, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication. The foregoing description has provided by way of exemplary and nonlimiting examples a full and informative description of the exemplary embodiment of this invention. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still fall within the scope of this invention as defined in the appended claims. 3GPP 3rd Generation Partnership Project AMR-WB IO Adaptive Multi Rate Wideband Inter-Operable BLI Bitrate Level Indication CMR Codec mode request DTX Discontinuous transmission EVS Enhanced Voice Services FT Frame type (index) ISM Independent Streams with Metadata (i.e., type of Object-Based Audio) IVAS Immersive Voice and Audio Services kbps MASA kilobits per second Metadata-Assisted Spatial Audio MC Multichannel OMASA Object-based audio with MASA (combined input format) OS BA Object-based audio with SBA (combined input format) PI Processing information (audio) RTCP Real-Time Transport Control Protocol RTP Real-Time Transport Protocol SBA Scene-Based Audio SDP Session Description Protocol SSB Supplemental Signaling Bits ToC Table of Content UE User equipment

Claims

1. An apparatus comprising means configured to:obtain, from a further apparatus, at least one bitstream comprising at least one component or stream of an immersive audio bitstream;identify at least one of the at least one component or stream of the immersive audio bitstream to be muted;transmit, to a further apparatus, a mute request indicating the identified at least one component or stream to be muted;obtain, from the further apparatus, the at least one bitstream, wherein the at least one bitstream is modified with respect to the identified at least one component or stream of the at least one immersive audio bitstream to be muted, wherein the modification in the at least one bitstream, is at least one of:the at least one component or stream having reduced energy;the at least one component or stream not being received;the at least one component or stream comprises a silence descriptor indicator for indicating the at least one component or stream is to be muted;the at least one component or stream having a reduced bit rate; andat least one other component or stream having an increased bit rate.

2. The apparatus as claimed in claim 1, wherein the means configured to obtain, from the further apparatus, the bitstream, is configured to depacketize and decode the at least one immersive audio bitstream from the bitstream.

3. The apparatus as claimed in claim 2, wherein the depacketized and decoded at least one immersive audio bitstream comprises at least one decoded IVAS frame, the at least one decoded IVAS frame comprising one of:at least one component having been the identified at least one component to be muted such that the at least one component has reduced energy; orat least one stream having been the identified at least one stream to be muted such that the at least one stream has reduced energy.

4. The apparatus as claimed in any of claims 1 to 3, wherein the silence descriptor indicator for indicating the at least one component or stream comprises at least one silence descriptor frame within the at least one component or stream.

5. The apparatus as claimed in any of claims 1 to 4, wherein the means is further configured to render from the obtained immersive audio bitstream at least one audio signal for playback wherein the at least one component or stream of an immersive audio bitstream to be muted, wherein the means configured to render is configured to generate the at least one audio signal for playback comprising at least one of:the identified at least one component or stream of the immersive audio bitstream with a significantly decreased signal energy compared to a signal energy of the identified at least one component or stream of the immersive audio bitstream prior to the mute request;a selectively muted at least one component or stream; and comfort noise to selectively replace the at least one component or stream.

6. The apparatus as claimed in any of claims 1 to 5, wherein the mute request comprises metadata comprising an identifier field for indicating the identified at least one component or stream to be muted.

7. The apparatus as claimed in any of claims 1 to 6, wherein the means configured to obtain, from the further apparatus, the bitstream or modified bitstream is configured to obtain, from the further apparatus, metadata for identifying a status of the at least one component or stream.

8. The apparatus as claimed in claim 7, wherein the metadata for identifying a status of the at least one component or the stream is configured to indicate which component or stream is muted.

9. The apparatus as claimed in any of claims 6 to 8, wherein the metadata is a processing information frame.

10. The apparatus as claimed in any of claims 1 to 9, wherein the means configured to transmit to the further apparatus the mute request indicating the identified at least one component or stream to be muted is configured to perform at least one of:transmit the mute request within an IVAS RTP payload header;transmit the mute request within a RTP header extension;transmit the mute request within a RTCP payload;transmit the mute request using an IVAS E-byte; and transmit the mute request using IVAS PI data.

11. The apparatus as claimed in any of claims 1 to 10, wherein the means configured to identify at least one component or stream of an immersive audio bitstream is to be muted is configured to obtain at least one mute input for identifying the at least one component or stream of the immersive audio bitstream.

12. An apparatus comprising means configured to:transmit to a further apparatus, at least one bitstream comprising at least one component or stream of an immersive audio bitstream;obtain, from a further apparatus, a mute request identifying at least one component or stream from the at least one component or stream to be muted;encode, based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream;transmit, to the further apparatus, the modified at least one bitstream, wherein the modified at least one bitstream comprising the identified at least one component or stream, wherein the modification in the at least one bitstream, is at least one of:the at least one component or stream having reduced energy;the at least one component or stream not being transmitted;the at least one component or stream comprises a silence descriptor indicator for indicating the at least one component or stream is to be muted;the at least one component or stream having a reduced bit rate; and at least one other component or stream having an increased bit rate.

13. The apparatus as claimed in claim 12, wherein the means configured to encode, based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream is configured to packetize and encode the identified at least one component or stream.

14. The apparatus as claimed in any of claims 12 or 13, wherein the means configured to encode, based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream is configured to employ an IVAS encoder to generate at least one encoded IVAS frame, the at least one encoded IVAS frame comprising one of:at least one component, identified by the mute request; or at least one stream, identified by the mute request.

15. The apparatus as claimed in any of claims 12 to 14, wherein the means configured to encode, based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream is configured to:apply a gain control processing to the identified at least one component or stream;encode at least the gain control processed at least one component or stream to generate the modified at least one bitstream.

16. The apparatus as claimed in claim 15, wherein the means configured to encode at least the gain control processed at least one component or stream to generate the modified at least one bitstream is configured to generate an encoded IVAS frame comprising at least one silence descriptor frame based on a signal energy of the gain controlled at least one component or stream being below the IVAS encoder signal level threshold.

17. The apparatus as claimed in any of claims 15 to 16, wherein the means configured to encode at least the gain control processed at least one component or stream to generate a modified at least one bitstream is configured to adaptivelyencode the at least one component or stream and / or the at least one other component or stream based on at least one of:a signal level of the gain controlled at least one component or stream;a relative signal level of the gain controlled at least one component or stream and at least one other component or stream.

18. The apparatus as claimed in any of claims 12 to 17, wherein the mute request comprises metadata comprising an identifier field for indicating the identified at least one component or stream to be muted.

19. The apparatus as claimed in any of claims 12 to 18, wherein the means configured to transmit, to the further apparatus, the modified at least one bitstream is configured to transmit, to the further apparatus, metadata for identifying a status of the identified at least one component or stream.

20. The apparatus as claimed in claim 19, wherein the metadata for identifying a status of the identified at least one component or stream is configured to indicate the at least one component or at least one stream is muted.

21. The apparatus as claimed in any of claims 18 to 20, wherein the metadata is a processing information frame.

22. The apparatus as claimed in any of claims 12 to 21, wherein the means configured to receive from the further apparatus the mute request indicating the identified at least one component or stream to be muted is configured to, at least one of:receive the mute request within an IVAS RTP payload header;receive the mute request within a RTP header extension;receive the mute request within a RTCP payload;receive the mute request using an IVAS E-byte; andreceive the mute request using IVAS PI data.

23. The apparatus as claimed in any of claims 12 to 22, wherein the means is further configured to negotiate with the further apparatus support for handling the mute request indicating the identified the at least one component or stream to be muted.

24. A method for an apparatus, the method comprising:obtaining, from a further apparatus, at least one bitstream comprising at least one component or stream of an immersive audio bitstream;identifying at least one of the at least one component or stream of the immersive audio bitstream to be muted;transmitting, to a further apparatus, a mute request indicating the identified at least one component or stream to be muted;obtaining, from the further apparatus, the at least one bitstream, wherein the at least one bitstream is modified with respect to the identified at least one component or stream of the at least one immersive audio bitstream to be muted, wherein the modification in the at least one bitstream, is at least one of:the at least one component or stream having reduced energy;the at least one component or stream not being received;the at least one component or stream comprises a silence descriptor indicator for indicating the at least one component or stream is to be muted;the at least one component or stream having a reduced bit rate; and at least one other component or stream having an increased bit rate.

25. A method for an apparatus, the method comprising:transmitting to a further apparatus, at least one bitstream comprising at least one component or stream of an immersive audio bitstream;obtaining, from a further apparatus, a mute request identifying at least one component or stream from the at least one component or stream to be muted;encoding, based on the mute request, at least the identified at least one component or stream to generate a modified at least one bitstream;transmitting, to the further apparatus, the modified at least one bitstream, wherein the modified at least one bitstream comprising the identified at least onecomponent or stream, wherein the modification in the at least one bitstream, is at least one of:the at least one component or stream having reduced energy;the at least one component or stream not being transmitted;5 the at least one component or stream comprises a silence descriptorindicator for indicating the at least one component or stream is to be muted;the at least one component or stream having a reduced bit rate; and at least one other component or stream having an increased bit rate.66Application No: GB2406053.5Claims searched: 1-25Examiner: Contract Unit ExaminerDate of search: 31 October 2024Patents Act 1977: Search Report under Section 17Documents considered to be relevant:Category Relevant to claims Identity of document and passage or figure of particular relevance X 1,2, 4, 5, 11-13,24, 25 US2013 / 044893 Al (MAUCHLY J WILLIAM ET AL) figures 1, 4, paragraph [0036], paragraph [0028], paragraph [0033] A - US2023 / 306975 Al (FUCHS GUILLAUME ET AL) paragraph [0007] A - 3GPP DRAFT, 2024, "3rd Generation Partnership Project; Technical Specification Group Services and System Aspects; Codec for Immersive Voice and Audio Services; Detailed Algorithmic Description inc. RTP payload format and SDP parameter definitions (Release 18)" Sections 5.6.6.6 and 8.2.2 A - US2023 / 267938 Al (MUNDT HARALD ET AL) the whole document A - GB2596138 A (NOKIA TECHNOLOGIES OY) the whole document A - EP3923280 Al (NOKIA TECHNOLOGIES OY) the whole document A - 3GPP STANDARD, vol SA WG4, 2024, "3rd Generation Partnership Project; Technical Specification Group Services and System Aspects; Codec for Immersive Voice and Audio Services; Error concealment of lost packets (Release 18)", pages 1-9 URL: https: / / ftp.3gpp.org / Specs / archive / 26_series / 26.255 / 26255-i00.zip the whole documentCategories:X Document indicating lack of novelty or inventive step A Document indicating technological background and / or state of the art. Y Document indicating lack of inventive step if P Document published on or after the declared priority date but combined with one or more other documents of same category. before the filing date of this invention. & Member of the same patent family E Patent document published on or after, but with priority date earlier than, the filing date of this application.Field of Search:International Classification:Subclass Subgroup Valid From None

Citation Information

Patent Citations

  • Apparatus and methods

    GB202313472D0

  • Immersive conversational audio

    GB2635735A

  • Adapting multi-source inputs for constant rate encoding

    EP3923280A1

  • Decoder spatial comfort noise generation for discontinuous transmission operation

    GB2596138A

  • System and method for muting audio associated with a source

    US20130044893A1