Immersive communication sessions
Patent Information
- Application Number
- PCT/EP2026/057888
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-27
- Filing Date
- 2026-03-20
- Publication Date
- 2026-10-01
Smart Images

Figure EP2026057888_01102026_PF_FP_ABST
Abstract
Description
[0001] TITLE
[0002] Immersive Communication Sessions
[0003] TECHNOLOGICAL FIELD
[0004] Examples of the disclosure relate to immersive communication sessions. Some relate to session negotiations for immersive communication sessions.
[0005] BACKGROUND
[0006] In traditional mono telephony, User Equipment (UE) is expected to remove all background noise and ambience while maintaining sufficient speech quality. However, in immersive communication sessions it can be beneficial to retain some of the background ambience.
[0007] BRIEF SUMMARY
[0008] According to various, but not necessarily all, examples of the disclosure there is provided an apparatus for immersive audio communication comprising:
[0009] at least one processor; and
[0010] at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to perform at least:
[0011] determining at least one audio processing mode that is supported by the apparatus; transmitting a list comprising the at least one audio processing mode;
[0012] receiving a response indicating at least one preferred audio processing mode wherein the at least one preferred audio processing mode is based on the transmitted list; and
[0013] using a preferred audio processing mode indicated in the response for an immersive audio communication session.
[0014] The different audio processing modes can comprise at least one of: different audio suppression modes; different speech enhancement modes; modes with different levels of amplification; modes with different levels of compression; modes with different levels of equalization; modes with different levels of bass roll-off; modes with different levels of suppression; modes with different levels of noise suppression; and modes with different levels of ambience suppression.
[0015] When a first audio processing mode is used processed audio signals can comprise speech and ambient sound and when a second audio processing mode is used the ambient sound can be suppressed in the processed audio signals.The respective audio processing modes can affect the content of a transmitted bitstream such that when a first audio processing mode is used the transmitted bitstream comprises speech and ambient sound and when a second audio processing mode is used ambient sound is suppressed in the transmitted bitstream to enhance the intelligibility of speech compared to the first audio processing mode.
[0016] The apparatus can support multiple different audio processing modes with different levels of suppression of the ambient sound.
[0017] The at least one audio processing mode can be determined at least in part based on one of: negotiated bitrate parameter; offered bitrate parameter; ambient noise in the environment of the apparatus; capture device used by the apparatus; and output device used by the apparatus.
[0018] The list can be transmitted using session description protocol (SDP) signalling. The list can be indicated using an audio processing mode negotiation parameter.
[0019] The processor and memory can also be configured to cause the apparatus to perform;
[0020] receiving a list comprising at least one audio processing mode;
[0021] determining at least one preferred audio processing mode from the received list;
[0022] transmitting a response indicating the determined at least one preferred audio processing mode of the apparatus; and
[0023] using a preferred audio processing mode indicated in the sent response for the immersive audio communication session
[0024] The immersive audio communication session can use the same audio processing mode in both directions.
[0025] The respective audio processing modes can affect the content of a transmitted bitstream and the immersive audio communication session can be established using a first audio processing mode for bitstreams transmitted by the apparatus and a second audio processing mode for bitstreams received by the apparatus.
[0026] The received response can indicate an order of preference for the at least one audio processing mode.
[0027] According to various, but not necessarily all, examples of the disclosure there may be provided a method comprising:determining at least one audio processing mode that is supported by the apparatus; transmitting a list comprising the at least one audio processing mode;
[0028] receiving a response indicating at least one preferred audio processing mode wherein the at least one preferred audio processing mode is based on the transmitted list; and
[0029] using a preferred audio processing mode indicated in the response for an immersive audio communication session.
[0030] According to various, but not necessarily all, examples of the disclosure there may be provided a computer program comprising computer program instructions for causing an apparatus to perform at least the following or for performing at least the following:
[0031] determining at least one audio processing mode that is supported by the apparatus; transmitting a list comprising the at least one audio processing mode;
[0032] receiving a response indicating at least one preferred audio processing mode wherein the at least one preferred audio processing mode is based on the transmitted list; and
[0033] using a preferred audio processing mode indicated in the response for an immersive audio communication session.
[0034] According to various, but not necessarily all, examples of the disclosure there may be provided an apparatus for immersive audio communication comprising:
[0035] at least one processor; and
[0036] at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to perform at least:
[0037] receiving a list comprising at least one audio processing mode that is supported for an immersive audio communication session;
[0038] determining at least one preferred audio processing mode from the received list;
[0039] transmitting a response indicating a determined at least one preferred audio processing mode of the apparatus; and
[0040] using a preferred audio processing mode indicated in the response for the immersive audio communication session.
[0041] The preferred at least one audio processing mode can be determined based on factors comprising at least one of: ambient noise in the environment of the apparatus; capture device used by the apparatus; output device used by the apparatus; and available bitrate levels.The transmitted response can indicate an order of preference for the at least one audio processing mode.
[0042] The processer and the memory can be arranged to cause the apparatus to perform determining that at least one of the factors has changed and transmitting a request to change the audio processing mode used for the immersive audio communication session.
[0043] The different audio processing modes can comprise at least one of: different audio suppression modes; different speech enhancement modes; modes with different levels of amplification; modes with different levels of compression; modes with different levels of equalization; modes with different levels of bass roll-off; modes with different levels of suppression; modes with different levels of noise suppression; and modes with different levels of ambience suppression.
[0044] When a first audio processing mode is used processed audio signals can comprise speech and ambient sound and when a second audio processing mode is used the ambient sound can be suppressed in the processed audio signals.
[0045] The respective audio processing modes can affect the content of a received bitstream such that when a first audio processing mode is used the received bitstream comprises speech and ambient sound and when a second audio processing mode is used ambient sound is suppressed in the received bitstream to enhance the intelligibility of speech compared to the first audio processing mode.
[0046] The apparatus can support multiple different audio processing modes with different levels of suppression of the ambient sound.
[0047] The list can be received using session description protocol (SDP) signalling. The list can be indicated using an audio processing mode negotiation parameter.
[0048] The processor and memory can also be configured to cause the apparatus to perform;
[0049] determining at least one audio processing mode;
[0050] transmitting a list comprising at least one audio processing mode;
[0051] receiving a response indicating at least one preferred audio processing mode wherein the at least one preferred audio processing mode is based on the transmitted list; and
[0052] using a preferred audio processing mode indicated in the received response for the immersive audio communication session.The immersive audio communication session can use the same audio processing mode in both directions.
[0053] The respective audio processing modes can affect the content of a transmitted bitstream and the immersive audio communication session can be established using a first audio processing mode for bitstreams received by the apparatus and a second audio processing mode for bitstreams transmitted by the apparatus.
[0054] According to various, but not necessarily all, examples of the disclosure there may be provided a method comprising:
[0055] determining at least one audio processing mode that is supported by the apparatus; transmitting a list comprising the at least one audio processing mode;
[0056] receiving a response indicating at least one preferred audio processing mode wherein the at least one preferred audio processing mode is based on the transmitted list; and
[0057] using a preferred audio processing mode indicated in the response for an immersive audio communication session.
[0058] According to various, but not necessarily all, examples of the disclosure there may be provided a computer program comprising computer program instructions for causing an apparatus to perform at least the following or for performing at least the following:
[0059] determining at least one audio processing mode that is supported by the apparatus; transmitting a list comprising the at least one audio processing mode;
[0060] receiving a response indicating at least one preferred audio processing mode wherein the at least one preferred audio processing mode is based on the transmitted list; and
[0061] using a preferred audio processing mode indicated in the response for an immersive audio communication session.
[0062] According to various, but not necessarily all, embodiments there is provided an apparatus comprising: at least one processor; and
[0063] at least one memory;
[0064] the at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to perform at least a part of one or more methods described herein.
[0065] According to various, but not necessarily all, embodiments there is provided an apparatus comprising means for performing at least part of one or more methods described herein. The description of a function and / oraction should additionally be considered to also disclose any means suitable for performing that function and / or action. Functions and / or actions described herein can be performed in any suitable way using any suitable method.
[0066] According to various, but not necessarily all, embodiments there is provided examples as claimed in the appended claims.
[0067] While the above examples of the disclosure and optional features are described separately, it is to be understood that their provision in all possible combinations and permutations is contained within the disclosure. It is to be understood that various examples of the disclosure can comprise any or all the features described in respect of other examples of the disclosure, and vice versa. Also, it is to be appreciated that any one or more or all the features, in any combination, may be implemented by / comprised in / performable by an apparatus, a method, and / or instructions as desired, and as appropriate. The description of a function should additionally be considered to also disclose any means suitable for performing that function
[0068] BRIEF DESCRIPTION
[0069] Some examples will now be described with reference to the accompanying drawings in which:
[0070] FIG. 1 shows an example packet structure;
[0071] FIGS. 2A and 2B show example methods;
[0072] FIG. 3 shows an example set up for an immersive audio communication session;
[0073] FIG. 4 shows an example method;
[0074] FIG. 5 shows an example set up for an immersive audio communication session;
[0075] FIG. 6 shows an example method;
[0076] FIGS. 7A and 7B show example reference signals;
[0077] FIG. 8 shows results from an example transmission;
[0078] FIG. 9 shows results from an example transmission;
[0079] FIG. 10 shows the impact of an audio processing mode on speech intelligibility; and
[0080] FIG. 11 shows an example controller.
[0081] The figures are not necessarily to scale. Certain features and views of the figures can be shown schematically or exaggerated in scale in the interest of clarity and conciseness. For example, the dimensions of some elements in the figures can be exaggerated relative to other elements to aid explication. Correspondingreference numerals are used in the figures to designate corresponding features. For clarity, all reference numerals are not necessarily displayed in all figures.
[0082] DETAILED DESCRIPTION
[0083] Immersive audio communication sessions can be implemented using codecs such as 3GPP IVAS (Immersive Voice and Audio Services) or other spatial audio codecs. The spatial audio codecs can support at least three degrees of rotation freedom (yaw, pitch, roll) for all spatial inputs. The IVAS codec, or other suitable spatial audio codecs, can be used in a variety of scenarios. The scenario cannot always be known beforehand.
[0084] The spatial audio codecs can have many supported input formats and output formats. The main supported input formats for IVAS are stereo, multichannel (MC), object-based audio (for example, independent streams with metadata (ISM)), scene-based audio (SBA), and Metadata-assisted spatial audio (MASA). In addition, the following combinations are supported: Objects with MASA (OMASA) and Objects with SBA (OSBA). IVAS furthermore includes the EVS (enhanced voice services) codec for mono input operation. The IVAS output formats comprise mono, stereo, multi-channel (including custom loudspeaker layouts), FOA (first order Ambisonic), H0A2 (higher order Ambisonic 2), H0A3, and binaural.
[0085] The IVAS codec bitstream contains information on the contents of each bitstream frame. This information includes the input format and additional information on the present format, for example sub format signaling, (for instance, order of Ambisonics, channel layout of MC, number of transport channels, or any other suitable information). This format information is in addition to the encoded audio signal and possible metadata. Furthermore, IVAS bitstream frames are designed to be standalone so that an IVAS decoder can start decoding from any valid IVAS frame (containing full data for the format specified in the bitstream) and produce good quality output.
[0086] 3GPP DaCAS (Diverse audio CApturing System for UEs) is a current 3GPP SA4 work item that aims to define immersive audio capture example solutions for converting raw / compensated microphone signals into the relevant IVAS codec input formats. This will be based on definition of a set of target devices or target device types with descriptions of the overall microphone configurations. For example, the number of microphones, their relative positions in the device, and other relevant parameters should be defined.
[0087] Real-time transport protocol (RTP) is used to carry multiple multimedia formats. This permits the transport of new formats without revising the RTP standard. To this end, the information required by a specific application of the protocol is not included in the generic RTP header. For a class of applications (such as audio or video),an RTP profile can be defined. For a media format (such as, a specific video coding format), an associated RTP payload format can be defined. Every instantiation of RTP in a particular application can require a profile and payload format specifications.
[0088] The profile defines the codecs used to encode the payload data and their mapping to payload format codes in the protocol field Payload Type (PT) of the RTP header.
[0089] For example, the RTP profile for audio and video conferences with minimal control is defined in RFC 3551. The profile defines a set of static payload type assignments, and a dynamic mechanism for mapping between a payload format, and a PT value using Session Description Protocol (SDP). The latter mechanism is used for newer video codecs such as RTP payload format for H.264 Video defined in RFC 6184 or RTP Payload Format for High Efficiency Video Coding (HEVC) defined in RFC 7798.
[0090] An RTP session can be established for respective multimedia streams. Audio and video streams can use separate RTP sessions. This enables a receiver to selectively receive components of a particular stream. The RTP specification recommends even port numbers for RTP, and the use of the next odd port number for the associated real-time transport control protocol (RTCP) session. A single port can be used for RTP and RTCP in applications that multiplex the protocols.
[0091] Each RTP stream consists of RTP packets, which in turn comprise an RTP header and payload pairs.
[0092] Fig. 1 shows an example IVAS RTP packet structure 100. The packet structure 100 comprises an RTP header 102 and an IVAS payload 104.
[0093] The RTP Header 102 (and any possible RTP Header Extension) follow the common RTP design.
[0094] The IVAS payload 104 comprises a payload header 106, frame data 108 and processing information (PI) data 110.
[0095] The payload header 106 comprises different types of header bytes: ToC (Table of Content) and E-bytes (Extra bytes,). The ToC bytes describe the content of the frame data 108 (by indicating the size / bitrate for the data frames). The ToC bytes also differentiate IVAS frames from EVS frames, in case the IVAS is operating in mono EVS mode. The E-bytes signal additional information, such as Codec Mode Requests (CMR) which indicate a request to change the bitrate (and possibly other configurations like format andbandwidth) of an incoming IVAS stream. E-bytes can also be used to explicitly indicate the presence of the PI data 110 at the end of the IVAS payload 104.
[0096] The frame data 108 comprises the IVAS data frames (and EVS data frames in a case of using IVAS in mono mode). The data frames represent the encoded IVAS bitstreams. The bitstreams comprise the encoded IVAS audio data with possible additional metadata. The bitstreams can also comprise information about the input / coded format and sub-format of the encoded data (e.g., multichannel 5.1 format). EVS frames do not include such format data.
[0097] The PI data 110 comprises the processing information data and related headers. PI data can be used to transmit any (non-audio) data that can be used to assist the rendering or processing of the audio data, such as, scene and device orientation data, or any other suitable information.
[0098] In some examples the PI data 110 can also be used to request something from other participants in a communication session. For instance, the PI data 110 could be used to mute an incoming stream or increase noise suppression. Feedback data can also be transmitted, for example, the head orientation of a listener.
[0099] In traditional mono telephony, User Equipment (UE) is expected to remove all background noise and ambience while maintaining sufficient speech quality. However, in an immersive audio communication session, the background ambience is not necessarily expected to be removed, or is not expected to be completely removed. In addition, the UEs might support multiple different audio processing modes. The different audio processing modes can provide different levels of noise suppression or other types of processing. The different audio processing modes can enable the UE to capture and deliver varying levels of background ambience along with the speech. When considering an IVAS session between a first UE and a second UE, where both UEs are able to offer and receive a flexible amount of background ambience along the speech, several problems could arise.
[0100] A first problem can arise if the first UE is in an environment comprising a lot of background ambience. The first UE can decide whether to transmit either a full immersive capture signal comprising speech and background ambience, or only the suppressed capture comprising speech. Furthermore, the second UE is in a noisy environment and wants to receive only speech without ambience. Without further knowledge on the preference of the second UE, the first UE might decide to transmit a full immersive signal. This can decrease the speech intelligibility and quality of experience perceived by the user of the second UE.Alternatively, when the second UE is in a quiet or controlled environment, the second UE would prefer to receive a fully immersive sound field comprising both speech and background ambience from the first UE. However, as the first UE has no knowledge of the preference of the second UE, it might decide to suppress all the background ambience, this would not allow the second UE to experience the environment in which the first UE is located. This prevents the second UE from achieving the highest possible immersion.
[0101] A second problem could arise if the first UE is using integrated handsfree loudspeakers or other similar playback device. Due to the playback interface, the first UE is not able to mitigate acoustic echo caused by the playback and capture, when the received signal comprises anything other than speech. Thus, when using the handsfree loudspeakers or other similar devices for playback the first UE would want to receive a signal comprising only speech. That is the first UE would want to receive a signal in which the background ambience is suppressed.
[0102] However, if the first UE uses a headset interface to output the received signal, it can handle all kinds of signals in terms of acoustic echo cancellation. Thus, when the first UE uses the headset as the playback interface, the first UE would prefer to receive a fully immersive signal comprising both speech and associated background ambience.
[0103] The second UE can have a flexible capability to suppress the background ambience completely and transmit signals comprising only speech or transmit fully immersive signal comprising both speech and background ambience. The first UE would want to receive either a speech signal with attenuated background ambience or a fully immersive signal with speech and background ambience, depending on its applied playback interface. Currently there is no mechanism to enable this.
[0104] A third problem could arise if the first UE can transmit either a fully immersive capture signal comprising both speech and complex ambient sound field or a suppressed signal comprising only speech without any background ambience. In some scenarios the first UE can be located in a sonically rich environment with a lot of background ambience at high level (relative to the speech level). If the first UE offers encoding at the lowest bitrates, for example, due to the network conditions, then the preference of the second UE would take the low bitrate into account. For example, due to the offered low encoding bitrate, the second UE would preferably want to receive only the speech signal, where all the background ambience is attenuated, as the overall quality and speech intelligibility might be decreased due to the combination of low bitrate encoding and presence of captured complex background ambience.In examples of the disclosure this problem of mismatch between the expected and the actual amount of background ambience is addressed by providing the UEs with a method to indicate their preferences on the amount of received ambience level and to indicate their capabilities in terms of providing varying amount of ambience or other different audio processing modes. These indications can be provided during a session setup. This can enable the expected amount of background ambience to match with the actual, transmitted amount of background ambience between the UEs. This can increase the perceived quality and intelligibility which can improve the experienced immersion.
[0105] Figs. 2A and 2B show example methods that can be used in examples of the disclosure to address this issue of the mismatch between the expected and the actual amount of background ambience. Fig. 2A shows an example method that can be implemented by an apparatus such as a media sender or any other suitable apparatus. Fig. 2B shows an example method that can be implemented by a corresponding apparatus such as media receiver or any other suitable apparatus. The respective apparatus could be a user equipment (UE) or any other device that enables immersive audio communication.
[0106] The respective apparatus can be for establishing an immersive audio communication session. The immersive audio communication session could be an IVAS session or any other suitable type of session.
[0107] The immersive audio communication session can be established with between a media sender and a media receiver. The media sender and the media receiver can be UEs or could be any other suitable type of apparatus.
[0108] Establishing the immersive audio communication session can comprise an offer-answer-acknowledge process. Any signals received after that process can be considered to be sent within the immersive communication session.
[0109] The audio processing mode that is used might change at the start of the session or at any other time during the session.
[0110] At block 200 the method comprises determining at least one audio processing mode that is supported by the apparatus for the immersive audio communication session. It is possible for an apparatus to support different audio processing modes in different communication sessions. In some examples the different audio processing modes can be different audio suppression modes that provide different levels of ambient or background noise. The different audio processing modes can comprise different audio suppression modes;different speech enhancement modes; modes with different levels of amplification; modes with different levels of compression; modes with different levels of equalization; modes with different levels of bass roll-off; modes with different levels of suppression; modes with different levels of noise suppression; modes with different levels of ambience suppression; and / or any other suitable audio processing modes.
[0111] In some examples when a first audio processing mode is used processed audio signals comprise speech or other desired sounds and ambient sound so that users hear both the speech and the ambient sound when this mode is used. A second audio processing mode can suppress the ambient sound to enhance the intelligibility of speech compared to the first audio processing mode. The apparatus may be able to support multiple different audio processing modes with different levels of suppression of the ambient sound. For example an audio processing mode can provide a medium amount of attenuation of the ambient sound and another audio processing mode can provide a higher amount of attenuation of the ambient sound.
[0112] The at least one audio processing mode that is determined at block 200 can be determined at least in part based on one of: a negotiated bitrate parameter; offered bitrate parameter; ambient noise in the environment of the apparatus; and output device used by the apparatus, and / or any other suitable factor.
[0113] At block 202 the apparatus transmits a list comprising the at least one audio processing mode. The list can be transmitted to another apparatus such as a media receiver.
[0114] The list can be transmitted using session description protocol (SDP) signalling or any other suitable signalling.
[0115] The list can be indicated using an audio processing mode parameter such as a cpt parameter or any other suitable means.
[0116] At block 204 the apparatus receives a response to the transmitted list. The response indicates at least one preferred audio processing mode for the immersive audio communication session. The at least one preferred audio processing mode is based on the transmitted list. That is, the at least one preferred audio processing mode can be selected from the transmitted list.
[0117] In some examples the received response can indicate multiple preferred audio processing modes. In such examples the received response can indicate an order of preference for the multiple preferred audio processing modes. The order of preference could be indicated by the order of the list or by any other suitable method.At block 206 the apparatus uses a preferred audio processing mode indicated in the response for the immersive audio communication session.
[0118] In some examples when a first audio processing mode is used processed audio signals comprise speech and ambient sound and when a second audio processing mode is used the ambient sound is suppressed in the processed audio signals. In some examples the respective audio processing modes affect the content of a received bitstream such that when a first audio processing mode is used the received bitstream comprises speech and ambient sound and when a second audio processing mode is used ambient sound is suppressed in the received bitstream to enhance the intelligibility of speech compared to the first audio processing mode.
[0119] The method shown in Fig. 2A shows how an apparatus such as a media sender can select an audio processing mode for use in the transmit direction. The apparatus can also follow an analogous method to select an audio processing mode to use in the receive direction. For example, the same apparatus that implements the method of Fig.2A can also receive a list comprising at least one audio processing mode that is supported by a media receiver and determine at least one preferred audio processing mode from the received list that is supported by the apparatus. The apparatus can then transmit a response indicating the determined at least one preferred audio processing mode of the apparatus. The preferred audio processing mode indicated in the sent response can then be used for the immersive audio communication session.
[0120] In some examples the immersive audio communication session can be established using the same audio processing mode in both directions. In other examples different audio processing modes can be used in different directions. That is, the immersive audio communication session can use a first audio processing mode for bitstreams sent from the media sender and a second audio processing mode for bitstreams received by the media sender.
[0121] Fig. 2B shows a corresponding method that can be implemented by an apparatus such as a media receiver. The apparatus that implements the method of Fig. 2B could be in a communication session with a corresponding apparatus that implements the method of Fig. 2B.
[0122] The immersive audio communication session can be established with a media receiver. In this case the media receiver could be another UE or could be any other suitable type of apparatus.At block 210 the apparatus receives a list comprising at least one audio processing mode that is supported for an immersive audio communication session. The list can be received from a corresponding apparatus such as a media sender implementing the method of Fig. 2A.
[0123] The list can be received using session description protocol (SDP) signalling or any other suitable signalling. In some examples the list can be indicated using an audio processing mode negotiation parameter.
[0124] The different audio processing modes comprise different audio suppression modes, different speech enhancement modes, modes with different levels of amplification, modes with different levels of compression, modes with different levels of equalization, modes with different levels of bass roll-off, modes with different levels of suppression, modes with different levels of noise suppression, modes with different levels of ambience suppression, or any other suitable audio processing modes.
[0125] The apparatus can support multiple different audio processing modes with different levels of suppression of the ambient sound.
[0126] The apparatus determines, at block 212, at least one preferred audio processing mode from the received list for the immersive audio communication session.
[0127] The preferred at least one audio processing mode can be determined based on any suitable factors such as, ambient noise in the environment of the apparatus, capture device used by the apparatus, output device used by the apparatus, available bitrate levels, or any other suitable factors.
[0128] The apparatus transmits, at block 214, a response indicating a determined at least one preferred audio processing mode of the apparatus. In some examples the transmitted response can indicate an order of preference for the at least one audio processing mode.
[0129] At block 216 the apparatus uses a preferred audio processing mode indicated in the response for the immersive audio communication session.
[0130] In some examples when a first audio processing mode is used processed audio signals comprise speech and ambient sound and when a second audio processing mode is used the ambient sound is suppressed in the processed audio signals. In some examples the respective audio processing modes affect the content of a received bitstream such that when a first audio processing mode is used the received bitstream comprisesspeech and ambient sound and when a second audio processing mode is used ambient sound is suppressed in the received bitstream to enhance the intelligibility of speech compared to the first audio processing mode.
[0131] In some examples the apparatus can determine, at any point during the immersive audio communication session, that at least one of the factors used to determine the audio processing mode has changed. When this is determined the apparatus can transmit a request to change the audio processing mode used for the immersive audio communication session.
[0132] The method shown in Fig. 2B shows how an apparatus such as a media receiver can select an audio processing mode for use in the receiving direction. The apparatus can also follow an analogous method to select an audio processing mode to use in the transmit direction. For example, the same apparatus that implements the method of Fig.2B can also determine at least one audio processing mode that is supported by the apparatus; transmit a list comprising at least one audio processing mode; receive a response indicating at least one preferred audio processing mode wherein the at least one preferred audio processing mode is based on the transmitted list; and use a preferred audio processing mode indicated in the received response for the immersive audio communication session.
[0133] In some examples the immersive audio communication session can be established using the same audio processing mode in both directions. In other examples different audio processing modes can be used in different directions. That is, the immersive audio communication session can use a first audio processing mode for bitstreams sent from the media sender and a second audio processing mode for bitstreams received by the media sender.
[0134] Fig. 3 shows an example set up for an immersive audio communication session. The setup comprises a first UE 300A and a second UE 300B. The first UE 300A could be an apparatus that can implement the method of Fig.2A and the second UE 300B could be an apparatus that can implement the method of Fig. 2B The UEs 300A, 300B can be any IVAS devices. The UEs 300A, 300B can be implementing DaCAS capture or any other suitable spatial audio capture. The UEs 300A, 300B can have one or more audio processing modes available. The audio processing modes can be audio suppression modes or any other suitable type of audio processing modes.
[0135] Each of the UEs 300A, 300B can support at least one audio processing mode but possibly more than one audio processing mode. The UEs 300A, 300B can support audio processing modes in both the transmit and receive directions.The respective UEs 300A, 300B can support the same audio processing modes or in some examples the UEs 300A, 300B can support at least some different audio processing modes.
[0136] The respective audio processing modes can enable the UEs 300A, 300B to apply different signal processing techniques to the captured and transmitted audio signals. An example signal processing technique that can be applied by the UEs 300A, 300B is suppression of different audio signal content, such as speech, background ambience, noise, or any other suitable type of audio signal content. The audio signal content that is suppressed is typically background ambience and / or noise but other types of audio content could be suppressed in other examples. Furthermore, the applied amount of suppression to audio signal content can be varied or dynamically adjusted. Different levels or variations of applied audio suppression can be associated with different audio suppression modes.
[0137] In the example shown in Fig. 3 the first UE 300A supports multiple audio processing modes in the transmit direction. This can enable the first UE 300A to offer different degrees of attenuation of a particular type of audio content such as background ambience and / or noise to be performed for the captured audio signal.
[0138] The first UE 300A can also support multiple audio processing modes in the receive direction. This can enable the first UE 300A to receive and output audio signals with different degrees of audio processing. In this example this can enable the first UE 300A to receive and output audio signals with different degrees of attenuation of a particular type of audio content such as background ambience and / or noise.
[0139] In some examples the UE 300A can support the same audio processing modes in both the transmit direction and the receive direction. In other examples the UE 300A can support different sets of audio processing modes in the different directions. For example the set of audio processing modes that is supported in the transmit direction can comprise one or more audio processing modes that is not supported in the receive direction and / or the set of audio processing modes that is supported in the receive direction can comprise one or more audio processing modes that is not supported in the transmit direction.
[0140] In the example of Fig. 3 the first UE 300A can support a transparent mode and a suppressed mode for the immersive audio communication session. In the transparent mode all types of audio content are captured and transmitted. The undesired type of audio content is not attenuated or suppressed. There would be no suppression, or very little suppression, for any captured and transmitted signal component (for example speech, background ambience, noise, or any other suitable components). In the suppressed mode theundesired audio content is targeted to be suppressed and removed as much as possible from the captured and transmitted signal. If a suppressed transmission mode is decided, the receiver expects that the audio stream comprises only, or predominantly, the desired content. The desired content could be speech or any other suitable type of content.
[0141] In the example of Fig. 3 the first UE 300A determines which audio processing modes are supported and preferred to be used by the first UE 300A for the immersive audio communication session and transmits, at block 302, an offer to the second UE 300B. The offer can comprise a list comprising at least one audio processing mode that is supported by the first UE 300A in the transmit direction and also the receive direction. The list can be indicated using a cp t parameter or by using any other suitable means. In this case the offer can indicate that the first UE 300A only supports the suppressed mode in both the transmit direction and the receive direction.
[0142] The second UE 300B receives the offer. At block 304 the second UE 300B determines at least one preferred audio processing mode indicated in the offer that is also supported by the second UE 300B for the immersive audio communication session. The second UE 300B can determine at least one preferred audio processing mode that is supported by the second UE 300B in both a transmit direction and a receive direction.
[0143] At block 306 the second UE 300B transmits a response to the offer. The response indicates at least one preferred audio processing mode of the second UE 300B that was determined at block 304. In this case the answer can indicate that the second UE 300B also supports the suppressed mode in both the transmit direction and the receive direction.
[0144] At block 308 the immersive communication session is established. As each of the UEs 300A, 300B has indicated that a suppressed mode is supported the audio stream would be a suppressed audio stream comprising predominantly speech in each direction.
[0145] In the example of Fig. 3 the first UE 300A has requested a suppressed mode. An analogous method would be used if the first UE 300A had requested a transparent mode or any other audio processing mode.
[0146] In the example of Fig. 3 the first UE 300A is a media sender and the second UE 300B is a media receiver. The UEs 300A, 300B can function in both a transmit mode and a receive mode so that the operations can be reversed. That is, the second UE3 300B can also transmit an offer to the first UE 300A indicating theavailable audio processing modes and the first UE 300A can then provide a reply selecting an audio processing mode from the available audio processing modes.
[0147] The example of Fig. 3 shows a basic scenario in which the same audio processing modes can be used in both the transmit and the receive directions. In other examples a first audio processing mode could be used in the transmit direction and a second, different audio processing mode could be used in the receive direction.
[0148] To enable different audio processing modes to be used in different directions different parameters can be used to indicate the audio processing modes for the respective directions. The different parameters can enable the audio processing modes for each direction to be negotiated independently. For instance, a cpt-send parameter can be used to indicate the audio processing modes indicated in the transmit direction and a cpt-recv parameter can be used to indicate the audio processing modes indicated in the receive direction. In some examples the respective parameters can indicate different sets of audio processing modes. The sets can be overlapping so that at least one audio processing mode is supported in both the transmit direction and the receive direction. In some examples the sets can be non-overlapping so that different audio processing modes are supported in the transmit direction compared to the receive direction. In some examples the respective parameters can indicate the same sets of audio processing modes so that if an audio processing mode is supported in the transmit direction it will be supported in the receive direction. In such cases the cpt parameter described above could be used.
[0149] The use of the different parameters can enable the respective UEs 300A, 300B to update their list of supported audio processing modes in relation to the offered modes for each direction. For example a UE 300A, 300B can update its list of supported audio processing modes in the receive direction based on a received offer relating to the available audio processing modes in the receive direction.
[0150] Fig. 4 shows an example method that can be used to enable different audio processing modes to be used in different directions in some examples of the disclosure. In this example the first UE 300A is setting up an immersive audio communication session. The immersive audio communication session can be an IVAS session or any other suitable type of communication session.
[0151] The first UE 300A can determine the audio processing modes that are supported by the first UE 300A for the immersive audio communication session. In this case the first UE 300A can only support a suppressed mode in the transmit direction. This could be due to any suitable reason such as implementation, playbackinterface, environmental conditions, or any other relevant reason. In this case suppressed mode might mean that the first UE 300A only transmits a speech signal with the other audio signal components attenuated.
[0152] However, in this example the system is asymmetric and the first UE 300A can support both the suppressed mode and the transparent mode in the receive direction. This means that the first UE 300A can receive an audio stream complying with either of the audio processing modes. The second UE 300B can support the transparent mode and the suppressed mode in both the transmit direction and the receive direction.
[0153] In the method of Fig. 4, at block 400, the first UE 300A transmits an offer to the second UE 300B. The offer indicates the audio processing modes that are supported by the first UE 300A in both the transmit direction and the receive direction. In this case the cpt- send parameter indicates that only the suppressed mode is supported in the transmit direction and the cpt-recv parameter indicates that both the transparent mode and the suppressed mode are supported in the receive direction.
[0154] At block 402 the second UE 300B can determine that both the transparent mode and the suppressed mode are supported in the transmit direction and also the receive direction for the immersive audio communication session. This can be indicated by the cpt parameter comprising a list of both the suppressed mode and the transparent mode.
[0155] At block 404 the second UE 300B receives the offer from the first UE 300A. As the cpt- send parameter in the offer indicates that the first UE 300A only supports the suppressed mode the second UE 300B will update its list associated with receive direction to only include modes that are also supported by the first UE 300A. In this case the second UE 300B will update its list associated with receive direction to only include the suppressed mode so as to enable the communication session to be established.
[0156] The list associated with receive direction would not need to be changed because the cpt-recv parameter from the first UE 300A indicates that both the transparent mode and the suppressed mode are supported in the receive direction.
[0157] At block 406 the second UE 300B transmits an answer to the first UE 300A. The answer comprises a cpt-recv parameter and a cpt- send parameter. The cpt-recv parameter only lists the suppressed mode but the cpt- send parameter lists both the transparent and suppressed modes.Fig. 5 shows an example set up for an immersive audio communication session. The setup comprises a first UE 300A and a second UE 300B similar to those shown in Fig. 3. The set up shown in Fig. 5 could be used to implement the scenario shown in Fig. 4 where the first UE 300A only supports the suppressed mode in the transmit direction but supports both the suppressed and transparent modes in the receive direction for the immersive audio communication session. The second UE 300B supports the suppressed mode and the transmit mode in both the transmit and receive direction for the immersive audio communication session.
[0158] In the example of Fig. 5 the first UE 300A determines which audio processing modes are supported by the first UE 300A for the respective directions for the immersive audio communication session and transmits, at block 500, an offer to the second UE 300B. The offer can comprise a list indicating the audio processing modes that are supported in the transmit direction and a list indicating the audio processing modes that are supported in the receive direction. In this case the audio processing modes that are supported in the transmit direction are indicated using a cpt- send parameter and the audio processing modes that are supported in the receive direction are indicated using a cpt-recv parameter. The cpt- send parameter in the offer indicates that the first UE300A can support the suppressed mode in the transmit direction. The cpt-recv parameter in the offer indicates that the first UE300A can support the suppressed mode and the transparent in the receive direction.
[0159] The second UE 300B receives the offer. At block 502 the second UE 300B determines at least one preferred audio processing mode for the transmit direction and at least one preferred audio processing mode for the receive direction that are indicated in the offer and also supported by the second UE 300B for the immersive audio communication session. In this case the second UE 300B has no restrictions and can determine that it can support the suppressed mode in the receive direction and that it can support both the suppressed mode and the transparent mode in the transmit direction.
[0160] At block 504 the second UE 300B transmits a response to the offer. The response indicates the preferred audio processing modes of the second UE 300B that were determined at block 502. In this case the answer can indicate that the second UE 300B can support the suppressed mode in the receive direction and that it can support both the suppressed mode and the transparent mode in the transmit direction. The cpt- send parameter in the answer indicates that the second UE300B can support the suppressed mode and the transparent mode in the transmit direction. The cpt-recv parameter in the answer indicates that the second UE300B can support the suppressed mode in the receive direction.At block 506 the immersive audio communication session is established. In this case a suppressed audio stream comprising predominantly speech would be sent from the first UE 300A to the second UE 300B. The audio stream that is transmitted from the second UE 300B to the first UE 300A could be a suppressed audio stream or a transparent audio stream.
[0161] In the example of Fig. 5 the second UE 300B does not have any restrictions on the audio processing mode for the immersive audio communication session. In other examples the second UE 300B could have restrictions. For example, it could be restricted so that it can only support the suppressed mode in the transmit direction. In such cases this would be determined at block 502 and the answer that is sent at block 504 would reflect the restrictions. That is, the cpt- send parameter in the answer from the second UE 300B would only list the suppressed mode.
[0162] Where a UE 300A, 300B can support more than one audio processing mode the preference of the supported modes can be indicated using any suitable means. In some examples the order in which the audio processing modes are listed can indicate the order of preference. In such cases, applied to the example of Fig. 5 if the transparent mode is listed first in the cpt- send parameter in the answer from the second UE 300B then this would be the preferred mode. The communication session could be established using transparent audio but could change to suppressed mode during the session.
[0163] In the examples described in Figs. 3 to 5 there are two different audio processing modes available. In other examples there can be more than two audio processing modes available. For example, in addition to the aforementioned suppressed and transparent modes, one or more additional audio processing mode(s) could be allowed and supported. Such audio processing modes could provide different levels of suppression or attenuation or could suppress or attenuate different parts of the audio or could provide different types of audio processing. For instance, an additional audio processing mode could indicate significantly less attenuation of the undesired signal component compared to the suppressed mode, while still providing significantly more attenuation than transparent mode.
[0164] In some examples the signaled audio processing mode associated with suppressive algorithms could target to suppress components of the audio signal other than undesired audio components. For example, a signaled parameter could indicate how much speech (or other desired audio component) is attenuated. That is a transparent mode would indicate that the speech component (or other desired audio component) is not attenuated while suppressed mode would indicate that speech component (or other desired audiocomponent) is maximally attenuated. In some examples the target audio signal component could be the whole audio signal that is, it does not need to be a certain component of the audio signal that is targeted.
[0165] In some examples the signaled audio suppression mode associated with suppressive algorithms could target to suppress the background noise to some predefined minimum signal to noise ratio (SNR). For example, in low noise environments no suppression is used, in medium noise scenarios some amount of suppression is applied to achieve the requested SNR. In high noise environments maximum noise suppression is applied in order to be as close as possible with the requested SNR.
[0166] In some embodiments the signaled audio suppression mode could include an indication on which audio component is enhanced or unmodified. For example, an indication of desired type of audio content that should not be suppressed could be included to the signaled audio suppression mode or level.
[0167] In some embodiments the audio processing modes can be associated with a signal processing technique other than noise suppression. For example, the negotiated capture mode could indicate the supported amount of enhancement, amplification, compression, equalization, bass roll-off or any other suitable processing that is applied to a targeted type of audio content.
[0168] In some embodiments a UE associated with the negotiation of the communication session can limit or modify the list of offered and answered audio processing modes based on relevant factors or circumstance such as the negotiated / offered bitrate parameter (ibr , ibr-send, ibr-recv). For example, a UE might support only a suppressed mode, when low bitrate encoding is offered, while at mid to high bitrates all possible audio suppression modes might be supported.
[0169] In some examples a UE can determine the list of supported audio processing modes to include in an offer or answer in dependence on the negotiated / offered coded format parameter (cf , cf-send, cf-recv). For example, when only ISM coded format is offered, a UE can indicate that it only supports the suppressed mode. Alternatively, when MASA coded format is offered, a UE can indicate that it only supports the transparent mode. Furthermore, the support of both the suppressed and transparent capture modes could be offered when OMASA (combined ISM and MASA) coded format is offered. The combinations of offered coded formats and supported audio processing modes described here are merely examples and implementations of the disclosure are not limited to the presented combinations.In the case of OMASA with multiple ISM objects, separate ISM can comprise differently processed object audio. For example, the first ISM could comprise signal processing optimized, or substantially optimized, for speech intelligibility (suppressed, compressed, equalized, amplified) and the second ISM could contain processing optimized, or substantially optimized, for naturalness (no suppression, full dynamics, no equalization).
[0170] In some examples the order of the offered audio processing modes can indicate the UE’s preference on which audio processing mode it would prefer to use. In such examples, when a cpt parameter comprises transparent mode and suppressed mode in that order, then the transparent mode would be determined to be the preferred audio processing mode. The suppressed mode would be determined to be supported but not preferred. In some other examples the order of the listed audio processing modes does not indicate preference.
[0171] In some examples, the audio processing modes could have different names, such as suppression mode, transparency mode, speech suppression level, suppression level, target SNR, natural mode, speech enhanced or something similar. The related SDP parameter cpt could also have a different name such as, capture-processing-mode, cp-mode, suppression-level, supp-level, supp-mode, suppression-mode, sm, si, audio-processing-mode, apm, suppression-processing-mode, spm, cpm or any suitable name.
[0172] Fig. 6 shows an example method that can be used in some examples of the disclosure. This shows a method between a first UE 300A and a second UE 300B. In this case the first UE 300A can support the transparent mode and the suppressed mode in both the transmit and the receive directions for the immersive audio communication session. The second UE 300B can only support the suppressed audio mode in the transmit and the receive directions for the immersive audio communication session.
[0173] At block 600 the first UE 300A makes an SDP offer the second UE 300B. In the offer the first UE 300A indicates the capability of the first UE 300A to support the transparent mode and the suppressed mode in both the transmit and the receive directions.
[0174] At block 602 the second UE 300B provides an answer to the SDP offer. The second UE 300B is only capable to support the suppressed mode and so in the answer, the second UE 300B only indicates the suppressed mode.At block 604 the first UE 300A accepts the answer from the second UE 300B and a session is established between the first UE 300A and the second UE 300B. The suppressed mode is used in both the transmit and receive directions in accordance with the answer from the second UE 300B.
[0175] After the session has been established, RTP traffic can begin between the first UE 300A and the second UE 300B at block 606.
[0176] Table 1 shows an example negotiation for an implementation of the disclosure. In this example the first UE 300A offers cpt mode suppressed for both transmit and receive directions for the communication session. In the answer, the second UE 300B accepts the cpt mode suppressed for the session.
[0177] SDP Offer (from UE A to UE B)
[0178] m=audio 49152 RTP / AVP 96
[0179] a=rtpmap : 96 IVAS / 16000
[0180] a=fmtp : 96 cpt=Suppressed
[0181] SDP Answer (from UE B to UE A)
[0182] m=audio 49152 RTP / AVP 96
[0183] a=rtpmap : 96 IVAS / 16000
[0184] a=fmtp : 96 cpt=Suppressed
[0185]
[0186] Table 1
[0187] Table 2 shows another example negotiation for an implementation of the disclosure. In this example the first UE offers cpt transparent mode and suppressed mode for both transmit and receive directions. The audio processing modes are listed in preference order so that transparent mode is preferred over suppressed mode. The second UE 300B makes a decision to only support the suppressed mode in the transmit direction and the transparent mode in the receive direction and indicates this in the session answer. The decision that is made by the second UE 300B can be based on the capabilities of the second UE 300B, the capturing or playback environment of the second UE 300B or any other relevant factors.
[0188] SDP Offer (from UE A to UE B)
[0189] m=audio 49152 RTP / AVP 96
[0190] a=rtpmap : 96 IVAS / 16000
[0191] a=fmtp : 96 cpt=Transparent , Suppressed
[0192] SDP Answer (from UE B to UE A)
[0193]
[0194] m=audio 49152 RTP / AVP 96
[0195] a=rtpmap : 96 IVAS / 16000
[0196] a=fmtp : 96 cpt-send=Suppressed; cpt-recv=Transparent
[0197]
[0198] Table 2
[0199] Table 3 shows another example negotiation for an implementation of the disclosure. In this example the first UE 300A offers the suppressed mode in the transmit direction and the transparent mode in the receive direction. The second UE 300B echoes the offered audio processing modes in the answer. That is, the second UE 300B indicates that the transparent mode can be used in the transmit direction of the second UE 300B and that the suppressed mode can be used in the receive direction of the second UE 300B.
[0200] SDP Offer (from UE A to UE B)
[0201] m=audio 49152 RTP / AVP 96
[0202] a=rtpmap : 96 IVAS / 16000
[0203] a=fmtp : 96 cpt-send=Suppressed; cpt-recv=Transparent
[0204] SDP Answer (from UE B to UE A)
[0205] m=audio 49152 RTP / AVP 96
[0206] a=rtpmap : 96 IVAS / 16000
[0207] a=fmtp : 96 cpt-send=Transparent ; cpt-recv=Suppressed
[0208]
[0209] Table 3
[0210] Examples of the disclosure can reduce a mismatch between the expected and the actual amount of background ambience by providing a mechanism for negotiating which audio processing modes are used for the immersive audio communication session. As an example, a first UE 300A and a second UE 300B could be establishing an immersive audio session (such as an IVAS call) between the UEs 300A, 300B. The first UE 300A is located in an environment comprising lots of background ambience, for example, in a rock concert or in a sports event. The first UE 300A is able to transmit full immersive capture comprising speech and background ambience sounds. The first UE 300A is also able to transmit only the suppressed capture comprising the speech of the user of the first UE 300A.
[0211] The second UE 300B is located in a noisy environment, for example, in a crowd or walking next to city traffic. The second UE 300B prefers to receive only the speech signal from the first UE 300A without the additional ambience sounds in order to hear better what the user of the first UE 300A is saying.In the initial session offer from first UE 300A to second UE 300B, the first UE 300A is not aware of the environment or preferences of the second UE 300B. Without the examples of the disclosure the first UE 300A is not able to indicate different suppression modes or levels in the session offer. Therefore, without the examples of the disclosure the first UE 300A is forced to offer only fully immersive capture to the second UE 300B. Furthermore, without the examples of the disclosure the second UE 300B is not able to indicate a preference to receive only speech signals from the first UE 300A. Consequently, a fully immersive session between the UEs 300A, 300B would be established, where the ambient background sounds from the first UE 300A are transmitted to the second UE 300B in addition to the speech signal of the user of first UE 300A. With the ambient background sounds present in the transmitted audio, the speech intelligibility and experienced quality by the user of the second UE 300B are decreased.
[0212] With the examples of the disclosure, the first UE 300A is able to indicate different audio processing modes in the initial session offer. The first UE 300A can offer, for example, fully immersive and suppressed capture modes. Also, the second UE 300B is able to indicate their preference (suppressed capture) in the session answer. Consequently, suppressed audio is transmitted from the first UE 300A to the second UE 300B during the session. The background ambience sounds from the capture of the first UE 300A are suppressed from the transmitted audio and it is easier for the user of the second UE 300B to focus on the speech of the user of the first UE 300A. The speech intelligibility and overall quality of experience are increased for the user of the second UE 300B compared to the situation without the examples of the disclosure where there is no ability to negotiate a suppressed session.
[0213] The results of inspections of the captured and transmitted signals by a media sender with different negotiated audio processing modes are shown in Figs 7A to 10. In this case the media sender is a first UE 300A and the signals that are inspected would be the signals received by a media receiver (the second UE 300B).
[0214] Fig. 7A shows the reference speech signal that is captured and transmitted by the first UE 300A. The upper plot shows the reference speech signal in the time domain and the lower plot shows the spectrogram of the reference speech signal.
[0215] Fig. 7B shows the reference background noise signal that is captured and transmitted by the first UE 300A. The upper plot shows the reference background noise signal in the time domain and the lower plot shows the spectrogram of the reference background noise signal.
[0216] In this example the first UE 300A supports both transparent and suppressed audio processing modes.Fig. 8 shows a signal that is transmitted by the first UE 300A in the transparent mode. The upper plot shows the signal in the time domain and the lower plot shows the spectrogram of the signal.
[0217] It can be seen in Fig. 8 that the signal comprises both speech and background noise, due to the negotiated audio suppression mode. Similar patterns can be seen to those in the reference speech signal. Background noise is also visually present in the transmitted signal. Furthermore, by analyzing the short-time objective intelligibility (STOI) metric against the reference signal, we obtain a value of:
[0218] u ST CH — n q
[0219]
[0220] i v / i Transparent o
[0221] Which indicates that the speech is somewhat understandable (STOI range is from -1 to 1, where -1 is the worst intelligibility and 1 is the highest intelligibility). The transmitted signal is not optimal in terms of speech intelligibility, however, if the second UE 300B is in a quiet environment, or uses a headset, the intelligibility can be considered to be acceptable. Furthermore, with this negotiated audio processing mode, the signal transmitted by the first UE 300A also comprises full immersive background, which can be preferred over optimal speech intelligibility. With the use of examples of the disclosure, the second UE 300B can decide to make a trade-off between the speech intelligibility in order to receive immersive background in suitable situations. The examples of the disclosure provide a mechanism that enables the second UE 300B to signal to the first UE 300A when it is able to receive such signals.
[0222] Fig. 9 shows a signal that is transmitted by the first UE 300A in the suppressed mode. For example, a suppressed mode could be negotiated between the first UE 300A and the second UE 300B. The upper plot shows the signal in the time domain and the lower plot shows the spectrogram of the signal.
[0223] It can be seen in Fig. 9 that the transmitted signal closely represents the reference speech signal as shown in Fig. 7A. Fig. 9 shows that most of the background noise can be removed from the captured signal. Furthermore, the analyzed STOI against the reference signal is:
[0224] STOISuppressed0.89
[0225] The obtained value indicates that the intelligibility of the transmitted signal is excellent. The second UE 300B could be in a noisy environment, or could be using integrated loudspeaker output, or otherwise could prefer to receive maximal speech intelligibility with the cost of lost immersiveness. In such scenarios the examples of the disclosure allow the second UE 300B to negotiate with the first UE 300A that it would prefer to receive audio according to the suppressed mode. This indication can be made at the beginning of the session initiation.Fig. 10 shows the impact of an audio processing mode on speech intelligibility. Due to internal or external reasons, a UE may want to negotiate a different audio processing mode in order to obtain optimal combination of speech intelligibility and immersiveness. As with different audio suppression modes, the perceived immersiveness typically decreases along the increased amount of suppression.
[0226] Fig. 11 shows an example controller 1100. The controller 1100 could be provided within an apparatus such as a media sender or a media receiver, or any other suitable entity that enables immersive communication sessions. Implementation of the controller 1100 may be as controller circuitry. The controller 1100 may be implemented in hardware alone, have certain aspects in software including firmware alone or can be a combination of hardware and software (including firmware). The controller 1100 can provide an apparatus for implementing the disclosure of could be provided as part of an apparatus that implements the disclosure.
[0227] As illustrated in Fig. 11 the controller 1100 can be implemented using instructions that enable hardware functionality, for example, by using executable instructions of a computer program 1106 in a general-purpose or special-purpose processor 1102 that may be stored on a machine readable storage medium (disk, memory etc.) to be executed by such a processor 1102.
[0228] The processor 1102 is configured to read from and write to the memory 1104. The processor 1102 may also comprise an output interface via which data and / or commands are output by the processor 1102 and an input interface via which data and / or commands are input to the processor 1102.
[0229] The memory 1104 stores instructions, program 1106, or code that controls the operation of the apparatus when loaded into the processor 1102. The instructions, program 1106, or code provide the logic and routines that enables the apparatus to perform the methods illustrated in the Figs. The processor 1102 by reading the memory 1104 is able to load and execute the instructions, program 1106, or code.
[0230] In some examples where the controller 1100 is provided within an apparatus such as a media sender, the controller therefore comprises means for:
[0231] determining 200 at least one audio processing mode that is supported by the apparatus; transmitting 202 a list comprising the at least one audio processing mode;
[0232] receiving 204 a response indicating at least one preferred audio processing mode wherein the at least one preferred audio processing mode is based on the transmitted list; andusing 206 a preferred audio processing mode indicated in the response for an immersive audio communication session.
[0233] In some examples where the controller 1100 is provided within an apparatus such as a media receiver, the controller therefore comprises means for:
[0234] receiving 210 a list comprising at least one audio processing mode that is supported for an immersive audio communication session;
[0235] determining 212 at least one preferred audio processing mode from the received list; transmitting 214 a response indicating a determined at least one preferred audio processing mode of the apparatus; and
[0236] using 216 a preferred audio processing mode indicated in the response for the immersive audio communication session.
[0237] The instructions, program 1106, or code may arrive at the apparatus via any suitable delivery mechanism 1108. The delivery mechanism 1108 may be, for example, a machine-readable medium, a computer-readable medium, a non-transitory computer-readable storage medium, a computer program product, a memory device, a record medium such as a Compact Disc Read-Only Memory (CD-ROM) or a Digital Versatile Disc (DVD) or a solid-state memory, an article of manufacture that comprises or tangibly embodies the computer program 1106. The delivery mechanism may be a signal configured to reliably transfer the computer program 1106. The apparatus may propagate or transmit the computer program 1106 as a computer data signal.
[0238] The term “non-transitory” as used herein, is a limitation of the medium itself that is, tangible, not a signal ) as opposed to a limitation on data storage persistency (for example, RAM vs. ROM).
[0239] The computer program 1106 can comprise computer program instructions for causing an apparatus to perform at least the following or for performing at least the following:
[0240] determining 200 at least one audio processing mode that is supported by the apparatus; transmitting 202 a list comprising the at least one audio processing mode;
[0241] receiving 204 a response indicating at least one preferred audio processing mode wherein the at least one preferred audio processing mode is based on the transmitted list; and
[0242] using 206 a preferred audio processing mode indicated in the response for an immersive audio communication session.The computer program 1106 can comprise computer program instructions for causing an apparatus to perform at least the following or for performing at least the following:
[0243] receiving 210 a list comprising at least one audio processing mode that is supported for an immersive audio communication session;
[0244] determining 212 at least one preferred audio processing mode from the received list; transmitting 214 a response indicating a determined at least one preferred audio processing mode of the apparatus; and
[0245] using 216 a preferred audio processing mode indicated in the response for the immersive audio communication session.
[0246] The computer program instructions may be comprised in a computer program, a non-transitory computer readable medium, a computer program product, a machine-readable medium. In some but not necessarily all examples, the computer program instructions may be distributed over more than one computer program.
[0247] Although the memory 1104 is illustrated as a single component / circuitry it may be implemented as one or more separate components / circuitry some or all of which may be integrated / removable and / or may provide permanent / semi-permanent / dynamic / cached storage.
[0248] Although the processor 1102 is illustrated as a single component / circuitry it may be implemented as one or more separate components / circuitry some or all of which may be integrated / removable. The processor 1102 may be a single core or multi-core processor.
[0249] References to “computer-readable storage medium”, “computer program product”, “tangibly embodied computer program” etc. or a “controller”, “computer”, “processor” etc. should be understood to encompass not only computers having different architectures such as single / multi- processor architectures and sequential (Von Neumann) / parallel architectures but also specialized circuits such as field-programmable gate arrays (FPGA), application specific circuits (ASIC), signal processing devices and other processing circuitry including quantum processing circuitry. References to computer program, instructions, code etc. should be understood to encompass software for a programmable processor or firmware such as, for example, the programmable content of a hardware device whether instructions for a processor, or configuration settings for a fixed-function device, gate array or programmable logic device etc.
[0250] As used in this application, the term “circuitry” can refer to one or more or all of the following:(a) hardware-only circuitry implementations (such as implementations in only analog and / or digital circuitry) and
[0251] (b) combinations of hardware circuits and software, such as (as applicable):
[0252] (i) a combination of analog, digital and / or quantum hardware circuit(s) with software / firmware and
[0253] (ii) any or all portions of hardware processor(s) (including digital and / or quantum processor(s)) with software, and memory(ies) that work together to cause an apparatus, such as a mobile device, computing device or server, to perform various functions and
[0254] (c) any or all portions of hardware circuit(s) such as a microprocessor(s) and / or quantum processor(s) that requires software (for example, firmware) for operation, but the software may not be present when it is not needed for operation.
[0255] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in a server, a cellular network device, or other computing or network device.
[0256] The blocks illustrated in the Figs, can represent steps in a method and / or sections of code in the computer program 1106. The illustration of a particular order to the blocks does not necessarily imply that there is a required or preferred order for the blocks and the order and arrangement of the block can be varied. Furthermore, it can be possible for some blocks to be omitted.
[0257] Where a structural feature has been described, it may be replaced by means for performing one or more of the functions of the structural feature whether that function or those functions are explicitly or implicitly described.
[0258] The apparatus can be provided in an electronic device, for example, a mobile terminal, according to an example of the present disclosure. It should be understood, however, that a mobile terminal is merely illustrative of an electronic device that would benefit from examples of implementations of the present disclosure and, therefore, should not be taken to limit the scope of the present disclosure to the same. While in certain implementation examples, the apparatus can be provided in a mobile terminal, other types of electronic devices, such as, but not limited to: mobile communication devices, hand portable electronicdevices, wearable computing devices, portable digital assistants (PDAs), pagers, mobile computers, desktop computers, televisions, gaming devices, laptop computers, cameras, video recorders, GPS devices and other types of electronic systems, can readily employ examples of the present disclosure. Furthermore, devices can readily employ examples of the present disclosure regardless of their intent to provide mobility.
[0259] The term ‘comprise’ is used in this document with an inclusive not an exclusive meaning. That is any reference to X comprising Y indicates that X may comprise only one Y or may comprise more than one Y. If it is intended to use ‘comprise’ with an exclusive meaning then it will be made clear in the context by referring to ‘comprising only one...’ or by using ‘consisting.’
[0260] In this description, the wording ‘connect’, ‘couple’ and ‘communication’ and their derivatives mean operationally connected / coupled / in communication. It should be appreciated that any number or combination of intervening components can exist (including no intervening components), i.e. , to provide direct or indirect connection / coupling / communication. Any such intervening components can include hardware and / or software components.
[0261] As used herein, the term "determine / determining" (and grammatical variants thereof) can include, not least: calculating, computing, processing, deriving, measuring, investigating, identifying, looking up (for example, looking up in a table, a database, or another data structure), ascertaining and the like. Also, "determining" can include receiving (for example, receiving information), accessing (for example, accessing data in a memory), obtaining and the like. Also, " determine / determining" can include resolving, selecting, choosing, establishing, and the like.
[0262] In this description, reference has been made to various examples. The description of features or functions in relation to an example indicates that those features or functions are present in that example. The use of the term ‘example’ or ‘for example’ or ‘can’ or ‘may’ in the text denotes, whether explicitly stated or not, that such features or functions are present in at least the described example, whether described as an example or not, and that they can be, but are not necessarily, present in some of or all other examples. Thus ‘example’, ‘for example’, ‘can’, or ‘may’ refers to a particular instance in a class of examples. A property of the instance can be a property of only that instance or a property of the class or a property of a sub-class of the class that includes some but not all the instances in the class. It is therefore implicitly disclosed that a feature described with reference to one example but not with reference to another example, can where possible be used in that other example as part of a working combination but does not necessarily have to be used in that other example.As used herein, “at least one of the following: ” and “at least one of ” and similar wording, where the list of two or more elements are joined by “and” or “or” mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements.
[0263] Although examples have been described in the preceding paragraphs with reference to various examples, it should be appreciated that modifications to the examples given can be made without departing from the scope of the claims.
[0264] Features described in the preceding description may be used in combinations other than the combinations explicitly described above.
[0265] Although functions have been described with reference to certain features, those functions may be performable by other features whether described or not.
[0266] The description of a feature, such as an apparatus or a component of an apparatus, configured to perform a function, or for performing a function, should additionally be considered to also disclose a method of performing that function. For example, description of an apparatus configured to perform one or more actions, or for performing one or more actions, should additionally be considered to disclose a method of performing those one or more actions with or without the apparatus.
[0267] Although features have been described with reference to certain examples, those features may also be present in other examples whether described or not.
[0268] The term ‘a’, ‘an’ or ‘the’ is used in this document with an inclusive not an exclusive meaning. That is any reference to X comprising a / an / the Y indicates that X may comprise only one Y or may comprise more than one Y unless the context clearly indicates the contrary. If it is intended to use ‘a’, ‘an’ or ‘the’ with an exclusive meaning then it will be made clear in the context. In some circumstances the use of ‘at least one’ or ‘one or more’ may be used to emphasis an inclusive meaning but the absence of these terms should not be taken to infer any exclusive meaning.
[0269] The presence of a feature (or combination of features) in a claim is a reference to that feature or (combination of features) itself and to features that achieve substantially the same technical effect (equivalent features). The equivalent features include, for example, features that are variants and achieve substantially the sameresult in substantially the same way. The equivalent features include, for example, features that perform substantially the same function, in substantially the same way to achieve substantially the same result.
[0270] In this description, reference has been made to various examples using adjectives or adjectival phrases to describe characteristics of the examples. Such a description of a characteristic in relation to an example indicates that the characteristic is present in some examples exactly as described and is present in other examples substantially as described.
[0271] As used herein, the terms “the at least one” and “the one or more” mean “any one of the at least one” and “any one of the one or mor” respectively.
[0272] The above description describes some examples of the present disclosure however those of ordinary skill in the art will be aware of possible alternative structures and method features which offer equivalent functionality to the specific examples of such structures and features described herein above and which for the sake of brevity and clarity have been omitted from the above description. Nonetheless, the above description should be read as implicitly including reference to such alternative structures and method features which provide equivalent functionality unless such alternative structures or method features are explicitly excluded in the above description of the examples of the present disclosure.
[0273] Whilst endeavoring in the foregoing specification to draw attention to those features believed to be of importance the Applicant may seek protection via the claims in respect of any patentable feature or combination of features hereinbefore referred to and / or shown in the drawings whether or not emphasis has been placed thereon.
[0274] l / we claim:
Claims
35CLAIMS1. An apparatus for immersive audio communication comprising:at least one processor; andat least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to perform at least:determining at least one audio processing mode that is supported by the apparatus; transmitting a list comprising the at least one audio processing mode;receiving a response indicating at least one preferred audio processing mode wherein the at least one preferred audio processing mode is based on the transmitted list; andusing a preferred audio processing mode indicated in the response for an immersive audio communication session.
2. An apparatus as claimed in claim 1 , wherein the different audio processing modes comprise at least one of:different audio suppression modes;different speech enhancement modes;modes with different levels of amplification;modes with different levels of compression;modes with different levels of equalization;modes with different levels of bass roll-off;modes with different levels of suppression;modes with different levels of noise suppression; andmodes with different levels of ambience suppression.
3. An apparatus as claimed in any preceding claim, wherein when a first audio processing mode is used processed audio signals comprise speech and ambient sound and when a second audio processing mode is used the ambient sound is suppressed in the processed audio signals.
4. An apparatus as claimed in claim 3, wherein the respective audio processing modes affect the content of a transmitted bitstream such that when the first audio processing mode is used the transmitted bitstream comprises speech and ambient sound.
365. An apparatus as claimed in any of claim 3 or 4, wherein the respective audio processing modes affect the content of a transmitted bitstream such that when the second audio processing mode is used ambient sound is suppressed in the transmitted bitstream to enhance the intelligibility of speech compared to the first audio processing mode.
6. An apparatus as claimed in any of claims 3 to 5, wherein the apparatus is configured to support multiple different audio processing modes with different levels of suppression of the ambient sound.
7. An apparatus as claimed in any preceding claim, wherein the at least one audio processing mode is determined at least in part based on one of:a negotiated bitrate parameter;offered bitrate parameter;ambient noise in the environment of the apparatus;capture device used by the apparatus; andoutput device used by the apparatus.
8. An apparatus as claimed in any preceding claim, wherein the list is transmitted using session description protocol (SDP) signaling.
9. An apparatus as claimed in any preceding claim, wherein the list is indicated using an audio processing mode negotiation parameter.
10. An apparatus as claimed in any preceding claim, wherein processor and memory are further configured to cause the apparatus to:receive a list comprising at least one audio processing mode;determine at least one preferred audio processing mode from the received list;transmit a response indicating the determined at least one preferred audio processing mode of the apparatus; anduse a preferred audio processing mode indicated in the sent response for the immersive audio communication session11. An apparatus as claimed in any preceding claim, wherein the immersive audio communication session is one of:using the same audio processing mode in both directions; andestablished using a first audio processing mode for bitstreams transmitted by the apparatus and a second audio processing mode for bitstreams received by the apparatus.
12. An apparatus as claimed in any preceding claim, wherein the received response indicates an order of preference for the at least one audio processing mode.
13. A method comprising:determining at least one audio processing mode that is supported by an apparatus; transmitting a list comprising the at least one audio processing mode;receiving a response indicating at least one preferred audio processing mode wherein the at least one preferred audio processing mode is based on the transmitted list; andusing a preferred audio processing mode indicated in the response for an immersive audio communication session.
14. A method as claimed in claim 13, wherein when a first audio processing mode is used processed audio signals comprise speech and ambient sound and when a second audio processing mode is used the ambient sound is suppressed in the processed audio signals.
15. A method as claimed in claim 14, wherein the respective audio processing modes affect the content of a transmitted bitstream such that when the first audio processing mode is used the transmitted bitstream comprises speech and ambient sound.
16. A method as claimed in any of claim 14 or 15, wherein the respective audio processing modes affect the content of a transmitted bitstream such that when the second audio processing mode is used ambient sound is suppressed in the transmitted bitstream to enhance the intelligibility of speech compared to the first audio processing mode.
17. A method as claimed in any of claims 13 to 16, comprising supporting multiple different audio processing modes with different levels of suppression of the ambient sound.
18. A method as claimed in any of claims 13 to 17, wherein the list is at least one of: transmitted using session description protocol (SDP) signaling; andindicated using an audio processing mode negotiation parameter.
19. A method as claimed in any of claims 13 to 18, further comprising:receiving a list comprising at least one audio processing mode;determining at least one preferred audio processing mode from the received list;transmitting a response indicating the determined at least one preferred audio processing mode of the apparatus; andusing a preferred audio processing mode indicated in the sent response for the immersive audio communication session20. A method as claimed in any of claims 13 to 19, wherein the immersive audio communication session is one of:using the same audio processing mode in both directions; andestablished using a first audio processing mode for bitstreams transmitted by the apparatus and a second audio processing mode for bitstreams received by the apparatus.
21. A method as claimed in any of claims 13 to 20, wherein the received response indicates an order of preference for the at least one audio processing mode.
22. An apparatus for immersive audio communication comprises means for:determining at least one audio processing mode that is supported by the apparatus; transmitting a list comprising the at least one audio processing mode;receiving a response indicating at least one preferred audio processing mode wherein the at least one preferred audio processing mode is based on the transmitted list; andusing a preferred audio processing mode indicated in the response for an immersive audio communication session.
23. A computer program comprising computer program instructions for causing an apparatus to perform at least the following or for performing at least the following:determining at least one audio processing mode that is supported by the apparatus; transmitting a list comprising the at least one audio processing mode;receiving a response indicating at least one preferred audio processing mode wherein the at least one preferred audio processing mode is based on the transmitted list; andusing a preferred audio processing mode indicated in the response for an immersive audio communication session.