Immersive communication sessions
A second encoding mode in immersive communication sessions dynamically enhances focused participants' audio by reallocating bitrate and muting non-focused participants, improving user experience and audio quality.
Patent Information
- Application Number
- GB2024008115
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-07
- Publication Date
- 2025-12-10
AI Technical Summary
Existing immersive communication sessions, such as audio teleconferences, lack the ability to dynamically focus on specific participants, leading to suboptimal audio quality and user experience.
Implementing a second encoding mode that enhances the audio of focused participants by reallocating bitrate, applying higher gains, muting or canceling audio from non-focused participants, and encoding focused audio as separate objects, based on user requests.
Enhances the audio quality of focused participants, improving the user experience by reducing distractions from non-focused participants, and efficiently managing bitrate allocation.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
TECHNOLOGICAL FIELD Examples of the disclosure relate to immersive communication sessions. Some relate to enabling focus requests in immersive audio sessions. BACKGROUND Immersive communication sessions such as audio teleconferences can enable voices or audio sources to be rendered from different directions. This can enable participants in the communication session to engage in more natural communication because the rendering to different directions can make it fell more like a real life situation. BRIEF SUMMARY According to various, but not necessarily all, examples of disclosure there is provided an apparatus comprising means for: encoding audio streams for a telecommunication session using a first encoding mode wherein the telecommunication session comprises three or more participants; receiving a request to focus on at least one of the participants; and enabling focussing on the requested participant by switching from the first encoding mode to a second encoding mode wherein the second encoding mode enhances the participant for which focus has been requested relative to the participants for which focus has not been requested such that different enhancements are provided to different participants based, at least in part, on an origin of the request. In the second encoding mode, the enhancement of the participant for which focus has been requested may be provided to a participant that made the request but not to participants that did not make the request. The means may be for receiving at least a first request to focus on a first participant and a second request to focus on a second participant. The means may be for performing, in response to receiving a second request to focus on a second participant, one of: enabling focus on the participant with the highest priority; using the second encoding mode that enables focus on both the first participant and the second participant; or enabling user selection of the participant to focus on. A request to focus on a participant may be received from one of: a sending participant; or a receiving participant. An encoding mode may be defined by one or more parameters, the parameters comprising one or more of: an input format; bitrate; gains applied to respective components of the audio stream; muting respective component of the audio stream; or mixes of object audio streams. The second encoding mode may allocate a higher bitrate to the one or more participants for which focus has been requested than to the one or more participants for which focus has not been requested. Enabling focussing on a requested participant may comprise reallocating bitrate from one or more participants for which focus has not been requested to the one or more participants for which focus has been requested. The second encoding mode may comprise applying a lower gain to participants for which focus has not been requested than to participants for which focus has been requested. The second encoding mode may comprise muting participants for which focus has not been requested. The second encoding mode may comprise cancelling audio from participants for which focus has not been requested in the audio from the participants for which focus has been requested. The second encoding mode may comprise encoding the audio from the participants for which focus has been requested as a separate object to the audio from participants for which focus has not been requested. The second encoding mode may comprise using discontinuous transmission for the audio from participants for which focus has not been requested. The request may be received using real-time transport protocol signalling. The apparatus may comprise, or may be comprised within one or more of: a participant device, a server device or a teleconferencing server. According to various, but not necessarily all, examples of disclosure there is provided a method comprising: encoding audio streams for a telecommunication session using a first encoding mode wherein the telecommunication session comprises three or more participants; receiving a request to focus on at least one of the participants; and enabling focussing on the requested participant by switching from the first encoding mode to a second encoding mode wherein the second encoding mode enhances the participant for which focus has been requested relative to the participants for which focus has not been requested such that different enhancements are provided to different participants based, at least in part, on an origin of the request. According to various, but not necessarily all, examples of disclosure there is provided a computer program comprising instructions which, when executed by an apparatus, cause the apparatus to perform: encoding audio streams for a telecommunication session using a first encoding mode wherein the telecommunication session comprises three or more participants; receiving a request to focus on at least one of the participants; and enabling focussing on the requested participant by switching from the first encoding mode to a second encoding mode wherein the second encoding mode enhances the participant for which focus has been requested relative to the participants for which focus has not been requested such that different enhancements are provided to different participants based, at least in part, on an origin of the request. According to various, but not necessarily all, embodiments there is provided an apparatus comprising at least one processor; and at least one memory including computer program code; the at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to perform at least a part of one or more methods described herein. According to various, but not necessarily all, embodiments there is provided an apparatus comprising means for performing at least part of one or more methods described herein. The description of a function and / or action should additionally be considered to also disclose any means suitable for performing that function and / or action. Functions and / or actions described herein can be performed in any suitable way using any suitable method. According to various, but not necessarily all, embodiments there is provided examples as claimed in the appended claims. While the above examples of the disclosure and optional features are described separately, it is to be understood that their provision in all possible combinations and permutations is contained within the disclosure. It is to be understood that various examples of the disclosure can comprise any or all the features described in respect of other examples of the disclosure, and vice versa. Also, it is to be appreciated that any one or more or all the features, in any combination, may be implemented by / comprised in / performable by an apparatus, a method, and / or computer program instructions as desired, and as appropriate. The description of a function should additionally be considered to also disclose any means suitable for performing that function BRIEF DESCRIPTION Some examples will now be described with reference to the accompanying drawings in which: FIG. 1 shows an example immersive communication session; FIG. 2 shows an example method; FIG. 3 shows an example use case; FIG. 4 shows an example method; FIGS. 5A to 5C show an example use case; FIGS. 6A to 6B show an example use case; and FIG. 7 shows an example apparatus. The figures are not necessarily to scale. Certain features and views of the figures can be shown schematically or exaggerated in scale in the interest of clarity and conciseness. For example, the dimensions of some elements in the figures can be exaggerated relative to other elements to aid explication. Corresponding reference numerals are used in the figures to designate corresponding features. For clarity, all reference numerals are not necessarily displayed in all figures. DETAILED DESCRIPTION Fig. 1 schematically shows an example immersive communication session 100. The example communication session can make use of the immersive voice and audio services (IVAS) codec. Other codecs could be used in other examples. The telecommunication session 100 is between multiple participants 102. The participants 102 can be users of participant devices 104. Different types of participant devices 104 can be used by different participants 102. In some examples a participant device 104 can be shared by multiple participants 102. The participant devices 104 can comprise any suitable type of devices. The participant devices 104 could comprise teleconferencing devices, mobile telephones, personal computers or any other suitable type of devices that can be configured to capture audio and provide playback of audio signals to one or more participants 102. In the example of Fig. 1 the telecommunication session 100 is between multiple participants 102. In this case three participants 102_1, 102_2, 102_3 are located in the same room and are sharing the participant device 104_1. The shared participant device 104_1 is a teleconferencing device. The teleconferencing devices comprises multiple microphones 106 and can capture spatial audio. This can enable information relating to the relative positions of the participants 102_1,102_2,102_3 to be captured in the audio signals. In the example of Fig. 1 the teleconferencing device comprises three microphones 106 so that each of the participants 102_1, 102_2, 102_3 using the teleconferencing device has their own microphone 106. The fourth participant 102_4 is located remotely to the other participants 102_1, 102_2, 102_3. For example, the fourth participant 102_4 is not in the same room as the other participants 102_1, 102_2, 102_3. The participant device 104_2 that is used by the fourth participant 102_4 can comprise a mobile telephone, a personal computer or any other suitable type of device. Other types of participant device 104 can be used in other examples. The respective participant devices 104 can be connected via any suitable communication network 108 so as to enable the telecommunication session 100 between the participants 102. The communication network 108 can comprise wired and / or wireless networks. The communication network 108 can comprise one or more servers. The servers, or any other suitable devices within the communication session 100, are configured to receive upstream signals from the participant devices and provide audio streams to the participant devices 104. The server, or any other suitable devices within the communication session 100, can be configured to perform any suitable mixing or processing of the received upstream signals to generate the audio stream. In the example of Fig. 1 the transmission of the audio can be handled as separate objects that are passed between the first participant device 104_1 and the second participant device 104_2. Other processes for handling the audio can be used in other examples. During a telecommunication session 100 a participant 102 might decide that they want to focus on one or more other participants 102 in the telecommunication session 100. For example, the fourth participant 102_4 might want to hear the first participant 102_1 more clearly than the other participants 102_2, 102_3 in the room. Examples of the disclosure provide methods and systems that can be used to enable participants 102 to focus on one or more other participants in the telecommunication session 100. Fig. 2 shows an example method that can be used to implement examples of the disclosure. The method can be implemented by a participant device 104, a server device, a teleconferencing server or any other suitable type of device or combination of devices. At block 200 the method comprises encoding audio streams for a telecommunication session 100 using a first encoding mode. The telecommunication session 100 comprises three or more participants 102. The participants 102 can use participant devices 104 to join the telecommunication session 100. The participants 102 and participant devices 104 can be arranged as shown in Fig. 1 or could be arranged in any suitable configuration. The telecommunication session 100 can comprise any number of participants 102 and / or participant devices 104 in examples of the disclosure. In some examples the audio streams can be encoded using IVAS or any other suitable codec. The first encoding mode can comprise any suitable encoding mode. In some examples the first encoding mode can provide an equal focus for each of the participants. An encoding mode can be defined by one or more parameters. The parameters can comprise one or more of: an input format; bitrate; gains applied to respective components of the audio stream; muting respective component of the audio stream; mixes of object audio streams, and any other suitable parameters or combinations of parameters. Different input formats that can be used in different encoding modes can comprise stereo, multi-channel audio, scene-based audio (SBA, Ambisonics), metadata-assisted spatial audio (MASA), object-based audio (Independent Stream with Metadata (ISM)), and combinations of object-based audio with MASA (OMASA) and object-based audio with SBA (OSBA) or any other suitable formats. At block 202 the method comprises receiving a request to focus on at least one of the participants 102. For instance, in the example of Fig. 1, the remote participant 102_4 can makea request that they want to focus on just one of the participants 102_1, 102_2, 102_3 using the shared participant device 104_1. The focus request can be made by any of the participants 102 within a telecommunication session. This could be a sending participant or a receiving participant. For example, a receiving participant 102 can request to focus on any of the other participants in the telecommunication session 100. A sending participant could request that they become the focus for other participants 102 in the telecommunication session 100. Any suitable means can be used to make the focus request. For example, the respective participant devices 104 can comprise user interfaces that can enable the participants 102 to make user inputs indicating the focus request. The focus request can be made implicitly or explicitly. An explicit focus request can comprise a participant 102 specifically indicating which participants 102 are to be focused on. An implicit focus request can be made indirectly by making a user input that causes a different action. For example, a sending participant 102 could mute all other participants 102 which can cause the action of muting other microphones, but also cause the focussing to the non-muted participant 102 by switching to the second encoding mode. The focus request can be received using any suitable signalling. For example, the focus request can be received using Real-time Transport Protocol (RTP) signalling. The request can be indicated using a Processing Information (PI) frame such as a MUTE PI frame or other suitable mechanism. At block 204 the method comprises enabling focussing on the requested participant 102. The focussing can comprise switching from the first encoding mode to a second encoding mode. The second encoding mode enhances the participant 102 for which focus has been requested relative to the participants 102 for which focus has not been requested. This enables different enhancements to be provided to different participants 102 based, at least in part, on an origin of the request. For example, the second encoding mode can provide a higher audio quality for a focussed participant 102 than for a non-focussed participant 102. In the second encoding mode, the enhancement of the participant 102 for which focus has been requested can be provided to a participant 102 that made the request but not to participants 102 that did not make the request. For example, in the telecommunication session 100 there could be multiple remote participants 102. In such cases if one of the remote participants 102 requested to focus on a participant 102 then this participant will receive the audio stream using the second encoding mode but other remote participants 102 that have not made a focus request will receive the audio stream using the first encoding mode. In some examples the second encoding mode allocates a higher bitrate to the one or more participants 102 for which focus has been requested than to the one or more participants 102 for which focus has not been requested. In some examples enabling focussing on a requested participant 102 comprises reallocating bitrate from one or more participants 102 for which focus has not been requested to the one or more participants 102 for which focus has been requested. This can enable the higher bitrate to be allocated to the one or more participants 102 for which focus has been requested. In some examples the second encoding mode can provide a focus on the participants 102 for which focus has been requested by providing less enhancement on the participants 102 for which focus has not been requested. For example, the second encoding mode can comprise applying a lower gain to participants 102 for which focus has not been requested than to participants for which focus has been requested and / or the second encoding mode can comprise muting participants 102 for which focus has not been requested. In some examples the second encoding mode can comprise cancelling audio from participants 102 for which focus has not been requested in the audio from the participants 102 for which focus has been requested. For instance, if the focus request has been to focus on a participant 102 in a shared room then the microphone used by the focussed participant 102 would also capture sound (noise) from other participants 102 in the room. This noise could be cancelled in the audio stream. In some examples the second encoding mode can comprise encoding the audio from the participants 102 for which focus has been requested as a separate object to the audio from participants 102 for which focus has not been requested. For example, OMASA can be used for the focused participants 102 and MASA can be used for non-focused participants 102. In some examples the second encoding mode can comprise using discontinuous transmission (DTX) for the audio from participants for which focus has not been requested. In some examples multiple focus requests can be made. In such examples a first request to focus on a first participant 102 can be received and a second request to focus on a second participant 102 can also be received. The respective requests can be received from the same participant 102 or from different participants 102. When multiple focus requests are received the method comprises blocks for handling any conflicts between the respective requests. Example methods for handling conflicting focus requests can comprise, enabling focus on the participant 102 with the highest priority, using a second encoding mode that enables focus on both the first participant 102 and the second participant 102, enabling user selection of the participant 102 to focus on, or any other suitable method. Where the conflict is resolved by enabling focus on the participant 102 with the highest priority, any suitable means can be used to assign levels of priority to respective participants 102. For example, a higher priority can be given to a request made by the presenter or host of the telecommunication session 100, than to requests made by other participants 102. In some examples a higher priority can be given to requests to focus on the presenter or host of the telecommunication session 100 (or other specified participant 102) than requests to focus on other participants 102. In some examples the requests for focus can be accepted in response to a user input from one or more of the participants 102. For example, the host or presenter (or other specified participant 102) could decide whether a focus request should be accepted. If multiple focus requests are received the host or presenter (or other specified participant 102) could decide which, if any, of the requests should be accepted. In some examples where multiple focus requests are received the second encoding mode can enable focus on multiple participants. For example, enhancements would be provided for the multiple participants 102 for which focus has been requested but not for the other participants for which a focus request has not been accepted. Fig. 3 shows an example use case for examples of the disclosure. Fig. 3 shows the same telecommunication session 100 as shown in Fig. 1. Corresponding reference numbers are used for corresponding features. In this example the fourth participant 102_4 makes a request 300 to focus on one of the participants 102 in the shared room. In this case the fourth participant 102_4 makes a request 300 to focus on the first participant 102_1. In this example the request 300 can be handled by the teleconferencing device. The request could be handled by other devices in other examples. Before the request is made the audio streams are encoded using a first encoding mode. In the first encoding mode none of the participants 102 are enhanced relative to any of the other participants 102. After the request 300 has been accepted the audio streams are encoded using a second encoding mode. In the second encoding mode the first participant 102_1 is enhanced relative to the other participants 102_2, 102_3. In this example the teleconferencing device can enable the focusing by muting the microphones of the other participants 102_2, 102_3. In some examples the teleconferencing device can also cancel any noise from the other participants 102_2, 102_3 that is captured by the non-muted microphone. In some examples, once the microphones of the other participants 102_2, 102_3 have been muted the audio objects from these participants 102_2, 102_3 can be changed to a low bitrate mode such as discontinuous transmission (DTX). In some examples, any bitrate that is saved by this change can be used to improve the audio object quality for the first participant 102_1. A focused audio stream can then be transmitted to the participant device 104_2 of the remote participant 102_4. This provides an improved user experience for the remote participant 102_4 who can now completely focus on their participant 102_1 of interest, (in this case the first participant 102_1) without disturbance from other voices from the other participants 102_2, 102_3. In some examples the telecommunication session 100 can use IVAS. In such examples an IVAS encoder (and decoder) is configured to run on the first participant device 104_1 and an IVAS encoder (and decoder) is configured to run on the second participant device 104_2. In this example the first participant device 104_1 receives a single ISM as an input and the second participant device 104_2 receives three ISMs as an input. The focus request 300 can be made by transmitting a PI frame such as MUTE PI data frame as part of the IVAS RTP payload. The focus request 300 can be made to the first participant device 104_1 by sending it a MUTE PI data frame for streams corresponding to the second participant 102_2 and the third participant 102_3. That is, the MUTE PI data frame is sent for participants 102 for which a focus request is note made. The first participant device 104_1 switches to the second encoding mode in which the audio object (ISM) streams from the second participant 102_2 and the third participant 102_3 are muted. The first participant device 104_1 can also cancel the corresponding audio from the remaining ISM input. The IVAS encoder now receives three ISMs as an input, however, two of the ISMs are muted. The first participant device 104_1 can therefore enable a higher bitrate allocation for the remaining ISM. The remote participant 102_4 therefore receives a higher-quality rendering of the first participant 102_1 without disturbance from other voices. For example, MUTE PI data frame or MUTE_STREAM PI data frame can be used to send a request to mute an incoming IVAS audio in a single stream or a combined stream (e.g., ISMs, OMASA) or an incoming IVAS stream in a multi-streaming session. The example below includes two stream identifiers which indicate a request to mute the incoming streams with ID2 and ID3. The F-bit indicates if another ID field follows. For example, for ID2 the F-bit is set to 1 (indicating that another ID field follows) and for ID3 the F-bit is set to 0 (indicating that ID3 is the last ID field in the MUTE_STREAM PI data). The MUTE_STREAM PI data frames can be zero-padded to force byte-alignment. o i 0123456789012345 +-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+ |F| ID2 |F| ID3 |F|0000000| For example, in above example, ID2 can correspond to second participant 102_2 and ID3 can correspond to third participant 102_3. Fig. 4 shows an example method that can be used to enable focus requests during a telecommunication session 100. The method can be used for a telecommunication session 100 as shown in Figs. 1 and 3. Similar methods can be used for other arrangements of telecommunication sessions. In the example of Fig. 4 different parts of the method can be performed by different entities within the telecommunication session 100. For example, blocks 400, 404 and 410 can be performed by a server or other suitable device, block 402, can be performed by a second participant device 104_2 and blocks 406 and 408 can be performed by a first participant device 104_1. Other arrangements of the blocks, and / or other types of devices can be used in other examples. At block 400 a telecommunication session 100 is established. In this example telecommunication session 100 comprises three participants 102_1, 102_2, 102_3 sharing the first participant device 104_1 and a remote participant 102_4 using the second participant device 104_2. At block 402 the remote participant 102_4 decides that they want to focus on the first participant 102_1. For example, the first participant 102_1 could be the main presenter or speaker in the telecommunication session 100 and the remote participant 102_4 could decide that they would like to hear them more clearly. The remote participant 102_4 can make an input to the second participant device 104_2 indicating that they would like to focus on the first participant 102_1. Any suitable means can be used to enable the user input. At block 404 the request to focus on the first participant 102_1 is sent from the second participant device 104_2 to the first participant device 104_1. The request can be sent via the server or any other suitable communication means. After the first participant device 104_1 receives the request the first participant device 104_1 can enable focusing on the first participant 102_1. This can be performed by switching to a second encoding mode. In the example of Fig. 4 the focusing is enabled at block 406 by muting the microphones 106 of the non-focused participants 102_2, 102_4 and at block 408 by using a low bit rate mode for the non-focused participants 102_2, 102_4 and at block 410 by increasing bitrate for focused participants 102_1. In some examples the focusing can be enabled by using one or more of blocks 406 to 410. At block 412 the focused stream is sent from the first participant device 104_1 to the second participant device 104_2. This can then enable focused audio to be provided to the remote participant 102_4. This can enable the remote participant 102_4 to hear their chosen participant more clearly and / or with improved perceived quality. Figs.5A to 5C show an example use case in which a participant 102 in a telecommunication session 100 can request to focus on one or more other participants 102. In the examples of Figs. 5A to 5C the telecommunication session 100 comprises a fourway session between four participants 102_1, 102_2, 102_3, 102_4. Each of the participants 102_1, 102_2, 102_3, 102_4 is using their own participant device 104_1, 104_2, 104_3, 104_4. The respective participant devices 104_1, 104_2, 104_3, 104_4 can comprise mobile telephones, personal computers or any other suitable type of device. Each of the participant devices 104_1, 104_2, 104_3, 104_4 transmits an audio signal 502_1, 502_2, 502_3, 502_4 to the server 500. The respective audio signals 502 comprise participant audio, for example, audio captured by the microphones of the respective participant devices 104. The server 500 provides a mix 504_1, 504_2, 504_3, 504_4 to each of the participant devices 104_1, 104_2, 104_3, 104_4. The mix comprises the audio from the other participants 102. As an example, the first participant device 104_1 transmits an audio signal 502_1 comprising the audio of the first participant 102_1 and receives a mix 504_1 comprising audio from the second participant 102_2, the third participant 102_3 and the fourth participant 102_4. In some examples the server 500 can be an IVAS encoding device. In such examples the transmission formats used for the audio signals 502 that are sent from the participant devices 104 to the server 500 can be ISM (audio object). Other formats can be used in other examples. The transmission format that is used for the signals sent from the server 500 to the participant devices 104 can comprise OMASA. The OMASA transmission format flexibly supports spatial captures or mixtures and separate (for example, separated) objects to be coded together with automatic bitrate adjustment between component signals. The bitrate adjustment can be based on the waveform properties or any other suitable parameters. Other formats can be used in other examples. In the initial state of the telecommunication session 100, as shown in Fig. 5A, there are no active focus requests. In this case all of the signals are transmitted to each other participant 102. In the mix signals 504 none of the participants 102 are enhanced relative to any of the other participants 102. In this case an OMASA signal that is sent from the server 500 to the respective participant devices 104 can comprise a mixture of other users as MASA format and object signals that are muted. The muting of the object signals enables more bitrate to be used for the MASA part. In Fig. 5B the fourth participant 102_4 decides that they would like to focus on the first participant 102_1 as indicated by the arrow 510. The fourth participant 102_4 can make a user input indicating that they would like to focus on the first participant 102_1. The user input can be made using the fourth participant device 104_4 or any other suitable device. The request to focus on the first participant 102_1 can be sent from the fourth participant device 104_4 to the server 500. When the server 500 receives the focus request the server switches from the first mode of encoding to second mode of encoding. Instead of sending a mix 504_4 of all the participants to the fourth participant 102_4 the server 500 now sends a focused audio stream 512. The focused audio stream 512 can comprise enhanced audio from the first participant 102_1. The focused audio can be obtained by muting the non- focused participants 104_2 and 104_3 or in any other suitable way. In examples where OMASA is used for the signals sent from the server 500 to the participant devices 104 the server 500 adjusts the encoding so that the stream from the first participant device 104_1 is encoded as a separate object in OMASA. That is the signal from the first participant device 104_1 is not mixed into the MASA part but is forwarded into one of the object parts. The server 500 can mix the signals from the other participant devices 104_2 and 104_3 in a normal way, that is without any attenuation or can mix the signals from the other participant devices 104_2, 104_3 with attenuation. In some examples the signals from the other participant devices 104_2 and 104_3 could be completely discarded. The server 500 can distribute bitrate between different parts of the OMASA signal. This can enable the bitrate used for the focused part to be increased. If the focused participant 102_1 is provided as a separated object in OMASA and the other participants 102_2, 102_3 are downmixed to the MASA part then gain control signaling can be provided in an RTP payload. The gain control signaling can set a default gain for the MASA part that attenuates it in playback. This can enable the focusing on the selected participant 102 to be achieved in an efficient way. As the audio from the other participants 102_2, 102_3 is still transmitted in the MASA part of the focused stream this can enable the fourth participant 102_4 to adjust the level of the other participants 102_2 and 102_3 with more granularity. This can enable fine-tuning of the rendering for optimal or personalized experiences. For example, gain information can be provided using a suitable PI data frame. At a later point in time during the telecommunication session 100 the second participant 102_2 decides that they would like all of the other participants 102_1, 102_3, 102_ 4 to focus on them. For example, the second participant 102_2 might have important information that they would like to share with the rest of the participants 102_1, 102_3, 102_ 4. In this case the second participant 102_2 can make a user input to indicate that they would like all of the other participants 102_1, 102_3, 102_ 4 to focus on them. A focus request can be sent from the second participant device 104_2 to the server 500. The server 500 can handle the focus requests. For the first participant device 104_1 and the third participant device 104_3 there is only one relevant focus request (the request to focus on the second participant 102_2). Therefore, for each of the first participant device 104_1 and the third participant device 104_3 the server 500 provides a focused audio stream 520 that comprises enhanced audio from the second participant 102_2. The same focused audio stream 520 can be provided to both the first participant device 104_1 and the third participant device 104_3. The same processes used to generate the focused audio stream 512 in Fig. 5B can be used to generate the focused audio streams 520 in Fig. 5C. The fourth participant device 104_4 has two relevant focus requests. These are the initial request from the fourth participant device 104_4 to focus on the first participant 102_1 and the later request from the second participant device 104_2 to focus on the second participant 102_2. This causes a conflict between the two requests. In the example of Fig. 5C the server 500 resolves this conflict by providing a focused audio stream 522 that focuses on both the first participant 102_1 and the second participant 102_2. In this case the focused audio stream 522 that is provided to the fourth participant device 104_4 is different to the focused audio streams 520 that are provided to the first participant device 104_1 and the third participant device 104_2. In some cases the focused audio stream 522 can be generated by providing both the audio from the first participant 102_1 and the audio from the second participant 102_2 as separate object parts of OMASA. In this case this would leave the audio from the third participant 102_3 as the only part in MASA. In this case it can be more efficient to also assign the audio from the third participant 102_3 to another separate object part of OMASA. The server 500 can distribute the bitrate automatically to provide optimal quality for transmission of OMASA. Depending on the configuration of the server 500 or the telecommunications session 100, the audio from the third participant 102_3 can be attenuated or muted before encoding. Gain metadata parameters (such as those transmitted using RTP) could be also affected. The focusing of the audio streams could be obtained in other ways in other examples of the disclosure. For example, multi-format encoding with multiple IVAS encoders could be used and focusing could be achieved by using DTX for audio from non-focused participants 102. In the example of Figs. 5A to 5C a conflict between focus requests is resolved by providing a focused audio stream 522 that focuses on both the first participant 102_1 and the second participant 102_2. This provides an output that is consistent with the indications of the respective participants 102. This results in the participants 102 receiving the signals that they consider to be important and that other participants 102 consider to be important. However, this approach can require a high demand for computational resources for constructing the relevant streams and also may require a higher overall bitrate or can lose in quality if multiple signals are concurrently active at a given bitrate. Other methods for the resolving conflicting focus requests can be used in other examples. For instance, in some examples the sever 500 can enable the focus request for the participant 102 with the highest priority and discard the other focus requests. In such examples a higher priority focus request would override a lower priority focus request. The priority of the respective focus requests can be determined using any suitable criteria. An example order of priority from higher to lower could be: - a request from the presenter or host of the telecommunication session 100 to focus to them a request from a participant 102 other than the presenter or host of the telecommunication session 100 to focus to them - a request from a participant to focus on another participant 102. In some examples the requests from a participant 102 other than the presenter or host of the telecommunication session 100 to focus to them can be permitted by the presenter or host of the telecommunication session 100. In the example of Fig. 5C the request to focus on the second participant 102_2 could be the higher priority because they could be the host or presenter of the telecommunication session 100. In such circumstances the server 500 would accept the request to focus on the second participant 102_2 but would discard the request to focus on the first participant 102_1. This would result in the fourth participant device 104_4 being provided with the focused audio stream 520 that only focusses on the second participant 102_2. In this case all of the other participant devices 104_1, 104_3, 104_4 would receive the same focused audio stream 520. In some examples, in order to resolve a conflict of focus requests, the server 500 could enable user selection of the request that is to be accepted. In the example of Fig. 5C the fourth participant 102_4 could select between retaining the focus on the first participant 102_1 or switching to focus on the second participant 102_2. In some examples the second participant 102_2 could make the selection. In some examples a user could make a user input to control or adjust the balance of focus streams. This can enable a participant 102 to adjust how a focused participant is enhanced relative to the non-focused participants 102. Enabling a user selection of the request that is to be accepted can help to reduce demands on computation resources. Figs.6A to 6B show another example use case in which a participant 102 in a telecommunication session 100 can request to focus on one or more other participants 102. In the examples of Figs. 6A to 6B the telecommunication session 100 comprises a first teleconference room 600_1 and a second teleconference room 600_2. The teleconference rooms 600 are also monitored by multiple remote participants 102_X. In the example of Figs. 6A and 6B five remote participants 102_X are shown but there could be any number. The first teleconference room 600_1 comprises a first participant device 104_1. The first participant device 104_1 can be a teleconferencing device or any other suitable device. The first participant device 104_1 can be shared by multiple participants 102 in the first teleconference room 600_1. In the example of Figs. 6A and 6B there are three participants 102_1, 102_2 102_3 in the first teleconference room 600_1. Similarly, the second teleconference room 600_2 comprises a second participant device 104_2. The second participant device 104_2 can be a teleconferencing device or any other suitable device. The second participant device 104_2 can be shared by multiple participants 102 in the second teleconference room 600_2. In the example of Figs. 6A and 6B there are two participants 102_4, 102_5 in the second teleconference room 600 2. The participant devices 104_1, 104_2 that are shared between multiple participants 102 can comprise a spatial microphone 106_1 and a close up microphone 106_2. The spatial microphone 106_1 can comprise an array of microphones that are configured to capture spatial audio. The close up microphone 106_2 can be used for focus audio. Each of the remote participants 102_X has their own participant device 104_X. the participant devices 104_X used by the remote participants 102_X could comprise mobile phones, personal computers or any other suitable type of devices. When the telecommunication session 100 is established the first participant device 104_1 transmits an audio signal 602 comprising the audio of the first teleconference room 600_1 to the server 500 and receives an audio signal 604 comprising a teleconference mix from the server 500. Similarly, the second participant device 104_2 transmits an audio signal 606 comprising the audio of the second teleconference room 600_2 to the server 500 and receives an audio signal 608 comprising a teleconference mix from the server 500. In this example the remote participant devices 104_X only receive an audio signal 610 comprising a teleconference mix. In the initial state of the telecommunication session 100, as shown in Fig. 6A, there are no active focus requests. In the teleconference mix signals (e.g. audio signals 604, 608, 610) none of the participants 102 are enhanced relative to any of the other participants 102. The spatial microphones 106_1 in the respective teleconference rooms can be enabled and used to capture the audio from the room. The close up microphones 106_2 can be disabled or inactive or suppressed. In Fig. 6B the first participant 102_1 decides that they would like all of the other participants 102 to focus on them. The first participant 102_1 can make a request for focusing by making a user input in the first participant device 104_1 or in any other suitable way. In response to the focus request the spatial microphone 106_1 in the first participant device 104_1 is suppressed and the close up microphone 106_2 can be enabled. The audio stream can be changed to use more bitrate for object transmission that for spatial stream transmission. The audio signal 622 that is transmitted from the first participant device 104_1 to the server 500 is focused on the first participant 102_1. That is the first participant 102_1 is enhanced relative to the other participants 102_2 and 102_3. In some examples the server 500 can also receive an indication of the focus request. This can be implicit in the modified audio stream 622 or could be received using any suitable signaling. In some examples the server can detect that there has been a focus request from changes in the received signals, for example, changes in the received IVAS packets. The first participant device 104_1 still receives the audio signal 604 comprising a teleconference mix from the server 500. This has not changed compared the scenario in Fig. 6A because there are no other focus requests on the other participants 102. Similarly, the audio signal 606 that is transmitted from the second participant device 104_2 to the server 500 has not changed compared the scenario in Fig. 6A because there are no focus requests on the participants 102_4 in the second teleconference room 600_2. However, the audio signal 628 that is received by the second participant device 104_2 from the server 500 has changed because it is now focused on the first participant 102_1. The audio signal 620 that is received by the remote participant devices 104_X from the server 500 has also changed because it now comprises a part focused on the first participant 102_1 and a part from the second teleconference room 600_2. The use case of Figs. 6A and 6B can be implemented using the IVAS codec. Such examples can use dual encoder transmission (that is, two IVAS encoder instances at each location) as each of the first participant device 104_1, the second participant device 104_2 and the server 500. When there are no active focus requests the format used for the audio signals 602, 606 transmitted from the teleconference rooms 600 can comprise SBA for transmitting HOA-microphone capture in ideal form. The audio from the separate close up microphone 106_2 can be transmitted as ISM (audio object with spatial position in the scene, for example relative to the SBA spatial scene). The transmission of audio signals 602, 606 from the teleconference rooms 600 can use constant bitrate. The bitrate can be 128 kbps or any other suitable bitrate. The bitrate of 128 kbps can provide good quality. The audio signals 604, 608 that are received by the teleconference rooms 600 from the server 500 can comprise multi-channel audio. The multi-channel audio can be directly suitable for reproduction using a loudspeaker system available in the teleconference room 600. The audio signals 610, 620 that are received by the remote participant devices 104_X from the server 500 can be provided as split rendered binaural rendering. The split rendered binaural rendering can allow lightweight consumption. In the initial state, as shown in Fig. 6A, the bitrate (128 kbps) in the audio signal 602,604 from each teleconference room 600 can be allocated to SBA format encoding using the first IVAS encoder while the ISM format encoding using the second IVAS encoder can be done with DTX. A gain control can be applied to the ISM channel to trigger the DTX operation. These configurations can be negotiated according to IVAS SDP negotiation parameters. After the focus request has been made the second IVAS encoder instance, in the first participant device 104_1 is changed to use a higher bitrate (such as 128 kbps) for encoding the ISM input. The gain control is not applied to the ISM channel anymore. Gain control is started for the SBA input. The second IVAS encoder instance is switched to DTX-on operation. This changes the audio signal 622 that is transmitted from the first participant device 104_1 to the server 500 to be focused on the first participant 102_1. In this case the audio signal 622 is completely focused on the first participant 102_1. When the server 500 determines that a focus request has been made the server 500 provides a focused stream 620 to the remote participant devices 104_X. The server 500 can also indicate that the focus request has been implemented. The focused stream, 620 that is provided to the remote participant devices 104_X can comprise binaural rendering of the focused audio from the first teleconference room 600_1. The focused stream, 620 can also comprise a normal, attenuated or muted signal from the second teleconference room 600_2. In the examples described herein any changes in bitrate assignment or applied gains to control the focus can be performed in a time smoothed way. This can help to avoid sudden changes that could be perceived as artifacts. The time smoothing can be achieved by averaging values over time or by using hysteresis operations or any other suitable means. The gain control can be applied in various different ways such as pre-encoding muting or gain reduction of non-focused audio signals after a focus request has been made. For high-quality IVAS encoding and transmission, (for example, efficient bitrate allocation) where the IVAS encoder is operated with ISM inputs, the gain of a focused ISM is maintained at a normal level. Alternatively, the gain of a focused ISM can be gain controlled to maintain a suitable maximum volume level. The gain of a non-focused ISM is significantly lowered or completely muted. In some examples DTX operation can be activated for the non-focused ISMs. In some examples the gain-control approach can be as follows: Focused audio is maintained at normal gain at all times Non-focused audio is muted when at least one focused audio is active If no focused audio input is active (for example, an external voice activity detection is zero), non-focused audio is faded back in (this can be at lower gain than under normal conditions) - As soon as at least one focused audio becomes active again (for example an external voice activity detection is one), any non-focused audio is again muted This approach provides an IVAS encoder active signal for transmission so that the listener will always get some audio as long as at least one person is talking. However, a focused participant 102 immediately overrides any non-focused participant 102 in away that allows for optimal, or substantially optimal, bit allocation for the desired scenario. For example, in case of OMASA input, the so-called separated object in IVAS OMASA encoding is always allocated as (one of) the focused participants. Fig. 7 schematically illustrates an apparatus 700 that can be used to implement examples of the disclosure. In this example the apparatus 700 comprises a controller 702. The apparatus 700 can be a chip or a chipset. In some examples the apparatus 700 can be provided within a user device, a server device, a teleconferencing server or within any other suitable device within a telecommunication system. In the example of Fig. 7 the implementation of the apparatus 700 can be as controller circuitry. In some examples the apparatus 700 can be implemented in hardware alone, have certain aspects in software including firmware alone or can be a combination of hardware and software (including firmware). As illustrated in Fig. 7 the apparatus 700 can be implemented using instructions that enable hardware functionality, for example, by using executable instructions of a computer program 708 in a general-purpose or special-purpose processor 702 that can be stored on a computer readable storage medium (disk, memory etc.) to be executed by such a processor 702. The processor 702 is configured to read from and write to the memory 704. The processor 702 can also comprise an output interface via which data and / or commands are output by the processor 702 and an input interface via which data and / or commands are input to the processor 702. The memory 704 is configured to store a computer program 706 comprising computer program instructions (computer program code) that controls the operation of the apparatus 700 when loaded into the processor 1702. The computer program instructions, of the computer program 706, provide the logic and routines that enables the apparatus 700 to perform the methods illustrated in the Figs. The processor 702 by reading the memory 704 is able to load and execute the computer program 706. The apparatus 700 therefore comprises: at least one processor 702; and at least one memory 704 including computer program code, the at least one memory 704 and the computer program code configured to, with the at least one processor 702, cause the apparatus 700 at least to perform: encoding 200 audio streams for a telecommunication session using a first encoding mode wherein the telecommunication session comprises three or more participants 102; receiving 202 a request to focus on at least one of the participants 102; and enabling 204 focussing on the requested participant 102 by switching from the first encoding mode to a second encoding mode wherein the second encoding mode enhances the participant 102 for which focus has been requested relative to the participants 102 for which focus has not been requested such that different enhancements are provided to different participants 102 based, at least in part, on an origin of the request. As illustrated in Fig. 7 the computer program 706 can arrive at the apparatus 700 via any suitable delivery mechanism 708. The delivery mechanism 708 can be, for example, a machine readable medium, a computer-readable medium, a non-transitory computer-readable storage medium, a computer program product, a memory device, a record medium such as a Compact Disc Read-Only Memory (CD-ROM) or a Digital Versatile Disc (DVD) or a solid state memory, an article of manufacture that comprises or tangibly embodies the computer program 706. The delivery mechanism 708 can be a signal configured to reliably transfer the computer program 706. The apparatus 700 can propagate or transmit the computer program 706 as a computer data signal. In some examples the computer program 706 can be transmitted to the apparatus 700 using a wireless protocol such as Bluetooth, Bluetooth Low Energy, Bluetooth Smart, 6LoWPan (IPv6 over low power personal area networks) ZigBee, ANT+, near field communication (NFC), Radio frequency identification, wireless local area network (wireless LAN) or any other suitable protocol. The computer program 706 comprises computer program instructions that when executed by an apparatus 700 cause the apparatus 700 to perform at least the following: encoding 200 audio streams for a telecommunication session using a first encoding mode wherein the telecommunication session comprises three or more participants 102; receiving 202 a request to focus on at least one of the participants 102; and enabling 204 focussing on the requested participant 102 by switching from the first encoding mode to a second encoding mode wherein the second encoding mode enhances the participant 102 for which focus has been requested relative to the participants 102 for which focus has not been requested such that different enhancements are provided to different participants 102 based, at least in part, on an origin of the request. The computer program instructions can be comprised in a computer program 706, a non-transitory computer readable medium, a computer program product, a machine readable medium. In some but not necessarily all examples, the computer program instructions can be distributed over more than one computer program 706. Although the memory 704 is illustrated as a single component / circuitry it can be implemented as one or more separate components / circuitry some or all of which can be integrated / removable and / or can provide permanent / semi-permanent / dynamic / cached storage. Although the processor 702 is illustrated as a single component / circuitry it can be implemented as one or more separate components / circuitry some or all of which can be integrated / removable. The processor 702 can be a single core or multi-core processor. References to “computer-readable storage medium”, “computer program product”, “tangibly embodied computer program” etc. or a “controller”, “computer”, “processor” etc. should be understood to encompass not only computers having different architectures such as single / multi- processor architectures and sequential (Von Neumann) / parallel architectures but also specialized circuits such as field-programmable gate arrays (FPGA), application specific circuits (ASIC), signal processing devices and other processing circuitry. References to computer program, instructions, code etc. should be understood to encompass software for a programmable processor or firmware such as, for example, the programmable content of a hardware device whether instructions for a processor, or configuration settings for a fixed-function device, gate array or programmable logic device etc. As used in this application, the term “circuitry” can refer to one or more or all of the following: (a) hardware-only circuitry implementations (such as implementations in only analog and / or digital circuitry) and (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions and (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g. firmware) for operation, but the software cannot be present when it is not needed for operation. This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit for a mobile device or a similar integrated circuit in a server, a cellular network device, or other computing or network device. The blocks illustrated in the Figs, can represent steps in a method and / or sections of code in the computer program 706. The illustration of a particular order to the blocks does not necessarily imply that there is a required or preferred order for the blocks and the order and arrangement of the blocks can be varied. Furthermore, it can be possible for some blocks to be omitted. Where a structural feature has been described, it may be replaced by means for performing one or more of the functions of the structural feature whether that function or those functions are explicitly or implicitly described. The apparatus can be provided in an electronic device, for example, a mobile terminal, according to an example of the present disclosure. It should be understood, however, that a mobile terminal is merely illustrative of an electronic device that would benefit from examples of implementations of the present disclosure and, therefore, should not be taken to limit the scope of the present disclosure to the same. While in certain implementation examples, the apparatus can be provided in a mobile terminal, other types of electronic devices, such as, but not limited to: mobile communication devices, hand portable electronic devices, wearable computing devices, portable digital assistants (PDAs), pagers, mobile computers, desktop computers, televisions, gaming devices, laptop computers, cameras, video recorders, GPS devices and other types of electronic systems, can readily employ examples of the present disclosure. Furthermore, devices can readily employ examples of the present disclosure regardless of their intent to provide mobility. The term ‘comprise’ is used in this document with an inclusive not an exclusive meaning. That is any reference to X comprising Y indicates that X may comprise only one Y or may comprise more than one Y. If it is intended to use ‘comprise’ with an exclusive meaning then it will be made clear in the context by referring to ‘comprising only one...’ or by using ‘consisting.’ In this description, the wording ‘connect’, ‘couple’ and ‘communication’ and their derivatives mean operationally connected / coupled / in communication. It should be appreciated that any number or combination of intervening components can exist (including no intervening components), i.e., to provide direct or indirect connection / coupling / communication. Any such intervening components can include hardware and / or software components. As used herein, the term "determine / determining" (and grammatical variants thereof) can include, not least: calculating, computing, processing, deriving, measuring, investigating, identifying, looking up (for example, looking up in a table, a database, or another data structure), ascertaining and the like. Also, "determining" can include receiving (for example, receiving information), accessing (for example, accessing data in a memory), obtaining and the like. Also, "determine / determining" can include resolving, selecting, choosing, establishing, and the like. In this description, reference has been made to various examples. The description of features or functions in relation to an example indicates that those features or functions are present in that example. The use of the term ‘example’ or ‘for example’ or ‘can’ or ‘may’ in the text denotes, whether explicitly stated or not, that such features or functions are present in at least the described example, whether described as an example or not, and that they can be, but are not necessarily, present in some of or all other examples. Thus ‘example’, ‘for example’, ‘can’, or ‘may’ refers to a particular instance in a class of examples. A property of the instance can be a property of only that instance or a property of the class or a property of a sub-class of the class that includes some but not all the instances in the class. It is therefore implicitly disclosed that a feature described with reference to one example but not with reference to another example, can where possible be used in that other example as part of a working combination but does not necessarily have to be used in that other example. As used herein, “at least one of the following: ” and “at least one of ” and similar wording, where the list of two or more elements are joined by “and” or “or” mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements. Although examples have been described in the preceding paragraphs with reference to various examples, it should be appreciated that modifications to the examples given can be made without departing from the scope of the claims. Features described in the preceding description may be used in combinations other than the combinations explicitly described above. Although functions have been described with reference to certain features, those functions may be performable by other features whether described or not. The description of a feature, such as an apparatus or a component of an apparatus, configured to perform a function, or for performing a function, should additionally be considered to also disclose a method of performing that function. For example, description of an apparatus configured to perform one or more actions, or for performing one or more actions, should additionally be considered to disclose a method of performing those one or more actions with or without the apparatus. Although features have been described with reference to certain examples, those features may also be present in other examples whether described or not. The term ‘a’, ‘an’ or ‘the’ is used in this document with an inclusive not an exclusive meaning. That is any reference to X comprising a / an / the Y indicates that X may comprise only one Y or may comprise more than one Y unless the context clearly indicates the contrary. If it is intended to use ‘a’, ‘an’ or ‘the’ with an exclusive meaning then it will be made clear in the context. In some circumstances the use of ‘at least one’ or ‘one or more’ may be used to emphasis an inclusive meaning but the absence of these terms should not be taken to infer any exclusive meaning. The presence of a feature (or combination of features) in a claim is a reference to that feature or (combination of features) itself and to features that achieve substantially the same technical effect (equivalent features). The equivalent features include, for example, features that are variants and achieve substantially the same result in substantially the same way. The equivalent features include, for example, features that perform substantially the same function, in substantially the same way to achieve substantially the same result. In this description, reference has been made to various examples using adjectives or adjectival phrases to describe characteristics of the examples. Such a description of a characteristic in relation to an example indicates that the characteristic is present in some examples exactly as described and is present in other examples substantially as described. The above description describes some examples of the present disclosure however those of ordinary skill in the art will be aware of possible alternative structures and method features which offer equivalent functionality to the specific examples of such structures and features described herein above and which for the sake of brevity and clarity have been omitted from the above description. Nonetheless, the above description should be read as implicitly including reference to such alternative structures and method features which provide equivalent functionality unless such alternative structures or method features are explicitly excluded in the above description of the examples of the present disclosure. Whilst endeavoring in the foregoing specification to draw attention to those features believed to be of importance the Applicant may seek protection via the claims in respect of any patentable feature or combination of features hereinbefore referred to and / or 5 shown in the drawings whether or not emphasis has been placed thereon. l / we claim: 10
Claims
1. An apparatus comprising means for:encoding audio streams for a telecommunication session using a first encoding mode wherein the telecommunication session comprises three or more participants;receiving a request to focus on at least one of the participants; andenabling focussing on the requested participant by switching from the first encoding mode to a second encoding mode wherein the second encoding mode enhances the participant for which focus has been requested relative to the participants for which focus has not been requested such that different enhancements are provided to different participants based, at least in part, on an origin of the request.
2. An apparatus as claimed in claim 1 wherein, in the second encoding mode, the enhancement of the participant for which focus has been requested is provided to a participant that made the request but not to participants that did not make the request.
3. An apparatus as claimed in any preceding claim wherein the means are for receiving at least a first request to focus on a first participant and a second request to focus on a second participant.
4. An apparatus as claimed in claim 3 wherein the means are for performing, in response to receiving a second request to focus on a second participant, one of:enabling focus on the participant with the highest priority;using the second encoding mode that enables focus on both the first participant and the second participant; orenabling user selection of the participant to focus on.
5. An apparatus as claimed in any preceding claim wherein a request to focus on a participant is received from one of:a sending participant; ora receiving participant.
6. An apparatus as claimed in any preceding claim wherein an encoding mode is defined by one or more parameters, the parameters comprising one or more of:an input format;bitrate;gains applied to respective components of the audio stream;muting respective component of the audio stream; or mixes of object audio streams.
7. An apparatus as claimed in any preceding claim wherein the second encoding mode allocates a higher bitrate to the one or more participants for which focus has been requested than to the one or more participants for which focus has not been requested.
8. An apparatus as claimed in any preceding claim wherein enabling focussing on a requested participant comprises reallocating bitrate from one or more participants for which focus has not been requested to the one or more participants for which focus has been requested.
9. An apparatus as claimed in any preceding claim wherein the second encoding mode comprises applying a lower gain to participants for which focus has not been requested than to participants for which focus has been requested.
10. An apparatus as claimed in any preceding claim wherein the second encoding mode comprises muting participants for which focus has not been requested.
11. An apparatus as claimed in any preceding claim wherein the second encoding mode comprises cancelling audio from participants for which focus has not been requested in the audio from the participants for which focus has been requested.
12. An apparatus as claimed in any preceding claim wherein the second encoding mode comprises encoding the audio from the participants for which focus has been requested as a separate object to the audio from participants for which focus has not been requested.
13. An apparatus as claimed in any preceding claim wherein the second encoding mode comprises using discontinuous transmission for the audio from participants for which focus has not been requested.
14. An apparatus as claimed in any preceding claim wherein the request is received using real-time transport protocol signalling.
15. An apparatus as claimed in any preceding claim wherein the apparatus comprises, or is comprised within one or more of: a participant device, a server device or a teleconferencing server.
16. A method comprising:encoding audio streams for a telecommunication session using a first encoding mode wherein the telecommunication session comprises three or more participants;receiving a request to focus on at least one of the participants; andenabling focussing on the requested participant by switching from the first encoding mode to a second encoding mode wherein the second encoding mode enhances the participant for which focus has been requested relative to the participants for which focus has not been requested such that different enhancements are provided to different participants based, at least in part, on an origin of the request.
17. A computer program comprising instructions which, when executed by an apparatus, cause the apparatus to perform:encoding audio streams for a telecommunication session using a first encoding mode wherein the telecommunication session comprises three or more participants;receiving a request to focus on at least one of the participants; andenabling focussing on the requested participant by switching from the first encoding mode to a second encoding mode wherein the second encoding mode enhances the participant for which focus has been requested relative to the participants for which focus has not been requested such that different enhancements are provided to different participants based, at least in part, on an origin of the request.
Citation Information
Patent Citations
System and method for a conference server architecture for low delay and distributed conferencing applications
EP1966917B1
System and method for a conference server architecture for low delay and distributed conferencing applications
US20080158339A1
System and method for a conference server architecture for low delay and distributed conferencing applications
US20160255307A1