Immersive communication sessions
By determining and selecting codec operating points based on participant computational capabilities, immersive communication sessions achieve high immersion levels for all participants, addressing the issue of reduced quality due to device limitations.
Patent Information
- Application Number
- PCT/EP2025/058848
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-24
- Filing Date
- 2025-04-01
- Publication Date
- 2025-10-30
AI Technical Summary
Existing communication sessions fail to account for the computational capabilities of participant devices, resulting in reduced immersion levels due to the limitations of less capable devices, leading to suboptimal rendering outcomes.
An apparatus and method for determining and selecting codec operating points in immersive communication sessions that consider the computational capabilities of participants, enabling high immersion by choosing operating points that align with the available resources and preferences of all participants.
This approach ensures that all participants in a communication session can experience high levels of immersion by optimizing codec operating points based on their computational capabilities, avoiding the limitations seen in sessions where less capable devices restrict output quality.
Smart Images

Figure EP2025058848_30102025_PF_FP_ABST
Abstract
Description
[0001] TITLE
[0002] Immersive Communication Sessions
[0003] TECHNOLOGICAL FIELD
[0004] Examples of the disclosure relate to immersive communication sessions. Some relate to selecting operating points for immersive communication sessions so as to provide high immersion for participants in the communication session.
[0005] BACKGROUND
[0006] Communication sessions between two or more participants can have multiple operating points available. The operating points can comprise a specific point within the operation characteristic of participant devices within the communication sessions. The respective operating points can comprise characteristics relating to encoding, decoding and rendering of an audio bitstream.
[0007] BRIEF SUMMARY
[0008] According to various, but not necessarily all, examples of the disclosure there is provided an apparatus for an immersive communication session, the apparatus comprising means for: determining a set of codec operating points indicated by a first participant in the communication session; obtaining a subset of the set of codec operating points wherein the subset is indicated to the first participant from at least one further participant in the communication session; selecting from the subset at least one codec operating point for the communication session which results in rendering with high immersion for at least one participant in the communication session wherein the selection is based, at least in part, on the computational capability of at least one participant in the communication session.
[0009] The codec operating point may be selected to provide high immersion for the first participant and one or more further participants. Multiple categories of immersion may be available and selecting a codec operating point which results in rendering with high immersion comprises selecting a codec operating point in a category with highest available immersion.
[0010] The different categories of immersion may comprise at least some of: mono rendering; stereo rendering; spatial rendering; binaural rendering; binaural rendering with headtracking and with room reverb; binaural rendering with headtracking and without room reverb; binaural rendering without headtracking and with room reverb; binaural rendering without headtracking and without room reverb; multichannel loudspeaker rendering, multichannel 5.1 loudspeaker rendering; multichannel 5.1.2 loudspeaker rendering; multichannel 5.1.4 loudspeaker rendering; multichannel 7.1 loudspeaker rendering; multichannel 7.1.4 loudspeaker rendering; multichannel loudspeaker rendering with horizontal speakers; multichannel loudspeaker rendering with elevated speakers; multichannel loudspeaker rendering with frontal speakers;
[0011] IVAS supported output rendering formats.
[0012] An immersion parameter of the at least one further participant may be indicated to the first participant.
[0013] The selection of the codec operating point may be based, at least in part, on the computational capability of the first participant and one or more further participants.
[0014] The computational capability of a participant may be indicated using one or more computational capability parameters.
[0015] The computational capability of a participant may be inferred from information provided the by the at least one participant.
[0016] The set of codec operating points may be determined based on one or more negotiation parameters.
[0017] The one or more negotiation parameters may comprise at least one of: bitrate, bandwidth, input format, discontinuous transmission; output format, coded format, computational capability. The selected codec operating point may use one or more format conversions for an audio stream.
[0018] The first participant may be a sending participant and the at least one further participant may be a receiving participant.
[0019] The subset of the set of codec operating points may be selected by multiple further participants to enable selection of at least one codec operating point based, at least in part, on the computational complexity of the multiple further participants.
[0020] The codec operating point may be selected to enable rendering of an audio stream by the multiple further participants without transcoding the audio stream.
[0021] The communication session may comprise a bidirectional immersive audio call.
[0022] According to various, but not necessarily all, examples of the disclosure there is provided a method comprising: determining a set of codec operating points indicated by a first participant in the communication session; obtaining a subset of the set of codec operating points wherein the subset is indicated to the first participant from at least one further participant in the session; selecting from the subset at least one codec operating point for the communication session which results in rendering with high immersion for at least one participant in the communication session wherein the selection is based, at least in part, on the computational capability of at least one participant in the session.
[0023] According to various, but not necessarily all, examples of the disclosure there is provided a computer program comprising instructions which, when executed by an apparatus, cause the apparatus to perform: determining a set of codec operating points indicated by a first participant in the communication session; obtaining a subset of the set of codec operating points wherein the subset is indicated to the first participant from at least one further participant in the session; selecting from the subset at least one codec operating point for the communication session which results in rendering with high immersion for at least one participant in the communication session wherein the selection is based, at least in part, on the computational capability of at least one participant in the session.
[0024] According to various, but not necessarily all, embodiments there is provided an apparatus comprising at least one processor; and at least one memory including computer program code; the at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to perform at least a part of one or more methods described herein.
[0025] According to various, but not necessarily all, embodiments there is provided an apparatus comprising means for performing at least part of one or more methods described herein. The description of a function and / or action should additionally be considered to also disclose any means suitable for performing that function and / or action. Functions and / or actions described herein can be performed in any suitable way using any suitable method.
[0026] According to various, but not necessarily all, embodiments there is provided examples as claimed in the appended claims.
[0027] While the above examples of the disclosure and optional features are described separately, it is to be understood that their provision in all possible combinations and permutations is contained within the disclosure. It is to be understood that various examples of the disclosure can comprise any or all the features described in respect of other examples of the disclosure, and vice versa. Also, it is to be appreciated that any one or more or all the features, in any combination, may be implemented by / comprised in / performable by an apparatus, a method, and / or computer program instructions as desired, and as appropriate. The description of a function should additionally be considered to also disclose any means suitable for performing that function BRIEF DESCRIPTION
[0028] Some examples will now be described with reference to the accompanying drawings in which:
[0029] FIG. 1 shows an example communication session:
[0030] FIG. 2 shows an example session negotiation;
[0031] FIG. 3 shows an example method;
[0032] FIG. 4 shows an example session negotiation;
[0033] FIG. 5 shows an example communication session;
[0034] FIG. 6 shows an example method;
[0035] FIG. 7 shows an example method;
[0036] FIG. 8 shows an example method;
[0037] FIG. 9 shows an example communication session; and
[0038] FIG. 10 shows an example controller.
[0039] The figures are not necessarily to scale. Certain features and views of the figures can be shown schematically or exaggerated in scale in the interest of clarity and conciseness. For example, the dimensions of some elements in the figures can be exaggerated relative to other elements to aid explication. Corresponding reference numerals are used in the figures to designate corresponding features. For clarity, all reference numerals are not necessarily displayed in all figures.
[0040] DETAILED DESCRIPTION
[0041] Fig. 1 shows an example communication session 100 that can implement examples of the disclosure. The communication session 100 comprises a bidirectional immersive audio call between a first user equipment (UE) 102_1 and a second UE 102_2. The UEs 102 are participants in the communication session 100. In other examples the communication session 100 can comprise more than two participants.
[0042] The UEs 102 can comprise any suitable types of devices. The UEs 102 can comprise personal communication devices such as mobile phones, devices within a teleconferencing system and / or any other types of communication devices. The UEs 102 can comprise microphones and / or any other suitable means for capturing audio. The UEs 102 can comprise, or can be coupled to, one or more playback devices or other means for playing back rendered audio. The means for playing back audio can comprise loudspeaker or headsets or any other suitable means.
[0043] The UEs 102 can also comprise apparatus or controllers that can be as shown in Fig. 10. The UEs 102 can comprise immersive audio codecs. The UEs 102 can make use of the immersive voice and audio services (IVAS) codec. Other codecs could be used in other examples.
[0044] The communication session 100 comprises a bidirectional immersive communication session. The bidirectional immersive communication session comprises two independent directions. The first direction is from the first UE 102_1 to the second UE102_2 as indicated by the arrow 104_1 . The second direction is from the second UE 102_2 to the first UE 102_1 as indicated by the arrow 104_2.
[0045] During the communication session 100 an audio stream can be processed and encoded by the first UE 102_1 and provided for rendering to the second UE 102_2. Similarly an audio stream can be processed and encoded by the second UE 102_2 and provided for rendering to the first UE 102_1. The respective audio streams can be captured by microphones associated with the respective UEs 102 and / or can be obtained as an input.
[0046] The different UEs 102 in the communication session 100 can have different computational capabilities. The computational capability of a UE 102 can be defined in terms of a level. The different levels enable rendering with different complexity and memory requirements.
[0047] For instance, in UEs 102 that use an IVAS codec, the different levels can be built on each other so that a higher level comprises all features and functionalities of lower levels but with some additional features. For example, level one could be a core level, level two can comprise all of the features of level one and some additional features. Level three can comprise all of the features of level two and some additional features. Level three can be the highest level. The highest level can comprise the full set of IVAS codec features and functionalities. A UE 102 or any other type of device configured for IVAS shall support at least functionality level one.
[0048] The levels can be defined as follows:
[0049] The following level-dependent limits apply for IVAS codec operations (encoder / decoder / renderer total) excluding Jitter Buffer Management and other supplementary operations:
[0050] Level 1 : o Complexity <= 3 * EVS o RAM <= 3 * EVS
[0051] Level 2: o Complexity <= 6 * EVS o RAM <= 6 * EVS
[0052] Level 3: o Complexity <= 10 * EVS o RAM <= 10 * EVS
[0053] The highest level provides the full functionality of IVAS. At the lower levels, reduced functionality is provided. Each increasing complexity level can provide full support of all the lower levels.
[0054] The EVS interoperability mode of IVAS should not require substantially increased complexity or memory compared to standard EVS. For example, IVAS can include a bit-exact implementation of the EVS codec standard.
[0055] The following level-independent ROM and PROM constraints apply: o ROM, PROM <= 10 * EVS
[0056] In the above level indications, “EVS” stands for the standard baseline functional implementation of the 3GPP EVS codec described in TS 26.445 specification. For example, the phrase “Complexity <= 3 * EVS” indicates that the measured or evaluated computational complexity of the IVAS codec should be less than or equal to three times the computational complexity of the standard baseline EVS implementation. The standard fixed point EVS implementations are described, for example, in 3GPP TS 26.442 and TS 26.452 specifications. Other definitions or indications of levels or general complexity-based device capabilities are possible. In this example three different levels of computational capabilities are defined. Other numbers of levels could be used in other examples.
[0057] In the example of Fig. 1 the first UE 102_1 has a lower computational capability than the second UE 102_2. The first UE 102_1 has a lower level for computational capability than the second UE 102_2. The first UE 102_1 could be a level 1 device while the second UE 102_2 could be a level 3 device. Other levels of the UEs 102 could be used in other examples.
[0058] Fig. 2 shows an example transmission of data between the two UEs 102. The two UEs 102 can be configured in a communication session 100 as shown in Fig. 1. In this case the different computational capabilities of the UEs 102 is not taken into account during the session negotiation. In this example, as in Fig. 1 , the first UE 102_1 operates in level 1 and the second UE 102_2 operates in level 3.
[0059] In this example the UEs 102 operate in their respective levels during the communication session 100 because their respective computational capabilities are not taken into account during session negotiation. Therefore the second UE 102_2 provides a level 3 encoded bitstream 200 to the first UE 102_1 . The bit stream is level 3 in terms of parameter combinations such as bitrate, bandwidth, input format, discontinuous transmission, output format, coded format, and / or any other suitable parameters. The decoding and rendering of the level 3 bitstream are computationally heavy for the first UE 102_1. In order to save computational resources the first UE 102_1 restricts the output format to mono. Therefore the first UE 102_1 is limited to providing a less immersive output. In this case the first UE 102_1 only provides a mono output 202.
[0060] The decoding of the level 3 bitstream and providing the mono output is still computationally heavy for the first UE 102_1 . As a result, the first UE 102_1 can only provide a mono encoded bitstream 204 to the second UE 102_2. The mono encoded bitstream 204 is decoded and rendered to mono output in the second UE 102_2 because there is no immersive output defined for mono encoded bitstreams. Therefore, the second UE 102_2 also provides a less immersive output. In this case the second UE 102_2 also only provides a mono output 206.
[0061] This example demonstrates the possible effects to the outputs in both UEs 102, if the computation capabilities of the UEs 102 are not considered during a session negotiation. In some cases, as shown in Fig. 2, both UEs 102 in the communication session provide less immersive outputs due to the restrictions of the computational capabilities of just one of the UEs 102 in the communication session.
[0062] Examples of the disclosure enable an immersive codec operating point to be selected for a communication session where the selection of the operating point accounts for the computational capabilities of at least one of the UEs 102 in the communication session. This can enable high levels of immersion to be provided by the UEs 102 within the communication session. This can help to avoid situations such as those shown in Fig. 2 where less immersive outputs are provided due to the computational capabilities of one or more of the UEs 102 in the communication session.
[0063] Fig. 3 shows an example method that can be used to implement examples of the disclosure. The example method could be implemented by a participant in a communication session 100 such as sending UE 102, a server or by any other suitable device in a communication system. The communication session 100 can be an immersive communication session. The communication session 100 can comprise a bidirectional immersive audio call and / or any other suitable type of communication session 100.
[0064] The method of Fig. 3 can be performed as part of a session negotiation or can be performed at any other suitable time. The session negotiation can occur during establishment of a communication session 100 between the two or more participants of the communication session 100. The session negotiation can be performed using session description offer answer model or using any other suitable procedure.
[0065] The method, or parts of the methods, can be implemented by encoder and decoder / renderer, encoder and decoder only, encoder and decoder and Tenderer, encoder only, decoder only, decoder / renderer, decoder and (separate) Tenderer, or any other suitable combination of devices.
[0066] The example method comprises, at block 300, determining a set of codec operating points. The codec operating points are indicated by a first participant in the communication session 100. The first participant can be a UE 102. The first participant can be a sending UE 102.
[0067] The codec operating points comprise specific points within the operation characteristics of the participants within the communication session. The respective operating points can comprise characteristics relating to encoding, decoding and rendering of an audio bitstream, and / or any other suitable characteristics.
[0068] The set of codec operating points can comprise any number of operating points. The set of codec operating points can comprise two or more operating points.
[0069] The codec operating points can be indicated as part of an offer from the first participant. The offer can be made as part of the session negotiation.
[0070] In some examples the set of codec operating points is determined based on one or more negotiation parameters. The one or more negotiation parameters can comprise bitrate, bandwidth, input format, discontinuous transmission; output format, coded format, computational capability, and / or any other suitable parameters. The codec operating points that are indicated in the set can be limited to operating points that operate within the negotiation parameters.
[0071] The set of codec operating points can provide an indication of the preferences of the first participant.
[0072] At block 302 the method comprises obtaining a subset of the set of codec operating points. The subset can comprise one or more of the codec operating points from the set of codec operating points that is determined at block 300. In some examples the subset can comprise the whole of the original set. In such cases the subset would comprise the same operating points as the original set.
[0073] The subset of the codec operating points can provide an indication of the operating points that can be met by further participant. This can take into account the computational capability of the further participant.
[0074] The subset is indicated to the first participant from at least one further participant in the communication session. The at least one further participant can select the subset of the operating points from the original set of codec operating points. The subset of codec operating points can then be indicated to the first participant. The at least one further participant can be another UE 102. The at least one further participant can be a receiving UE 102 or any other suitable device within the communication session 100.
[0075] At block 304 the method comprises selecting at least one codec operating point for the communication session. The codec operating point is selected from the subset. The codec operating point that is selected results in rendering with high immersion for at least one participant in the communication session 100. The codec operating point can be selected to provide high immersion for the first participant and one or more further participants.
[0076] The selection of the codec operating point is based, at least in part, on the computational capability of at least one participant in the communication session 100. In some examples the selection of the codec operating point is based, at least in part, on the computational capability of the first participant and one or more further participants.
[0077] The computational capability provides an indication of the computational resources that are available at a participant such as a UE 102. The computational capability can comprise an indication of the complexity or memory availability or any other suitable computation characteristic of a participant.
[0078] The computational capability of the participants can be defined in terms of levels. The levels can be as described above or can be configured in any other suitable manner. In the example described above three levels were indicated. Other numbers of levels could be used in other examples.
[0079] In some examples the computational capability of a participant can be indicated using one or more computational capability parameters. For example, a computational capability parameter can indicate a level of a participant, or any other suitable characteristic. The computational capability can be indicated using a session description protocol (SDP) parameter or by using any other suitable signalling. The computational capability parameter can be taken into account when the codec operating point is being selected.
[0080] In some examples the computational capability of a participant can be inferred from information provided by a participant. For example, rather than receiving an explicit indication of a computational complexity level the computational capability can be determined from information that is used during the session negotiation. The information that is used during the session negotiation that can also be used to infer the computational capability can comprise information relating to bitrate, bandwidth, input formats, discontinuous transmission; output format, coded format, and / or any other suitable parameters.
[0081] Providing high immersion can mean enabling an output to be provided in a format that has some immersive characteristics and avoiding formats or operating points that do not have immersive characteristics or that have less immersive characteristics.
[0082] In some examples multiple categories or steps of immersion are available. In such examples selecting a codec operating point which results in rendering with high immersion comprises selecting a codec operating point in a category with highest available immersion. Selecting a codec operating point which results in rendering with high immersion can comprise selecting a codec operating point that enables rendering from a more immersive category rather than using a less immersive category.
[0083] The different categories of immersion in terms of rendering can comprise; mono rendering, stereo rendering, spatial rendering, binaural rendering, binaural rendering with headtracking and with room reverb, binaural rendering with headtracking and without room reverb, binaural rendering without headtracking and with room reverb, binaural rendering without headtracking and without room reverb, multichannel loudspeaker rendering, multichannel 5.1 loudspeaker rendering, multichannel 5.1.2 loudspeaker rendering, multichannel 5.1.4 loudspeaker rendering, multichannel 7.1 loudspeaker rendering, multichannel 7.1.4 loudspeaker rendering, multichannel loudspeaker rendering with only horizontal speakers, multichannel loudspeaker rendering with elevated speakers, multichannel loudspeaker rendering with only frontal speakers, IVAS supported output rendering formats and / or any other suitable categories.
[0084] The different categories or steps of immersion can be arranged in order or the amount of immersion that can be provided by each category or step. This can enable higher and lower categories or steps to be defined. In the lower steps or categories lower amounts of immersion are available while in the higher steps of categories higher amounts of immersion are available.
[0085] In a case where the respective steps or categories comprise mono operation, stereo operation and spatial operation the ordering of the steps or categories would be:
[0086] Category / Step 1 : mono Category / Step 2: stereo Category / Step 3: spatial.
[0087] For example, the operation can be an encoder operation or property relating to encoder operation, for example immersion step of an input format and / or a decoder / renderer operation, for example rendering. Other types of categories and steps could be used in other examples.
[0088] For example, one decoder / renderer specific ordering of the steps or categories can be: Category / Step 1 : mono (rendering) Category / Step 2: stereo (rendering) Category / Step 3: binaural (rendering) Category / Step 4: multi-channel loudspeaker (rendering). In examples of the disclosure, a participant might be unable to support a preferred immersive category. In examples of the invention a different category can be selected where the different category can still provide some immersive characteristics. This could provide a higher quality of experience than simply reverting to a mono output. For example, if immersion category 3 (or spatial rendering) is not appropriate due to the computational capability of a participant, then an operating point can be selected from the highest immersion category that is available. This could enable an operating point that provides immersion category 2 (or stereo rendering) to be selected rather than switching to mono outputs as shown in Fig. 2. In the use case shown in Fig. 2 the methods of Fig. 3, that take the computational capability into account, could enable outputs to be provided in a stereo format rather than a mono format which would provide higher levels of immersion and improved quality of experience for a user.
[0089] In some examples of the disclosure the codec operating point can be selected so as to minimize, or substantially minimize, any reductions in the immersive category or step for one or more of the participants in the communication session 100.
[0090] In some examples the selected codec operating point can use one or more format conversions for an audio stream. The format conversions can be selected to enable higher immersive categories to be used. For example, the input format can be converted to a different format before encoding so that the decoder receives the audio stream in a converted format. The converted format might be less complex to decode than the original format and so might enable a high level of immersion to be provided by a decoding device. For example, the decoding device might be able to avoid dropping down to a lower immersive category level because it receives the audio stream in a format that is simpler to decode.
[0091] In the above described examples, the subset of codec operating points can be selected by one further participant or by multiple further participants and indicated to the first participant. If the subset of codec operating points is selected by multiple further participants this can enable the codec operating point to be selected based, at least in part, on the computational complexity of the multiple further participants. In such examples, the codec operating point can be selected to enable rendering of an audio stream by the multiple further participants without transcoding the audio stream. In an example use case a first UE 102_1 and a second UE 102_2 can be configured in a communication session 100 similar to that shown in Fig. 1. In this example, the first UE 102_1 and the second UE 102_2 can support IVAS encoder formats metadata- assisted spatial audio 2 (MASA2 (stereo-MASA)), first-order Ambisonics (FOA), stereo, and mono (enhanced voice services (EVS)). These formats can be offered as part of the session negotiation procedure.
[0092] The formats can be categorized in terms of the steps or categories of immersion that can be provided with the respective formats. Table 1 shows the categorization.
[0093] Using spatial encoding with a MASA2 format or an FOA format can enable spatial, stereo or mono outputs. That is, the immersion category or step for the decoding rendering could be any of the available categories or steps. If a stereo input format is used this can enable, stereo or mono outputs. That is, the immersion category or step can be the same as, or lower than, the input immersion category or step. If a mono input format is used this can only provide mono outputs.
[0094] Considering the second direction in the example of Fig. 4, which can be from the second UE 102_2 to the first UE 102_1. The second UE 102_2 has a higher computational capability than the first UE 102_1 . In the example being considered the second UE 102_2 has computational capability level 2 and the first UE 102_1 has computational capability level 1. As part of the session negotiation the second UE 102_2 receives an indication that the first UE 102_1 has computational capability level 1. This indication can be received explicitly as an SDP parameter or could be inferred from information comprised within other session negotiation parameters.
[0095] The computational capability can be measured in terms of weighted millions operations per second (WMOPS). EVS complexity is approximately 100 WMOPS, which is an example value used in the examples below. Therefore, from the definitions given above:
[0096] - Level 1 : 3 x EVS = 300 WMOPS
[0097] - Level 2: 6 x EVS = 600 WMOPS
[0098] - Level 3: 10 x EVS = 1000 WMOPS
[0099] For purposes of this example, it is assumed that all EVS operations are covered at least at Level 3 (here, 1000 WMOPS). In different embodiments, there may also be more granular level definitions (for example, sub levels) at 3.3 x EVS or 4.5 EVS.
[0100] As mentioned above the first UE 102_1 has computational capability level 1. Using complexity estimates for IVAS operation it can be observed that mono inputs, stereo inputs and MASA1 inputs are fully supported at level 1. That is, these inputs have complexity estimates that are below 300 WMOPS for all bandwidths and bitrates. MASA2 input is supported at level 1 to all binaural rendering outputs and also to stereo and mono. FOA inputs at super wide band (SWB) and below is supported up to a bitrate or 80 kbps to binaural output without room reverb and below (stereo, mono) at level 1. Alternatively, instead of limiting the bitrate the output could be limited to stereo or mono for an FOA input.
[0101] The second UE 102_2 has computational complexity level 2. For level 2, the computational capability limit is 600 WMOPS. This can enable FOA input to be fully supported to binaural output without room reverb. Alternatively, instead of limiting to a type of binaural output a different limitation could be used.
[0102] The second UE 102_2 has level 2 computational capability and so any of the offered formats (MASA2, FOA, stereo, mono) could be used. However, the first UE 102_1 has level 1 computational capability so there are operating points for both MASA2 and FOA input formats that the first UE 102_1 might not be able to decoder / render at high immersion quality. This could result in a limitation of the decoding / rendering at the first UE 102_1 if the second UE 102_2 freely chooses any Level 2 operation for FOA. For example if the second UE 102_2 chooses to encode using FOA with fullband (FB) bandwidth at 32 kbps, only stereo decoding operations would be available for the Level 1 operation at the first UE 102_1.
[0103] To enable selection of a codec operating point the decoding operations that would be available if the second UE 102_2 chooses to encode using MASA2 can be compared to those that would be available if the second UE 102_2 chooses to encode using FOA.
[0104] This shows that computational complexity levels for binaural outputs are lower for MASA2 than for FOA2. In this case binaural with reverb is available in level 1 for a MASA2 input. Therefore, it is beneficial to choose the MASA2 input instead of the FOA input for direction 2 encoding / transmission by the second UE 102_2 because this can provide high levels of immersion.
[0105] In this example, if MASA2 is chosen as the input format for the second UE 102_2 there is no reduction in the immersion step or category. This improves on the selection of FOA which would result in the reduction from binaural to stereo output at the first UE 102_1. This improvement is enabled because the second UE 102_2 selects the input format with knowledge of the computational capability of the first UE 102_1 .
[0106] In order to enable the selection of the operating point or input format with knowledge of the computational capability of other UEs 102 constraints can be added during a session negotiation. The constraints can relate to parameters such as bandwidth, bitrate, and can be used to enable selection of a codec operating point that is mutually acceptable to multiple participants within the communication session. In examples of the disclosure the session negotiation can comprise the first UE 102_1 and the second UE 102_2 negotiating the codec operating point to be used. The procedure used to select the codec operating point can comprise a first UE 102_1 indicating a set of parameters. The set of parameters can comprise parameters such as input audio formats, computational capability, output formats, bitrates, bandwidths and / or any other suitable parameters. In this case the input audio formats for the first UE 102_1 would be MASA, SBA, stereo, mono. This could also be the order of preference for the input formats. For example, MASA could be preferred over SBA. This could also include specific types of formats such as specifying that MASA is MASA2 and / or that SBA is FOA. The computational capability for the first UE 102_1 is level 1 and the output format is binaural. The binaural output format could be a specific type of binaural output. For example it can be specified whether it should be with or without head tracking and / or whether it should be with or without room reverb. The bit rate can be up to 32 kbps and the bandwidths could be FB or SWB.
[0107] The second UE 102_2 can also indicate a corresponding set of parameters. The corresponding set of parameters can also comprise parameters such as input audio formats, computational capability, output formats, bitrates, bandwidths and / or any other suitable parameters. In this case the input audio formats for the second UE 102_2 would be SBA, MASA, stereo, mono. This could also be the order of preference for the input formats. For example, SBA could be preferred over MASA. This could also include specific types of formats such as specifying that MASA is MASA2 and / or that SBA is FOA. The computational capability for the second UE 102_2 is level 2 and the output format is binaural. The binaural output format could be a specific type of binaural output. For example it can be specified whether it should be with or without head tracking and / or whether it should be with or without room reverb. The bit rate can be up to 32kbps and the bandwidths could be FB or SWB.
[0108] In examples, the corresponding parameters for UEs 102 can be dependent on directions. For example, the first UE 102_1 can indicate a preferred list of input formats or coded formats that the first UE 102_1 encodes and transmits to the second UE 102_2 (such as, sending parameters). The first UE 102_1 can then also indicate a preferred list of input formats or coded formats that the first UE 102_1 receives from the second UE 102_2 and decodes (such as, receiving parameters). The directionality can be indicated with a separate set of parameters, for example, cf-send and cf-recv for preferred list of sending and receiving coded formats, respectively. For a common sending and receiving parameter, a parameter of cf (coded format) could be used, for example. Similar directionality can be specified for all the suitable parameters.
[0109] The first UE 102_1 can determine a set of codec operating points that use level 1 encoding and level 2 decoding with rendering to binaural audio. The codec operating points are constrained to those that conform to the relevant computational capability level requirements, and the other parameters indicated by the UEs 102.
[0110] The second UE 102_2 can determine a set of codec operating points that use level 2 encoding and level 1 decoding with rendering to binaural audio. The codec operating points are constrained to those that conform to the relevant computational capability level requirements, and the other parameters indicated by the UEs 102.
[0111] The sets of codec operating points can be modified to subsets by taking into account the parameters indicated by the other UE 102. The respective codec operating points can be selected from the indicated sets or subsets. The codec operating points can be selected so as to provide high immersion. In this example, the preferred outputs are binaural and so the codec operating points can be selected to enable binaural. If binaural is not available then the codec operating can be selected so that the output is in an immersion step or category that reduces or minimizes the drop in immersion steps or categories.
[0112] Fig. 4 shows an example session negotiation between a first UE 102_1 and a second UE 102_2. The UEs 102 can be participants in a communication session. The UEs 102 can comprise mobile device or any other suitable type of devices. In this example the computational capability of both of the UEs 102 is taken into account during the negotiation of the session parameters.
[0113] The first UE 102_1 has characteristics that are relevant for the selection of a codec operating point. The characteristics can comprise preferred input formats, preferred output formats, computational complexity and / or any other suitable characteristics. In this example, the preferred IVAS encoder input formats for the first UE 102_1 are metadata-assisted spatial audio (MASA), scene based audio (SBA), stereo, mono, in that order. The first UE 102_1 also has a preferred output format of binaural. The first UE 102_1 also has a computational complexity of level 1.
[0114] The second UE 102_2 also has characteristics that are relevant for the selection of a codec operating point. The characteristics of the second UE 102_2 can also comprise preferred input formats, preferred output formats, computational complexity and / or any other suitable characteristics. In this example, the preferred IVAS encoder input formats for the second UE 102_2 are SBA, MASA, stereo, mono, in that order. The second UE 102_2 also has a preferred output format of binaural. In this example the second UE 102_2 has a computational complexity of level 2.
[0115] As part of the session negotiation the first UE 102_1 sends an offer message 400 to the second UE 102_2. The offer message 400 can be sent at the beginning of the session negotiation. The offer message 400 can be sent using any suitable signaling. In some examples the offer message 400 can be sent using session initiation protocol (SIP) or session description protocol (SDP) signaling. The offer message 400 can comprise an indication of the characteristics of the first UE 102_1. For example, it can comprise an indication of the preferred input formats, the order of preference of the input formats, the preferred output format, the computation complexity and / or any other suitable characteristics such as bitrate and bandwidth.
[0116] I n response to the offer message 400 the second U E 102_2 sends an answer message 402. The answer message 402 can be sent using any suitable signaling. In some examples the answer message 402 can be sent using SIP or SDP signaling. The answer message 402 can comprise an indication of the characteristics of the second UE 102_2. For example, it can comprise an indication of the preferred input formats, the order of preference of the input formats, the preferred output format, the computation complexity and / or any other suitable characteristics such as bitrate and bandwidth.
[0117] In some examples, the answer message 402 can comprise an indication of immersion step (IS). In some of these examples, if the IS parameter is included in the answer message 402 then the preferred output format might not be included as part of the answer message 402. The IS parameter can a simplified indication of the preferred output format. The IS parameter can describe the quality of experience (QoE) relating to immersion achievable on the second UE 102_2, for example, mono, stereo, binaural, multi-channel loudspeaker (rendering).
[0118] The first UE 102_1 and the second UE 102_2 can negotiate 404 a codec operating point for the communication session. The negotiation can comprise the first UE 102-1 determining 406 a set of codec operating points. The set of codec operating points can be indicated or offered as part of the session negotiation. The points within the set of codec operating points can take into account the characteristics of the respective UEs 102, including the computational capability.
[0119] The set of codec operating points can be indicated to the second UE 102_2. The second UE 102_2 can then, at block 408, select a subset codec operating points from the indicated set. For example, the second UE 102_2 can modify the set of codec operating points by removing operating points that are not suitable for the second UE 102_2. For example an operating point could be removed if the computational requirements are above the level of the UE 102_2. The subset of the codec operating points can then be indicated to the first UE 102_1 .
[0120] An operating point can then be selected from the subset of codec operating points. This selection can take into account the computational complexity of the respective UEs 102. The computational complexity can be taken into account when determining the set and subsets of codec operating points. The computational complexity can be taken into account by not including points that exceed the computational complexity of a UE 102 within the set or subset of operating points. The operating point can be selected to provide high levels of immersion for one or both of the UEs 102. A high level of immersion can be provided by a preferred output or an output that has an immersion category or step that is close to the immersion category or step of the preferred output.
[0121] In some examples the signaling can differ to that shown in the example of Fig. 4. For example, the set of codec operating points for the first UE 102_1 can be determined before the initial offer message 400. The subset for the operating points can be determined for second UE 102_2 after the initial offer 400 has been received by the second UE 102_2. The subset can be included in the answer message 402 to the first UE 102_1 . The offer-answer messaging between the UEs 102 determines the selected codec operating points for the session. After the first UE 102_1 receives the answer message 402 from the second UE 102_2, the operating points for the session are known, and an RTP session can be established between the UEs 102 and the audio payload transport 410 can begin.
[0122] The messages that are exchanged by the UEs 102 during the negotiation of the operating point can be sent using SIP or SDP signaling.
[0123] Once the codec operating point has been selected a real-time transport protocol session (RTP) between the first UE 102_1 and the second UE 102_2 is established between the UEs 102. The RTP session enables the audio payload transport 410.
[0124] Fig. 5 shows another example communication session 100 that can implement examples of the disclosure. This is similar to the communication session 100 shown in Fig. 1 in that it comprises a bidirectional immersive audio call between a first user equipment (UE) 102_1 and a second UE 102_2.
[0125] In the example of Fig. 1 computational capabilities were defined for the individual participants or UEs 102 taking into account the encoding and decoding / rendering in a single participant or UE 102. In the example of Fig. 1 the first UE 102_1 had a level 1 computational complexity and the second UE 102_2 had a level 3 computational complexity.
[0126] The computational capabilities can be defined in other ways. In the example of Fig. 5 the computational capabilities are defined in terms of an audio transmission path rather than the individual capabilities of the UEs 102. The audio transmission path comprises encoding in one UE 102 or participant and decoding / rendering in a different UE 102 or participant.
[0127] In the example of Fig. 5 the first direction 104_1 from the first UE 102_1 to the second UE 102_2 has a computational capability of level 1 and the second direction 104_2 from the second UE 102_2 to the first UE 102_1 has a computational capability of level 3. Other configurations and computational capabilities could be used in other examples.
[0128] The computational capability can be determined for a transmission path. For example, it can be considered for the encoder and the decoder / renderer. In some examples the computational capability can be divided into respective components such as optional pre-processing, encoder complexity, decoder complexity, Tenderer complexity, additional processing (such as jitter buffer management), or optional post processing or any other suitable components.
[0129] Fig. 6 shows an example method that can be used in some examples of the disclosure. The example of Fig. 6 shows determining one or more sets of encoder operating points for the encoding or sending participant in a communication session. In the example of Fig. 6 three sets are determined by the encoding participant. This can be understood as three examples of one set or one example of three sets. This then results in three subsets for the decoding / rendering participant. The method of Fig. 6 can be implemented by a sending participant or any other suitable device in the communication session.
[0130] To determine the set or sets of encoder operating points the capability of the sending participant is determined. The capability of the sending participant can comprise all possible encoder operations and decoder operations available for the sending participant. The encoder and decoder operations can also be available in different steps of immersion.
[0131] In the example of Fig. 6 four different encoder and decoder operations are shown. For example, these operations are based on the audio capture that is implemented on the device. These operations are labelled as operation A, operation B, operation C, and operation D in Fig. 6. Other numbers of encoder and decoder operations can be used in other examples. The different encoder operations can be based on different parameters or combinations of parameters such as bitrate and bandwidth. Each encoder operation is associated with one or more corresponding decoder operations. In this case three immersion steps are defined where step 1 enables mono rendering, step 2 enables stereo rendering and step 3 enables spatial rendering. Other numbers and types of immersion steps can be used in other examples. For example, steps that enable binaural rendering could be used. The steps could enable different types of binaural rendering such as with or without reverb and / or with or without head tracking.
[0132] The set of encoder operating points that is to be indicated or offered as part of the session negotiation can be determined by taking into account one or more session negotiation parameters for the sending participant. The session negotiation parameters can comprise bitrate, bandwidth, input format, discontinuous transmission, output format, coded format and / or any other suitable parameters. For example, in Fig. 6, each of encoder operation A, B, C, and D can correspond with one input format. For example, each encoder operation can furthermore include at least one bitrate, at least one bandwidth and / or any other suitable parameter. In examples of the disclosure the computational capability of the sending participant is also taken into account. This computational capability can be considered as the computational capability of the sending participant or of the transmission path used by the sending participant.
[0133] One or more of the encoder operations can be restrained when the negotiation parameters, including the computational complexity are taken into account. For instance, in the example of Fig. 6 some of the encoder operations are reduced (for encoder operation A). The reduction can be based on the computation level or any other suitable parameter or combination of parameters. For example, the reduction can relate to bitrates and / or maximum bandwidth of encoder operation A and / or any other suitable parameter.
[0134] After the encoder operations have been reduced to account for the negotiation parameters this provides a set of encoder operating points. The set of encoder operating points can be indicated or offered as part of a session negotiation with one or more receiving participants. In other example the encoder operations might not be reduced, for example, they could be checked and / or verified, for example, in terms of computational capability or complexity limits and / or arranged into an order of preference. Fig. 7 shows another example method that can be used in some examples of the disclosure. The method of Fig. 7 can follow on from the method of Fig. 6. The method of Fig. 7 can be implemented by a receiving participant or any other suitable device in the communication session. In the example of Fig. 7 the receiving participant can select a subset of encoder operating points from the original set of operating points that are offered or indicated by the sending participant. The subset can be selected taking into account negotiation parameters from the receiving participant. This enables the computational capability of both the participants to be taken into account.
[0135] As shown in Fig. 7 the set of codec operating points from the sending participant are offered or indicated to the receiving participant. The session negotiation parameters of the receiving participant are then taken into account. The session negotiation parameters can comprise bitrate, bandwidth, input format, discontinuous transmission, output format, coded format. In examples of the disclosure the computational capability of the sending participant is also taken into account. This computational capability can be considered as the computational capability of the receiving participant or of the transmission path used by the sending participant.
[0136] The receiving participant can select a subset of codec operating points from the set of codec operating points that is indicated by the sending participant. The subset can be selected by determining the immersion steps that are possible for each encoder operation taking into account the session negotiation parameters and the computational capability.
[0137] The process of selecting a subset of codec operating points can comprise taking the decoder output format into account. One or more session negotiation parameters can be used to take the decoder output format into account. For instance, if it is known that the decoder output is limited to a low immersion step it might not be appropriate to optimize for a higher immersion step.
[0138] In the example of Fig. 7 three example subsets 700, 702, 704 are shown. In the first subset 700 encoder operation A has been reduced by the sending participant. This subset 700 has also been further limited by the receiving participant due to the computational capability, and other relevant parameters, of the receiving participant. If an operating point is selected from this subset 700 then encoder operation B with immersion step 3 would preferably be selected because this enables selecting an encoder input operation (including encoder input format) that provides the highest immersion step and also the highest immersion step at the decoder / renderer according to the capabilities (including computational complexity) of both participants. In this case the operating points in the subset 700 have been reordered in terms of preference.
[0139] In the second subset 702 the encoder operation A is not reduced by the sending participant but it is limited by the receiving participant due to the computational capability, and other relevant parameters, of the receiving participant. If an operating point is selected from this subset 702 then encoder operation B with immersion step 3 would still preferably be selected because this also enables selecting an encoder input operation that provides a high immersion step and also a high immersion step at the decoder / renderer according to the capabilities (including computational complexity) of both participants. In this case the operating points in the subset 702 have been reordered in terms of preference.
[0140] In the third subset 704 the encoder operation A is not reduced by the sending participant but it is limited by the receiving participant due to the computational capability, and other relevant parameters, of the receiving participant. In this third subset 704 the encoder operation B is also reduced due to the immersion step and computational capability of the receiving participant, for example, the receiving device cannot support binaural rendering or spatial rendering via loudspeakers, and only mono and stereo outputs are supported.
[0141] If an operating point is selected from this subset 704 then encoder operation A with immersion step 2 would preferably be selected because this provides the highest level of immersion that is available for both the sending participant and the receiving participant once the computational capabilities (and, for example output format) are taken into account. In this case the points in the subset do not need to be reordered in terms of preference because there is no strong quality of experience benefit. Therefore the order of preference from the sending participant does not need to be overruled or otherwise changed.
[0142] The examples of the subsets therefore enable an encoder operating point to be selected that provides a high immersion step for both sending and receiving participants. This provides a high quality of experience for both the sending participant and the receiving participant.
[0143] Fig. 8 shows another example method that can be used in some examples of the disclosure. The method of Fig. 8 can follow on from the method of Fig. 6. The method of Fig. 8 can be implemented by a receiving participant or any other suitable device in the communication session. In the example of Fig. 8 the receiving participant can select or indicate a preference order for a subset of encoder operating points from the original set of operating points that are offered or indicated by the sending participant. In this example the subsets are ordered so that the operating points that provide the highest quality of experience are given a higher preference. The quality of experience can be correlated with the immersion steps that can be provided within an operating point.
[0144] The operating points that are given the highest preference can be those that maintain the highest immersion steps. That is, the operating points that are given the highest preference can be those that avoid reducing to a lower immersion step. For example, the highest preference can be assigned to the operating points that avoid switching from spatial rendering to stereo rendering or that avoid switching from stereo rendering to mono rendering.
[0145] In the examples of Figs. 6 to 8 there are three different immersion steps but there can be more immersion steps in other examples. This can provide a higher level of granularity in the selection of the operating points.
[0146] In some examples the encoder operating points that are included in the set can comprise a format conversion for the encoder input format of an audio stream. The format conversions can be selected to enable higher immersive categories to be used for the output formats. This could be used in scenarios where the negotiation encoder input formats are provided at a high immersion step such as MASA or SBA. In such cases it can be determined to use MASA1 instead of MASA2 which can require a format conversion. Or in some cases if the original format is HOA2 it can be determined to use FOA or planar FOA instead of the original input format. In some examples an original input format of HOA3 can be converted to MASA2 which would provide significant complexity savings. Other types of formats and conversions between the formats could be used in other examples, the input format can be converted to a different format before encoding so that the decoder receives the audio stream in a converted format. The converted format might be less complex to decode than the original format and so might enable a high level of immersion to be provided by a decoding device.
[0147] Fig. 9 shows another example communication session 100 that can implement examples of the disclosure. In this example the communication session 100 comprises more than two participants. In the example of Fig. 9 the communication session 100 comprises a multi-party immersive audio call between a first user equipment (UE) 102_1 , a second UE 102_2, a third UE 102_3 and a fourth UE 102_4. The UEs 102 are participants in the communication session 100.
[0148] The UEs 102 are configured to exchange audio signals via a server 900. For example, the UEs 102 are configured to send upstream signals to a server 900 and to receive an audio stream from the server 900. The server 900 could be an edge server, a multipoint control unit (MCU) server, a selective forwarding unit (SFU), or any other suitable type of server or device.
[0149] In Fig. 9 the server 900 is shown as a single entity. The server 900 could comprise multiple entities or components in some examples. The respective entities or components could be distributed within a network.
[0150] The server 900 can be configured to mix the upstream signals received from the UEs 102 to generate an audio stream. The server 900 can then provide the audio streams for the respective UEs 102 in the communication session 100. When examples of the disclosure are used in multi-party communication sessions the subsets of the set of codec operating points can be selected by multiple UEs 102 to enable selection of at least one codec operating point based, at least in part, on the computational complexity of multiple further participants. For example, the first UE 102_1 can indicate a set of codec operating points and then each of the other UEs 102_2, 102_3, 102_4 would select a subset from the originally indicated set.
[0151] When examples of the disclosure are used in multi-party communication sessions such as those shown in Fig. 9 the session negotiation can be configured to select a codec operating point so that all of the UEs 102 can be served at a level that provides high immersion but is also compatible with UE 102 with the lowest computational capability in the communication session.
[0152] In some examples, using the examples of the disclosure to select a suitable operating point can enable the server 900 to provide the immersive audio without transcoding. In other examples transcoding might be needed. If transcoding is needed, then selecting the operating point using the examples of the disclosure can enable transcoding to be performed for the fewest number of participants which can result in maximal reuse of the received immersive audio bitstream.
[0153] Fig. 10 schematically illustrates an apparatus 1000 that can be used to implement examples of the disclosure. In this example the apparatus 1000 comprises a controller 1002. The controller 1002 can be a chip or a chipset. In some examples the controller 1002 can be provided within a UE or within any other suitable device within a telecommunication system.
[0154] In the example of Fig. 10 the implementation of the controller 1002 can be as controller circuitry. In some examples the controller 1002 can be implemented in hardware alone, have certain aspects in software including firmware alone or can be a combination of hardware and software (including firmware).
[0155] As illustrated in Fig. 10 the controller 1002 can be implemented using instructions that enable hardware functionality, for example, by using executable instructions of a computer program 1008 in a general-purpose or special-purpose processor 1004 that can be stored on a computer readable storage medium (disk, memory etc.) to be executed by such a processor 1004.
[0156] The processor 1004 is configured to read from and write to the memory 1006. The processor 1004 can also comprise an output interface via which data and / or commands are output by the processor 1004 and an input interface via which data and / or commands are input to the processor 1004.
[0157] The memory 1006 is configured to store a computer program 1008 comprising computer program instructions (computer program code 1010) that controls the operation of the controller 1002 when loaded into the processor 1004. The computer program instructions, of the computer program 1008, provide the logic and routines that enables the controller 1002 to perform the methods illustrated in the Figs. The processor 1004 by reading the memory 1006 is able to load and execute the computer program 1008.
[0158] The apparatus 1000 therefore comprises: at least one processor 1004; and at least one memory 1006 including computer program code 1010, the at least one memory 1006 and the computer program code 1010 configured to, with the at least one processor 1004, cause the apparatus 1000 at least to perform: determining 300 a set of codec operating points indicated by a first participant in the communication session; obtaining 302 a subset of the set of codec operating points wherein the subset is indicated to the first participant from at least one further participant in the communication session; selecting 304 from the subset at least one codec operating point for the communication session which results in rendering with high immersion for at least one participant in the communication session wherein the selection is based, at least in part, on the computational capability of at least one participant in the communication session.
[0159] As illustrated in Fig. 10 the computer program 1008 can arrive at the controller 1002 via any suitable delivery mechanism 1012. The delivery mechanism 1012 can be, for example, a machine readable medium, a computer-readable medium, a non-transitory computer-readable storage medium, a computer program product, a memory device, a record medium such as a Compact Disc Read-Only Memory (CD-ROM) or a Digital Versatile Disc (DVD) or a solid state memory, an article of manufacture that comprises or tangibly embodies the computer program 1008. The delivery mechanism can be a signal configured to reliably transfer the computer program 1008. The controller 1002 can propagate or transmit the computer program 1008 as a computer data signal. In some examples the computer program 1008 can be transmitted to the controller 1000 using a wireless protocol such as Bluetooth, Bluetooth Low Energy, Bluetooth Smart, 6LoWPan (IPv6 over low power personal area networks) ZigBee, ANT+, near field communication (NFC), Radio frequency identification, wireless local area network (wireless LAN) or any other suitable protocol.
[0160] The computer program 1008 comprises computer program instructions that when executed by an apparatus 1000 cause the apparatus 1000 to perform at least the following: determining 300 a set of codec operating points indicated by a first participant in the communication session; obtaining 302 a subset of the set of codec operating points wherein the subset is indicated to the first participant from at least one further participant in the communication session; selecting 304 from the subset at least one codec operating point for the communication session which results in rendering with high immersion for at least one participant in the communication session wherein the selection is based, at least in part, on the computational capability of at least one participant in the communication session.
[0161] The computer program instructions can be comprised in a computer program 1008, a non-transitory computer readable medium, a computer program product, a machine readable medium. In some but not necessarily all examples, the computer program instructions can be distributed over more than one computer program 1008.
[0162] Although the memory 1006 is illustrated as a single component / circuitry it can be implemented as one or more separate components / circuitry some or all of which can be integrated / removable and / or can provide permanent / semi-permanent / dynamic / cached storage.
[0163] Although the processor 1004 is illustrated as a single component / circuitry it can be implemented as one or more separate components / circuitry some or all of which can be integrated / removable. The processor 1004 can be a single core or multi-core processor.
[0164] References to “computer-readable storage medium”, “computer program product”, “tangibly embodied computer program” etc. or a “controller”, “computer”, “processor” etc. should be understood to encompass not only computers having different architectures such as single / multi- processor architectures and sequential (Von Neumann) / parallel architectures but also specialized circuits such as field- programmable gate arrays (FPGA), application specific circuits (ASIC), signal processing devices and other processing circuitry. References to computer program, instructions, code etc. should be understood to encompass software for a programmable processor or firmware such as, for example, the programmable content of a hardware device whether instructions for a processor, or configuration settings for a fixed-function device, gate array or programmable logic device etc.
[0165] As used in this application, the term “circuitry” can refer to one or more or all of the following:
[0166] (a) hardware-only circuitry implementations (such as implementations in only analog and / or digital circuitry) and
[0167] (b) combinations of hardware circuits and software, such as (as applicable):
[0168] (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and
[0169] (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions and
[0170] (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g. firmware) for operation, but the software can not be present when it is not needed for operation. This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit for a mobile device or a similar integrated circuit in a server, a cellular network device, or other computing or network device.
[0171] The blocks illustrated in the Figs, can represent steps in a method and / or sections of code in the computer program 1008. The illustration of a particular order to the blocks does not necessarily imply that there is a required or preferred order for the blocks and the order and arrangement of the blocks can be varied. Furthermore, it can be possible for some blocks to be omitted.
[0172] Where a structural feature has been described, it may be replaced by means for performing one or more of the functions of the structural feature whether that function or those functions are explicitly or implicitly described.
[0173] The apparatus can be provided in an electronic device, for example, a mobile terminal, according to an example of the present disclosure. It should be understood, however, that a mobile terminal is merely illustrative of an electronic device that would benefit from examples of implementations of the present disclosure and, therefore, should not be taken to limit the scope of the present disclosure to the same. While in certain implementation examples, the apparatus can be provided in a mobile terminal, other types of electronic devices, such as, but not limited to: mobile communication devices, hand portable electronic devices, wearable computing devices, portable digital assistants (PDAs), pagers, mobile computers, desktop computers, televisions, gaming devices, laptop computers, cameras, video recorders, GPS devices and other types of electronic systems, can readily employ examples of the present disclosure. Furthermore, devices can readily employ examples of the present disclosure regardless of their intent to provide mobility.
[0174] The term ‘comprise’ is used in this document with an inclusive not an exclusive meaning. That is any reference to X comprising Y indicates that X may comprise only one Y or may comprise more than one Y. If it is intended to use ‘comprise’ with an exclusive meaning then it will be made clear in the context by referring to ‘comprising only one...’ or by using ‘consisting.’
[0175] In this description, the wording ‘connect’, ‘couple’ and ‘communication’ and their derivatives mean operationally connected / coupled / in communication. It should be appreciated that any number or combination of intervening components can exist (including no intervening components), i.e., to provide direct or indirect connection / coupling / communication. Any such intervening components can include hardware and / or software components.
[0176] As used herein, the term "determine / determining" (and grammatical variants thereof) can include, not least: calculating, computing, processing, deriving, measuring, investigating, identifying, looking up (for example, looking up in a table, a database, or another data structure), ascertaining and the like. Also, "determining" can include receiving (for example, receiving information), accessing (for example, accessing data in a memory), obtaining and the like. Also, " determine / determining" can include resolving, selecting, choosing, establishing, and the like.
[0177] In this description, reference has been made to various examples. The description of features or functions in relation to an example indicates that those features or functions are present in that example. The use of the term ‘example’ or ‘for example’ or ‘can’ or ‘may’ in the text denotes, whether explicitly stated or not, that such features or functions are present in at least the described example, whether described as an example or not, and that they can be, but are not necessarily, present in some of or all other examples. Thus ‘example’, ‘for example’, ‘can’, or ‘may’ refers to a particular instance in a class of examples. A property of the instance can be a property of only that instance or a property of the class or a property of a sub-class of the class that includes some but not all the instances in the class. It is therefore implicitly disclosed that a feature described with reference to one example but not with reference to another example, can where possible be used in that other example as part of a working combination but does not necessarily have to be used in that other example. As used herein, “at least one of the following:” and “at least one of ” and similar wording, where the list of two or more elements are joined by “and” or “or” mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements.
[0178] Although examples have been described in the preceding paragraphs with reference to various examples, it should be appreciated that modifications to the examples given can be made without departing from the scope of the claims.
[0179] Features described in the preceding description may be used in combinations other than the combinations explicitly described above.
[0180] Although functions have been described with reference to certain features, those functions may be performable by other features whether described or not.
[0181] The description of a feature, such as an apparatus or a component of an apparatus, configured to perform a function, or for performing a function, should additionally be considered to also disclose a method of performing that function. For example, description of an apparatus configured to perform one or more actions, or for performing one or more actions, should additionally be considered to disclose a method of performing those one or more actions with or without the apparatus.
[0182] Although features have been described with reference to certain examples, those features may also be present in other examples whether described or not.
[0183] The term ‘a’, ‘an’ or ‘the’ is used in this document with an inclusive not an exclusive meaning. That is any reference to X comprising a / an / the Y indicates that X may comprise only one Y or may comprise more than one Y unless the context clearly indicates the contrary. If it is intended to use ‘a’, ‘an’ or ‘the’ with an exclusive meaning then it will be made clear in the context. In some circumstances the use of ‘at least one’ or ‘one or more’ may be used to emphasis an inclusive meaning but the absence of these terms should not be taken to infer any exclusive meaning. The presence of a feature (or combination of features) in a claim is a reference to that feature or (combination of features) itself and to features that achieve substantially the same technical effect (equivalent features). The equivalent features include, for example, features that are variants and achieve substantially the same result in substantially the same way. The equivalent features include, for example, features that perform substantially the same function, in substantially the same way to achieve substantially the same result.
[0184] In this description, reference has been made to various examples using adjectives or adjectival phrases to describe characteristics of the examples. Such a description of a characteristic in relation to an example indicates that the characteristic is present in some examples exactly as described and is present in other examples substantially as described.
[0185] The above description describes some examples of the present disclosure however those of ordinary skill in the art will be aware of possible alternative structures and method features which offer equivalent functionality to the specific examples of such structures and features described herein above and which for the sake of brevity and clarity have been omitted from the above description. Nonetheless, the above description should be read as implicitly including reference to such alternative structures and method features which provide equivalent functionality unless such alternative structures or method features are explicitly excluded in the above description of the examples of the present disclosure.
[0186] Whilst endeavoring in the foregoing specification to draw attention to those features believed to be of importance the Applicant may seek protection via the claims in respect of any patentable feature or combination of features hereinbefore referred to and / or shown in the drawings whether or not emphasis has been placed thereon. l / we claim:
Claims
CLAIMS1. An apparatus for an immersive communication session, the apparatus comprising means for: determining a set of codec operating points indicated by a first participant in the communication session; obtaining a subset of the set of codec operating points wherein the subset is indicated to the first participant from at least one further participant in the communication session; and selecting from the subset at least one codec operating point for the communication session which results in rendering with high immersion for at least one participant in the communication session wherein the selection is based, at least in part, on the computational capability of at least one participant in the communication session.
2. An apparatus as claimed in claim 1 , wherein the at least one codec operating point is selected to provide high immersion for the first participant and one or more further participants.
3. An apparatus as claimed in any preceding claim, wherein multiple categories of immersion are available and selecting a codec operating point which results in rendering with high immersion comprises selecting the codec operating point in a category with highest available immersion.
4. An apparatus as claimed claim 3, wherein the multiple categories of immersion comprise at least some of: mono rendering; stereo rendering; spatial rendering; binaural rendering; binaural rendering with headtracking and with room reverb; binaural rendering with headtracking and without room reverb; binaural rendering without headtracking and with room reverb; binaural rendering without headtracking and without room reverb;multichannel loudspeaker rendering; multichannel 5.1 loudspeaker rendering; multichannel 5.1.2 loudspeaker rendering; multichannel 5.1.4 loudspeaker rendering; multichannel 7.1 loudspeaker rendering; multichannel 7.1.4 loudspeaker rendering; multichannel loudspeaker rendering with horizontal speakers; multichannel loudspeaker rendering with elevated speakers; multichannel loudspeaker rendering with frontal speakers; and IVAS supported output rendering formats.
5. An apparatus as claimed in any preceding claim, wherein an immersion parameter of the at least one further participant is indicated to the first participant.
6. An apparatus as claimed in any preceding claim, wherein the selection of the at least one codec operating point is based, at least in part, on the computational capability of the first participant and one or more further participants.
7. An apparatus as claimed in claim 6, wherein the computational capability of a participant is at least one of: indicated using one or more computational capability parameters; and inferred from information provided the by the at least one participant.
8. An apparatus as claimed in any preceding claim, wherein the set of codec operating points is determined based on one or more negotiation parameters.
9. An apparatus as claimed in claim 8, wherein the one or more negotiation parameters comprise at least one of: bitrate, bandwidth, input format, discontinuous transmission, output format, coded format, computational capability.
10. An apparatus as claimed in any preceding claim, wherein the selected at least one codec operating point uses one or more format conversions for an audio stream.
11. An apparatus as claimed in any preceding claim, wherein the first participant is a sending participant and the at least one further participant is a receiving participant.
12. An apparatus as claimed in any preceding claim, wherein the subset of the set of codec operating points is selected by multiple further participants to enable selection of the at least one codec operating point based, at least in part, on the computational complexity of the multiple further participants.
13. An apparatus as claimed in claim 12, wherein the at least one codec operating point is selected to enable rendering of an audio stream by the multiple further participants without transcoding the audio stream.
14. An apparatus as claimed in any preceding claim, wherein the communication session comprises a bidirectional immersive audio call.
15. A method comprising: determining a set of codec operating points indicated by a first participant in the communication session; obtaining a subset of the set of codec operating points wherein the subset is indicated to the first participant from at least one further participant in the session; and selecting from the subset at least one codec operating point for the communication session which results in rendering with high immersion for at least one participant in the communication session wherein the selection is based, at least in part, on the computational capability of at least one participant in the session.
16. An apparatus for an immersive communication session comprises: at least one processor; and at least one memory including computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to: determine a set of codec operating points indicated by a first participant in the communication session; obtain a subset of the set of codec operating points wherein the subset is indicated to the first participant from at least one further participant in the communication session; andselect from the subset at least one codec operating point for the communication session which results in rendering with high immersion for at least one participant in the communication session wherein the selection is based, at least in part, on the computational capability of at least one participant in the communication session.
17. An apparatus as claimed in claim 16, wherein the apparatus is caused to indicate an immersion parameter of the at least one further participant to the first participant.
18. An apparatus as claimed in any of claim 16 or 17, wherein the selection of the at least one codec operating point is based, at least in part, on the computational capability of the first participant and one or more further participants.
19. An apparatus as claimed in claim 18, wherein the computational capability of a participant is at least one of: indicated using one or more computational capability parameters; and inferred from information provided the by the at least one participant.
20. An apparatus as claimed in any of claims 16 to 19, wherein the apparatus is caused to determine the set of codec operating points based on one or more negotiation parameters.
21. An apparatus as claimed in claim 20, wherein the one or more negotiation parameters comprise at least one of: bitrate; bandwidth; input format; discontinuous transmission; output format; coded format; and computational capability.
22. An apparatus as claimed in any of claims 16 to 21 , wherein the apparatus is caused to select the at least one codec operating point based on one or more format conversions for an audio stream.
23. An apparatus as claimed in any of claims 16 to 22, wherein the apparatus is caused to select the subset of the set of codec operating points based on multiple further participants to enable selection of the at least one codec operating point based, at least in part, on the computational complexity of the multiple further participants.
24. An apparatus as claimed in claim 23, wherein the apparatus is caused to select the at least one codec operating point to enable rendering of an audio stream based on the multiple further participants without transcoding the audio stream.
25. An apparatus as claimed in any of claims 16 to 24, wherein the communication session comprises a bidirectional immersive audio call.
Citation Information
Patent Citations
Method for performing wi-fi display service and device for same
EP3104551A1
Adapting multi-source inputs for constant rate encoding
EP3923280A1