Device configuration based on configuration information for ivas

WO2026167536A1PCT designated stage Publication Date: 2026-08-13NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-02-03
Publication Date
2026-08-13

Smart Images

  • Figure IB2026051009_13082026_PF_FP_ABST
    Figure IB2026051009_13082026_PF_FP_ABST
Patent Text Reader

Abstract

An apparatus for immersive audio capture configuration, the apparatus comprising: at least one processor; and at least one memory storing instructions that when executed by the at least one processor, cause the apparatus at least to: obtain at least one of: codec configuration parameter; or renderer configuration parameter; configure at least one immersive audio capture parameter based on the obtained at least one of: the codec configuration parameter; or the renderer configuration parameter; and generate at least one immersive format audio signal based on the at least one configured immersive audio capture parameter.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] DEVICE CONFIGURATION BASED ON CONFIGURATION INFORMATION FOR IVAS

[0002] Field

[0003] The present application relates to apparatus and methods for a device configuration based on configuration information within Immersive Voice and Audio Services (IVAS) applications.

[0004] Background

[0005] Spatial audio can be captured and represented in various ways. For mobile communications, e.g., in the scope of 3GPP services, the smartphone is an example capture device and form factor.

[0006] The average size of a mono-block type mobile phone in recent years is over 15cm in length, over 7cm in width, and less than 1 cm in depth (or thickness). Higher-end devices with better capabilities are larger than the average, but not necessarily in terms of thickness.

[0007] Spatial audio capture capability in modern smartphones is based on integration of multiple MEMS microphones, where there can be variants using, e.g., dual microphone arrays, triple microphone arrays, or quad microphone arrays. In some examples, there can be more microphones, which benefits spatial audio capture, but 2 to 4 microphones are generally sufficient for the three levels of capture immersion discussed herein.

[0008] Additionally a spatial or immersive audio experience over headphones is possible by binaural reproduction (headphone playback) of a binaural recording or through a dedicated binauralization processing using a suitable rendering algorithm. As a result of binaural processing, a listener can thus experience an audio scene as if it surrounded them in real life.

[0009] However, the above experiences are currently static in nature, in that the scene follows a listener as they move about in real-life or rotate their head.Summary

[0010] There is provided according to a first aspect an apparatus comprising: at least one processor; and at least one memory storing instructions that when executed by the at least one processor, cause the apparatus at least to: obtain at least one of: codec configuration parameter; or renderer configuration parameter; configure at least one immersive audio capture parameter based on the obtained at least one of: the codec configuration parameter; or the renderer configuration parameter; and generate at least one immersive format audio signal based on the at least one configured immersive audio capture parameter.

[0011] According to a second aspect there is provided an apparatus comprising means configured to: obtain at least one of: codec configuration parameter; or renderer configuration parameter; configure at least one immersive audio capture parameter based on the obtained at least one of: the codec configuration parameter; or the renderer configuration parameter; and generate at least one immersive format audio signal based on the at least one configured immersive audio capture parameter.

[0012] According to a third aspect there is provided a method comprising: obtaining at least one of: codec configuration parameter; or renderer configuration parameter; configuring at least one immersive audio capture parameter based on the obtained at least one of: the codec configuration parameter; or the renderer configuration parameter; and generating at least one immersive format audio signal based on the at least one configured immersive audio capture parameter.

[0013] According to a fourth aspect there is provided a computer program comprising instructions [or a computer readable medium comprising program instructions] for causing an apparatus to perform at least the following: obtaining at least one of: codec configuration parameter; or renderer configuration parameter; configuring at least one immersive audio capture parameter based on the obtained at least one of: the codec configuration parameter; or the renderer configuration parameter; and generating at least one immersive format audio signal based on the at least one configured immersive audio capture parameter.

[0014] According to a fifth aspect there is provided a non-transitory computer readable medium comprising program instructions for causing an apparatus toperform at least the following: obtaining at least one of: codec configuration parameter; or renderer configuration parameter; configuring at least one immersive audio capture parameter based on the obtained at least one of: the codec configuration parameter; or the renderer configuration parameter; and generating at least one immersive format audio signal based on the at least one configured immersive audio capture parameter.

[0015] According to a sixth aspect there is provided a computer readable medium comprising program instructions for causing an apparatus to perform at least the following: obtaining at least one of: codec configuration parameter; or renderer configuration parameter; configuring at least one immersive audio capture parameter based on the obtained at least one of: the codec configuration parameter; or the renderer configuration parameter; and generating at least one immersive format audio signal based on the at least one configured immersive audio capture parameter.

[0016] Summary of the Figures

[0017] For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings in which:

[0018] Fig.1 shows schematically example systems within which embodiments may be implemented;

[0019] Fig.2 shows schematically a flow diagram of operations of the example system as shown in Fig.1 according to some embodiments;

[0020] Fig.3 shows an example IVAS RTF packet format configuration according to some embodiments;

[0021] Fig.4 shows schematically an example capture processor as shown in Fig .1 in further detail according to some embodiments;

[0022] Fig.5 shows schematically a flow diagram of operations of the example capture processor as shown in Fig.4 according to some embodiments;

[0023] Fig.6 shows schematically an example encoder as shown in Fig .1 in further detail according to some embodiments;

[0024] Fig.7 shows schematically a flow diagram of operations of the example encoder as shown in Fig.6 according to some embodiments;Figs.8 and 9 show example erroneous direction detection scenarios where the head of the listener is facing towards and at ninety degrees from an audio source;

[0025] Fig.10 shows schematically an example flow diagram of a direction ambiguity determination operation implementable by the capture processor according to some embodiments;

[0026] Fig.11 shows example capture processing analysis time-frequency configurations according to some embodiments;

[0027] Fig.12 shows a schematic representation of an example microphone arrangement for a spatial audio capture smartphone with 2, 3, or 4 microphones;

[0028] Fig.13 shows schematically further capture and processing time-frequency resolutions according to some embodiments;

[0029] Figs.14 and 15 show example improvements of the scenario shown with respect to Fig .9 following the application of the example embodiments;

[0030] Fig.16 shows example scene orientation limitations for a device with low front-back determination performance according to some embodiments; and Fig.17 shows schematically a device suitable for implementing the apparatus described herein.

[0031] Embodiments of the Application

[0032] Binaural rendering can be enhanced by employing head-tracking data, e.g., provided by the headphone device or other suitable tracking mechanisms. When head rotation is considered as part of the rendering, the experience can be more life-like, i.e., the reproduced scene remains in place and for example a source to the front of the listener moves to the left when the listener rotates their head to their right.

[0033] The communication from the capture end to the rendering or reproduction end can employ a codec following an IVAS standard by 3GPP which addresses the need for immersive audio services or experiences. IVAS specifically targets immersive audio for 3GPP conversational applications with support for a new parametric spatial audio format, Metadata- Assisted Spatial Audio (MASA), which is specifically developed for mobile phone capture. IVAS comes with a completeframework for mobile communications including Discontinuous Transmission (DTX) with Voice / Signal Activity Detection (VAD / SAD) and Comfort noise Generation (CNG), packet loss concealment (PLC), integrated Jitter Buffer Management (JBM), integrated renderer with default binaural filter sets (HRIR / BRIR), and split rendering capabilities enabling head-tracked binaural audio even on lightweight devices. Such immersive services include uses for example in immersive voice and audio for applications such as virtual reality (VR), augmented reality (AR) and mixed reality (MR) as well as spatial voice communication including teleconferencing. The IVAS codec can furthermore operate with low latency to enable conversational services as well as support high error robustness under various transmission conditions.

[0034] The following describes in further detail suitable apparatus and possible mechanisms for the provision of efficient capture apparatus and methods for IVAS audio and with respect to capture apparatus and method configuration based on codec or renderer configuration information.

[0035] The concept as discussed in the embodiments and examples herein is one of negotiation or communication of information between capture apparatus and rendering apparatus (or between capture apparatus and another entity within the signal flow network, e.g. a server) which assists in the determination of captured audio signals. This information can comprise the codec or renderer configuration parameters.

[0036] In some embodiments this information can be employed within the capture apparatus to provide a suitable format spatial audio signal (such as an IVAS MASA spatial audio signal) with a suitable configuration, for example frequency band / number of channels etc. These spatial audio signals can then be efficiently encoded by the IVAS encoder and is able to then be decoded and rendered by the renderer such that the embodiments attempt to result in a good output quality rendered audio signal.

[0037] An example system within which embodiments may be implemented is shown in Fig.1 .

[0038] Fig .1 , for example, shows an example immersive call or teleconferencing system within which some embodiments can be implemented. In this examplethere is shown two sites or rooms, Room A 100 and Room B 102. Room A 100 comprises a ‘talker’ or user, Talker 1 103. Room B 102 comprises one ‘talker’ or user, Talker RX 141.

[0039] In the following example within room A 100 is a suitable audio communication system 110, such as a teleconference apparatus (or more generally telecommunications apparatus 110 which can operate as capture apparatus) configured to spatially capture and encode the audio environment and furthermore is configured to render a spatial audio signal to the room. The apparatus 110 can in some embodiments be implemented by a user equipment (UE), such as a mobile communication device, or a smartphone operating within a communications system, e.g., a cellular communications system or accessing any suitable access network. The user equipment in some embodiments comprises a smartphone such as shown in Fig.12. In some further or additional use cases the apparatus 110 can be implemented in a vehicle.

[0040] Within each of the other spaces, e.g. rooms, may be a suitable audio communication device, such as a teleconference apparatus (or more generally telecommunications apparatus such as apparatus 120 within room B, which in the following examples can be designated as rendering apparatus) configured to render a spatial audio signal to the room. In some embodiments the apparatus 120 within room B can furthermore be configured to capture and encode at least a mono audio and optionally configured to spatially capture and encode the audio environment. The apparatus 120 can in some embodiments be implemented by a user equipment (UE), such as a mobile communication device, or a smartphone operating within a communication system, e.g., a cellular communications system or accessing any suitable access network. The user equipment in some embodiments comprises a smartphone such as shown in Fig.12. In some further or additional use cases the apparatus 120 can be implemented in a vehicle.

[0041] In the following examples each room is provided with the means to spatially capture, encode spatial audio signals, receive spatial audio signals and render these to a suitable listener. It would be understood that there may be other embodiments where the system comprises some apparatus configured to only capture and encode audio signals (in other words the apparatus is a ‘transmit’only or purely capture apparatus), and other apparatus configured to only receive and render audio signals (in other words the apparatus is a ‘receive’ only or purely rendering apparatus). In such embodiments the system within which embodiments may be implemented may comprise apparatus with varying abilities to capture / render audio signals.

[0042] The telecommunications apparatus (for each space, site or room) 110, 120 in this example can be configured to call into a teleconference controlled by and implemented over a server 111.

[0043] In some embodiments the communications or teleconferencing system comprises a (peer-to-peer) communications system (rather than the server based system shown in Fig.1) within which some embodiments can be implemented. Thus, for example, two or more UEs can be configured to interact directly with each other (for example to implement an immersive audio phone call between users) without the need for a central server.

[0044] For example, the communications or teleconferencing, e.g., the immersive audio phone call, between users can be an IMS (IP (Internet Protocol) Multimedia Subsystem) call over 4G (Generation) (LTE (Long Term Evolution)), 5G (NR (New Radio)), or 6G link or any other suitable link, for example, WLAN (Wireless Local Area Network).

[0045] The telecommunications apparatus or UEs 110, 120 can be configured to spatially capture and encode the audio environment and furthermore can be configured to render a spatial audio signal to the space. In this example only the communications path from the Room A 100 to the Room B 102 is shown for simplicity but a duplex or multipoint communication system comprising multiple signalling paths can be implemented using the methods as described herein without significant inventive input.

[0046] The teleconference apparatus (for each site or room) 110, 120 is further configured to communicate with each other to implement a teleconference function.

[0047] In some embodiments the user equipment 110 comprises a microphone array 115. The microphone array 115, can be, for example an array such as shown in Fig.12 which shows schematically example microphone integration fora spatial audio capture smartphone comprising 2, 3, or 4 microphones. The example smartphone 1230 is shown comprising a side 1218 which comprises a screen viewable by the user 1207 and a side 1216 which comprises the main camera module. There is also shown when the smartphone is in landscape orientation as shown in Fig.12 and with the side 1218 with the screen facing the user with a right side 1214, a left side 1212, a top side 1220, and a bottom side 1210. In the following examples with respect to the capturing of the audio signal a ‘front’ from the perspective of the user is with respect to the main camera module side (such as the side 1216), however as a user capturing an audio scene, e.g., on a video call capturing the environment such as a street performance or event, is using the main camera module and the microphone array. However, it is appreciated that in some embodiments, for example when the user is capturing the device with the ‘selfie’ or screen facing camera the front could be the side with the screen (such as the side 1218).

[0048] Furthermore Fig.12 shows some example microphone arrangements or configurations. For example, the smartphone 1230 can comprise an example dual mic array arrangement, where the placement of the two microphones is on rightside 1201 (Dual: R microphone 1201) and left side 1212 (Dual: L microphone 1203) of the device bezel. With two such microphones, a stereo audio signal can be captured that provides information to separate left and right directions. This information could for example be able to estimate a position for an audio or sound source on a 180-degree plane. This estimated position is only a 180-degree plane because it is assumed that the audio source is located to the front of the smartphone as there is a front-back ambiguity where it is not possible to determine whether the source is from the front or back of the smartphone with respect to the two microphone positions separated by an axis through L and R channels. On the other hand, considering rendering and presentation of spatial audio, it is generally preferable to have audio sources located in the front rather than back of the listener if full 360 capture and presentation is not achievable.

[0049] Furthermore, a triple mic array arrangement is shown where the dual mic array is supplemented with a rear-facing mic (the triple 1205 microphone which can for example be located as part of a camera bump integration). With threemicrophones and in this orientation, the front-back ambiguity can be resolved, and it is generally possible to derive information to provide an estimated position within a 360-degree plane. In other words, using three microphone audio signals to provide an actual azimuth position for the audio source that could be virtually placed in the audio scene.

[0050] A quad mic array configuration can further introduce a further microphone quad microphone 1207 on the left side 1212. Thus for example as shown in Fig .12 the quad microphone 1207 along with the Dual: L microphone 1203 effectively presents an arrangement with dual mics at the bottom bezel, which is closest to user’s mouth in typical handset (device on ear) and handheld hands-free (device in hand, at maximum at arm’s length in front of user) use cases when the device is held in the portrait orientation. With four microphones, an estimated position can estimate audio source positions also in elevation, in other words it is possible to describe the whole 3D space correctly at the capture point.

[0051] Furthermore Fig.1 shows a capture (front end) processor 107 configured to receive the audio signals 114 from the microphone array 115 and perform processing on these audio signals 114 to generate a suitable input format, for example a MASA spatial audio format, for the (IVAS) encoder 101. In some embodiments as described in further detail herein the capture (front end) processor 107 is configured by a controller 105.

[0052] An example IVAS MASA format generated by the capture processor 107 can be based on 1 or 2 audio channels and associated metadata, as described by the following tables (which are referenced by the references found in and further described in 3GPP TS 26.258 Annex A):

[0053] Table A.1 : MASA format descriptive common metadata parameters

[0054]

[0055]

[0056] Table A.2a: MASA format spatial metadata parameters (dependent of number of directions)

[0057]

[0058] Table A.2b: MASA format spatial metadata parameters (independent of number of directions)

[0059]

[0060]

[0061] Furthermore, MASA spatial metadata can be provided in each 20-ms frame for each analyzed / provided direction (1 or 2) for 4 subframes and 24 frequency bands, which are defined by the table below:

[0062] Table A.3. MASA spatial metadata frequency bands

[0063]

[0064]

[0065] Additionally, as shown in Fig.1 the UE 110 can comprise a controller 105 configured to control the capture (front end) processor 107 and in some embodiments the (IVAS) encoder 101. In some further embodiments the controller 105 can furthermore control the operation of the microphone array 115, for example, where at least some of the capture processor 107 functionality is incorporated or integrated into the microphone array (for example selection or activation of microphones to provide audio signals, or configuration of integrated analogue-to-digital conversion and / or frequency domain conversion).

[0066] As shown in Fig.1, the apparatus, UEs 110, 120 and server 111 can comprise suitable encoder and decoder functionality. For example, the apparatus 110 is shown comprising the (IVAS) encoder 101 , the server 111 is shown comprising an (IVAS) decoder and encoder 121 and the apparatus 120 is shown comprising an (IVAS) decoder 131. In such a manner the audio signals representing the user or talker 1 103 can be encoded by the encoder 101 which generates a bitstream 106 to be passed to the server 111. The server 111 can then decode, (optionally then mix, e.g., with other objects and otherwise process the audio signals) and encode then to generate the bitstream 108 to be passedto the apparatus 120. The apparatus 120 can then decode the audio signals and present them to the user or listener ‘RX’ 141.

[0067] Although this example shows a teleconference application the encoder / decoder functionality can be applied to the streaming of any suitable media.

[0068] The IVAS decoder / renderer for each of the UE 120, such as the teleconference apparatus 102 can be furthermore configured to handle multiple input streams that may each originate from a different encoder.

[0069] Fig.1 furthermore shows the room B apparatus or UE2 comprising a suitable (IVAS) renderer 135 configured to receive the decoded spatial audio signals 124 output from the (IVAS) decoder 131 and generate or render a suitable output audio signal, for example a binaural audio signal output, to be provided to suitable output apparatus, which in Fig .1 is shown as headphones (or a headset comprising suitable audio transducers) 143 worn by listener 141. In some embodiments the listener position and orientation can be tracked, for example by suitable headset sensors and this information passed back to the (IVAS) renderer 135 which uses this information in generating or rendering the output audio signals. Although this example shows the listener using headphones any suitable output or output device can be employed in other embodiments, for example UE2 can be configured to generate a multichannel audio signal format (for example 7.1 multichannel audio signals) to be output to a multi-speaker arrangement in the room. A suitable output device (102, 143 or a combination) can be, e.g., AR (augmented reality) or VR (virtual reality) headset. In another use case, the suitable output device (102, 143 or a combination) can be a vehicle having the multi-speaker arrangement with an optional display.

[0070] Furthermore, the UE 120 can comprise a controller 133 configured to control the operations of the (IVAS) decoder 131 and (IVAS) renderer 135. In some embodiments the controller 133 in the renderer apparatus 120 is configured to communicate with or negotiate with a controller 105 within the capture apparatus 110, in this example UE1 110. In some embodiments the controllers 133 and 105 can communicate directly with each other (or as shown in Fig.1 via a server 111 or central controller) as shown by bitstreams 150, 152. Further, insome other or additional embodiments, the (IVAS) encoder 101 and the (IVAS) decoder 131 can communicate directly with each other without an involvement of the (IVAS) decoder & (IVAS) encoder 121. In some embodiments this communication can be incorporated within the bitstreams 106 and 108. Further, in some other or additional embodiments, the server 111 can be omitted in the communication between the Room A apparatus (UE1) 110 and the Room B apparatus (UE2) 120. In this case, the (IVAS) encoder 101 and the (IVAS) decoder 131 can communicate directly with each other via the Bitstream as depicted on 106 and 108. In similar case, also the controllers 105 and 133 can communicate directly with each other via the Bitstream as depicted on 150 and 152, or alternatively via the Bitstream as depicted on 106 and 108.

[0071] In some embodiments this communication can be the information identifying at least one codec configuration parameter or rendering configuration parameter. For example, the codec configuration parameters can comprise, for example: coded format(s), bitrate, and signal bandwidth. Also, the renderer configuration parameters can, for example, comprise: type of rendering (loudspeaker, binaural) and availability of head-tracking capability.

[0072] In some embodiments the information is exchanged as part of a codec negotiation (e.g., SDR (Session Description Protocol) offer-answer). The information can however be provided by any suitable signaling during the call, e.g., RTF PI data frame indicating activation of head-tracking.

[0073] The negotiation or communication between the capture apparatus and the rendering or reproduction apparatus for both the audio signal bitstreams 106, 108 and the control information 150, 152 can thus in some embodiments employ RTP (Real-Time Transport Protocol) for an end-to-end, real-time transfer of streaming media and provides facilities for jitter compensation and detection of packet loss and out-of-order delivery. RTP allows data transfer to multiple destinations through IP (Internet Protocol) multicast or to a specific destination through IP unicast. The majority of the RTP implementations are built on top of the User Datagram Protocol (UDP). Other transport protocols may also be utilized. RTP is used in together with other protocols such as H.323 and Real Time Streaming Protocol (RTSP).The RTP specification describes two protocols: RTP and RTCP. RTP is used for the transfer of multimedia data, and its companion protocol (RTCP) is used to periodically send control information and QoS (Quality of Service) parameters.

[0074] RTP sessions are initiated between client and server or between client and another client (or a multi-party topology) using a signalling protocol, such as H.323, the Session Initiation Protocol (SIP), or RTSP. These protocols use the Session Description Protocol (SDP), such as defined by RFC 8866 to specify parameters for the sessions.

[0075] The IVAS codec algorithm which can be employed is described in further detail in 3GPP TS 26.253 (Codec for Immersive Voice and Audio Services; Detailed Algorithmic Description incl. RTP payload format and SDP parameter definitions).

[0076] The Immersive codec operation can for example be initialized based on codec negotiation, where, e.g., the immersive format(s) for the call are offered and selected. IVAS codec negotiation is governed by the relevant specifications, e.g., TS 26.253 Annex A and TS 26.114. Currently, in 3GPP Rel-18, there are many SDP parameters and new SDP parameters can be expected to be added in future updates to allow for enhanced operation of the features.

[0077] In addition to a used format (which determines whether spatial audio capture is needed) the immersive audio bitrate and bandwidth information is useful for the optimization of the spatial audio capture algorithm.

[0078] These information elements are summarized from TS 26.253 Annex A: ibr: Specifies the range of source codec bitrate for IVAS Immersive mode in the session, in kilobits per second, for the direction specified by the session directionality attribute or the suffix. The ibr parameter can either have: a single bitrate (ibr1); or a hyphen-separated pair of two bitrates (ibr1-ibr2). If a single value is included, this bitrate, ibr1, is used. If a hyphen-separated pair of two bitrates is included, ibr1 and ibr2 are used as the minimum bitrate and the maximum bitrate respectively. ibr1 shall be smaller than ibr2. ibr 1 and ibr2 have a value from the set in Table 4.2-2 of the TS 26.253 Annex A. If none of theseparameters is present, all bitrates consistent with the IVAS codec capabilities are allowed in the session.

[0079] ibr-send / ibr-recv: ibr parameter in send or receive direction.

[0080] ibw: Specifies the audio bandwidth for IVAS Immersive modes to be used in the session, for the direction specified by the session directionality attribute or the suffix, ibw has a value from the set: wb, swb, fb, wb-swb, and wb-fb. wb, swb, and fb represent wideband, super-wideband, and fullband respectively, and wb-swb, and wb-fb represent all bandwidths from wideband to super-wideband, and fullband respectively. If none of these parameters is present, all bandwidths consistent with the negotiated bitrate(s) are allowed in the session.

[0081] ibw-send / ibw-recv: ibw parameter in send or receive direction.

[0082] cf: Specifies the IVAS Immersive mode coded-format (cf) transmitted in the IVAS Immersive mode frames in the session. IVAS coded format corresponds to the format represented in the IVAS Immersive mode coded frames, which is generally the input format to the encoder. The cf parameter is a list of supported comma-separated IVAS Immersive mode coded formats in the order of preference, using the identifiers from Table A.4.1-1 of TS 26.253 Annex A (column "Identifier"). Selection of the format is application-specific and out of scope of this document. EVS (Enhanced Voice Services) frames in the session are in mono format; switching to mono shall be possible.

[0083] cf-send / cf-recv: cf parameter in send or receive direction. If the cf-recv parameter is not present and not otherwise specified by cf, all IVAS coded formats consistent with the negotiated bitrate(s) are allowed in the session in receive direction.

[0084] Table 4.2-2 from TS 26.253:

[0085] Table 4.2-2 Supported audio bandwidth per input audio format and bitrate

[0086]

[0087]

[0088]

[0089]

[0090] Modified Table A.4.1-1 from TS 26.253 Annex A:

[0091] Table A.4.1-1 : IVAS coded-format (“Clause” column removed)

[0092]

[0093] In the above table Mono is not listed as an IVAS Immersive mode coded-format as EVS is defined to be supported and shall be used for mono.

[0094] For the offer-answer model it is further considered:

[0095] ibr: When the same bitrate or bitrate range is defined for the send and the receive directions, ibr should be used but ibr-send and ibr-recv may also be used, ibr can be used even if the session is negotiated to be sendonly, recvonly, or inactive. For sendonly session, ibr and ibr-send can be interchangeably used. For recvonly session, ibr and ibr-recv can be interchangeably used. When ibr is not offered for a payload type, the answerer may include ibr for the payload type in the SDP answer. When ibr is offered for a payload type and the payload type is accepted, the answerer shall include ibr in the SDP answer which shall be identical to or a subset of ibr for the payload type in the SDP offer.ibr-send: When ibr-send is not offered for a payload type, the answerer may include ibr-recv for the payload type in the SDP answer. When ibr-send is offered for a payload type and the payload type is accepted, the answerer shall include ibr-recv in the SDP answer, and the ibr-recv shall be identical to or a subset of ibr-send for the payload type in the SDP offer.

[0096] ibr-recv: When ibr-recv is not offered for a payload type, the answerer may include ibr-send for the payload type in the SDP answer. When ibr-recv is offered for a payload type and the payload type is accepted, the answerer shall include ibr-send in the SDP answer, and the ibr-send shall be identical to or a subset of ibr-recv for the payload type in the SDP offer.

[0097] ibw: When the same bandwidth or bandwidth range is defined for the send and the receive directions, ibw should be used but ibw-send and ibw-recv may also be used, ibw can be used even if the session is negotiated to be sendonly, recvonly, or inactive. For sendonly session, ibw and ibw-send can be interchangeably used. For recvonly session, ibw and ibw-recv can be interchangeably used. When ibw is not offered for a payload type, the answerer may include ibw for the payload type in the SDP answer. When ibw is offered for a payload type and the payload type is accepted, the answerer shall include ibw in the SDP answer, which shall be identical to or a subset of ibw for the payload type in the SDP offer.

[0098] ibw-send: When ibw-send is not offered for a payload type, the answerer may include ibw-recv for the payload type in the SDP answer. When ibw-send is offered for a payload type and the payload is accepted, the answerer shall include ibw-recv in the SDP answer, and the ibw-recv shall be identical to or a subset of ibw-send for the payload type in the SDP offer.

[0099] ibw-recv When ibw-recv is not offered for a payload type, the answerer may include ibw-send for the payload type in the SDP answer. When ibw-recv is offered for a payload type and the payload is accepted, the answerer shall include ibw-send in the SDP answer, and the ibw-send shall be identical to or a subset of ibw-recv for the payload type in the SDP offer.

[0100] cf: When the same IVAS Immersive mode coded formats are defined for the send and the receive directions, cf should be used but cf-send and cf-recvmay also be used. For sendonly session, cf and cf-send can be interchangeably used. For recvonly session, cf and cf-recv can be interchangeably used.

[0101] cf-send: The SDR offer shall contain the cf-send parameter and list at least one but may list several IVAS Immersive mode coded formats. The SDR answer shall include at least one IVAS Immersive mode coded format in cf-recv or and should respond with the one most preferred coded format from the list in the SDR offer. If more than one format is present in the SDR answer, the first format shall be used at the start of a session and may only be modified by the adaptation mechanisms present in this specification. When cf-send is offered for a payload type and the payload type is accepted, the answerer shall include cf-recv in the SDR answer, and the cf-recv shall be identical to or a subset of the cf-send parameter for the payload type in the SDR offer. If cf-recv is not offered for a payload type, cf-send in the answer may indicate any coded format.

[0102] cf-recv When cf-recv is offered for a payload type and the payload type is accepted, the answerer shall include cf-send in the SDR answer, and the cf-send shall be identical to or a subset of the cf-recv parameter for the payload type in the SDR offer.

[0103] In addition, SDR parameters relating to PI frame usage (pi-types, pi-types-send, pi-types-recv) as discussed below can be employed to determine whether PI data frames can be used during the call.

[0104] Although the above relates to IVAS, it would be appreciated that this information can be applied to other immersive codecs.

[0105] For a class of applications (e.g., audio, video, control), an RTP profile may be defined. For a media format (e.g., a specific video coding format), an associated RTP payload format may be defined. Every instantiation of RTP in a particular application may therefore require a profile and payload format specifications.

[0106] The profile is configured to define the codec used to encode the payload data and the mapping to payload format codes in the protocol field Payload Type (PT) of the RTP header.

[0107] An RTP session can be established for each multimedia stream. Audio, control and video streams may be implemented which use separate RTPsessions, enabling a receiver to selectively receive components of a particular stream. The RTP specification can furthermore be configured to recommend port numbers for RTP, and furthermore to recommend the use of the next odd port number for the associated RTCP session. A single port can be used for RTP and RTCP in applications that multiplex the protocols.

[0108] Each RTP stream can comprise RTP packets, and the RTP packet in turn can comprise a RTP header and payload pair.

[0109] Fig.3 shows an example IVAS RTP packet structure (following TS 26.253 Annex A) 300. The RTP Header (and possible RTP Header Extension) 301 follows a known RTP design. The structure 300 further comprises a IVAS payload 311. The IVAS payload 311 comprises a payload header 303 section, a (IVAS) frame data 205 section and an optional PI (processing information) data 307 section.

[0110] The payload header 303 section can comprise different types of header bytes, for example: ToC (Table of Content) and E-bytes (Extra bytes, or further bytes which can be used to define or signal aspects of the payload).

[0111] The ToC bytes can be employed to describe the content of the frame data section (by indicating the size / bitrate for the data frames). The ToC bytes can also differentiate IVAS frames from EVS frames, in situations where the IVAS is operating in mono EVS mode.

[0112] The E-bytes can signal additional information, such as Codec Mode Requests (CMR) which indicate a request to change or signal at least one codec configuration parameter or rendering configuration parameter. The E-bytes may also be used to explicitly indicate the presence of the PI data section at the end of the pay load.

[0113] The frame data 305 section can include the IVAS data frames (and EVS data frames in a case of employing IVAS in mono mode). The data frames represent the encoded IVAS bitstreams. The bitstream includes the encoded IVAS audio data with possible additional metadata. The bitstream can also include initialization data or information, for example information about the input / coded format and sub-format of the encoded data (e.g., multichannel 5.1 format). Any EVS frames do not include such format data.The PI data 307 section can include (PI) processing information data and related headers. The PI data can be used to transmit any (non-audio) data that can be used to assist the rendering or processing of the audio data, such as, for example, scene and device orientation data. The PI data can also be used to request something from the other session participant, such as, for example to mute an incoming stream or increase noise suppression. Feedback data can also be transmitted, for example, the head orientation of a listener. As such the PI data can be used to transmit between the apparatuses the least one codec configuration parameter or rendering configuration parameter.

[0114] Although current IVAS Rel-18 specifications do not specify reverse direction PI types, they can be employed to provide renderer / codec configuration parameters as discussed herein. For example, the following table describes two potential reverse direction PI types.

[0115]

[0116] The above reverse direction PI types enable the capture device and capture processor, to know which type of output is targeted (e.g., mono, stereo, multi-channel loudspeaker playback, binaural) and / or whether head-tracking capability is active. Such information can be used to configure the capture processor in order to provide an improved spatial audio input format content to the (IVAS) encoder in order to get best possible audio quality.

[0117] For example, for some device / transmission scenario / use case it can be beneficial to provide only stereo if no headtracking is utilized and an immersive format when headtracking is used. Thus, for example, a triggering of format switching could be based on such information.

[0118] With respect to Fig.2 is shown example operations with respect to a summary of the operations of the system as shown in Fig.1 according to some embodiments.Thus, for example, as shown by 201 is the operation of communicating / negotiating between the at least two apparatus: codec configuration parameter (information) or renderer configuration parameter (information). Although in this example there is one negotiation / communication operation acting as an initialization operation it would be understood that in some embodiments this communication / negotiation operation can occur during a communication potentially requiring an updating of the parameter, or modification of a communication session or restart of the communication session.

[0119] Then, as shown by 203, is the operation of receiving or otherwise obtaining microphone array audio signals (and other microphone array related information such as microphone positions, captured audio signal bitrate etc).

[0120] Furthermore, as shown by 205, is the operation of determining control information (which may be designated immersive audio capture parameters) based on the obtained at least one of: the codec configuration parameter; or the renderer configuration parameter.

[0121] As shown in 207 is the operation of performing front-end processing of the microphone array audio signals to generate (MASA) spatial audio signals based on the control information (immersive audio capture parameter).

[0122] Then, as shown by 209, is the operation of encoding the spatial audio signals (based on the control information) to generate an encoded immersive stream.

[0123] The encoded immersive stream can then be output (or stored), as shown by 211, and optionally with other microphone array information such as the microphone array position).

[0124] Following this operation can be the complimentary operation of receiving or retrieving the encoded immersive stream as shown by 221. In some embodiments there further is the operation of receiving or otherwise obtaining microphone information such as microphone array positions and rendering information such as listener position and / or orientation.

[0125] Then, as shown by 223 is the operation of decoding the bitstream to obtain the spatial audio signal.Then there is the operation, as shown by 225, of rendering an output (binaural) audio signal based on the decoded bitstream (spatial audio signal) and optionally the rendering information and microphone information

[0126] Finally, there is the operation, as shown by 227, of presenting the output (binaural) audio signal.

[0127] With respect to Fig.4 is shown schematically an example capture (front end) processor 107 suitable for employing in some embodiments, Fig.5 shows the operation of the example capture (front end) processor 107 according to some embodiments.

[0128] Thus, for example, the capture (front end) processor 107 is configured to receive or otherwise obtain the microphone array audio signals 114 and furthermore the immersive audio capture parameter 112.

[0129] The microphone array audio signals 114 and the immersive audio capture parameter 112 can be passed to a transport signal generator 401 that is configured to generate suitable (MASA) transport audio signals 402 by processing the microphone array audio signals 114, the processing based on the immersive audio capture parameter 112. For example, in some embodiments the transport signal generator 401 is configured to select at least one of the microphone array audio signals 114 to output as the transport audio signals. In some embodiments the number of channels, the selection criteria etc used to generate the transport audio signals is determined based on the immersive audio capture parameter 112 (and optionally based on other microphone information such as microphone array position). In some embodiments the transport signal generator 401 is configured to mix, downmix, or otherwise combine microphone audio signals to generate the transport audio signals 402. For example, the immersive audio capture parameter 112 may provide information that there are expected only two transport audio signals, and with the microphone array position information that the smartphone is in landscape mode select the audio signals from the Dual: R microphone 1201 and Dual: L microphone 1203 to be output.

[0130] The microphone array audio signals 114 and the immersive audio capture parameter 112 can also be passed to a metadata determiner 403 configured to generate suitable (MASA) spatial metadata 404 (as discussed above) byprocessing and / or analysing the microphone array audio signals 114, the processing of the microphone array audio signals 114 is based on the immersive audio capture parameter 112. For example, as is detailed later herein, there can be a situation where a three or four microphone equipped user equipment (e.g. the capture apparatus 110) generates for multiple frequency bands a direction parameter associated with an estimate position of an audio source. Errors in such direction parameter can in some situations be tolerated but in other situations can result in rendered or output audio signals which create an unnatural effect. In some embodiments therefore the metadata determiner 403 can implement an ambiguity detection and correction processing operation for some renderer or codec configurations (for example where the listener indicates that they are equipped with a headset with headtracking capability) but not for other configurations.

[0131] The capture stream output 116 can therefore comprise the generated transport audio signals 402 and the spatial metadata 404.

[0132] With respect to Fig.5 the example operations of the capture (front end) processor 107 according to some embodiments is shown.

[0133] Thus, for example, as shown by 501 is the operation of receiving the microphone array audio signals and the immersive audio capture parameter.

[0134] Then, as shown by 503, is the operation of generating transport audio signals based on the microphone array audio signals and the immersive audio capture parameter.

[0135] Furthermore, as shown by 505, is the operation of generating spatial metadata based on the microphone array audio signals and the immersive audio capture parameter.

[0136] Finally, as shown by 507, is the operation of outputting (MASA) spatial audio signals comprising the spatial metadata and the transport audio signals.

[0137] With respect to Fig.6 is shown schematically an example (IVAS) encoder 101 suitable for employing in some embodiments, Fig.7 shows the operation of the example (IVAS) encoder 101 according to some embodiments.

[0138] The (MASA) spatial audio signals comprising the spatial metadata 404 and transport audio signals 402, and control information (for example a determinedbitrate 610 and optionally the immersive audio capture parameter 112) can be passed to the (IVAS) encoder 101 which is configured to output the bitstream 106.

[0139] In some embodiments the encoder comprises a bitrate allocator / encoder controller 601 configured to receive the bitrate 610 and optionally the immersive audio capture parameter 112 and generate suitable audio bitrates 612 and metadata bitrates 614 which are passed to the transport audio signal encoder 603 and metadata encoder 605.

[0140] The transport audio signal encoder 603 is configured to receive the transport audio signals 402 and the audio bitrate 612 and generate encoded transport audio signals 602 which can be passed to the multiplexer 607.

[0141] The metadata encoder 605 is configured to receive the spatial metadata 404 and the metadata bitrate 614 and generate encoded metadata 604 which can be passed to the multiplexer 607.

[0142] The multiplexer 607 is configured to receive or obtain the encoded transport audio signals 602 and the encoded metadata 604 and multiplex them to form the bitstream 106.

[0143] With respect to Fig.7 the example operations of the encoder 101 according to some embodiments is shown.

[0144] Thus, for example, as shown by 701 is the operation of receiving the spatial audio signals comprising the spatial metadata and transport audio signals and control information (for example bitrate and optionally immersive audio capture parameter).

[0145] Then, as shown by 703, is the operation of determining bitrate allocations for the transport audio signals and the spatial metadata based on control information.

[0146] Furthermore, is the operation, as shown by 705, of encoding the transport audio signals and spatial metadata based on the bitrate allocations.

[0147] Finally, there is the operation, as shown by 707, of multiplexing the encoded transport audio signals and the encoded spatial metadata to generate the encoded bitstream and outputting the encoded bitstream.

[0148] Having established example apparatus suitable for implementing some embodiments with respect to capture processing microphone audio signals basedon codec or Tenderer configuration information the following discusses example scenarios within which such apparatus can be applied.

[0149] For example, as discussed above a typical current smartphone can have multiple omnidirectional MEMS microphones integrated, for example, at the top and bottom bezels of the device - right-side 1201 (Dual: R microphone 1201) and left side 1212 (Dual: L microphone 1203). In one example IVAS MASA spatial audio capture scenario, e.g., while streaming a video with spatial audio of a band playing music in a park, or during a video call, the device is in a landscape orientation and the Dual: R microphone 1201 audio signal and Dual: L microphone 1203 audio signal are selected by the transport signal generator 401 to generate stereo channels of a stereo-MASA transport audio signal.

[0150] In addition, the metadata determiner 403 uses the audio signals from these two omnidirectional microphones to generate MASA spatial metadata, but are unable to provide the information for front-back separation (and resolve the front-back ambiguous estimate).

[0151] There can be sounds to the front and to the back. The front and back sounds create about the same left-right correlation and other audio features as picked up by the left-hand side and right-hand side microphones, which means that the sounds emitted from front and back appear similar for the capture algorithms based on these channels only. This, of course, includes sounds that even arrive at the left and right microphones substantially at the same time, respectively, for example in a case where there are simultaneous sound sources at front and back.

[0152] As also shown in the example on Fig.12 a ‘thin’ smartphone device, can be equipped with at least a third microphone that is displaced slightly from the axis between left and right channel microphones. Based on those one or more microphones, the metadata determiner 403 can be configured to decide whether a sound comes from the front or back of the smartphone or, for example, both the front and the back at the same time.

[0153] However, this information can be very noisy, especially at lower frequencies, due to the very small separation on the front-back axis. In practical scenarios, there can also be other complicating factors such as shadowing by theuser’s hands, etc. that can be significantly time-varying and differs between use cases and individual user preferences (for example how the user likes to hold the smartphone).

[0154] Considering that the microphones are not directly located on the front-back axis but there can be an angle on the azimuth as well as on the elevation and different devices have different placing, a single algorithm or method to resolve the ambiguity can be difficult to optimize to cover all possible devices and use cases.

[0155] In many current implementations, the left-right microphone pair audio signals are employed to determine a relative azimuth value as well as the direct-to-total ratio MASA spatial metadata. However, the front-back ambiguity is unresolved at this stage with a default direction being the ‘main camera’ direction (or active camera direction) as this is what the user is ‘looking’ at (or where the user appears). In some situations, front-back information from the at least third microphone can then be generated to resolve the ambiguity (or provide a binary decision to flip direction instead of the default side).

[0156] In addition, in some implementations the metadata determiner 403 is configured to modify the direct-to-total ratio value to account for some fluctuations of the direction information.

[0157] Considering the above, many current smartphones where spatial audio capture can be implemented will have some errors in the front-back data during the capture. These errors are then part of the spatial audio signals generated as MASA format (or synthesized into ambisonics in other examples) that is fed to the IVAS encoder. The (IVAS) encoder 101 , depending on the bitrate and signal characteristics, modifies the data further, quantizes and encodes it, and creates a bitstream. This bitstream is received (potentially with some packet loss and jitter) and is decoded by the (IVAS) decoder 131. When the spatial audio is then rendered by the renderer 135, there can be perceived issues for the listener due to the erroneous front-back data and how this interacts with the IVAS coding.

[0158] For example, as shown in Fig.8, when a listener 807 is listening to a binaural rendering of a talker who is on the front, some sounds can be erroneously rendered coming from the back instead. For example, in Fig.8 there are illustrated4 temporal instances 821 and a number of frequency bands 831 (low band, mid band, high band) for a source that is in front of the listener 830. However, the high bands 851 for 3 time instances 841 have erroneously 820 been determined to be emitted from behind the capture device (and thus listener). In general, this is not particularly bad in this case, as the listener’s ears act a little like the left-hand side and right-hand side microphones as explained above and the renderer front lefthand 813, and behind left-hand 803 signals combine and the renderer front righthand 811 , and behind right-hand 801 signals combine. For example, the listener perceives a centrally located source or sources. While the listener may have difficulty to tell whether the source(s) appears in the front or the back, the experience remains fairly natural.

[0159] However, more perceivable errors are introduced when head-tracking is implemented, and the listener turns their head. For example, as illustrated in Fig .9, the listener 907 turns their head by about 90 degrees on the azimuth during rendering and playback of the same error segment as above. Now the renderings to listener’s left ear via 911 , 901 and right ear via 913, 903 are significantly different since the correct source 830 and the incorrect source 820 are positioned differently relative to the listener. This can be perceived as a very disturbing impairment by the listener, since it is now apparent that there is a significant difference between the binaural signals to the left and right ear. For example, a source (such as talker) that should appear from one direction is clearly present in two directions simultaneously. Furthermore, such behavior can be very timevarying, e.g., providing a sense of source position jumping around.

[0160] In addition, there could in some cases be some asynchrony between audio channels and the generated metadata (e.g., for some implementation reasons) as both audio and metadata is required to be framed correctly for an encoder. Such asynchrony could make the above front-back problem more pronounced.

[0161] Consider the listener without head-tracking or simply facing front, the transport audio signals are in this case quite close to the target binaural signals. In this case, the renderer does not need to perform processing, e.g., modify the stereo signals as part of the rendering (e.g., creation of prototypes). When headtracking is used and listener is turning their head, e.g., at 90 degrees, the transportsignals can be very different to the target binaural signals. In this case, the renderer is required perform significant processing of the audio signals. As this is implemented based on the metadata, any asynchrony can result in audible difference, where a front-back error is then more pronounced.

[0162] As such this can result in a significant issue relating to front-back errors in capture and their interaction with the IVAS codec and rendering.

[0163] The embodiments herein can attempt to produce the following improvements:

[0164] 1. Reducing / removing front-back capture errors

[0165] 2. Reduce any device orientation effect on audio capture and format generation, where different configurations of the capture microphones and generated format (e.g., mono-MASA, stereo-MASA) need to be considered for high-quality spatial audio reproduction

[0166] 3. Capture of complex sound scenes with low-bitrate coding, where reproduced scene can become chaotic due to fluctuating directions of sources

[0167] The embodiments can therefore relate to spatial audio capture and generating an immersive audio encoder input format. In this example, with respect to the front-back capture error situation the apparatus and methods attempt to adapt spatial audio directional information relating to determining front and back direction discrepancies at least partly based on at least one of an audio encoder configuration or a binaural rendering configuration in order to improve the perceptual quality for the listener.

[0168] In some embodiments the metadata determiner 403 can be configured to determine at least one representative or combined spatial direction based on energy-weighted direct-to-total energy ratios in each captured frame of spatial audio. The at least one representative spatial direction can then be employed to weight the front-back ambiguity determination for each individual direction where a front-back discrepancy may be detected. Although the following uses the term representative spatial direction it could be otherwise be designated an averagedirection, weighted average direction, energy weighted average direction, or combined direction.

[0169] The front-back direction, direct / diffuse energy, and / or surround coherence parameter value of the direction associated with the front-back discrepancy can then be modified.

[0170] Furthermore, the metadata determiner 403 can be configured to control these operations based at least partially the immersive audio capture parameter 112 or similar generated from at least one of codec configuration or renderer configuration.

[0171] The obtaining can, as discussed above be part of a negotiation or communication, for example a codec negotiation (SDR offer-answer) or any other suitable mechanism.

[0172] Based on codec configuration, the IVAS internal MASA band information can be used as part of the spatial audio analysis to reduce computational complexity and provide higher accuracy and more stable spatial representation of the codec input data and, therefore in the end, of the presentation at rendering and also employed as described herein for analysis correction.

[0173] In some embodiments, with respect to the analysis correction, there can be more than one adaption direction per time-frequency tile. This adaptation can be any of the following:

[0174] generation of a second direction;

[0175] reduction of two directions into one joint direction; or adaptation of at least one of the two directions including: front-back direction, direct / diffuse energy, spread coherence, and / or surround coherence parameter values.

[0176] In some embodiments some scene orientation signaling is limited according to codec configuration or renderer configuration.

[0177] In such a manner it can be possible to attempt to optimize the audio quality of the spatial audio capture output, where the specific audio codec configuration and audio renderer configuration affecting the spatial audio rendering in the erroneous cases is determined and employed to assist the spatial audio capture operations. These spatial audio capture processing operations can for examplecomprise one or more of modifying the capture analysis based on configuration parameters (such as coded format, bitrate, bandwidth), and use of binaural headtracking (or multi-channel loudspeaker layout).

[0178] In some embodiments the capture processing attempts to overcome the issue discussed above with respect to defining a representative direction (or ambiguity decision) based on combining energy-weighted direct-to-total energy ratios for directions across a number of analysis bands or, alternatively, identify a number of frequency bands that differ from those associated to generated audio format as given in Table A.3 (of TS 26.253, as shown above) for one or more time instances, where the time instance can be, e.g., a subframe.

[0179] A representative direction can thus be understood as a general direction as found by the spatial analysis. The representative direction can be determined, e.g., by combining over at least two frequency bands or time instances (subframes or frames) at least two directions values by averaging their directions by combining their energy-weighted direct-to-total energy ratios. This can be implemented for all of the bands / time instances covered by the above range or only for those whose direction value is within a pre-determined threshold of a targeted representative direction, where the targeted representative direction can be, e.g., the representative direction of band N, where, e.g., bands N-1 and N+1 are further combined with band N.

[0180] In some embodiments the capture processor is configured to find and correct (potential) errors of direction information in the front-back ambiguity determination or finding or correcting errors in any front-back discrepancy, where a direction, e.g., associated with a TF (Time-Frequency) tile, would be expected to be a front or back direction relative, e.g., to a representative direction but is instead a back or front direction, respectively.

[0181] Fig .10 shows schematically a flow diagram of the operations of the capture processor according to some embodiments when applied to the specific scenario of front-back ambiguity determination.

[0182] For example, in some embodiments an immersive audio call is negotiated (according to relevant specifications) between at least one sending side (for example UE1 110) and one receiving side (for example UE2 120). As part of thenegotiation, as shown by 1001 , at least one codec configuration or rendering configuration parameter is obtained.

[0183] Then as shown by 1003, at least three audio signals are provided, for example from a microphone array, or obtained otherwise.

[0184] Then based on a suitable spatial audio capture algorithm (for example by employing the capture processor shown in Fig.1) and potentially controlled or adapted by the at least one codec and / or renderer configuration parameter at least one transport audio signal is obtained as shown by 1005. In some embodiments the generated at least one transport audio signal is adapted or modified from the received or obtained audio signals based on at least one codec and / or renderer configuration parameter.

[0185] Furthermore, is the operation, as shown by 1007, of performing spatial audio scene analysis to determine spatial audio metadata per frequency bands or time instances including at least one of: direction, direct energy (for example direct-to-total energy ratio, energy-weighted direct-to-total energy ratio), diffuse energy (for example diffuse-to-total energy ratio), spread coherence, surround coherence.

[0186] As shown by 1009, is the operation of determining at least one combined direction over at least two frequency bands or time instances by combining (energy-weighted) direct energy values relating to the bands or instances.

[0187] Then as shown by 1011 is the operation of adapting the determined at least one direction, direct energy, diffuse energy, spread coherence, or surround coherence parameter value based on determined direction value being opposite to the associated combined direction value in front-to-back sense relative to the array associated with the at least three obtained audio signals, where at least the associated azimuth values are within a threshold angle.

[0188] In some examples, the correction or adaptation for a detected front-back discrepancy is carried out for each TF tile separately. In other examples, the discrepancies are considered at least in part together, i.e. , the outcome for at least one discrepancy TF tile can affect the outcome for at least one other discrepancy TF tile.Having checked the front-to-back changes the final operation, as shown in 1013, is one of forming an adapted spatial audio format / representation based on the determined and adapted spatial audio parameters and obtained audio signals and provide these to the encoder for encoding.

[0189] In some embodiments the capture processor is configured to perform the direction determination by evaluating or determining a broadband direction based on the combined energy-weighted direct-to-total energy ratios. In other words, the whole frequency bandwidth is considered and only one representative direction for the whole bandwidth (or only a single combined direction) is determined.

[0190] The whole frequency band in this example can be based on the ibw parameter obtained as codec configuration parameter.

[0191] Alternatively, in some embodiments a broadband or combined direction can be evaluated for each frequency band separately by determining a representative or combined direction for each band by excluding the frequency band itself from the calculation.

[0192] The representative (broadband or combined) direction can then be compared against each of the determined TF tile directions to determine whether there is a potential error in the determined TF tile direction. In some embodiments the potential error can be indicated by a comparison resulting in a determined TF tile direction which is close to the representative direction in azimuth but mirrored. For example, where both the representative direction and determined TF tile direction have a direction about 60 degrees right, but one is towards the front and one is towards the back is a front-back discrepancy that may or may not indicate an erroneous front-back determination for the TF tile direction.

[0193] This erroneous front-back determination can furthermore identify whether the discrepancy is one of two types of potential TF tile front-back discrepancies. The first type or example is a weak discrepancy (low energy) where the directional discrepancy is associated with a weak or low energy direction, the second type or example is a strong (high-energy) content discrepancy where the directional discrepancy is associated with a strong or high energy direction.

[0194] In some embodiments the capture processor can be configured to force the weak discrepancy case to the side of the representative direction. In otherwords, a determined TF tile direction with associated low energy (for example below a determined ratio or absolute value) with a direction approximately opposite the representative or combined direction is modified to the side of the representative direction.

[0195] In examples where the potential front-back discrepancy is related to strong content (for example with an energy above a determined ratio or absolute value), the capture processor can be configured to not ‘correct’ the discrepancy or in other words not to move the direction to match the representative direction side (for example move it to the front or to back if the example were reversed).

[0196] In some embodiments, as discussed above, the further parameters associated with the direction can be also modified.

[0197] In some embodiments the capture processor can be configured to attempt to reduce perceptual impairment from a suspected front-back discrepancy by making the sound more diffuse. For example, this can be useful if the related sound is very strong and is not moved by some energy-weighting based comparison.

[0198] In this case, the direct-to-total energy ratio, regardless of whether the capture processor changes the direction, associated with the strong content direction is reduced, while increasing the diffuse-to-total energy ratio.

[0199] This change alone would have the effect of making the sound appear more diffuse and around the user. For example, the perception of such processing could result in a sound with too much externalized non-directional content, e.g., feeling of presence of echo. These unwanted processing effects can, in some embodiments, be mitigated by changing or modifying a surround coherence value associated with the diffuse energy. For example, surround coherence based on analysis could be 0.0 or null. This is because current smartphone captures are not generally capable of properly determining a surround coherence value. In some embodiments the surround coherence value for the TF tile could be set to a relatively high value, e.g., 1.0 or 0.75 or any suitable pre-determined value. In some examples, the closest TF tiles (either in frequency and / or time) could be set to a lower non-zero value, e.g., 0.5 or 0.25, in order to smooth the effect further.These operations can make the associated diffuse sound more coherent and appear more similar (with respect to the two output channels in binaural output and as heard by the two ears and, for example, reduce the externalization and the echoic nature of the energy ratio processing described above).

[0200] In some embodiments the determination or generation of the (combination or) representative direction can employ equal energy-weighting (in that the weighting for each of the directions is based on the energy value associated with the direction) or unequal weighting across frequency bands (for example the weights can be further based on the frequency of each band to emphasize directions associated with the human vocal range to prevent any voice related components from having an ‘wrong’ direction).

[0201] In some embodiments the evaluation of the representative (or combined) direction for each subband group can be performed separately according to a bitrate-specific MASA subband grouping. In such embodiments the codec configuration information including the coded format (input format) and current bitrate and / or negotiated bitrate range is assumed to have been received or obtained (and for example signaled to the capture processor via the controller). In some further embodiments the current audio bandwidth and / or negotiated bandwidth range can also be considered when performing the evaluation.

[0202] For example, the following table (which echoes Table 5.5-3 from TS 26.253) provides the IVAS-internal coding configurations relating to MASA format and MASA spatial metadata parameters including the whole supported range of bitrates:

[0203]

[0204]

[0205] For example, in the scenario of a 1 -direction analysis, all the metadata bands as given in Table A.3 of TS 26.258 (shown earlier) are available only starting at bitrate of 256 kbps. For 2-direction analysis, all the metadata bands for both directions (where the final coded order of primary and secondary directions is selected by the IVAS algorithm, not the capture system) are available only at the highest bitrate of 512 kbps. On the other hand, the bands for the 2nddirection at lower bit rates are at adaptive positions and therefore can be unpredictable from outside of the encoder. Since most services are expected to operate in the range of 13.2-128 kbps (or even lower), considering the IVAS MASA band usage and the current IVAS bitrate is highly beneficial for the capture algorithm.

[0206] In examples, IVAS voice service can utilize a fixed bitrate only or allow adaptation to radio channel for certain bitrate ranges. For example, we can consider IVAS voice service at one of following configurations:

[0207] - ibr = 13.2, ibw = swb

[0208] - ibr = 13.2 - 24.4, ibw = swb

[0209] - ibr = 13.2 - 48, ibw = swb

[0210] The following example tables presented below provide a mapping of the actual MASA bands as considered by the IVAS encoder in its processing at various bitrates. The capture processor can therefore take an audio input with metadata according to Table A.3 (26.258) and combine metadata according to the band indices given here. The indexing in the following two tables starts from 0 (as it is based on code implementation), the indexing in the above Table A.3starts from 1 . Thus, the band indices for comparison purposes are off by one. For example, according to the table below, IVAS utilizes 5 bands at 13.2 kbps, where the 4thband has a lowest frequency corresponding to 7thMASA format band and highest frequency corresponding to 14thMASA format band, whereas Table A.3 presented earlier shows that the two frequency values from band 8 and band 15: 2800 Hz and 6000 Hz.

[0211] A one-direction input utilizes values in the first of the following tables. Two- direction inputs utilize a combination of both of the following tables such that the first (selected) direction uses the first of the following tables (at one bitrate step lower), while the second (selected) direction uses the second of the following two tables.

[0212] In the following examples the codec bandwidth can be taken into account by the capture processor by providing only non-default values in the bandwidth range given by the codec. The analysis can also ignore any effect from audio bandwidth outside the bandwidth used by the codec. This bandwidth based analysis can, for example, be tuned based on the specific embodiments and implementation. For example, if ibw=swb, band 24 can not be used. However, this band could be used during analysis and the analysis correction to reinforce the desired value (e.g., front direction) in a lower band (e.g., band 23, 22, etc.) but not used to reinforce the corresponding potential front-back discrepancy value (e.g., back direction).

[0213] MASA metadata coding bands per bitrate in IVAS, 1 direction (1stdirection)

[0214]

[0215]

[0216] MASA metadata coding bands per bitrate in IVAS, 2 directions (2nddirection)

[0217]

[0218]

[0219] From these tables “MASA metadata coding bands per bitrate in IVAS (1stdirection” and “MASA metadata coding bands per bitrate in IVAS, 2 directions (2nddirection)” for both 1 -direction and 2-direction MASA inputs the capture processor can combine data before IVAS encoding to attempt to filter out potential errors that could be worse in the encoder. For MASA or similar 1 -direction determinations, this combination can be implemented at bitrates of 192 kbps and below. For MASA or similar 2-direction format inputs, the combination operation can be implemented for all inputs below 512 kbps.

[0220] In some embodiments for the 1 -direction format signals, the desired bands can be derived based on the first of the two tables above. For 2-direction format inputs, the desired bands can be derived based on the first of the two tables above using a bitrate 1 level lower, for example, 80 kbps for coding at 96 kbps, starting from 64 kbps. At bitrates 13.2-48 kbps, the regular line can be used. In addition, in all cases, the bands selected can depend on whether metadata with 20-msinherent resolution (joined st mode) or higher inherent temporal resolution, e.g., 5-ms resolution allowed by the input format is used.

[0221] For example, Fig .11 shows an example format structure, where the codec configuration parameter for the bitrate (e.g., ibr) indicates a bitrate of 24.4 kbps for the current frame. The capture processor can be configured to target an effective resolution according to the right-hand side of Fig.11 as shown by the frequency band resolution 1111 (bands 1 ,2, 3, 4, 5) and the time resolution 1103 (1 ,2,3,4). This can be implemented by using the corresponding frequency bands for analysis and / or adaptation, and the corresponding values are duplicated for the left-hand side structure as shown by the frequency band resolution 1101 (bands 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24) and the time resolution 1103 (1 ,2,3,4) as required.

[0222] For example, after analysis and adaptation is completed for the target resolution, in subframe 1 , frequency band 5, the parameter values are inserted or copied to a high-resolution data structure as follows:

[0223] Copy subframe 1 , frequency band 5 values to subframe 1 , frequency bands 16-24.

[0224] The same duplication or mapping operation can be employed for temporal resolution changes. However, for temporal copying or mapping the “joined sf mode” columns of the table are employed instead of the “regular mode” columns. For 24.4 kbps IVAS encoding, this ‘mode’ switch does not affect the resulting mapping or copying.

[0225] As can be seen in Fig .11 , some of the coded bands are very large, e.g., at 24.4 kbps band 5 covers everything over 6 kHz audio bandwidth. Therefore, correcting the potential front-back discrepancies by taking directly into account the affected band in the codec domain (effective TF resolution according to codec internal processing) can ensure that large frequency bands do not shift incorrectly in the rendering. By correcting a deviation only in the capture or format domain (effective TF resolution according to MASA metadata format specification), a correct perception through the end-to-end system cannot be guaranteed.In some embodiments the capture processor when performing the determination of the representative or combined direction can be for more than one time instance to attempt to reduce fluctuations over time.

[0226] For example, more than one successive subframe or frame of audio input and determined parameter values can be employed to determine the representative or combined direction.

[0227] In one example, if the representative direction changes significantly from one subframe to next subframe or from one frame to next frame, a lower temporal weighting can be applied. This allows a reduction in false positive front-back direction determinations while introducing fewer false negative front-back direction determinations during scenes where the sources are truly changing.

[0228] For example, Fig.13 shows an example where the representative or combined direction can be implemented with weighting across example capture bands for an example IVAS MASA WB (Wide-Band) bandwidth (corresponding to 16-kHz sampling rate) input generation for one of the low bit rate operations (e.g., 24.4 kbps or below). For example, the capture processor can be configured by the controller when it has been determined that the codec parameters (for example provided by the SDP parameters e.g., ibr=16.4 and ibw=wb) operate with SWB (Super Wide-Band) bandwidth (corresponding to 32-kHz sampling rate).

[0229] In Fig .13 are shown in three TF band structures. The first TF band structure relates to the audio signals generated by the microphones and shows a frequency band structure 1301 of equal capture bands 1 to 40 and four time instances or sub-frames 1303. For example, these may be 40 400-Hz frequency bands up to 16 kHz and four 5-ms subframes.

[0230] The second TF band structure represents targeted MASA format content based on the above parameters and has frequency bands 1311 (shown by frequency bands 1-5) over the same four time instances or sub-frames 1303. Thus, as shown in Fig 13 the capture processor can be configured to employ or use only a determined number of capture bands (e.g. capture bands 1 to 22) in the determination, and thus ignore capture bands 23-40, since they relate to SWB content well beyond the 8-kHz bandwidth of the WB input. In some embodimentsthe spatial audio analysis bands may be different than the IVAS CLDFB (Complex Low-Delay Filter Bank) bands that are based on a 400-Hz prototype, e.g., different filter banks, e.g., based on STFT (Short-Time Fourier Transform), can be used for the analysis.

[0231] For example, the dark-shaded capture bands are discarded and the white and mid-shaded bands employed based on the following:

[0232] a mid-shaded capture band is considered only for the target band to which it belongs during the analysis and the adaptation process; and

[0233] a white or non-shaded capture band affect also target bands outside the frequency range associated with the capture band itself.

[0234] For example, capture band 6 is employed to only affect band 3, while capture band 7 is employed to affect both band 3 and band 4 through the weighting process. Similarly, in some embodiments there can be temporal weighting which is not specifically illustrated in this example.

[0235] The third TF band structure can be the output of the capture processor and is the input MASA format provided to the encoder based on the content generated according to embodiments. It can thus comprise frequency bands 1 to 24 1321 over the same four time instances or subframes 1303. This can be shown by the augmentation of the second TF band structure by inserting zero-valued (i.e., default-valued) data for the frequencies above 8 kHz (which are shown by the bands 21-24).

[0236] With respect to Fig.14, is shown the effect of the application of the embodiments described herein to the example shown in Fig.9. In this example all of the bands for the time sub-frames have been correctly assigned as shown by the correct source direction 1403 and then provided to the listener 1407 by the left channel 1411 and right channel 1413 audio signal components.

[0237] In some embodiments the front-back determination operations for a 1-direction metadata representation can be extended to a two-direction or more generally multiple direction implementation. For example, the following describes embodiments where there is a utilization of a two-direction metadata representation based on one- or two-direction spatial audio capture analysis.As indicated by Table A.1 above the MASA format supports one or two analyzed directions per frame. In some embodiments the capture processor can be configured to generate or modify a second direction when a front-back discrepancy is determined.

[0238] For example, in some situations moving a directional sound that appears very strong relative to the scene can result in a reduction in perceived quality, and a better possibility could be to soften the directional sound and play it back from two directions instead (e.g., front and back).

[0239] This determination may be, e.g., a binary decision relating to at least one energy (ratio) comparison based with at least one representative direction and associated energy (ratio), where the representative direction is on the front or on the back and the tested direction is on the back or on the front, respectively, within a threshold angle.

[0240] The threshold angle can at least be an azimuth angle, but in some embodiments the threshold angle can be an angle or solid angle defining a three-dimensional angle (for example using both azimuth and elevation or other angular notation). An example threshold angle can be 10 degrees however this is an example only. Furthermore, in some embodiments the angle can differ in value between azimuth or elevation dimensions, for example the threshold angle can be 10 degrees in azimuth but 5 degrees in elevation.

[0241] It is noted that two-direction output according to some embodiments is employed at 64 kbps and higher. At all lower bitrates, the IVAS encoder will combine the directions (i.e., simplify the spatial scene such that it can be quantized) based on the assumption that the input scene is correct. Thus, as the input scene in this case is a best-effort error correction, it can be beneficial to apply the simplification already before the IVAS encoder is given the input signal.

[0242] In other words, a 1 -direction processing as described above is preferably employed when the encoder bitrate is below 64 kbps.

[0243] In some embodiments, the capture processor is configured to determine a front-back discrepancy based on the methods employed above in a one-direction analysis. In other words, only one direction is analyzed per each TF tile. In addition, in some embodiments the capture processor is configured to determinethat at least one (potentially) erroneous direction has an energy level (either ratio or absolute value) greater than a determined threshold (i.e., is significantly loud). In the examples described above the capture processor can be configured not to move the direction.

[0244] Instead, in some embodiments, a second direction analysis is triggered, where the second direction analysis comprises:

[0245] setting “zero values” for the second directions / direct energies etc., e.g., by copying the direction index from the first analyzed direction data and giving them zero direct-to-total energies in the second direction. Alternatively, the energies could be, e.g., set by putting half of the direct- to-total energy value for each of 1stand 2nddirection. Thus, the introduced second direction has no effective contribution at this stage;

[0246] for the at least one front-back discrepancy TF tile, an update of the direction is then performed by setting or giving the second direction a new value that is the front-back mirrored value of the first direction. The direct- to-total energy can be, e.g., split (for example halved) between the two directions, or employing this split value as a baseline energy ratio value, where a further weighting or modification of the energy value of one of the two directions is implemented based on a confidence level associated with the front-back error analysis.

[0247] In some embodiments a two-direction capture analysis is performed, and there are thus two analyzed directions for at least one of the TF tiles. In these embodiments it is assumed there are two analyzed directions for the TF tile that can exhibit front-back discrepancies.

[0248] Multi-direction analysis can be performed over the whole scene or, for example on a per sector basis. The capture processor when attempting to perform front-back corrections in these embodiments further is configured to take this information into account. In the following example it is assumed that both directions are analyzed over the full 360-degree azimuth (3-mic) or the full 3D space (4-mic).As first step of the two-direction analysis embodiments, the capture processor can be configured to determine whether the two directions are directly showing front-back discrepancy.

[0249] In some embodiments where the analysis is showing that the two directions are showing front-back discrepancy then the directions can be left as is, or a small adjustment for a preferred direction by means of direct-to-total energy adaptation can be performed as described in the above examples.

[0250] Furthermore, in some embodiments where the two directions do not show front-back discrepancy, this can be because of two possible cases.

[0251] In the first case, the two directions are substantially the same directions. In this case, they can be at least temporarily combined, and 1 -direction error correction is performed according to the example embodiment described above.

[0252] In the second case, at least one direction is outside the representative or combined direction consideration. In this case, the direction or directions ignored, and the analysis performed for any remaining direction that may have front-back discrepancy.

[0253] In one further example, the capture processor is configured to adapt or modify a spread coherence value for at least one determined output direction. In these embodiments it is determined that the two directions exhibit front-back discrepancy and the direct-to-total energy ratios associated with these directions are significantly similar. In this case, a determined direction value can in one example be directly to the side of the device (according to the originally determined azimuth values), and a spread coherence value could spread the energy to the side and to the front-back directions or towards the two front-back directions only.

[0254] Fig .15 shows an example of how a second direction can be heard in a two-direction situation. In this example, the correct source direction 1530 produces the outputs 1511 and 1513 and the ‘erroneous’ source direction 1520 contributes outputs 1501 and 1503. In such a manner the introduction of the second direction reduces the listener annoyance of false positive determined back-direction sounds but can contain some erroneous sound.A further advantage with respect to the two-direction embodiments is that false negative cases, e.g., a real back-direction source being “corrected” into a front-direction by the algorithm will not appear as severe because of the introduction of the second direction component. In the example shown in Fig.15, the sound may predominantly appear from one direction (front or back), however in general the perception can be somewhat vague, yet still spatial, and therefore not annoying.

[0255] In some embodiments the capture apparatus can comprise at least one further sensor, for example a video camera, and the capture processor can obtain inputs such as camera angle and / or identified object (sound source) based on camera angle as a further input to assist in the determination or weight front-back analysis error or a front-back determination, e.g., according to examples above.

[0256] For example, when a video call is placed, the directions on the side of the active camera are preferred, i.e., they receive higher weights or are favored when adjusting direct-to-total ratios between two directions where a first direction is on the camera side and a second direction is on the other side.

[0257] The capture processor can be implemented based on any suitable means. For example, in some embodiments the processor can be implemented by a machine learning (ML) based processor. Thus, in such implementations, the backpropagation / update of the representative direction calculation weights can be employed based on addition of bands / time instances where, for example, confidence of estimate does not reach a pre-determined threshold based on a first set of bands / time instances in a multi-stage ML-based estimation.

[0258] ML-based methods can be employed to detect, e.g., a person or other likely sound source from a captured video. This information can then be employed, at least in part, to determine one or more representative directions in a scene.

[0259] Another suitable application of ML in some embodiments is a detection of speech. In other words, some embodiments can employ a classifier of sounds and have different weights or otherwise different processing operations can be employed for different sound classes. Thus, in speech detection examples, typical speech frequencies are maintained to be in the same direction. This can furtherbe even enhanced by, e.g., phoneme-level detection or based on speaker identity detection.

[0260] As described earlier, in some embodiments the capture processor can be controlled or configured based on the renderer configuration, e.g., at least one renderer configuration parameter. For example, when the listener is not using head-tracking (e.g., rendering / playback device is not capable of that), the renderer apparatus can provide signaling to the capture apparatus for example using the corresponding value for the reverse PI type HEADTRACK ACTIVE or any such suitably defined reverse PI type. In such embodiments this information can be used by the capture processor to process the captured audio signals to attempt to improve perception quality. For example, the processing can be controlled in such a manner that scene orientation signaling input on a capture device with low front-back determination quality is limited, when no head-tracking capability is available at renderer. For example, scene orientation can be controlled on a capture device via a suitable control, e.g., a user input or an orientation control feature on an application. This scene orientation can be signaled to the receiver via RTP using SCENE ORIENTATION PI data. Thus, for applications where, e.g., continuous scene orientation updates are not considered necessary, the allowed values for SCENE ORIENTATION PI data can be limited. For example, extreme scene orientation changes in left-right sense (towards 90-degree rotations) are in this example not allowed.

[0261] For example, Fig.16 shows example scene orientation limitations according to some embodiments. Example allowed angle (in left-and-right sense) could be, e.g., 30 degrees azimuth. Thus, for example, as shown in the first example on Fig.16 there can be a full scene orientation control where orientation can occur across the full range 1605 and, e.g., the user 1601 can rotate the scene around the capture device 1603. The rotation input can be signaled using SCENE ORIENTATION PI data, and the rotation can be applied at the renderer. Furthermore, as shown in the second example on Fig.16 there can be an extreme limitation of scene orientation control allowed where only front-back switching 1607 is allowed. For example, no scene orientation is generally allowed, but a 180-degree scene rotation can be made, e.g., based on camera selection in avideo call. Thus, SCENE ORIENTATION PI data can only receive values corresponding to 0 degrees or 180 degrees on azimuth in this example. Also as shown in the third example on Fig .16 there can be partial scene orientation control where the front-back switching is allowed and some azimuth rotation is allowed (as shown by regions 1611) but angles are set such that the user cannot rotate close to 90 degrees from the front and back directions as shown by the shaded area 1609.

[0262] In some examples it could be negotiated similar limitations for receive-side use.

[0263] In such embodiments as described herein improvements of the analysis and merging of spatial feature data using capture processing or capture analysis. The effect of which can result in outputs at the renderer where, e.g., at lower bit rates, the sound scene can become more stable (e.g., appear less busy and jumpy).

[0264] The problem described above can be observed in multi-channel loudspeaker rendering where at least surround / back channels are available, e.g., 5.1 , in addition to head-tracked binaural rendering or static binaural rendering with orientation signaling (embodiment 5). Thus, output format of multi-channel audio (i.e., more than stereo loudspeakers) can be similarly considered as a renderer configuration input for the methods of the concept.

[0265] To provide a high-level summary of the above there can be apparatus or methods configured to perform the following:

[0266] 1. An immersive audio call is negotiated (according to relevant specifications) between at least one ‘capture apparatus’ sending side and one ‘rendering apparatus’ receiving side device. As part of the negotiation at least one codec configuration or rendering configuration parameter is obtained. o For example, codec configuration parameters comprise at least one of: coded format(s), bitrate, and signal bandwidth.

[0267] o For example, renderer configuration parameters comprise at least one of: type of rendering (LS, binaural) and availability of headtracking capability in case of binaural rendering.2. A spatial audio capture apparatus or algorithm is adapted (based on the at least one codec and / or renderer configuration parameter) to produce: 3. An immersive audio format (which can be one of the IVAS specification input formats) whose specific spatial audio content is modified based on the adaptation of the spatial audio capture algorithm in order to improve the rendering of the immersive audio scene based on

[0268] o Selected codec configuration and the algorithmic steps that depend on properties of the immersive audio format input, and o Selected renderer configuration, where the resulting reproduced audio quality depends on (the decoded representation of) the immersive audio format input content, the preceding codec algorithmic steps, and the user interaction / input to the renderer.

[0269] Thus, in some embodiments the adaptation or modification applied (for example by the capture or front end processor) in the spatial audio capture / immersive format generation affects the determination of the direction parameter (e.g., direction index) and / or direct-to-total energy ratio associated with a determined direction, and / or diffuse-to-total energy ratio, based on indications relating to the codec configuration.

[0270] In other words, the codec configuration identifies what types of algorithmic steps are foreseen, e.g., based on bitrate and / or the renderer configuration, e.g., whether head-tracking capability is available and used. The effect can be in at least one time-frequency (TF) tile (or subband) for at least one direction of a 1 - or 2-direction capture analysis and spatial audio representation.

[0271] In some embodiments, as indicated above the modification or adaptation can affect determination of spread coherence value associated with a determined second direction, where the determined second direction and the determined spread coherence value can be based on determining two first directions associated with significantly similar azimuth values relative to the left-right axis such that a first of the two first directions is front of the capture device and a second of the two first directions is back of the capture device and the direct-to-total energy ratios associated with the two first directions are significantly similar.In some embodiments, the adaptation can affect determination of surround coherence value associated with the diffuse energy, where a determined direction value differing from determined representative direction can result in use of weighting the associated energies / energy ratios such that diffuse energy estimate can be maintained or increased, and the associated surround coherence value can be increased.

[0272] With respect to Fig.17 an example electronic device is shown. The device may be any suitable electronics device or apparatus. For example, in some embodiments the device 1900 is a mobile device, user equipment, tablet computer, computer, consumer electronic device, still / video camera device, mobile communication device, audio / video / image playback or recording apparatus, vehicle, etc. or any combination thereof.

[0273] In some embodiments the device 1900 comprises at least one processor or central processing unit 1907. The processor 1907 can be configured to execute various program codes such as the methods such as described herein.

[0274] In some embodiments the device 1900 comprises at least one memory 1911. In some embodiments the at least one processor 1907, e.g. CPU (Central processing Unit) and / or GPU (Graphical Processing Unit), is coupled to the at least one memory 1911. The memory 1911 can be any suitable storage means. In some embodiments the memory 1911 comprises a program code section for storing one or more program codes, or program instructions, implementable upon the one or more processor 1907. Furthermore, in some embodiments the memory 1911 can further comprise a stored data section for storing data, for example data that has been processed or to be processed in accordance with the embodiments as described herein. The implemented program code stored within the program code section and the data stored within the stored data section can be retrieved by the processor 1907 whenever needed via the memory-processor coupling.

[0275] In some embodiments the device 1900 comprises a user interface 1905. The user interface 1905 can be coupled in some embodiments to the processor 1907. In some embodiments the processor 1907 can control the operation of the user interface 1905 and receive inputs from the user interface 1905. In some embodiments the user interface 1905 can enable a user to input commands tothe device 1900, for example via a keypad. In some embodiments the user interface 1905 can enable the user to obtain information from the device 1900. For example, the user interface 1905 may comprise a display configured to display information from the device 1900 to the user. The user interface 1905 can in some embodiments comprise a touch screen or touch interface capable of both enabling information to be entered to the device 1900 and further displaying information to the user of the device 1900. The user interface 1905 in some embodiments can comprise one or more loudspeakers, and one or more microphones.

[0276] In some embodiments the device 1900 comprises an input / output port 1909. The input / output port 1909 in some embodiments comprises a transceiver. The transceiver in such embodiments can be coupled to the processor 1907 and configured to enable a communication with other apparatus or electronic devices, for example via wireless and / wireless communications networks. The transceiver or any suitable transceiver or transmitter and / or receiver means can in some embodiments be configured to communicate with other electronic devices or apparatus via a wire or wired coupling.

[0277] The transceiver can communicate with further apparatus by any suitable communications protocol. For example, in some embodiments the transceiver can use a suitable mobile telecommunications protocol, such as 3G / 4G / 5G / 6G or any further generation protocol, a short-range wireless communication protocol, such as a wireless local area network (WLAN) protocol, for example IEEE 802.X, a Bluetooth, or infrared data communication pathway (IRDA), or any combination thereof.

[0278] The transceiver input / output port 1909 may be configured to receive the signals and in some embodiments obtain the focus parameters as described herein.

[0279] In some embodiments the device 1900 may be employed to generate a suitable audio signal using the processor 1907 executing suitable code. The input / output port 1909 may be coupled to any suitable audio output for example to a multichannel speaker system and / or headphones (which may be a headtracked or a non-tracked headphones) or similar.The various example embodiments may be implemented in hardware or special purpose circuits, software, logic, circuitry or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the examples are not limited thereto. While various aspects of the examples may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.

[0280] The example embodiments may be implemented by computer software executable by a data processor UE, such as a mobile user device or a consumer electronic device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD.

[0281] The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processors may be of any type suitable to the local technical environment, and may include one or more of general-purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), gate level circuits and processors based on multi-core processor architecture, as non-limiting examples.The example embodiments may be practiced in various components such as one or more integrated circuit modules or circuitry. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.

[0282] As used in this application, the term “circuitry” may refer to one or more or all of the following:

[0283] (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry) and

[0284] (b) combinations of hardware circuits and software, such as (as applicable):

[0285] (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and

[0286] (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and I hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.”

[0287] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.

[0288] Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre-stored design modules.Once the design for a semiconductor circuit has been completed, the resultant design, in a standardized electronic format (e.g., Opus, GDSII, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication.

[0289] The foregoing description has provided by way of exemplary and nonlimiting examples a full and informative description of the exemplary embodiments. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of the exemplary embodiments will still fall within the scope of this concept as defined in the appended claims.

[0290] 3GPP 3rdGeneration Partnership Project

[0291] CMR Codec mode request

[0292] EVS Enhanced Voice Services

[0293] ISM Independent Streams with Metadata

[0294] (i.e., type of Object-Based Audio)

[0295] IVAS Immersive Voice and Audio Services

[0296] MASA Metadata-Assisted Spatial Audio

[0297] MC Multichannel

[0298] OMASA Object-based audio with MASA (combined input format) OSBA Object-based audio with SBA (combined input format) PI Processing information (audio)

[0299] RTCP Real-Time Transport Control Protocol

[0300] RTP Real-Time Transport Protocol

[0301] RTP HE Real-Time Transport Protocol Header Extension

[0302] SBA Scene-Based Audio

[0303] SDP Session Description Protocol

[0304] ToC Table of Content

Claims

1. CLAIMS:1 . An apparatus comprising:at least one processor; andat least one memory storing instructions that when executed by the at least one processor, cause the apparatus at least to:obtain at least one of:codec configuration parameter; orrenderer configuration parameter;configure at least one immersive audio capture parameter based on the obtained at least one of: the codec configuration parameter; or the Tenderer configuration parameter; andgenerate at least one immersive format audio signal based on the at least one configured immersive audio capture parameter.

2. The apparatus as claimed in claim 1 , further caused to: obtain at least one input audio signal, wherein the apparatus caused to generate the at least one immersive format audio signal is further caused to process the at least one input audio signal based on the at least one configured immersive audio capture parameter to generate the at least one immersive format audio signal.

3. The apparatus as claimed in claim 2, wherein caused to obtain the at least one input audio signal is further caused to receive at least one microphone audio signal.

4. The apparatus as claimed in claim 3, wherein caused to process the at least one input audio signal based on the at least one configured immersive audio capture parameter to generate the at least one immersive format audio signal is further caused to:determine, for at least one direction, at least one direction parameter based on an analysis of the at least one microphone audio signal;determine at least one representative direction parameter from the at least one direction parameter at least based on the at least one configured immersive audio capture parameter; andmodify at least one of the at least one direction parameter based on the at least one representative direction parameter to correct any direction discrepancy between the at least one of the at least one direction parameter and the at least one representative direction parameter.

5. The apparatus as claimed in claim 4, wherein caused to determine the at least one representative direction parameter from the at least one direction parameter at least based on the at least one configured immersive audio capture parameter is further caused to:determine a selection of time and / or frequency bands based on the at least one configured immersive audio capture parameter; andcombine or average the at least one direction parameter for the selection of time and / or frequency band to determine the at least one representative direction parameter.

6. The apparatus as claimed in any of claims 4 or 5, wherein caused to process the at least one input audio signal based on the at least one configured immersive audio capture parameter to generate the at least one immersive format audio signal is further caused to determine, based on the at least one representative direction parameter, at least one of:a direct-to-total energy ratio associated with the at least one direction; an energy-weighted direct-to-total energy ratio associated with the at least one direction;a diffuse-to-total energy ratio; ora surround coherence value.

7. The apparatus as claimed in any of claims 4 to 6, wherein caused to determine at least one representative direction parameter from the at least onedirection parameter at least based on the at least one configured immersive audio capture parameter is further caused to:analyse the at least one representative direction parameter over at least two instances; andgenerate at least one temporal weighting function based on the analysis of the at least one representative direction parameter over the at least two instances, wherein caused to modify the at least one of the at least one direction parameter based on the at least one representative direction parameter is further caused to modify the at least one of the at least one direction parameter further based on the at least one temporal weighting function.

8. The apparatus as claimed in any of claims 1 to 7, wherein caused to configure the at least one configured immersive audio capture parameter based on the obtained at least one of: the codec configuration parameter; or the renderer configuration parameter is further caused to determine for the immersive format audio signal at least one ofa number of frequency bands for the immersive format audio signal, at least one frequency band range of a frequency band for the immersive format audio signal,at least one frequency band starting frequency of a frequency band for the immersive format audio signal, orat least one frequency band ending frequency of a frequency band for the immersive format audio signal.

9. The apparatus as claimed in any of the claims 1 to 8, wherein caused to configure the at least one configured immersive audio capture parameter based on the obtained at least one of: the codec configuration parameter; or the renderer configuration parameter is further caused to determine for the immersive format audio signal at least one ofa bit rate for the immersive format audio signal, ora number of directions for the immersive format audio signal.

10. The apparatus as claimed in any of claims 1 to 9, wherein the at least one codec configuration parameter comprises at least one ofa codec format,an encoding bitrate, oran encoding bandwidth.

11. The apparatus as claimed in any of claims 1 to 10, wherein the at least one renderer configuration parameter comprises at least one ofan output rendering type parameter for at least one further apparatus comprising a Tenderer,a head-tracking capability of at least one further apparatus comprising a Tenderer, oran active head-tracking operation of at least one further apparatus comprising a renderer.

12. The apparatus as claimed in any of claims 1 to 11 , further caused to obtain the at least one of the codec configuration parameter, or the renderer configuration parameter in at least one ofa real-time transport protocol payload,a real-time transport protocol header extension, ora real-time transport control protocol payload.

13. The apparatus as claimed in any of claims 1 to 12, further caused to obtain the at least one of: the codec configuration parameter; and the renderer configuration parameter for at least one ofduring a session setup orafter a session setup.

14. A method comprising:obtaining at least one of:codec configuration parameter; orrenderer configuration parameter;configuring at least one immersive audio capture parameter based on the obtained at least one of: the codec configuration parameter; or the renderer configuration parameter; andgenerating at least one immersive format audio signal based on the at least one configured immersive audio capture parameter.

15. A computer program comprising instructions which, when executed by an apparatus, cause the apparatus to perform at least the following:obtaining at least one of:codec configuration parameter; orrenderer configuration parameter;configuring at least one immersive audio capture parameter based on the obtained at least one of: the codec configuration parameter; or the renderer configuration parameter; andgenerating at least one immersive format audio signal based on the at least one configured immersive audio capture parameter.