Scene orientation modification
Patent Information
- Application Number
- PCT/EP2026/057592
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-25
- Filing Date
- 2026-03-18
- Publication Date
- 2026-10-01
Smart Images

Figure EP2026057592_01102026_PF_FP_ABST
Abstract
Description
SCENE ORIENTATION MODIFICATIONFIELD
[0001] The present application relates to a method, apparatus, system and computer program for scene orientation modification based on configuration information and not limited to scene orientation modification within Immersive Voice and Audio Services (IVAS) applications.BACKGROUND
[0002] Spatial audio can be captured and represented in various ways. For mobile communications, e.g., in the scope of 3GPP services, the smartphone is an example capture device and form factor.
[0003] Additionally a spatial or immersive audio experience over headphones is possible by binaural reproduction (headphone playback) of a binaural recording or through a dedicated binauralization processing using a suitable rendering algorithm. As a result of binaural processing, a listener can thus experience an audio scene as if it surrounded them in real life.
[0004] 3GPP IVAS (Immersive Voice and Audio Services) is a new conversational audio codec standardized in 3GPP SA4 and part of Release 18. This is a versatile codec with many supported input formats and output formats. The main supported input formats for IVAS are stereo, multichannel (MC), objectbased audio (ISM), scene-based audio (SBA), and Metadata-assisted spatial audio (MASA). In addition, the following combinations are supported: Objects with MASA (OMASA) and Objects with SBA (OSBA). IVAS furthermore includes the EVS codec for backwards compatible mono input operation. The IVAS output formats include mono, stereo, multi-channel (including custom loudspeaker layouts), FOA, HOA2, HOA3, and binaural. In addition, so-called pass-through operation is possible allowing, e.g., MASA output for MASA input. As a spatial audio codec supporting at least three degrees of rotation freedom (yaw, pitch, roll) for all spatial inputs, the IVAS codec is expected to be used in a variety of scenarios, all of which cannot be known beforehand.
[0005] The IVAS codec algorithm is described in 3GPP TS 26.253 (Codec for Immersive Voice and Audio Services; Detailed Algorithmic Description incl. RTP payload format and SDP parameter definitions). The IVAS codec floating-point C code is provided in 3GPP TS 26.258. IVAS codec fixed-point C code is currently being standardized for Release 19.
[0006] IVAS MASA (metadata-assisted spatial audio) is one of the immersive formats supported by the IVAS standard. It uses audio signal(s) together with corresponding spatial metadata (containing, e.g., directions and direct-to-total energy ratios in frequency bands). The MASA stream can, e.g., be obtained bycapturing spatial audio with microphones of, e.g., a mobile device, where the set of spatial metadata is estimated based on the microphone signals. The MASA stream can be obtained also from other sources, such as specific spatial audio microphones (such as Ambisonics), studio mixes (e.g., 5.1 mix) or other content by means of a suitable format conversion.
[0007] MASA format is specified in 3GPP TS 26.258 Annex A.
[0008] 3GPP DaCAS (Diverse audio CApturing System for UEs) is a current 3GPP SA4 work item that aims to define immersive audio capture example solutions for converting raw / compensated microphone signals into the relevant IVAS codec input formats. This will be based on definition of a set of target devices or target device types with descriptions of the overall microphone configurations. For example, number of microphones, their relative positions in the device, etc. should be defined.SUMMARY
[0009] According to a first aspect, there is provided an apparatus for immersive audio capture, the apparatus comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the system at least to: obtain scene orientation data for a spatial audio scene, the scene orientation data corresponding to an improved spatial balance for rendering a stereo audio signal, wherein the scene orientation data is based on analysis of spatial audio capture source positions; determine rotation information based on the scene orientation data, the rotation information for providing the improved spatial balance for rendering the stereo audio signal; and enable the rotation information for a rotation of the spatial audio scene.
[0010] The apparatus caused to enable the rotation information for a rotation of the spatial audio scene may be caused to transmit the rotation information to a further apparatus for application of the rotation of the spatial audio scene based on the rotation information.
[0011] The apparatus caused to enable the rotation information for a rotation of the spatial audio scene may be caused to rotate the spatial audio scene based on the rotation information.
[0012] The apparatus may be further caused to: obtain input audio signals from at least two microphones; process the input audio signals to rotate the spatial audio scene based on the rotation information to determine at least two audio signals; and encode the at least two audio signals to generate a transport audio signal component of an immersive audio bitstream.
[0013] The apparatus caused to process the input audio signals to rotate the spatial audio scene based on the rotation information to determine at least two audio signals may be caused to: determine at least one first format audio signal based on the input audio signals; and process the at least one first format audio signal based on the rotation information to determine the at least two audio signals.
[0014] The at least one first format audio signal may be at least one of: a MASA format audio signal; an Ambisonics audio signal; a multichannel audio signal; and a stereo audio signal.
[0015] The apparatus may be further caused to: determine at least one first format spatial metadata parameter based on the input audio signals; and process the at least one first format spatial metadata parameter based on the rotation information to determine at least one spatial metadata parameter; and encode the at least one spatial metadata parameter to generate a spatial metadata component of an immersive audio bitstream.
[0016] The at least one first format spatial metadata parameter may comprise at least one of: a direction and orientation parameter; and an energy ratio parameter.
[0017] The apparatus caused to process the at least one first format audio signal based on the rotation information to determine the at least two audio signals may be caused to process and / or mix individual channel parts of the at least one first format audio signal to determine the at least two audio signals.
[0018] The at least two audio signals may comprise two transport audio signals, and may be: stereo downmix audio signals; or binaural stereo downmix audio signals.
[0019] The apparatus may be further caused to transmit and / or store the rotation information for the rotation of the spatial audio scene based on the scene orientation data to a further apparatus, the further apparatus for providing the improved spatial balance for the rendering of the stereo audio signal based on the rotation of the spatial audio scene based on the rotation information.
[0020] The apparatus may be further caused to: obtain input audio signals from at least two microphones; determine at least one transport audio signal based on the input audio signals; analyse the input audio signals to obtain spatial metadata; encode the at least one transport audio signal and spatial metadata as an immersive audio bitstream; and generate an auxiliary bitstream comprising the rotation information for a rotation of the spatial audio scene.
[0021] The apparatus may be further caused to obtain at least one bitstream, the bitstream comprising at least one encoded audio signal and encoded spatial metadata, wherein the apparatus caused to obtain scene orientation data for a spatial audio scene is caused to obtain the scene orientation data from the at least one bitstream.
[0022] The apparatus may be further caused to decode the encoded spatial metadata, wherein the apparatus caused to obtain the scene orientation data from the at least one bitstream may be caused to obtain the scene orientation data based on the decoded spatial metadata.
[0023] The apparatus may be further caused to decode the encoded audio signal, wherein the apparatus caused to obtain the scene orientation data from the at least one bitstream is caused to analyse the decoded audio signal.
[0024] The apparatus caused to analyse spatial audio capture source positions to determine the scene orientation data may be further caused to: obtain the spatial audio capture source positions; divide the audio scene into a number of audio scene portions, each portion spanning a range of directions of the audio scene; identify for at least one audio scene portion a dominant spatial audio capture source to generate survivor spatial audio capture source directions; and identify one of: a smallest angle between the spatial audio capture source directions farthest from each other; or a smallest angle between the spatial audio capture source directions that includes a listener front direction.
[0025] The scene orientation data for a spatial audio scene may comprise at least one of: a coded format; an encoding bitrate; or an encoding bandwidth.
[0026] The scene orientation data may comprise at least one renderer configuration parameter, wherein the at least one renderer configuration parameter may comprise at least one of: an output rendering type parameter for a renderer; a head-tracking capability of a renderer; or an active head-tracking operation of a renderer.
[0027] The apparatus may be further caused to obtain the scene orientation data in at least one of: a realtime transport protocol payload; a real-time transport protocol header extension; or a real-time transport control protocol payload.
[0028] The apparatus may be further caused to obtain the scene orientation data for at least one of: during a session setup; or after a session setup.
[0029] The stereo audio signal may comprise a binaural stereo audio signal.
[0030] According to a second aspect, there is provided an apparatus for immersive audio capture, the apparatus comprising means configured to: obtain scene orientation data for a spatial audio scene, the scene orientation data corresponding to an improved spatial balance for rendering a stereo audio signal, wherein the scene orientation data is based on analysis of spatial audio capture source positions; determine rotation information based on the scene orientation data, the rotation information for providing the improved spatial balance for rendering the stereo audio signal; and enable the rotation information for a rotation of the spatial audio scene.
[0031] The means configured to enable the rotation information for a rotation of the spatial audio scene may be configured to transmit the rotation information to a further apparatus for application of the rotation of the spatial audio scene based on the rotation information.
[0032] The means configured to enable the rotation information for a rotation of the spatial audio scene may be configured to rotate the spatial audio scene based on the rotation information.
[0033] The means may be further configured to: obtain input audio signals from at least two microphones; process the input audio signals to rotate the spatial audio scene based on the rotation information todetermine at least two audio signals; and encode the at least two audio signals to generate a transport audio signal component of an immersive audio bitstream.
[0034] The means configured to process the input audio signals to rotate the spatial audio scene based on the rotation information to determine at least two audio signals may be configured to: determine at least one first format audio signal based on the input audio signals; and process the at least one first format audio signal based on the rotation information to determine the at least two audio signals.
[0035] The at least one first format audio signal may be at least one of: a MASA format audio signal; an Ambisonics audio signal; a multichannel audio signal; and a stereo audio signal.
[0036] The means may be further configured to: determine at least one first format spatial metadata parameter based on the input audio signals; and process the at least one first format spatial metadata parameter based on the rotation information to determine at least one spatial metadata parameter; and encode the at least one spatial metadata parameter to generate a spatial metadata component of an immersive audio bitstream.
[0037] The at least one first format spatial metadata parameter may comprise at least one of: a direction and orientation parameter; and an energy ratio parameter.
[0038] The means configured to process the at least one first format audio signal based on the rotation information to determine the at least two audio signals may be further configured to process and / or mix individual channel parts of the at least one first format audio signal to determine the at least two audio signals.
[0039] The at least two audio signals may comprise two transport audio signals, and may be: stereo downmix audio signals; or binaural stereo downmix audio signals.
[0040] The means may be further configured to transmit and / or store the rotation information for the rotation of the spatial audio scene based on the scene orientation data to a further apparatus, the further apparatus for providing the improved spatial balance for the rendering of the stereo audio signal based on the rotation of the spatial audio scene based on the rotation information.
[0041] The means may be further configured to: obtain input audio signals from at least two microphones; determine at least one transport audio signal based on the input audio signals; analyse the input audio signals to obtain spatial metadata; encode the at least one transport audio signal and spatial metadata as an immersive audio bitstream; and generate an auxiliary bitstream comprising the rotation information for a rotation of the spatial audio scene.
[0042] The means may be further configured to obtain at least one bitstream, the bitstream comprising at least one encoded audio signal and encoded spatial metadata, wherein the means configured to obtain scene orientation data for a spatial audio scene may be configured to obtain the scene orientation data from the at least one bitstream.
[0043] The means may be further configured to decode the encoded spatial metadata, wherein the means configured to obtain the scene orientation data from the at least one bitstream may be configured to obtain the scene orientation data based on the decoded spatial metadata.
[0044] The means may be further configured to decode the encoded audio signal, wherein the apparatus configured to obtain the scene orientation data from the at least one bitstream may be configured to analyse the decoded audio signal.
[0045] The means configured to analyse spatial audio capture source positions to determine the scene orientation data may be further configured to: obtain the spatial audio capture source positions; divide the audio scene into a number of audio scene portions, each portion spanning a range of directions of the audio scene; identify for at least one audio scene portion a dominant spatial audio capture source to generate survivor spatial audio capture source directions; and identify one of: a smallest angle between the spatial audio capture source directions farthest from each other; or a smallest angle between the spatial audio capture source directions that includes a listener front direction.
[0046] The scene orientation data for a spatial audio scene may comprise at least one of: a coded format; an encoding bitrate; or an encoding bandwidth.
[0047] The scene orientation data may comprise at least one renderer configuration parameter, wherein the at least one renderer configuration parameter may comprise at least one of: an output rendering type parameter for a renderer; a head-tracking capability of a renderer; or an active head-tracking operation of a renderer.
[0048] The means may be further configured to obtain the scene orientation data in at least one of: a realtime transport protocol payload; a real-time transport protocol header extension; or a real-time transport control protocol payload.
[0049] The means may be further caused to obtain the scene orientation data for at least one of: during a session setup; or after a session setup.
[0050] The stereo audio signal may comprise a binaural stereo audio signal.
[0051] According to a third aspect, there is provided a method for an apparatus for immersive audio capture, the method comprising: obtaining scene orientation data for a spatial audio scene, the scene orientation data corresponding to an improved spatial balance for rendering a stereo audio signal, wherein the scene orientation data is based on analysis of spatial audio capture source positions; determining rotation information based on the scene orientation data, the rotation information for providing the improved spatial balance for rendering the stereo audio signal; and enabling the rotation information for a rotation of the spatial audio scene.
[0052] Enabling the rotation information for a rotation of the spatial audio scene may comprise transmitting the rotation information to a further apparatus for application of the rotation of the spatial audio scene based on the rotation information.
[0053] Enabling the rotation information for a rotation of the spatial audio scene may comprise rotating the spatial audio scene based on the rotation information.
[0054] The method may further comprise: obtaining input audio signals from at least two microphones; processing the input audio signals to rotate the spatial audio scene based on the rotation information to determine at least two audio signals; and encoding the at least two audio signals to generate a transport audio signal component of an immersive audio bitstream.
[0055] Processing the input audio signals to rotate the spatial audio scene based on the rotation information to determine at least two audio signals may comprise: determining at least one first format audio signal based on the input audio signals; and processing the at least one first format audio signal based on the rotation information to determine the at least two audio signals.
[0056] The at least one first format audio signal may be at least one of: a MASA format audio signal; an Ambisonics audio signal; a multichannel audio signal; and a stereo audio signal.
[0057] The method may further comprise: determining at least one first format spatial metadata parameter based on the input audio signals; and processing the at least one first format spatial metadata parameter based on the rotation information to determine at least one spatial metadata parameter; and encoding the at least one spatial metadata parameter to generate a spatial metadata component of an immersive audio bitstream.
[0058] The at least one first format spatial metadata parameter may comprise at least one of: a direction and orientation parameter; and an energy ratio parameter.
[0059] Processing the at least one first format audio signal based on the rotation information to determine the at least two audio signals may further comprise processing and / or mixing individual channel parts of the at least one first format audio signal to determine the at least two audio signals.
[0060] The at least two audio signals may comprise two transport audio signals, and may be: stereo downmix audio signals; or binaural stereo downmix audio signals.
[0061] The method may further comprise transmitting and / or storing the rotation information for the rotation of the spatial audio scene based on the scene orientation data to a further apparatus, the further apparatus for providing the improved spatial balance for the rendering of the stereo audio signal based on the rotation of the spatial audio scene based on the rotation information.
[0062] The method may further comprise: obtaining input audio signals from at least two microphones; determining at least one transport audio signal based on the input audio signals; analysing the input audio signals to obtain spatial metadata; encoding the at least one transport audio signal and spatial metadata asan immersive audio bitstream; and generating an auxiliary bitstream comprising the rotation information for a rotation of the spatial audio scene.
[0063] The method may further comprise obtaining at least one bitstream, the bitstream comprising at least one encoded audio signal and encoded spatial metadata, wherein obtaining scene orientation data for a spatial audio scene may comprise obtaining the scene orientation data from the at least one bitstream.
[0064] The method may further comprise decoding the encoded spatial metadata, wherein obtaining the scene orientation data from the at least one bitstream may comprise obtaining the scene orientation data based on the decoded spatial metadata.
[0065] The method may further comprise decoding the encoded audio signal, wherein obtaining the scene orientation data from the at least one bitstream may comprise analysing the decoded audio signal.
[0066] Analysing spatial audio capture source positions to determine the scene orientation data may further comprise: obtaining the spatial audio capture source positions; dividing the audio scene into a number of audio scene portions, each portion spanning a range of directions of the audio scene; identifying for at least one audio scene portion a dominant spatial audio capture source to generate survivor spatial audio capture source directions; and identifying one of: a smallest angle between the spatial audio capture source directions farthest from each other; or a smallest angle between the spatial audio capture source directions that includes a listener front direction.
[0067] The scene orientation data for a spatial audio scene may comprise at least one of: a coded format; an encoding bitrate; or an encoding bandwidth.
[0068] The scene orientation data may comprise at least one renderer configuration parameter, wherein the at least one renderer configuration parameter may comprise at least one of: an output rendering type parameter for a renderer; a head-tracking capability of a renderer; or an active head-tracking operation of a renderer.
[0069] The method may further comprise obtaining the scene orientation data in at least one of: a real-time transport protocol payload; a real-time transport protocol header extension; or a real-time transport control protocol payload.
[0070] The method may further comprise obtaining the scene orientation data for at least one of: during a session setup; or after a session setup.
[0071] The stereo audio signal may comprise a binaural stereo audio signal.
[0072] According to a fourth aspect, there is an apparatus for immersive audio capture, the apparatus comprising: obtaining circuitry configured to obtain scene orientation data for a spatial audio scene, the scene orientation data corresponding to an improved spatial balance for rendering a stereo audio signal, wherein the scene orientation data is based on analysis of spatial audio capture source positions; determining circuitry configured to determine rotation information based on the scene orientation data, the rotation information forproviding the improved spatial balance for rendering the stereo audio signal; and enabling circuitry configured to enable the rotation information for a rotation of the spatial audio scene.
[0073] According to a fifth aspect, there is an apparatus for immersive audio capture comprising: means for obtaining scene orientation data for a spatial audio scene, the scene orientation data corresponding to an improved spatial balance for rendering a stereo audio signal, wherein the scene orientation data is based on analysis of spatial audio capture source positions; means for determining rotation information based on the scene orientation data, the rotation information for providing the improved spatial balance for rendering the stereo audio signal; and means for enabling the rotation information for a rotation of the spatial audio scene
[0074] According to a sixth aspect, there is provided a non-transitory computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform at least the method according to any of the preceding aspects.
[0075] According to a seventh aspect there is provided a computer program comprising instructions [or a computer readable medium comprising instructions] for causing an apparatus to for immersive audio capture, caused to perform at least the following: obtaining scene orientation data for a spatial audio scene, the scene orientation data corresponding to an improved spatial balance for rendering a stereo audio signal, wherein the scene orientation data is based on analysis of spatial audio capture source positions; determining rotation information based on the scene orientation data, the rotation information for providing the improved spatial balance for rendering the stereo audio signal; and enabling the rotation information for a rotation of the spatial audio scene.
[0076] In the above, many different embodiments have been described. It should be appreciated that further embodiments may be provided by the combination of any two or more of the embodiments described above.DESCRIPTION OF FIGURES
[0077] Embodiments will now be described, by way of example only, with reference to the accompanying Figures in which:
[0078] Fig.1 shows schematically example systems within which embodiments may be implemented;
[0079] Figs.2a to 2d show example binaural experiences for head-tracking and non-head-tracking situations;
[0080] Figs.3a and 3b show schematically example potential source differentiation issues with non-headtracking binaural rendering;
[0081] Fig.4 shows a schematic representation of an example microphone arrangement for a spatial audio capture smartphone with 2, 3, or 4 microphones;
[0082] Fig.5 shows an example stereo capture scenario employing a smartphone supporting multimicrophone spatial audio capture;
[0083] Fig.6 shows schematically example systems within which embodiments may be implemented;
[0084] Fig.7 shows the example system as shown in Fig.6 in further detail within which embodiments may be implemented;
[0085] Fig.8 shows an example spatial audio input generation system according to some embodiments;
[0086] Figs.9a to 9c shows an example of source behaviour for a head rotation within head-tracking and non-head-tracking situations;
[0087] Figs.10a to 10c shows a further example of source behaviour for a head rotation within head-tracking and non-head-tracking situations;
[0088] Figs. Ha to 11d show improved spatial rendering and source behaviour in in non-head-tracking examples according to some embodiments;
[0089] Fig.12 shows an example IVAS RTP packet format configuration according to some embodiments;
[0090] Figs.13a to 13f show an example method for determining the spread of the audio sources according to some embodiments;
[0091] Fig.14 shows a further example spatial audio input generation system according to some embodiments;
[0092] Fig.15 shows a flow diagram of an example implementation of spatial audio capture to determine stereo audio input according to some embodiments;
[0093] Fig.16 shows a flow diagram of a further example implementation of spatial audio capture to determine stereo audio input according to some embodiments;
[0094] Fig.17 shows an example spatial audio synthesis apparatus according to some embodiments;
[0095] Fig.18 shows a further example spatial audio synthesis apparatus according to some embodiments;
[0096] Fig.19 shows another example spatial audio synthesis apparatus with signalling according to some embodiments;
[0097] Fig.20 shows a flow diagram of an example implementation of spatial audio synthesis as shown in Fig.19 according to some embodiments;
[0098] Fig.21 a and 21 b show a comparison of an example stereo listening experience when implementing some embodiments;
[0099] Fig.22a and 22b show an example non-head-tracking listening experience when implementing some embodiments;
[0100] Fig.23a and 23b show a further example non-head-tracking listening experience when implementing some embodiments;
[0101] Fig.24a and 24b show an example head-tracking listening experience when implementing some embodiments;
[0102] Fig.25 shows a graph of example rotation effects for a mono-source placed in frontof a listener when implementing some embodiments; and
[0103] Fig.26 shows apparatus suitable for implementing some embodiments.DETAILED DESCRIPTION
[0104] The concept as discussed in the embodiments and examples herein is one of providing an enhanced spatial balance (or improved externalization or improved spatial positioning of sources for externalization) when rendering an output audio signal. For example, rendering a binaural stereo audio signal to a listener in spatial audio systems or decoding a stereo audio signal and providing this to the listener.
[0105] In summary, the embodiments as discussed in detail hereafter feature apparatus and methods for obtaining scene orientation data for a spatial audio scene, the scene orientation data corresponding to an improved spatial balance for rendering a stereo audio signal, wherein the scene orientation data is based on analysis of spatial audio capture source positions.
[0106] In these embodiments the apparatus or methods can then determine rotation information based on the scene orientation data, the rotation information for providing the improved spatial balance for rendering the stereo audio signal.
[0107] Then the embodiments enable the rotation information for a rotation of the spatial audio scene.
[0108] The enabling of the rotation information for a rotation of the spatial audio scene can in some embodiments be transmitting the rotation information to a further apparatus for application of the rotation of the spatial audio scene based on the rotation information.
[0109] The rotation information for a rotation of the spatial audio scene can control a rotation of the spatial audio scene based on the rotation information.
[0110] In some embodiments the apparatus / methods are further caused to obtain input audio signals from at least two microphones and process the input audio signals to rotate the spatial audio scene based on the rotation information to determine at least two audio signals which can then be encoded to generate a transport audio signal component of an immersive audio bitstream.
[0111] In some embodiments the apparatus or methods determine at least one first format audio signal based on the input audio signals and process the at least one first format audio signal based on the rotation information to determine the at least two audio signals.
[0112] In such embodiments the at least one first format audio signal is at least one of: a MASA format audio signal; an Ambisonics audio signal; a multichannel audio signal; and a stereo audio signal.
[0113] The apparatus / methods furthermore in some embodiments determine at least one first format spatial metadata parameter based on the input audio signals; and process the at least one first format spatial metadata parameter based on the rotation information to determine at least one spatial metadata parameter. These processed spatial metadata parameters can then be encoded to generate a spatial metadata component of an immersive audio bitstream.
[0114] The at least one first format spatial metadata parameter can comprise at least one of: a direction and orientation parameter; and an energy ratio parameter.
[0115] In some embodiments the processing can be a processing and / or mixing individual channel parts of the at least one first format audio signal to determine the at least one two audio signals.
[0116] The at least two audio signals can be two transport audio signals, and are: a stereo downmix audio signals; or a binaural stereo downmix audio signals.
[0117] In some examples the rotation information is transmitted and / or stored to a further apparatus, the further apparatus for providing the improved spatial balance for the rendering of the stereo audio signal based on the rotation of the spatial audio scene based on the rotation information.
[0118] In such examples the input audio signals can be obtained from at least two microphones and at least one transport audio signal is obtained based on the input audio signals. In such embodiments there can be an analysis of the input audio signals to obtain spatial metadata and then an encoding of the transport audio signal and spatial metadata as an immersive audio bitstream.
[0119] Furthermore the apparatus or method is configured to generate an auxiliary bitstream comprising the rotation information for a rotation of the spatial audio scene.
[0120] The apparatus or methods furthermore could obtain at least one bitstream, the bitstream comprising at least one encoded audio signal and encoded spatial metadata, wherein the apparatus caused to obtain scene orientation data for a spatial audio scene is caused to obtain the scene orientation data from the at least one bitstream.
[0121] The encoded spatial metadata could in some examples be decoded, wherein the apparatus caused to obtain the scene orientation data from the at least one bitstream is caused to obtain the scene orientation data based on the decoded spatial metadata.
[0122] The apparatus or methods in some embodiments can be further caused to decode the encoded audio signal, and analyse the decoded audio signal.
[0123] The scene orientation data for a spatial audio scene can comprise at least one of: a coded format, an encoding bitrate, or an encoding bandwidth.
[0124] The scene orientation data comprises at least one renderer configuration parameter, wherein the at least one renderer configuration parameter comprises at least one of: an output rendering type parameter for a renderer, a head-tracking capability of a renderer, or an active head-tracking operation of a renderer.
[0125] The obtaining of the scene orientation data can be from at least one of: a real-time transport protocol payload, a real-time transport protocol header extension, or
[0126] a real-time transport control protocol payload.
[0127] The scene orientation data can be obtained during a session setup or after a session setup. While spatial audio reproduction using headphones is foreseen to be based mainly on head-tracked presentation in the future, the majority of new headphone devices today still do not provide head-tracking capability.
[0128] It can furthermore be assumed that many current devices will continue to be used also in the future.
[0129] Furthermore, users may wish to turn off head-tracking capability in some cases, e.g., depending on what they will do in addition to listening to spatial audio reproduction.
[0130] Furthermore, IVAS UEs targeting different use cases and based on different form factors will have varying capabilities, e.g., based on IVAS Levels defined in 3GPP TS 26.250, and there can be UEs with no head-tracking rendering capability.
[0131] Thus, future spatial audio reproduction systems should consider methods which provide a high-quality rendered audio signal output over headphones with and without head-tracking capability. This is relevant in immersive voice communications and user-generated content (UGC) creation and streaming, which are example cases the IVAS codec standard is addressing.
[0132] Fig.1, for example, shows an example immersive call or teleconferencing system within which some embodiments can be implemented. In this example there is shown two sites or rooms, Room A 100 and Room B 102. Room A 100 comprises a ‘talker’ or user, Talker 1 103. Room B 102 comprises one ‘talker’ or user, Talker RX 141.
[0133] In the following example within room A 100 is a suitable audio communication system 110, such as a teleconference apparatus (or more generally telecommunications apparatus 110 which can operate as capture apparatus) configured to spatially capture and encode the audio environment and furthermore is configured to render a spatial audio signal to the room. The apparatus 110 can in some embodiments be implemented by a user equipment (UE), such as a mobile communication device, or a smartphone operating within a communications system, e.g., a cellular communications system or accessing any suitable access network. The user equipment in some embodiments comprises a smartphone such as shown in Fig.4. In some further or additional use cases the apparatus 110 can be implemented in a vehicle.
[0134] Within each of the other spaces, e.g. rooms, may be a suitable audio communication device, such as a teleconference apparatus (or more generally telecommunications apparatus such as apparatus 120 within room B, which in the following examples can be designated as rendering apparatus) configured to render a spatial audio signal to the room. In some embodiments the apparatus 120 within room B can furthermore be configured to capture and encode at least a mono audio and optionally configured to spatially capture and encode the audio environment. The apparatus 120 can in some embodiments be implementedby a user equipment (UE), such as a mobile communication device, or a smartphone operating within a communication system, e.g., a cellular communications system or accessing any suitable access network. The user equipment in some embodiments comprises a smartphone such as shown in Fig.4. In some further or additional use cases the apparatus 120 can be implemented in a vehicle.
[0135] In the following examples each room is provided with the means to spatially capture, encode spatial audio signals, receive spatial audio signals, and render these in a suitable way to a listener. It would be understood that there may be other embodiments where the system comprises some apparatus configured to only capture and encode audio signals (in other words the apparatus is a ‘transmit’ only or purely capture apparatus), and other apparatus configured to only receive and render audio signals (in other words the apparatus is a ‘receive’ only or purely rendering apparatus). In such embodiments the system within which embodiments may be implemented may comprise apparatus with varying abilities to capture / render audio signals.
[0136] The telecommunications apparatus (for each space, site, or room) 110, 120 in this example can be configured to call into a teleconference controlled by and implemented over a server 111.
[0137] In some embodiments the communications or teleconferencing system comprises a (peer-to-peer) communications system (rather than the server-based system shown in Fig.1) within which some embodiments can be implemented. Thus, for example, two or more UEs can be configured to interact directly with each other (for example to implement an immersive audio phone call between users) without the need for a central server.
[0138] For example, the communications or teleconferencing, e.g., the immersive audio phone call, between users can be an IMS (IP (Internet Protocol) Multimedia Subsystem) call over 4G (Generation) (LTE (Long Term Evolution)), 5G (NR (New Radio)), or 6G link or any other suitable link, for example, WLAN (Wireless Local Area Network).
[0139] The telecommunications apparatus or UEs 110, 120 can be configured to spatially capture and encode the audio environment and furthermore can be configured to render a spatial audio signal to the space. In this example only the communications path from the Room A 100 to the Room B 102 is shown for simplicity but a duplex or multipoint communication system comprising multiple signalling paths can be implemented using the methods as described herein without significant inventive input.
[0140] The teleconference apparatus (for each site or room) 110, 120 is further configured to communicate with each other to implement a teleconference function.
[0141] In some embodiments the user equipment 110 comprises a microphone array 115. The microphone array 115, can be, for example an array such as shown in Fig.4 which shows schematically example microphone integration for a spatial audio capture smartphone comprising 2, 3, or 4 microphones. The example smartphone 400 is shown comprising a side 422 which typically comprises a screen viewable bythe user 410 and a side 420 which typically comprises the main camera module. There is also shown, from the viewpoint of a user operating the smartphone in landscape orientation as shown in Fig.4 and with the side 422 with the display module (screen) facing the user and the side 420 comprising a main camera module 440 facing away from the user, a right side 430, a left side 428, a top side 424, and a bottom side 426. In the following examples with respect to the capturing of the audio signal a ‘front’ from the perspective of the user is with respect to the main camera module 440 side (such as the side 420), as typically a user capturing an audio scene, e.g., on a video call capturing the environment such as a street performance or event, is using the main camera module and the microphone array. However, it is appreciated that in some embodiments, for example when the user is capturing the device with the ‘selfie’ or screen facing camera the front could be the side with the screen (such as the side 422).
[0142] Furthermore Fig.4 shows some example microphone arrangements or configurations. For example, the smartphone 400 can comprise an example dual mic array arrangement, where the placement of the two microphones is on right-side 430 (Dual: R microphone 405) and left side 428 (Dual: L microphone 403) of the device bezel. With two such microphones, a stereo audio signal can be captured that provides information to separate left and right directions. This information could for example be able to estimate a position for an audio or sound source on a 180-degree plane. This estimated position is only a 180-degree plane because it is assumed that the audio source is located to the front of the smartphone as there is a front-back ambiguity where it is not possible to determine whether the source is from the front or back of the smartphone with respect to the two microphone positions separated by an axis through L and R channels. On the other hand, considering rendering and presentation of spatial audio, e.g., in an immersive communications scenario, it is generally preferable to have audio sources located in the front rather than back of the listener if full 360 capture and presentation is not achievable.
[0143] Furthermore, a triple mic array arrangement is shown where the dual mic array is supplemented with a rear-facing mic. In the example shown in Fig.4 the triple 407 microphone is located as part of a camera module protrusion (or camera bump or bulge). With three microphones and in this orientation, the front-back ambiguity can be resolved, and it is generally possible to derive information to provide an estimated position within a 360-degree plane. In other words, using three microphone audio signals to provide an actual azimuth position for the audio source that could be virtually placed in the audio scene.
[0144] A quad mic array configuration can further introduce a further microphone quad microphone 409 on the left side 403. Thus for example as shown in Fig.4 the quad microphone 409 along with the Dual: L microphone 403 effectively presents an arrangement with dual mics at the ‘bottom’ bezel when the device is used in portrait orientation, and which is closest to user’s mouth in typical handset (device on ear) and handheld hands-free (device in hand, at maximum at arm’s length in front of user) use cases.
[0145] With four microphones, an estimated position can generally estimate audio source positions also in elevation, in other words it is generally possible to describe the whole 3D space correctly at the capture point.
[0146] Furthermore Fig.1 shows a capture (front end) processor 107 configured to receive the audio signals 114 from the microphone array 115 and perform processing on these audio signals 114 to generate a suitable input format, for example a MASA spatial audio format or an Ambisonics (SBA) spatial audio format, for the (IVAS) encoder 101. In some embodiments as described in further detail herein the capture (front end) processor 107 is configured by a controller 105.
[0147] An example IVAS MASA format generated by the capture processor 107 can be based on 1 or 2 audio channels and associated metadata, as described by the following tables (which are referenced by the references found in and further described in 3GPP TS 26.258 Annex A):Table A.1: MASA format descriptive common metadata parameters Field Bits DescriptionFormat 64 Defines the MASA format for IVAS. Eight 8-bit ASCII characters: descriptor 01001001, 01010110, 01000001, 01010011,01001101, 01000001, 01010011, 01000001Values stored as 8 consecutive 8-bit unsigned integers.Channel 16 Combined following fields stored in two bytes.audio format Value stored as a single 16-bit unsigned integer.Number of (1) Number of directions described by the spatial metadata.directions Each direction is associated with a set of direction dependent spatial metadata.Range of values: [1, 2]Number of (1) Number of transport channels in the format.channels Range of values: [1, 2]Source format (2) Describes the original format from which MASA was created.(Variable (12) Further description fields based on the values of ‘Number of channels’ description) and ‘Source format’ fields.When all bits are not used, zero padding is applied.Table A.2a: MASA format spatial metadata parameters (dependent of number of directions) Field Bits DescriptionDirection 16 Direction of arrival of the sound at a time-frequency parameter interval. index Spherical representation at about 1 -degree accuracy.Range of values: “covers all directions at about 1° accuracy”Values stored as 16-bit unsigned integers.Direct-to-total 8 Energy ratio for the direction index (i.e., time-frequency subframe). energy ratio Calculated as energy in direction / total energy.Range of values: [0.0, 1.0]Values stored as 8-bit unsigned integers with uniform spacing of mapped values.Spread 8 Spread of energy for the direction index (i.e., time-frequency subframe). coherence Defines the direction to be reproduced as a point source or coherently around the direction.Range of values: [0.0, 1.0]Values stored as 8-bit unsigned integers with uniform spacing of mapped values.Table A.2b: MASA format spatial metadata parameters (independent of number of directions) Field Bits Description8 Energy ratio of non-directional sound over surrounding directions.Calculated as energy of non-directional sound / total energy.Range of values: [0.0, 1.0](Parameter is independent of number of directions provided.) Values stored as 8-bit unsigned integers with uniform spacing of mapped values.Surround 8 Coherence of the non-directional sound over the surrounding coherence directions.Range of values: [0.0, 1.0](Parameter is independent of number of directions provided.) Values stored as 8-bit unsigned integers with uniform spacing of mapped values.Remainder- 8 Energy ratio of the remainder (such as microphone noise) sound energy to- to fulfil requirement that sum of energy ratios is 1.total energy Calculated as energy of remainder sound / total energy.ratio Range of values: [0.0, 1.0](Parameter is independent of number of directions provided.)Values stored as 8-bit unsigned integers with uniform spacing of mapped values.
[0148] Furthermore, MASA spatial metadata can be provided in each 20-ms frame for each analyzed / provided direction (1 or 2) for 4 subframes and 24 frequency bands, which are defined by the table below:Table A.3. MASA spatial metadata frequency bandsBand LF (Hz) HF (Hz) BW (Hz)1 0 400 4002 400 800 4003 800 1200 4004 1200 1600 4005 1600 2000 4006 2000 2400 4007 2400 2800 4008 2800 3200 4009 3200 3600 40010 3600 4000 40011 4000 4400 40012 4400 4800 40013 4800 5200 40014 5200 5600 40015 5600 6000 40016 6000 6400 40017 6400 6800 40018 6800 7200 40019 7200 7600 40020 7600 8000 40021 8000 10000 200022 10000 12000 200023 12000 16000 400024 16000 24000 8000
[0149] Additionally, as shown in Fig.1 the UE 110 can comprise a controller 119 configured to control the capture (front end) processor 117 and in some embodiments the (IVAS) encoder 101. In some further embodiments the controller 119 can furthermore control the operation of the microphone array 115, for example, where at least some of the capture processor 117 functionality is incorporated or integrated into the microphone array (for example selection or activation of microphones to provide audio signals, detection of blocked microphones or microphone ports, or configuration of integrated analogue-to-digital conversion and / or frequency domain conversion).
[0150] As shown in Fig.1, the apparatus, UEs 110, 120 and server 111 can comprise suitable encoder and decoder functionality. For example, the apparatus 110 is shown comprising the (IVAS) encoder 101, the server 111 is shown comprising an (IVAS) decoder and encoder 121 and the apparatus 120 is shown comprising an (IVAS) decoder 131. In such a manner the audio signals representing the user or talker 1 103 can be encoded by the encoder 101 which generates a bitstream 106 to be passed to the server 111. The server 111 can then decode, (optionally then mix, e.g., with other objects and otherwise process the audio signals) and encode then to generate the bitstream 108 to be passed to the apparatus 120. The apparatus 120 can then decode the audio signals and present them to the user or listener ‘RX’ 141.
[0151] Although this example shows a teleconference application the encoder / decoder functionality can be applied to the (real-time) transmission or streaming of any suitable media.
[0152] The IVAS decoder / renderer for each of the UE 120, such as the teleconference apparatus 102 can be furthermore configured to handle multiple input streams that may each originate from a different encoder.
[0153] Fig.1 furthermore shows the room B apparatus or UE2 comprising a suitable (IVAS) renderer 135 configured to receive the decoded spatial audio signals 134 output from the (IVAS) decoder 131 and generate or render a suitable output audio signal, for example a binaural audio signal output, to be provided to suitable output apparatus, which in Fig.1 is shown as headphones (or more generally a headset comprising suitable audio transducers) 143 worn by listener 141. In some embodiments the listener position and orientation can be tracked, for example by suitable headset sensors and this information passed back to the (IVAS) renderer 135 (or UE2 120 more generally) which uses this information in generating or rendering the output audio signals. Although this example shows the listener using headphones any suitable output or output device can be employed in other embodiments, for example UE2 can be configured to generate a multichannel audio signal format (for example 7.1 multichannel audio signals) to be output to a multi-speaker arrangement in the room. A suitable output device can be, e.g., AR (augmented reality) or VR (virtual reality) headset. In another use case, the suitable output device can be a vehicle having the multi-speaker arrangement with an optional display.
[0154] Furthermore, the UE 120 can comprise a controller 133 configured to control the operations of the (IVAS) decoder 131 and (IVAS) renderer 135. In some embodiments the controller 133 in the rendererapparatus 120 is configured to communicate with or negotiate with a controller 119 within the capture apparatus 110, in this example UE1. In some embodiments the controllers 133 and 119 can communicate directly with each other (or as shown in Fig.1 via a server 111 or central controller) as shown by bitstreams 124. In some embodiments this communication can be incorporated within the bitstreams 106 and 108.
[0155] In some embodiments the communication or negotiation between UEs can be implemented by use of RTP (Real-time transport protocol). RTP is designed to carry a multitude of multimedia formats, which permit the transport of new formats without revising the RTP standard. To this end, the information required by a specific application of the protocol is not included in the generic RTP header. For a class of applications (e.g., audio, video), an RTP profile may be defined. For a media format (e.g., a specific video coding format), an associated RTP payload format may be defined. Every instantiation of RTP in a particular application may require a profile and payload format specifications
[0156] The profile defines the codecs used to encode the payload data and their mapping to payload format codes in the protocol field Payload Type (PT) of the RTP header.
[0157] For example, the RTP profile for audio and video conferences with minimal control is defined in RFC 3551. The profile defines a set of static payload type assignments, and a dynamic mechanism for mapping between a payload format, and a PT value using Session Description Protocol (SDP). The latter mechanism is used for newer video codecs such as RTP payload format for H.264 Video defined in RFC 6184 or RTP Payload Format for High Efficiency Video Coding (HEVC) defined in RFC 7798.
[0158] IVAS RTP payload format, originally provided as part of the IVAS Release 18 specifications, is currently being enhanced with the latest state described in TS 26.253 Annex A.
[0159] For a class of applications (e.g., audio, video, control), an RTP profile may be defined. For a media format (e.g., a specific video coding format), an associated RTP payload format may be defined. Every instantiation of RTP in a particular application may therefore require a profile and payload format specifications.
[0160] An RTP session can be established for each multimedia stream. Audio, control, and video streams may be implemented which use separate RTP sessions, enabling a receiver to selectively receive components of a particular stream. The RTP specification can furthermore be configured to recommend port numbers for RTP, and furthermore to recommend the use of the next odd port number for the associated RTCP session. A single port can be used for RTP and RTCP in applications that multiplex the protocols.
[0161] Each RTP stream can comprise RTP packets, and the RTP packet in turn can comprise a RTP header and payload pair.
[0162] Fig.12 shows an example IVAS RTP packet structure (following TS 26.253 Annex A) 1201. The RTP Header (and possible RTP Header Extension) 1203 follows a known RTP design. The structure 1201 furthercomprises a IVAS payload 1211. The IVAS payload 1211 comprises a payload header 1205 section, a (IVAS) frame data 1207 section and an optional PI (processing information) data 1209 section.
[0163] The payload header 1205 section can comprise different types of header bytes, for example: ToC (Table of Content) and E-bytes (Extra bytes, or further bytes which can be used to define or signal aspects of the payload).
[0164] The ToC bytes can be employed to describe the content of the frame data section (by indicating the size / bitrate for the data frames). The ToC bytes can also differentiate IVAS frames from EVS frames, in situations where the IVAS is operating in mono EVS mode.
[0165] The E-bytes can signal additional information, such as Codec Mode Requests (CMR) which indicate a request to change or signal at least one codec configuration parameter or rendering configuration parameter. The E-bytes may also be used to explicitly indicate the presence of the PI data section at the end of the payload.
[0166] The frame data 1207 section can include the IVAS data frames (and EVS data frames in a case of employing IVAS in mono mode). The data frames represent the encoded IVAS bitstreams. The bitstream includes the encoded IVAS audio data with possible additional metadata. The bitstream can also include initialization data or information, for example information about the input / coded format and sub-format of the encoded data (e.g., multichannel 5.1 format). Any EVS frames do not include such format data.
[0167] The PI data 1209 section can include (PI) processing information data and related headers. The PI data can be used to transmit any (non-audio) data that can be used to assist the rendering or processing of the audio data.
[0168] One example PI data is as follows:TypeForward direction PI type Description SDP indication Size (bytes) bitsDescribes theorientation of a spatial00000 SCENE_ORIENTATION fsco 8audio scene in unitquaternions.
[0169] The SCENEJDRIENTATION PI data describes the desired orientation of the capturer side scene. The frontal direction of the scene points towards the indicated capture frontal direction. For example, the scene orientation is described in 3GPPTS 26.253 Annex A clause 7.4.2.1 (Scene orientation). In the PI data, the azimuth and elevation values of the scene orientation can be transformed into (unit) quaternions. A frontal direction of zero azimuth and zero elevation corresponds to (w=0, x=1, y=0, z=0) in quaternions, see 3GPP TS 26.253 clause 7.4.3.1 for the IVAS coordinate system.
[0170] In other words, this PI data can therefore provide the capability to describe any offset or dynamic changes of the desired front or default front, as the intent or point of interest etc. from the capture side may be different than what the default orientation (according to the device front that generally corresponds to the format front) of the format indicates.
[0171] One reason to carry this data to the decoder / renderer is to perform only one rotation per rendering, if possible, because some spatial audio formats can be lossy when rotated, i.e., one cannot, e.g., rotate 30 degrees left (and render) and then rotate 30 degrees right (and render) and have the original signal. For example, one rotation can relate to a scene rotation, while another rotation can relate to a head-tracking data input. Preferably, these rotations are combined during rendering, and only one combined rotation is applied. Some spatial audio formats can provide lossless rotation.
[0172] For example, a microphone array can be placed on the table of a conference room. There may be several participants in the conference room. The array can capture, e.g., a MASA spatial audio format. When a certain talker is talking, an automatic algorithm (e.g., camera selection for a webcam system seen in some current implementations) could provide an orientation data to indicate that the front is now in the direction of the active talker. In some embodiments this data could be set on a separate user interface (Ul), e.g., to orientate the scene towards the main talker.
[0173] The SCENEJDRIENTATION PI data is applied to the audio frame with the same timestamp. The latest received SCENEJDRIENTATION PI data is used until a new SCENEJDRIENTATION PI data is received.
[0174] In some examples, another mechanism, e.g., RTP header extension mechanism, can be used to transmit same or similar data as shown above for the PI data frame mechanism.
[0175] A first step towards head-tracked rendering is the binauralization of the spatial audio. For example, this can be achieved by use of suitable HRIRs / HRTFs in the rendering. Figs.2a and 2b show an example scenario which illustrates a difference in experience between traditional stereo audio over headphones and binauralized headphone presentation.
[0176] In stereo playback, as shown in Fig.2a, a listener 203 wearing headphones 201 with left and right transducer channels is experiencing a conventional stereo audio signal rendering. In this example the audio sources ‘bird’ 207, ‘speaker T 205, and ‘speaker 2’ 209 are panned between the channels, and the sound seems to originate from inside the head.
[0177] In extreme cases, left and right channels for a stereo signal including its playback can furthermore be completely uncorrelated.
[0178] The binauralization rendering as shown in Fig.2b shows the same listener 203 wearing headphones 201 with left and right transducer channels. In this example the rendering producing the binauralized stereo audio signals presents the audio sources ‘bird’ 207, ‘speaker T 205, and ‘speaker 2’ 209 as being outside ofthe listener’s head. In this example the audio source ‘bird’ 207 is presented as being from approximately in front of the listener, ‘speaker 1’ 205 to the right of the listener, and ‘speaker 2’ 209 to the left of the listener. For example, the rendering provides directions for the sound sources and offers at least some level of externalization of the playback, e.g., based on the suitable HRIRs / HRTFs used in the rendering. The level of externalization may depend on, e.g., “roomification” of the binauralized audio. For example, early reflections and room reverberation effects will help to externalize the audio further.
[0179] Very “dry” binaural rendering, e.g., without any reverberation as part of the original captured audio scene or provided through the binaural rendering processing, despite use of HRIRs / HRTFs, can in some examples produce poor externalization.
[0180] Furthermore, binaural rendering can be enhanced by employing head-tracking data, e.g., provided by the headphone device or other suitable tracking mechanisms. When head rotation is considered as part of the rendering, the experience can be more life-like, i.e., the reproduced scene remains in place and for example a source to the front of the listener moves to the left when the listener rotates their head to their right.
[0181] Figs.2c and 2d further shows the effect of employing head-tracking to the rendering operation when the listener rotates their head (to the left).
[0182] Thus Fig.2c shows an example where the listener rotates their head approximately 90 degrees to the left but the rendering does not employ head-tracking and thus the audio sources, ‘bird’ 207, ‘speaker T 205, and ‘speaker 2’ 209 rotate with the head motion such all of the sources move relative to the audio scene by the head rotation. This type of motion in a VR or AR environment can produce an unnatural experience, for example any image of the ‘bird’ will not be located at the same position of the audio position.
[0183] Whereas Fig.2d shows an example of the head motion when employing head-tracking as the audio sources are presented independent of any head rotation. Thus as the listener turns their attention to the ‘speaker 209’ audio source the position of the audio source is maintained and does not rotate in the audio scene.
[0184] Additionally, Figs.3a and 3b shows a potential issue when employing binaural stereo audio rendering when head-tracking is not enabled. In such situations there is a potential accuracy issue in audio scene presentation of the spatial audio signals. As seen in Fig.3a and 3b, the rendering of the ‘bird’ 207 source being approximately to the front of the head results in potential confusion relating to source position on the front-back and the up-down axes. In other words, as shown in Fig.3b the listener would hear the ‘bird’ 207 source but would be unable to distinguish where the source is located as it could be in front 207 of the listeners head (the ‘correct’ location) or behind 307 the listeners head (the ‘incorrect’ location) and the listener has not ability to turn their head to discover which is the correct location (as the source is always presentedin the same manner. This is because of the listeners ears being on the left-right axis the left-right axis is only being able to be distinguished well.
[0185] This could be a significant problem for future spatial audio communications as there is currently no inherent guaranteed control of the audio scene orientation and the audio source positions relative to the user.
[0186] While in produced content source positions can be controlled, in communications and user generated content (UGC) of user captured content (such as a user capturing spatial audio of a music concert for a media channel, such as you-tube), they can be arbitrary.
[0187] These issues could be solved by head-tracking, as discussed herein, however as discussed above head-tracking is not always available.
[0188] Furthermore, when considering practical spatial audio communications (or UGC streaming) scenarios, it can be appreciated that a user could hold their recording / capture device in any arbitrary position and orientation relating to a sound scene being captured and furthermore the sound scene can itself be arbitrary, meaning relevant sound sources can be in any position relative to each other.
[0189] In other words, the sound scene can be inherently balanced or non-balanced in the spatial sense (grouping of sound sources based on their positions relative to the capture point in a way that they are easy to distinguish and follow provided that the capture provides, e.g., suitable spatial representation, the stereo pair is positioned at suitable angle, etc.).
[0190] For example as shown in Fig.5 the user 501 is attempting to capture an audio scene in stereo audio format comprising audio or sound sources 507 (to the left of the user) and 509 (to the right of the user) but doing so holding the capture device, a mobile phone 503, in portrait orientation and therefore comprising microphone positions 505 on the ‘top’ and ‘bottom’ edges with poor left-right axis discrimination and resulting in a rendering or presentation of the stereo audio format via headphones 553 of the listener 551 with very poor stereo separation of the sound sources. However, a capture device with a full 4-microphone 3D capture configuration producing an immersive audio format, e.g., MASA format or Ambisonics (SBA) format, as shown by the quad microphone arrangement shown in Fig.4 would result in an immersive experience as there would be better source separation.
[0191] Additionally a ‘front’ direction with respect to capture scene is often arbitrary and meaningless in terms of not directing the capture device towards a specific audio source or direction. In other words, the capture is generally not controlled or designed / set up for each capture scenario in mind, and the user may simply start or answer a call, i.e., “push record”.
[0192] As such device and sound scene orientations can change during the call or streaming instance, whilst the audio format typically has a front, e.g., C channel in 5.1 or 0-orientation in MASA, the sound scene recorded may not take this into account in any way (as explained above) and, furthermore, a spatial audio capturing device (e.g., a multi-microphone smartphone) can have one or more audio capturing configurationsthat provide the support for one or more input audio formats of the multi-mode immersive audio codec, where the capture formats and the input formats are not necessarily the same thus potentially requiring a downmix or format conversion, where some of the existing spatial audio information may be lost.
[0193] Furthermore the recording situation is dynamic where, e.g., the sound sources may be people moving or the user may hold the capture device in their hand and turn around or walk, etc.
[0194] Considering the above, there is therefore currently no single definition or truth of what, e.g., a stereo capture in a given real scene is. A live communications scenario is not equivalent to “canned” studio content that is carefully designed with artistic intent in mind. Instead, a communications capture has the main target of providing intelligible, and with the introduction of IVAS also more pleasing and immersive, communications.
[0195] Under these conditions, stereo capture, or stereo input generation, only can be very limiting. Specifically, when employing spatial audio capture (e.g., a MASA capture) for an arbitrary spatial sound scene, a lot of spatial information is obtained. However, when the situation or case is one which attempts to provide a specific output format, e.g., only a stereo audio input for the multi-mode codec (e.g., the IVAS codec), current methods remove the spatial audio information that would make the stereo experience better. For example, when there is the scenario where the scene comprises two talkers that are on the “front” and “back” of a stereo pair, they will both appear in the center instead of “left” and “right” of the scene. This reduces the stereo experience (making it more like mono), and will have negative implications on intelligibility of the communications, where the two users expect to be spatially captured.
[0196] Thus, in the following embodiments there is described apparatus and methods which attempt to improve the stereo capture and transmission based on additional spatial cues that are available.
[0197] In other words, a practical capture of a spatial sound scene for a multi-mode immersive audio codec can have sound source groupings that may be unbalanced in the spatial sense, e.g., in the left-right separation sense, (i.e., sources having positions relative to capture point in a way that they may be difficult to distinguish and follow given the default scene orientation rendering). In contrast, a studio recording or other such “canned” immersive content is generally designed to sound in some particular way by arranging the source positions, setting up capture devices in certain way, etc.
[0198] This aspect of practical capture becomes problematic, when a spatial scene is encoded in stereo or encoded spatially (e.g., captured as MASA and encoded as MASA) but it is presented to the receiving user in a more limited way than intended. For example, a MASA stream is primarily intended to be presented via head-tracked binaural rendering or via multi-channel loudspeaker setup. Instead, if it is presented as stereo or as non-head-tracked binaural audio, the relative positions of the captured sources become critical, since the positions dictate at least in part how immersive the presentation is, what is the quality of experience, and finally how happy the listener will be with the immersive service deployment. Such arbitrary variations and degradations of quality of experience should be avoided.
[0199] Therefore, the apparatus and methods as discussed in the following embodiments are configured to allow a user, for example a listener, to distinguish spatial audio sources and their relative directions in arbitrary scene orientations under binaural headphone presentation of encoded spatial audio stream (e.g., MASA stream) without head-tracking capability.
[0200] This enables the communication to maintain intelligibility of conversational spatial audio and UGC presentation for future spatial audio systems, such as those enabled by the 3GPP Immersive Voice and Audio Services standard.
[0201] Furthermore, in embodiments where the spatial scene is encoded as a spatial audio signal, then a user is not required to continuously control the spatial audio scene presentation (e.g., its component balance or orientation) using a user interface or other method for mode selection, e.g., utilizing renderer-side external orientation input (3GPP TS 26.253, clause 7.4.5). Instead, such mechanisms as described herein aim to provide a guaranteed quality of experience for immersive voice services.
[0202] The concept as discussed in the following embodiments is thus one which relates to immersive audio capture for (multi-mode) immersive audio coding, transmission, and rendering on a multi-microphone audio capture device. In these embodiments there is first described apparatus and methods which aim to improve the robustness of stereo input audio generation against dynamic spatial sound scene source directions for improved stereo listening experience.
[0203] For example, in some embodiments, this can be achieved by analysis of spatial audio capture source positions to generate scene orientation data input corresponding to an improved spatial balance or left-right separation for the targeted (binaural) stereo channels and pre-rotating the spatial audio scene according to said scene orientation to provide a (binaural) stereo downmix as stereo input for the multi-mode immersive audio encoder.
[0204] Second, in some embodiments there can be provided apparatus and methods to improve spatial audio reproduction perception in binaural headphone rendering of immersive audio streams (containing, e.g., transport audio signals and spatial metadata) upon obtaining information that head-tracking at receiving device is not offered.
[0205] In such embodiments this can be achieved by analysis of spatial audio capture to generate spatial (or scene) orientation data input corresponding to improved spatial balance or left-right separation for the targeted (binaural) stereo channels and transmitting the generated spatial orientation data as auxiliary rendering support data to receiving device.
[0206] Furthermore, in some embodiments the transmitted spatial (or scene) orientation data is used for non-head-tracked binaural rendering with improved spatial balance or left-right separation of audio sources.
[0207] The above is achieved for IVAS encoder inputs based on obtaining renderer configuration information (e.g., SDP offer-answer or PI data describing use / non-use of head-tracking) and utilizing this in(DaCAS) spatial audio capture and analysis to generate scene orientation input (for example the SCENEJ3RIENTATI0N PI data as described above) that is transmitted using the IVAS RTP payload format together with associated immersive audio stream, e.g., MASA bitstream.
[0208] These embodiments are particularly targeted for improved headphone listening of spatial audio content, e.g., spatial communications, and where head-tracking is not available. In such embodiments there can be provided an improved spatial source separation in (binauralized) headphone listening, e.g., by increasing spatial balance or left-right separation of sources (and thus also reducing effect of front-back confusion). This can be particularly suitable for spatial audio scenes that have been compressed, e.g., due to low-bit rate communications, such as relevant for 3GPP IVAS.
[0209] In some embodiments these methods can be applied also as post processing step at the decoder to determine a desired rendering.
[0210] Fig.6 shows an example block diagram of parametric spatial audio input generation, coding, and synthesis framework which can be implemented within the apparatus shown in Fig.1 according to some embodiments.
[0211] The input to the system are input audio signals 600, which can be from various sources including: Two or more microphones integrated or connected onto a mobile device (e.g., a smartphone); other microphone arrays, e.g., B-format microphone, planar microphone array or Eigenmike; Ambisonic signals, e.g., first-order Ambisonics (FOA), higher-order Ambisonics (HOA); Loudspeaker surround mix and / or objects; Artificially created spatial mix, e.g., from audio or VR teleconference bridge; or any combinations of the above.
[0212] The input audio signals 600 are provided to the analysis processor 603 and to the transport signal generator 601.
[0213] The analysis processor 603 is configured to estimate spatial metadata (in frequency bands) from the input audio signals 600. For all of the aforementioned input types, there exists known methods to generate spatial metadata (e.g., directions and direct-to-total energy ratios) in frequency bands, and these methods are not explained here in detail.
[0214] However, an example implementation for the analysis processor 603 can be that as described herein. The analysis processor 603 for example in some embodiments is configured to apply a timefrequency transform to the input audio signals to generate frequency bands representations of the input audio signals.
[0215] Then, when the input audio signals are provided by a mobile phone microphone array (such as shown in Fig.4) the analysis processor 603 is configured to estimate delay-values between microphone pairs that maximize the inter-microphone correlation, and determine a corresponding direction value to that delay(as is described in further detail in UK patent application GB1619573.7), and formulating a ratio parameter based on the correlation value.
[0216] When the input audio signals are a FOA signal, the analysis processor 603 is configured to determine an intensity vector, based on which the direction parameter is formulated, and compare the intensity vector length to the overall sound field energy estimate to determine the ratio parameter. This method is known in the literature as Directional Audio Coding (DirAC).
[0217] When the input audio signal is a HOA signal, the analysis processor 603 is configured to divide the HOA signal into multiple sectors, in each of which the method above is utilized. This sector-based method is known in the literature as higher order DirAC (HO-DirAC). In this case, there can be more than one simultaneous direction parameter per frequency band corresponding to the multiple sectors.
[0218] When the input audio signals are loudspeaker surround mix and / or objects, the analysis processor 603 is configured to convert the signal into a FOA / HOA signal(s) and to analyse direction and ratio parameters as above.
[0219] UK patent application GB2574238 describes an example of mixing in the MASA domain (based on merging of spatial audio parameters from at least two separately captured / generated spatial audio inputs, where the spatial audio input may correspond to any of the capture / generation operations described above).
[0220] Various other methods (and also other spatial metadata sets) exist in the literature and can be employed by the Analysis processor 603 to generate the spatial metadata (which can be determined in frequency bands (time-frequency (TF) tiles)). The spatial metadata 604 can comprise directions and ratios in frequency bands but can comprise any of the metadata types listed in the background section (or any other suitable parameter which assist in the characterization of the audio scene being captured or designed).
[0221] The transport signal generator 601 furthermore is configured to receive the input audio signals 600 and generate transport audio signals 602. In some embodiments the transport signal generator 601 is configured to generate a stereo or mono audio signal. The generation of transport audio signals 602 can be any suitable and known method, examples including:
[0222] When the input audio signal is a mobile phone microphone array, selecting a left-right microphone pair, and applying any suitable processing to the signal pair, such as automatic gain control, microphone noise removal, wind noise removal, and equalization.
[0223] When the input audio signal is a FOA / HOA signal, formulating directional beam signals towards left and right directions, such as two opposing cardioid signals.
[0224] When the input audio signal is a loudspeaker surround mix and / or objects, generating a downmix signal that combines left side channels to left downmix channel, and same for right side, and adds centre channels to both transport channels with a suitable gain.
[0225] In some embodiments the transport signal generator 601 is configured to only bypass the input audio signal. For example in some situations where the analysis and synthesis occurs at the same device at a single processing step, without intermediate processing. Furthermore the number of transport channels can also be any suitable number, and in some embodiments more than one or two transport channel audio signals.
[0226] The transport audio signals 602 and the spatial metadata 604 can be passed to the encoder 605. The encoder 605 is configured to encode the transport audio signals 602 and the spatial metadata 604. In addition, the encoded transport audio signals 602 and the spatial metadata 604 can be multiplexed to a single data stream 606 (e.g., an IVAS bitstream). The data stream 606 can be transmitted or stored.
[0227] The data stream can furthermore be passed to or retrieved by the decoder 607. The decoder 607 is configured to decode (and possibly demultiplex) the data stream 606 into transport audio signals 608 and spatial metadata 610, which are forwarded to the synthesis processor 609.
[0228] The synthesis processor 609 is configured to receive the decoded transport audio signals 608 and decoded spatial metadata 610, and determine or generate output audio signals 612. In some embodiments the output audio signals 612 are generated as any suitable output format, e.g., multichannel loudspeaker signals or binaural audio signals.
[0229] The synthesis processor 609 is configured to render the spatial or output audio signals based on suitable method, for example such as described in further detail in PCT application WO2019086757A1, and are not explained here in further detail. However, as a simplified example, the rendering by the synthesis processor 609 according to some embodiments can be that performed for a loudspeaker output as follows:
[0230] The transport audio signals 608 are divided by the synthesis processor 609 to direct and ambient streams based on the direct-to-total and diffuse-to-total energy ratios (or any equivalent energy information).
[0231] The direct stream is rendered by the synthesis processor 608 based on the direction parameter(s) using amplitude panning.
[0232] The ambient stream is rendered by the synthesis processor 608 using decorrelation.
[0233] The direct and the ambient streams are combined by the synthesis processor 608 to generate the output audio signals. The output signals can, e.g., be reproduced using a multichannel loudspeaker setup or headphones.
[0234] It should be noted that the processing blocks shown in Fig.6 can be in same or different processing entities. For example, as shown in Fig.7 in some embodiments, microphone signals from a mobile device are processed with a spatial audio capture 701 system. The spatial audio capture system in some embodiments comprises the analysis processor 703 (which in this example is configured to receive renderer configuration data 700 and output spatial orientation data 702), and the transport signal generator 601, and the resultingspatial metadata 604 and transport audio signals 602 (e.g., in the form of a MASA input stream) are forwarded to an encoder 605 (e.g., an IVAS encoder), and which implements the encoding as described above.
[0235] In some embodiments, the input audio signals 600 (e.g., 5.1 signals) are directly forwarded to an encoder (e.g., an IVAS encoder), which is configured to implement the analysis processor 603, the transport signal generator 601, and encoder 605 functionalities.
[0236] In some further embodiments, the input audio signals 600 can be two (or more) input audio signals, where the first audio signal(s) is processed according to the example system as shown in Fig.6 (resulting, e.g., in a MASA stream as an input for the encoder) and a second audio signal(s) is directly forwarded to an encoder (e.g., an IVAS encoder), which contains the functions of the analysis processor, the transport signal generator, and the encoder. The audio input signals may then be encoded in the encoder independently or they may, e.g., be combined in the parametric domain according to what may be called, e.g., MASA mixing (for example as further described, e.g., in UK patent application GB2574238).
[0237] Correspondingly, in some embodiments a decoder (e.g., an IVAS decoder) can comprise a decoder and the synthesis processor functionality is implemented in a different system (e.g., an external renderer); or the decoder (e.g., an IVAS decoder) can implement the functionality of both the decoder and the synthesis processor (e.g., integrated IVAS renderer). In some embodiments, the decoder block comprises more than one decoder instance and therefore is configured to process more than one incoming data stream in parallel.
[0238] As shown in the example of Fig.7 the spatial audio capture 701, and in examples specifically the analysis processor 703, is configured to receive as an input renderer configuration data 700 and generates spatial orientation data 702 in addition to previously considered outputs, the spatial metadata 604 and transport audio signals 602 from the transport signal generator 601).
[0239] In these embodiments the spatial orientation data 702 is here not passed to the (IVAS) encoder 605 but is transmitted to the receiver / renderer (for example via the negotiation information) as auxiliary data, as described below.
[0240] The implementation of these embodiments can be prefaced by discussion of potential scenarios. In these embodiments while there is shown 2D examples it would be understood that they can be extended to 3D scenarios. Furthermore, it is worth noting that horizontal rotation (yaw) is actually often entirely sufficient for conversational use cases. Considering other 3D rotations (pitch, roll) may sometimes be useful, e.g., for UGC and professionally produced spatial audio content.
[0241] Fig.9a to 9c presents an example of user’s head rotation in a first example audio scene with and without head-tracking. This audio scene is characterized by two main audio sources (A 951 and B 953). For example, the sources may represent two talkers in a conversational spatial audio scene.
[0242] Thus as shown in Fig.9a an audio scene when the user 901 has a first head position or orientation 900 (which is shown by the arrow 955 representing an audio scene orientation) and the audio or soundsource A 951 is approximately -80 degrees (or to the right) of the first head orientation 900 and the audio or sound source B 953 is approximately +60 degrees (to the left) of the first head orientation 900.
[0243] Furthermore as shown in Fig.9b is shown the audio scene when the user 901 has a second head position or orientation 910 and head tracking has been applied (which is approximately 60 degrees to the left) and the audio or sound source A 951 is approximately -80-60=-140 degrees (to the right) of the second head orientation 910 and the audio or sound source B 953 is approximately 60-60=0 degrees of the first head orientation 910. Thus when head-tracking is available for the user, the user can, e.g., turn left (applying a head rotation of, say, 60 degrees) to better concentrate on certain audio component (e.g., source B 953).
[0244] Additionally as shown in Fig.9c is shown the audio scene when the user 901 has a second head position or orientation without head tracking 920 (which is approximately 60 degrees to the left of the first head position). But with no head tracking the audio scene has also rotated 60 degrees to the right as shown by arrow 957 and the audio or sound source A 951 remains approximately -80 degrees (to the right) of the second head orientation and the audio or sound source B 953 is approximately 60 degrees (to the left) of the second head orientation. However, in this example the corresponding experience without head-tracking is also reasonable due to good left-right separation provided by the audio scene composition.
[0245] Figs.10a to 10c presents an example of user’s head rotation in a second example audio scene with and without head-tracking. This audio scene is characterized by two main audio sources (C 913 and D 911). For example, the sources may represent the same two talkers in a conversational spatial audio scene.
[0246] Thus as shown in Fig.10a an audio scene when the user 901 has a first head position or orientation (which is in line with the arrow 955 representing an audio scene orientation) and the audio or sound source C 913 is approximately 10 degrees to the left of the first head orientation (in other words almost ahead of the user) and the audio or sound source D 911 is approximately 170 degrees to the left of the first head orientation 900 (in other words almost behind the user).
[0247] Furthermore as shown in Fig.10b is shown the audio scene when the user 901 has a second head position or orientation and head tracking has been applied (which is approximately 60 degrees to the left) and the audio or sound source C 913 is approximately 10-60—50 degrees (to the right) of the second head orientation and the audio or sound source D 911 is approximately 170-60=110 degrees (to the left) of the first head orientation 910. Thus, when head-tracking is available for the user, the user can, e.g., turn left (applying a head rotation of, say, 60 degrees) to better improve any front-back confusion to bring the audio sources C and D more aligned with the left-right axis through user’s head motion.
[0248] Additionally, as shown in Fig.10c is shown the audio scene when the user 901 has a second head position or orientation without head tracking (which is approximately 60 degrees to the left of the first head position). But with no head tracking the audio scene has also rotated 60 degrees to the right as shown by arrow 957 and the audio or sound source C 913 remains approximately 10 degrees to the right of the secondhead orientation and the audio or sound source D 911 is approximately 170 degrees (to the left) of the second head orientation. In this example the non-head-tracked experience suffers from front-back confusion and poor spatial separation. There are very limited cues for the user to understand the relative positions of the two sources, and there may therefore be significant spatial masking. This leads to reduced intelligibility and increased listener fatigue and frustration. Thus, some of the advantages of spatial audio are lost.
[0249] As shown in Figs.9a-9c and 10a-10c there is a significant difference for the user without headtracking when considering the first example spatial audio scene and the second example spatial audio scene. The first example Figs.9a-9c provide a high-quality experience (although limited vs. same experience with head-tracking), while the second example Figs.10a-10c provides an example that is more similar to traditional mono / stereo experience. This effect is because of the exact composition of the scene and the (arbitrary) orientation relative to the user (i.e., listening position / orientation).
[0250] In some embodiments the synthesis processor (or a suitable point in the system) is configured to improve the experience by applying an automatic modification of the audio scene. For example, this modification can be applied during rendering (e.g., in “Synthesis processor) in order to benefit from the leftright separation also for the user without head-tracking.
[0251] As will be explained herein, this audio modification can be employed based on suitable analysis of the audio scene (or spatial capture analysis) and for example is configured to control when head-tracking information itself is not received.
[0252] Specifically, such as shown in Fig.7, the audio modification can be based on:Obtaining at spatial audio capture 701 of the transmitting UE information (e.g., renderer configuration data 700, such as head-tracking capability parameter in IVAS SDP offer-answer or head-tracking availability / usage data, e.g., HEAD_TRACKING reverse direction PI type in IVAS RTP payload); andUsing this renderer configuration data 700 in the spatial audio capture 701 to determine spatial orientation data that is then provided to receiving UE using, e.g., SCENEJDRIENTATION forward direction PI type as part of the IVAS RTP payload.
[0253] In some embodiments the idea of intelligent format converting downmix generation is introduced.
[0254] Thus, based on the example shown in Fig.7 the analysis and the determination of the spatial orientation data 702 can be implemented based on the renderer configuration data 700. For example, the renderer configuration data 700 in some embodiments can indicate whether the renderer is implementing head-tracking or no head-tracking (or whether the renderer or headset has head-tracking capability).
[0255] The renderer configuration data 700 or more generally information from the receiver or renderer can take any suitable form including but not limited to:of=STEREO SDP parameter indicating stereo output formattf=STEREO SDP parameter indicating stereo target formatht=NO SDP parameter indicating no head-tracking capabilityHEAD_TRACKING reverse direction PI type (or similar functionality, e.g., using PI data frames, RTP head extension mechanism, or other signaling) indicating no head-tracking support or usage at receiving UE
[0256] As discussed herein, in some embodiments as also discussed later it is possible to modify the encoder input (i.e., pre-rotate it) before feeding it to the encoder instead of transmitting the associated orientation data. An advantage of transmitting the spatial orientation data instead of applying a pre-rotation is that by not modifying the spatial audio, a higher quality level can be maintained.
[0257] The receiving user may utilize other orientation controls, e.g., a III to rotate the scene, and this would result in two independent rotations of the scene thus degrading the quality: the first rotation would modify the input, which would then be quantized / compressed, even heavily, depending on the codec bitrate, and a second rotation would further modify the (now compressed and decoded) signals. By transmitting the desired rotation information, a single rotation can always be performed as it is possible to combine the corresponding orientation data.
[0258] A suitable method, e.g., one described below is then used to generate “Spatial orientation data” that improves the left-right separation of the spatial audio content for a bi naurally rendered output with no headtracking data being applied. This data is transmitted via RTP using the SCENEJDRIENTATION forward direction PI type.
[0259] In these embodiments, as shown by Fig.8 a spatial audio capture 801 comprises the transport signal generator 601 configured to receive the input audio signals 600 and generate the transport audio signals 602.
[0260] Furthermore, is shown an analysis processor 803. The analysis processor 803 is configured to receive the input audio signals 600 and can furthermore receive coder configuration data 810. The analysis processor 803 in some embodiments comprises a spatial audio analyzer and metadata generator 813. The spatial audio analyzer and metadata generator 813 is configured to obtain the input audio signals 600 and generate the spatial metadata 604 in a manner as described above.
[0261] Furthermore, the analysis processor 803 comprises a scene orientation determiner 815 which is configured to receive at least some of the spatial metadata information from the spatial audio analyzer and metadata generator 813 and further receive the coder configuration data 810. The coder configuration data 810 can in some embodiments be, e.g., a stereo or a related parameter indicating that stereo input is required for the encoder 805 instead of the (parametric) spatial audio input, e.g., MASA. The scene orientationdeterminer 815 is configured to then generate spatial orientation data 802 based on the spatial metadata and coder configuration data 810.
[0262] The spatial audio capture 801 furthermore comprises an intelligent spatial-to-stereo downmixer 817 configured to receive the spatial orientation data 802 from the scene orientation determiner 815 and performs an intelligent spatial-to-stereo downmix according to this orientation data to produce the downmix transport audio signals 804. In some embodiments this is output to the encoder 805 and can be employed as a stereo audio input.
[0263] In some embodiments the downmix operation can be based on the IVAS parametric binaural renderer, where external orientation input is based on the spatial orientation data.
[0264] The transport audio signals 602, spatial metadata 604 and downmix transport audio signals 804 can be passed to the encoder 805 which is configured to generate the data stream 806.
[0265] In implementing these embodiments, the stereo input quality from a spatial capture device can be improved by applying an automatic spatial audio scene orientation modification prior to binaural rendering of the spatial audio representation to obtain a (binaural) stereo signal that therefore benefits from improved spatial balance or left-right separation.
[0266] In some embodiments, this audio modification can be provided based on rendering tools used for head-tracking by introducing novel spatial capture analysis and control when only stereo encoding is available.
[0267] Thus in such embodiments, the audio modification and improved (binaural) stereo generation can be based on:Obtaining at spatial audio capture 801 of the transmitting UE information (e.g., coder configuration data 810, such as coded format parameter in IVAS SDP offer-answer, andEmploying the information in the spatial audio capture to derive spatial orientation data 802 that is then used to pre-rotate a spatial audio scene for stereo audio input generation, e.g., using an intelligent spatial-to-stereo downmixer 817.
[0268] The obtaining of the spatial orientation data 802, for example via the scene orientation determiner 815 as shown in the embodiments represented with respect to Fig.8 or the spatial orientation data 702 from the analysis processor 703 as shown in the embodiments represented with respect to Fig.7 can be based on firstly finding a spread of the audio sources that indicates a preferred left-right orientation for the rendering target (stereo or binaural stereo audio input). The preferred orientation is based on finding either the smallest angle - or alternatively the most forward-facing angle - spanning over the (dominant) audio sources in the scene as given by the spatial audio capture 701 / 801. It is understood that the terms scene orientation data and spatial orientation data are interchangeable.
[0269] Figs.11 a to 11 d show example determinations for the spatial orientation data 802 or the spatial orientation data 702. According to the method, angles between the source positions are determined.
[0270] Thus for example Fig.11a shows the example shown in Fig.10a where a user 901 has a first head position or orientation (which is in line with the arrow 955 representing an audio scene orientation) and the audio or sound source C 913 is approximately 10 degrees to the left of the first head orientation (in other words almost ahead of the user) and the audio or sound source D 911 is approximately 170 degrees to the left of the first head orientation 900 (in other words almost behind the user). Furthermore, is shown the angle 1100 between the two sources.
[0271] Fig.11b shows furthermore a mid-point 1101 between the two sources. The mid-point 1101 can for example be denoted a reference direction that relates to the angle and the sources. The reference direction can in some embodiments be other than the mid-point of the angle. For example, the reference direction can be a weighted or perceptually motivated average of the source directions, for example a weighted average where the weights are based on the energy or energy ratios. In the examples as shown in Figs.11a and 11b thus show a spread angle 1100 which is 160 degrees with a reference angle 1101 of approximately 90 degrees.
[0272] Fig.11 c and 11 d further show the effect of audio scene modification to the example shown in Figs.11 a and 11 b based a modification of the reference direction and / or the source directions or angles.
[0273] In some embodiments the aim of the audio scene modification is to improve or maximize the spatial balance or left-right separation for the scene by applying a rotation that utilizes knowledge of the source spread - the angle between the ‘dominant’ sources and the reference direction.
[0274] For example in general, especially for dynamic scenes, a first modification is one which aims to minimize the modification.
[0275] An example of a ‘minimized’ modification is shown with respect to Fig.11 c. The modification in this example is one where the scene of Fig.11a is rotated to the right (clockwise as shown by the arrow 1131) such that audio source D reaches the user’s left-right axis on the left-hand side. On that side, the scene thus takes full advantage of the separation. Thus, in this example the reference direction 1101 (and the audio sources) are rotated by -80 degrees (clockwise) such that the audio source D 911 is now at 90 degrees (to the left) and the audio source C 931 is now approximately -70 degrees (to the right).
[0276] In some embodiments the modification aims to bring the reference direction to user’s front, which provides a natural listening experience. As shown in Fig.11d the audio scene rotation as shown by the arrow 1141 is one where the reference angle 1101 is now in alignment with the head direction or orientation 955. In this example the reference direction 1101 (and the audio sources) are rotated by -90 degrees (clockwise) such that the audio source D 911 is now at 80 degrees (to the left) and the audio source C 931 is now approximately -80 degrees (to the right).
[0277] The determination of the spread (and / or the reference angle or direction) can be based on any suitable method. For example, Figs.13a to 13f show an example method which finds the spread of the sources.
[0278] For example, Fig.13a shows the user 1300 and the audio scene orientation or user head orientation or direction 1301 which is represented as an arrow on a ring surrounding the user.
[0279] Furthermore, as shown in Fig.13b there is shown an example distribution of audio sources A 951, B 953, C 931 and D 911 (which are distributed in a manner similar as described above).
[0280] Then can be defined and applied overlapping sub-angles that cover the whole angle / space. For example, a minimum number of sub-angles to cover yaw rotation is three 120-degree sub-angles but as shown in Figs.13c to 13f there are shown four 180-degree sub-angles which overlap by 90 degrees with both adjacent sub-angles. However, any number of sub-angles and / or overlaps can be defined and employed.
[0281] Then for each of the sub-angles a dominant source direction is determined. For example a dominant source can be the source with the highest energy or energy ratio in each of the overlapping sections.
[0282] For example, as shown in Fig.13c a first ‘front’ sub-angle 1320 is defined. The dominant source in the ‘front’ sub-angle 1320 is audio source A 951 (but the sub-angle also covers source C 931 and B 953).
[0283] Then as shown in Fig.13d a second ‘left’ sub-angle 1330 is defined. The dominant source in the ‘left’ sub-angle 1330 is audio source C 931 (but the sub-angle also covers source D 911 and B 953).
[0284] Furthermore, as shown in Fig.13e a third ‘rear’ sub-angle 1340 is defined. The dominant source in the ‘rear’ sub-angle 1340 is audio source D 911 (and the sub-angle does not cover any other source).
[0285] Finally, as shown in Fig.13f a fourth ‘right’ sub-angle 1350 is defined. The dominant source in the ‘right’ sub-angle 1350 is audio source A 951 (and the sub-angle does not cover any other source).
[0286] These selected or surviving directions can then be used to define the spread angle and furthermore the reference direction or orientation. By employing the overlapping sections, it is possible to consider all spatial directions, in other words, the method does not simply look for K most dominant directions over the whole space, which may not capture all of the spatially useful information.
[0287] However, the method is configured to determine the N dominant directions that are (somewhat) sparse. Next, N <= M, where M is the number of the overlapping sections. Here, N can be equal to M in special cases, such as certain two directions having exactly the same energy. Furthermore, in embodiments, a dominant direction needs to fulfil the condition of being above an energy threshold (e.g., directional energy relative to total energy).
[0288] It would be appreciated that variances to this example method can furthermore be applied based on the audio scene format.
[0289] For example, when a HOA scene is received, analysis can be applied based on HO-DirAC to derive the dominant directions in different sectors.
[0290] In some embodiments the coverage of the sub-angles may not be even. For example, analysis of a 5.1 audio scene could provide a higher weight to the front region that has three loudspeakers (L-C-R). For example, this weighted coverage could be achieved in some embodiments by specific coverage sub-angles ranges for various angles.
[0291] In some other embodiments the spread (and / or the reference angle or direction) can be based on an alternative approach of finding for the most dominant direction and then require each subsequent direction to differ by a suitable angle from the previously found directions.
[0292] Furthermore, in some embodiments the processing is based on information that stereo or binaural stereo input is requested for the (IVAS) immersive audio session.
[0293] For example in some embodiments, such as shown in the example shown in Fig.8, the information (coder configuration data 810) could take various forms depending on the context. In some embodiments, the IVAS SDP offer-answer with cf or cf-recv (or cf-sub or cf-sub-recv) parameter indicating Stereo is considered, which is then selected as the audio input format for the IVAS immersive audio session in the send direction of the IVAS UE (e.g., implementing DaCAS audio capture). In these embodiments stereo can be stereo or binaural stereo (a pre-binauralized stereo signal that does not support head-tracking).
[0294] In some embodiments, the apparatus and methods (for example the scene orientation determiner 815 or analysis processor 703 / 803 as part of the spatial audio capture 701 / 801) can be configured to spatial orientation data 702 / 802 that improves the left-right separation of the spatial audio content for stereo or binaurally rendered stereo where no head-tracking data can anymore be applied. This orientation data is then utilized to pre-rotate the scene for a balanced spatial-to-stereo downmix.
[0295] As discussed above in some embodiments the analysis processor is configured to receive the input audio signals, where the spatial audio analyzer and metadata generator (e.g., in case of parametric spatial audio capture) determines the pertinent directional data. Then based on the generated metadata, the scene orientation determiner is configured to determine the audio scene target orientation. For example, the target orientation can be understood as finding a rotation to the audio scene such that effect as shown in Fig.11c and 11 d are achieved.
[0296] In some embodiments this rotation is not applied at a single time instant as this would lead to jittery and unstable perception of the scene rendering (source positions, or in case of stereo, source panning). Instead, the analyser can be configured to update the target in each processed frame (time instant) and determines a new rotation towards this target such that over time (e.g., over several frames) the target orientation is reached. The rotation value can be called, e.g., spatial orientation data that is sent to intelligent spatial-to-stereo downmixer that is configured to rotate the audio scene accordingly and downmixes the audio signals to stereo or binaural stereo. For example, this data could indicate clockwise rotation of 15 degreesalong the azimuth in one example frame. It would be understood that this rotation smoothing can be applied to all of the scene rotation approaches and not only to the intelligent spatial-to-stereo downmixer example.
[0297] Fig.14 shows a further example wherein spatial audio capture 1401 in this example comprises an analysis processor 703 which comprises the scene orientation determiner 1405, the scene orientation determiner 1415 is similar to the scene orientation determiner 805, but receives the renderer configuration data 1400 (rather than coder configuration data shown in Fig.8) and outputs the spatial orientation data 802 to an auxiliary data packer 1407. The auxiliary data packer 1407 encodes or transmits the spatial orientation data into an auxiliary data stream and passes this to a receiver or renderer where this spatial orientation data is then used to change the scene orientation.
[0298] Fig.15 presents example method steps for generating improved or balanced stereo audio signals or improved spatial balance according to some embodiments.
[0299] It is understood that the same overall orientation determination principles can be used for different audio capture solutions, e.g., depending on the capture audio format and spatial audio capture implementation. To provide a suitable example the following implementation is described in context of a metadata-assisted spatial audio (MASA) format.
[0300] Thus firstly, as shown by Fig.15 by 1501 is the operation of obtaining input audio signals.
[0301] Furthermore there is obtained session information on the audio format as shown by Fig.15 by 1503. In some embodiments this session information can be the renderer configuration data 700 or coder configuration data 700. As described herein this data can comprise information as to the output format (binaural stereo audio or stereo audio).
[0302] In some embodiments the method is configured to check or determine whether the output spatial audio is a binaural stereo audio signal format or not. This check is shown in Fig.15 by 1505. The following operations are then performed when the output format is a binaural format.
[0303] In some embodiments as shown by 1507 in Fig.15, there is performed an analysis of the spatial audio scene to determine whether the spatial balance or left-right separation is suitable. For example, the audio signals can be in a frequency domain S(i, b, n), where i is the channel index, b is the frequency bin, and n is the temporal frame. This can be a result of the captured audio signals being converted into the frequency domain or obtaining the audio signals in a frequency domain. The following examples show analysis in the frequency domain but in some embodiments the analysis is performed on frequency band filtered time domain audio signals.
[0304] The analysis in some embodiments is configured to obtain or generate spatial metadata containing at least a direction (0(fc, ri), (p(k, ri)) (azimuth, elevation) and a direct-to-total energy ratio r k, ri), where k is the frequency band. In some embodiments the input audio signals are any input format and the audio signals and the spatial metadata can be determined in spatial audio capture block.
[0305] In some embodiments the analysis is further configured to generate or obtain bandwise energies which can for example be determined based on the followingkfc.highE(k, ri) = |S(i, b, n) |2I frfc,10W
[0306] where bk>iowis the lowest bin of the band k and bk;highis the highest bin.
[0307] In some embodiments the frequency bands in the energy estimation match the frequency bands of the spatial metadata.
[0308] In some embodiments the analysis further comprises processing the obtained spatial metadata, for example, the direction (0(fc,n), (fc, n)) is converted to a Cartesian coordinate vector weighted by the energy E k, r) and the direct-to-total energy ratio r k, )x(k,ri) = E(k, ri) r(k, ri) cos 6 k, n) cos (p(k, ri)y(k,ri) = E(k, ri) r(k, ri) sin 6 k, n) cos (p(k, ri)z(k, ri) = E(k,ri)r(k,ri) sin < >(fc,n)v(k, ri) = [x(k,ri),y(k,ri),z(k, ri)]T
[0309] These vectors can then be summed over a suitable frequency range and time period. In this example, the frequency range is the whole audio spectrum (i.e., all the frequency bands from 0 to K - 1), but in some alternative implementations the estimation can also be done in smaller frequency ranges. The time period is in this example from N to N2, resulting in the following equationw2K-lVN'12 = V('k’N1k=O
[0310] The energies are summed over the same periodw2K-lENI2=’ E(k, ri)N1k=O
[0311] Using the two equations above, a spatial average vector is computed for the direction as
[0312] In some embodiments there can be more than one direction 0d(k, ri), cf>d(k, ri)). For example, the MASA format supports provision and transmission of two simultaneous directions, and as such the analysis or spatial audio capture can implement two-direction data.
[0313] In some embodiments the input audio signals can correspond to format combinations such as, e.g., OMASA (MASA + object(s)) or more generally microphone array signals and separate object signals. Analysis of such inputs results in additional directional components (e.g., the objects) that can be considered in the same vein. When the calculation is repeated, or a similar calculation employed, for each of thedirections and objects, the analysis is able to determine at least two spatial average vectors, vd vl2. The vectors take into account the signal energy making them already perceptually motivated.
[0314] In some embodiments, additional perceptual weighting can be used (in addition to energy and direct-to-total energy ratio).
[0315] In some embodiments the analysis further is configured to determine the scene spread based on the received or analysed directions.
[0316] This scene spread analysis can follow the overlapping sub-angle method discussed above or any other suitable method. For example, for each sub-angle the method can be configured, where there is more than one source direction to drop at least one direction that is not considered relevant, for example removing less energetic audio sources). For example, if one survivor per sub-angle is allowed (Nsurv= 1), following simple pseudo-code can be used at each frame:For seg_ovlapDetermine Ndircandidate source directions in current overlapping segment Designate Nsurvhighest-energy source directions as survivorsRemove any non-survivor source directions, where each removed source direction is not more considered in current frame (i.e., for remaining sub-angles)End for
[0317] The spread can then be determined as one of:the smallest angle between the two survivors farthest from each other, or the smallest angle between the two survivors farthest from each other that includes the front of the listener.
[0318] Where in some embodiments the latter being the preferred implementation.
[0319] Once the spread is determined, then the reference direction is determined, based on how and which target orientation is defined.
[0320] In some embodiments since the spatial average vectors, vd vl2are perceptually motivated, the vector mean of the survivor vectors can be used as the reference direction.
[0321] It is noted that this reference direction can dynamically move or jump significantly over time depending, e.g., on the source dynamics present in the scene. Such motion or jumping is an unwanted property to drive the rotation of the whole scene. Thus, in some embodiments a suitable smoothing of the reference direction (or any suitable parameter relating to the reference direction or any equivalent measure) can be applied.
[0322] It makes no significant difference whether the smoothing is applied to the reference direction or to the rotation information (spatial orientation data) which is generated from the reference direction. In thefollowing examples the smoothing is applied to the reference direction. This also allows, for example the spatial orientation data to be transmitted or sent to a separate module (e.g., for transmission), which can then use it reliably without knowledge of the preceding smoothing or requirement for smoothing is required.
[0323] In some embodiments the spatial orientation data for modification of the scene orientation is then determined as shown in Fig.15 by 1509. As described above this spatial orientation data can comprise a rotation angle based on the analysis (for example to bring the reference direction into alignment with the ‘forward’ direction or a minimum rotation angle to bring one of the sources to the side of the user.
[0324] Furthermore, as shown in Fig.15 by 1511 the method is configured to pre-rotate the spatial audio scene according to the determined spatial orientation data. In some embodiments this involve modifying the metadata parameters.
[0325] Furthermore, is the operation of generating or determining a downmix of the spatial audio scene to a (binaural) stereo audio signal and provide this downmix audio signal to the encoder according to the session information as shown in Fig.15 by 1513. In some embodiments the downmixing can be based on the IVAS parametric binauralizer and parametric stereo renderer, TS 26.253, clause 7.2.2.3, depending on whether binaural stereo or stereo is generated.
[0326] Rendering control relating to application of the “Spatial orientation data” can be carried out, e.g., according to IVAS external orientation input handling, TS 26.253, clause 7.4.5, with head-tracking set to default values and only spatial orientation data being considered.
[0327] Fi gs.21 a and 21 b show example advantages of the invention implementation. Fig.21 a for example shows the earlier example where the user 2101 is listening to the source C 2111 located at 10 degrees to the left of the user’s front and source D 2113 located at 170 degrees to the left of the user’s front results in a poor stereo separation (or spatial balance) on the left-right axis. Whereas Fig.21b shows the pre-rotated audio scene where the rotated source C 2121 and rotated source D 2123 now has significant stereo separation or in other words has better spatial balance.
[0328] In some embodiments rather than modifying the audio signals (for example by the intelligent spatial-to-stereo downmixer 817) the RTP payload format and specifically SCENEJDRIENTATION forward direction PI type is employed to transmit spatial orientation data that is generated as part of the spatial audio capture based on suitable renderer configuration data to modify the orientation of the audio scene at the renderer. This can be implemented by the apparatus such as shown in Fig.14. Although the following embodiments employ a RTP payload format to transmit the spatial orientation data any suitable packet or bitstream method generating method can be employed.
[0329] In some embodiments the renderer or receiver is thus configured to obtain the spatial orientation data. In such embodiments the synthesis processor is configured to generate modified audio signals thatprovide a spatially improved or more balanced reproduction perception for the listener without head-tracking capability based on the obtained spatial orientation data.
[0330] An example flow diagram showing the operations of the embodiments where the determination of the spatial orientation data is implemented at the capture device and the audio modification is implemented at the renderer is shown with respect to Fig.16. Thus firstly, as shown by Fig.16 by 1601 is the operation of obtaining input audio signals.
[0331] Furthermore, there is obtained renderer or receiver configuration information as shown by Fig.16 by 1603. For example the information signals whether the output format is a binaural stereo audio signals and / or whether the renderer or receiver is employing head-tracking or not.
[0332] In some embodiments the method is configured to check or determine whether the output spatial audio is generated based on headtracking or not. This check is shown in Fig.16 by 1605. The following operations are then performed when the output format is a non-headtracking mode. In a manner similar to as shown in Fig.15 the determination or check can be any suitable check such as whether the output spatial audio signal is a binaural stereo audio signal or not.
[0333] In some embodiments as shown by 1607 in Fig.16, there is performed an analysis of the spatial audio scene to determine whether spatial balance or left-right separation is suitable. This analysis can be any suitable method, such as those described in the embodiments above with overlapping sub-angles or their alternatives.
[0334] In some embodiments the spatial orientation data for modification of the scene orientation is then determined as shown in Fig.16 by 1609. As described above this spatial orientation data can comprise a rotation angle based on the analysis (for example to bring the reference direction into alignment with the ‘forward’ direction or a minimum rotation angle to bring one of the sources to the side of the user.
[0335] Furthermore, as shown in Fig.16 by 1611 the method is configured to transmit this spatial orientation data associated with the transmitted spatial audio signals (encoded by the IVAS format).
[0336] Then in some embodiments as shown in Fig.16 by 1613 is the operation of rendering a spatial audio signal to a stereo or binaural stereo audio signal using the received spatial orientation data (for example including the operation of rotating the spatial audio scene according to the determined spatial orientation data).
[0337] In some embodiments, rather than implementing the analysis of the capture or input audio signals and / or implementing the pre-rotation of the spatial scene prior to encoding, the analysis leading to the determination of the spatial orientation data and the modification of the audio signal to improve stereo separation (forexamplewhen head tracking is not employed) is implemented only at the receiver or renderer. For example such an example comprises a synthesis pre-processor, where the above method (of scene rotation) is implemented.
[0338] In some embodiments the data stream is analysed at the receiver / renderer to determine spatial orientation data. In such embodiments the spatial orientation data can then be transmitted back to the capture device. The spatial orientation data can be signalled in some embodiments to the capture device as a defined reverse direction PI type in a IVAS RTP payload. The spatial orientation data when received by the capture device is then used to apply a modification to the captured audio signal, for example an audio scene rotation. In such embodiments the renderer or receiver does not necessarily apply and further scene orientation processing.
[0339] In such embodiments, with respect to the capture apparatus or capture side the SCENEJDRIENTATION PI data is not generated (or any other signaling of the spatial orientation data), as the processing is based on an analysis of the data stream, e.g., IVAS bitstream following the decoder such as shown in Figs.17 and 18.
[0340] In some embodiments the analysis methods described earlier (with respect to the identification of dominant sources in a series of sub-angles, the determination of range of surviving sources and a direction based on the sources) may not be necessarily be employed. This is because the received spatial audio datastream, for some parameterized formats such as encoded MASA, comprise metadata which comprises directional information. As such the analysis in some embodiments can be performed from decoding the spatial metadata and determining source range and reference angles from these decoded direction parameters.
[0341] Other examples of parametric approaches with various levels of parameterization include transmission of FOA / HOA content using DirAC and / or SPAR (Spatial Audio Reconstruction, see TS 26.253) technologies and their various derivatives.
[0342] Fig.17, for example shows the introduction of a synthesis pre-processor 1701 following the decoder 607. The synthesis pre-processor 1701 is configured to receive the (decoded) transport audio signals 608 and (decoded) spatial metadata 610 from the decoder.
[0343] In some embodiments the synthesis pre-processor 1701, comprises an audio scene analyser 1703 which is configured to receive the (decoded) transport audio signals 608 and (decoded) spatial metadata 610 and from these generate spatial orientation data 1706 and pass the received transport audio signals 1718 and spatial metadata to an audio scene rotator 1705 with the spatial orientation data 1706.
[0344] The synthesis pre-processor 1701 can further comprise an audio scene rotator 1705, which is configured to generate modified or processed transport audio signals 1718 from the transport audio signals 1718 / 608 based on the spatial orientation data 1706. Furthermore the audio scene rotator 1705 is configured to generate modified or processed spatial metadata 1720 from the spatial metadata 1710 / 610 based on the spatial orientation data 1706. These processed transport audio signals 1718 and processed spatial metadata 1720 can be passed to the synthesis processor 609 from which is generated the output audio signals 612.
[0345] In some embodiments the synthesis pre-processor is implemented as part of an external renderer implementation that also contains the synthesis processor that can be, e.g., the IVAS external renderer or any other suitable renderer.
[0346] This framework particularly relates to certain spatial audio coding methods, e.g., MASA as well as relevant format combinations, e.g., MASA + object(s).
[0347] In some embodiments the audio scene analyser 1703 is configured to determine or obtain the pertinent directional data, for example, in case of MASA from the spatial metadata at least one transmitted direction and then determines the audio scene target orientation for the spatial orientation data.
[0348] For example, the target orientation can be understood as finding a rotation to the audio scene such that effect shown in Figs. Hc or 11 d are achieved.
[0349] In a manner similar to that described above, in some embodiment the scene rotation is implemented with a smoothing function to precent a jittery and unstable perception of the scene rendering. For example the spatial orientation data target orientation can be achieved via a series of intermediate frames (time instant) towards the target rotation such that over time (e.g., over several frames) the target orientation is reached.
[0350] With respect to Fig.18 a further implementation embodiment of the synthesis pre-processor 1801 is shown. In these embodiments the synthesis pre-processor 1801 is configured to operate on decoded audio signals, e.g., FOA component signals 1800 from the decoder 607. In such embodiments the synthesis preprocessor 1801 further comprises an audio scene analyser which is configured to implement embodiments such as described above, except the analysis and the audio signal modification can be performed on the decoded audio 1800 rather than the input audio signals as described herein. The analysis can employ any suitable method, e.g., DirAC analysis where the analysis is based on the decoded audio format.
[0351] Furthermore, in such embodiments the synthesis pre-processor 1801 can be configured to perform the audio scene rotation and forward processed audio signals 1810 to the synthesis processor 1809 for rendering. In some embodiments the synthesis pre-processor 1801 is configured to pass the audio (in an unmodified form) and further output the spatial orientation data to the synthesis processor 1809, where the synthesis processor 1809 is configured to perform the scene rotation operations.
[0352] Fig.19 shows the example renderer as shown in Fig.17 with the additional explicit signaling of the head-tracking mode 1900 to the synthesis pre-processor 1701. As described earlier, in some embodiments this signaling of the head-tracking mode 1900 represents one possible renderer configuration information aspect, with other information elements being the output format (stereo or binaural stereo).
[0353] An example flow diagram showing the operations of the embodiments where the determination of the spatial orientation data is implemented at the capture device and the audio modification is implementedat the renderer is shown with respect to Fig.20. Thus firstly, as shown by Fig.20 by 2001 is the operation of decoding the spatial audio signals.
[0354] Furthermore, there is obtained renderer or receiver configuration information as shown by Fig.20 by 2003. For example, the information signals whether the output format is a binaural stereo audio signals and / or whether the renderer or receiver is employing head-tracking or not.
[0355] In some embodiments the method is configured to check or determine whether the output spatial audio is generated based on headtracking or not. This check is shown in Fig.20 by 2005. The following operations are then performed when the output format is a non-headtracking mode. In a manner similar to as shown in Figs.15 or 16 the determination or check can (also) be any suitable check such as whether the output spatial audio signal is a binaural stereo audio signal or not.
[0356] In some embodiments as shown by 2007 in Fig.20, there is performed an analysis of the spatial audio scene to determine whether spatial balance or left-right separation is suitable. This analysis can be any suitable method, such as those described in the embodiments above with overlapping sub-angles or their alternatives or analysis of the metadata.
[0357] In some embodiments the spatial orientation data for modification of the scene orientation is then determined as shown in Fig.20 by 2009. As described above this spatial orientation data can comprise a rotation angle based on the analysis (for example to bring the reference direction into alignment with the ‘forward’ direction or a minimum rotation angle to bring one of the sources to the side of the user.
[0358] Furthermore, as shown in Fig.20 by 2011 the method is configured to modift the spatial audio signals based on this spatial orientation data.
[0359] Then in some embodiments as shown in Fig.20 by 2013 is the operation of rendering a spatial audio signal to a stereo or binaural stereo audio signal using the modified spatial orientation data.
[0360] Figs.22a, 22b, 23a, 23b, 24a and 24b show example scenarios where the application of spatial orientation data to modify the spatial audio signals.
[0361] Thus for example Fig.22a shows the example where a user 2201 has a first head position or orientation (which is in line with the arrow 2203 representing an audio scene orientation), an audio or sound source A 2211 approximately at -80 degrees (to the right) of the first head orientation, the audio or sound source C 2213 is approximately 10 degrees to the left of the first head orientation (in other words almost ahead of the user) and the audio or sound source D 2215 is approximately 170 degrees to the left of the first head orientation 2203 (in other words almost behind the user). Furthermore is shown the reference angle 2251 between the three sources and is approximately at -90 degrees to the right of the first head orientation.
[0362] Fig.22b shows a rotation of the audio scene such that the reference angle 2251 and the first head orientation are aligned (an approximate rotation of +90 degrees). In this arrangement the three sources are arranged such the audio sources are approximately balanced on the left-right axis.
[0363] In other words Figs.22a and 22b present an original audio scene and rotation to the mid-point of the left-right span of the sources (as discussed above).
[0364] Fi gs.23a and 23b show the difference between the alignment of the reference angle 2251 and the first head orientation 2303 example as shown in Fig.22b and the minimum rotation approach as shown in Fig.23b where the rotation is approximately 80 degrees. In other words Fig.23a and 23b show a comparison of the rotation to the mid-point of the span (Fig.23a) against the rotation to the corresponding energy midpoint (Fig.23b). This alternative can have certain advantages for real scenes
[0365] Figs.24a and 24b show a further example of the effect of the rotation of the head in a nonheadtracking situation. Fig.24a for example shows the effect of the modification based on the minimum rotation example as shown in Fig.23b and Fig.24b which shows the later effect of the head rotation to a second head orientation 2403. In other words Fig.24a and 24b illustrate head-rotation with the latter rotation being applied at the renderer.
[0366] In some embodiments, the SCENEJDRIENTATION forward direction PI data is transmitted according to examples above.
[0367] In some alternative implementations, the spatial orientation data is always analysed and transmitted (as PI data), and no renderer configuration data is considered. This approach has the disadvantage that it consumes more overall channel capacity.
[0368] Fig.25 furthermore shows an example audio waveform comparison. In Fig.25 a (mono) source that is placed spatially in front of a user, for example in similar way as source C. When that simple spatial scene is rendered into stereo, the resulting stereo signal collapses inside the listener’s head (two identical signals for left and right channels). When the scene is pre-rotated as part of the downmix or rendering process according to the invention, there is a difference between the resulting left and right channels. If we next consider a larger number of sources, e.g., a second source (such as shown in the above examples), then for most scenes there can be found a suitable rotation that allows for the good spatial balance or left-right separation that makes it easier for the listener to identify the sources, listen to a specific talker, etc.
[0369] In some embodiments at least some of the functions of the analysis processor and the scene orientation determiner can be implemented using ML methods. For example, a suitable neural network can be trained based on spatial metadata or a combination of spatial metadata and corresponding transport audio signals, e.g., MASA input format to predict spatial orientation data candidates. The training reference can be, e.g., corresponding expert-determined ground-truth spatial orientation data.
[0370] With respect to Fig.26 an example electronic device is shown. The device may be any suitable electronics device or apparatus. For example, in some embodiments the device 3000 is a mobile device, user equipment, tablet computer, computer, consumer electronic device, still / video camera device, mobilecommunication device, audio / video / image playback or recording apparatus, vehicle, etc. or any combination thereof.
[0371] In some embodiments the device 3000 comprises at least one processor or central processing unit 3007. The processor 3007 can be configured to execute various program codes such as the methods such as described herein.
[0372] In some embodiments the device 3000 comprises at least one memory 3011. In some embodiments the at least one processor 3007, e.g. CPU (Central processing Unit) and / or GPU (Graphical Processing Unit), is coupled to the at least one memory 3011. The memory 3011 can be any suitable storage means. In some embodiments the memory 3011 comprises a program code section for storing one or more program codes, or program instructions, implementable upon the one or more processor 3007. Furthermore, in some embodiments the memory 3011 can further comprise a stored data section for storing data, for example data that has been processed or to be processed in accordance with the embodiments as described herein. The implemented program code stored within the program code section and the data stored within the stored data section can be retrieved by the processor 3007 whenever needed via the memory-processor coupling.
[0373] In some embodiments the device 3000 comprises a user interface 3005. The user interface 3005 can be coupled in some embodiments to the processor 3007. In some embodiments the processor 3007 can control the operation of the user interface 3005 and receive inputs from the user interface 3005. In some embodiments the user interface 3005 can enable a user to input commands to the device 3000, for example via a keypad. In some embodiments the user interface 3005 can enable the user to obtain information from the device 3000. For example, the user interface 3005 may comprise a display configured to display information from the device 3000 to the user. The user interface 3005 can in some embodiments comprise a touch screen or touch interface capable of both enabling information to be entered to the device 3000 and further displaying information to the user of the device 3000. The user interface 3005 in some embodiments can comprise one or more loudspeakers, and one or more microphones.
[0374] In some embodiments the device 3000 comprises an input / output port 3009. The input / output port 3009 in some embodiments comprises a transceiver. The transceiver in such embodiments can be coupled to the processor 3007 and configured to enable a communication with other apparatus or electronic devices, for example via wireless and / wireless communications networks. The transceiver or any suitable transceiver or transmitter and / or receiver means can in some embodiments be configured to communicate with other electronic devices or apparatus via a wire or wired coupling.
[0375] The transceiver can communicate with further apparatus by any suitable known communications protocol. For example, in some embodiments the transceiver can use a suitable mobile telecommunications protocol, such as 3G / 4G / 5G / 6G or any further generation protocol, a short-range wireless communicationprotocol, such as a wireless local area network (WLAN) protocol, for example IEEE 802. X, a Bluetooth, or infrared data communication pathway (IRDA), or any combination thereof.
[0376] The transceiver input / output port 3009 may be configured to receive the signals and in some embodiments obtain the focus parameters as described herein.
[0377] In some embodiments the device 3000 may be employed to generate a suitable audio signal using the processor 3007 executing suitable code. The input / output port 3009 may be coupled to any suitable audio output for example to a multichannel speaker system and / or headphones (which may be a head-tracked or a non-tracked headphones) or similar.
[0378] In general, the various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic, circuitry or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto. While various aspects of the invention may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
[0379] The embodiments of this invention may be implemented by computer software executable by a data processor UE, such as a mobile user device or a consumer electronic device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD.
[0380] The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processors may be of any type suitable to the local technical environment, and may include one or more of general-purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), gate level circuits and processors based on multi-core processor architecture, as non-limiting examples.
[0381] Embodiments of the inventions may be practiced in various components such as one or more integrated circuit modules or circuitry. The design of integrated circuits is by and large a highly automatedprocess. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.
[0382] As used in this application, the term “circuitry” may refer to one or more or all of the following:(a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry) and(b) combinations of hardware circuits and software, such as (as applicable):(i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.”
[0383] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.
[0384] Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre-stored design modules. Once the design for a semiconductor circuit has been completed, the resultant design, in a standardized electronic format (e.g., Opus, GDSII, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication.
[0385] The foregoing description has provided by way of exemplary and non-limiting examples a full and informative description of the exemplary embodiment of this invention. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still fall within the scope of this invention as defined in the appended claims.3GPP 3rdGeneration Partnership ProjectCMR Codec mode requestEVS Enhanced Voice ServicesISM Independent Streams with Metadata(i.e., type of Object-Based Audio)IVAS Immersive Voice and Audio ServicesMASA Metadata-Assisted Spatial AudioMC MultichannelOMASAObject-based audio with MASA (combined input format) OSBA Object-based audio with SBA (combined input format) PI Processing information (audio)RTCP Real-Time Transport Control Protocol
Claims
CLAIMS1. An apparatus for immersive audio capture, the apparatus comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the system at least to:obtain scene orientation data for a spatial audio scene, the scene orientation data corresponding to an improved spatial balance for rendering a stereo audio signal, wherein the scene orientation data is based on analysis of spatial audio capture source positions;determine rotation information based on the scene orientation data, the rotation information for providing the improved spatial balance for rendering the stereo audio signal; andenable the rotation information for a rotation of the spatial audio scene.
2. The apparatus as claimed in claim 1, caused to enable the rotation information for a rotation of the spatial audio scene is caused to transmit the rotation information to a further apparatus for application of the rotation of the spatial audio scene based on the rotation information.
3. The apparatus as claimed in claim 1, caused to enable the rotation information for a rotation of the spatial audio scene is caused to rotate the spatial audio scene based on the rotation information.
4. The apparatus as claimed in any of claims 1 or 3, further caused to:obtain input audio signals from at least two microphones;process the input audio signals to rotate the spatial audio scene based on the rotation information to determine at least two audio signals; andencode the at least two audio signals to generate a transport audio signal component of an immersive audio bitstream.
5. The apparatus as claimed in claim 4, caused to process the input audio signals to rotate the spatial audio scene based on the rotation information to determine at least two audio signals is caused to:determine at least one first format audio signal based on the input audio signals; andprocess the at least one first format audio signal based on the rotation information to determine the at least two audio signals.
6. The apparatus as claimed in claim 5, wherein the at least one first format audio signal is at least one of:a MASA format audio signal;an Ambisonics audio signal;a multichannel audio signal; anda stereo audio signal.
7. The apparatus as claimed in any of claims 5 or 6, further caused to:determine at least one first format spatial metadata parameter based on the input audio signals; andprocess the at least one first format spatial metadata parameter based on the rotation information to determine at least one spatial metadata parameter; andencode the at least one spatial metadata parameter to generate a spatial metadata component of an immersive audio bitstream.
8. The apparatus as claimed in claim 7, wherein the at least one first format spatial metadata parameter comprises at least one of:a direction and orientation parameter; andan energy ratio parameter.
9. The apparatus as claimed in any of claims 5 to 8, caused to process the at least one first format audio signal based on the rotation information to determine the at least two audio signals is caused to process and / or mix individual channel parts of the at least one first format audio signal to determine the at least two audio signals.
10. The apparatus as claimed in any of claims 4 to 9, wherein the at least two audio signals comprises two transport audio signals, and are:stereo downmix audio signals; orbinaural stereo downmix audio signals.
11. The apparatus as claimed in claim 1, further caused to transmit and / or store the rotation information for the rotation of the spatial audio scene based on the scene orientation data to a further apparatus, the further apparatus for providing the improved spatial balance for the rendering of the stereo audio signal based on the rotation of the spatial audio scene based on the rotation information.
12. The apparatus as claimed in claim 11, further caused to:obtain input audio signals from at least two microphones;determine at least one transport audio signal based on the input audio signals;analyse the input audio signals to obtain spatial metadata;encode the at least one transport audio signal and spatial metadata as an immersive audio bitstream; and generate an auxiliary bitstream comprising the rotation information for a rotation of the spatial audio scene.
13. The apparatus as claimed in any of claims 1, 11 or 12, further caused to:obtain at least one bitstream, the bitstream comprising at least one encoded audio signal and encoded spatial metadata, wherein the apparatus caused to obtain scene orientation data for a spatial audio scene is caused to obtain the scene orientation data from the at least one bitstream.
14. The apparatus as claimed in claim 13, is further caused to decode the encoded spatial metadata, wherein the apparatus caused to obtain the scene orientation data from the at least one bitstream is caused to obtain the scene orientation data based on the decoded spatial metadata.
15. The apparatus as claimed in claim 13, is further caused to decode the encoded audio signal, wherein the apparatus caused to obtain the scene orientation data from the at least one bitstream is caused to analyse the decoded audio signal.
16. The apparatus as claimed in any of claims 1 to 15, caused to analyse spatial audio capture source positions to determine the scene orientation data is further caused to:obtain the spatial audio capture source positions;divide the audio scene into a number of audio scene portions, each portion spanning a range of directions of the audio scene;identify for at least one audio scene portion a dominant spatial audio capture source to generate survivor spatial audio capture source directions; andidentify one of:a smallest angle between the spatial audio capture source directions farthest from each other; ora smallest angle between the spatial audio capture source directions that includes a listener front direction.
17. The apparatus as claimed in any of claims 1 to 16, further caused to obtain the scene orientation data in at least one of:a real-time transport protocol payload;a real-time transport protocol header extension; ora real-time transport control protocol payload.
18. The apparatus as claimed in any of claims 1 to 17, further caused to obtain the scene orientation data for at least one of:during a session setup; orafter a session setup.
19. The apparatus as claimed in any of claims 1 to 18, wherein the stereo audio signal comprises a binaural stereo audio signal.
20. A method for an apparatus for immersive audio capture, the method comprising:obtaining scene orientation data for a spatial audio scene, the scene orientation data corresponding to an improved spatial balance for rendering a stereo audio signal, wherein the scene orientation data is based on analysis of spatial audio capture source positions;determining rotation information based on the scene orientation data, the rotation information for providing the improved spatial balance for rendering the stereo audio signal; andenabling the rotation information for a rotation of the spatial audio scene.
21. An apparatus for immersive audio capture, the apparatus comprising means configured to:obtain scene orientation data for a spatial audio scene, the scene orientation data corresponding to an improved spatial balance for rendering a stereo audio signal, wherein the scene orientation data is based on analysis of spatial audio capture source positions;determine rotation information based on the scene orientation data, the rotation information for providing the improved spatial balance for rendering the stereo audio signal; andenable the rotation information for a rotation of the spatial audio scene.