A method and apparatus for describing anchors for immersive audio
A universal description convention for anchor references in AR audio systems addresses placement inconsistencies by aligning content creator and consumption context cues, enhancing rendering accuracy and quality across AR platforms.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NOKIA TECHNOLOGIES OY
- Filing Date
- 2025-10-14
- Publication Date
- 2026-05-07
AI Technical Summary
Existing augmented reality (AR) audio rendering systems face challenges in accurately placing audio elements due to inconsistencies between content creator specified anchors and listening space descriptions, leading to incorrect rendering and poor subjective quality.
Implement a universal description convention for anchor references that includes representation methods and taxonomy identification, ensuring consistent interpretation across different AR platforms by aligning content creator and consumption context cues.
Ensures accurate placement of audio elements by aligning content creator intents with consumption context, improving rendering consistency and subjective quality across various AR devices.
Smart Images

Figure EP2025079513_07052026_PF_FP_ABST
Abstract
Description
[0001] A METHOD AND APPARATUS FOR DESCRIBING ANCHORS FOR IMMERSIVE AUDIO
[0002] Field
[0003] The present application relates to method and apparatus for describing anchors for immersive audio, but not exclusively for method and apparatus for describing anchors for immersive audio within augmented reality 6 degrees-of-freedom rendering applications.
[0004] Background
[0005] Augmented Reality (AR) applications (and other similar virtual scene creation applications such as Mixed Reality (MR or XR) and Virtual Reality (VR)) where a virtual scene is represented to a user wearing a head mounted device (HMD) have become more complex and sophisticated over time. The application may comprise data which comprises a visual component (or overlay) and an audio component (or overlay) which is presented to the user. These components may be provided to the user dependent on the position and orientation of the user (for a 6 degree-of-freedom application) within an Augmented Reality (AR) scene.
[0006] Scene information for rendering an AR scene typically comprises two parts. One part is the virtual scene information which may be described during content creation (or by a suitable capture apparatus or device) and represents the scene as captured (or initially generated). The virtual scene may be provided in an encoder input format (EIF) data format. The EIF and (captured or generated) audio data is used by an encoder to generate the scene description and spatial audio metadata (and audio signals), which can be delivered via the bitstream to the rendering (playback) device or apparatus. The scene description for an AR or VR scene is thus specified by the content creator at least partially during a content creation phase. In the case of VR, the scene is specified in its entirety and it is rendered exactly as specified in the content creator bitstream.
[0007] The second part of the AR audio scene rendering is related to the physical listening space (or physical space) of the listener (or end user). The scene or listener space information may be obtained during the AR rendering (when the listener is consuming the content). Thus there is a fundamental aspect of AR which is different from VR, which means the acoustic properties of the audio scene and potentially information about audio sources (such as their position) are known (for AR) only during content consumption and cannot be known or optimized during content creation. Such audio sources known only at content consumption time can be referred as locally captured audio sources.
[0008] Fig.1 a shows an example AR scene where a virtual scene is located within a physical listening space. In this example there is a user 107 who is located within a physical listening space 101. Furthermore in this example the user 109 is experiencing a six-degree-of-freedom (6DOF) virtual scene 113 with virtual scene elements. In this example the virtual scene 113 elements are represented by two audio elements or objects, a first object 103 (guitar player) and second object 105 (drummer), a virtual occlusion element (e.g., represented as a virtual partition 117) and a virtual room 115 (e.g., with walls which have a size, a position, acoustic materials which are defined within the virtual scene description). A Tenderer (which in this example is a hand held electronic device or apparatus 111 ) is configured to perform the rendering so that the auralization is plausible for the user’s physical listening space (e.g., position of the walls and the acoustic material properties of the wall). The rendering is presented to the user 107 in this example by a suitable headphone or headset 109.
[0009] The position of the audio elements or objects and scene geometry elements is known during rendering time with the help of “anchors” which are embedded within the listening space description of the listening space, which is obtained during content consumption. The expectation is that the anchors referred to in the content creator bitstream find a corresponding match in the listening space description. This description has been specified in the MPEG Audio group as LSDF (Listener Space Description Format) as the means for providing listener space information to the Tenderer. The LSDF format can be flexible and different from what is described in the MPEG audio group ISO / IEC JTC1 SC29 WG06 but the information about the anchor and acoustic parameters is provided to the Tenderer by any suitable interface. The LSDF as, e.g., a file or listening space information (including anchor position or acoustic parameters of the listening space) is available only during rendering. Consequently, forming of the complete audio scene and the derivation of at least some acoustic modelling parameters is done in the Tenderer in case of AR scenarios.
[0010] Thus for AR scenes, the content creator bitstream carries information about which audio elements and scene geometry elements correspond to which anchors in the listening space. Consequently, the positions of the audio element positions, reflecting elements, occluding elements, etc. are known only during rendering. Furthermore, the acoustic modeling parameters are known only during rendering. Summary
[0011] There is provided according to a first aspect an apparatus comprising means configured to: obtain information associated with an audio scene, wherein the information is made available to a rendering device using a bitstream; obtain at least one audio signal, the at least one audio signal is associated with at least one audio object within the audio scene; obtain at least one anchor reference, wherein the at least one anchor reference comprises an anchor reference description information, such that an audio scene rendering based on the at least one audio signal reproduces the audio scene based on the anchor reference description information; obtain at least one anchor parameter for identifying the at least one anchor position and / or orientation; and transmit the bitstream comprising: the at least one audio signal; the at least one anchor reference; and the at least one anchor parameter.
[0012] The means configured to obtain at least one anchor reference associated with the at least one audio object may be further configured to define a universal description convention which can be interpreted consistently and correctly by the apparatus and the rendering device.
[0013] The means configured to obtain at least one anchor parameter may be configured to determine anchor reference description information based on the universal description convention.
[0014] The means configured to transmit the bitstream may be configured to transmit the bitstream to a further apparatus comprising at least the rendering device, wherein the rendering device may be configured to render an output audio signal based on the at least one audio signal and located and / or orientated within the audio scene based on the at least one anchor parameter reference correctly interpreted based on the at least one anchor reference. According to a second aspect there is provided an apparatus comprising means configured to: obtain a bitstream, wherein the bitstream comprises: information associated with an audio scene; an encoded at least one anchor reference; an encoded at least one audio signal; and an encoded at least one anchor parameter; obtain the encoded at least one anchor reference, the encoded at least one audio signal; and the encoded at least one anchor parameter from the obtained bitstream; decode the encoded at least one audio signal to obtain at least one audio signal, the at least one audio signal is associated with at least one audio object within the audio scene; decode the encoded at least one anchor reference to obtain at least one anchor reference associated with the at least one audio object, wherein the at least one anchor reference comprises an anchor reference description information, such that an audio scene rendering based on the at least one audio signal reproduces the audio scene based on the anchor reference description information; decode the encoded at least one anchor parameter to obtain at least one anchor parameter for identifying at least one anchor position and / or orientation relative to at least one origin position and / or orientation; and render an output signal based on the at least one audio signal, located and / or orientated within the audio scene based on the at least one anchor parameter being reproduced based on the at least one anchor reference.
[0015] The means may be configured to identify the at least one anchor position and / or orientation within the audio scene based on the anchor reference description information and the at least one anchor parameter.
[0016] The means may be configured to select the at least one anchor based on the at least one anchor reference description information and at least one candidate anchor obtained from listening space information.
[0017] The anchor reference description information may be configured to one of: describe an anchor reference definition; and identify a location where the anchor reference definition is accessible.
[0018] The audio scene may comprise one of: an augmented reality audio scene; a virtual reality audio scene; and a mixed reality audio scene.
[0019] The anchor reference may further be employed to identify at least one of: at least one acoustic environment; at least one acoustic portal; and at least one scene geometry element in a physical space.
[0020] The at least one anchor reference and the at least one anchor parameter may be components of a parameter.
[0021] According to a third aspect there is provided a method for an apparatus, the method comprising: obtaining information associated with an audio scene, wherein the information is made available to a rendering device using a bitstream; obtaining at least one audio signal, the at least one audio signal is associated with at least one audio object within the audio scene; obtaining at least one anchor reference associated with the at least one audio object, wherein the at least one anchor reference comprises an anchor reference description information, such that an audio scene rendering based on the at least one audio signal reproduces the audio scene based on the anchor reference description information; obtaining at least one anchor parameter for identifying at least one anchor position and / or orientation; and transmit the bitstream comprising: the at least one audio signal; the at least one anchor reference; and the at least one anchor parameter.
[0022] Obtaining the at least one anchor reference associated with the at least one audio object may further comprise defining a universal description convention which can be interpreted consistently and correctly by the apparatus and the rendering device.
[0023] Obtaining the at least one anchor reference further may comprise determining anchor reference description information based on the universal description convention.
[0024] Transmitting the bitstream may comprise transmitting the bitstream to a further apparatus comprising at least the rendering device, wherein the rendering device is configured to render an output audio signal based on the at least one audio signal and located and / or orientated within the audio scene based on the at least one anchor parameter reference correctly interpreted based on the at least one anchor reference.
[0025] According to a fourth aspect there is provided a method for an apparatus, the method comprising: obtaining a bitstream, wherein the bitstream comprises: information associated with an audio scene; an encoded at least one anchor reference; an encoded at least one audio signal; and an encoded at least one anchor parameter; and obtain the encoded at least one anchor reference, the encoded at least one audio signal; and the encoded at least one anchor parameter from the obtained bitstream; decoding the encoded at least one audio signal to obtain at least one audio signal, the at least one audio signal is associated with at least one audio object within the audio scene; decoding the encoded at least one anchor reference to obtain at least one anchor reference associated with the at least one audio object, wherein the at least one anchor reference comprises an anchor reference description information, such that an audio scene rendering based on the at least one audio signal reproduces the audio scene based on the anchor reference description information; decoding the encoded at least one anchor parameter to obtain at least one anchor parameter for identifying at least one anchor position and / or orientation relative to at least one origin position and / or orientation; and rendering an output signal based on the at least one audio signal, located and / or orientated within the audio scene based on the at least one anchor parameter reference being reproduced based on the at least one anchor reference.
[0026] The method may further comprise identifying the at least one anchor position and / or orientation within the audio scene based on the anchor reference description information and the at least one anchor parameter. The method may further comprise selecting the at least one anchor based on the at least one anchor reference description information and at least one candidate anchor obtained from listening space information.
[0027] The anchor reference description information may be configured to one of: describe an anchor reference definition; and identify a location where the anchor reference definition is accessible.
[0028] The audio scene may comprise one of: an augmented reality audio scene; a virtual reality audio scene; and a mixed reality audio scene.
[0029] The anchor reference may further be employed to identify at least one of: at least one acoustic environment; at least one acoustic portal; and at least one scene geometry element in a physical space.
[0030] The at least one anchor reference and the at least one anchor parameter ay be components of a parameter.
[0031] According to a fifth aspect there is provided an apparatus comprising at least one processor and at least one memory including a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to: obtain information associated with an audio scene, wherein the information is made available to a rendering device using a bitstream; obtain at least one audio signal, the at least one audio signal is associated with at least one audio object within the audio scene; obtain at least one anchor reference, wherein the at least one anchor reference comprises an anchor reference description information, such that an audio scene rendering based on the at least one audio signal reproduces the audio scene based on the anchor reference description information; obtain at least one anchor parameter for identifying the at least one anchor position and / or orientation; and transmit the bitstream comprising: the at least one audio signal; the at least one anchor reference; and the at least one anchor parameter.
[0032] The apparatus caused to obtain at least one anchor reference associated with the at least one audio object may be further caused to define a universal description convention which can be interpreted consistently and correctly by the apparatus and the rendering device.
[0033] The apparatus caused to obtain at least one anchor parameter may be caused to determine anchor reference description information based on the universal description convention. The apparatus caused to transmit the bitstream may be caused to transmit the bitstream to a further apparatus comprising at least the rendering device, wherein the rendering device may be configured to render an output audio signal based on the at least one audio signal and located and / or orientated within the audio scene based on the at least one anchor parameter reference correctly interpreted based on the at least one anchor reference.
[0034] According to a sixth aspect there is provided an apparatus comprising at least one processor and at least one memory including a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to: obtain a bitstream, wherein the bitstream comprises: information associated with an audio scene; an encoded at least one anchor reference; an encoded at least one audio signal; and an encoded at least one anchor parameter; obtain the encoded at least one anchor reference, the encoded at least one audio signal; and the encoded at least one anchor parameter from the obtained bitstream; decode the encoded at least one audio signal to obtain at least one audio signal, the at least one audio signal is associated with at least one audio object within the audio scene; decode the encoded at least one anchor reference to obtain at least one anchor reference associated with the at least one audio object, wherein the at least one anchor reference comprises an anchor reference description information, such that an audio scene rendering based on the at least one audio signal reproduces the audio scene based on the anchor reference description information; decode the encoded at least one anchor parameter to obtain at least one anchor parameter for identifying at least one anchor position and / or orientation relative to at least one origin position and / or orientation; and render an output signal based on the at least one audio signal, located and / or orientated within the audio scene based on the at least one anchor parameter being reproduced based on the at least one anchor reference.
[0035] The apparatus may be caused to identify the at least one anchor position and / or orientation within the audio scene based on the anchor reference description information and the at least one anchor parameter.
[0036] The apparatus may be caused to select the at least one anchor based on the at least one anchor reference description information and at least one candidate anchor obtained from listening space information. The anchor reference description information may be configured to one of: describe an anchor reference definition; and identify a location where the anchor reference definition is accessible.
[0037] The audio scene may comprise one of: an augmented reality audio scene; a virtual reality audio scene; and a mixed reality audio scene.
[0038] The anchor reference may further be employed to identify at least one of: at least one acoustic environment; at least one acoustic portal; and at least one scene geometry element in a physical space.
[0039] The at least one anchor reference and the at least one anchor parameter may be components of a parameter.
[0040] According to a seventh aspect there is provided an apparatus comprising: means for obtaining information associated with an audio scene, wherein the information is made available to a rendering device using a bitstream; means for obtaining at least one audio signal, the at least one audio signal is associated with at least one audio object within the audio scene; means for obtaining at least one anchor reference associated with the at least one audio object, wherein the at least one anchor reference comprises an anchor reference description information, such that an audio scene rendering based on the at least one audio signal reproduces the audio scene based on the anchor reference description information; means for obtaining at least one anchor parameter for identifying at least one anchor position and / or orientation; and means for transmitting the bitstream comprising: the at least one audio signal; the at least one anchor reference; and the at least one anchor parameter.
[0041] According to an eighth aspect there is provided an apparatus comprising: means for obtaining a bitstream, wherein the bitstream comprises: information associated with an audio scene; an encoded at least one anchor reference; an encoded at least one audio signal; and an encoded at least one anchor parameter; and means for obtaining the encoded at least one anchor reference, the encoded at least one audio signal; and the encoded at least one anchor parameter from the obtained bitstream; means for decoding the encoded at least one audio signal to obtain at least one audio signal, the at least one audio signal is associated with at least one audio object within the audio scene; means for decoding the encoded at least one anchor reference to obtain at least one anchor reference associated with the at least one audio object, wherein the at least one anchor reference comprises an anchor reference description information, such that an audio scene rendering based on the at least one audio signal reproduces the audio scene based on the anchor reference description information; means for decoding the encoded at least one anchor parameter to obtain at least one anchor parameter for identifying at least one anchor position and / or orientation relative to at least one origin position and / or orientation; and rendering an output signal based on the at least one audio signal, located and / or orientated within the audio scene based on the at least one anchor parameter reference being reproduced based on the at least one anchor reference.
[0042] According to a ninth aspect there is provided a computer program comprising instructions [or a computer readable medium comprising program instructions] for causing an apparatus to perform at least the following: obtaining information associated with an audio scene, wherein the information is made available to a rendering device using a bitstream; obtaining at least one audio signal, the at least one audio signal is associated with at least one audio object within the audio scene; obtaining at least one anchor reference associated with the at least one audio object, wherein the at least one anchor reference comprises an anchor reference description information, such that an audio scene rendering based on the at least one audio signal reproduces the audio scene based on the anchor reference description information; obtaining at least one anchor parameter for identifying at least one anchor position and / or orientation; and transmit the bitstream comprising: the at least one audio signal; the at least one anchor reference; and the at least one anchor parameter.
[0043] According to a tenth aspect there is provided a computer program comprising instructions [or a computer readable medium comprising program instructions] for causing an apparatus to perform at least the following: obtaining a bitstream, wherein the bitstream comprises: information associated with an audio scene; an encoded at least one anchor reference; an encoded at least one audio signal; and an encoded at least one anchor parameter; and obtaining the encoded at least one anchor reference, the encoded at least one audio signal; and the encoded at least one anchor parameter from the obtained bitstream; decoding the encoded at least one audio signal to obtain at least one audio signal, the at least one audio signal is associated with at least one audio object within the audio scene; decoding the encoded at least one anchor reference to obtain at least one anchor reference associated with the at least one audio object, wherein the at least one anchor reference comprises an anchor reference description information, such that an audio scene rendering based on the at least one audio signal reproduces the audio scene based on the anchor reference description information; decoding the encoded at least one anchor parameter to obtain at least one anchor parameter for identifying at least one anchor position and / or orientation relative to at least one origin position and / or orientation; and rendering an output signal based on the at least one audio signal, located and / or orientated within the audio scene based on the at least one anchor parameter reference being reproduced based on the at least one anchor reference.
[0044] According to an eleventh aspect there is provided a non-transitory computer readable medium comprising program instructions for causing an apparatus to perform at least the following: obtaining information associated with an audio scene, wherein the information is made available to a rendering device using a bitstream; obtaining at least one audio signal, the at least one audio signal is associated with at least one audio object within the audio scene; obtaining at least one anchor reference associated with the at least one audio object, wherein the at least one anchor reference comprises an anchor reference description information, such that an audio scene rendering based on the at least one audio signal reproduces the audio scene based on the anchor reference description information; obtaining at least one anchor parameter for identifying at least one anchor position and / or orientation; and transmit the bitstream comprising: the at least one audio signal; the at least one anchor reference; and the at least one anchor parameter.
[0045] According to a twelfth aspect there is provided a non-transitory computer readable medium comprising program instructions for causing an apparatus to perform at least the following: obtaining a bitstream, wherein the bitstream comprises: information associated with an audio scene; an encoded at least one anchor reference; an encoded at least one audio signal; and an encoded at least one anchor parameter; and obtaining the encoded at least one anchor reference, the encoded at least one audio signal; and the encoded at least one anchor parameter from the obtained bitstream; decoding the encoded at least one audio signal to obtain at least one audio signal, the at least one audio signal is associated with at least one audio object within the audio scene; decoding the encoded at least one anchor reference to obtain at least one anchor reference associated with the at least one audio object, wherein the at least one anchor reference comprises an anchor reference description information, such that an audio scene rendering based on the at least one audio signal reproduces the audio scene based on the anchor reference description information; decoding the encoded at least one anchor parameter to obtain at least one anchor parameter for identifying at least one anchor position and / or orientation relative to at least one origin position and / or orientation; and rendering an output signal based on the at least one audio signal, located and / or orientated within the audio scene based on the at least one anchor parameter reference being reproduced based on the at least one anchor reference.
[0046] According to a thirteenth aspect there is provided an apparatus comprising: obtaining circuity configured to obtain information associated with an audio scene, wherein the information is made available to a rendering device using a bitstream; obtaining circuity configured to obtain at least one audio signal, the at least one audio signal is associated with at least one audio object within the audio scene; obtaining circuity configured to obtain at least one anchor reference associated with the at least one audio object, wherein the at least one anchor reference comprises an anchor reference description information, such that an audio scene rendering based on the at least one audio signal reproduces the audio scene based on the anchor reference description information; obtaining circuity configured to obtain at least one anchor parameter for identifying at least one anchor position and / or orientation; and transmitting circuitry configured to transmit the bitstream comprising: the at least one audio signal; the at least one anchor reference; and the at least one anchor parameter.
[0047] According to a fourteenth aspect there is provided an apparatus comprising: obtaining circuitry configured to obtain a bitstream, wherein the bitstream comprises: information associated with an audio scene; an encoded at least one anchor reference; an encoded at least one audio signal; and an encoded at least one anchor parameter; and obtaining circuitry configured to obtain the encoded at least one anchor reference, the encoded at least one audio signal; and the encoded at least one anchor parameter from the obtained bitstream; decoding circuitry configured to decode the encoded at least one audio signal to obtain at least one audio signal, the at least one audio signal is associated with at least one audio object within the audio scene; decoding circuitry configured to decode the encoded at least one anchor reference to obtain at least one anchor reference associated with the at least one audio object, wherein the at least one anchor reference comprises an anchor reference description information, such that an audio scene rendering based on the at least one audio signal reproduces the audio scene based on the anchor reference description information; decoding circuitry configured to decode the encoded at least one anchor parameter to obtain at least one anchor parameter for identifying at least one anchor position and / or orientation relative to at least one origin position and / or orientation; and rendering circuitry configured to render an output signal based on the at least one audio signal, located and / or orientated within the audio scene based on the at least one anchor parameter reference being reproduced based on the at least one anchor reference.
[0048] According to an fifteenth aspect there is provided a computer readable medium comprising program instructions for causing an apparatus to perform at least the following: obtaining information associated with an audio scene, wherein the information is made available to a rendering device using a bitstream; obtaining at least one audio signal, the at least one audio signal is associated with at least one audio object within the audio scene; obtaining at least one anchor reference associated with the at least one audio object, wherein the at least one anchor reference comprises an anchor reference description information, such that an audio scene rendering based on the at least one audio signal reproduces the audio scene based on the anchor reference description information; obtaining at least one anchor parameter for identifying at least one anchor position and / or orientation; and transmit the bitstream comprising: the at least one audio signal; the at least one anchor reference; and the at least one anchor parameter.
[0049] According to an sixteenth aspect there is provided a computer readable medium comprising program instructions for causing an apparatus to perform at least the following: obtaining a bitstream, wherein the bitstream comprises: information associated with an audio scene; an encoded at least one anchor reference; an encoded at least one audio signal; and an encoded at least one anchor parameter; and obtaining the encoded at least one anchor reference, the encoded at least one audio signal; and the encoded at least one anchor parameter from the obtained bitstream; decoding the encoded at least one audio signal to obtain at least one audio signal, the at least one audio signal is associated with at least one audio object within the audio scene; decoding the encoded at least one anchor reference to obtain at least one anchor reference associated with the at least one audio object, wherein the at least one anchor reference comprises an anchor reference description information, such that an audio scene rendering based on the at least one audio signal reproduces the audio scene based on the anchor reference description information; decoding the encoded at least one anchor parameter to obtain at least one anchor parameter for identifying at least one anchor position and / or orientation relative to at least one origin position and / or orientation; and rendering an output signal based on the at least one audio signal, located and / or orientated within the audio scene based on the at least one anchor parameter reference being reproduced based on the at least one anchor reference. An apparatus comprising means for performing the actions of the method as described above.
[0050] An apparatus configured to perform the actions of the method as described above.
[0051] A computer program comprising program instructions for causing a computer to perform the method as described above.
[0052] A computer program product stored on a medium may cause an apparatus to perform the method as described herein.
[0053] An electronic device may comprise apparatus as described herein.
[0054] A chipset may comprise apparatus as described herein.
[0055] Embodiments of the present application aim to address problems associated with the state of the art.
[0056] Summary of the Figures
[0057] For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings in which:
[0058] Fig.1 a shows schematically a suitable environment showing an example of a combination of virtual scene elements within a physical listening space;
[0059] Fig.1 b shows schematically a system of rendering according to some embodiments;
[0060] Fig.2 shows schematically an example of the anchor mechanism suitable for implementation with respect to rendering an augmented reality scene according to some embodiments;
[0061] Fig.3 shows a flow diagram of the operation of a suitable Tenderer according to some embodiments;
[0062] Fig.4 shows a flow diagram of the operation of encoding anchor information according to some embodiments;
[0063] Fig.5 shows an example taxonomy of concepts which can be used as an anchor and respective anchor specific identifiers according to some embodiments;
[0064] Fig.6 shows a flow diagram of the operation of decoding anchor information and obtaining anchor position according to some embodiments;
[0065] Figs.7a and 7b show schematically systems of suitable for implementing some embodiments; and
[0066] Fig.8 show schematically an example device suitable for implementing the apparatus shown. Embodiments of the Application
[0067] The following describes in further detail suitable apparatus and possible mechanisms for describing an anchor with a universal description convention for creating correspondence between the content creator specified anchor and the listening space anchor information which enables a common presentation engine or a Tenderer entity or a player to correctly update the spatiotemporal coordinates of the content creator specified anchor.
[0068] As discussed earlier the listening space information can be represented as LSDF (Annex C of ISO / IEC 23090-34, Immersive audio reference software MPEG-I Immersive Audio Augmented Reality Listener Space Description Format, Version 3), N00250 document which is an output of ISO / IEC JTC1 SC29 WG6.
[0069] Fig.1 b illustrates a figure from annex C of ISO / IEC 23090-34, N00250, wherein the (virtual) scene information 151 is described during content creation and represents the scene as captured (or initially generated). This scene information can be captured or otherwise provided in an encoder input format (EIF) data format 153. The EIF and (captured or generated) audio data is used by an encoder to generate a bitstream 155 comprising the scene description and spatial audio metadata (and audio signals), which can be delivered via the bitstream to the Tenderer 171 device or apparatus.
[0070] As discussed above the second part of the AR audio scene rendering system is related to the physical listening space (or physical space) 161 of the listener (or end user). The scene or listener space information 161 may be obtained during the AR rendering (when the listener is consuming the content). This listening space information includes geometric and acoustic properties of the listening space such as the space geometry, its reverberation parameters and acoustic material properties of the walls. Furthermore, the listening space information also includes information for location information for audio elements in the scene in a particular location in the listening space via the mechanism of Anchors. The listening space information can be provided via the LSDF 163, which can, e.g., be a file. The LSDF file is a placeholder for a suitable interface that can be implemented by any AR consumption device to provide listening space information to the audio Tenderer 171 .
[0071] An anchor within an AR / MR / XR context can be labeled as <ARAnchor> and can be used to indicate a real-world position for the EIF to reference. Furthermore the <ARAnchor> data structure can be defined by the following:
[0072] The mechanism of the Anchor can furthermore be shown with respect to Fig.2 wherein in LSDF 201 there is defined a listening space 203 with an origin point 205 and the anchor point 207. This is combined with the Anchor point in the EIF 221 which comprises an anchor point with reference 223 and associated elements, element 1 225 and element 2 227. Thus when combined the listening space 233 comprises the origin 235 the anchor point 237 and superimposed elements element 1 245 and element 2 247. The EIF document which is Annex B of ISO / IEC 23090-34, Immersive audio reference software (MPEG-I Encoder Input Format) can be found as N00249 as output of ISO / IEC JTC1 SC29 WG06.
[0073] The MPEG-I immersive audio specification (ISO / IEC 23090-4) which is currently in the Draft International Standard (DIS) stage describes the anchors in the following syntax table (Table 66 of ISO / IEC 23090-4 DIS).
[0074]
[0075] The semantics of the Anchor related elements can be as follows: anchorsCount This value is the number of anchors in this payload anchorld This value represents the unique identifier for this anchor. anchorLsdfRef This string is the reference for the LSDF anchor which is known to the Tenderer via the LSDF. The LSDF anchor indicates the position of the anchor in the listener space.
[0076] Thus in summary the anchorLsdfRef is matched with the corresponding ARAnchor ID in the LSDF. There are systems, such as described hereafter, that can sense the surroundings to find semantic information in the scene which can subsequently be used for finding the relevant Anchor objects in the listening space. Furniture recognition with Meta Quest 3
[0077] (https: / / www.meta.com / en-gb / help / quest / articles / whats-new / release- notes / ?intern_source=blog&intern_content=meta-quest-v64-update-passthrough- improvements-external-microphones):
[0078] The Meta Quest 3 is able to scan the environment and understand the surroundings which enables segmenting certain objects of interest. With v64 though, at the end of mixed reality room scanning a Meta Quest 3 is configured to create a labelled rectangular cuboid bounding box around objects such as:
[0079] Doors;
[0080] Windows;
[0081] Beds;
[0082] Tables;
[0083] Sofas;
[0084] Storage (cabinets, shelves, etc); and
[0085] Screens (TVs and monitors).
[0086] Google’s ARCore Scene Semantics API is detailed for example in https: / / developers.google.eom / ar / develop / c / scene-semantics or https: / / developers.google.com / ar / develop / java / scene-semantics:
[0087] Where there is defined semantic label quality tiers such as
[0088] Main scene components: Sky; Building; Tree; Road; and Vehicle.
[0089] Major scene details: sidewalk; terrain; structure; and water.
[0090] Minor scene details: object; and person.
[0091] Furthermore Apple Roomplan labels support are as follows (https: / / machinelearning.apple.com / research / roomplan): Storage; Sofa; Table; Chair; Bed; Refrigerator; Oven; Stove; Dishwasher; Washer or dryer; Fireplace; Sink; Bathtub; Toilet; Stairs and TV.
[0092] For AR scene creation and consumption there are two components. The first component is describing during content creation the Anchor object which is expected in the AR scene. This can be anything, ranging from a physical feature such as a floor, a static object (e.g., window, table, poster, etc.), a dynamic object (e.g., a face or even a particular face). The second component is obtaining the position (and / or orientation) of the specific Anchor object in the listening space.
[0093] In a real-world deployment of AR playback or content consumption system, the Anchor is defined during a content creation phase of the work-flow from content creation to content consumption, but the listening space information is obtained with the knowledge of the scene to be consumed or agnostic of the scene to be consumed. This is an implementation specific choice thus a flexible approach for implementing the listening space information is required. The specified Anchor during content creation is carried in the scene specific bitstream to the consumption device.
[0094] According to the current MPEG-I immersive audio specification, the Anchor specific to the scene is specified as a string.
[0095] The listening space description generation is not normative or specific to a single method. This could lead to large variability in the generation of the listening space description information or the LSDF described in Annex C of ISO / IEC 23090-34.
[0096] The method for representation of the information obtained from the scene sensing methods to obtain the position of the Anchor object is specific to AR device platforms. Consequently, the current Anchor specification method is insufficient and under specified. For example, if the scenario where the content creator uses English language to describe the desired Anchor object “picture_frame” but the listening space description uses a different representation “frame” when it detects a window, the two may not match or even match incorrectly. This would result in placement of the Anchor in the wrong position and would be presented in the scene in an incorrect manner.
[0097] In case of different languages or different semantic concepts, the interoperability challenges make it harder to determine whether the scene is rendered correctly and harder to effect control to maintain a correct rendering. For example there is the possibility of false positive or false negative anchor detection and incorrect placement of anchors in the AR scene. In summary the interoperability challenges could thus lead to improper placement of the audio scene elements which will not be according to how the scene was authored or intended and leading to inconsistencies within the scene which in turn results in incorrect scene rendering.
[0098] The aim of the embodiments as discussed in further detail hereafter is to address the problem of reliance or dependance on simple string matching (between the content creator specific bitstream and listening space description generated by the AR consumption device) and to aim to produce apparatus and methods to enable an interoperable standard with consistent high-quality subjective performance.
[0099] The high-quality subjective performance may be adversely impacted due to incorrect placement of audio scene elements leading to rendered audio output being inconsistent with authoring intent. For example, an important audio object if placed at the origin of the coordinate system in absence of finding the correct Anchor can result in audio object location that may be in a place that does not result in any audible audio or in some other scenarios the audio object related to the Anchor may be put in a location that overlaps with another audio element in the scene thus masking the other audio element in the rendered audio output. In case the player has a policy to not render any audio for an anchor that is not detected, this may result in a silent scene or rendered audio is silence which is not the intent of the scene creator. Thus, leading to poor subjective quality and inconsistency across different users of the same content.
[0100] In the current disclosure a trackable is anything that can be detected and tracked in the real world. For example such as defined in the document https: / / docs.unity3d.com / Packages / com.unity.xr.arfoundation%403.0 / manual / trackable- managers.html. The trackables are certain objects or features that can be labelled. Some examples are finding the floor, wall, an object with a certain point cloud model, etc.
[0101] The concept as discussed in further by the following embodiments is related to rendering an AR scene where apparatus and methods are proposed for describing an anchor with a universal description convention for creating correspondence between the content creator specified anchor and the listening space anchor information. Thus when embodiments are implemented a common presentation engine or a Tenderer entity or a player is able to correctly update the spatiotemporal coordinates of the content creator specified anchor which results in selecting the right object in the listening space (i.e., obtain position for the correct anchor from among the multiple objects detected in the listening space) by aligning the representation method of the tracking system with the content creator specified anchor representation method.
[0102] This is achieved by in some embodiments by extending the anchor reference definition in the content creator specified anchor which comprises at least one of the following: representation method index which describes the convention for the representation (e.g., URN or Universal Resource Name); location (e.g., URL or Universal Resource Locator) for the representation convention describes the location to obtain the representation convention.
[0103] In some embodiments, the representation convention may be configured to represent trackables. In some further embodiments, the anchor reference definition may carry taxonomy identification information and the taxonomy element index corresponding to the specified anchor.
[0104] In some embodiments, the anchor reference definition may carry multiple representation method references to enable interoperability across multiple AR platforms for the scene specific bitstream.
[0105] These embodiments can be implemented for apparatus or methods employing AR or in general MR (mixed reality) / XR (extended reality) content rendering where part of the scene information is known during content creation and another part is known only during content consumption.
[0106] The embodiments aim, for such scenes, to achieve a correct understanding between the content creation intent and content consumption context. This is achieved by including content consumption context related cues which are of interest in the content creation intent information and which is scene specific bitstream information. This can, for example, be achieved by the following: representing content consumption context related cues, such as anchors, using a method that is correctly and consistently interpreted by the content creator specified anchor information and content consumption anchor information; including the Anchor information in the content creator generated scene specific bitstream using the method that can be correctly interpreted by the Tenderer or media player; and generating listening space information based on the scene specific anchor information to ensure the right Anchors are detected and tracked, resulting in correct coordinates (position and orientation) being delivered to the audio Tenderer for performing the audio scene rendering. In some implementations the correct Anchor’s position and orientation is selected from the prior generated content consumption space information.
[0107] Thus, the embodiments facilitate the correct interpretation of the Anchor information by the Tenderer and the content consumption space context generation entity.
[0108] With respect to Fig.3 is shown a flow diagram showing the overview of the embodiments with respect to representation of Anchor information in the bitstream with universal description convention to enable interoperability with a wide range of AR consumption devices. For example, as shown in Fig.3 by 301 is the operation of representing anchor information during content creation with a universal description convention so that it can be interpreted consistently and correctly by other entities.
[0109] Then as shown by 303 in Fig.3 is the operation of including anchor information with one or more universal description conventions including representation identifiers of the universal description convention identifier(s).
[0110] Following this is shown in Fig.3 by 305 is the operation of generating listening space information with one or more universal description conventions including representation identifiers of the universal description convention identifier(s).
[0111] Additionally as shown in Fig.3 by 307 is the operation of selecting the corresponding Anchor position from the listening space information obtained from the player or context sensing subsystem or provide anchor information in a format that is compatible with the supported universal description conventions by the player or context sensing subsystem.
[0112] Finally in this summary, as shown in Fig.3 by 309 is the operation of determining correct position and orientation information for the anchor in the listening space.
[0113] With respect to Fig.4 is shown in further detail the operation of representing anchor information during content creation with the universal description convention so that it can be interpreted consistently and correctly by other entities.
[0114] The first operation involves obtaining the anchor information that is required to be specified via the scene authoring or content creator specified authoring format as shown in Fig.4 by 401 . An example of such anchor information can be a “table” which will have an associated virtual audio object which will be placed on top of the table to represent a virtual Boombox.
[0115] The next operation, as shown in Fig.4 by 403, is where the content creator is configured to obtain desired (or consider the use of) one or more anchor representation methods which may be used.
[0116] Following this, as shown in Fig.4 by 405, is the operation of including, based on the desired representation methods, respective representation method URNs for the anchor information to be included in the anchor reference syntax element in scene plus payload or any suitable element in the bitstream. This ensures that the representation method is recognized by the AR consumption device implementations which generate the listening space information and track the relevant objects in the listening space. Subsequently, as shown in Fig.4 by 407 is the identification of the anchor specific identifier for the given representation method. In some examples the representation method employs a selection of words from the English language as a dictionary, from which the desired anchor can be identified by a single word.
[0117] Finally, as shown in Fig.4 by 409 is the operation of encoding the anchor specified identifier and the representation method in the scene specific bitstream.
[0118] In some embodiments a representation method involving a classification scheme, each class may have a separate identifier. For example, Fig.5 shows an example simple taxonomy representation (with Anchor specific identifiers for each concept that can be used as an anchor). Thus for example a first layer or level comprises the space 501 with an identifier of 0. The space can be further defined by the next layer which has identifiers indoor 511 with an identifier 10, or outdoor 513 with an identifier 11 . Then with respect to the indoor 511 there are further sub-layers, for example furniture 525 with identifier 20, floor 523 with identifier 21 and wall 521 with identifier 22. Additionally is shown the Furniture 525 identifier further being defined by a sub-layer showing Furniture 531 with identifier 30, Furniture 533 with identifier 31 and Furniture 535 with identifier 32.
[0119] This example is a suggestion of a possible taxonomy and any suitable representation method and Anchor specific identifier definition scheme could be used. In another implementation embodiment, the Anchor description could be a URI which could be defined for a supported Anchor description approach to carry in the MPEG-I immersive audio bitstream or used by the AR sensing and tracking subsystem to indicate the detected anchor objects in listening space information. For example, when anchorLsdfRef is carried by MPEG-I bitstream then this can definescheme: anchor-description (e.g., anchords) authority: Approach (Implementation approaches used by implementations e.g., Company X approach)- query or path: Category (Broader category to enable easy disambiguation e.g., Home, Living room, Furniture, Indoor, Shopping Mall, Museum, etc.) fragment: AnchorName or Anchorldx (a specific Anchor object within the Approach and broad category e.g., Picture frame which is in Indoor and according to Approach 1 )
[0120] URI = anchords: / / Approach?Category#AnchorName / Anchorldx
[0121] In different implementations, the URI could also indicate if the anchor being selected should be static or dynamic. For example, one entity could assign URI equal to anchords: / / Approach?Dynamic#Anchorldx The presence of Anchors in the listening space is directly related to the listening space where the AR scene is being consumed. Consequently, the determination of Anchor position in the AR scene is performed during playback consumption, when the user is in the listening space physically.
[0122] For example there can be employed suitable apparatus to implement the method steps such as shown with respect to Fig.6.
[0123] Thus as shown by 601 in Fig.6 is the operation of retrieving or receiving the AR scene specific bitstream from the content delivery server or a local file (for example within an AR scene player). Subsequently the scene specific bitstream can be received by the audio Tenderer.
[0124] Furthermore, as shown by 603 in Fig.6, is an operation of parsing the received scene specific bitstream to determine the type of scene (AR or VR). This can for example be implemented in an audio Tenderer, or suitable apparatus. In the example of an AR scene, the Tenderer can also expect or determine the presence of anchor information in the bitstream.
[0125] Then, as shown by 605, is an operation of reading, for example within the Tenderer, the anchor information in the bitstream to obtain the representation method and anchor specific identifier for each of the anchors specified in the bitstream.
[0126] In some embodiments there can be a first type of approach (referred to as a type 1 or T1 approach and shown by the right side of Fig.6.
[0127] The a priori listening space generation (T1 ) approach comprises a priori generation of the listening space information. This approach is applicable for scenarios where the AR scene consumption device is sensing on a continual basis, the surroundings in general or in particular, listening space. Thus, the device maintains a current version of listening space description. In different implementation embodiment, the listening space information can also be obtained via a location based service which updates the listening space information based on the AR scene consumption device.
[0128] Thus, there is an operation, as shown by 617 of Fig.6 of generating AR anchors information comprising representation method, anchor specific identifier and position information for each candidate Anchor in the listening space. This is delivered to the Tenderer as data (via a file or an interface, etc.). For example, the AR scene consumption device has generated a list of candidate AR anchors listing each of the AR anchors detected and tracked by the AR device. Depending on the implementation choices, the AR playback device can implement one or more Anchor representation methods and subsequently identify each of the detected and tracked anchors with the respective Anchor specific identifier for each of the supported methods. This information can be made available to the Tenderer via a file, an interface or any other suitable mechanism.
[0129] Furthermore, as shown by 619 of Fig.9 is the operation of receiving AR anchor information comprising representation method and anchor specific identifier information to the Tenderer. In other words the candidate AR anchors which are included in the listening space information are received by the Tenderer via a suitable interface. This comprises anchor positions in addition to the representation method as well as Anchor specific identifier.
[0130] Following this is the operation, as shown by 621 of Fig.9, of selecting relevant anchors based on the anchor information from the scene specific bitstream and candidate anchors from the listening space information. For example, in some embodiments, the Tenderer identifies the relevant anchors in the listening space by matching the anchor information comprising anchor representation and anchor specific identifier. This can be obtained by finding the matching Anchor specific identifier in the bitstream as well as listening space information. In some implementation embodiments, there may be further implementation specific logic in the Tenderer to find the appropriate match. This can be due to non-availability or multiple availability of the anchor in the listening space.
[0131] Finally as shown in Fig.6 by 631 is the operation of updating a Tenderer scene state with the Anchor specific position information. In other words updating the audio Tenderer scene state with the position information for the detected and trackable anchors. In some embodiments, for static anchors, this position update happens only once during initialization. In some embodiments, for dynamic or moving anchors, the position information is updated frequently. Furthermore the position of the audio elements or other scene elements which is specified as part of the Anchor transform can be updated in accordance with the Anchor position
[0132] In some embodiments the posteriori listening space generation (T2) approach comprises a posteriori generation of the listening space information. This approach is applicable for scenarios where the AR scene consumption device is sensing the listening space to detect and track the anchors described in the scene specific bitstream. Thus, the apparatus is configured to initiate listening space information generation after selection of the AR scene to be consumed.
[0133] In this (T2) approach, following the operation of extracting the representation method and anchor specific identifier to obtain anchor information from the scene specific bitstream is shown by 607 in Fig.6 the operation of sending anchor representation, anchor Specific Identifier to the player or presentation engine or AR sensing subsystem.
[0134] In other words, sending the AR anchor information comprising representation method and anchor specific identifier from the Tenderer to the player or presentation engine or AR sensing subsystem in the AR consumption device. In this operation the bitstream specified anchor is indicated to the listening space sensing system to determine if the object of interest (anchor) is detected and trackable.
[0135] Then is the operation of receiving the AR Anchor position information at the Tenderer as shown by 609 in Fig.6. Thus the AR consumption device or player or presentation engine responds with the successfully detected and trackable anchors with the positions to the Tenderer. In some implementation embodiments, the AR sensing and tracking subsystem may consider additional information such as the quality or confidence level to filter false positive detections.
[0136] Finally as shown in Fig.6 by 631 is the operation of updating a Tenderer scene state with the Anchor specific position information (in other words both approaches are similar from this point).
[0137] In some embodiments the bitstream syntax employed to describe Anchor information with universal description convention can, for example be shown as the following
[0138]
[0139] In this example anchorReferenceType indicates the type of anchor reference information being included in the bitstream. The different types of information can be indicated by a table such as:
[0140] For anchorReferenceType equal to 1 , URN can be for example, urn:mpegi:2024:anchorlabel:language / anchorname.
[0141] For example, urn:mpegi:2024:anchorlabel:english / Television.
[0142] Furthermore the num_method_bytes if anchorReferenceType is equal to 2, the num_method_bytes indicates the number of bytes carried in addition to the URI which indicates the information about the Anchor description. For example, some bytes may indicate the object type and other bytes may indicate further information about the color or proximity of the Anchor or any other such information. For other values of anchorReferenceType this value is not carried in the bitstream. Additionally in some embodimens the num_method_byte_value is carried in the bitstream if anchorReferenceType is equal to 2., the num_method_byte_value may be used to indicate the anchor specific identifier, other attributes (e.g., color, proximity, etc. of the physical object specified as the anchor).
[0143] The identifier anchorSpecificldentifier indicates the specific Anchor for the specified taxonomy in the anchorLsdfRef when the anchorReferenceType is equal to 3. The anchorLsdfRef indicates the representation method type.
[0144] The listening space description furthermore, in some embodiments, can carry information regarding the anchor description method (representation method) in addition to the Anchor identifier. The combination of ARAnchor and ARAnchorConvention enables the Tenderer to interpret the Anchor information in the listening space description in a consistent manner. In order to enable consistent end to end interoperability between the scene specific bitstream and the listening space description, it is important to ensure that both these entities (bitstream creation and listening space representation need to follow the same convention regarding Anchor description.
[0145] Figs.7a and 7b show example system implementations corresponding to T1 and T2 approaches as described in Fig.6 according to some embodiments. For both Fig.7a and 7b the system can comprise a content creator 700 which can be implemented on any suitable computer or processing device. The content creator 700 comprises an (MPEG-I Immersive audio) encoder 701 which is configured to receive the audio scene description 702 which includes anchors and the audio signals or data 704. The audio scene description 702 can be provided in the MPEG-I Immersive audio Encoder Input Format (EIF) or in other suitable format. Generally, the audio scene description contains an acoustically relevant description of the contents of the audio scene, and contains, for example, the scene geometry as a mesh or voxel, acoustic materials, acoustic environments with reverberation parameters, positions of sound sources, and other audio element related parameters such as whether reverberation is to be rendered for an audio element or not. The MPEG-I Immersive audio encoder 701 is configured to output encoded data with representation specified anchors 706.
[0146] The content creator 700 furthermore in some embodiments comprises a MPEG-H encoder 703 and generates MPEG-H audio bitstream 708.
[0147] The MPEG-H audio bitstream 708 and 6DoF audio bitstream with representation specified anchors 706 in some embodiments can be streamed to end-user devices or made available for download or stored.
[0148] Additionally the system comprises a server 710 configured to obtain the bitstreams, and store them (in bitstream storage 711 ) and supply them (for example as a six degrees of freedom (6-DoF) audio bitstream 712) to the player 720.
[0149] The relevant (6DoF audio) bitstream 712 is retrieved by the player 720. In some embodiments other implementation options are feasible such as broadcast, multicast.
[0150] Fig.7a for example shows the scenario where the LSDF generation (listening space information generation) is performed independently of the scene. Furthermore, difference AR scene consumption platforms can utilize their own methods to implement the listening space generation, independent of the content creation. The method expects leveraging the commonly available AR platform AR anchor representation methods and using those with additional information to identify the representation method. This approach has the benefit of performing AR sensing continuously, thus the latency to detect Anchors in the listening space is near instantaneous resulting in improved experience for the end user. The cost of this instantaneous information is additional computational load when the device is not consuming the AR scene. This approach for continuous sensing can be used by the device for other utilities such as context sensing for other applications thus dividing the load to multiple applications. Consequently, in this case, the bitstream specified Anchor information is compared only within the Tenderer with the listening space information provided by the player or presentation engine or the playback device AR subsystem.
[0151] The player 720 in some embodiments comprises a playback device 721 configured to obtain or receive the 6D0F audio bitstream (with representation specified anchors) 712, and furthermore can be configured to receive or otherwise obtain the 6 DoF tracking information (listener orientation or position information) 740 from a suitable listener user interface, for example from the head mounted device (HMD) 741 . These can for example be generated by sensors within the HMD 741 or from sensors in the environment sensing the orientation or position of the listener.
[0152] In some embodiments the playback device 721 comprises a bitstream parser 723 configured to obtain the encoded bitstream 712 and decode these in an opposite or inverse operation to the encoders 701 , 703 to generate audio and metadata 724 (for example as an MPEG-I bitstream) which can be passed to a MPEG-I Immersive audio Tenderer 725.
[0153] In some embodiments the playback device 721 comprises the MPEG-I Immersive audio Tenderer 725 configured to implement the rendering operations as described above and hereafter and generate audio output signals 742 which can be output to the head mounted device 741 .
[0154] The playback device 721 can be implemented in different form factors depending on the application. In some embodiments the playback device is equipped with its own listener position tracking apparatus or receives the listener position information from an external apparatus. The playback device can in some embodiments be also equipped with headphone connector to deliver output of the rendered binaural audio to the headphones. Additionally the player 720 comprises a listening space description generator 730. The listening space description generator 730 communicates with the MPEG-I audio Tenderer 725 such that it able to receive the anchor representation information 728 from the audio Tenderer 725 and send the detected anchors 726 to the audio Tenderer 725.
[0155] Fig.7b illustrates the scenario where the AR anchor information is delivered from the bitstream to the AR player subsystem which is handling the generation of the listening space information. In this case the process of detecting and tracking the AR anchor is triggered only when the AR scene is selected for playback and bitstream is retrieved. In this case, the dependency on the AR playback device is to have a common vocabulary to describe the AR Anchor information and provide the information for the same to the Tenderer.
[0156] The player 720 in this approach comprises a playback device 721 configured to obtain or receive the 6DoF audio bitstream (with representation specified anchors) 712, and furthermore can be configured to receive or otherwise obtain the 6 DoF tracking information (listener orientation or position information) 740 from a suitable listener user interface, for example from the head mounted device (HMD) 741 . These can for example be generated by sensors within the HMD 741 or from sensors in the environment sensing the orientation or position of the listener.
[0157] In some embodiments the playback device 721 comprises a bitstream parser 723 configured to obtain the encoded bitstream 712 and decode these in an opposite or inverse operation to the encoders 701 , 703 to generate audio and metadata 754 which can be passed to a MPEG-I Immersive audio Tenderer 725.
[0158] In some embodiments the playback device 721 comprises the MPEG-I Immersive audio Tenderer 725 configured to implement the rendering operations as described above and hereafter and generate audio output signals 742 which can be output to the head mounted device 741 .
[0159] The playback device 721 can be implemented in different form factors depending on the application. In some embodiments the playback device is equipped with its own listener position tracking apparatus or receives the listener position information from an external apparatus. The playback device can in some embodiments be also equipped with headphone connector to deliver output of the rendered binaural audio to the headphones. Additionally the player 720 comprises a listening space scene sensor 751. The listening space scene sensor 751 communicates with the MPEG-I audio Tenderer 725 such that it able to receive Bitstream Anchors with representation information 748 from the audio Tenderer 725 and send the position for detected anchors 746 to the audio Tenderer 725. Additionally the listening space scene sensor 751 is configured generate the interactive anchor selection indicator 744.
[0160] The rendering is the same for the AR scene after the AR anchor position has been determined. In different implementation embodiments, the two approaches or methods can be used in cascade (Type 1 followed by Type 2 on a need basis, if an anchor is not detected via a Type 1 approach). Furthermore, a hybrid method combining the two can be implemented, where the Type 1 approach of listening space Anchor detection and sensing is performed with the help of scene specific anchor information. This hybrid method may be especially useful in case the expected AR scenes are known a prior based on listener position information (e.g., location based advertising, etc.).
[0161] In such embodiments an example scenario can be envisaged without loss of generality. A content creator authors an AR scene which envisages rendering an audio track corresponding to the position of the specified anchor object in the content consumption space. For this the scene creator specifies “picture” as the Anchor description.
[0162] In examples where the above embodiments are not implemented, if the object “picture” is not recognized because someone implemented the picture feature with a label “photograph” due to absence of any consistent description method specified in the anchor description. This may lead to not detecting the anchor object which can result in no audio being rendered.
[0163] Furthermore in some embodiments, the anchor information can be described in a conformant manner (in compliance with the indicated representation method), and the picture placed in the content consumption space is augmented with the audio track envisioned by the content creator.
[0164] The information (or anchor information) can in some embodiments be obtained from a bitstream parameter or payload. In other words the information can be read from the bitstream parameter or payload or be provided in the bitstream parameter or as part of bitstream parameter. Thus how the bitstream arrives at the Tenderer or how it is made available to the Tenderer is a question of implementation. In some embodiments, the anchor or anchor reference can point in addition to the audio elements, the listening space features such as acoustic environments, acoustic portals, scene geometry elements in the physical space. In other words the anchor or anchor reference can be employed to identify any suitable acoustic feature such as at least one acoustic environment, at least one acoustic portal, and at least one scene geometry element in a physical space.
[0165] With respect to Figure 8 an example electronic device which may represent any of the apparatus shown above. The device may be any suitable electronics device or apparatus. For example in some embodiments the device 1400 is a mobile device, user equipment, tablet computer, computer, audio playback apparatus, etc.
[0166] In some embodiments the device 1400 comprises at least one processor or central processing unit 1407. The processor 1407 can be configured to execute various program codes such as the methods such as described herein.
[0167] In some embodiments the device 1400 comprises a memory 1411. In some embodiments the at least one processor 1407 is coupled to the memory 1411. The memory 1411 can be any suitable storage means. In some embodiments the memory 1411 comprises a program code section for storing program codes implementable upon the processor 1407. Furthermore in some embodiments the memory 1411 can further comprise a stored data section for storing data, for example data that has been processed or to be processed in accordance with the embodiments as described herein. The implemented program code stored within the program code section and the data stored within the stored data section can be retrieved by the processor 1407 whenever needed via the memory-processor coupling.
[0168] In some embodiments the device 1400 comprises a user interface 1405. The user interface 1405 can be coupled in some embodiments to the processor 1407. In some embodiments the processor 1407 can control the operation of the user interface 1405 and receive inputs from the user interface 1405. In some embodiments the user interface 1405 can enable a user to input commands to the device 1400, for example via a keypad. In some embodiments the user interface 1405 can enable the user to obtain information from the device 1400. For example the user interface 1405 may comprise a display configured to display information from the device 1400 to the user. The user interface 1405 can in some embodiments comprise a touch screen or touch interface capable of both enabling information to be entered to the device 1400 and further displaying information to the user of the device 1400. In some embodiments the user interface 1405 may be the user interface for communicating with the position determiner as described herein.
[0169] In some embodiments the device 1400 comprises an input / output port 1409. The input / output port 1409 in some embodiments comprises a transceiver. The transceiver in such embodiments can be coupled to the processor 1407 and configured to enable a communication with other apparatus or electronic devices, for example via a wireless communications network. The transceiver or any suitable transceiver or transmitter and / or receiver means can in some embodiments be configured to communicate with other electronic devices or apparatus via a wire or wired coupling.
[0170] The transceiver can communicate with further apparatus by any suitable known communications protocol. For example in some embodiments the transceiver can use a suitable universal mobile telecommunications system (UMTS) protocol, a wireless local area network (WLAN) protocol such as for example IEEE 802. X, a suitable short-range radio frequency communication protocol such as Bluetooth, or infrared data communication pathway (IRDA).
[0171] The transceiver input / output port 1409 may be configured to receive the signals and in some embodiments determine the parameters as described herein by using the processor 1407 executing suitable code.
[0172] It is also noted herein that while the above describes example embodiments, there are several variations and modifications which may be made to the disclosed solution without departing from the scope of the present invention.
[0173] In general, the various embodiments may be implemented in hardware or special purpose circuitry, software, logic or any combination thereof. Some aspects of the disclosure may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the disclosure is not limited thereto. While various aspects of the disclosure may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
[0174] As used in this application, the term “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry) and
[0175] (b) combinations of hardware circuits and software, such as (as applicable):
[0176] (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and
[0177] (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and
[0178] (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.”
[0179] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware.
[0180] The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.
[0181] The embodiments of this disclosure may be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Computer software or program, also called program product, including software routines, applets and / or macros, may be stored in any apparatus-readable data storage medium and they comprise program instructions to perform particular tasks. A computer program product may comprise one or more computer-executable components which, when the program is run, are configured to carry out embodiments. The one or more computer-executable components may be at least one software code or portions of it.
[0182] Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD. The physical media is a non-transitory media.
[0183] The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processors may be of any type suitable to the local technical environment, and may comprise one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), FPGA, gate level circuits and processors based on multi core processor architecture, as non-limiting examples.
[0184] Embodiments of the disclosure may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.
[0185] The scope of protection sought for various embodiments of the disclosure is set out by the independent claims. The embodiments and features, if any, described in this specification that do not fall under the scope of the independent claims are to be interpreted as examples useful for understanding various embodiments of the disclosure.
[0186] The foregoing description has provided by way of non-limiting examples a full and informative description of the exemplary embodiment of this disclosure. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this disclosure will still fall within the scope of this invention as defined in the appended claims. Indeed, there is a further embodiment comprising a combination of one or more embodiments with any of the other embodiments previously discussed.
Claims
CLAIMS:1 . An apparatus comprising at least one processor and at least one memory including a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to: obtain information associated with an audio scene, wherein the information is made available to a rendering device using a bitstream; obtain at least one audio signal, the at least one audio signal is associated with at least one audio object within the audio scene; obtain at least one anchor reference associated with the at least one audio object, wherein the at least one anchor reference comprises an anchor reference description information, such that an audio scene rendering based on the at least one audio signal reproduces the audio scene based on the anchor reference description information; obtain at least one anchor parameter for identifying at least one anchor position and / or orientation; and transmit the bitstream comprising: the at least one audio signal; the at least one anchor reference; and the at least one anchor parameter.
2. The apparatus as claimed in claim 1 , caused to obtain the at least one anchor reference associated with the at least one audio object is further caused to define a universal description convention which can be interpreted consistently and correctly by the apparatus and the rendering device.
3. The apparatus as claimed in claim 2, caused to obtain the at least one anchor reference is caused to determine anchor reference description information based on the universal description convention.
4. The apparatus as claimed in any of claims 1 to 3, caused to transmit the bitstream is caused to transmit the bitstream to a further apparatus comprising at least the rendering device, wherein the rendering device is configured to render an output audio signal based on the at least one audio signal and located and / or orientated within the audio scene37based on the at least one anchor parameter reference correctly interpreted based on the at least one anchor reference.
5. An apparatus comprising at least one processor and at least one memory including a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to: obtain a bitstream, wherein the bitstream comprises: information associated with an audio scene; an encoded at least one anchor reference; an encoded at least one audio signal; and an encoded at least one anchor parameter; obtain the encoded at least one anchor reference, the encoded at least one audio signal; and the encoded at least one anchor parameter from the obtained bitstream; decode the encoded at least one audio signal to obtain at least one audio signal, the at least one audio signal is associated with at least one audio object within the audio scene; decode the encoded at least one anchor reference to obtain at least one anchor reference associated with the at least one audio object, wherein the at least one anchor reference comprises an anchor reference description information, such that an audio scene rendering based on the at least one audio signal reproduces the audio scene based on the anchor reference description information; decode the encoded at least one anchor parameter to obtain at least one anchor parameter for identifying at least one anchor position and / or orientation relative to at least one origin position and / or orientation; and render an output signal based on the at least one audio signal, located and / or orientated within the audio scene based on the at least one anchor parameter being reproduced based on the at least one anchor reference.
6. The apparatus as claimed in claim 5, caused to identify the at least one anchor position and / or orientation within the audio scene based on the anchor reference description information and the at least one anchor parameter.
7. The apparatus as claimed in any of claims 5 or 6, is caused to select the at least one anchor based on the at least one anchor reference description information and at least one candidate anchor obtained from listening space information.
8. The apparatus as claimed in any of claims 1 to 7, wherein the anchor reference description information is configured to one of: describe an anchor reference definition; and identify a location where the anchor reference definition is accessible.
9. The apparatus as claimed in any of claims 1 to 8, wherein the audio scene comprises one of: an augmented reality audio scene; a virtual reality audio scene; and a mixed reality audio scene.
10. The apparatus as claimed in any of claims 1 to 9, wherein the anchor reference further can employed to identify at least one of: at least one acoustic environment; at least one acoustic portal; and at least one scene geometry element in a physical space.11 . The apparatus as claimed in any of claims 1 to 10, wherein the at least one anchor reference and the at least one anchor parameter are components of a parameter.
12. A method for an apparatus, the method comprising: obtaining information associated with an audio scene, wherein the information is made available to a rendering device using a bitstream; obtaining at least one audio signal, the at least one audio signal is associated with at least one audio object within the audio scene; obtaining at least one anchor reference associated with the at least one audio object, wherein the at least one anchor reference comprises an anchor reference description information, such that an audio scene rendering based on the at least one audio signal reproduces the audio scene based on the anchor reference description information; obtaining at least one anchor parameter for identifying at least one anchor position and / or orientation; and transmit the bitstream comprising: the at least one audio signal; the at least one anchor reference; andthe at least one anchor parameter.
13. The method as claimed in claim 12, wherein obtaining the at least one anchor reference associated with the at least one audio object further comprises defining a universal description convention which can be interpreted consistently and correctly by the apparatus and the rendering device.
14. The method as claimed in claim 13, wherein obtaining the at least one anchor reference further comprises determining anchor reference description information based on the universal description convention.
15. The method as claimed in any of claims 12 to 14, wherein transmitting the bitstream comprises transmitting the bitstream to a further apparatus comprising at least the rendering device, wherein the rendering device is configured to render an output audio signal based on the at least one audio signal and located and / or orientated within the audio scene based on the at least one anchor parameter reference correctly interpreted based on the at least one anchor reference.
16. A method for an apparatus, the method comprising: obtaining a bitstream, wherein the bitstream comprises: information associated with an audio scene; an encoded at least one anchor reference; an encoded at least one audio signal; and an encoded at least one anchor parameter; and obtain the encoded at least one anchor reference, the encoded at least one audio signal; and the encoded at least one anchor parameter from the obtained bitstream; decoding the encoded at least one audio signal to obtain at least one audio signal, the at least one audio signal is associated with at least one audio object within the audio scene; decoding the encoded at least one anchor reference to obtain at least one anchor reference associated with the at least one audio object, wherein the at least one anchor reference comprises an anchor reference description information, such that an audio scene rendering based on the at least one audio signal reproduces the audio scene based on the anchor reference description information; decoding the encoded at least one anchor parameter to obtain at least one anchor parameter for identifying at least one anchor position and / or orientation relative to at least one origin position and / or orientation; andrendering an output signal based on the at least one audio signal, located and / or orientated within the audio scene based on the at least one anchor parameter reference being reproduced based on the at least one anchor reference.
17. The method as claimed in claim 16, further comprising identifying the at least one anchor position and / or orientation within the audio scene based on the anchor reference description information and the at least one anchor parameter.
18. The method as claimed in any of claims 16 or 17, further comprising selecting the at least one anchor based on the at least one anchor reference description information and at least one candidate anchor obtained from listening space information.
19. The method as claimed in any of claims 12 to 18, wherein the anchor reference description information is configured to one of: describe an anchor reference definition; and identify a location where the anchor reference definition is accessible.
20. The method as claimed in any of claims 12 to 19, wherein the audio scene comprises one of: an augmented reality audio scene; a virtual reality audio scene; and a mixed reality audio scene.
21. The method as claimed in any of claims 12 to 20, wherein the anchor reference further can employed to identify at least one of: at least one acoustic environment; at least one acoustic portal; and at least one scene geometry element in a physical space.
22. The method as claimed in any of claims 12 to 21 , wherein the at least one anchor reference and the at least one anchor parameter are components of a parameter.
23. An apparatus comprising means configured to:obtain information associated with an audio scene, wherein the information is made available to a rendering device using a bitstream; obtain at least one audio signal, the at least one audio signal is associated with at least one audio object within the audio scene; obtain at least one anchor reference associated with the at least one audio object, wherein the at least one anchor reference comprises an anchor reference description information, such that an audio scene rendering based on the at least one audio signal reproduces the audio scene based on the anchor reference description information; obtain at least one anchor parameter for identifying at least one anchor position and / or orientation; and transmit the bitstream comprising: the at least one audio signal; the at least one anchor reference; and the at least one anchor parameter.
24. An apparatus comprising means configured to: obtain a bitstream, wherein the bitstream comprises: information associated with an audio scene; an encoded at least one anchor reference; an encoded at least one audio signal; and an encoded at least one anchor parameter; obtain the encoded at least one anchor reference, the encoded at least one audio signal; and the encoded at least one anchor parameter from the obtained bitstream; decode the encoded at least one audio signal to obtain at least one audio signal, the at least one audio signal is associated with at least one audio object within the audio scene; decode the encoded at least one anchor reference to obtain at least one anchor reference associated with the at least one audio object, wherein the at least one anchor reference comprises an anchor reference description information, such that an audio scene rendering based on the at least one audio signal reproduces the audio scene based on the anchor reference description information; decode the encoded at least one anchor parameter to obtain at least one anchor parameter for identifying at least one anchor position and / or orientation relative to at least one origin position and / or orientation; and42render an output signal based on the at least one audio signal, located and / or orientated within the audio scene based on the at least one anchor parameter being reproduced based on the at least one anchor reference.43
Citation Information
Patent Citations
A method and apparatus for ar rendering adaption
GB2608847A
A Method and Apparatus for Scene Dependent Listener Space Adaptation
US20240048936A1
Real objects anchoring in a scene description
WO2025068063A1