Audio-visual presentation apparatus and method of operating same

Through mappers and category coordinate system transformation in audio-visual presentation devices, the problem of audio and visual scenes responding to user head movement in virtual reality and augmented reality is solved, providing flexible spatial perception and improving immersive experience.

CN120298630APending Publication Date: 2025-07-11KONINKLIJKE PHILIPS NV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510370784.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-10-13
Filing Date
2021-10-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In existing virtual and augmented reality applications, the presentation of audio and visual scenes is difficult to respond to user head movements, resulting in reduced immersive experiences, and existing methods often fail to provide natural and consistent spatial perception.

Method used

Through the audio-visual presentation device, the input posture is mapped to a rendering coordinate system fixed to the user's head movement with the mapper. Combined with the transformation of the coordinate system in different categories, the presentation mode of audio and visual items is dynamically adjusted to adapt to the user's head movement and provide flexible spatial perception and immersive experience.

Benefits of technology

It realizes flexible presentation of audio and visual projects, adapts to user head movement, improves the naturalness and immersion of the user experience, and reduces system complexity and resource requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298630A_ABST
    Figure CN120298630A_ABST
Patent Text Reader

Abstract

An audiovisual presentation device comprises a receiver (201) to receive an audiovisual item and a receiver (209) to receive metadata comprising an input pose provided with reference to an input coordinate system and a presentation category flag indicating a presentation category. A receiver (213) receives user head movement data, and a mapper (211) maps the input gesture to a presentation gesture in a presentation coordinate system in response to the user head movement data. A renderer (203) renders the audiovisual item using the rendering gesture. Each presentation category is linked with a different coordinate system transformation from a real world coordinate system to a category coordinate system, at least one coordinate system in the different coordinate system transformation being variable relative to the real world coordinate system and the presentation coordinate system. The mapper selects a presentation category for the audiovisual item in response to the presentation category flag, and maps the input gesture to a presentation gesture corresponding to a fixed gesture in a category coordinate system for varying user head movement, the category coordinate system is determined from the coordinate system transformation of the presentation category.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the application with the application number 202180070194.0 and the invention name "Audio-visual presentation device and its operation method", which was filed on April 13, 2023. Technical Field

[0002] The present invention relates to an audio-visual presentation device and its operation method, and more particularly but not exclusively to the use of these to support, for example, augmented / virtual reality applications. Background Art

[0003] In recent years, with the continuous development and launch of new services and methods for utilizing and consuming audio-visual content, the types and scope of experiences based on audio-visual content have increased significantly. Specifically, many spatial and interactive services, applications, and experiences are being developed to give users a more engaging and immersive experience.

[0004] Examples of such applications are virtual reality (VR), augmented reality (AR), and mixed reality (MR) applications that are rapidly becoming mainstream, where many technical solutions are aimed at the consumer market. Many standards are also being developed by many standardization bodies. Such standardization activities are actively developing standards for various aspects of VR / AR / MR systems, including, for example, streaming, dissemination, presentation, etc.

[0005] VR applications tend to provide a user experience corresponding to the user in different worlds / environments / scenarios, while AR (including mixed reality MR) applications tend to provide a user experience corresponding to the user in the current environment but with additional information or virtual objects or information added. Thus, VR applications tend to provide a fully immersive synthetically generated world / scenario, while AR applications tend to provide a partially synthetic world / scenario superimposed on the real scenario where the user physically exists. However, the terms are often used interchangeably and have a high degree of overlap. Hereinafter, the term virtual reality / VR will be used to represent both virtual reality and augmented reality.

[0006] As an example, a popular service is to provide images and audio in such a way that the user can actively and dynamically interact with the system to change the presentation parameters so that this will suit the movement and change of the user's position and orientation. An attractive feature in many applications is the ability to change the effective viewing position and viewing direction of the viewer, for example, allowing the viewer to move and "look around" in the scene being presented.

[0007] This feature can particularly allow to provide a virtual reality experience to a user. This can allow the user to move around (relatively) freely in a virtual environment and dynamically change its position and the place it is looking at. Generally, such virtual reality applications are based on a three-dimensional model of a scene, where the model is dynamically evaluated to provide a view for a specific request. For computers and consoles, this method is well-known in, for example, gaming applications (e.g., in the first-person shooter genre).

[0008] Particularly for virtual reality applications, it is also desirable that the presented images are three-dimensional images. In fact, in order to optimize the viewer's immersion, it is generally preferred to make the presented scene of the user experience a three-dimensional scene. In fact, the virtual reality experience should preferably allow the user to choose his / her own position, camera viewpoint, and moment relative to the virtual world.

[0009] In addition to the virtual presentation, most VR / AR applications also provide a corresponding audio experience. In many applications, the audio preferably provides a spatial audio experience, where the audio source is perceived as arriving from a position corresponding to the position of the corresponding object in the virtual scene (including both currently visible objects and currently invisible (e.g., behind the user) objects). Thus, the audio and video scenes are preferably perceived as being consistent, and where both provide a complete spatial experience.

[0010] For audio, until now it has mainly focused on headphone reproduction using binaural audio rendering techniques. In many cases, headphone reproduction achieves a highly immersive and personalized experience for the user. Using head tracking, the rendering can respond to the user's head movements, which highly increases the immersion.

[0011] For the purpose of immersive voice and audio services (IVAS), the 3GPP consortium has developed a so-called IVAS codec (3GPP SP-170611’New WID on EVS Codec Extension for Immersive Voice and Audio Services’). The codec includes a renderer that converts various audio streams into a form suitable for reproduction at the receiving end. In particular, the audio can be presented in a binaural format for reproduction via headphones or a head-mounted VR device with built-in headphones.

[0012] In many such applications, the rendering device can receive input data describing a three-dimensional audio and / or visual scene, and the renderer can be arranged to render the data such that the user is provided with an audiovisual experience that provides a perception of a three-dimensional scene.

[0013] However, it is a challenging expectation to provide a suitable experience in many applications, and in particular to adapt the presentation in response to head movement such that it is challenging to provide the desired experience to the user.

[0014] For example, it is known that human perception of the direction and distance of a sound source is due not only to the (usually different) delays and filtering of the sound from the source to the two ears, but also largely due to how these change as the head moves (such as rotates). Similarly, the parallax and similar movement of visual objects provide strong three-dimensional visual cues. Unconsciously, as we move and sway our heads (usually slightly) in our daily lives, the sound changes in a similarly slight but distinct way, and this significantly increases the immersive "around us" auditory / visual experience to which we are accustomed.

[0015] Experiments using headphone reproduction have shown that even if the sound paths from the sound source to the ears are adequately modeled by filters by keeping these sound paths stationary (i.e., by lacking head movement-related changes), the immersive experience is reduced and the sound may tend to seem "inside our heads".

[0016] Accordingly, some applications have been developed to present audio sources and / or visual objects at positions perceived as fixed relative to the real world in order to create the impression of an immersive virtual world. However, this is a challenging optimal operation and may not always result in the desired user experience. In some applications, a three-dimensional scene that follows head movement and thus appears fixed relative to the user's head is provided. This may be a desired experience in many applications, but may provide an unnatural experience in other applications and may, for example, not allow for an immersive experience of "being in" the virtual scene. US10015620B2 discloses another example of presenting audio relative to a reference orientation representing the user. Other examples of presenting audio and video for virtual experiences can be found in WO20 / 012067A1, WO2019 / 141900A1, and US2019 / 215638.

[0017] However, although such applications can provide a suitable user experience in many embodiments, they tend not to provide an optimal or even desired user experience for some applications.

[0018] Accordingly, an improved method for presenting audiovisual items (particularly for virtual / augmented / mixed reality experiences / applications) would be advantageous. Specifically, a method that allows for improved operation, increased flexibility, reduced complexity, convenient implementation, improved user experience, more consistent perception of the audio and / or visual scene, improved customization, improved personalization; improved virtual reality experience and / or improved performance and / or operation would be advantageous. Summary of the Invention

[0019] Accordingly, the present invention seeks to alleviate, mitigate or eliminate one or more of the disadvantages mentioned above, preferably singly or in any combination.

[0020] According to one aspect of the present invention, there is provided an audio-visual presentation apparatus, comprising: a first receiver arranged to receive an audio-visual item; a metadata receiver arranged to receive metadata, the metadata including an input pose for each of at least some of the audio-visual items and a presentation category flag for each of at least some of the audio-visual items, the input pose being provided with reference to an input coordinate system, and the presentation category flag indicating a presentation category from a set of presentation categories; a receiver arranged to receive user head movement data indicative of a movement of a user's head; a mapper arranged to map the input pose to a presentation pose in a presentation coordinate system in response to the user head movement data, the presentation coordinate system being fixed relative to the head movement; a presenter arranged to present the audio-visual item using the presentation pose; wherein each presentation category is linked to a coordinate system transformation from a real-world coordinate system to a category coordinate system, the coordinate system transformation being different for different categories, and at least one category coordinate system is variable relative to the real-world coordinate system and the presentation coordinate system; and the mapper is arranged to select a first presentation category from the set of presentation categories for the first audio-visual item in response to the presentation category flag for the first audio-visual item, and map the input pose for the first audio-visual item to a presentation pose in the presentation coordinate system, the presentation pose corresponding to a fixed pose in a first category coordinate system for a varying user head movement, the first category coordinate system being determined according to a first coordinate system transformation for the first presentation category.

[0021] In many embodiments, the method can provide an improved user experience and can specifically provide an improved user experience for many virtual reality (including augmented and mixed reality) applications, specifically including social or shared experiences. The method can provide a very flexible approach where the presentation operation and spatial perception of audio-visual items can be individually adapted to individual audio-visual items. The method can, for example, allow some audio-visual items to be presented as appearing completely fixed relative to the real world, some audio-visual items to be presented as appearing completely fixed to the user (following the user's head movement), and some audio-visual items to be presented as appearing fixed for some movements in the real world and following the user for other movements. In many embodiments, the method can allow for flexible presentation where the audio-visual items are perceived as substantially following the user, but still provide an out-of-headspace experience for the audio-visual items.

[0022] This method can reduce complexity and resource requirements in many embodiments and can allow source - side control of rendering operations in many embodiments.

[0023] The rendering category can indicate that the audiovisual item represents an audio source with spatial properties that are fixed with respect to head orientation or not fixed with respect to head orientation (corresponding to listener - pose - related positions and listener - pose - independent positions, respectively). The rendering category can indicate whether the audio element is narrative.

[0024] In many embodiments, the mapper can also be arranged to select a second rendering category from the set of rendering categories for the second audiovisual item in response to a rendering - category flag for the second audiovisual item, and map an input pose for the first audiovisual item to a rendering pose in the rendering coordinate system, the rendering pose corresponding to a fixed pose in a second - category coordinate system for changing user head movements, the second - category coordinate system being determined according to a second coordinate - system transformation for the second rendering category. The mapper can be similarly arranged to perform such operations on third, fourth, fifth, etc. audiovisual items.

[0025] An audiovisual item can be an audio item and / or a visual / video / image / scene item. An audiovisual item can be a visual or audio representation of a scene object in a scene represented by the audiovisual item. In some embodiments, the term audiovisual item can be replaced by the term audio item (or element). In some embodiments, the term audiovisual item can be replaced by the term visual item (or scene object).

[0026] In many embodiments, the renderer can be arranged to generate an output binaural audio signal for a binaural rendering device by applying binaural rendering to an audiovisual item (as an audio item) using the rendering pose.

[0027] The term pose can represent position and / or orientation. In some embodiments, the term "pose" can be replaced by the term "position". In some embodiments, the term "pose" can be replaced by the term "orientation". In some embodiments, the term "pose" can be replaced by the term "position and orientation".

[0028] The receiver can be arranged to receive real - world user head - movement data that references a real - world coordinate system indicating the user's head movement.

[0029] The renderer is arranged to use the rendering pose to render the audiovisual item, wherein the audiovisual item is referenced / located in the rendering coordinate system.

[0030] The mapper may be arranged to map an input pose of the first audiovisual item to a presented pose in the presentation coordinate system, the presented pose corresponding to a fixed pose in a first category coordinate system of the head movement data for different indicated changing user head movements.

[0031] In some embodiments, the presentation category flag indicates a source type, such as an audio type of an audio item or a scene object type of a visual element.

[0032] In many embodiments, this can provide an improved user experience. The presentation category flag may indicate an audio format from a set of audio formats, including at least one audio format from the following group: speech audio; music audio; foreground audio; background audio; narration audio; and narrator audio.

[0033] According to an optional feature of the present invention, a second coordinate system transformation for a second category is such that the category coordinate system of the second category is aligned with the user head movement.

[0034] In many embodiments, this can provide an improved user experience and / or improved performance and / or facilitated implementation. In particular, it can support that some audiovisual items may appear fixed relative to the user's head while other audiovisual items are not.

[0035] According to an optional feature of the present invention, a third coordinate system transformation for a third category is such that the category coordinate system of the third category is aligned with the real-world coordinate system.

[0036] In many embodiments, this can provide an improved user experience and / or improved performance and / or facilitated implementation. In particular, it can support that some audiovisual items may appear fixed relative to the real world while other audiovisual items are not.

[0037] According to an optional feature of the present invention, the first coordinate system transformation depends on the user head movement data.

[0038] In many embodiments, this can provide an improved user experience and / or improved performance and / or facilitated implementation. In particular, in some applications, it can provide a very beneficial experience where the audiovisual item pose follows the overall movement of the user rather than smaller / faster head movements, thus providing an improved user experience with an improved out-of-head experience.

[0039] According to an optional feature of the present invention, wherein the first coordinate system transformation depends on an average head pose.

[0040] According to an optional feature of the present invention, the first coordinate system transformation aligns the first category coordinate system with the average head pose.

[0041] According to an optional feature of the present invention, different coordinate system transformations for different presentation categories depend on the user head movement data, and the dependencies of the first coordinate system transformation and the different coordinate system transformations on the user head movement have different time-average properties.

[0042] According to an optional feature of the present invention, the audiovisual presentation device further includes a receiver for receiving user torso pose data indicating the pose of the user's torso, and the first coordinate system transformation depends on the user torso pose data.

[0043] In many embodiments, this can provide an improved user experience and / or improved performance and / or facilitated implementation. In particular, in some applications, it can provide a very advantageous experience, in which the audiovisual item pose follows the overall movement of the user rather than smaller / faster head movements, thus providing an improved user experience with an improved out-of-head experience.

[0044] According to an optional feature of the present invention, the first coordinate system transformation depends on the average torso pose.

[0045] According to an optional feature of the present invention, the first coordinate system transformation aligns the first category coordinate system with the user torso pose.

[0046] According to an optional feature of the present invention, the audiovisual presentation device further includes a receiver for receiving device pose data indicating the pose of an external device, and the first coordinate system transformation depends on the device pose data.

[0047] In many embodiments, this can provide an improved user experience and / or improved performance and / or facilitated implementation. In particular, in some applications, it can provide a very advantageous experience, in which the audiovisual item pose follows the overall movement of the user rather than smaller / faster head movements, thus providing an improved user experience with an improved out-of-head experience.

[0048] According to an optional feature of the present invention, the first coordinate system transformation depends on the average device pose.

[0049] According to an optional feature of the present invention, the first coordinate system transformation aligns the first category coordinate system with the device pose.

[0050] According to an optional feature of the present invention, the mapper is arranged to select the first presentation category in response to user movement parameters indicating the movement of the user.

[0051] According to an optional feature of the invention, the mapper is arranged to determine a coordinate system transformation between the real-world coordinate system and the coordinate system of the user head movement data in response to user movement parameters indicative of the movement of the user.

[0052] According to an optional feature of the invention, the mapper is arranged to determine the user movement parameters in response to the user head movement data.

[0053] According to an optional feature of the invention, at least some of the presentation category flags are indicative of whether the audiovisual item for the at least some presentation category flags is a narrative audiovisual item or a non-narrative audiovisual item.

[0054] According to an optional feature of the invention, the audiovisual item is an audio item, and the presenter is arranged to generate an output binaural audio signal for a binaural presentation device by applying binaural presentation to the audio item using the presentation pose.

[0055] According to another aspect of the invention, there is provided a method of presenting an audiovisual item, the method comprising: receiving an audiovisual item; receiving metadata including an input pose for each of at least some of the audiovisual items in the audiovisual item and a presentation category flag for each of at least some of the audiovisual items in the audiovisual item, the input pose being provided with reference to an input reference coordinate system, and the presentation category flag indicating a presentation category from a set of presentation categories; receiving user head movement data indicative of movement of the user's head; mapping the input pose to a presentation pose in a presentation coordinate system in response to the user head movement data, the presentation coordinate system being fixed relative to the head movement; presenting the audiovisual item using the presentation pose; wherein each presentation category is linked to a coordinate system transformation from a real-world coordinate system to a category coordinate system, the coordinate system transformation being different for different categories, and at least one category coordinate system is variable relative to the real-world coordinate system and the presentation coordinate system; and the method includes selecting a first presentation category from the set of presentation categories for the first audiovisual item in response to the presentation category flag for the first audiovisual item and mapping the input pose for the first audiovisual item to a presentation pose in the presentation coordinate system, the presentation pose corresponding to a fixed pose in a first category coordinate system for a changing user head movement, the first category coordinate system being determined according to a first coordinate system transformation for the first presentation category.

[0056] These and other aspects, features and advantages of the invention will become apparent with reference to the embodiments described hereinafter and will be elucidated with reference to the embodiments described hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Embodiments of the present invention will be described with reference to the accompanying drawings by way of example only, in which:

[0058] Figure 1 An example of a client-server based virtual reality system is illustrated;

[0059] Figure 2 An example of elements of an audiovisual presentation device according to some embodiments of the present invention is illustrated;

[0060] Figure 3 An example of a possible presentation method of the audiovisual presentation device through Figure 2 is illustrated;

[0061] Figure 4 An example of a possible presentation method of the audiovisual presentation device through Figure 2 is illustrated;

[0062] Figure 5 An example of a possible presentation method of the audiovisual presentation device through Figure 2 is illustrated;

[0063] Figure 6 An example of a possible presentation method of the audiovisual presentation device through Figure 2 is illustrated; and

[0064] Figure 7 An example of a possible presentation method of the audiovisual presentation device through Figure 2 is illustrated. DETAILED DESCRIPTION

[0065] The following description will focus on embodiments in which audiovisual items, including both audio items and visual items, are provided by a presentation that includes both an audio presentation and a visual presentation. However, it should be understood that the methods and principles described can also be applied individually and separately to, for example, the presentation of only audio items or the video / visual / image presentation of only visual items (such as visual objects in a video scene).

[0066] The description will also focus on virtual reality applications, but it will be understood that the methods described can be used in many other applications, including augmented and mixed reality applications.

[0067] Virtual reality (including augmented and mixed reality) experiences that allow users to move around in a virtual or augmented world are becoming increasingly popular, and services that meet such needs are being developed. In many such ways, virtual and audio data can be dynamically generated to reflect the current pose of the user (or viewer).

[0068] In this field, the terms orientation and pose are used as common terms for position and / or direction / orientation. For example, the combination of the position and direction / orientation of an object, a camera, a head, or a view can be referred to as a pose or an orientation. Thus, an orientation or pose indication can include up to six values / components / degrees of freedom, where each value / component typically describes an individual property of the position / location or orientation / direction of the corresponding object. Of course, in many cases, for example, if one or more components are considered fixed or irrelevant, the orientation or pose can be represented by fewer components (e.g., if all objects are considered to be at the same height and have a horizontal orientation, four components can provide a complete representation of the pose of the object). Hereinafter, the term pose is used to refer to a position and / or orientation that can be represented by one to six values (corresponding to the maximum possible degrees of freedom).

[0069] Many VR applications are based on poses with the maximum degrees of freedom, i.e., three degrees of freedom for each of position and orientation, resulting in a total of six degrees of freedom. Thus, a pose can be represented by a set or vector of six values representing these six degrees of freedom, and thus, a pose vector can provide a three-dimensional position and / or a three-dimensional direction indication. However, it should be understood that in other embodiments, a pose can also be represented by fewer values.

[0070] A system or entity that provides the maximum degrees of freedom for a viewer is typically referred to as having six degrees of freedom (6DoF). Many systems and entities provide only orientation or position, and these are typically referred to as having three degrees of freedom (3DoF).

[0071] Typically, virtual reality applications generate three-dimensional output in the form of separate view images for the left and right eyes. These can then be presented to the user via suitable devices, such as the individual left and right eye displays of a typical VR headset. In other embodiments, one or more view images can be presented, for example, on an autostereoscopic display, or in fact, in some embodiments, only a single two-dimensional image can be generated (e.g., using a conventional two-dimensional display).

[0072] Similarly, for a given viewer / user / listener pose, an audio representation of the scene can be provided. The audio scene is typically presented to provide a spatial experience where the audio sources are perceived to originate from the desired locations. Since the audio sources can be static in the scene, a change in the user pose will result in a change in the relative position of the audio sources with respect to the user's pose. Thus, the spatial perception of the audio sources can change to reflect the new position with respect to the user. The audio presentation can be adapted accordingly based on the user pose.

[0073] Viewer or user pose input can be determined in different ways in different applications. In many embodiments, the physical movement of the user can be directly tracked. For example, a camera monitoring the user area can detect and track the user's head (or even eyes (eye tracking)). In many embodiments, the user can wear a VR headset that can be tracked by external and / or internal devices. For example, the headset can include an accelerometer and a gyroscope that provide information about the movement and rotation of the headset and thus the head. In some examples, the VR headset can emit signals or include (e.g., visual) identifiers that enable external sensors to determine the position and orientation of the VR headset.

[0074] In some systems, a VR application can be locally provided to a viewer by a stand-alone device, e.g., without using any remote VR data or processing or even without any access to any remote VR data or processing. For example, a device such as a game console can include a storage device for storing scene data, an input for receiving / generating viewer pose, and a processor for generating corresponding images and ( / or) audio based on the scene data.

[0075] In other systems, VR / scene data can be provided from a remote device or server.

[0076] For example, a remote device can generate audio data representing an audio scene and can transmit audio components / objects / signals or other audio elements corresponding to different audio sources in the audio scene and position information indicating their positions (which can change dynamically, e.g., for moving objects). Audio elements can include elements associated with specific positions, but can also include elements for more distributed or diffuse audio sources. For example, audio elements representing general (non-localized) background sounds, environmental sounds, diffuse reverberation, etc. can be provided.

[0077] The local VR device can then appropriately present the audio elements, e.g., by applying appropriate binaural processing that reflects the relative positions of the audio sources for the audio components.

[0078] Similarly, a remote device can generate visual / video data representing a visual / video scene and can transmit visual scene components / objects / signals or other visual elements corresponding to different objects in the visual scene and position information indicating their positions (which can change dynamically, e.g., for moving objects). Visual items can include elements associated with specific positions, but can also include video items for more distributed sources.

[0079] In some embodiments, visual items may be provided as individual and separate items, such as, for example, descriptions of individual scene objects (e.g., size, texture, opacity, reflectivity, etc.). Alternatively or additionally, visual items may be represented as part of an overall model of the scene, e.g., including descriptions of different objects and their relationships to each other.

[0080] For VR services, in some embodiments, a central server may correspondingly generate audiovisual data representing a three-dimensional scene, and may specifically represent audio through a plurality of audio items that can be presented by a local client / device and represent the visual scene through a plurality of video items that can be presented by a local client / device.

[0081] Figure 1 An example of a VR system is illustrated, where a central server 101 communicates with a plurality of remote clients 103 via a network 105 (e.g., the Internet), for example. The central server 101 may be arranged to support a potentially large number of remote clients 103 simultaneously.

[0082] In many cases, this approach can provide an improved trade-off, for example, between complexity and the resource requirements (communication requirements, etc.) of different devices. For example, the scene data may be transmitted only once or relatively infrequently, where the local rendering device (remote client 103) receives the viewer's pose and processes the scene data locally to render audio and / or video to reflect changes in the viewer's pose. This approach can provide an efficient system and an attractive user experience. It can, for example, substantially reduce the required communication bandwidth while providing a low-latency real-time experience, while allowing the scene data to be stored, generated, and maintained centrally. It can, for example, be suitable for applications where VR experiences are provided to multiple remote devices.

[0083] Figure 2 Elements of an audiovisual presentation device that can provide improved audiovisual presentations in many applications and scenarios are illustrated. Specifically, the audiovisual presentation device can provide improved presentations for many VR applications, and the audiovisual presentation device can specifically be arranged to perform processing and rendering for Figure 1 the VR client 103.

[0084] Figure 2 The audio device of is arranged to present a three-dimensional scene by presenting spatial audio and video to provide a three-dimensional perception of the scene. The specific description of the audiovisual presentation device will focus on applications that provide the described method to both audio and video, but it should be understood that in other embodiments, the method may be applied only to audio or video / visual processing, and in fact, in some embodiments, the presentation device may include only functions for presenting audio or only functions for presenting video, i.e., the audiovisual presentation device may be any audio presentation device or any video presentation device.

[0085] The audio-visual presentation device includes a first receiver 201, which is arranged to receive audio-visual items from a local or remote source. In a specific example, the first receiver 201 receives data describing the audio-visual item from a server 101. The first receiver 201 may be arranged to receive data describing a virtual scene. The data may include data providing a visual description of the scene and may include data providing an audio description of the scene. Thus, the audio scene description and the virtual scene description may be provided by the received data.

[0086] The audio item may be encoded audio data, such as an encoded audio signal. The audio item may be different types of audio elements including different types of signals and components, and in fact in many embodiments, the first receiver 201 may receive audio data defining different types / formats of audio. For example, the audio data may include audio represented by audio channel signals, individual audio objects, scene-based audio (such as higher-order ambisonics (HOA), etc.). The audio may be represented, for example, as encoded audio of a given audio component to be presented.

[0087] The first receiver 201 is coupled to a renderer 203, which continues to present the scene based on the received data describing the audio-visual item. In the case of encoded data, the renderer 203 may also be arranged to decode the data (or in some embodiments, the decoding may be performed by the first receiver 201).

[0088] Specifically, the renderer 203 may include an image renderer 205, which is arranged to generate an image corresponding to the current viewing pose of the viewer. For example, the data may include spatial 3D image data (such as an image of the scene and depth or model description), and accordingly, the visual renderer 203 may generate stereoscopic images (images for the user's left and right eyes), as would be known to those skilled in the art. The image may be presented to the user, for example, via the individual left and right eye displays of a VR headset.

[0089] The renderer 203 also includes an audio renderer 207, which is arranged to present an audio scene by generating an audio signal based on the audio item. In this example, the audio renderer 207 is a binaural audio renderer that generates binaural audio signals for the user's left and right ears. The binaural audio signals are generated to provide a desired spatial experience and are typically reproduced by a headset or headphones that may specifically be part of a headset worn by the user, which also includes a left eye display and a right eye display.

[0090] Thus, in many embodiments, the audio rendering by the audio renderer 207 is a binaural rendering process that uses an appropriate binaural transfer function to provide the desired spatial effect for the user wearing the earphones. For example, the audio renderer 207 may be arranged to generate audio components that are perceived to arrive from a specific location using binaural processing.

[0091] Binaural processing is known to provide a spatial experience by using virtual localization of sound sources with individual signals for the listener's ears. With appropriate binaural rendering processing, the signals required at the eardrums to make the listener perceive that the sound is coming from any desired direction can be calculated, and the signals can be rendered such that they provide the desired effect. These signals are then re-generated at the eardrums using either headphones or a crosstalk cancellation method (suitable for rendering over closely spaced loudspeakers). Binaural rendering can be considered a method for generating signals for the listener's ears so as to trick the human auditory system into perceiving that the sound is coming from a desired location.

[0092] Binaural rendering is based on binaural transfer functions, which vary from person to person due to the acoustic properties of the head, ears, and reflective surfaces such as shoulders. For example, binaural filters can be used to create binaural recordings simulating multiple sources at various positions. This can be achieved by convolving each sound source with, for example, a head-related impulse response (HRIR) pair corresponding to the position of the sound source.

[0093] A well-known method for determining the binaural transfer function is binaural recording. It is a method of recording sounds using a specialized microphone arrangement and is intended for replay using headphones. The recording is made by placing microphones in the ear canals of an object or using a mannequin head with embedded microphones, a bust including the pinna (outer ear). The use of such a mannequin head including the pinna provides a spatial impression very similar to that of a person listening to the recording being present during the recording.

[0094] By measuring the response of a microphone placed in or near the human ear to a sound source at a specific location in, for example, 2D or 3D space, an appropriate binaural filter can be determined. Based on such measurements, a binaural filter reflecting the acoustic transfer function to the user's ears can be generated. The binaural filter can be used to create binaural recordings simulating multiple sources at various positions. This can be achieved, for example, by convolving each sound source with the measured impulse response pair for the desired position of the sound source. To create the illusion of a sound source moving around the listener, a large number of binaural filters with a certain spatial resolution, such as 10 degrees, are typically required.

[0095] The head-related binaural transfer function can be represented, for example, as a head-related impulse response (HRIR), or equivalently a head-related transfer function (HRTF), or a binaural room impulse response (BRIR), or a binaural room transfer function (BRTF). The transfer function (e.g., estimated or assumed) from a given location to the listener's ears (or eardrums) can be given, for example, in the frequency domain, in which case it is commonly referred to as an HRTF or BRTF; or in the time domain, in which case it is commonly referred to as an HRIR or BRIR. In some scenarios, the head-related binaural transfer function is determined to include aspects or properties of the acoustic environment and in particular the room in which the measurements are made, while in other examples only user characteristics are considered. Examples of the first type of function are BRIR and BRTF.

[0096] The audio renderer 207 can accordingly include a storage device having binaural transfer functions for a usually high number of different locations, where each binaural transfer function provides information on how an audio signal should be processed / filtered so as to be perceived as originating from that location. Applying binaural processing individually to multiple audio signals / sources and combining the results can be used to generate an audio scene with multiple audio sources located at appropriate positions in the sound field.

[0097] The audio renderer 207 can select and retrieve the stored binaural transfer function that most closely matches the desired location (or in some cases can be generated by interpolating between multiple neighboring binaural transfer functions) for a given audio element to be perceived as originating from a given location relative to the user's head. It can then apply the selected binaural transfer function to the audio signal of the audio element, thereby generating an audio signal for the left ear and an audio signal for the right ear.

[0098] The output stereo signal in the form of the generated left and right ear signals is suitable for headphone presentation and can be amplified to generate a drive signal that is fed to the user's headset. The user will then perceive the audio element as originating from the desired location.

[0099] It should be appreciated that in some embodiments, the audio item can also be processed to, for example, add acoustic environment effects. For example, the audio item can be processed to add reverberation or, for example, decorrelation / diffusivity. In many embodiments, such processing can be performed on the generated binaural signals rather than directly on the audio element signals.

[0100] Thus, the audio renderer 207 can be arranged to generate an audio signal such that a given audio element is presented such that a user wearing headphones perceives the audio element as being received from the desired location. Other audio items can, for example, possibly be distributed and diffuse and can be presented as such.

[0101] It should be appreciated that many algorithms and methods for presenting spatial audio, for example using headphones and specifically for binaural presentation, will be known to those skilled in the art, and any suitable method can be used without departing from the present invention.

[0102] The audiovisual presentation device further includes a second receiver 209, which is a metadata receiver arranged to receive metadata for the audiovisual item. The metadata specifically includes location data for one or more of the audiovisual items. The metadata may include an input location indicating the location of one or more of the audiovisual items.

[0103] The received audiovisual data may include audio and / or visual data describing the scene. The audiovisual data specifically includes audiovisual data for a set of audiovisual items corresponding to audio sources and / or visual objects in the scene. Some audio items may represent localized audio sources in the scene associated with a specific location and / or orientation in the scene (where the location and / or orientation may change dynamically for a moving object). The visual data may include data describing the scene objects, thereby allowing visual representations of these to be generated and represented in the image / video presented to the user (usually a 3D image using a separate display of the headset).

[0104] Generally, an audio element may represent audio generated by a specific scene object in the virtual scene, and thus may represent an audio source at a location corresponding to the location of the scene object (e.g., human speech). In this case, the same location data / indicator may be included and used for the audio item and the corresponding visual scene object (and similarly for orientation).

[0105] Other elements may represent more distributed or diffuse audio sources, such as, for example, ambient or background noise that may be diffuse. As another example, some audio elements may represent, in whole or in part, the non-spatial localized component of audio from a localized audio source, such as, for example, the diffuse reverberation from a spatially well-defined audio source.

[0106] Similarly, some visual scene objects may have an extended location, and for example, the location data may indicate the center or reference location of the scene object.

[0107] The metadata may include pose data that indicates the location and / or orientation of the audiovisual item, and specifically indicates the location and / or orientation of the audio source and / or visual scene object or element. The pose data may include, for example, absolute location and / or orientation data defining the location of each or at least some of the items.

[0108] The pose is provided with reference to an input coordinate system, i.e., the input coordinate system is the reference coordinate system for the pose indication provided in the metadata received by the second receiver. The input coordinate system is typically a coordinate system fixed with reference to the represented / presented scene. For example, the scene can be a virtual (or real) scene where the audio source and scene objects are at positions provided with reference to the scene, i.e., the input coordinate system is typically the scene coordinate system of the scene represented by the audiovisual data.

[0109] To provide a user representation of the scene, the scene will be presented from the viewer's or user's pose, i.e., the scene will be presented as it is perceived at a given viewer / user pose in the scene, and where the audio and visual presentation provides the audio and images that would be perceived for that viewer pose.

[0110] The presentation by the renderer 203 is performed with respect to a presentation coordinate system that is fixed with respect to the user's head and head movement. The reproduction of the presented audio signal is typically a head-mounted or mounted reproduction device, such as head-mounted headphones / earphones and an individual display for the eyes. Typically, the reproduction is by a head-mounted kit device that includes audio and video reproduction means. The renderer is arranged to generate the presented audiovisual items with reference to the reproduction device / means, and it is assumed that the user's pose with reference to the reproduction device / means is constant / fixed. For example, the position of the audio source is determined with respect to the head-mounted headphones (i.e., the position in the presentation coordinate system), and the appropriate HRTF filter for that position is retrieved and used to present the audio signal such that the audio source is perceived as arriving from the desired relative position in the presentation coordinate system. Similarly, for a given scene object, the relative position with respect to the display is determined (i.e., the position in the presentation coordinate system), and the image corresponds to views from the left-eye pose and the eye pose respectively with respect to that position.

[0111] Thus, the presentation coordinate system can be considered as a coordinate system fixed to the user's head, and specifically as a presentation coordinate system independent of head movement or actually changes in the user's pose. The reproduction device (whether audio, visual, or both audio and visual) is assumed / considered to be fixed with respect to the user's head and thus fixed with respect to the presentation coordinate system.

[0112] The presentation coordinate system fixed with respect to user head movement can be considered to correspond to a reproduction coordinate system fixed with respect to the reproduction device used to reproduce the presented audiovisual items. The term presentation coordinate system can be equivalent to the reproduction device / means coordinate system and can thus be replaced. Similarly, the term "presentation coordinate system fixed with respect to user head movement" can be equivalent to "reproduction device / means coordinate system fixed with respect to the reproduction device used to reproduce the presented audiovisual items" and can be replaced by it.

[0113] When a renderer refers to a rendering coordinate system to perform rendering based on a pose and provides a pose of an audiovisual item with reference to an input coordinate system, the audiovisual rendering device includes a mapper 211 that is arranged to map an input position in the input coordinate system to a rendering position in the rendering coordinate system.

[0114] The audiovisual rendering device includes a head movement data receiver 213 for receiving user head movement data indicating a movement of a user's head. The user head movement data may indicate a movement of the user's head in the real world and is typically provided with reference to a real world coordinate system. The user head movement data may indicate an absolute or relative movement of the user's head in the real world and may specifically reflect an absolute or relative change in the user's pose with respect to the real world. The coordinate system head movement data may indicate a change (or no change) in the head pose (orientation and / or position) and may also be referred to as head pose data.

[0115] It should be understood that many different possible methods for detecting and representing head movement are known and any suitable method may be used without departing from the present invention. The head movement data receiver 213 may specifically receive the head movement data from a VR headset or a VR head movement detector, as known in the art.

[0116] The mapper 211 is coupled to the head movement data receiver 213 and receives the user head movement data. The mapper is arranged to perform the mapping between the input position in the input coordinate system and the rendering position in the rendering coordinate system in response to the user head movement data. For example, the mapper 211 may continuously process the user head movement data to continuously track the current user pose in the real world coordinate system. Then, the mapping between the input pose and the rendering pose may be based on the user pose.

[0117] For example, in many applications, it is desired to provide the user with an experience as if he were present in the three-dimensional scene being represented. Therefore, it is desired that the presented audio and images reflect the user pose following the movement of the user's head. Therefore, it is desired that the audiovisual items be presented such that they are perceived as being fixed with respect to the real world, as this enables the reproduction of real world movement in the presentation of the (usually virtual) scene.

[0118] In such a case, the mapping from the input pose to the rendered pose is such that the audiovisual items appear fixed relative to the real world, i.e., they are rendered as being perceived as fixed relative to the real world. Thus, the same input pose is mapped to different rendered poses to reflect changes in the user's head pose. For example, if the user turns his head, say, by 30°, the real-world scene is referenced to the user's rotation of -30°. The mapper 211 can perform the corresponding change such that the mapping from the input pose to the rendered pose is modified to include an additional 30° rotation relative to the situation before the user's head rotation. Thus, the audiovisual items will be in different poses in the rendered coordinate system but will be perceived as being in the same real-world pose. Thus, the mapping can be changed dynamically such that the audiovisual items are perceived as fixed relative to the real world and thus provide a very natural experience.

[0119] For example, to create the illusion of an immersive virtual world, three-dimensional audio and / or visual rendering is typically controlled via head tracking, where the rendering is compensated for head pose, specifically including changes in head orientation (such as 3 spatial degrees of freedom or 3-DOF, such as yaw, pitch, roll). The rendering is such that the audiovisual items are perceived as fixed relative to the user. Compared to static rendering, the effect of this head tracking and subsequent rendering adaptation is a high sense of realism and extra-diegetic perception of the rendered content.

[0120] However, another approach is to map the input pose to a rendered pose that is fixed relative to the rendered coordinate system. This can be done, for example, by the mapper 211 applying a fixed mapping from the input pose to the rendered pose, where the mapping is independent of head movement data, and specifically, where changes in the user's pose do not cause a change in the mapping between the input pose and the rendered pose. The effect of such a mapping is effectively to make the perceived scene move with the head, i.e., it is static relative to the user's head. Although this may seem unnatural for most scenarios, it can be advantageous in some scenarios. For example, it can provide a desired experience for music or for listening to sounds that are not part of the scene (such as a narrator, for example).

[0121] Different methods can be applied to different audiovisual items. In MPEG terminology, the terms "fixed with respect to head orientation" or "not fixed with respect to head orientation" are used to refer to audio items that are to be rendered to either fully follow or ignore user movement.

[0122] For example, an audio item may be considered "head-independent", which means that it is an audio element designed to have a fixed position in a (virtual or real) environment, and thus its rendering dynamically adapts to the user's head orientation (changes). Another audio item may be considered "head-fixed", which means that it is an audio item designed to have a fixed position relative to the user's head. Such an audio item can be presented independently of the listener's pose. Thus, the rendering of such an audio item does not take into account the user's head orientation (changes), in other words, such an audio item is an audio element whose relative position does not change when the user turns their head (e.g., non-spatial audio such as ambient noise or music designed to follow the user without changing relative position).

[0123] In the described system, the second receiver 209 is arranged to receive metadata that also includes a rendering category flag for at least some of the audiovisual items. The rendering category flag indicates a rendering category from a set of rendering categories, and the rendering of the audiovisual item is performed according to the rendering category indicated for the audiovisual item. Different rendering categories may define different rendering parameters and operations.

[0124] The rendering category flag can be any indication that can be used to select a rendering category from a set of rendering categories. In many embodiments, it can be data provided only for the purpose of selecting a rendering category and / or can be data that directly specifies a category. In other embodiments, the rendering category flag can be an indication that can also provide additional information or provide some description of the corresponding audiovisual item. In some embodiments, the rendering category flag can be a parameter considered when selecting a rendering category, and other parameters can also be considered.

[0125] As a specific example, in some embodiments, an audio item may be encoded audio data, such as an encoded audio signal, where the audio item may be different types of audio items including different types of signals and components, and in fact, in many embodiments, the metadata receiver 201 can receive metadata that defines different types / formats of audio. For example, the audio data may include audio represented by audio channel signals, individual audio objects, higher-order ambisonics (HOA), etc. The metadata can be included as part of the audio item or separately from the audio item that describes the audio type of each audio item. This metadata can be a rendering category flag and can be used to select an appropriate rendering category for the audio item.

[0126] Rendering categories are specifically associated with different reference coordinate systems, and specifically, each rendering category is linked to a coordinate system transformation from the real-world coordinate system to the category coordinate system. Specifically, for each category, a coordinate system transformation can be defined that transforms the real-world coordinate system (such as specifically, the real-world coordinate system to which the head movement data is referenced) into a different coordinate system given by the transformation. Since different categories have different coordinate system transformations, they will be linked to different category reference systems.

[0127] The coordinate system transformation is typically a dynamic coordinate system transformation for one, some, or all categories. Thus, the coordinate system transformation is typically not a fixed or static coordinate system transformation, but can vary over time and depend on different parameters. For example, as will be described in more detail later, the coordinate system transformation can depend on parameters that change dynamically, such as for example user torso movement, external device movement, and / or actually even head movement data. Thus, in many embodiments, the coordinate system transformation of a category is a time-varying coordinate system transformation that depends on user movement parameters. The user movement parameters can indicate the movement of the user relative to the real-world coordinate system.

[0128] The mapping performed by the mapper 211 for a given audiovisual item depends on the category coordinate system of the rendering category to which the audiovisual item is indicated as belonging. Specifically, based on the rendering category flag, the mapper 211 can determine the rendering category intended for presenting the audiovisual item. Then, the mapper can determine the coordinate system transformation linked to the selected category. Then, the mapper 211 can continue to perform the mapping from the input pose to the rendering pose such that these correspond to fixed poses in the category coordinate system resulting from the selected coordinate system transformation.

[0129] Thus, the category coordinate system can be considered as a reference coordinate system with respect to which the audiovisual item is presented as being fixed. The category coordinate system can also be referred to as a reference coordinate system or a fixed reference coordinate system (for a given category).

[0130] In many embodiments, one rendering category can correspond to a rendering where the audio source and scene objects represented by the audiovisual item are fixed relative to the real world as previously described. In such embodiments, the coordinate system transformation is such that the category coordinate system of the category is aligned with the real-world coordinate system. For such a category, the coordinate system transformation can be a fixed coordinate system transformation and can be, for example, a unity one-to-one mapping of the real-world coordinate system. Thus, the category coordinate system can effectively be the real-world coordinate system or, for example, a fixed static translation, scale, and / or rotation.

[0131] In many embodiments, a rendering category may correspond to a rendering where the audio source and scene objects represented by the audiovisual item are fixed relative to head movement (i.e., relative to the rendering coordinate system). In such an embodiment, the coordinate system is transformed such that the category coordinate system of the category aligns with the user's head / reproduction device / rendering coordinate system. For such a category, the coordinate system transformation may be a coordinate system transformation that fully follows the head movement of the user. For example, any rotation of the head is followed by a corresponding rotation in the coordinate system transformation, and any change in the position of the user's head is followed by the same change in the coordinate system transformation. Thus, according to such a rendering category, the coordinate system transformation is dynamically modified to follow the head movement data such that the resulting category coordinate system aligns with the rendering coordinate system, thereby resulting in a fixed mapping from the input coordinate system to the rendering coordinate system as previously described.

[0132] Although not required, in many embodiments, the rendering category may correspondingly include a category where the audiovisual item is rendered as fixed relative to the real-world coordinate system and a category where the audiovisual item is rendered as fixed relative to the rendering coordinate system. However, in the described system, one or more of the rendering categories include a rendering category where the listening item is fixed in a coordinate system that is neither the real-world coordinate system nor the rendering coordinate system, i.e., a rendering category that provides a rendering of the audiovisual item that is neither fixed in the real world nor fixed relative to the user's head. Thus, at least one category coordinate system is variable relative to the real-world coordinate system and the rendering coordinate system. Specifically, the coordinate system is different from the rendering coordinate system and the real-world coordinate system, and in fact the differences between these and the category coordinate system are not constant but can change.

[0133] Thus, at least one rendering category may provide a rendering that is neither fixed to the real world nor fixed to the user. More precisely, in many embodiments, it may provide an intermediate experience.

[0134] For example, the coordinate system transformation can be such that the corresponding coordinate system is fixed relative to the real world, except when an update criterion is met. However, if the criterion is met, the coordinate system transformation can be adapted to provide different relationships between the real world coordinate system and the category coordinate system. For example, the mapping can be such that the audiovisual is presented as fixed relative to the real world coordinate system, i.e., the audiovisual item appears to be in a fixed position. However, if the user rotates their head by more than a given amount, the relationship between the real world coordinate system and the category coordinate system is changed to compensate for the rotation. For example, as long as the user's head moves less than, say, 20°, the audiovisual item is presented such that it is in a fixed position. However, if the user moves their head by more than 20°, the category coordinate system is rotated 20° relative to the real world coordinate system. This can provide the experience that, as long as the movement is small enough, the user perceives a natural three-dimensional experience of the presented audiovisual item. However, for large head movements, the presentation of the audiovisual item is re-aligned with the modified head position.

[0135] As a specific example, the audio source corresponding to the narrator can initially be presented as being directly in front of the user. For small user movements, the audio is presented such that the narrator is perceived as being stationary in the same position. This provides a natural experience and perception, and in particular provides an out-of-head perception of the narrator. However, if the user rotates their head by more than, for example, 20° from the original direction towards the narrator audio source, the system adapts the mapping to re-position the narrator audio source in front of the user's new orientation. For small movements around this point, the narrator audio source is presented at this new fixed position (relative to the real world coordinate system). If the movement again exceeds a given threshold with respect to this new audio source position, the update of the category coordinate system and thus the perceived position of the narrator audio source can be performed again. Thus, a narrator can be provided for the user who, for smaller movements, is fixed relative to the real world, but for larger movements, follows the user. In the example described, it can allow the narrator to be perceived as being fixed and provide appropriate spatial cues regarding head movement, but always substantially in front of the user (even if, for example, the user turns a full 180°).

[0136] In this method, the metadata includes presentation category identifiers for a plurality of audiovisual items, thereby allowing the source side to control flexible presentation at the receiving side, where the presentation is specifically adapted to individual audiovisual items. For example, different spatial presentations and perceptions can be applied to items corresponding to, for example, background music, narration, audio sources corresponding to specific objects fixed in the scene, conversations, etc.

[0137] In some embodiments, the coordinate system transformation for the presentation category depends on user head movement data. Thus, in some embodiments, changing at least one parameter of the coordinate system transformation depends on user head movement data.

[0138] In many embodiments, the coordinate system transformation may depend on user head pose attributes or parameters determined from user head movement data. For example, as previously described, if the user head pose indicates a rotation beyond a certain amount, the coordinate system transformation may be adapted to include a rotation corresponding to that amount. As another example, the mapper 211 may detect that the user has maintained a (sufficiently) constant pose for longer than a given duration, and if so, the coordinate system transformation may be adapted to position the audiovisual item in a given position in the coordinate system, i.e., having a particular position relative to the user (e.g., directly in front of the user).

[0139] In some embodiments, the coordinate system transformation depends on the average head pose. In particular, in some embodiments, the coordinate system transformation may be such that the category coordinate system is aligned with the average head pose. In some embodiments, the coordinate system transformation may be such that the category coordinate system is fixed relative to the average head pose.

[0140] The average head pose may be determined, for example, by low-pass filtering the head pose measurements by a low-pass filter with a suitable cut-off frequency (such as specifically by applying an unweighted average over a window of a suitable duration).

[0141] In some embodiments, for one or more presentation categories, the reference for presentation may be selected to be the average head orientation h. This has the effect that the audiovisual item follows slower, longer-lasting head orientation changes, such that the sound source appears to remain at the same point relative to the head (e.g., in front of the face), but rapid head movements will cause the audiovisual item to appear fixed relative to the real world (and thus also in the virtual world) rather than relative to the head. Thus, typical small and rapid head movements during daily life will still produce an immersive and out-of-head illusion while still allowing an overall perception of the audiovisual item following the user.

[0142] In some embodiments, the adjustment and tracking may be made non-linear such that if the head rotates significantly, the average head orientation reference is, for example, "clipped" so as not to deviate from the instantaneous head orientation by more than a certain maximum angle. For example, if the maximum value is 20 degrees, an out-of-head experience is achieved as long as the head "swings" within those + / - 20 degrees. If the head rotates rapidly and exceeds the maximum value, the reference will follow the head orientation (with a maximum 20-degree lag), and once the movement stops, the reference becomes stable again.

[0143] In some embodiments, at least one presentation category is associated with a coordinate system transformation that depends on the user's torso pose. In such embodiments, the audiovisual presentation device may include a torso pose receiver 215 arranged to receive user torso pose data indicative of the user's torso pose.

[0144] The torso pose may be determined, for example, by a dedicated inertial sensor unit positioned or worn on the torso. As another example, the torso pose may be determined by sensors in a smart device (such as a smartphone) when it is worn in a pocket. As yet another example, coils may be placed on the user's torso and head respectively, and the movement of the head relative to the torso may be determined based on changes in the coupling between these.

[0145] In such embodiments, changing at least one parameter of the coordinate system transformation depends on the torso pose data.

[0146] In many embodiments, the coordinate system transformation may depend on torso pose data properties or parameters determined from the torso pose data.

[0147] For example, if the torso pose data indicates that the torso has rotated by more than a certain amount, the coordinate system transformation may be adapted to include a rotation corresponding to that amount. As another example, the mapper 211 may detect that the user has maintained a (sufficiently) constant torso pose for longer than a given duration, and if so, the coordinate system transformation may be adapted to position the audiovisual item in a given position in the presentation coordinate system, where that position corresponds to the torso pose.

[0148] In particular, in some embodiments, the coordinate system transformation may align the category coordinate system with the user's torso pose. In some embodiments, the coordinate system transformation may be such that the category coordinate system is fixed relative to the user's torso pose. Thus, in some embodiments, the audiovisual item may be presented as following the user's torso, and thus may provide a perception and experience where the audiovisual item follows the user's movement of their entire body, but appears fixed with respect to head movement relative to the torso. This may provide both the desired experience of exocentric perception and the presentation of an audiovisual item that follows the user.

[0149] In some embodiments, the coordinate system transformation may depend on the average torso pose. The average torso pose may be determined, for example, by low-pass filtering head pose measurements with a low-pass filter having a suitable cut-off frequency (such as specifically by applying an unweighted average over a window of a suitable duration).

[0150] Thus, in some embodiments, one or more of the presentation categories may employ a coordinate system transformation that is aligned with the instantaneous or average chest / torso orientation tAligned presentation provides a reference. In this way, the audiovisual item can appear to stay in front of the user's body rather than in front of the user's face. By rotating the head relative to the chest / trunk, the audiovisual item can still be perceived from all directions, greatly increasing the immersive out-of-head experience again.

[0151] In some embodiments, the coordinate system transformation for the presentation category depends on the device pose data indicating the pose of the external device. In such embodiments, the audiovisual presentation device may include a device pose receiver 217, which is arranged to receive the device pose data indicating the device torso pose.

[0152] The device can be, for example, (hypothetically) a device worn, carried, attached to, or otherwise fixed relative to the user. In many embodiments, the external device can be, for example, a mobile phone or a personal device, such as a smartphone in a pocket, a body-mounted device, or a handheld device (e.g., a smart device for viewing visual VR content).

[0153] Many devices include gyroscopes, accelerometers, GPS receivers, etc., which allow the relative or absolute orientation of the device. The device can then determine the current relative or absolute orientation and send it to the device pose receiver 217 using a suitable communication (which is usually wireless). For example, the communication can be via a WiFi or Bluetooth connection.

[0154] In such embodiments, changing at least one parameter of the coordinate system transformation depends on the device pose data.

[0155] In many embodiments, the coordinate system transformation can depend on the device pose data attributes or parameters determined from the device pose data.

[0156] For example, if the device pose data pose indicates that the device has rotated by more than a certain amount, the coordinate system transformation can be adapted to include a rotation corresponding to that amount. As another example, the mapper 211 can detect that the device has maintained a (sufficiently) constant torso pose for longer than a given duration, and if so, the coordinate system transformation can be adapted to position the audiovisual item at a given position in the presentation coordinate system, where the position corresponds to the device pose.

[0157] In particular, in some embodiments, the coordinate system transformation can be such that the category coordinate system is aligned with the device pose. In some embodiments, the coordinate system transformation can be such that the category coordinate system is fixed relative to the device pose. Thus, in some embodiments, the audiovisual item can be presented to follow the device pose, and thus a perception and experience of the audiovisual item following the device movement can be provided. In many actual user scenarios, the device can provide a good indication of the user pose. For example, a body-worn device or a smartphone in a pocket, for instance, can provide a good reflection of the overall movement of the user. It can provide a good reference for determining relative head movement, and thus can provide an experience that combines a realistic response to head movement while allowing the audiovisual item to follow the greater movement of the user.

[0158] Furthermore, using an external device as a reference can be highly practical and provide a reference that results in a desired user experience. The method can be based on a device that is typically already worn or carried by the user and includes the required functionality for determining and transmitting the device pose. For example, most people currently carry a smartphone that already includes an accelerometer, etc. for determining the device pose and a communication device (such as Bluetooth) suitable for transmitting the device pose data to the audiovisual presentation device.

[0159] In some embodiments, the coordinate system transformation can depend on the average device pose. The average device pose can be determined, for example, by low-pass filtering the device pose measurements by a low-pass filter with a suitable cut-off frequency (such as specifically by applying an unweighted average over a window of a suitable duration).

[0160] Thus, in some embodiments, one or more in the presentation can employ a coordinate system transformation that provides a reference for a presentation aligned with the instantaneous or average device orientation. In this way, the audiovisual item can appear to remain in a fixed position relative to the device, such that it moves when the device moves, but remains fixed relative to head movement, thus providing a more natural feel and an immersive out-of-head experience.

[0161] Thus, the method can provide a method in which metadata can be used to control the presentation of the audiovisual items such that these can be individually controlled to provide different user experiences for different audiovisual items. The experiences include one or more options that provide the perception of an audiovisual item that is neither completely fixed in the real world nor completely follows the user (for head fixation). Specifically, an intermediate experience can be provided in which the audiovisual item is presented to be fixed relative to the real world to some extent and follows the movement of the user to some extent.

[0162] It should be understood that in some embodiments, possible presentation categories may be determined in advance, where each category is associated with a predetermined coordinate system transformation. In such embodiments, the audiovisual presentation device may store the coordinate system transformation for each category, or equivalently, directly store the mapping corresponding to the coordinate system transformation when appropriate. The mapper 211 may be arranged to retrieve the stored coordinate system transformation (or mapping) for the selected presentation category and apply the coordinate system transformation (or mapping) when performing the mapping for the audiovisual item.

[0163] For example, for a first audiovisual item, the presentation category identifier may indicate that it should be presented according to a first category. The first category may be a category for presenting an audiovisual item fixed to a head pose, and thus the mapper may retrieve a mapping that provides a fixed one-to-one mapping between the input pose and the presentation pose. For a second audiovisual item, the presentation category identifier may indicate that it should be presented according to a second category. The second category may be a category for presenting an audiovisual item fixed to the real world, and thus the mapper may retrieve a coordinate system transformation that adapts the mapping such that head movement is compensated, resulting in a presentation pose corresponding to a fixed position in real space. For a third audiovisual item, the presentation category identifier may indicate that it should be presented according to a third category. The third category may be a category for presenting an audiovisual item fixed to a device pose or a torso pose. The mapper may retrieve a coordinate system transformation or mapping that adapts the mapping such that head movement relative to the device or torso pose is compensated, resulting in the presentation of an audiovisual item fixed relative to the device or torso pose.

[0164] It should be understood that although different presentation categories are associated with different coordinate system transformations such that the mapping results in a presentation position fixed relative to the category coordinate system, the mapper 211 does not need to explicitly determine such a coordinate system transformation or the category coordinate system. Instead, in a typical embodiment, a mapping function is defined for an individual presentation category such that the resulting presentation pose is fixed relative to the category coordinate system. For example, a mapping function that is a function of the head pose relative to the device (or torso) pose may be used to directly map the input position to a presentation position fixed relative to the category coordinate system, which is fixed relative to the device (or torso) pose.

[0165] In some embodiments, metadata may include data that partially or fully characterizes, describes, and / or defines one or more of a plurality of presentation categories. For example, in addition to a presentation category flag, the metadata may include data describing a coordinate system transformation and / or mapping function to be applied to one or more categories. For example, the metadata may indicate that a first presentation category requires a fixed mapping from an input location to a presentation location, a second category requires a mapping that fully compensates for head movement such that items appear fixed relative to the real world, and a third presentation category that should compensate the mapping for head movement relative to an average head movement such that an intermediate experience is perceived, where for smaller and faster head movements, the items appear fixed, but for slow average movements, also appear to follow the user.

[0166] In different embodiments, different methods and data may be used as presentation category flags. In some embodiments, each category may be associated with a category number, for example, and the presentation category flag may directly provide the number of the category to be used for the audiovisual item.

[0167] In many embodiments, the presentation category flag may indicate an attribute or characteristic of the audiovisual item, and this may be mapped to a specific presentation category flag.

[0168] In some embodiments, the presentation category flag may specifically indicate whether the audiovisual item is a narrative audiovisual item or a non-narrative audiovisual item. A narrative audiovisual item may be an audiovisual item belonging to a scene of a movie or story being presented; in other words, a narrative audiovisual item originates from a source within a movie, story, etc. (e.g., an actor in a script, a bird with its sound in a nature movie, etc.). A non-narrative audiovisual item may be an item originating outside of the movie or story (e.g., an audio commentary by the director, mood music, etc.). In many cases, according to MPEG terminology, a narrative audiovisual item may correspond to "not fixed with respect to head orientation", and a non-narrative audiovisual item may correspond to "fixed with respect to head orientation".

[0169] In some embodiments, there may actually be only two presentation categories, and specifically, one may correspond to the presentation of audiovisual items indicated as narrative, and one may correspond to the presentation of audiovisual items indicated as non-narrative.

[0170] For example, in some applications and systems, narrative signaling may be sent downstream to an audiovisual presentation device, and the desired presentation behavior may depend on this signaling, as exemplified with respect to Figure 3 as illustrated:

[0171] · By using head tracking with a real-world orientation as a reference, it is desired that the diegetic sound source D appears to remain stable in its position in the virtual world V and should thus be presented as being fixed relative to the real world. In a given example of a movie application, if the actor's voice is presented directly in front of the user and the user rotates their head 50 degrees to the left, the sound will be presented to the right headphone projected 50 degrees, so that it appears to remain in the same virtual position.

[0172] · Non-diegetic sound sources N can alternatively be presented independently of head orientation. In other words, the audio remains in a "hard-coupled" fixed position relative to the head (e.g., in front of it) and rotates with the head. This is achieved by not applying a head-orientation-related mapping to the audio but using a fixed mapping from the input position to the presentation position. In the movie example, the director's commentary audio can be presented precisely in front of the user and any head movement will have no effect on it, i.e., the sound remains in front of the head.

[0173] Figure 2 The audiovisual presentation device is arranged to provide a more flexible method, where at least one optional presentation category allows the presentation of audiovisual items that allows the presentation of audiovisual items to be fixed relative to the real / virtual world for some movements and to follow the user for other movements. The method can be specifically applied to non-diegetic audiovisual items.

[0174] In a specific example, alternative or additional options for presenting non-diegetic sound sources can include one or more of the following:

[0175] · The reference for presentation can be chosen to be the average head orientation / pose h (as Figure 4 shown). This has the following effect: non-diegetic sound sources follow slower and longer-lasting head-orientation changes, such that the sound source appears to remain at the same point relative to the head (e.g., in front of the face), but rapid head movements will cause the non-diegetic sound source to appear fixed in the virtual world (rather than fixed relative to the head). Thus, typical small and rapid head movements during daily life will still produce an immersive and out-of-head illusion.

[0176] As an improvement in some embodiments, since it is generally desired that non-diegetic audio remains at least closely in the same virtual position, the tracking can be made non-linear such that if the head rotates significantly, the average head-orientation reference is, for example, "clipped"

[0177] Deviate by no more than a certain maximum angle relative to the instantaneous head orientation. For example, if the maximum value is 20 degrees, as long as the head "swings" within those + / - 20 degrees, an out-of-head experience is achieved. If the head rotates rapidly and exceeds the maximum value, the reference will follow the head orientation (with a maximum 20-degree lag), and once the movement stops, the reference becomes stable again.

[0178] · The reference for head tracking can be selected to be the instantaneous or average chest / trunk orientation / pose t (as Figure 5 shown). In this way, non-diegetic content will appear to stay in front of the user's body rather than in front of the user's face. By rotating the head relative to the chest / trunk, non-diegetic content can still be heard from all directions, greatly increasing the immersive out-of-head experience once again.

[0179] · The reference for head tracking can be selected to be the instantaneous or average orientation / pose of an external device (such as a mobile phone or a body-worn device). In this way, non-diegetic content will appear to stay in front of the device rather than in front of the user's face. By rotating the head relative to the device, non-diegetic content can still be heard from all directions, greatly increasing the immersive out-of-head experience once again.

[0180] In some embodiments, the presentation can be arranged to operate in different modes according to user movement. For example, if the user movement meets the movement criteria, the audiovisual presentation device can operate in a first mode, and if the user movement does not meet the movement criteria, the audiovisual presentation device can operate in a second mode. In this example, the two modes can provide different category coordinate systems for the same presentation category flag, that is, depending on the user movement, the audiovisual presentation device can present a given audiovisual item using a given presentation category flag fixed with reference to different coordinate systems.

[0181] In some embodiments, the mapper 211 can be arranged to select the presentation category of a given presentation category flag in response to user movement parameters indicating the user's movement. Thus, different presentation categories can be selected for a given audiovisual item and presentation category flag value according to the user movement parameters. Specifically, if the user movement parameters meet a first criterion, a given link between the possible presentation category flag values and the set of presentation categories can be used to select the presentation category of the received presentation category flag. However, if the criterion is not met (or, for example, a different criterion is met), the mapper 211 can use a different link between the possible presentation category flag values and the same or a different set of presentation categories to select the presentation category for the received presentation category flag.

[0182] The method can, for example, enable the audiovisual presentation device to provide different presentations and experiences for mobile users and stationary users.

[0183] In some embodiments, the selection of the presented category may also depend on other parameters, such as for example user settings or configuration settings of an application (e.g., an app on a mobile device).

[0184] In another approach, the mapper may be arranged to determine a coordinate system transformation between a real-world coordinate system that is a reference for a coordinate system transformation of the selected category in response to user movement parameters indicating the movement of the user and the coordinate system with reference to which user head movement data is provided.

[0185] Thus, in some embodiments, an adaptation of the presentation may be introduced by a coordinate system transformation of the presented category with respect to a reference, which may vary with respect to a (typically real-world) coordinate system indicating head movement. For example, based on user movement, some compensation may be applied to the head movement data, e.g., to compensate for some movement of the user as a whole. As an example, if the user is, for example, on a boat, the user head movement data may not only indicate the movement of the user relative to the body or the body relative to the boat, but may also reflect the movement of the boat. This may be undesirable, and thus the mapper 211 may compensate the head movement data for user movement parameter data that reflects the component of user movement caused by the movement of the boat. The resulting modified / compensated head movement data is then given with respect to a coordinate system in which the movement of the boat has been compensated, and the coordinate system transformation of the selected category may be directly applied to this modified coordinate system to achieve a desired presentation and user experience.

[0186] It should be understood that the user movement parameters may be determined in any suitable manner. For example, in some embodiments, it may be determined by a dedicated sensor providing relevant data. For example, an accelerometer and / or a gyroscope may be attached to a vehicle carrying the user, such as a car or a boat.

[0187] In many embodiments, the mapper may be arranged to determine the user movement parameters in response to the user head movement data itself. For example, a long-term average or a base movement analysis identifying, for example, periodic components (e.g., corresponding to the waves of a moving boat) may be used to determine user parameters indicating the user movement parameters.

[0188] In some embodiments, the coordinate system transformations of two different presented categories may depend on the user head movement data, but wherein the dependence has different time-averaging properties. For example, one presented category may be associated with a coordinate system transformation that depends on the average head movement but has a relatively low average time (i.e., a relatively high cut-off frequency of an average low-pass filter), while the other presented category is also associated with a coordinate system transformation that depends on the average head movement but has a higher average time (i.e., a relatively low cut-off frequency of an average low-pass filter).

[0189] As an example, the mapper 211 can evaluate the head movement data against criteria that reflect the consideration that the user is in a static environment. For example, the low-pass filtered position change can be compared to a given threshold, and if it is below the threshold, the user can be considered to be in a static environment. In this case, the rendering can be the same as in the specific examples described above for narrative and non-narrative audio items. However, if the position change is above the threshold, the user can be considered to be in a moving environment, such as during a stroll, or using a vehicle such as a car, train, or plane. In this case, the following rendering methods can be applied (see also Figure 6 and 7 ):

[0190] · The average head orientation h can be used as a reference to render the non-narrative sound source N, to keep the non-narrative sound source mainly in front of the face, but still allow small head movements to create an "out-of-head" experience. In other examples, for instance, the user's torso or device pose can be used as a reference. Thus, this method can correspond to the method that can be used for narrative sound sources in a static situation.

[0191] · The average head orientation h can also be used as a reference but with a further offset corresponding to a longer-term average head orientation h ' to render the narrative sound source D. Thus, in this case, the narrative virtual sound source will appear to be in a fixed position relative to the virtual world V, but the virtual world V will appear to "move with the user", i.e., it will remain in approximately the same orientation relative to the user. The averaging (or other filtering process) used to obtain h ' is typically chosen to be much slower than the averaging used for h . Thus, the non-narrative source will follow faster head movements, while the narrative content (and the virtual world V as a whole) will take some time to re-orient with the user's head.

[0192] It will be appreciated that, for clarity, the above description has described embodiments of the invention with reference to different functional circuits, units, and processors. However, it will be apparent that any suitable distribution of functions between different functional circuits, units, or processors can be used without departing from the invention. For example, functions illustrated as being performed by separate processors or controllers can be performed by the same processor. Thus, the reference to a particular functional unit or circuit is only to be regarded as a reference to a suitable device for providing the described function, and does not indicate a strict logical or physical structure or organization.

[0193] The present invention can be implemented in any suitable form, including hardware, software, firmware, or any combination thereof. Optionally, the present invention may be at least partially implemented as computer software running on one or more data processors and / or digital signal processors. The elements and components of embodiments of the present invention can be physically, functionally, and logically implemented in any suitable manner. In fact, the functions can be implemented in a single unit, in multiple units, or as part of other functional units. Thus, the present invention can be implemented in a single unit, or can be physically and functionally distributed among different units, circuits, and processors.

[0194] Although the present invention has been described in connection with some embodiments, it is not intended to limit the present invention to the specific forms set forth herein. On the contrary, the scope of the present invention is limited only by the claims. Additionally, although it may seem that features are described in connection with specific embodiments, those skilled in the art will recognize that the various features of the described embodiments can be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.

[0195] Furthermore, although listed separately, multiple devices, elements, circuits, or method steps can be implemented by, for example, a single circuit, unit, or processor. Additionally, although the individual features may be included in different claims, these features can be advantageously combined, and inclusion in different claims does not mean that the combination of features is infeasible and / or disadvantageous. Inclusion of a feature in a class of claims does not mean a limitation to that class, but rather indicates that the feature is equally applicable to other claim classes when appropriate. Moreover, the order of features in the claims does not imply any particular order in which the features must operate, and in particular, the order of the individual steps in method claims does not mean that the steps must be performed in that order. Rather, the steps can be performed in any suitable order. Additionally, singular references do not exclude pluralities. Thus, references to "a", "an", "first", "second", etc. do not exclude pluralities. The reference numerals in the claims are provided merely to make the examples clear and should not be construed as limiting the scope of the claims in any way.

Claims

1. An audiovisual presentation device, comprising: A first receiver (201) arranged to receive audiovisual items; A metadata receiver (209) arranged to receive metadata, the metadata including an input pose for each of at least some of the audiovisual items in the audiovisual items and a presentation category flag for each of at least some of the audiovisual items in the audiovisual items, the input pose being provided with reference to an input coordinate system, and the presentation category flag indicating a presentation category from a set of presentation categories; A receiver (213) arranged to receive user head movement data indicating the movement of a user's head; A mapper (211) arranged to map the input pose to a presentation pose in a presentation coordinate system in response to the user head movement data, the presentation coordinate system being fixed relative to the head movement; A presenter (203) arranged to present the audiovisual items using the presentation pose; Wherein, At least one of the audiovisual items represents one or more distributed or decentralized audio sources; Each presentation category is linked to a coordinate system transformation from a real-world coordinate system to a category coordinate system, the coordinate system transformation being different for different presentation categories, and at least one category coordinate system being variable relative to the real-world coordinate system and the presentation coordinate system; and The mapper is arranged to select a first presentation category for the first audiovisual item from the set of presentation categories in response to the presentation category flag for the first audiovisual item, and map the input pose for the first audiovisual item to a presentation pose in the presentation coordinate system, the presentation pose corresponding to a fixed pose in a first category coordinate system for a changing user head movement, the first category coordinate system being determined according to a first coordinate system transformation for the first presentation category.

2. The audiovisual presentation device according to claim 1, wherein, A second coordinate system transformation for a second category causes the category coordinate system for the second category to align with the user head movement.

3. The audiovisual presentation device according to claim 1 or 2, wherein, A third coordinate system transformation for a third category causes the category coordinate system for the third category to align with the real-world coordinate system.

4. The audiovisual presentation device according to any one of the preceding claims, wherein, The first coordinate system transformation depends on the user head movement data.

5. The audiovisual presentation device according to claim 4, wherein, The first coordinate system transformation depends on an average head pose.

6. The audiovisual presentation device according to claim 5, wherein, The first coordinate system transformation aligns the first category coordinate system with the average head pose.

7. The audiovisual presentation device according to any one of claims 4-6, wherein, The different coordinate system transformations for different presentation categories depend on the user head movement data, and the dependence of the first coordinate system transformation and the different coordinate system transformations on the user head movement has different time-average properties.

8. The audiovisual presentation device according to any one of the preceding claims, further comprising a receiver (215) arranged to receive user torso pose data indicating the pose of a user's torso, and the first coordinate system transformation depends on the user torso pose data.

9. The audiovisual presentation apparatus according to any one of the preceding claims, further comprising a receiver (217), the receiver being arranged to receive device pose data indicating the pose of an external device, and the first coordinate system transformation depending on the device pose data.

10. The audiovisual presentation device according to any one of the preceding claims, wherein, The mapper (211) is arranged to select the first presentation category in response to user movement parameters indicating movement of the user.

11. The audiovisual presentation device according to any one of the preceding claims, wherein, The mapper (211) is arranged to determine a coordinate system transformation between the real-world coordinate system and the coordinate system of the user head movement data in response to user movement parameters indicating movement of the user.

12. The audiovisual presentation device according to claim 10 or 11, wherein, The mapper (211) is arranged to determine the user movement parameters in response to the user head movement data.

13. The audiovisual presentation device according to any one of the preceding claims, wherein, At least some of the presentation category flags indicate whether the audiovisual item for the at least some of the presentation category flags is a narrative audiovisual item or a non-narrative audiovisual item.

14. The audiovisual presentation device according to any one of the preceding claims, wherein, The audiovisual item is an audio item, and the presenter (211) is arranged to generate an output binaural audio signal for a binaural presentation device by applying a binaural presentation to the audio item using the presentation pose.

15. A method of presenting an audiovisual item, the method comprising: Receiving an audiovisual item; Receiving metadata including an input pose for each of at least some of the audiovisual items in the audiovisual item and a presentation category flag for each of at least some of the audiovisual items in the audiovisual item, the input pose being provided with reference to an input coordinate system, and the presentation category flag indicating a presentation category from a set of presentation categories; Receiving user head movement data indicating movement of the user's head; Mapping the input pose to a presentation pose in a presentation coordinate system that is fixed relative to the head movement in response to the user head movement data; Presenting the audiovisual item using the presentation pose; Wherein, At least one of the audiovisual items represents one or more distributed or decentralized audio sources; Each presentation category is linked to a coordinate system transformation from a real-world coordinate system to a category coordinate system, the coordinate system transformation being different for different presentation categories, and at least one category coordinate system being variable relative to the real-world coordinate system and the presentation coordinate system; and The method includes: selecting a first presentation category for the first audiovisual item from the set of presentation categories in response to the presentation category flag for the first audiovisual item, and mapping the input pose for the first audiovisual item to a presentation pose in the presentation coordinate system, the presentation pose corresponding to a fixed pose in a first category coordinate system for varying user head movement, the first category coordinate system being determined according to a first coordinate system transformation for the first presentation category.

Citation Information

Patent Citations

  • Head tracking

    US10015620B2

  • Near-field binaural rendering

    US20190215638A1

  • Associated spatial audio playback

    WO2019141900A1

  • Spatial audio augmentation

    WO2020012067A1