Audiovisual rendering device and its operation method

The visual rendering apparatus and method address the challenge of adapting audiovisual rendering to user head movements by transforming poses into rendering poses, enhancing immersion and consistency in virtual and augmented reality applications.

JP7714030B2Active Publication Date: 2025-07-28KONINKLIJKE PHILIPS NV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023522375
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-10-13
Filing Date
2021-10-11
Publication Date
2025-07-28
Estimated Expiration
2041-10-11

AI Technical Summary

Technical Problem

Existing audiovisual rendering technologies struggle to adapt to user head movements, leading to suboptimal immersive experiences in virtual and augmented reality applications, particularly in maintaining consistent spatial audio and visual cues.

Method used

A visual rendering apparatus and method that utilizes a mapper to transform input poses into rendering poses based on different coordinate systems, allowing flexible rendering categories that adjust to user head movements, ensuring some items appear fixed relative to the real world or the user's head, and others track head movements.

Benefits of technology

Enhances user experience by providing flexible and consistent spatial audiovisual rendering, reducing complexity and resource requirements, and improving immersion in virtual and augmented reality applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007714030000001
    Figure 0007714030000001
  • Figure 0007714030000002
    Figure 0007714030000002
  • Figure 0007714030000003
    Figure 0007714030000003
Patent Text Reader

Abstract

The audiovisual rendering device includes a receiver 201 that receives an audiovisual item and a receiver 209 that receives metadata including an input pose relative to an input coordinate system and a rendering category indication indicating a rendering category. A receiver 213 receives user head movement data, and a mapper 211, in response to the user head movement data, maps the input pose to a rendering pose in the rendering coordinate system. A renderer 203 renders the audiovisual item using the rendering pose. Each rendering category is linked to a different coordinate system transformation from a real-world coordinate system to a category coordinate system, and at least one category coordinate system is variable with respect to the real-world coordinate system and the rendering coordinate system. The mapper, in response to the rendering category indication, selects a rendering category for the audiovisual item and maps the input pose to a rendering pose corresponding to a fixed pose in the category coordinate system for varying user head movement. The category coordinate system is determined from the coordinate system transformation of the rendering category.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to audiovisual rendering devices and methods of operation and particularly, but not exclusively, to their use to support augmented / virtual reality applications and the like. [Background technology]

[0002] In recent years, the variety and range of experiences based on audiovisual content has grown significantly, and new services and methods for using and consuming such content are continually being developed and introduced. In particular, many spatially interactive services, applications, and experiences have been developed to provide users with more engaging and immersive experiences.

[0003] Examples of such applications include virtual reality (VR), augmented reality (AR), and mixed reality (MR) applications, which are rapidly becoming mainstream, with many solutions targeting the consumer market. Also, various standards are being developed by various standards bodies. Such standardization activities are actively developing standards for various aspects of VR / AR / MR systems, e.g., streaming, broadcasting, and rendering.

[0004] VR applications tend to provide a user experience that corresponds to the user being in a different world / environment / scene, while AR (including mixed reality MR) applications tend to provide a user experience that corresponds to the user being in a current environment with additional information or virtual objects or information added. Thus, VR applications tend to provide fully immersive artificial worlds / scenes, while AR applications tend to provide partially artificial worlds / scenes overlaid on a real scene in which the user is physically present. However, these terms are often used interchangeably and overlap significantly. In the following, the term virtual reality / VR will be used to refer to both virtual reality and augmented reality.

[0005] As an example, what is becoming increasingly popular is a service that provides images and sounds so that users can interact with the system actively and dynamically in order to change the rendering parameters to adapt to changes in the movement, position, and orientation of the user. In many applications, functions that change the viewer's de facto viewpoint and line of sight, for example, functions that enable the viewer to move within the presented scene and "look around", are very attractive.

[0006] Such functions can, in particular, provide users with a virtual reality experience. This can enable the user to move around (relatively) freely within the virtual environment and dynamically change their position and line of sight. Usually, such virtual reality applications are based on a three-dimensional model of the scene, and the model is dynamically evaluated to provide a specific required view. This approach is well known in game applications such as the first-person shooting game genre for computers and consoles.

[0007] Also, especially in the case of virtual reality applications, it is desirable for the presented image to be a three-dimensional image. In practice, in order to optimize the viewer's sense of immersion, it is usually preferable for the user to experience the presented scene as a three-dimensional scene. In fact, a virtual reality experience preferably enables the user to select their position, the camera's viewpoint, and the time point with respect to the virtual world.

[0008] In addition to visual rendering, most VR / AR applications further provide a corresponding audio experience. In many applications, the audio preferably provides a spatial audio experience where the sound source is recognized as arriving from a position corresponding to the position of the corresponding object (including both the object currently in view and the object not currently in view (e.g., behind the user)) within the visual scene. Therefore, the audio scene and the video scene should be consistent and preferably be perceived as both providing a complete spatial experience.

[0009] Regarding audio, conventionally, focus has been on headphone playback using binaural audio rendering technology. In many scenarios, headphone playback can provide users with a very immersive and personalized experience. Using head tracking enables rendering according to the movement of the user's head, significantly improving the immersion.

[0010] For IVAS (Immersive Voice and Audio Services), the 3GPP consortium is developing a so-called IVAS codec (3GPP SP-170611 'New WID on EVS Codec Extension for Immersive Voice and Audio Services'). This codec includes a renderer that converts various audio streams into a format suitable for playback on the receiving side. Specifically, for playback using headphones or a head-mounted VR device with built-in headphones, the audio can be made in binaural format.

[0011] In many such applications, the rendering device may receive input data representing three-dimensional audio and / or visual scenes. The renderer may be configured to render this data so that an audiovisual experience providing a perception of the three-dimensional scene is provided to the user.

[0012] However, providing an appropriate experience is difficult in many applications, and in particular, it is difficult to adapt the rendering according to the movement of the head so that a desirable experience is provided to the user.

[0013] For example, human perception of the direction and distance of a sound source depends not only on the (usually different) delays and filtering of the sound from the sound source to both ears, but also greatly on how these change when the head is moved, for example rotated. Similarly, the parallax and similar movements of visual objects provide strong three-dimensional visual cues. Unconsciously, we move our heads (usually slightly) or jostle them slightly in our daily lives, so that the sound also changes, albeit slightly but clearly, which greatly contributes to the immersive "around us" auditory / visual experience we are accustomed to.

[0014] In headphone playback experiments, even if the sound path from the sound source to the ears is properly modeled by filters, making these static (i.e., by having no changes related to head movement) can reduce the sense of immersion and the sound may be felt to be "in the head".

[0015] Therefore, several applications have been developed that render sound sources and / or visual objects at positions that are perceived to be fixed relative to the real world in order to create the impression of an immersive virtual world. However, it is difficult to execute this optimally and it does not always result in the desired user experience. In some applications, a three-dimensional scene that follows the movement of the head is displayed, thus appearing to be fixed relative to the user's head. This may be a desirable experience in many applications, but in other applications it may provide an unnatural experience and, for example, may not achieve the immersive experience of "being present" within the virtual scene. US10015620B2 discloses another example where sound is rendered with respect to a reference orientation representing the user.

[0016] However, while such applications may provide an appropriate user experience in many embodiments, some applications tend not to provide an optimal, and thus desirable, user experience.

[0017] Therefore, an improved method for rendering audiovisual items, particularly audiovisual items for virtual / augmented / composite reality experiences / applications, would be beneficial. In particular, methods that enable improved performance, increased flexibility, reduced complexity, easier implementation, improved user experience, more consistent perception of audio and / or visual scenes, improved customization, improved personalization, improved virtual reality experience, and / or improved performance and / or operation would be beneficial. SUMMARY OF THE INVENTION

[0018] Accordingly, the present invention aims to suitably mitigate, reduce, or eliminate one or more of the above drawbacks, either alone or in any combination.

[0019] According to one aspect of the present invention, there is provided a visual rendering apparatus comprising: a first receiver for receiving audiovisual items; a metadata receiver for receiving metadata including an input pose and a rendering category display for each of at least some of the audiovisual items, wherein the input pose is provided with reference to an input coordinate system and the rendering category display indicates a rendering category from a set of rendering categories; a receiver for receiving user head movement data indicating movement of the user's head; a mapper for mapping the input pose to a rendering pose in a rendering coordinate system in response to the user head movement data, wherein the rendering coordinate system is fixed relative to the movement of the head; and a renderer for rendering the audiovisual items using the rendering pose, wherein each rendering category is linked to a coordinate system transformation from a real-world coordinate system to a category coordinate system, the coordinate system transformation being different for each category, at least one category coordinate system being variable with respect to the real-world coordinate system and the rendering coordinate system, and the mapper selects a first rendering category of a first audiovisual item from the set of rendering categories in response to the rendering category display of the first audiovisual item, and maps the input pose of the first audiovisual item to a rendering pose in the rendering coordinate system corresponding to a fixed pose in a first category coordinate system for changing the movement of the user's head, the first category coordinate system being determined from a first coordinate system transformation of the first rendering category.

[0020] This approach can provide an improved user experience in many embodiments, specifically improving the user experience of many virtual reality (including augmented and mixed reality) applications, particularly those involving social or shared experiences. This approach can provide a very flexible approach where the rendering operation and spatial recognition of visual items can be individually adapted to individual visual items. This approach can, for example, render some visual items such that they appear to be completely fixed with respect to the real world, render some visual items such that they appear to be completely fixed to the user (tracking the user's head movements), and render some visual items such that they appear to be fixed with respect to the real world for some movements and tracking the user for other movements. This approach can, in many embodiments, enable flexible rendering where visual items are perceived to substantially track the user, yet still provide a spatial extra-cranial localization experience of the visual items.

[0021] This approach reduces complexity and resource requirements in many embodiments and enables source-side control of the rendering operation in many embodiments.

[0022] The rendering category can indicate whether the visual item represents a sound source with spatial characteristics fixed to the head orientation or not fixed to the head orientation (corresponding to listener pose-dependent position and listener pose-independent position, respectively). The rendering category can indicate whether the audio element is diegetic or not.

[0023] In many embodiments, the mapper may further select, in response to a rendering category display of a second audiovisual item, a second rendering category of the second audiovisual item from a set of rendering categories, and map an input pose of the second audiovisual item to a rendering pose in a rendering coordinate system corresponding to a fixed pose in a second category coordinate system for changing the movement of the user's head, where the second category coordinate system may be determined from a second coordinate system transformation of the second rendering category. The mapper may similarly be configured to perform such an operation on third, fourth, fifth, and other audiovisual items.

[0024] The audiovisual item may be an audio item and / or a visual / video / image / scene item. The audiovisual item may be a visual or audio representation of a scene object of a scene represented by the audiovisual item. In some embodiments, the term "audiovisual item" may be replaced with the term "audio item (or element)". In some embodiments, the term "audiovisual item" may be replaced with the term "visual item (or scene object)".

[0025] In many embodiments, the renderer may be configured to generate an output binaural audio signal for a binaural rendering device by applying binaural rendering to an audiovisual item (which is an audio item) using the rendering pose.

[0026] The term "pose" may represent position and / or orientation. In some embodiments, the term "pose" may be replaced with the term "position". In some embodiments, the term "pose" may be replaced with the term "orientation". In some embodiments, the term "pose" may be replaced with the term "position and orientation".

[0027] The receiver may receive real-world user head movement data based on a real-world coordinate system indicating the movement of the user's head.

[0028] The renderer may render the visual item using a rendering pose, and the visual item is referenced within / arranged within a rendering coordinate system.

[0029] The mapper may map the input pose of the first visual item to a rendering pose within a rendering coordinate system corresponding to a fixed pose within a first category coordinate system of head movement data indicating different changing head movements of the user.

[0030] In some embodiments, the rendering category display indicates a source type, e.g., voice type for a voice item or scene object type for a visual element.

[0031] This may result in an improved user experience in many embodiments. The rendering category display may indicate a certain sound source type among a sound source type set including at least one sound source type selected from the group of uttered audio, music audio, foreground audio, background audio, voice-over audio, and narrator audio.

[0032] According to an optional feature of the present invention, the second coordinate system transformation of the second category aligns the category coordinate system of the second category with the movement of the user's head.

[0033] This may result in an improved user experience and / or improved performance and / or easier implementation in many embodiments. In particular, it may support that some visual items are felt to be fixed relative to the user's head while other items are not.

[0034] According to an optional feature of the present invention, the third coordinate system transformation of the third category aligns the category coordinate system of the third category with the real-world coordinate system.

[0035] This can lead to an improvement in user experience and / or performance and / or ease of implementation in many embodiments. In particular, it can support that some visual items may seem to be fixed with respect to the real world, while other items may not seem so.

[0036] According to an optional feature of the present invention, the first coordinate system transformation depends on user head movement data.

[0037] This can lead to an improvement in user experience and / or performance and / or ease of implementation in many embodiments. In particular, in some applications, it provides a very advantageous experience where the pose of a visual item follows the overall movement of the user but not small / fast movements of the head, thereby providing an improved user experience along with an improved extra-head localization experience.

[0038] According to an optional feature of the present invention, the first coordinate system transformation depends on the average head pose.

[0039] According to an optional feature of the present invention, the first coordinate system transformation aligns the first category coordinate system with the average head pose.

[0040] According to an optional feature of the present invention, different coordinate system transformations for different rendering categories depend on user head movement data, and the dependencies of the first coordinate system transformation and the different coordinate system transformations on the user's head movement have different time-averaging characteristics.

[0041] According to an optional feature of the present invention, the visual rendering device further comprises a receiver for receiving user body pose data indicating the user body pose, and the first coordinate system transformation depends on the user body pose data.

[0042] This can result in an improvement in user experience and / or an improvement in performance and / or an ease of implementation in many embodiments. In particular, in some applications, it provides a very advantageous experience where the pose of the visual item follows the overall movement of the user but not small / fast movements of the head, thereby providing an improved user experience along with an improved out-of-head localization experience.

[0043] According to an optional feature of the present invention, the first coordinate system transformation depends on the average torso pose.

[0044] According to an optional feature of the present invention, the first coordinate system transformation aligns the first category coordinate system with the user torso pose.

[0045] According to an optional feature of the present invention, the visual rendering device further comprises a receiver that receives device pose data indicating the pose of an external device, and the first coordinate system transformation depends on the device pose data.

[0046] This can result in an improvement in user experience and / or an improvement in performance and / or an ease of implementation in many embodiments. In particular, in some applications, it provides a very advantageous experience where the pose of the visual item follows the overall movement of the user but not small / fast movements of the head, thereby providing an improved user experience along with an improved out-of-head localization experience.

[0047] According to an optional feature of the present invention, the first coordinate system transformation depends on the average device pose.

[0048] According to an optional feature of the present invention, the first coordinate system transformation aligns the first category coordinate system with the device pose.

[0049] According to an optional feature of the present invention, the mapper selects a first rendering category in response to user movement parameters indicating the movement of the user.

[0050] According to an optional feature of the present invention, the mapper determines a coordinate system transformation between the real-world coordinate system and the coordinate system of the user head movement data in response to user movement parameters indicating the movement of the user.

[0051] According to an optional feature of the present invention, the mapper determines user movement parameters in response to data of user head movement.

[0052] According to an optional feature of the present invention, at least some of the rendering category displays indicate whether the visual items of at least some of the rendering category displays are digital visual items or non-digital visual items.

[0053] According to an optional feature of the present invention, the visual item is an audio item, and the renderer generates an output binaural audio signal for a binaural rendering device by applying binaural rendering to the audio item using a rendering pose.

[0054] According to another aspect of the present invention, there is provided a method of rendering a visual item, the method comprising receiving a visual item, and receiving metadata including an input pose and a rendering category display for each of at least a portion of the visual item, wherein the input pose is provided with reference to an input coordinate system, and the rendering category display indicates a rendering category from a set of rendering categories, receiving user head movement data indicating movement of the user's head, and in response to the user head movement data, mapping the input pose to a rendering pose within a rendering coordinate system, wherein the rendering coordinate system is fixed relative to the movement of the head, rendering the visual item using the rendering pose, each rendering category being linked to a coordinate system transformation from a real-world coordinate system to a category coordinate system, the coordinate system transformation being different for each category, and at least one category coordinate system being variable relative to the real-world coordinate system and the rendering coordinate system, the method comprising selecting a first rendering category of a first visual item from the set of rendering categories in response to the rendering category display of the first visual item, and mapping the input pose of the first visual item to a rendering pose within the rendering coordinate system corresponding to a fixed pose within a first category coordinate system for changing the movement of the user's head, the first category coordinate system being determined from a first coordinate system transformation of the first rendering category.

[0055] The above and other aspects, features, and advantages of the present invention will be described and will become apparent with reference to the embodiments described hereinafter.

Brief Description of the Drawings

[0056] Hereinafter, embodiments which are merely examples of the present invention will be described with reference to the following drawings.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

DETAILED DESCRIPTION OF THE INVENTION

[0057] The following description focuses on embodiments in which visual items, including both audio items and visual items, are presented by rendering that includes both audio rendering and visual rendering. However, it should be understood that the methods and principles described may be applied separately, for example, to only the audio rendering of audio items or only the video / visual / image rendering of visual items (e.g., visual objects within a video scene).

[0058] Also, although the description focuses on virtual reality applications, it will be understood that the methods described can be used in many other applications, including augmented and mixed reality applications.

[0059] Virtual reality (including augmented and mixed reality) experiences in which a user can move around within a virtual or augmented world are becoming increasingly popular, and services have been developed to meet such demand. In many such methods, visual and audio data can be dynamically generated to reflect the current pose of the user (or viewer).

[0060] In this technical field, the terms "placement" and "pose" are used as general terms representing position and / or orientation / direction. For example, a combination of the position and orientation / direction of an object, a camera, a head, or a view can be called a pose or a placement. Therefore, an indicator of a placement or a pose can have up to six values / components / degrees of freedom. Each value / component typically describes an individual characteristic of the position / location or orientation / direction of the corresponding object. Of course, in many situations, a placement or a pose may be represented by fewer components, for example, when one or more components are considered constant or irrelevant (e.g., when all objects are considered to have the same height and a horizontal orientation, it may be possible to fully represent the pose of an object with four components). Hereinafter, the term "pose" is used to refer to a position and / or orientation that can be represented by one to six values (corresponding to the maximum possible degrees of freedom).

[0061] Many VR applications are based on a pose having the maximum degrees of freedom (i.e., three degrees of freedom each for position and orientation, for a total of six degrees of freedom). Therefore, a pose can be represented by a set or vector of six values representing six degrees of freedom, and thus, a pose vector can provide an indicator of three-dimensional position and / or three-dimensional orientation. However, it should be understood that in other embodiments, a pose can be represented by a smaller number of values.

[0062] A system or entity based on providing the viewer with the maximum degrees of freedom is typically said to have six degrees of freedom (6DoF). Many systems and entities provide only orientation or position, and these are typically said to have three degrees of freedom (3DoF).

[0063] Typically, a virtual reality application generates a three-dimensional output in the form of separate view images for the left and right eyes. These images can then be presented to the user by appropriate means, such as the (usually individual) left and right displays of a VR headset. In other embodiments, one or more view images may be presented, for example, on a naked-eye stereoscopic display, or in some embodiments, only one two-dimensional image may be generated (e.g., using a conventional two-dimensional display).

[0064] Similarly, for a given viewer / user / listener pose, an audio representation of the scene can be provided. Typically, the audio scene is rendered to provide a spatial experience where the sound source is perceived to be emanating from a desired location. Since the sound source may be static within the scene, as the user's pose changes, the relative position of the sound source with respect to the user's pose changes. Thus, the spatial perception of the sound source can change to reflect the new position with respect to the user. Accordingly, the audio rendering can be adapted according to the user's pose.

[0065] Viewer or user pose input can be determined in various ways in various applications. In many embodiments, the user's physical movement can be directly tracked. For example, a camera surveying the user area can detect and track the user's head (or eyes (eye tracking)). In many embodiments, the user can wear a VR headset that can be tracked by external and / or internal means. For example, the headset can include an accelerometer and a gyroscope that provide information about the movement and rotation of the headset (and thus the head). In some examples, the VR headset may transmit a signal or include an identifier (e.g., visual) that enables an external sensor to determine the position and orientation of the VR headset.

[0066] In some systems, VR applications can be provided locally to viewers by a stand-alone device that, for example, does not use (or in some cases does not even have access to) any remote VR data or processing. For example, a device such as a game console can include a store for storing scene data, an input device for receiving / generating the viewer's pose, and a processor for generating corresponding images and / or audio from the scene data.

[0067] In other systems, VR / scene data can be provided from a remote device or server.

[0068] For example, a remote device can generate audio data representing an audio scene and transmit audio components / objects / signals corresponding to a plurality of different sound sources within the audio scene, or other audio elements, along with position information indicating their positions (which can change dynamically in the case of moving objects). The audio elements can include elements associated with a specific position, but can also include elements for more distributed or diffuse sound sources. For example, audio elements representing general (non-localized) background sound, ambient sound, reverberation, etc. can be provided.

[0069] In that case, the local VR device can appropriately render the audio elements, for example, by applying appropriate binaural processing that reflects the relative positions of the sound sources of the audio components.

[0070] Similarly, a remote device can generate visual / video data representing a visual scene and transmit visual components / objects / signals corresponding to a plurality of different objects within the visual scene, or other visual elements, along with position information indicating their positions (which can change dynamically in the case of moving objects). The visual items can include elements associated with a specific position, but can also include video items for more distributed sources.

[0071] In some embodiments, the visual items may be provided as individual discrete items, for example, provided as descriptions of individual scene objects (e.g., dimensions, textures, opacities, reflectivities). Alternatively or additionally, the visual items may be represented as part of an overall model of the scene, for example, may include descriptions of multiple different objects and their interrelationships.

[0072] Thus, in the case of a VR service, the central server in some embodiments can generate audiovisual data representing a three-dimensional scene, specifically, can represent sound by a plurality of audio items renderable by a local client / device, and can represent a visual scene by a plurality of video items.

[0073] FIG. 1 shows an example of a VR system in which a central server 101 communicates with a plurality of remote clients 103 via a network 105 such as the Internet. The central server 101 can be configured to support a large number of remote clients 103 that may exist simultaneously.

[0074] Such an approach can, in many scenarios, for example, improve the trade-off between complexity and the resource and communication requirements of various devices. For example, the scene data can be transmitted once or at a relatively low frequency, and the local rendering device (remote client 103) can receive the viewer's pose, locally process the scene data, and render audio and / or video to reflect the change in the viewer's pose. This approach can provide an efficient system and an attractive user experience. This can, for example, provide a real-time experience with low latency, and while enabling the centralized storage, generation, and management of scene data, can significantly reduce the required communication bandwidth. For example, it may be suitable for applications where a VR experience is provided to multiple remote devices.

[0075] Figure 2 shows elements of a visual rendering apparatus that can provide improved visual rendering in many applications and scenarios. In particular, the visual rendering apparatus may provide improved rendering for many VR applications, and the visual rendering apparatus may be specifically configured to execute the processing and rendering of the VR client 103 of FIG. 1.

[0076] The audio device of FIG. 2 is configured to render a three-dimensional scene by rendering spatial audio and video to provide a three-dimensional perception of the scene. In a specific description of the visual rendering apparatus, the focus is on applications that provide the described techniques for both audio and video, but it should be understood that in other embodiments, the techniques may be applied to only audio or only video / visual processing. In fact, in some embodiments, the rendering apparatus may include only a function for rendering audio or only a function for rendering video, that is, the visual rendering apparatus may be any audio rendering apparatus or any video rendering apparatus.

[0077] The visual rendering apparatus includes a first receiver 201 configured to receive visual items from a local or remote source. In this specific example, the first receiver 201 receives data representing visual items from the server 101. The first receiver 201 may be configured to receive data representing a virtual scene. The data may include data providing a visual description of the scene and may also include data providing an audio description of the scene. Thus, the received data may provide an audio scene description and a visual scene description.

[0078] The audio item may be encoded audio data, for example, an encoded audio signal. The audio item can be various types of audio elements, such as various types of signals and components. In fact, in many embodiments, the first receiver 201 may receive audio data that defines various audio types / formats. For example, the audio data may include audio represented by an audio channel signal, individual audio objects, scene-based audio (e.g., HOA (Higher Order Ambisonics)), etc. The audio can be represented, for example, as encoded audio for a given audio component to be rendered.

[0079] The first receiver 201 is coupled to a renderer 203, and the renderer renders a scene based on received data representing audiovisual items. In the case of encoded data, the renderer 203 may also be configured to decode the data (alternatively, in some embodiments, decoding may be performed by the first receiver 201).

[0080] Specifically, the renderer 203 may include an image renderer 205 configured to generate an image corresponding to the current viewing pose of the viewer. For example, the data includes spatial 3D image data (e.g., an image of a scene and depth or model description), and the visual renderer 203 can generate stereoscopic images (images for the user's left and right eyes) from this data, as is known to those skilled in the art. These images can be presented to the user, for example, by individual left-eye and right-eye displays of a VR headset.

[0081] The renderer 203 further includes an audio renderer 207 configured to render an audio scene by generating an audio signal based on an audio item. In this example, the audio renderer 207 is a binaural audio renderer that generates binaural audio signals for the user's left and right ears. The binaural audio signals are generated to provide a desired spatial experience and are typically reproduced by headphones or earphones, which may specifically be part of a headset worn by the user. The headset also includes a left-eye display and a right-eye display.

[0082] Accordingly, in many embodiments, the audio rendering by the audio renderer 207 is a binaural rendering process that uses an appropriate binaural transfer function to provide a desired spatial effect to a user wearing headphones. For example, the audio renderer 207 may be configured to use binaural processing to generate audio components that are perceived as arriving from a specific location.

[0083] Binaural processing is known to be used to provide a spatial experience by virtually placing sound sources using individual signals for each ear of a listener. By using an appropriate binaural rendering process, the signals required at the eardrums for a listener to perceive sound from any direction can be calculated, and the signals can be rendered to provide a desired effect. These signals are re-formed at the eardrums using headphones or crosstalk cancellation (suitable for rendering with speakers in close proximity to each other). Binaural rendering can be considered a technique that tricks the human auditory system into generating signals for the listener's ears such that the sound is perceived as coming from a desired location.

[0084] Binaural rendering is based on binaural transfer functions that vary from person to person due to the acoustic characteristics of the head, ears, and reflecting surfaces (e.g., shoulders). For example, binaural recordings can be created that simulate multiple sources at various locations using binaural filters. This can be achieved by convolving each sound source with a pair such as the head-related impulse response (HRIR) corresponding to the position of the sound source.

[0085] A well-known method for determining the binaural transfer function is binaural recording. This is a recording method using a dedicated microphone arrangement, and playback using headphones is assumed. The recording is performed by placing a microphone in the external auditory canal of the subject or using a dummy head with a built-in microphone, i.e., a half-body figure including the auricle (outer ear). The use of such a dummy head including the auricle provides a very similar spatial impression as if the person listening to the recording was present at the scene during the recording.

[0086] For example, an appropriate binaural filter can be determined by measuring the response from a sound source at a specific position in a 2D or 3D space to a microphone placed inside or near a human ear. Based on such measurements, a binaural filter reflecting the acoustic transfer function to the user's ear can be generated. Binaural recordings can be created that simulate multiple sources at various locations using binaural filters. This can be achieved, for example, by convolving each sound source with a pair of measured impulse responses at the desired sound source position. To create an illusion that the sound source is moving around the listener, usually a large number of binaural filters with a specific spatial resolution (e.g., 10 degrees) are required.

[0087] The head binaural transfer function can be expressed, for example, as a head impulse response (HRIR), or equivalently as a head-related transfer function (HRTF), a binaural room impulse response (BRIR), or a binaural room transfer function (BRTF). The transfer function (e.g., estimated or assumed) from a given location to the listener's ear (or eardrum) is given, for example, in the frequency domain, in which case it is usually called an HRTF or BRTF, and when given in the time domain, it is usually called an HRIR or BRIR. In some cases, the head binaural transfer function is determined to include the acoustic environment, particularly the characteristics or properties of the room in which the measurement is made, while in other cases only the characteristics of the user are considered. Examples of the first type of function are BRIR and BRTF.

[0088] Accordingly, the audio renderer 207 comprises a store having binaural transfer functions for a plurality of different locations, which are usually numerous. Each binaural transfer function provides information on how the audio signal should be processed / filtered for the audio signal to be perceived as having originated from that location. By applying binaural processing individually to a plurality of audio signals / sound sources and combining the results, an audio scene can be generated that includes a plurality of sound sources located at appropriate positions within the soundstage.

[0089] The audio renderer 207 can select and retrieve (or, in some cases, interpolate between a plurality of nearby binaural transfer functions) the stored binaural transfer function that best matches the desired location for a given audio element perceived to have originated from a given location relative to the user's head. The audio renderer can then apply the selected binaural transfer function to the audio signal of the audio element, thereby generating an audio signal for the left ear and an audio signal for the right ear.

[0090] The output stereo signal generated in the form of a left ear signal and a right ear signal is suitable for rendering on headphones and can be amplified to generate a drive signal for supply to the user's headset. Then, the user recognizes that the audio element is originating from the desired location.

[0091] It should be understood that in some embodiments, audio items can be processed, for example, to add acoustic environment effects. For example, the audio item can be processed to add reverberation or decorrelation / diffusivity, etc. In many embodiments, this processing can be performed on the generated binaural signal rather than directly on the audio element signal.

[0092] Thus, the audio renderer 207 can be configured to generate an audio signal such that a user wearing headphones perceives that the audio element is received from a desired position. Other audio items can be, for example, dispersed and diffused and rendered as such, depending on the case.

[0093] It will be understood that many algorithms and techniques for the rendering of spatial audio, particularly binaural rendering (e.g., using headphones), are known to those skilled in the art and any suitable technique can be used without detracting from the present invention.

[0094] The audiovisual rendering device further comprises a second receiver 209 which is a metadata receiver configured to receive metadata of the audiovisual item. The metadata particularly includes one or more position data of the audiovisual item. The metadata can include an input position indicating one or more positions of the audiovisual item.

[0095] The received audiovisual data may include audio and / or visual data that describe a scene. The feed kernel data specifically includes visual data regarding a set of visual items corresponding to sound sources and / or visual objects within the scene. Some audio items may represent a sound source located within the scene that is associated with a specific position and / or orientation within the scene (in the case of a moving object, the position and / or orientation may change dynamically). The visual data may include data that describes scene objects that generate these visual representations and enable them to be represented in an image / video presented to the user (typically, a 3D image using an individual display of a headset).

[0096] Often, an audio element may represent audio generated by a specific scene object within the virtual scene, and thus may represent a sound source at a position corresponding to the position of the scene object (e.g., a person's voice). In such cases, the same position data / display may be included and used for both the audio item and the corresponding visual scene object (similarly for orientation).

[0097] Other elements may represent a plurality of more dispersed or diffused sound sources, such as ambient noise or background noise that can spread. As another example, some audio elements may fully or partially represent spatially non-localized components of the audio from a sound source that is spatially clearly defined, such as a reverberation from a spatially defined sound source.

[0098] Similarly, some visual scene objects may have an extended position, for example, the position data may indicate the center or reference position of the scene object.

[0099] The metadata may include pose data indicating the position and / or orientation of the audiovisual items, specifically the position and / or orientation of the sound sources and / or visual scene objects or elements. The pose data may include, for example, absolute position and / or orientation data that define the position of each item or at least some of the items.

[0100] The pose is provided with reference to an input coordinate system, i.e., this input coordinate system is the reference coordinate system for the pose display provided within the metadata received by the second receiver. The input coordinate system is typically fixed with reference to the scene to be represented / rendered. For example, the scene may be a virtual (or actual) scene in which a sound source and scene objects are present at positions provided with reference to the scene. That is, the input coordinate system is typically the scene coordinate system of the scene represented by the audiovisual data.

[0101] To provide a user representation of the scene, the scene is rendered from the pose of the viewer or user. That is, the scene is rendered as it would be perceived at a particular viewer / user pose within the scene, and the audio and visual rendering provide the audio and images that would be perceived at that viewer's pose.

[0102] The rendering by the renderer 203 is performed with respect to a rendering coordinate system fixed with respect to the user's head and head movement. The playback of the rendered audio signal is typically a playback device worn or attached to the head, such as headphones / earphones and individual displays for each eye. Usually, the playback is performed by a headset device having audio and video playback means. The renderer is configured to generate audiovisual items rendered with reference to the playback device / means, and it is assumed that the user's pose is fixed / constant with reference to the playback device / means. For example, the position of the sound source is determined relative to the headphones (i.e., the position within the rendering coordinate system), the appropriate HRTF filter is retrieved, and it is used to render the audio signal such that the sound source is recognized as arriving from the required relative position within the rendering coordinate system. Similarly, for a given scene object, the relative position with respect to the display (i.e., the position within the rendering coordinate system) is determined, and the images corresponding to the views from the left and right eyes for this position are determined respectively.

[0103] Therefore, the rendering coordinate system is considered to be a coordinate system fixed with respect to the user's head, and specifically, it can be considered as a rendering coordinate system that does not depend on the movement or pose change of the user's head. The playback device (audio, visual, or both audio and visual) is assumed / considered to be fixed with respect to the user's head and thus with respect to the rendering coordinate system.

[0104] The rendering coordinate system that is fixed with respect to the movement of the user's head can be considered to correspond to the playback coordinate system fixed with respect to the playback device for playing back the rendered audiovisual item. The term rendering coordinate system may be synonymous with the coordinate system of the playback device / means, and can be replaced by these terms. Similarly, the term "rendering coordinate system fixed with respect to the movement of the user's head" may be synonymous with the "playback device / means coordinate system fixed with respect to the playback device for playing back the rendered audiovisual item" and can be replaced by this.

[0105] Since the renderer performs rendering based on the pose with reference to the rendering coordinate system and the pose for the audiovisual item is provided with reference to the input coordinate system, the audiovisual rendering device includes a mapper 211 configured to map the input position in the input coordinate system to the rendering position in the rendering coordinate system.

[0106] The audiovisual rendering device includes a head movement data receiver 213 for receiving user head movement data indicating the movement of the user's head. The user head movement data can indicate the movement of the user's head in the real world and is usually provided with reference to the coordinate system of the real world. The head movement data can indicate the absolute or relative movement of the user's head in the real world, and specifically, can reflect the absolute or relative change in the user's pose with respect to the coordinate system of the real world. The head movement data can indicate the change (or no change) in the pose (orientation and / or position) of the head and can also be referred to as head pose data.

[0107] Many different techniques are known for detecting and representing head movement, and it will be understood that any suitable technique can be used without detracting from the present invention. Specifically, the head movement data receiver 213 can receive head movement data from a VR headset or VR head movement detector, as is known in the art.

[0108] The mapper 211 is coupled to the head movement data receiver 213 and receives user head movement data. The mapper is configured to perform a mapping between an input position in an input coordinate system and a rendering position in a rendering coordinate system in response to the user head movement data. For example, the mapper 211 can continuously process the user head movement data to continuously track the current pose of the user in a real-world coordinate system. Then, a mapping between the input pose and the rendering pose can be performed based on the user pose.

[0109] For example, in many applications, it is desirable to provide the user with an experience as if they were within the three-dimensional scene being presented. Therefore, it is desirable for the rendered audio and images to reflect the pose of the user that follows the movement of the user's head. Thus, it is desirable for the audiovisual items to be rendered such that they are perceived as being fixed relative to the movement in the real world. This is because the movement in the real world will be reproduced in the rendering of the (usually virtual) scene.

[0110] In such a case, the mapping of the input pose to the rendering pose is a mapping as if the visual item were fixed with respect to the real world, i.e., it is rendered so as to be recognized as fixed with respect to the real world. Thus, the same input pose is mapped to different rendering poses so that changes in the pose of the user's head are reflected. For example, when the user rotates the head by 30°, the real-world scene is based on the user rotated by -30°. The mapper 211 can perform corresponding changes so that the mapping from the input pose to the rendering pose is modified to include an additional 30° rotation with respect to the situation before the rotation of the user's head. As a result, the visual item will be in a different pose in the rendering coordinate system but is recognized as being in the same real-world pose. Thus, the mapping can be dynamically changed so that the visual item is recognized as being fixed with respect to the real world, and thus a very natural experience is provided.

[0111] For example, to create the illusion of an immersive virtual world, the three-dimensional audio and / or visual rendering is typically controlled by head tracking. The rendering is corrected with respect to the pose of the head, and the pose of the head includes in particular changes in the orientation of the head (such as the three degrees of freedom in space (3-DoF) of yaw, pitch, roll, etc.). The rendering is such that the visual item is recognized as being fixed with respect to the user. This head tracking and the effect of rendering adaptation brought about by head tracking are the high sense of reality and exocentric recognition of the rendered content as compared to static rendering.

[0112] However, as another approach, there is a method of mapping an input pose to a rendering pose that is fixed with respect to the rendering coordinate system. This can be implemented, for example, by the mapper 211 applying a fixed mapping from the input pose to the rendering pose, where the mapping is independent of head movement data, specifically, a change in the user's pose does not change the mapping between the input pose and the rendering pose. The effect of such a mapping is effectively that the recognized scene moves with the head, i.e., it is static with respect to the user's head. This may seem unnatural in most scenes, but in some cases it can be advantageous. For example, it can provide a desirable experience when listening to music or sounds that are not part of the scene, such as a narrator.

[0113] It is possible to apply different methods to different audiovisual items. In MPEG terminology, the terms "fixed to the head orientation" or "not fixed to the head orientation" are used to refer to audio items that are rendered either completely following the user's movement or ignoring it.

[0114] For example, an audio item can be considered "not fixed to the head", which means that the rendering is dynamically adapted to the (change in) orientation of the user's head because it is an audio element intended to have a fixed position in the (virtual or real) environment. Another audio item can be considered "fixed to the head", which means that it is an audio item intended to have a fixed position with respect to the user's head. Such an audio item can be rendered independently of the listener's pose. Therefore, in the rendering of such an audio item, the (change in) orientation of the user's head is not considered, i.e., such an audio item is an audio element whose relative position does not change even if the user rotates their head (e.g., ambient sound or non-spatial audio such as music intended to track the user without changing the relative position).

[0115] In the described system, the second receiver 209 is configured to receive metadata that further includes a rendering category display for at least some of the audiovisual items. The rendering category display indicates a rendering category from a set of rendering categories, and the rendering of the audiovisual item is performed according to the rendering category indicated for the audiovisual item. Different rendering categories may define different rendering parameters and operations.

[0116] The rendering category display can be any display that can be used to select a rendering category from a set of rendering categories. In many embodiments, it may be data provided only for selecting a rendering category and / or data that directly specifies one category. In other embodiments, the rendering category display may be a display that provides additional information or provides a description of the corresponding audiovisual item. In some embodiments, the rendering category display is one parameter considered when selecting a rendering category, and other parameters may also be considered.

[0117] As a specific example, in some embodiments, an audio item can be encoded audio data, for example, an encoded audio signal where the audio item can be a plurality of different types of audio items including a plurality of different types of signals and components. In practice, in many embodiments, the metadata receiver 201 can receive metadata that defines a plurality of different types / formats of audio. For example, the audio data can include audio represented by audio channel signals, individual audio objects, HOA (Higher Order Ambisonics), etc. The metadata can be included as part of the audio item or separately from the audio item that describes the audio type of each audio item. This metadata can be a rendering category display and can be used to select an appropriate rendering category for the audio item.

[0118] The rendering categories are associated with different specific reference coordinate systems, and each rendering category is linked to a specific coordinate system transformation from the real-world coordinate system to the category coordinate system. Specifically, for each category, the real-world coordinate system, for example, the real-world coordinate system that serves as a reference in providing head movement data, can be transformed into different coordinate systems given by the transformation. Since the coordinate system transformation is different for different categories, it will be linked to different category reference systems.

[0119] The coordinate system transformation is usually a dynamic coordinate system transformation for one, some, or all of the categories. Therefore, the coordinate system transformation is usually not a fixed or static coordinate system transformation and can change over time according to various parameters. For example, as will be explained in more detail later, the coordinate system transformation can depend on dynamically changing parameters such as the movement of the user's torso, the movement of an external device, and / or in some cases head movement data. Thus, in many embodiments, the coordinate system transformation of a category is a time-varying coordinate system transformation that depends on user movement parameters. User movement parameters can indicate the movement of the user relative to the real-world coordinate system.

[0120] The mapping performed by the mapper 211 for a given audiovisual item depends on the category coordinate system of the rendering category to which the audiovisual item is shown to belong. Specifically, based on the rendering category display, the mapper 211 can determine the rendering category to be used for rendering the audiovisual item. Next, the mapper can determine the coordinate system transformation linked to the selected category. Thereafter, the mapper 211 can perform the mapping from the input pose to the rendering pose so as to correspond to a fixed pose within the category coordinate system obtained from the selected coordinate system transformation.

[0121] Accordingly, the category coordinate system can be regarded as a reference coordinate system serving as a reference for rendering such that the audiovisual item is fixed. The category coordinate system may also be referred to as the reference coordinate system or the fixed reference coordinate system (for a given category).

[0122] In many embodiments, one rendering category may correspond to a rendering in which the sound source and scene objects represented by the audiovisual item are fixed with respect to the real world as described above. In such embodiments, the coordinate system transformation is such that the category coordinate system for the category is aligned with the real-world coordinate system. For such a category, the coordinate system transformation may be a fixed coordinate system transformation, for example, a single one-to-one mapping of the real-world coordinate system. Accordingly, the category coordinate system can in fact be the real-world coordinate system, or, for example, a fixed static translation, scaling, and / or rotation.

[0123] In many embodiments, one rendering category may correspond to a rendering in which the sound source and scene objects represented by the audiovisual item are fixed with respect to the movement of the head, i.e., with respect to the rendering coordinate system. In such embodiments, the coordinate system transformation is such that the category coordinate system for the category is aligned with the user's head / playback device / rendering coordinate system. For such a category, the coordinate system transformation can be a coordinate system transformation that fully follows the movement of the user's head. For example, when the head is rotated, a corresponding rotation of the coordinate system transformation occurs, and when the position of the user's head changes, the same change occurs in the coordinate system transformation. Accordingly, according to such a rendering category, the coordinate system transformation is dynamically modified to follow the head movement data so that the resulting category coordinate system is aligned with the rendering coordinate system. Thereby, as described above, a fixed mapping from the input coordinate system to the aforementioned rendering coordinate system is obtained.

[0124] Although not required, in many embodiments, the rendering category may include a category in which the audiovisual item is fixed and rendered with respect to a real-world coordinate system and a category in which the audiovisual item is fixed and rendered with respect to a rendering coordinate system. However, in the system being described, one or more rendering categories include a rendering category in which the audiovisual item is fixed to a coordinate system that is neither the real-world coordinate system nor the rendering coordinate system. That is, the rendering of the audiovisual item is provided with a rendering category that is not fixed to either the real world or the user's head. Thus, at least one category coordinate system is variable with respect to the real-world coordinate system and the rendering coordinate system. Specifically, the coordinate system is different from the rendering coordinate system and the real-world coordinate system, and the difference between these and the category coordinate system is not constant and may vary.

[0125] Thus, at least one rendering category may provide a rendering that is not fixed to either the real world or the user. Rather, in many embodiments, an intermediate experience may be provided.

[0126] For example, the coordinate system transformation can be such that the corresponding coordinate system is fixed with respect to the real world, except when an update criterion is met. However, when the criterion is met, the coordinate system transformation can provide a different relationship between the real-world coordinate system and the category coordinate system. For example, this mapping can be such that the visual is rendered fixed with respect to the real-world coordinate system, i.e., the visual item is felt to be in a fixed position. However, when the user rotates their head by more than a certain amount, the relationship between the real-world coordinate system and the category coordinate system is changed to compensate for this rotation. For example, as long as the user's head movement is less than 20°, the visual item is rendered as being in a fixed position. However, when the user's head movement exceeds 20°, the category coordinate system is rotated 20° with respect to the real-world coordinate system. This can provide the user with an experience of perceiving a natural three-dimensional experience with respect to the rendered visual item, as long as the movement is small enough. However, when the head movement is large, the rendering of the visual item is readjusted to match the changed head position.

[0127] As a specific example, the sound source corresponding to the narrator can be presented so as to be initially placed in front of the user. When the user's movement is small, the voice is rendered so that the narrator is recognized as being stationary at the same position. This provides a natural experience and perception, and in particular provides an out-of-head localization perception of the narrator. However, when the user rotates their head by, for example, 20° or more from the original direction towards the narrator sound source, the system adjusts the mapping and relocates the narrator sound source to the front of the user's new orientation. For small movements around this point, the sound source of the narrator is rendered at this new (with respect to the real-world coordinate system) fixed position. If the movement with respect to this new sound source position exceeds a given threshold again, the category coordinate system can be updated again and the recognized position of the narrator sound source can be updated again. In this way, a narrator that is fixed with respect to small movements but follows the user with respect to large movements can be provided to the user. In the example described, the narrator is made to be recognized as fixed and appropriate spatial cues are provided for head movements, but it is possible for the narrator to always be located substantially in front of the user (for example, even when the user rotates 180° completely).

[0128] In this approach, the metadata includes a plurality of rendering category indicators for a plurality of audiovisual items. Thereby, the source side can control flexible rendering at the receiving side, and the rendering is adapted specifically to individual audiovisual items. For example, different spatial renderings and perceptions can be applied to items corresponding to background music, narration, sound sources corresponding to specific objects fixed within a scene, conversations, etc.

[0129] In some embodiments, the coordinate system transformation for the rendering category depends on user head movement data. Accordingly, in some embodiments, at least one parameter that changes the coordinate system transformation depends on user head movement data.

[0130] In many embodiments, the coordinate system transformation may depend on user head pose characteristics or parameters determined from user head movement data. For example, as described above, if the user head pose exhibits a rotation exceeding a certain amount, the coordinate system transformation may be adapted to include a rotation corresponding to that amount. As another example, mapper 211 may be able to detect whether the user has maintained a given pose for longer (sufficiently) than a given duration. If so, the coordinate system transformation may be adapted to place the audiovisual item at a given position within the rendering coordinate system (i.e., a specific position relative to the user (such as directly in front of the user)).

[0131] In some embodiments, the coordinate system transformation depends on the average head pose. In particular, in some embodiments, the coordinate system transformation may be such that the categorical coordinate system is aligned with the average head pose. In some embodiments, the coordinate system transformation may be such that the categorical coordinate system is fixed relative to the average head pose.

[0132] The average head pose may be determined, for example, by low-pass filtering the head pose measurements with a low-pass filter having an appropriate cut-off frequency, specifically, by applying an unweighted average over a window of appropriate duration.

[0133] In some embodiments, the reference for rendering may be selected as the average head orientation h for one or more rendering categories. Thereby, the audiovisual item follows the slower and longer-lasting changes in the head orientation, and the sound source is felt to remain in the same place relative to the head (e.g., in front of the face), but when the head movement is fast, the audiovisual item appears to be fixed relative to the real world rather than the head (and thus feels fixed in the virtual world as well). Thus, typical small and fast head movements in daily life create an immersive out-of-head localization illusion while allowing the overall perception that the audiovisual item is following the user.

[0134] In some embodiments, the adaptation and tracking may be non-linear. For example, when the head rotates significantly, the reference of the average head orientation can be "clipped" so as not to deviate beyond a certain maximum angle with respect to the instantaneous head orientation. For example, if this maximum value is 20°, as long as the head "moves incrementally" within the range of + / - 20°, the out-of-head localization experience is realized. When the head rotates quickly and exceeds the maximum value, the reference follows the head orientation (lagging by a maximum of 20°), and when the movement stops, the reference is fixed again.

[0135] In some embodiments, at least one rendering category is associated with a coordinate system transformation that depends on the user's body pose. In such embodiments, the audiovisual rendering device may include a body pose receiver 215 configured to receive user body pose data indicating the user's body pose.

[0136] The body pose can be determined, for example, by a dedicated inertial sensor unit arranged or worn on the body. As another example, the body pose can be determined by sensors of a smart device such as a smartphone worn in a pocket. As yet another example, coils can be arranged on the user's body and head respectively, and based on the change in the coupling between the two, the movement of the head relative to the body can be determined.

[0137] In such embodiments, at least one parameter that changes the coordinate system transformation depends on the body pose data.

[0138] In many embodiments, the coordinate system transformation may depend on the body pose data or parameters determined from the body pose data.

[0139] For example, if the torso pose data indicates a rotation of the torso exceeding a certain amount, the coordinate system transformation can be adapted to include a rotation corresponding to that amount. As another example, mapper 211 can detect whether the user has maintained a substantially constant torso pose for longer than a given duration. If so, the coordinate system transformation can be adapted to place the audiovisual item at a given position corresponding to the torso pose within the rendering coordinate system.

[0140] In particular, in some embodiments, the coordinate system transformation can be such that the category coordinate system is adapted to the user's torso pose. In some embodiments, the coordinate system transformation can be such that the category coordinate system is fixed relative to the user's torso pose. Thus, in some embodiments, the audiovisual item can be rendered to follow the user's torso, and thus, a perception and experience can be provided such that the audiovisual item follows the user's full body movement but appears fixed with respect to head movement relative to the torso. This can provide both a desirable out-of-head localization perception and a rendering of an audiovisual item that follows the user.

[0141] In some embodiments, the coordinate system transformation depends on the average torso pose. The average torso pose can be determined, for example, by low-pass filtering the head pose measurements with a low-pass filter having an appropriate cut-off frequency, specifically, by applying an unweighted average over a window of appropriate duration.

[0142] Thus, in some embodiments, one or more of the rendering categories can use a coordinate system transformation that provides a rendering criterion aligned with the instantaneous or average chest / torso orientation t By doing so, the audiovisual item can be made to feel as if it remains in front of the user's body rather than in front of the user's face. Still, by rotating the head relative to the chest / torso, the audiovisual item can be perceived as coming from various directions, greatly contributing to an immersive out-of-head experience.

[0143] In some embodiments, the coordinate system transformation for a rendering category depends on device pose data indicating the pose of an external device. In such embodiments, the audiovisual rendering device may include a device pose receiver 217 configured to receive device pose data indicating the device body pose.

[0144] The device can be, for example, a device that (hypothetically) a user wears, carries, attaches to the user, or otherwise secures. In many embodiments, the external device can be, for example, a mobile phone or a personal device, such as a smartphone in a pocket, a body-worn device, or a handheld device (e.g., a smart device used to view visual VR content).

[0145] Many devices include a gyro, an accelerometer, a GPS receiver, etc., and can determine the relative or absolute orientation of the device. In that case, the device can determine its current relative or absolute orientation and transmit it to the device pose receiver 217 using a suitable communication, usually wireless. For example, the communication can be via a WiFi or Bluetooth connection.

[0146] In such embodiments, at least one parameter that varies the coordinate system transformation depends on the device pose data.

[0147] In many embodiments, the coordinate system transformation can depend on the device pose data or parameters determined from the device pose data.

[0148] For example, if the device pose data indicates a rotation of the device beyond a certain amount, the coordinate system transformation can be adapted to include a rotation corresponding to that amount. As another example, the mapper 211 can detect whether the device has maintained a sufficiently constant body pose for a given duration. If so, the coordinate system transformation can be adapted to place the audiovisual item at a given position corresponding to the device pose within the rendering coordinate system.

[0149] In particular, in some embodiments, the coordinate system transformation can be such that the categorical coordinate system is aligned with the device pose. In some embodiments, the coordinate system transformation can be such that the categorical coordinate system is fixed relative to the device pose. Thus, in some embodiments, the audiovisual item can be rendered to follow the pose of the device, and thus, a perception and experience can be provided where the audiovisual item follows the movement of the device. In many practical user scenarios, the device can provide an excellent representation of the reference of the user pose. For example, a body-worn device, or a smartphone in a pocket for example, may well reflect the movement of the whole user. This can provide an excellent reference for judging relative head movement and thus can provide an experience that combines both a realistic reaction to head movement and the ability of the audiovisual item to follow the user's larger movements.

[0150] Furthermore, using an external device as a reference can be very practical and can provide a reference that results in a desirable user experience. This approach may often be based on a device that is already worn or carried by the user and has the functionality necessary to determine and transmit the device pose. For example, nowadays most people carry a smartphone that has, for example, an accelerometer for determining the device pose and communication means (e.g., Bluetooth) suitable for transmitting the device pose data to the audiovisual rendering device.

[0151] In some embodiments, the coordinate system transformation can depend on the average device pose. The average device pose can be determined, for example, by low-pass filtering the device pose measurements with a low-pass filter having an appropriate cut-off frequency, specifically, by applying an unweighted average over a window of appropriate duration.

[0152] Thus, in some embodiments, one or more of the rendering categories may use a coordinate transformation that provides a rendering criterion adapted to the instantaneous or average device orientation. By doing so, the audiovisual item may be felt to remain in a fixed position relative to the device, so that the item moves as the device moves, but remains fixed relative to head movement, thereby providing a more natural and immersive out-of-head localization experience.

[0153] Thus, this approach provides a technique for controlling the rendering of audiovisual items using metadata, and can control the audiovisual items individually to provide different user experiences for each audiovisual item. The experience includes one or more options that provide a perception of the audiovisual item that is neither completely fixed to the real world nor completely follows the user (fixed to the head). Specifically, the audiovisual item can provide an intermediate experience that is somewhat fixed relative to the real world and somewhat follows the user's movement.

[0154] It will be appreciated that in some embodiments, the possible rendering categories, each associated with a given coordinate transformation, may be determined in advance. In such embodiments, the audiovisual rendering device can save the coordinate transformation for each category, or equivalently, directly save the mapping corresponding to the coordinate transformation when appropriate. The mapper 211 can be configured to retrieve the saved coordinate transformation (or mapping) for the selected rendering category and apply it when performing the mapping of the audiovisual item.

[0155] For example, in the case of the first audiovisual item, the rendering category indicator may indicate that it should be rendered according to the first category. The first category may be for rendering an audiovisual item fixed to the head pose, and thus the mapper can extract a mapping that provides a fixed one-to-one mapping between the input pose and the rendering pose. In the case of the second audiovisual item, the rendering category indicator may indicate that it should be rendered according to the second category. The second category may be for rendering an audiovisual item fixed to the real world, and thus the mapper can extract a coordinate system transformation that adjusts the mapping so that the head movement is compensated and the rendering pose corresponds to a fixed position in the real space. In the case of the third audiovisual item, the rendering category indicator may indicate that it should be rendered according to the third category. The third category may be for rendering an audiovisual item fixed to the device pose or the body pose. The mapper can extract a coordinate system transformation or a mapping that adjusts the mapping so that the rendering of the audiovisual item fixed to the device or the body pose is obtained by compensating for the head movement relative to the device or the body pose.

[0156] Through mapping, different rendering categories are associated with different coordinate system transformations so that a rendering position fixed to the category coordinate system is brought about. However, it should be understood that the mapper 211 does not need to explicitly determine such a coordinate system transformation or the category coordinate system. Instead, in a typical embodiment, a mapping function is defined for each individual rendering category such that the resulting rendering pose is fixed relative to the category coordinate system. For example, a mapping function that is a function of the head pose relative to the device (or body) pose may be used to directly map the input position to a rendering position that is fixed relative to the category coordinate system that is fixed relative to the device (or body) pose.

[0157] In some embodiments, the metadata may include data that partially or fully characterizes, describes, and / or defines one or more of a plurality of rendering categories. For example, in addition to the rendering category display, the metadata may include data that describes a coordinate system transformation and / or a mapping function to be applied to one or more categories. For example, the metadata may indicate that in a first rendering category, a fixed mapping from an input position to a rendering position is required, in a second category, a mapping that fully compensates for head movement so that the item appears to be fixed relative to the real world is required, and in a third rendering category, it is recognized as an intermediate experience where the item appears fixed for smaller and faster head movements but follows the user for slower average movements, such that a mapping needs to be compensated for the amount of head movement relative to the average head movement.

[0158] In different embodiments, different techniques and data may be used as the rendering category display. In some embodiments, each category may be associated with, for example, a category number, and the rendering category display may directly provide the category number to be used for the audiovisual item.

[0159] In many embodiments, the rendering category display may indicate a characteristic or feature of the audiovisual item, which may be mapped to a particular rendering category display.

[0160] In some embodiments, the rendering category display can specifically indicate whether the audiovisual item is a diegetic audiovisual item or a non-diegetic audiovisual item. A diegetic audiovisual item can be an item belonging to a scene such as a movie or a story being screened. In other words, a diegetic audiovisual item originates from a source within a movie, a story, etc. (e.g., an actor in a play, a bird in a natural video and its chirping sound, etc.). A non-diegetic audiovisual item can be an item originating from outside the movie or the story (e.g., the director's audio commentary, mood music, etc.). In many scenarios, in accordance with the MPEG usage, a diegetic audiovisual item can correspond to "not fixed to the head orientation", and a non-diegetic audiovisual item can correspond to "fixed to the head orientation".

[0161] In some embodiments, there may only be two rendering categories, specifically corresponding to the rendering of audiovisual items shown to be diegetic and the rendering of audiovisual items shown to be non-diegetic, respectively.

[0162] For example, in some applications and systems, a diegetic signal is transmitted downstream to an audiovisual rendering device, and as illustrated with respect to FIG. 3, the desired rendering operation may depend on this signal. · By using head tracking based on the real-world orientation, it is desirable for the diegetic source D to be felt to stay firmly at the position of that source within the virtual world V, and thus it is necessary to render it so as to be felt to be fixed based on the real world. In a given example of a movie application, the actor's voice is rendered straight ahead of the user, and when the user rotates the head 50° to the left, the sound is rendered to the headphones so as to be projected at 50° to the right, thereby making it feel as if it stays at the same virtual position. · On the one hand, the non-diegetic sound source N can be rendered regardless of the head orientation. In other words, the sound remains at a fixed position “hard-coupled” to the head (e.g., in front of the head) and rotates with the head. This is achieved by using a fixed mapping from the input position to the rendering position, rather than applying a mapping that depends on the head orientation to the sound. In the movie example, the director's commentary audio can be rendered right in front of the user's eyes and is not affected by head movement (i.e., the sound remains in front of the head).

[0163] The visual rendering device of FIG. 2 is configured to provide a more flexible approach, in which at least one selectable rendering category enables the rendering of visual items to follow the user for some movements and be fixed with respect to the real / virtual world for other movements, as if it were fixed for some movements. This approach can be particularly applied to non-diegetic visual items.

[0164] In this specific example, alternative or additional options for rendering the non-diegetic sound source may include one or more of the following. ·The rendering reference can be selected as the average head orientation / pose h (as shown in FIG. 4). As a result, non-diegetic sound sources follow slower and longer-lasting changes in head orientation, and while the sound source is felt to remain in the same place relative to the head (e.g., in front of the face), when the head moves quickly, the non-diegetic sound source is felt to be fixed within the virtual world (rather than relative to the head). Thus, typical small and fast head movements in daily life still produce an immersive out-of-head localization illusion. As an improvement in some embodiments, since non-diegetic audio preferably remains at the same virtual position as much as possible, the tracking may be non-linear. For example, when the head rotates significantly, the reference of the average head orientation can be "clipped" so as not to deviate beyond a certain maximum angle relative to the instantaneous head orientation. For example, if this maximum value is 20°, as long as the head "moves incrementally" within the range of + / -20°, an out-of-head localization experience is achieved. When the head rotates quickly and exceeds the maximum value, the reference follows the head orientation (lagging up to 20°), and when the movement stops, the reference is fixed again. ·The head tracking reference can be selected as the average chest / trunk orientation / pose t (as shown in FIG. 5). By doing so, non-diegetic content is felt to remain in front of the user's body rather than in front of the user's face. Still, by rotating the head relative to the chest / trunk, non-diegetic content can be made to sound from various directions, greatly contributing to an immersive out-of-head localization experience. ·The head tracking reference can be selected as the instantaneous or average orientation / pose of an external device such as a mobile phone or a body-worn device. By doing so, non-diegetic content is felt to remain in front of the device rather than in front of the user's face. Still, by rotating the head relative to the device, non-diegetic content can be made to sound from various directions, greatly contributing to an immersive out-of-head experience.

[0165] In some embodiments, the rendering may be configured to operate in different modes in response to user movement. For example, if the user's movement meets a movement criterion, the audiovisual rendering device operates in a first mode; otherwise, it operates in a second mode. In this example, the two modes can provide different category coordinate systems for the same rendering category display. That is, depending on the user's movement, the audiovisual rendering device can render a given audiovisual item in a given rendering category display fixed with respect to different coordinate systems.

[0166] In some embodiments, the mapper 211 may be configured to select a rendering category for a given rendering category display in response to user movement parameters indicating the user's movement. Thus, depending on the user movement parameters, different rendering categories can be selected for a given audiovisual item and rendering category display. Specifically, if the user movement parameters meet a first criterion, a given link between possible rendering category display values and a set of rendering categories can be used to select the rendering category of the received rendering category display. However, if the criterion is not met (or, for example, if different criteria are met), the mapper 211 can use a different link between possible rendering category indication values and the same or a different set of rendering categories to select the rendering category of the received rendering category display.

[0167] This approach can enable, for example, an audiovisual rendering device to provide different renderings and experiences for mobile users and fixed-type users.

[0168] In some embodiments, the selection of the rendering category may also depend on other parameters, such as user settings or configuration settings by an application (e.g., an app on a mobile device).

[0169] In another approach, the mapper can be configured to determine a coordinate system transformation between a real-world coordinate system used as a reference for coordinate system transformation of a selected category and a coordinate system that is a reference in providing user head movement data provided in response to user movement parameters indicating the user's movement.

[0170] Accordingly, in some embodiments, rendering compliance can be introduced by a coordinate system transformation of the rendering category with respect to a reference that can vary with respect to a (typically real-world) coordinate system indicating head movement. For example, based on the user's movement, compensation can be applied to the head movement data, for example, offsetting the overall movement of the user. As an example, when the user is on a boat, the user head movement data may not only indicate the user's movement relative to the body or the body's movement relative to the boat, but may also reflect the movement of the boat. Since this may be undesirable, the mapper 211 can compensate the head movement data with respect to user movement parameter data reflecting the component of the user's movement due to the movement of the boat. The resulting modified / compensated head movement data is given with respect to a coordinate system in which the movement of the boat is compensated, and the desired rendering and user experience can be realized by directly applying the transformation of the selected category coordinate system to this modified coordinate system.

[0171] It will be appreciated that the user movement parameters can be determined in any suitable manner. For example, in some embodiments, they can be determined by a dedicated sensor that provides relevant data. For example, an accelerometer and / or gyroscope can be attached to a vehicle such as an automobile or boat carrying the user.

[0172] In many embodiments, the mapper can be configured to determine user movement parameters in response to the user head movement data itself. For example, user parameters indicating user movement parameters may be determined using analysis of underlying motion, such as long-term averaging or identification of periodic components (e.g., corresponding to waves moving the boat).

[0173] In some embodiments, the coordinate system transformation for two different rendering categories depends on the user's head movement data, but the dependencies have different temporal averaging characteristics. For example, one rendering category depends on the average of the head movement, but is associated with a coordinate system transformation with a relatively short averaging time (i.e., a relatively high cutoff frequency for the averaging low-pass filter), while another rendering category also depends on the average of the head movement, but may be associated with a coordinate system transformation with a longer averaging time (i.e., a relatively low cutoff frequency for the averaging low-pass filter).

[0174] As an example, mapper 211 may evaluate the head movement data against a reference that reflects the consideration that the user is located in a static environment. For example, the low-pass filtered position change is compared to a given threshold, and if it is below the threshold, the user may be considered to be in a static environment. In this case, the rendering may be performed as in the specific examples described above for diegetic and non-diegetic audio items. However, if the position change exceeds the threshold, the user may be considered to be in a moving environment, for example, walking, or using a means of transportation such as a car, train, or airplane. In this case, the following rendering techniques may be applied (see also FIGS. 6 and 7). · The non-diegetic sound source N may be rendered using the average head orientation h as a reference. This makes it possible to create an "out-of-head localization" experience with small head movements while keeping the non-diegetic sound source mainly in front of the face. In other examples, for example, the user's torso or the pose of the device may be used as a reference. Thus, this technique may be compatible with the technique used for diegetic sound sources when they are static. · The diegetic sound source D also uses the average head orientation h as a reference, but may be rendered using an additional offset corresponding to a longer-term average head orientation h'. Thus, in this case, the diegetic virtual sound source is felt to be in a fixed position relative to the virtual world V, but this virtual world V is felt to "move with the user", i.e., it maintains an orientation that is more or less fixed relative to the user. The averaging (or other filtering) used to obtain h' from the instantaneous head orientation is typically selected to be significantly slower than that used for h. Thus, non-diegetic sources follow faster head movements, while diegetic content (and the virtual world V as a whole) takes time to change with the user's head orientation.

[0175] For clarity, the above description has been presented in terms of different functional circuits, units, and processors in connection with embodiments of the invention. However, it will be understood that the functions may be appropriately distributed among different functional circuits, units, or processors without detracting from the invention. For example, functions described as being performed by a plurality of separate processors or controllers may be performed by the same processor or controller. Thus, references to specific functional units or circuits are not intended to imply a particular logical or physical structure or organization, but rather to an appropriate means for providing the described function.

[0176] The present invention can be implemented in any suitable form including hardware, software, firmware, or any combination thereof. The present invention may be at least partially implemented as computer software operating on one or more data processors and / or digital signal processors. The elements and components of the embodiments of the present invention can be physically, functionally, and logically implemented in any suitable manner. In practice, the functions may be implemented as a single unit, multiple units, or part of other functional units. Accordingly, the present invention may be implemented as a single unit or physically and functionally distributed among different multiple units, circuits, and processors.

[0177] Although the present invention has been described in connection with some embodiments, the present invention is not limited to the specific forms described in the specification. The scope of the present invention is limited only by the appended claims. Further, even if a certain feature may seem to be described in connection with a particular embodiment, those skilled in the art will recognize that the various features of the above embodiments can be combined according to the present invention. In the claims, terms such as "comprising," "including," etc. do not exclude the presence of other elements or steps.

[0178] Furthermore, even if individually enumerated, a plurality of means, elements, circuits, or method steps may be implemented, for example, by a single circuit, unit, or processor. Further, even if individual features are included in different claims, these may be suitably combined, and being included in different claims does not mean that the combination of features is impossible and / or not advantageous. Also, just because a feature is included within one claim category does not mean that the feature is limited to this category, and the feature may equally be applied to other claim categories as appropriate. Further, the order of features in the claims does not refer to a specific order in which the features should act, and in particular, the order of individual steps in a method claim does not mean that the steps must be performed in that order. The steps may be performed in any suitable order. Also, a singular expression does not exclude a plural. Thus, a singular expression does not exclude a plurality. Reference signs within the claims are merely examples for clarity and do not limit the scope of the claims in any way.

Claims

1. A first receiver that receives audiovisual items, A metadata receiver that receives metadata including an input pose for each of at least some of the audiovisual items and a rendering category display for each of at least some of the audiovisual items, wherein the input pose is provided with reference to an input coordinate system, and the rendering category display indicates a certain rendering category among a set of rendering categories, the metadata receiver; A receiver that receives user head movement data indicating the movement of the user's head; A mapper that maps the input pose to a rendering pose within a rendering coordinate system in response to the user head movement data, wherein the rendering coordinate system is fixed with respect to the movement of the head, the mapper; A renderer that renders the audiovisual item using the rendering pose, an audiovisual rendering device comprising: Each rendering category is linked to a coordinate system transformation from a real-world coordinate system to a category coordinate system, the coordinate system transformation being different for each category, and at least one category coordinate system is variable with respect to the real-world coordinate system and the rendering coordinate system; The mapper selects, in response to a rendering category display of a first audiovisual item, a first rendering category of the first audiovisual item from the set of rendering categories, and maps the input pose of the first audiovisual item to a rendering pose within the rendering coordinate system corresponding to a fixed pose within a first category coordinate system for changing the movement of the user's head, the first category coordinate system being determined from a first coordinate system transformation of the first rendering category, the audiovisual rendering device.

2. The audiovisual rendering device according to claim 1, wherein a second coordinate system transformation of a second category aligns the category coordinate system of the second category with the movement of the user's head.

3. The audiovisual rendering device according to claim 1 or 2, wherein a third coordinate system transformation of a third category aligns the category coordinate system of the third category with the real-world coordinate system.

4. The audiovisual rendering device according to any one of claims 1 to 3, wherein the first coordinate system transformation depends on the user head movement data.

5. The visual rendering device according to claim 4, wherein the first coordinate system transformation depends on the average head pose.

6. The visual rendering device according to claim 5, wherein the first coordinate system transformation aligns the first category coordinate system with the average head pose.

7. The visual rendering device according to any one of claims 4 to 6, wherein different coordinate system transformations for different rendering categories depend on the user head movement data, and the dependencies of the first coordinate system transformation and the different coordinate system transformations on the movement of the user's head have different time-average characteristics.

8. The visual rendering device according to any one of claims 1 to 7, further comprising a receiver that receives user body pose data indicating a user body pose, wherein the first coordinate system transformation depends on the user body pose data.

9. The visual rendering device according to claim 8, wherein the first coordinate system transformation aligns the first category coordinate system with the user body pose.

10. The visual rendering device according to any one of claims 1 to 9, further comprising a receiver that receives device pose data indicating the pose of an external device, wherein the first coordinate system transformation depends on the device pose data.

11. The visual rendering device according to claim 10, wherein the first coordinate system transformation aligns the first category coordinate system with the device pose.

12. The visual rendering device according to any one of claims 1 to 11, wherein the mapper selects the first rendering category in response to user movement parameters indicating the movement of the user.

13. The visual rendering device according to any one of claims 1 to 12, wherein the mapper determines a coordinate system transformation between the real-world coordinate system and the coordinate system of the user head movement data in response to user movement parameters indicating the movement of the user.

14. The visual rendering device according to claim 12 or 13, wherein the mapper determines the user movement parameters in response to the user head movement data.

15. The visual rendering device according to any one of claims 1 to 14, wherein at least some of the rendering category displays indicate whether the visual items of the at least some of the rendering category displays are digital visual items or non-digital visual items.

16. The visual rendering device according to any one of claims 1 to 15, wherein the visual item is an audio item, and the renderer generates an output binaural audio signal for a binaural rendering device by applying binaural rendering to the audio item using the rendering pose.

17. A method of rendering a visual item, the method comprising: receiving the visual item; receiving metadata including an input pose for each of at least some of the visual items and a rendering category display for each of at least some of the visual items, wherein the input pose is provided with reference to an input coordinate system and the rendering category display indicates a certain rendering category among a set of rendering categories; receiving user head movement data indicating movement of the user's head; mapping the input pose to a rendering pose within a rendering coordinate system in response to the user head movement data, wherein the rendering coordinate system is fixed with respect to the movement of the head; rendering the visual item using the rendering pose. Each rendering category is linked to a coordinate system transformation from a real-world coordinate system to a category coordinate system, the coordinate system transformation being different for each category, and at least one category coordinate system is variable with respect to the real-world coordinate system and the rendering coordinate system. The method includes a step of selecting, in response to a rendering category display of a first audiovisual item, a first rendering category of the first audiovisual item from a set of the rendering categories, and mapping an input pose of the first audiovisual item to a rendering pose in a rendering coordinate system corresponding to a fixed pose in a first category coordinate system for changing a movement of a user's head, wherein the first category coordinate system is determined from a first coordinate system transformation of the first rendering category, the method.

18. A computer program including computer program code means for performing all steps of the method according to claim 17 when the program is executed on a computer.

Citation Information

Patent Citations

  • Distance panning with near / far rendering

    JP2019523913A

  • Audio device and audio processing method

    JP2022502886A

  • US2018/91922A1

  • Associated spatial audio playback

    WO2019141900A1

  • Spatial audio augmentation

    WO2020012067A1