Audiovisual rendering apparatus and method of operation therefor
The system addresses the challenge of adapting audiovisual rendering to user head movement by using coordinate system transformations for each rendering category, enhancing immersion and user experience in virtual and augmented reality applications.
Patent Information
- Application Number
- JP2025053714
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-10-13
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-10-11
AI Technical Summary
Existing virtual and augmented reality applications struggle to provide an optimal user experience due to difficulties in adapting audiovisual rendering to the movement of the user's head, leading to reduced immersion and unnatural experiences.
A system that includes a receiver for audiovisual items, metadata for rendering categories, and a mapper that adjusts the rendering pose based on user head movement data, using coordinate system transformations specific to each rendering category to maintain or vary the position of audiovisual items relative to the user's head or real-world environment.
This approach enhances user experience by providing flexible and immersive audiovisual rendering, reducing complexity, and improving the perception of spatial audio and visual scenes, while allowing for individual customization and personalization.
Smart Images

Figure 2025102862000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a visual and auditory rendering apparatus and an operation method thereof, and more particularly, although not limited thereto, to the use of these for supporting extended / virtual reality applications and the like.
Background Art
[0002] In recent years, the diversity and scope of experiences based on visual and auditory content have increased significantly, and new services and methods for utilizing and consuming such content have been continuously developed and introduced. In particular, many spatial interactive services, applications, and experiences have been developed to provide users with more engaging and immersive experiences.
[0003] Examples of such applications include virtual reality (VR) applications, augmented reality (AR) applications, and mixed reality (MR) applications that are rapidly becoming mainstream, and many solutions are targeting the consumer market. Also, various standards have been developed by various standardization bodies. Such standardization activities are actively developing standards for various aspects of VR / AR / MR systems, such as streaming, broadcasting, and rendering.
[0004] VR applications tend to provide a user experience corresponding to users in different worlds / environments / scenes, while AR (including mixed reality MR) applications tend to provide a user experience corresponding to users in the current environment with additional information or virtual objects or information added. Therefore, VR applications tend to provide a completely immersive artificial world / scene, while AR applications tend to provide a partially artificial world / scene superimposed on the real scene where the user physically exists. However, these terms are often used synonymously and largely overlap. Hereinafter, the term virtual reality / VR is used to represent both virtual reality and augmented reality.
[0005] As an example, what is becoming increasingly popular is a service that provides images and sounds so that users can interact with the system actively and dynamically in order to change the rendering parameters to adapt to changes in the movement, position, and orientation of the user. In many applications, functions that change the viewer's de facto perspective and line of sight, for example, functions that enable the viewer to move within the presented scene and "look around", are very attractive.
[0006] Such functions can, in particular, provide users with a virtual reality experience. This can enable the user to move around (relatively) freely within the virtual environment and dynamically change their position and line of sight. Usually, such virtual reality applications are based on a three-dimensional model of the scene, and the model is dynamically evaluated to provide a specific required view. This approach is well known in game applications such as the first-person shooting game genre for computers and consoles.
[0007] Also, especially in the case of virtual reality applications, it is desirable for the presented image to be a three-dimensional image. In practice, in order to optimize the viewer's sense of immersion, it is usually preferred that the user experiences the presented scene as a three-dimensional scene. In fact, a virtual reality experience preferably enables the user to select their position, the camera's perspective, and the point in time with respect to the virtual world.
[0008] In addition to visual rendering, most VR / AR applications further provide a corresponding audio experience. In many applications, the audio preferably provides a spatial audio experience where the sound source is recognized as arriving from a position corresponding to the position of the corresponding object (including both the object currently in view and the object not currently in view (e.g., behind the user)) within the visual scene. Therefore, the audio scene and the video scene should be consistent and preferably be perceived as both providing a complete spatial experience.
[0009] Regarding audio, conventionally, focus has been on headphone playback using binaural audio rendering technology. In many scenarios, headphone playback can provide users with a very immersive and personalized experience. Using head tracking enables rendering according to the movement of the user's head, significantly improving the immersion.
[0010] For IVAS (Immersive Voice and Audio Services), the 3GPP consortium is developing a so-called IVAS codec (3GPP SP-170611 ‘New WID on EVS Codec Extension for Immersive Voice and Audio Services’). This codec includes a renderer that converts various audio streams into a format suitable for playback on the receiving side. Specifically, for playback using headphones or a head-mounted VR device with built-in headphones, the audio can be put into a binaural format.
[0011] In many such applications, the rendering device may receive input data representing three-dimensional audio and / or visual scenes. The renderer may be configured to render this data such that an audiovisual experience providing a perception of a three-dimensional scene is provided to the user.
[0012] However, providing an appropriate experience is difficult in many applications, and in particular, it is difficult to adapt the rendering according to the movement of the head so that a desired experience is provided to the user.
[0013] For example, human perception of the direction and distance of a sound source depends not only on the (usually different) delays and filtering of the sound from the sound source to both ears, but also strongly on how these change when the head is moved, for example rotated. Similarly, the parallax and similar motion of visual objects provide strong three-dimensional visual cues. Unconsciously, we move our heads (usually slightly) or jiggle them slightly in our daily lives, so that the sound also changes, albeit slightly but clearly, which greatly contributes to the immersive "around us" auditory / visual experience we are familiar with.
[0014] In headphone playback experiments, even if the sound path from the sound source to the ear is properly modeled by a filter, making these static (i.e., by having no changes related to head movement) can reduce the sense of immersion and the sound may be felt as "in the head".
[0015] Therefore, several applications have been developed that render sound sources and / or visual objects at positions that are perceived to be fixed relative to the real world in order to create the impression of an immersive virtual world. However, it is difficult to execute this optimally and it does not always result in a desirable user experience. In some applications, a three-dimensional scene that follows the movement of the head is displayed, thus appearing to be fixed relative to the user's head. This may be a desirable experience in many applications, but in other applications it may provide an unnatural experience, for example, it may not achieve the immersive experience of "being present" within the virtual scene. US10015620B2 discloses another example where the sound is rendered with respect to a reference orientation representing the user.
[0016] However, while such applications may provide an appropriate user experience in many embodiments, some applications tend not to provide an optimal, and thus desirable, user experience.
[0017] Therefore, an improved method for rendering audiovisual items, particularly audiovisual items for virtual / augmented / composite reality experiences / applications, would be beneficial. In particular, methods that enable improved performance, increased flexibility, reduced complexity, easier implementation, improved user experience, more consistent perception of audio and / or visual scenes, improved customization, improved personalization, improved virtual reality experience, and / or improved performance and / or operation would be beneficial. SUMMARY OF THE INVENTION
[0018] Accordingly, the present invention aims to suitably mitigate, reduce, or eliminate one or more of the above drawbacks, either alone or in any combination.
[0019] According to one aspect of the present invention, a first receiver for receiving audiovisual items, a metadata receiver for receiving metadata including an input pose and a rendering category display for each of at least some of the audiovisual items, wherein the input pose is provided with reference to an input coordinate system, and the rendering category display indicates a certain rendering category among a set of rendering categories, a receiver for receiving user head movement data indicating the movement of the user's head, and in response to the user head movement data, a mapper for mapping the input pose to a rendering pose within a rendering coordinate system, wherein the rendering coordinate system is fixed with respect to the movement of the head, a renderer for rendering the audiovisual item using the rendering pose, and an audiovisual rendering device comprising: each rendering category is linked to a coordinate system transformation from a real-world coordinate system to a category coordinate system, the coordinate system transformation is different for each category, at least one category coordinate system is variable with respect to the real-world coordinate system and the rendering coordinate system, and the mapper selects a first rendering category of the first audiovisual item from the set of rendering categories in response to the rendering category display of the first audiovisual item, and maps the input pose of the first audiovisual item to a rendering pose within the rendering coordinate system corresponding to a fixed pose within a first category coordinate system for changing the movement of the user's head, and the first category coordinate system is determined from a first coordinate system transformation of the first rendering category. An audiovisual rendering device is provided.
[0020] This approach can provide an improved user experience in many embodiments, specifically improving the user experience of many virtual reality (including augmented and mixed reality) applications, particularly those involving social or shared experiences. This approach can provide a very flexible approach where the rendering operation and spatial recognition of visual items can be individually adapted to individual visual items. This approach can, for example, render some visual items so that they seem to be completely fixed with respect to the real world, render some visual items so that they seem to be completely fixed to the user (tracking the user's head movement), and render some visual items so that they seem to be fixed with respect to the real world for some movements and tracking the user for other movements. This approach can, in many embodiments, enable flexible rendering where visual items are perceived to substantially follow the user, yet still provide a spatial out-of-head localization experience of the visual items.
[0021] This approach reduces complexity and resource requirements in many embodiments and enables source-side control of the rendering operation in many embodiments.
[0022] The rendering category can indicate whether the visual item represents a sound source having spatial characteristics fixed to the head orientation or spatial characteristics not fixed to the head orientation (corresponding to listener pose-dependent position and listener pose-independent position, respectively). The rendering category can indicate whether the audio element is diegetic or not.
[0023] In many embodiments, the mapper may further select a second rendering category of the second visual item from the set of rendering categories in response to a rendering category display of the second visual item, and map an input pose of the second visual item to a rendering pose in a rendering coordinate system corresponding to a fixed pose in a second category coordinate system for changing the movement of the user's head, where the second category coordinate system may be determined from a second coordinate system transformation of the second rendering category. The mapper can similarly be configured to perform such operations on third, fourth, fifth, and other visual items.
[0024] The visual item may be an audio item and / or a visual / video / image / scene item. The visual item may be a visual or audio representation of a scene object of a scene represented by the visual item. In some embodiments, the term "visual item" can be replaced with the term "audio item (or element)". In some embodiments, the term "visual item" can be replaced with the term "visual item (or scene object)".
[0025] In many embodiments, the renderer may be configured to generate an output binaural audio signal for a binaural rendering device by applying binaural rendering to a visual item (which is an audio item) using the rendering pose.
[0026] The term "pose" may represent a position and / or an orientation. In some embodiments, the term "pose" can be replaced with the term "position". In some embodiments, the term "pose" can be replaced with the term "orientation". In some embodiments, the term "pose" can be replaced with the term "position and orientation".
[0027] The receiver may receive real-world user head movement data based on a real-world coordinate system indicating the movement of the user's head.
[0028] The renderer may render the visual item using the rendering pose, and the visual item is referenced within / arranged within the rendering coordinate system.
[0029] The mapper may map the input pose of the first visual item to a rendering pose within the rendering coordinate system corresponding to a fixed pose within the first category coordinate system of the head movement data indicating the different and varying head movements of the user.
[0030] In some embodiments, the rendering category display indicates a source type, such as a voice type for a voice item or a scene object type for a visual element.
[0031] This may result in an improvement in the user experience in many embodiments. The rendering category display may indicate a certain sound source type among a sound source type set including at least one sound source type selected from the group of utterance audio, music audio, foreground audio, background audio, voiceover audio, and narrator audio.
[0032] According to an optional feature of the present invention, the second coordinate system transformation of the second category aligns the category coordinate system of the second category with the movement of the user's head.
[0033] This may result in an improvement in the user experience and / or an improvement in performance and / or an ease of implementation in many embodiments. In particular, it may support that some visual items are felt to be fixed relative to the user's head while other items are not.
[0034] According to an optional feature of the present invention, the third coordinate system transformation of the third category aligns the category coordinate system of the third category with the real-world coordinate system.
[0035] This can bring about an improvement in user experience and / or an improvement in performance and / or an ease of implementation in many embodiments. In particular, it can support that some visual items may seem to be fixed with respect to the real world, while other items may not seem so.
[0036] According to an optional feature of the present invention, the first coordinate system transformation depends on user head movement data.
[0037] This can bring about an improvement in user experience and / or an improvement in performance and / or an ease of implementation in many embodiments. In particular, in some applications, it provides a very advantageous experience where the pose of a visual item follows the overall movement of the user but not small / fast movements of the head, thereby providing an improved user experience along with an improved extra - cranial localization experience.
[0038] According to an optional feature of the present invention, the first coordinate system transformation depends on the average head pose.
[0039] According to an optional feature of the present invention, the first coordinate system transformation aligns the first category coordinate system with the average head pose.
[0040] According to an optional feature of the present invention, different coordinate system transformations for different rendering categories depend on user head movement data, and the dependencies of the first coordinate system transformation and the different coordinate system transformations on the user's head movement have different time - averaging characteristics.
[0041] According to an optional feature of the present invention, the visual rendering device further comprises a receiver for receiving user body pose data indicating the user body pose, and the first coordinate system transformation depends on the user body pose data.
[0042] This can result in an improvement in user experience and / or performance and / or ease of implementation in many embodiments. In particular, in some applications, it provides a highly advantageous experience where the pose of the visual item follows the user's overall movement but not small / fast head movements, thereby providing an improved user experience along with an improved out-of-head localization experience.
[0043] According to an optional feature of the present invention, the first coordinate system transformation depends on the average torso pose.
[0044] According to an optional feature of the present invention, the first coordinate system transformation aligns the first category coordinate system with the user torso pose.
[0045] According to an optional feature of the present invention, the visual rendering device further comprises a receiver that receives device pose data indicating the pose of an external device, and the first coordinate system transformation depends on the device pose data.
[0046] This can result in an improvement in user experience and / or performance and / or ease of implementation in many embodiments. In particular, in some applications, it provides a highly advantageous experience where the pose of the visual item follows the user's overall movement but not small / fast head movements, thereby providing an improved user experience along with an improved out-of-head localization experience.
[0047] According to an optional feature of the present invention, the first coordinate system transformation depends on the average device pose.
[0048] According to an optional feature of the present invention, the first coordinate system transformation aligns the first category coordinate system with the device pose.
[0049] According to an optional feature of the present invention, the mapper selects a first rendering category in response to user movement parameters indicating the user's movement.
[0050] According to an optional feature of the present invention, the mapper determines a coordinate system transformation between a real-world coordinate system and a coordinate system of user head movement data in response to user movement parameters indicating the movement of the user.
[0051] According to an optional feature of the present invention, the mapper determines user movement parameters in response to data of user head movement.
[0052] According to an optional feature of the present invention, at least some of the rendering category displays indicate whether the visual items of at least some of the rendering category displays are digital visual items or non-digital visual items.
[0053] According to an optional feature of the present invention, the visual item is an audio item, and the renderer generates an output binaural audio signal for a binaural rendering device by applying binaural rendering to the audio item using a rendering pose.
[0054] According to another aspect of the present invention, there is provided a method of rendering a visual item, the method comprising receiving the visual item; receiving metadata including an input pose and a rendering category display for each of at least a part of the visual item, wherein the input pose is provided with reference to an input coordinate system and the rendering category display indicates a rendering category among a set of rendering categories; receiving user head movement data indicating movement of the user's head; in response to the user head movement data, mapping the input pose to a rendering pose within a rendering coordinate system, the rendering coordinate system being fixed with respect to the movement of the head; and rendering the visual item using the rendering pose, each rendering category being linked to a coordinate system transformation from a real-world coordinate system to a category coordinate system, the coordinate system transformation being different for each category, at least one category coordinate system being variable with respect to the real-world coordinate system and the rendering coordinate system, the method including, in response to the rendering category display of the first visual item, selecting a first rendering category of the first visual item from the set of rendering categories and mapping the input pose of the first visual item to a rendering pose within the rendering coordinate system corresponding to a fixed pose within a first category coordinate system for changing the movement of the user's head, the first category coordinate system being determined from a first coordinate system transformation of the first rendering category.
[0055] The above and other aspects, features, and advantages of the present invention will be described and will become apparent with reference to the embodiments described below.
Brief Description of the Drawings
[0056] Hereinafter, embodiments which are merely examples of the present invention will be described with reference to the following drawings.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
[0057] The following description focuses on embodiments in which an audiovisual item, including both audio and visual items, is presented by a rendering that includes both audio and visual renderings, but it should be understood that the techniques and principles described may also be applied separately, for example to only audio renderings of an audio item, or only video / visual / image renderings of a visual item (e.g., a visual object in a video scene).
[0058] Also, although the description focuses on virtual reality applications, it will be appreciated that the described techniques can be used in many other applications, including augmented and mixed reality applications.
[0059] Virtual reality (including augmented and mixed reality) experiences that allow users to move around in virtual or augmented worlds are becoming increasingly popular, and services are being developed to meet such demand. In many such approaches, visual and audio data can be dynamically generated to reflect the current pose of the user (or viewer).
[0060] In this technical field, the terms "placement" and "pose" are used as general terms representing position and / or orientation / direction. For example, a combination of the position and orientation / direction of an object, a camera, a head, or a view can be called a pose or a placement. Thus, an indicator of a placement or a pose can have up to six values / components / degrees of freedom. Each value / component typically describes an individual characteristic of the position / location or orientation / direction of the corresponding object. Of course, in many situations, a placement or a pose may be represented by fewer components, for example, when one or more components are considered to be constant or irrelevant (e.g., when all objects are considered to have the same height and a horizontal orientation, it may be possible to fully represent the pose of an object with four components). Hereinafter, the term "pose" is used to refer to a position and / or orientation that can be represented by one to six values (corresponding to the maximum possible degrees of freedom).
[0061] Many VR applications are based on a pose having the maximum degrees of freedom (i.e., three degrees of freedom each for position and orientation, for a total of six degrees of freedom). Thus, a pose can be represented by a set or vector of six values representing six degrees of freedom, and thus, a pose vector can provide an indicator of three-dimensional position and / or three-dimensional orientation. However, it should be understood that in other embodiments, a pose may be represented by a smaller number of values.
[0062] A system or entity based on providing the viewer with the maximum degrees of freedom is typically said to have six degrees of freedom (6DoF). Many systems and entities provide only orientation or position, and these are typically said to have three degrees of freedom (3DoF).
[0063] Typically, a virtual reality application generates a three-dimensional output in the form of separate view images for the left and right eyes. These images can then be presented to the user by appropriate means, such as the (usually individual) left and right displays of a VR headset. In other embodiments, one or more view images may be presented, for example, on a naked-eye stereoscopic display, or in some embodiments, only one two-dimensional image may be generated (e.g., using a conventional two-dimensional display).
[0064] Similarly, for a given viewer / user / listener pose, an audio representation of the scene can be provided. Typically, the audio scene is rendered to provide a spatial experience where the sound source is perceived to be emanating from the desired location. Since the sound source may be static within the scene, as the user's pose changes, the relative position of the sound source with respect to the user's pose changes. Thus, the spatial perception of the sound source can change to reflect the new position relative to the user. Accordingly, the audio rendering can be adapted according to the user's pose.
[0065] Viewer or user pose input can be determined in various ways in various applications. In many embodiments, the user's physical movements can be directly tracked. For example, a camera surveying the user area can detect and track the user's head (or eyes (eye tracking)). In many embodiments, the user can wear a VR headset that can be tracked by external and / or internal means. For example, the headset can include an accelerometer and a gyroscope that provide information about the movement and rotation of the headset (and thus the head). In some examples, the VR headset may transmit a signal or include an identifier (e.g., visual) that enables an external sensor to determine the position and orientation of the VR headset.
[0066] In some systems, VR applications can be provided locally to viewers by a stand-alone device that, for example, does not use (or, in some cases, even have access to) any remote VR data or processing. For example, a device such as a game console can include a store for storing scene data, an input device for receiving / generating the viewer's pose, and a processor for generating corresponding images and / or audio from the scene data.
[0067] In other systems, VR / scene data can be provided from a remote device or server.
[0068] For example, a remote device can generate audio data representing an audio scene and transmit audio components / objects / signals corresponding to a plurality of different sound sources within the audio scene, or other audio elements, together with position information indicating their positions (which can change dynamically in the case of moving objects). The audio elements can include elements associated with a specific position, but can also include elements for more distributed or diffuse sound sources. For example, audio elements representing general (non-localized) background sound, ambient sound, reverberation, etc. can be provided.
[0069] In that case, the local VR device can appropriately render the audio elements, for example, by applying appropriate binaural processing that reflects the relative positions of the sound sources of the audio components.
[0070] Similarly, a remote device can generate visual / video data representing a visual scene and transmit visual components / objects / signals corresponding to a plurality of different objects within the visual scene, or other visual elements, together with position information indicating their positions (which can change dynamically in the case of moving objects). The visual items can include elements associated with a specific position, but can also include video items for more distributed sources.
[0071] In some embodiments, the visual items may be provided as individual discrete items, e.g., provided as descriptions of individual scene objects (e.g., dimensions, textures, opacities, reflectivities). Alternatively or additionally, the visual items may be represented as part of an overall model of the scene, which may include, for example, descriptions of multiple different objects and their interrelationships.
[0072] Thus, in the case of VR services, the central server in some embodiments can generate audiovisual data representing a three-dimensional scene, specifically, represent sound by a plurality of audio items renderable by a local client / device, and represent a visual scene by a plurality of video items.
[0073] FIG. 1 shows an example of a VR system in which a central server 101 communicates with a plurality of remote clients 103 via a network 105 such as the Internet. The central server 101 can be configured to support a plurality of remote clients 103 that may be numerous simultaneously.
[0074] Such an approach can, in many scenarios, improve the trade-off between, for example, complexity and the resource and communication requirements of various devices. For example, the scene data can be sent once or at a relatively low frequency, and the local rendering device (remote client 103) can receive the viewer's pose, locally process the scene data, and render audio and / or video to reflect the change in the viewer's pose. This approach can provide an efficient system and an appealing user experience. This can, for example, provide a low-latency real-time experience, enable centralized storage, generation, and management of scene data, and significantly reduce the required communication bandwidth. For example, it may be suitable for applications where a VR experience is provided to multiple remote devices.
[0075] Figure 2 shows elements of a visual rendering apparatus that can provide improved visual rendering in many applications and scenarios. In particular, the visual rendering apparatus may provide improved rendering for many VR applications, and the visual rendering apparatus may be specifically configured to perform the processing and rendering of the VR client 103 of FIG. 1.
[0076] The audio device of FIG. 2 is configured to render a three-dimensional scene by rendering spatial audio and video to provide a three-dimensional perception of the scene. In a specific description of the visual rendering apparatus, the focus is on an application that provides a described technique for both audio and video, but it should be understood that in other embodiments, the technique may be applied only to audio or video / visual processing. In fact, in some embodiments, the rendering apparatus may include only a function for rendering audio or only a function for rendering video, that is, the visual rendering apparatus may be any audio rendering apparatus or any video rendering apparatus.
[0077] The visual rendering apparatus includes a first receiver 201 configured to receive visual items from a local or remote source. In this specific example, the first receiver 201 receives data representing visual items from the server 101. The first receiver 201 may be configured to receive data representing a virtual scene. The data may include data providing a visual description of the scene and may also include data providing an audio description of the scene. Thus, the received data may provide an audio scene description and a visual scene description.
[0078] The audio item may be encoded audio data, such as an encoded audio signal. The audio item can be various types of audio elements, such as various types of signals and components. In fact, in many embodiments, the first receiver 201 may receive audio data defining various audio types / formats. For example, the audio data may include audio represented by an audio channel signal, individual audio objects, scene-based audio (e.g., HOA (Higher Order Ambisonics)), etc. The audio may be represented, for example, as encoded audio for a given audio component to be rendered.
[0079] The first receiver 201 is coupled to a renderer 203, and the renderer renders a scene based on received data representing the audiovisual item. In the case of encoded data, the renderer 203 may also be configured to decode the data (or, in some embodiments, the decoding may be performed by the first receiver 201).
[0080] Specifically, the renderer 203 may include an image renderer 205 configured to generate an image corresponding to the viewer's current viewing pose. For example, the data includes spatial 3D image data (e.g., an image of the scene and depth or model description), and the visual renderer 203 can generate stereoscopic images (images for the user's left and right eyes) from this data, as is known to those skilled in the art. These images can be presented to the user, for example, by individual left-eye and right-eye displays of a VR headset.
[0081] The renderer 203 further includes an audio renderer 207 configured to render an audio scene by generating an audio signal based on an audio item. In this example, the audio renderer 207 is a binaural audio renderer that generates binaural audio signals for the user's left and right ears. The binaural audio signals are generated to provide a desired spatial experience and are typically reproduced by headphones or earphones that can be part of a headset worn by the user. The headset also includes a left-eye display and a right-eye display.
[0082] Accordingly, in many embodiments, the audio rendering by the audio renderer 207 is a binaural rendering process that uses an appropriate binaural transfer function to provide a desired spatial effect to a user wearing headphones. For example, the audio renderer 207 can be configured to use binaural processing to generate audio components that are perceived as arriving from a specific location.
[0083] Binaural processing is known to be used to provide a spatial experience by virtually placing sound sources using individual signals for each ear of the listener. By using an appropriate binaural rendering process, the signals required at the eardrums for the listener to perceive sound from any direction can be calculated, and the signals can be rendered to provide the desired effect. These signals are reformed at the eardrums using headphones or crosstalk cancellation (suitable for rendering with speakers in close proximity to each other). Binaural rendering can be considered a technique that tricks the human auditory system into generating signals for the listener's ears such that the sound is perceived as coming from the desired location.
[0084] Binaural rendering is based on binaural transfer functions that vary from person to person due to the acoustic characteristics of the head, ears, and reflecting surfaces (e.g., shoulders). For example, binaural recordings can be created that simulate multiple sources at various locations using binaural filters. This can be achieved by convolving each sound source with a pair such as the head-related impulse response (HRIR) corresponding to the position of the sound source.
[0085] A well-known method for determining the binaural transfer function is binaural recording. This is a recording method using a dedicated microphone arrangement and is intended for playback with headphones. The recording is performed by placing microphones in the ear canals of the subject or using a dummy head with built-in microphones, i.e., a half-body model including the pinna (outer ear). The use of such a dummy head including the pinna provides a very similar spatial impression as if the person listening to the recording was present at the scene during recording.
[0086] For example, by measuring the response from a sound source at a specific position in 2D or 3D space to a microphone placed in or near a human ear, an appropriate binaural filter can be determined. Based on such measurements, a binaural filter reflecting the acoustic transfer function to the user's ear can be generated. Binaural recordings can be created that simulate multiple sources at various locations using binaural filters. This can be achieved, for example, by convolving each sound source with a pair of measured impulse responses at the desired sound source positions. To create the illusion that the sound source is moving around the listener, usually a large number of binaural filters with a specific spatial resolution (e.g., 10 degrees) are required.
[0087] The head binaural transfer function can be expressed, for example, as a head impulse response (HRIR), or equivalently as a head-related transfer function (HRTF), a binaural room impulse response (BRIR), or a binaural room transfer function (BRTF). The transfer function (e.g., estimated or assumed) from a given position to the listener's ears (or eardrums) is given, for example, in the frequency domain, in which case it is usually called an HRTF or BRTF, and in the time domain, it is usually called an HRIR or BRIR. In some cases, the head binaural transfer function is determined to include the acoustic environment, particularly the characteristics or properties of the room in which the measurement is made, while in other cases only the characteristics of the user are considered. Examples of the first type of function are BRIR and BRTF.
[0088] Accordingly, the audio renderer 207 comprises a store having binaural transfer functions for a plurality of different positions, which are usually numerous. Each binaural transfer function provides information on how the audio signal should be processed / filtered so that the audio signal is perceived to have originated from that position. By applying binaural processing individually to a plurality of audio signals / sound sources and combining the results, an audio scene can be generated that includes a plurality of sound sources arranged at appropriate positions within the sound stage.
[0089] The audio renderer 207 can select and retrieve (or, in some cases, interpolate between a plurality of nearby binaural transfer functions) the stored binaural transfer function that most closely matches the desired position for a given audio element perceived to have originated from a given position relative to the user's head. The audio renderer can then apply the selected binaural transfer function to the audio signal of the audio element, thereby generating an audio signal for the left ear and an audio signal for the right ear.
[0090] The output stereo signals generated in the form of left ear and right ear signals are suitable for rendering on headphones and can be amplified to generate drive signals for supply to the user's headset. Then, the user perceives that the audio element is originating from the desired position.
[0091] It should be understood that the audio item, in some embodiments, can be processed, for example, to add acoustic environment effects. For example, the audio item can be processed to add reverberation or decorrelation / diffusivity, etc. In many embodiments, this processing can be performed on the generated binaural signal rather than directly on the audio element signal.
[0092] Thus, the audio renderer 207 can be configured to generate an audio signal such that a user wearing headphones perceives that the audio element is received from a desired position. Other audio items can be, for example, dispersed and diffused and rendered as such, depending on the case.
[0093] Many algorithms and techniques for rendering spatial audio, particularly binaural rendering (e.g., using headphones), are known to those skilled in the art, and it will be understood that any suitable technique can be used without detracting from the present invention.
[0094] The audiovisual rendering device further comprises a second receiver 209, which is a metadata receiver configured to receive metadata of the audiovisual item. The metadata particularly includes one or more position data of the audiovisual item. The metadata can include an input position indicating one or more positions of the audiovisual item.
[0095] The received audiovisual data may include audio and / or visual data that describes a scene. The feed core data specifically includes audiovisual data regarding a set of audiovisual items corresponding to sound sources and / or visual objects within the scene. Some audio items may represent sound sources located within the scene that are associated with a specific position and / or orientation within the scene (in the case of a moving object, the position and / or orientation may change dynamically). The visual data may include data that describes scene objects that generate these visual representations and enable them to be represented by an image / video presented to the user (typically, a 3D image using an individual display of a headset).
[0096] Often, an audio element may represent audio generated by a specific scene object within a virtual scene, and thus may represent a sound source at a position corresponding to the position of the scene object (e.g., a person's voice). In such cases, the same position data / display may be included and used for both the audio item and the corresponding visual scene object (similarly for orientation).
[0097] Other elements may represent a plurality of more dispersed or diffused sound sources, such as ambient noise or background noise that can spread. As another example, some audio elements may fully or partially represent spatially undefined components of sound from a position-specific sound source, such as reverberation spreading from a spatially well-defined sound source.
[0098] Similarly, some visual scene objects may have an extended position. For example, the position data may indicate the center or reference position of the scene object.
[0099] The metadata may include pose data indicating the position and / or orientation of the audiovisual item, specifically the position and / or orientation of the sound source and / or visual scene object or element. The pose data may include, for example, absolute position and / or orientation data that defines the position of each item or at least some of the items.
[0100] The pose is provided with reference to an input coordinate system, i.e., this input coordinate system is the reference coordinate system for the pose display provided within the metadata received by the second receiver. The input coordinate system is typically fixed with reference to the scene to be represented / rendered. For example, the scene may be a virtual (or actual) scene in which there are sound sources and scene objects at positions provided with reference to the scene. That is, the input coordinate system is typically the scene coordinate system of the scene represented by the audiovisual data.
[0101] To provide a user representation of the scene, the scene is rendered from the pose of the viewer or user. That is, the scene is rendered as it would be perceived at a particular viewer / user pose within the scene, and the audio and visual rendering provide the audio and images that would be perceived at that viewer's pose.
[0102] The rendering by the renderer 203 is performed with respect to a rendering coordinate system fixed with respect to the user's head and head movement. The playback of the rendered audio signal is typically a playback device worn or attached to the head, such as headphones / earphones and individual displays for each eye. Usually, the playback is performed by a headset device having audio and video playback means. The renderer is configured to generate audiovisual items rendered with reference to the playback device / means, and it is assumed that the user's pose is fixed / constant with reference to the playback device / means. For example, the position of the sound source is determined relative to the headphones (i.e., the position within the rendering coordinate system), an appropriate HRTF filter is retrieved, and it is used to render the audio signal such that the sound source is recognized as arriving from the required relative position within the rendering coordinate system. Similarly, for a given scene object, the relative position with respect to the display (i.e., the position within the rendering coordinate system) is determined, and the images corresponding to the views from the left and right eyes for this position are determined respectively.
[0103] Therefore, the rendering coordinate system is considered to be a coordinate system fixed with respect to the user's head, and specifically, it can be considered to be a rendering coordinate system that does not depend on the movement or pose change of the user's head. The playback device (audio, visual, or both audio and visual) is assumed / considered to be fixed with respect to the user's head, and thus with respect to the rendering coordinate system.
[0104] The rendering coordinate system fixed with respect to the movement of the user's head can be considered to correspond to the playback coordinate system fixed with respect to the playback device for playing back the rendered audiovisual item. The term rendering coordinate system may be synonymous with the coordinate system of the playback device / means, and may be replaced by these terms. Similarly, the term "rendering coordinate system fixed with respect to the movement of the user's head" may be synonymous with "playback device / means coordinate system fixed with respect to the playback device for playing back the rendered audiovisual item", and may be replaced by this.
[0105] Since the renderer performs rendering based on the pose with reference to the rendering coordinate system and the pose for the audiovisual item is provided with reference to the input coordinate system, the audiovisual rendering device includes a mapper 211 configured to map the input position in the input coordinate system to the rendering position in the rendering coordinate system.
[0106] The audiovisual rendering device includes a head movement data receiver 213 for receiving user head movement data indicating the movement of the user's head. The user head movement data may indicate the movement of the user's head in the real world and is usually provided with reference to the coordinate system of the real world. The head movement data may indicate the absolute or relative movement of the user's head in the real world, and specifically, may reflect the absolute or relative change in the user's pose with respect to the coordinate system of the real world. The head movement data may indicate the change (or no change) in the pose (orientation and / or position) of the head and may also be referred to as head pose data.
[0107] Many different techniques are known for detecting and representing head movement, and it will be understood that any suitable technique can be used without detracting from the present invention. Specifically, the head movement data receiver 213 can receive head movement data from a VR headset or VR head movement detector, as is known in the art.
[0108] The mapper 211 is coupled to the head movement data receiver 213 and receives user head movement data. The mapper is configured to perform a mapping between an input position in an input coordinate system and a rendering position in a rendering coordinate system in response to the user head movement data. For example, the mapper 211 can continuously process the user head movement data to continuously track the current pose of the user in the real-world coordinate system. Then, a mapping between the input pose and the rendering pose can be performed based on the user pose.
[0109] For example, in many applications, it is desirable to provide the user with an experience as if they were within the three-dimensional scene being presented. Therefore, it is desirable for the rendered audio and images to reflect the pose of the user that follows the movement of the user's head. Thus, it is desirable for the audiovisual items to be rendered such that they are perceived as being fixed relative to the movement in the real world. This is because the movement in the real world will be reproduced in the rendering of the (usually virtual) scene.
[0110] In such a case, the mapping of the input pose to the rendering pose is a mapping as if the visual item were fixed with respect to the real world, i.e., it is rendered so as to be recognized as fixed with respect to the real world. Accordingly, the same input pose is mapped to different rendering poses so that changes in the pose of the user's head are reflected. For example, when the user rotates the head by 30°, the real-world scene is based on the user rotated by -30°. The mapper 211 can perform corresponding changes so that the mapping from the input pose to the rendering pose is modified to include an additional 30° rotation with respect to the situation before the rotation of the user's head. As a result, the visual item will be in a different pose in the rendering coordinate system but is recognized as being in the same real-world pose. Accordingly, the mapping can be dynamically changed so that the visual item is recognized as being fixed with respect to the real world, thus providing a very natural experience.
[0111] For example, to create the illusion of an immersive virtual world, the three-dimensional audio and / or visual rendering is typically controlled by head tracking. The rendering is corrected with respect to the pose of the head, and the pose of the head includes in particular changes in the orientation of the head (such as the three degrees of freedom in space (3-DoF) of yaw, pitch, roll, etc.). The rendering is such that the visual item is recognized as being fixed with respect to the user. This head tracking and the effect of the rendering adaptation brought about by the head tracking are the high sense of reality and exocentric recognition of the rendered content as compared to static rendering.
[0112] However, as another approach, there is a method of mapping an input pose to a rendering pose that is fixed with respect to the rendering coordinate system. This can be implemented, for example, by the mapper 211 applying a fixed mapping from the input pose to the rendering pose, where the mapping is independent of head movement data. Specifically, changes in the user's pose do not change the mapping between the input pose and the rendering pose. The effect of such a mapping is effectively that the recognized scene moves with the head, i.e., it is static with respect to the user's head. This may feel unnatural in most scenes, but can be advantageous in some cases. For example, it can provide a desirable experience when listening to music or sounds that are not part of the scene, such as a narrator.
[0113] It is possible to apply different methods to different audiovisual items. In MPEG terminology, the terms "fixed to the head orientation" or "not fixed to the head orientation" are used to refer to audio items that are rendered either completely following the user's movement or ignoring it.
[0114] For example, an audio item can be considered "not fixed to the head", which means that the rendering is dynamically adapted to the (change in) orientation of the user's head because it is an audio element intended to have a fixed position in the (virtual or real) environment. Another audio item can be considered "fixed to the head", which means that it is an audio item intended to have a fixed position with respect to the user's head. Such an audio item can be rendered independently of the listener's pose. Therefore, in the rendering of such an audio item, the (change in) orientation of the user's head is not considered, i.e., such an audio item is an audio element whose relative position does not change even if the user rotates their head (e.g., non-spatial audio such as ambient sound or music intended to track the user without changing the relative position).
[0115] In the described system, the second receiver 209 is configured to receive metadata that further includes a rendering category display for at least some of the audiovisual items. The rendering category display indicates a rendering category from a set of rendering categories, and the rendering of the audiovisual item is performed according to the rendering category indicated for the audiovisual item. Different rendering categories may define different rendering parameters and operations.
[0116] The rendering category display can be any display that can be used to select a rendering category from a set of rendering categories. In many embodiments, it may be data provided only for selecting a rendering category and / or data that directly specifies one category. In other embodiments, the rendering category display may be a display that provides additional information or provides a description of the corresponding audiovisual item. In some embodiments, the rendering category display is one parameter considered when selecting a rendering category, and other parameters may also be considered.
[0117]
[0118] As a specific example, in some embodiments, an audio item can be encoded audio data, for example, an encoded audio signal where the audio item can be a plurality of different types of audio items including a plurality of different types of signals and components. In practice, in many embodiments, the metadata receiver 201 can receive metadata that defines a plurality of different types / formats of audio. For example, the audio data can include audio represented by an audio channel signal, an individual audio object, HOA (Higher Order Ambisonics), etc. The metadata can be included as part of the audio item or separately from the audio item that describes the audio type of each audio item. This metadata can be a rendering category display and can be used to select an appropriate rendering category for the audio item.Rendering categories are associated with different specific reference coordinate systems, and each rendering category is linked to a specific coordinate system transformation from the real-world coordinate system to the category coordinate system. Specifically, for each category, the real-world coordinate system, for example, the real-world coordinate system that serves as a reference in providing head movement data, can be transformed into different coordinate systems given by the transformation. Since different categories result in different coordinate system transformations, they are linked to different category reference systems.
[0119] The coordinate system transformation is usually a dynamic coordinate system transformation for one, some, or all of the categories. Therefore, the coordinate system transformation is usually not a fixed or static coordinate system transformation and can change over time according to various parameters. For example, as will be explained in more detail later, the coordinate system transformation can depend on dynamically changing parameters such as the movement of the user's torso, the movement of an external device, and / or in some cases, head movement data. Thus, in many embodiments, the coordinate system transformation of a category is a coordinate system transformation that changes over time depending on user movement parameters. User movement parameters can indicate the movement of the user relative to the real-world coordinate system.
[0120] The mapping performed by mapper 211 for a given audiovisual item depends on the category coordinate system of the rendering category to which the audiovisual item is shown to belong. Specifically, based on the rendering category display, mapper 211 can determine the rendering category to be used for rendering that audiovisual item. Next, the mapper can determine the coordinate system transformation linked to the selected category. Thereafter, mapper 211 can perform the mapping from the input pose to the rendering pose so as to correspond to a fixed pose within the category coordinate system obtained from the selected coordinate system transformation.
[0121] Thus, the category coordinate system can be regarded as a reference coordinate system serving as a reference for rendering so that the audiovisual item is fixed. The category coordinate system may also be referred to as the reference coordinate system or the fixed reference coordinate system (for a given category).
[0122] In many embodiments, one rendering category may correspond to a rendering in which the sound source and scene objects represented by the audiovisual item are fixed with respect to the real world as described above. In such embodiments, the coordinate system transformation is such that the category coordinate system for the category is aligned with the real-world coordinate system. For such a category, the coordinate system transformation can be a fixed coordinate system transformation, for example, a single one-to-one mapping of the real-world coordinate system. Thus, the category coordinate system can in fact be the real-world coordinate system, or for example, a fixed static translation, scaling, and / or rotation.
[0123] In many embodiments, one rendering category may correspond to a rendering in which the sound source and scene objects represented by the audiovisual item are fixed with respect to the movement of the head, i.e., with respect to the rendering coordinate system. In such embodiments, the coordinate system transformation is such that the category coordinate system for the category is aligned with the user's head / playback device / rendering coordinate system. For such a category, the coordinate system transformation can be a coordinate system transformation that fully follows the movement of the user's head. For example, when the head is rotated, a corresponding rotation of the coordinate system transformation occurs, and when the position of the user's head changes, the same change occurs in the coordinate system transformation. Thus, according to such a rendering category, the coordinate system transformation is dynamically modified to follow the head movement data so that the resulting category coordinate system is aligned with the rendering coordinate system. Thereby, as described above, a fixed mapping from the input coordinate system to the aforementioned rendering coordinate system is obtained.
[0124] Although not required, in many embodiments, the rendering category can include a category in which the audiovisual item is fixed and rendered with respect to the real-world coordinate system and a category in which the audiovisual item is fixed and rendered with respect to the rendering coordinate system. However, in the systems described, one or more rendering categories include a rendering category in which the audiovisual item is fixed to a coordinate system that is neither the real-world coordinate system nor the rendering coordinate system. That is, the rendering of the audiovisual item is provided with a rendering category that is not fixed to either the real world or the user's head. Thus, at least one category coordinate system is variable with respect to the real-world coordinate system and the rendering coordinate system. Specifically, the coordinate system is different from the rendering coordinate system and the real-world coordinate system, and the difference between these and the category coordinate system is not constant and can vary.
[0125] Thus, at least one rendering category can provide a rendering that is not fixed to either the real world or the user. Rather, in many embodiments, an intermediate experience may be provided.
[0126] For example, the coordinate system transformation may be such that the corresponding coordinate system is fixed with respect to the real world, except when an update criterion is met. However, when the criterion is met, the coordinate system transformation may provide a different relationship between the real-world coordinate system and the categorical coordinate system. For example, this mapping may be such that the visual is rendered fixed with respect to the real-world coordinate system, i.e., the visual item is felt to be in a fixed position. However, when the user rotates their head by more than a certain amount, the relationship between the real-world coordinate system and the categorical coordinate system is changed to compensate for this rotation. For example, as long as the user's head movement is less than 20°, the visual item is rendered as being in a fixed position. However, when the user's head movement exceeds 20°, the categorical coordinate system is rotated 20° with respect to the real-world coordinate system. This may provide the user with an experience of perceiving a natural three-dimensional experience with respect to the rendered visual item, as long as the movement is small enough. However, when the head movement is large, the rendering of the visual item is readjusted to match the changed head position.
[0127] As a specific example, the sound source corresponding to the narrator can be presented so as to be initially placed in front of the user. When the user's movement is small, the voice is rendered so that the narrator is recognized as being stationary at the same position. This provides a natural experience and perception, and in particular provides an out-of-head localization perception of the narrator. However, when the user rotates their head by, for example, 20° or more from the original direction towards the narrator sound source, the system adjusts the mapping and relocates the narrator sound source to the front of the user's new orientation. For small movements around this point, the sound source of the narrator is rendered at this new (with respect to the real-world coordinate system) fixed position. If the movement with respect to this new sound source position exceeds a given threshold again, the category coordinate system can be updated again, and the recognized position of the narrator sound source can be updated again. In this way, a narrator that is fixed with respect to small movements but follows the user with respect to large movements can be provided to the user. In the example described, the narrator is made to be recognized as being fixed and appropriate spatial cues are provided for head movements, but it is possible that the narrator can always be positioned substantially in front of the user (for example, even when the user rotates 180° completely).
[0128] In this approach, the metadata includes a plurality of rendering category indicators for a plurality of audiovisual items. Thereby, the source side can control flexible rendering at the receiving side, and the rendering is adapted specifically to individual audiovisual items. For example, different spatial renderings and perceptions can be applied to items corresponding to background music, narration, sound sources corresponding to specific objects fixed within a scene, conversations, etc.
[0129] In some embodiments, the coordinate system transformation for the rendering category depends on user head movement data. Accordingly, in some embodiments, at least one parameter that changes the coordinate system transformation depends on user head movement data.
[0130] In many embodiments, the coordinate system transformation may depend on user head pose characteristics or parameters determined from user head movement data. For example, as described above, if the user head pose exhibits a rotation exceeding a certain amount, the coordinate system transformation may be adapted to include a rotation corresponding to that amount. As another example, mapper 211 can detect whether the user has maintained a given pose for longer (sufficiently) than a given duration. If so, the coordinate system transformation may be adapted to place the audiovisual item at a given position within the rendering coordinate system (i.e., a specific position relative to the user (such as directly in front of the user)).
[0131] In some embodiments, the coordinate system transformation depends on the average head pose. In particular, in some embodiments, the coordinate system transformation may be such that the categorical coordinate system is aligned with the average head pose. In some embodiments, the coordinate system transformation may be such that the categorical coordinate system is fixed relative to the average head pose.
[0132] The average head pose can be determined, for example, by low-pass filtering the head pose measurements with a low-pass filter having an appropriate cut-off frequency, specifically by applying an unweighted average over a window of appropriate duration.
[0133] In some embodiments, the reference for rendering can be selected as the average head orientation h for one or more rendering categories. This causes the audiovisual item to follow slower and longer-lasting changes in the head orientation, such that the sound source feels like it remains in the same place relative to the head (e.g., in front of the face), but when the head movement is fast, the audiovisual item appears to be fixed relative to the real world rather than the head (and thus feels fixed in the virtual world as well). Thus, typical small and fast head movements in daily life create an immersive out-of-head localization illusion while allowing the overall perception that the audiovisual item is following the user.
[0134] In some embodiments, the adaptation and tracking may be non-linear. For example, when the head rotates significantly, the reference of the average head orientation may be "clipped" so as not to deviate beyond a certain maximum angle with respect to the instantaneous head orientation. For example, if this maximum value is 20°, as long as the head "moves incrementally" within the range of + / - 20°, the out-of-head localization experience is realized. When the head rotates quickly and exceeds the maximum value, the reference follows the head orientation (lagging by a maximum of 20°), and when the movement stops, the reference is fixed again.
[0135] In some embodiments, at least one rendering category is associated with a coordinate system transformation that depends on the user's body pose. In such embodiments, the audiovisual rendering device may include a body pose receiver 215 configured to receive user body pose data indicating the user's body pose.
[0136] The body pose can be determined, for example, by a dedicated inertial sensor unit arranged or worn on the body. As another example, the body pose can be determined by sensors of a smart device such as a smartphone worn in a pocket. As yet another example, coils can be arranged on the user's body and head respectively, and based on the change in the coupling between the two, the movement of the head relative to the body can be determined.
[0137] In such embodiments, at least one parameter that changes the coordinate system transformation depends on the body pose data.
[0138] In many embodiments, the coordinate system transformation may depend on the body pose data or parameters determined from the body pose data.
[0139] For example, when the torso pose data indicates a torso rotation exceeding a certain amount, the coordinate system transformation can be adapted to include a rotation corresponding to that amount. As another example, mapper 211 can detect whether the user has maintained a substantially constant torso pose for longer than a given duration. If so, the coordinate system transformation can be adapted to place the audiovisual item at a given position corresponding to the torso pose within the rendering coordinate system.
[0140] In particular, in some embodiments, the coordinate system transformation can be such that the category coordinate system is adapted to the user's torso pose. In some embodiments, the coordinate system transformation can be such that the category coordinate system is fixed relative to the user's torso pose. Thus, in some embodiments, the audiovisual item can be rendered to follow the user's torso, and thus, a perception and experience can be provided such that the audiovisual item follows the user's full body movement but appears fixed with respect to head movement relative to the torso. This can provide both the desirable experiences of exocentric perception and rendering of an audiovisual item that follows the user.
[0141] In some embodiments, the coordinate system transformation depends on the average torso pose. The average torso pose can be determined, for example, by low-pass filtering the head pose measurements with a low-pass filter having an appropriate cut-off frequency, specifically, by applying an unweighted average over a window of appropriate duration.
[0142] Thus, in some embodiments, one or more of the rendering categories can use a coordinate system transformation that provides rendering criteria aligned with the instantaneous or average chest / torso orientation t By doing so, the audiovisual item can be made to feel as if it remains in front of the user's body rather than in front of the user's face. Still, by rotating the head relative to the chest / torso, the audiovisual item can be perceived as coming from various directions, greatly contributing to an immersive exocentric experience.
[0143] In some embodiments, the coordinate system transformation for a rendering category depends on device pose data indicating the pose of an external device. In such embodiments, the audiovisual rendering device may comprise a device pose receiver 217 configured to receive device pose data indicating the device body pose.
[0144] The device can be, for example, a device that (hypothetically) a user wears, carries, attaches to the user, or otherwise secures. In many embodiments, the external device can be, for example, a mobile phone or a personal device, such as a smartphone in a pocket, a body-worn device, or a handheld device (e.g., a smart device used to view visual VR content).
[0145] Many devices include a gyro, an accelerometer, a GPS receiver, etc., and can determine the relative or absolute orientation of the device. In that case, the device can determine its current relative or absolute orientation and transmit it to the device pose receiver 217 using a suitable communication, usually wireless. For example, the communication can be via a WiFi or Bluetooth connection.
[0146] In such embodiments, at least one parameter that varies the coordinate system transformation depends on the device pose data.
[0147] In many embodiments, the coordinate system transformation can depend on the device pose data or parameters determined from the device pose data.
[0148] For example, if the device pose data indicates a rotation of the device beyond a certain amount, the coordinate system transformation can be adapted to include a rotation corresponding to that amount. As another example, the mapper 211 can detect whether the device has maintained a sufficiently constant body pose for a given duration. If so, the coordinate system transformation can be adapted to place the audiovisual item at a given position corresponding to the device pose within the rendering coordinate system.
[0149] In particular, in some embodiments, the coordinate system transformation may be such that the category coordinate system is adjusted to the device pose. In some embodiments, the coordinate system transformation may be such that the category coordinate system is fixed relative to the device pose. Thus, in some embodiments, the audiovisual item may be rendered to follow the pose of the device, and thus a perception and experience may be provided where the audiovisual item follows the movement of the device. In many practical user scenarios, the device may provide an excellent representation of the reference of the user pose. For example, a body-worn device, or a smartphone in a pocket for example, may well reflect the movement of the entire user. This may provide an excellent reference for determining relative head movement and thus may provide an experience that combines both a realistic response to head movement and the ability of the audiovisual item to follow the user's larger movements.
[0150] Furthermore, using an external device as a reference can be very practical and can provide a reference that results in a desirable user experience. This approach may often be based on a device that is already worn or carried by the user and has the functionality necessary to determine and transmit the device pose. For example, currently most people carry a smartphone that has an accelerometer for determining the device pose and communication means (e.g., Bluetooth) suitable for transmitting the device pose data to the audiovisual rendering device.
[0151] In some embodiments, the coordinate system transformation may depend on the average device pose. The average device pose may be determined, for example, by low-pass filtering the device pose measurements by a low-pass filter having an appropriate cut-off frequency, specifically by applying an unweighted average over a window of appropriate duration.
[0152] Thus, in some embodiments, one or more of the rendering categories may use a coordinate transformation that provides a rendering criterion aligned with the instantaneous or average device orientation. By doing so, the audiovisual item may be perceived as remaining in a fixed position relative to the device, such that the item moves as the device moves, but remains fixed relative to head movement, thereby providing a more natural and immersive out-of-head localization experience.
[0153] Thus, this approach provides a method of controlling the rendering of audiovisual items using metadata, and can control the audiovisual items individually to provide different user experiences for each audiovisual item. The experience includes one or more options that provide a perception of the audiovisual item that is neither completely fixed to the real world nor completely follows the user (fixed to the head). Specifically, the audiovisual item can provide an intermediate experience that is somewhat fixed relative to the real world and somewhat follows the user's movement.
[0154] It will be appreciated that in some embodiments, the possible rendering categories, each associated with a predetermined coordinate transformation, may be determined in advance. In such embodiments, the audiovisual rendering device can save the coordinate transformation for each category, or equivalently, directly save the mapping corresponding to the coordinate transformation, if appropriate. The mapper 211 can be configured to retrieve the saved coordinate transformation (or mapping) for the selected rendering category and apply it when performing the mapping of the audiovisual item.
[0155] For example, in the case of the first visual item, the rendering category indicator may indicate that it should be rendered according to the first category. The first category may be for rendering a visual item fixed to a head pose, and thus the mapper can extract a mapping that provides a fixed one-to-one mapping between the input pose and the rendering pose. In the case of the second visual item, the rendering category indicator may indicate that it should be rendered according to the second category. The second category may be for rendering a visual item fixed to the real world, and thus the mapper can extract a coordinate system transformation that adjusts the mapping so that the head movement is compensated and the rendering pose corresponds to a fixed position in the real space. In the case of the third visual item, the rendering category indicator may indicate that it should be rendered according to the third category. The third category may be for rendering a visual item fixed to a device pose or a torso pose. The mapper can extract a coordinate system transformation or a mapping that adjusts the mapping so that rendering of a visual item fixed to the device or torso pose is obtained by compensating for the head movement relative to the device or torso pose.
[0156] Through mapping, different rendering categories are associated with different coordinate system transformations so that a rendering position fixed with respect to the category coordinate system is obtained. However, it should be understood that the mapper 211 does not need to explicitly determine such a coordinate system transformation or the category coordinate system. Instead, in an exemplary embodiment, a mapping function is defined for each individual rendering category such that the resulting rendering pose is fixed with respect to the category coordinate system. For example, a mapping function that is a function of the head pose with respect to the device (or torso) pose may be used to directly map the input position to a rendering position that is fixed with respect to the category coordinate system that is fixed with respect to the device (or torso) pose.
[0157] In some embodiments, the metadata may include data that partially or fully characterizes, describes, and / or defines one or more of a plurality of rendering categories. For example, in addition to the rendering category display, the metadata may include data that describes a coordinate system transformation and / or a mapping function to be applied to one or more categories. For example, the metadata may indicate that in a first rendering category, a fixed mapping from an input position to a rendering position is required, in a second category, a mapping that fully compensates for head movement so that the item appears fixed relative to the real world is required, and in a third rendering category, it may be necessary to compensate for head movement by an amount relative to the average head movement such that for smaller and faster head movements the item appears fixed, but for slower average movements it is recognized as an intermediate experience where the item appears to follow the user.
[0158] In different embodiments, different techniques and data may be used as the rendering category display. In some embodiments, each category may be associated with, for example, a category number, and the rendering category display may directly provide the category number to be used for the audiovisual item.
[0159] In many embodiments, the rendering category display may indicate a characteristic or feature of the audiovisual item, which may be mapped to a particular rendering category display.
[0160] In some embodiments, the rendering category display can specifically indicate whether the audiovisual item is a diegetic audiovisual item or a non-diegetic audiovisual item. A diegetic audiovisual item can be an item belonging to a scene such as a movie or a story being screened. In other words, diegetic audiovisual items originate from sources within a movie, story, etc. (e.g., actors in a play, birds in a natural video and their chirping sounds, etc.). Non-diegetic audiovisual items can be items originating from outside the movie or story (e.g., the director's audio commentary, mood music, etc.). In many scenarios, in accordance with the MPEG usage, diegetic audiovisual items can correspond to "not fixed to the head orientation" and non-diegetic audiovisual items can correspond to "fixed to the head orientation".
[0161] In some embodiments, there may only be two rendering categories, specifically corresponding to the rendering of audiovisual items shown to be diegetic and the rendering of audiovisual items shown to be non-diegetic, respectively.
[0162] For example, in some applications and systems, diegetic signals are transmitted downstream to the audiovisual rendering device, and as illustrated with respect to FIG. 3, the desired rendering operations may depend on this signal. · By using head tracking based on the real-world orientation, it is desirable for the diegetic source D to be felt to stay firmly at the position of that source within the virtual world V, and thus it is necessary to render it as if it were fixed based on the real world. In a given example of a movie application, an actor's voice is rendered straight in front of the user, and when the user rotates their head 50° to the left, the sound is rendered to the headphones so as to be projected 50° to the right, thereby making it feel as if it stays at the same virtual position. · On the one hand, the non-diegetic sound source N can be rendered regardless of the head orientation. In other words, the sound stays at a fixed position “hard-coupled” to the head (e.g., in front of the head) and rotates with the head. This is achieved by using a fixed mapping from the input position to the rendering position, rather than applying a mapping that depends on the head orientation to the sound. In the movie example, the director's commentary audio can be rendered right in front of the user's eyes and the head movement does not affect it (i.e., the sound stays in front of the head).
[0163] The visual rendering device of FIG. 2 is configured to provide a more flexible approach, in which at least one selectable rendering category enables rendering visual items to follow the user for some movements, as if fixed with respect to the real / virtual world for other movements. This approach can be particularly applied to non-diegetic visual items.
[0164] In this specific example, alternative or additional options for rendering the non-diegetic sound source may include one or more of the following. ·The rendering reference can be selected as the average head orientation / pose h (as shown in FIG. 4). This causes the non-diegetic sound source to follow slower and longer-lasting changes in head orientation, making the sound source feel as if it remains in the same location relative to the head (e.g., in front of the face), while when the head movement is fast, the non-diegetic sound source feels fixed within the virtual world (rather than relative to the head). Thus, typical small and fast head movements in daily life still produce an immersive out-of-head localization illusion. As an improvement in some embodiments, since non-diegetic audio preferably remains at the same virtual position as much as possible, the tracking can be non-linear. For example, when the head rotates significantly, the reference of the average head orientation can be "clipped" so as not to deviate beyond a certain maximum angle relative to the instantaneous head orientation. For example, if this maximum value is 20°, as long as the head "moves incrementally" within the range of + / -20°, an out-of-head localization experience is achieved. When the head rotates quickly and exceeds the maximum value, the reference follows the head orientation (lagging by a maximum of 20°), and when the movement stops, the reference is fixed again. ·The head tracking reference can be selected as the average chest / trunk orientation / pose t (as shown in FIG. 5). By doing so, the non-diegetic content feels as if it remains in front of the user's body rather than in front of the user's face. Still, by rotating the head relative to the chest / trunk, the non-diegetic content can be made to sound as if it comes from various directions, greatly contributing to an immersive out-of-head localization experience. ·The head tracking reference can be selected as the instantaneous or average orientation / pose of an external device such as a mobile phone or a body-worn device. By doing so, the non-diegetic content feels as if it remains in front of the device rather than in front of the user's face. Still, by rotating the head relative to the device, the non-diegetic content can be made to sound as if it comes from various directions, greatly contributing to an immersive out-of-head experience.
[0165] In some embodiments, the rendering can be configured to operate in different modes in response to the user's movement. For example, if the user's movement meets a movement criterion, the audiovisual rendering device operates in a first mode; otherwise, it operates in a second mode. In this example, the two modes can provide different category coordinate systems for the same rendering category display. That is, depending on the user's movement, the audiovisual rendering device can render a given audiovisual item in a given rendering category display fixed with respect to different coordinate systems.
[0166] In some embodiments, the mapper 211 can be configured to select a rendering category for a given rendering category display in response to user movement parameters indicating the user's movement. Thus, depending on the user movement parameters, different rendering categories can be selected for a given audiovisual item and rendering category display. Specifically, if the user movement parameters meet a first criterion, a given link between possible rendering category display values and a set of rendering categories can be used to select the rendering category of the received rendering category display. However, if the criterion is not met (or, for example, if a different criterion is met), the mapper 211 can use a different link between possible rendering category indication values and the same or a different set of rendering categories to select the rendering category of the received rendering category display.
[0167] This approach can enable, for example, the audiovisual rendering device to provide different renderings and experiences for mobile users and fixed-type users.
[0168] In some embodiments, the selection of the rendering category can also depend on other parameters, such as user settings or configuration settings by an application (e.g., an app on a mobile device).
[0169] In another approach, the mapper can be configured to determine a coordinate system transformation between a real-world coordinate system used as a reference for coordinate system transformation of a selected category and a coordinate system that is a reference in providing user head movement data provided in response to user movement parameters indicating the user's movement.
[0170] Thus, in some embodiments, rendering adaptation can be introduced by a coordinate system transformation of a rendering category with respect to a reference that can vary with respect to a (typically real-world) coordinate system indicating head movement. For example, based on the user's movement, compensation can be applied to the head movement data, for example, offsetting the user's overall movement. As an example, when the user is on a boat, the user head movement data may not only indicate the user's movement relative to the body or the body's movement relative to the boat, but may also reflect the movement of the boat. Since this may be undesirable, the mapper 211 can compensate the head movement data with respect to user movement parameter data reflecting the component of the user's movement due to the movement of the boat. The resulting corrected / compensated head movement data is provided with respect to a coordinate system in which the movement of the boat is compensated, and by directly applying the transformation of the selected category coordinate system to this corrected coordinate system, a desirable rendering and user experience can be achieved.
[0171] It will be appreciated that the user movement parameters can be determined in any suitable manner. For example, in some embodiments, they can be determined by a dedicated sensor that provides relevant data. For example, an accelerometer and / or gyroscope can be attached to a vehicle such as an automobile or a boat carrying the user.
[0172] In many embodiments, the mapper can be configured to determine user movement parameters in response to the user head movement data itself. For example, user parameters indicating user movement parameters may be determined using an analysis of underlying motion, such as a long-term average or identifying periodic components (e.g., corresponding to waves moving the boat).
[0173] In some embodiments, the coordinate system transformation of two different rendering categories depends on the user's head movement data, but the dependencies have different temporal averaging characteristics. For example, one rendering category depends on the average of the head movement, but is associated with a coordinate system transformation with a relatively short averaging time (i.e., a relatively high cutoff frequency of the averaging low-pass filter), while another rendering category, which also depends on the average of the head movement, may be associated with a coordinate system transformation with a longer averaging time (i.e., a relatively low cutoff frequency of the averaging low-pass filter).
[0174] As an example, mapper 211 may evaluate the head movement data against a reference that reflects the consideration that the user is located in a static environment. For example, the low-pass filtered position change is compared to a given threshold, and if it is below the threshold, the user may be considered to be in a static environment. In this case, rendering may be performed as in the specific examples described above for diegetic and non-diegetic audio items. However, if the position change exceeds the threshold, the user may be considered to be in a moving environment, for example, walking, or using a means of transportation such as a car, train, or airplane. In this case, the following rendering techniques may be applied (see also FIGS. 6 and 7). · The non-diegetic sound source N can be rendered using the average head orientation h as a reference. This makes it possible to create an "out-of-head localization" experience with small head movements while keeping the non-diegetic sound source mainly in front of the face. In other examples, for example, the user's torso or the pose of the device may be used as a reference. Thus, this technique can correspond to the technique used for diegetic sound sources when static. · The diegetic sound source D also uses the average head orientation h as a reference, but it may be rendered using an additional offset corresponding to a longer-term average head orientation h'. Thus, in this case, the diegetic virtual sound source appears to be in a fixed position relative to the virtual world V, but this virtual world V appears to "move with the user", i.e., it maintains an orientation that is more or less fixed relative to the user. The averaging (or other filtering) used to obtain h' from the instantaneous head orientation is typically chosen to be significantly slower than that used for h. Thus, non-diegetic sources follow faster head movements, while diegetic content (and the virtual world V as a whole) takes time to change with the user's head orientation.
[0175] For clarity, the above description has been presented in terms of different functional circuits, units, and processors related to embodiments of the invention. However, it will be understood that the functions may be appropriately distributed among different functional circuits, units, or processors without detracting from the invention. For example, functions described as being performed by a plurality of separate processors or controllers may be performed by the same processor or controller. Thus, references to specific functional units or circuits are not intended to denote a strict logical or physical structure or organization, but rather references to suitable means for providing the described functions.
[0176] The present invention can be implemented in any suitable form including hardware, software, firmware, or any combination thereof. The present invention may be implemented at least in part as computer software running on one or more data processors and / or digital signal processors. The elements and components of embodiments of the present invention can be physically, functionally, and logically implemented in any suitable manner. In practice, the functions may be implemented as a single unit, multiple units, or as part of other functional units. Thus, the present invention may be implemented as a single unit or may be physically and functionally distributed among different multiple units, circuits, and processors.
[0177] Although the present invention has been described in connection with some embodiments, the present invention is not limited to the specific forms described in the specification. The scope of the present invention is limited only by the appended claims. Further, even if a certain feature may seem to be described in connection with a particular embodiment, those skilled in the art will recognize that the various features of the above embodiments can be combined according to the present invention. In the claims, terms such as comprising, including, etc. do not exclude the presence of other elements or steps.
[0178] Furthermore, even if listed individually, a plurality of means, elements, circuits, or method steps may be implemented, for example, by a single circuit, unit, or processor. Additionally, even if individual features are included in different claims, these may be suitably combined, and being included in different claims does not mean that the combination of features is impossible and / or not advantageous. Also, just because a feature is included within one claim category does not mean that the feature is limited to that category, and the feature may equally be applied to other claim categories as appropriate. Moreover, the order of features in a claim does not refer to a specific order in which the features should act, and in particular, the order of individual steps in a method claim does not mean that the steps must be performed in that order. The steps may be performed in any appropriate order. Also, singular expressions do not exclude plurals. Thus, singular expressions do not exclude plurals. Reference signs within the claims are merely examples for clarity and do not limit the scope of the claims in any way.
Claims
1. A first receiver that receives audiovisual items, A metadata receiver that receives metadata including an input pose and a rendering category display for each of at least some of the audiovisual items, wherein the input pose is provided with reference to an input coordinate system, and the rendering category display indicates a certain rendering category among a set of rendering categories, the metadata receiver, A receiver that receives user head movement data indicating the movement of the user's head, A mapper that maps the input pose to a rendering pose within a rendering coordinate system in response to the user head movement data, wherein the rendering coordinate system is fixed with respect to the movement of the head, the mapper, A renderer that renders the audiovisual item using the rendering pose, an audiovisual rendering device comprising: Each rendering category is linked to a coordinate system transformation from a real-world coordinate system to a category coordinate system, the coordinate system transformation being different for each category, and at least one category coordinate system is variable with respect to the real-world coordinate system and the rendering coordinate system, The mapper selects a first rendering category of the first audiovisual item from the set of rendering categories in response to the rendering category display of the first audiovisual item, and maps the input pose of the first audiovisual item to a rendering pose within the rendering coordinate system corresponding to a fixed pose within a first category coordinate system for changing the movement of the user's head, the first category coordinate system being determined from a first coordinate system transformation of the first rendering category, the audiovisual rendering device.
2. The audiovisual rendering device according to claim 1, wherein the second coordinate system transformation of the second category aligns the category coordinate system of the second category with the movement of the user's head.
3. The audiovisual rendering device according to claim 1 or 2, wherein the third coordinate system transformation of the third category aligns the category coordinate system of the third category with the real-world coordinate system.
4. The audiovisual rendering device according to any one of claims 1 to 3, wherein the first coordinate system transformation depends on the user head movement data.
5. The visual rendering device according to claim 4, wherein the first coordinate system transformation depends on the average head pose.
6. The visual rendering device according to claim 5, wherein the first coordinate system transformation aligns the first category coordinate system with the average head pose.
7. For different coordinate system transformations of different rendering categories, they depend on the user head movement data, and the dependencies of the first coordinate system transformation and the different coordinate system transformations on the movement of the user's head have different time-averaging characteristics. The visual rendering device according to any one of claims 4 to 6.
8. The visual rendering device according to any one of claims 1 to 7, further comprising a receiver for receiving user body pose data indicating the user body pose, wherein the first coordinate system transformation depends on the user body pose data.
9. The visual rendering device according to claim 8, wherein the first coordinate system transformation aligns the first category coordinate system with the user body pose.
10. The visual rendering device according to any one of claims 1 to 9, further comprising a receiver for receiving device pose data indicating the pose of an external device, wherein the first coordinate system transformation depends on the device pose data.
11. The visual rendering device according to claim 10, wherein the first coordinate system transformation aligns the first category coordinate system with the device pose.
12. The visual rendering device according to any one of claims 1 to 11, wherein the mapper selects the first rendering category in response to user movement parameters indicating the movement of the user.
13. The visual rendering device according to any one of claims 1 to 12, wherein the mapper determines a coordinate system transformation between the real-world coordinate system and the coordinate system of the user head movement data in response to user movement parameters indicating the movement of the user.
14. The visual rendering device according to claim 12 or 13, wherein the mapper determines the user movement parameters in response to the user head movement data.
15. The visual rendering device according to any one of claims 1 to 14, wherein at least some of the rendering category displays indicate whether the visual items of the at least some of the rendering category displays are digital visual items or non-digital visual items.
16. The visual item is an audio item, and the renderer generates an output binaural audio signal for a binaural rendering device by applying binaural rendering to the audio item using the rendering pose, according to any one of claims 1 to 15.
17. A method of rendering a visual item, the method comprising: receiving the visual item; receiving metadata including an input pose and a rendering category display for each of at least some of the visual items, the input pose being provided with reference to an input coordinate system, and the rendering category display indicating a certain rendering category among a set of rendering categories; receiving user head movement data indicating movement of the user's head; mapping the input pose to a rendering pose within a rendering coordinate system in response to the user head movement data, the rendering coordinate system being fixed with respect to the movement of the head; rendering the visual item using the rendering pose; and each rendering category is linked to a coordinate system transformation from a real-world coordinate system to a category coordinate system, the coordinate system transformation being different for each category, and at least one category coordinate system is variable with respect to the real-world coordinate system and the rendering coordinate system. The method includes the step of selecting, in response to a rendering category display of a first audiovisual item, a first rendering category of the first audiovisual item from a set of the rendering categories, and mapping an input pose of the first audiovisual item to a rendering pose in a rendering coordinate system corresponding to a fixed pose in a first category coordinate system for changing a movement of a user's head, wherein the first category coordinate system is determined from a first coordinate system transformation of the first rendering category, the method. Claim 18 A computer program, which when executed on a computer, includes computer program code means for performing all the steps of the method according to claim 17.
Citation Information
Patent Citations
Audio apparatus and method of audio processing
EP3617871A1
Method and binaural sound system for conveying binaural information to a user
JP2009543479A
Coordinated tracking for binaural audio rendering
US20180091922A1
Associated spatial audio playback
WO2019141900A1
Spatial audio augmentation
WO2020012067A1