Audio-visual presentation apparatus and method of operating same
Through the mapper and category coordinate system transformation in the audio-visual presentation device, the problem of insufficient response to user head movement in virtual reality and augmented reality is solved, the immersive experience and consistency are improved, and the presentation effect of audio and visual scenes is improved.
Patent Information
- Application Number
- CN202510371324.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-13
- Filing Date
- 2021-10-11
- Publication Date
- 2025-07-11
AI Technical Summary
In the virtual reality and augmented reality applications, it is difficult to effectively respond to user head movements, resulting in reduced immersive experience, inconsistent audio and visual scenes, and poor user experience.
Through the audio-visual presentation device, the input pose is mapped into the presentation coordinate system using a mapper, and combined with the transformation of coordinate system in different categories, the flexible presentation of audio-visual projects is achieved, adapting to user head movement, and maintaining spatial perception consistency between audio and visual scenes.
Improves the flexibility and immersion of the user experience, reduces system complexity and resource requirements, provides more consistent audio and visual scene perception, and supports social and shared experiences.
Smart Images

Figure CN120298632A_ABST
Abstract
Description
[0001] This application is a divisional application of the application with the application number 202180070194.0 and the invention name of "Audio-visual presentation device and its operation method" submitted on April 13, 2023. Technical Field
[0002] The present invention relates to an audio-visual presentation device and its operation method, and in particular but not exclusively to using these to support, for example, augmented / virtual reality applications. Background Art
[0003] In recent years, with the continuous development and launch of new services and methods for utilizing and consuming audio-visual content, the types and scope of experiences based on audio-visual content have increased significantly. Specifically, many spatial and interactive services, applications, and experiences are being developed to give users a more engaging and immersive experience.
[0004] Examples of such applications are virtual reality (VR), augmented reality (AR), and mixed reality (MR) applications that are rapidly becoming mainstream, where many technical solutions target the consumer market. Many standards are also being developed by many standardization bodies. Such standardization activities are actively developing standards for various aspects of VR / AR / MR systems, including, for example, streaming, dissemination, presentation, etc.
[0005] VR applications tend to provide a user experience corresponding to the user in different worlds / environments / scenarios, while AR (including mixed reality MR) applications tend to provide a user experience corresponding to the user in the current environment but with additional information or virtual objects or information added. Therefore, VR applications tend to provide a fully immersive synthetically generated world / scenario, while AR applications tend to provide a partially synthetic world / scenario superimposed on the real scenario where the user physically exists. However, the terms are often used interchangeably and have a high degree of overlap. Hereinafter, the term virtual reality / VR will be used to represent both virtual reality and augmented reality.
[0006] As an example, a popular service is to provide images and audio in such a way that the user can actively and dynamically interact with the system to change the presentation parameters so that it will be suitable for the movement and change of the user's position and orientation. An attractive feature in many applications is the ability to change the viewer's effective viewing position and viewing direction, for example, allowing the viewer to move and "look around" in the scene being presented.
[0007] This feature can in particular allow a virtual reality experience to be provided to a user. This can allow the user to move around relatively freely in the virtual environment and dynamically change their position and the place they are looking at. Generally, such virtual reality applications are based on a three-dimensional model of a scene, where the model is dynamically evaluated to provide a view for a particular request. For computers and consoles, this method is well-known in, for example, gaming applications (such as in the first-person shooter genre).
[0008] Particularly for virtual reality applications, it is also desirable that the presented images are three-dimensional images. In fact, in order to optimize the viewer's sense of immersion, it is generally preferred to make the presented scene of the user experience a three-dimensional scene. In fact, the virtual reality experience should preferably allow the user to select their own position, camera viewpoint, and moment relative to the virtual world.
[0009] In addition to the virtual presentation, most VR / AR applications also provide a corresponding audio experience. In many applications, the audio preferably provides a spatial audio experience, where the audio source is perceived as arriving from a position corresponding to the position of the corresponding object in the virtual scene (including both currently visible objects and currently invisible objects (e.g., behind the user)). Thus, the audio and video scenes are preferably perceived as being consistent, and where both provide a complete spatial experience.
[0010] For audio, until now it has mainly focused on headphone reproduction using binaural audio presentation techniques. In many cases, headphone reproduction achieves a highly immersive and personalized experience for the user. Using head tracking, the presentation can respond to the user's head movements, which highly increases the sense of immersion.
[0011] For the purpose of immersive voice and audio services (IVAS), the 3GPP consortium has developed a so-called IVAS codec (3GPP SP-170611’New WID on EVS Codec Extension for Immersive Voice and Audio Services’). The codec includes a renderer that converts various audio streams into a form suitable for reproduction at the receiving end. In particular, the audio can be presented in a binaural format for reproduction via headphones or a head-mounted VR device with built-in headphones.
[0012] In many such applications, the presentation device can receive input data describing a three-dimensional audio and / or visual scene, and the renderer can be arranged to present the data such that the user is provided with an audiovisual experience that provides a perception of a three-dimensional scene.
[0013] However, it is challenging to provide a suitable experience in many applications, and in particular, adapting the presentation in response to head movement such that the desired experience is provided to the user is challenging.
[0014] For example, it is known that human perception of the direction and distance of a sound source is due not only to the (usually different) delays and filtering of the sound from the source to the two ears, but also to a large extent due to how these change as the head moves (such as rotates). Similarly, the parallax and similar movements of visual objects provide strong three-dimensional visual cues. Unconsciously, as we move and sway our heads (usually slightly) in our daily lives, the sound changes in a similar but distinct way, and this significantly increases the immersive "around us" auditory / visual experience to which we are accustomed.
[0015] Experiments using headphone reproduction have shown that even when the sound paths from the sound source to the ears are adequately modeled by filters by keeping these sound paths stationary (i.e., by lacking head movement-related changes), the immersive experience is reduced, and the sound may tend to seem "inside our heads".
[0016] Therefore, to create the impression of an immersive virtual world, some applications have been developed that present audio sources and / or visual objects at positions perceived as fixed relative to the real world. However, this is a challenging optimum operation and may not always result in the desired user experience. In some applications, a three-dimensional scene that follows head movement and thus appears fixed relative to the user's head is provided. This may be a desired experience in many applications, but may provide an unnatural experience in other applications and may, for example, not allow an immersive experience of "being in" the virtual scene. US10015620B2 discloses another example of presenting audio relative to a reference orientation representing the user. Other examples of presenting audio and video for virtual experiences can be found in WO20 / 012067A1, WO2019 / 141900A1, and US2019 / 215638.
[0017] However, although such applications can provide a suitable user experience in many embodiments, they tend not to provide an optimum or even desired user experience for some applications.
[0018] Therefore, an improved method for presenting audiovisual items (especially for virtual / augmented / mixed reality experiences / applications) would be advantageous. Specifically, a method that allows for improved operation, increased flexibility, reduced complexity, convenient implementation, improved user experience, more consistent perception of the audio and / or visual scene, improved customization, improved personalization; improved virtual reality experience and / or improved performance and / or operation would be advantageous. SUMMARY OF THE INVENTION
[0019] Accordingly, the present invention seeks to alleviate, mitigate or eliminate, preferably singly or in any combination, one or more of the drawbacks mentioned above.
[0020] According to one aspect of the present invention, there is provided an audio-visual presentation apparatus comprising: a first receiver arranged to receive an audio-visual item; a metadata receiver arranged to receive metadata including an input pose for each of at least some of the audio-visual items and a presentation category flag for each of at least some of the audio-visual items, the input pose being provided with reference to an input coordinate system, and the presentation category flag indicating a presentation category from a set of presentation categories; a receiver arranged to receive user head movement data indicative of a movement of a user's head; a mapper arranged to map the input pose to a presentation pose in a presentation coordinate system in response to the user head movement data, the presentation coordinate system being fixed relative to the head movement; a presenter arranged to present the audio-visual item using the presentation pose; wherein each presentation category is linked to a coordinate system transformation from a real-world coordinate system to a category coordinate system, the coordinate system transformation being different for different categories, and at least one category coordinate system is variable relative to the real-world coordinate system and the presentation coordinate system; and the mapper is arranged to select a first presentation category from the set of presentation categories for the first audio-visual item in response to the presentation category flag for the first audio-visual item, and map the input pose for the first audio-visual item to a presentation pose in the presentation coordinate system, the presentation pose corresponding to a fixed pose in a first category coordinate system for a changing user head movement, the first category coordinate system being determined according to a first coordinate system transformation for the first presentation category.
[0021] In many embodiments, the method can provide an improved user experience and can specifically provide an improved user experience for many virtual reality (including augmented and mixed reality) applications, specifically including social or shared experiences. The method can provide a very flexible approach in which the presentation operation and spatial perception of the audio-visual item can be individually adapted to an individual audio-visual item. The method can, for example, allow some audio-visual items to be presented as appearing completely fixed relative to the real world, some audio-visual items to be presented as appearing completely fixed to the user (following the user's head movement), and some audio-visual items to be presented as appearing fixed for some movements in the real world and following the user for other movements. In many embodiments, the method can allow flexible presentation in which the audio-visual item is perceived as substantially following the user, but still provides an out-of-headspace experience of the audio-visual item.
[0022] This method can reduce complexity and resource requirements in many embodiments and can allow source - side control of presentation operations in many embodiments.
[0023] The presentation category can indicate that the audiovisual item represents an audio source with spatial properties that are fixed for a head orientation or not fixed for a head orientation (corresponding to listener - pose - related positions and listener - pose - independent positions respectively). The presentation category can indicate whether the audio element is dramatic.
[0024] In many embodiments, the mapper can also be arranged to select a second presentation category from the set of presentation categories for the second audiovisual item in response to a presentation - category flag for the second audiovisual item, and map the input pose for the first audiovisual item to a presentation pose in the presentation coordinate system, the presentation pose corresponding to a fixed pose in a second - category coordinate system for changing user head movements, the second - category coordinate system being determined according to a second - coordinate transformation for the second presentation category. The mapper can be similarly arranged to perform such operations on third, fourth, fifth, etc. audiovisual items.
[0025] An audiovisual item can be an audio item and / or a visual / video / image / scene item. An audiovisual item can be a visual or audio representation of a scene object in a scene represented by the audiovisual item. In some embodiments, the term audiovisual item can be replaced by the term audio item (or element). In some embodiments, the term audiovisual item can be replaced by the term visual item (or scene object).
[0026] In many embodiments, the renderer can be arranged to generate an output binaural audio signal for a binaural presentation device by applying binaural presentation to an audiovisual item (as an audio item) using the presentation pose.
[0027] The term pose can represent position and / or orientation. In some embodiments, the term "pose" can be replaced by the term "position". In some embodiments, the term "pose" can be replaced by the term "orientation". In some embodiments, the term "pose" can be replaced by the term "position and orientation".
[0028] The receiver can be arranged to receive real - world user head - movement data that references a real - world coordinate system indicating the user's head movement.
[0029] The renderer is arranged to use the presentation pose to present the audiovisual item, where the audiovisual item is referenced / located in the presentation coordinate system.
[0030] The mapper may be arranged to map an input pose of the first audiovisual item to a rendered pose in the rendered coordinate system, the rendered pose corresponding to a fixed pose in a first category coordinate system of the head movement data for different indicated changes in user head movement.
[0031] In some embodiments, the rendered category flag indicates a source type, such as an audio type of an audio item or a scene object type of a visual element.
[0032] In many embodiments, this may provide an improved user experience. The rendered category flag may indicate an audio format from a set of audio formats, including at least one audio format from the following group: speech audio; music audio; foreground audio; background audio; narration audio; and narrator audio.
[0033] According to an optional feature of the invention, a second coordinate system transformation for a second category is such that the category coordinate system of the second category is aligned with the user head movement.
[0034] In many embodiments, this may provide an improved user experience and / or improved performance and / or facilitated implementation. In particular, it may support that some audiovisual items may appear fixed relative to the user's head while others are not.
[0035] According to an optional feature of the invention, a third coordinate system transformation for a third category is such that the category coordinate system of the third category is aligned with the real-world coordinate system.
[0036] In many embodiments, this may provide an improved user experience and / or improved performance and / or facilitated implementation. In particular, it may support that some audiovisual items may appear fixed relative to the real world while others are not.
[0037] According to an optional feature of the invention, the first coordinate system transformation depends on the user head movement data.
[0038] In many embodiments, this may provide an improved user experience and / or improved performance and / or facilitated implementation. In particular, in some applications, it may provide a very beneficial experience where the audiovisual item pose follows the overall movement of the user rather than smaller / faster head movements, thus providing an improved user experience with an improved out-of-head experience.
[0039] According to an optional feature of the invention, wherein the first coordinate system transformation depends on an average head pose.
[0040] According to an optional feature of the invention, the first coordinate system transformation aligns the first category coordinate system with the average head pose.
[0041] According to an optional feature of the present invention, different coordinate system transformations for different presentation categories depend on the user head movement data, and the dependencies of the first coordinate system transformation and the different coordinate system transformations on the user head movement have different time-averaging properties.
[0042] According to an optional feature of the present invention, the audiovisual presentation device further includes a receiver for receiving user torso pose data indicating the pose of the user's torso, and the first coordinate system transformation depends on the user torso pose data.
[0043] In many embodiments, this can provide an improved user experience and / or improved performance and / or facilitated implementation. In particular, in some applications, it can provide a very beneficial experience where the audiovisual item pose follows the overall movement of the user rather than smaller / faster head movements, thereby providing an improved user experience with an improved out-of-head experience.
[0044] According to an optional feature of the present invention, the first coordinate system transformation depends on the average torso pose.
[0045] According to an optional feature of the present invention, the first coordinate system transformation aligns the first category coordinate system with the user torso pose.
[0046] According to an optional feature of the present invention, the audiovisual presentation device further includes a receiver for receiving device pose data indicating the pose of an external device, and the first coordinate system transformation depends on the device pose data.
[0047] In many embodiments, this can provide an improved user experience and / or improved performance and / or facilitated implementation. In particular, in some applications, it can provide a very beneficial experience where the audiovisual item pose follows the overall movement of the user rather than smaller / faster head movements, thereby providing an improved user experience with an improved out-of-head experience.
[0048] According to an optional feature of the present invention, the first coordinate system transformation depends on the average device pose.
[0049] According to an optional feature of the present invention, the first coordinate system transformation aligns the first category coordinate system with the device pose.
[0050] According to an optional feature of the present invention, the mapper is arranged to select the first presentation category in response to user movement parameters indicating the movement of the user.
[0051] According to an optional feature of the present invention, the mapper is arranged to determine a coordinate system transformation between the real-world coordinate system and the coordinate system of the user's head movement data in response to user movement parameters indicative of the movement of the user.
[0052] According to an optional feature of the present invention, the mapper is arranged to determine the user movement parameters in response to the user's head movement data.
[0053] According to an optional feature of the present invention, at least some of the presentation category flags indicate whether the audiovisual item for the at least some presentation category flags is a narrative audiovisual item or a non-narrative audiovisual item.
[0054] According to an optional feature of the present invention, the audiovisual item is an audio item, and the presenter is arranged to generate an output binaural audio signal for a binaural presentation device by applying binaural presentation to the audio item using the presentation pose.
[0055] According to another aspect of the present invention, there is provided a method of presenting an audiovisual item, the method comprising: receiving an audiovisual item; receiving metadata including an input pose for each of at least some of the audiovisual items in the audiovisual item and a presentation category flag for each of at least some of the audiovisual items in the audiovisual item, the input pose being provided with reference to an input coordinate system, and the presentation category flag indicating a presentation category from a set of presentation categories; receiving user head movement data indicative of the movement of the user's head; mapping the input pose to a presentation pose in a presentation coordinate system in response to the user head movement data, the presentation coordinate system being fixed relative to the head movement; presenting the audiovisual item using the presentation pose; wherein each presentation category is linked to a coordinate system transformation from a real-world coordinate system to a category coordinate system, the coordinate system transformation being different for different categories, and at least one category coordinate system is variable relative to the real-world coordinate system and the presentation coordinate system; and the method includes selecting a first presentation category from the set of presentation categories for the first audiovisual item in response to the presentation category flag for the first audiovisual item and mapping the input pose for the first audiovisual item to a presentation pose in the presentation coordinate system, the presentation pose corresponding to a fixed pose in the first category coordinate system for a changing user head movement, the first category coordinate system being determined according to a first coordinate system transformation for the first presentation category.
[0056] These and other aspects, features and advantages of the present invention will become apparent with reference to the embodiments described hereinafter and will be elucidated with reference to the embodiments described hereinafter. Description of the Drawings
[0057] Embodiments of the present invention will be described with reference to the accompanying drawings by way of example only, in which:
[0058] Figure 1 An example of a client-server based virtual reality system is illustrated;
[0059] Figure 2 An example of elements of an audiovisual presentation device according to some embodiments of the present invention is illustrated;
[0060] Figure 3 Illustrated by Figure 2 Examples of possible presentation methods of the audiovisual presentation device;
[0061] Figure 4 Illustrated by Figure 2 Examples of possible presentation methods of the audiovisual presentation device;
[0062] Figure 5 Illustrated by Figure 2 Examples of possible presentation methods of the audiovisual presentation device;
[0063] Figure 6 Illustrated by Figure 2 Examples of possible presentation methods of the audiovisual presentation device; and
[0064] Figure 7 Illustrated by Figure 2 Examples of possible presentation methods of the audiovisual presentation device. Detailed Description
[0065] The following description will focus on embodiments in which audiovisual items, including both audio items and visual items, are provided by a presentation that includes both an audio presentation and a visual presentation. However, it should be understood that the methods and principles described can also be applied individually and separately to, for example, the presentation of only audio items or the video / visual / image presentation of only visual items (such as visual objects in a video scene).
[0066] The description will also focus on virtual reality applications, but it will be understood that the methods described can be used in many other applications, including augmented and mixed reality applications.
[0067] Virtual reality (including augmented and mixed reality) experiences that allow users to move around in a virtual or augmented world are becoming increasingly popular, and services that meet such needs are being developed. In many such ways, virtual and audio data can be dynamically generated to reflect the current pose of the user (or viewer).
[0068] In this field, the terms orientation and pose are used as common terms for position and / or direction / orientation. For example, the combination of the position and direction / orientation of an object, a camera, a head, or a view can be referred to as a pose or an orientation. Thus, an orientation or pose indication can include up to six values / components / degrees of freedom, where each value / component typically describes an individual property of the position / location or orientation / direction of the corresponding object. Of course, in many cases, for example, if one or more components are considered to be fixed or irrelevant, the orientation or pose can be represented by fewer components (e.g., if all objects are considered to be at the same height and have a horizontal orientation, four components can provide a complete representation of the pose of the object). Hereinafter, the term pose is used to refer to a position and / or orientation that can be represented by one to six values (corresponding to the maximum possible degrees of freedom).
[0069] Many VR applications are based on poses with the maximum degrees of freedom, i.e., three degrees of freedom for each of position and orientation, resulting in a total of six degrees of freedom. Thus, a pose can be represented by a set or vector of six values representing these six degrees of freedom, and thus, a pose vector can provide a three-dimensional position and / or a three-dimensional direction indication. However, it should be understood that in other embodiments, a pose can also be represented by fewer values.
[0070] A system or entity that provides the maximum degrees of freedom for a viewer is typically referred to as having six degrees of freedom (6DoF). Many systems and entities provide only orientation or position, and these are typically referred to as having three degrees of freedom (3DoF).
[0071] Typically, virtual reality applications generate three-dimensional output in the form of separate view images for the left and right eyes. These can then be presented to the user via suitable devices, such as the individual left and right eye displays of a typical VR headset. In other embodiments, one or more view images can be presented, for example, on an autostereoscopic display, or in fact, in some embodiments, only a single two-dimensional image can be generated (e.g., using a conventional two-dimensional display).
[0072] Similarly, for a given viewer / user / listener pose, an audio representation of the scene can be provided. The audio scene is typically presented to provide a spatial experience where the audio sources are perceived to originate from desired positions. Since the audio sources can be static in the scene, a change in the user pose will result in a change in the relative position of the audio sources with respect to the user's pose. Thus, the spatial perception of the audio sources can change to reflect the new position with respect to the user. The audio presentation can be adapted accordingly based on the user pose.
[0073] Viewer or user pose input can be determined in different ways in different applications. In many embodiments, the physical movement of the user can be directly tracked. For example, a camera monitoring the user area can detect and track the user's head (or even eyes (eye tracking)). In many embodiments, the user can wear a VR headset that can be tracked by external and / or internal devices. For example, the headset can include an accelerometer and a gyroscope that provide information about the movement and rotation of the headset and thus the head. In some examples, the VR headset can emit signals or include (e.g., visual) identifiers that enable an external sensor to determine the position and orientation of the VR headset.
[0074] In some systems, a VR application can be provided locally to a viewer by a stand-alone device, e.g., without using any remote VR data or processing or even without any access to any remote VR data or processing. For example, a device such as a game console can include a storage device for storing scene data, an input for receiving / generating the viewer pose, and a processor for generating corresponding images and ( / or) audio based on the scene data.
[0075] In other systems, VR / scene data can be provided from a remote device or server.
[0076] For example, a remote device can generate audio data representing an audio scene and can transmit audio components / objects / signals or other audio elements corresponding to different audio sources in the audio scene and position information indicating their positions (which can change dynamically, e.g., for moving objects). The audio elements can include elements associated with specific positions, but can also include elements for more distributed or diffuse audio sources. For example, audio elements representing general (non-localized) background sounds, ambient sounds, diffuse reverberation, etc. can be provided.
[0077] The local VR device can then appropriately present the audio elements, e.g., by applying appropriate binaural processing that reflects the relative positions of the audio sources for the audio components.
[0078] Similarly, a remote device can generate visual / video data representing a visual / video scene and can transmit visual scene components / objects / signals or other visual elements corresponding to different objects in the visual scene and position information indicating their positions (which can change dynamically, e.g., for moving objects). The visual items can include elements associated with specific positions, but can also include video items for more distributed sources.
[0079] In some embodiments, visual items may be provided as individual and separate items, such as, for example, descriptions of individual scene objects (e.g., size, texture, opacity, reflectivity, etc.). Alternatively or additionally, visual items may be represented as part of an overall model of the scene, e.g., including descriptions of different objects and their relationships to each other.
[0080] For VR services, in some embodiments, a central server may correspondingly generate audio-visual data representing a three-dimensional scene, and may specifically represent audio by a plurality of audio items that can be presented by a local client / device and represent the visual scene by a plurality of video items that can be presented by a local client / device.
[0081] Figure 1 An example of a VR system is illustrated, in which a central server 101 communicates with a plurality of remote clients 103 via a network 105 (e.g., the Internet), for example. The central server 101 may be arranged to support potentially a large number of remote clients 103 simultaneously.
[0082] In many cases, this approach can provide, for example, an improved trade-off between complexity and the resource requirements (communication requirements, etc.) of different devices. For example, the scene data may be transmitted only once or relatively infrequently, where the local rendering device (remote client 103) receives the viewer's pose and processes the scene data locally to render audio and / or video to reflect changes in the viewer's pose. This approach can provide an efficient system and an attractive user experience. It can, for example, substantially reduce the required communication bandwidth while providing a low-latency real-time experience, while allowing the scene data to be stored, generated, and maintained centrally. It can, for example, be suitable for applications where VR experiences are provided to multiple remote devices.
[0083] Figure 2 Elements of an audio-visual presentation device that can provide improved audio-visual presentations in many applications and scenarios are illustrated. Specifically, the audio-visual presentation device can provide improved presentations for many VR applications, and the audio-visual presentation device can be specifically arranged to perform processing and rendering for Figure 1 the VR client 103.
[0084] Figure 2 The audio device of is arranged to present a three-dimensional scene by presenting spatial audio and video to provide a three-dimensional perception of the scene. The specific description of the audio-visual presentation device will focus on applications that provide the described method to both audio and video, but it should be understood that in other embodiments, the method may be applied only to audio or video / visual processing, and in fact, in some embodiments, the presentation device may include only functions for presenting audio or only for presenting video, i.e., the audio-visual presentation device may be any audio presentation device or any video presentation device.
[0085] The audio-visual presentation device includes a first receiver 201, which is arranged to receive audio-visual items from a local or remote source. In a specific example, the first receiver 201 receives data describing the audio-visual item from a server 101. The first receiver 201 may be arranged to receive data describing a virtual scene. The data may include data providing a visual description of the scene and may include data providing an audio description of the scene. Thus, the audio scene description and the virtual scene description may be provided by the received data.
[0086] The audio item may be encoded audio data, such as an encoded audio signal. The audio item may be different types of audio elements including different types of signals and components, and in fact in many embodiments, the first receiver 201 may receive audio data defining different types / formats of audio. For example, the audio data may include audio represented by audio channel signals, individual audio objects, scene-based audio (such as higher-order ambisonics (HOA), etc.). The audio may be represented, for example, as encoded audio of a given audio component to be presented.
[0087] The first receiver 201 is coupled to a renderer 203, which continues to present the scene based on the received data describing the audio-visual item. In the case of encoded data, the renderer 203 may also be arranged to decode the data (or in some embodiments, the decoding may be performed by the first receiver 201).
[0088] Specifically, the renderer 203 may include an image renderer 205, which is arranged to generate an image corresponding to the current viewing pose of the viewer. For example, the data may include spatial 3D image data (such as an image of the scene and depth or model description), and accordingly, the visual renderer 203 may generate stereoscopic images (images for the user's left and right eyes), as would be known to those skilled in the art. The images may be presented to the user, for example, via the individual left and right eye displays of a VR headset.
[0089] The renderer 203 also includes an audio renderer 207, which is arranged to present an audio scene by generating an audio signal based on the audio item. In this example, the audio renderer 207 is a binaural audio renderer that generates binaural audio signals for the user's left and right ears. The binaural audio signals are generated to provide a desired spatial experience and are typically reproduced by a headset or earphones, which may specifically be part of a headset worn by the user, and the headset also includes a left eye display and a right eye display.
[0090] Thus, in many embodiments, the audio rendering by the audio renderer 207 is a binaural rendering process that uses a suitable binaural transfer function to provide the desired spatial effect for the user wearing the headphones. For example, the audio renderer 207 may be arranged to generate audio components that are perceived as arriving from a specific location using binaural processing.
[0091] Binaural processing is known to provide a spatial experience by using virtual localization of sound sources with individual signals for the listener's ears. With appropriate binaural rendering processing, the signals required at the eardrums to make the listener perceive that the sound is coming from any desired direction can be calculated, and the signals can be rendered such that they provide the desired effect. These signals are then re-generated at the eardrums using either headphones or a crosstalk cancellation method (suitable for rendering over closely spaced speakers). Binaural rendering can be considered a method for generating signals for the listener's ears so as to trick the human auditory system into perceiving that the sound is coming from a desired location.
[0092] Binaural rendering is based on binaural transfer functions, which vary from person to person due to the acoustic properties of the head, ears, and reflective surfaces such as shoulders. For example, binaural filters can be used to create binaural recordings simulating multiple sources at various positions. This can be achieved by convolving each sound source with, for example, a head-related impulse response (HRIR) pair corresponding to the position of the sound source.
[0093] A well-known method for determining the binaural transfer function is binaural recording. It is a method of recording sound using a specialized microphone arrangement and is intended for replay using headphones. The recording is made by placing microphones in the ear canals of an object or using a mannequin head with embedded microphones, a bust including the pinna (outer ear). The use of such a mannequin head including the pinna provides a very similar spatial impression as if the person listening to the recording were present during the recording.
[0094] By measuring the response of a microphone placed in or near the human ear to a sound source at a specific location in, for example, 2D or 3D space, an appropriate binaural filter can be determined. Based on such measurements, a binaural filter can be generated that reflects the acoustic transfer function to the user's ears. Binaural filters can be used to create binaural recordings simulating multiple sources at various positions. This can be achieved, for example, by convolving each sound source with the measured impulse response pair for the desired position of the sound source. To create the illusion of a sound source moving around the listener, a large number of binaural filters with a certain spatial resolution, such as 10 degrees, are typically required.
[0095] The head-related binaural transfer function can be represented, for example, as a head-related impulse response (HRIR), or equivalently a head-related transfer function (HRTF), or a binaural room impulse response (BRIR), or a binaural room transfer function (BRTF). The transfer function (e.g., estimated or assumed) from a given location to the listener's ears (or eardrums) can be given, for example, in the frequency domain, in which case it is typically referred to as an HRTF or BRTF; or in the time domain, in which case it is typically referred to as an HRIR or BRIR. In some scenarios, the head-related binaural transfer function is determined to include aspects or properties of the acoustic environment and in particular the room in which the measurements are made, while in other examples only user characteristics are considered. Examples of the first type of function are BRIR and BRTF.
[0096] The audio renderer 207 can accordingly include a storage device having binaural transfer functions for a typically high number of different locations, where each binaural transfer function provides information on how an audio signal should be processed / filtered so as to be perceived as originating from that location. Applying binaural processing individually to multiple audio signals / sources and combining the results can be used to generate an audio scene with multiple audio sources positioned at appropriate locations in the sound field.
[0097] The audio renderer 207 can select and retrieve the stored binaural transfer function that most closely matches the desired location (or in some cases can be generated by interpolating between multiple neighboring binaural transfer functions) for a given audio element to be perceived as originating from a given location relative to the user's head. It can then apply the selected binaural transfer function to the audio signal of the audio element, thereby generating an audio signal for the left ear and an audio signal for the right ear.
[0098] The output stereo signal in the form of the generated left and right ear signals is suitable for headphone presentation and can be amplified to generate a drive signal that is fed to the user's headset. The user will then perceive the audio element as originating from the desired location.
[0099] It should be appreciated that in some embodiments, the audio item can also be processed to, for example, add acoustic environment effects. For example, the audio item can be processed to add reverberation or, for example, decorrelation / diffusivity. In many embodiments, such processing can be performed on the generated binaural signals rather than directly on the audio element signals.
[0100] Thus, the audio renderer 207 can be arranged to generate an audio signal such that a given audio element is presented such that a user wearing headphones perceives the audio element as being received from the desired location. Other audio items can, for example, possibly be distributed and diffuse and can be presented as such.
[0101] It should be appreciated that many algorithms and methods for presenting spatial audio, for example using headphones and specifically for binaural presentation, will be known to those skilled in the art and any suitable method can be used without departing from the present invention.
[0102] The audiovisual presentation apparatus further includes a second receiver 209, which is a metadata receiver arranged to receive metadata for the audiovisual item. The metadata particularly includes position data for one or more of the audiovisual items. The metadata may include an input position indicating the position of one or more of the audiovisual items.
[0103] The received audiovisual data may include audio and / or visual data describing the scene. The audiovisual data specifically includes audiovisual data for a set of audiovisual items corresponding to audio sources and / or visual objects in the scene. Some audio items may represent localized audio sources in the scene associated with a specific position and / or orientation in the scene (where the position and / or orientation may change dynamically for a moving object). The visual data may include data describing the scene objects, thereby allowing the generation and representation of visual representations of these in the image / video presented to the user (usually a 3D image using a separate display of a head-mounted kit).
[0104] Generally, an audio element may represent audio generated by a specific scene object in the virtual scene and may thus represent an audio source at a position corresponding to the position of the scene object (e.g., human speech). In this case, the same position data / indications may be included and used for the audio item and the corresponding visual scene object (and similarly for the orientation).
[0105] Other elements may represent more distributed or diffuse audio sources, such as, for example, ambient or background noise that may be diffuse. As another example, some audio elements may represent fully or partially the non-spatial localized component of audio from a localized audio source, such as, for example, the diffuse reverberation from a spatially well-defined audio source.
[0106] Similarly, some visual scene objects may have an extended position and, for example, the position data may indicate the center or reference position of the scene object.
[0107] The metadata may include pose data that indicates the position and / or orientation of the audiovisual item and specifically indicates the position and / or orientation of the audio source and / or visual scene object or element. The pose data may, for example, include absolute position and / or orientation data defining the position of each or at least some of the items.
[0108] The pose is provided with reference to an input coordinate system, i.e., the input coordinate system is the reference coordinate system for the pose indication provided in the metadata received by the second receiver. The input coordinate system is typically a coordinate system fixed with reference to the represented / presented scene. For example, the scene can be a virtual (or real) scene where the audio source and the scene objects are at the positions provided with reference to the scene, i.e., the input coordinate system is typically the scene coordinate system of the scene represented by the audiovisual data.
[0109] To provide a user representation of the scene, the scene will be presented from the viewer's or user's pose, i.e., the scene will be presented as it is perceived in the given viewer / user pose in the scene, and where the audio and visual presentation provides the audio and images that would be perceived for that viewer pose.
[0110] The presentation by the presenter 203 is performed with respect to a presentation coordinate system that is fixed relative to the user's head and head movement. The reproduction of the presented audio signal is typically a head-mounted or mounted reproduction device, such as head-mounted headphones / earphones and an individual display for the eyes. Typically, the reproduction is by a head-mounted kit device that includes audio and video reproduction means. The presenter is arranged to generate the presented audiovisual item with reference to the reproduction device / means, and it is assumed that the user's pose with reference to the reproduction device / means is constant / fixed. For example, the position of the audio source is determined relative to the head-mounted headphones (i.e., the position in the presentation coordinate system), and the appropriate HRTF filter for that position is retrieved and used to present the audio signal such that the audio source is perceived to arrive from the desired relative position in the presentation coordinate system. Similarly, for a given scene object, the relative position relative to the display (i.e., the position in the presentation coordinate system) is determined, and the image corresponds to the views from the left-eye pose and the eye pose respectively relative to that position.
[0111] Therefore, the presentation coordinate system can be considered as a coordinate system fixed to the user's head, and specifically as a presentation coordinate system independent of head movement or actually changes in the user's pose. The reproduction device (whether audio, visual, or both audio and visual) is assumed / considered to be fixed relative to the user's head and thus relative to the presentation coordinate system.
[0112] The presentation coordinate system fixed relative to the user's head movement can be considered to correspond to a reproduction coordinate system fixed relative to the reproduction device for reproducing the presented audiovisual item. The term presentation coordinate system can be equivalent to the reproduction device / means coordinate system and can thus be replaced. Similarly, the term "presentation coordinate system fixed relative to the user's head movement" can be equivalent to "reproduction device / means coordinate system fixed relative to the reproduction device for reproducing the presented audiovisual item" and can be replaced by it.
[0113] When a renderer refers to a rendering coordinate system to perform rendering based on a pose and provides a pose of an audiovisual item with reference to an input coordinate system, the audiovisual rendering device includes a mapper 211 which is arranged to map an input position in the input coordinate system to a rendering position in the rendering coordinate system.
[0114] The audiovisual rendering device includes a head movement data receiver 213 for receiving user head movement data indicating movement of a user's head. The user head movement data may indicate movement of the user's head in the real world and is typically provided with reference to a real world coordinate system. The user head movement data may indicate absolute or relative movement of the user's head in the real world and may specifically reflect an absolute or relative change in the user's pose with respect to the real world. The coordinate system head movement data may indicate a change (or no change) in the head pose (orientation and / or position) and may also be referred to as head pose data.
[0115] It should be understood that many different possible methods for detecting and representing head movement are known and any suitable method may be used without departing from the present invention. The head movement data receiver 213 may specifically receive head movement data from a VR headset or a VR head movement detector, as known in the art.
[0116] The mapper 211 is coupled to the head movement data receiver 213 and receives the user head movement data. The mapper is arranged to perform the mapping between the input position in the input coordinate system and the rendering position in the rendering coordinate system in response to the user head movement data. For example, the mapper 211 may continuously process the user head movement data to continuously track the current user pose in the real world coordinate system. Then, the mapping between the input pose and the rendering pose may be based on the user pose.
[0117] For example, in many applications, it is desirable to provide the user with an experience as if he were present in the three-dimensional scene being represented. Therefore, it is desirable for the presented audio and images to reflect the user pose following the movement of the user's head. Therefore, it is desirable for the audiovisual item to be presented such that they are perceived as being fixed with respect to the real world, as this enables the reproduction of real world movement in the presentation of the (usually virtual) scene.
[0118] In such a case, the mapping from the input pose to the rendered pose is such that the audiovisual items appear fixed relative to the real world, i.e., they are rendered as being perceived as fixed relative to the real world. Thus, the same input pose is mapped to different rendered poses to reflect changes in the user's head pose. For example, if the user turns his head by, say, 30°, the real-world scene is referenced to a -30° rotation of the user. The mapper 211 can perform the corresponding change such that the mapping from the input pose to the rendered pose is modified to include an additional 30° rotation relative to the situation before the user's head rotation. Thus, the audiovisual items will be in different poses in the rendered coordinate system but will be perceived as being in the same real-world pose. Thus, the mapping can be changed dynamically such that the audiovisual items are perceived as fixed relative to the real world and thus a very natural experience is provided.
[0119] For example, to create the illusion of an immersive virtual world, three-dimensional audio and / or visual rendering is typically controlled by head tracking, where the rendering compensates for head pose, specifically including head orientation changes (such as yaw, pitch, roll) in three spatial degrees of freedom or 3-DOF. The rendering is such that the audiovisual items are perceived as fixed relative to the user. Compared to static rendering, the effect of this head tracking and subsequent rendering adaptation is a high sense of realism and exocentric perception of the rendered content.
[0120] However, another approach is to map the input pose to a rendered pose that is fixed relative to the rendered coordinate system. This can be done, for example, by the mapper 211 applying a fixed mapping from the input pose to the rendered pose, where the mapping is independent of head movement data and, specifically, where changes in the user's pose do not result in a change in the mapping between the input pose and the rendered pose. The effect of this mapping is effectively to make the perceived scene move with the head, i.e., it is static relative to the user's head. Although this may seem unnatural for most scenes, it can be advantageous in some scenes. For example, it can provide the desired experience for music or for listening to sounds that are not part of the scene (such as a narrator, for example).
[0121] Different methods can be applied to different audiovisual items. In MPEG terminology, the terms "fixed with respect to head orientation" or "not fixed with respect to head orientation" are used to refer to audio items that are to be rendered to either fully follow or ignore user movement.
[0122] For example, an audio item may be considered "head-independent", which means that it is an audio element designed to have a fixed position in a (virtual or real) environment, and thus its rendering dynamically adapts to the (changing) head orientation of the user. Another audio item may be considered "head-fixed", which means that it is an audio item designed to have a fixed position relative to the user's head. Such audio items can be presented independently of the listener's pose. Thus, the rendering of such audio items does not take into account the (change in) head orientation of the user. In other words, such audio items are audio elements whose relative position does not change when the user turns their head (e.g., non-spatial audio such as ambient noise or music that is designed to follow the user without changing relative position).
[0123] In the described system, the second receiver 209 is arranged to receive metadata that also includes a presentation category flag for at least some of the audiovisual items. The presentation category flag indicates a presentation category from a set of presentation categories, and the rendering of the audiovisual item is performed according to the presentation category indicated for the audiovisual item. Different presentation categories may define different presentation parameters and operations.
[0124] The presentation category flag can be any indication that can be used to select a presentation category from a set of presentation categories. In many embodiments, it can be data provided solely for the purpose of selecting a presentation category and / or can be data that directly specifies a category. In other embodiments, the presentation category flag can be an indication that can also provide additional information or provide some description of the corresponding audiovisual item. In some embodiments, the presentation category flag can be a parameter considered when selecting a presentation category, and other parameters can also be considered.
[0125] As a specific example, in some embodiments, an audio item may be encoded audio data, such as an encoded audio signal, where the audio item can be different types of audio items including different types of signals and components, and in fact, in many embodiments, the metadata receiver 201 can receive metadata that defines different types / formats of audio. For example, the audio data may include audio represented by audio channel signals, individual audio objects, high-order ambisonics (HOA), etc. The metadata can be included as part of the audio item or separately from the audio item that describes the audio type of each audio item. This metadata can be a presentation category flag and can be used to select an appropriate presentation category for the audio item.
[0126] Rendering categories are specifically associated with different reference coordinate systems, and specifically, each rendering category is linked to a coordinate system transformation from the real-world coordinate system to the category coordinate system. Specifically, for each category, a coordinate system transformation can be defined that transforms the real-world coordinate system (such as specifically, the real-world coordinate system to which the head movement data is referenced) into a different coordinate system given by the transformation. Since different categories have different coordinate system transformations, they will be linked to different category reference systems.
[0127] The coordinate system transformation is typically a dynamic coordinate system transformation for one, some, or all categories. Thus, the coordinate system transformation is typically not a fixed or static coordinate system transformation, but can vary over time and depend on different parameters. For example, as will be described in more detail later, the coordinate system transformation can depend on parameters that change dynamically, such as for example user torso movement, external device movement, and / or indeed even head movement data. Thus, in many embodiments, the coordinate system transformation of a category is a time-varying coordinate system transformation that depends on user movement parameters. User movement parameters can indicate the movement of the user relative to the real-world coordinate system.
[0128] The mapping performed by the mapper 211 for a given audiovisual item depends on the category coordinate system of the rendering category to which the audiovisual item is indicated as belonging. Specifically, based on the rendering category flag, the mapper 211 can determine the rendering category intended for presenting the audiovisual item. Then, the mapper can determine the coordinate system transformation linked to the selected category. Then, the mapper 211 can continue to perform the mapping from the input pose to the rendering pose such that these correspond to fixed poses in the category coordinate system resulting from the selected coordinate system transformation.
[0129] Thus, the category coordinate system can be considered as a reference coordinate system with respect to which the audiovisual item is presented as being fixed. The category coordinate system can also be referred to as a reference coordinate system or a fixed reference coordinate system (for a given category).
[0130] In many embodiments, one rendering category can correspond to a rendering where the audio source and the scene object represented by the audiovisual item are fixed relative to the real world as previously described. In such embodiments, the coordinate system transformation is such that the category coordinate system of the category is aligned with the real-world coordinate system. For such a category, the coordinate system transformation can be a fixed coordinate system transformation and can be, for example, a one-to-one mapping that is a unity of the real-world coordinate system. Thus, the category coordinate system can effectively be the real-world coordinate system or, for example, a fixed static translation, scaling, and / or rotation.
[0131] In many embodiments, a rendering category may correspond to a rendering in which the audio source and the scene object represented by the audiovisual item are fixed relative to the movement of the head (i.e., relative to the rendering coordinate system). In such an embodiment, the coordinate system is transformed such that the category coordinate system of the category is aligned with the user's head / reproduction device / rendering coordinate system. For such a category, the coordinate system transformation may be a coordinate system transformation that fully follows the movement of the user's head. For example, any rotation of the head is followed by a corresponding rotation in the coordinate system transformation, and any change in the position of the user's head is followed by the same change in the coordinate system transformation. Thus, according to such a rendering category, the coordinate system transformation is dynamically modified to follow the head movement data such that the resulting category coordinate system is aligned with the rendering coordinate system, thereby resulting in a fixed mapping from the input coordinate system to the rendering coordinate system as previously described.
[0132] Although not required, in many embodiments, the rendering category may correspondingly include a category in which the audiovisual item is rendered as fixed relative to the real-world coordinate system and a category in which the audiovisual item is rendered as fixed relative to the rendering coordinate system. However, in the described system, one or more of the rendering categories include a rendering category in which the listening item is fixed in a coordinate system that is neither the real-world coordinate system nor the rendering coordinate system, i.e., a rendering category that provides a rendering of the audiovisual item that is neither fixed in the real world nor fixed relative to the user's head. Thus, at least one category coordinate system is variable relative to the real-world coordinate system and the rendering coordinate system. Specifically, the coordinate system is different from the rendering coordinate system and the real-world coordinate system, and in fact the differences between these and the category coordinate system are not constant but can change.
[0133] Thus, at least one rendering category may provide a rendering that is neither fixed to the real world nor fixed to the user. More precisely, in many embodiments, it may provide an intermediate experience.
[0134] For example, the coordinate system transformation can be such that the corresponding coordinate system is fixed relative to the real world, except when an update criterion is met. However, if the criterion is met, the coordinate system transformation can be adapted to provide a different relationship between the real world coordinate system and the category coordinate system. For example, the mapping can be such that the audiovisual is presented as fixed relative to the real world coordinate system, i.e., the audiovisual item appears to be in a fixed position. However, if the user rotates their head by more than a given amount, the relationship between the real world coordinate system and the category coordinate system is changed to compensate for the rotation. For example, as long as the user's head moves less than, say, 20°, the audiovisual item is presented such that it is in a fixed position. However, if the user moves their head by more than 20°, the category coordinate system is rotated 20° relative to the real world coordinate system. This can provide the experience that as long as the movement is small enough, the user perceives a natural three-dimensional experience of the presented audiovisual item. However, for large head movements, the presentation of the audiovisual item is realigned with the modified head position.
[0135] As a specific example, an audio source corresponding to a narrator can initially be presented as being directly in front of the user. For small user movements, the audio is presented such that the narrator is perceived as being stationary in the same position. This provides a natural experience and perception, and in particular provides an out-of-head perception of the narrator. However, if the user rotates their head by more than, for example, 20° from the original direction towards the narrator audio source, the system adapts the mapping to reposition the narrator audio source in front of the user's new orientation. For small movements around this point, the narrator audio source is presented at this new fixed position (relative to the real world coordinate system). If the movement again exceeds a given threshold with respect to this new audio source position, the category coordinate system and thus the perceived position of the narrator audio source can be updated again. Thus, a narrator can be provided for the user who, for smaller movements, is fixed relative to the real world, but for larger movements, follows the user. In the example described, it can allow the narrator to be perceived as being fixed and provide appropriate spatial cues regarding head movement, but always be essentially in front of the user (even if, for example, the user turns a full 180°).
[0136] In this method, the metadata includes presentation category identifiers for a plurality of audiovisual items, thereby allowing the source side to control flexible presentation at the receiving side, where the presentation is specifically adapted to individual audiovisual items. For example, different spatial presentations and perceptions can be applied to items corresponding to, for example, background music, narration, audio sources corresponding to specific objects fixed in the scene, conversations, etc.
[0137] In some embodiments, the coordinate system transformation for the presentation category depends on user head movement data. Thus, in some embodiments, changing at least one parameter of the coordinate system transformation depends on user head movement data.
[0138] In many embodiments, the coordinate system transformation may depend on user head pose attributes or parameters determined from user head movement data. For example, as previously described, if the user head pose indicates a rotation beyond a certain amount, the coordinate system transformation may be adapted to include a rotation corresponding to that amount. As another example, the mapper 211 may detect that the user has maintained a (sufficiently) constant pose for longer than a given duration, and if so, the coordinate system transformation may be adapted to position the audiovisual item at a given position in the coordinate system, i.e., having a specific position relative to the user (e.g., directly in front of the user).
[0139] In some embodiments, the coordinate system transformation depends on the average head pose. In particular, in some embodiments, the coordinate system transformation may be such that the category coordinate system is aligned with the average head pose. In some embodiments, the coordinate system transformation may be such that the category coordinate system is fixed relative to the average head pose.
[0140] The average head pose may be determined, for example, by low-pass filtering the head pose measurements with a low-pass filter having a suitable cut-off frequency (such as specifically by applying an unweighted average over a window of a suitable duration).
[0141] In some embodiments, for one or more presentation categories, the reference for presentation may be selected to be the average head orientation h. This has the effect that the audiovisual item follows slower, longer-lasting head orientation changes, such that the sound source appears to remain at the same point relative to the head (e.g., in front of the face), but rapid head movements will cause the audiovisual item to appear fixed relative to the real world (and thus also appear fixed in the virtual world) rather than relative to the head. Thus, typical small and rapid head movements during daily life will still produce an immersive and out-of-head illusion while still allowing an overall perception of the audiovisual item following the user.
[0142] In some embodiments, the adjustment and tracking may be made non-linear such that if the head rotates significantly, the average head orientation reference is, for example, "clipped" so as not to deviate from the instantaneous head orientation by more than a certain maximum angle. For example, if the maximum value is 20 degrees, an out-of-head experience is achieved as long as the head "swings" within those + / - 20 degrees. If the head rotates rapidly and exceeds the maximum value, the reference will follow the head orientation (with a maximum 20-degree lag), and once the movement stops, the reference becomes stable again.
[0143] In some embodiments, at least one presentation category is associated with a coordinate system transformation that depends on the user's torso pose. In such embodiments, the audiovisual presentation device may include a torso pose receiver 215 arranged to receive user torso pose data indicative of the user's torso pose.
[0144] The torso pose may be determined, for example, by a dedicated inertial sensor unit positioned or worn on the torso. As another example, the torso pose may be determined by sensors in a smart device (such as a smartphone) when it is worn in a pocket. As yet another example, coils may be placed on the user's torso and head respectively, and the movement of the head relative to the torso may be determined based on changes in the coupling between these.
[0145] In such embodiments, changing at least one parameter of the coordinate system transformation depends on the torso pose data.
[0146] In many embodiments, the coordinate system transformation may depend on torso pose data properties or parameters determined from the torso pose data.
[0147] For example, if the torso pose data indicates that the torso has rotated by more than a certain amount, the coordinate system transformation may be adapted to include a rotation corresponding to that amount. As another example, the mapper 211 may detect that the user has maintained a (sufficiently) constant torso pose for longer than a given duration, and if so, the coordinate system transformation may be adapted to position the audiovisual item at a given location in the presentation coordinate system, where the location corresponds to the torso pose.
[0148] In particular, in some embodiments, the coordinate system transformation may align the category coordinate system with the user's torso pose. In some embodiments, the coordinate system transformation may be such that the category coordinate system is fixed relative to the user's torso pose. Thus, in some embodiments, the audiovisual item may be presented as following the user's torso, and thus may provide a perception and experience where the audiovisual item follows the user's movement of their entire body, but appears fixed with respect to head movement relative to the torso. This may provide both the desired experience of exocentric perception and the presentation of an audiovisual item that follows the user.
[0149] In some embodiments, the coordinate system transformation may depend on the average torso pose. The average torso pose may be determined, for example, by low-pass filtering head pose measurements with a low-pass filter having a suitable cut-off frequency (such as specifically by applying an unweighted average over a window of a suitable duration).
[0150] Thus, in some embodiments, one or more of the presentation categories may employ a coordinate system transformation that is aligned with the instantaneous or average chest / torso orientation tAligned presentation provides a reference. In this way, the audiovisual item can appear to stay in front of the user's body rather than in front of the user's face. By rotating the head relative to the chest / trunk, the audiovisual item can still be perceived from all directions, again greatly increasing the immersive out-of-head experience.
[0151] In some embodiments, the coordinate system transformation for the presentation category depends on the device pose data indicating the pose of the external device. In such an embodiment, the audiovisual presentation device may include a device pose receiver 217, which is arranged to receive the device pose data indicating the device trunk pose.
[0152] The device can be, for example (hypothetically), a device worn, carried, attached to, or otherwise fixed relative to the user. In many embodiments, the external device can be, for example, a mobile phone or a personal device, such as, for example, a smartphone in a pocket, a body-mounted device, or a handheld device (e.g., a smart device for viewing visual VR content).
[0153] Many devices include gyroscopes, accelerometers, GPS receivers, etc., which allow the relative or absolute orientation of the device. Then, the device can determine the current relative or absolute orientation and send it to the device pose receiver 217 using a suitable communication (which is usually wireless). For example, the communication can be via a WiFi or Bluetooth connection.
[0154] In such an embodiment, changing at least one parameter of the coordinate system transformation depends on the device pose data.
[0155] In many embodiments, the coordinate system transformation can depend on the device pose data attributes or parameters determined from the device pose data.
[0156] For example, if the device pose data indicates that the device has rotated by more than a certain amount, the coordinate system transformation can be adapted to include a rotation corresponding to that amount. As another example, the mapper 211 can detect that the device has maintained a (sufficiently) constant trunk pose for longer than a given duration, and if so, the coordinate system transformation can be adapted to position the audiovisual item at a given position in the presentation coordinate system, where the position corresponds to the device pose.
[0157] In particular, in some embodiments, the coordinate system transformation can be such that the category coordinate system is aligned with the device pose. In some embodiments, the coordinate system transformation can be such that the category coordinate system is fixed relative to the device pose. Thus, in some embodiments, the audiovisual item can be presented to follow the device pose, and thus a perception and experience of the audiovisual item following the device movement can be provided. In many actual user scenarios, the device can provide a good indication of the user pose. For example, a body-worn device or a smartphone in a pocket, for example, can provide a good reflection of how the user moves overall. It can provide a good reference for determining relative head movement and thus can provide an experience that combines a realistic response to head movement while allowing the audiovisual item to follow the user's greater movement.
[0158] In addition, using an external device as a reference can be highly practical and provide a reference that leads to the desired user experience. The method can be based on a device that is typically already worn or carried by the user and includes the required functionality for determining and transmitting the device pose. For example, most people currently carry a smartphone that already includes an accelerometer, etc. for determining the device pose and a communication device (such as Bluetooth) suitable for transmitting the device pose data to the audiovisual presentation device.
[0159] In some embodiments, the coordinate system transformation can depend on the average device pose. The average device pose can be determined, for example, by low-pass filtering the device pose measurements with a low-pass filter having a suitable cut-off frequency (such as specifically by applying an unweighted average over a window of a suitable duration).
[0160] Thus, in some embodiments, one or more in the presentation can employ a coordinate system transformation that provides a reference for a presentation aligned with the instantaneous or average device orientation. In this way, the audiovisual item can appear to remain in a fixed position relative to the device, such that it moves when the device moves, but remains fixed relative to head movement, thus providing a more natural feeling and an immersive out-of-head experience.
[0161] Thus, the method can provide a method in which metadata can be used to control the presentation of audiovisual items such that these can be individually controlled to provide different user experiences for different audiovisual items. The experience includes providing one or more options for the perception of an audiovisual item that is neither completely fixed in the real world nor completely follows the user (for head fixation). Specifically, an intermediate experience can be provided in which the audiovisual item is presented to be fixed relative to the real world to some extent and follows the user's movement to some extent.
[0162] It should be understood that, in some embodiments, possible presentation categories may be determined in advance, where each category is associated with a predetermined coordinate system transformation. In such embodiments, the audiovisual presentation device may store the coordinate system transformation for each category, or equivalently directly store the mapping corresponding to the coordinate system transformation when appropriate. The mapper 211 may be arranged to retrieve the stored coordinate system transformation (or mapping) for the selected presentation category and apply the coordinate system transformation (or mapping) when performing the mapping for the audiovisual item.
[0163] For example, for a first audiovisual item, the presentation category identifier may indicate that it should be presented according to the first category. The first category may be a category for presenting an audiovisual item fixed to the head pose, and thus the mapper may retrieve a mapping that provides a fixed one-to-one mapping between the input pose and the presentation pose. For a second audiovisual item, the presentation category identifier may indicate that it should be presented according to the second category. The second category may be a category for presenting an audiovisual item fixed to the real world, and thus the mapper may retrieve a coordinate system transformation that adapts the mapping such that head movement is compensated, resulting in a presentation pose corresponding to a fixed position in real space. For a third audiovisual item, the presentation category identifier may indicate that it should be presented according to the third category. The third category may be a category for presenting an audiovisual item fixed to the device pose or the torso pose. The mapper may retrieve a coordinate system transformation or mapping that adapts the mapping such that head movement relative to the device or torso pose is compensated, resulting in the presentation of an audiovisual item fixed relative to the device or torso pose.
[0164] It should be understood that although different presentation categories are associated with different coordinate system transformations such that the mapping results in a presentation position fixed relative to the category coordinate system, the mapper 211 does not need to explicitly determine such a coordinate system transformation or the category coordinate system. Instead, in a typical embodiment, a mapping function is defined for an individual presentation category such that the resulting presentation pose is fixed relative to the category coordinate system. For example, a mapping function that is a function of the head pose relative to the device (or torso) pose may be used to directly map the input position to a presentation position fixed relative to the category coordinate system, which is fixed relative to the device (or torso) pose.
[0165] In some embodiments, metadata may include data that partially or fully characterizes, describes, and / or defines one or more of a plurality of presentation categories. For example, in addition to a presentation category flag, the metadata may include data describing a coordinate system transformation and / or mapping function to be applied to one or more categories. For example, the metadata may indicate that a first presentation category requires a fixed mapping from an input location to a presentation location, a second category requires a mapping that fully compensates for head movement such that items appear fixed relative to the real world, and a third presentation category that should compensate the mapping for head movement relative to an average head movement such that an intermediate experience is perceived, where for smaller and faster head movements the items appear fixed, but for slow average movements also appear to follow the user.
[0166] In different embodiments, different methods and data may be used as presentation category flags. In some embodiments, each category may be associated with, for example, a category number, and the presentation category flag may directly provide the number of the category to be used for the audiovisual item.
[0167] In many embodiments, the presentation category flag may indicate an attribute or characteristic of the audiovisual item, and this may be mapped to a specific presentation category flag.
[0168] In some embodiments, the presentation category flag may specifically indicate whether the audiovisual item is a narrative audiovisual item or a non - narrative audiovisual item. A narrative audiovisual item may be an audiovisual item that belongs to, for example, a scene of a movie or story being presented; in other words, a narrative audiovisual item originates from a source within a movie, story, etc. (e.g., an actor in a script, a bird with its sound in a nature movie, etc.). A non - narrative audiovisual item may be an item that originates outside of the movie or story (e.g., a director's audio commentary, mood music, etc.). In many cases, according to MPEG terminology, a narrative audiovisual item may correspond to "not fixed with respect to head orientation", and a non - narrative audiovisual item may correspond to "fixed with respect to head orientation".
[0169] In some embodiments, there may actually be only two presentation categories, and specifically, one may correspond to the presentation of audiovisual items indicated as narrative, and one may correspond to the presentation of audiovisual items indicated as non - narrative.
[0170] For example, in some applications and systems, narrative signaling may be sent downstream to an audiovisual presentation device, and the desired presentation behavior may depend on this signaling, as illustrated with respect to Figure 3 as follows:
[0171] · By using head tracking with a real-world orientation as a reference, it is desired that the narrative source D appears to remain stable in its position within the virtual world V and should thus be presented as being fixed relative to the real world. In a given example of a movie application, if the actor's voice is presented directly in front of the user and the user rotates their head 50 degrees to the left, the sound will be presented to the headphone that projects 50 degrees to the right, making it appear to remain in the same virtual position.
[0172] · Non-narrative sound sources N can alternatively be presented independently of head orientation. In other words, the audio remains in a "hard-coupled" fixed position relative to the head (e.g., in front of it) and rotates with the head. This is achieved by not applying a head-orientation-related mapping to the audio but instead using a fixed mapping from the input position to the presentation position. In the movie example, the director's commentary audio can be presented precisely in front of the user, and any head movement will have no effect on it, i.e., the sound remains in front of the head.
[0173] Figure 2 The audiovisual presentation device is arranged to provide a more flexible method, where at least one optional presentation category allows the presentation of audiovisual items that allows the presentation of audiovisual items to be fixed relative to the real / virtual world for some movements and to follow the user for other movements. This method can be specifically applied to non-narrative audiovisual items.
[0174] In a specific example, alternative or additional options for presenting non-narrative sound sources can include one or more of the following:
[0175] · The reference for presentation can be chosen to be the average head orientation / pose h (as Figure 4 shown). This has the following effect: Non-narrative sound sources follow slower and longer-lasting head-orientation changes, such that the sound source appears to remain at the same point relative to the head (e.g., in front of the face), but rapid head movements will make the non-narrative sound source appear fixed in the virtual world (rather than fixed relative to the head). Thus, typical small and rapid head movements during daily life will still produce an immersive and out-of-head illusion.
[0176] As an improvement to some embodiments, since it is generally desired that non-narrative audio remains at least closely in the same virtual position, the tracking can be made non-linear, such that if the head rotates significantly, the average head-orientation reference is, for example, "clipped"
[0177] Deviate from the instantaneous head orientation by no more than a certain maximum angle. For example, if the maximum value is 20 degrees, as long as the head "swings" within those + / - 20 degrees, an out-of-head experience is achieved. If the head rotates rapidly and exceeds the maximum value, the reference will follow the head orientation (with a maximum lag of 20 degrees), and once the movement stops, the reference becomes stable again.
[0178] · The reference for head tracking can be selected to be the instantaneous or average chest / trunk orientation / pose t (as Figure 5 shown). In this way, non-diegetic content will appear to stay in front of the user's body rather than in front of the user's face. By rotating the head relative to the chest / trunk, non-diegetic content can still be heard from all directions, greatly increasing the immersive out-of-head experience again.
[0179] · The reference for head tracking can be selected to be the instantaneous or average orientation / pose of an external device (such as a mobile phone or a body-worn device). In this way, non-diegetic content will appear to stay in front of the device rather than in front of the user's face. By rotating the head relative to the device, non-diegetic content can still be heard from all directions, greatly increasing the immersive out-of-head experience again.
[0180] In some embodiments, the presentation can be arranged to operate in different modes according to user movement. For example, if the user movement meets the movement criteria, the audiovisual presentation device can operate in a first mode, and if the user movement does not meet the movement criteria, the audiovisual presentation device can operate in a second mode. In this example, the two modes can provide different category coordinate systems for the same presentation category flag, that is, depending on the user movement, the audiovisual presentation device can present a given audiovisual item using a given presentation category flag fixed with reference to different coordinate systems.
[0181] In some embodiments, the mapper 211 can be arranged to select the presentation category of a given presentation category flag in response to user movement parameters indicating the user's movement. Therefore, different presentation categories can be selected for a given audiovisual item and presentation category flag value according to the user movement parameters. Specifically, if the user movement parameters meet a first criterion, a given link between the possible presentation category flag values and the set of presentation categories can be used to select the presentation category of the received presentation category flag. However, if the criterion is not met (or, for example, a different criterion is met), the mapper 211 can use a different link between the possible presentation category flag values and the same or different set of presentation categories to select the presentation category for the received presentation category flag.
[0182] The method can, for example, enable the audiovisual presentation device to provide different presentations and experiences for mobile users and stationary users.
[0183] In some embodiments, the selection of the presented category may also depend on other parameters, such as, for example, user settings or configuration settings of an application (e.g., an app on a mobile device).
[0184] In another approach, the mapper may be arranged to determine a coordinate system transformation between a real-world coordinate system that serves as a reference for a coordinate system transformation of the selected category in response to user movement parameters indicating the movement of the user and the coordinate system with respect to which user head movement data is provided.
[0185] Thus, in some embodiments, the adaptation of the presentation can be introduced by a coordinate system transformation of the presented category with respect to a reference, which can vary with respect to a (typically real-world) coordinate system indicating head movement. For example, based on user movement, some compensation can be applied to the head movement data, e.g., to compensate for some movement of the user as a whole. As an example, if the user is, for example, on a boat, the user head movement data can not only indicate the movement of the user relative to the body or the body relative to the boat, but can also reflect the movement of the boat. This may be undesirable, and thus the mapper 211 can compensate the head movement data for user movement parameter data that reflects the component of the user movement caused by the movement of the boat. The resulting modified / compensated head movement data is then given with respect to a coordinate system that has compensated for the movement of the boat, and the coordinate system transformation of the selected category can be directly applied to this modified coordinate system to achieve the desired presentation and user experience.
[0186] It should be understood that the user movement parameters can be determined in any suitable manner. For example, in some embodiments, it can be determined by a dedicated sensor that provides relevant data. For example, an accelerometer and / or a gyroscope can be attached to a vehicle carrying the user, such as a car or a boat.
[0187] In many embodiments, the mapper may be arranged to determine user movement parameters in response to the user head movement data itself. For example, long-term averaging or underlying movement analysis that identifies, for example, periodic components (e.g., corresponding to waves of a moving boat) can be used to determine user parameters indicating user movement parameters.
[0188] In some embodiments, the coordinate system transformations of two different presented categories can depend on the user head movement data, but wherein the dependence has different time-averaging properties. For example, one presented category can be associated with a coordinate system transformation that depends on the average head movement but has a relatively low average time (i.e., a relatively high cut-off frequency of an average low-pass filter), while the other presented category is also associated with a coordinate system transformation that depends on the average head movement but has a higher average time (i.e., a relatively low cut-off frequency of an average low-pass filter).
[0189] As an example, the mapper 211 can evaluate the head movement data against criteria that reflect the consideration that the user is in a static environment. For example, the low-pass filtered position change can be compared to a given threshold, and if it is below the threshold, the user can be considered to be in a static environment. In this case, the rendering can be the same as in the specific examples described above for narrative and non-narrative audio items. However, if the position change is above the threshold, the user can be considered to be in a moving environment, such as during a stroll, or using a vehicle such as a car, train, or plane. In this case, the following rendering methods can be applied (see also Figure 6 and 7 ):
[0190] · The average head orientation h can be used as a reference to render the non-narrative sound source N, so as to keep the non-narrative sound source mainly in front of the face, but still allow small head movements to create an "out-of-head" experience. In other examples, for instance, the user's torso or the device's pose can be used as a reference. Thus, this method can correspond to the method that can be used for narrative sound sources in static situations.
[0191] · The average head orientation h can also be used as a reference but with a further offset corresponding to a longer-term average head orientation h ' to render the narrative sound source D. Thus, in this case, the narrative virtual sound source will appear to be in a fixed position relative to the virtual world V, but the virtual world V will appear to "move with the user", that is, it maintains an almost fixed orientation relative to the user. The averaging (or other filtering process) used to obtain h ' is typically chosen to be much slower than the averaging used for h . Thus, the non-narrative source will follow faster head movements, while the narrative content (and the virtual world V as a whole) will take some time to re-orient with the user's head.
[0192] It will be appreciated that, for clarity, the above description has described embodiments of the present invention with reference to different functional circuits, units, and processors. However, it will be apparent that any suitable functional distribution between different functional circuits, units, or processors can be used without departing from the present invention. For example, functions illustrated as being performed by separate processors or controllers can be performed by the same processor. Thus, the reference to a particular functional unit or circuit is only considered as a reference to a suitable device for providing the described function, and does not indicate a strict logical or physical structure or organization.
[0193] The present invention can be implemented in any suitable form, including hardware, software, firmware, or any combination thereof. Optionally, the present invention can be at least partially implemented as computer software running on one or more data processors and / or digital signal processors. The elements and components of the embodiments of the present invention can be physically, functionally, and logically implemented in any suitable manner. In fact, the functions can be implemented in a single unit, in multiple units, or as part of other functional units. Thus, the present invention can be implemented in a single unit, or can be physically and functionally distributed among different units, circuits, and processors.
[0194] Although the present invention has been described in connection with some embodiments, it is not intended to limit the present invention to the specific forms set forth herein. On the contrary, the scope of the present invention is limited only by the claims. Additionally, although features may seem to be described in connection with specific embodiments, those skilled in the art will recognize that the various features of the described embodiments can be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.
[0195] Furthermore, although listed separately, multiple devices, elements, circuits, or method steps can be implemented by, for example, a single circuit, unit, or processor. Additionally, although the individual features may be included in different claims, these features can be advantageously combined, and the inclusion in different claims does not mean that the combination of the features is not feasible and / or disadvantageous. The inclusion of a feature in a class of claims does not mean a limitation to that class, but rather indicates that the feature is equally applicable to other claim classes where appropriate. Moreover, the order of the features in the claims does not mean any particular order in which the features must operate, and in particular, the order of the individual steps in method claims does not mean that the steps must be performed in that order. Rather, the steps can be performed in any suitable order. Additionally, a singular reference does not exclude a plurality. Thus, references to "a", "an", "first", "second", etc. do not exclude a plurality. The reference numerals in the claims are provided merely to make the examples clear and should not be construed as limiting the scope of the claims in any way.
Claims
1. An audiovisual presentation device, comprising: a first receiver (201) arranged to receive audiovisual items; a metadata receiver (209) arranged to receive metadata, the metadata including an input pose for each of at least some of the audiovisual items and a presentation category flag for each of at least some of the audiovisual items, the input pose being provided with reference to an input coordinate system, and the presentation category flag indicating a presentation category from a set of presentation categories; a receiver (213) arranged to receive user head movement data indicating movement of a user's head; a mapper (211) arranged to map the input pose to a presentation pose in a presentation coordinate system in response to the user head movement data, the presentation coordinate system being fixed relative to the head movement; a presenter (203) arranged to present the audiovisual items using the presentation pose; wherein, for at least one of the audiovisual items, the corresponding input pose represents a center or reference position; each presentation category is linked to a coordinate system transformation from a real-world coordinate system to a category coordinate system, the coordinate system transformation being different for different presentation categories, and at least one category coordinate system is variable relative to the real-world coordinate system and the presentation coordinate system; and the mapper is arranged to select a first presentation category for the first audiovisual item from the set of presentation categories in response to the presentation category flag for the first audiovisual item, and to map the input pose for the first audiovisual item to a presentation pose in the presentation coordinate system, the presentation pose corresponding to a fixed pose in a first category coordinate system for varying user head movement, the first category coordinate system being determined according to a first coordinate system transformation for the first presentation category.
2. The audiovisual presentation device according to claim 1, wherein, A second coordinate system transformation for a second category causes the category coordinate system for the second category to align with the user head movement.
3. The audiovisual presentation device according to claim 1 or 2, wherein, A third coordinate system transformation for a third category causes the category coordinate system for the third category to align with the real-world coordinate system.
4. The audiovisual presentation device according to any one of the preceding claims, wherein, The first coordinate system transformation depends on the user head movement data.
5. The audiovisual presentation device according to claim 4, wherein, The first coordinate system transformation depends on an average head pose.
6. The audiovisual presentation device according to claim 5, wherein, The first coordinate system transformation aligns the first category coordinate system with the average head pose.
7. The audiovisual presentation device according to any one of claims 4-6, wherein, The different coordinate system transformations for different presentation categories depend on the user head movement data, and the dependencies of the first coordinate system transformation and the different coordinate system transformations on the user head movement have different time-averaging properties.
8. The audiovisual presentation device according to any one of the preceding claims, further comprising a receiver (215) arranged to receive user torso pose data indicating a user's torso pose, and the first coordinate system transformation depends on the user torso pose data.
9. The audiovisual presentation device according to any one of the preceding claims, further comprising a receiver (217), the receiver being arranged to receive device pose data indicating the pose of an external device, and the first coordinate system transformation depending on the device pose data.
10. The audiovisual presentation device according to any one of the preceding claims, wherein, The mapper (211) is arranged to select the first presentation category in response to user movement parameters indicating movement of the user.
11. The audiovisual presentation device according to any one of the preceding claims, wherein, The mapper (211) is arranged to determine a coordinate system transformation between the real-world coordinate system and the coordinate system for the user head movement data in response to user movement parameters indicating movement of the user.
12. The audiovisual presentation device according to claim 10 or 11, wherein, The mapper (211) is arranged to determine the user movement parameters in response to the user head movement data.
13. The audiovisual presentation device according to any one of the preceding claims, wherein, At least some of the presentation category flags indicate whether the audiovisual item for the at least some of the presentation category flags is a narrative audiovisual item or a non-narrative audiovisual item.
14. The audiovisual presentation device according to any one of the preceding claims, wherein, The audiovisual item is an audio item, and the renderer (211) is arranged to generate an output binaural audio signal for a binaural presentation device by applying a binaural presentation to the audio item using the presentation pose.
15. A method of presenting an audiovisual item, the method comprising: Receiving an audiovisual item; Receiving metadata including an input pose for each of at least some of the audiovisual items in the audiovisual item and a presentation category flag for each of at least some of the audiovisual items in the audiovisual item, the input pose being provided with reference to an input coordinate system, and the presentation category flag indicating a presentation category from a set of presentation categories; Receiving user head movement data indicating movement of the user's head; Mapping the input pose to a presentation pose in a presentation coordinate system that is fixed relative to the head movement in response to the user head movement data; Presenting the audiovisual item using the presentation pose; wherein, for at least one of the audiovisual items, the corresponding input pose represents a center or reference position; each presentation category is linked to a coordinate system transformation from the real-world coordinate system to a category coordinate system, the coordinate system transformation being different for different presentation categories, and at least one category coordinate system being variable relative to the real-world coordinate system and the presentation coordinate system; and the method includes: selecting a first presentation category for the first audiovisual item from the set of presentation categories in response to the presentation category flag for the first audiovisual item, and mapping the input pose for the first audiovisual item to a presentation pose in the presentation coordinate system, the presentation pose corresponding to a fixed pose in a first category coordinate system for varying user head movement, the first category coordinate system being determined according to a first coordinate system transformation for the first presentation category.
Citation Information
Patent Citations
Head tracking
US10015620B2
Near-field binaural rendering
US20190215638A1
Associated spatial audio playback
WO2019141900A1
Spatial audio augmentation
WO2020012067A1