Apparatus and method for audio-visual rendering
The apparatus and method improve extended reality rendering in vehicles by differentiating between intentional and involuntary user movements, reducing the impact of vehicle motion on the rendered scene, and enhancing user experience and processing efficiency.
Patent Information
- Application Number
- JP2025504219
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-29
- Filing Date
- 2023-07-26
- Publication Date
- 2025-08-01
AI Technical Summary
Existing audio-visual rendering technologies for extended reality applications in mobile environments, such as vehicles, face challenges in accurately adapting to complex movements, leading to suboptimal user experiences due to conflicts between sensory inputs and increased complexity in processing and resource usage.
An apparatus and method that includes a receiver for audio-visual data, a vehicle motion signal, a relative user motion signal, a predictor to generate predicted relative user motion, a residual signal generator, and a viewing pose determiner to differentiate between intentional and involuntary user movements, enabling improved rendering by accounting for vehicle motion.
Enhances user experience by reducing the impact of vehicle motion on the rendered scene, improving flexibility and efficiency in processing, and providing a more accurate and adaptable audio-visual experience, particularly in vehicles.
Smart Images

Figure 2025524950000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an apparatus and method for audio - visual (audiovisual) rendering, and more particularly, although not exclusively, to the rendering of audiovisual signals for extended - reality applications for a user who is subject to the movement of a vehicle.
Background Art
[0002] In recent years, the diversity and scope of image and video applications have increased significantly as new services and methods for using and consuming video are continuously developed and introduced.
[0003] For example, one increasingly popular service is to provide an image sequence in such a way that the viewer can interact actively and dynamically with the system to change the rendering parameters. A very attractive feature in many applications is the ability to change the effective viewing position and viewing direction of the viewer, such as being able to move around within the presented scene.
[0004] Such features particularly enable a virtual reality experience to be provided to the user. Thereby, the user can move around (relatively) freely within the virtual environment and dynamically change their position and the place they are looking at. Usually, such virtual reality (VR) applications are based on a three - dimensional model of the scene, and the model is dynamically evaluated to provide the required specific view. This approach is well - known, for example, from gaming applications in the classification of first - person shooters for computers and consoles. Other examples include augmented reality (AR) or mixed reality (MR) applications. Such applications are generally referred to as extended - reality (XR) applications.
[0005] An important feature in many (e.g., XR) applications is to determine the user's movement in the real world and adapt the audiovisual representation of virtual features to reflect the user's movement.
[0006] The detection of the user's movement can include what is called outer-to-inner (outside-in) tracking and inner-to-outer (inside-out) tracking. Outside-in VR tracking uses a camera or other sensor directed at a target to be tracked (e.g., a headset) that is placed in a stationary position and moves freely within a predetermined area covered by the sensor.
[0007] Inside-out tracking is different from outside-in tracking in that the sensor is attached to the target (e.g., incorporated into a headset). The image data captured by multiple camera sensors is used to reconstruct 3D features in the visual surrounding world. Assuming this world is static, the headset is positioned relative to this world system. In addition to the camera, other sensors such as accelerometers can be used to improve the accuracy and robustness of the headset pose estimation.
[0008] When viewing A / V content in a stationary environment using a VR headset or headphones, the rotation of the head and optionally translation (e.g., 3DoF or 6DoF (degrees of freedom)) (e.g., using one or more cameras) are continuously measured. The measured motion is fed back into the rendering process to account for the user's motion relative to the real world. For example, if the user rotates their head to the left, the associated objects will rotate to the right. If the user moves slightly laterally, the associated objects are rendered such that the user can view the objects from the side, perhaps even look around the objects, to view objects that were previously hidden. An important prerequisite for such a system to operate correctly is that the objects in the world used as a reference for determining the user's motion do not move themselves (referred to as independent motion in the context of Structure from Motion algorithms). For example, if the world reference system moves but the user does not, the user will observe a virtual scene that moves while the user is not moving. This would be an undesirable effect. Therefore, the internal and external tracking needs to be robust to objects that move independently within the real-world scene surrounding the user.
[0009] It has been proposed to provide XR applications and services that are not only suitable for use in a mobile environment but can also be adapted to reflect or correct for motion. However, this comes with many challenges and difficulties, especially as the desired behavior and experience can depend on the specific requirements and preferences of individual applications.
[0010] In some cases, when the user is moving in a mobile environment such as a train or car, it may be desirable not to classify the vehicle as an independently moving object. For example, the visual feature structure inside the vehicle can be considered to define the reference system, and the motion of the world outside the vehicle (relative motion with respect to the vehicle) can be ignored.
[0011] When the vehicle is moving purely at a constant speed, it would be relatively easy to achieve this. However, when the vehicle changes speed or vibrates, this can cause a difference in movement between the passenger and the vehicle's internal reference system (usually determined by the vehicle's appearance and geometry), creating a complex relationship that can significantly complicate the operation of the application. Usually, such a difference in movement is detected by internal and external tracking devices and can be used to render the virtual scene. Due to such effects, the user experience can be limited and / or not optimal, for example, resulting in a presentation that does not exactly match the perceived movement.
[0012] Furthermore, the processing and algorithms considering such more complex movements tend not to be optimal in terms of complexity, accuracy, use of computing resources, etc. Summary of the Invention Problems to be Solved by the Invention
[0013] Therefore, an improved approach for audio-visual rendering suitable for users affected by the movement of the vehicle would be advantageous. In particular, an approach that enables improvement in processing, flexibility, user experience, reduction in complexity, facilitation of implementation, improvement in rendering quality, improvement and / or facilitation of rendering, improvement and / or facilitation of adaptation to the movement of the user and the vehicle, and / or improvement in performance and / or operation would be advantageous.
[0014] Therefore, the present invention preferably aims to reduce, alleviate, or eliminate one or more of the above-mentioned drawbacks, either alone or in any combination. Means for Solving the Problems
[0015] According to one aspect of the present invention, there is provided an apparatus for audio-visual rendering, the apparatus comprising: a receiver configured to receive audio-visual data representing a scene; a first source supplying a vehicle motion signal indicating the motion of a vehicle; a second source supplying a relative user motion signal indicating the motion of a user relative to the vehicle; a predictor configured to generate a predicted relative user motion signal by applying a prediction model to the vehicle motion signal; a residual signal generator configured to generate a residual user motion signal indicating at least one component of the difference between the relative user motion signal and the predicted relative user motion signal; a viewing pose determiner configured to determine a viewing pose (viewpoint) depending on the residual user motion signal and the predicted relative user motion signal, wherein the dependence of the viewing pose on the residual user motion signal is different from the dependence of the viewing pose on the predicted relative user motion signal; and a renderer configured to render an audio-visual signal regarding the viewing pose from the audio-visual data.
[0016] The present invention enables improved generation of audio-visual signals for a user affected by the motion of a vehicle (automobile) in many applications and scenarios. The present invention can provide an improved user perception of a scene in many scenarios and applications, for example, a more desirable user experience. This approach enables a differentiated adaptation of the rendered audio-visual experience, in particular, for different types of motion. This approach enables a user experience that, for example, is less affected by the motion of the vehicle and can be felt more similar to an experience in a scenario where the user is not affected by the motion of the vehicle.
[0017] In some scenarios, the present approach can provide, for example, a more flexible and / or improved correction for the effects of vehicle movement. The present approach can provide an experience that can correct or adapt an audiovisual signal rendered in response to vehicle movement, for example, such that motion sickness (which typically results from a conflict between sensory inputs of different senses, such as the sense of balance and vision) or discomfort can be reduced, while an intentional movement of the user can be determined and taken into account in the rendering.
[0018] For example, the present approach can enable the adaptation of a VR or XR experience, in which case a user's movement can be reflected in the audiovisual cues (images and / or sounds) provided to the user, while the movement of a vehicle and its effect on the user can be corrected.
[0019] In many scenarios, efficient, low-complexity, and / or low-resource-demanding processing can be achieved.
[0020] This approach can be used, for example, as a specific example, in a situation where a user attempts to use a VR application that displays 3D images to the user while moving in a vehicle. While the 3D image can show a view of a 3D scene, when the user moves their head, the view of the scene can change accordingly. For example, when the user turns their head to the left, the 3D image is updated to show the view of the scene that a person within the scene would see by turning their head to the left. However, when the user is inside a moving vehicle, the movement of the user's head would seemingly reflect another component. One is the intentional movement actively executed by the user, such as moving the head horizontally. However, another component is the movement of the user's head caused by the movement of the vehicle. For example, proceeding over a speed bump can cause a sudden vertical movement of the passengers' heads inside the vehicle relative to the vehicle. This movement is involuntary and can be determined simply by the mechanical dynamics of the vehicle (such as the seat suspension, etc.) and the user's body (such as the elasticity provided by the user's neck). Thus, the movement of the vehicle affects the user and can move the user's head relative to the vehicle. When such movement is detected by a VR headset and the image is adjusted accordingly, the user will experience an unrealistic virtual world experience. The device for audiovisual rendering can provide improved performance and an improved user experience in many such / similar scenarios and cases. The device can be configured to distinguish between movements intentionally caused by the user (e.g., by turning the head) or movements caused by the movement of the vehicle without the user desiring such head movement. In this case, the rendering of the scene, particularly the viewing posture of the user within the scene, can be reflected to have different dependencies for movements that are intentional / voluntary and movements that occur simply as a function of the movement of the vehicle and are unintentional / involuntary.
[0021] In many embodiments, the viewing pose determiner can be configured to determine a viewing pose signal depending on a residual user motion signal and a predicted relative user motion signal, and the dependence of the viewing pose on the residual user motion signal is different from the dependence of the viewing pose on the predicted relative user motion signal. The renderer can be configured to render an audiovisual signal for the viewing pose (of the viewing pose signal) from the audiovisual data.
[0022] The user's motion can specifically indicate the motion of the user's head (or perhaps the eyes). The motion can be the time derivative of the pose. The signal can be one or more values that have a temporal aspect / are time-dependent / change over time.
[0023] According to a feature as an option of the present invention, the prediction model includes temporal filtering of the vehicle motion signal.
[0024] Thereby, in many embodiments, particularly advantageous processing and / or rendering are provided. The feature enables a prediction that is particularly suitable for adapting the rendering of the audiovisual signal reflecting the user's motion to a user exposed to the motion of the vehicle. The temporal filtering can be high-pass filtering of the vehicle motion signal.
[0025] According to a feature as an option of the present invention, the predictor is configured to apply temporal filtering to the relative user motion signal to generate a filtered relative user motion signal, and to predict a predicted relative user motion signal depending on the filtered relative user motion signal.
[0026] Thereby, in many embodiments, particularly advantageous processing and / or rendering are provided. The feature enables a more accurate prediction of the relative user motion directly resulting from, for example, the motion of the vehicle.
[0027] According to a feature of the present invention as an option, the prediction model includes a biomechanical model.
[0028] Thereby, in many embodiments, particularly advantageous processing and / or rendering can be provided. This feature enables, for example, a more accurate prediction of the relative user movement directly resulting from the movement of the vehicle.
[0029] According to a feature of the present invention as an option, the viewing posture determiner is configured to apply a weighting different from that for the residual user movement signal to the predicted relative user movement signal.
[0030] This feature can bring particularly advantageous effects in many embodiments and scenarios, and in particular, can perform improved adaptation and discrimination of different movements.
[0031] According to a feature of the present invention as an option, the viewing posture determiner is configured to apply a time filtering different from that for the residual user movement signal to the predicted relative user movement signal.
[0032] This feature can bring particularly advantageous effects in many embodiments and scenarios, and in particular, can perform improved adaptation and discrimination of different movements.
[0033] According to a feature of the present invention as an option, the viewing posture determiner determines a first viewing posture contribution degree from the predicted relative user movement signal, and determines a second viewing posture contribution degree from the residual user movement signal, and is configured to generate a viewing posture by combining the first viewing posture contribution degree and the second viewing posture contribution degree.
[0034] This feature can provide advantageous processing in many scenarios.
[0035] According to a feature as an option of the present invention, the viewing posture determination device extracts a first motion component from the predicted relative user motion signal by attenuating the time frequency, and determines the viewing posture so as to include a contribution from the first motion component.
[0036] This feature can provide advantageous processing in many scenarios.
[0037] According to a feature as an option of the present invention, the viewing posture determiner detects a user gesture motion component in the residual user motion signal, extracts a first motion component from the residual user motion signal depending on the user gesture motion component, and determines the viewing posture so as to include a contribution from the user gesture motion component.
[0038] This feature can provide advantageous processing in many scenarios.
[0039] According to a feature as an option of the present invention, the predictor is configured to determine the predicted relative user motion signal in response to the correlation between the relative user motion signal and the vehicle motion signal.
[0040] This feature can provide advantageous processing in many scenarios.
[0041] According to a feature as an option of the present invention, at least one of the predicted relative user motion signal and the residual user motion signal indicates a plurality of posture components, and the viewing posture determiner is configured to determine the viewing posture with different dependencies on at least a first posture component and a second posture component among the plurality of posture components.
[0042] This feature can provide advantageous processing in many scenarios.
[0043] According to a feature of the present invention as an option, the receiver is configured to receive processing data indicating the processing of at least one of the predicted relative user movement signal and the residual user movement signal, and the viewing posture determiner is configured to determine the viewing posture according to the processing instruction of the processing data.
[0044] This feature can provide advantageous processing in many scenarios.
[0045] According to a feature of the present invention as an option, the predictor is configured to generate a predicted relative user movement signal depending on the residual user movement signal.
[0046] This feature can provide particularly advantageous processing and / or rendering in many embodiments. This feature enables, for example, a more accurate prediction of the relative user movement directly resulting from the movement of the vehicle.
[0047] According to one aspect of the present invention, a method for audio-visual rendering is provided, the method comprising: receiving audio-visual data representing a scene; supplying a vehicle movement signal indicating the movement of the vehicle; supplying a relative user movement signal indicating the movement of the user with respect to the vehicle; generating a predicted relative user movement signal by applying a prediction model to the vehicle movement signal; generating a residual user movement signal indicating at least one component of the difference between the relative user movement signal and the predicted relative user movement signal; determining a viewing posture depending on the residual user movement signal and the predicted relative user movement signal, wherein the dependence of the viewing posture on the residual user movement signal is different from the dependence of the viewing posture on the predicted relative user movement signal; and rendering an audio-visual signal related to the viewing posture from the audio-visual data.
[0048] These and other aspects, features, and advantages of the present invention will become apparent from the embodiments described below and will be explained with reference to these embodiments.
Brief Description of the Drawings
[0049]
Figure 1
Figure 2
Figure 3
Embodiments for Carrying Out the Invention
[0050] Hereinafter, embodiments of the present invention will be described by way of example only with reference to the drawings.
[0051] The following description focuses on an example of audiovisual rendering of an extended reality application for a user affected by the movement of a vehicle, such as a user in a car, truck, airplane, ship, train, etc. Audiovisual rendering is the rendering of audio and / or video, and specifically can be the rendering of an audiovisual signal including both audio data and video data.
[0052] FIG. 1 shows an example of a rendering device configured to generate and render an audiovisual signal representing a given viewing posture (view pose) and a scene perceived from various actually changing viewing postures. Specifically, the rendering device can receive a motion input that indicates (directly or indirectly) a change in the user's posture as a function of time, determine a viewing posture signal (time series of viewing postures), and render / generate audio and / or video data representing a scene from these viewing postures.
[0053] In this field, the terms "placement" and "pose" are used as ordinary terms related to position and / or direction / orientation. For example, a combination of the position and direction / orientation of an object, a camera, a head, or a view can be referred to as a pose or a placement. Thus, an indication of a placement or a pose can have six values / components / degrees of freedom (6DoF), and each value / component typically describes an individual characteristic of the position / location or orientation / direction of the corresponding object. Of course, in many situations, a placement or a pose can be considered and represented with fewer components if, for example, one or more components are fixed or irrelevant (e.g., if all objects are considered to be at the same height and have a horizontal orientation, the pose of the object can be fully represented with four components). In the following, the term "pose" is used to indicate a position and / or orientation that can be represented by one to six values (corresponding to the maximum possible degrees of freedom). The term "pose" can be replaced by the term "placement". The term "pose" can be replaced by the terms "position and / or orientation". The term "pose" can be replaced by the term "position and direction" (when the pose provides information on both position and direction), the term "position" (when the pose provides information on position (presumably position only)), or the term "orientation" (when the pose provides information on orientation (presumably orientation only)).
[0054] Motion can be a series of poses and, in particular, can be a time series of pose changes / variations. Motion can represent one or more components of a pose. Motion can be a sequence / change of orientation and / or position. Motion (of an object) can be a change over time of the orientation and / or position (of that object). Motion (of an object) can be a change over time of the pose (of that object).
[0055] When the rendering device is configured to receive audiovisual data representing a scene and to render from this data an audiovisual signal that reflects the audio and / or visual perception of the scene from a viewing pose with respect to the scene. The device can generate an audiovisual signal representing a scene from a viewing pose that depends on the movement of the user within and with respect to the vehicle. The device may attempt to distinguish between movements intentionally / voluntarily caused by the user and movements resulting from involuntary movements of the vehicle rather than being spontaneous. In this case, the viewing pose from which the audiovisual signal is generated may depend differently on voluntary / intentional / known movements and involuntary / unknown movements resulting from movements of the vehicle that the user has not intentionally caused. Such an approach can provide a significantly improved user experience, for example, for VR applications used by the user within a vehicle.
[0056] In some embodiments, the audiovisual data may be audio data, while the generated audiovisual signal may be an audio signal that provides an auditory representation of the scene from the viewing pose. In some embodiments, the audiovisual data may be visual / video data, and the generated audiovisual signal may be a visual / video signal that provides a visual representation of the scene from the viewing pose. In many embodiments, the rendering device may be configured to generate an audiovisual signal that includes both audio data and visual data representing both the visual and auditory perception of the scene from the viewing pose (thus, the received audiovisual data representing the scene may be both video data and audio data).
[0057] The rendering device includes a receiver 101 configured to receive audio-visual data representing a scene. The receiver 101 can provide a three-dimensional image and / or audio data that provides a representation of a three-dimensional scene. The audio-visual data received by the receiver 101 will hereinafter also be referred to as scene data for the sake of brevity.
[0058] The receiver 101 may include, for example, a storage unit or memory in which scene data is stored and from which the scene data can be retrieved. In other embodiments, the receiver (image data source) 101 can receive or retrieve audio-visual data from any external (or other internal source). In many embodiments, the receiver 101 can receive, for example, from a video and / or audio capture system including a video camera and a microphone (array) that captures a real-world scene in real time.
[0059] Scene data can provide an appropriate representation of the scene using an appropriate image or typically a video format / representation and / or an appropriate three-dimensional audio format / representation. In some embodiments, the receiver 101 can receive scene data from various sources and / or in various formats, and from this scene data, data suitable for rendering can be generated, for example, by converting or processing the received scene data. For example, an image and depth can be received from a remote camera and processed to generate a video representation according to a given format. In some embodiments, the receiver 101 can be configured to generate three-dimensional image data by evaluating a model of the scene.
[0060] Scene data is provided according to a suitable three-dimensional data format / representation. The representation can include, for example, multi-view and depth, multi-layer (multi-plane, multi-sphere), mesh model, and / or point cloud (point set) representation. Further, in the examples described, the scene data can specifically be video data including a time component. Similarly, the scene data can include a suitable three-dimensional audio representation such as, for example, a multi-channel representation, an audio object-based spatial audio representation, etc.
[0061] Receiver 101 is coupled to renderer 103, which is configured to generate an audiovisual signal that provides a representation / view / audio from the viewing pose of the scene. In many embodiments, renderer 103 is configured to generate an audiovisual signal that includes both audio and one or more images. In particular, in many XR applications, as the user changes their pose, a continuous series of images and audio representing the scene from different view images will be generated.
[0062] A number of different techniques, algorithms, and approaches for generating such audio and / or images / videos are known to those skilled in the art, and it will be understood that any suitable approach for renderer 103 can be used without detracting from the present invention.
[0063] From the perspective of an image, the operation of renderer 103 will be described below with respect to the generation of a single image. However, it is understood that in many embodiments, the image can be part of a series of images and, in particular, can be a frame of a video sequence. In fact, the approach described can be applied to generate multiple (in many cases all) frames / images of an output video sequence.
[0064] It is understood that a large number of stereoscopic video sequences can be generated, including a video sequence for the right eye and a video sequence for the left eye. Thus, when an image is presented to a user, for example, via an AR / VR headset, this will make the three-dimensional scene appear as seen from the viewing posture. In other examples, the image is provided to the user using a tablet with an inclination sensor, and a single visual video sequence is generated. In yet other examples, multiple images are woven together or tiled for presentation on an autostereoscopic display.
[0065] The renderer 103 can perform appropriate image synthesis / generation processing to generate a view image for a specific representation / format of three-dimensional image data (or, in fact, in some embodiments, the three-dimensional image data can be converted to a different format in which the view image is generated).
[0066] For example, in the case of multi-view + depth representation, the renderer 103 can be configured to perform a view shift or projection of the received multi-view image typically based on depth information. As is known to those skilled in the art, this usually includes techniques such as pixel shift (changing the position of pixels to reflect an appropriate difference corresponding to a change in parallax), occlusion (masking) removal (usually based on embedding from other images), and pixel combination from different images.
[0067] It will be understood that many algorithms and approaches are known for synthesizing images from different three-dimensional image data formats and representations, and the renderer 103 can use any appropriate approach. Examples of suitable view synthesis algorithms are: "A review on image-based rendering", Yuan HANG, Guo-Ping ANG, Virtual Reality & Intelligent Hardware, Vol 1, Issue 1, February 2019, pp. 39-54, https: / / doi.org / 10.3724 / SP. J.2096-5796.2018.0004; "A Review of Image-Based Rendering Techniques", Shum; Kang, Proceedings of SPIE - The International Society for Optical Engineering 4067:2-13, May 2000, DOI:10.1137 / 12.386541; or Wikipedia article on 3D rendering: https: / / en.wikipedia.org / wiki / 3D_rendering; can be found in
[0068] Similarly, from an audio perspective, a number of different spatial audio rendering approaches are known, including, for example, surround sound algorithms, room / head-related transfer function (HRTF)-based algorithms, etc. Examples thereof are, for example: "Sound Rendering" by Takala; Tapio; James, Hahn. SIGGRAPH Comput. Graph. Vol. 26. pp. 211-220. doi:10.1145 / 133994.134063. ISBN 978-0897914796; "Techniques for Low Cost Spatial Audio" by Burgess; David A, Proceedings of the 5th Annual ACM Symposium on User Interface Software and Technology" pp. 53-59, doi:10.1145 / 142621.142628; "Psychoacoustic Music Sound Field Synthesis" by Ziemer, Tim, Current Research in Systematic Musicology. Vol. 7. Cham: Springer. p. 287. doi: 10.1007 / 978-3-030-23033-3; can be found in
[0069] In the example of FIG. 1, the rendering device is configured to dynamically modify and update the viewing posture in which the renderer 103 renders an image and / or audio so as to reflect the movement / motion of the user. The rendering device is further configured to be suitable for a user in a vehicle and to provide an audiovisual output signal that can improve the user experience and operation for such a scenario. In particular, the rendering device is configured to consider a plurality of movements when determining the viewing posture in which the audiovisual signal is generated, rather than determining the viewing posture to render the audiovisual signal based on a single posture / movement.
[0070] The rendering device includes a first motion information source 105 configured to supply a vehicle motion signal indicating the motion of the vehicle. The first vehicle motion signal can provide information regarding changes in at least one position or orientation component of the vehicle as a function of time.
[0071] In some embodiments, the first motion information source 105 can be coupled to or have one or more sensors of the vehicle that can determine the motion of the vehicle (guaranteed to be attached to the vehicle or located in the same place as the vehicle). For example, the first motion information source 105 can have or be coupled to a satellite navigation receiver that can continuously supply the position and possibly the orientation of the vehicle. As another example, the first motion information source 105 can have an inertial guidance system disposed on the vehicle that can determine the attitude from the measured acceleration.
[0072] The vehicle motion signal can reflect the vehicle's attitude as a function of time. The vehicle motion signal indicates the current position and / or orientation of the vehicle and reflects how these change over time. The vehicle motion signal can be a digital signal including a time-sampled motion signal, in which case each sample reflects one or more coordinates of the attitude (including one or more position coordinates and / or azimuth coordinates). For example, in some embodiments, the vehicle motion signal can have time samples, in which case each time sample includes one, more, or all of the forward / backward, left / right, up / down, pitch, yaw, roll attitude values.
[0073] In some embodiments, the first motion information source 105 can directly receive the vehicle motion signal from an external information source (e.g., a part of the vehicle) and supply this signal directly to other functions of the rendering device or after appropriate processing.
[0074] The rendering device includes a second motion information source 107, which is configured to supply a relative user motion signal indicating the motion of the user with respect to the vehicle.
[0075] In some embodiments, the second motion information source 107 can be coupled to or have one or more sensors that detect the user and derive an indicator of the user's motion with respect to the vehicle. For example, the second motion information source 107 can have or be coupled to a group of cameras that monitor the interior of the vehicle. The user can be detected in the image, and from this information, the position and / or orientation of the user (or, for example, only the head / face of the user) can be determined. As another example, distance sensors (e.g., based on infrared or ultrasonic signals) are arranged inside the vehicle so that the distance to the user from these distance sensors can be determined, and the position can be detected from this distance. As another example, the user wears one or more sensors that detect signals from appropriate transmission elements (e.g., ultrasonic transmitters) attached inside the vehicle so that the distance to these transmitters can be determined, thereby determining the attitude of the viewer with respect to the vehicle.
[0076] The relative user movement signal can reflect the user's posture with respect to the vehicle as a function of time. The relative user movement signal indicates the current position and / or orientation of the user with respect to the vehicle's posture and can reflect how these change over time. The relative user movement signal can be a digital signal including a time-sampled movement signal, in which case each sample can reflect one or more coordinates of the posture (including one or more position coordinates and / or azimuth coordinates). For example, in some embodiments, the relative user movement signal can have time samples where each sample includes one, more or all of the forward / backward, left / right, up / down, pitch, yaw, roll posture values.
[0077] It will be understood that similar characteristics apply (as appropriate) to other posture or movement signals derived from the vehicle movement signal and / or the relative user movement signal and processed by the rendering device.
[0078] In some embodiments, the second movement information source 107 can directly receive the relative user movement signal from an external source (e.g., a part of the vehicle) and supply this signal directly or after appropriate processing to other functions of the rendering device.
[0079] Thus, in this approach, the movement signal from which the audiovisual signal is generated is not based on a single posture measurement of the user, but takes into account two different movements, namely the movement of the vehicle and the movement of the user with respect to the vehicle.
[0080] The rendering device is further configured to split the relative user movement signal into different components, which are then processed in different ways such that their relative influence / contribution to the determination of the viewing posture for the renderer 103 is different.
[0081] The first motion information source 105 is coupled to a predictor 109, which is configured to predict a relative user motion signal by applying a prediction model to the vehicle motion signal. The prediction model is applied to the received vehicle motion signal, and from this signal, the prediction model predicts / estimates which of them is the relative motion of the user (relative to the vehicle).
[0082] In this way, the predictor 109 can generate a predicted relative user motion signal that reflects the relative user motion that can be expected to result from the motion of the vehicle. The predictor can generate a prediction of the relative user motion that reflects the relative user motion caused by the motion of the vehicle, specifically, it can reflect the relative user motion correlated with the motion of the vehicle. The predicted relative user motion can reflect the involuntary motion of the user resulting from the influence of the vehicle's motion on the user.
[0083] The predictor 109 is further coupled to a residual signal generator 111, which generates a residual user motion signal that at least reflects the component of the difference between the relative user motion signal and the predicted relative user motion signal. The residual signal generator 111 is thus also connected to the second motion information source 107 and receives the relative user motion signal from the second motion information source. In many embodiments, the residual user motion signal is simply generated as the relative user motion signal by subtracting the predicted relative user motion signal. For example, for each time point (i.e., each sample), subtraction for each attitude component can be performed.
[0084] In this way, the rendering device is configured to generate two motion signals, namely, a predicted relative user motion signal that can indicate the estimated / predicted motion of the user resulting from the motion of the vehicle, and a residual user motion signal that reflects the remaining motion of the user, i.e., the motion of the user that may not directly result from the motion of the vehicle. The predicted relative user motion signal can indicate the predicted / correlated / involuntary motion of the user, while the residual user motion signal can indicate the unpredicted / uncorrelated / intentional / voluntary motion of the user.
[0085] Predictor 109 and residual signal generator 111 are coupled to viewing pose determiner 113, which is configured to determine one or more viewing poses. Viewing pose determiner 113 is coupled to renderer 103, which is supplied with a viewing pose and generates an audiovisual signal therefrom, specifically, a series of view images for a changing viewing pose and / or spatial audio of the scene perceived by the user in that viewing pose (or viewing poses) can be generated.
[0086] Viewing pose determiner 113 may be configured to generate a viewing pose signal by repeatedly generating a viewing pose and including it in the viewing pose signal. For example, at each sample / time point, a new viewing pose may be determined and supplied to renderer 103 for generation of samples (images and / or audio) of the audiovisual signal.
[0087] Viewing pose determiner 113 is configured to determine a viewing pose based on a predicted relative user movement signal and a residual user movement signal. These two signals can be inherently combined. However, the viewing pose dependencies on these two signals are different. The contributions from the predicted relative user movement signal and the residual user movement signal are different, and scaling and / or filtering between these two signals are often different.
[0088] As a specific example, in many embodiments, the two signals may be filtered and / or scaled before being added. For example, the predicted relative user movement signal may be filtered to attenuate high frequencies and scaled before being added to the residual user movement signal. As a result of such an approach, the viewing pose closely follows the residual user movement signal, but only the mainly low-frequency contribution from the predicted relative user movement signal is reduced. As a result, the viewing pose is not directly following the movement of the vehicle, but is mainly determined to follow the user's movement that is the result of other, usually intentional movements (e.g., movements resulting from the user turning their head, etc.).
[0089] The predicted relative user movement signal may strictly reflect the movement generated by the vehicle to the user and may directly correspond to the movement of the vehicle. Depending on the prediction model, the predicted relative user movement signal may also include movements resulting from involuntary corrections by the user. For example, when the user is jolted up and down by a bumpy road and the vehicle makes a sudden direction change, the user is pushed sideways.
[0090] Furthermore, the residual user movement signal may reflect active corrections by the user or movements resulting from further predictions. For example, the residual user movement signal may reflect the user's movement resulting from the user's prediction when approaching a sharp curve in the road or a road bump or a traffic signal. As a result, the user will usually tense their muscles. If the prediction is timely and accurate enough, the user can completely correct the movement that would have been caused by the vehicle. When the user is wearing a VR headset, this component may not be so dominant, but there may be other inducements such as screams or a co-passenger squeezing the hand. Also, the user may be startled when the vehicle suddenly starts to change direction on the road, but it is natural to compensate for the vector of the force exerted by the vehicle and tilt in the opposite direction of the vector. Furthermore, the residual user movement signal may also reflect spontaneous movements by the user, such as the user turning their head to look around.
[0091] In this approach, the movement of the user's head relative to the world coordinate system can be separated / decomposed / split into at least two components, namely a component corresponding to the predicted relative movement and a component reflecting the remainder of the movement. Then, separate corrections for these at least two components can be performed to determine a viewing pose that can be used, for example, to generate a 3D virtual audio-visual user representation.
[0092] As another example, the predicted relative movement is corrected only when it corresponds to a large movement such as the vehicle making a sudden change of direction. In all other cases, the predicted relative movement is not corrected in the calculation of the user's pose.
[0093] As yet another example, the predicted relative movement is corrected depending on what is happening in the virtual scene (virtual content).
[0094] As yet another example, the separate correction depends on whether the user is looking at a pure virtual scene or an augmented scene. In the latter case, depending on whether the user is looking inside or outside the vehicle, the user will see overlay graphic content superimposed on the inside of the vehicle or on the road. Each of these requires a different correction for the individual conditions.
[0095] Thus, by considering a plurality of specific movement components of the user's movement and having different dependencies of the viewing pose on these components, an improved effect and a more appropriate and flexible user experience can be realized.
[0096] To determine a relative user movement signal predicted from the vehicle movement signal, different approaches and prediction models can be used by the predictor 109 in different embodiments.
[0097] In some embodiments, the prediction model can have or include temporal filtering of the vehicle motion signal. In particular, the prediction model can attenuate some frequencies relative to other frequencies.
[0098] The filtering is temporal filtering rather than spatial filtering, and thus affects the temporal characteristics / changes of the vehicle motion signal. In the time domain, the (prediction) filter can be a mathematical operation in which future values of a discrete-time signal are estimated as a linear function of previous samples. In fact, this linear function can represent a filter with a specific frequency transfer function where some frequencies are attenuated relative to other frequencies.
[0099] For example, typically, some very low frequencies of vehicle motion tend to cause corresponding user motion. For example, when the vehicle turns or accelerates, the user typically follows the vehicle motion very closely. When the vehicle travels over a speed bump, this causes an up-and-down motion, and its low frequency is typically followed by a passenger sitting in the vehicle. Thus, for such low-frequency motion, the user motion tends to follow the vehicle motion, and thus the vehicle motion tends to have relatively little impact on the relative user motion with respect to the vehicle. However, for higher frequencies, due to attenuation by the seat, the human body, etc., the vehicle motion will not be transmitted to the user motion. As a result, for these higher frequencies, the vehicle motion can be directly converted into relative user motion (e.g., with a sign inversion depending on the applied coordinate system). In this example, the predictor 109 can accordingly apply a prediction model in the form of a high-pass filter. The resulting high-pass vehicle motion signal can be used as the predicted relative user motion signal.
[0100] As another example, when the vehicle is traveling on a long, gentle curve in the road, the user will experience a constant force on one side of the body. However, the relative movement of the user will indicate that the user is resisting the force and not moving at all. In this case, no movement correction is required.
[0101] In some embodiments, the predictor 109 may further be configured to apply temporal filtering to the relative user movement signal to generate a filtered relative user movement signal. In this case, the prediction of the predicted relative user movement signal may be responsive to the filtered relative user movement signal.
[0102] Thus, in some embodiments, the prediction model may be based on both the vehicle movement signal and the relative user movement signal inputs, specifically, based on these filtered versions.
[0103] Specifically, if the vehicle makes only inaccurate attitude (position and rotation) measurements, these measurements can be made more accurate by observing the relative user movement signal. The latter is typically measured by a camera in the headset, which observes changes relative to features inside the vehicle. By calculating the correlation between the predictions made via the vehicle sensors and the relative movement of the user, it is possible to segment which part of the relative movement over time is mainly determined by the vehicle. For those segments with high correlation over time, the change in the vehicle attitude parameters can be made more accurate by assuming that the user's movement was determined only by the vehicle (and thus the correlation should have been 1).
[0104] In some embodiments, the prediction model can have a biomechanical model. The biomechanical model can provide a mechanical model of the human body, for example, a model showing how the movement of the human torso or seat area causes the movement of the user's head.
[0105] A suitable model could be, for example, a vertical bar 0.6 to 1.0 m in length having the mass of a typical human upper body representing a person sitting in the rear seat. Assuming the vehicle is traveling along a curve in the road, the curvature of the curve, the speed of the vehicle, and the mass of the upper body will determine the force exerted on the upper body. Assuming the person's muscles are completely relaxed, the person's upper body will fall until they notice they are about to fall. At least with respect to the first unanticipated part (start) of the curve, the falling motion will follow a simple model of a bar with length and mass falling. If the person senses the motion of the bar (usually within 1 second), they will attempt to correct it with muscle action and correct to an upright sitting position or maintain an upright sitting position.
[0106] As an example, a biomechanical model can be used to reflect additional motion of the user (in many embodiments, particularly the user's head) that does not directly exist in the vehicle motion signal itself but occurs directly and involuntarily from the vehicle motion. For example, a vehicle changing direction on a road can cause continuous motion of a human head and thus a headset, but that motion can be motion resisting the change in direction and thus different from the vehicle motion signal. For example, the user may tend to turn their head or torso inward to correct for the centrifugal force experienced by the user with respect to the vehicle.
[0107] An example of relative head motion can be shown by FIG. 2, which shows examples of three different situations during a vehicle direction change.
[0108] First, the vehicle is moving straight and the passenger / user is sitting straight (a). Thus, in this case, the motion of the user's head can remain substantially fixed with respect to the vehicle itself.
[0109] The car then begins to turn, and during the turn itself the passenger / user is swung to one side by the force of the turning car (b). Thus, the physical force resulting from the car turning ("centrifugal force" on the car) pushes the user towards the side of the car that is on the outside of the turn. Because the user is sitting inside the car, their head may tilt / pivot towards the outside of the car, by angle α in the example of Figure 2. Thus, simply by initiating the turn, the user's head has been tilted relative to the car (without the user having made a voluntary decision to make this relative movement).
[0110] After the turn is complete, the car will be going straight again and the user's head will return to the same straight up position as before the turn, as there will no longer be any physical forces pointing outwards from the car (c).
[0111] Thus, the user's head moves significantly relative to the car during a turn, a movement that is caused by but distinct from the car's motion, and this behavior can be modeled by a biomechanical model.
[0112] In some embodiments, the predictor 109 may be configured to determine a predicted relative user motion signal dependent on a correlation between the relative user motion signal and the vehicle motion signal. In some embodiments, the prediction model may include a correlation between the relative user motion signal and the vehicle motion signal. The prediction model may be a linear prediction model that accounts for the correlation between the relative user motion signal and the vehicle motion signal.
[0113] More specifically, to generate an empirical prediction model based on the correlation, passengers are instructed to sit relaxed in the back seat of a car wearing headsets that block all external visual signals. The car then drives around several specific road curves at different speeds, and accelerometers attached to the passenger headsets measure the tilt effect as a function of the passenger's biophysical parameters (total length, upper body height, mass, etc.). After sufficient measurements are made, a linear or nonlinear model can be fitted that can be used for real-time predictions in the future.
[0114] In some embodiments, the predictor 109 may be configured to generate a predicted relative user motion signal dependent on the residual user motion signal.
[0115] For example, the prediction model may be a feedback or loop model that can adapt one or more parameters to reduce the level or amplitude of the residual user motion signal. In some embodiments, the residual user motion signal may be fed back to the predictor 109, which may modify parameters of the model that generates the predicted relative user motion signal to reduce the amplitude or level of the residual user motion signal. This typically maximizes (or at least increases) the level of the predicted relative user motion signal. Such feedback processing may therefore provide improved prediction and more accurate determination of motion that may result directly from vehicle motion. In this way, the residual user motion signal may be used, in effect, as an error signal to adapt the prediction model.
[0116] It will be appreciated that various ways of adapting a prediction based on an error signal are known. For example, the prediction model may consist of a linear prediction filter whose coefficients are adapted (e.g., without a predetermined range) based on a least mean squares (LMS)-based linear adaptation algorithm.
[0117] As another example, the prediction may be periodically reset based on an error model.
[0118] The viewing pose determiner 113 may be configured in different embodiments to apply different approaches to determining the viewing pose(s), and the dependency on the predicted relative user motion signal and the residual user motion signal may depend on the particular preferences and requirements of each individual embodiment.
[0119] In many embodiments, the viewing pose determiner 113 may process the predicted relative user motion signal to generate a first motion component. Similarly, the viewing pose determiner 113 may process the residual user motion signal to generate a second motion component. The processing of the two motion signals may be different. Each of the above motion components describes a motion and may be represented as an absolute pose in the scene coordinate system, or perhaps as a relative pose with respect to a series of poses such as a fixed pose and / or a predetermined path through the scene. The viewing pose determiner 113 may then synthesize the first and second motion components into a viewing pose signal that includes the viewing pose at which rendering is to be performed.
[0120] Thus, the viewing pose determiner 113 may generate a first viewing pose contribution from the predicted relative user motion signal represented by the first motion component. The determiner may further generate a second viewing pose contribution from the residual user motion signal represented by the second motion component. The determiner can then generate the viewing pose at which rendering is to be performed as a combination of these separate contributions.
[0121] The above combination can simply be the sum of the contributions / components, or more generally, in many embodiments, it can be a weighted sum. For example, the relative poses of the first and second components can be added to form a single viewing pose offset signal that can be added to, for example, a fixed pose, and a changing pose signal (i.e., changing as a function of time). In some embodiments, more complex combinations can also be performed, including introducing a time offset between the signals.
[0122] The viewing pose determiner 113 is configured to apply different processing to the predicted relative user motion signal and the residual user motion signal when generating the viewing pose (signal). Thus, the effects of different types of motion on the resulting viewing pose are distinguished and can be individually adapted to provide the desired effects and user experience, specifically, for example, to correct for involuntary user motion caused by vehicle motion.
[0123] In many embodiments, the viewing pose determiner 113 can be configured to assign a different weight to the predicted relative user movement signal than to the relative user movement signal when determining the viewing pose. The processing of one movement relative to the other movement (or equivalently, the weighting in combination) can include different scalings / gains for the two movements.
[0124] For example, in some embodiments, the scaling or gain for the predicted relative user movement signal can be significantly reduced relative to the scaling or gain for the residual user movement signal. Thus, an effect can be brought about that the involuntary movement resulting from the vehicle movement signal is significantly attenuated relative to the user's voluntary movement. This effect can, for example, mitigate or reduce the influence of being in the vehicle, so that an improved VR experience can be provided to the user.
[0125] In some embodiments, the viewing pose determiner 113 can be configured to apply a different temporal filtering to the predicted relative user movement signal than to the relative user movement signal. The first and second movement components can be generated, for example, such that different frequencies are attenuated.
[0126] In particular, in some embodiments, the viewing pose determiner 113 can be configured to extract a first motion component from the predicted relative user motion signal by attenuating the time frequency of the predicted relative user motion signal. For example, very low frequencies are filtered out, such that, for example, motions corresponding to gentle accelerations are reduced or removed. For example, low frequency motions of the vehicle such as long direction changes (<0.1 Hz) can be attenuated because they are not perceived by the user. Similarly, high frequency motions such as vehicle shakes (>10 Hz) can also be completely filtered out. This is because such shakes can be annoying to the user and, if this motion is not transmitted from the physical world to the virtual world, will not cause motion-induced nausea. In the case of medium frequency motions due to the vehicle, no filtering is performed and that motion will be provided to the user. Further, no filtering is applied to the residual user motion signal, such that the viewing pose will fully follow this motion.
[0127] As a result, an experience can be provided where the virtual user view fully follows the user's voluntary motions but does not follow gentle accelerations or direction changes of the vehicle or high frequency shakes. However, the medium frequencies of the vehicle motion signal can still be reflected in the viewing pose, such that the view of the scene perceived by the user can still include some motion corresponding to the vehicle motion. If this motion is not reflected in the view presented to the user, symptoms / nausea caused by the motion may occur (this typically results from a conflict between sensory inputs of different senses such as the sense of balance and vision).
[0128] Thus, such an approach can provide a significantly improved user experience.
[0129] Of course, in other applications, other approaches can also be used. For example, depending on the VR experience provided (e.g., the story presented), or depending on how susceptible the user is to nausea in a particular frequency band, some filtering of the medium frequencies can be applied.
[0130] As another example, the VR experience can be designed to incorporate some or all of the low-speed direction changes, and the filtering coefficient corresponding to the low-speed direction change can be filtered or not filtered based on the state of the VR experience.
[0131] In some embodiments, the determination of the second motion component may include extracting a portion of the motion from the residual user motion signal. For example, while it may be desirable for a particular motion to be included in the change of the viewing posture, it may still be undesirable for a motion that is not directly predictable from the vehicle motion signal to be included in the presentation to the user.
[0132] In particular, in some embodiments, the viewing posture determiner 113 can be configured to detect the user gesture motion component within the residual user motion signal. The determiner can then extract the first motion component from the residual user motion signal by at least partially removing the user gesture motion component. Specifically, the second motion component can be generated from the residual user motion signal by subtracting the gesture motion component.
[0133] The motion of the user's gesture can specifically be a predetermined motion of the user. For example, the viewing posture determiner 113 can store a plurality of parameters of the predetermined motions of the user, each of which is described by a group of parameters that can vary within a range. In this case, the viewing posture determiner 113 can correlate the residual user motion signal with these predetermined motions of the user while varying the parameters to achieve the closest possible fit. When the correlation exceeds an appropriate level, the viewing posture determiner 113 can consider that such a predetermined motion of the user is included in the residual user motion signal as being executed by the user. The determiner then subtracts this predetermined motion of the user (regarding the identified parameters) from the residual user motion signal to generate the second motion component, which is then combined with the first motion component.
[0134] Thus, with this approach, specific predefined gesture movements of the user can be detected and removed from the user's movements included in the viewing gesture actions.
[0135] In some embodiments, the viewing gesture determiner 113 can be configured to detect the presence of a predefined user movement / user gesture by correcting the residual user movement signal for movements that are thought to result from the movement signal of the vehicle. For example, it can be preceded by correcting the residual user movement signal with a movement signal that reflects the influence of the vehicle's movement on a specific user movement in the residual user movement signal of the predefined user movement. For example, the correction can be based on a predicted relative user movement signal.
[0136] As a specific example, the residual user movement is supplied to the gesture detector and tracker. When a gesture is detected, the movement corresponding to the gesture is tracked. For example, if the gesture includes the movement of both arms, the trajectory of the arm joints is estimated by the tracker. This movement is then converted into a representation that matches the residual user movement. Thus, this signal is the component related to the gesture among the residual user movements, and the difference signal (the residual among the residuals) can be the other components.
[0137] In such an example, the second movement component can be related to the movement of a predefined user movement / gesture movement, for example, the movement of the hands and feet making the gesture.
[0138] Such an approach can be advantageous in many scenarios. For example, hand gestures made by the user inside the vehicle will be affected by the movement of the vehicle. In the case of gesture control in a virtual world, these gestures can be advantageously corrected by subtracting the movements caused by the vehicle.
[0139] In many embodiments, the movement and pose are multi-dimensional and include, for example, three position coordinates and, often, three orientation coordinates, thereby providing a full 6DoF experience. However, in some embodiments, the pose / movement may be represented by fewer dimensions, such as, for example, only three or even two spatial position coordinates.
[0140] In scenarios where multiple components may be present, the dependencies of the different pose components may differ when determining the viewing pose. Thus, the viewing pose determiner 113 may be configured to determine the viewing pose with different dependencies for at least a first pose component and a second pose component. In many embodiments, this applies when both the first component and the second component are position components or both are orientation components.
[0141] For example, the lateral dependency on the user may be different from the vertical direction. As a result, the distinction between the predicted user movement and the residual user movement will vary depending on whether it relates to the user's vertical movement in the seat or the lateral movement in the seat. For example, vertical movement in a vehicle is usually easy to convert into a virtual scene, and passing this component to the viewing pose improves realism. Passing the lateral movement is likely to cause nausea and requires more filtering.
[0142] In some embodiments, the rendering device (renderer) may further be configured to receive processing data indicating the processing of the predicted relative user movement signal and / or the residual user movement signal. The processing data may in particular define the dependency between the movement signal and the viewing pose. In this case, the viewing pose determiner 113 can adapt its processing to determine the viewing pose according to the processing instructions of the processing data.
[0143] The above processing data may be supplied in a message that is part of the received audiovisual signal, which in particular also includes audiovisual data for the scene. The data can be received, for example, as a bitstream from a remote server.
[0144] For example, the data can be packed with the audiovisual signal to optimize the user experience when inside a moving vehicle. Such processing messages can indicate, for example, for each posture component, how it should be filtered. For example, the processing message can be: - For each movement component, for example, involuntary, voluntary, anticipatory, etc.) * Optionally, for each axis * Optionally, for each device class - The message is required to have only sub - sets of components present. - The filtering value can be an attenuation coefficient, but can also include a cut - off frequency, or even be FIR or IIR filter coefficients. For example, it is possible to supply a message that provides the following processing data indicating attenuation rates for various types of posture coordinates and movement components.
[0145] [Table 1]
[0146] Such messages can also include information on how each scriptable movement event should be converted into movement in the movement response of the virtual scene. Only sub - sets of movement events must be present in the message. The above response can be a mapping or can include a filtering value. The filtering value can be an attenuation coefficient as shown in the table below, can also include a cut - off frequency, or even be an IIR filter coefficient.
[0147] [Table 2]
[0148] The following describes specific examples of available approaches. This description may be considered an example of a rendering of a view image from a viewing pose generated by the viewing pose determiner 113.
[0149] In this example, the viewing pose used to render the audiovisual signal and (virtual scene object) is associated with a common world coordinate system. In the following specific example, the viewing pose is referred to as a virtual camera (the image is generated to reflect the image that would be captured by the virtual camera in the viewing pose).
[0150] Thus, mapping a scene object point to a virtual camera can be represented as follows, where the conversion to the common world system is followed by the conversion to the virtual camera:
Equation
[0151] In this equation, M0 is the so-called model matrix that positions the virtual scene object in the virtual world coordinate system. This matrix changes as the object moves through the virtual space. Matrix T w converts a point from the virtual world to the real world.
[0152] Matrix T w usually results from an initialization step and remains constant during processing. Under user control, calibration, reset, or setup can be performed. For example, at initialization, T w can be defined by actively placing visual markers in the scene or by automatically detecting visual feature structures and placing virtual axes somewhere with respect to their positions. Cameras, depth sensors, and / or other sensors can be used to assist with this initial setup.
[0153] The view matrix V captures the position of the user's eyes in real-world space and represents the user's movement. The view matrix V is essentially the pose of the human head or eyes. Virtual reality headsets are usually stereo, so have different view matrices for the left and right eyes. The view matrix is usually estimated from sensors (cameras, depth sensors, etc.) mounted in or on the headset (inside-outside tracking).
[0154] When the vehicle changes motion, this results in different torques [m·s] acting on different body parts. The human body is flexible, the distance from the car seat to the human head is large, and a typical human head is heavy, so the moment of inertia (angular mass) is large. In the absence of early muscle predictions (based on visual input), slight vehicle motion can result in large head movements relative to the vehicle (world). This is an example of undesirable motion caused by vehicle motion.
[0155] Therefore, any vehicle movement that causes a difference in motion between the headset and the vehicle, i.e., relative user movement, will cause a change observed by the sensors in the headset. The 4x4 view matrix measured is the user's pose, and contains the user's view (rotation and translation) of the world scene:
number
[0156] Thus, in this example, the matrix V corresponds to a combination of intended and unintended motion, and the multiplied transformation matrix, i.e., the intended motion V intended T acting on vehicle These matrices are time-varying, and in this approach, the matrix T vehicle is determined by prediction and corresponds to the predicted relative user motion signal, while the intended motion V intended is determined as a residual signal, i.e., a relative user motion signal. The system can determine the viewing pose to compensate for vehicle motion / vibration, for example. In particular, the virtual world is
number
number
[0157] In this approach, the predictor 109 calculates, inter alia, the vehicle transformation matrix T vehicle We can try to estimate its affine inverse T -1 vehicle is calculated and used by the viewing pose determiner 113 to obtain the corrected view matrix V compensated The view pose represented by
[0158] As a specific example, the matrix T corresponding to the predicted relative user motion signal vehicle can be determined by filtering the view matrix V, where V is the view matrix corresponding to the vehicle motion signal combined with the relative user motion signal.
[0159] Although vehicle suspension systems are highly effective, high temporal frequency vehicle motion can still occur. Therefore, T is measured as a high-pass component of V measured by the inside and outside tracking sensors in the headset. vehicle In this way, Tvehicle The predicted relative user motion signal, denoted by , can be predicted as the high-pass component of V, which represents the vehicle motion signal. In this case, the vehicle-induced high-frequency component of the pose measured from the headset camera relative to the interior of the car is directly used to estimate the component induced by the vehicle motion.
[0160] For completeness, we demonstrate how we can handle a 4x4 homogeneous transformation matrix containing a 3x3 rotation matrix and a 3x1 translation vector. vehicle As a concrete example of determining , we first split a 4x4 view matrix into an orientation quaternion (a four-dimensional vector) and a translation vector:
number
[0161] Note that alternative representations of rotation exist and may be preferred in some cases. To measure the change in rotation and translation over time, the differential rotation and differential translation can be calculated as follows:
number
[0162] where k and k-1 correspond to two discrete time steps (e.g., 1 / 30 seconds apart). A recursive filter is then used to output a filtered rotation that ignores large instantaneous rotations assumed to be induced by vehicle motion:
number
[0163] Similarly, the translation vector is:
number
[0164] In this way, T vehicle of:
number
[0165] Therefore, the above process assumes that the effect of vehicle motion on headset motion can be derived from the high frequency motion components measured by sensors in the headset relative to the (visual) features inside the car.
[0166] In the following, we eliminate the high frequency effects caused by the vehicle. compensated =T -1 vehicle A residual user motion signal, denoted by V, may be determined. Finally, a rendering mapping may be performed according to compensated reflects the view pose determined by view pose determiner 113):
number
[0167] The data signal generating device and the rendering device may in particular be implemented by one or more suitably programmed processors, an example of a suitable processor is given below.
[0168] Figure 3 is a block diagram illustrating an exemplary processor 300 according to an embodiment of the present disclosure. Processor 300 may be used to configure one or more processors that implement the rendering apparatus of Figure 1. Processor 300 may be any suitable processor type, including, but not limited to, a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable array (FPGA) (wherein the FPGA is programmed to form a processor), a graphics processing unit (GPU), an application specific integrated circuit (ASIC) (wherein the ASIC is designed to form a processor), or a combination thereof.
[0169] Processor 300 may include one or more cores 302. Core 302 may include one or more arithmetic logic units (ALUs) 304. In some embodiments, core 302 may include a floating point logic unit (FPLU) 306 and / or a digital signal processing unit (DSPU) 308 in addition to or instead of ALU 304.
[0170] The processor 300 may include one or more registers 312 communicatively coupled to the core 302. The registers 312 may be implemented using dedicated logic gate circuits (e.g., flip-flops) and / or any memory technology. In some embodiments, the registers 312 may be implemented using static memory. The registers may provide data, instructions, and addresses to the core 302.
[0171] In some embodiments, processor 300 can include one or more levels of cache memory 310 communicatively coupled to core 302. Cache memory 310 can supply computer-readable instructions to core 302 for execution. Cache memory 310 can supply data for processing by core 302. In some embodiments, the computer-readable instructions can be supplied to cache memory 310 by local memory, such as local memory attached to external bus 316. Cache memory 310 can be implemented by any suitable cache memory type, such as metal-oxide semiconductor (MOS) memory, such as static random access memory (SRAM), dynamic random access memory (DRAM), and / or some other suitable memory technology.
[0172] Processor 300 can include a controller 314, which can control inputs to processor 300 from components included in other processors and / or systems, and / or outputs from processor 300 to components included in other processors and / or systems. Controller 314 can control data paths in ALU 304, FPLU 306, and / or DSPU 308. Controller 314 can be implemented as one or more state machines, data paths, and / or dedicated control logic. Gates of controller 314 can be implemented as individual gates, FPGAs, ASICs, or any other suitable technology.
[0173] Registers 312 and cache 310 can communicate with controller 314 and core 302 via internal connections 320A, 320B, 320C, and 320D. The internal connections can be implemented as a bus, multiplexer, crossbar switch, and / or any other suitable connection technology.
[0174] The input and output of the processor 300 may be provided via a bus 316 including one or more conductors. The bus 316 may be communicatively coupled to one or more components of the processor 300, such as the controller 314, the cache 310 and / or the register 312. The bus 316 may be coupled to one or more components of the system.
[0175] The bus 316 may be coupled to one or more external memories. The external memory may include a read-only memory (ROM) 332. The ROM 332 can be a mask ROM, an electronically programmable read-only memory (EPROM), or of any other suitable technology. The external memory may include a random access memory (RAM) 333. The RAM 333 can be a static RAM, a battery-backed static RAM, a dynamic RAM (DRAM), or of any other suitable technology. The external memory may include an electrically erasable programmable read-only memory (EEPROM (registered trademark)) 335. The external memory may include a flash memory 334. The external memory may include a magnetic storage device such as a disk 336. In some embodiments, the external memory may be included within the system.
[0176] The present invention can be implemented in any suitable form including hardware, software, firmware, or any combination thereof. Optionally, the present invention can be at least partially implemented as computer software executed on one or more data processors and / or digital signal processors. The elements and components of an embodiment of the present invention can be physically, functionally, and logically implemented in any suitable manner. In fact, the functions can be implemented within a single unit, within multiple units, or as part of other functional units. Thus, the present invention can be implemented within a single unit, or physically and functionally distributed among different units, circuits, and processors.
[0177] The terms "in response to" or "depending on" can be replaced, for example, with the terms "responsive to", "as a function of", or "based on".
[0178] Although the present invention has been described in connection with several embodiments, it is not intended to be limited to the specific form set forth herein. Rather, the scope of the present invention is limited only by the appended claims. Moreover, while features may appear to be described in connection with particular embodiments, those skilled in the art will recognize that various features of the described embodiments can be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps. A claimed method is a method excluding a method of performing mental acts per se.
[0179] Furthermore, although individually listed, a plurality of means, elements, circuits, or method steps may be implemented by, for example, a single circuit, unit, or processor. Furthermore, although individual features may be included in different claims, they can, possibly to advantage, be combined, and their inclusion in different claims does not imply that a combination of these features is not feasible and / or advantageous. Furthermore, the inclusion of a feature in one claim does not imply limitation to this claim, but rather indicates that the feature may be applied to other claim classes as well, where appropriate. Furthermore, the order of features in the claims does not imply any particular order in which the features must be performed, and in particular the order of individual steps in a method claim does not imply that the steps must be performed in this order. Rather, the steps may be performed in any suitable order. Furthermore, singular reference does not exclude a plurality; therefore, singular reference does not exclude a plurality. Reference signs in the claims are provided merely as a clarifying example and are not to be construed as limiting the scope of the claims in any way.
Claims
1. A receiver that receives audiovisual data representing a scene, a first source that supplies a vehicle motion signal indicating the motion of a vehicle, a second source that supplies a relative user motion signal indicating the motion of a user with respect to the vehicle, a predictor that generates a predicted relative user motion signal by applying a prediction model to the vehicle motion signal, a residual signal generator that generates a residual user motion signal indicating at least one component of the difference between the relative user motion signal and the predicted relative user motion signal, a viewing pose determiner that determines a viewing pose depending on the residual user motion signal and the predicted relative user motion signal, wherein the dependence of the viewing pose on the residual user motion signal is different from the dependence of the viewing pose on the predicted relative user motion signal, and a renderer that renders an audiovisual signal regarding the viewing pose from the audiovisual data A device for audiovisual rendering having the above components.
2. The device according to claim 1, wherein the prediction model includes temporal filtering of the vehicle motion signal.
3. The device according to claim 2, wherein the predictor applies temporal filtering to the relative user motion signal to generate a filtered relative user motion signal, and predicts the predicted relative user motion signal depending on the filtered relative user motion signal.
4. The device according to any one of claims 1 to 3, wherein the prediction model includes a biomechanical model.
5. The device according to any one of claims 1 to 4, wherein the viewing pose determiner applies a different weighting to the predicted relative user motion signal than to the residual user motion signal.
6. The device according to any one of claims 1 to 5, wherein the viewing pose determiner applies a different temporal filtering to the predicted relative user motion signal than to the residual user motion signal.
7. The device according to any one of claims 1 to 6, wherein the viewing pose determiner determines a first viewing pose contribution degree from the predicted relative user motion signal, and determines a second viewing pose contribution degree from the residual user motion signal, and generates the viewing pose by combining the first viewing pose contribution degree and the second viewing pose contribution degree.
8. The apparatus according to any one of claims 1 to 7, wherein the viewing posture determination device extracts a first motion component from the predicted relative user motion signal by attenuating time frequencies, and determines the viewing posture so as to include a contribution from the first motion component.
9. The apparatus according to any one of claims 1 to 8, wherein the viewing posture determiner detects a user gesture motion component in the residual user motion signal, extracts a first motion component from the residual user motion signal depending on the user gesture motion component, and determines the viewing posture so as to include a contribution from the user gesture motion component.
10. The apparatus according to any one of claims 1 to 9, wherein the predictor determines the predicted relative user motion signal in response to a correlation between the relative user motion signal and the vehicle motion signal.
11. The apparatus according to any one of claims 1 to 10, wherein at least one of the predicted relative user motion signal and the residual user motion signal indicates a plurality of posture components, and the viewing posture determiner determines the viewing posture with different dependencies on at least a first posture component and a second posture component among the plurality of posture components.
12. The apparatus according to any one of claims 1 to 11, wherein the receiver receives processing data indicating processing of at least one of the predicted relative user motion signal and the residual user motion signal, and the viewing posture determiner determines the viewing posture according to a processing instruction of the processing data.
13. The apparatus according to any one of claims 1 to 12, wherein the predictor generates the predicted relative user motion signal depending on the residual user motion signal.
14. Receiving audiovisual data representing a scene; Supplying a vehicle motion signal indicating a motion of the vehicle; Supplying a relative user motion signal indicating a motion of the user with respect to the vehicle; Generating a predicted relative user motion signal by applying a prediction model to the vehicle motion signal; Generating a residual user motion signal indicating at least one component of a difference between the relative user motion signal and the predicted relative user motion signal; Determining a viewing posture depending on the residual user motion signal and the predicted relative user motion signal, wherein a dependency of the viewing posture on the residual user motion signal is different from a dependency of the viewing posture on the predicted relative user motion signal; Rendering an audiovisual signal regarding the viewing posture from the audiovisual data; A method for audiovisual rendering, comprising:
15. A computer program comprising computer program code means for performing all steps of the method according to claim 14 when executed on a computer.
Citation Information
Patent Citations
Methods and apparatus for compensating for vehicular motion
US20170050743A1
Technologies for motion-compensated virtual reality
US20180096501A1
Virtual reality headset with relative motion head tracker
US9459692B1