Image representation of a scene

By aligning the scene coordinate system with the viewer coordinate system and optimizing image representation with an offset processor, the problems of poor image quality and high resource consumption in existing virtual reality technologies are solved, and a more efficient and flexible virtual reality experience is achieved.

CN113366540BActive Publication Date: 2025-06-13KONINKLIJKE PHILIPS NV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080011735.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-01-30
Filing Date
2020-01-19
Publication Date
2025-06-13
Estimated Expiration
2040-01-19

AI Technical Summary

Technical Problem

Existing virtual reality technologies in generating and drawing images are limited by predetermined models and high computing or communication resource requirements, resulting in poor user experience and degraded image quality.

Method used

By receiving the image representation of the scene, the viewer pose is determined, and the view image is drawn by aligning the scene coordinate system with the viewer coordinate system. The device includes an offset processor for determining an offset relative to the viewer's eye position according to the orientation of the viewer's posture, ensuring an improvement in image quality.

Benefits of technology

This achieves improved image quality within a certain range, especially when the head rotates and moves, providing a more flexible and efficient virtual reality experience, reducing data rates and complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113366540B_ABST
    Figure CN113366540B_ABST
Patent Text Reader

Abstract

An apparatus includes a receiver (301) for receiving an image representation of a scene. A determiner (305) determines a viewer pose with respect to a viewer coordinate system for a viewer. An aligner (307) aligns a scene coordinate system with the viewer coordinate system by aligning a scene reference position with a viewer reference position in the viewer coordinate system. A renderer (303) renders view images for different viewer poses in response to the image representation and the alignment of the scene coordinate system with the viewer coordinate system. An offset processor (309) determines the viewer reference position in response to an aligned viewer pose, wherein the viewer reference position depends on an orientation of the aligned viewer pose and has an offset with respect to a viewer eye position for the aligned viewer pose. The offset includes an offset component in a direction opposite to a viewing direction of the viewer eye position.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the image representation of a scene, and in particular but not exclusively to the generation of an image representation as part of a virtual reality application and the image rendering based on the image representation. Background Art

[0002] In recent years, with the continuous development and introduction of new services and the ways of using and consuming videos, the types and scopes of image and video applications have increased significantly.

[0003] For example, an increasingly popular service is to provide an image sequence in a way that allows a viewer to actively and dynamically interact with the system to change rendering parameters. In many applications, a very attractive feature is the ability to change the viewer's effective viewing position and viewing direction, thus for example allowing the viewer to move and "look around" in the presented scene.

[0004] Such a feature can in particular allow a virtual reality experience to be provided to the user. This can allow the user to move (relatively) freely in the virtual environment and dynamically change his position and where he is looking. Generally, such virtual reality applications are based on a three-dimensional model of the scene, where the model is dynamically evaluated to provide a specific requested view. Such an approach is well known, for example from game applications for computers and consoles (e.g., first-person shooter games).

[0005] It is also desirable that the presented image is a three-dimensional image, especially for virtual reality applications. In fact, in order to optimize the viewer's immersion, users generally prefer to experience the presented scene as a 3D scene. In fact, the virtual reality experience should preferably allow the user to select his / her own position, camera viewpoint, and moment relative to the virtual world.

[0006] Generally, virtual reality applications are inherently limited because they are based on a predefined model of the scene and usually on an artificial model of the virtual world. It is generally desirable to provide a virtual reality experience based on real-world capture. However, in many cases, such an approach is limited by or tends to require the construction of a virtual model of the real world based on real-world capture. Then the virtual reality experience is generated by evaluating the model.

[0007] However, current methods tend to be suboptimal and often have high computational or communication resource requirements and / or provide a suboptimal user experience (e.g., due to degraded quality or restricted freedom).

[0008] In many applications, such as virtual reality applications, a scene can be represented by an image representation, e.g., by one or more images representing a particular viewing pose of the scene. In some cases, such images can provide a wide-angle view of the scene and can cover, for example, a full 360° view or cover a full view sphere.

[0009] In many applications (and particularly for virtual reality applications), an image data stream is generated based on data representing a scene, such that the image data stream reflects the (virtual) position of a user within the scene. Such an image data stream is typically generated dynamically and in real-time so as to reflect the movement of the user within the virtual scene. The image data stream can be provided to a renderer, which renders images to the user based on the image data of the image data stream. In many applications, providing the image data stream to the renderer is via a bandwidth-constrained communication link. For example, the image data stream can be generated by a remote server and transmitted (e.g., via a communication network) to a rendering device. However, for most such applications, it is important to maintain a reasonable data rate for efficient communication.

[0010] It has been proposed to provide a virtual reality experience based on a 360° video stream, in which a full 360° view of a scene is provided by a server for a given viewer position, thereby allowing a client to generate views in different directions. In particular, one of the promising virtual reality (VR) applications is omnidirectional video (e.g., VR360 or VR180). This approach tends to incur a high data rate, and thus the number of viewpoints providing a full 360° view sphere is typically limited to a low number.

[0011] As a specific example, virtual reality glasses have entered the market. These glasses allow a viewer to experience a captured 360-degree (panoramic) video. These 360-degree videos are typically pre-captured using camera stitching, where individual images are stitched together to form a single spherical map. In some such embodiments, an image representing a full spherical view from a given viewpoint can be generated and transmitted to a driver, which is arranged to generate an image corresponding to the current view of the user for the glasses.

[0012] In many systems, an image representation of a scene can be provided, where the image representation includes images of one or more capture points / viewpoints in the scene and typically also includes the depth at the capture point / viewpoint. In many such systems, a renderer can be arranged to dynamically generate a view that matches the current local viewer pose. In such a system, the viewer pose can be dynamically determined and a view can be dynamically generated to match the viewer pose. Such an operation requires the viewer pose to be aligned or mapped to the image representation. This is typically done by positioning the viewer at a given optimal or default position in the scene / image representation at the start of the application and then tracking the viewer's movement relative to that optimal or default position. The optimal or default position is typically chosen to correspond to the position where the image representation includes image data, i.e., the capture or anchor position.

[0013] However, as the viewer pose changes from that position, view interpolation and synthesis are required, which tend to introduce degradations and artifacts, thus reducing the image quality.

[0014] Accordingly, an improved method for processing and generating an image representation of a scene would be advantageous. In particular, a system and / or method that allows for improved operation, increased flexibility, enhanced virtual reality experience, reduced data rate, increased efficiency, facilitated distribution, reduced complexity, facilitated implementation, reduced storage requirements, improved image quality, improved rendering, improved user experience, and / or improved performance and / or operation would be advantageous. SUMMARY OF THE INVENTION

[0015] Accordingly, the present invention seeks to alleviate, mitigate, or eliminate one or more of the above disadvantages, preferably individually or in any combination.

[0016] According to one aspect of the present invention, there is provided an apparatus for rendering an image, the apparatus comprising: a receiver for receiving an image representation of a scene, the image representation being provided with respect to a scene coordinate system including a reference position; a determiner for determining a viewer pose for a viewer, the viewer pose being provided with respect to a viewer coordinate system; an aligner for aligning the scene coordinate system with the viewer coordinate system by aligning the scene reference position with a viewer reference position in the viewer coordinate system; a renderer for rendering view images for different viewer poses in response to the image representation and the alignment of the scene coordinate system with the viewer coordinate system; the apparatus further comprising: an offset processor arranged to determine the viewer reference position in response to a first viewer pose which is the viewer pose after alignment has been performed, the viewer reference position depending on the orientation of the first viewer pose and having an offset relative to the viewer's eye position for the first viewer pose, the offset including an offset component in a direction opposite to the viewing direction of the viewer's eye position; wherein the receiver is arranged to receive an image data signal, the image data signal including the image representation and further including an offset indication; and wherein the offset processor is arranged to determine the offset in response to the offset indication.

[0017] The present invention can provide improved operation and / or improved performance in many embodiments. In particular, the present invention can provide improved image quality for a range of viewer poses.

[0018] In many embodiments, the method can provide an improved user experience. For example, the method can allow flexible, efficient, and / or high-performance virtual reality (VR) applications in many scenarios. In many embodiments, the method can allow or enable VR applications with a significantly improved trade-off between image qualities for different viewer poses.

[0019] The method can be particularly applicable to, for example, broadcast video services that support adjustment of movement and head rotation at the receiving end.

[0020] The image representation may include one or more images of the scene. Each image of the image representation may be associated with and linked to a viewing or capturing pose for the scene. The viewing or capturing pose may be provided with reference to the scene coordinate system. The scene reference position may be any position in the scene coordinate system. The scene reference position is independent of the viewer pose. The scene reference position may be a predetermined and / or fixed position. The scene reference position may remain unchanged between at least some successive alignments.

[0021] The rendering of the view image can occur after alignment. By being opposite to the viewing direction for the aligned viewer pose, the offset component can be in the direction opposite to the viewing direction of the viewer's eye position.

[0022] The first viewer pose can also be referred to as the aligned viewer pose (which is the viewer pose for which alignment has been performed).

[0023] The alignment / first viewer pose can indicate / represent / describe the viewer's eye position and the viewing direction of the viewer's eye position. The alignment / first viewer pose includes data that allows the determination of the viewer's eye position and the viewing direction of the viewer's eye position.

[0024] When aligning the scene coordinate system with the viewer coordinate system, the offset indication can indicate the target offset applied between the scene reference position and the viewer's eye position. The target offset can include an offset component in the direction opposite to the viewing direction of the viewer's eye position.

[0025] According to an optional feature of the present invention, the offset component is not less than 2 cm.

[0026] This can provide particularly advantageous operation in many embodiments. In many scenarios, it can allow for a sufficiently high quality improvement for many viewer poses.

[0027] In some embodiments, the offset component is not less than 1 cm, 4 cm, 5 cm or even 7 cm. It has been found that a larger offset can improve the quality of the image generated for the viewing pose corresponding to head rotation, while possibly reducing the quality of the forward view image, but usually to a much lower extent.

[0028] According to an optional feature of the present invention, the offset component does not exceed 12 cm.

[0029] This can provide particularly advantageous operation in many embodiments. In many scenarios, it can provide an improved image quality trade-off for the images generated for different viewer poses.

[0030] In some embodiments, the offset component does not exceed 8 cm or 10 cm.

[0031] According to an optional feature of the present invention, the receiver (301) is arranged to receive an image data signal, the image data signal including the image representation and also including an offset indication; and wherein, the offset processor is arranged to determine the offset in response to the offset indication.

[0032] This can provide advantageous operation in many systems and scenarios (such as especially for broadcast scenarios). It can allow for the simultaneous execution of offset optimization for multiple different rendering devices.

[0033] According to an optional feature of the present invention, the offset processor is arranged to determine the offset in response to an error metric for at least one viewer pose, the error metric depending on a candidate value of the offset.

[0034] This can provide improved operation and / or improved performance in many embodiments. In particular, it can allow for improved quality trade-offs for different viewer poses in many embodiments and can allow for dynamic optimization in many embodiments. It can also allow for low complexity and / or efficient / low resource operation in many embodiments.

[0035] In many embodiments, the offset can be determined as the offset that causes a (combined) minimum error metric for one or more viewer poses.

[0036] In some embodiments, the error metric can represent an error measure or value for a continuous range of candidate values. For example, the error metric can be represented as a function of the candidate offset, and the offset can be determined as the candidate offset at which the function is minimized.

[0037] In some embodiments, only a plurality of discrete candidate offset values are considered, and the offset processor can determine an error metric / measure / value for each of these candidate values. The offset processor can then determine the offset as the candidate offset at which the lowest error metric is found (e.g., after combining error metrics for multiple viewer poses).

[0038] According to an optional feature of the present invention, the offset processor is arranged to determine the error metric for the candidate value in response to a combination of error metrics for a plurality of viewer poses.

[0039] In many embodiments, this can provide improved operation and / or improved performance.

[0040] In some embodiments, the offset processor is arranged to determine an error metric for one viewer pose in a range of viewer poses in response to an error metric for a range of gaze directions.

[0041] According to an optional feature of the present invention, the error metric for the viewer pose and the candidate value of the offset includes an image quality metric of a view image for the viewer pose synthesized from at least one image of the image representation, the at least one image having a position relative to the viewer pose that depends on the candidate value.

[0042] In many embodiments, this can provide improved operation and / or improved performance.

[0043] According to an optional feature of the present invention, the error metric for the viewer pose and the candidate value of the offset includes an image quality metric of a view image for the viewer pose synthesized from at least two images of the image representation, the at least two images having a reference position relative to the viewer pose depending on the candidate value.

[0044] In many embodiments, this can provide improved operation and / or improved performance.

[0045] According to an optional feature of the present invention, the image representation includes an omnidirectional image representation.

[0046] In particular, the present invention can provide improved performance for an image representation based on omnidirectional images (e.g., especially omnidirectional stereo (ODS) images).

[0047] According to an optional feature of the present invention, the offset includes an offset component in a direction perpendicular to the viewing direction of the viewer's eye position.

[0048] In many embodiments, this can provide improved operation and / or improved performance.

[0049] The direction perpendicular to the viewing direction of the viewer's eye position may be a horizontal direction. In particular, it may be in a direction corresponding to the direction from one eye to the other eye of the viewer pose. The vertical component may be in a direction corresponding to the vertical direction in the scene coordinate system.

[0050] According to an optional feature of the present invention, the offset includes a vertical component.

[0051] In many embodiments, this can provide improved operation and / or improved performance. The vertical component may be in a direction perpendicular to the plane formed by the viewing direction of the viewer's eye position and the direction between the two eyes of the viewer pose. The vertical component may be in a direction corresponding to the vertical direction in the scene coordinate system.

[0052] According to one aspect of the present invention, there is provided an apparatus for generating an image signal, the apparatus comprising: a receiver for receiving a plurality of images representing a scene seen from one or more poses; a representation processor for generating image data providing an image representation of the scene, the image data including the plurality of images, and the image representation being provided with respect to a scene coordinate system including a scene reference position; an offset generator for generating an offset indication indicating an offset to be applied between the scene reference position and the viewer's eye position when the scene coordinate system is aligned with the viewer coordinate system, the offset including an offset component in a direction opposite to the viewing direction of the viewer's eye position; and an output processor for generating the image signal to include the image data and the offset indication.

[0053] According to one aspect of the present invention, there is provided a method of rendering an image, the method comprising: receiving an image representation of a scene, the image representation being provided with respect to a scene coordinate system including a reference position; determining a viewer pose for a viewer, the viewer pose being provided with respect to a viewer coordinate system; aligning the scene coordinate system with the viewer coordinate system by aligning the scene reference position with a viewer reference position in the viewer coordinate system; rendering view images for different viewer poses in response to the image representation and the alignment of the scene coordinate system with the viewer coordinate system; the method further comprising: determining the viewer reference position in response to a first viewer pose, the viewer reference position depending on the orientation of the first viewer pose and having an offset with respect to the viewer's eye position for the first viewer pose, the offset including an offset component in a direction opposite to the viewing direction of the viewer's eye position; wherein receiving the image representation of the scene includes receiving an image data signal including the image representation and further including an offset indication; and further including determining the offset in response to the offset indication.

[0054] According to one aspect of the present invention, there is provided a method for generating an image signal, the method comprising: receiving a plurality of images representing a scene seen from one or more poses; generating image data providing an image representation of the scene, the image data including the plurality of images, and the image representation being provided with respect to a scene coordinate system including a scene reference position; generating an offset indication indicating an offset to be applied between the scene reference position and the viewer's eye position when the scene coordinate system is aligned with the viewer coordinate system, the offset including an offset component in a direction opposite to the viewing direction of the viewer's eye position; and generating the image signal to include the image data and the offset indication.

[0055] These and other aspects, features, and advantages of the present invention will be apparent and elucidated with reference to the (one or more) embodiments described below. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Embodiments of the present invention will be described only by way of example and with reference to the accompanying drawings, in which:

[0057] Figure 1 An example of an arrangement for providing a virtual reality experience is illustrated;

[0058] Figure 2 An example of elements of a device according to some embodiments of the present invention is illustrated;

[0059] Figure 3 An example of elements of a device according to some embodiments of the present invention is illustrated; and

[0060] Figure 4 An example of a configuration for an image representation of a scene is illustrated;

[0061] Figure 5 An example of an omnidirectional stereoscopic image representation of a scene is illustrated;

[0062] Figure 6 An example of an omnidirectional stereoscopic image representation of a scene is illustrated;

[0063] Figure 7 An example of an omnidirectional stereoscopic image with a depth map is illustrated;

[0064] Figure 8 An example of determining a viewer reference position relative to a viewer pose is illustrated;

[0065] Figure 9 An example of determining a viewer reference position relative to a viewer pose is illustrated;

[0066] Figure 10 An example of determining a viewer reference position relative to a viewer pose is illustrated;

[0067] Figure 11 An example of determining a viewer reference position relative to a viewer pose is illustrated;

[0068] Figure 12 An example of determining an offset for a viewer reference position relative to a viewer pose is illustrated;

[0069] Figure 13 An example of determining an offset for a viewer reference position relative to a viewer pose is illustrated;

[0070] Figure 14Illustrates an example of determining an offset of a viewer reference position relative to a viewer pose; and

[0071] Figure 15 Illustrates an example of determining an offset of a viewer reference position relative to a viewer pose. Detailed Description

[0072] Virtual experiences that allow users to move around in a virtual world are becoming increasingly popular, and services are being developed to meet such needs. However, providing efficient virtual reality services is very challenging, especially if the experience is based on the capture of a real-world environment rather than on a fully virtual, generated artificial world.

[0073] In many virtual reality applications, the viewer pose input is determined to reflect the pose of the virtual viewer in the scene. The virtual reality device / system / application then generates one or more images corresponding to the view and viewport of the scene for the viewer corresponding to the viewer pose.

[0074] Typically, virtual reality applications generate three-dimensional outputs in the form of separate view images for the left and right eyes. These outputs can then be presented to the user in a suitable manner (e.g., separate left-eye and right-eye displays of a VR headset, typically). In other embodiments, the images can be presented, for example, on an autostereoscopic display (in which case a large number of view images can be generated for the viewer pose), or in some embodiments, actually only a single two-dimensional image can be generated (e.g., using a conventional two-dimensional display).

[0075] The viewer pose input can be determined in different ways in different applications. In many embodiments, the user's body movements can be directly tracked. For example, cameras in the surveyed area can detect and track the user's head (and even eyes). In many embodiments, the user can wear a VR headset that can be tracked by external means and / or internal means. For example, the headset can include accelerometers and gyroscopes that provide information about the movement and rotation of the headset and thus the head. In some examples, the VR headset can transmit signals or include (e.g., visual) identifiers that enable external sensors to determine the movement of the VR headset.

[0076] In some systems, the viewer pose can be provided manually, for example, by the user manually controlling a joystick or similar manual input. For example, the user can manually move the virtual viewer around in the scene by controlling a first analog joystick with one hand, and manually control the direction the virtual viewer is looking by manually moving a second analog joystick with the other hand.

[0077] In some applications, a combination of manual and automatic methods can be used to generate the input viewer pose. For example, a head-mounted device can track the orientation of the head and the viewer's movement / position in the scene can be controlled by the user using a joystick.

[0078] The generation of the image is based on a suitable representation of the virtual world / environment / scene. In some applications, a complete 3D model of the scene can be provided and the view of the scene from a particular viewer pose can be determined by evaluating the model.

[0079] In many practical systems, the scene can be represented by an image representation including image data. The image data can typically include images associated with one or more capture or anchor poses and, in particular, can include images for one or more viewports, where each viewport corresponds to a particular pose. An image representation including one or more images can be used, where each image represents the view of a given viewport for a given viewing pose. Such viewing poses or positions for which image data is provided are often referred to as anchor poses or positions or capture poses or positions (since the image data can typically correspond to images that would be captured by a camera positioned in the scene with a position and orientation corresponding to the capture pose).

[0080] Many typical VR applications can then continue to provide a view image corresponding to the viewport of the scene for the current viewer pose, where the image is dynamically updated to reflect changes in the viewer pose and the image is generated based on the image data representing the (possibly) virtual scene / environment / world. The application can do this by performing view synthesis and view shifting algorithms known to those skilled in the art.

[0081] In this field, the terms placement and pose are used as general terms for position and / or orientation. For example, the combination of the position and orientation of an object, camera, head, or view can be referred to as a pose or placement. Thus, a placement or pose indication can include six values / components / degrees of freedom, where each value / component typically describes an independent property of the position / location or orientation / direction of the corresponding object. Of course, in many cases, a placement or pose may be considered or represented with fewer components, for example, when one or more components are considered fixed or irrelevant (e.g., when all objects are considered to be at the same height and have a horizontal orientation, four components can provide a complete representation of the object pose). Hereinafter, the term pose is used to refer to a position and / or orientation that can be represented by one to six values (corresponding to the maximum possible degrees of freedom).

[0082] Many VR applications are based on poses with the maximum degrees of freedom, i.e., three degrees of freedom for each of position and orientation, resulting in a total of six degrees of freedom. A pose can thus be represented by a set or vector of six values representing the six degrees of freedom, and thus the pose vector can provide three-dimensional position and / or three-dimensional orientation indication. However, it should be understood that in other embodiments, a pose can be represented by fewer values.

[0083] A pose can be at least one of orientation and position. Pose values can indicate at least one of orientation values and position values.

[0084] A system or entity that provides the maximum degrees of freedom for a viewer is typically referred to as having six degrees of freedom (6DoF). Many systems and entities provide only orientation or position, and these systems and entities are typically referred to as having three degrees of freedom (3DoF).

[0085] In some systems, a VR application can be provided locally to a viewer by, for example, a stand-alone device that does not use (or even have access to) any remote VR data or processing. For example, a device such as a game console can include a storage device for storing scene data, an input unit for receiving / generating a viewer's pose, and a processor for generating corresponding images based on the scene data.

[0086] In other systems, a VR application can be implemented and executed remotely from the viewer. For example, a device local to the user can detect / receive movement / pose data, which is transmitted to a remote device that processes the data to generate a viewer's pose. The remote device can then generate a suitable view image for the viewer's pose based on scene data describing the scene. The view image is then transmitted to a device local to the viewer, where the view image is presented. For example, the remote device can directly generate a video stream (usually a stereoscopic / 3D video stream) that is directly presented by the local device. Thus, in such an example, the local device may not perform any VR processing other than transmitting movement data and presenting the received video data.

[0087] In many systems, functionality can be distributed across local and remote devices. For example, a local device can process received input and sensor data to generate viewer poses that are continuously transmitted to a remote VR device. The remote VR device can then generate corresponding view images and transmit these view images to the local device for presentation. In other systems, the remote VR device may not directly generate view images, but instead can select relevant scene data and transmit it to the local device, which can then generate the presented view images. For example, the remote VR device can identify the closest capture point and extract the corresponding scene data (e.g., spherical images and depth data from that capture point) and transmit it to the local device. The local device can then process the received scene data to generate an image for a particular current viewing pose. The viewing pose will typically correspond to a head pose, and a reference to the viewing pose can generally be equivalently considered to correspond to a reference to the head pose.

[0088] In many applications, especially for broadcast services, the source can transmit scene data in the form of an image (including video) representation of a scene that is independent of the viewer pose. For example, an image representation of a single view sphere for a single capture location can be transmitted to multiple clients. The individual clients can then locally synthesize view images corresponding to the current viewer pose.

[0089] An application that has drawn particular attention is one that supports a limited amount of movement such that the presented view is updated to follow small movements and rotations corresponding to a substantially static viewer who only makes small head movements and head rotations. For example, a seated viewer can turn his head and move his head slightly, and the presented view / image is adjusted to follow these pose changes. Such an approach can provide a highly immersive video experience, for example. A viewer watching a sports event may feel as if he is present at a particular location in the arena.

[0090] Such limited degrees of freedom applications have the advantage that they are able to provide an improved experience while not requiring an accurate representation of the scene seen from many different positions, thus significantly reducing the capture requirements. Similarly, the amount of data that needs to be provided to the renderer can be greatly reduced. In fact, in many scenarios, it is only necessary to provide an image for a single viewpoint and typically also depth data to the local renderer from which the desired view can be generated.

[0091] The method may be particularly suitable for applications that need to transmit data from a source to a destination over a bandwidth - limited communication channel, such as broadcast or client - server applications.

[0092] Figure 1Illustrated is an example of a VR system in which a remote VR client device 101 communicates with a VR server 103 via a network 105 (e.g., the Internet), for example. The server 103 may be arranged to support potentially a large number of client devices 101 simultaneously.

[0093] The VR server 103 can support a broadcast experience, for example, by transmitting image data and depth for a particular viewpoint, and then the client device is arranged to process this information to locally synthesize a view image corresponding to the current pose.

[0094] Figure 2 Illustrated are example elements of an exemplary embodiment of the VR server 103.

[0095] The apparatus includes a receiver 201 arranged to receive a plurality of images representing a scene seen from one or more poses.

[0096] In this example, the receiver 201 is coupled to a source 203 that provides the plurality of images. In particular, the source 203 may be a local memory storing the images, or the source 203 may be, for example, a suitable capture unit, such as a set of cameras.

[0097] The receiver 201 is coupled to a processor referred to as a representation processor 205, which is fed the plurality of images. The representation processor 205 is arranged to generate an image representation of the scene, the image representation including image data derived from the plurality of images.

[0098] The representation processor 205 is coupled to an output processor 207 that generates an image signal including the image representation, so the image signal specifically includes the image data of the plurality of images. In many embodiments, the output processor 207 may be arranged to encode the images and include them in a suitable data stream (e.g., a data stream generated according to a suitable standard).

[0099] The output processor 207 may also be arranged to send or broadcast the image signal to a remote client / device, and in particular, may transmit the image signal to the client device 101.

[0100] Figure 3 Illustrated are some elements of an apparatus for rendering an image according to some embodiments of the present invention. The apparatus will be described in the context of the Figure 1 system, where the apparatus is specifically the client device 101.

[0101] The client device 101 includes a data receiver 301 which is arranged to receive an image signal from the server 103. It should be understood that any suitable method and format for communication can be used without departing from the present invention.

[0102] The data receiver 301 is coupled to a renderer 303 which is arranged to generate view images for different viewports / viewer poses.

[0103] The client device 101 further includes a viewing pose determiner 305 which is arranged to dynamically determine the current viewer pose. In particular, the viewing pose determiner 305 can receive data from a head-mounted device that reflects the movement of the head-mounted device. The viewing pose determiner 305 can be arranged to determine the viewing pose based on the received data. In some embodiments, the viewing pose determiner 305 can receive, for example, sensor information (such as accelerometer and gyroscope data) and thereby determine the viewing pose. In other embodiments, the head-mounted device can directly provide the viewing pose data.

[0104] The viewing pose is fed to the renderer 303 which continues to generate view images corresponding to the views of the scene seen by the two eyes of a viewer in the current viewer pose. The view images are generated according to any suitable image generation and synthesis algorithm based on the received image representation. The specific algorithm will depend on the specific image representation as well as the preferences and requirements of the individual embodiments.

[0105] As an example, the image representation can include one or more images, each of which corresponds to a scene viewed from a given viewpoint and a given direction. Thus, for each of a plurality of capture or anchor poses, an image corresponding to the viewport of the scene seen from that viewpoint and that direction is provided. In many embodiments, a plurality of cameras can be positioned, for example, in a line and capture the scene from different positions along this line and all aimed in the same direction, such as in the Figure 4 example. It should be understood that in other embodiments, other capture configurations can be employed, including more or fewer anchor points and / or images.

[0106] In many embodiments, the image representation can be according to a particular existing 3D image format known as omnidirectional stereo (ODS). For ODS, rays are created for the left-eye image and the right-eye image such that the origin of these rays lies on a circle whose diameter is typically equal to the interpupillary distance (e.g., ~6.3 cm). For ODS, narrow-angle image portions are captured for opposite directions corresponding to the tangents of the view circle and at regular angular distances around the view circle (see Figure 5 ).

[0107] Thus, for an ODS, an image is generated for the left eye, where each pixel column corresponds to a position on the unit circle and reflects the light rays at that position in the tangential direction to the ODS visual circle. The positions on the ODS visual circle are different for each column and typically a relatively large number of equally spaced positions are defined on the ODS visual circle to cover the entire 360° field of view, where each column corresponds to one position. Thus, a single ODS image captures the complete 360° field of view, where each column corresponds to a different position on the ODS visual circle and corresponds to a different light direction.

[0108] The ODS includes a right-eye image and a left-eye image. As Figure 6 shown, for a given column in these images, the left-eye image and the right-eye image will reflect the light rays at the relative positions on the ODS visual circle. Thus, the ODS image format provides a 360° view and stereoscopic information based on only two images.

[0109] It should be understood that although the following description considers an ODS that provides two stereoscopic images (corresponding to the left-eye image and the right-eye image), a corresponding format that provides only one image can also be used, i.e., an omnidirectional single-image format can be used. This image can still be provided relative to the visual circle to provide a view from one eye, as it rotates around, for example, the center point between the two eyes.

[0110] Alternatively, by making the radius of the visual circle approach zero, a single-image format is achieved. In this case, the left view is the same as the right view, so only one image is required to represent both.

[0111] Generally, omnidirectional video / images can be provided as, for example, omnidirectional stereoscopic video / images or as omnidirectional single video / images.

[0112] For a given orientation (viewpoint), an image can be generated by combining narrow-angle image portions for viewing directions that match within the viewport of the given orientation. Thus, a given view image is formed by combining narrow-angle image portions corresponding to captures in different directions (but different narrow-angle image portions are from different positions on the circle). Thus, the view image includes captures from different positions on the visual circle rather than only from a single viewpoint. However, if the visual circle represented by the ODS is small enough (relative to the content of the scene), its influence can be reduced to an acceptable level. Additionally, since the captures along a given direction can be reused for multiple different viewing orientations, the amount of image data required is significantly reduced. The view images for the viewer's two eyes are typically generated by captures in opposite directions for the appropriate tangents.

[0113] Figure 6Illustrated is an example of an ideal head rotation that can be supported by ODS. In this example, the head rotates such that both eyes move along a circle with a diameter equal to the interpupillary distance. Assuming this corresponds to the width of the ODS visual circle, view images for different orientations can be simply determined by selecting appropriate narrow-angle image portions corresponding to different viewing point orientations.

[0114] However, for a standard ODS, the viewer will perceive stereoscopic vision rather than motion parallax. Even with a small movement of the viewer (about a few centimeters), the absence of motion parallax tends to provide an unpleasant experience. For example, if the viewer moves such that the eyes no longer exactly fall on the ODS visual circle, then generating a view image based simply on selecting and combining appropriate narrow-angle image portions will result in the generated view image being the same as the case where the user's eyes remain on the visual circle, and thus will not exhibit the parallax that should occur when the user moves his head, which will result in the perception of being unable to move relative to the real world.

[0115] To solve this problem and allow the generation of motion parallax based on ODS data, the ODS format can be extended to include depth information. A narrow-angle depth map portion can be added for each narrow-angle image portion. Figure 7 Illustrated is an example of an ODS image with an associated depth map. This depth information can be used to perform a viewpoint shift such that the generated image corresponds to a new position outside (or inside) the visual circle (e.g., each view image or narrow-angle image portion can be processed using known images and a depth-based viewpoint shift algorithm). For example, a 3D mesh can be created for each eye, and motion parallax can be introduced using ODS data drawing based on the meshes and textures for the left and right eyes.

[0116] However, whether the image representation is based on, for example, multiple images for different capture poses or ODS data, generating view images for poses different from the anchor pose that provides the image data often introduces artifacts and errors, resulting in potential degradation of image quality.

[0117] The renderer 303 is arranged to generate a view image for the current viewing pose based on the received image representation. In particular, a right-eye image and a left-eye image can be generated for a stereoscopic display (e.g., a head-mounted device), or multiple view images can be generated for the views of an autostereoscopic display. It should be understood that many different algorithms and techniques are known for generating view images from the provided scene images, and any suitable algorithm can be used according to a particular embodiment.

[0118] However, the key operation for rendering is that the determined viewer pose must be aligned with the image representation with each other.

[0119] An image representation is provided with reference to a given coordinate system for a scene. The image representation includes image data associated with specific anchor poses, and these anchor poses are provided relative to the scene coordinate system.

[0120] Similarly, a viewer pose is provided with reference to a viewer coordinate system. For example, the viewer pose data provided by the viewing pose determiner 305 can indicate changes in the position and rotation of the viewer's head, and this data is provided relative to the coordinate system.

[0121] Thus, inherently, since the image representation and the viewer pose are provided independently and generated separately, the image representation and the viewer pose will be provided relative to two different coordinate systems. In order to draw an image for the viewer pose based on the image representation, it is therefore necessary to link / align these coordinate systems with each other. In many cases, the coordinates used can have the same scale and in particular can be provided relative to a coordinate system having a scale that matches the real-world scale. For example, two anchor positions can be defined or described as being 1 meter apart, thereby reflecting that they have been captured by cameras that are 1 meter apart in the real world, or that they provide views that should be interpreted as being 1 meter apart. Similarly, the viewer pose data can indicate how the user moves his head in the real world, for example, how many centimeters the user has moved his head.

[0122] However, even when it is known (or assumed) that the scales of the coordinate systems are the same, it is necessary to align the relative positions of the coordinate systems with each other. Conventionally, this is typically done at the start of the application by aligning a reference pose (in particular, the position) in the scene coordinate system with the current viewer pose. Thus, when the application starts, the viewer is effectively positioned at a nominal or default start position in the scene. Subsequently, this alignment is used such that the viewer position in the scene coordinate system is determined by tracking the relative changes indicated by the viewer pose in the scene coordinate system (which the renderer can use when drawing the view). Thus, an indication that the viewer pose changes 2 centimeters to the left in the viewer coordinate system will correspond to a shift 2 centimeters to the left in the scene coordinate system.

[0123] The initial alignment of the two coordinate systems is thus fixed in the conventional case and independent of the actual viewer pose. It is typically initialized by aligning the viewer at a position in the scene coordinate system corresponding to the maximum quality (for example, at the default position where the image representation includes the image data). The user starts experiencing at a predetermined start position in the scene and then tracks any changes in the viewer pose.

[0124] However, the inventors have recognized that while such methods may be suitable for many applications, they are not suitable for all applications. The inventors have also recognized that in many embodiments, a more adaptive method in which alignment depends on the viewer pose and in particular on the viewer orientation can provide improved operation. This can be particularly advantageous for many restricted movement services in which the user is restricted (e.g., restricted to small head movements and head rotations). In such services, performing realignment at certain times (e.g., when the viewer pose indicates that the movement since the last alignment meets the criteria) can be particularly attractive, and the adaptive methods described below can be particularly advantageous in such scenarios.

[0125] Accordingly, Figure 3 the apparatus of includes an aligner 307 arranged to align the scene coordinate system with the viewer coordinate system. The alignment can in particular align the coordinate systems by aligning / linking a scene reference position in the scene coordinate system with a viewer reference position in the viewer coordinate system. Accordingly, the alignment can cause the scene reference position in the scene coordinate system to be juxtaposed with the viewer reference position in the viewer coordinate system, i.e., the scene reference position and the viewer reference position are the same position.

[0126] The reference position in the scene coordinate system can be any suitable position, and in particular a fixed and invariant scene reference position. In particular, the scene reference position can be a position independent of the viewer pose and can generally be a predetermined reference position. In embodiments where alignment is performed repeatedly, the scene reference position can be constant between (at least two) successive alignment operations.

[0127] In many embodiments, the scene reference position can be defined relative to an anchor position. For example, the reference position can be the anchor position or an average position for the anchor position, for example. In some embodiments, the scene reference position can correspond to a position where the image representation includes image data and / or a position where optimal image quality can be achieved.

[0128] In contrast to the scene reference position that is independent of the viewer pose, the viewer reference position depends on a first viewer pose that is the viewer pose for which the alignment is performed. This first viewer pose will hereinafter be referred to as the alignment viewer pose. The alignment viewer pose can be any viewer pose for which alignment is desired to be performed. In many cases, the alignment viewer pose can be the current viewer pose at the time when the alignment is performed. The alignment viewer pose can generally be the current viewer pose and can in particular be the pose indicated by the viewer pose data when the (re)alignment is initialized. Accordingly, when alignment occurs, the viewer reference position is dynamically determined and the viewer reference position depends on the viewer pose and in particular on the orientation of the viewer pose.

[0129] Thus, Figure 3 the apparatus of Figure 3 includes an offset processor 309 that is coupled to an aligner 307 and a viewing pose determiner 305. The offset processor 309 is arranged to determine a viewer reference position based on an aligned viewer pose and in particular on the current viewer pose at the time when the alignment is performed. The aligned viewer pose is the viewer pose for which the alignment between two coordinate systems is performed and will be considered the current viewer pose (at the time when the alignment is performed) in the following description.

[0130] The offset processor 309 is in particular arranged to determine the viewer reference position such that it depends on the orientation of the aligned / current viewer pose. The offset processor 309 is arranged to generate the viewer reference position such that it is offset with respect to the viewer eye position for the current viewer pose. Additionally, the offset has a component in a direction opposite to the viewing direction of the viewer eye position.

[0131] Thus, for a given viewer pose, the viewer reference position is determined such that it does not coincide with any eye position for that viewer pose and it is also not located to the side of the eye position (or on the line between the two eyes). Instead, the viewer reference position is located “behind” the eye position for the current viewer pose. Thus, in contrast to conventional systems where the two coordinate systems are aligned such that for example the viewer pose position always coincides with the reference position, Figure 3 the apparatus of

[0130] determines an offset viewer reference position relative to the viewer pose (and its (one or more) eye positions). The offset is in a direction behind the viewing direction for the current viewer position, thereby causing an offset between the two coordinate systems. Additionally, since the offset has a rear component, the offset depends on the orientation of the viewer pose and different viewer poses will cause different viewer reference positions, thereby causing different alignment offsets.

[0132] In particular, for two viewer poses representing the same viewer position but different viewer orientations, the determined viewer reference positions will be different. Equivalently, if two viewer poses representing different orientations represent different positions, the same viewer reference position can only be obtained for two viewer poses representing different orientations.

[0133] As an example, Figure 8 illustrates a simple example where the current viewer pose indicates an eye position 801 having a view / viewport 803 in a given direction 805. In this case, the viewer reference position 807 is determined to have a rear offset 809 from the eye position 801. Figure 8Two examples with different orientations are illustrated. It can be clearly seen that the viewer reference position depends on the rotation (if the eye positions 801 in the two examples are considered to be in the same position, the position determined as the viewer reference position 807 will be significantly different).

[0134] Figure 9 A corresponding example is illustrated in which the viewer pose 901 indicates the midpoint between the user's two eyes. An offset 903 in the rearward direction provides the viewer reference position 905. Figure 10 Two overlapping examples are shown, in which the viewer poses represent the same position but different orientations. It can be clearly seen that the corresponding viewer reference positions are different.

[0135] Therefore, when the viewer reference position is aligned with the scene reference position, different orientations will cause different alignments, especially between the anchor pose and the viewer pose. This is shown as an ODS circle image representation in Figure 11 and as an example of an image representation obtained from images of three capture poses in Figure 12 In the latter example, the orientation of the capture configuration has not changed, but it should be understood that in some embodiments, this can be rotated, for example, to correspond to the orientation differences between the viewer poses.

[0136] The inventors have recognized that in many practical applications, such alignment of orientation-related offsets can provide a significant quality improvement. In particular, it has been found that a significant quality improvement is provided in applications where the viewer is limited to relatively small head movements and rotations. In such cases, the adaptive change of the offset can effectively shift the image data away from the nominal and default positions conventionally used. This may require performing some additional view shift, which may reduce the image quality for that particular viewing pose. However, the inventors have recognized that it can generally improve the image quality for other nearby poses, especially for rotational movements. In fact, by applying an appropriate offset, the view shift required for some other poses can be reduced, thus allowing the quality of these poses to be improved. The inventors have also recognized that an adaptive rearward offset can generally provide a quality improvement for non-aligned viewer poses (i.e., poses that do not correspond to the nominal default pose), and the quality degradation for non-aligned viewer poses significantly exceeds any quality degradation for aligned viewer poses. The inventors have particularly recognized that an adaptive rearward offset can take advantage of the following: the effect on the quality of view shift in the forward (and backward) direction is significantly lower than the effect on the quality of side view shift (e.g., due to the de-occlusion effect).

[0137] The method can thus provide an improved trade-off by sacrificing some quality degradation for the nominal viewer pose in exchange for a significant quality improvement over a range of other viewer poses. In particular, for applications with limited head movement, this can typically provide a significantly improved user experience and can also reduce the requirements for realignment, since a given required image quality can be supported over a wider range of poses.

[0138] The exact size of the offset can vary in different embodiments and can be determined dynamically in many embodiments. However, in many embodiments, the rearward offset component and thus the offset in the opposite direction of the viewing direction is not less than 1 cm, 2 cm, or 4 cm in the viewer coordinate system. The viewer coordinate system is typically a real-world scale coordinate system that reflects the actual changes in the viewer's pose and, in particular, the position in the real world. The rearward offset component can thus correspond to at least 1 cm, 2 cm, or 4 cm in the real-world scale for the viewer's pose.

[0139] Such a minimum offset value can typically ensure that the offset has a significant impact and can achieve a significant quality improvement for many actual viewer poses and pose changes.

[0140] In many embodiments, in the viewer coordinate system, the rearward offset component and thus the offset in the opposite direction of the viewing direction does not exceed 8 cm, 10 cm, or 12 cm. The rearward offset component can correspond to not more than 8 cm, 10 cm, or 12 cm in the real-world scale for the viewer's pose.

[0141] Such a minimum offset value can typically ensure that the offset is low enough to ensure that the offset does not unreasonably affect the quality of the generated image for the aligned viewer pose.

[0142] In particular, selecting an offset with such parameters can result in a viewer reference position that is typically located within the viewer's head.

[0143] The offset value can be considered to balance the forward viewing quality and the side viewing quality. Due to the smaller viewpoint shift, a smaller offset value results in a smaller forward viewing error. However, due to de-occlusion, head rotation can cause a rapid quality degradation. The impact of the degradation and the impact of the offset on side viewing in practice typically tend to be much greater than the impact on forward viewing. Therefore, the benefit of using the offset tends to provide a significantly improved benefit for head rotation while only causing a slight degradation to forward viewing.

[0144] Therefore, when it is expected that the viewer will mainly look forward, a smaller offset value (1 - 4 cm) is useful. Larger offset values (8 - 12 cm) up to the offset corresponding to the axis of rotation of the head (projecting upwards from the neck) result in a quality that is less dependent on head rotation. For this reason, when it is expected that the viewer will look around more, a larger offset value is appropriate. The medium offset value corresponds to the optimal value we expect for typical viewer behavior, where there is a bias towards looking forward, thus favoring providing more quality for forward viewing, and this method is applied to reduce the degradation of side viewing.

[0145] In many embodiments, the backward offset component and thus the offset in the opposite direction of the viewing direction is not less than 1 / 10, 1 / 5, 1 / 3, or 1 / 2 of the (nominal / default / assumed) inter - pupillary distance. In many embodiments, the backward offset component and thus the offset in the opposite direction of the viewing direction does not exceed 1 times, 1.5 times, or 2 times the (nominal / default / assumed) inter - pupillary distance.

[0146] In some embodiments, the normalization scale for the offset can be considered to be in the range of [0, 1], where it is 0 at the camera (e.g., no position change due to rotation) and 1 at the neck / head rotation position. A normalized distance of 1 is typically considered to correspond to, for example, 8, 10, or 12 cm.

[0147] In different embodiments, different methods can be used to determine the offset. For example, in some embodiments, a predetermined and fixed offset can simply be applied in all cases. Thus, the offset can be an inherent part of the alignment process and can be hard - coded into the algorithm. However, in many embodiments, the offset can be determined dynamically during operation.

[0148] In some embodiments, the offset can be generated on the server side and can be provided to the rendering device together with the image representation, and in particular as part of the image signal.

[0149] In particular, Figure 2The VR server 103 can include an offset generator 209 arranged to determine an offset between a viewer reference position and a viewer eye position for a viewer pose for which alignment has been performed. When aligning the scene coordinate system with the viewer coordinate system, the offset generator 209 can thus determine the offset to be applied between the scene reference position and the viewer eye position. As previously described, the offset includes an offset component in a direction opposite to the viewing direction of the viewer eye position (i.e., in a direction opposite to the viewing direction for the aligned viewer pose). An offset indication can then be generated to at least partially represent the offset, and the offset indication can be included in the image signal transmitted to the client device 101. The client device 101 can then continue to determine the offset to be applied based on the offset indication, and then, the client device 101 can use the offset when aligning the scene coordinate system with the viewer coordinate system. In particular, the client device 101 can apply the offset to the viewer eye position to determine the viewer reference position. In many embodiments (e.g., for broadcast applications), such a method can be particularly advantageous. The method can particularly allow an optimization process to be performed to determine an optimal offset according to suitable criteria, and then the optimized offset can be transmitted to all client devices capable of using it when performing alignment.

[0150] Thus, in such an embodiment, the client device 101 can receive an image data signal that includes an offset indication in addition to the image representation. The client device 101 can then directly (or, e.g., after modification) use the offset.

[0151] Thus, the offset can be determined dynamically, e.g., based on an optimization process. Such a process can be performed at the source or can be performed, e.g., by a rendering device (i.e., by the client device 101 in the previous example).

[0152] The offset generator 209 can use any suitable method to generate an appropriate offset. For example, in some embodiments, user input can be used to select a preferred offset, e.g., by enabling the user to dynamically adjust the offset and then generate images for different viewer poses to select the preferred offset. In other embodiments, an automated process can be used to evaluate a range of candidate values for the offset, and the value that results in the highest quality within a suitable range of viewer poses can be used.

[0153] The server can in particular simulate operations performed on the rendering side to evaluate the impact of different offsets. For example, the server can evaluate different candidate values by simulating the processing that would be performed by the renderer, and then use a suitable quality metric to evaluate the resulting quality. In fact, when determining the offset to be indicated in the image data stream, the methods, algorithms, and considerations described with respect to the renderer / receiver / client and in particular with respect to the offset processor 309 can also be performed by the offset generator 209.

[0154] It should also be understood that any suitable means can be used to provide an indication of the offset. In many embodiments, the offset to be applied can be described (e.g., by providing a two-dimensional or three-dimensional vector indication) with a complete description. Such an offset vector can, for example, indicate the direction and magnitude of the offset to be applied. When determining the viewer reference position, the renderer can then position the vector at the aligned eye position and with a direction determined relative to the eye viewing direction for the aligned viewer pose.

[0155] In some embodiments, the offset indication can include only a partial representation of the offset. For example, in many embodiments, the offset indication can include only a distance / magnitude indication and / or a direction indication. In such cases, the renderer can, for example, determine suitable values for these items. Thus, in some embodiments, the determination of the offset can be based on decisions made on both the server side and the client side. For example, the direction can be given by the offset indication and determined by the server, while the magnitude / distance is determined by the renderer. The method for determining the offset by the offset processor 309 will be described below. It should be understood that the principles and methods described can equally be performed by the offset generator 209, and the resulting offset is included in the image data stream. Thus, in the following, references to the offset processor 309 can be considered to apply equally (mutatis mutandis) to the offset generator 209.

[0156] The method is based on the offset processor 309 determining an error metric (value / measure) for one or more viewer poses for different candidate values of the offset. Then it can select the candidate value that results in the lowest error metric and use that value for the offset when performing the alignment.

[0157] In many embodiments, the error metric for a given candidate offset can be determined in response to a combination of error metrics for multiple viewer poses (e.g., in particular for a range of viewer poses). For example, for different viewer poses, when the offset is set to the candidate offset, an error metric can be determined that indicates the degradation of the image quality when generating a view image for that viewing pose from the image representation. Then the resulting error metrics for different viewer poses can be averaged to provide a combined error metric for the candidate value. Then the combined error metrics for different candidate values can be compared.

[0158] In many embodiments, an error metric for a candidate value can be generated by combining error metrics from multiple different locations and / or for multiple different poses.

[0159] In some embodiments, an error metric for one viewer pose for a candidate value can be responsive to error metrics for multiple fixation directions and typically for a range of fixation directions. Thus, not only can different locations and / or orientations be considered, but also for a given pose, the quality at different fixation directions (i.e., in different directions of the viewport for a given viewing pose) can be considered.

[0160] As an example, the offset processor 309 can determine error metrics for rear offsets of 2 cm, 4 cm, 6 cm, 8 cm, 10 cm, and 12 cm. The error metric can be determined, for example, for each of the corresponding candidate offsets as an indication of the image quality of the rendered image for a set of viewer poses (e.g., viewer poses corresponding to the current user having turned his head exactly 15°, 30°, and 45° to the left or right). The candidate offset that results in the lowest error metric can then be selected.

[0161] In different embodiments, the exact error metric considered may vary. Generally, a suitable error metric indicates a degradation in the quality of view synthesis based on the image representation for that pose and candidate offset. Thus, the error metric for a viewer pose and candidate value includes an image quality metric of the view image for the viewer pose synthesized based on the image representation (at least one image thereof).

[0162] In embodiments where the offset is determined on the server side based on image generation according to a scene model, an image for a given candidate offset and viewer pose can be generated and compared to the corresponding image generated directly from the model.

[0163] In other embodiments, more indirect measures can be used. For example, the distance from the viewer pose for which the image data is synthesized to the capture pose (with a given candidate offset) for which the image data is provided can be determined. The greater this distance, the greater the error metric is considered to be.

[0164] In some embodiments, the determination of the error metric can include considering more than one capture pose. Thus, in particular, the ability to generate a synthesized image for a given viewer pose based on different anchor pose images can be considered in cases where the image representation includes more than one capture pose. For example, for an image representation such as Figure 12 the synthesis can be based on any one of the anchor images, and in fact, the synthesized image for a given pixel can be generated by considering two or more of the anchor images in the anchor images. For example, interpolation can be used.

[0165] Thus, in some embodiments, the error metric for the viewer pose and the candidate value includes an image quality metric of a view image for the viewer pose, the view image being synthesized from at least two images of an image representation, wherein the at least two images have reference positions relative to the viewer pose that depend on the candidate offset value.

[0166] Refer to Figures 13 - 15 a specific example of a method for selecting an offset will be described. In this example, the image representation is a single ODS stereo image.

[0167] Figure 13 The viewer's head is illustrated as an ellipse, with the positions / viewpoints of the two eyes facing forward. The nominal head pose corresponding to the current / aligned viewer pose is illustrated as coinciding with the y-axis. The figure also shows some possible head rotations (rotated 45° and 90° to the left respectively) (the corresponding right rotations are illustrated only by the eye positions).

[0168] Figure 13 It is also illustrated how alignment is conventionally performed such that the ODS visual circle is positioned to coincide with the viewer's eyes.

[0169] Figure 14 An example of how an error metric can be generated for a viewing pose corresponding to a head rotation α is illustrated. In this example, the resulting left-eye position xl, yl can be determined, and for this position, the gaze direction β can be determined. The distance D from the eye positions xj, xy to the tangent of the ODS visual circle at the position xl', yl' is determined, where the ODS visual circle includes the image data of the light rays in the gaze direction. This distance D indicates the image degradation that occurs for the pixels in the gaze direction β for the head rotation α when generating an image based on the ODS view image with the ODS visual circle positioned as shown. Thus, a combined error metric can be generated by averaging the corresponding error metrics over a certain range of gaze directions (e.g., -45° < β < +45°) and a certain range of head rotation angles (e.g., -90° < β < +90°) (in some embodiments, different positions of the head can also be considered). The resulting combined error metric can thus indicate the average magnitude of the "viewpoint shift distance" required when generating a view image for a head rotation based on the ODS stereo image positioned as shown.

[0170] In many embodiments, this method can be performed for both the left and right eyes and the metrics for the left and right eyes can be combined.

[0171] As Figure 15As shown, this can be done for different positions of the ODS viewing circle, and in particular for different offsets in the y - direction, and the offset / position that finds the minimum error metric can be selected as the offset for alignment.

[0172] In some embodiments, the offset is directly in the backward direction, thus directly in the opposite direction of the viewing direction for the current / aligned viewer pose. Relative to the viewer coordinate system (and thus relative to the viewer), the backward direction is in the direction of the intersection (or parallel to the intersection) of the sagittal or median plane of the viewer with the transverse or axial plane (see, for example,

[0173] https: / / en.wikipedia.org / wiki / Anatomical_plane). This direction (e.g., in ultrasound) is also referred to as the axial (forward, into the body) direction.

[0174] However, in some embodiments and scenarios, the offset can also include a lateral component, i.e., the offset can include an offset component in a direction perpendicular to the viewing direction for the aligned viewer pose. This lateral component can still be in the horizontal direction, and in particular can be in a direction corresponding to the direction between the viewer's two eyes.

[0175] The lateral component can be in the direction of the intersection (or parallel to the intersection) of the coronal or frontal plane of the viewer with the transverse or axial plane relative to the viewer coordinate system (and thus relative to the viewer). This direction (e.g., in ultrasound) is also referred to as the lateral direction.

[0176] The lateral offset component can provide a more flexible method, which can provide improved quality within a range of viewer poses in many embodiments. For example, if the current viewer pose for which alignment is known or assumed is asymmetric with respect to a predicted or assumed further movement of the viewer, improved average quality for future viewer poses can be achieved by laterally offsetting the viewer reference position towards a position that is more likely to be the viewer's future position.

[0177] In some embodiments, the offset can include a vertical component.

[0178] The vertical component can be in the direction of the intersection (or parallel to the intersection) of the coronal or frontal plane of the viewer with the sagittal or median plane relative to the viewer coordinate system (and thus relative to the viewer). This direction (e.g., in ultrasound) is also referred to as the transverse direction.

[0179] For example, typical scenarios have more cases below the line of sight and then above the line of sight. When the viewer is more likely to look down and then up, when the viewer looks down, it is beneficial to have a small upward offset to reduce the amount of occlusion removal.

[0180] As another example, it can be considered that a viewer can move his head by nodding up and down. This will cause the position and orientation of the eyes to shift, thereby affecting the viewpoint shift. The previously described method can be enhanced to further consider determining the vertical offset. For example, when determining the error metric for a given candidate offset value, this can include evaluating the viewer pose corresponding to different up and down head rotations. Similarly, candidate offset values having a vertical component can be evaluated.

[0181] Such a method can provide increased flexibility for user movement and can provide improved overall quality for an increased range of movement.

[0182] It should be understood that, for clarity, the above description has described embodiments of the present invention with reference to different functional circuits, units, and processors. However, it is apparent that any suitable functional distribution between different functional circuits, units, or processors can be used without departing from the present invention. For example, functions illustrated as being performed by separate processors or controllers can be performed by the same processor or controller. Thus, the reference to a particular functional unit or circuit is only considered as a reference to a suitable module for providing the described function, rather than indicating a strict logical or physical structure or organization.

[0183] The present invention can be implemented in any suitable form, including hardware, software, firmware, or any combination of these items. The present invention can optionally be at least partially implemented as computer software running on one or more data processors and / or digital signal processors. The elements and components of embodiments of the present invention can be physically, functionally, and logically implemented in any suitable manner. In fact, the function can be implemented in a single unit, in multiple units, or as part of other functional units. For this reason, the present invention can be implemented in a single unit or can be physically and functionally distributed between different units, circuits, and processors.

[0184] Although the present invention has been described in connection with some embodiments, this is not intended to limit the present invention to the specific forms set forth herein. Rather, the scope of the present invention is limited only by the claims. Additionally, although features may seem to be described in connection with specific embodiments, those skilled in the art will recognize that the various features of the described embodiments can be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.

[0185] In addition, although listed separately, multiple modules, elements, circuits or method steps can be implemented by, for example, a single circuit, unit or processor. In addition, although individual features may be included in different claims, these features may be advantageously combined, and the inclusion of these features in different claims does not mean that the combination of these features is not feasible and / or not advantageous. Moreover, including features in a claim category does not mean a restriction on this category, but rather indicates that the feature is equally applicable to other claim categories when appropriate. In addition, the order of features in the claims does not mean that these features must work in any particular order, and in particular, the order of the individual steps in the method claims does not mean that these steps must be performed in this order. Instead, these steps can be performed in any suitable order. In addition, singular references do not exclude plural numbers. Therefore, references to "one", "one", "first", "second", etc. do not exclude multiple. The figure marks in the claims are only provided for clear illustration of examples and should not be interpreted as limiting the scope of the claims in any way.

[0186] According to some embodiments, it may be provided that:

[0187] A device for drawing an image, the device comprising:

[0188] A receiver (301) for receiving an image representation of a scene, the image representation being provided relative to a scene coordinate system, the scene coordinate system comprising a reference position;

[0189] a determiner (305) for determining a viewer pose for the viewer, the viewer pose being provided relative to a viewer coordinate system;

[0190] an aligner (307) for aligning the scene coordinate system with the viewer coordinate system by aligning the scene reference position with a viewer reference position in the viewer coordinate system;

[0191] a renderer (303) for rendering view images for different viewer postures in response to the image representation and the alignment of the scene coordinate system with the viewer coordinate system;

[0192] The device also includes:

[0193] An offset processor (309) is arranged to determine the viewer reference position in response to an alignment viewer gesture, the viewer reference position being dependent on the orientation of the alignment viewer gesture and having an offset relative to a viewer eye position for the alignment viewer gesture, the offset comprising an offset component in a direction opposite to a viewing direction of the viewer eye position.

[0194] An apparatus for generating an image signal, the apparatus comprising:

[0195] A receiver (201) for receiving a plurality of images representing a scene seen from one or more poses;

[0196] A representation processor (205) for generating image data providing an image representation of the scene, the image data including the plurality of images, and the image representation being provided with respect to a scene coordinate system including a scene reference position;

[0197] An offset generator (209) for generating an offset indication indicating an offset between the scene reference position and a viewer reference position in a viewer coordinate system;

[0198] An output processor (207) for generating the image signal to include the image data and the offset indication.

[0199] A method of rendering an image, the method comprising:

[0200] Receiving an image representation of a scene, the image representation being provided with respect to a scene coordinate system including a reference position;

[0201] Determining a viewer pose for a viewer, the viewer pose being provided with respect to a viewer coordinate system;

[0202] Aligning the scene coordinate system with the viewer coordinate system by aligning the scene reference position with the viewer reference position in the viewer coordinate system;

[0203] Rendering view images for different viewer poses in response to the image representation and the alignment of the scene coordinate system with the viewer coordinate system;

[0204] The method further comprises:

[0205] Determining the viewer reference position in response to an aligned viewer pose, the viewer reference position depending on the orientation of the aligned viewer pose and having an offset with respect to the viewer eye position for the aligned viewer pose, the offset including an offset component in a direction opposite to the viewing direction of the viewer eye position.

[0206] A method for generating an image signal, the method comprising:

[0207] Receiving a plurality of images representing a scene seen from one or more poses;

[0208] Generate image data that provides an image representation of the scene, the image data including the plurality of images, and the image representation being provided relative to a scene coordinate system that includes a scene reference position;

[0209] Generate an offset indication that indicates an offset between the scene reference position and a viewer reference position in a viewer coordinate system;

[0210] Generate the image signal to include the image data and the offset indication.

[0211] The above-described apparatus and method may be combined separately with each of the features of the dependent claims or in any combination.

Claims

1. An apparatus for rendering an image, the apparatus comprising: a receiver for receiving an image representation of a scene, the image representation being provided with respect to a scene coordinate system that includes a scene reference position; a determiner (305) for determining a viewer pose for a viewer, the viewer pose being provided with respect to a viewer coordinate system; an aligner (307) for aligning the scene coordinate system with the viewer coordinate system by aligning the scene reference position with a viewer reference position in the viewer coordinate system; a renderer (303) for rendering view images for different viewer poses in response to the image representation and the alignment of the scene coordinate system with the viewer coordinate system; and an offset processor (309) arranged to determine the viewer reference position in response to a first viewer pose, wherein the first viewer pose is the viewer pose for performing the alignment, wherein the viewer reference position depends on the orientation of the first viewer pose, wherein the viewer reference position has an offset relative to the viewer eye position of the first viewer pose, and wherein the offset includes an offset component in a direction opposite to the viewing direction of the viewer eye position; wherein the receiver is arranged to receive an image data signal that includes the image representation and an offset indication; and wherein the offset processor (309) is arranged to determine the offset in response to the offset indication.

2. The apparatus according to claim 1, wherein the offset component is not less than 2 cm.

3. The apparatus according to claim 1 or 2, wherein the offset component does not exceed 12 cm.

4. The apparatus according to claim 1 or 2, wherein the offset processor (309) is arranged to determine the offset in response to an error metric for at least one viewer pose, the error metric depending on a candidate value of the offset.

5. The apparatus according to claim 4, wherein the offset processor (309) is arranged to determine the error metric for the candidate value in response to a combination of error metrics for a plurality of viewer poses.

6. The apparatus according to claim 4, wherein the error metric for the viewer pose and the candidate value of the offset includes an image quality metric of a view image for the viewer pose synthesized from at least one image of the image representation, the at least one image having a position relative to the viewer pose that depends on the candidate value.

7. The apparatus according to claim 4, wherein the error metric for the viewer pose and the candidate value of the offset includes an image quality metric of a view image for the viewer pose synthesized from at least two images of the image representation, the at least two images having a reference position relative to the viewer pose that depends on the candidate value.

8. The apparatus according to claim 1 or 2, wherein the image representation includes an omnidirectional image representation.

9. The apparatus according to claim 1 or 2, Wherein, the offset includes an offset component in a direction perpendicular to the viewing direction of the viewer's eye position.

10. The apparatus according to claim 1 or 2, Wherein, the offset includes a vertical component.

11. An apparatus for generating an image signal, the apparatus comprising: a receiver for receiving a plurality of images representing a scene seen from one or more poses; a representation processor (205) for generating image data providing an image representation of the scene, the image data including the plurality of images, and the image representation being provided with respect to a scene coordinate system including a scene reference position; an offset generator (209) for generating an offset indication, wherein the offset indication indicates an offset, wherein the offset is applied between the scene reference position and the viewer's eye position when the scene coordinate system is aligned with the viewer coordinate system, and wherein the offset includes an offset component in a direction opposite to the viewing direction of the viewer's eye position; an output processor (207) for generating the image signal to include the image data and the offset indication.

12. The apparatus according to claim 11, Wherein, the offset generator (209) is arranged to determine the offset in response to an error metric for at least one viewer pose, the error metric depending on a candidate value of the offset.

13. The apparatus according to claim 12, Wherein, the offset generator (209) is arranged to determine an error metric for a candidate value in response to a combination of error metrics for a plurality of viewer poses.

14. A method of rendering an image, the method comprising: receiving an image representation of a scene, the image representation being provided with respect to a scene coordinate system including a scene reference position; determining a viewer pose for a viewer, the viewer pose being provided with respect to a viewer coordinate system; aligning the scene coordinate system with the viewer coordinate system by aligning the scene reference position with a viewer reference position in the viewer coordinate system; rendering view images for different viewer poses in response to the image representation and the alignment of the scene coordinate system with the viewer coordinate system; and determining the viewer reference position in response to a first viewer pose, wherein the first viewer pose is the viewer pose performing the alignment, wherein the viewer reference position depends on the orientation of the first viewer pose, and wherein the viewer reference position has an offset relative to the viewer's eye position of the first viewer pose, and wherein the offset includes an offset component in a direction opposite to the viewing direction of the viewer's eye position; wherein receiving the image representation of the scene includes receiving an image data signal, wherein the image data signal includes the image representation and an offset indication; and receiving the image representation of the scene further includes determining the offset in response to the offset indication.

15. A method for generating an image signal, the method comprising: Receive a plurality of images representing a scene as seen from one or more poses; Generate image data providing an image representation of the scene, the image data including the plurality of images, and the image representation being provided relative to a scene coordinate system including a scene reference position; Generate an offset indication, wherein the offset indication indicates an offset that is applied between the scene reference position and the viewer's eye position when the scene coordinate system is aligned with the viewer coordinate system, and wherein the offset includes an offset component in a direction opposite to the viewing direction of the viewer's eye position; Generate the image signal to include the image data and the offset indication.

16. A computer program product comprising computer program code modules which, when the program is run on a computer, are adapted to carry out all the steps of claim 14 or 15.

Citation Information

Patent Citations

  • Stereo viewing

    WO2015155406A1