Method for displaying stereoscopic image on self-stereoscopic display device
By rendering stereo images with fixed-position rendering points in the automatic stereo display device, the incorrect rendering and crosstalk problems caused by saccade are solved, and the accuracy of image display and viewing experience are improved.
Patent Information
- Application Number
- CN202380082817.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-08
- Filing Date
- 2023-12-08
- Publication Date
- 2025-07-11
AI Technical Summary
In automatic stereo display devices, incorrect rendering of stereo image and perceived crosstalk caused by saccade, especially when fast eye movements, resulting in errors in image display and inaccurate position perception.
The rendering points with fixed positions (left rendering points and right rendering points) are located in or near the rotation center of the viewer's eyes. The stereoscopic image is rendered by tracking the positions of these rendering points to avoid the impact of saccade on image rendering.
Reduces stereoscopic image rendering errors, improves viewers' perception of the 3D scene relative to the real world position, and improves the quality of viewing experience.
Smart Images

Figure CN120303609A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for displaying a stereoscopic image of a 3D scene to a viewer on an autostereoscopic display device. The present invention also relates to an autostereoscopic display device for displaying a stereoscopic image of a 3D scene to a viewer thereof. Background Art
[0002] Autostereoscopic displays are playing an increasingly important role in virtual reality and augmented reality applications. One of their most notable features is that they allow viewers to perceive three-dimensional images without wearing special glasses or other wearable devices, and they can still perceive the stereoscopic effect even if the viewer moves relative to the display.
[0003] The key to this technology is the use of an eye tracker combined with a screen containing a lenticular lens array or a parallax barrier. In this way, an autostereoscopic display is able to simultaneously direct the left-eye image to the viewer's left eye and the right-eye image to the viewer's right eye, thereby achieving a stereoscopic image. The resulting stereoscopic image provides a sense of depth, and elements in the image may appear in front of the display or behind the display (away from the display).
[0004] Like almost all electronic devices that rely on input data to function, autostereoscopic display systems face latency issues. Latency is usually understood as the time delay between user input and system response, also known as input lag. In autostereoscopic display systems, this basically means that there is a delay between the viewer's head / eye movement (user input) and the system's response to the displayed content. Latency manifests itself as incorrect rendering of stereo images and may even cause crosstalk, where parts of the left eye image are also directed to the viewer's right eye and vice versa. If the latency exceeds a critical threshold, the user's performance and experience will be affected, usually resulting in disrupted surround effects and / or crosstalk. For example, displayed objects cannot be perceived in the correct position, at least for a short period of time.
[0005] Delay in autostereoscopic display is usually obtained by extrapolating the historical position data of the viewer's pupil to obtain a prediction of the future position (the extrapolation can also include velocity and acceleration data). These predictions are used for rendering stereo images to ensure that the display system can respond in time (delay compensation) when the viewer's head moves quickly, thereby avoiding or reducing rendering errors.
[0006] In addition to the movement of the user's overall head, the pupil itself may also move relative to the autostereoscopic display (e.g., when the head remains stationary). This involves rotational movement of the eyeball. In some cases, this movement is very rapid and is called a saccade. It is this saccade that causes difficulties in latency compensation because the saccade is usually completed before the system can react. Moreover, this sudden movement does not follow the head movement model usually used to extrapolate past pupil positions to predict future positions. For example, the latency that usually needs to be compensated is in the range of 60 - 130 milliseconds, so it is necessary to predict the pupil position at 60 - 130 milliseconds in the future, while saccades usually occur within 20 - 60 milliseconds. Therefore, saccades are either not captured at all or cause incorrect pupil position prediction when captured by the eye tracker because the extremely fast eye movement speed measured during saccades may lead to unrealistic predictions. For example, the movement of the pupil position is extrapolated as a continuous movement, resulting in overshoot and jitter of the position.
[0007] Although the pupil displacement caused by eyeball rotation can be considered relatively small, especially compared to the free movement of the overall head, it is not negligible. Just the eyeball rotation, if not considered in a timely manner, can also lead to significant rendering errors. This problem is particularly evident especially in cases where the displayed object needs to be aligned with real-world objects. For example, when a virtual object is stationary on a stationary real-world object, and the viewer's eyes make a saccade relative to this combination and the eye tracker fails to react in time, the viewer will notice that the virtual object has moved somewhat (although small but significant) relative to the stationary real-world object, even though the virtual object should also be stationary. This effect is most obvious when the virtual object is close to the viewer's eyes. In addition, when the viewer's eyes make a saccade, the viewer may also experience crosstalk, when the light that should be displayed for the left eye shines on the right eye after a saccade of the right eye, and vice versa.
[0008] Therefore, it is necessary to address the impact of saccades in the latency compensation of the system. However, no satisfactory solution has been found to solve this problem so far. Summary of the Invention
[0009] Therefore, the object of the present invention is to solve the problem of incorrect rendering of stereoscopic images caused by saccades; and to solve the problem of perceived crosstalk caused by saccades. Another object is to reduce or even eliminate image display errors caused by the latency of the autostereoscopic display device, especially the image display of incorrect positions of 3D scenes relative to the real world. More generally, the object of the present invention is to improve the viewing experience of the viewer on the autostereoscopic display device, including improving the viewer's perception of the position of 3D scenes relative to the real world.
[0010] It has now been found that by adopting different eye tracking methods, one or more of the above objectives can be achieved.
[0011] Accordingly, the present invention relates to a method for displaying a stereoscopic image of a three-dimensional scene to a viewer of an autostereoscopic display device, the stereoscopic image being displayed by the autostereoscopic display device and consisting of a left-eye image for the viewer's left eye and a right-eye image for the viewer's right eye; the method comprising:
[0012] - providing 3D scene data representative of the 3D scene;
[0013] - rendering the stereoscopic image from the three-dimensional scene data, taking into account the viewing position of the viewer relative to the autostereoscopic display device, so that the viewer experiences the 3D scene from a viewing angle corresponding to his viewing position relative to the 3D scene;
[0014] wherein the method further comprises:
[0015] - determining the position of a left rendering point in the viewer's left eye relative to the autostereoscopic display device;
[0016] - determining the position of a right rendering point in the viewer's right eye relative to the autostereoscopic display device;
[0017] wherein the left rendering point and the right rendering point have fixed positions relative to the viewer's entire head and are located at the center of rotation of their respective eyes, or at a distance of 8.0 mm or less from the center of rotation;
[0018] - rendering the stereoscopic image from the 3D scene data using the position of the left rendering point relative to the autostereoscopic display device and the position of the right rendering point relative to the autostereoscopic display device;
[0019] - displaying the rendered stereoscopic image.
[0020] The present invention also relates to an autostereoscopic display device for displaying a stereoscopic image of a 3D scene to a viewer, the stereoscopic image consisting of a left image for the viewer's left eye and a right image for the viewer's right eye, the autostereoscopic display device comprising:
[0021] - a left and right rendering point tracking system configured to track
[0022] ○ the position of a left rendering point in the viewer's left eye relative to the autostereoscopic display device;
[0023] ○ the position of a right rendering point in the viewer's right eye relative to the autostereoscopic display device;
[0024] wherein the left rendering point and the right rendering point have fixed positions relative to the viewer's entire head and are located at the center of rotation of their respective eyes, or at a distance of 8.0 mm or less from the center of rotation;
[0025] - A display section, such as a screen, which is configured to display a left-eye image for a viewer's left eye and a right-eye image for the viewer's right eye, the display section comprising:
[0026] - An array of display pixel elements, which is configured to generate a display output;
[0027] - A lenticular device, which is disposed on the array, the lenticular device including lenticular lens regions capable of guiding display outputs from different display pixel elements to different spatial positions within the viewing field of the autostereoscopic display device, thereby enabling the display of a stereoscopic image including a left-eye image and a right-eye image;
[0028] - A rendering module, which is configured to render a stereoscopic image based on three-dimensional scene data, taking into account the viewing position of the viewer relative to the autostereoscopic display device, such that the viewer experiences the 3D scene from a viewing perspective corresponding to his viewing position relative to the 3D scene;
[0029] wherein the rendering module is further configured to render the stereoscopic image using the position of the left rendering point relative to the autostereoscopic display device and the position of the right rendering point relative to the autostereoscopic display device. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 Schematically shows a sequence of top views of a setup for rendering an image according to a conventional method.
[0031] Figure 2 Schematically shows a sequence of top views of a setup for rendering an image according to the method of the present invention.
[0032] Figure 3 is a visual representation of events occurring in the real world, reflecting the direction of the eyeball along the eyeball time axis and the image display lag due to latency along the display time axis. DETAILED DESCRIPTION
[0033] The drawings do not limit the present invention to the specific embodiments disclosed therein, and the elements in the description are shown schematically for purposes of simplicity and clarity and are not necessarily drawn to scale, with the emphasis being on clearly illustrating the principles of the invention. For example, the relative sizes between the screen of the autostereoscopic display device, the virtual objects presented on the screen, and the eyeballs perceiving the virtual objects cannot be deduced from the drawings.
[0034] In the present invention, the term "viewer" refers to a person who is capable of viewing the content presented by the autostereoscopic display device in the real world. Throughout the text, male terms such as "he" and "his" are used when referring to the viewer, solely for purposes of clarity and brevity, and are understood to be equally applicable to the female terms "she" and "her".
[0035] Throughout the text, the term "3D" is used for the sake of brevity. The term "3D" is equivalent to the term "three-dimensional". For example, the terms "3D-scene" and "3D-object" denote "three-dimensional scene" and "three-dimensional object", respectively.
[0036] In the present invention, the term "rendering" refers to the process of generating or deriving a left-eye image and a right-eye image from a three-dimensional scene from the perspective of a specific virtual stereoscopic camera based on the available 3D scene data by calculation. When viewed in combination with the left-eye image and the right-eye image, the viewer perceives the 3D scene as a 3D image from the perspective of the specific virtual stereoscopic camera.
[0037] Therefore, rendering in the present invention can be regarded as generating image data through a virtual stereoscopic camera, which represents a stereoscopic view of the 3D scene from a perspective corresponding to a specific position of the virtual stereoscopic camera relative to the 3D scene. The position of the virtual stereoscopic camera is at the position of the viewer's eyes, the left virtual camera is at the position of the viewer's left eye, and the right virtual camera is at the position of the viewer's right eye. Therefore, the movement of the viewer's head in the real world causes a change in the position of the virtual stereoscopic camera relative to the 3D scene. The rendering module accordingly continuously generates a subsequent pair of stereoscopic images of the left-eye image and the right-eye image from the moving viewpoint of the virtual stereoscopic camera. The tracking of the eye position ultimately provides the content recorded by the virtual stereoscopic camera by determining the position of the virtual stereoscopic camera relative to the 3D scene, and can be displayed as a stereoscopic image.
[0038] In the present invention, the term "3D-scene" refers to an environment constructed in a three-dimensional manner, which contains one or more elements having one or more known attributes, such as attributes selected from the following group: shape, size, surface orientation, surface properties, relative position (e.g., relative to other elements), and physical behavior. The behavior of light and / or sound in the 3D scene can also be known, resulting in, for example, shadows and surface reflections appearing in the rendered image. The 3D scene is usually a virtual environment, but can also be a (recorded) real environment.
[0039] In the present invention, the term "3D-scene data" refers to information representing a 3D scene, such as 3D features of 3D scene elements and their relative relationships in the 3D scene (i.e., the 3D mapping of the 3D scene). The 3D scene data can include data of one or more attributes of the scene elements, which are selected from the following group: shape, size, surface orientation, surface properties, relative position (e.g., relative to other elements), and physical behavior. According to the method of the present invention, the 3D scene data is used to generate a stereoscopic image corresponding to a specific perspective relative to the 3D scene, and this process is called "rendering". Such rendering can be performed for any desired perspective of the scene. Then, the rendered stereoscopic image can be displayed on a autostereoscopic display device and perceived by the viewer as a 3D image from a specific perspective.
[0040] 3D scene data can be stored in a memory section associated with the autostereoscopic display device or generated in real time, for example by inputting real-time recordings (usually recordings of the real world) into the autostereoscopic display device. It can also be a combination of both, for example when the viewer or someone else makes real-time modifications to the stored three-dimensional scene data.
[0041] In this description, it has been pointed out that the latency of the autostereoscopic display device can be compensated in the usual way, typically by extrapolating the historical viewing positions into the future and using these extrapolated values to render the stereoscopic images. This means that the apparent latency, i.e., the latency perceived by the viewer, is reduced here. The latency of the autostereoscopic display device itself is not affected by such methods.
[0042] The method according to the invention applies a known principle: displaying stereoscopic images of a 3D scene to the viewer in such a way that the viewer can experience the 3D scene from a perspective corresponding to his viewing position relative to the 3D scene. In addition, when the viewer's position relative to the 3D scene changes, the displayed stereoscopic images are updated accordingly to reflect this change. In this way, when the viewer makes a lateral movement relative to the display device, for example in the case where the 3D scene is fixed on the screen of the autostereoscopic display device, the viewer can experience the motion parallax of the 3D scene (i.e., the objects in the foreground appear to move more noticeably than the objects in the background). The viewer is also able to examine the displayed 3D objects from different angles, thus perceiving the so-called "circumstantial effect". Ideally, as pursued by the present invention, the motion parallax of the three-dimensional scene is consistent with that of the real world. For example, when a virtual object is positioned on a real object between the screen and the viewer, the relative positions of the two objects do not change when observed from different angles.
[0043] However, in order for the viewer to experience the 3D scene from a perspective corresponding to his viewing position, it is not necessarily required that the three-dimensional scene be fixed on the screen. In some embodiments, the 3D scene can be fixed in the real world and not affected by the movement of the screen relative to the real world; in this case, the screen acts as a virtual window or a movable frame through which the viewer views the 3D scene.
[0044] In other embodiments, the 3D scene can move relative to the screen, which allows the viewer to examine the three-dimensional scene from different angles while sitting in front of the screen. In this case, the viewer may use a joystick, a mouse or other input devices to control the movement of the 3D scene, such as its rotation or translation relative to the screen.
[0045] Known methods for adapting a three - dimensional scene view to the viewer's perspective usually present the desired image information accurately to each eye by tracking the pupil position. However, all display devices applying this method have a significant problem - when the pupil movement is very fast, the rendering of the stereoscopic image will be significantly inaccurate. This means that the time required for the pupil to move to a new position sufficient to cause a significant error in image display is shorter than the latency of the display device. This latency is usually referred to as the response time of the display device, that is, the time delay between the moment when the viewer's pupil reaches a new position and the moment when the eye sees the display image that has taken this new pupil position into account.
[0046] For a given latency, slower pupil movements do not cause image display errors, such as normal head movements or pupil movements when the pupil is in the so - called smooth pursuit mode. On the other hand, relatively fast pupil movements cannot be taken into account in time. Such fast pupil movements are almost always caused by saccades rather than by the regular movement of the whole head. The speed of saccades is usually too fast for the stereoscopic display device to take them into account in time.
[0047] Figure 3 Illustrated by an example scenario, the figure shows the eye time axis, reflecting the change in the orientation of the eye at time t 眼球 (the lower time axis in the figure), and the display time axis, reflecting the display of the stereoscopic image at time t 显示 (the upper time axis in the figure). Due to a system latency of 80 milliseconds, the display time axis lags behind the eye time axis. Between the two time axes, a schematic top view of the eye, including the pupil, is provided, which shows different pupil positions at certain intervals of the eye time axis to illustrate the changing orientation of the eye over time (such as the saccade occurring between 20 and 80 milliseconds).
[0048] In this specific example, the display time axis is delayed by 80 milliseconds relative to the eye time axis. This delay is caused by the latency of the stereoscopic display device. A latency of 80 milliseconds has the following effect: the orientation of the eye at t 眼球 = 0 milliseconds is used to render the image that the viewer sees at t 眼球 = 80 milliseconds (i.e., t 显示 = 0 milliseconds). And the 60 - millisecond saccade that ends at t 眼球 = 80 milliseconds is not taken into account in time. Therefore, at t 眼球 = 80 milliseconds, the orientation of the eye does not appear in the image rendered at this time (t 显示 = 0 milliseconds). At the same time, the eye orientation at t 眼球 = 80 milliseconds is used to render the image that will be seen at 160 milliseconds (i.e., t 显示An image that is presented to a viewer (e.g., with a latency of 80 milliseconds). Thus, it is only 80 milliseconds after the saccade is completed that the viewer will see an image that is consistent with the eye orientation resulting from that saccade (assuming no new saccades occur during this period).
[0049] In addition, during the occurrence of a saccade, the image never aligns with the eye orientation (the degree of misalignment increases gradually from zero at t 眼球 = 20 milliseconds and reaches a maximum at t 眼球 = 80 milliseconds). Thus, within the time period from t 眼球 = 20 milliseconds to t 眼球 = 160 milliseconds, the viewer will see an incorrectly rendered image.
[0050] Figure 1 shows incorrect rendering due to the pupil position not being considered in a timely manner. Each figure represents a setup from a top-down perspective, in which the image is rendered to the viewer's pupil. Each image includes a triangular object and a square object, which the viewer perceives as being in front of the screen of the autostereoscopic display device (the bottom of each figure shows an eye with a pupil, and the top shows the screen that the eye is looking at). The four figures correspond to different time points during a saccade in the setup. These time points correspond to the events visualized on the eye timeline in Figure 3 . The eye is initially pointed at the triangle (t 眼球 = 0 milliseconds). When the eye points at the square object at t 眼球 = 80 milliseconds (i.e., after the saccade is completed), it can be seen that the rendering of the square object at this time is still based on the pupil position at t 眼球 = 0 milliseconds. It is only at t 眼球 = 160 milliseconds that the rendered image is rendered based on the correct pupil position.
[0051] The present invention is not aimed at solving latency itself (i.e., true latency reduction), such as by improving the hardware or software of the display device, but rather provides a method in which image rendering is independent of the position of the pupil relative to the head, so that the occurrence of a saccade no longer affects the rendering. In this alternative method, image rendering is performed at a rendering point in the viewer's eye, and the position of this rendering point relative to the overall head is (substantially) fixed. Therefore, the position of the rendering point in the eye is not affected by whether the eye rotates in the eye socket. According to the present invention, the two rendering points are located at or near the center of their respective eyes (usually no more than 5.0 millimeters from the center).
[0052] According to common theories, traditional rendering of stereoscopic images on the pupil provides the best user perception of 3D scenes because it is stable and fixed in the real world. This means that the pupil position obtained by the eye tracker is regarded as the viewing position for rendering the stereoscopic image. However, surprisingly, the deviation method employed in the present invention (i.e., selecting a rendering point located at or near the center of the eyeball) hardly affects the viewer's perception of the position of the 3D scene relative to the real world. This is probably because the deviation method applied in the present invention does not significantly affect the light rays passing through the lens, whether these light rays pass through the normal direction of the lens or through a direction slightly deviated from the normal, because these light rays do not (or hardly) refract. Moreover, it is this direction that is most important to the viewer because the viewer's line of sight is mainly focused on the direction of these light rays that hardly refract. Although all light rays in non-normal directions theoretically cause some incorrect rendering, it should be admitted that directions close to the normal direction (e.g., within 10° deviation from the normal) do not cause the user to perceive obvious image degradation. Only in the peripheral area far from the viewing point may there be a significant error in the position of the 3D scene relative to the real world, but usually the viewer does not notice these errors because the human eye does not have many light receptors in the peripheral area of the visual field. Therefore, the perceptual details in the peripheral area are much less than those seen by the viewer through the central area of the visual field, which has more light receptors.
[0053] The beneficial effect of this incorrect rendering in a specific peripheral area of the line of sight is that when the line of sight points to these areas after a saccade, this incorrect rendering immediately disappears (assuming that the change in pupil position caused by the overall movement of the head is relatively small compared to the change in pupil position caused by the saccade). This is because according to the method of the present invention, the rendering is performed on a rendering point located at the center (or near it) of the eyeball. Regardless of the orientation of the eye, the light rays in the normal direction passing through the eye lens always pass through this rendering point - the same is true for the viewer's line of sight light rays. In other words, all areas of the image rendered according to the present invention are rendered as if the viewer is simultaneously looking at all these areas. The new fixation point after each saccade is thus automatically rendered correctly. After a saccade, the user does not perceive a change in the rendered position after the actual display delay because the rendered position is not affected by the saccade. This is exactly why the method of the present invention excludes the influence of saccades without having a negative impact on the viewer's perception of the image at the viewing point.
[0054] The rendering according to the present invention is in Figure 2 similar to Figure 1Visualized in the following manner. Each figure shows a setup from a top-down perspective, in which the image is rendered at the center of the viewer's eye (rendering point). Each image includes a triangular object and a square object, which the viewer perceives as being in front of the screen (the eye with the pupil is shown at the bottom of each figure, and the screen that the eye is looking at is shown at the top). The four figures correspond to different time points during a saccade that occurs in the same setup. These time points correspond to Figure 3 the events visualized on the eye timeline. The eye is initially pointing at the triangle (t 眼球 = 0 milliseconds). When the eye points at the square object, at t 眼球 = 80 milliseconds (i.e., after the saccade is complete), it can be seen that at this time the rendering of the square object is based on the rendering point at the center of the eye, which remains in the same position as the eye rotates. In other words, the rotation of the eye does not cause incorrect rendering at the viewing point (the square object). However, it can be seen that at t 眼球 = 0 milliseconds, the rendering of the square object is indeed incorrect. As explained above, this incorrect rendering occurs in the peripheral region of the visual field, and thus is more acceptable than incorrect rendering at the viewing point.
[0055] Therefore, rendering occurs at two rendering points, the positions of which are defined relative to the overall head. The left eye is associated with the left rendering point, and the right eye is associated with the right rendering point. Each rendering point is located within the transparent inner part of its respective eye, i.e., the vitreous body. Each rendering point itself is not specifically defined or demarcated by specific physical features. It is a point within the volume element around the center of rotation of its respective eye. Each rendering point has a fixed position relative to the viewer's head and is located at the center of rotation of its respective eye (or at a position 8.0 millimeters or less away from the center of rotation).
[0056] It should be noted that the center of rotation of the eye may vary slightly relative to the eye socket (and thus relative to the head) due to the actual orientation of the eye, because the eye and / or the eye socket do not necessarily form a perfect sphere. Therefore, in reality, the eye may have multiple centers of rotation. Since the change in the position of the center of rotation is very small, the viewer does not perceive these changes, and thus the perceived quality of the stereoscopic image displayed will be the same as the quality of the method of the present invention when considering variable centers of rotation. Therefore, if the eye has multiple centers of rotation, any one of them can be selected to define the rendering point.
[0057] The rendering point is generally located at a position within 8.0 millimeters from the center of rotation. More preferably, the rendering point is located at a relatively short distance from the center of rotation of its respective eye, such as 7.0 millimeters or less, 6.0 millimeters or less, 5.0 millimeters or less, 4.0 millimeters or less, 3.0 millimeters or less, 2.0 millimeters or less, or 1.0 millimeters or less. In a preferred embodiment, their positions coincide with the centers of rotation of their respective eyes.
[0058] During rendering, the position of the left rendering point relative to the autostereoscopic display device and the position of the right rendering point relative to the autostereoscopic display device are used. These positions can be obtained in the following ways: (1) After initially determining the position of the rendering point in the eye (or the position relative to the pupil), tracking the binocular eyes of a specific viewer; or (2) When the positions of one or more other facial features of the viewer relative to the left rendering point and the right rendering point are known (e.g., through initial determination), tracking one or more other facial features of a specific viewer.
[0059] Thus, in one embodiment, the position of the left rendering point relative to the autostereoscopic display device and the position of the right rendering point relative to the autostereoscopic display device are determined by the following steps:
[0060] 1) Determine one or more facial features of the viewer that do not include the left rendering point and the right rendering point;
[0061] 2) Determine the positions of the left rendering point and the right rendering point relative to these facial features;
[0062] 3) Determine the positions of these facial features relative to the autostereoscopic display device;
[0063] 4) Use the positions determined in steps 2) and 3) to determine the position of the left rendering point relative to the autostereoscopic display device and / or the position of the right rendering point relative to the autostereoscopic display device.
[0064] Facial features that can be used in this method are preferably selected from the following group: nose, mouth, ears, wrinkles, and eyebrows.
[0065] The method of the present invention uses an autostereoscopic display device. This device is usually fixed in the real world during use, such as a desktop device or a wall-mounted device. For example, the autostereoscopic display device can be a television, a (desktop) computer with a monitor, a laptop computer, or a theater display system. It can also be a portable device, such as a mobile phone, a tablet computer, or a game console.
[0066] In the method of the present invention, the autostereoscopic display device is preferably a device including a pixel array and equipped with lenticular lenses, which can direct the pixel outputs (i.e., light rays) belonging to the left-eye image to the left eye of the viewer and can direct the pixel outputs (i.e., light rays) belonging to the right-eye image to the right eye of the viewer. The combined output of the pixels in the pixel array constitutes the overall output of the display, i.e., the display output.
[0067] This autostereoscopic display device preferably further includes a tracking system configured to determine the positions of the left and right rendering points of the viewer relative to the autostereoscopic display device. Then, the obtained position data is used to render a stereoscopic image and control the pixels.
[0068] Thus, in one embodiment, the method of the present invention includes a autostereoscopic display device, which includes:
[0069] - A left and right rendering point tracking system configured to track the positions of the left and right rendering points relative to the autostereoscopic display device; and
[0070] - A display section configured to display a left-eye image for a left-eye viewer and a right-eye image for a right-eye viewer, the display section including
[0071] ○ An array of display pixel elements for generating a display output; and
[0072] ○ A lenticular lens device mounted on the array, wherein the lenticular lens device includes lenticular lens regions capable of directing the display outputs from different display pixel elements to different spatial positions in the field of view of the autostereoscopic display device, thereby displaying a stereoscopic image composed of the left-eye image and the right-eye image.
[0073] Typically, such an autostereoscopic display device further includes a rendering module configured to render a stereoscopic image based on 3D scene data, taking into account the position of the left rendering point relative to the autostereoscopic display device and the position of the right rendering point relative to the autostereoscopic display device.
[0074] The position data of the left and right rendering points relative to the autostereoscopic display device can also be used as an input for weaving the left-eye image and the right-eye image into the array of display pixel elements, i.e., selecting the correct display pixel elements to display these two images so that the stereoscopic image is presented to the viewer as expected. Thus, in the method of the present invention, the display process of rendering the stereoscopic image may include weaving the left-eye image and the right-eye image into the array of display pixel elements, wherein the weaving process includes:
[0075] - Considering the position of the left rendering point relative to the autostereoscopic display device and the position of the right rendering point relative to the autostereoscopic display device respectively, selecting the display pixel elements for the left-eye image pixel output and selecting the display pixel elements for generating the right-eye image pixel output;
[0076] - Controlling the selected display pixel elements accordingly to display the stereoscopic image to the viewer.
[0077] Typically, the method of the present invention is executed multiple times in a sequence. In this way, the new position of the viewer relative to the autostereoscopic display device can be considered, enabling the viewer to perceive the movement of the displayed 3D scene. For example, the method is executed at least 10 times, at least 100 times, at least 1000 times, at least 10,000 times, at least 100,000 times or at least 1,000,000 times.
[0078] For each new position of the viewer relative to the 3D scene, the stereoscopic images are re - rendered, giving the viewer the impression that they are truly moving relative to the 3D scene (alternatively, when the viewer can control the position and orientation of the 3D scene relative to the screen of the autostereoscopic display device, the 3D scene is indeed moving relative to the viewer), especially when the re - rendering is performed at a suitable frequency. To obtain a realistic viewing experience, the rendering is typically performed at a frequency of at least 10 times per second. Preferably, the frequency is at least 20 times per second, more preferably 30 times per second, and even more preferably 50 times per second. For example, the frequency can be between 50 - 250 times per second, between 55 - 150 times per second, or between 60 - 120 times per second. In particular, it can be 60Hz, 120Hz, 144Hz, 165Hz, or 240Hz. Preferably, its frequency is the same as the refresh rate of the screen itself.
[0079] As described above, the method of the present invention excludes saccades. This basically means that when (1) the first stereoscopic image is displayed by the autostereoscopic display device and (2) a saccade occurs before the second subsequent stereoscopic image is displayed, the second stereoscopic image will not be different from the first stereoscopic image. In other words, the autostereoscopic display device implementing the method of the present invention does not respond to saccades. Additionally, after a saccade occurs, the viewer immediately experiences a good stereoscopic image rendering effect at the fixation point.
[0080] Therefore, the stereoscopic image rendering in the method of the present invention only considers the movement of the overall head (rather than the movement of the eyeball relative to other parts of the head). This movement is not as rapid as the movement of the pupil during a saccade and can be predicted using known methods. The input for such prediction consists of the historical head position and orientation, as well as the velocities and / or accelerations of the left and right rendering points relative to the autostereoscopic display device. Through such prediction, the apparent latency of the autostereoscopic display device can be significantly reduced (in other words, the latency of the autostereoscopic display device can be significantly compensated), thereby improving the viewing experience of the viewer. Therefore, the method of the present invention may include reducing the apparent latency of the autostereoscopic display device (or in other words, may include compensating for the latency of the autostereoscopic display device).
[0081] In this way, compared to rendering stereoscopic images by using the pupil position instead of the rendering points according to the present invention, the stereoscopic images are rendered more accurately and with less jitter within the viewer's fixation area.
[0082] In particular, the method of the present invention may include the following steps:
[0083] - Obtaining data on the velocities and / or accelerations of the left and right rendering points relative to the autostereoscopic display device;
[0084] - Using the obtained data to predict the positions of the left and right rendering points relative to the autostereoscopic display device;
[0085] - Render a stereoscopic image based on the predicted positions using the 3D scene data and display it to the viewer;
[0086] - Optionally, use the predicted positions to weave the left-eye image and the right-eye image into an array of display pixel elements.
[0087] In some cases, it is not necessary to predict the positions of the left and right rendering points. Since the distance between these two rendering points is fixed, only the velocity and / or acceleration of one of the rendering points needs to be predicted when the following two conditions are met: (1) the head orientation is known and can be taken into account; and (2) the relative positions of the two rendering points in the head are known.
[0088] In the art, it is known how to compensate for the latency of a autostereoscopic display device. For example, this can be done by using the history of multiple left-eye and right-eye viewing positions of the viewer and fitting a model to these viewing positions to extrapolate the positions at future times.
[0089] The 3D scene can in principle be any imaginable 3D scene. Since the 3D scene needs to be perceivable from different viewing positions and the perspectives corresponding to these viewing positions, there must be a mapping of the 3D scene that includes more than just the stereoscopic images taken from a specific perspective. For this reason, the 3D scene data preferably represents an artificially created 3D scene because such a scene typically includes an extensive or complete 3D mapping of the 3D scene - such a 3D mapping in principle allows the generation of stereoscopic images from any viewpoint (i.e., from any virtual stereoscopic camera position).
[0090] However, in principle, the 3D scene data can also represent a 3D scene in the real world. For example, a real scene can be recorded from multiple different camera positions. The 3D scene data obtained from the recordings at different camera positions enables the real scene to be perceived from different perspectives (i.e., it allows the generation of stereoscopic images from multiple viewpoints). Therefore, in this case, stereoscopic images can also be generated from any viewpoint (i.e., from any virtual stereoscopic camera position).
[0091] The present invention further relates to an autostereoscopic display device configured to perform the method of the present invention (i.e., to perform one of the methods described above).
[0092] In a preferred embodiment, the autostereoscopic display device includes a latency compensation module (in other words, an apparent latency reduction module) configured to predict the position of the left rendering point relative to the autostereoscopic display device and the position of the right rendering point relative to the autostereoscopic display device, wherein the rendering module is configured to use the predicted positions to render the stereoscopic image.
[0093] In another embodiment, the autostereoscopic display device is configured to determine the positions of the left and right rendering points relative to the autostereoscopic display device by determining the positions of one or more facial features of the viewer, which facial features have known positions relative to the two rendering points. Thus, the autostereoscopic display device of the present invention can:
[0094] - be configured to determine the positions of one or more facial features (excluding the left and right rendering points) of the viewer relative to the autostereoscopic display device; and
[0095] - include a memory storing stored data representing the positions of the left and right rendering points relative to the one or more facial features.
[0096] The autostereoscopic display device of the present invention generally includes a weaving module configured to weave the left-eye image and the right-eye image into an array of display pixel elements by:
[0097] - selecting display pixel elements that produce left-eye image pixel outputs and selecting display pixel elements that produce right-eye image pixel outputs;
[0098] - correspondingly controlling the selected display pixel elements to display the stereoscopic image to the viewer.
[0099] In a preferred embodiment, when selecting the display pixel elements that produce left-eye image pixel outputs and right-eye image pixel outputs, the positions of the left rendering point relative to the autostereoscopic display device and the position of the right rendering point relative to the autostereoscopic display device are taken into account.
[0100] The autostereoscopic display device of the present invention can be a device that is substantially stationary during actual use, such as a desktop device or a wall-mounted device. For example, the autostereoscopic display device can be a television, a (desktop) computer with a display, a laptop computer, or a cinema display system. It can also be a portable device, such as a mobile phone, a tablet computer, or a gaming console.
Claims
1. A method for presenting a stereoscopic image of a 3D scene to a viewer via an autostereoscopic display device, the stereoscopic image consisting of a left-eye image viewed by the viewer's left eye and a right-eye image viewed by the viewer's right eye; The method includes: - Providing 3D scene data representing the 3D scene; - Rendering the stereoscopic image based on the 3D scene data by creating the left-eye image and the right-eye image from the 3D scene data and taking into account the viewing position of the viewer relative to the autostereoscopic display device so that the viewer can experience the 3D scene from a perspective corresponding to their viewing position relative to the 3D scene, wherein the method further includes: - Determining the position of a left rendering point within the viewer's left eyeball relative to the autostereoscopic display device; - Determining the position of a right rendering point within the viewer's right eyeball relative to the autostereoscopic display device; The left rendering point and the right rendering point have fixed positions relative to the viewer's entire head and are located at the center of rotation of their respective eyeballs or 5.0 mm or less from the center of rotation; - Using the position of the left rendering point relative to the autostereoscopic display device and the position of the right rendering point relative to the autostereoscopic display device to render the stereoscopic image based on the 3D scene data; - Displaying the rendered stereoscopic image to the viewer.
2. The method according to claim 1, wherein the distance of the left rendering point and the right rendering point from the center of rotation of their respective eyeballs is 3.0 mm or less, preferably 1.0 mm or less.
3. The method according to claim 1 or 2, wherein the position of the left rendering point relative to the autostereoscopic display device and the position of the right rendering point relative to the autostereoscopic display device are determined by the following steps: 1) Identifying one or more facial features of the viewer, the facial features not including the left rendering point and the right rendering point; 2) Determining the position of the left rendering point and the right rendering point relative to the one or more facial features; 3) Determining the position of the one or more facial features relative to the autostereoscopic display device; 4) Using the positions determined in steps 2) and 3) to determine the position of the left rendering point relative to the autostereoscopic display device and / or the position of the right rendering point relative to the autostereoscopic display device.
4. The method according to any one of claims 1-3, wherein the autostereoscopic display device includes: - A left and right rendering point tracking system configured to track the position of the left rendering point relative to the autostereoscopic display device and the position of the right rendering point relative to the autostereoscopic display device; - A display section configured to display the left-eye image to the viewer's left eye and the right-eye image to the viewer's right eye, the display section including: ○ An array of display pixel elements for generating a display output; and ○ A lenticular lens device disposed on the array, wherein the lenticular lens device includes a lenticular lens area capable of guiding display outputs from different display pixel elements to different spatial positions in the field of view of the autostereoscopic display device, thereby displaying a stereoscopic image composed of a left-eye image and a right-eye image.
5. The method according to claim 4, wherein displaying the rendered stereoscopic image includes weaving the left-eye image and the right-eye image into an array of display pixel elements, wherein the weaving includes: - Respectively considering the position of the left rendering point relative to the autostereoscopic display device and the position of the right rendering point relative to the autostereoscopic display device, selecting the display pixel elements that produce the left-eye image pixel output and selecting the display pixel elements that produce the right-eye image pixel output; - Controlling the selected display pixel elements accordingly to display the stereoscopic image.
6. The method according to any one of claims 1 - 5, wherein the method is repeatedly executed multiple times, for example, at least 10 times, at least 100 times, at least 1,000 times, or at least 10,000 times.
7. The method according to any one of claims 1 - 6, wherein the method is repeatedly executed at a frequency of at least 50 times per second, preferably in the range of 55 - 150 times per second.
8. The method according to any one of claims 1 - 7, wherein the method includes compensating for the latency of the autostereoscopic display device.
9. The method according to any one of claims 1 - 8, wherein the method includes: - obtaining data on the speed and / or acceleration of the left rendering point and / or the right rendering point relative to the autostereoscopic display device; - using the obtained data to predict the position of the left rendering point relative to the autostereoscopic display device and the position of the right rendering point relative to the autostereoscopic display device; - using the predicted positions to render a stereoscopic image based on 3D scene data and display it to the viewer; - optionally, using the predicted positions to display the rendered stereoscopic image.
10. The method according to any one of claims 1 - 9, wherein the 3D scene data represents an artificially created 3D scene.
11. The method according to any one of claims 1 - 9, wherein the 3D scene data represents a 3D scene in the real world.
12. An autostereoscopic display device for presenting stereoscopic images of a 3D scene to a viewer, the autostereoscopic display device comprising: - a left rendering point and right rendering point tracking system configured to track: ○ The position of the left rendering point within the viewer's left eye ball relative to the autostereoscopic display device; ○ The position of the right rendering point within the viewer's right eye ball relative to the autostereoscopic display device; The left rendering point and the right rendering point have fixed positions relative to the entire head of the viewer and are located at the center of rotation of their respective eyeballs, or 5.0 mm or less from the center of rotation; - a display section configured to display a left eye image to the viewer's left eye and a right eye image to the viewer's right eye, the display section comprising: ○ An array of display pixel elements for generating a display output; ○ A lenticular lens device disposed on the array, wherein the lenticular lens device includes a lenticular lens area capable of guiding display outputs from different display pixel elements to different spatial positions in the viewing field of the autostereoscopic display device, thereby displaying a stereoscopic image composed of a left-eye image and a right-eye image; - a rendering module configured to render a stereoscopic image based on 3D scene data, taking into account the viewing position of the viewer relative to the autostereoscopic display device, so that the viewer can experience the 3D scene from a viewing angle corresponding to their viewing position relative to the 3D scene; wherein the rendering module is further configured to render a stereoscopic image using the position of the left rendering point relative to the autostereoscopic display device and the position of the right rendering point relative to the autostereoscopic display device.
13. The autostereoscopic display device according to claim 12, the autostereoscopic display device further comprising a latency compensation module configured to predict the position of the left rendering point relative to the autostereoscopic display device and the position of the right rendering point relative to the autostereoscopic display device, wherein the rendering module is configured to render a stereoscopic image using the predicted positions.
14. The autostereoscopic display device according to claim 12 or 13, wherein the autostereoscopic display device: - is configured to determine the position of one or more facial features of the viewer relative to the autostereoscopic display device, the one or more facial features not including the left rendering point and the right rendering point; - includes a memory storing stored data representing the positions of the left rendering point and the right rendering point relative to the one or more facial features.
15. The autostereoscopic display device according to any one of claims 12 - 14, wherein the autostereoscopic display device includes a weaving module configured to weave a left-eye image and a right-eye image into an array of display pixel elements by: - Selecting display pixel elements that produce left-eye image pixel outputs and display pixel elements that produce right-eye image pixel outputs, taking into account the position of the left rendering point relative to the autostereoscopic display device and the position of the right rendering point relative to the autostereoscopic display device, respectively; - Controlling the selected display pixel elements accordingly to display a stereoscopic image to a viewer.