Adjustable Near Eye Display
The stereoscopic near-eye display assembly addresses the issue of vergence accommodation conflict by using a diffractive optical element and preprocessing techniques to extend the depth of field, resulting in improved user comfort and image quality.
Patent Information
- Application Number
- JP2024570579
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-04-20
- Filing Date
- 2024-04-23
- Publication Date
- 2025-06-19
- Estimated Expiration
- 2044-04-23
AI Technical Summary
Conventional near-eye display (NED) devices struggle to provide a comfortable and natural viewing experience due to vergence accommodation conflict (VAC), which results in user discomfort and compromised image quality.
The implementation of a stereoscopic near-eye display assembly that includes a 2D display, a lens assembly with a diffractive optical element (DOE) providing different phase delays, and a preprocessing unit that applies image-region position-dependent preprocessing to extend the depth of field and reduce VAC.
This solution effectively reduces user discomfort caused by VAC while maintaining high image quality and reducing the complexity of the NED device, providing an adjustment-invariant 3D presentation across an extended depth of field.
Smart Images

Figure 2025518729000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a near-eye display (NED).
Background Art
[0002] Stereoscopic near-eye displays (NEDs) are particularly applicable in various virtual reality (VR) and augmented reality (AR) applications. Due to the visual discomfort experienced by some users, there is a problem in ensuring a comfortable but natural or near-natural viewing experience through the use of wearable NED devices.
[0003] One of the important problems in ensuring a comfortable viewing experience involves dealing with vergence accommodation conflict (VAC), which can occur when the NED device is applied to display an object such that its vergence distance does not match its accommodation distance. When viewing VR or AR content using conventional NED devices, such a situation frequently occurs due to displaying objects of three-dimensional (3D) images having respective positions in the image space at distances different from the distance to the position of the (virtual) image plane of the NED device (i.e., at their respective vergence distances).
[0004] NED devices equipped with technologies for dealing with VAC, such as focus-variable, multi-focus, light irradiation field, holographic, and Maxwellian, are known in the art, but they all suffer from their own trade-offs, such as limited eye boxes, limited image quality (e.g., low spatial resolution, speckle noise), and / or complex device requirements (e.g., bulky optical systems and / or high-speed optical systems, accurate eye tracking devices). As a result, there is a continuing need for NED devices that can reduce VAC in a way that provides a high level of user comfort without sacrificing image quality and device complexity.
Summary of the Invention
[0005] An object of the present invention is to provide a technique for reducing or eliminating the user's discomfort caused by the convergence adjustment competition (VAC) in the NED device without significantly degrading the resulting perceived image quality.
[0006] According to an exemplary embodiment, a stereoscopic near-eye display (NED) assembly for an NED device including a pair of NED assemblies is provided, the NED assembly including a two-dimensional (2D) display for rendering an image for viewing by a viewer's eyes, a lens assembly disposed at a predefined distance from the 2D display to enable viewing of the image rendered through the 2D display, the lens assembly including a diffractive optical element (DOE) configured to provide different phase delays through a plurality of positions of its aperture, and a preprocessing unit configured to derive a preprocessed image based on the received image through application of a preprocessing procedure that causes different transmission characteristics through different sub-apertures of the lens aperture of the lens assembly and applies image-region position-dependent preprocessing in the derivation of different sub-regions of the image region of the preprocessing image and to supply the preprocessed image for rendering through the 2D display, the preprocessing procedure and the lens assembly being configured to display the preprocessed image as being sharp for a plurality of predefined adjustment depths existing within an extended depth of field DoF of the NED assembly when viewed through the lens assembly, facilitating the provision of an adjustment-invariant 3D presentation based on the preprocessed image through operation of the NED device.
[0007] According to another exemplary embodiment, there is provided a preprocessing procedure for a stereoscopic NED assembly according to the foregoing exemplary embodiment and a device for deriving a phase delay profile, the device being configured to apply respective learning models to derive together at least one artificial neural network (ANN) that functions as the phase delay profile that determines respective phase delays for a plurality of positions of the DOE via the preprocessing procedure and an iterative learning procedure based on a plurality of training images, the device in a plurality of iterative rounds: for each of the plurality of training images, for each of the respective iterative rounds, select one or more sub-apertures, determine one or more sub-regions of each of the image regions of the respective training images that spatially correspond to the selected one or more sub-apertures; process the determined one or more sub-regions of the respective training images by the at least one ANN to determine one or more sub-regions of the preprocessed image that spatially correspond thereto using the current weights of the at least one ANN, taking into account each position within the image region; process the one or more sub-regions of the preprocessed image by a learning model configured to perform, by a display and an eye model, the one or more sub-regions of the preprocessed image into one or more sub-regions of a simulated retinal image that spatially correspond thereto according to a current phase delay profile and taking into account an adjustment depth selected for each of the respective iterative rounds, and determine a difference between the one or more sub-regions of the simulated retinal image and one or more sub-regions of a corresponding reference retinal image that spatially correspond thereto by a predefined loss function; and update the weights of the at least one ANN and a part of the phase delay profile that spatially corresponds to the one or more selected sub-apertures based on each difference determined for the plurality of training images.
[0008] According to another exemplary embodiment, a preprocessing procedure for a stereoscopic NED assembly according to the foregoing exemplary embodiment and a method for deriving a phase delay profile are provided, the method including performing an iterative learning procedure based on a plurality of training images by applying respective learning models to jointly derive the preprocessing procedure and at least one ANN that functions as the phase delay profile defining respective phase delays for a plurality of positions of the DOE, the method for a plurality of iterative rounds, for each of the plurality of training images, selecting one or more sub-apertures for the respective iterative round, determining respective sub-regions of an image region of the respective training image that spatially correspond to the selected one or more sub-apertures, processing the determined one or more sub-regions of the respective training image by the at least one ANN to determine one or more sub-regions of a preprocessed image that spatially correspond thereto using current weights of the at least one ANN taking into account respective positions within the image region, processing the one or more sub-regions of the preprocessed image by a display and an eye model according to a current phase delay profile and considering an adjustment depth selected for each iterative round to one or more sub-regions of a simulated retinal image that spatially correspond thereto, and determining a difference between the one or more sub-regions of the simulated retinal image and one or more sub-regions of a corresponding reference retinal image that spatially correspond thereto by a predefined loss function, processing through a learning model including the steps, and updating the weights of the at least one ANN and a part of the phase delay profile that spatially corresponds to the one or more selected sub-apertures based on respective differences determined for the plurality of training images.
[0009] According to another exemplary embodiment, a computer program is provided, the computer program including computer-readable program code configured to cause the method according to the foregoing exemplary embodiment to be performed when the program code is executed on one or more computing devices.
[0010] The computer program according to the above exemplary embodiments can be embodied as a computer program product including at least one non-transitory computer-readable medium storing program code, for example, on a volatile or non-volatile computer-readable recording medium. When executed by one or more computing devices, the computing devices are caused to execute at least the method according to the above exemplary embodiments.
[0011] The exemplary embodiments of the invention presented in this patent application should not be construed as imposing limitations on the applicability of the appended claims. The verb "comprising" and its derivatives are used in this patent application as open limitations that do not exclude the presence of features not recited. The features described below can be freely combined with each other unless otherwise specified.
[0012] Some features of the invention are recited in the appended claims. However, aspects of the invention will be best understood from the following description of some exemplary embodiments when read in conjunction with the accompanying drawings, with additional objects and advantages as well as with respect to both its structure and its manner of operation.
[0013] Embodiments of the invention are shown by way of example and not by way of limitation in the figures of the accompanying drawings.
Brief Description of the Drawings
[0014]
Figure 1A
Figure 1B
Figure 2A
Figure 2B
Figure 3A
Figure 3B
Figure 4
Figure 5
Figure 6
DETAILED DESCRIPTION OF THE INVENTION
[0015] Figures 1A and 1B schematically show a viewer's eye 110 viewing some characteristics of a conventional stereoscopic near-eye display (NED) assembly 101. The NED assembly 101 includes a two-dimensional (2D) display 102 and a magnifying lens 103, which are positioned in front of the viewer's eye 110 during operation of the NED assembly 101 for viewing an image rendered on the 2D display 102 through the magnifying lens 103. The 2D display 102 is positioned at a fixed position relative to the magnifying lens 103, and the position of the 2D display 102 is closer to the magnifying lens 103 than its focal length. As a result, the 2D display 102 is mapped to a (virtual) image plane 104 located at a fixed distance behind the 2D display map 102. The NED device can include a pair of NED assemblies 101 (i.e., each NED assembly 101 for each eye 110 of the viewer) and a display controller configured to supply each image of a stereoscopic image pair for rendering on each 2D display 102 of the pair of NED assemblies 101 to provide a three-dimensional (3D) presentation of the scene captured in the stereoscopic image pair.
[0016] Theoretically, assuming an optical system without aberration, the visual content displayed on the 2D display 102 can be shown with a resolution up to the diffraction limit in the image plane 104. At a depth offset from the image plane 104 (i.e., the distance from the 2D display 102), due to defocus blur (or defocus aberration), a sharp decline in the frequency response is seen with an increase in depth. Thus, if the viewing depth of the viewer's eye 110 coincides with the image plane 104, the viewer observes a sharp image at a high resolution, but the greater the offset between the viewing depth of the eye 110 and the image plane 104, the more blurred the image observed by the viewer becomes. In this regard, the example of FIG. 1A shows a scenario where the viewing depth 105a of the eye 110 coincides with the image plane 105 and the viewer observes a sharp image, and the example of FIG. 1B shows a scenario where the viewing depth 105b is offset from the image plane 104 and thus results in a blurred image. The viewing depth may also be referred to as the accommodation distance.
[0017] The defocus blur is typically the main cause of the viewing depth of the eye 110 in a viewing situation where the eye 110 tends to adjust to the distance at which the image appears sharpest. As a result, in the configurations shown through the respective examples of FIGS. 1A and 1B, regardless of the convergence distance of the object displayed in the 3D presentation of the displayed image data, the viewing depth tends to coincide with or fall very close to the image plane 104, whereby a sharp image of the object is perceived in the scenario of FIG. 1A and a blurred image of the object is perceived in the scenario of FIG. 1B. Another factor affecting the viewing depth is binocular disparity, which mainly drives the convergence distance but also affects the viewing depth. The present disclosure uses the coupling between the viewing depth and the convergence distance to create a viewing situation where the viewing depth of the eye 110 is mainly driven by the binocular disparity, to eliminate or at least significantly reduce the retinal blur in the image displayed by the NED assembly 101, thereby aligning the viewing depth with the convergence distance, and as a result, describes a method for dealing with VAC. Alternatively, this method can be considered as an extension of the depth of field (DoF) of the NED assembly 101.
[0018] The improved volumetric NED assembly according to the present disclosure utilizes wavefront encoding to extend the DoF to provide a substantially uniform frequency response across a plurality of (virtual) image planes within the depth range of the NED assembly, thereby substantially providing an accommodation invariant (AI) NED assembly. As an example in this regard, FIG. 2A schematically shows some characteristics of the accommodation invariant NED assembly 201 according to the present disclosure, along with the viewer's eye 110. In the following, the accommodation invariant NED assembly 201 will be mainly, and for simplicity, referred to as the NED assembly 201. The NED assembly 201 includes a 2D display 202 and a lens assembly 203 disposed at a predefined distance from the 2D display 202. A pair of NED assemblies 201 can be provided as respective elements of an NED device (such as an AR headset or a VR headset) to view respective images of a stereoscopic image pair that are each displayed via a respective 2D display 202 through a respective lens assembly 203, for example, to perceive a 3D presentation of a scene captured in the stereoscopic image pair. In the following, for simplicity and clarity, the description mainly refers to the NED assembly 201 in the singular, but this description is implicitly applicable to both NED assemblies 201 of the NED device.
[0019] In the following, when referring to the position of the 2D display 202 in relation to other elements of the NED assembly 201 and various concepts related to the characteristics of the NED assembly 201, the term display plane is used to refer to the position of the 2D display 202. Similarly, the term lens plane is applied to refer to the position of the lens assembly 203 relative to other elements of the NED assembly 201 and various concepts related to the characteristics of the NED assembly 201.
[0020] The lens assembly 203 includes at least a diffractive optical element (DOE), and the lens assembly 203 is configured to provide depth of field (DoF) extension by implementing wavefront encoding that focuses the pre-processed image displayed on the 2D display 202 such that a substantially uniform frequency response is provided across the depth range of the NED assembly 201. According to an example, the lens assembly 203 can include separate optical sub-assemblies for providing imaging / enlargement and DoF extension respectively, but according to another example, the lens assembly 203 can include a single optical component configured to provide both imaging / enlargement and DoF extension. In an example where separate optical sub-assemblies are applied for imaging / enlargement and DoF extension, imaging / enlargement can be provided by, for example, a refractive lens, a Fresnel lens or a pancake lens, and DoF extension can be provided by a DOE, but in an example of a single optical component configured to provide both imaging / enlargement and DoF extension, a DOE further configured to provide imaging / enlargement can be applied.
[0021] In this regard, the illustration of FIG. 2A provides a non-limiting example in which the lens assembly 203 comprises a refractive lens 203a and a DOE 203b arranged on the same optical axis, the refractive lens 203a providing said imaging / enlargement, and the DOE 203b providing DoF extension via wavefront coding. The refractive lens 203a can include, for example, a magnifying lens, while the DOE 203b can be provided, for example, as an optical element that can be integrated with the refractive lens 203a or as an element physically separated from the refractive lens 203 and arranged at (and in proximity to) a predefined distance from the refractive lens 203a. The optical characteristics of the lens assembly 203 formed by the refractive lens 203a and the DOE 203b will provide a substantially uniform frequency response over the depth range of the NED assembly 201, which is shown in FIG. 2A via the (virtual) image planes 204-1, 204-2, 204-3, 204-4, and 204-5 corresponding to each of said adjustment depths within the depth range of the NED assembly 201. In this regard, the image planes 204-1, 204-2, 204-3, 204-4, 204-5 represent a plurality of image planes (corresponding to each of said adjustment depths), which can be collectively referred to via the reference numeral 204, while any individual image plane among the plurality of image planes 204 can be referred to via the reference numeral 204-k.
[0022] In the following, various aspects regarding the lens assembly 203 and the DOE therein are described (similarly) with reference to the non-limiting example of FIG. 2A and the DOE 203b included therein, although these aspects can be readily generalized to lens assemblies 203 and / or DOEs having characteristics different from those described for the non-limiting example of FIG. 2A. Further, although the description herein refers to a plurality of individual (virtual) image planes 204-k, this is a selection made for illustrative purposes, while in the actual implementation of the NED assembly 201, the plurality of image planes 204 can include a predefined set of individual (virtual) image planes 204-k that are each at a respective distance from the display plane, or it is worth noting that the plurality of image planes 204 can include (virtual) image planes 204-k that can exist at any distance from the display plane within the depth range of the NED assembly 201.
[0023] FIG. 2B shows a block diagram of some components of a display controller 210 applicable for preprocessing an input image for rendering on a 2D display 202 in a manner that facilitates providing a viewer with a regulation-invariant 3D presentation of an image supplied as an input to the display controller 210 when viewed through a lens assembly 203. The display controller 210 may be provided as an element of the NED assembly 201 and / or as an element of a NED device that utilizes a pair of NED assemblies 201. The display controller 210 includes a control unit 211 and an image preprocessing unit 212, and the image preprocessing unit 212 may be configured to operate at least partially under the control of the control unit 211, and the control unit 211 may be configured to control at least some aspects of the operation of the 2D display 202 regarding how an image is displayed. The operation of the control unit 211 can be at least partially controlled via, for example, a control input provided thereto via a user interface of the NED device. By way of example, the display controller 210 may be implemented via the operation of a computing device including one or more processors and one or more memories for storing one or more computer programs, and the one or more computer programs are configured to cause the computing device to operate as the display controller 210 according to the present disclosure when executed by the one or more processors.
[0024] The display controller 210, for example, the image preprocessing unit 212 therein, can receive a time-series pair of stereoscopic images and derive a corresponding time-series pair of preprocessed images based on the time-series pair of stereoscopic images, and these can be applied to provide a pose-invariant 3D presentation of the scene captured in the time-series pair of stereoscopic images by rendering them on the 2D displays 201 of the two NED assemblies 201 of the NED device. Therefore, for each of the two NED assemblies 201 of the NED device, the image preprocessing unit 212 receives the respective time-series input images I(t), and based on the respective received time-series input images I(t), the respective corresponding time-series preprocessed images I d (t) can be configured to be derived.
[0025] In this regard, any single image of the time-series input images I(t) may be referred to as the input image I, and any single image of the time-series preprocessed images I d (t) may be referred to as the preprocessed image I d . Here, the input image I refers to the pixel values of the underlying input image, while the preprocessed image I d refers to the pixel values of the underlying preprocessed image. In a non-limiting example, the input image I and the corresponding preprocessed image I d can each include an RGB image that provides respective pixel values separately for each pixel position of the respective images I, I d in the red, green, and blue color channels. In some examples, the time-series pair of stereoscopic images may be accompanied by a corresponding time-series depth map D(t), that is, each time-series pair of stereoscopic images may be accompanied by a corresponding depth map D that provides respective depth information for each pixel of the corresponding pair of stereoscopic images. The image preprocessing unit 212 is based on the time-series input images I(t) to obtain the corresponding time-series preprocessed images I dIn the process of deriving (t), the depth information received in the time-series depth map D(t) can be applied. In the following, various aspects related to the operation of the image preprocessing unit 212 of the display controller 210 are mainly described through reference to one of the two NED assemblies 210 of the NED device. Further, a single input image I is received, and based on the received input image I, the corresponding preprocessed image I d is described through reference to deriving it, but the processing of a single input image I can be easily generalized to perform the corresponding processing on a plurality of images of the respective time-series input images I(t) for each of the two NED assemblies 201 of the NED device. In this regard, regarding the aspect of deriving the preprocessed image I d based on the received input image I, it can also be considered that the received input image I is processed into the corresponding preprocessed image I d .
[0026] The image preprocessing unit 212 applies a predefined preprocessing procedure that matches the optical characteristics of the lens assembly 203, particularly the internal DOE, in order to facilitate an adjustment-invariant 3D presentation on the image plane 204-k within the depth range of the NED assembly 201 that exhibits an extended DoF compared to conventional stereo NEDs, so as to process the input image I into the corresponding preprocessed image I d . The preprocessing procedure can be regarded as encoding the input image I into the corresponding preprocessed image I d , and the lens assembly 203 can be regarded as providing an extended DoF reconstruction of the preprocessed image I d within the depth range of the NED assembly 201. In this regard, at least some characteristics of the preprocessing procedure and the diffraction characteristics of the DOE can be designed together to ensure a sufficient degree of adjustment invariance for the operation of the NED assembly 201.
[0027] The preprocessing procedure processes the input image I into the corresponding preprocessed image I dIt can be configured to process. In other words, the preprocessing procedure can be executed by the method of the display plane position recognition preprocessing, and as a result, different preprocessing characteristics can be provided for different sub-regions of the display plane. Such non-uniform processing of the sub-regions of the image regions at different positions across the display plane functions to compensate for different light transmission characteristics for each position of the display plane due to differences in the transmission characteristics through the lens assembly 203 at different positions of the lens plane. The display plane position recognition preprocessing divides, for example, the image region into a plurality of sub-regions and preprocesses the image I in a manner that depends on its position P within the image region. d It can be provided by deriving each of a plurality of sub-regions of the image region of.
[0028] In this regard, the display plane position recognition preprocessing results from the fact that the pupil 110a of the viewer's eye 110 is smaller than the lens aperture of the lens assembly 203, which has at least the following consequences:
[0029] - Only a sub-part of the light from each pixel displayed on the display plane and transmitted through the lens assembly 203 enters through the pupil 110a and encounters the retina of the eye 110;
[0030] - Only sub-parts of the lens and DOE apertures are involved in transmitting the light from a particular pixel of the display plane through the lens assembly 203 and the pupil 110a to the retina of the eye 110, and this sub-part can be called the (virtual) sub-aperture associated with each pixel.
[0031] Regarding the latter point above, the position of the sub-aperture through which light from a particular pixel of the display plane enters the retina depends on the display plane position of each pixel, the distance between the display plane and the lens assembly 203, and the (assumed) distance between the lens assembly 203 and the viewer's eye 110. In particular, the position of the sub-aperture (e.g., its center point) within the lens plane and the angle of incidence (θ, φ) of the light received through it at the pupil 110a vary with the pixel position (ξ, η) on the display plane. Thus, light from different positions on the display plane reaches the retina of the eye 110 via different paths and through different (virtual) sub-apertures of the lens aperture, thereby resulting in sub-aperture-dependent light transmission characteristics through the lens assembly 203. These light transmission characteristics can be modeled via the point spread function (PSF) of the display at the reference plane 206, which depends on the position of each sub-aperture on the lens plane (e.g., at the distance to the central axis of the lens assembly 203), but the display plane position recognition preprocessing procedure can be configured to take into account these differences in light transmission characteristics, thereby contributing to maintaining a wide field of view (FoV) without sacrificing the depth of field (DoF).
[0032] Each example of FIGS. 3A and 3B shows the sub-aperture for two different pixel positions on the display plane, with the sub-apertures shown as respective ellipses superimposed on the refractive lens 203a. In both examples, the preprocessed image I dis imaged at a reference plane 206 at a distance z from the lens plane. The example of FIG. 3A shows that light is received from a first pixel position (ξ1, η1) that lies on the central axis of the lens assembly 203. As a result, the light from the first pixel position (ξ1, η1) is imaged at the position (x1, y1) of the reference plane 206, with an incident angle of zero, and thus reaches the pupil 110a through a first sub-aperture centered on the central axis of the lens assembly 203 at the position (s1, t1) of the lens plane. The example of FIG. 3B shows that light is received from a second pixel position (ξ2, η2) that is offset from the central axis of the lens assembly 203. As a result, the light from the second pixel position (ξ2, η2) is imaged at the position (x2, y2) of the reference plane 206 and reaches the pupil 110a with a non-zero incident angle (θ2, φ2), and thus passes through a second sub-aperture that is offset from the central axis of the lens assembly 203 and centered on the position (s2, t2) of the lens plane.
[0033] Thus, the preprocessing procedure can have a shape and size that approximate a typical (e.g., average) pupil 110a and can be designed considering a plurality of sub-apertures that together cover the lens aperture without any gaps therebetween. Along the lines of the foregoing description, the sub-apertures of the lens plane are mapped, at least conceptually, to corresponding sub-regions of the display plane (e.g., corresponding sub-groups of pixels of the display plane), and the sub-regions of the display plane are further mapped to corresponding sub-regions of the image region of the preprocessing image I d . In a preprocessing procedure designed considering a plurality of sub-apertures of the lens aperture that map to corresponding sub-regions of the image region of the preprocessing image I d , the processing applied to the input image I by the preprocessing procedure to derive a particular sub-region of the preprocessing image I d can vary from sub-region to sub-region, thereby leading the preprocessing procedure to provide different preprocessing characteristics for deriving different sub-regions of the image region.
[0034] The preprocessing procedure processes the input image I to obtain the corresponding preprocessed image I, for example, by considering sub-openings of the lens opening of the lens assembly 203, as will be described in more detail below. d It can be provided via the application of at least one artificial neural network (ANN) configured (e.g., trained) to do so, thereby obtaining a preprocessing procedure for display plane position recognition. As a non-limiting example in this regard, at least one ANN used by the image preprocessing unit 212 to implement the preprocessing procedure can be provided as at least one convolutional neural network (CNN) or at least one derivative of a CNN. However, a CNN (or its derivative) functions as a non-limiting example of an applicable ANN, and in other examples, different types of ANNs can be used instead without departing from the scope of the NED assembly 201 according to the present disclosure.
[0035] According to an example of applying at least one ANN, different processes for different sub-regions of the image region of the corresponding preprocessed image I d can be provided via a single ANN trained to process each sub-region of the image region of the preprocessed image I separately according to its position within the image region. The training of such an ANN is described via the examples provided below. In the example, such an ANN can obtain the input image I and an indication of the image sub-region of interest as inputs and provide the corresponding sub-region of the preprocessed image I d as an output. In a variation of this approach, the input to the single ANN can further comprise depth information regarding the sub-region being considered. Thus, the entire preprocessed image I d can be derived by applying a single ANN individually to each of the sub-regions of the image region and combining the sub-regions thus obtained into the preprocessed image I d d d d
[0036] According to another example of applying at least one ANN, different processes for different sub-regions of the image region of the corresponding preprocessed image I d can be provided such that each is for the preprocessed image I dIt can be provided through the application of a plurality of ANNs trained to process respective sub-regions of the image region. The training of such a plurality of ANNs is illustrated by the examples provided below. In the examples, each of such ANNs takes in as input each sub-region of the input image I and can provide as output the corresponding sub-region of the pre-processed image I d In a variation of this approach, the input to each ANN can further include depth information regarding the sub-region processed by each ANN. Ultimately, by applying a plurality of ANNs to the corresponding sub-regions of the image region and combining the sub-regions thus obtained into the pre-processed image I d the entire pre-processed image I d can be derived.
[0037] In accordance with the approach described above, by way of non-limiting example, the DOE is the pre-processed image I displayed on the 2D display 202 dAn optical element may be provided that is configured to perform wavefront encoding for the purpose of providing a substantially uniform frequency response over the depth range of the NED assembly 201. In this regard, the DOE may be configured to introduce a phase delay that depends on the lens plane position into the transmitted light in the lens assembly 203. In other words, the DOE may be configured to provide different phase delays at different positions (e.g., sub-parts) of the lens plane in order to facilitate providing an adjustment-invariant 3D presentation of the input image I over the depth range of the NED assembly 201. The phase delays passing through different positions of the lens plane can be represented as Φ(s,t), and the tuple (s,t) serves to indicate 2D coordinates in the lens plane. As schematically shown for the DOE 203b in the example of FIG. 2A, the variation of the phase delay through different sub-parts of the DOE can be provided via the non-uniform thickness of the optical element. As an example, such non-uniform thickness can be implemented via a profile defined by a height map d(s,t) that defines the thickness of the DOE 203b over the position of the lens plane. In this regard, the phase delay Φ(s,t) can be converted into the height map d(s,t) taking into account the optical properties of the material applied to implement the lens element functioning as the DOE 203b, and the height map d(s,t) can be converted into the phase delay Φ(s,t).
[0038] In other examples, the DOE included in the lens assembly 203 may be different from the DOE 203b in the example of FIG. 2A. As an example in this regard, the DOE may be provided by the application of a phase modulation metamaterial configured, for example, as a metasurface, and the desired phase delay over the position of the lens aperture of the lens assembly 203 can be provided by correspondingly controlling at least one aspect (e.g., orientation and / or size) of the nanostructure of the metasurface.
[0039] According to a non - limiting example, the phase delay Φ(s,t) may be rotationally symmetric with respect to the central axis of the lens assembly 203. In the scenario according to the example of FIG. 2A, this can be achieved by applying a height map d(s,t) that is rotationally symmetric with respect to the central axis of the lens assembly 203 (and thus with respect to the central axis of the DOE 203b). This is a choice that simplifies the design of the DOE while taking into account the fact that the defocus aberration of the lens assembly 203 and the display PSF across the lens plane position is typically rotationally symmetric with respect to the central axis of the lens assembly 203. According to other examples, the phase delay Φ(s,t) may not exhibit symmetry with respect to the central axis of the lens assembly 203 (or with respect to any other axis), which typically increases the complexity of the design of the DOE (and the pre - processing procedure) while enhancing the possibility of controlling the phase delay Φ(s,t) across the position of the lens plane. Considering the DOE 203b in the framework of the example of FIG. 2A, FIG. 4 shows a non - limiting example of a height map d(s,t) that is symmetric with respect to the central axis of the lens assembly 203 and can thus be applied to define the DOE 203b as a symmetry element with respect to its central axis.
[0040] In accordance with the above - described guidelines, the phase - delay characteristics of the DOE across the position of the lens plane of the lens assembly 203 can be made to match the pre - processing characteristics of the display - plane position recognition provided via the operation of the pre - processing procedure in order to facilitate driving the depth of accommodation of the viewer's eye 110 in a desired manner. In this regard, the specific characteristics of the pre - processing procedure and the phase delay Φ(s,t) of the DOE can be designed via a common procedure to ensure such matching characteristics.
[0041] As described above, in various examples, the preprocessing procedure can be provided by the application of an ANN, but the weights of the ANN that performs the preprocessing procedure can be determined by a training procedure executed before the application of the preprocessing procedure in the image preprocessing unit 212 of the display controller 210. The match between the preprocessing characteristics provided by the preprocessing procedure and the phase delay characteristics of the DOE can be provided by modeling and optimizing the phase delay Φ(s,t) as part of the training procedure applied to determine the weights of the ANN that serves to provide the preprocessing procedure.
[0042] FIG. 5 shows a learning model 300 according to an example. The learning model 300 performs an iterative learning procedure that depends on supervised learning to determine the weights of the ANN that serves to perform the preprocessing procedure of the image processing unit 212. At the same time, it determines the phase delay of each of a plurality of positions on the lens plane (s,t) of the lens assembly 203, and thus can be applied to determine a phase delay profile that can be applied as the phase delay Φ(s,t) of the DOE. Along the lines of the approach described above, the phase delay Φ(s,t) can be applied to determine the corresponding height map d(s,t) of the DOE 203b when implemented as an element having a variable height across the DOE aperture, for example according to the example of FIG. 2A. The learning procedure can apply a training data set that typically includes several thousand images of a plurality of training images I t and can include color images (e.g., RGB images) representing natural scenes that can be selected in consideration of the intended use of the NED assembly 201 and / or the NED device that utilizes the NED assembly 201. According to an example, before starting the iterative learning procedure, the weights of the ANN and / or the phase delay profile can be initialized with (pseudo) random values, while in another example, the initial values of the weights of the ANN and / or the phase delay profile can comprise the respective values obtained by a previous learning procedure.
[0043] The training data set can include images, and each training image I tcan define each pixel value that can be provided as an input to the learning model 300 during the iterative learning procedure (as shown in FIG. 5). In some examples, the training data set includes each depth map D t (t), which can likewise be provided as an input to the learning model 300 (not shown in the figure of FIG. 5) along with the corresponding training image I t during the learning procedure.
[0044] According to an example, the learning procedure can be executed via the operation of a computing device comprising one or more processors and one or more memories for storing one or more computer programs, the one or more computer programs being configured to cause the computing device to execute the learning procedure described in this disclosure when executed by the one or more processors. According to another example, the learning procedure can be executed via the use of a plurality (i.e., two or more) computing devices of the above-described type with necessary modifications.
[0045] The learning model 300 applies an eye model 302 for processing each training image I t to the corresponding reference retinal image I r which models directly (or naturally) viewing each training image I t without considering the effect of the NED assembly 201 caused by the combination of the preprocessing procedure and the lens assembly 203. The reference retinal image I r is derived for a given accommodation state of the eye, the eye being within the depth range of the NED assembly 201 and assumed to be adjusted to a reference plane representing the conjugate plane of the retina of the viewer's eye 110. The reference retinal image I r functions as the ground truth image of the corresponding training image I t within the framework of the learning model 300. The eye model 302 can be configured to model the chromatic aberration of the (typical) optical system of the eye and the diffraction-limited resolution of the eye 110 due to the finite-sized pupil 110a, whereby the reference retinal image I ris derived as a cause of at least a part of the optical limit of the eye 110. This can be achieved, for example, by convolution of the corresponding training image with the diffraction-limited PSF for the viewer representing the chromatic aberration of the eye 110. The corresponding training image I t -derived reference retinal image I r remains the same throughout the iterative rounds, and thus the learning procedure, before entering the iterative learning procedure or only during the first iterative round of the iterative learning procedure, for each training image I of the training dataset t corresponding reference retinal image I r can be determined based on.
[0046] In each iterative round of the learning procedure, each training image I t passes through the learning model 300, which results in performing the following operations based on each training image I t respectively. -Select one or more sub-apertures of the lens plane in each iterative round and determine each one or more sub-regions of the image region of the corresponding training image I t spatially corresponding to the selected one or more sub-apertures, -Considering each position P n within the image region, using the current weights of at least one ANN 304, process the determined one or more sub-regions of each training image I t into one or more sub-regions spatially corresponding to the preprocessed image I d respectively, -The one or more sub-regions of the preprocessed image I d by the display and eye model 306 are processed into one or more sub-regions spatially corresponding to the simulated retinal image I^ r according to the current phase delay profile of the display model and considering the adjustment depth z selected for each iterative round, -By the loss function 308, the one or more sub-regions of the simulated retinal image I^ r and the reference retinal image Ir Determine the difference between one or more spatially corresponding sub-regions of r .
[0047] At the end of each iteration round, at least one weight of ANN304 and at least a part of the phase delay profile that spatially corresponds to one or more sub-apertures selected for each iteration round are based on the determined difference between each one or more sub-regions of the simulated retinal image I^ r and the spatially corresponding one or more sub-regions of the reference retinal image I r r . This can be updated. As an example in this regard, the update can be performed via the use of gradient descent. The learning procedure is until one or more predefined convergence criteria related to the difference between the simulated retinal image I^ r and the corresponding reference retinal image I r are met, and / or until a predefined number of iteration rounds are executed. When the learning procedure converges and / or a predefined number of iteration rounds are completed, the weight of at least one ANN304 at the end of the final iteration round can be adopted as the weight of at least one ANN applied to perform the preprocessing procedure in the image preprocessing unit 212 of the display controller 210, and the phase delay profile at the end of the final iteration round can be adopted as the phase delay Φ(s,t) of the DOE of the NED assembly 201.
[0048] According to an example, instead of considering the phase delay profile itself, the learning model 300 can directly consider the parameters that define the phase delay profile. In an example using a DOE characterized by material properties and thickness (e.g., DOE 203b according to the example of FIG. 2A), the parameters considered in the process of the learning procedure by the application of the learning model 300 can include a height map d(s,t) over the position of the lens plane of the lens assembly 203. In another example, the phase delay profile of an optical element functioning as a DOE can be considered via the respective coefficients of a set of predefined phase functions in an appropriate signal space, e.g., Zernike polynomials. In an example where the DOE is provided via the use of metamaterials / metasurfaces, the parameters that define the phase delay considered in the process of the learning procedure can include one or more geometric properties (such as radius and / or orientation, etc.) of the nanostructures implementing the metasurface over the position (s,t) of the lens plane. Irrespective of the method of defining the phase delay profile in the process of the learning procedure, the phase delay profile can be defined by defining respective phase delays for the position (s,t) of the lens plane, and the position (s,t) can be from a substantially uniform grid over the lens plane or the position (s,t).
[0049] As described above, in each iteration round, the learning model 300 considers one or more sub-apertures of the lens plane and one or more sub-regions of the image region of the training image I that spatially correspond to the one or more sub-apertures considered in each iteration round. t In this regard, as described above with reference to FIGS. 3A and 3B for example, each sub-aperture of the lens plane map is mapped to a corresponding sub-region of the image region (and thus to the corresponding sub-region of the display plane), and the mapping depends on the spatial relationship between the respective positions of the 2D display 202 and the lens assembly 203 and their distances from the pupil 110a. In this regard, the illustration of FIG. 5 shows the respective one or more incident angles (θ n at the pupil 110a that map to the corresponding one or more sub-regions of the image region at the position P n, φ n suggests defining one or more sub - apertures via n .
[0050] According to an example, a single sub - aperture and a spatially corresponding sub - region of the training image I t are considered in each iteration round. In this regard, for a particular iteration round of the learning procedure, the single sub - aperture of the lens plane applied to all training images I t can be (substantially) randomly selected from a plurality of predefined sub - apertures. However, in another example, the sub - aperture applied to a particular iteration round can be varied from one iteration round to another by selecting one of the plurality of predefined sub - apertures according to a predefined rule. The plurality of sub - apertures available for consideration in the learning procedure can have respective predefined positions within the lens plane, they can have predefined shapes and sizes, and this aperture approximates that of a typical (e.g., average) pupil 110a. As an example, the plurality of sub - apertures can together cover the entire lens aperture without any gaps between them. As a non - limiting example in this regard, the sub - aperture can have a rectangular or hexagonal shape, the sub - apertures may not overlap, or they may partially overlap.
[0051] According to another example, a set of two or more sub - apertures of the image region of the training image I t and corresponding one or more sub - regions can be applied in each iteration round, and the set of two or more sub - apertures applied in a particular iteration round can be (substantially) randomly selected from a plurality of predefined sub - apertures, or the set of two or more sub - apertures applied to a particular iteration round can be selected from a plurality of predefined sub - apertures according to a predefined rule. As a particular example of the latter, the set of two or more sub - apertures selected for a particular (and each) iteration round can include a plurality of sub - apertures that together cover the entire lens aperture.
[0052] Each training image I tFor at least one ANN304 in a given iteration round, the input to each ANN304 can include the pixel values of each training image I within the image sub-region considered in each iteration round, along with information identifying the image region position P t of the considered image sub-region, while the output of at least one ANN304 can include the corresponding pixel values of the corresponding image sub-region of the preprocessed image I n . The position P d of the image sub-region within the image region can be determined, for example, via the corresponding incident angles (θ, φ) in the pupil 110a (as shown in FIG. 5), via the corresponding pixel positions (ξ, η) on the display plane, or via the corresponding positions (s, t) of the lens plane. As described above, in some examples, the learning procedure can further include depth information obtained via the depth map D n (t) associated with the training image I t . Thus, the input to at least one ANN304 can further include the depth information of the image sub-region considered in each iteration round, which can be obtained from the depth map D t associated with the considered training image I t . In accordance with the foregoing guidelines, at least one ANN304 can include, for example, at least one CNN or at least one modified CNN t .
[0053] The display and eye model 306 can include a physics-based differentiable simulation model that, in combination with the eye model, simulates the transmission of the preprocessed image I d from the display plane through the lens assembly 203 to derive the corresponding simulated retinal image I^ r at the reference plane. In this regard, the simulated retinal image I^ d derived from the corresponding preprocessed image I rmodels the effects of the lens assembly 203 along with the (typical) aberrations of the eye's optical system and the diffraction-limited resolution of the eye, thereby each training image I t viewed through the NED assembly 201. In accordance with the approach described above, the modeling of the lens assembly 203 models the optical characteristics of one or more sub-apertures of the respective lens plane selected for each iteration round, taking into account the adjustment depth z selected for each iteration round, via the application of a current phase delay profile, although the eye modeling can be performed in a similar manner as described above for the eye model 302 with the necessary changes. In particular, the modeling of the optical characteristics can include modeling the display PSF according to the respective lens plane positions of the one or more sub-apertures selected for each iteration round and according to the adjustment depth z selected for each iteration round. The one or more sub-apertures to which the lens aperture is applied can be defined, for example, via the corresponding incident angles (θ, φ) in the pupil 110a (as shown in FIG. 5), via the corresponding pixel positions (ξ, η) on the display plane, or via the corresponding positions (s, t) of the lens plane that define the center points of the sub-apertures considered in each iteration round.
[0054] According to an example, the adjustment depth z applied to all training images in a particular iteration round of the learning procedure can be selected (substantially) randomly from a set of predefined adjustment depths that cover the (desired) depth range of the NED assembly 201, or from a predefined range of adjustment depths, although in another example, the adjustment depth applied to a particular iteration round can be changed from one iteration round to another by selecting the predefined adjustment depth or by selecting the adjustment depth from a predefined range according to a predefined rule.
[0055] The loss function 308 measures the difference between the simulated retinal image I^ r and the corresponding reference retinal image I r considering the two images I^r and r can be determined as an objective error metric such as L1 loss or mean squared error (MSE) derived based on the error (or difference) per pixel between. Additionally or alternatively, the difference can be determined as a perceptual error metric such as a structural similarity metric (SSIM) aimed at predicting the perceived error (or difference) between the two images I^ r and r I. According to an example, the loss function 308 can further include applying a so-called neural contrast sensitivity function (NCSF) before determining the difference between each image I^ r and r I. The use of NCSF can contribute by enabling an improved trade-off between the spatial resolution and DoF provided by the NED assembly 201.
[0056] The above description of various examples regarding the application of the learning model 300 in the iterative learning procedure refers to the processing of the overall image, but the images processed by the preprocessing procedure can include each RGB image including three separate color channels (as described above), and as a result, the training model 300 can be applied to consider the three color channels separately from each other during the iterative learning procedure.
[0057] The above example regarding the (implicit) learning procedure assumes that the operation of at least one ANN304 includes training a single ANN configured to process each sub-region of the image region of the preprocessed image I n depending on its position P d in the image region. Thus, as described above, the input to a single ANN during the learning procedure can include the pixel values of each training image I t in the image sub-region considered in each iteration round, along with the information P n identifying the image sub-region being considered, while the output of a single ANN is the preprocessed image I dcan include the corresponding pixel values of the corresponding image sub-regions (the input to a single ANN can optionally further include respective depth information).
[0058] According to other examples, the operation of at least one ANN 304 can include training a plurality of ANNs each configured to process a corresponding sub-region of an image region of a pre-processed image I d can include training a plurality of ANNs each configured to process a corresponding sub-region of an image region of a pre-processed image I. Thus, in each iteration, one or more ANNs corresponding to each of the one or more image sub-regions considered in each iteration are trained and / or updated. In accordance with the foregoing approach, during the learning procedure, the input to each ANN is the respective training image I within each of the respective image sub-regions considered in each iteration round t can include the pixel values thereof, but the output of each ANN can include the corresponding pixel values of the corresponding image sub-regions of the pre-processed image I d can include the corresponding pixel values of the corresponding image sub-regions (the input to each ANN can optionally further include respective depth information).
[0059] FIG. 6 shows a block diagram of some components of a device 400 that can be used to implement at least some aspects of the display controller 210 or to perform the iterative learning procedure described above. Although described herein with reference to a single device 400, at least some aspects of the display controller 210 or the foregoing iterative learning procedure can be implemented by the coordinated operation of two or more devices 400 configured to provide cloud-based computing services.
[0060] Device 400 includes a processor 410 and a memory 420. The memory 420 can store data and computer program code 425. Device 400 can be further configured with the processor 410 and a part of the computer program code 425 to provide a user interface for receiving input from a user and / or providing output to the user, and can further include communication means 430 for wired or wireless communication with other devices and / or user I / O (input / output) components 440. In particular, the user I / O components can include user input means such as one or more keys or buttons, a keyboard, a touch screen or a touch pad. The user I / O components can include output means such as a display or a touch screen. The components of device 400 are communicatively coupled to each other via a bus 450 that enables transfer of data and control information between the components.
[0061] A part of the memory 420 and the computer program code 425 stored therein can be further configured to operate device 400 as a display controller 210 using the processor 410, or (if applicable) to execute the iterative learning procedure described above. The processor 410 is configured to read from and write to the memory 420. Although the processor 410 is shown as a single component each, it can be implemented as one or more separate processing components each. Similarly, the memory 420 is shown as a single component each, but can be implemented as one or more separate components each, some or all of which may be integrated / removable and / or can provide persistent / semi-persistent / dynamic / cache storage.
[0062] The computer program code 425 can include computer-executable instructions that implement at least some aspects of the display controller 210 or, when loaded into the processor 410 (if applicable), execute the aforementioned iterative learning procedure. By way of example, the computer program code 425 can include a computer program consisting of one or more sequences of one or more instructions. The processor 410 can load and execute the computer program by reading the one or more sequences of one or more instructions contained therein from the memory 420. When executed by the processor 410, the one or more sequences of one or more instructions can be configured to operate the device 400 as a display controller 210 or, if applicable, execute the iterative learning procedure described above. Accordingly, the device 400 can comprise at least one processor 410 and at least one memory 420 including computer program code 425 for one or more programs, and the at least one memory 420 and the computer program code 425 are configured to cause the device 400 to execute at least some aspects of the display controller 210 or the aforementioned iterative learning procedure (if applicable) using the at least one processor 410.
[0063] The computer program code 425 can be provided, for example, as a computer program product including at least one computer-readable non-transitory medium having the computer program code 425 stored thereon, which, when executed by the processor 410, causes the device 400 to execute at least some aspects of the display controller 210 or, if applicable, the aforementioned iterative learning procedure. The computer-readable non-transitory medium can include a memory device or a recording medium that tangibly embodies the computer program. As another example, the computer program can be provided as a signal configured to reliably transfer the computer program.
[0064] References to processors in this specification should not be understood to include only programmable processors, but also include dedicated circuits such as field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and signal processors. The features described in the foregoing description may be used in combinations other than those explicitly described.
Claims
1. A stereoscopic near-eye display (NED) assembly (201) for a NED device, comprising a pair of NED assemblies (201), the NED assemblies (201) comprising: a two-dimensional (2D) display (202) for rendering an image for viewing by a viewer's eye (110); a lens assembly (203) arranged at a predefined distance from the 2D display (202) to enable viewing of the image rendered via the 2D display (202), the lens assembly (203) comprising a diffractive optical element (DOE) (203b) configured to provide different phase retardations through multiple positions of its aperture; and Due to different transmission characteristics through different sub-apertures of the lens aperture of the lens assembly (203), the pre-processed image (I d ) on the basis of the received image (I) through application of a pre-processing procedure configured to apply an image region location dependent pre-processing in the derivation of different sub-regions of the image region of d ) and outputs the pre-processed image (I) for rendering via the 2D display (202). d a display controller (210) comprising a pre-processing unit (212) configured to provide The pre-processing procedure and the lens assembly (203) are adapted to process the pre-processed image (I) as perceived as sharp for a number of pre-defined accommodation depths that exist within an extended depth of field (DoF) of the NED assembly (201) when viewed through the lens assembly (203). d ), and through the operation of the NED device, the pre-processed image (I d ) to facilitate providing an accommodation invariant 3D representation.
2. The NED assembly (201) of claim 1, wherein the pre-processing section (212) comprises at least one artificial neural network (ANN) configured to perform the pre-processing procedure.
3. The NED assembly (201) of claim 2, wherein the at least one ANN comprises at least one convolutional neural network (CNN).
4. The at least one ANN processes the preprocessed image (I d 4. The NED assembly (201) of claim 1, further comprising a single ANN trained to process each of a plurality of sub-regions of the image region of said image.
5. The at least one ANN includes a plurality of ANNs, each of which is adapted to process the pre-processed image (I d 4. The NED assembly (201) of claim 1, wherein the NED assembly (201) is trained to process each one of a plurality of sub-regions of the image region of said image region.
6. 6. The NED assembly (201) of claim 1, wherein the DOE (203b) includes an optical element having thicknesses that are determined separately for the multiple positions of its aperture to provide different phase delays through the multiple positions.
7. 2. A pre-processing procedure and an apparatus for deriving a phase delay profile for a stereoscopic near-eye display (NED) assembly (201) according to claim 1, said apparatus comprising: t and applying each learning model to jointly derive at least one artificial neural network (ANN) (304) that serves as the phase delay profile defining respective phase delays for a plurality of positions of the DOE (203 b) via an iterative learning procedure based on Each of the plurality of training images is Selecting one or more subapertures for each iteration round, and dividing each of the training images (I t determining one or more sub-regions of each of the image regions of The at least one ANN (304) selects each of the training images (I t ) to process the determined sub-region or sub-regions of the image region to determine their respective positions (P n ) and using the current weights of the at least one ANN (304), a preprocessed image (I d determining one or more spatially corresponding sub-regions of the The display and eye model (306) generates the pre-processed image (I d ) according to the current phase delay profile and taking into account the accommodation depth (z) selected for the respective iteration round, to generate a simulated retinal image (I r ) into one or more spatially corresponding sub-regions of the image; The simulated retinal image (Î) is computed by a predefined loss function (308). r ) and the corresponding reference retinal image (I r determining a difference between one or more spatially corresponding sub-regions of the image; The plurality of training images (I t updating the weights of the at least one ANN (304) and a portion of the phase delay profile that spatially corresponds to the one or more selected subapertures based on the respective differences determined for An apparatus configured to perform the steps of:
8. The apparatus of claim 7 , configured to update the weights and the phase delay profile of the at least one ANN using a gradient descent method.
9. The corresponding reference retinal image (I r ) and an eye model (302) representing one or more optical limitations of the eye is applied to the respective training images (I t 9. The apparatus of claim 7 or 8, configured to derive the signal by applying
10. below, selecting one of a plurality of predefined adjustment depths as the adjustment depth (z) for the respective iteration round; selecting the adjustment depth (z) for each iteration round from among a predefined range of adjustment depths; 10. The apparatus of claim 7, configured to perform one of the following:
11. configured to select one of a plurality of predefined subapertures for each iteration round; each predefined subaperture having a respective position within the lens aperture and a respective predefined shape and size; 11. An apparatus according to claim 7, wherein the subapertures together cover an entire aperture of the DOE.
12. configured to select two or more of a plurality of predefined subapertures for each iteration round; each predefined subaperture having a respective position within the lens aperture and a respective predefined shape and size; 11. An apparatus according to claim 7, wherein the sub-apertures together entirely cover an aperture of the DOE.
13. The display and eye model (306) performs a phase delay on the pre-processed image (I) by applying a phase delay according to the current phase delay profile defined for the one or more sub-apertures selected for the respective iteration round. d ) based on the one or more sub-regions of the pre-processed retinal image (I r 13. The apparatus of claim 7, configured to derive the spatially corresponding one or more sub-regions of the aperture (a) of the aperture array (b) and to model optical properties of the one or more sub-apertures selected for the respective iteration round taking into account the accommodation depth (z) selected for the respective iteration round.
14. 14. The apparatus of claim 13, wherein applying the optical characteristics comprises applying a point spread function, PSF, selected according to the respective positions of the one or more subapertures selected for the respective iteration round, taking into account the accommodation depth (z) selected for the respective iteration round.
15. 15. The apparatus of claim 7, wherein the at least one ANN comprises at least one Convolutional Neural Network (CNN).
16. The at least one ANN (304) computes the preprocessed image (I d 16. The apparatus of claim 7, further comprising a single ANN trained to process each of a plurality of sub-regions of the image region of said first image.
17. The at least one ANN (304) comprises a plurality of ANNs, each of which is adapted to process the pre-processed image (I d 16. The apparatus of claim 7, wherein the apparatus is trained to process each of a plurality of sub-regions of the image region of said plurality of pixels.
18. 2. A method for deriving a pre-processing procedure and a phase delay profile for a stereoscopic near-eye display (NED) assembly (201) according to claim 1, said method comprising: applying respective learning models to derive the pre-processing procedure and at least one artificial neural network, ANN, (304) that serves as the phase delay profile defining respective phase delays for a plurality of positions of the DOE (203b) based on a plurality of training images (I). t ), the method comprising, for a number of iteration rounds, performing an iterative learning procedure based on Each of the plurality of training images is subjected to the following: Selecting one or more subapertures for each iteration round, and dividing each of the training images (I t determining one or more sub-regions of each of the image regions of The at least one ANN (304) recognizes each of the training images (I t ) to process the determined sub-region or sub-regions of the image region to determine their respective positions (P n ) and using the current weights of the at least one ANN (304), a preprocessed image (I d determining one or more spatially corresponding sub-regions of the The pre-processed image (I d ) according to the current phase delay profile and taking into account the accommodation depth (z) selected for the respective iteration round, to generate a simulated retinal image (I r ) into one or more spatially corresponding sub-regions of the image; The simulated retinal image (Î) is computed by a predefined loss function (308). r ) and the corresponding reference retinal image (I r ) through a learning model (300) that includes determining differences between one or more spatially corresponding sub-regions of the image; The plurality of training images (I t updating the weights of the at least one ANN (304) and a portion of the phase delay profile that spatially corresponds to the one or more selected subapertures based on the respective differences determined for The method includes:
19. A computer program (425) comprising computer readable program code configured to cause the computer program (425) to perform the method of claim 18 when executed on one or more computing devices (400).
Citation Information
Patent Citations
Near-to-eye display module based on orthogonal characteristic pixel blocks
CN115128811A
Image-enhanced depth sensing using machine learning
JP2021517685A
Optical approach to overcoming vergence-accommodation conflict
US20190258054A1
Enhanced eye tracking techniques based on neural network analysis of images
WO2021247435A1