Adjustable near-eye display
The stereoscopic NED assembly with a DOE and ANN-based preprocessing addresses VAC, ensuring sharp images across various depths for a comfortable and high-quality 3D experience.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- TAMPERE UNIV FOUND SR
- Filing Date
- 2024-04-23
- Publication Date
- 2026-06-04
AI Technical Summary
Conventional near-eye displays (NEDs) cause visual discomfort due to congestion-accommodation conflict (VAC) when the display objects' distances differ from the accommodation distance, leading to blurred images and compromised image quality.
A stereoscopic NED assembly with a diffractive optical element (DOE) and a preprocessing procedure using artificial neural networks (ANNs) to extend the depth of field (DoF) and align accommodation depth with convergence distance, ensuring sharp images across a range of depths.
The solution provides a comfortable, high-quality 3D viewing experience by reducing retinal blurring and extending the depth of field, aligning accommodation depth with convergence distance.
Smart Images

Figure 0007870362000001 
Figure 0007870362000002 
Figure 0007870362000003
Abstract
Description
[Technical Field]
[0001] This invention relates to a near-eye display (NED). [Background technology]
[0002] Stereoscopic near-eye displays (NEDs) have particular applications in various virtual reality (VR) and augmented reality (AR) scenarios. Due to visual discomfort experienced by some users, challenges arise in ensuring a comfortable, natural, or near-natural viewing experience through the use of wearable NED devices.
[0003] One of the key challenges in ensuring a comfortable viewing experience involves addressing congestion-accommodation competition (VAC), which can occur when NED equipment is applied to display objects in a way that its congestion distance does not match its accommodation distance. When viewing VR or AR content using conventional NED equipment, such situations frequently occur due to displaying objects in a three-dimensional (3D) image that have their respective positions in image space at distances different from the distance to the (virtual) image plane of the NED equipment (i.e., accommodation distance) (i.e., at their respective congestion distances).
[0004] NED devices equipped with technologies to address VAC, such as variable focus, multifocal, field of view, holographic, and Maxwellian systems, are known in the art, but they are all subject to their own trade-offs, such as a limited eyebox, limited image quality (e.g., low spatial resolution, speckle noise), and / or complex device requirements (e.g., bulky and / or high-speed optics, precise eye-tracking devices). Consequently, there is a continuing demand for NED devices that can mitigate VAC in a way that provides high user comfort without compromising image quality and device complexity. [Overview of the Initiative]
[0005] The object of the present invention is to provide a technology for reducing or eliminating user discomfort arising from congestion regulation competition (VAC) within NED equipment without significantly degrading the resulting perceived image quality.
[0006] According to an exemplary embodiment, a stereoscopic near-eye display (NED) assembly for an NED device is provided, comprising a pair of NED assemblies, wherein the NED assembly comprises a two-dimensional (2D) display for rendering an image for viewing by the viewer's eye, a lens assembly positioned at a predefined distance from the 2D display to enable viewing the rendered image through the 2D display, the lens assembly comprising a diffractive optical element (DOE) configured to provide different phase delays through a plurality of positions of its aperture, and causing different transmission characteristics through different sub-apers of the lens aperture of the lens assembly, A display controller comprising a preprocessing unit configured to derive a preprocessed image based on a received image by applying a preprocessing procedure configured to apply image region position-dependent preprocessing in the derivation of different sub-regions of an image region of a preprocessed image, and to supply the preprocessed image for rendering via a 2D display, wherein the preprocessing procedure and the lens assembly are configured to display the preprocessed image as perceived to be sharp for a plurality of predefined accommodative depths that exist within the extended depth of field (DoF) of the NED assembly when viewed through the lens assembly, thereby facilitating the provision of accommodative-invariant 3D presentation based on the preprocessed image via the operation of the NED device.
[0007] According to another exemplary embodiment, a preprocessing procedure and apparatus for deriving a phase delay profile for a stereoscopic NED assembly according to the above exemplary embodiment are provided, wherein the apparatus is configured to apply respective learning models to derive together the preprocessing procedure and at least one artificial neural network (ANN) which serves as the phase delay profile that defines the respective phase delays for a plurality of positions of the DOE via an iterative learning procedure based on a plurality of training images, wherein the apparatus selects one or more subapertures for each of the plurality of training images in a plurality of iterative rounds, determines one or more sub-regions for each of the image regions of the respective training images that spatially correspond to the selected one or more subapertures, and processes the determined one or more sub-regions of the respective training images by the at least one ANN to determine the respective positions within the image region Taking this into consideration, the learning model is configured to perform the following: determine one or more spatially corresponding subregions of the preprocessed image using the current weights of the at least one ANN; process the one or more subregions of the preprocessed image into one or more spatially corresponding subregions of the simulated retinal image according to the current phase delay profile and taking into consideration the selected accommodation depth for each iteration round by the display and eye model; and process the one or more subregions of the simulated retinal image into one or more spatially corresponding subregions of the corresponding reference retinal image by a predefined loss function; and update the weights of the at least one ANN and the portion of the phase delay profile spatially corresponding to the one or more selected subapex based on the respective differences determined for the plurality of training images.
[0008] According to another exemplary embodiment, a method is provided for deriving a preprocessing procedure and a phase delay profile for a stereoscopic NED assembly according to the above exemplary embodiment, the method comprising performing an iterative learning procedure based on a plurality of training images by applying a respective learning model to derive together the preprocessing procedure and at least one ANN that serves as the phase delay profile, which defines the respective phase delays for a plurality of positions of the DOE, the method comprising: for each of the plurality of training images for a plurality of iteration rounds, selecting one or more subapertures for each of the respective iteration rounds; determining one or more subregions for each of the image regions of the respective training image that spatially correspond to the selected one or more subapertures; and processing the determined one or more subregions of the respective training image with the at least one ANN to determine the respective positions within the image region The process includes: determining one or more spatially corresponding subregions of a preprocessed image using the current weights of the at least one ANN, processing the one or more subregions of the preprocessed image into one or more spatially corresponding subregions of a simulated retinal image according to the current phase delay profile and taking into account the selected accommodation depth for each iteration round, and processing them through a learning model that includes the steps of: determining the difference between the one or more subregions of the simulated retinal image and the one or more spatially corresponding subregions of a corresponding reference retinal image by a predefined loss function; and updating the weights of the at least one ANN and the portion of the phase delay profile spatially corresponding to the one or more selected subapex based on the respective differences determined for the plurality of training images.
[0009] According to another exemplary embodiment, a computer program is provided, which includes computer-readable program code configured to cause the method according to the above exemplary embodiment to be performed when the program code is executed on one or more computing devices.
[0010] The computer program according to the exemplary embodiment described above can be embodied as a computer program product comprising, for example, at least one computer-readable non-temporary medium on which program code is stored on a volatile or non-volatile computer-readable recording medium, and when executed by one or more computing devices, causes the computing devices to perform at least the method according to the exemplary embodiment described above.
[0011] The exemplary embodiments of the invention presented in this patent application should not be construed as limiting the applicability of the appended claims. The verb “to provide” and its derivatives are used in this patent application as an open limitation, without prejudice to the existence of features not listed. The features described below may be freely combined with each other unless otherwise specified.
[0012] Some features of the present invention are described in the appended claims. However, aspects of the present invention, along with their additional objectives and advantages, will be best understood from the following description of some exemplary embodiments, when read in conjunction with the appended drawings, with respect to both their structure and their method of operation.
[0013] Embodiments of the present invention are shown in the accompanying drawings for illustrative purposes only and not for limiting purposes. [Brief explanation of the drawing]
[0014] [Figure 1A] Several embodiments of stereoscopic near-eye display assemblies are schematically shown as examples. [Figure 1B] Several embodiments of stereoscopic near-eye display assemblies are schematically shown as examples. [Figure 2A] Several embodiments of stereoscopic near-eye display assemblies are schematically shown as examples. [Figure 2B] A block diagram of some of the components of a near-eye display assembly is shown. [Figure 3A] This diagram schematically illustrates how light is received by the eye from the pixels of a display through a sub-aperture in the lens plane. [Figure 3B] This diagram schematically illustrates how light is received by the eye from the pixels of a display through a sub-aperture in the lens plane. [Figure 4] An example shows a map of the heights of dispersed optical elements. [Figure 5] An example block diagram of some components of a learning model is shown. [Figure 6] An example block diagram of some components of the device is shown. [Modes for carrying out the invention]
[0015] Figures 1A and 1B schematically illustrate some characteristics of a conventional stereoscopic near-eye display (NED) assembly 101 along with the viewer's eye 110. The NED assembly 101 comprises a two-dimensional (2D) display 102 and a magnifying lens 103, which are positioned in front of the viewer's eye 110 when the NED assembly 101 is operating to view an image rendered on the 2D display 102 via the magnifying lens 103. The 2D display 102 is positioned at a fixed position relative to the magnifying lens 103, and the position of the 2D display 102 is closer to the magnifying lens 103 than its focal length. As a result, the 2D display 102 is mapped to a (virtual) image plane 104 positioned at a fixed distance behind the 2D display map 102. The NED device may comprise a pair of NED assemblies 101 (i.e., each NED assembly 101 for each of the viewer's eyes 110) and a display controller configured to supply each image of the stereoscopic image pair for rendering onto each of the pair of NED assemblies 101 on a 2D display 102 in order to provide a three-dimensional (3D) presentation of the scene captured by the stereoscopic image pair.
[0016] Theoretically, assuming an aberration-free optical system, the visual content displayed on the 2D display 102 can be shown with a resolution up to the diffraction limit in the image plane 104. At depths offset from the image plane 104 (i.e., distance from the 2D display 102), a sharp decrease in the frequency response is observed with increasing depth due to defocus blur (or defocus aberration). Therefore, the viewer observes a sharp image with high resolution if the accommodation depth of the eye 110 coincides with the image plane 104, but the greater the offset between the accommodation depth of the eye 110 and the image plane 104, the more blurred the image observed by the viewer becomes. In this regard, the example in Figure 1A shows a scenario in which the accommodation depth 105a of the eye 110 coincides with the image plane 105 and the viewer observes a sharp image, while the example in Figure 1B shows a scenario in which the accommodation depth 105b is offset from the image plane 104 and therefore results in a blurred image. The accommodation depth may also be referred to as the accommodation distance.
[0017] The aforementioned defocus blur is the primary driver of the accommodation depth of the eye 110 in viewing situations where the eye 110 tends to adjust to the distance at which the image appears sharpest. As a result, in the configurations shown through the respective examples in Figures 1A and 1B, the accommodation depth tends to coincide with or be very close to the image plane 104, regardless of the convergence distance of the object displayed in the 3D presentation of the displayed image data, thereby resulting in a sharp image of the object being perceived in the scenario of Figure 1A and a blurred image of the object being perceived in the scenario of Figure 1B. Another factor influencing the accommodation depth is binocular disparity, which primarily drives the convergence distance but also affects the accommodation depth. This disclosure describes a technique for addressing VAC by using a coupling between the accommodation depth and the convergence distance to create a viewing situation in which the accommodation depth of the eye 110 is primarily driven by the binocular parallax, thereby eliminating or at least significantly reducing retinal blurring in the image displayed by the NED assembly 101, thereby aligning the accommodation depth with the convergence distance, and consequently addressing VAC. Alternatively, this technique can be considered an extension of the depth of field (DoF) of the NED assembly 101.
[0018] The improved stereoscopic NED assembly according to this disclosure utilizes wavefront coding to extend the DoF to yield a substantially uniform frequency response across multiple (virtual) image planes within the depth range of the NED assembly, thereby substantially providing an accommodative-invariant (AI) NED assembly. As an example in this regard, Figure 2A schematically illustrates some characteristics of the accommodative-invariant NED assembly 201 according to this disclosure, along with the viewer's eye 110. Hereinafter, the accommodative-invariant NED assembly 201 will mainly be referred to simply as the NED assembly 201. The NED assembly 201 comprises a 2D display 202 and a lens assembly 203 positioned at a predefined distance from the 2D display 202. A pair of NED assemblies 201 may be provided as respective elements of an NED device (such as an AR headset or a VR headset) that allows each of the two NED assemblies 201 to be positioned in front of each of the viewer's eyes 110 to view each image of the stereoscopic image pair, which is displayed via each of the 2D displays 202 via the respective lens assemblies 203, in order to perceive a 3D presentation of a scene captured in a stereoscopic image pair. Hereafter, for the sake of brevity and clarity, the description will primarily refer to the NED assembly 201 in the singular, although this description implicitly applies to both NED assemblies 201 of the NED device.
[0019] In the following, the term "display plane" is used to refer to the position of the 2D display 202 in relation to other elements of the NED assembly 201 and various concepts related to the characteristics of the NED assembly 201. Similarly, the term "lens plane" is used to refer to the position of the lens assembly 203 relative to other elements of the NED assembly 201 and various concepts related to the characteristics of the NED assembly 201.
[0020] The lens assembly 203 comprises at least a diffractive optical element (DOE), and is configured to provide DoF extension by implementing wavefront coding that focuses a preprocessed image displayed on the 2D display 202 such that a substantially uniform frequency response is provided over the depth range of the NED assembly 201. By example, the lens assembly 203 may include separate optical subassemblies for providing imaging / magnification and DoF extension, respectively, but by another example, the lens assembly 203 may include a single optical component configured to provide both imaging / magnification and DoF extension. In the example where separate optical subassemblies are applied for imaging / magnification and DoF extension, imaging / magnification may be provided by, for example, a refractive lens, a Fresnel lens, or a pancake lens, and DoF extension may be provided by a DOE, but in the example of a single optical component configured to provide both imaging / magnification and DoF extension, a DOE further configured to provide imaging / magnification may be applied.
[0021] In this regard, the illustration in Figure 2A provides a non-limiting example comprising a refractive lens 203a and a DOE 203b arranged on the same optical axis, wherein the refractive lens 203a provides the imaging / magnification and the DOE 203b provides DoF extension via waveform coding. The refractive lens 203a may include, for example, a magnifying lens, while the DOE 203b may be provided, for example, as an optical element that can be integrated with the refractive lens 203a, or as an element that is physically separated from the refractive lens 203 and positioned at (and close to) a predefined distance from the refractive lens 203a. The optical properties of the lens assembly 203 formed by the refractive lenses 203a and DOE 203b provide a substantially uniform frequency response over the depth range of the NED assembly 201, which is shown in Figure 2A via (virtual) image planes 204-1, 204-2, 204-3, 204-4, and 204-5 corresponding to each of the aforementioned adjustment depths within the depth range of the NED assembly 201. In this regard, image planes 204-1, 204-2, 204-3, 204-4, and 204-5 represent a plurality of image planes (corresponding to each of the plurality of aforementioned adjustment depths), which may be referred to together via reference numeral 204, but any individual image plane among the plurality of image planes 204 may be referred to via reference numeral 204-k.
[0022] Hereinafter, various embodiments of the lens assembly 203 and the DOE therein will be described with reference to the non-limiting example in Figure 2A and the DOE 203b contained therein, although these embodiments can be readily generalized to lens assemblies 203 and / or DOEs having different characteristics from those described in the non-limiting example in Figure 2A. Furthermore, it is worth noting that while the description herein refers to a plurality of separate (virtual) image planes 204-k, this is a selection made for illustrative purposes, whereas in actual implementations of the NED assembly 201, the plurality of image planes 204 may include a predefined set of separate (virtual) image planes 204-k at respective distances from the display plane, or the plurality of image planes 204 may include (virtual) image planes 204-k that may exist at any distance from the display plane within the depth range of the NED assembly 201.
[0023] Figure 2B shows a block diagram of some components of a display controller 210 applicable to preprocessing an input image for rendering on a 2D display 202 in a manner that facilitates providing the viewer with an accommodatively invariant 3D presentation of the image supplied as input to the display controller 210 when viewed through the lens assembly 203. The display controller 210 may be provided as an element of an NED assembly 201 and / or as an element of an NED apparatus utilizing a pair of NED assemblies 201. The display controller 210 comprises a control unit 211 and an image preprocessing unit 212, the image preprocessing unit 212 may be configured to operate at least partially under the control of the control unit 211, and the control unit 211 may be configured to control at least some aspects of the operation of the 2D display 202 in terms of how the image is displayed. The operation of the control unit 211 may be controlled at least partially through control inputs provided therein, for example, through a user interface of the NED apparatus. For example, the display controller 210 may be implemented through the operation of a computing device comprising one or more processors and one or more memories for storing one or more computer programs, the one or more computer programs configured to cause the computing device to operate as the display controller 210 according to the Disclosure when executed by the one or more processors.
[0024] The display controller 210, for example, the image preprocessing unit 212 therein, can receive a time-series pair of stereoscopic images and derive a corresponding time-series preprocessed image pair based on the time-series pair of stereoscopic images, which can be applied to provide an adjustable-invariant 3D presentation of the scene captured in the time-series pair of stereoscopic images by rendering them on the respective 2D displays 201 of the two NED assemblies 201 of the NED apparatus. Thus, for each of the two NED assemblies 201 of the NED apparatus, the image preprocessing unit 212 receives the respective time-series input image I(t) and, based on the respective received time-series input image I(t), derives the respective corresponding time-series preprocessed image I(t) which is displayed via the 2D display 202 of the respective NED assembly 201 to provide the presentation of the 3D scene. d It can be configured to derive (t).
[0025] In this regard, any single image in the time-series input image I(t) is sometimes referred to as the input image I, and the time-series preprocessed image I d (t) Any single image is preprocessed into image I d This is sometimes referred to as the preprocessing image I. Here, input image I refers to the pixel values of the underlying input image, while preprocessing image I refers to the preprocessing image I. d This refers to the pixel values of the underlying preprocessed image. In a non-limiting example, the input image I and the corresponding preprocessed image I d These are images I and I respectively. d This may include separate RGB images that provide separate pixel values for each pixel position in the red, green, and blue color channels. In some examples, the time-series stereoscopic image pairs may be accompanied by a corresponding time-series depth map D(t), that is, each time-series stereoscopic image pair may be accompanied by a corresponding depth map D that provides depth information for each pixel of the corresponding stereoscopic image pair. The image preprocessing unit 212 processes the corresponding time-series preprocessed image I based on the time-series input image I(t). dIn the process of deriving (t), the depth information received in the time-series depth map D(t) can be applied. In the following, various aspects related to the operation of the image pre-processing unit 212 of the display controller 210 are mainly described through reference to one of the two NED assemblies 210 of the NED device. Further, a single input image I is received, and a corresponding pre-processed image I d is described through reference to deriving it, but the processing of a single input image I can be easily generalized to perform the corresponding processing on a plurality of images of respective time-series input images I(t) for each of the two NED assemblies 201 of the NED device. In this regard, regarding the aspect of deriving the pre-processed image I d based on the received input image I, it can also be considered that the received input image I is processed into the corresponding pre-processed image I d .
[0026] The image pre-processing unit 212 applies a pre-defined pre-processing procedure that matches the optical characteristics of the lens assembly 203, particularly the internal DOE, in order to facilitate adjusted-invariant 3D presentation on the image plane 204-k within the depth range of the NED assembly 201 that exhibits an extended DoF compared to conventional stereo NED, thereby processing the input image I into the corresponding pre-processed image I d . The pre-processing procedure can be regarded as encoding the input image I into the corresponding pre-processed image I d , and the lens assembly 203 can be regarded as providing an extended DoF reconstruction of the pre-processed image I d within the depth range of the NED assembly 201. In this regard, at least some characteristics of the pre-processing procedure and the diffraction characteristics of the DOE can be designed together to ensure a sufficient degree of adjustment invariance for the operation of the NED assembly 201.
[0027] The pre-processing procedure processes the input image I into the corresponding pre-processed image I dIt can be configured to process in such a way. In other words, the preprocessing procedure can be performed in the manner of the display plane position recognition preprocessing, and as a result, different preprocessing characteristics can be provided for different subregions of the display plane. Such non-uniform processing of subregions of the image region at different positions across the display plane serves to compensate for different light transmission characteristics at each position on the display plane due to differences in transmission characteristics through the lens assembly 203 at different positions on the lens plane. The display plane position recognition preprocessing can, for example, divide the image region into a plurality of subregions and preprocess the image I in a manner that relies on its position P within the image region. d This can be provided by deriving each of the multiple sub-regions of the image region.
[0028] In this regard, the display plane position recognition preprocessing arises from the fact that the pupil 110a of the viewer's eye 110 is smaller than the lens aperture of the lens assembly 203, which has at least the following result:
[0029] -Only a sub-portion of light from each pixel displayed on the display plane and transmitted through the lens assembly 203 enters the pupil 110a and encounters the retina of the eye 110;
[0030] -Only the lens and sub-parts of the DOE aperture are involved in transmitting light from a particular pixel of the display plane to the retina of the eye 110 through the lens assembly 203 and the pupil 110a, and these sub-parts can be called (virtual) sub-apertures associated with each pixel.
[0031] Regarding the latter point mentioned above, the position of the sub-aperture through which light from a particular pixel on the display plane enters the retina depends on the display plane position of each pixel, the distance between the display plane and the lens assembly 203, and the (assumed) distance between the lens assembly 203 and the viewer's eye 110. In particular, the position of the sub-aperture in the lens plane (e.g., its center point) and the angle of incidence (θ,φ) of the light received through it by the pupil 110a vary with the pixel position (ξ,η) on the display plane. Thus, light from different positions on the display plane reaches the retina of the eye 110 via different paths and through different (virtual) sub-apertures of the lens aperture, resulting in sub-aperture-dependent light transmission characteristics through the lens assembly 203. These light transmission characteristics can be modeled via a point image distribution function (PSF) of the display on a reference plane 206, which depends on the position of each sub-aperture on the lens plane (for example, at a distance to the central axis of the lens assembly 203). However, the display plane position recognition preprocessing procedure can be configured to take into account these differences in light transmission characteristics, thereby contributing to maintaining a wide field of view (FoV) without compromising the degree of view (DoF).
[0032] The examples in Figures 3A and 3B show the respective sub-apertures at two different pixel positions on the display plane, with the sub-apertures shown as ellipses superimposed on the refractive lens 203a. In both examples, preprocessed image I dThe light is imaged at a reference plane 206 at a distance z from the lens plane. An example in Figure 3A shows that light is received from a first pixel position (ξ1,η1) located on the central axis of the lens assembly 203. As a result, the light from the first pixel position (ξ1,η1) is imaged at position (x1,y1) on the reference plane 206, with an angle of incidence of zero, and therefore passes through a first sub-aperture centered on the central axis of the lens assembly 203 at position (s1,t1) on the lens plane, and reaches the pupil 110a. The example in Figure 3B shows that light is received from a second pixel position (ξ2,η2) offset from the central axis of the lens assembly 203. As a result, the light from the second pixel position (ξ2,η2) is imaged at position (x2,y2) on the reference plane 206, reaches the pupil 110a at a non-zero angle of incidence (θ2,φ2), and therefore passes through a second sub-aperture offset from the central axis of the lens assembly 203 and centered at position (s2,t2) on the lens plane.
[0033] Therefore, the preprocessing procedure can have a shape and size that approximates a typical (e.g., average) pupil 110a, and can be designed considering multiple sub-apertures that completely cover the lens aperture together without any gaps between them. In line with the preceding explanation, the sub-apertures of the lens plane are mapped, at least conceptually, to corresponding sub-regions of the display plane (e.g., corresponding subgroups of pixels of the display plane), and the sub-regions of the display plane are mapped to the preprocessed image I d The image region is further mapped to the corresponding sub-region. Preprocessed image I d In a preprocessing procedure designed to consider multiple sub-apers of a lens aperture that map to corresponding sub-regions of the image region, the preprocessed image I d The processing applied to the input image I by the preprocessing procedure to derive a specific subregion may differ for each subregion, thereby providing different preprocessing characteristics for the preprocessing procedure to derive different subregions of the image region.
[0034] The preprocessing procedure involves processing the input image I to take into account the sub-apertures of the lens aperture of the lens assembly 203, as will be described in more detail below, to obtain the corresponding preprocessed image I. d This can be provided through the application of at least one artificial neural network (ANN) configured (e.g., trained) to obtain a display plane position recognition preprocessing procedure. In a non-limiting example in this regard, at least one ANN used by the image preprocessing unit 212 to carry out the preprocessing procedure may be provided as at least one convolutional neural network (CNN) or at least one derivative of a CNN. However, a CNN (or its derivative) serves as a non-limiting example of an applicable ANN, and in other examples, different types of ANNs may be used instead without departing from the scope of the NED assembly 201 by this disclosure.
[0035] According to an example of applying at least one ANN, the corresponding preprocessed image I d Different processing of different sub-regions of an image region is performed on pre-processed image I, depending on its position within the image region. d Each sub-region of the image region may be provided via a single ANN trained to process them separately. Training such an ANN is illustrated below through an example. In the example, such an ANN takes an input image I and an instruction for the image sub-region of interest as input, and preprocesses the image I d The corresponding sub-region can be provided as output. In a variation of this approach, the input to a single ANN may further include depth information about the sub-region under consideration. Thus, preprocessed image I d The overall process involves applying a single ANN individually to each sub-region of the image region, and then preprocessing the resulting sub-regions into image I. d It can be derived by combining them.
[0036] In another example where at least one ANN is applied, the corresponding preprocessed image I d Different processing of different sub-regions of an image region is performed on each pre-processed image I dThis can be provided through the application of multiple ANNs trained to process each sub-region of the image region. Training such multiple ANNs is illustrated by the example provided below. In the example, each of such ANNs takes each sub-region of the input image I as input and preprocesses the image I. d The corresponding sub-region can be provided as output. In a variation of this approach, the input to each ANN may further include depth information about the sub-region processed by each ANN. In the end, multiple ANNs are applied to the corresponding sub-regions of the image region, and the sub-regions thus obtained are preprocessed image I d By combining them, preprocessed image I d The whole can be derived.
[0037] Following the policy described above, in a non-limiting example, the DOE displays preprocessed image I on the 2D display 202. dThe optical element may be configured to perform wavefront coding aimed at providing a substantially uniform frequency response over the depth range of the NED assembly 201 based on the following. In this regard, the DOE may be configured to cause the lens assembly 203 to introduce a phase delay in the transmitted light that depends on the position of the lens plane. In other words, the DOE may be configured to provide a phase delay that may differ at different positions (e.g., sub-parts) of the lens plane in order to facilitate the provision of accommodatively invariant 3D presentation of the input image I over the depth range of the NED assembly 201. The phase delay through different positions of the lens plane can be represented as Φ(s,t), where the tuple (s,t) acts to indicate the 2D coordinates on the lens plane. As schematically shown for DOE203b in the example of Figure 2A, the variation in phase delay through different sub-parts of the DOE may be provided via a non-uniform thickness of the optical element. For example, such a non-uniform thickness may be implemented via a profile defined by a height map d(s,t) that defines the thickness of DOE203b over the position of the lens plane. In this regard, the phase delay Φ(s,t) can be converted to a height map d(s,t), taking into account the optical properties of the material applied to implement the lens element functioning as a DOE203b, and the height map d(s,t) can be converted to a phase delay Φ(s,t).
[0038] In other examples, the DOE included in the lens assembly 203 may differ from DOE203b in the example in Figure 2A. As an example in this regard, the DOE may be provided by the application of a phase-modulated metamaterial configured, for example, as a metasurface, and the desired phase delay across the position of the lens aperture of the lens assembly 203 may be provided by correspondingly controlling at least one aspect (e.g., orientation and / or size) of the nanostructure of the metasurface.
[0039] In a non-restrictive example, the phase delay Φ(s,t) may be rotationally symmetric with respect to the central axis of the lens assembly 203. In the scenario shown in Figure 2A, this can be achieved by applying a height map d(s,t) that is rotationally symmetric with respect to the central axis of the lens assembly 203 (and therefore with respect to the central axis of the DOE 203b). This is a choice that simplifies the DOE design while simultaneously considering the fact that the defocus aberrations of the lens assembly 203 and the display PSF across the lens plane positions are typically rotationally symmetric with respect to the central axis of the lens assembly 203. In another example, the phase delay Φ(s,t) does not have to exhibit symmetry with respect to the central axis of the lens assembly 203 (or any other axis), which typically increases the complexity of the DOE (and preprocessing procedure) design while increasing the possibility of controlling the phase delay Φ(s,t) across the lens plane positions. Considering DOE203b in the framework of the example in Figure 2A, Figure 4 shows a non-restrictive example of a height map d(s,t) that can be applied to define DOE203b as a symmetric element with respect to the central axis of the lens assembly 203.
[0040] In accordance with the above policy, the phase delay characteristics of the DOE across the position of the lens plane of the lens assembly 203 can be matched with display plane position recognition preprocessing characteristics provided through the operation of a preprocessing procedure to facilitate driving the accommodation depth of the viewer's eye 110 in a desired manner. In this regard, specific characteristics of the preprocessing procedure and the phase delay Φ(s,t) of the DOE can be designed through a joint procedure to ensure such matched characteristics.
[0041] As mentioned above, in various examples, the preprocessing procedure may be provided by the application of an ANN, but the weights of the ANN that performs the preprocessing procedure may be determined by a training procedure performed before the application of the preprocessing procedure in the image preprocessing unit 212 of the display controller 210. The agreement between the preprocessing characteristics provided by the preprocessing procedure and the phase delay characteristics of the DOE may be provided by modeling and optimizing the phase delay Φ(s,t) as part of a training procedure applied to determine the weights of the ANN that plays a role in providing the preprocessing procedure.
[0042] Figure 5 shows an example of a learning model 300, which performs an iterative learning procedure that relies on supervised learning to determine the weights of the ANN, which is responsible for performing the preprocessing procedure of the image processing unit 212, and at the same time can be applied to determine the phase delay of each of several positions on the lens plane (s,t) of the lens assembly 203, and thus determine a phase delay profile that can be applied as the phase delay Φ(s,t) of the DOE. Following the policy described above, the phase delay Φ(s,t) can be applied to determine the corresponding height map d(s,t) of the DOE 203b, for example, if it is implemented as an element with a variable height across the DOE aperture, as in the example of Figure 2A. The learning procedure is performed using multiple training images I t A training dataset typically containing several thousand images can be applied, and the training dataset includes color images (e.g., RGB images) representing natural scenes that may be selected considering the intended use of the NED assembly 201 and / or the NED device utilizing the NED assembly 201. For example, before starting the iterative learning procedure, the ANN weights and / or phase delay profile may be initialized with (pseudo) random values, while in another example, the initial values of the ANN weights and / or phase delay profile may be the respective values obtained from a previous learning procedure.
[0043] The aforementioned training dataset may include images, and each training image I tThis allows us to define each pixel value that can be provided as input to the learning model 300 during the iterative learning procedure (as shown in Figure 5). In some examples, the training dataset is divided into each depth map D t (t) may further include the corresponding training image I in the course of the learning procedure. t It can also be provided as input to the learning model 300 (not shown in Figure 5).
[0044] For example, the learning procedure may be performed through the operation of a computing device comprising one or more processors and one or more memories for storing one or more computer programs, the one or more computer programs being configured, when executed by the one or more processors, to cause the computing device to perform the learning procedure described herein. In another example, the learning procedure may be performed through the use of multiple (i.e., two or more) computing devices of the above-described types, with any necessary modifications.
[0045] The learning model 300 uses each training image I t Eye model 302 for processing the corresponding reference retinal image I r This applies to each training image I, without considering the effect of the NED assembly 201 caused by the combined effect of the preprocessing procedure and the lens assembly 203. t Models direct (or natural) viewing. Reference retinal image I r This is derived for a given accommodative state of the eye, assuming that the eye is within the depth range of the NED assembly 201 and is accommodating to a reference plane representing the conjugate plane of the retina of the viewer's eye 110. Reference retinal image I r This corresponds to the training image I within the framework of the learning model 300. t It functions as a ground truth image. The eye model 302 can be configured to model the chromatic aberration of the (typical) optical system of the eye and the diffraction-limited resolution of the eye 110 due to the finite-sized pupil 110a, thereby a reference retinal image I rThis is derived as the cause of at least some optical limitations of eye 110. This can be achieved, for example, by convolution of a corresponding training image using the diffraction-limited PSF for the viewer, which represents the chromatic aberration of eye 110. Corresponding training image I t Reference retinal image I derived based on r This remains the same throughout the iteration rounds, and therefore the learning procedure is performed before entering the iterative learning procedure, or only during the first iteration round of the iterative learning procedure, for each training image I of the training dataset. t Based on the corresponding reference retinal image I r This could include making a decision.
[0046] In each iteration round of the learning procedure, each training image I t This goes through the learning model 300, which is each training image I t This results in the following actions being performed. - Select one or more sub-apertures of the lens plane in each iterative round, and each training image I spatially corresponding to the selected one or more sub-apertures t Determine one or more sub-regions of each image region, - Each position P within the image region n Taking this into consideration, use the current weights of at least one ANN304 to train each image I by at least one ANN304. t The determined one or more subregions are preprocessed into image I d Process into one or more spatially corresponding sub-regions, -Preprocessed image I by display and eye model 306 d One or more subregions of the above are simulated retinal image I^ according to the current phase delay profile of the display model and considering the selected accommodation depth z for each iteration round. r Process in one or more spatially corresponding sub-regions, -The simulated retinal image I^ is obtained by loss function 308. r The aforementioned one or more subregions and reference retinal image Ir Determine the difference between the spatially corresponding one or more sub-regions.
[0047] At the end of each iteration round, the weights of at least one ANN304 and at least a portion of the phase delay profiles spatially corresponding to one or more subapertures selected for each iteration round are used to simulate the retinal image I^ r Each of one or more subregions and reference retinal image I r The update is based on the determined difference between the spatially corresponding one or more sub-regions of the image. As an example in this regard, the update can be performed via the use of gradient descent. The learning procedure involves a simulated retinal image I^ r and the corresponding reference retinal image I r The process can run until one or more predefined convergence criteria related to the difference between and are met, and / or until a predefined number of iterations have been performed. Once the learning process converges and / or the predefined number of iterations have been completed, the weights of at least one ANN304 from the last iteration can be adopted as the weights of at least one ANN applied to the image preprocessing unit 212 of the display controller 210 to perform a preprocessing procedure, and the final phase delay profile from the last iteration can be adopted as the phase delay Φ(s,t) of the DOE of the NED assembly 201.
[0048] For example, the learning model 300 can directly consider the parameters that define the phase delay profile, rather than considering the phase delay profile itself. In an example using a DOE characterized by material properties and thickness (e.g., DOE 203b in the example in Figure 2A), parameters considered during the learning procedure by applying the learning model 300 may include a height map d(s,t) over the position of the lens plane of the lens assembly 203. In another example, the phase delay profile of an optical element acting as a DOE can be considered via the coefficients of each of a predefined set of phase functions in a suitable signal space, e.g., a Zernike polynomial. In an example where the DOE is provided via the use of a metamaterial / metasurface, parameters that define the phase delay considered during the learning procedure may include one or more geometric properties (such as radius and / or orientation) of the nanostructure implementing the metasurface over the position (s,t) of the lens plane. Regardless of the method used to define the phase delay profile during the learning process, the phase delay profile can be defined by defining a phase delay for each position (s,t) on the lens plane, where (s,t) may be from the lens plane or a substantially uniform grid across the position (s,t).
[0049] As explained above, in each iteration round, the learning model 300 spatially corresponds to one or more sub-apertures of the lens plane and the training image I considered in each iteration round. t Consider one or more sub-regions of the image region. In this regard, as previously described with reference to Figures 3A and 3B, for example, each sub-aperture of the lens plane map maps to a corresponding sub-region of the image region (and therefore to a corresponding sub-region of the display plane), and the mapping depends on the spatial relationship between the respective positions of the 2D display 202 and the lens assembly 203 and their distances from the pupil 110a. In this regard, the illustration in Figure 5 shows position P n The image region in the image region is mapped to one or more corresponding sub-regions, and each of the one or more incident angles (θ) in the pupil 110a is mapped to one or more sub-regions. n,φ n This suggests defining one or more sub-openings via ).
[0050] For example, a single sub-aperture and training image I t The spatially corresponding subregions of are considered in each iteration round. In this regard, all training images I in a particular iteration round of the learning procedure are considered. t A single sub-aperture of the lens plane to which the sub-aperture is applied can be selected (substantially) randomly from a plurality of predefined sub-apertures, but in another example, a sub-aperture applied to a particular iteration round can be varied from one iteration round to another by selecting one of a plurality of predefined sub-apertures according to a predefined rule. The plurality of sub-apertures available for consideration in the learning procedure can each have a predefined position in the lens plane, and they can have a predefined shape and size, and this aperture approximates that of a typical (e.g., average) pupil 110a. As an example, the plurality of sub-apertures can together cover the entire lens aperture without any gaps between them. In a non-limiting example thereof, the sub-apertures may have a rectangular or hexagonal shape, and the sub-apertures may not overlap, or they may partially overlap.
[0051] In another example, training image I t A set of two or more sub-apertures and corresponding one or more sub-regions of the image region may be applied in each iteration round, and the set of two or more sub-apertures applied in a particular iteration round may be selected (substantially) randomly from a plurality of predefined sub-apertures, or the set of two or more sub-apertures applied in a particular iteration round may be selected from a plurality of predefined sub-apertures according to a predefined rule. As a particular example of the latter, the set of two or more sub-apertures selected for a particular (and each) iteration round may include a plurality of sub-apertures that together cover the entire lens aperture.
[0052] Each training image I tRegarding this, at least one input to ANN304 in a given iteration round is each training image I within the image subregion considered in each iteration round. t The pixel value of the image region position P of the image subregion being considered. n It can include information that identifies the, while at least one output of ANN304 is preprocessed image I d It can include the corresponding pixel values of the corresponding image subregions. Position P of the image subregion within the image region. n This can be determined, for example, via the corresponding incident angle (θ,φ) at the pupil 110a (as shown in Figure 5), via the corresponding pixel position (ξ,η) on the display plane, or via the corresponding position (s,t) on the lens plane. As mentioned above, in some examples, the learning procedure involves training image I t Depth map D associated with (t) t (t) can further include depth information obtained via (t), and therefore at least one input to ANN304 can further include depth information of the image subregion being considered in each iteration round, which is the training image I being considered. t Associated depth map D t It can be obtained from. In accordance with the aforementioned policy, at least one ANN304 may include, for example, at least one CNN or at least one modified CNN.
[0053] The display and eye model 306, in combination with the eye model, process the preprocessed image I from the display plane through the lens assembly 203. d By simulating transmission, the corresponding simulated retinal image I^ on the reference plane is obtained. r This may include a physically based differentiable simulation model that derives the following: d Simulated retinal image I^ derived from rThis models the effect of lens assembly 203 along with the aberrations of the (typical) optical system of the eye and the diffraction-limited resolution of the eye, thereby enabling the use of each training image I via NED assembly 201. t This models viewing the display. Following the approach described above, the modeling of the lens assembly 203 involves modeling the optical properties of one or more sub-apertures of the lens plane selected for each iteration round, taking into account the adjustment depth z selected for each iteration round, via the application of a current phase delay profile, while eye modeling can be performed in a similar manner to that described above for the eye model 302, with necessary modifications. In particular, the modeling of the optical properties may include modeling the display PSF depending on the respective lens plane positions of the one or more sub-apertures selected for each iteration round, and depending on the adjustment depth z selected for each iteration round. The one or more sub-apertures to which the lens aperture is applied may be defined, for example, via the corresponding angle of incidence (θ,φ) at the pupil 110a (as shown in Figure 5), via the corresponding pixel position (ξ,η) on the display plane, or via the corresponding position (s,t) on the lens plane that defines the center point of the sub-apertures considered in each iteration round.
[0054] For example, the accommodation depth z applied to all training images in a particular iteration round of the learning procedure can be selected (substantially) randomly from a set of predefined accommodation depths covering a (desired) depth range of the NED assembly 201, or from accommodation depths within a predefined range. In another example, the accommodation depth applied to a particular iteration round can be changed from one iteration round to another by selecting the predefined accommodation depth or by selecting the accommodation depth from a predefined range according to a predefined rule.
[0055] The loss function 308 is calculated based on the simulated retinal image I^ r and the corresponding reference retinal image I r The difference between the two images I^r 、I r can be determined as an objective error metric such as L1 loss or mean squared error (MSE) derived based on the error (or difference) per pixel between. Additionally or alternatively, the difference can be determined as a perceptual error metric such as a structural similarity measure (SSIM) aimed at predicting the perceived error (or difference) between the two images I^ r 、I r 、I. According to an example, the loss function 308 can further include applying a so-called neural contrast sensitivity function (NCSF) before determining the difference between each image I^ r 、I r 、I. The use of NCSF can contribute by enabling an improved trade-off between the spatial resolution and the DoF provided by the NED assembly 201.
[0056] The above description of various examples regarding the application of the learning model 300 in the iterative learning procedure refers to the processing of the overall image. However, the images processed by the preprocessing procedure can include each RGB image including (as described above) three separate color channels. As a result, the training model 300 can be applied to consider the three color channels separately from each other during the iterative learning procedure.
[0057] The above example regarding the前述の (暗黙的な) learning procedure assumes that the operation of at least one ANN304 includes training a single ANN configured to process each sub-region of the image region of the preprocessed image I n depending on its position P d in the image region. Thus, as described above, the input to the single ANN during the learning procedure can include the pixel values of each training image I t in the image sub-region considered in each iteration round, together with the information P n identifying the image sub-region being considered, while the output of the single ANN is the preprocessed image I dcan include the corresponding pixel values of the corresponding image sub-regions (the input to a single ANN can optionally further include respective depth information).
[0058] According to other examples, the operation of at least one ANN 304 can include training a plurality of ANNs each configured to process a corresponding sub-region of an image region of a pre-processed image I d Thus, in each iteration, one or more ANNs corresponding to each of the one or more image sub-regions considered in each iteration are trained and / or updated. Along the foregoing lines, during the learning procedure, the input to each ANN is the pixel values within each of the respective image sub-regions considered in each iteration round of the respective training image I t but the output of each ANN can include the corresponding pixel values of the corresponding image sub-regions of the pre-processed image I d (the input to each ANN can optionally further include respective depth information).
[0059] FIG. 6 shows a block diagram of some components of a device 400 that can be used to implement at least some aspects of the display controller 210 or to execute the iterative learning procedure described above. Although described herein with reference to a single device 400, at least some aspects of the display controller 210 or the foregoing iterative learning procedure can be implemented by the coordinated operation of two or more devices 400 configured to provide cloud-based computing services.
[0060] The device 400 comprises a processor 410 and memory 420. The memory 420 can store data and computer program code 425. The device 400 may further comprise communication means 430 for wired or wireless communication with other devices and / or user I / O (input / output) components 440, which may be configured together with the processor 410 and a portion of the computer program code 425 to provide a user interface for receiving input from a user and / or providing output to the user. In particular, user I / O components may include user input means such as one or more keys or buttons, a keyboard, a touchscreen or touchpad. User I / O components may include output means such as a display or touchscreen. The components of the device 400 are coupled to communicate with one another via a bus 450 that enables the transfer of data and control information between components.
[0061] Memory 420 and a portion of the computer program code 425 stored therein may be further configured, using the processor 410, to cause the device 400 to operate as a display controller 210, or (if applicable) to perform the iterative learning procedure described above. The processor 410 is configured to read from and write to memory 420. Although each processor 410 is shown as a single component, it may be implemented as one or more separate processing components. Similarly, although each memory 420 is shown as a single component, it may be implemented as one or more separate components, some or all of which may be integrated / removable and / or may provide persistent / semi-persistent / dynamic / cache storage.
[0062] The computer program code 425 may include computer executable instructions that implement at least some aspects of the display controller 210, or (if applicable) perform the iterative learning procedure described above when loaded into the processor 410. For example, the computer program code 425 may include a computer program consisting of one or more instructions in one or more sequences. The processor 410 can load and execute the computer program by reading one or more instructions in one or more sequences contained therein from memory 420. When executed by the processor 410, one or more instructions in one or more sequences can be configured to cause the device 400 to operate as the display controller 210, or (if applicable) to perform the iterative learning procedure described above. Accordingly, the device 400 may comprise at least one processor 410 and at least one memory 420 containing computer program code 425 for one or more programs, the at least one memory 420 and the computer program code 425 being configured to cause the device 400 to execute at least some embodiments of the display controller 210 or the aforementioned iterative learning procedure (if applicable) using at least one processor 410.
[0063] The computer program code 425 may be provided, for example, as a computer program product including at least one computer-readable non-temporary medium on which the computer program code 425 is stored, and when executed by the processor 410, the computer program code 425 causes the device 400 to execute at least some embodiments of the display controller 210 or (where applicable) the aforementioned iterative learning procedure. The computer-readable non-temporary medium may include a memory device or recording medium that tangibly embodies the computer program. In another example, the computer program may be provided as a signal configured to reliably transport the computer program.
[0064] References to processors in this specification should not be understood to encompass only programmable processors, but also specialized circuits such as field-programmable gate arrays (FPGAs), application-specific circuits (ASICs), and signal processors. The features described above may be used in combinations other than those explicitly stated.
Claims
1. A stereoscopic NED assembly (201) for an NED device, comprising a pair of near-eye display (NED) assemblies (201), wherein the NED assemblies (201) are A two-dimensional (2D) display (202) for rendering an image to be seen by the viewer's eye (110), A lens assembly (203) positioned at a predefined distance from the 2D display (202) to enable viewing of the image rendered via the 2D display (202), wherein the lens assembly (203) comprises a diffractive optical element (DOE) (203b) configured to provide different phase delays through a plurality of positions within its aperture, and Preprocessed image (I d The preprocessed image (I) is obtained by applying a preprocessing procedure configured to apply image region position-dependent preprocessing in order to take into account different transmission characteristics through different sub-apertures of the lens aperture of the lens assembly (203) in the derivation of different sub-regions of the image region (I), based on the received image (I). d ) is derived and the preprocessed image (I) is rendered via the 2D display (202) for rendering. d A display controller (210) is provided, which includes a pre-processing unit (212) configured to supply ), The preprocessing procedure and the lens assembly (203) are such that the preprocessed image (I) is perceived as sharp for a plurality of predefined adjustment depths that exist within the extended depth of field (DoF) of the NED assembly (201) when viewed through the lens assembly (203). d ) is configured to display the preprocessed image (I) via the operation of the NED device. d NED assembly (201) facilitates the provision of adjustable invariant 3D presentation based on ).
2. The NED assembly (201) according to claim 1, wherein the preprocessing unit (212) includes at least one artificial neural network (ANN) configured to perform the preprocessing procedure.
3. The NED assembly (201) according to claim 2, wherein the at least one ANN includes at least one convolutional neural network (CNN).
4. The at least one ANN, depending on its position within the image region, processes the preprocessed image (I d The NED assembly (201) according to claim 2, comprising a single ANN trained to process each of a plurality of sub-regions of the image region of ).
5. The at least one ANN is a plurality of ANNs, and each ANN is the preprocessed image (I d The NED assembly (201) according to claim 2, which is trained to process each of a plurality of sub-regions of the aforementioned image region.
6. The NED assembly (201) according to any one of claims 1 to 5, wherein the DOE (203b) includes an optical element having a thickness determined separately for each of the multiple positions in order to provide different phase delays through the multiple positions within its aperture.
7. A preprocessing procedure and apparatus for deriving a phase delay profile for a stereoscopic near-eye display (NED) assembly (201) according to claim 1, wherein the apparatus comprises the preprocessing procedure and a plurality of training images (I t The instrument is configured to apply each learning model to jointly derive at least one artificial neural network (ANN) (304) which serves as the phase delay profile that defines the respective phase delays for multiple positions of the DOE (203b) via an iterative learning procedure based on the following, in multiple iterative rounds the instrument Each of the aforementioned training images is selected as follows: Select one or more sub-openings of the lens aperture for each of the respective iterative rounds, and determine each one or more sub-regions of the image region of the respective learning image (I t ), which spatially corresponds to the one or more selected sub-openings. The at least one ANN (304) generates the respective training images (I t The process of the one or more sub-regions determined above in the image region, and the respective positions (P n Taking into consideration the above, the current weights of at least one ANN(304) are used to preprocess the image (I d Determining one or more spatially corresponding sub-regions of ) The display and eye model (306) are used to process the preprocessed image (I d One or more subregions of ) are simulated according to the current phase delay profile and taking into account the selected accommodation depth (z) for each of the iterative rounds, to obtain a simulated retinal image (I^ r Processing in one or more spatially corresponding sub-regions of ) and The simulated retinal image (I^) is obtained by the predefined loss function (308). r ) one or more subregions and the corresponding reference retinal image (I r This is done via a learning model (300) configured to perform the task of determining the difference between one or more spatially corresponding sub-regions of ), and The aforementioned plurality of training images (I t Based on the differences determined for each of the above, update the weights of the at least one ANN(304) and the portion of the phase delay profile that spatially corresponds to the one or more selected sub-apertures. A device configured to perform the following actions.
8. The apparatus according to claim 7, configured to update the weights and phase delay profile of the at least one ANN using a gradient descent method.
9. The corresponding reference retinal image (I r ) an eye model (302) representing one or more optical limits of the eye is used for each of the learning images (I t The apparatus according to claim 7, configured to be derived by applying to ).
10. below, For each of the aforementioned iteration rounds, the adjustment depth (z) is one of several predefined Select one of the adjustment depths. The adjustment depth (z) for each of the aforementioned iteration rounds is selected from a predefined range of adjustment depths. The apparatus according to claim 7, configured to perform one of the following:
11. For each of the aforementioned iterative rounds, it is configured to select one of a plurality of predefined sub-apertures of the lens aperture, Each predefined sub-aperture has its own position within the lens aperture, and its own predefined shape and size. The apparatus according to claim 7, wherein the plurality of sub-openings together cover the entire opening of the DOE.
12. For each of the aforementioned repeating rounds, it is configured to select two or more of the multiple predefined sub-apertures of the lens aperture, Each predefined sub-aperture has its own position within the lens aperture, and its own predefined shape and size. The apparatus according to claim 7, wherein the plurality of sub-openings all cover the opening of the DOE together.
13. The display and eye model (306) apply a phase delay according to the current phase delay profile defined for the one or more sub-apertures selected for each iteration round, thereby processing the preprocessed image (I d Based on the one or more subregions of the above, the simulated retinal image (I^ r The apparatus according to any one of claims 7 to 12, configured to derive one or more spatially corresponding sub-regions of ) and to model the optical properties of the one or more sub-apertures selected for each iterative round, taking into account the adjustment depth (z) selected for each iterative round.
14. The apparatus according to claim 13, wherein applying the optical properties includes applying a point image distribution function, PSF, selected according to the respective positions of the one or more sub-apers selected for each iterative round, taking into account the adjustment depth (z) selected for each iterative round.
15. The apparatus according to any one of claims 7 to 12, wherein the at least one ANN includes at least one convolutional neural network (CNN).
16. The at least one ANN(304) is, depending on its position within the image region, the preprocessed image (I d The apparatus according to any one of claims 7 to 12, comprising a single ANN trained to process each of a plurality of sub-regions of the image region of the aforementioned image region.
17. The at least one ANN (304) is a plurality of ANNs, and each ANN is the preprocessed image (I d The apparatus according to any one of claims 7 to 12, which is trained to process each of the multiple sub-regions of the aforementioned image region.
18. A preprocessing procedure and a method for deriving a phase delay profile for a stereoscopic near-eye display (NED) assembly (201) according to claim 1, wherein the method applies a learning model to derive together the preprocessing procedure and at least one artificial neural network, ANN, (304) which serves as the phase delay profile that defines the respective phase delays for a plurality of positions of the DOE (203b), thereby obtaining a plurality of training images (I t The method includes performing an iterative learning procedure based on ), wherein for multiple iterative rounds, Each of the aforementioned training images is, For each of the aforementioned iteration rounds, one or more sub-apertures of the lens aperture are selected, and each of the training images (I) spatially corresponding to the selected one or more sub-apertures is selected. t To determine one or more sub-regions of each image region of ) The respective training images (I t The process of the one or more sub-regions determined above in the image region, and the respective positions (P n Taking into consideration the above, the current weights of at least one ANN(304) are used to preprocess the image (I d Determining one or more spatially corresponding sub-regions of ) The preprocessed image (I d One or more subregions of ) are simulated according to the current phase delay profile and taking into account the selected accommodation depth (z) for each of the iterative rounds, to obtain a simulated retinal image (I^ r Processing in one or more spatially corresponding sub-regions of ) and The simulated retinal image (I^) is obtained by the predefined loss function (308). r ) one or more subregions and the corresponding reference retinal image (I r This involves processing via a learning model (300) which includes determining the difference between one or more spatially corresponding sub-regions of ), and The aforementioned plurality of training images (I t Based on the differences determined for each of the above, update the weights of the at least one ANN(304) and the portion of the phase delay profile that spatially corresponds to the one or more selected sub-apertures. A method that includes this.
19. Computer program (425) comprising computer-readable program code configured to cause one or more computing devices (400) to perform the method described in claim 18 when executed on one or more computing devices (400).