Accommodation invariant near-eye display for ar applications
The near-eye display assembly with optical waveguides and DOEs addresses the vergence-accommodation conflict by ensuring focus over a predefined depth range, enhancing the viewing experience in AR applications.
Patent Information
- Application Number
- PCT/FI2024/050673
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-26
- Filing Date
- 2024-12-11
- Publication Date
- 2025-10-02
AI Technical Summary
Conventional near-eye displays (NEDs) for augmented reality (AR) applications suffer from visual discomfort due to the vergence-accommodation conflict (VAC), where the depth of vergence does not match the accommodation depth, leading to distorted viewing experiences.
A near-eye display assembly that employs an optical waveguide with diffractive optical elements (DOEs) and a preprocessing procedure to provide accommodation invariant rendering, ensuring the virtual image is in focus over a predefined depth range while maintaining an undistorted view of the real world.
The solution addresses the VAC by aligning accommodation depth with vergence depth, providing an extended depth of field (EDoF) and improving the perceived quality of 3D virtual images, reducing retinal defocus blur and enhancing the viewing experience.
Smart Images

Figure FI2024050673_02102025_PF_FP_ABST
Abstract
Description
[0001] Accommodation invariant near-eye display for AR applications
[0002] TECHNICAL FIELD
[0003] The present invention relates to near-eye displays (NEDs) suitable for augmented reality (AR) applications.
[0004] BACKGROUND
[0005] Stereoscopic near-eye displays (NEDs) are of particular interest for application in various virtual reality (VR) and augmented reality (AR) applications. A challenge arises in ensuring comfortable but yet natural or near-to-natural viewing experience via usage of a wearable NED apparatus due to visual discomfort experienced by some users.
[0006] One of key challenges in ensuring comfortable viewing experience involves addressing the vergence-accommodation conflict (VAC), which may arise when the NED apparatus is applied to display an object such that its vergence depth does not match its accommodation depth. In viewing VR or AR content using a conventional NED apparatus such situations occur frequently due to displaying objects of a three-dimensional (3D) image having their respective positions within the image space at depths (i.e. at respective vergence depths) that are different from the depths of the (virtual) image plane of the NED apparatus (i.e. the accommodation depth).
[0007] A NED apparatus that is designed to address the VAC may be referred to as an accommodation invariant (Al) NED apparatus. Al NED apparatuses known in the art that are provided for VR and AR applications make use of various techniques for addressing the VAC, such as varifocal, multifocal, light field, holographic and Maxwellian. All these techniques are subject to their own trade-offs, such as limited eyebox, limited image quality (e.g. low spatial resolution, speckle noise) and / or complex device requirements (e.g. bulky and / or fast optics, accurate eye tracker). Moreover, particular further challenges may arise in NEDs designed for use in NED apparatuses provided for AR applications, where the display design also needs to account for keeping the user’s view to real-world clear and undistorted.
[0008] SUMMARY
[0009] It is an object of the present invention to provide a technique for providing accommodation invariant NED apparatus for AR applications that enables accommodation invariant rendering of the AR content without substantially distorting the view to the real world though the NED.
[0010] According to an example embodiment, a near-eye display (NED) assembly for a stereoscopic NED apparatus including a pair of NED assemblies is provided, the NED assembly comprising: an image projection portion projecting collimated light that represents a preprocessed image; an optical waveguide comprising an incoupler arranged to receive the light projected from the image projection portion and an outcoupler arranged to release the light received via the incoupler; a diffractive optical element (DOE) arranged on a first side of the optical waveguide to allow for viewing the light released from the optical waveguide therethrough, wherein the DOE is arranged to apply a phase modulation function that results in a phase delay that varies with a spatial position of an aperture of the DOE; and a preprocessing portion arranged to derive the preprocessed image based on an input image via application of a preprocessing procedure that is arranged to apply an image-area-position dependent preprocessing in derivation of different image sub-areas of the preprocessed image to account for the variation in the phase delay through spatially corresponding sub-areas of the aperture of the DOE, wherein said preprocessing procedure and the phase modulation function are arranged to jointly provide a virtual image for perception through the DOE substantially in focus over a predefined depth range.
[0011] According to another example embodiment, an apparatus for deriving a preprocessing procedure and a phase modulation function for a NED assembly according to the example embodiment described in the foregoing is provided, the apparatus arranged to carry out an iterative learning procedure that comprises processing a plurality of training images through a learning arrangement that represents the arrangement of the preprocessing procedure and the DOE and that comprises respective learning models for determining weights for at least one artificial neural network (ANN) that serves as the preprocessing procedure and for determining a phase delay profile serves to implement the phase modulation function of the DOE, wherein the phase delay profile is determined for a predefined reference wavelength.
[0012] According to an exemplifying aspect of the present disclosure, a computer program is provided, the computer program comprising computer readable program code configured to carry out the iterative learning procedure described in the foregoing when said program code is executed on one or more computing apparatuses.
[0013] The computer program according to the above-described example embodiment may be embodied on a volatile or a non-volatile computer- readable record medium, for example as a computer program product comprising at least one computer readable non-transitory medium having the program code stored thereon, which, when executed by one or more computing apparatuses, causes the computing apparatuses at least to perform the method according to the example embodiment described in the foregoing.
[0014] The exemplifying embodiments of the invention presented in this patent application are not to be interpreted to pose limitations to the applicability of the appended claims. The verb "to comprise" and its derivatives are used in this patent application as an open limitation that does not exclude the existence of also unrecited features. The features described hereinafter are mutually freely combinable unless explicitly stated otherwise.
[0015] Some features of the invention are set forth in the appended claims. Aspects of the invention, however, both as to its construction and its method of operation, together with additional objects and advantages thereof, will be best understood from the following description of some example embodiments when read in connection with the accompanying drawings. BRIEF DESCRIPTION OF FIGURES
[0016] The embodiments of the invention are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings, where
[0017] Figure 1 schematically illustrates some aspects of a near-eye display assembly according to an example;
[0018] Figure 2 schematically illustrates some aspects of a near-eye display assembly according to an example;
[0019] Figure 3A schematically illustrates some aspects of a near-eye display assembly according to an example;
[0020] Figure 3B illustrates a block diagram of some components of a near-eye display assembly;
[0021] Figure 4 illustrates a height map for a dispersive optical element according to an example;
[0022] Figure 4A illustrates a block diagram that represents an eye model according to an example;
[0023] Figure 4B illustrates a block diagram that represents a learning model according to an example:
[0024] Figure 4C illustrates a block diagram that represents a learning model according to an example;
[0025] Figure 5 schematically illustrates some aspects of a near-eye display assembly according to an example;
[0026] Figure 6A illustrates a block diagram that represents an eye model according to an example;
[0027] Figure 6B illustrates a block diagram that represents a learning model according to an example: Figure 6C illustrates a block diagram that represents a learning model according to an example;
[0028] Figure 7 schematically illustrates some aspects of a near-eye display assembly according to an example;
[0029] Figure 8 schematically illustrates some aspects of a near-eye display assembly according to an example;
[0030] Figure 9 schematically illustrates some aspects of a near-eye display assembly according to an example; and
[0031] Figure 10 illustrates a block diagram of some components of an apparatus according to an example.
[0032] DESCRIPTION OF SOME EMBODIMENTS
[0033] Figure 1 schematically illustrates some aspects of an exemplifying near-eye display (NED) assembly 100 for an augmented reality (AR) application together with an eye 120 of a viewer. The NED assembly 100 comprises an image projection unit 102 and an optical waveguide 106, where the optical waveguide 106 is to be arranged in front of the eye 120 of the viewer upon operation of the NED assembly 100 such that a pupil 120a of the eye 120 is positioned within an eyebox that resides substantially at a viewing distance from the optical waveguide 106 on a first side of the optical waveguide 106. In this regard, the eyebox consists of a predefined range of positions (i.e. an area) on a plane that is at the viewing distance from the optical waveguide 106. This plane may be also referred to as a pupil plane. Consequently, the viewer is able to see his / her environment - i.e. the real world - through the optical waveguide 106, whereas the view to the real world may be complemented by AR content displayed via the optical waveguide 106. The AR content is provided in an input image I, whereas the image projection unit 102 is arranged to project light that represents the input image I towards the optical waveguide 106, which serves to release the light that represents the input image I in a controlled manner from the first side of the optical waveguide 106. At the same time the optical waveguide 106 transmits the external light that enters the optical waveguide 106 from its second side that is opposite to the first side towards the eyebox on the first side. Hence, the eye 120 positioned at the pupil plane within the eyebox receives both the light that represents the input image I and the external light scattered from objects of the real world, which substantially results in embedding a virtual image of the AR content into the viewer’s view to the real word.
[0034] It is worth noting that the schematic illustration of the NED assembly 100 in Figure 1 , as well as the subsequent illustrations of other NED assemblies in subsequent figures, are highly schematic ones that serve to illustrate a conceptual structure of the respective NED assembly and operational relationship between its elements, while on the hand these illustrations do not aim at indicating respective sizes of elements of the respective NED apparatus in relation to each other or their exact positions with respect to each other. As a particular example in this regard, in the interest of graphical clarity, the pupil 120a is illustrated in size that is significantly smaller than the distance between two consecutive outcoupling locations of the illustrated ray (i.e. the exit pupil replication distance): in a real-life implementation, this distance is typically closer to the size of the pupil 120a, i.e. it is typically larger than the pupil 120a by no more than the size of the incoupler 106a (an entrance pupil).
[0035] In the example of Figure 1 the image projection unit 102 comprises a two- dimensional (2D) display 102a and a collimating lens 102b, whereas the optical waveguide 106 is shown with an incoupler 106a and an outcoupler 106b. In particular, the light that represents the input image I rendered on the 2D display 102a is projected though the collimating lens 102b towards the incoupler 106a, which is arranged to couple light directed thereto into the optical waveguide 106 for propagation therethrough due to total internal reflection (TIR). The outcoupler 106b is arranged to release the light propagating through the optical waveguide 106 towards the eyebox at a substantially uniform intensity along the length of the outcoupler 106b, thereby conveying the virtual image represented by the light received via the incoupler 106a to the viewer. Hence, the optical waveguide 106 serves to expand the image received thereat, which results in expanding size of the eyebox accordingly.
[0036] The optical waveguide 106 applied in the NED assembly 100 may be one known in the art. In an example, each of the incoupler 106a and the outcoupler 106b comprises a respective diffractive structure, such as a surface relief grating or a holographic grating, arranged to implement the respective optical transformation for light directed thereto. In another example, each of the incoupler 106a and the outcoupler 106b comprises a respective refractive structure, such as a mirror or a prism, arranged to implement the respective optical transformation for light directed thereto. In some examples, the optical waveguide 106 is substantially planar element (as implied via the schematic illustration of Figure 1 ), whereas in other examples the optical waveguide 106 has a curved shape with a concave surface on its first side.
[0037] The image projection 102 unit may be also referred to as a light engine of the NED assembly 100. It is also worth noting that the image projection unit 102 according to the example of Figure 1 serves as a non-limiting example and in other examples a different image projection unit known in the art may be applied. As an example in this regard, the image projection 102 unit may comprise a 2D display or a light emitter of other kind arranged to emit light that represents the input image I together with a transmissive or reflective spatial light modulator (SLM) arranged to direct the light that represents the input image I towards the incoupler 106a of the optical waveguide 106. The input image I comprises, for example, a RGB image that defines respective pixel values for each pixel position of the input image I separately in red, green and blue color channels.
[0038] A stereoscopic NED apparatus comprises a pair of NED assemblies 100, i.e. one NED assembly 100 for each eye of the viewer, and a display controller arranged to supply a respective input image / of a stereoscopic image pair for viewing via respective waveguides 106 of the two NED assemblies 100 to convey a three-dimensional (3D) virtual image of the AR content represented by the stereoscopic image pair. The display controller may further serve to control one or more (other) aspects of operation of the respective image projection units 102 of the two NED assemblies 100 of the stereoscopic NED apparatus.
[0039] The NED assembly 100 according to the example of Figure 1 basically enables viewing 3D virtual images (i.e. the AR content) that are focused at infinity due to respective light rays representing the pixels of the input image I displayed via the optical waveguide 106 propagating in parallel to each other from the outcoupler 106b towards the eyebox. When using the pair of NED assemblies 100 to display 3D virtual images representing an object having its vergence depth (i.e. vergence distance) different from infinity, the eye 120 nevertheless tends to accommodate at infinity, where the displayed virtual image appears sharpest. This results in the vergence-accommodation conflict (VAC) due to difference in the vergence depth and the accommodation depth. The VAC is found to be a major source of visual discomfort experienced by many users viewing 3D virtual images via the NED apparatus. On the other hand, the external light scattered from real-world objects visible through the two NED assemblies 100 pass through their respective waveguides 106 unmodified and hence their actual depths (cf. ‘vergence depths’) match their focal depths (cf. ‘accommodation depths’) and, consequently, no VAC arises from viewing the real-world objects.
[0040] Figure 2 schematically illustrates some aspects of a NED assembly 100’, which includes the elements already described with reference to the NED assembly 100 and further includes a first lens 108 arranged on the first side of the optical waveguide 106 between the optical waveguide 106 and the eyebox at the pupil plane and a second lens 110 arranged on the second side of the optical waveguide 106. The first lens 108 is a negative lens (concave lens, diverging lens) that serves to change the focal depth of the virtual image representing the AR content to a desired vergence depth that depends on optical characteristics of the first lens 108, whereas the second lens 110 is a positive lens (convex lens, converging lens) that serves as a corrective lens that compensates for the optical power of the first lens 108: since the first lens 108 also modifies the focal depth of the real-world objects visible through the optical waveguide 106 in the same manner as it modifies the focal depth of the virtual image displayed via the optical waveguide 106, the second lens 110 is applied to provide the viewer with a substantially undistorted view to the real world through the optical components of the NED assembly 100’.
[0041] Along the lines described in the foregoing for the NED assembly 100, a stereoscopic NED apparatus may employ a pair of NED assemblies 100’ and the display controller arranged to supply the respective input image I of a stereoscopic image pair for viewing via respective waveguides 106 of the two NED assemblies 100’ to embed the 3D virtual image of the AR content represented by the stereoscopic image pair into the viewer’s view to the real world.
[0042] The respective optical arrangements of the NED assemblies 100 and 100’ enable displaying a 3D virtual image at a virtual image plane that substantially corresponds to the focal depth defined by the optical components of the respective NED assembly: in the NED assembly 100 the virtual image plane resides at infinity, whereas in the NED assembly 100’ the virtual image plane is defined via optical characteristics of the first lens 108. Moreover, in the NED assemblies 100 and 100’ the virtual image may be displayed at a resolution up to the diffraction limit at the virtual image plane. At vergence depths (i.e. virtual distances from the pupil plane) that are offset from the virtual image plane, a rapid decrease in frequency response with increasing offset is observed due to the defocus blur (or defocus aberration). Hence, the viewer observes a sharp image at a high resolution if an accommodation depth of the eye 120 coincides with the virtual image plane, whereas the image observed by the viewer becomes increasingly blurred with increasing offset between the accommodation depth of the eye 120 and the virtual image plane.
[0043] The defocus blur is a major driver of the accommodation depth of the eye 120 in a viewing situation such that the eye 120 typically tends to accommodate at a depth where the image appears sharpest. Consequently, in the optical arrangement of the NED assemblies 100, 100’ the accommodation depth tends to coincide with the virtual image plane or fall in its immediate vicinity, regardless of the actual vergence depth of the virtual image displayed via the optical waveguide 106. This leads into the viewer perceiving the virtual image as a sharp one only in scenarios where the vergence depth is relatively close to the focal depth of the NED assembly 100, 100’. Another factor having an effect on the accommodation depth is the binocular disparity: even though the binocular disparity primarily drives the vergence depth, it also affects the accommodation depth.
[0044] The present disclosure describes an approach where the coupling between the accommodation depth and the vergence depth is employed to eliminate or at least significantly reduce retinal defocus blur in the virtual images displayed via the optical waveguide 106. This results in creating a viewing situation where the accommodation depth of the eye 120 is predominantly driven by the binocular disparity, thereby substantially aligning the accommodation depth with the vergence depth and hence addressing the VAC. Consequently, a NED assembly making use of this operating principle serves as an accommodation invariant NED assembly that provides an extended depth of field (EDoF) that significantly improves the perceived quality of the 3D virtual images viewed by a NED apparatus that employs a pair of such NED assemblies. Such a NED apparatus may be referred to as an accommodation invariant NED apparatus.
[0045] Figure 3A schematically illustrates some aspects of an accommodation invariant NED assembly 200a according to a first embodiment. The illustration of Figure 3 further shows the eye 120 of the viewer positioned within the eyebox at the pupil plane. Along the lines described for the NED assemblies 100 and 100’, the NED assembly 200a comprises the image projection unit 102 and the optical waveguide 106. The image projection unit 102 and the optical waveguide 106 as well as their arrangement with respect to each other are similar to those described in the foregoing with references to the NED assembly 100 illustrated in Figure 1 . In other words, the optical waveguide 106 receives the light projected from the image projection unit 102 via the incoupler 106a and releases the light received via the incoupler 106a from the first side of the optical waveguide 106 towards the eyebox, while at the same time the optical waveguide 106 transmits the external light that enters the optical waveguide 106 from its second side that is opposite to the first side towards the eyebox on the first side.
[0046] A stereoscopic NED apparatus may employ a pair of NED assemblies 200a for viewing respective images of a stereoscopic image pair via their respective optical waveguides 106 to provide the viewer with the 3D virtual image of the AR content represented by the stereoscopic image pair embedded into the view to the real world. The difference to the NED apparatus that rely e.g. on one of the NED assemblies 100, 100’ described in the foregoing is that the NED apparatus making use of a pair of NED assemblies 200a enables displaying the 3D virtual image in an accommodation invariant manner, thereby addressing the VAC.
[0047] In the NED assembly 200a, the light projected from the image projection unit 102 represents a preprocessed image ld, which is derived from a respective input image I of a stereoscopic image pair via application of a predefined preprocessing procedure. In this regard, the NED assembly 200a further includes an image preprocessing portion 232 arranged to apply the preprocessing procedure to the respective input image I to derive the corresponding preprocessed image ld. Alternatively, the image preprocessing portion 232 may be provided as an element of the NED apparatus that comprises the pair of NED assemblies 200a, where the image preprocessing portion 232 is arranged to derive the respective preprocessed image / dfor each of the two NED assemblies 200a.
[0048] The NED assembly 200a further comprises a diffractive optical element (DOE) 208a arranged on the first side of the optical waveguide 106 to allow for viewing the light released via the outcoupler 106b through the DOE 208a. In other words, the DOE 208a is arranged between the optical waveguide 106 and the eyebox and the light released from the optical waveguide 106 is transmitted towards the eyebox through the DOE 208a. The NED assembly 200a further comprises a second DOE 210a arranged on the second side of the optical waveguide 106such that the external light enters the optical waveguide 106 through the second DOE 210a, passes through the optical waveguide 106 and transmits through the DOE 208a towards the eyebox. The optical waveguide 106, DOE 208a and the second DOE 210a may be jointly referred to as optical components of the NED assembly 200a.
[0049] Each of the DOE 208a and the second DOE 210a are optical elements separate from the optical waveguide 106. According to an example, the DOE 208a is arranged directly adjacent to the optical waveguide 106 substantially without a gap therebetween, whereas in another example the DOE 208a is arranged at a first distance from the optical waveguide with an airgap therebetween. Along similar lines, according to an example, the second DOE 210a is arranged directly adjacent to the optical waveguide 106 substantially without a gap therebetween, whereas in another example the second DOE 210a is arranged at a second distance from the optical waveguide with an airgap therebetween. Each of the first distance and second distance, if applicable, is a respective fixed predefined (non-zero) distance chosen in consideration of the physical and optical design of the NED assembly 200a. In various examples, each of the first distance and the second distance may be in a range from a fraction of a millimeter to a few millimeters.
[0050] Each of the DOE 208a and the second DOE 210a is arranged to modify the phase of the light transmitted therethrough such that the respective phase delay introduced by the respective one of the DOE 208a and the second DOE 210a is different through different positions through the respective one of the DOE 208a and the second DOE 210a. The preprocessing procedure is matched with the phase modification characteristics of the DOE 208a to enable the accommodation invariant viewing of a 3D virtual image represented by the underlying stereoscopic image pair, whereas the second DOE 210a is matched with the DOE 208a to facilitate providing substantially undistorted view to the real world through the optical components of the NED assembly 200a. Consequently, the preprocessing procedure and the phase modification implemented by the DOE 208a jointly provide the virtual image for perception through the DOE 208a substantially in focus at accommodation depths within a predefined depth range, thereby providing 3D virtual image at an extended DoF when viewed through the pair of NED assemblies 200a of the NED apparatus from the viewing distance. The accommodation depth may be denoted as z, whereas the depth range of the NED assembly 200a may span over accommodation depths from z+to z- and it may include a plurality of accommodation depths within a depth range [z+, r].
[0051] The accommodation depths within the depth range [z+, z~] and the respective endpoints of the depth range [z+, z~] may be expressed as diopters. Consequently, in the dioptric domain, the accommodation depth may be denoted as ZD and the depth range may include a plurality of accommodation depths ZD within a depth range [Z+D, Z’D]. Considering the accommodation depths ZD in diopters is a convenient choice in view of the defocus blur changing linearly in diopters and the optical components of the NED assembly 200a are typically arranged to provide images of substantially equal sharpness for the plurality of accommodation depths within the depth range [z+, r]. In an example, the NED assembly 200a is arranged to cover a continuous range of accommodation depths ZD within the depth range [Z+D, Z’D] (i.e. ZD e [Z+D, Z’D]), whereas in another example the NED assembly 200a is arranged to consider a predefined set of accommodation depths ZD within the depth range [Z+D, Z’D] (i.e. ZD e {Z1D, Z2D, . . . , ZND}) that cover the depth range according to a predefined grid. The former example may be referred to as a continuous EDoF, whereas the later example may be referred to as a multifocal EDoF. As an example of implementing the multifocal EDoF, the predefined grid may substantially uniformly cover the depth range [Z+D, Z’D] via defining the step size between adjacent accommodation depths to be small enough (e.g. at or around 0.3 diopters) to ensure granularity that is fine enough to go unnoticed by human vision. In the following description, however, the symbols z and [z+, z~] are predominantly applied to refer, respectively, to the accommodation depth and to the depth range and these symbols serve to represent the corresponding measures regardless of the manner of expressing and / or defining these measures (e.g. in dioptric domain or in any other domain). In other words, the NED assembly 200a enables providing accommodation invariant representation of the 3D virtual image over an extended depth range via a joint contribution of the preprocessing procedure and the DOE 208a: the phase modification provided by the DOE 208a serves to implement the extension of the DoF within the predefined (desired) depth range [z+, ], whereas the preprocessing procedure operates to modify the underlying input image I in a manner that enables perceiving the resulting 3D virtual image as substantially sharp across the predefined depth range [z+, r], In yet other words, the preprocessing procedure may be considered to provide encoding of the input image I into the corresponding preprocessed image ldin consideration of the phase modification characteristics of the DOE 208a, whereas the DOE 208a may be considered to reconstruct the preprocessed image ldinto the 3D virtual image having the EDoF that spans over the predefined depth range [z+, z~] of the NED assembly 200a.
[0052] Moreover, the second DOE 210a serves to compensate for the phase delay characteristics provided by the DOE 208a in order to keep the view to the real world through the optical components of the NED assembly 200a substantially undistorted. In particular, since the DOE 208a serves to extend the DoF of the 3D virtual image, also the external light received through the optical waveguide 106 is subjected to the similar extension of DoF, whereas the second DOE 210a serves to compensate for the extended DoF to enable perceiving the real-world objects as sharp through the NED assembly 200a only at (or in vicinity of) depths that match their actual physical locations with respect to the eyebox.
[0053] The DOE 208a defines an aperture through which the light arriving from the direction of the optical waveguide 106 is transmitted towards the eyebox. The DOE 208a is arranged to implement a predefined first phase modulation function that results in a predefined phase delay that varies with a spatial position within the aperture of the DOE 208a. In other words, the first phase modulation function defines a respective phase delay for a plurality of predefined spatial positions within the aperture of the DOE 208a. According to an example, the DOE 208a may be provided as a lens-like optical element arranged to implement the first phase modulation function via having a non- uniform thickness, which is defined separately for the plurality of spatial positions of its aperture to provide a respective desired phase delay through the respective spatial positions of the aperture. In another example, the DOE 208a may be provided via application of phase-modulating metamaterial arranged e.g. as a meta-surface, where a nanostructure forming the metasurface is arranged to provide a respective desired phase delay through a plurality of positions of the meta-surface. In latter example, the desired phase modulation function may be provided via controlling at least one aspect (e.g. orientation and / or size) of a nanostructure of the meta-surface at different positions of its ‘aperture’ accordingly. In the context of the present disclosure, also the meta-surface implementation may be considered as an optical element due to its function in shaping optical characteristics of the light transmitted therethrough.
[0054] Referring back to the first phase modulation function provided by the DOE 208a, the phase delay through a position (s, t) of the aperture of the DOE 208a at a reference wavelength may be denoted as < / >e(s, t). If using the lens-like optical element to implement the DOE 208a, along the lines described in the foregoing, the first phase modulation function may be provided via a non- uniform thickness of the lens-like optical element via selecting its thickness at a position (s, t) such that it provides a desired phase delay <^(s, t) therethrough at the reference wavelength. The thickness of the lens-like optical element serving to provide the DOE 208a across the positions (s, t) of its aperture may be defined as a height map de(s, t). While in theory the phase delay <^(s, t) may be a continuous function of the position (s, t) of the aperture, in a practical implementation the phase delay < / >e(s, t) may be discretized such that the aperture positions (s, t) are sampled along the ’s axis’ by Asand along the ‘t axis’ by At. The phase delay < / >e(s, t) providing the desired phase modulation function at the reference wavelength may be defined via a training procedure described in the following. In case of employing the lens-like optical element as the DOE 208a, the phase delay < / >e(s, t) may be translated into the height map de(s, t) in consideration of optical characteristics of the material applied for implementing the lens-like optical element serving as the DOE 208a.
[0055] The example described above refers to the phase delaye(s, t) through the DOE 208a at the reference wavelength. Typically, the phase delay through the DOE 208a depends on the wavelength and, consequently, the phase delay for wavelengths other than the reference wavelength may be different. In this regard, assuming an RGB image as the input image I, the reference wavelength may be e.g. a wavelength corresponding to a predefined one of red, green and blue channels, e.g. the wavelength at which the spectrum of the green channel of the 2D display 102a has a peak.
[0056] The second DOE 210a defines an aperture through which the external light enters the optical waveguide 106. The second DOE 210a is arranged to implement a second predefined phase modulation function that results in a predefined phase delay that varies with a spatial position within the aperture of the second DOE 210a. In other words, the second phase modulation function defines a respective phase delay for a plurality of predefined spatial positions within the aperture of the second DOE 210a. In one example, the second DOE 210a may be provided as a lens-like optical element, whereas in another example the second DOE 210a may be provided via application of phase-modulating metamaterial arranged e.g. as a meta-surface, along the lines described in the foregoing for the DOE 208a, mutatis mutandis. The phase delay through a position (s, t) of the aperture of the second DOE 210a at the reference wavelength may be denoted asc(s, t). In case a lens-like optical element is applied to implement the second DOE 210a, the phase modulation function may be provided via a non-uniform thickness of the lenslike optical element such that its thickness at a position (s, t) is selected such that it provides a desired phase delayc(s, t) therethrough at the reference wavelength. The phase delayc(s, t) may be translated into a height map dc(s, t) in consideration of optical characteristics of the material applied for implementing the lens-like optical element serving as the second DOE 210a. According to a non-limiting example, the phase delay e(s, t) may be rotationally symmetric with respect to the center axis of the DOE 208a. In scenarios where the DOE 208a is implemented as a lens element, this may be accomplished via application of a height map de(s, t) that is rotationally symmetric with respect to the center axis of the DOE 208a. This is a choice that simplifies the design of the DOE 208a while at the same time accounting for the fact that defocus aberration of the lens element and the display PSF across the positions of the lens aperture are typically rotationally symmetric with respect to the center axis. According to other examples, the phase delay e(s, t) may not exhibit symmetry with respect to the center axis of the DOE 208a (or with respect to any other axis), which typically leads to increased complexity of design of the DOE 208a (and the preprocessing procedure) while enhancing the possibility to control the phase delay < / e(s, t) across positions through the lens aperture. While described herein with references to the DOE 208a and the phase delaye(s, t) it serves to implement, similar considerations apply to the second DOE 210a and the phase delayc(s, t) it serves to implement as well, mutatis mutandis.
[0057] Along the lines described in the foregoing, the predefined preprocessing procedure applied by the image preprocessing portion 232 is arranged to process the input image I into the corresponding preprocessed image ld, where the preprocessing procedure is matched with the phase modification characteristics of the DOE 208a in order to facilitate accommodation invariant 3D presentation at the underlying stereoscopic image pair. The aspect of processing the input image I into the corresponding preprocessed image ld, may also be considered as deriving the preprocessed image ldbased on the corresponding received input image / . In this regard, the preprocessing procedure processes the input image I into the corresponding preprocessed image ldin a manner that is different for different locations of the display plane of the 2D display 102a to account for the variation in phase modification characteristics through different locations of the aperture of the DOE 208a and for different positions of the eye 120 within the eyebox. In other words, the preprocessing procedure is carried out in a display-plane-position-dependent manner, which may result in providing different preprocessing characteristics for different locations of the display plane and, consequently, for different locations of the image area. In particular, the non-uniform processing of subareas of the display plane is matched with the phase modification characteristics through the corresponding positions of the aperture of the DOE 208a in consideration of different eye positions within the eyebox. Such display-plane-position-dependent processing of the input image I may hence result in deriving a plurality of sub-areas of the image area of the preprocessed image ldin a manner that depends on its position P within the image area (and hence on its position within the display plane).
[0058] For any position of the pupil 120a within the eyebox, light released from a certain sub-area of the optical waveguide transfers to the pupil 120a and further to the retina via a certain sub-area of the aperture of the DOE 208a, whereas the spatial correspondence between the certain sub-area of the optical waveguide 106 and the corresponding sub-area of the DOE 208a is defined via a spatial relationship between the respective elements of the NED assembly 200a. In this regard, a location of the sub-area of the aperture of the DOE 208a through which the light released from a certain sub-area of the optical waveguide 106 enters the retina depends on the position of the pupil 120a within the eyebox, on the position of the respective sub-area of the optical waveguide 106, on the (first) distance between the optical waveguide 106 and the DOE 208a, and on the distance between the DOE 208a and the pupil plane. Therefore, light transmission from the optical waveguide 106 through the DOE 208a to different pupil positions within the eyebox may be considered via a plurality of eyebox sub-areas of predefined shape and size, where each eyebox sub-area is associated with (e.g. mapped to) a respective plurality of sub-areas of the aperture of the DOE 208a and where each sub-area of the aperture of the DOE 208a further maps to a spatially corresponding sub-area of the optical waveguide 106. Moreover, due to the operating principle of the optical waveguide 106, each sub-area of the optical waveguide 106 directly maps to a spatially corresponding sub-area on the display plane of the 2D display 102a, which in turn represents a spatially corresponding image sub- area of the preprocessed image ld. Herein, the plurality of sub-areas of the aperture of the DOE 208a associated with a given position of the pupil 120a within the eyebox may be referred to as a respective plurality of sub-apertures of the DOE 208a.
[0059] In a most straightforward approach, the above-discussed spatial correspondence between image sub-areas and the sub-apertures the DOE 208a is not explicitly accounted for in the training procedure applied to derive the preprocessing procedure and the phase delaye(s, t) for the DOE 208a. However, this approach nevertheless results in the preprocessing procedure and the phase delaye(s, t) that implicitly account for the different light transmission characteristics from different positions of the optical waveguide 106 to the retina of the eye 120.
[0060] In another approach, the above-discussed spatial correspondence between image sub-areas and the sub-apertures of the DOE 208a is partially accounted for via choosing a single sub-area of the eyebox (e.g. one that is co-centered with the eyebox) and carrying out the training to derive the preprocessing procedure and the phase delaye(s, t) for the DOE 208a via expressly considering the respective plurality of sub-apertures of the DOE 208a and the corresponding (plurality of) image sub-areas that are associated with the chosen eyebox sub-area. In a further approach, the above-discussed spatial correspondence is accounted for to a larger extent via consideration of a plurality of sub-areas of the eyebox via carrying out the training to derive the preprocessing procedure and the phase delaye(s, t) for the DOE 208a via expressly considering, for each eyebox sub-area under consideration, the respective plurality of sub-apertures of the DOE 208a and the corresponding (plurality of) image sub-areas that are associated with the respective eyebox sub-area.
[0061] In this regard, the position of an eyebox sub-area (e.g. its center point) within the eyebox may be indicated as a position (x, y) within the eyebox, the location of a sub-aperture of the DOE 208a (e.g. its center point) may be indicated as a position (s, t) within the aperture of the DOE 208a, and a position of an image sub-area (e.g. its center point) within the image area may be indicated as a (pixel) position (f, q). The position (s, t) within the aperture of the DOE 208a may be, alternatively, expressed as a corresponding angle of incidence (0, of light received at the pupil plane via the position (s, t) of the aperture of the DOE 208a.
[0062] Hence, light originating from different positions of the display plane of the 2D display 102a and released from corresponding positions of the outcoupler 106b arrive at the retina of the eye 120 via different paths and via different subapertures of the DOE 208a, which results in sub-aperture-dependent light transmission characteristics from each pixel position of the 2D display 102 through the DOE 208a to the eye 120 of the viewer that further depends on the position of the pupil 120a within the eyebox. Therefore, the matching between the preprocessing portion and the phase modification characteristics of the DOE 208a implicitly results in the preprocessing procedure providing different preprocessing characteristics for different image locations of the preprocessed images lddue to variations in the phase delayc(s, t) across the aperture of the DOE 208a even in scenarios where the sub-aperture- dependency of the light transmission through the DOE 208a is not explicitly accounted for in the design procedure of the preprocessing procedure and the DOE 208a.
[0063] In scenarios where the sub-aperture-dependency of the light transmission is expressly considered in the design procedure the preprocessing procedure and the DOE 208a, these light transmission characteristics may be modeled via a display point spread function (PSF) at a reference plane in the 3D virtual image, where the PSF is dependent on the position of the respective subaperture within the aperture of the DOE 208a (e.g. on its distance to the center axis of the DOE 208a), whereas the display-plane-position-dependent preprocessing procedure applied to the input image I is arranged to account for these differences in the light transmission characteristics in order to convey 3D virtual images that are perceived sharp at the plurality of accommodation depths within the depth range [z+, z~] of the NED assembly 200a. Such a match between the preprocessing procedure and the DOE 208a contributes towards keeping a broad field of view (FoV) while providing the EDoF.
[0064] In some examples, the preprocessing procedure and the DOE 208a may be jointly designed via consideration of a plurality of eyebox sub-areas and a corresponding plurality of sub-apertures of the DOE 208a for each eyebox subarea. In this regard, each eyebox sub-area may have a shape and size that approximate those of a typical (e.g. average) pupil 120a and they jointly cover the eyebox in its entirety without any gaps therebetween either in a nonoverlapping or in a partially overlapping manner. In this regard, the eyebox sub-areas preferably have a substantially circular shape, whereas in some examples sub-areas of hexagonal or rectangular shape may be applied instead. The sub-apertures of the DOE 208a have a shape and size that typically follow either those of the entrance pupil (the incoupler 106a) or those of the eyebox sub-areas, whichever is more limiting (i.e. smaller) in terms of size.
[0065] Along the lines described in the foregoing, the sub-apertures of the DOE 208a at least conceptually map to corresponding sub-areas of the outcoupler 106b (and to the corresponding sub-groups of pixels of the display plane of the 2D display 102a), whereas the sub-areas of the outcoupler 106a further map to corresponding sub-areas of an image area of the preprocessed image ld(i.e. image sub-areas of the preprocessed image ld). With the preprocessing procedure and the phase delay profile < / e(s, t) of the DOE 208a designed in consideration of the plurality of eyebox sub-areas and the corresponding plurality of sub-apertures of the DOE 208a that map to the corresponding plurality of image sub-areas of the preprocessed image ld, the resulting preprocessing procedure implicitly accounts for different light transmission characteristics through different sub-apertures of the DOE 208a. Consequently, the preprocessing applied to the input image I to derive a certain image sub-area of the preprocessed image / dmay be different from one image sub-area to another, thereby resulting in the preprocessing procedure that may involve providing different preprocessing characteristics for derivation of different image sub-areas of the preprocessed image ld.
[0066] The preprocessing procedure may be provided, for example, via application of at least one artificial neural network (ANN) arranged (e.g. trained) to process the input image I into the corresponding preprocessed image ldin a manner that accounts for different pupil positions within the eyebox and for the corresponding plurality of sub-apertures of the DOE 208a, thereby resulting in the display-plane-position-dependent preprocessing procedure. As nonlimiting examples in this regard, the at least one ANN employed by the image preprocessing portion 232 to implement the preprocessing procedure may be provided as at least one convolutional neural network (CNN) or as at least one derivative of a CNN. However, the CNN (or a derivative thereof) serves as a non-limiting example of an applicable ANN and in other examples an ANN of different kind may be employed instead without departing from the scope of the NED assembly 200a according to the present disclosure.
[0067] According to an example of applying the at least one ANN, the different processing for different image sub-areas of the preprocessed image / dmay be provided via a single ANN that is trained to separately derive each image subarea of the preprocessed image ldin dependence of its position within the display plane. Training of such an ANN is described via examples provided in the following. In an example, such an ANN may take the input image I and an indication of the image sub-area of interest as input and provide the corresponding image sub-area of the preprocessed image ldas output. In a variation of this approach, the input to the single ANN may further comprise depth information pertaining to the image sub-area under consideration. Consequently, the preprocessed image ldin its entirety may be derived via applying the single ANN separately to each image sub-area of the input image / to derive the corresponding image sub-area of the preprocessed image / dand combining the image sub-areas so obtained into the preprocessed image ld.
[0068] According to another example of applying the at least one ANN, the different processing for the different image sub-areas the corresponding preprocessed image ldmay be provided via application of a plurality of ANNs, each trained to derive a respective image sub-area of the preprocessed image ld. Training of such plurality of ANNs is described via examples provided in the following. In an example, each of such ANNs may take the respective image sub-area of input image I as input and provide the corresponding image sub-area of the preprocessed image ldas output. In a variation of this approach, the input to each of the ANNs may further comprise depth information pertaining to the sub-area processed by the respective ANN. Consequently, the preprocessed image ldin its entirety may be derived via applying the plurality of ANNs to the corresponding image sub-areas of the input image I to derive the corresponding image sub-areas of the preprocessed image / dand combining the image sub-areas so obtained into the preprocessed image ld.
[0069] The above-discussed spatial correspondence between the plurality of subapertures of the DOE 208a and the corresponding plurality of sub-areas of the optical waveguide 106 also extends to spatially corresponding sub-areas of the second DOE 210a: for a given eyebox sub-area, each of the plurality of sub-apertures of the DOE 208a map to a spatially corresponding sub-area of the optical waveguide 106, which further maps to a spatially corresponding sub-area of the aperture of the second DOE 210a. The sub-areas of the aperture of the second DOE 210a may be also referred to as respective subapertures of the second DOE 210a. This extended spatial correspondence may be accounted for in the training procedure applied to derive the phase delayc(s, t) for the second DOE 210a together with the preprocessing procedure and the phase delaye(s, t) for the DOE 208a in a manner described above for joint training of the preprocessing procedure and the phase delaye(s, t) for the DOE 208a, mutatis mutandis. As in case of the DOE 208a, the location of a sub-aperture of the second DOE 210a (e.g. its center point) may be indicated as a position (s, t) within the aperture of the second DOE 210a, while it may be, alternatively, expressed as a corresponding angle of incidence (0, of light received at the pupil plane via the position (s, t) of the aperture of the second DOE 210a. Figure 3B illustrates a block diagram of some components of a display controller 230 which is applicable for preprocessing the input image I into the preprocessed image ld. Along the lines described in the foregoing, the display controller 230 may be provided as an element of the NED assembly 200a and / or as an element of the NED apparatus making use of the pair of NED assemblies 200a. The display controller 230 comprises a control portion 231 and the image preprocessing portion 232, where the image preprocessing portion 232 implements the preprocessing procedure. The control portion 231 may control at least some aspects of operation of the image preprocessing portion 232 and it may also control one or more aspects of operation of the respective image projection units 102 of the two NED assemblies 200a of the stereoscopic NED apparatus in terms of projecting the light that represents the respective preprocessed images ldtowards the incouplers 106a of the respective optical waveguides 106 of the two NED assemblies 200a of the NED apparatus. Operation of the control portion 231 may be at least partially controlled via control input provided thereto e.g. via a user interface the NED apparatus.
[0070] According to an example, the display controller 230 may be implemented via operation of a computing apparatus that comprises one or more processors and one or more memories for storing one or more computer programs, where the one or more computer programs are arranged to cause the computing apparatus to operate as the display controller 230 according to the present disclosure when executed by the one or more processors. Further details of using a computing apparatus for implementing the display controller is provided in the following with references to Figure 10.
[0071] The present disclosure predominantly refers to the preprocessing procedure being applied to process a single input image I to derive a corresponding preprocessed image ld. This is, however, a choice made in the interest of brevity and editorial clarity of the description, whereas in the course of its operation the display controller 230, e.g. the image preprocessing portion 232 therein, may receive a time series of stereoscopic image pairs and derive a corresponding time series of preprocessed image pairs based on the time series of stereoscopic image pairs. In particular, the image processing portion 232 may be applied to receive a respective time series of input images l(t) for each of the two NED assemblies 200a of the NED apparatus and to derive, based on the respective received time series of input images l(t), a respective corresponding time series of preprocessed images ld(t) to be supplied for the respective image projection units 102 of the two NED assemblies 200a.
[0072] In this regard, any single image of the time series of input images l(t) may be referred to as the input image I, whereas any single image of the time series of preprocessed images ld(t) may be referred to as the corresponding preprocessed image ld. Further in this regard, the input image I refers to pixel values of the underlying input image, whereas the preprocessed image ldrefers to pixel values of the underlying preprocessed image. In a non-limiting example, each of the input image / and the corresponding preprocessed image ldmay comprise a respective RGB image that defines respective pixel values for each pixel position of the respective image I, ldseparately in red, green and blue color channels. In some examples, the time series of stereoscopic image pairs may be accompanied by a corresponding time series of depth maps D(t), i.e. each stereoscopic image pair of the time series may be accompanied by a corresponding depth map D that provides respective depth information for each pixel of the corresponding stereoscopic image pair. The image preprocessing portion 232 may apply the depth information received in the time series of depth maps D(t) in the process of deriving the corresponding time series of preprocessed images ld(t) based on the time series of input images l(t).
[0073] As described in the foregoing, the preprocessing procedure may be provided via application of at least one ANN, whereas weights of the at least one ANN that implement the preprocessing procedure may be determined via a training procedure carried out prior to application of the preprocessing procedure in the image preprocessing portion 232. The match between the preprocessing characteristics provided by the preprocessing procedure and the phase delay characteristics of the DOE 208a may be provided via modeling and optimizing the phase delaye(s, t) as part of the training procedure applied for determining the weights of the ANN that serves to provide the preprocessing procedure. Along similar lines, the match between respective optical characteristics of the DOE 208a and the second DOE 210a may be provided via modeling and optimizing the respective phase delayse(s, t) andc(s, t) in a coordinated manner as part of the training procedure. The training procedure may be also referred to as a learning procedure, whereas the examples described in the following predominantly apply the term learning procedure.
[0074] The learning procedure relies on supervised learning carried out based on training images processed through a first learning model that represents the at least one ANN serving as the preprocessing procedure and the light transmission through the DOE 208a to the retina of the eye 120 and through a second learning model that represents the light transmission through the second DOE 210a, through the optical waveguide 106 and through the DOE 208a to the retina of the eye 120. Moreover, the training images are also processed by an eye model that represents one or more optical limitations of the eye 120 to determine the respective ground-truth (GT) images for the supervised learning procedure. The same or similar eye model may be also applied as part of the first and second learning models. The first and second learning models together with the eye model may be collectively referred to as a learning arrangement that represents the NED assembly 200a and that is applicable for jointly determining the weights of the at least one ANN, the phase delay profilee(s, t) of the DOE 208a and the phase delay profilec(s, t) of the second DOE 210a. Along the lines described in the foregoing, the phase delay profilese(s, t) andc(s, t) are those defined for the reference wavelength, e.g. for a wavelength corresponding to a predefined one of red, green and blue channels.
[0075] In particular, the learning arrangement may be applied to carry out the learning procedure as an iterative procedure that relies on a training dataset that includes a plurality of training images / f, typically in the order of thousands of images, depicting natural scenes that may be chosen in view of intended usage of the NED assembly 200a and / or the NED apparatus making use of the NED assembly 200a. According to an example, prior to starting the iterative learning procedure, the weights of the at least one ANN and the phase delay profiles < / >e(s, t) and < / >c(s, t) may be initialized using (pseudo) random values, whereas in another example the initial values for the weights of the at least one ANN and the phase delay profiles < / >e(s, t) and < / >c(s, t) may comprise respective values obtained via another learning procedure carried out earlier.
[0076] Each training image / fof the training dataset defines respective pixel values for a plurality of color channels. In one example, each training image / fcomprises a respective RGB image that defines respective pixel values for each pixel position of the respective training image / fseparately in red, green and blue color channels. In another example, each training image / fcomprises a respective hyperspectral image that defines respective pixel values for each pixel position of the respective training image / fseparately in a plurality of predefined color channels within visible light spectrum. In some examples, the training dataset further comprises a respective depth map Df, for each of the training images / f, which may be provided as input to the learning model 300 together with the corresponding training image / fin the course of the learning procedure.
[0077] The learning procedure may be carried out, for example, via operation of a computing apparatus that comprises one or more processors and one or more memories for storing one or more computer programs, where the one or more computer programs are arranged to cause the computing apparatus to carry out the learning procedure described in the present disclosure when executed by the one or more processors. According to another example, the learning procedure may be carried out via usage of a plurality of (i.e. two or more) computing apparatuses of the kind described above, mutatis mutandis, where the plurality of computing apparatuses may be arranged to provide a cloud computing service. Figure 4A illustrates the eye model 302, which may be applied to process each training image / finto a corresponding reference retinal image lrthat models direct (or natural) viewing of the respective training image / fand that serves as the GT image corresponding to the respective training image / f. The reference retinal image lris derived in consideration of the accommodation depth z selected for the respective iteration round. The eye model 302 may be arranged to account for one or more optical limitations of the eye, for example chromatic aberrations of (typical) eye optics and / or the diffraction-limited resolution of the eye 120 due to a finite size of the pupil 120a. This may be accomplished, for example, via convolution of the training image with a diffraction-limited PSF that represents the chromatic aberrations of the eye 120 positioned within the eyebox.
[0078] Along the lines described in the foregoing, the first learning model of the learning arrangement models the at least one ANN and the phase delay profile < / >e(s, t) of the DOE 208a and the first learning model may be applicable for processing each training image / finto a corresponding simulated retinal virtual image / rv, whereas the first learning model may further determine a respective first error measure Evthat is descriptive of a difference between simulated retinal virtual image / rvand the corresponding reference retinal image lr. Figure 4B illustrates an example of the first learning model, which comprises an ANN 304 that represents the at least one ANN serving as the preprocessing procedure and a virtual display and eye model 306 that represents the light transmission from the optical waveguide 106 through the DOE 208a to the retina of the eye 120. Application of the first learning model involves the following aspects:
[0079] - The ANN 304 processes each training image / finto a corresponding preprocessed image ldusing current weights of the ANN 304. In this regard, the ANN 304 derives pixel values of the preprocessed image ldbased on pixel values of the corresponding training image / fvia usage of the current weights of the ANN 304. Along the lines described in the foregoing, the applied ANN may comprise e.g. a CNN or a modified CNN.
[0080] - The virtual display and eye model 306 processes each preprocessed image ldinto a corresponding simulated retinal virtual image / rvat the reference plane in accordance with current phase delay profilee(s, t). In this regard, the virtual display and eye model 306 derives pixel values of the simulated retinal virtual image / rvbased on pixel values of the corresponding preprocessed image ldaccording to the current phase delay profilee(s, t) in consideration of the accommodation depth z selected for the respective iteration round. The virtual display and eye model 306 may comprise a differentiable simulation model that simulates transmission of light that represents the preprocessed image / dfrom the outcoupler 106b through the DOE 208a in consideration of the current phase delay profilee(s, t) together with the eye model of the kind described in the foregoing with references to the eye model 302.
[0081] - A first loss function 308 determines a respective error measure Evbetween each simulated retinal virtual image and the corresponding retinal reference image lr. The error measure Evmay be determined as an objective error measure such as the L1 -loss or the mean squared error (MSE) derived based on pixel-wise difference between the images / rvand lrunder consideration or as a subjective error measure such as the structural similarity measure (SSIM) that aims at predicting perceived error (or difference) between the images / rvand lrunder consideration.
[0082] Along the lines described in the foregoing, the second learning model of the learning arrangement models the phase delay profilec(s, t) of the second DOE 210a, the light transmission through the optical waveguide 106 and the phase delay profilee(s, t) of the DOE 208a and the second learning model may be applicable for processing each training image / finto a corresponding simulated retinal real-world image / rr, whereas the second learning model may further determine a respective second error measure Erthat is descriptive of a difference between simulated retinal real-world image / rrand the corresponding reference retinal image lr. Figure 4C illustrates an example of the second learning model, which comprises a see-through display and eye model 310 that represents the light transmission through the second DOE 210a, the optical waveguide 106 and the DOE 208a to the retina of the eye 120. Application of the second learning model involves the following aspects:
[0083] - The see-through display and eye model 310 processes each input image / finto a corresponding simulated retinal real-world image / rrat the reference plane in accordance with current phase delay profilese(s, t) andc(s, t). In this regard, the see-through display and eye model 310 derives pixel values of the simulated retinal real-world image / rrbased on pixel values of the corresponding training image / faccording to the current phase delay profilese(s, t) andc(s, t) in consideration of the accommodation depth z selected for the respective iteration round. The see-through display and eye model 310 may comprise respective differentiable simulation models that model transmission of external light through the second DOE 210a in consideration of the current phase delay profilec(s, t) and transmission of the external light received through the second DOE (and through the optical waveguide 106) through the DOE 208a in consideration of the current phase delay profilee(s, t) together with the eye model of the kind described in the foregoing with references to the eye model 302.
[0084] - A second loss function 312 determines a respective error measure Erbetween each simulated retinal real-world image / rrand the corresponding retinal reference image lr. Like the error measure Ev, also the error measure Ermay be determined as an objective error measure such as the L1 -loss or the mean squared error (MSE) derived based on pixel-wise difference between the images / rrand lrunder consideration or as a subjective error measure such as the structural similarity measure (SSIM) that aims at predicting perceived error (or difference) between the images / rrand lrunder consideration. In the following, two exemplifying scenarios for using the first and second learning models to determine the preprocessing procedure, the phase delay profilee(s, t) of the DOE 208a and the phase delay profile < / >c(s, t) of the second DOE 210a for the NED assembly 200a are described.
[0085] In a first scenario, at each iteration round of the learning procedure, each training image / fis processed through the eye model 302 into the corresponding reference retinal image lr(i.e. the corresponding GT image) before or as part of the first iteration round of the learning procedure. Moreover, at each iteration round, each training image I1is processed through the first training model to determine the corresponding error measure Evand through the second training model to determine the corresponding error measure Er. Moreover, for each training image / f, the respective error measures Evand Erare combined into a joint error measure E, which may be derived e.g. as a weighted sum of the respective error measures error measures E = aEv+ (1- a)Er, where a has a predefined value in range between 0 and 1 , e.g. 0.5.
[0086] Continuing with the first scenario, at the end of each iteration round, the weights of the ANN 304 and the phase delay profiles < / >e(s, t) and < / >c(s, t) are updated based on the respective combined error measures E determined for the plurality of training images / f. As an example in this regard, the updating may be carried out via usage of a gradient descent method known in the art. The delay profile < / >e(s, t) is considered in both the first and second training models and it may be updated for both the first and second training models based on a weighted average of the respective partial derivatives dE I da>eobtained in the first and second training models to ensure keeping the respective instances of the delay profile <Pes, t) appearing in the first and second training models aligned with each other.
[0087] Still referring to the first scenario, the iterative learning procedure may be carried out until one or more predefined convergence criteria that pertain to the combined error measures E are met and / or until a predefined number of iteration rounds have been carried out. Once the learning procedure has been completed, the weights of the at least one ANN 304 at the end of the final iteration round may be adopted as the weights of the at least one ANN applied to implement the preprocessing procedure in the image preprocessing portion 232 and the phase delay profilese(s, t) andc(s, t) at the end of the final iteration round may be adopted as the respective phase delay profiles to be implemented by the DOE 208a and the second DOE 210a, respectively.
[0088] In a second scenario, the learning procedure is carried out in two phases. At each iteration round of the first phase of the learning procedure, each training image / fis processed through the eye model 302 into the corresponding reference retinal image lr(i.e. the corresponding GT image) and each training image / fis processed through the first training model to determine the corresponding error measure Ev. At the end of each iteration round of the first phase, the weights of the ANN 304 and the phase delay profilee(s, t) are updated, e.g. via usage of the gradient descent method, based on the respective error measures Evdetermined for the plurality of training images / f. The criterion for completing the first phase of iterative learning procedure may be similar to that applied in the first scenario, mutatis mutandis, whereas the weights of the at least one ANN 304 at the end of the final iteration round of the first phase may be adopted as the weights of the at least one ANN applied to implement the preprocessing procedure in the image preprocessing portion 232 and the phase delay profilee(s, t) at the end of the final iteration round of the first phase may be adopted as the respective phase delay profile to be implemented by the DOE 208a.
[0089] Continuing with the second scenario, at each iteration round of the second phase of the learning procedure, each training image / fis processed through the eye model 302 into the corresponding reference retinal image lr(i.e. the corresponding GT image) and each training image / fis processed through the second training model to determine the corresponding error measure Er. At the end of each iteration round of the second phase, the phase delay profilec(s, t) is updated, e.g. via usage of the gradient descent method, based on the respective error measures Erdetermined for the plurality of training images / f. In contrast, the delay profilee(s, t) obtained at the end of the first phase is applied throughout the second phase and it is not updated at the end of iteration rounds of the second phase. Termination criteria of the second phase may be similar to that described above for the first scenario, mutatis mutandis, whereas the phase delay profilec(s, t) at the end of the final iteration round of the second phase may be adopted as the respective phase delay profile to be implemented by the second DOE 210a.
[0090] The first and second scenarios for using the first and second learning models of the learning arrangement implicitly assume an approach where each iteration round of the iterative learning procedure considers the plurality of training images / fin their entirety. In another example, the learning procedure may be carried out such that for each iteration round one or more subapertures of the DOE 208a and the spatially corresponding sub-apertures of the second DOE 210a are selected for consideration. In this regard, one of two approaches envisaged above may be applied:
[0091] - In a first approach, a (single) predefined sub-area of the eyebox is assumed, whereas the learning procedure is carried out in consideration of the plurality of sub-apertures of the first DOE 208a associated with the predefined eyebox sub-area. In this regard, the predefined eyebox sub-area may be one that is co-centered with the eyebox area, whereas one or more of the plurality of sub-apertures of the DOE 208a associated with the predefined eyebox sub-area are selected for consideration at a given iteration round.
[0092] - In a second approach, a plurality of sub-areas of the eyebox are assumed, whereas the learning procedure selects, for each iteration round, one of the plurality of eyebox sub-areas for consideration at the respective iteration round. Moreover, the respective iteration round is carried out in consideration of the plurality of sub-apertures of the first DOE 208a associated with the eyebox sub-area selected for consideration in the respective iteration round via selecting one or more of the plurality of sub-apertures of the first DOE 208a for consideration in the respective iteration round.
[0093] Along the lines described in the foregoing, the one or more sub-apertures of the DOE 208a selected for the given iteration directly map to the spatially corresponding sub-apertures of the second DOE 210a and they are likewise chosen for consideration in the respective iteration round. Consequently, only those sub-areas of the training images / fthat spatially correspond to the selected one or more sub-apertures of the DOE 208a and the second DOE 210a are processed through the first and second training models and the respective error measures Evand the Erare derived only in consideration of the one or more image sub-areas that spatially correspond to the selected one or more sub-apertures of the DOE 208a and the second DOE 208b. The information that identifies the one or more sub-apertures of the DOE 208a and the second DOE 208b under consideration at a given iteration round and / or the corresponding image sub-areas may be input e.g. to elements of the learning arrangement, e.g. to the ANN 304, to the virtual display and eye model 306 and to the see-through display and eye model 310.
[0094] The image sub-areas under consideration at the given iteration round may be denoted as Pn, whereas due to the above-discussed spatial relationships between the image area (pixel) positions (f, q) and the positions (s, t) of the respective apertures of the DOE 208a and the second DOE 210a, this information may be alternatively defined via the corresponding angle of incidence (0n, of light received at the pupil plane via the corresponding position ( , qn) of the optical waveguide 106 that indicates e.g. a center point of the respective sub-area of the optical waveguide 106 and that directly maps to the corresponding pixel positions ,n) of the image area. The position of the sub-area of the eyebox under consideration may be defined via the corresponding position (xn, yn) that indicates e.g. the center point of the eyebox sub-area under consideration. These indications of the spatial positions within the image area, within the respective aperture of the DOE 208a and the second DOE 210a and within the eyebox are input to the models applied as part of the learning arrangement to extent they are necessary in view of the applied manner of carrying out the training procedure.
[0095] The sub-apertures of the DOE 208a and the second DOE 208b for consideration in a given iteration round may be selected, for example, (substantially) randomly or in accordance with a predefined rule from the plurality of predefined sub-apertures of the DOE 208a and the second DOE 208b. In this regard, as discussed in the foregoing, the plurality of subapertures that are available for consideration in the learning procedure may have respective predefined positions within the respective apertures of the DOE 208a and the second DOE 210a, they may have a predefined shape and size and that approximate those of a typical (e.g. an average) pupil 120a, and they may jointly cover the respective apertures of the DOE 208a and the second DOE 210a substantially in their entirety.
[0096] The examples pertaining to the learning procedure provided in the foregoing refer to learning the one or more ANNs modeled by the ANN 304 in general. In various examples, this may involve learning a single ANN or learning a plurality of ANNs. In one example, the ANN 304 may be employed to learn a single ANN for processing the entire image area of the input image / f. In another example, the ANN 304 may be employed to learn a single ANN for processing any one of a plurality of image sub-areas of the input image I in dependence of its position in the image area. In a further example, the ANN 304 may be employed to learn a respective ANN for a plurality of image subareas of the input image / .
[0097] Along the lines described in the foregoing, each training image / fof the training dataset defines respective pixel values for a plurality of color channels and the training images / fmay be provided e.g. as respective RGB images or as respective hyperspectral images, whereas application of the training network is described above with references to the training images / fin general. In case of using RGB images as the training images / f, each color channel may be processed through the learning arrangement separately from each other. In case of using hyperspectral images as the training images / f, all color channels may be considered in the eye model 302 and in the second learning model, whereas only those color channels of hyperspectral images that correspond to respective wavelengths of red, green and blue colors may be considered in the first learning model. In this regard, the RGB images or a corresponding color channels of the hyperspectral images are well suited as training data for the ANN 304 and the virtual display and eye model 306, whereas hyperspectral images may provide an improved model of real-world scenes for the see- through display and eye model 310.
[0098] Along the lines described in the foregoing, each of the eye model 302, the virtual display and eye model 306 and the see-through display and eye model 310 consider the accommodation depth z selected for a given iteration round. In a scenario that aims at providing a continuous EDoF discussed in the foregoing, the accommodation depth z considered in a given iteration round of the learning procedure may be chosen, for example, (substantially) randomly from the depth range [z+, r], whereas in another example in this regard the accommodation depth to be applied for the given iteration round may be varied from one iteration round to another via choosing the accommodation depth z from the depth range [z+, z~] according to a predefined rule such that depth range [z+, z~] is covered in a substantially uniform manner (in the dioptric domain) in the course of the learning procedure. In a scenario that aims at providing a multifocal EDoF discussed in the foregoing, the accommodation depth z considered in a given iteration round of the learning procedure may be chosen from a set of predefined accommodation depths within the depth range [z+, z~] substantially randomly or according to a predefined rule.
[0099] In the course of the learning procedure, the virtual display and eye model 306 considers the accommodation depth z as the desired accommodation depth of the 3D virtual image to be displayed via operation of the optical waveguide 106 based on the respective training image / f, whereas each of the eye model 302 and the see-through display and eye model 310 consider the depth z as an actual (focal) depth of a real-world object represented by the respective training image / f. In some examples, the learning procedure may further consider accommodation depths z outside the depth range [z+, z~] in a subset of iteration rounds. In iteration rounds for which the accommodation depth z is outside the depth range [z+, r], the first training model may be omitted and at the end of such an iteration round the delay profilese(s, t) andc(s, t) may be updated based on the second training model only.
[0100] Figure 5 schematically illustrates some aspects of an accommodation invariant NED assembly 200b according to a second embodiment, whereas the illustration of Figure 5 further shows the eye 120 of the viewer positioned within the eyebox at the pupil plane. The NED assembly 200b is similar to the NED assembly 200a described in the foregoing except for a DOE 208b replacing the DOE 208a and for omission of the second DOE 210a altogether. Hence, the DOE 208b is arranged on the first side of the optical waveguide 106 between the optical waveguide 106 and the eyebox. The discussion regarding the spatial arrangement of the DOE 208a in relation to the optical waveguide 106 provided in the following also applies to the spatial arrangement of the DOE 208b in relation to the optical waveguide 106 as well, mutatis mutandis. Omission of the second DOE 210a requires and results in optical characteristics of the DOE 208b being different from those applied in the DOE 208a of the NED assembly 200a due to the DOE 208b of the NED assembly 200b being responsible both for extending the DoF of the virtual images displayed via the optical waveguide 106 and for ensuring perceivable quality and clarity of the view to the real-world through the optical waveguide 106.
[0101] Since the NED assembly 200b does not apply the second DOE 210a at the second side of the optical waveguide 106, the view through the NED assembly 200b to the real world is also involves a similar extension of DoF as applied to the virtual image that represents the AR content illustrated in the input image / . While in some scenarios this may be considered as a shortcoming of the NED assembly 200b in comparison to the NED assembly 200a, this nevertheless enables a more compact design of the NED assembly 200b, thereby contributing towards improved user comfort via reducing the size and weight of the NED assembly 200b and the NED apparatus making use of a pair of NED assemblies 200.
[0102] Moreover, the extended DoF provided for the real-world objects may also be applied as an advantage in terms of providing vision correction for near-sighted and / or far-sighted users via designing the preprocessing procedure and the DOE 208b to provide a suitable depth range [z+. z~] that compensates for nearsightedness and / or short-sightedness in a desired manner. As a concrete example in this regard, the extended depth range [z+, z~] may span from 0 diopters (i.e. infinity) to 3 diopters to provide vision correction for near-sighted persons that require correction up to -3 diopters, whereas in another example the extended depth range [z+, z~] may span from -2 diopters to 2 diopters in order to provide vision correction for both near-sighted and far-sighted persons that require correction in a range from -2 to 2 diopters.
[0103] Like in the NED assembly 200a described in the foregoing via various examples, also in the NED assembly 200b the preprocessing procedure and the phase modulation function implemented by the DOE 208b are matched to each other to enable providing the extended DoF that spans over the predefined depth range [z+, r]. The iterative learning procedure described in the foregoing for the NED assembly 200a is also applicable for determining the weights of the at least one ANN that serves as the preprocessing procedure and the phase delay profile < / >e(s, t) of the DOE 208b with the eye model 302, the first learning model and the second learning model replaced with the respective models described in the following with references to Figures 6A, 6B and 6C.
[0104] Figure 6A illustrates an example of the eye model 302’, which may be applied to process each training image / finto a corresponding reference retinal image f that models direct (or natural) viewing of the respective training image / fand that serves as the GT image corresponding to the respective training image / f. The reference retinal image lris derived in consideration of the accommodation depth z selected for the respective iteration round and further in consideration of a focal depth z0selected for the respective iteration round, whereas other characteristics of the eye model 302’ are similar to those of the eye model 302.
[0105] Figure 6B illustrates an example of a modified first learning model, which comprises the ANN 304’ that represents the at least one ANN serving as the preprocessing procedure and a virtual display and eye model 306’ that represents the light transmission from the optical waveguide 106 through the DOE 208b to the retina of the eye 120.
[0106] Figure 6C illustrates an example of a modified second learning model, which comprises a see-through display and eye model 310’, which derives pixel values of the simulated retinal real-world image / rrbased on pixel values of the corresponding training image / faccording to the current phase delay profile e(s, t) in consideration of the accommodation depth z selected for the respective iteration round and further in consideration of the focal depth z0selected for the respective iteration round. Hence, the difference to the see- through display and eye model 310 is that also the selected focal depth z0is considered in finding the phase delay profile < / >e(s, t), whereas the phase delay c(s, t) is not considered at all due to absence of the second DOE 210a from the NED assembly 200b.
[0107] Hence, each of the eye model 302’ and the see-through display and eye model 310’ consider the accommodation depth z and the focus depth z0selected for a given iteration round. The accommodation depth z for the given iteration round may be chosen as described in the foregoing for the first embodiment, whereas the focal depth z0for the given iteration round may be chosen in dependence of the accommodation depth z chosen for the same iteration round. In one example, the focal depth z0considered in the given iteration round may be chosen substantially randomly from a predefined range of distances, whereas in another example the selection may be carried out according to a predefined rule e.g. such that the predefined range is covered in a substantially uniform manner and / or that both scenarios where the focal depth z0is different from the accommodation depth z and scenarios where the focus depth z0is the same as the focus depth z0are covered. The eye model 302’ and the modified first and second learning models may be applied in various ways to determine the preprocessing procedure and the phase delay profilee(s, t) of the DOE 208b. A first scenario for making use of these models is similar to the first scenario described in the foregoing in context of the first embodiment apart from applying and updating the phase delay profilec(s, t), mutatis mutandis.
[0108] In a second scenario, the learning procedure may be carried out in two phases. At each iteration round of the first phase of the learning procedure, each training image / fis processed through the eye model 302’ into the corresponding reference retinal image / r(i.e. the corresponding GT image) and each training image / fis processed through the modified first and second learning models to determine the corresponding error measures Evand Erwith the modified first learning model further modified such that the ANN 304 is omitted and each training image / fis directly passed to the virtual display and eye model 306’. At the end of each iteration round of the first phase, the phase delay profilee(s, t) is updated, e.g. via usage of the gradient descent method, based on the respective error measures Evand Erdetermined for the plurality of training images / f. The criterion for completing the first phase of iterative learning procedure may be similar to that applied in first embodiment, mutatis mutandis, whereas the phase delay profilee(s, t) at the end of the final iteration round of the first phase may be adopted as the respective phase delay profile to be implemented by the DOE 208b.
[0109] At each iteration round of the second phase, each training image is processed through the eye model 302’ into the corresponding reference retinal image lr(i.e. the corresponding GT image) and each training image / fis processed through the modified first learning model to corresponding error measure Ev. At the end of each iteration round of the second phase, the weights of the ANN 304 are updated, e.g. via usage of the gradient descent method, based on the respective error measures Evdetermined for the plurality of training images / f. Termination criteria of the second phase may be similar to that described above for the first embodiment, mutatis mutandis, whereas the weights of the at least one ANN 304 at the end of the final iteration round of the first phase may be adopted as the weights of the at least one ANN applied to implement the preprocessing procedure in the image preprocessing portion 232.
[0110] Figure 7 schematically illustrates some aspects of an accommodation invariant NED assembly 200c according to a third embodiment, whereas the illustration of Figure 7 further shows the eye 120 of the viewer positioned within the eyebox at the pupil plane. The NED assembly 200c is similar to the NED assembly 200b described in the foregoing, with the DOE 208b replaced with a DOE 208c and complemented with the first lens 108 and the second lens 110 of the kind described in the forgoing in context of the NED assembly 100’. Hence, the first lens 108 is a negative lens (concave lens, diverging lens) that serves to change the focal depth of the virtual image to a desired vergence depth that depends on optical characteristics of the first lens 108, whereas the second lens 110 is a positive lens (convex lens, converging lens) that serves as a corrective lens that compensates for the optical power of the first lens 108: since the first lens 108 also modifies the focal depth of the real-world objects visible through the optical waveguide 106 in the same manner as it modifies the focal depth of the virtual image displayed via the optical waveguide 106, the second lens 110 is applied to provide the viewer with a substantially undistorted view to the real world through the optical components of the NED assembly 200c.
[0111] Like the DOE 208b of the NED assembly 200b, the DOE 208c of the NED assembly 200c is arranged to both extend the DoF of the virtual images displayed via the optical waveguide 106 and to ensure perceivable quality and clarity of the view to the real-world through the optical waveguide 106, whereas the lenses 106 and 108 serve to bring the focal depth of the virtual images closer to the viewer without substantially changing the view to the real-world objects through the NED assembly 200c. In this regard, the optical design of the DOE 208c is similar to that of the DOE 208b due to the each of the DOE 208b and the DOE 208c being responsible both for extending the DoF of the virtual images displayed via the optical waveguide 106 and for ensuring perceivable quality and clarity of the view to the real-world through the optical waveguide 106. However, the DOE 208b and the preprocessing procedure designed for the NED assembly 200b are not directly applicable to serve as the corresponding elements of NED assembly 200c due to the optical path from the optical waveguide 106 through the DOE 208c to the eye 120 of the viewer being different from the optical path from the optical wave guide 106 through the DOE 208b, which also results in different mapping from subapertures of the DOE 208c to the corresponding spatial positions (f, q) of the optical waveguide 106 and the spatially corresponding pixel positions ( q) of the display plane of the 2D display 102a.
[0112] As in the case of the NED assembly 200b, also in the NED assembly 200c the extended DoF provided for the real-world objects may also be applied as an advantage in terms of providing vision correction for near-sighted and / or farsighted users via designing the preprocessing procedure and the DOE 208b to provide a suitable depth range [z+.z~] that compensates for near-sightedness and / or short-sightedness in a desired manner.
[0113] Nevertheless, the learning arrangement and the iterative learning procedure described in the foregoing for the NED assembly 200b are also applicable for determining the weights of the at least one ANN that serves as the preprocessing procedure implemented by the image preprocessing portion 232 and the phase delay profile < / >e(s, t) of the DOE 208c for the NED assembly 200c, when the different optical path from the optical waveguide 106 to the retina of the eye 120 of the viewer is accounted for in the virtual display and eye model 306’ and in the see-through display and eye model 310’.
[0114] Figure 8 schematically illustrates some aspects of an accommodation invariant NED assembly 200d according to a fourth embodiment together with the eye 120 of the viewer positioned within the eyebox at the pupil plane. The NED assembly 200d is similar to the NED assembly 200b described in the foregoing except for the manner of implementing the DOE: while the NED assembly 200b includes the DOE 208b arranged on the first side of the optical waveguide 106 and provided separately therefrom, in the NED assembly 200d a DOE 208d is integrated to the outcoupler 106b of the optical waveguide 106.
[0115] In this design, the functionality of the DOE 208d only serves to modify the light received in the optical waveguide via the incoupler 106a while the functionality of the DOE 208d does not modify optical characteristics of the external light received through optical waveguide 106. Therefore, in the NED assembly 200d the DOE 208d is only responsible for extending the DoF of the virtual images displayed via the optical waveguide 106 and hence has a role that is similar to the DOE 208a of the NED assembly 200a. However, due to integration of the DOE 208d to the outcoupler 106b, the light transmission path from the optical waveguide 106 towards the eyebox is different from that of the NED assembly 200b and, consequently, the optical characteristics of the DOE 208d in the NED assembly 200d are different from those of the DOE 208b in the NED assembly 200b.
[0116] Despite differences in placement of the DOE 208d with respect to the optical waveguide 106 and the role of the DOE 208d, the learning arrangement and the iterative learning procedure(s) described in the foregoing for the NED assembly 200b are also applicable for determining the weights of the at least one ANN that serves as the preprocessing procedure implemented by the image preprocessing portion 232 and the phase delay profilee(s, t) of the DOE 208d for the NED assembly 200d, apart from consideration of the modified second learning model (due to the assumption of the DOE 208d not modifying the external light received through the optical waveguide 106). Moreover, the virtual display and eye model 306’ in the learning arrangement needs to be modified to account for the different optical path from the optical waveguide 106 to the retina of the eye 120 of the viewer.
[0117] In a variation of the NED assembly 200d, optical characteristics corresponding to those of the lens 108 applied in the NED assembly 200c are further integrated to the outcoupler 106b of the optical waveguide 106, whereas this variation further involves the second lens 110 arranged at the second side of the optical waveguide as in the NED assembly 200c, mutatis mutandis. This variation of the NED assembly 200d further requires accounting for the optical characteristic of the lens 108 in the iterative learning procedure carried out on basis of the modified first learning model.
[0118] Figure 9 schematically illustrates some aspects of an accommodation invariant NED assembly 200e according to a fifth embodiment together with the eye 120 of the viewer positioned within the eyebox at the pupil plane. The NED assembly 200e bears some structural similarity with the NED assembly 200a: the NED assembly 200e comprises a polarization selective DOE 208e that is positioned with respect to the optical waveguide 106 in a similar manner as the DOE 208a of the NED assembly 200a and a polarization filter 21 Oe positioned with the optical waveguide 106 in a similar manner as the second DOE 210a of the NED assembly 200a. Moreover, in the NED assembly 200e the incoupler 106a of the optical waveguide 106a is arranged to incouple the light projected thereto via operation of the image projection unit 102 at a first polarization state, whereas the outcoupler 106c is arranged to release the light received via the incoupler 106a at the first polarization state. The polarization filter 21 Oe is arranged to prevent external light with the first polarization state from entering the optical waveguide 106.
[0119] Like the DOE 208a of the NED assembly 200a, the polarization selective DOE 208e of the NED assembly 200e is arranged on the first side of the optical waveguide 106 between the optical waveguide 106 and the eyebox. The spatial arrangement of the polarization selective DOE 208e in relation to the optical waveguide 106 is similar to the placement of the DOE 208a of the NED assembly 200a in relation to the optical waveguide, whereas positioning of the polarization filter 21 Oe is similar to that of the second DOE 210a of the NED assembly 200a, mutatis mutandis.
[0120] The phase modification characteristics of the polarization selective DOE 208e are similar to those of the DOE 208a, the DOE 208b, the DOE 208c and the DOE 208d in that the DOE 208a introduces the phase delay profilee(s, t) though position (s, t) of its aperture. However, the polarization selective DOE 208e is arranged to apply the phase delay profilee(s, t) only to light having the first polarization state, whereas light at polarization that is orthogonal to the first polarization state is transmitted through the polarization selective DOE 208e substantially unmodified. The polarization selective DOE 208e may be implemented, for example, via application of phase-modulating metamaterial arranged into a meta-surface that provides a respective desired phase delay through a plurality of positions of the meta-surface as described in the foregoing, mutatis mutandis, where the meta-surface is further arranged to implement the polarization selectivity described in the foregoing.
[0121] The role of the polarization filter 21 Oe has a role similar to that of the second DOE 210a of the NED assembly 200a in that the polarization filter 21 Oe serves to ensure substantially undistorted view to the real world through the optical components of the NED assembly 200e. In this regard, the polarization filter 21 Oe is arranged to prevent light having the first polarization state from passing therethrough. Hence, the polarization filter 21 Oe prevents external light at the first polarization state from entering the optical waveguide 106 while it passes through light in a polarization state that is substantially orthogonal to the first polarization state, thereby rendering the polarization selective DOE 208e substantially transparent to the external light that passes the polarization filter 21 Oe and enters the polarization selective DOE 208e trough the optical waveguide 106..
[0122] Along the lines described in the forgoing for the NED assemblies 200a, 200b, 200c and 200d, the phase modification characteristics of the polarization selective DOE 208e are matched with the preprocessing procedure of the image preprocessing portion 232 such that the preprocessed image / dobtained via operation of the preprocessing portion and displayed via the optical waveguide 106 is perceivable through the polarization selective DOE 208e substantially in focus at accommodation depths within the predefined depth range [z+, z-]. In this regard, the learning arrangement and the iterative learning procedure(s) described in the foregoing for the NED assemblies 200b, 200d are also applicable for determining the weights of the at least one ANN that serves as the preprocessing procedure implemented by the image preprocessing portion 232 and the phase delay profile < / >e(s, t) of the polarization selective DOE 208e for the NED assembly 200e, when the virtual display and eye model 306’ and in the see-through display and eye model 310’ are modified to account for the different optical path from the optical waveguide 106 to the retina of the eye 120 of the viewer.
[0123] Figure 10 illustrates a block diagram of some components of an apparatus 400 that may be employed to implement at least some aspects of the display controller 230 and / or to carry out the iterative learning procedure described in the foregoing. Although described herein with references to a single apparatus 400, the at least some aspects of the display controller 230 and / or the iterative learning procedure described in the foregoing may be implemented via joint operation of two or more apparatuses 400 arranged to provide a cloud-based computing service.
[0124] The apparatus 400 comprises a processor 410 and a memory 420. The memory 420 may store data and computer program code 425. The apparatus 400 may further comprise communication means 430 for wired or wireless communication with other apparatuses and / or user I / O (input / output) components 440 that may be arranged, together with the processor 410 and a portion of the computer program code 425, to provide the user interface for receiving input from a user and / or providing output to the user. In particular, the user I / O components may include user input means, such as one or more keys or buttons, a keyboard, a touchscreen or a touchpad, etc. The user I / O components may include output means, such as a display or a touchscreen. The components of the apparatus 400 are communicatively coupled to each other via a bus 450 that enables transfer of data and control information between the components.
[0125] The memory 420 and a portion of the computer program code 425 stored therein may be further arranged, with the processor 410, to cause the apparatus 400 to operate as the display controller 230 and / or to carry out the iterative learning procedure described in the foregoing (as applicable). The processor 410 is configured to read from and write to the memory 420. Although the processor 410 is depicted as a respective single component, it may be implemented as respective one or more separate processing components. Similarly, although the memory 420 is depicted as a respective single component, it may be implemented as respective one or more separate components, some or all of which may be integrated / removable and / or may provide permanent / semi-permanent / dynamic / cached storage.
[0126] The computer program code 425 may comprise computer-executable instructions that implement at least some aspects of the display controller 230 or carry out the iterative learning procedure described in the foregoing (as applicable) when loaded into the processor 410. As an example, the computer program code 425 may include a computer program consisting of one or more sequences of one or more instructions. The processor 410 is able to load and execute the computer program by reading the one or more sequences of one or more instructions included therein from the memory 420. The one or more sequences of one or more instructions may be configured to, when executed by the processor 410, cause the apparatus 400 to operate as the display controller 230 and / or to carry out the iterative learning procedure described in the foregoing (as applicable). Hence, the apparatus 400 may comprise at least one processor 410 and at least one memory 420 including the computer program code 425 for one or more programs, the at least one memory 420 and the computer program code 425 configured to, with the at least one processor 410, cause the apparatus 400 to perform at least some aspects of the display controller 230 or the iterative learning procedure described in the foregoing (as applicable).
[0127] The computer program code 425 may be provided e.g. as a computer program product comprising at least one computer-readable non-transitory medium having the computer program code 425 stored thereon, which computer program code 425, when executed by the processor 410 causes the apparatus 400 to perform at least some aspects of the display controller 230 and / or to carry out the iterative learning procedure described in the foregoing (as applicable). The computer-readable non-transitory medium may comprise a memory device or a record medium that tangibly embodies the computer program. As another example, the computer program may be provided as a signal configured to reliably transfer the computer program.
[0128] Reference(s) to a processor herein should not be understood to encompass only programmable processors, but also dedicated circuits such as field- programmable gate arrays (FPGA), application specific circuits (ASIC), signal processors, etc. Features described in the preceding description may be used in combinations other than the combinations explicitly described.
Claims
Claims1. A near-eye display, NED, assembly (200a, 200b, 200c, 200d, 200e) for a stereoscopic NED apparatus including a pair of NED assemblies, the NED assembly comprising: an image projection portion (102) projecting collimated light that represents a preprocessed imagean optical waveguide (106) comprising an incoupler (106a) arranged to receive the light projected from the image projection portion (102) and an outcoupler (106b) arranged to release the light received via the incoupler (106a); a diffractive optical element, DOE, (208a, 208b, 208c, 208d, 208e) arranged on a first side of the optical waveguide (106) to allow for viewing the light released from the optical waveguide (106) therethrough, wherein the DOE (208 a, 208b, 208c, 208d, 208e) is arranged to apply a phase modulation function that results in a phase delay that varies with a spatial position of an aperture of the DOE (208 a, 208b, 208c, 208d, 208e); and a preprocessing portion (232) arranged to derive the preprocessed image ( / d) based on an input image ( / ) via application of a preprocessing procedure that is arranged to apply an image-area-position dependent preprocessing in derivation of different image sub-areas of the preprocessed image ( / d) to account for the variation in the phase delay through spatially corresponding sub-areas of the aperture of the DOE (208 a, 208b, 208c, 208d, 208e), wherein said preprocessing procedure and the phase modulation function are arranged to jointly provide a virtual image for perception through the DOE (208a, 208b, 208c, 208d, 208e) substantially in focus over a predefined depth range.
2. A NED assembly (200a, 208b, 208c, 208d, 208e) according to claim 1 , wherein the preprocessing portion (232) comprises at least one artificial neural network, ANN, arranged to implement the preprocessing procedure.
3. A NED assembly (200a, 208b, 208c, 208d, 208e) according to claim 1 or 2, wherein the at least one ANN comprises a single ANN trained to derive each of a plurality of image sub-areas of the preprocessed image ( / d) in dependence of its position within the image area.
4. A NED assembly (200a, 208b, 208c, 208d, 208e) according to claim 1 or 2, wherein the at least one ANN comprises a plurality of ANNs, each ANN trained to derive a respective one of a plurality of image sub-areas of the preprocessed image ( / d).
5. A NED assembly (200a, 200b, 200c, 200e) according to any of claims 1 to 4, wherein the first DOE (208a, 208b, 208c, 208e) is an optical component provided separately from the optical waveguide (106).
6. A NED assembly (200c) according to claim 5, further comprising a first lens (108) arranged between the optical waveguide (106) and the DOE (208c), wherein the first lens (108) is a negative lens arranged to diverge the light directed towards the DOE (208c); and a second lens (110) arranged on a second side of the optical waveguide (106) that is opposite to the first side to allow for external light to enter the optical waveguide (106) therethrough, wherein the second lens (110) is a positive lens that compensates for the optical power of the first lens (108).
7. A NED assembly (200e) according to claim 5, wherein the incoupler (106a) is arranged to incouple the projected light into the optical waveguide (106) at a first polarization state and the outcoupler (106b) is arranged to release the light received via the incoupler (106a) at the first polarization state, wherein the DOE (208) comprises a polarization selective DOE (208e) arranged to apply the first phase modulation function to light having the first polarization state, and wherein the NED assembly (200e) further comprises a polarization filter (21 Oe) arranged on a second side of the optical waveguide (106) that is opposite to the first side at a second predefined distance from the optical waveguide (106) to allow for external light to enter the optical waveguide (106) therethrough, wherein the polarization filter (21 Oe) is arranged to prevent light at the first polarization state from passing therethrough.
8. A NED assembly (200a, 200b, 200c, 208e) according to any of claims 5 to 8, wherein the DOE (208a, 208b, 208c, 208e) comprises one of the following: a lens-like optical element having a non-uniform thickness that is defined separately for a plurality of predefined spatial positions within its aperture to provide a respective desired phase delay through the respective spatial positions of the aperture; or phase-modulating metamaterial arranged into a meta-surface, where a nanostructure forming the meta-surface is arranged to provide a respective desired phase delay through a plurality of spatial positions of the meta-surface.
9. A NED assembly (200a) according to claim 5, further comprising a second DOE (210) arranged on a second side of the optical waveguide (106) that is opposite to the first side to allow for external light to enterthe optical waveguide (106) therethrough, wherein the second DOE (210) is arranged to apply a second phase modulation function that results in a phase delay that varies with a spatial position of an aperture of the second DOE (210) and that is arranged to compensate for the first phase modulation function applied by the first DOE (208).
10. A NED assembly (200a) according to claim 9 wherein each of the DOE (208a) and the second DOE (210a) comprises one of the following: a respective lens-like optical element having a non-uniform thickness that is defined separately for a plurality of predefined spatial positions within its aperture to provide a respective desired phase delay through the respective spatial positions of the aperture; or phase-modulating metamaterial arranged into a respective meta-surface, where a nanostructure forming the meta-surface is arranged to provide a respective desired phase delay through a plurality of spatial positions of the meta-surface.
11. A NED assembly (200d) according to any of claims 1 to 4, wherein the first DOE (210) is integrated to the outcoupler (106b, 208d)12. An apparatus for deriving a preprocessing procedure and a phase modulation function for a NED assembly (200a, 200b, 200c, 200d, 200e) according to any of claims 5 to 8 or 11 , the apparatus arranged to carry out an iterative learning procedure that comprises processing a plurality of training images ( / f) through a learning arrangement that represents the arrangement of the preprocessing procedure and the DOE (208b, 208c, 208d, 208e) and that comprises respective learning models for determining weights for at least one artificial neural network, ANN, that serves as the preprocessing procedure and for determining a phase delay profile (e(s, f)) serves to implement the phase modulation functionof the DOE (208b, 208c, 208d, 208e), wherein the phase delay profile (e(s, t)) is determined for a predefined reference wavelength.
13. An apparatus according to claim 12, wherein the apparatus is arranged to carry out the following at a plurality of iteration rounds: process each training image ( / f) through a first learning model (304’, 306’) that is arranged to model the at least one ANN and the phase delay profile (< / e(s, t)) of the DOE (208b, 208c, 208e) to derive a corresponding simulated retinal virtual image ( / rv) and to determine (308’) a respective first error measure (Ev) that is descriptive of a difference between the simulated retinal virtual image ( / rv) and a corresponding reference retinal image (f), process each training image ( / f) through a second learning model (310’) that is arranged to model light transmission through the optical waveguide (106) and the phase delay profile (e(s, / )) of the DOE (208b, 208c, 208e) to derive a corresponding simulated retinal real-world image ( / rr) and to determine (312’) a respective second error measure (Er) that is descriptive of a difference between the simulated retinal real-world image ( / rr) and a corresponding reference retinal image ( / '), and update the weights of the at least one ANN and the phase delay profile (e(s, t)) based on respective first error measures (Ev) and / or respective second error measures (Er) determined for the plurality of training images ( / ')■14. An apparatus according to claim 12, wherein the apparatus is arranged to carry out the following at a plurality of iteration rounds: process each training image ( / f) through a first learning model (304’, 306’) that is arranged to model the at least one ANN and the phase delay profile (< / e(s, t)) of the DOE (208d) to derive a corresponding simulated retinal virtual image ( / rv) and to determine (308’) a respective error measure (Ev)that is descriptive of a difference between the simulated retinal virtual image ( / rv) and a corresponding reference retinal image ( / '), and update the weights of the at least one ANN and the phase delay profile (e(s, t)) based on respective error measures (Ev) determined for the plurality of training images ( / f).
15. An apparatus according to claim 13 or 14, arranged to select one of a plurality of accommodation depths and one of a plurality of focal depths for each iteration and to apply the first learning model (304’, 306’) and the second learning model (310’) in consideration of the selected accommodation depth and the selected focal depth.
16. An apparatus according to claim 15, wherein the selected accommodation depth and the selected focal depth are different for a first plurality of iteration rounds, and wherein the selected accommodation depth and the selected focal depth are the same for a second plurality of iteration rounds.
17. An apparatus for deriving a preprocessing procedure and a phase modulation function for a NED assembly (200a) according to claim 9 or 10, the apparatus arranged to carry out an iterative learning procedure that comprises processing a plurality of training images ( / f) through a learning arrangement that represents the arrangement of the preprocessing procedure, the DOE (208a) and the second DOE (21 Oe) and that comprises respective learning models for determining weights for at least one artificial neural network, ANN, that serves as the preprocessing procedure, for determining a phase delay profile (e(s, f)) that serves to implement the phase modulation function of the DOE (208a), and for determining a second phase delay profile (c(s, f)) that serves to implement the second phase modulation function of the secondDOE (210a), wherein the phase delay profile (e(s, / )) and the second phase delay profile (c(s, / )) are determined for a predefined reference wavelength.
18. An apparatus according to claim 17, wherein the apparatus is arranged to carry out the following at a plurality of iteration rounds: process each training image ( / f) through a first learning model (304, 306) that is arranged to model the at least one ANN and the phase delay profile (< / e(s, t)) of the DOE (208a) to derive a corresponding simulated retinal virtual image ( / rv) and to determine (308) a respective first error measure (Ev) that is descriptive of a difference between the simulated retinal virtual image ( / rv) and a corresponding reference retinal image ( / '), process each training image ( / f) through a second learning model (310) that is arranged to model the second phase delay profile (e(s, / )) of the second DOE (210a), light transmission through the optical waveguide (106), and the phase delay profile (e(s, / )) of the DOE (208a) to derive a corresponding simulated retinal real-world image ( / rr) and to determine (312) a respective second error measure (Er) that is descriptive of a difference between the simulated retinal real-world image ( / rr) and a corresponding reference retinal image ( / '), and update the weights of the ANN, the phase delay profile (e(s, / )) and the second phase delay profile (c(s, / )) based on respective first error measures (Ev) and / or respective second error measures (Er) determined for the plurality of training images ( / f).
19. An apparatus according to claim 18, arranged to select one of a plurality of accommodation depths for each iteration and to apply the first learning model (304, 306) and the second learning model (310) in consideration of the selected accommodation depth.
20. An apparatus according to any of claims 13 to 19, wherein the learning arrangement further comprises an eye model (302, 302’) that represents one or more optical limitations of an eye, and wherein the apparatus is arranged to apply the eye model (302, 302’) to derive, based on each training image ( / f), the corresponding reference retinal image ( / Q.
Citation Information
Patent Citations
Capping machine disposable strong bottle and weak bottle
KR102619386B1
Eye tracker illumination through a waveguide
US11669159B2
Holographic virtual reality display
US11892629B2
Varifocal display with fixed-focus lens
US20200301239A1
Waveguide-type display device
US20220390749A1