Multi-instance light emission system and method for retinal scanning near-eye display
The multi-instance lighting system in retinal scanning near-eye displays expands the eyebox and enhances depth perception by utilizing multiple light-emitting units and redirectors, addressing size and image quality issues.
Patent Information
- Application Number
- JP2025061750
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-05
- Filing Date
- 2025-04-03
- Publication Date
- 2026-02-18
AI Technical Summary
Existing retinal scanning near-eye displays face challenges in expanding the eyebox without increasing the device's physical size and resolving issues like focal rivalry and vergence accommodation conflict, which affect image quality and depth perception.
A multi-instance lighting system with multiple light-emitting units and light redirectors on a transparent substrate provides collimated optical signals to the retina, allowing for a wider eyebox and accurate depth perception without bulky optical components.
The system achieves a significantly larger eyebox and improved depth perception by using multiple light instances with different angles, reducing device size and weight, and minimizing reliance on complex mechanical and optical mechanisms.
Smart Images

Figure 2026027174000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a retinal scanning augmented reality device capable of displaying virtual images with a three-dimensional (3D) perception in a real space. More specifically, the present invention relates to a retinal scanning augmented reality device having an enlarged effective visual field area (eyebox), and a method for enlarging the effective visual field area in a retinal scanning augmented reality device. [Background technology]
[0002] One of the major challenges in designing head-wearable AR / VR devices is to expand the observer's field of view (FOV) while maintaining sufficient image quality and minimizing the device's physical size. One way to expand the FOV is to provide multiple viewing positions for each eye, allowing the observer to receive image information from various positions or orientations. The range of viewing positions from which the image provided by the device is visible to the observer is sometimes referred to as the "eyebox." The size and shape of the eyebox can significantly affect the observer's experience, and this is particularly true for retinal scanning head-wearable AR / VR devices. For example, if the eyebox is too small, the observer may not be able to see the image generated by the head-wearable AR / VR device even if the observer's line of sight (or visual axis) deviates even relatively slightly from the direction of the incident image. Expanding the eyebox (i.e., increasing the range or number of viewing positions from which the image provided by the head-wearable AR / VR device is visible) is often achieved by optical means. However, expanding the eyebox involves adding bulky optical components to the head-wearable AR / VR device. Therefore, it is desirable to design a system and method for expanding the eyebox without sacrificing the viewer's experience and without increasing the physical size of the head-wearable AR / VR device. Summary of the Invention [Problem to be solved by the invention]
[0003] Referring to FIG. 1, many novel near-eye displays have been proposed that can provide a viewer with multiple viewing positions (or viewpoints). The idea is to create multiple instances from a single pixel and project these instances to multiple positions (viewpoints 151, 152, 153, 161, 162, and 163). When all pixels of an image frame are projected to their respective viewpoints, the viewer receives visual information from multiple pixels, allowing them to view the complete image rendered by the near-eye display. Therefore, even when the viewer's eyes rotate from one viewing position, they can perceive the complete image from another viewing position. However, this method requires the near-eye display to render multiple complete images close to each other in a narrow space in front of the pupil. In many cases, one viewpoint may undesirably overlap with an adjacent viewpoint, resulting in a double image when the viewer's vision receives image information from two viewpoints in the image frame. [Means for solving the problem]
[0004] This invention introduces a novel approach to solve the problem of eyebox limitations in retinal scanning display systems. In this invention, each monocular pixel is composed of multiple light instances with different incident angles. This design allows the retina to receive the light instances regardless of eye orientation. As a result, the observer can perceive the monocular pixel in a wider eyebox, thereby expanding the effective field of view.
[0005] The present invention achieves eyebox enlargement without relying on complex mechanical and optical mechanisms, significantly reducing the size and weight of retinal scanning near-eye displays. As a result, this advancement makes retinal scanning displays more commercially viable for end consumers.
[0006] The present invention also offers significant advantages over the prior art by resolving focal rivalry and vergence accommodation conflict (VAC) issues in virtual and mixed reality displays. In the fields of augmented and mixed reality, depth perception and 3D effects in virtual images are often achieved using parallax imaging technology. Typically, parallax images of partial binocular virtual images are displayed separately to each eye on a screen fixed at a fixed distance from the observer's eyes. However, this distance often differs from the depth perception of the virtual image. In addition, the different distances of the object and the screen from the observer prevent them from focusing on both simultaneously. The present invention solves these issues by providing realistic and accurate depth perception, enhancing the superposition of real and virtual objects for each individual user.
[0007] Unlike conventional methods for expanding the eyebox using light splitters and / or motors to change the orientation and position of optical components in a retinal scanning display, the present invention uses multiple light-emitting units and light redirectors mounted on a transparent substrate. Each light-emitting unit, together with its corresponding light redirector, can provide an optical signal with a unique optical path to the retina. Using multiple light-emitting units allows for multiple light emission directions, providing the viewer with a nearly unlimited eyebox. The present invention is also significantly lighter and more compact than conventional methods.
[0008] A multi-instance light-emitting system for a retinal scanning near-eye display according to the present invention comprises a plurality of light-emitting units for emitting collimated optical signals to a first eye or a second eye of a viewer, respectively, and at least one light redirector for redirecting light emission of the plurality of light-emitting units such that optical paths of a first collimated optical signal and a third collimated optical signal from the plurality of collimated optical signals have optical paths or optical path extensions that converge to form a first convergence point, and a second collimated optical signal and a fourth collimated optical signal from the plurality of collimated optical signals have optical paths or optical path extensions that converge to form a second convergence point. The first convergence point is located on the optical path after entering the pupil of the first eye, and the second convergence point is located on the optical path after entering the pupil of the second eye. The first collimated optical signal and the third collimated optical signal contain substantially the same image information, and the second collimated optical signal and the fourth collimated optical signal contain substantially the same image information. The first collimated optical signal, the second collimated optical signal, the third collimated optical signal, and the fourth collimated optical signal have different optical paths.
[0009] According to the present invention, the first collimated optical signal and the third collimated optical signal form a first monocular image, the second collimated optical signal and the fourth collimated optical signal form a second monocular image, an observer, upon receiving the first monocular image and the second monocular image, perceives a partial binocular image having depth perception, and the depth of the partial binocular image perceived by the observer is modulated by changing the distance between the first convergence point and the second convergence point based on the observer's interpupillary distance. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 shows a method for expanding the eyebox according to the prior art.
[0011] [Figure 2] Figure 2 illustrates the principles of binocular vision for depth perception.
[0012] [Figure 3A] FIG. 3A illustrates the principle for rendering depth perception according to the present invention.
[0013] [Figure 3B] FIG. 3B illustrates the principle for rendering depth perception according to the invention.
[0014] [Figure 4A] FIG. 4A shows a first embodiment according to the present invention.
[0015] [Figure 4B] FIG. 4B shows a first embodiment according to the present invention.
[0016] [Figure 5] FIG. 5 is another view showing the first embodiment according to the present invention.
[0017] [Figure 6A] Figure 6A shows two light instances with a convergence point exactly on the retina.
[0018] [Figure 6B] Figure 6B shows two light instances with a convergence point on one side of the retina.
[0019] [Figure 6C] FIG. 6C is another diagram showing two light instances with a convergence point on one side of the retina.
[0020] [Figure 7] FIG. 7 illustrates the principle of the eyebox and FOV expansion according to the present invention.
[0021] [Figure 8] FIG. 8 is a diagram illustrating the principle of the eyebox and FOV expansion according to the present invention.
[0022] [Figure 9A]FIG. 9A shows one embodiment of the present invention with three light instances per monocular pixel.
[0023] [Figure 9B] FIG. 9B is another diagram illustrating an embodiment of the present invention having three light instances per monocular pixel.
[0024] [Figure 10] Figure 10 illustrates the principle of simultaneously rendering two binocular pixels with different depth perception with two light instances per monocular pixel.
[0025] [Figure 11A] FIG. 11A illustrates the principle of simultaneously rendering two binocular pixels with different depth perception with two light instances per monocular pixel.
[0026] [Figure 11B] Figure 11B illustrates the principle of simultaneously rendering two binocular pixels with different depth perception with two light instances per monocular pixel.
[0027] [Figure 12A] FIG. 12A shows an alternative embodiment according to the present invention.
[0028] [Figure 12B] FIG. 12B is another diagram illustrating an alternative embodiment according to the present invention.
[0029] [Figure 13A] FIG. 13A shows two light instances emitted within the persistence window.
[0030] [Figure 13B] FIG. 13B is another diagram showing two light instances emitted within the persistence window.
[0031] [Figure 14]FIG. 14 shows factors to be considered for arranging a light emitting unit according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0032] The terms used in the following description are intended to be interpreted in the broadest reasonable manner, even when used in conjunction with a detailed description of specific embodiments of the technology. Although certain terms may be emphasized below, terms intended to be interpreted in a restrictive manner shall be specifically defined as such in this detailed description section.
[0033] The multi-instance lighting system according to the present invention is particularly advantageous for near-eye displays (NEDs). The NEDs referred to in this invention are visual display technologies designed to be worn near both eyes, typically within 5 centimeters of the eyes. In this case, the NEDs take the form of eyeglasses or head-wearable devices. In some other embodiments, the physical dimensions of the present invention may be minimized so that it can be implemented on a contact lens.
[0034] The multi-instance light-emitting system for a retinal scanning near-eye display according to the present invention can directly emit light signals toward the observer's retina to form a resolved image on the retina. In other words, the observer's eyes do not need to focus on a specific image display to perceive the image. Conventional retinal scanning near-eye displays suffer from a smaller eyebox and viewing angle compared to waveguide-based near-eye displays. To solve this problem, complex optical systems must be implemented. The present invention discloses a system and method for a novel retinal scanning near-eye display that can create realistic depth perception while using a simple optical design, while simultaneously solving the problems of the limited eyebox and viewing angle in the prior art.
[0035] A multi-instance lighting system for a retinal scanning near-eye display according to the present invention can render a virtual image that an observer can perceive at a specific position in a three-dimensional real space. The multi-instance lighting system includes a plurality of lighting units, each configured to emit a collimated light signal toward at least one of the observer's eyes at a specific angle relative to the observer's coronal plane. The multi-instance lighting system also includes a plurality of light redirectors for redirecting the light emission of the lighting units. In one embodiment, light signals from multiple lighting units (e.g., two, three, or four) are used to generate a pixel image of an entire image frame. In other embodiments, it is also possible to generate a pixel image of an image frame (an image frame may include, for example, 1280 x 720 pixels) using only one light signal emitted from a single lighting unit. In some embodiments, the light signal emitted from the lighting units is collimated by at least one light redirector. The present invention does not require that the lighting units emit collimated light signals. The light redirector can be used to collimate the light signal. Importantly, each light-emitting unit only needs to generate the entire image frame or a partial image of the virtual object perceived by the observer, and the light signal entering the pupil needs to be collimated.
[0036] For purposes of explanation, in some parts of this specification, each light-emitting unit is designated to emit a light signal for one pixel of an image frame. A light-emitting unit may include at least one light emitter (e.g., LED, micro-LED, OLED, LCD, etc.) for generating a collimated or non-collimated light signal (for one pixel) having various wavelengths / colors (e.g., by mixing different colored lights from different light emitters). In some embodiments, a light-emitting unit may include other optical elements (e.g., lenses, collimators, color filters, waveguides, etc.) for modifying the basic optical properties of the emitted light signal. In the present invention, regardless of the number of light-emitting units or the number and type of optical elements included in the light-emitting units, each light-emitting unit can be individually controlled to emit a light signal to the retina. Thus, when there are multiple light-emitting units, each light-emitting unit can be independently controlled to emit a light signal.
[0037] The multiple light redirectors may be integrated into the light-emitting units (e.g., each light-emitting unit includes a light redirector therein) or may be separate from the light-emitting units. In some embodiments, a multi-instance light-emitting system for a retinal scanning near-eye display may further include, in addition to the optical elements for modifying the basic optical properties of the emitted light signals described above, multiple light redirectors for independently and individually redirecting the light emission direction of each of the multiple light-emitting units, or the light redirectors may redirect the light signals emitted from the multiple light-emitting units. In some embodiments, the light redirectors may be microlenses, liquid crystal spatial light modulators (LCSLMs), planar metalenses, or the like. Each of the light redirectors may be provided in one light-emitting unit. However, depending on the embodiment, one light redirector may be shared by multiple light-emitting units, or vice versa. When the multiple light redirectors are liquid crystal spatial light modulators, the liquid crystal spatial light modulators may include multiple liquid crystal cells. When at least one light-emitting unit emits a light signal, the light signal emitted from the at least one light-emitting unit is collimated and propagates in a predetermined direction by changing the driving voltage of one of the liquid crystal cells corresponding to the light-emitting unit (a known technique for changing the phase of liquid crystals by changing the driving voltage of a liquid crystal cell is omitted here). In this case, the light redirector can dynamically change the light emission direction of the multiple light-emitting units by modulating the driving voltage as needed, thereby changing the emission angle of each light-emitting unit. To render a virtual image or object moving relative to a viewer, the light redirector may change the light emission direction of the multiple light-emitting units so that the collimated light signal is directed toward the viewer's eyes at a time-varying angle relative to the viewer's frontal plane. Detailed methods for rendering objects with different depth positions will become more apparent later in this disclosure.When the multiple light redirectors are multiple metalenses, each metalenses may comprise nanostructures such that an optical signal incident on one of the light emitting units and received by the metalenses may be collimated and / or travel in a direction different from its original emission direction upon interacting with the metalenses (it is well known in the art that flat metalenses can act to cause light to travel in a different direction, and therefore will not be further described here). The metalenses described herein may function to refract and / or collimate light so as to change the direction of an optical signal emitted from a corresponding light emitting unit. For example, multiple metalenses may be provided, with each light emitting unit having at least one corresponding metalense capable of directing emitted light in a particular direction or to a particular location on the retina. It is worth noting that in this embodiment, at least two or three light emitting units may be provided with metalenses that direct the emitted optical signal to the same location on the retina. Other methods of implementing a light redirector are possible. The present disclosure is not limited to the foregoing examples. As long as the light redirector is capable of changing the direction of light emitted from the corresponding light emitting unit, it is considered to be within the scope of the present disclosure. In some embodiments, one light redirector may be responsible for modifying a corresponding lighting unit, although, depending on the embodiment, multiple light redirectors may be responsible for modifying corresponding lighting units, or one light redirector may be responsible for modifying multiple corresponding lighting units.
[0038] The multiple light-emitting units (in addition to the light redirectors) may be positioned in front of the observer's eyes and emit collimated light signals directly onto the observer's retina. The multiple light-emitting units do not form a perceptible image along the light path. A resolved, focused image is generated only on the observer's retina. Further, by way of example, different light-emitting units may emit different light signals, each forming a different pixel of an image frame on the observer's retina. Each pixel may exhibit a different color. The color of the light signal may be the result of mixing multiple colors of light emitted from light emitters within a single light-emitting unit. That is, in some embodiments, the light-emitting unit may include at least one light emitter. However, the present invention is not limited to the above embodiments.
[0039] As long as each of the light-emitting units can be controlled to emit a light signal to the observer's retina at a specific angle relative to the observer's frontal plane (directional light emission), or as long as each of the light-emitting units can be controlled to emit a light signal to a specific position on the observer's retina, it should be considered within the scope of the present invention. It is worth mentioning that the retinal scanning near-eye display provides a directional and / or collimated light signal to the retina, unlike conventional waveguide-based augmented reality (AR) / virtual reality (VR) near-eye displays (which provide scattered light signals to the observer).
[0040] The multi-instance light-emitting system for a retinal scanning near-eye display according to the present invention can emit light signals directly to the retina to render a resolved, sharp virtual image on the retina, allowing a viewer to perceive an image of a virtual object without fixating and focusing on a screen or display panel. In some cases, the present invention can even enable a viewer to perceive a sharp image without relying on the focusing effect of the eye's lens (focus-free). As a result, even people with eye damage can perceive the image generated by the present invention as long as part of the retina is functioning normally.
[0041] The following describes a basic method utilized to render a virtual image with a three-dimensional perception according to the present invention. To create a complete image frame, multiple collimated right optical signals and multiple collimated left optical signals are emitted to a first retina (e.g., the left retina) and a second retina (e.g., the right retina), respectively. Each of the right collimated optical signals and its corresponding left collimated optical signal are received by the right and left retinas, and the image of each of the right collimated optical signals and its corresponding left collimated optical signal are fused by the human brain, allowing the observer to perceive multiple partial binocular images (e.g., binocular pixels) of a virtual image or image frame (comprising multiple binocular images / pixels). The complete virtual image or image frame is formed by the combination of the partial binocular images (e.g., pixels). Each of the multiple collimated right optical signals has a corresponding collimated left optical signal. In some embodiments, the right collimated optical signal and its corresponding left collimated optical signal have substantially the same image information (e.g., the same color / wavelength or intensity). In other embodiments, they may have slightly different image information (e.g., disparity image information). Hereinafter, it is assumed that a right collimated light signal has substantially the same image information as its corresponding left collimated light signal. This means that both the right collimated light signal and its corresponding left collimated light signal contain image information of the same binocular pixels, i.e., the right collimated light signal and its corresponding left collimated light signal render images of the same region of a virtual image or image frame. In some cases, this also means that both the right collimated light signal and its corresponding left collimated light signal have the same intensity or color. When the right collimated light signal and its corresponding left collimated light signal are received by a first retina and a second retina at specific positions on the first retina and the second retina, respectively, they form a partial binocular image perceivable at a specific spatial location. It is well known that the horizontal and vertical positions of a partial binocular image in three-dimensional space perceived by an observer are related to the horizontal and vertical positions on the first and second retinas at which the right and left collimated light signals are received, respectively.Additionally, in accordance with the present invention, it has been found that the depth position of a partial binocular image perceived by an observer correlates with the distance between the locations on the retina at which a right collimated light signal and a corresponding left collimated light signal are received.
[0042] In the present invention, directional optical signals (i.e., collimated light) are preferred for virtual image rendering because they direct optical signals to target locations on the observer's retina, thereby manipulating the distance between the locations on the retina where a right collimated optical signal and a corresponding left collimated optical signal are received. Conventional optical signals generated by waveguide-based near-eye displays are less desirable because the optical signals emitted from such devices are scattered and require the eye's lens to focus the scattered light. This can lead to a mismatch between the perceived depth of the virtual image and the eye's focal position on the screen / display.
[0043] In this application, for the convenience of explaining the principles of human vision and retinal scanning, the retinas of the observer's first and second eyes are represented as matrices, with each matrix element corresponding to a specific horizontal and vertical position on the retina. Furthermore, for simplicity, only a few light instances are illustrated in the figures. It is worth noting that the phrase "light instance" used herein collectively refers to a single emission of a light pulse or light signal and its light path propagating through space during a specific period of time. In the figures, light instances are illustrated as the light paths of light signals (shown as straight lines). Referring to FIG. 2, according to natural vision, when a human views an object, the visual axes of the left and right eyes (defined as imaginary lines extending from the center of the retina (or fovea) to the center of the pupil) point toward the object, and the visual axes of both eyes converge at the object's location. As a result, most of the light emitted from the object is received by the observer's central retinal region (the fovea), where most of the photoreceptor cells are concentrated. The eye's lens then focuses the object based on the distance between the object and the eye. When fixating an object, the human brain interprets the depth of the object based in part on the convergence angle between the visual axes of the two eyes. That is, human binocular depth perception is based in part on the convergence angle between the visual axes of the two eyes. As the convergence angle increases, the depth perception perceived and interpreted by human vision decreases, which means that the object is located closer to the observer. On the other hand, as the convergence angle decreases, the depth perception perceived and interpreted by human vision increases, which means that the object is located farther from the observer.
[0044] To obtain accurate depth perception, the visual axis of the eye must be reoriented (i.e., the eye must be rotated toward the object) so that the image of the object reaches a region close to the center of the retina (the fovea). If light from the object is received in a region outside the central region of the retina (meaning the visual axis of the eye is not directed toward the object), depth perception is reduced. However, the observer can still vaguely perceive the object's position and depth. Therefore, if the observer wants to see the object clearly, they need to orient their eyes toward the object so that the light from the object reaches a region close to the fovea, as shown in Figure 2.
[0045] The above-described method can be implemented to display depth perception in a retinal scanning near-eye display. Referring to FIG. 3A, the first and second retinas of an observer are shown as two matrices, where L(2,2) and R(2,2) represent the central regions of the retinas. Consider a case where the observer's eyes fixate on a virtual object P1 at a first moment. The virtual object P1 may be a partial binocular image (e.g., a binocular pixel) of the virtual object, or a complete binocular image (composed of multiple binocular pixels). At the first moment, light instances R1 and L1 from the virtual object P1 reach the first and second retinas (the center of the retina) at L(2,2) and R(2,2). The convergence angle between the observer's eyes is CA1, which is the same as the optical convergence angle OCA1 between the light instances R1 and L1. Note that in both FIG. 3A and FIG. 3B, the solid lines represent the optical paths of the light signals and do not represent the visual axes of the eyes. In general, the convergence angle between optical axes may not be the same as the optical convergence angle (the term "optical convergence angle" is defined as the convergence angle between the optical paths of collimated optical signals). After a first moment, the virtual object P1 approaches the observer. At a second moment, light instances R1' and L1' from the virtual object P1 reach the first and second retinas L(1,2) and R(3,2) (which are not retinal centers) at an optical convergence angle OCA2. Because the light instances R1' and L1' are not received by L(2,2) and R(2,2), the observer must rotate both eyes so that the visual axis follows the virtual object P1 to clearly perceive the image. As a result, the convergence angle CA1 of the visual axis increases to CA2, and during this time, the light instances R1' and L1' are received by L(2,2) and R(2,2), as shown in Figure 3B. The depth perceived by the observer when fixating the virtual object P1 is correlated with the distance between the positions where the two retinas receive the light instance. More specifically, the distance d1 between the positions (R(2,2) and L(2,2)) where the two eyes receive the light instance from the virtual object P1 at a first moment is shorter than the distance d2 between the positions (R(2,2) and L(2,2)) where the two eyes receive the light instance from the virtual object P1 at a second moment.Apparently, modulating the distance between the locations where the retina receives the left and right light signals (i.e., R1 and L1 and R1' and L1') allows us to perceive different depths of a virtual image. Specifically, rendering a virtual object closer to the observer requires a larger vergence angle of the visual axis. Inducing an observer to fixate a virtual object at a larger vergence angle requires a relatively larger distance between the locations where the retina receives the left and right light signals (thus requiring the observer's eyes to rotate more nasally to compensate for the larger distance between the light signals, so that the light signals representing the virtual object are received at the fovea). On the other hand, rendering a virtual object farther from the observer requires a smaller vergence angle of the visual axis. Inducing an observer to fixate a virtual object at a smaller vergence angle requires a relatively smaller distance between the locations where the retina receives the left and right light signals. It should be noted that the left and right optical signals (collimated) do not need to be projected directly onto the center of the retina, but can be projected onto areas other than the center of the retina, and by directing the observer's gaze toward the center of the retina to receive the left and right optical signals, depth perception is then formed based on the change in the orientation of the eyes (related to the angle of convergence) when fixating the image. However, since the interpupillary distance may differ for each observer, it is also important to take the user's interpupillary distance into consideration when adjusting the distance between the positions where the retina receives the left and right optical signals.
[0046] One of the key advantages of the present invention is that by utilizing a collimated optical signal as a light source, high-resolution, focused images can be generated directly on the observer's retina, without the need for a traditional screen or display for the observer to fixate on. In the present invention, the visual location of each pixel in real space is rendered by projecting a collimated optical signal for each pixel onto a specific location on the observer's retina.
[0047] The following describes a method for rendering a virtual object having a three-dimensional contour surface. This method can also be applied to the simultaneous rendering of image frames having multiple virtual objects at different depths. Referring to FIG. 4A, this figure shows a first binocular pixel BP1 and a second binocular pixel BP2 at different depths being rendered simultaneously. In this example, the observer's eyes are initially fixated on the first binocular pixel BP1. The first binocular pixel BP1 is composed of a first collimated optical signal S1 and a second collimated optical signal S2, while the second binocular pixel BP2 is composed of a third collimated optical signal S3 and a fourth collimated optical signal S4. The first collimated optical signal S1 and the third collimated optical signal S3 are projected to the observer's left eye, and the second collimated optical signal S2 and the fourth collimated optical signal S4 are projected to the observer's right eye. Because both eyes of the observer are fixating on the first binocular pixel BP1, the first collimated optical signal S1 arrives at a position L(2,2) on the left retina corresponding to the center of the left retina, and the second collimated optical signal S2 arrives at a position R(2,2) on the right retina corresponding to the center of the right retina. The first collimated optical signal S1, the second collimated optical signal S2, the third collimated optical signal S3, and the fourth collimated optical signal S4 are emitted by the first light-emitting unit 101, the second light-emitting unit 102, the third light-emitting unit 103, and the fourth light-emitting unit 104, respectively. The radiation directions of the first light-emitting unit 101, the second light-redirecting unit 202, the third light-redirecting unit 103, and the fourth light-redirecting unit 204 are affected by the first light-redirecting device 201, the second light-redirecting device 202, the third light-redirecting device 203, and the fourth light-redirecting device 204, respectively. In order for the centers of the retinas (foveae) of both eyes to receive the first collimated optical signal S1 and the second collimated optical signal S2, the observer's eyes must be rotated so that the convergence angle between the visual axes of the two eyes is equal to the convergence angle CA1. At this time, the observer interprets the visual depth of the first binocular pixel BP1 as z1 based on the convergence angle CA1.In other words, to guide the observer to fixate the first binocular pixel BP1 with a convergence angle CA1, the distance between the receiving positions of the first collimated optical signal S1 and the second collimated optical signal S2 on the retina must be d1. At this time, the third collimated optical signal S3 and the fourth collimated optical signal S4 are received at positions L(1,2) and R(3,2) on the retina. The observer may only vaguely perceive the depth and position of the second binocular pixel BP2.
[0048] Referring to FIG. 4B, if an observer aims to fixate the second binocular pixel BP2 and directs both eyes toward the second binocular pixel BP2, the third collimated optical signal S3 arrives at position L(2,2) on the left retina, which corresponds to the center of the left retina, and the fourth collimated optical signal S4 arrives at position R(2,2) on the right retina, which corresponds to the center of the right retina. In order for the centers of the retinas (foveae) of both eyes to receive the third collimated optical signal S3 and the fourth collimated optical signal S4, the observer's eyes must rotate so that the convergence angle between the visual axes of the eyes is equal to the convergence angle CA2. In this case, the observer interprets the visual depth of the second binocular pixel BP2 as z2 based on the convergence angle CA2. In other words, to guide the observer to fixate on the second binocular pixel BP2 with a convergence angle CA2, the distance between the receiving positions of the third collimated optical signal S3 and the fourth collimated optical signal S4 on the retina must be d2. In this case, the first collimated optical signal S1 and the second collimated optical signal S2 are received at positions L(3,2) and R(1,2) on the retina. The observer may only vaguely perceive the depth and position of the first binocular pixel BP1.
[0049] In some embodiments, the first binocular pixel BP1 and the second binocular pixel BP2 may be two separate pixels from an image frame generated by a near-eye display according to the present invention. In other embodiments, the first binocular pixel BP1 and the second binocular pixel BP2 may be two separate pixels forming a portion of an image of a virtual object having a 3D surface contour. The image frame or virtual object may be composed of multiple binocular pixels (e.g., 1440 x 1080 pixels). In any case, each binocular pixel may have a unique depth position. It is clear that binocular pixels can be rendered according to this method so that each of them has a unique perceptible depth to the observer. It is worth noting that, although each of the binocular pixels in an image frame or virtual object may be rendered simultaneously, in some embodiments, not all binocular pixels of an image frame or virtual object may be rendered simultaneously. The observer can see the binocular pixels simultaneously because they only need to be rendered within the duration of human vision. Either case should be considered within the scope of the present invention.
[0050] Furthermore, according to the present invention, all binocular pixels are presented to the observer simultaneously (or within the duration of human vision). The observer can freely decide which part of the image frame or virtual object to fixate at any time and can always perceive a clear depth perception of that part of the image frame or virtual object. This is because, when the observer rotates their eyes to fixate on the image generated by the collimated optical signal, the eyes rotate at a desired convergence angle, and as a result, the collimated optical signal is projected at a predetermined position on the observer's retina so as to generate a corresponding depth perception for the image similar to natural vision. This feature is particularly advantageous because it can reduce the dependency on eye-tracking mechanisms in retinal scanning near-eye displays.
[0051] The following describes a method for increasing the eyebox in a retinal scanning near-eye display. It is worth mentioning that in conventional retinal scanning near-eye displays, which suffer from the problem of a small eyebox, a single light instance is used to render a monocular pixel (i.e., a pixel for the left or right eye). The monocular pixel projected to the left eye and the monocular pixel projected to the right eye are fused by the human brain to generate a binocular pixel. If both eyes rotate away from the light path of one of the light instances, the light can no longer enter both eyes of the observer, and the observer cannot see the image of the binocular pixel. In the present invention, each partial binocular image (e.g., binocular pixel) of an image frame or virtual object is formed by multiple light instances. Furthermore, each light instance may have a different light path and enter both eyes of the observer at a different angle relative to the observer's frontal plane. Furthermore, the aforementioned depth rendering technique enables a multi-instance lighting system to have the capability of 3D effect rendering, which has previously been considered difficult and overlooked.
[0052] Referring to FIG. 5 , a multi-instance light-emitting system for a retinal scanning near-eye display according to the present invention may include a first light-emitting unit 101, a second light-emitting unit 102, a third light-emitting unit 103, and a fourth light-emitting unit 104 for emitting a first collimated light signal S1, a second collimated light signal S2, a third collimated light signal S3, and a fourth collimated light signal S4, respectively. While only four light-emitting units are shown in the figure for illustrative purposes, those skilled in the art will understand that in actual implementations of the present invention, there may be four or more light-emitting units (e.g., several thousand or more). In this example, the first collimated light signal S1 and the third collimated light signal S3 are emitted to the left eye, and the second collimated light signal S2 and the fourth collimated light signal S4 are emitted to the right eye. However, the present invention is not limited to this configuration. In other embodiments, the direction of light emission may be adjusted so that the collimated light signals are emitted to different eyes at different times. The multi-instance emission system may further include at least one light redirector 201, 202, 203, and 204 for redirecting light emission directions of the plurality of light emitting units 101, 102, 103, and 104 such that the optical paths of a first collimated optical signal S1 and a third collimated optical signal S3 among the plurality of collimated optical signals from the plurality of light emitting units 101, 102, 103, and 104 (and the light redirectors 201, 202, 203, and 204) converge to form a first convergence point CP1, and the optical paths of a second collimated optical signal S2 and a fourth collimated optical signal S4 among the plurality of collimated optical signals converge to form a second convergence point CP2, where the first collimated optical signal S1, the second collimated optical signal S2, the third collimated optical signal S3, and the fourth collimated optical signal S4 have different optical paths. This is particularly important because it allows the eyes to receive image information from different optical signals when the eyes are oriented differently: the first collimated optical signal S1 and the third collimated optical signal S3 contain substantially the same image information, and the second collimated optical signal S2 and the fourth collimated optical signal S4 contain substantially the same image information.This means that the first collimated optical signal S1 and the third collimated optical signal S3 carry the same image information of a first monocular image MP1 (e.g., monocular pixels provided to the left eye), and the second collimated optical signal S2 and the fourth collimated optical signal S4 carry the same image information of a second monocular image MP2 (e.g., monocular pixels provided to the right eye). When a viewer receives both the first monocular image MP1 and the second monocular image MP2, the viewer perceives a partial binocular image (e.g., binocular pixels of an image frame or a virtual object).
[0053] Referring to FIG. 5, the first convergence point CP1 is located on the optical path after entering the pupil of the first eye, and the second convergence point CP2 is located on the optical path after entering the pupil of the second eye. This means that the intersection of the optical paths of the first collimated optical signal S1 and the third collimated optical signal S3 appears behind the pupils of both eyes. In one embodiment, the first convergence point CP1 and the second convergence point CP2 are located substantially on the retina of the left eye and the right eye, respectively, so that the observer can see an integrated resolution image of the first monocular image MP1 and the second monocular image MP2. However, because the shape of the eye is not completely the same for all users, and furthermore, the shape of the eye is usually not circular, it is difficult to accurately converge the first collimated optical signal S1 and the third collimated optical signal S3 (or the second collimated optical signal S2 and the fourth collimated optical signal S4) on the surface of the retina (as shown in FIG. 6A). Alternatively, referring to Figures 6B and 6C, as long as the convergence point (the intersection of optical path S1 and optical path S3, or the intersection of optical path S2 and optical path S4) is located on one side of the retina (e.g., behind or in front of the retina) within a tolerance range (less than ±2 mm), or as long as diplopia does not occur when both the first collimated optical signal S1 and the third collimated optical signal S3 are received by the retina, it should be considered to be within the scope of the present invention.
[0054] Referring to FIG. 6A , in the present invention, each monocular image projected to the left and right eyes is rendered by multiple light instances. The light instances that render the same monocular image are intended to be projected at the same position on the retina to create a unified, focused image of the monocular image. Referring to FIGS. 6B and 6C , in some embodiments, the light instances that render the same monocular image are projected at positions close to each other on the retina so that the spatial separation (on the retina) between the light instances of the same monocular image is within an acceptable range, so that the observer still perceives the two light instances as one. The incident light instances form light points on the retina (as shown in the figure). It is desirable that the light points of the light instances on the retina are non-resolvable (to avoid diplopia and distortion of the monocular image). The definition of a resolved image is clearly defined by the Rayleigh criterion and is well known to those skilled in the art, so a detailed description is omitted here. In some embodiments, to limit the amount of energy received by the same location on the retina over an extended period of time (to avoid retinal damage from light), the light instances of the monocular image (e.g., the first collimated light signal S1 and the third collimated light signal S3) may not be emitted simultaneously, but instead may be emitted alternately or intermittently so that the retina does not receive excessive energy from a single monocular image, although light instances of the same monocular image may need to be projected within the persistence of vision period.
[0055] Referring to Figure 7, assume that light instances 1, 2, and 3, containing the same image information of a monocular image, are initially emitted to the central region of the retina while the eye is looking straight ahead. As the observer's eye rotates, the convergence points of the light instances are received at different positions on the retina (although the spatial location of the convergence points relative to the environment remains the same). As a result, the observer perceives the monocular image as moving within the field of view (FOV) as the eye rotates. This is consistent with natural vision, where a visible object moves relative to the field of view as the eye rotates. Furthermore, because there are multiple light instances forming a single monocular image, the observer's eyes can receive the image information of the monocular image at different eye orientations. This expands the eyebox and effective field of view. Conventionally, when only one light instance is used per monocular image (e.g., a monocular pixel), if the pupil rotates excessively from the light instance, the light instance of the light signal will no longer enter the pupil, and the monocular image will disappear from the observer's field of view. On the other hand, unlike conventional retinal scanning near-eye display systems, the present invention allows the observer to receive image information from the optical signals even when the pupils of the eyes are pointing in different directions.
[0056] By rendering a monocular image (e.g., a monocular pixel) using multiple light instances having different angles of incidence on the retina, the observer can receive information about the monocular image from various directions. As a result, the eyebox of the present invention can be significantly enlarged compared to the prior art. Furthermore, because the observer's eyes can view the monocular image from various directions, the monocular image is always present in the field of view even during eye rotation, and the observer can see the monocular image moving smoothly within the field of view in response to eye rotation, similar to natural vision.
[0057] Referring to FIG. 8 , in one embodiment of the present invention, three light instances (L1, L3, L5) are used to render a single monocular image (e.g., a monocular pixel) for one eye. Each of the light instances is incident on the pupil at a different angle of incidence relative to the observer's frontal plane. In some cases, the angle of incidence (θ ) of the first light instance L1 relative to the frontal plane may be adjusted to increase the field of view (FOV) of the multi-instance lighting system.L1 ) and the angle of incidence (θ L5 ) can be reduced. However, this may prevent the first light instance L1 and the fifth light instance L5 from entering the eye when the eye is looking straight ahead. In this case, the eye can still receive monocular image information from light instance L3. When the eye is rotated to the left, light instances L3 and L5 cannot enter the eye, but the retina of the eye will still be able to receive light instance L1. The position at which the retina receives the light instances of the monocular image changes from the fovea to the left side of the retina (thus, the eye perceives the monocular image as having moved to the right of the FOV). When the eye is rotated to the right, light instances L1 and L3 cannot enter the eye, but the retina of the eye will still be able to receive light instance L5. The position at which the retina receives the light instances of the monocular image changes from the fovea to the right side of the retina (thus, the eye perceives the monocular image as having moved to the left of the FOV). Therefore, according to the present invention, an observer can view a monocular image regardless of the orientation of the eyes. This method can be applied to both eyes of an observer.
[0058] It is worth noting that the optical paths of light instances may be perturbed due to geometrical shape and eye rotation. However, the optical path extensions of different light instances still converge with each other at the convergence point, as shown in Figure 8. Therefore, the location of the convergence point can be determined based on the optical path extensions of the light instances anyway.
[0059] As mentioned above, the depth perception of a partial binocular image formed by projecting collimated optical signals to both eyes is known to correlate with the convergence angle between the visual axes of the two eyes when the observer fixates on the partial binocular image. From the perspective of a head-wearable display system, the convergence angle of the observer's eyes (correlated with the angle at which the visual axes of the two eyes must face each other for the optical signal to reach the fovea) can be manipulated by changing the distance between the first and second convergence points of the optical signals. The depth perception of a partial binocular image (e.g., binocular pixels from the entire image frame or an image of a virtual object) can be manipulated by changing the relative distance between the first and second convergence points (as shown in Figure 4B) and rotating the eyes so that the observer receives the optical signal / image at the center of the retina, resulting in a 3D perception depending on the final convergence angle of the visual axes. Generally, the temporal increase in depth change of a partial binocular image can be adjusted by temporally decreasing the distance between the first and second convergence points, and the temporal decrease in depth change of a partial binocular image can be adjusted by temporally increasing the distance between the first and second convergence points. However, in practice, because the shape and parameters of observers' eyes vary from person to person, several additional factors must be considered, such as the observer's interpupillary distance (IPD) and the actual orientation of the eyes when a particular observer fixates at a particular depth. For example, the positions of the first and second convergence points on the retina must be adapted to the observer's IPD. Even if the distance between the first and second convergence points is the same, observers with different IPDs may experience different depth perception. According to the present invention, the time-varying depth of the partial binocular image when the observer perceives the right and left optical signals is realized by adjusting the distance between the first and second convergence points based on the observer's IPD, so that the positions of the convergence points can be projected to the correct positions on the retina for each observer. Furthermore, because the retina is not flat but curved, in some embodiments, the positions of the first and second convergence points on the retina also need to be adapted and modified according to eye rotation to render accurate depth perception of the binocular image.It should be noted that the depth perception referred to in this invention refers to the depth perceived by an observer when directing both eyes toward the spatial position where the partial binocular image (virtual image) or binocular image (also a virtual image) is located. This is characterized by the observer rotating the center of the retina (or fovea) of the eye so that the convergence point of the light instance is on the fovea (or substantially near the fovea). In this case, the visual axis of the two eyes refers to the rendered position of the binocular image or partial binocular image. More specifically, the rendered position (the position seen by the observer) of the binocular image or partial binocular image is the position where the visual axes of the two eyes intersect. When the observer looks away from the partial binocular image or binocular image, the depth perception decreases. However, this is consistent with natural vision.
[0060] Therefore, to accurately render the depth perception of partial or full binocular images, an initial calibration process can be performed to determine the relationship between the distance between the convergence points of monocular images and the depth (or convergence angle) perceived by a particular user. For example, to know the distance between the pupils (and approximate fovea) of the two eyes, the observer's IPD must first be determined. For example, the IPD may be determined when the observer is looking straight ahead (when the visual axis is looking straight ahead). The observer is then presented with partial or full binocular images rendered with various distances between the first and second convergence points (meaning that these partial binocular images have different perceptible depths for the observer). In one embodiment, the observer may be asked to fixate one image at a time, and an eye-tracking device is used to determine the corresponding orientation of the eyes (when the observer is fixating on a particular partial binocular image with a particular depth). As a result, a relationship between the distance between the first and second convergence points and the corresponding convergence angle of the eyes (related to the depth perceived by the observer) can be determined. This relationship can be used to render accurate depth perception for each individual user. In another embodiment, calibration can be performed using real objects in the environment with the assistance of a distance measurement device. For example, the distance measurement device can measure the actual depth of real objects around the user, and the multi-instance lighting system can attempt to render partial binocular images with depths that match the depths of the real objects. The observer can be asked to fixate the partial binocular images and, based on the observer's preferences, adjust the perceived depth of the partial binocular images until the perceived depth of the partial binocular images and the depth of the real objects overlap with each other. As the perceived depth of the partial binocular images is adjusted, the multi-instance lighting system changes the distance between the convergence points of the light instances rendering the partial binocular images. If the observer determines that the perceived depth of the partial binocular images matches the depth of the real objects, the distance between the convergence points can be recorded. This allows the relationship between perceived depth and the distance between points of convergence to be determined.There are other ways to calibrate a multi-instance lighting system to render accurate depth perception for a user without departing from the spirit of the present invention.
[0061] The above description can now be applied to rendering partial binocular pixels with depth perception. Referring to Figures 9A and 9B, in some embodiments, the partial binocular pixel is rendered by two monocular pixels (a first left monocular pixel MP1 and a first right monocular pixel MP2). While only six light-emitting units are shown in the figures for illustrative purposes, those skilled in the art will understand that in actual implementations of the present invention, there may be six or more (e.g., several thousand or more) light-emitting units. In this example, the first left monocular pixel MP1 and the first right monocular pixel MP2 are each composed of three light instances. Specifically, the first left monocular pixel MP1 is rendered by a first collimated light signal S1, a third collimated light signal S3, and a fifth collimated light signal S5 emitted from the first light-emitting unit 101, the third light-emitting unit 103, and the fifth light-emitting unit 105, respectively. Similarly, the first right monocular pixel MP2 is rendered by the second collimated light signal S2, the fourth collimated light signal S4, and the sixth collimated light signal S6 emitted from the second light-emitting unit 102, the fourth light-emitting unit 104, and the sixth light-emitting unit 106, respectively. The light instance of the first left monocular pixel MP1 forms a first convergence point CP1 on the left retina, and the light instance of the first right monocular pixel MP2 forms a second convergence point CP2 on the left retina. Note that the light instances of the first left monocular pixel MP1 and the first right monocular pixel MP2 enter the observer's eyes at different angles relative to the observer's frontal plane. This feature allows both eyes to see the images of the first left monocular pixel MP1 and the first right monocular pixel MP2 regardless of the orientation of the observer's eyes. For example, as shown in FIG. 9A, assume that the observer's eyes are initially looking straight ahead and the visual axes of the eyes are facing forward. Due to the orientation of the two eyes, only two of the three light instances of the first left monocular pixel MP1 (i.e., S1 and S3) can enter the pupil of the left eye, and only two of the three light instances of the first right monocular pixel MP2 (i.e., S2 and S4) can enter the pupil of the right eye.Note that the optical path extension of the fifth collimated optical signal S5 still converges with the optical paths of the first and third collimated optical signals S1 and S3, and the optical path extension of the sixth collimated optical signal S6 still converges with the optical paths of the second and fourth collimated optical signals S2 and S4. Because the light instances are not received by the fovea, the observer can only vaguely see the image and depth perception of the partial binocular pixel (formed by the fusion of the first left monocular pixel MP1 and the first right monocular pixel MP2). Referring to FIG. 9B, assume that when the observer rotates their eyes to fixate on the partial binocular pixel, the convergence point of the light instances reaches the fovea. The image and depth perception of the binocular pixel can be clearly perceived by the observer. Note that in the present invention, the collimated optical signal emitted from the light-emitting unit is configured to provide a resolved image on the observer's retina. As a result, the observer can perceive an image without using the eye's lens to focus the light signals, similar to a Maxwell near-eye display. Because the visual axes rotate to allow the fovea to receive the light instances, the depth of the binocular pixel perceived by the observer is correlated with the convergence angle between the visual axes. The convergence angle between the visual axes is related to the spatial distance d1 between the first convergence point CP1 and the second convergence point CP2. Note that in Figures 9A and 9B, d1 is constant, and only the eye orientation changes. As mentioned above, when the spatial distance d1 is relatively and moderately large, the eyes must face each other more closely for the convergence points to reach the fovea, resulting in a larger convergence angle between the visual axes, and the perceived image appears closer, and vice versa. In Figure 9B, only collimated light signals S3, S5, S4, and S6 can enter the pupils of both eyes, while collimated light signals S1 and S2 are blocked by both eyes. Because multiple light instances of the same monocular pixel are implemented, the observer can see the image.
[0062] As described above, according to some embodiments of the present invention, a partial binocular image (e.g., a binocular pixel) is formed by multiple light instances projected onto two separate convergence points (which can be considered as a convergence point pair). In a binocular image (e.g., an image frame or an image of a virtual object) composed of multiple partial binocular images, each pair of a first convergence point (e.g., for the left eye) and a second convergence point (e.g., for the right eye) in each partial binocular image has a different distance between the first and second convergence points. Referring to FIG. 10, two partial binocular pixels PBP1 and PBP2 are rendered. Although only eight light-emitting units are shown in the figure for illustrative purposes, those skilled in the art will understand that in actual implementations of the present invention, there will be more than eight light-emitting units (e.g., several thousand or more). The first partial binocular pixel PBP1 is rendered by a pair of monocular pixel 1 (MP1) and monocular pixel 2 (MP2), and the second partial binocular pixel PBP2 is rendered by a pair of monocular pixel 3 (MP3) and monocular pixel 4 (MP4). In this example, each monocular pixel is composed of two light instances. Furthermore, the pair of convergence points CP1 and CP2 for the first binocular pixel PBP1 has a separation distance d1, and the pair of convergence points CP3 and CP4 for the second binocular pixel PBP2 has a separation distance d2. According to the above-described method for depth rendering, it is clear that the perceived depths for the partial binocular pixels PBP1 and PBP2 are different. The depths of the first partial binocular pixel PBP1 and the second partial binocular pixel PBP2 perceived by an observer depend on the angle by which the observer's eyes need to rotate in order for the fovea of each observer to receive the first and second convergence points (which corresponds to the convergence angle, and differs between observers with different IPDs).
[0063] The following describes an exemplary embodiment for rendering multiple binocular pixels with different depths of an image frame or a virtual object image according to the present invention. Referring to FIGS. 11A and 11B, two partial binocular images of an image frame (a first partial binocular pixel PBP1 and a second partial binocular pixel PBP2) are shown. Note that the first partial binocular pixel PBP1 and the second partial binocular pixel PBP2 belong to the same image frame and are therefore rendered substantially simultaneously. A viewer may freely fixate either of the two pixels. The first partial binocular pixel PBP1 is rendered by fusing a first left monocular pixel MP1 with a first right monocular pixel MP2, and the second partial binocular pixel PBP2 is rendered by fusing a second left monocular pixel MP3 with a second right monocular pixel MP4. In this example, each monocular pixel may be composed of two light instances. For example, the first left monocular pixel MP1 is rendered by the first collimated optical signal S1 and the third collimated optical signal S3 emitted from the light-emitting units 101 and 103, respectively. The first and third collimated optical signals S1 and S3 are emitted to the first eye (e.g., the left eye) of the observer. The first right monocular pixel MP2 is rendered by the second collimated optical signal S2 and the fourth collimated optical signal S4 emitted from the light-emitting units 102 and 104, respectively. The second and fourth collimated optical signals S2 and S4 are emitted to the second eye (e.g., the right eye) of the observer. Similarly, the second left monocular pixel MP3 is rendered by the fifth collimated optical signal S5 and the seventh collimated optical signal S7 emitted from the light-emitting units 105 and 107, respectively. The fifth and seventh collimated optical signals S5 and S7 are emitted to the first eye (e.g., the left eye) of the observer. The second right monocular pixel MP4 is rendered by the sixth collimated optical signal S6 and the eighth collimated optical signal S8 emitted from the light-emitting units 106 and 108, respectively. The sixth and eighth collimated optical signals S6 and S8 are emitted to the second eye (e.g., the right eye) of the observer.The first collimated optical signal S1 and the third collimated optical signal S3 contain substantially the same image information, the fifth collimated optical signal S5 and the seventh collimated optical signal S7 contain substantially the same image information, the second collimated optical signal S2 and the fourth collimated optical signal S4 contain substantially the same image information, and the sixth collimated optical signal S6 and the eighth collimated optical signal S8 contain substantially the same image information.
[0064] The first, second, third, fourth, fifth, sixth, seventh, and eighth collimated optical signals (S1, S2, S3, S4, S5, S6, S7, and S8) may be emitted simultaneously or non-simultaneously. If the collimated optical signals are not emitted simultaneously, they may be emitted within a persistence of vision period to ensure that the rendered image does not disappear from the viewer's field of view. In this embodiment, the emission directions of the first, second, third, fourth, fifth, sixth, seventh, and eighth collimated optical signals (S1, S2, S3, S4, S5, S6, S7, and S8) are such that the optical paths of the first collimated optical signal S1 and the third collimated optical signal S3, or their optical path extensions, converge to form a first convergence point CP1, and the optical paths of the second collimated optical signal S2 and the fourth collimated optical signal S4, or their optical path extensions, converge to form a first convergence point CP1. The optical paths of the sixth collimated optical signal S6 and the eighth collimated optical signal S8, or their optical path extensions, are adjusted by at least one optical direction changer (e.g., 201, 202, etc.) so that they converge to form a second convergence point CP2, the optical paths of the fifth collimated optical signal S5 and the seventh collimated optical signal S7, or their optical path extensions, converge to form a third convergence point CP3, and the optical paths of the sixth collimated optical signal S6 and the eighth collimated optical signal S8, or their optical path extensions, converge to form a fourth convergence point CP4. The optical paths of the sixth collimated optical signal S6 and the eighth collimated optical signal S8, or their optical path extensions, converge to form the fourth convergence point CP4. As in the previous embodiment (e.g., FIG. 5), the first and third convergence points CP1 and CP3 are located on the optical path after entering the pupil of a first eye (e.g., the left eye), and the second and fourth convergence points CP2 and CP4 are located on the optical path after entering the pupil of a second eye (e.g., the right eye). The first, second, third, and fourth convergence points CP1, CP2, CP3, and CP4 are continuously positioned on the retinas of the eyes, as described in FIG. 5 and the description associated with FIG. 5.
[0065] Referring to FIG. 11A, in this embodiment, the depth of the first partial binocular pixel PBP1 is different from the depth of the second partial binocular pixel PBP2. Assume that both eyes of an observer first fixate on the first partial binocular pixel PBP1. A light instance of the first partial binocular pixel PBP1 is received by the fovea. Note that the visual axis of the two eyes is an imaginary line extending from the fovea to the center of the pupil, and is shown by a dotted line in the figure. The physical location in space of the first partial binocular pixel PBP1 perceived by the observer appears to be at the position where the visual axes intersect. At this time, the left and right eyes are fixating on the first partial binocular pixel PBP1 at a convergence angle CA1. As described above, the depth of the first partial binocular pixel PBP1 perceived by the observer may be adjusted by setting the distance between the first convergence point CP1 and the second convergence point CP2 (in this case, d1 as shown in the figure) based on the observer's interpupillary distance. The depth of the second partial binocular pixel PBP2 perceived by the observer is adjusted by setting the distance between the third convergence point CP3 and the fourth convergence point CP4 (d2 in this case, as shown) based on the observer's interpupillary distance. It is worth noting that at this time, the observer can also see the image of the second partial binocular pixel PBP2, but the depth perception is vague and unclear. The observer can vaguely recognize that the second partial binocular pixel PBP2 is relatively closer to the observer than the first partial binocular pixel PBP1.
[0066] Referring to FIG. 11B, assume that an observer is fixating on the second partial binocular pixel PBP2 and intends to move their eyes toward the second partial binocular pixel PBP2. Then, a light instance of the second left monocular pixel MP3 arrives at a position L(2,2) on the left retina, which corresponds to the center (or fovea) of the left retina, and a light instance of the second right monocular pixel MP4 arrives at a position R(2,2) on the right retina, which corresponds to the center (or fovea) of the right retina. Note that, in order for the center (fovea) of the retina to receive either the light instance of the second left monocular pixel MP3 or the light instance of the second right monocular pixel MP4, the observer's eyes must be rotated so that the convergence angle between the visual axes of the two eyes is equal to the convergence angle CA2. In this case, the observer interprets the visual depth of the second partial binocular pixel PBP2 based on the convergence angle CA2. In other words, to guide the observer to fixate on the second partial binocular pixel PBP2 having a convergence angle CA2, the distance between the positions where the third convergence point CP3 and the fourth convergence point CP4 of the light instance are received on the retina must be d2. At this time, the light instance of the first partial binocular pixel PBP1 is received at positions L(3,2) and R(1,2) on the retina. The observer may only have a vague understanding of the depth and position of the first partial binocular pixel PBP1. Note that the distances d1 and d2 are constant in both Figures 11A and 11B.
[0067] 11A or 11B , in some embodiments, when the first partial binocular pixel PBP1 and the second partial binocular pixel PBP2 belong to an image of the same virtual object and it is desired to display the virtual object moving away or toward the observer, the observer's perceived depth of the first partial binocular pixel PBP1 may be adjusted by varying the distance d1 between the first and second convergence points CP1 and CP2 based on the observer's interpupillary distance. Similarly, the observer's perceived depth of the second partial binocular pixel PBP2 is adjusted by varying the distance d2 between the third and fourth convergence points CP3 and CP4 based on the observer's interpupillary distance. In this embodiment (displaying a moving virtual object), the distance d1 between the first and second convergence points CP1 and CP2 and the distance d2 between the third and fourth convergence points CP3 and CP4 are simultaneously changed. Note that when the image including the first partial binocular pixel PBP1 and the second partial binocular pixel PBP2 does not move relative to the observer's viewpoint, the distances d1 and d2 are always constant. The distances d1 and d2 are physical distances between the convergence points in real space.
[0068] It is also worth mentioning that in the above embodiments, when a liquid crystal spatial light modulator (LCSLM) is used as the light redirector or the multiple light redirectors are liquid crystal spatial light modulators, the liquid crystal spatial light modulator may include multiple liquid crystal cells, and when at least one light-emitting unit emits a light signal, the driving voltage of one of the liquid crystal cells corresponding to the at least one light-emitting unit can be changed so that the incident light signal from the at least one light-emitting unit is collimated and travels in a predetermined direction (the known technology of changing the driving voltage of the liquid crystal cell to change the phase of the liquid crystal is omitted here). Thus, the light redirector can dynamically change the light emission direction of the multiple light-emitting units at any time by adjusting the driving voltage. The light redirector can change the direction of light emission of the light-emitting units 101-106 so that (for example) either the first collimated light signal S1, the third collimated light signal S3, the fifth collimated light signal S5, or the seventh collimated light signal S7 is directed toward a first eye at a time-varying angle relative to the observer's frontal plane, and either the second collimated light signal S2, the fourth collimated light signal S4, the sixth collimated light signal S6, or the eighth collimated light signal S8 is directed toward a second eye at a second time-varying angle relative to the observer's frontal plane. This allows each of the light-emitting units to have the ability to emit light in different directions, resulting in a reduction in the total number of light-emitting units required for implementation. This can significantly reduce the complexity of manufacturing and signal control.
[0069] In the above-described embodiments, each monocular image (or pixel) is composed of multiple light instances with different optical paths, which enter the observer's pupil at various angles relative to the frontal plane. If each monocular image (or pixel) is composed of three light instances, with the light instances coming from the center, left, and right of the eye (as shown in FIG. 8 ), the eyebox can be maximized, and the observer may be able to see the monocular image when turning left or right. In this embodiment, no eye-tracking mechanism is required to track eye orientation. This is a significant advantage over conventional retinal scanning near-eye displays, which operate by projecting an image directly onto the retina of the user's eye. Integrated eye-tracking monitors the position and movement of the user's eye in real time. As the user's gaze moves, the system adjusts the position of the projector and various auxiliary optical elements to ensure that the optical signal enters the pupil, thereby maintaining image visibility to the observer. Conventional eye-tracking mechanisms are composed of complex and heavy mechanical components to drive auxiliary optical elements to desired positions. It is clear that the above-described embodiments of the present invention eliminate the need for an eye-tracking mechanism. The present invention significantly improves the ease of use of AR glasses for the average consumer, something that has not been achieved with conventional technology.
[0070] Although constructing a monocular image with multiple light instances can significantly expand the eyebox of a retinal scanning near-eye display, extreme eye orientations may prevent all light instances of a monocular pixel from reaching the pupil. Addressing this issue may require incorporating a device for tracking eye orientation and alternative light-emitting unit control schemes, as described later in this specification. In the aforementioned embodiments, each collimated light signal is emitted from a respective light-emitting unit. For example, in FIG. 9A , first, second, third, fourth, fifth, and sixth collimated light signals S1, S2, S3, S4, S5, and S6 are emitted at a given moment from the first light-emitting unit 101, the second light-emitting unit 102, the third light-emitting unit 103, the fourth light-emitting unit 104, the fifth light-emitting unit 105, and the sixth light-emitting unit 106, respectively. When both eyes are facing leftward, the first, third, and fifth collimated optical signals S1, S3, and S5 may not be incident on the left eye. Similarly, when both eyes are facing rightward, the second, fourth, and sixth collimated optical signals S2, S4, and S6 may not be incident on the right eye. Referring to FIGS. 12A and 12B, both figures show another embodiment of the present invention. In this embodiment, one of the collimated optical signals is emitted from one set of light-emitting units at a first moment, and is emitted from a different set of light-emitting units at a second moment after the orientation of the eyes changes. In FIG. 12A, both eyes are fixating on any virtual image or real object in the surrounding environment, and the eyes are facing in corresponding directions. At this time, the first and third collimated optical signals S1 and S3 are projected to the left eye, and the second and fourth collimated optical signals S2 and S4 are projected to the right eye. Similar to the previous embodiment, the first and third collimated light signals S1, S3 are two light instances of a left monocular pixel, and the second and fourth collimated light signals S2, S4 are two light instances of a right monocular pixel. The left and right monocular pixels are then perceived by the observer to form binocular pixels of an image frame or virtual object. The orientation of the two eyes is known by an eye tracking device (not shown).As shown in FIG. 12A , first, second, third, and fourth collimated optical signals S1, S2, S3, and S4 are emitted from the first, second, third, and fourth light-emitting units 101, 102, 103, and 104, respectively, at a first moment. Note that the fifth light-emitting unit 105 and the sixth light-emitting unit 106 do not emit optical signals at the first moment. Referring to FIG. 12B , assuming that the eyes turn to the right at a second moment, the first collimated optical signal S1 originally emitted from the first light-emitting unit 101 and the fourth collimated optical signal S4 originally emitted from the fourth light-emitting unit 104 may not be able to enter the pupil due to the change in the orientation of the eyes. The eye tracking device detects the orientation of the eyes, and the multi-instance light-emitting system determines the appropriate light-emitting units to emit the first collimated optical signal S1 and the fourth collimated optical signal S4 based on the positions of the light-emitting units and the orientation of the eyes. For example, if the system determines that the fifth light-emitting unit 105 is available and suitable for emitting the first collimated light signal S1 (meaning that the fifth light-emitting unit 105 can emit a light signal that can enter the pupil and reach the same convergence point of the previous first collimated light signal S1 and the third collimated light signal S3), the light-emitting unit 101 may be turned off, and the fifth light-emitting unit 105 becomes the light source that radiates the first collimated light signal S1, as shown in FIG. 12B. Similarly, when the system determines that the sixth light-emitting unit 106 is available and suitable for emitting the fourth collimated optical signal S4 (meaning that the sixth light-emitting unit 106 can emit an optical signal that can enter the pupil and reach the same convergence point of the previous second collimated optical signal S2 and fourth collimated optical signal S4), the fourth light-emitting unit 104 may be turned off, and the sixth light-emitting unit 106 becomes the light source that emits the fourth collimated optical signal S1, as shown in Figure 12B. Because the first and fourth light-emitting units 101 and 104 are turned off, the power consumption of the entire system does not increase.Generally speaking, as the orientation of the eyes changes, any of the first collimated optical signal S1, the second collimated optical signal S2, the third collimated optical signal S3, or the fourth collimated optical signal S4 can be emitted by a light-emitting unit other than the first light-emitting unit 101, the second light-emitting unit 102, the third light-emitting unit 103, or the fourth light-emitting unit 104. Note that the spatial positions of the first convergence point (e.g., the convergence point of the first and third collimated optical signals S1 and S3) and the second convergence point (e.g., the convergence point of the second and fourth collimated optical signals S2 and S4) remain substantially the same. That is, the distance d1 between the convergence points remains constant, thereby maintaining the depth of the image. Only the relative position with respect to the observer's retina changes, thereby moving the position of the rendered image within the observer's FOV. This embodiment is highly advantageous because it can further expand the FOV of the multi-instance lighting system.
[0071] In natural vision, when the observer's eyes rotate from one orientation to another, the object perceived by the observer changes its position within the field of view in response to the rotation of the eyes. For example, when the observer's eyes turn to the left, the object appears to move to the right relative to the field of view, while the object in real space remains at the same three-dimensional coordinates relative to the environment. Similarly, in an augmented reality (AR) or mixed reality (MR) environment generated by the multi-instance lighting system of the present invention, a virtual object may be configured to be fixed relative to the real space or the AR / MR environment. However, when the observer's eyes rotate or their orientation changes, the virtual object may appear to move relative to the observer's field of view. All of these functions can be achieved by the multi-instance lighting system, as described above. In the above-described embodiment, to display a static virtual object relative to the real space or AR / MR environment, an image of the virtual object is projected onto the observer's left and right retinas, and the convergence point of the light instances is at a fixed position relative to the real three-dimensional space. As the observer's eyes rotate and the orientation of the visual axes of the left and right eyes changes, the convergence point of the light instances of the virtual image moves to a new position on the retina as the eyes rotate. Furthermore, because multiple light instances are used per pixel for each eye, the eyes can receive light from various orientations. As a result, the virtual object appears to move to a different position within the observer's field of view. That is, the spatial positions of the first and second convergence points CP1 and CP2 remain substantially the same even after the orientation of the first or second eye changes relative to an initial point in time. The present invention allows the observer to experience a user experience closer to natural vision than was possible with conventional technologies. Furthermore, the present invention can generate a partial binocular image at any given spatial position in depth. The observer's visual axes can point directly at the position where the virtual image is rendered, allowing the eyes to fixate and focus on that position. As a result, any depth perception can be generated without using a display screen, eliminating focal competition and binocular vergence-adaptation conflict.
[0072] In the present invention, each monocular pixel is composed of multiple light instances with different incident angles, so the retina can receive the light instances regardless of the orientation of the eyes. Even if the change in the orientation of the eyes exceeds the original viewing angle provided by the light instances, the near-eye system can drive inactive light-emitting units at other positions in the light-emitting array to emit light instances with the same image information from different angles. The observer can perceive the monocular pixels with a much wider eyebox, which means the effective field of view is expanded.
[0073] It should be noted that the light instances of a monocular pixel image do not necessarily have to be emitted simultaneously. In other words, the light-emitting units do not need to remain "on" (e.g., continuously emitting light signals) all the time. To reduce power consumption, the light-emitting units may emit light intermittently. However, the time between the "power-on" states of the light-emitting units must fall within the persistence window so that the image of the monocular pixel does not disappear to the observer. For example, if a monocular pixel is composed of two light instances (see FIGS. 13A and 13B), the light-emitting units emitting the two light instances do not need to be "on" simultaneously (i.e., emitting two collimated light signals simultaneously). As shown in FIG. 13A, both the first light-emitting unit 101 and the third light-emitting unit 103 are responsible for projecting the light instances of the monocular pixel (i.e., the first collimated light signal S1 and the third collimated light signal S3). The optical path extensions of both the first collimated light signal S1 and the third collimated light signal S3 form a first convergence point on the retina. At a first moment (FIG. 13A), only the first light-emitting unit 101 is powered to emit a light instance of the monocular pixel (first collimated light signal S1), and the third light-emitting unit 103 is not emitting light. At a second moment (FIG. 13B), only the third light-emitting unit 103 is powered to emit a light instance of the monocular pixel (third collimated light signal S3), and the first light-emitting unit 101 is not emitting light. In this embodiment, neither light-emitting unit is always in the "on" state. However, the time interval between the "power-on" states (or the time interval between the "power-off" states) of any of the light-emitting units is shorter than the persistence window. As a result, the viewer can always see an image.
[0074] When the light-emitting units for emitting the light instances of a monocular pixel do not emit light simultaneously, the convergence point of the light instances can be determined by calculating the light path length of the light instances and estimating the convergence point of the light paths. The most important factor is that different instances of a monocular pixel must be received at approximately the same position on the retina. As described above, this requires that the convergence point be near the retinal surface. Referring again to Figures 6A to 6C, in the present invention, each light signal (or light instance) is a light ray with a small cross-sectional area. Therefore, when a light instance illuminates the retinal surface, a small light spot area may be formed. In order for a light instance of a single monocular pixel to be perceived by an observer as a single integrated image of that monocular pixel, the distance between the centers of the light spots must meet certain criteria. Note that the shape of the light spots may be circular, elliptical, or even rectangular. For example, in some embodiments, the distance between the pattern centers (the positions with the maximum light intensity) may be set to be smaller than the smallest dimension of the light spot so that the observer's retina perceives the two light instances as the same monocular pixel. The minimum dimension defined here refers to the minimum distance from the center of the light spot to the periphery (or boundary). In some embodiments, the boundary may be defined as the position where the light intensity is 15% of the maximum intensity. In some cases, a more general approach may be adopted to determine whether two light instances have a convergence point near the retina. Unless the observer can perceive two (or three) maximum light intensities from the two (or three) light instances (meaning the observer can only see one maximum light), the two light instances are considered to have a convergence point near the retina.
[0075] The light-emitting units and light redirectors of the present invention may be mounted on a substantially transparent substrate, allowing ambient light to enter the observer's eye. In a multi-instance light-emitting system for a retinal scanning near-eye display, the substrate does not function as a screen on which the observer fixates to display an image, but rather as a carrier for the light-emitting units and light redirectors. Images are formed only on the retina. The substrate comprises multiple light-emitting units spaced apart from one another. The layout and arrangement of the light-emitting units may vary depending on the embodiment and will not be described in detail here. However, for a general guideline regarding the spacing between the light-emitting units, see FIG. 14. In FIG. 14, the eye's rotation locus is assumed to be substantially circular, centered on a center of rotation D. The eye's radius is r, and it is assumed that the eye rotates θ radians around point D. Points B, A, and C correspond to the positions of the pupil center when the eye is not rotated and when the eye rotates θ radians to the left and right, respectively. Finally, x is the distance between the eye's pupil (point B) and the light-emitting unit 103. The light-emitting units 101 and 105 are provided near the light-emitting unit 103 and to the side of the light-emitting unit 103. It is assumed that the distance between the light-emitting units 101 and 105 and the pupil is approximately x when the pupil rotates to points A and C, respectively. E is a point on the surface (circumference) of the eye that is located directly opposite point B. In this case, the distance l between the light-emitting units is expressed by the following equation: l=(2r+x)θ / 2 Note that when the eyeball rotates θ radians, the displacement (distance) between B and C or A and B is rθ. The above is merely an example of a method and formula for setting the distance between light-emitting units. The present invention is not limited to the above example.
[0076] The foregoing description of the embodiments is provided to enable any person skilled in the art to make and use the subject matter. The methods described herein may be performed in any order. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the novel principles and subject matter disclosed herein may be applied to other embodiments without the exercise of any innovative faculty. The claimed subject matter is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein. Additional embodiments are contemplated within the spirit and true scope of the disclosed subject matter. Thus, it is intended that the present invention cover modifications and variations that come within the scope of the appended claims and their equivalents.
Claims
1. 1. A multi-instance lighting system for a retinal scanning near-eye display, comprising: a plurality of light emitting units for emitting light signals to a first eye or a second eye of an observer, respectively; at least one light redirector for changing the light emission directions of the plurality of light emitting units or collimating the optical signals such that optical paths of a first collimated optical signal and a third collimated optical signal from the plurality of collimated optical signals have optical paths or optical path extensions that converge to form a first convergence point, and a second collimated optical signal and a fourth collimated optical signal from the plurality of collimated optical signals have optical paths or optical path extensions that converge to form a second convergence point; the first convergence point is located on an optical path after the light has entered the pupil of the first eye, and the second convergence point is located on an optical path after the light has entered the pupil of the second eye; the first collimated optical signal and the third collimated optical signal contain substantially the same image information, and the second collimated optical signal and the fourth collimated optical signal contain substantially the same image information; the first collimated optical signal, the second collimated optical signal, the third collimated optical signal, and the fourth collimated optical signal have different optical paths. Multi-instance lighting system.
2. 2. The system of claim 1, wherein the first collimated optical signal and the third collimated optical signal form a first monocular image, the second collimated optical signal and the fourth collimated optical signal form a second monocular image, the observer perceives a partial binocular image having depth perception upon receiving the first monocular image and the second monocular image, and the depth of the partial binocular image perceived by the observer is adjusted by changing the distance between the first convergence point and the second convergence point based on the observer's interpupillary distance.
3. 3. The system of claim 2, wherein the first and second convergence points are continuously located on the retinas of the first and second eyes, or the first and second convergence points are located on one side of the retinas.
4. 4. The system of claim 3, wherein an increase in depth of partial binocular image change over time is adjusted by decreasing the distance between the first convergence point and the second convergence point over time, and a decrease in depth of partial binocular image change over time is adjusted by increasing the distance between the first convergence point and the second convergence point over time.
5. 2. The system of claim 1, wherein the first collimated optical signal, the second collimated optical signal, the third collimated optical signal, and the fourth collimated optical signal are emitted by a first light-emitting unit, a second light-emitting unit, a third light-emitting unit, and a fourth light-emitting unit, respectively, at a first moment in time, and any of the first collimated optical signal, the second collimated optical signal, the third collimated optical signal, or the fourth collimated optical signal is emitted by a light-emitting unit other than the first light-emitting unit, the second light-emitting unit, the third light-emitting unit, or the fourth light-emitting unit after an orientation of the first eye or the second eye relative to the first moment in time changes.
6. 6. The system of claim 5, wherein the spatial positions of the first and second convergence points remain substantially the same after an orientation of the first eye or the second eye relative to the first moment changes.
7. The system of claim 1 , wherein each of the plurality of lighting units comprises a light emitter and a light redirector of the at least one light redirector.
8. The system of claim 1 , wherein the at least one light redirector is configured to dynamically change the light emission direction of the plurality of light emitting units.
9. 2. The system of claim 1, wherein the at least one light redirector redirects the light emission direction of the plurality of light emitting units such that either the first collimated light signal or the third collimated light signal is directed toward the first eye at a first angle that varies with time relative to the observer's frontal plane, and the at least one light redirector redirects the light emission direction of the plurality of light emitting units such that either the second collimated light signal or the fourth collimated light signal is directed toward the second eye at a second angle that varies with time relative to the observer's frontal plane.
10. 1. A multi-instance lighting method for rendering binocular image frames for a retinal scanning near-eye display, comprising: emitting a first collimated light signal and a third collimated light signal from each of the plurality of light emitting units toward a first eye of an observer; emitting a second collimated optical signal and a fourth collimated optical signal from each of the plurality of light emitting units toward a second eye of the observer; emitting a fifth collimated optical signal and a seventh collimated optical signal from the plurality of light emitting units toward a first eye of an observer; emitting a sixth collimated light signal and an eighth collimated light signal from the plurality of light emitting units toward a second eye of the observer, the first collimated optical signal, the second collimated optical signal, the third collimated optical signal, the fourth collimated optical signal, the fifth collimated optical signal, the sixth collimated optical signal, the seventh collimated optical signal, and the eighth collimated optical signal are emitted simultaneously or within a period of visual persistence if not simultaneously; radiation directions of the first collimated optical signal, the second collimated optical signal, the third collimated optical signal, the fourth collimated optical signal, the fifth collimated optical signal, the sixth collimated optical signal, the seventh collimated optical signal, and the eighth collimated optical signal are adjusted by at least one light redirector such that optical paths of the first collimated optical signal and the third collimated optical signal have optical paths or optical path extensions that converge to form a first convergence point, the second collimated optical signal and the fourth collimated optical signal have optical paths or optical path extensions that converge to form a second convergence point, the fifth collimated optical signal and the seventh collimated optical signal have optical paths or optical path extensions that converge to form a third convergence point, and the sixth collimated optical signal and the eighth collimated optical signal have optical paths or optical path extensions that converge to form a fourth convergence point; the first convergence point and the third convergence point are disposed on an optical path after the light has entered the pupil of the first eye, and the second convergence point and the fourth convergence point are disposed on an optical path after the light has entered the pupil of the second eye, the first collimated optical signal and the third collimated optical signal form a first left monocular image, the second collimated optical signal and the fourth collimated optical signal form a first right monocular image, the fifth collimated optical signal and the seventh collimated optical signal form a second left monocular image, and the sixth collimated optical signal and the eighth collimated optical signal form a second right monocular image, the observer perceiving a first partial binocular image having depth perception upon receiving the first left monocular image and the first right monocular image, and the observer perceiving a second partial binocular image having depth perception upon receiving the second left monocular image and the second right monocular image; A multi-instance lighting method, wherein the binocular image frame includes the first partial binocular image and the second partial binocular image, and a depth of the first partial binocular image is different from a depth of the second partial binocular image.
11. 11. The method of claim 10, wherein a depth of the first partial binocular image perceived by the observer is adjusted by changing a distance between the first convergence point and the second convergence point based on the interpupillary distance of the observer, and a depth of the second partial binocular image perceived by the observer is adjusted by changing a distance between the third convergence point and the fourth convergence point based on the interpupillary distance of the observer, and the distance between the first convergence point and the second convergence point and the distance between the third convergence point and the fourth convergence point are changed simultaneously.
12. The method of claim 11 , wherein the binocular image frame includes an image of a virtual object, the image of the virtual object including the first partial binocular image and the second partial binocular image.
13. 11. The method of claim 10, wherein the first collimated optical signal and the third collimated optical signal contain substantially the same image information, the fifth collimated optical signal and the seventh collimated optical signal contain substantially the same image information, the second collimated optical signal and the fourth collimated optical signal contain substantially the same image information, and the sixth collimated optical signal and the eighth collimated optical signal contain substantially the same image information.
14. The method of claim 10 , wherein the first convergence point, the second convergence point, the third convergence point, and the fourth convergence point are continuously positioned on the retina of the first eye and the second eye.
15. 15. The method of claim 14, wherein the first collimated optical signal, the second collimated optical signal, the third collimated optical signal, the fourth collimated optical signal, the fifth collimated optical signal, the sixth collimated optical signal, the seventh collimated optical signal, and the eighth collimated optical signal are emitted by a first light-emitting unit, a second light-emitting unit, a third light-emitting unit, a fourth light-emitting unit, a fifth light-emitting unit, a sixth light-emitting unit, a seventh light-emitting unit, and an eighth light-emitting unit, respectively, at a first instant in time, and any of the first collimated optical signal, the second collimated optical signal, the third collimated optical signal, the fourth collimated optical signal, the fifth collimated optical signal, the sixth collimated optical signal, the seventh collimated optical signal, or the eighth collimated optical signal is emitted by a light-emitting unit different from the plurality of light-emitting units after an orientation of the first eye or the second eye relative to the first instant in time changes.
16. 16. The system of claim 15, wherein the spatial positions of the first and second convergence points remain substantially the same after an orientation of the first eye or the second eye relative to the first moment in time changes.