A wearable virtual image module for overlaying virtual images onto real-time images.

The system superimposes a virtual image onto a real-time image using collimated light signals, addressing the limitation of 2D information in conventional systems by enabling precise 3D overlay and alignment, enhancing surgical precision and efficiency.

JP2026056527APending Publication Date: 2026-04-01WOOMY INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2026-04-01

AI Technical Summary

Technical Problem

Conventional visualization support systems in healthcare, particularly during medical procedures, are limited to providing additional visual information in 2D images, requiring healthcare professionals to switch between real-time and processed images, which complicates the observation of superimposed 3D information.

Method used

A system and method for superimposing a virtual image onto a real-time image by projecting collimated light signals onto the observer's eyes, allowing for the generation and calibration of a virtual image at a specific depth and position to overlay it on the real-time image, with adjustable magnification and automatic alignment.

Benefits of technology

Enables healthcare professionals to view additional processed information, such as 3D images, directly superimposed on real-time images, improving surgical precision and efficiency by maintaining alignment and adjusting for individual observer characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026056527000001_ABST
    Figure 2026056527000001_ABST
Patent Text Reader

Abstract

This system provides a way to overlay virtual images onto real-time images or physical objects. [Solution] The system comprises a right collimated light signal generator 170 for generating a right collimated light signal, and a left collimated light signal generator 175 for generating a left collimated light signal corresponding to the right collimated light signal, which is directed towards the other retina of the observer, wherein the right collimated light signal and the left collimated light signal form binocular pixels of a virtual image having a first depth, and the first depth perceived by the observer is modified by changing the convergence angle between the optical path extensions of the right collimated light signal and the corresponding left collimated light signal projected onto the observer's eye, based on the interpupillary distance, wherein the first depth corresponds to the depth position of the convergence point of the optical path extensions of the right collimated light signal and the corresponding left collimated light signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to a method and system for superimposing a virtual image on a real-time image, and more particularly, to a method and system for superimposing a depth virtual image generated by projecting a plurality of right collimated optical signals and corresponding left optical signals onto an observer's eyes on a real-time image.

Background Art

[0002] In recent years, numerous visualization support systems and methods have been developed to assist healthcare professionals during examinations and surgeries, including ophthalmic surgery. During medical procedures, visualization support systems can provide additional visual information such as patient medical records, photographs, magnetic resonance imaging (MRI), X-rays, computed tomography (CT), and optical coherence tomography (OCT) for surgical parameters. In some cases, this additional visual information may be processed images of the patient, such as CT images with several markers. Visualization support systems are often used in conjunction with other medical devices capable of providing real-time images of the patient. Healthcare professionals may receive additional information provided by visualization support systems separately from real-time images. For example, the additional information may be displayed separately on a monitor rather than on a surgical microscope that can observe the patient's real-time image. Monitors typically only provide 2D images. However, during medical procedures, healthcare professionals may want to observe additional visual information (e.g., pre-processed images of the patient) superimposed on the patient's real-time image. Furthermore, conventional visualization support systems can only provide additional visual information in 2D images. Therefore, the ability to generate 3D images to provide additional visual information superimposed on real-time images of patients has become a major concern in the medical industry. For example, in ophthalmic examinations and surgeries, medical professionals perform surgery while observing real-time optical images of the patient's eye through the eyepiece of an ophthalmic microscope. However, surgeons cannot simultaneously observe the processed retinal image of the patient's eye through the microscope during the procedure; they must turn their heads to look at another monitor and then look at the microscope again. Thus, the need remains to incorporate additional visual information of the patient, provided by visualization support systems, into the real-time optical images observed by medical professionals. [Overview of the project] [Problems that the invention aims to solve]

[0003] The object of this disclosure is to provide a system and method for superimposing a virtual image onto a real-time image. The system for superimposing a virtual image onto a real-time image or a real object comprises a real-time image module and a virtual image module. The real-time image module includes an enlargement assembly that generates a real-time image of an object at a first position and a first depth at a predetermined magnification. [Means for solving the problem]

[0004] The virtual image module generates a virtual image by projecting a right collimated light signal onto the observer's right eye and a corresponding left collimated light signal onto the observer's left eye. The right collimated light signal and its corresponding left collimated light signal are perceived by the observer, and the virtual image is displayed at a second position and a second depth. The second depth is related to the angle between the right collimated light signal and its corresponding left collimated light signal projected onto the observer's eye. In one embodiment, the second depth is approximately the same as the first depth. The virtual image is superimposed on the real-time image to provide the observer with more information. Therefore, in one embodiment, the virtual image is a processed image of a real object.

[0005] The magnification of the real-time image is adjustable. After the real-time image has been enlarged, the virtual image may be manually or automatically enlarged to maintain the original overlay between the virtual image and the real-time image. An automatic mode for overlay may be selected.

[0006] To overlay a virtual image onto a real-time image, the system must first be calibrated for the observer. Because observers' eyes have different physical characteristics, such as interpupillary distance, the system must be calibrated specifically for the observer so that when the right and left collimated light signals are projected onto the observer's eyes, the observer can perceive the virtual image displayed at a second position and a second depth.

[0007] The process of overlaying a virtual image onto a real-time image includes (a) selecting a first point on the real-time image as a first landmark; (b) displaying the real-time image at a predetermined magnification at a first position and a first depth; and (c) projecting a virtual image by projecting a right collimated light signal onto the observer's right eye and a corresponding left collimated light signal onto the observer's left eye, so that the corresponding first landmark on the virtual image overlaps with the first landmark on the real-time image, so that the observer can perceive the virtual image at a second position and a second depth. In one embodiment, the depth of the first landmark on the real-time image is approximately the same as the depth of the corresponding first landmark on the virtual image. A second or third landmark may be used in a similar manner to achieve more accurate overlay.

[0008] Further features and advantages of this disclosure are described below, some of which will become apparent from the description or be understood through the practice of this disclosure. The true purpose and other advantages of this disclosure are realized and achieved by the structures and methods particularly indicated in the specification, claims and accompanying drawings. It should be understood that both the above summary description and the following detailed description are illustrative and descriptive, and are intended to further illustrate the invention as described in the claims. [Brief explanation of the drawing]

[0009] [Figure 1A] Figure 1A is a schematic diagram showing an embodiment of the system according to the present invention. [Figure 1B] Figure 1B is a schematic diagram showing another embodiment of the system according to the present invention. [Figure 1C] Figure 1C is a schematic diagram showing a collimator in the virtual image module of the system according to the present invention. [Figure 2] Figure 2 is a block diagram showing an embodiment of a system having various modules according to the present invention. [Figure 3A]Figure 3A is a schematic diagram (part 1) showing a possible embodiment of the system according to the present invention. [Figure 3B] Figure 3B is a schematic diagram (part 2) showing a possible embodiment of the system according to the present invention. [Figure 3C] Figure 3C is a schematic diagram showing an embodiment of a portable AR device. [Figure 3D] Figure 3D is another schematic diagram showing an embodiment of a portable AR device. [Figure 4] Figure 4 is a schematic diagram illustrating an embodiment of the relationship between an object, a real-time image, and a virtual image according to the present invention. [Figure 5] Figure 5 is a photograph in which a virtual image of the retina is superimposed on a real-time image according to the present invention. [Figure 6] Figure 6 is a flowchart showing an embodiment of the process for overlaying a virtual image onto a real-time image according to the present invention. [Figure 7] Figure 7 is a flowchart showing another embodiment of the process of overlaying a virtual image onto a real-time image according to the present invention. [Figure 8] Figure 8 is a schematic diagram showing an embodiment of the virtual image module according to the present invention. [Figure 9] Figure 9 is a schematic diagram showing the relationship between the virtual binocular pixels and their corresponding right and left pixels according to the present invention. [Figure 10] Figure 10 is a schematic diagram showing the optical path from the optical signal generator according to the present invention to the beam splitter and the observer's retina. [Figure 11] Figure 11 is a schematic diagram showing virtual binocular pixels formed by the right collimated light signal and the left collimated light signal according to the present invention. [Figure 12] Figure 12 is a table showing an embodiment of the lookup table according to the present invention. [Figure 13] Figure 13 is a schematic diagram illustrating a method for superimposing virtual binocular pixels onto a real object according to the present invention. [Modes for carrying out the invention]

[0010] The terms used in the following description are intended to be interpreted in the broadest and most reasonable way, even when used in conjunction with detailed descriptions of specific embodiments of the technology. Certain terms may be emphasized below, but terms intended to be interpreted more restrictively are specifically defined in this detailed description section.

[0011] The present invention relates to a system and method for superimposing a virtual image onto a real-time image or a real object. By superimposing a virtual image with depth onto a real-time image or a real object, more information about the real-time image or real object can be provided to the observer, such as for surgical guidance, instructions, and navigation. A real-time image is an image that reflects changes in a real object in real time. A real-time image may be a two-dimensional (2D) image or a three-dimensional (3D) image. In one embodiment, the real-time image is generated by light reflected or emitted from a real object and is observed, for example, by a microscope or telescope. In some embodiments, the real-time image may be reflected or emitted from a real object and observed with the naked eye without the use of a microscope or telescope. In another embodiment, the real-time image is generated by a display receiving an image of an object that would have been captured in real time by a camera, for example, an image on a display from an endoscope. The real-time image may also be a real image or a virtual image. A virtual image with depth is generated by projecting light signals onto the observer's eyes. The depth of the virtual image is related to the angle between the right collimated light signal and the corresponding left collimated light signal projected onto the observer's eyes. The virtual image may be a 2D image or a 3D image. When the virtual image is superimposed on a real-time image or a real object, a portion of the virtual image will overlap with the real-time image.

[0012] A system for superimposing a virtual image on a real-time image or a real object includes a real-time image module and a virtual image module. Generally, the real-time image module may include any optical element having light reflection, transmission, or refraction characteristics. An observer can observe a real-time image of a real object at a second position with a second depth. In some embodiments, the real-time image module may include a magnification assembly that magnifies the real-time image of the real object at the second position and the second depth by a predetermined magnification. Magnification is a process of magnifying the apparent size rather than the physical size of an object. This magnification is quantified by a calculated value also referred to as "magnification factor", which is the ratio of the apparent size of the object (in the real-time image) to the observed size of the real object without magnification. The magnification factor is adjustable and may be any positive number such as 0.5, 1, and 10. When the magnification factor is less than 1, it refers to a reduction in size and may also be referred to as reduction or demagnification.

[0013] The virtual image module generates a virtual image by projecting a right collimated light signal onto the observer's right eye and a corresponding left collimated light signal onto the observer's left eye, respectively. The right collimated light signal and the corresponding left collimated light signal are perceived by the observer to display a virtual image at a first position and a first depth. The first depth is related to the angle between the right collimated light signal projected onto the observer's eye and the corresponding left collimated light signal. In one embodiment, the first depth is approximately the same as the second depth.

[0014] The virtual image is superimposed on a real-time image or a real object to provide the observer with more information. Therefore, in one embodiment, the virtual image is a processed image of a real object. For example, the real object is the brain, and the real-time image is a real-time generated or processed image of the brain. However, in some embodiments, the observer may directly observe an image of the real object (e.g., without magnification, or when the image of the real object observed by the observer is not captured and reproduced in any way by a photodetector and display device). When the observer directly observes the real object, the observer may see the virtual image superimposed directly onto the real object. The virtual image may be a CT or MRI image of the brain taken before surgery, with the location of the brain tumor to be removed marked. The marked virtual image is displayed superimposed on a real-time image of the brain or the brain during surgery to help the surgeon locate the brain tumor to be removed. In this case, to accurately identify the surgical site, the first depth of the virtual image (the marked CT or MRI image) is approximately the same as the second depth of the real-time image or real object, i.e., the actual brain as captured by a surgical microscope. The virtual image may further include some text information, markers, pointers, etc., for guidance and explanation to support diagnosis and treatment. Furthermore, image overlay may allow observers to compare past images of real objects displayed in the virtual image with the current state of the real objects displayed in the real-time image, potentially enabling prediction of disease progression or treatment outcomes.

[0015] The magnification of the real-time image is adjustable. In one embodiment, such adjustment can be performed manually by turning a knob, changing the objective lens, controlling a virtual switch, or giving verbal instructions. After the real-time image has been magnified, the virtual image may be magnified manually or automatically to maintain the original superposition between the virtual image and the real-time image. An automatic superposition mode may also be selected.

[0016] To overlay a virtual image on a real-time image or a real object, it is first necessary to calibrate the system for the observer. Since the physical characteristics of the observer's eyes, such as the interpupillary distance (IPD), are different for each observer, the system needs to be calibrated specifically for the observer so that when the right collimated light signal and the left collimated light signal are projected onto the observer's eyes, the observer can perceive the virtual image displayed at the second position and the second depth. For example, the distance between the right and left eyepieces of a microscope needs to be adjusted according to the observer's interpupillary distance, and the angle between the right collimated light signal and the corresponding left collimated light signal needs to be adjusted so that the virtual image is accurately recognized by the observer at the second depth.

[0017] The process of overlaying a virtual image on a real-time image includes: (a) selecting a first point on the real-time image as a first landmark; (b) displaying the real-time image at a second position and a second depth at a predetermined magnification; and (c) projecting the virtual image by projecting the right collimated light signal onto the observer's right eye and the corresponding left collimated light signal onto the observer's left eye so that the observer can perceive the virtual image at the first position and the first depth, and making the corresponding first landmark on the virtual image overlap with the first landmark on the real-time image. As described above, the first depth is related to the angle between the right collimated light signal projected onto the observer's eye and the corresponding left collimated light signal. In one embodiment, the depth of the first landmark on the real-time image is approximately the same as the depth of the corresponding first landmark on the virtual image. To achieve a more accurate overlay, a second landmark or a third landmark may be used in a similar manner.

[0018] The process of superimposing a virtual image onto a real object includes (a) selecting a first point on the real object as a first landmark, and (b) projecting a virtual image by projecting a right collimated light signal onto the observer's right eye and a corresponding left collimated light signal onto the observer's left eye, so that the corresponding first landmark on the virtual image overlaps with the first landmark on the real object, so that the observer can perceive the virtual image at a first position and a first depth. As described above, the first depth is related to the angle between the right collimated light signal and the corresponding left collimated light signal projected onto the observer's eye. In one embodiment, the depth of the first landmark on the real object is approximately the same as the depth of the corresponding first landmark on the virtual image. A second or third landmark may be used in a similar manner to achieve more accurate superimposition.

[0019] As shown in Figures 1A and 1B, the system 100 for superimposing a virtual image 165 onto a real-time image 115 includes a real-time image module 110 and a virtual image module 160. The real-time image module 110 may include a magnification assembly 120 for generating magnified real-time images of an object 105, such as a brain, for both eyes of the observer. The magnification assembly 120 may include a plurality of optical units and assemblies, such as various lenses including an objective lens 113. In another embodiment, the magnification assembly 120 may process and magnify the real-time image of the real object 105 using electronic circuits. The magnification of the real-time image module may be determined before observation and adjustable during observation. The magnification may be 1 / 2, 1, 3, 10, 100x, etc. Magnification adjustment may be performed via a user interface that communicates with the real-time image module. The real-time image module may have one set of optical units and assemblies for generating real-time images for both eyes of the observer, or two sets of optical units and assemblies for generating real-time images for the observer's right eye and left eye, respectively. The real-time image module 110 may further include a prism assembly 130 for changing the direction of light, beam splitters 140, 145 for splitting light, an observation tube 150 for guiding light, and eyepieces 152, 154 for further magnifying the image. Again, the real-time image may be generated from light reflected or emitted from a real object 105, such as a real-time image generated by a microscope, including a surgical microscope. In another embodiment, the real-time image may be generated by an image acquisition device and display device, such as an endoscope and an associated display. Depending on the image size and resolution, the real-time image may actually or conceptually contain 921,600 pixels in a 1280 × 720 array. Each pixel may have a slightly different position and depth from its adjacent pixels. A representative pixel, such as a first landmark, may be selected for the real-time image.Landmarks, such as the first and second landmarks, are typically unique points with identifiable features that are easily recognizable by an observer in a real-time image, such as a center point or the intersection of two specific blood vessels. A landmark may be a single pixel or may consist of multiple adjacent pixels. In one embodiment, the position and depth of a representative pixel may be used as the second position and second depth of the real-time image.

[0020] The virtual image module 160, configured to connect to the real-time image module 110, includes a right collimated light signal generator 170 and a left collimated light signal generator 175. The right collimated light signal generator 170 generates multiple right collimated light signals for the virtual image and is often positioned close to the right side of the real-time image module. Similarly, the left collimated light signal generator 175 generates multiple left collimated light signals for the virtual image and is often positioned close to the left side of the real-time image module. The right collimated light signals are then redirected towards one of the observer's eyes by a right beam splitter 140. Similarly, the left collimated light signals are redirected towards the other of the observer's eyes by a left beam splitter 145. The redirected right collimated light signals and their corresponding redirected left collimated light signals are perceived by the observer to display a virtual image in a second depth. Depending on the image size and resolution, the virtual image may actually contain 921,600 virtual binocular pixels in a 1280 × 720 array. The virtual binocular pixels may have slightly different positions and depths from their adjacent pixels. A representative virtual binocular pixel, such as a first landmark, may be selected for the virtual image. In one embodiment, the position and depth of the representative virtual binocular pixel, i.e., a second position and a second depth, may be used for the virtual image. After the observer's eye receives the reoriented right collimated light signal and the corresponding reoriented left collimated light signal of the representative virtual binocular pixel, the observer perceives the representative virtual binocular pixel with a second depth related to the angle between such reoriented right collimated light signal and the corresponding reoriented left collimated light signal.

[0021] The light beam of the real-time image may pass through the right beam splitter 140 and the left beam splitter 145 towards the observer's eye. Therefore, the right beam splitter 140 and the left beam splitter 145 are shared to some extent by both the real-time image module and the virtual image module. In one embodiment, the light signal generated from the virtual image module can be redirected towards the observer's eye by rotating the beam splitter originally installed in the real-time image module by an appropriate angle to share the real-time image with other observers.

[0022] As shown in Figures 1B and 1C, the virtual image module 160 may further include a right focus adjustment unit 182 between the right collimated optical signal generator 170 (or right collimator 180, if available) and the right beam splitter 140, and a left focus adjustment unit 187 between the left collimated optical signal generator 175 (or right collimator 185, if available) and the left beam splitter 145, in order to improve the clarity of the virtual image for the observer. The left and right focus adjustment units may include optical units such as various lenses, including convex lenses. In one embodiment in which a convex lens is used as a focus adjustment unit, the focal position of the light beam changes by adjusting its distance from the optical signal generator, assuming that the distance between the optical signal generator and the beam splitter is the same. The closer the focal position of the light beam is to the retina, the clearer the virtual image for the observer. Since the axial length of the observer's eye may change, the preferred focal position of the light beam, and therefore the distance between the optical signal generator and the focus adjustment unit, will change accordingly. In other words, for observers with a long axial length of the eye, the focusing unit needs to be moved further away from the optical signal generator so that the focal point of the light beam is closer to the observer's retina. If a collimator is available, the focusing unit is placed between the collimator and the beam splitter. After passing through the collimator, the light beam from the optical signal generator becomes nearly parallel, and then converges after passing through the focusing unit. Also, since the focusing unit does not change the angle of incidence of the light beam, the depth of the virtual image is not affected.

[0023] As partially shown in Figure 1C, the virtual image module 160 may further include a right collimator 180 and a left collimator 185 that narrow the optical beams of multiple optical signals, for example, by aligning the direction of motion of the optical beams in a specific direction or by reducing the spatial cross-sectional area of ​​the optical beams. The right collimator 180 may be located between the right collimated optical signal generator 170 and the right beam splitter 140, and the left collimator 185 may be located between the left collimated optical signal generator 175 and the left beam splitter 145. The collimators may be curved mirrors or lenses.

[0024] The virtual image module 160 may also include a control module 190 that controls the virtual image signals to the right collimated optical signal generator 175 and the left collimated optical signal generator 175. The control module 190 communicates with the virtual image module 160 and adjusts the right collimated optical signal and its corresponding left collimated optical signal so that the virtual image can be automatically modified based on changes in the real-time image and superimposed on the real-time image. Changes in the real-time image include changes in field of view, magnification, or position. For example, if the magnification of the real-time image is adjusted from 3x to 10x, the control module 190 processes the image signal to enlarge the virtual image to the same size and continues to superimpose the virtual image on the real-time image using at least a first landmark. The control module 190 includes one or more processors, but to perform complex signal processing, the control module 190 may use an external server 250 for calculations.

[0025] The virtual image may be stored in the memory module 195. In one embodiment, the virtual image is a processed image, such as an X-ray, ultrasound, CT, or MRI image, of a real object with some mark or highlight in the region of interest. The virtual image may further include text information or pointers for guidance or explanation. For example, the virtual image may be a pre-captured and processed image of the retina of a patient with bleeding blood vessels marked for laser sealing. The system 100 can superimpose such a virtual image onto a real-time image of the same retina obtained from a slit-lamp microscope. The control module 190 may acquire the virtual image stored in the memory module 195 and then generate virtual image signals for the right collimated light signal generator 170 and the left collimated light signal generator 175 as needed.

[0026] As shown in Figure 2, in addition to the real-time image module 110 and the virtual image module 160, the system 100 may further include a recording module 210 for recording either or both real-time and virtual images, an object measurement module 220 for measuring the position and depth of real objects, a surgical module 230 for performing physical surgery on real objects 105, and a user interface 240 for an observer to communicate with various modules of the system 100 and control various functions of the system 100. All modules of the system 100 can communicate electronically with each other via wired or wireless connections. Wireless methods include WiFi, Bluetooth, near-field communication (NFC), the Internet, telecommunications, and radio frequency (RF). The real-time image module 110, the virtual image module 160, and the recording module 210 may communicate with each other optically via light beams and optical signals. The observer may observe real-time and virtual images through the system 100 and then control the system 100 through physical interaction with the user interface 240. System 100 may communicate optically with the real object 105 by receiving a light beam reflected or emitted from the real object, or by projecting a light beam onto the real object. System 100 may also perform physical interactions with the real object 105, such as performing laser surgery on the real object.

[0027] As described above, the system 100 may further include a recording module 210 for recording either or both real-time images and virtual images. In one embodiment, the recording module 210 may be positioned between the right beam splitter 140 and the left beam splitter 145 to record real-time images, which are light beams from real objects reflected by the right beam splitter and the left beam splitter, respectively, during surgery. The recording module 210 may include a digital camera or a charge-coupled device (CCD) for capturing images. In another embodiment, the recording module 210 may be positioned adjacent to the eyepiece to record light beams that pass through the eyepiece and reach the observer's eye, including both light beams that form the real-time image and light beams that form the virtual image. The recording module 210 may be connected to a control unit to directly record virtual image signals and associated information and parameters for later display.

[0028] As described above, system 100 may further include an object measurement module 220 that measures the position and depth of a real object. The real object measurement module 220, configured to connect to the system, may continuously or periodically measure the position and depth of the real object relative to the real object measurement module (or observer) and may transmit relevant information to the virtual image module to adjust the virtual image. Upon receiving such information, the control module 190 may process the virtual image signal based on the updated position and depth of the real object relative to the real object measurement module and the observer. As a result, the virtual image remains superimposed on the real image or the real object. The distance or relative position between the real object 105 and the real object measurement module 220 (or the observer's eye) may change over time. In some situations, the real object 105, such as a part of the human body like an eyeball, may move during surgery. In other situations, system 100 is worn by an observer, such as a surgeon, and the observer may move their head during surgery. Therefore, in order to continuously overlay the virtual image onto a real-time image or real object, it is necessary to measure and calculate the relative position and distance between the real object 105 and the observer's eye. The real object measurement module 220 may include a gyroscope, an indoor / outdoor Global Positioning System (GPS), and distance measuring components (e.g., emitters and sensors) to accurately track changes in the relative position and depth of the real object 105.

[0029] As described above, the system 100 may further include a surgical module 230 for physically performing surgery on a real object 105. The surgical module 230 may include a laser for removing tissue or blocking bleeding blood vessels, and / or a scalpel for cutting tissue. The surgical module 230 can work in conjunction with the real-time imaging module 110 to position the laser and / or scalpel toward a site of interest identified by an observer, such as a surgeon, as shown in the real-time image.

[0030] As described above, system 100 may further include a user interface 240 for the observer to control various functions of system 100, such as the real-time image magnification, the first position and first depth of the virtual image, the focus adjustment unit, the recording module 210, the real object measurement module 220, etc. The user interface 240 can be operated by voice, hand gestures, or finger / foot movements, and may take the form of a pedal, keyboard, mouse, knob, switch, stylus, button, stick, touchscreen, etc. The user interface 240 may communicate with other modules of system 100 (including the real-time image module 110, virtual module 160, recording module 210, real object measurement module 220, and surgical module 230) via wired or wireless means. Examples of wireless methods include WiFi, Bluetooth, near-field communication (NFC), the Internet, telecommunications, and radio frequency (RF). The observer may use the user interface 240, such as by manipulating a stick, to move a cursor to a region of interest on the real-time image, and then use the user interface 240, such as by pressing a pedal, to fire a laser beam toward the corresponding region of interest on the real object 105 to remove tissue or block bleeding blood vessels.

[0031] In one embodiment, the system 100 may be an AR microscope for surgery and / or diagnosis, such as an AR ophthalmoscope or an AR slit-lamp microscope. Figure 3A shows an example of a fixed AR surgical microscope 310 including a user interface pedal 320.

[0032] Figures 3B-3D show an example of a portable AR device 350 (a head-wearable device). Referring to Figure 3B, the portable AR device 350 includes a real-time image module 370 and a virtual image module 360. The real-time image module 370 allows real-time images from the surroundings or images of real objects to be incident on the user's eyes. In one embodiment of the present invention, the real-time image module 370 includes a combiner for the portable augmented reality device. The real-time image module 370 allows real-time images from the surroundings or images of real objects to be incident on the observer's eyes, while it may reflect virtual images generated by the virtual image module 360 ​​towards the observer's eyes. Referring to Figures 3C-3D, which show two exemplary embodiments of the portable AR device, the portable AR device may include at least a virtual image module 160 and a real-time image module 370. The real-time image module 370 may include a combiner 8 for allowing ambient light to be incident on the observer's eyes. In some embodiments, the combiner 8 may also reflect images provided by the virtual image module 160 towards the observer's eyes. In this embodiment, the virtual image module 160 may include a laser emitter and other optical elements such as lenses and reflectors. The real-time image module 370 may or may not have a function to magnify the real-time image. In some embodiments, the real-time image module 370 may further include a beam splitter, as shown in Figure 3C. In one embodiment, the real-time image module 370 does not have a magnification function, and the user can see the real object directly, thereby allowing the user to see the virtual object superimposed directly onto the real object.

[0033] As shown in Figure 4, the real object 105, the real-time image 115 generated by the real-time image module 110, and the virtual image 165 generated by the virtual image module 160 may have different positions and depths. In this embodiment, the virtual image 165 is a processed partial image of the real object 105. The virtual image module 160 may generate the virtual image 165 only for the field of view or region of interest of the real object. Images of the real object may be captured and processed, for example, by an artificial intelligence (AI) module to generate virtual images at very short time intervals, such as one second.

[0034] As mentioned above, depending on the resolution, the real object 105, the real-time image 115, and the virtual image 165 may conceptually or practically have a large number of pixels, such as 921,600 pixels in a 1280 × 720 array. In this embodiment, the position and depth of the real object 105, the real-time image 115, and the virtual image 165 are represented by the position and depth of the corresponding first landmark, respectively. The depth is measured based on the distance between the eyepiece 152 and the real object 105, the real-time image 115, or the virtual image 165. Thus, as shown in Figure 4, the real object 105 is located at the real object position L(o) and object depth D(o), the real-time image 115, which is an enlarged image of the real object 105, is located at the second position L(r) and second depth D(r), and the virtual image 165 is located at the first position L(v) and first depth D(v). Depending on the optical characteristics of the real-time image module, the depth of the real-time image 115 may be closer to or further away from the observer's eye. In this embodiment, the depth of the real-time image D(r) is greater than the depth of the real object D(o). However, in other embodiments, the depth of the real-time image D(r) may be less than or approximately the same as the depth of the real object D(o). The virtual image module 160 then generates a virtual image 165 at a depth D(v) closer to the eyepiece than the real-time image 115.

[0035] As shown in Figure 4, using L(r) and D(r) information, the virtual image module 160 of system 100 may superimpose a virtual image onto a real-time image by superimposing the corresponding first landmark LM1(v) on the virtual image onto the first landmark LM1(r) on the real-time image. For a more advanced superimposition, the virtual image module 160 of system 100 may further superimpose the corresponding second landmark LM2(v) on the virtual image onto the second landmark LM2(r) on the real-time image. In another embodiment where the superimposition goes beyond superimposition with respect to the location of the landmarks, the depth of the corresponding first landmark on the virtual image may be approximately the same as the depth of the first landmark on the real-time image. Similarly, the depth of the corresponding second landmark on the virtual image may be approximately the same as the depth of the second landmark on the real-time image. In order to accurately and completely superimpose a 3D virtual image onto a 3D real-time image, a third landmark on the real-time image is selected in addition to the first and second landmarks. The virtual image module then makes the position and depth of the corresponding third landmark in the virtual image approximately the same as the third landmark in the real-time image.

[0036] Figure 5 shows three images: a real-time image of the patient's retina, a processed virtual image of the retina, and an image superimposed on both. In one embodiment, an angiographic image of the patient's retina may be acquired and processed by a slit-lamp biomicroscope. The virtual image module 160 may then use such processed images to project a virtual image superimposed on a real-time image of the patient's retina during surgery, assisting in the identification and visualization of the edges of choroidal neovascularization membranes. AR / MR microscopes have the potential to significantly advance the diagnosis and treatment of various ophthalmic diseases.

[0037] As shown in Figure 6, the process of overlaying a virtual image onto a real-time image or real object involves four steps. In step 610, a first point on the real-time image or real object is selected as the first landmark by the observer, expert, computer, or system 100. For example, the observer may use a mouse to move the cursor or pointer visible through the eyepiece to select the first landmark on the real-time image or real object. As described above, landmarks, including the first, second, and third landmarks, are typically unique points with identifiable features that are easily recognizable to the observer in the real-time image or real object, such as a central point or the intersection of two specific blood vessels. Landmarks may be defined manually by an expert or automatically by a computer program. There are three basic types of landmarks: anatomical landmarks, mathematical landmarks, and pseudo-landmarks. Anatomical landmarks are biologically meaningful points in an organism. Anatomical features that are consistently present within tissues, such as folds, protrusions, tubes, and blood vessels, help indicate specific structures or locations. Anatomical landmarks may be used by surgical pathologists to orient specimens. Mathematical landmarks are points within a figure that are arranged according to some mathematical or geometrical property, such as points of high curvature or extreme values. A computer program may determine the mathematical landmarks used for automated pattern recognition. Pseudo-landmarks are construction points placed between anatomical or mathematical landmarks. A typical example is a set of equally spaced points between two anatomical landmarks to obtain more sample points from a shape. Pseudo-landmarks are useful for shape matching when a large number of points are required in the matching process. A landmark may be a single pixel or may consist of multiple pixels adjacent to each other.

[0038] In step 620, if the observer is viewing a real-time image, the real-time image of the real object is displayed at a predetermined magnification at a first position and a first depth. If the observer is directly observing the real object, neither the real object nor the real-time image is magnified. There are at least two types of real-time images. The first type of real-time image is generated by light reflected or emitted from the real object and is, for example, an image observed with a microscope or telescope. In this case, the first position and first depth may be determined by the optical features of the real-time image module. The observer may view the real-time image through an eyepiece. The second type of real-time image is generated by a display that receives an image of the object captured in real time by a camera, for example, an image on a display from an endoscope including a gastroscope, colonoscope, or proctoscope. The endoscope may have two separately positioned imaging devices to capture and generate 3D images. The real-time image may be a two-dimensional (2D) image or a three-dimensional (3D) image. Steps 610 and 620 are interchangeable.

[0039] In step 630, the virtual image module is calibrated for a specific observer. As previously mentioned, certain physical characteristics of each observer, such as interpupillary distance, may affect the position and depth of the virtual image perceived by the observer with the same right collimated light signal and its corresponding left collimated light signal. In one embodiment, the control module adjusts the virtual image signals based on the observer's IPD so that the right collimated light signal generator 170 and the left collimated light signal generator 175 project light signals at the appropriate position and angle, ensuring that the observer perceives the virtual image accurately at a first position and first depth.

[0040] In step 640, the virtual image module projects a virtual image so that the observer perceives the virtual image at a first position and a first depth by projecting a right collimated light signal to the observer's right eye and a corresponding left collimated light signal to the left eye, respectively. As a result, the corresponding first landmark on the virtual image overlaps with the first landmark on the real-time image or real object. In other words, the virtual image module projects a virtual image superimposed on the real-time image or real object. At least the position of the corresponding first landmark on the virtual image (first position) is approximately the same as the position of the first landmark on the real-time image or real object (second position). Generally, the virtual image is divided into multiple virtual binocular pixels depending on the resolution, for example, 921,600 virtual binocular pixels in a 1280 × 720 array. For each right collimated light signal and its corresponding left collimated light signal projected onto the observer's retina, the observer perceives the virtual binocular pixels at a specific position and depth. The depth is related to the angle between the right collimated light signal and its corresponding left collimated light signal projected onto the observer's eye. More specifically, the perceived depth of the virtual binocular pixel corresponds to the depth position of the convergence point of the optical path extensions of the right collimated light signal and its corresponding left collimated light signal. In other words, the perceived depth of the virtual binocular pixel is substantially equal to the depth position of the convergence point of the optical path extensions of the right collimated light signal and its corresponding left collimated light signal. When a first landmark on a real-time image or real object is at a second position and second depth, the virtual binocular pixel of the corresponding first landmark on the virtual image is projected so that it is perceived by the observer as being at the first position and first depth. In the initial superposition, the position of the corresponding first landmark on the virtual image (first position) is set to be approximately the same as the position of the first landmark on the real-time image or real object (second position), but the depth may be different. This superposition may be performed manually by an observer, or automatically by system 100 using shape recognition technology, including an artificial intelligence (AI) algorithm. To further improve the superposition, the first depth is set to be approximately the same as the second depth.When superimposing a virtual image onto a real-time image, if the real-time image is enlarged beyond the actual size of the real object, the virtual image must also be enlarged to the same size for superimposition. Furthermore, to further improve the superimposition, the field of view of the virtual image must match the field of view of the real-time image. The relationship between the optical signal generated by the optical signal generator and the depth perceived by the observer will be explained in detail below. When superimposing a virtual image onto a real object, this can be achieved by placing the convergence point near a portion of the real object.

[0041] In step 650, if the position, magnification, or field of view of the real-time image changes, the virtual image module modifies the virtual image to maintain the superimposition between the virtual image and the real-time image. Similarly, if the position or field of view of a real object changes, the virtual image module modifies the virtual image to maintain the superimposition between the virtual image and the real object. Changes in the position, magnification, and field of view of the real image or real object may be caused by observer actions or by the movement of the real object or the observer. System 100 constantly monitors the first position and first depth of the real-time image and the second position and second depth of the virtual image. If any change occurs in the real-time image or real object, the virtual image module modifies the virtual image signal to maintain the superimposition between the virtual image and the real image or real object.

[0042] As shown in Figure 7, the alternative process for overlaying a virtual image onto a real-time image comprises six steps. Some steps are identical or similar to those described in the previously described embodiment shown in Figure 6. Some steps are optional and can be further modified. In step 710, a first point, a second point, and a third point on the real-time image or a real object are selected by the observer, expert, computer, or system 100 as the first landmark, second landmark, and third landmark, respectively. Three landmarks are used here to achieve the most accurate overlay. In some surgeries, such as neurosurgery, very high precision is required, so three landmarks may be necessary to ensure that the virtual image is perfectly overlaid onto the real-time image or the real object. However, if necessary, this process may also include two landmarks. Step 720 may be the same as step 620, and step 730 may be the same as step 630. Step 740 follows the same principle as described in step 640. However, the positions and depths of the corresponding first, second, and third landmarks on the virtual image are approximately the same as the positions and depths of the first, second, and third landmarks on the real-time image or the real object, respectively. In step 750, the second position and second depth are repeatedly monitored or measured. The second position and second depth may be calculated based on the position and depth of the real object relative to the real object measurement module (or observer), as measured by the real object measurement module. As a result, the virtual image may remain superimposed on the real image or the real object. In step 760, the observer, for example, a surgeon, performs surgery on the real object using a laser or scalpel at the site of interest identified by the observer.

[0043] A virtual image module 160, a method for generating a virtual image 165 at a first position and a first depth, and a method for moving the virtual image as desired are described in detail below. PCT international application PCT / US20 / 59317, filed on 6 November 2020, entitled "System and method for displaying objects with depth," is incorporated herein by reference in its entirety. According to one embodiment, as shown in Figure 8, the virtual image module 160 includes a right collimated light signal generator 170 that generates multiple right collimated light signals such as 12 for RLS_1, 14 for RLS_2, and 16 for RLS_3; a right beam splitter 140 that receives the multiple right collimated light signals and redirects them toward the observer's right retina 54; a left collimated light signal generator 175 that generates multiple left collimated light signals such as 32 for LLS_1, 34 for LLS_2, and 36 for LLS_3; and a left beam splitter 145 that receives the multiple left collimated light signals and redirects them toward the observer's left retina 64. The observer has a right eye 50 including a right pupil 52 and a right retina 54, and a left eye 60 including a left pupil 62 and a left retina 64. The diameter of a human pupil is generally in the range of 2 to 8 mm, depending in part on ambient light. The normal pupil size of an adult is 2-4 mm in diameter in bright light and 4-8 mm in dim light. Multiple right collimated light signals are redirected by the right beam splitter 140, pass through the right pupil 52, and are finally received by the right retina 54. Right collimated light signal RLS_1 is the rightmost light signal that the observer's right eye can see on a given horizontal plane. Right collimated light signal RLS_2 is the leftmost light signal that the observer's right eye can see on the same horizontal plane. Upon receiving the redirected right collimated light signals, the observer recognizes multiple right pixels of the real object 105 in region A, which is enclosed by the extensions of the redirected right collimated light signals RLS_1 and RLS_2. Region A is referred to as the field of view (FOV) of the right eye 50. Similarly, multiple left collimated light signals are redirected by the left beam splitter 145, pass through the center of the left pupil 62, and are ultimately received by the left retina 64.The left collimated light signal LLS_1 is the rightmost light signal visible to the observer's left eye on a given horizontal plane. The left collimated light signal LLS_2 is the leftmost light signal visible to the observer's left eye on the same horizontal plane. When the observer receives the reoriented left collimated light signals, they perceive multiple left pixels of the real object 105 in region B, which is demarcated by the extensions of the reoriented left collimated light signals LLS_1 and LLS_2. Region B is referred to as the field of view (FOV) of the left eye 60. When multiple right and left pixels are displayed in region C, where regions A and B overlap, at least one right collimated light signal displaying one right pixel and the corresponding left collimated light signal displaying one left pixel are merged to display virtual binocular pixels with a specific depth in region C. This depth is related to the angle between the reoriented right collimated light signal and the reoriented left collimated light signal projected onto the observer's retina. Such angles are also called convergence angles.

[0044] As shown in Figures 8 and 9, the observer perceives a virtual image of the brain object 105 having multiple depths in a region C in front of the observer. The image of the brain object 105 includes a first virtual binocular pixel 72 displayed at a first depth D1 and a second virtual binocular pixel 74 displayed at a second depth D2. The first angle between the optical path extension of the first reoriented right collimated light signal 16' and the corresponding optical path extension of the first reoriented left collimated light signal 26' is θ1. The first depth D1 is associated with the first angle θ1. In particular, the first depth of the first virtual binocular pixel of the real object 105 can be determined by the first angle θ1 between the optical path extension line of the first reoriented right collimated light signal and the corresponding optical path extension line of the first reoriented left collimated light signal. As a result, the first depth D1 of the first virtual binocular pixel 72 can be approximately calculated using the following formula. Tan(θ / 2) = IPD / 2D The distance between the right pupil 52 and the left pupil 62 is the interpupillary distance (IPD). Similarly, the second angle between the second directional right collimated light signal 18' and the corresponding second directional left collimated light signal 38' is θ2. The second depth D2 is related to the second angle θ2. In particular, the second depth D2 of the second virtual binocular pixel of the real object 105 can be approximately determined by the second angle θ2 between the optical path extension of the second directional right collimated light signal and the optical path extension of the corresponding second directional left collimated light signal, using the same formula. Since the second virtual binocular pixel 74 is perceived by the observer as being farther away (i.e., at a greater depth) than the first virtual binocular pixel 72, the second angle θ2 is smaller than the first angle θ1.

[0045] Furthermore, the reoriented right collimated light signal 16' of RLG_2 and the corresponding reoriented left collimated light signal 36' of LLS_2 together display a first virtual binocular pixel 72 having a first depth D1, although the reoriented right collimated light signal 16' of RLG_2 may have the same or a different field of view as the corresponding reoriented left collimated light signal 36' of LLS_2. In other words, the first angle θ1 determines the depth of the first virtual binocular pixel, but the reoriented right collimated light signal 16' of RLG_2 may or may not have parallax with the corresponding reoriented left collimated light signal 36' of LLS_2. Therefore, the red, blue, and green (RGB) intensities and / or brightness of the right and left collimated light signals may be nearly the same or slightly different depending on the hue and field of view to better present some 3D effect.

[0046] As described above, multiple right collimated optical signals are generated by a right collimated optical signal generator, redirected by a right beam splitter, and then scanned directly onto the right retina to form a right retinal image on the right retina. Similarly, multiple left collimated optical signals are generated by a left collimated optical signal generator, redirected by a left beam splitter, and then scanned directly onto the left retina to form a left retinal image on the left retina. In the embodiment shown in Figure 9, the right retinal image 80 contains 36 right pixels in a 6x6 array, and the left retinal image 90 also contains 36 left pixels in a 6x6 array. In another embodiment, the right retinal image 80 contains 921,600 right pixels in a 1280x720 array, and the left retinal image 90 also contains 921,600 left pixels in a 1280x720 array. The virtual image module 160 is configured to generate multiple right collimated optical signals and their corresponding multiple left collimated optical signals, forming a right retinal image on the right retina and a left retinal image on the left retina, respectively. As a result, the observer perceives a virtual image with a specific depth in region C through image fusion.

[0047] Each left or right pixel is formed by a single collimated optical signal. Therefore, each left or right pixel has its own unique actual optical path in real space. This feature differs from conventional waveguide-type headwearable displays where the optical signal of each pixel is scattered without being collimated.

[0048] Referring to Figure 9, the first right collimated light signal 16 from the right collimated light signal generator 170 is received and reflected by the right beam splitter 140. The first redirected right collimated light signal 16' passes through the right pupil 52 and reaches the observer's right retina, displaying the right pixel R43. The corresponding left collimated light signal 36 from the left collimated light signal generator 175 is received and reflected by the left beam splitter 145. The first redirected light signal 36' passes through the left pupil 62 and reaches the observer's left retina, displaying the left retinal pixel L33. As a result of image fusion, the observer perceives a virtual image with multiple depths, where the depth is determined by the angle between the multiple redirected right collimated light signals and their corresponding multiple redirected left collimated light signals. The angle between the redirected right collimated light signal and its corresponding left collimated light signal is determined by the relative horizontal distance between the right and left pixels. Therefore, the depth of a virtual binocular pixel is inversely proportional to the relative horizontal distance between the right pixel and its corresponding left pixel that form the virtual binocular pixel. In other words, the deeper the virtual binocular pixel is perceived by the observer, the smaller the relative horizontal distance in the X-axis between the right and left pixels that form such a virtual binocular pixel. For example, as shown in Figure 9, the observer perceives the second virtual binocular pixel 74 as having a greater depth (i.e., being further from the observer) than the first virtual binocular pixel 72. Therefore, the horizontal distance between the second right pixel and the second left pixel in the retinal image is smaller than the horizontal distance between the first right pixel and the first left pixel. Specifically, the horizontal distance between the second right pixel R41 and the second left pixel L51 that constitute the second virtual binocular pixel is 4 pixels long. On the other hand, the distance between the first right pixel R43 and the first left pixel L33 that constitute the first virtual binocular pixel is 6 pixels long.

[0049] In one embodiment shown in Figure 10, the optical paths of multiple right collimated optical signals and multiple left collimated optical signals from an optical signal generator to the retina are shown. Multiple right collimated optical signals generated from the right collimated optical signal generator 170 are projected onto the right beam splitter 140 to form a right splitter image (RSI) 82. These multiple right collimated optical signals are redirected by the right beam splitter 140 and converge into a small right pupillary image (RPI) 84, passing through the right pupil 52, and then finally reaching the right retina 54 to form a right retinal image (RRI) 86. Each of the RSI, RPI, and RRI consists of i × j pixels. Each right collimated optical signal RLS(i,j) passes through the corresponding same pixels from RSI(i,j) to RPI(i,j) and then to RRI(x,y). For example, RLS(5,3) proceeds from RSI(5,3) to RPI(5,3) and then to RRI(2,4). Similarly, multiple left collimated optical signals generated from the left collimated optical signal generator 175 are projected onto the left beam splitter 145 to form the left splitter image (LSI) 92. These multiple left collimated optical signals are redirected by the left beam splitter 145 and converge into a small left pupillary image (LPI) 94, passing through the left pupil 62, and then finally reaching the left retina 64 to form the right retinal image (LRI) 96. Each of the LSI, LPI, and LRI consists of i × j pixels. Each left collimated optical signal LLS(i,j) passes through the same corresponding pixels from LCI(i,j) to LPI(i,j) and then to LRI(x,y). For example, LLS(3,1) proceeds from LCI(3,1) to LPI(3,1) and then to LRI(4,6). The (0,0) pixel is the topmost and leftmost pixel in each image. Pixels in the retinal image are inverted horizontally and vertically with respect to their corresponding pixels in the splitter image. Based on the appropriate relative position and angle of the optical signal generator and beam splitter, each optical signal has its own optical path from the optical signal generator to the retina. A combination of one right collimated optical signal displaying one right pixel on the right retina and a corresponding left collimated optical signal displaying one left pixel on the left retina creates a virtual binocular pixel with a specific depth perceived by the observer.Thus, virtual binocular pixels in space can be represented by pairs of right retinal pixels and left retinal pixels, or pairs of right splitter pixels and left splitter pixels.

[0050] The virtual image perceived by the observer within region C includes multiple virtual binocular pixels. To accurately describe the position of each virtual binocular pixel in space, each position in space is assigned a three-dimensional (3D) coordinate, such as an XYZ coordinate. In other embodiments, other 3D coordinate systems may be used. As a result, each virtual binocular pixel has 3D coordinates in the horizontal, vertical, and depth directions. The horizontal direction (or X-axis direction) is aligned with the interpupillary line. The vertical direction (or Y-axis direction) is aligned with the midline of the face and is perpendicular to the horizontal direction. The depth direction (or Z-axis direction) is perpendicular to the frontal plane and is perpendicular to both the horizontal and vertical directions. In this invention, the horizontal and vertical coordinates are collectively referred to as position.

[0051] Figure 11 shows the relationship between pixels in the right splitter image, pixels in the left splitter image, and virtual binocular pixels. As described above, pixels in the right splitter image correspond one-to-one with pixels (right pixels) in the right retinal image. Pixels in the left splitter image correspond one-to-one with pixels (left pixels) in the left retinal image. However, pixels in the retinal image are inverted horizontally and vertically compared to their corresponding pixels in the combiner image. However, if eyepieces 152 and 154 are available in system 100, the relationship between pixels in the splitter image and corresponding pixels in the retinal image may be further modified by the optical characteristics of the eyepieces. For a right retinal image containing 36 (6×6) right pixels and a left retinal image containing 36 (6×6) right pixels, assuming all light signals are within the observer's binocular field of view (FOV), region C contains 216 (6×6×6) virtual binocular pixels (represented as dots). The optical path extension of one diverted right collimated light signal intersects with the optical path extension of each diverted left collimated light signal in the same row of the image. Similarly, the optical path extension of one diverted left collimated light signal intersects with the optical path extension of each diverted right collimated light signal in the same row of the image. Thus, there are 36 (6x6) virtual binocular pixels in one layer, and there are six layers in space. Although shown as parallel lines in Figure 11, there is usually a small angle between two adjacent lines indicating that the optical path extensions intersect to form virtual binocular pixels. Right pixels and their corresponding left pixels that are at approximately the same height in each retina (i.e., in the same row of the right and left retinal images) tend to fuse earlier. As a result, right pixels pair with left pixels in the same row of the retinal image to form virtual binocular pixels.

[0052] As shown in Figure 12, a lookup table is created for each virtual binocular pixel to easily identify the pair of right and left pixels. For example, 216 virtual binocular pixels numbered from 1 to 216 consist of 36 (6x6) right pixels and 36 (6x6) left pixels. The first virtual binocular pixel VBP(1) represents the pair of right pixel RRI(1,1) and left pixel LRI(1,1). The second virtual binocular pixel VBP(2) represents the pair of right pixel RRI(2,1) and left pixel LRI(1,1). The seventh virtual binocular pixel VBP(7) represents the pair of right pixel RRI(1,1) and left pixel LRI(2,1). The 37th virtual binocular pixel VBP(37) represents the pair of right pixel RRI(1,2) and left pixel LRI(1,2). The 216th virtual binocular pixel VBP(216) represents the pair of right pixel RRI(6,6) and left pixel LRI(6,6). Thus, it is determined which pair of right and left pixels can be used to generate the corresponding right and left collimated light signals in order to display a particular virtual binocular pixel of the virtual image in the observer's space. Each row of virtual binocular pixels on the lookup table also contains a pointer that leads to a memory address storing the perceived depth (z) and perceived position (x,y) of the VBP. Additional information such as the size scale, the number of overlapping objects, and the depth in sequence depth can also be stored in the VBP. The size scale may be relative size information comparing a particular VBP to a standard VBP. For example, if the virtual image is displayed on a standard VBP 1m in front of the observer, the size scale may be set to 1. Consequently, for a particular VBP 90cm in front of the observer, the size scale may be set to 1.2. Similarly, for a particular VBP located 1.5m in front of the observer, the size scale may be set to 0.8. The size scale can be used to determine the size of the virtual image for display when the virtual image is moved from a first depth to a second depth. The size scale may be a magnification in this invention. The number of overlapping objects is the number of objects that overlap so that one object is completely or partially hidden behind another object.Sequence depth provides information about the depth order of various overlapping images. For example, if three images are overlapping, the sequence depth of the first image in the foreground may be set to 1, and the sequence depth of the second image hidden behind it may be set to 2. The number of overlapping images and their sequence depths may be used to determine which images to display and which parts of the images to display when the various overlapping images are moving.

[0053] A lookup table may be created by the following process: In the first step, an individual virtual map is obtained based on the individual's IPD created by the virtual image module during initialization or calibration. This map specifies the boundary of region C, where the observer can perceive a virtual image with depth through the fusion of the right and left retinal images. In the second step, for each depth along the Z axis (each point in the Z coordinate), the convergence angle is calculated to identify pairs of right and left pixels on the right and left retinal images, respectively, regardless of their X and Y coordinate positions. In the third step, the pairs of right and left pixels are moved along the X axis to identify the X and Z coordinates of each pair of right and left pixels at a given depth, regardless of their Y coordinate position. In the fourth step, the pairs of right and left pixels are moved along the Y axis to determine their Y coordinates. As a result, the three-dimensional coordinate systems, such as XYZ, of each pair of right and left pixels on the right and left retinal images, respectively, can be determined, and a lookup table can be created. Furthermore, the third and fourth steps are interchangeable.

[0054] The optical signal generators 170 and 175 may use as their light source a laser, mini-LEDs and micro-LEDs, organic light-emitting diodes ("OLEDs") or light-emitting diodes ("LEDs") including superluminescent diodes ("SLDs"), LCoS (Liquid Crystal on Silicon), liquid crystal displays ("LCDs"), or any combination thereof. In one embodiment, the optical signal generators 170 and 175 are a laser beam scanning projector (LBS projector), which may include a red light laser, a green light laser, and a blue light laser, a color modulator such as a dichroic combiner or a polarizing combiner, and a two-dimensional (2D) adjustable reflector such as a 2D electromechanical system ("MEMS") mirror. The 2D adjustable reflector can be replaced with two one-dimensional (1D) reflectors, such as two 1DMEMS mirrors. The LBS projector sequentially generates and scans optical signals one by one to form a 2D image at a predetermined resolution, for example, 1280 × 720 pixels per frame. In this way, one optical signal is generated for each pixel and projected simultaneously toward the beam splitters 140 and 145. For an observer to see such a 2D image with one eye, the LBS projector needs to sequentially generate, for example, 1280 × 720 optical signals for each pixel within the visual afterimage period, for example, 1 / 18 of a second. Therefore, the duration of each optical signal is approximately 60.28 nanoseconds.

[0055] In another embodiment, the optical signal generators 170 and 175 may be digital light processing projectors ("DLP projectors") capable of generating a 2D color image at once. Texas Instruments' DLP technology is one of several technologies that may be used to manufacture DLP projectors. For example, an entire 2D color image frame, which may consist of 1280 x 720 pixels, is projected simultaneously toward the splitters 140 and 145.

[0056] The beam splitters 140 and 145 receive and redirect multiple optical signals generated by the optical signal generators 170 and 175. In one embodiment, the beam splitters 140 and 145 reflect the multiple optical signals so that the redirected optical signals are on the same side as the incident optical signals of the beam splitters 140 and 145. In another embodiment, the beam splitters 140 and 145 refract the multiple optical signals so that the redirected optical signals are on a different side from the incident optical signals of the beam splitters 140 and 145. When the beam splitters 140 and 145 function as refractors, the reflectivity can vary over a wide range, such as 20% to 80%, depending in part on the output of the optical signal generators. Those skilled in the art know how to determine an appropriate reflectivity based on the characteristics of the optical signal generators and splitters. Also, in one embodiment, the beam splitters 140 and 145 are optically transparent to ambient light from the opposite side of the incident optical signal, so that an observer can simultaneously observe a real-time image. The degree of transparency can vary greatly depending on the application. For AR / MR applications, transparency of 50% or more is preferable, and in one embodiment it is approximately 75%. In addition to changing the direction of the light signals, the focus adjustment units 182 and 187 focus them so that multiple light signals can pass through the pupil and reach the retinas of both eyes of the observer.

[0057] The beam splitters 140 and 145 are made of a lens-like glass or plastic material and may be coated with a specific material such as metal to be partially transparent and partially reflective. One advantage of using reflective splitters instead of conventional waveguides to guide the optical signal to the observer's eye is that it eliminates problems of undesirable diffraction effects such as multiple shadows and color shifts.

[0058] Referring to Figure 13, this figure shows an example of displaying a virtual image superimposed on a real object 205. The virtual image referred to in this invention consists of a plurality of virtual binocular pixels, each virtual binocular pixel formed by a pair consisting of a right pixel and a left pixel. In Figure 13, for simplification, only the first virtual binocular pixel 72 and the second virtual binocular pixel 74 are shown. These two virtual binocular pixels are superimposed on the real object 205. The first virtual binocular pixel 72 is superimposed on the first part of the real object 205 having a depth D1 (both the first virtual binocular pixel 72 and the first part of the real object 205 have the same depth D1) from the perspective of the observer or the real object measurement module. The second virtual binocular pixel 74 is superimposed on the second part of the real object 205 having a depth D2 (both the second virtual binocular pixel 74 and the second part of the real object 205 have the same depth D2) from the perspective of the observer or the real object measurement module. As described above, the depth of a virtual binocular pixel is determined by the convergence angle between the optical path extensions of the collimated light signals that constitute the right and left pixels. Furthermore, the depth of a virtual binocular pixel is determined by the depth coordinates of the convergence point of the optical path extensions of the collimated light signals that constitute the right and left pixels. By using a collimated light source, when an observer fixates on a virtual binocular pixel, the visual axes of both eyes approximately coincide with the optical paths (or optical path extensions) of the right and left pixels. Therefore, the observer's fixation position coincides with the position of the virtual binocular pixel rendered in real space (the convergence point of the optical path extensions of the collimated light signals). This feature is extremely important for eliminating focal rivalry and vergence accommodation conflict. Moreover, this feature can also be applied when superimposing virtual binocular pixels onto real objects in real space.

[0059] In this invention, the right pixel is rendered by the right collimated light signal, and the left pixel is rendered by the left collimated light signal. The fusion of the right and left pixels generates a virtual binocular pixel. Each virtual binocular pixel in the virtual image may have a different depth. The depth of each virtual binocular pixel perceived by the observer is modified, as described above, by changing the convergence angle between the optical path extensions (16”, 18”, 36”, and 38”) of the right collimated light signal and its corresponding left collimated light signal projected onto the observer's eye, based on the interpupillary distance. The perceived depth of the virtual binocular pixel corresponds to the depth position (or coordinates) of the convergence point of the optical path extensions of the right collimated light signal and its corresponding left collimated light signal. More specifically, theoretically, the perceived depth position of the virtual binocular pixel is the depth position (or coordinates) of the convergence point of the optical path extensions of the right collimated light signal and its corresponding left collimated light signal. However, due to various variations and factors (e.g., variations in machine dimensions, human physiological factors, etc.), the perceived depth (coordinates) is actually substantially the same as, but not exactly the same as, the depth (coordinates) of the convergence point. In any case, since the depth and position of the virtual binocular image perceived by the observer in real space are determined by the position of the convergence point in real space, the superposition of virtual binocular pixels onto a real object in real space can be achieved by setting the projection angles of the right collimated light signal and the left collimated light signal so that the convergence point is close to the target portion of the real object, or preferably as close as possible to the superposition target position on the real object.

[0060] The present invention is particularly advantageous over conventional techniques due to its ability to render different depths for each binocular pixel in a virtual image. Since the surface of a real object may have a contour, conventional 3D rendering techniques cannot set the depth for each binocular pixel to match the contour of the real object, making it difficult to perfectly superimpose the virtual image onto the real object. However, in the present invention, by bringing the convergence point of the collimated light signals (left and right) as close as possible to the real object (for example, the surface of the real object), the depth for each virtual binocular pixel can be set to approach the target portion of the real object. Furthermore, when an observer simultaneously observes the virtual image superimposed on the real object, focus competition does not occur. This is because the convergence point of the optical path extensions of the binocular pixels lies on the real object, and the position of the convergence point of the observer's visual axis is the same as the convergence point of the optical path extensions of the binocular pixels.

[0061] The above-mentioned descriptions of the embodiments are provided to enable those skilled in the art to manufacture and use the subject matter. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the novel principles and subject matter disclosed herein may be applied to other embodiments without employing innovative capabilities. The subject matter described in the claims is not intended to be limited to the embodiments shown herein, but should be given the broadest scope consistent with the principles and novel features disclosed herein. Additional embodiments are considered to be within the spirit and true scope of the disclosed subject matter. Accordingly, the present invention is intended to cover modifications and variations that fall within the scope of the appended claims and their equivalents.

Claims

1. A right collimated light signal generator for generating a right collimated light signal directed towards one retina of the observer, The system includes a left collimated light signal generator for generating a left collimated light signal corresponding to the right collimated light signal, which is directed towards the other retina of the observer, The right collimated light signal and the left collimated light signal form binocular pixels of a virtual image having a first depth, and the first depth perceived by the observer is modified by changing the convergence angle between the optical path extensions of the right collimated light signal and the corresponding left collimated light signal projected onto the observer's eye, based on the interpupillary distance, and the first depth corresponds to the depth position of the convergence point of the optical path extensions of the right collimated light signal and the corresponding left collimated light signal. The virtual image is extended and superimposed onto a real object by setting the convergence point in proximity to a part of the real object having a second depth. The first depth and the magnification of the virtual image are changed according to the second depth. The aforementioned virtual image module is included in a portable augmented reality device. The portable augmented reality device comprises a real-time image module, and the real-time image module comprises a combiner for the portable augmented reality device. A virtual image module for generating virtual images with depth.

2. A virtual image module for generating a virtual image having the depth described in claim 1, wherein the first depth and the magnification of the virtual image are changed according to the second depth.

3. A virtual image module for generating a virtual image having depth according to claim 1, wherein the virtual image is a photograph, magnetic resonance image, X-ray image, computed tomography, and optical coherence tomography of an organ or tissue of the body provided by the virtual image module.

4. A virtual image module for generating a depth-based virtual image according to claim 3, wherein the virtual image is marked with locations, guidance, instructions, or navigation for performing surgery.

5. A virtual image module for generating a virtual image having depth according to claim 2, wherein a first point on the real object is selected as the first landmark for superimposing the virtual image onto the real object by superimposing a corresponding first landmark on the virtual image onto the real object.

6. A virtual image module for generating a virtual image having depth according to claim 5, wherein a second point on the real object is selected as the second landmark for superimposing the virtual image onto the real object by superimposing a corresponding second landmark on the virtual image onto the real object.

7. A virtual image module for generating a virtual image having depth according to claim 1, further comprising a control module that processes the right collimation light signal and the corresponding left collimation light signal so that the virtual image is modified to be superimposed on the real object based on the field of view and the position of the real object.

8. A virtual image module for generating a virtual image having depth according to claim 1, further comprising an object measurement module configured to measure the position of a part of the real object and the second depth.