Head-mounted display device and real-time guidance system

The HMD device addresses real-time tracking and rendering challenges by preprocessing image data with parallel processing modules, enhancing accuracy and efficiency in surgical guidance systems.

JP2026505464APending Publication Date: 2026-02-13ARSPECTRA SARL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025546628
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-05-05
Filing Date
2024-02-14
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Existing image-based guidance systems for precision operations, such as surgical procedures, face challenges in accurately tracking and rendering tool poses in real-time due to complex hardware setups and high data processing requirements, leading to inefficiencies and potential errors.

Method used

A head-mounted display (HMD) device with integrated data processing means, including parallel processing modules, preprocesses high-entropy image data to low-entropy data for real-time guidance, reducing latency and maintaining accuracy by leveraging miniaturized hardware components.

Benefits of technology

The HMD device provides low-latency, accurate, and ergonomic image-based guidance by preprocessing image data, enhancing precision and reducing the complexity of hardware systems, thus improving procedural efficiency and reducing the risk of errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026505464000001_ABST
    Figure 2026505464000001_ABST
Patent Text Reader

Abstract

A head-mounted display (HMD) device and an image-based guidance system are disclosed. The HMD has a sensor that images a scene as pixel data and a display that outputs a graphical user interface (GUI). The HMD has one or more parallel processing modules consisting of five pixel-respective data processing threads, each of which receives pixel data from the sensor, filters it to isolate pixel data representing an optical contrast agent applied to at least one target in the scene, and calculates two-dimensional (2D) pixel coordinate data from the isolated pixel data. The data processing means of the HMD converts the 2D pixel coordinate data into three-dimensional (3D) zero coordinate data representing the target or each target relative to the device by referencing a first geometric dataset, generates 3D guidance data by referencing one or more additional geometric datasets, each representing a respective target in the scene, and outputs the 3D guidance data to the GUI.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention is in the field of head-mounted display ("HMD") devices that implement image-based guidance functionality. [Background technology]

[0002] Image-based guidance systems improve user efficiency and accuracy during precision operations. An example operation is inserting a needle-like surgical tool into a specific location and orientation within a patient. Such procedures typically require high positional accuracy, and image-based guidance systems help achieve the necessary degree of precision by detecting, tracking, and rendering both the current and desired tool poses by referencing the operating environment on a display, e.g., a field of view including the surgical site, for comparison and adjustment by the system user.

[0003] The technical challenges are considerable, as the tool must be accurately detected in the image data amidst the scene clutter in the camera field of view, and similarly, the tool must be posed by reference to six mechanical degrees of freedom ("6DoF") of movement in three-dimensional space ("3D") with sub-millimeter accuracy. The calculated poses, respectively the actual and desired tool poses, must be accurately rendered on the display to guide the user during the procedure, all in substantially real time, i.e., with minimal latency between sensor input and display output.

[0004] Many image-based tool guidance systems exist, such as Kalstein®'s Surgical Navigation Systems YR02143™ and Medtronic®'s StealthStation™ S8 Surgical Navigation System. Such solutions track the spatial location and orientation of targets using fiducial markers, such as QR codes, or active markers, such as light-emitting diodes (LEDs), or passive markers, such as reflective dots applied to surgical tools and tool targets. The markers are detected in high-resolution images acquired by an optical system typically placed in fixed locations overlooking the surgical operating room, in a location that is least obtrusive to the surgeon and the surgical site. The three-dimensional (3D) positions of the markers are triangulated from their detection in the images, thereby determining the current 3D position and orientation of the tool and target in the optical system's reference coordinate system. The results are then rendered on a display monitor, and the tool user must continuously refer to the display monitor to adjust the tool's position and orientation toward the desired location. This serial comparison is subjective, error prone, and time inefficient.

[0005] To address and mitigate these shortcomings, alternative image-based tool guidance systems have been proposed by M. A. Lin et al. ("HoloNeedle: Augmented-reality Guidance System for Needle Placement Investigating the Advantages of 3D Needle Shape Reconstruction," IEEE Robotics and Automation Letters, July 2018) and T. Kuzhagaliyev et al. ("Augmented reality needle ablation guidance tool for irreversible electroporation in the pancreas," Proceedings SPIE 10576, Medical Imaging 2018: Image-Guided Procedures, Robotic Interventions and Modeling, March 2018). These systems are based on the use of an augmented reality (AR) head-mounted display (HMD) device to provide the HMD wearer with guidance regarding targets within their line of sight, such as the tool at the surgical site and its destination.

[0006] In such systems, the current and desired tool positions, possibly a pre-acquired scan of the surgical site, and the guidance alignment between the current and desired tool positions are rendered on the HMD's display, advantageously eliminating the need for the wearer to check a display monitor separate from the operating table, thus reducing the risk of error. However, the movement of the HMD AR headset relative to the fixed optical reference frame adds complexity, and the HMD's position and orientation must also be accurately determined and tracked, for example, using a QR code or other marker. In these systems, the AR HMD device is always used as a simple display, with image processing, pose extraction and calculations, and other complex guidance-related image and data processing performed by an external computing resource to which the HMD is tethered. This is because the AR HMD does not have sufficient data processing means to provide the amount of image processing, complex calculations, and rendering required to maintain real-time levels of latency.

[0007] In a more recent alternative proposed by BJ Park et al. ("Augmented reality improves procedural efficiency and reduces radiation dose for CT-guided lesion targeting: a phantom study using HoloLens 2," Sci Rep 10, 18620 (2020); https: / / doi.org / 10.1038 / s41598-020-75676-4), researchers demonstrated that using an AR HMD to provide holographic guidance for needle navigation improved the efficiency of needle-based procedures. However, this approach, based on the Vuforia™ software development kit, still only superimposes the desired 3D tool position onto a surface; the actual needle position is not tracked, and the alignment and accuracy of the current needle position relative to the desired position are not estimated or calculated. Therefore, the HMD wearer must align the physical needle tool in their own or another person's hand with the projected graphic. Improvements to such techniques that might be expected from advances in miniaturized processing components and power optimization are mitigated by the increasing amount and density of image data processed by cameras that tend to be equipped with ever-increasing frame rates and ever-wider fields of view in HMD models.

[0008] Ultimately, proven image-based navigation techniques rely on multiple separate hardware systems for optical acquisition and tracking, data processing, and rendering, resulting in expensive, complex systems with sophisticated data communication and synchronization requirements, all of which define multiple potential points of failure. Recent image-based navigation techniques that attempt to simplify such systems require trade-offs between accuracy, ergonomics, and ease of use in order to meet minimum latency requirements.

[0009] Therefore, there is a need for an HMD device that mitigates at least some of the shortcomings of these image-based guidance techniques of the prior art. [Prior art documents] [Non-patent literature]

[0010] [Non-Patent Document 1] MALin, “HoloNeedle: Augmented-reality Gudance System for Needle Placement Investigating the Advantages of 3D Needle Shape Reconstruction,” IEEE Robotics and Automation Letters, July 2018. [Non-patent document 2] T. Kuzhagaliyev, “Augmented reality needle ablation guidance tool for irreversible electroporation in the pancreas,” Proceedings SPIE 10576, Medical Imaging 2018: Image-Guided Procedures, Robotic Interventions and Modelling, March 2018. [Non-patent document 3] BJPark, “Augmented reality improves procedural efficiency and reduces radiation dose for CT-guided lesion targeting: a phantom study using HoloLens 2”, Sci Rep 10, 18620 (2020); https: / / doi.org / 10.1038 / s41598-020-75676-4 Summary of the Invention

[0011] Aspects of the present invention are set forth in the accompanying claims and are directed to various embodiments of a head-mounted display (HMD) device, various embodiments of a distributed imaging system based on a head-mounted display (HMD) device, and various embodiments of a method for distributing image data within a network using the system, respectively.

[0012] The concept of the present invention is to substantially reduce the amount of image data that a portable display device needs to process in order to provide image-based guidance in real time. Recent technological improvements in hardware acceleration of embedded systems, particularly in small, power-efficient parallel processing modules, are leveraged to preprocess high-entropy image data captured by high-resolution imaging sensors, reducing it to low-entropy image data that preserves semantic content useful for guidance purposes. Such low-entropy image data is input to other processing modules for geometric and rendering calculations, and the inventive technique provides a low-latency image processing technique particularly suited for lightweight head-mounted display ("HMD") devices.

[0013] Accordingly, in a first aspect, the present invention provides a head mounted display (HMD) device comprising imaging means for generating pixel data representative of a scene in use, display means for outputting a graphical user interface in use, power means, data storage means for storing a geometric data set, and data processing means operatively interfaced with the imaging means and display means, the data processing means filtering the pixel data with a predetermined value to separate first pixel data from second pixel data, the first pixel data representing an optical contrast agent applied to at least one target in the scene, and generating a secondary image from the separated first pixel data. and at least one parallel processing module consisting of data processing threads for each of a plurality of pixels adapted to calculate original (2D) pixel coordinate data, the data processing means being further adapted to: transform the calculated 2D pixel coordinate data into three-dimensional (3D) coordinate data representing the target or each target relative to the device by referencing a first geometric dataset representing at least one coordinate system having an origin at the device; generate 3D guidance data by referencing one or more further geometric datasets representing each target in the scene according to the 3D coordinate data; and output the 3D guidance data to said graphical user interface.

[0014] Thus, the present technology provides low-latency image processing techniques that are particularly suitable for lightweight head-mounted display ("HMD") devices. Referring to trade-offs between accuracy, ergonomics, and ease of use to meet minimum latency requirements, the present technology advantageously reduces the data processing overhead associated with increasingly high-resolution, increasingly wide-field-of-view cameras that are desirable for accurate target acquisition in real time by removing captured image data that is redundant for guidance purposes, while preserving semantic image data that is most important for guidance purposes.

[0015] In embodiments of the HMD device, the data processing means may be further adapted to convert the calculated 2D pixel coordinate data into 3D pixel coordinate data by triangulating the 2D pixel coordinate data with reference to the first geometric data set. Alternatively, the data processing means may be further adapted to convert the calculated 2D pixel coordinate data into 3D pixel coordinate data by solving for a 3D rotation and translation based on the 2D pixel coordinate data.

[0016] In an embodiment of the HMD device, the origin of at least one coordinate system located on the device may be selected from one of the aperture of the imaging means, the display unit of the display means, and the eyes of the HMD wearer. In a variation of such an embodiment, the first geometric data set may include a calibrated set of transformations between coordinate systems having origins at the aperture of the imaging means, the display unit, and the eyes of the HMD wearer, respectively.

[0017] Subject to the operational requirements of the embodiments and the technical capabilities of the components available to implement them, the imaging means may be implemented as a single imaging sensor, optionally as a pair of imaging sensors in a stereo configuration, a hybrid combination of high resolution and low resolution imaging sensors, a hybrid combination of imaging sensor and range sensor, e.g. time-of-flight, event-based or echolocation type.

[0018] In an embodiment of an HMD device, the at least one target in the scene may be a tool being used by or in proximity to the HMD wearer, and at least one of the one or more further geometric data sets includes a three-dimensional model representing the tool. Alternatively or additionally, the target may be a marker defining a location in the scene, and at least one of the one or more further geometric data sets includes a three-dimensional model representing the marker. Alternatively, the target may be a biological marker, for example, subcutaneous tissue made fluorescent by injection or ingestion of an optical contrast agent.

[0019] In a variant of such an embodiment particularly adapted for scenes including at least two targets, the data processing means may be further adapted, when generating the 3D guidance display data, to generate display data representing a path between the tool and the marker in the scene.

[0020] Embodiments of the HMD device may be devised for use with passive optical contrast agents and may therefore further comprise a switchable illumination source operably connected to power means for supplying, configured to excite the optical contrast agent in the scene.

[0021] In an embodiment of the HMD device, the data processing means may further comprise a graphics processing unit ("GPU") programmed to generate 3D guidance display data in accordance with the 3D guidance data by reference to one or more further geometric datasets, and the data processing means may further be adapted to output the 3D guidance display data to a graphical user interface.

[0022] An embodiment of the HMD device may be devised to increase the accuracy of the display, wherein the data processing means is further adapted to determine a mismatch between the generated 3D guidance display data and the eyes of the HMD wearer based on the distance measurement and the position of the wearer's eyes, and to adjust the position of the generated guidance display data in the graphical user interface according to the determined mismatch. The distance measurement may be performed based on the stereoscopic image data and / or may be performed using an optional distance sensor of the HMD device.

[0023] An embodiment of the HMD device may be designed to enhance the optical accuracy of the wearer and may therefore further comprise a pair of eye imaging sensors, each generating eye pixel data representative of a respective eye of an HMD wearer in use, and at least a second parallel processing module constructed and operative in accordance with the inventive principles disclosed herein. Accordingly, the second parallel processing may comprise a data processing thread for each of a plurality of eye pixels, adapted to filter the eye pixel data with a predetermined value to separate first eye pixel data from second eye pixel data, the first eye pixel data representing at least a portion of the or each wearer's eye, and to calculate two-dimensional (2D) eye pixel coordinate data from the separated first eye pixel data. Accordingly, the data processing means may be further adapted to triangulate the 2D eye pixel coordinate data received from the or each second data parallel processing module by reference to the first geometric dataset, thereby generating three-dimensional (3D) eye coordinate data representative of the focus of the wearer's eye relative to the device, transform the 3D coordinate data by reference to the 3D eye coordinate data, and generate 3D guidance data according to the transformed 3D coordinate data.

[0024] Variations of such embodiments may be devised to optimize the use of computing resources when generating the display data, and the data processing means may be further adapted to set the 2D eye coordinate data as a gaze point when generating the guidance display data, and the data processing means may be further adapted to output the generated guidance display data to a graphical user interface as foveated display data according to the gaze point.

[0025] In any of the foregoing embodiments, the or each parallel processing module may be selected from the group comprising: a field programmable gate array ("FPGA"), a graphics processing unit ("GPU"), a video processing unit ("VPU"), an application specific integrated circuit ("ASIC"), an image signal processor ("ISP"), a digital signal processor ("DSP"). Alternatively, or complementary, the data processing means may be selected from the group comprising a hybrid programmable parallel central processing unit and a configurable processor.

[0026] In another aspect, the present invention provides a system, an image-based guidance system, comprising at least one detectable target, one or more portions of which are comprised of an optical contrast agent; and a head mounted display (HMD) device substantially as described above, wherein the first pixel data represents the one or more portions of the detectable target, and the data processing means generates three dimensional (3D) coordinate data representing the one or more portions of the detectable target, and one or more further geometric data sets each representing a respective detectable target in the scene.

[0027] In system embodiments, the optical contrast agent may be an active agent that emits light waves, such as a light-emitting diode (LED). Alternatively, the optical contrast agent may be a passive agent, such as a fluorophore compound, and system embodiments may further include an illumination source configured to excite the optical contrast agent during use. In variations of such embodiments, the HMD device may include the illumination source to align the excited agent with the field of view of the HMD device's imaging sensor and minimize the number of hardware units in the system.

[0028] In system embodiments, each of the one or more portions of the detectable target may be a marker having a predetermined relative geometric relationship thereto. The or each detectable target may be, for example, a tool being used by or in proximity to the HMD wearer, such as a needle, biopsy syringe, or other surgical device, whether handled by the HMD wearer or by another user or robotic device adjacent to the HMD wearer, and at least one of the one or more additional geometric data sets includes a three-dimensional model representing the tool.

[0029] Alternatively or additionally, the or each detectable target may be, for example, a marker defining a location within the scene, e.g., a fiducial marker indicating the target destination of a surgical device on a patient's body, and at least one of the one or more additional geometric data sets may include a three-dimensional model representing the marker. One embodiment of such a marker may be a matrix barcode, also known as a quick response (“QR”) code or ArUco marker, one or more portions of which are composed of a passive or active optical contrast agent, or a combination of both passive and active optical contrast agents. Another embodiment of a marker defining a location within the scene may be biological tissue made fluorescent by an optical contrast agent after injection or ingestion, having a predetermined relative geometric relationship to one or more properties of the surrounding tissue, such as the distance of the fluorescent tissue relative to the surface of the tissue.

[0030] In a variation of such an embodiment particularly adapted to a scene including at least two detectable targets, the data processing means of the HMD, when generating the 3D guidance display data, may further be programmed to generate display data representing a path between two detectable targets in the scene.

[0031] In a further aspect, the present invention provides a method for guiding a detectable target using a head-mounted display (HMD) device, the method comprising: generating pixel data of a scene using an imaging sensor of the HMD device, wherein the detectable target is within the scene; filtering, using at least one parallel processing module of the HMD device configured with data processing threads for each of a plurality of pixels, the pixel data with a predetermined value to separate first pixel data from second pixel data, the first pixel data representing one or more portions of the detectable target comprised of an optical contrast agent; calculating two-dimensional (2D) pixel coordinate data from the separated first pixel data; converting, using at least one further processing unit of the HMD device, the 2D pixel coordinate data into three-dimensional (3D) coordinate data representing one or more portions of the detectable target by reference to a first geometric dataset representing at least one coordinate system having an origin in the HMD device; and generating 3D guidance display data according to the 3D coordinate data by reference to one or more further geometric datasets, each representing a respective detectable target in the scene; and outputting the 3D guidance display data to a graphical user interface on at least one display of the HMD device.

[0032] Other aspects of the invention are set out in the accompanying claims. [Brief explanation of the drawings]

[0033] The invention will be more clearly understood from the following description of embodiments thereof, given by way of example only, with reference to the accompanying drawings, in which: [Figure 1] FIG. 1 is a front view of one embodiment of a head-mounted display (HMD) device according to the present invention, including an imaging sensor. [Figure 2] 2 is a front view of an alternative embodiment of an HMD device according to the present invention, which configures the HMD device of FIG. 1 with an illumination source. [Figure 3]2 illustrates an exemplary hardware architecture of the HMD shown in FIG. 1, including an imaging sensor, data processing means, and memory. [Figure 4] 3 illustrates an exemplary hardware architecture of the HMD shown in FIG. 2. [Figure 5] 5 illustrates one embodiment of a detectable target viewable by the imaging sensor of FIGS. 1-4. [Figure 6] 1-4 depict further embodiments of detectable targets observable by the imaging sensor of FIGS. [Figure 7] 5 and / or 6 show in detail the data processing steps performed by the HMD device of FIGS. 1-4 to generate 3D guidance data for the target of FIG. 5, including steps of filtering pixel data and transforming 2D coordinate data. [Figure 8] 3 or 4 at run time when the steps of FIG. 7 are performed. [Figure 9] The data processing and associated data type flow are shown in Figures 7 and 8. [Figure 10] The steps of filtering pixel data in FIGS. 7 and 9 performed by the parallel processing units of FIGS. 3 and 4 are shown in more detail. [Figure 11] The steps for generating the 3D guidance data in Figures 7 and 9, which are performed by the further processing units in Figures 3 and 4, are shown in more detail. [Figure 12] An alternative embodiment of the steps of filtering pixel data of FIGS. 7 and 10 is shown in more detail. [Figure 13] 1 provides a front view of an alternative embodiment of an HMD device with a time-of-flight sensor. Detailed Description of the Drawings

[0034] Examples of specific embodiments contemplated by the inventors will now be described. In the following description and the accompanying drawings, numerous specific details are set forth to provide a thorough understanding, and like reference numerals refer to like features. It will be readily apparent to those skilled in the art that the present invention can be practiced without being limited to these specific details. In other instances, well-known methods and structures have not been described in detail to avoid unnecessarily obscuring the description.

[0035] 1, there is shown a first embodiment 10A of an HMD device, in this example an augmented reality ("AR") device, comprising imaging and display means according to the present invention. The AR HMD 10A comprises a wearer visor 20 including a main see-through portion 22 and video display portions 24A, 24B for each eye, positioned equidistantly in a central bridge portion that covers the wearer's nose in use. Perceptually, the display portions 24A, 24B implement a single video display that occupies a subset of the front of the HMD, allowing the wearer to observe both the surrounding physical environment in front of the HMD 10 and video content superimposed thereon.

[0036] Each video display portion 24A, 24B comprises a respective video display unit 26A, 26B, in this example a micro OLED panel having a minimum 60 Hz frame refresh rate and a resolution of 1920 x 1080 pixels, and the video display units 26A, 26B are positioned adjacent to the lower edge of the visor such that the see-through portion 22 extends upward to the upper edge of the visor to avoid visual obstruction when the VDU is displayed. The HMD further comprises first and second high-resolution imaging sensors 30A, 30B arranged in a stereoscopic configuration, each capturing visible light in the wavelength range of 400 nm to 700 nm over its field of view (FoV), typically 70 to 160 degrees or more, and outputting the captured images as a stream of pixel data at a resolution of at least 1920 x 1080 pixels and at a rate of 60 frames per second or more.

[0037] Referring to FIG. 2, in which like numerals refer to like features relative to FIG. 1, a second embodiment 10B of a head-mounted display ("HMD") device comprising imaging and display means in accordance with the present invention is shown, which further comprises an illumination source 32, such as a light-emitting diode ("LED") 32 emitting light in the wavelength range 800 nm to 2,500 nm, corresponding to near-infrared ("NIR") light, for excitation of targets or objects comprised of passive optical contrast agents, e.g., fluorophores, within the FoV of each of the HMD imaging sensors 30A-30B.

[0038] Those skilled in the art will also understand that the technical principles disclosed herein may be implemented in other HMD types, such as augmented reality monocular or contact lens devices, virtual reality ("VR") or mixed reality ("MR") closed-display devices, where the video display portion 24A, 24B for each eye perceptually implements a single video display portion that occupies substantially the entire inner front surface of the HMD. For such HMDs, each video display portion 24A, 24B may be comprised of an RGB low-persistence panel 26A, 26B with a minimum 60 Hz frame refresh rate and an individual resolution of 2048 x 1080 pixels per eye, and is perceived as a single video display with a resolution of 4096 x 2160 pixels.

[0039] Embodiments of the HMD according to the present invention may include fewer or additional imaging sensors as imaging means. The techniques of the present invention may be implemented using a single imaging sensor 30A and a target of known geometric shape including at least four detectable points, or using the two imaging sensors 30A-30B described above with at least three detectable points. Embodiments of the HMD according to the present invention may also, or instead, include other types of sensors, such as distance or depth sensors implementing time-of-flight technology, which are particularly useful for preventing display artifacts according to the principles described below. Next, exemplary hardware architectures of HMD devices 10A-10B according to the present invention will be described in further detail with reference to FIGS. 3 and 4, respectively, where like reference numerals refer to like features by way of non-limiting example.

[0040] All embodiments of the HMD according to the present invention comprise data processing means. Thus, in addition to the video display units 26A-26B and the imaging sensors 30A-30B, and optionally the NIR LEDs 32, the HMD according to the present invention includes data processing means consisting of at least one data processing unit 301 that functions as the main controller of the HMD and at least one parallel processing module 302 that preprocesses the pixel data generated by the imaging sensors 30A-30B. The CPU 301 is a general-purpose microprocessor, for example according to the Cortex™ architecture manufactured by ARM™, and the parallel processing module 302 is a field programmable gate array ("FPGA") semiconductor device, for example according to the Artix™ architecture manufactured by AMD™ or Xylinx™.

[0041] CPU 301 may further include or be associated with a dedicated graphical processing unit ("GPU") 321 that receives data and processing commands from CPU 301 to generate display data before the display data is output to displays 26A-26B. CPU 301 and FPGA 302 are coupled to memory means 303, which may include volatile random access memory (RAM), non-volatile random access memory (NVRAM), or a combination thereof, by a data input / output bus 304 through which they communicate and through which other components of HMDs 10A-10B are similarly connected to data input / output bus 304 to provide headset functionality and receive user commands.

[0042] The data connection between the imaging sensors 30A-30B, the CPU 301, and the FPGA 302 via a bus 304 or otherwise is a high frequency data communication interface, with at least the FPGA 302 (and preferably also the CPU 301) being located closest to the imaging sensors 30A-30B interconnections, i.e., in front of or within a portion of the HMD device, to minimize latency in data flow.

[0043] User input data may be received directly from one or more buttons, including at least an on / off switch, and / or a physical input interface 305, which may be part of the HMD casing configured for tactile interaction with the wearer's touch. User input data may also be received indirectly, such as gestures optically captured by optical sensors 30A-30B and / or spoken words captured as analog sound wave data by microphone 306, for which DSP module 307 implements analog-to-digital conversion functions, and CPU 301 then interprets both according to principles beyond the scope of this disclosure. Processed audio data is output to speaker unit 308, and all components are powered by electrical circuitry 309 interfaced with internal battery module 310, which is periodically recharged by electrical converter 311.

[0044] Embodiments of an HMD in accordance with the present invention may further include data connectivity capacity, shown in dotted lines as a wireless network interface card or module (WNIC) 322, also connected to data input / output bus 304 and electrical circuitry 308, suitable for interfacing the HMD with a wireless local area network ("WLAN") created by a local wireless router. Alternative or additional wireless data communication functionality may be provided by the same or another module, for example, implementing short-range data communication via Bluetooth™ and / or Near Field Communication (NFC) interoperability and data communication protocols.

[0045] The HMD devices of the present invention are used to guide an article relative to other articles and / or article destinations in a scene, for example, to guide a surgical device during a surgical procedure. The articles and article destinations must therefore be detectable targets, and embodiments of the present invention rely on such targets being composed of passive or active optical contrast agents so as to be captured as imaging data by the imaging sensors 30A-30B in use.

[0046] 5, a first embodiment of a detectable target is shown, in this example a surgical tool 50 having an elongated body 51 terminated at a first end by a needle 52 to facilitate subcutaneous insertion, and a user grip portion 54 distal to the needle and proximate a second, opposing end of the tool, the grip portion being terminated by a passive contrast coated geometric reference dot or indicator 55. Those skilled in the art will appreciate that the passive geometric indicator 55 can be used in place of an active indicator, such as an LED, that is equally compatible with the technology described herein.

[0047] The three-dimensional pose of any target detectable in the scene facing the HMDs 10A-10B needs to be determined, and therefore the exemplary surgical tool 50 includes at least two further geometric reference indicators 55, each attached to the grip portion 54 by a handle-like member 56, each oriented orthogonal to the other and to the major axis of the elongated body 51. Suitably, the three geometric reference indicators 55 collectively define a three-dimensional coordinate system N having its origin at the target 50, its axis coaxial with the major axis of the tool. The longitudinal dimension of the surgical tool 50 between its opposing ends 52, 55 is known, as are the respective dimensions between each further geometric reference indicator 55 and the grip portion surface, such that the three-dimensional geometry 58 of the tool is known and thus preset.

[0048] Referring to FIG. 6 , two additional embodiments of detectable targets that can be used as fiducial markers to indicate the destination of an item within a scene are shown. In the first example, surgical marker 60A has a square-shaped planar body 62 with a plurality of positioning leg members 64 extending from its underside, each of which may optionally terminate in a needle distal to body 62 for ease of subcutaneous insertion, or in a clamp, grip, or some other means of attachment to the patient. Geometric reference indicators 55, as previously described, are fixed to each corner of planar body 62, with the four geometric reference indicators 55 corresponding to and thus defining the major plane and orientation of marker 60A at any given time. Suitably, any three of the four geometric reference indicators 55 collectively define a three-dimensional coordinate system G having its origin at the geometric center of the major plane of target 60A, with two orthogonal axes coplanar with the major plane of the fiducial marker and a third axis orthogonal thereto.

[0049] In a second example, for ease of explanation, another surgical marker 60B has substantially the same configuration as before. However, rather than affixing physical indicators 55 to some or each corner of its planar body 62, in this embodiment, the upper major planar surface of the body is configured as a matrix barcode patterned according to ArUco™ technology, with several geometric portions 65, each coated with a passive optical contrast agent. While the white portions are coated with the example ArUco™ marker 60B, those skilled in the art will readily understand that the black portions may instead be coated with a passive optical contrast agent, and that the passive optical contrast agent may similarly substitute for one or more active indicators, such as LEDs. ArUco™ markers are known to encode more geometric and semantic information than classic matrix (QR) barcodes and dot-like passive or active indicators 55, and this increased precision can be exploited by the present technology with its added optical contrast properties.

[0050] Thus, again in this example, the geometric reference indicator 65, comprised of an optical contrast pattern, defines at least the major plane of the marker 60B and its orientation at any given time, and may encode further information, such as dimensional data of the marker. Suitably, the geometric reference indicator 65 in this example defines the same three-dimensional coordinate system G having its origin at the geometric center of the major plane of the target 60B, with two orthogonal axes coplanar with the major plane of the marker and a third axis orthogonal to it.

[0051] The lateral dimensions between the two corners of the main plane of the surgical markers 60A, 60B are known, and similarly the respective positions and dimensions of each leg 64 extending below the main plane are also known and / or may be encoded in a detectable pattern 65 thereon and decoded by an associated configuration of the HMD, whereby the three-dimensional shape 68 of the marker is known and therefore preset.

[0052] The basic and improved data processing configurations and functionality of HMDs 10A-10B of Figures 1-4 will now be described in accordance with one embodiment of the present invention with reference to Figures 7-11. The data structures stored in memory 303 and processed by FPGA 302 and CPU 301 are shown in Figure 8, with like numbers referring to like features, steps, and structures throughout.

[0053] The operating system 801 is initially loaded in step 701 upon initial powering on the HMDs 10A-10B to manage basic data processing, interdependencies, and interoperability of the HMD components 26A-26B, 30A-30B, 32 (if present), and 301-321 (and further including the WNIC 322, if present). The HMD OS may be based on Android™, distributed by Google™, Mountain View, California, USA. The OS includes device drivers for the HMD components, input subroutines for reading and processing input data, including direct user input to the physical interface device 305, and output subroutines for outputting display data to the displays 26A-26B. In particular, the OS 801 interfaces the output of the image sensors 30A-30B with the FPGA 302 within a low computational layer, illustratively at the kernel level for minimal latency. In embodiments of the HMDs 10A-10B that include a networking means 322, the OS 601 further includes a communications subroutine 802 for configuring the HMDs 10A-10B for bilateral network communications with remote terminals via the WNIC 322 that interfaces with a network router device.

[0054] In step 702, a set of instructions embodying a visualization application 803 is loaded as a subroutine of the OS 801 or as a separate application on a higher computational tier. The visualization application 803 interfaces with the FPGA 302 through the OS 801 via one or more application programming interfaces (APIs) 804. The visualization application 803 comprises and coordinates data processing subroutines that embody the various functions described herein, including updating and outputting the user interface 805 to the displays 26A, 26B in real time.

[0055] In addition to initializing the imaging sensors 30A-30B in step 703 to generate respective pixel data streams 806 including one or more targets 50, 60A-60B and their respective geometric reference indicators 55, 65, the user interface 805 itself is initialized in step 704, whereby the HMDs 10A-10B are configured to begin processing image data for display navigation in real time.

[0056] In step 705, FPGA 302 receives pixel data stream 806 and filters each input pixel according to a predetermined value, for example, a pixel brightness or brightness threshold corresponding to the captured optical contrast agent of geometric reference indicator 55 in an excited state, or a pixel position offset relative to its position in the previous capture.

[0057] With specific reference to FIG. 10, each pixel 900 in the pixel data stream 806 N 9, each of which represents a respective data processing thread or pipeline 920 implemented within the massively parallel architecture of the FPGA 302. N Each input block 910 N Each parallel pipeline 920 1-NThe image processing unit 30A receives the pixel data in step 811, performs one or more standard artifact removal operations, such as rectification and lens-related distortion removal, in step 812 to obtain corrected accurate pixel data. Then, in step 813, the image processing unit 30A is configured to separate the accurate pixel data according to a predetermined value between pixel data that meets or exceeds the predetermined value, and thus corresponds to a geometric reference indicator 55 in the pixel data stream, and pixel data below the predetermined value, which corresponds to any other aspect of the captured scene and is not of further interest and is therefore discarded. As shown in FIG. 9, the output of step 803 is very low-entropy binary image data that encodes only data representing each geometric reference indicator 55 within the field of view of the imaging sensors 30A-30B, and the respective corrected two-dimensional (2D) screen coordinates within the frame are calculated for extraction in step 814.

[0058] Thus, the output of FPGA 302 in step 705 is low-entropy data, as shown at 807, describing each pixel representing geometric reference indicator 55 and its corrected 2D position within the field of view of imaging sensors 30A-30B at that precise moment, the data being significantly smaller than full-resolution or low-resolution RGB or grayscale image data typically used in known image-based navigation systems.

[0059] In step 706, the CPU 301 receives the 2D pixel coordinate data 807 from the FPGA 302 and converts it to three-dimensional (3D) pixel coordinate data 809. This conversion may be performed using a variety of techniques, including, but not limited to, depending on whether the HMD 10 includes a single image sensor 30A or a pair of image sensors 30A-30B in a stereo configuration.

[0060] In the case of stereoscopic HMDs 10A, 10B, the transformation can be performed by triangulation techniques, in which the CPU 301 triangulates the 2D pixel coordinate data 807 from the FPGA 302 by reference to a first geometric data set 808 representing at least one coordinate system having an origin in the HMD 10A, 10B, thereby generating 3D coordinate data 809 representing the or each target 50, 60A, 60B relative to the HMD 10A, 10B. Here, the verb "triangulation" and equivalent adjectives and expressions should be understood in their ordinary sense in the field of computer vision, i.e., as the process of determining a point in 3D space given its projection onto a 2D image plane. Accordingly, those skilled in the art will understand that different techniques can be used to implement this particular process, such as a direct linear transformation, or more computationally efficient alternatives, depending on the quality and accuracy of the 2D data set 807 in HMD embodiments with high-resolution and high-frame-rate imaging sensors.

[0061] 11, the characteristics of HMDs 10A-10B and the intrinsic parameters of their components related to the optical and vision systems embodied therein are known from manufacturing and pre-configuration. Thus, the reference geometric data set 808 comprises one or more coordinate systems, each with its origin at a respective component of the HMD, which have also been pre-calibrated with six DoF transformations between them, and a 3×3 rotation matrix JPEG2026505464000002.jpg89 and 3x1 translation vector In this example, a three-dimensional coordinate system H has its origin at the image sensor 30A, a three-dimensional coordinate system S has its origin at the display 26A, and a three-dimensional coordinate system E has its origin at the right eye or left eye 1100 of the HMD wearer, and these are included in the first geometric data set 808.

[0062] Thus, when the triangulation of step 706 is performed in step 821 by referencing the first geometric data set 808, a six DoF transformation is performed between H and E, i.e. JPEG2026505464000004.jpg1124 Known between E and S, i.e. JPEG2026505464000005.jpg1123 is pre-calculated accordingly. Similarly, the transformation between the left image sensor 30A and the right image sensor 30B, thus The baseline distance between them is also known. The assumption that the intrinsic parameters of the pair of image sensors 30A and 30B are identical is JPEG2026505464000007.jpg2548, where: JPEG2026505464000008.jpg1018 corresponds to the focal length of the image sensor in the x and y dimensions, and similarly corresponds to the center of projection of the image sensor in the x and y dimensions, respectively. Supports JPEG2026505464000009.jpg1118.

[0063] A pair of image sensors 30A and 30B respectively JPEG2026505464000010.jpg1118 Given captured stereo image data, a set of JPEG2026505464000011.jpg1118 Geometric indicators 55 are detected and their corresponding 2D positions JPEG2026505464000012.jpg1130 and JPEG2026505464000013.jpg1130807 is extracted in step 705 and received as input in step 706. The position corresponding to the geometric indicator 55 JPEG2026505464000014.jpg1134 and each pixel 3D position corresponding to JPEG2026505464000015.jpg1134 JPEG2026505464000016.jpg1150, JPEG2026505464000017.jpg1134 is is given by JPEG2026505464000018.jpg20135.

[0064] The 3D position is then Convert H to E using JPEG2026505464000019.jpg1244. Next, JPEG2026505464000020.jpg1244 is the current eye position relative to S, i.e., JPEG2026505464000021.jpg1244 is taken into consideration and projected correctly onto S. This allows the eye 1100 and the HMD displays 30A-30B to be Unique matrix defined as JPEG2026505464000022.jpg1111 This is done by modeling it as an off-axis pinhole imaging sensor with JPEG2026505464000023.jpg2692, where: JPEG2026505464000024.jpg918 is the scaling factor that converts 3D points on the screen to pixel points.

[0065] Therefore, on S as three-dimensional (3D) coordinate data 809 The rendering position of JPEG2026505464000025.jpg1013 is It is determined as JPEG2026505464000026.jpg1998.

[0066] The above assumptions are based on the camera and the eyepiece. JPEG2026505464000027.jpg1022 and between the display and the eye The transformation between JPEG2026505464000028.jpg1022 is known and pre-calculated. In certain embodiments, these transformations can be calculated using an additional eye-tracking device (e.g., F), and the pre-calculated or pre-calibrated transformations are used to calculate the image quality between the camera and the tracking device. JPEG2026505464000029.jpg1022 and between the camera and the display device JPEG2026505464000030.jpg1022. At runtime, the tracking device estimates the eye position and outputs eye coordinate data, which is then used to communicate between the tracking device and the eye. It is used to calculate and update the transformation between JPEG2026505464000031.jpg1022, JPEG2026505464000032.jpg1022 is JPEG2026505464000033.jpg1022, JPEG2026505464000034.jpg1022 and JPEG2026505464000035.jpg1022 is calculated via JPEG2026505464000036.jpg1022 is JPEG2026505464000037.jpg1022 and Calculated via JPEG2026505464000038.jpg1022.

[0067] Alternatively, in the case of an HMD with a single imaging sensor 30A, the transformation may be performed by solver techniques by solving rotations and translations based on 2D pixel coordinate data, i.e., by aligning the 2D image data captured thereby to a 3D model describing the real world, based on a pose calculation problem such as the Perspective n-Point Problem ("PnP"), which aims to recover the position and orientation of an object.

[0068] This technique calculates the pose of the targets 50, 60, including the rotation matrix R and translation vector t between the world frame in which the targets 50, 60 are located and the frame of the imaging sensor 30A, from N feature points corresponding to 2D pixel coordinate data 807 as input, where N is greater than or equal to 3. The solution output by the solver minimizes the reprojection error between the 3D points of the targets and the input 2D point data, corresponding to 3D coordinate data 809 representing each target 50, 60A-60B relative to the HMDs 10A-10B.

[0069] Once the conversion in step 706 is complete, in step 707, the CPU 301 generates 3D guidance data according to the 3D coordinate data 809 by referencing one or more further geometric data sets, each representing a respective target 50, 60A-60B in the scene, examples of which are the pre-set geometric shapes 58, 68 shown in Figures 5 and 6.

[0070] Thus, in step 821, the CPU 301 calculates or "matches" the pose of the target 50, 60A-60B or each target 50, 60A-60B, provided that a quorum of at least three geometric indicators 55 is met for each detected target. Many techniques are known for implementing this step, the purpose of which is to compare the 3D coordinate data 809 of step 821 with predefined geometric shapes 58, 68 of targets detected in the scene for a geometric match, and, upon matching one or more geometric shapes, determine the orientation of the or each matched geometric shape in 6 DoF or pose at that precise moment. One example technique is Umeyama's least squares estimation technique (IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol. 13, Issue 4, April 1991). The CPU 301 may further calculate semantic, fiducial, or similar other meaningful indicator data that may be derived from the geometric information specific to the extracted poses and displayed on the user interface 805, for example, a guide line 900 extending between the detected target 50 and the detected markers 60A-60B, which have dimensions calculated by referencing known pre-set geometric shapes 58, 68, and which have 6DoF poses calculated by referencing their respective extracted poses.

[0071] In the next step 822, the CPU 301 calculates the transformation of the extracted pose of the first matched geometric shape 58, 68 relative to the origin of the field of view corresponding to the viewpoint of the user interface 805, i.e., relative to the "virtual camera" according to which the 3D object is rendered within the user interface 805. Next, in step 823, a question is asked as to whether poses of further matched geometric shapes 58, 68 and / or semantic or reference indicators still need to be extracted and transformed, and control returns to step 821 in the affirmative. Alternatively, or finally, once the pose and optional semantic or fiducial indicators of each matched geometric shape 58, 68 have been extracted and transformed, the computed transformation accordingly comprises 3D guidance data 810, i.e. data defining the pose of the or each target 50, 60A-60B detected within the field of view of the imaging sensors 30A-30B at that precise moment, and optionally data defining any additional semantic or fiducial indicators, and is rendered on the user interface 805 as quickly as the processing latency of the relevant HMD components allows.

[0072] In a next step 708, the 3D guidance data 810 is submitted to a rendering subroutine of the visualization application 804, which is processed by either the CPU 301 or the GPU 321, if present in the HMD architecture, advantageously freeing the CPU 301 from corresponding graphics processing overhead. 3D guidance display data 820 is generated from the 3D guidance data 810, for example, by mapping respective target geometry data sets 58, 68 to respective target pose data 810 and mapping semantic or reference graphic data 900 to semantic or reference indicator pose data 810. The rendering subroutine may further transform the 3D guidance data 810 according to any intervening updates to world coordinate space, since the coordinates of at least the imaging sensor 30A, and optionally other anchors such as the eyes 1100, may change significantly from the current display frame to the next, in order to maintain the relative positions of their respective physical locations.

[0073] The rendered three-dimensional guidance display data 820 is suitably output to the user interface 805 in step 709 as quickly as the processing latency of the involved HMD components allows, so that the HMD wearer is presented with a representation 58 of the surgical tool 50, a representation 68 of the fiducial marker 60A or 60B, and a fiducial indicator 900 extending between both representations, in their precise respective positions and orientations on both displays 26A-26B, and thus superimposed on the see-through field of view of the visor 20. In step 710, a question is asked as to whether operation of the HMD should be suspended, which is answered in the negative so long as the HMD remains in use, and the logic returns to the pixel isolation of step 705 until the HMD is switched off.

[0074] 12A, an alternative embodiment of the logic of the parallel processing pipeline 920 of the FPGA 302 can be modified to query 1210 whether tracking of the target 50, 60A-60B in the scene has been interrupted after the isolation of step 803. For example, this query may be embodied in a simple buffer of set pixel isolation states, i.e., the isolation state of the same pixel is compared between the previous and current frames, and a change in the state of pixel isolation between two successive frames indicates an interruption of tracking.

[0075] If the query at step 1210 is answered negatively, the logic maintains target tracking in pixel data stream 806 in step 1220. Alternatively, if the query at step 1210 is answered affirmatively, the logic reacquires the target in pixel data stream 806 in step 1230. In either case, the logic proceeds to extract 2D coordinate data in step 804.

[0076] 12B, an alternative embodiment of the HMD device 10C may be configured with a single imaging sensor 30A and a time-of-flight (ToF) sensor 1200, which facilitates detection and tracking of detectable targets 50, 60A-60B by determining their respective distances to the HMD device 10C. The single imaging sensor 30A provides a pixel stream 806 to the FPGA 302 for low-entropy filtering per step 705, as previously described. In this embodiment, the characteristics of the HMD 10C and the intrinsic parameters of its components are still known from manufacturing and pre-configuration, and therefore the reference geometric data set 808 includes an additional coordinate system D having its origin at the ToF sensor 1200. The triangulation of step 706 is again performed by reference to the first geometric data set 808 in step 706, so that the six DoF transformations are performed between H and D, i.e., JPEG2026505464000039.jpg1012, between D and E, i.e. JPEG2026505464000040.jpg1012, between E and S, i.e. JPEG2026505464000041.jpg1012 is known as the The 3D position is converted from H to E via JPEG2026505464000042.jpg1039. The range data for each target provided by ToF sensor 1200 is used by CPU 301 to estimate the position of each detected target relative to the HMD, and the generation of the three-dimensional guidance data in step 707 is facilitated by this position data. Thus, ToF sensor 1200 maintains the accuracy of the techniques disclosed herein using a single imaging sensor, while advantageously further reducing the overall amount of image data processed by the HMD architecture.

[0077] The CPU 301 of the HMD 10C may be further adapted to determine a mismatch between the generated 3D guidance display data 820 and the HMD wearer's eyes 1100 based on distance measurements and the wearer's eye positions, as taught by the applicant in GB 2588774, and to adjust the position of the generated 3D guidance display data 820 within the graphical user interface 805 according to the determined mismatch. For example, the distance measurement may be performed based on stereo image data captured by the image sensors 30A-30B of the stereoscopic HMDs 10A, 10B, and / or may be performed or augmented by the ToF sensor 1200 of the HMD 10C.

[0078] Yet another embodiment of the hardware-accelerated pixel data separation technique disclosed herein for AR navigation purposes is contemplated. The imaging sensors 30A-30B can be configured to perform optical filtering of the scene, filtering the geometric indicators 55 upon capture with reference to active or passive optical contrast agents applied to the targets 50, 60A-B, and providing the already separated data for performing steps 812 and 814 to the parallel processing module 302, thereby significantly accelerating the technique. If the filtering of step 803 is based on a thresholding technique, it can be performed with a different value-based comparator. In a first alternative, the luminance or luminous intensity values ​​in the pixel data 806 can be rounded or truncated with reference to a threshold or range. In a second alternative, the pixel data 806 can be selected within each pipeline 920 with reference to a threshold or range stored in the parallel module 302 or memory 303.

[0079] In either of these alternatives, and in the main embodiment described herein, separation may be further accelerated by spatial selectivity, whereby only a subset of pixel data 806 defined by reference to the resolution of imaging sensors 30A-30B, e.g., an area or diameter having a center pixel as its origin and expressed as a pixel count, or some other predetermined location within the captured image data, is input into parallel pipeline 920.

[0080] Surgical tools and markers are known that include active or passive magnetic transponder units or modules to aid in determining their location within the operating room. When such devices are configured with optical contrast agents, the triangulation of the two-dimensional pixel coordinate data in step 706 and the generation of the three-dimensional guidance data in step 707 can be facilitated by using position data for each target obtained from the HMD's magnetic sensors or from an external device or system in wireless data communication therewith via the WNIC 322, in substantially the same manner as using position data for each target determined from distances measured by the HMD's time-of-flight sensor 1200.

[0081] Those skilled in the art will appreciate that the hardware-accelerated pixel data separation techniques disclosed herein for AR navigation purposes, while described with reference to surgical applications as a non-limiting example, can be adapted to many other applications where accurate, low-latency visual guidance is desirable with maximum portability and minimal burden on the HMD wearer. Furthermore, those skilled in the art will also appreciate that the pixel data separation techniques disclosed herein can be adapted to other types of HMDs, such as video see-through virtual reality (VR) and, in particular, mixed reality (MR) HMDs, because the imaging sensors, matched field of view, and line-of-sight visualization are substantially similar for MR HMDs, including the ability to image the field of view in front of them as an alternative to the clear visor 20.

[0082] In this specification, the terms "comprise, comprises, comprised, and comprising" or any variation thereof and the terms "include, includes, included, and including" or any variation thereof are considered to be fully interchangeable and should all be given the broadest possible interpretation, and vice versa. The present invention is not limited to the embodiments described above, which may be modified both in arrangement and detail.

Claims

1. A head-mounted display (HMD) device, comprising: imaging means for generating pixel data representative of a scene in use; a display means for outputting a graphical user interface in use; a power means; a data storage means for storing a geometric data set; and a data processing means operatively interfaced with said imaging means and said display means, said data processing means comprising: filtering the pixel data with a predetermined value to separate first pixel data from second pixel data, the first pixel data representing an optical contrast agent applied to at least one target within the scene; and at least one parallel processing module configured with a data processing thread for each of a plurality of pixels, the data processing thread being adapted to calculate two-dimensional (2D) pixel coordinate data from the separated first pixel data; said data processing means converting said calculated 2D pixel coordinate data into three-dimensional (3D) coordinate data representing the or each target relative to said device by reference to a first geometric data set representing at least one coordinate system having an origin at said device; generating 3D guidance data according to the 3D coordinate data by reference to one or more further geometric data sets each representing a respective target in the scene; A head-mounted display (HMD) device further adapted to output the 3D guidance data to the graphical user interface.

2. 2. The head mounted display device of claim 1, wherein the data processing means is further adapted to transform the calculated 2D pixel coordinate data by triangulating the 2D pixel coordinate data with reference to the first geometric data set.

3. 3. The head-mounted display device according to claim 1, wherein the origin of the at least one coordinate system located on the device is selected from one of an aperture of the imaging means, a display unit of the display means, and an eye of the HMD wearer.

4. 4. The head-mounted display device of claim 3, wherein the first geometric data set includes a calibrated set of transformations between coordinate systems having origins at the aperture of the imaging means, the display unit, and the eyes of the HMD wearer, respectively.

5. 2. The head mounted display device of claim 1, wherein the data processing means is further adapted to transform the calculated 2D pixel coordinate data by solving for rotations and translations based on the 2D pixel coordinate data.

6. at least one target in the scene is a tool being used by or in proximity to the HMD wearer, and at least one of the one or more further geometric data sets includes a three-dimensional model representing the tool; and / or at least one target in the scene is a marker defining a position in the scene, and at least one of the one or more further geometric data sets includes a three-dimensional model representing the marker; 6. The head-mounted display device according to claim 1.

7. 7. The head-mounted display device of claim 6, wherein the scene includes at least two targets, and the data processing means, when generating the 3D guidance display data, is further programmed to generate display data representing a path between the two targets in the scene.

8. 8. A head mounted display device as described in any preceding claim, further comprising a switchable illumination source operatively connected to the power means for supply, configured to excite the optical contrast agent in the scene.

9. the data processing means further comprises a graphics processing unit ("GPU") programmed to generate 3D guidance display data in accordance with the 3D guidance data by reference to the one or more further geometric data sets; 9. A head mounted display device according to claim 1, wherein the data processing means is further adapted to output the 3D guidance display data to the graphical user interface.

10. The data processing means determining a mismatch between the generated 3D guidance display data and the eyes of the HMD user based on distance measurements and the position of the wearer's eyes; further adapted to adjust a position of the generated guidance display data within the graphical user interface according to the determined inconsistency; and optionally A head mounted display device according to claim 1 , wherein the distance measurement is performed based on stereo image data or by using an optional distance sensor of the HMD device.

11. The imaging means further generates eye pixel data representing each eye of the HMD wearer in use, and the device HMD comprises: filtering the eye pixel data with a predetermined value to separate first eye pixel data from second eye pixel data, the first eye pixel data representing at least a portion of an eye of the or each wearer; and at least a second parallel processing module configured with a data processing thread for each of a plurality of eye pixels adapted to calculate two-dimensional (2D) eye pixel coordinate data from the separated first eye pixel data; The data processing means converting the 2D eye pixel coordinate data received from the or each second data parallel processing module into three-dimensional (3D) eye coordinate data representative of a focal point of the wearer's eye relative to the device; transforming the 3D coordinate data by referencing the 3D eye coordinate data; A head-mounted display device according to claim 1 , further adapted to generate the 3D guidance data according to transformed 3D coordinate data.

12. The data processing means When generating the guidance display data, the 2D eye coordinate data is set as a gaze point; The head-mounted display device of claim 11 , further adapted to output the generated guidance display data to the graphical user interface as foveated display data according to the gaze point.

13. the parallel processing module is selected from the group comprising a field programmable gate array ("FPGA"), a graphics processing unit ("GPU"), a video processing unit ("VPU"), an application specific integrated circuit ("ASIC"), an image signal processor ("ISP"), a digital signal processor ("DSP"), or 13. A head mounted display device according to any one of claims 1 to 12, wherein the data processing means is selected from the group comprising hybrid programmable parallel central processing units and configurable processors.

14. 1. An image-based guidance system, comprising: at least one detectable target, one or more portions of which are comprised of an optical contrast agent; A head-mounted display (HMD) device, comprising: imaging means for generating pixel data representative of a scene in use; a display means for outputting a graphical user interface in use; a power means; a data storage means for storing a geometric data set; and a data processing means operatively interfaced with said imaging means and said display means, said data processing means comprising: filtering the pixel data with a predetermined value to separate first pixel data from second pixel data, the first pixel data representing the one or more portions of the detectable target; at least one parallel processing module configured with a data processing thread for each of a plurality of pixels, the data processing thread being adapted to calculate two-dimensional (2D) pixel coordinate data from the separated first pixel data; the data processing means transforms the 2D pixel coordinate data into three-dimensional (3D) coordinate data representing the one or more portions of the detectable target by reference to a first geometric data set representing at least one coordinate system having an origin at the HMD device; generating 3D guidance data according to the 3D coordinate data by reference to one or more further geometric data sets each representing a respective detectable target in the scene; An image-based guidance system further adapted to output the 3D guidance data to the graphical user interface.

15. The system of claim 14 , wherein the optical contrast agent is an active agent that emits light waves.

16. The system of claim 14 , wherein the optical contrast agent is a passive agent, and further comprising an illumination source configured to excite the optical contrast agent.

17. The system of claim 16 , wherein the HMD device comprises the illumination source.

18. 18. The system of claim 14, wherein each of the one or more portions of the detectable target is a marker having a predetermined relative geometric relationship thereto.

19. at least one detectable target is a tool being used by or in proximity to the HMD wearer, and at least one of the one or more further geometric data sets includes a three-dimensional model representing the tool; and / or 20. The system of claim 18, wherein at least one detectable target is a marker defining a position within the scene, and at least one of the one or more additional geometric data sets includes a three-dimensional model representing the marker.

20. 20. The system of claim 19, wherein the marker is a matrix barcode, one or more portions of which are comprised of the optical contrast agent.

21. 21. The system of claim 19 or 20, wherein the scene includes at least two detectable targets, and wherein the data processing means, when generating the 3D guidance display data, is further programmed to generate display data representing a path between the two detectable targets in the scene.

22. 1. A method of navigating a detectable target using a head mounted display (HMD) device, comprising: generating pixel data of a scene using an imaging sensor of the HMD device, the detectable target being within the scene; Using at least one parallel processing module of the HMD device, which is configured with a data processing thread for each of a plurality of pixels, filtering the pixel data with a predetermined value to separate first pixel data from second pixel data, the first pixel data representing one or more portions of the detectable target comprised of an optical contrast agent; calculating two-dimensional (2D) pixel coordinate data from the separated first pixel data; using at least one further processing unit of said HMD device, transforming the 2D pixel coordinate data into three-dimensional (3D) coordinate data representing the one or more portions of the detectable target by referencing a first geometric data set representing at least one coordinate system having an origin at the HMD device; generating 3D guidance display data in accordance with said 3D coordinate data by referencing one or more further geometric data sets, each representing a respective detectable target in said scene; and outputting the 3D guidance display data to a graphical user interface on at least one display of the HMD device.

23. 23. The method of claim 22, wherein the converting step further comprises triangulating the 2D pixel coordinate data by reference to the first geometric data set.

24. 23. The method of claim 22, wherein the transforming step further comprises solving for rotation and translation based on the 2D pixel coordinate data.