Head mounted display device and system for real time guidance
By embedding a parallel processing module in the HMD device to preprocess image data, the problems of insufficient real-time performance and accuracy in the existing technology are solved, achieving low-latency image processing and efficient guidance effects.
Patent Information
- Application Number
- CN202480024861.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-05-05
- Filing Date
- 2024-02-14
- Publication Date
- 2025-11-11
AI Technical Summary
Existing image-based tool guidance systems are inadequate in terms of real-time performance and accuracy, especially HMD devices, which face challenges in data processing and rendering latency, leading to complexity and high cost.
A small parallel processing module is embedded in a lightweight head-mounted display (HMD) to reduce the amount of data by preprocessing the high-entropy image data captured by the high-resolution imaging sensor, and retain only the low-entropy image data for geometric and rendering calculations, thereby achieving low-latency image processing.
It improves the real-time guidance accuracy and efficiency of HMD devices, reduces data processing redundancy, and lowers system complexity and cost.
Smart Images

Figure CN120937340A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of head-mounted display ('HMD') devices that realize image-based guided functionality. Background Technology
[0002] Image-based guidance systems improve the efficiency and accuracy of their users during precise manipulation. An example manipulation involves inserting a needle-like surgical instrument into a patient at a specific location with a specific orientation. Such procedures typically require high positional accuracy, where image-based guidance systems help achieve the necessary accuracy by detecting, tracking, and rendering the current and desired tool posture on a display, referencing a field of view that includes the manipulation environment (e.g., the surgical site), for the system user to compare and adjust.
[0003] The technical challenges are immense because the tool needs to be accurately detected in image data amidst scene clutter within the camera's field of view. Similarly, the tool pose is detected with sub-millimeter accuracy, referencing six degrees of freedom (6DoF) of mechanical motion in three-dimensional space ('3D'). The calculated poses (both the actual and desired tool poses) then need to be precisely rendered on the display to guide the user during procedural processes, all essentially in real-time, i.e., with minimal latency between sensor input and display output.
[0004] There are many image-based tool-guided systems, such as the surgical navigation system YR02143 manufactured by Kalstein®. TM and StealthStation manufactured by Medtronic® TM The S8 surgical navigation system utilizes reference markers (such as QR codes), active markers (such as light-emitting diodes (LEDs), or passive markers (such as reflective points applied to or near surgical instruments and target objects) to track the spatial position and orientation of the target. Markers are detected in high-resolution images acquired by an optical system, typically placed in a fixed position overlooking the operating room, in a location least conspicuous to the surgeon and surgical site. The three-dimensional (3D) position of the markers is triangulated based on their detection in the images, thereby determining the current 3D position and orientation of the tool and target in the optical system's reference coordinate system. The results are then rendered on a display monitor, which the tool user must constantly refer to to adjust the tool's position and orientation toward the desired location. This constant comparison is subjective, error-prone, and time-inefficient.
[0005] By MA Lin et al. (" HoloNeedle: Augmented-reality Guidance System for Needle Placement Investigating the Advantages of 3D Needle Shape Reconstruction"In July 2018, IEEE Robotics and Automation Letters" and T. Kuzhagaliyev et al. (" Augmented reality needle ablation guidance tool for irreversible electroporation in the pancreas In March 2018, Proceedings 10576, Medical Imaging 2018: Image-Guided Procedures, Robotic Interventions and Modeling, proposed an alternative image-based tool guidance system to attempt and mitigate these drawbacks. This system is based on an augmented reality (AR) head-mounted display (HMD) device to provide HMD wearers with guidance about targets within their field of vision, such as tools in surgical sites and their destinations.
[0006] In such systems, the current and desired positions of the tools (and sometimes a pre-obtained scan of the surgical site and alignment guidance between the current and desired positions) are rendered onto the HMD's display. This advantageously eliminates the need for the wearer to leave the operating table to query the display monitor, thus reducing the risk of errors. However, the motion of the HMD AR headset relative to the reference coordinate system of the fixed optical system adds complexity, where the HMD's position and orientation also need to be accurately determined and tracked, for example, using QR codes or other markers. In these systems, the AR HMD device is always used as a simple display, where image processing, pose extraction and calculation, and other complex guidance-related image and data processing are performed by external computing resources, with the HMD tethered to these external resources. This is because AR HMDs do not have sufficient onboard data processing components to provide the necessary amount of image processing, complex calculations, and rendering to maintain real-time latency levels.
[0007] In a more recent alternative proposed by BJ Park et al. (“ Augmented reality improves procedural efficiency and reduces radiation dose for CT-guided lesions targeting: a phantom study using HoloLens 2 In Sci Rep 10, 18620 (2020); https: / / doi.org / 10.1038 / s41598-020-75676-4, researchers demonstrated that using AR HMD to provide holographic guidance for needle navigation improves the efficiency of needle-based procedures. However, this Vuforia-based approach... TMThe software development kit (SDK) approach still merely overlays the desired position of the tool in 3D onto a surface, allowing the HMD wearer to align the physical needle tool in their hand or another person's hand with the projected graphic. Because the actual needle position is not tracked, the registration and accuracy of the current needle position relative to the desired position are not estimated or calculated. Improvements to this approach, which can be expected from advancements in miniaturized processing components and power optimization, are mitigated by equipping HMD models with cameras that have increasingly higher frame rates and ever-expanding fields of view, where the correspondingly increasing capacity and density of processed image data are processed.
[0008] In summary, proven image-based navigation technologies rely on multiple distinct hardware systems for optical capture and tracking, data processing, and rendering, resulting in expensive and complex systems with significant data communication and synchronization requirements, all of which define multiple potential points of failure. Recent image-based navigation technologies attempting to simplify such systems require trade-offs between accuracy, functional differentiation, and availability to meet minimum latency requirements.
[0009] Therefore, there is a need for an HMD device that at least mitigates some of the drawbacks of these existing image-based guidance technologies. Summary of the Invention
[0010] Various aspects of the invention are set forth in the appended claims, relating to various embodiments of head-mounted display (HMD) devices, various embodiments of distributed imaging systems based on head-mounted display (HMD) devices, and various embodiments of methods for distributing image data in a network using such systems.
[0011] The concept of this invention lies in significantly reducing the amount of image data that portable display devices need to process to provide real-time image-based guidance. Recent advancements in hardware acceleration of embedded systems, particularly in small and power-efficient parallel processing modules, are used to preprocess high-entropy image data captured by high-resolution imaging sensors to reduce it to low-entropy image data that retains semantic content useful for guidance purposes. This low-entropy image data is then fed into other processing modules for geometric and rendering calculations. The technology of this invention provides a low-latency image processing technique particularly suitable for lightweight head-mounted display ('HMD') devices.
[0012] Therefore, in a first aspect, the present invention provides a head-mounted display (HMD) device, comprising: an imaging unit that generates pixel data representing a scene in use; a display unit that outputs a graphical user interface in use; a power unit; a data storage unit storing geometric datasets; and a data processing unit operatively interfaced with the imaging unit and the display unit, wherein the data processing unit includes at least one parallel processing module configured with a plurality of pixel-specific data processing threads, adapted to filter pixel data with predetermined values to separate first pixel data from second pixel data, wherein the first pixel data represents an optical contrast agent applied to at least one target in the scene, and to calculate two-dimensional (2D) pixel coordinate data from the separated first pixel data; and wherein the data processing unit is further adapted to transform the calculated 2D pixel coordinate data into three-dimensional (3D) coordinate data representing the target or target relative to the device by referring to a first geometric dataset representing at least one coordinate system derived from the device, and to generate 3D guidance data based on the 3D coordinate data by referring to one or more additional geometric datasets, each geometric dataset representing a corresponding target in the scene; wherein the data processing unit is further adapted to output the 3D guidance data to the graphical user interface.
[0013] Therefore, the technology of the present invention provides a low-latency image processing technique, which is particularly suitable for lightweight head-mounted display ('HMD') devices. By balancing reference accuracy, ergonomy, and usability, and to meet minimum latency requirements, the technology of the present invention advantageously reduces the data processing overhead associated with increasingly higher resolution and wider field-of-view cameras by filtering out redundant captured image data for guidance purposes while retaining the semantic image data most important for guidance purposes. This is desirable for real-time, accurate target capture(s).
[0014] In embodiments of the HMD device, the data processing unit may further be adapted to transform the calculated 2D pixel coordinate data into 3D pixel coordinate data by performing triangulation on the 2D pixel coordinate data with reference to a first geometric dataset. Alternatively, the data processing unit may further be adapted to transform the calculated 2D pixel coordinate data into 3D pixel coordinate data by solving for 3D rotation and translation based on the 2D pixel coordinate data.
[0015] In embodiments of the HMD device, the origin of at least one coordinate system derived from the device can be selected from the aperture of the imaging element, the display unit of the display element, and one eye of the HMD wearer. In variations of such embodiments, the first geometric dataset may include a set of calibrated transformations derived from the coordinate systems of the aperture of the imaging element, the display unit, and the eye of the HMD wearer, respectively.
[0016] Depending on the operational requirements of the embodiments and the technical capabilities of the components that can be used to implement them, the imaging component can be implemented as a single imaging sensor, optionally a pair of imaging sensors arranged in a stereoscopic manner, a hybrid combination of high-resolution and low-resolution imaging sensors, or a hybrid combination of one or more imaging sensors and one or more distance sensors, such as time-of-flight, event-based, or echolocation types.
[0017] In embodiments of the HMD device, at least one target in the scene may be a tool used by or near the HMD wearer, and at least one of one or more additional geometric datasets includes a three-dimensional model representing the tool. Alternatively or additionally, the target may be a marker defining a location in the scene, and at least one of one or more additional geometric dataset values includes a three-dimensional model representing the marker. Also alternatively, the target may be a biomarker, such as subcutaneous tissue that fluoresces by injection or ingestion of an optical contrast agent.
[0018] In a variant of such an embodiment, which is particularly suited to scenarios involving at least two objectives, the data processing unit may be further adapted to generate display data representing paths between tools and markers in the scene when generating 3D guided display data.
[0019] Embodiments of the HMD device may be designed for use with a passive optical contrast agent, and may therefore further include a switchable illumination source operatively connected to a power component for supply, configured to excite the optical contrast agent in the scene.
[0020] In embodiments of the HMD device, the data processing unit may further include a graphics processing unit ('GPU') programmed to generate 3D guided display data based on the 3D guided data by referencing one or more additional geometric datasets; wherein the data processing unit may also be adapted to output the 3D guided display data to a graphical user interface.
[0021] Embodiments of the HMD device can be designed to enhance display accuracy, wherein the data processing unit is also adapted to determine the mismatch between the generated 3D guided display data and the wearer's eyes based on distance measurements and the position of the wearer's eyes, and to adjust the position of the generated guided display data in the graphical user interface according to the determined mismatch. Distance measurements can be performed based on stereoscopic image data and / or using an optional distance sensor of the HMD device.
[0022] Embodiments of the HMD device can be designed to enhance the optical accuracy of the wearer and can therefore further include a pair of eye imaging sensors, each generating eye pixel data representing the corresponding eye of the HMD wearer in use, and at least a second parallel processing module configured and operating according to the inventive principles disclosed herein. The second parallel processing may be correspondingly configured with multiple data processing threads corresponding to the eye pixels, adapted to filter the eye pixel data with predetermined values to separate first eye pixel data from second eye pixel data, wherein the first eye pixel data represents at least a portion of the eye of the wearer; and to calculate two-dimensional (2D) eye pixel coordinate data from the separated first eye pixel data. The data processing unit may also be correspondingly adapted to triangulate the 2D eye pixel coordinate data received from the second data parallel processing module with reference to a first geometric dataset, thereby generating three-dimensional (3D) eye coordinate data representing the wearer's eye relative to the focal point of the device, transforming the 3D coordinate data with reference to the 3D eye coordinate data, and generating 3D guidance data based on the transformed 3D coordinate data.
[0023] Variations of such embodiments can be designed to optimize the use of computing resources when generating display data, wherein the data processing unit is also adapted to set 2D eye coordinate data as the gaze point when generating guided display data; and wherein the data processing unit is also adapted to output the generated guided display data to a graphical user interface as display data recessed according to the gaze point.
[0024] For any of the foregoing embodiments, the said or each parallel processing module may be selected from the group consisting of a field-programmable gate array ('FPGA'), a graphics processing unit ('GPU'), a video processing unit ('VPU'), an application-specific integrated circuit ('ASIC'), an image signal processor ('ISP'), and a digital signal processor ('DSP'). Alternatively or additionally, the data processing unit may be selected from the group consisting of a hybrid programmable parallel central processing unit and a configurable processor.
[0025] In another aspect, the present invention provides a system, an image-based guidance system, comprising at least one detectable target, one or more portions of which are configured with an optical contrast agent; and a head-mounted display (HMD) device substantially as described above, wherein first pixel data represents one or more portions of the detectable target, and wherein a data processing unit generates three-dimensional (3D) coordinate data representing one or more portions of the detectable target, and wherein one or more additional geometric datasets, each representing a corresponding detectable target in the scene.
[0026] In embodiments of this system, the optical contrast agent can be an active agent that emits light waves, such as a light-emitting diode (LED). Alternatively, the optical contrast agent can be a passive agent, such as a fluorophore compound, wherein embodiments of the system may further include an illumination source configured to excite the optical contrast agent in use. In variations of such embodiments, the HMD device may include an illumination source to ensure that the excited agent coincides with the field of view of the HMD device's imaging sensor, and to minimize the number of hardware units in the system.
[0027] In embodiments of the system, each of one or more portions of a detectable target may be a marker having a predetermined relative geometric relationship with it. The detectable target may be, for example, a tool used by or near the HMD wearer, such as a needle, biopsy syringe, or other surgical device, whether operated by the HMD wearer or by another user or robotic device adjacent to the HMD wearer, and at least one of one or more additional geometric datasets includes a three-dimensional model representing the tool.
[0028] Alternatively or additionally, the or each detectable target may be, for example, a marker defining its location within a scene, such as a reference marker indicating the target destination of a surgical device on a patient's body, and at least one of one or more additional geometric datasets may include a three-dimensional model representing the marker. An embodiment of such a marker may be a matrix barcode, also known as a Quick Response (“QR”) code or ArUco marker, one or more of which is configured with a passive or active optical contrast agent, or a combination of passive and active optical contrast agents. Another embodiment of a marker defining its location within a scene may be biological tissue that has been fluorescently rendered by an optical contrast agent after injection or uptake, having a predetermined relative geometric relationship with one or more features of the surrounding tissue, such as the distance of the fluorescent tissue relative to the tissue surface.
[0029] In a variant of such an embodiment, which is particularly well-suited to scenarios including at least two detectable targets, the HMD’s data processing unit can also be programmed to generate display data representing the path between the two detectable targets in the scene when generating 3D guided display data.
[0030] In another aspect, the present invention provides a method for guiding a detectable target using a head-mounted display (HMD) device, comprising the steps of: generating pixel data of a scene using an imaging sensor of the HMD device, wherein the detectable target is in the scene; using at least one parallel processing module of the HMD device, wherein the module is configured with a plurality of pixel-specific data processing threads, filtering the pixel data with predetermined values to separate first pixel data from second pixel data, wherein the first pixel data represents one or more portions of the detectable target disposed of an optical contrast agent, and calculating two-dimensional (2D) pixel coordinate data from the separated first pixel data; using at least one additional processing unit of the HMD device, transforming the 2D pixel coordinate data into three-dimensional (3D) coordinate data representing one or more portions of the detectable target by referring to a first geometric dataset representing at least one coordinate system originating from the HMD device, generating 3D guided display data based on the 3D coordinate data by referring to one or more additional geometric datasets, each geometric dataset representing a corresponding detectable target in the scene, and outputting the 3D guided display data to a graphical user interface on at least one display of the HMD device.
[0031] Other aspects of the invention are set forth in the appended claims. Attached Figure Description
[0032] Referring to the accompanying drawings, the invention will be more clearly understood from the following description of embodiments given by way of example only, wherein - Figure 1 A front view of an embodiment of a head-mounted display (HMD) device including an imaging sensor according to the present invention is provided.
[0033] Figure 2 A front view of an alternative embodiment of the HMD device according to the present invention is provided, Figure 1 The HMD device is equipped with a lighting source.
[0034] Figure 3 It shows Figure 1 The example hardware architecture of the HMD shown includes an imaging sensor, a data processing unit, and memory.
[0035] Figure 4 It shows Figure 2 The example hardware architecture of the HMD shown is illustrated.
[0036] Figure 5 The diagram illustrates what can be obtained from Figures 1 to 4 An example of an imaging sensor observing a detectable target.
[0037] Figure 6 The diagram illustrates what can be obtained from Figures 1 to 4 Another embodiment of the detectable target observed by the imaging sensor.
[0038] Figure 7 Detailed description of the Figures 1 to 4 The data processing steps performed by the HMD device are used to generate information about... Figure 5 and / or Figure 6 The target 3D guiding data includes steps of filtering pixel data and transforming 2D coordinate data.
[0039] Figure 8 The diagram illustrates when execution Figure 7 During the steps, Figure 3 or Figure 4 The contents of the memory during runtime.
[0040] Figure 9 The diagram shows... Figure 7 and Figure 8 The data processing and related data type processes.
[0041] Figure 10 Further details are provided by Figure 3 and Figure 4 Filtering performed by parallel processing units Figure 7 and Figure 9 The steps for processing pixel data.
[0042] Figure 11 Further details are provided by Figure 3 and Figure 4 The other processing units perform the generation Figure 7 and Figure 9 The steps of 3D guide data in the process.
[0043] Figure 12 Further details Figure 7 and Figure 10 An alternative embodiment of the step of filtering pixel data.
[0044] Figure 13 A front view of an alternative embodiment of an HMD device including a time-of-flight sensor is provided. Detailed Implementation
[0045] Specific patterns conceived by the inventors will now be described by way of example. In the following description and figures, numerous specific details are set forth to provide a comprehensive understanding, wherein similar reference numerals denote similar features. It will be apparent to those skilled in the art that the invention can be practiced without being limited to these specific details. In other instances, well-known methods and structures have not been described in detail to avoid unnecessarily obscuring the description.
[0046] refer to Figure 1A first embodiment 10A of an HMD device including an imaging component and a display component according to the present invention is shown, in this example being an augmented reality ('AR') device. The AR HMD 10A includes wearer goggles 20, which include a main viewing portion 22 and corresponding video display portions 24A, 24B for the eyes. In use, the video display portions 24A, 24B are equidistantly located on a central bridge portion of the wearer's nose. The display portions 24A, 24B perceptually realize a single video display occupying a subset of the front of the HMD, allowing the wearer to observe the surrounding physical environment in front of the HMD 10 and the video content superimposed thereon.
[0047] Each video display portion 24A 24B comprises a corresponding video display unit 26A, 26B. In this example, the video display units 26A, 26B are miniature OLED panels with a minimum 60 Hz frame refresh rate and a 1920×1080 pixel resolution, located near the lower edge of the goggles so that the see-through portion 22 extends above it and to its upper edge without visual obstruction when the VDU is in positive display. The HMD also includes first and second high-resolution imaging sensors 30A, 30B arranged in a stereoscopic configuration. Each high-resolution imaging sensor captures visible light in a wavelength range typically between 400 and 700 nm within a field of view (FoV) typically between 70 and 160 degrees or even greater, and outputs the captured images as a pixel data stream at a resolution of at least 1920×1080 pixels and a rate of 60 frames per second or higher.
[0048] refer to Figure 2 Similar reference numerals in the figures indicate relative to Figure 1 Similar features are shown in a second embodiment 10B of a head-mounted display ('HMD') device including an imaging component and a display component according to the invention, the device further including an illumination source 32, such as a light-emitting diode ('LED') 32, which emits light in the wavelength range of 800 to 2500 nm corresponding to near-infrared ('NIR') light, for exciting one or more aspects of a target or object disposed in a corresponding FOV of the HMD imaging sensor 30A-B with a passive optical contrast agent, such as a fluorophore.
[0049] Those skilled in the art will also appreciate that the technical principles disclosed herein can be implemented in other HMD types, such as augmented reality monocular or contact lens devices, virtual reality ('VR') or mixed reality ('MR') enclosed display devices, wherein the corresponding video display portions 24A, 24B for each eye perceptually realize a single video display portion that substantially occupies the entire inner front surface of the HMD. For such an HMD, for a perceptual single video display with a resolution of 4096 × 2160 pixels, each video display portion 24A, 24B can consist of an RGB low-persistence panel with a minimum frame refresh rate of 60 Hz and an individual resolution of 2048 × 1080 pixels per eye.
[0050] Through the imaging component, embodiments of the HMD according to the present invention may include fewer or more imaging sensors. The technology of the present invention can be implemented with a single imaging sensor 30A and a target of known geometry containing at least four detectable points, or with two imaging sensors 30A-B having at least three detectable points as described above. Embodiments of the HMD according to the present invention may also, or alternatively, include other types of sensors, such as distance or depth sensors implementing time-of-flight technology, which, according to the principles described below, are particularly useful for preventing display artifacts. References are made below respectively... Figure 3 and Figure 4 An example hardware architecture of the HMD device 10A-B according to the present invention is described in more detail, wherein, by way of non-limiting example, similar numbers represent similar features.
[0051] All embodiments of the HMD according to the present invention include a data processing unit. Therefore, in addition to the video display units 26A-26B and the imaging sensors 30A-B, and the optional NIR LED 32, the HMD according to the present invention includes a data processing unit consisting of at least one data processing unit 301 and at least one parallel processing module 302. The data processing unit 301 acts as the main controller of the HMD, and the parallel processing module 302 preprocesses the pixel data generated by the imaging sensors 30A-B. The CPU 301 is, for example, a general-purpose microprocessor based on the Cortex™ architecture manufactured by ARM™, and the parallel processing module 302 is, for example, based on the Cortex™ architecture manufactured by AMD. TM Xylinx TM Artix manufactured TM Field-Programmable Gate Array ('FPGA') semiconductor devices.
[0052] The CPU 301 may also include or be associated with a dedicated graphics processing unit ('GPU') 321, which receives data and processing commands from the CPU 301 to generate display data before it is output to the displays 26A-26B. The CPU 301 and FPGA 302 are coupled to a memory unit 303 via a data input / output bus 304, which includes volatile random access memory (RAM), non-volatile random access memory (NVRAM), or a combination thereof. The CPU 301 and FPGA 302 communicate via the data input / output bus 304, and other components of the HMD 10A-B are similarly connected to the data input / output bus 304 to provide headset functionality and receive user commands.
[0053] The data connection between the imaging sensors 30A-B, CPU 301 and FPGA 302 via bus 304 or another bus is a high-frequency data communication interface, and at least FPGA 302 (but preferably also CPU 301) is positioned closest to the interconnect of the imaging sensors 30A-B, i.e., on the front or part of the HMD device, in order to minimize the latency of the data stream.
[0054] User input data can be received directly from physical input interface 305, which may be one or more buttons including at least an on / off switch, and / or part of an HMD housing configured for haptic interaction with the wearer's touch. User input data can also be received indirectly, such as gestures optically captured by optical sensors 30A-B and / or spoken words captured as analog sound wave data by microphone 306. DSP module 307 performs analog-to-digital conversion for this purpose, and CPU 301 then interprets the function according to principles outside the scope of this disclosure. Processed audio data is output to speaker unit 308, and power is supplied to all components by circuitry 309, which interfaces with internal battery module 310, where the battery is periodically recharged by power converter 311.
[0055] HMD embodiments may also include data connectivity capabilities, shown in dashed lines as a wireless network interface card or module (WNIC) 322, also connected to data input / output bus 304 and circuitry 308, and adapted to connect the HMD to a wireless local area network ('WLAN') interface generated by a local wireless router. Alternative or additional wireless data communication functionality may be provided by the same or another module, for example, enabling short-range data communication according to Bluetooth™ and / or Near Field Communication (NFC) interoperability and data communication protocols.
[0056] The HMD device of the present invention is used to guide items relative to other items and / or item destinations in a scene, such as guiding surgical devices during a surgical procedure. Therefore, the items and item destinations need to be detectable targets, and embodiments of the invention rely on configuring such targets with passive or active optical contrast agents to be captured as imaging data by the imaging sensors 30A-B in use.
[0057] refer to Figure 5 The illustration shows a first embodiment of a detectable target, in which the surgical tool 50 has an elongated body 51 terminating at a first end at a needle 52 for facilitating subcutaneous insertion, and a user grip portion 54 located away from the needle and near a second, opposite end of the tool, terminating at a geometric reference point or indicator 55 coated with a passive contrast agent. Those skilled in the art will appreciate that the passive geometric indicator 55 can replace an active indicator, such as an LED, and is equally compatible with the techniques described herein.
[0058] The 3D pose of any detectable target in a scenario facing HMD 10A-B needs to be determined; therefore, the example surgical tool 50 includes at least two additional geometric reference indicators 55, each attached to the gripping portion 54 via rod-like members 56, each rod oriented orthogonal to the principal axis of the other rod and the elongated body 51. Appropriately, the three geometric reference indicators 55 collectively define a 3D coordinate system N derived from the target 50, with its axes coaxial with the tool's principal axis. The longitudinal dimension of the surgical tool 50 between its opposite ends 52, 55 is known; similarly, the corresponding dimension between each additional geometric reference indicator 55 and the gripping portion surface is also known, thereby the 3D geometry 58 of the tool is known and accordingly preset.
[0059] refer to Figure 6 Two further embodiments of detectable targets are shown, which can be used as reference markers to indicate the destination of an item in a scene. In the first example, the surgical marker 60A has a planar body 62 of square shape, with a plurality of positioning foot members 64 extending from the underside of the body 62, wherein each foot member may optionally terminate by a needle located at the distal end of the body 62 for subcutaneous insertion, or by a clip, gripper, or some other component attached to the patient. Geometric reference indicators 55, as previously described, are fixed to each corner of the planar body 62, wherein the four geometric reference indicators 55 correspond to and thus define the principal plane of the marker 60A and its orientation at any given time. Appropriately, any three of the four geometric reference indicators 55 collectively define a three-dimensional coordinate system G derived from the geometric center of the principal plane of the target 60A, whose two orthogonal axes are coplanar with the principal plane of the reference marker, and whose third axis is orthogonal to it.
[0060] In the second example, for simplicity, another surgical marker 60B has essentially the same configuration as before. However, instead of fixing the physical indicator 55 to some or every corner of its planar body 62, in this embodiment, the top side of the main plane of the body is configured according to ArUco. TM The technology features a patterned matrix barcode with several geometric sections 65, each coated with a passive optical contrast agent. The white portion is coated with ArUco in this example. TM As indicated by marking 60B, but which will be readily apparent to those skilled in the art, the black portion may alternatively be coated with a passive optical contrast agent, and similarly, the passive optical contrast agent may replace one or more active indicators such as LEDs. ArUco is known. TM The marker encodes more geometric and semantic information than classic matrix (QR) barcodes and dot-matrix passive or active indicators 55, wherein this enhanced accuracy can be utilized by adding optical contrast characteristics through the techniques of the present invention.
[0061] Therefore, again in this example, the geometric reference indicator 65, composed of an optical contrast pattern, at least defines the principal plane of the mark 60B and its orientation at any given time, and can encode further information, such as the size data of the mark. Appropriately, the geometric reference indicator 65 in this example defines the same three-dimensional coordinate system G originating from the geometric center of the principal plane of the target 60B, whose two orthogonal axes are coplanar with the principal plane of the mark, and whose third axis is orthogonal to it.
[0062] The lateral dimensions of surgical markers 60A and 60B between the two corners of their main plane are known, as are the corresponding positions and dimensions of each foot member 64 extending below the main plane, and / or can be encoded in a detectable pattern 65 thereon and decoded by the relevant configuration of the HMD, thereby the three-dimensional geometry 68 of the markers is known and preset accordingly.
[0063] Now refer to Figures 7 to 11 According to embodiments of the present invention Figures 1 to 4 The basic and enhanced data processing configurations and functionalities of the HMD 10A-B, in which... Figure 8 The data structure stored in memory 303 and processed by FPGA 302 and CPU 301 is shown in the figure. Similar numbers throughout the text refer to similar features, steps and structures.
[0064] Upon initial power-up of the HMD 10A-B, the operating system 801 is initially loaded at step 701 to manage basic data processing, interdependencies, and interoperability of HMD components 26A-B, 30A-B, 32 (if present), and 301 to 321 (including, in addition, WNIC 322, if present). The HMD OS may be based on Android™ distributed by Google™ in Mountain View, California, USA. The OS includes device drivers for the HMD components, input subroutines for reading and processing input data, including direct user input to the physical interface device 305, and output subroutines for outputting display data to the displays 26A-B. Notably, in this example, OS 801 interfaces the output of the imaging sensors 30A-B to the FPGA 302 at a kernel level within a low computational layer for minimal latency. In an embodiment of the HMD 10A-B that includes networking component 322, OS601 also includes communication subroutine 802 to configure the HMD 10A-B to perform bilateral network communication with a remote terminal via WNIC 322, which interfaces with a network router device.
[0065] At step 702, a set of instructions embodying the visualization application 803 is loaded as a subroutine of OS 801 or as a different application in a higher computing layer. The visualization application 803 interfaces with the FPGA 302 via OS 801 through one or more application programming interfaces (APIs) 804. The visualization application 803 includes and coordinates data processing subroutines embodying the various functions described herein, including real-time updates and output of the user interface 805 to displays 26A, 26B.
[0066] In addition to initializing the imaging sensors 30A-B at step 703 to generate the corresponding pixel data stream 806, the user interface 805 itself is initialized at step 704 when including one or more targets 50, 60A-B and their corresponding geometric reference indicators 55, 65, thereby configuring the HMD 10A-B to begin real-time processing of image data for display guidance.
[0067] At step 705, FPGA 302 receives pixel data stream 806 and filters each input pixel according to a predetermined value, such as a pixel brightness or luminance threshold corresponding to the captured optical contrast agent of the geometric reference indicator 55 in an excited state, or a pixel position offset relative to the position in a previous capture.
[0068] For details, please refer to the following: Figure 10 Each pixel in the 806 pixel data stream has 900 pixels. N The data is fed into the corresponding data processing thread or pipeline 920 implemented within the massively parallel architecture of the FPGA 302.N The corresponding input box 910 N 920 per parallel pipeline 1-N Configured to receive pixel-related data at step 811, perform one or more standard artifact removal operations at step 812, such as correction and lens-related distortion removal, to obtain corrected, accurate pixel data, and then at step 813 separate accurate pixel data between pixel data that matches or exceeds a predetermined value and pixel data that falls below the predetermined value, which therefore corresponds to a geometric reference indicator 55 in the pixel data stream. This predetermined value corresponds to any other aspect of the captured scene that is of no further interest and is therefore discarded. Figure 9 As illustrated, the output of step 803 is very low-entropy binary image data that encodes only data representing each geometric reference indicator 55 in the field of view of the imaging sensors 30A-B and its corresponding and corrected two-dimensional (2D) screen coordinates in the frame, which is computed at step 814 for extraction.
[0069] As shown at 807, the output of FPGA 302 at step 705 is correspondingly low-entropy data, which describes each pixel representing the geometric reference indicator 55 and its corrected 2D position within the field of view of the imaging sensors 30A-B at that precise moment. This data is significantly less than the full-resolution or even low-resolution RGB or grayscale image data typically used in known image-based navigation systems.
[0070] At step 706, CPU 301 receives 2D pixel coordinate data 807 from FPGA 302 and transforms it into three-dimensional (3D) pixel coordinate data 809. This transformation can be implemented using various techniques, and as a non-limiting approach, the transformation can be implemented using various techniques, for example, depending on whether HMD 10 includes a single imaging sensor 30A or a pair of imaging sensors 30A-B arranged in a stereo configuration.
[0071] In the case of stereo HMDs 10A and 10b, the transformation can be achieved through triangulation techniques, wherein CPU 301 triangulates 2D pixel coordinate data 807 from FPGA 302 with reference to a first geometric dataset 808 representing at least one coordinate system derived from HMDs 10A-10b, thereby generating 3D coordinate data 809 representing the aforementioned or each target 50, 60A-B relative to HMDs 10A-10b. In this document, the verb “triangulation” and its equivalent adjectives and expressions should be understood in their common sense in the field of computer vision, namely, the process of determining a point in 3D space given its projection onto a 2D image plane. Therefore, those skilled in the art will understand that different techniques can be used to implement this particular process, such as direct linear transformation, or, in HMD embodiments with high-resolution and high-frame-rate imaging sensors (a more computationally efficient alternative), based on the quality and accuracy of the 2D dataset 807.
[0072] Specific reference Figure 11 The characteristics of the HMD10A-B and the intrinsic parameters of its components related to the optical and vision systems embodied therein are known during manufacturing and pre-setting. The reference geometric dataset 808 accordingly includes one or more coordinate systems, each derived from a corresponding component of the HMD. These coordinate systems are pre-calibrated, as is the 6 DoF transformation between them, which is represented by a 3×3 rotation matrix R and a 3×1 translation vector t. In this example, the three-dimensional coordinate system H originates from the imaging sensor 30A, the three-dimensional coordinate system S originates from the display 26A, and the three-dimensional coordinate system E originates from the right or left eye 1100 of the HMD wearer, and is included in the first geometric dataset 808.
[0073] Therefore, when the triangulation of step 706 is performed at step 821 with reference to the first geometric dataset 808, the distance between H and E is... And between E and S, that is The 6 DoF transform is known and has been pre-calculated accordingly. Similarly, the transform between the left and right imaging sensors 30A-B (and therefore the baseline distance b between them) is also known. Assuming the corresponding intrinsic parameters of this pair of imaging sensors 30A-B are identical, they are defined as follows: in These correspond to the focal lengths of the imaging sensor in the x and y dimensions, respectively, and This corresponds to the projection center of the imaging sensor in the x and y dimensions.
[0074] Given stereo image data captured by the pair of imaging sensors 30A-B respectively In step 705, a set of scenes is detected. Geometric indicators 55, and extract their corresponding 2D positions. and 807, and receive them as input at step 706. The corresponding position for each pixel of the geometry indicator 55. and 3D position It is given by the following formula, where : .
[0075] 3D position then via Transform from H to E. Consider the current eye position relative to S, i.e. This is then correctly projected onto S. This is achieved by modeling the eye 1100 and the HMD display 30A-B as having an intrinsic matrix. This is accomplished using an off-axis pinhole imaging sensor, and the intrinsic matrix is defined as: in It is a scaling factor that converts 3D points on the screen into pixels.
[0076] Therefore, as three-dimensional (3D) coordinate data 809, S on The rendering location is determined as follows: .
[0077] The above assumption is between the camera and the eye. And between the monitor and the eyes The transformations are known and pre-computed. In some embodiments, these transformations can be computed using an additional eye tracker (e.g., F), where the pre-computed or pre-calibrated transformations include those between the camera and the tracker. And between the camera and the monitor The transformation. During runtime, the tracker estimates the eye's position and outputs eye coordinate data, which is used to calculate and update the tracker and eye coordinates. The transformation between them, where via , and Calculate, and in via and calculate.
[0078] Alternatively, in the case of an HMD with a single imaging sensor 30A, the transformation can be achieved by solving for rotation and translation based on 2D pixel coordinate data, i.e., by using solver techniques, such as pose-based problems like the perspective n-point problem ('PnP'), which aims to recover the position and orientation of an object by aligning the 2D image data captured with a 3D model describing the real world.
[0079] This technique calculates the pose of targets 50 and 60 from N feature points, including a rotation matrix R and a translation vector t between the world coordinate system of targets 50 and 60 and the coordinate system of the imaging sensor 30A, where N ≥ 3, corresponding to the 2D pixel coordinate data 807 as input. The solution output by the solver is the solution that minimizes the reprojection error between the 3D points of the targets and the input 2D point data, corresponding to the 3D coordinate data 809 representing the 3D coordinate data 809 of each target 50, 60A-B relative to HMD 10A-B. After completing the transformation in step 706, in step 707, CPU 301 generates 3D guidance data based on 3D coordinate data 809 by referencing one or more additional geometric datasets. Each geometric dataset represents a corresponding target 50, 60A-B in the scene, and its example is... Figure 5 and 6 The preset geometries 58 and 68 are shown.
[0080] Therefore, at step 821, based on the quorum of at least three geometric indicators 55 for each detected target, CPU 301 calculates or 'matches' the pose of said or each target 50, 60A-B. Many techniques are known to implement this step, the purpose of which is to compare the 3D coordinate data 809 of step 821 with preset geometries 58, 68 of the detected targets in the scene for geometric matching, and, when matching one or more geometries, determine the orientation of said or each matched geometry in 6DoF or pose at that precise moment. An example technique is Umeyama's least squares estimation technique (IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol.13, Issue 4, April 1991). CPU 301 can additionally calculate semantic, baseline, or other meaningful indicator data, which can be derived from the geometric information inherent in the extracted pose and can be displayed in user interface 805, such as a guide line 900 extending between the detected target 50 and the detected markers 60A-B, the size of which is calculated with reference to known preset geometries 58, 68, and the 6DoF pose is calculated with reference to their respective extracted poses.
[0081] At the next step 822, the CPU 301 calculates the transformation of the extracted pose of the first matching geometry 58, 68 relative to the field of view origin corresponding to the viewpoint of the user interface 805, i.e., the transformation relative to the 'virtual camera' referenced by which the 3D object is rendered in the user interface 805. Next, at step 823, a question is posed regarding whether the poses of additional matching geometry 58, 68 and / or semantic or reference indicators should be extracted and transformed, where control returns affirmatively to step 821. Alternatively, or ultimately, when the corresponding poses of each matching geometry 58, 68 and optional semantic or reference indicator have been extracted and transformed, the calculated transformation accordingly includes 3D guiding data 810, i.e., data defining the pose of said or each target 50, 60A-B detected in the field of view of the imaging sensors 30A-B at that precise moment, and optionally data defining any additional semantic or reference indicators, which are rendered to the user interface 805 as rapidly as allowed by the processing latency of the HMD components involved.
[0082] In the next step 708, the 3D guidance data 810 is submitted to the rendering subroutine of the visualization application 804, which is processed by the CPU 301, or, if present in an HMD architecture, by the GPU 321, advantageously reducing the corresponding graphics processing overhead of the CPU 301. For example, 3D guidance display data 820 is generated from the 3D guidance data 810 by mapping (one or more) corresponding target geometry datasets 58, 68 to the corresponding target pose data 810, and mapping semantic or reference graphics data 900 to semantic or reference indicator pose data 810. The rendering subroutine can also transform the 3D guidance data 810 based on any intermediate updates to the world coordinate space, since at least the coordinates of the imaging sensor 30A, and optionally the coordinates of other (one or more) anchors such as the eye 1100, can change significantly from the current display frame to the next display frame in order to maintain the relative positions of their respective physical locations.
[0083] At step 709, the rendered 3D guided display data 820 is appropriately output to the user interface 805, still as quickly as allowed by the processing latency of the HMD components involved. This presents the HMD wearer with a representation 58 of the surgical tool 50, a representation 68 of reference markers 60A or 60B, and a reference indicator 900 extending between these two representations, superimposed on the field of view of the goggles 20 with accurate corresponding positions and orientations. At step 710, an inquiry is made regarding whether HMD operation should be interrupted; a negative answer is received as long as the HMD is still in use, and the logic returns to the pixel separation in step 705, and so on, until the HMD should be turned off.
[0084] refer to Figure 12 A. After separation in step 803, an alternative embodiment of the logic of the parallel processing pipeline 920 of FPGA 302 can be modified to query whether the tracking of targets 50, 60A-B in scene 1210 has been interrupted. For example, this query can be implemented with a simple buffer of the separation states of a set of pixels, i.e., where the separation states of the same pixels are compared between the previous frame and the current frame, and where a change in the pixel separation state between two consecutive frames indicates a tracking interruption.
[0085] When the question in step 1210 is answered negatively, the logic maintains target tracking within the pixel data stream 806 in step 1220. Alternatively, when the question in step 1210 is answered positively, the logic reacquires the target within the pixel data stream 806 at step 1230. In either case, the logic then proceeds to the 2D coordinate data extraction in step 804.
[0086] refer to Figure 12 B. An alternative embodiment of the HMD device 10C may be configured with a single imaging sensor 30A and a time-of-flight (ToF) sensor 1200, which facilitates the detection and tracking of detectable targets 50, 60A-B by determining the corresponding distances from the detectable targets 50, 60A-B to the HMD device 10C. As previously described, according to step 705, the single imaging sensor 30A supplies a pixel stream 806 to the FPGA 302 for low-entropy filtering. In this embodiment, the characteristics of the HMD 10C and the intrinsic parameters of its components remain known from manufacturing and are preset, wherein the reference geometry dataset 808 accordingly includes an additional coordinate system D derived from the ToF sensor 1200. Since the triangulation of step 706 is performed again at step 706 with reference to the first geometry dataset 808, the distance between H and D (i.e., The 6 DoF transform between D and E (i.e.) The 6 DoF transform of E and S, and the relationship between E and S (i.e. The 6 DoF transform of ), where the 3D position is now correspondingly changed via The transformation from H to E is performed. The CPU 301 uses the target-specific distance data supplied by the ToF sensor 1200 to estimate the position of each detected target relative to the HMD, and utilizes this position data to facilitate the generation of 3D guidance data at step 707. Thus, the ToF sensor 1200 maintains the accuracy of the techniques disclosed herein using a single imaging sensor, while still advantageously further reducing the total amount of image data processed by the HMD architecture.
[0087] The CPU 301 of the HMD 10C can also be adapted to determine the mismatch between the generated 3D guided display data 820 and the wearer's eyes 1100 based on distance measurement and the position of the wearer's eyes, and adjust the position of the generated 3D guided display data 820 in the graphical user interface 805 according to the determined mismatch, as taught by the applicant in GB 2588774 A. For example, the distance measurement can be performed based on stereoscopic image data captured by the image sensors 30A-B of the stereoscopic HMDs 10A and 10B, and / or can be performed or enhanced by the ToF sensor 1200 of the HMD 10C.
[0088] Consider further embodiments of the hardware-accelerated pixel data separation technique for AR navigation purposes disclosed herein. Imaging sensors 30A-B can be configured to perform optical filtering of the scene and, with reference to active or passive optical contrast agents applied to targets 50, 60A-B, filter geometry indicator 55 during acquisition, supplying the separated data to parallel processing module 302, on which steps 812 and 814 are performed, thereby significantly accelerating the technique. When based on a threshold method, the filtering in step 803 can be implemented using different value-based comparators. In a first alternative, the luminance or luminance value in pixel data 806 can be rounded or truncated with reference to a threshold or range. In a second alternative, pixel data 806 can be selected in each pipeline 920 with reference to a threshold or range stored in parallel module 302 or memory 303.
[0089] For any of these alternatives, as well as the main embodiments described herein, separation can be further accelerated by spatial selectivity, wherein only a subset of the pixel data 806 defined with reference to the resolution of the imaging sensors 30A-B (e.g., a range or diameter with the center pixel as its origin and represented as a pixel count, or some other predetermined location in the captured image data) is input to the parallel pipeline 920.
[0090] Surgical tools and markers are known, comprising active or passive magnetic transponder units or modules used to aid in determining their location within the operating room. When such a device is configured with an optical contrast agent, the triangulation of 2D pixel coordinate data at step 706 and the generation of 3D guidance data at step 707 can be facilitated essentially in the same manner as when target-related location data is determined using distance measured by the time-of-flight sensor 1200 of the HMD10C, by using target-related location data acquired from the HMD's magnetic sensor or from an external device or system communicating wirelessly with it via the WNIC 322.
[0091] Those skilled in the art will understand that the hardware-accelerated pixel data separation technique disclosed herein for AR navigation purposes has been described by way of non-limiting example with reference to the field of surgical applications and is adaptable to many other application areas where accurate, low-latency visual guidance is desired with maximum portability and minimal obstruction to the HMD wearer. Furthermore, those skilled in the art will also understand that the pixel data separation technique disclosed herein can be applied to other types of HMDs, such as video-based virtual reality (VR) and, in particular, mixed reality (MR) HMDs, because the imaging sensors, overlapping fields of view, and gaze visualization are essentially similar, including the ability to image its forward field of view as an alternative to the transparent goggles 20, for MR HMDs.
[0092] In this specification, the terms “comprise” (“comprise,” “comprise,” “comprise,” and “comprising”) or any variations thereof and the terms “include,” “includes,” “included,” and “including” or any variations thereof are considered to be completely interchangeable, and they should be given the broadest possible interpretation, and vice versa. The invention is not limited to the embodiments described above, but may vary in structure and detail.
Claims
1. A head-mounted display (HMD) device, comprising: Imaging components that generate pixel data representing the scene during use; The display component outputs a graphical user interface during use; A power unit, a data storage unit for storing geometric datasets, and a data processing unit operatively interfaced with the imaging unit and the display unit, wherein the data processing unit includes - At least one parallel processing module, configured with multiple data processing threads corresponding to each pixel, is suitable for - Pixel data is filtered using a predetermined value to separate first pixel data from second pixel data, wherein the first pixel data represents an optical contrast agent applied to at least one target in the scene, and Calculate two-dimensional (2D) pixel coordinate data from the separated first pixel data; as well as The data processing component is also adapted to - The calculated 2D pixel coordinate data is transformed into three-dimensional (3D) coordinate data representing the target or target relative to the device by referring to a first geometric dataset representing at least one coordinate system originating from the device. 3D guidance data is generated based on the 3D coordinate data by referencing one or more other geometric datasets, where each geometric dataset represents a corresponding target in the scene. The 3D guidance data is output to the graphical user interface.
2. The head-mounted display device according to claim 1, wherein, The data processing unit is also adapted to transform the calculated 2D pixel coordinate data by performing triangulation on the 2D pixel coordinate data with reference to the first geometric dataset.
3. The head-mounted display device according to claim 1 or 2, wherein, The origin of the at least one coordinate system derived from the device is selected from the aperture of the imaging component, the display unit of the display component, and one eye of the HMD wearer.
4. The head-mounted display device according to claim 3, wherein, The first geometric dataset includes a set of calibrated transformations derived from the coordinate systems of the imaging component's aperture, the display unit, and the HMD wearer's eye.
5. The head-mounted display device according to claim 1, wherein, The data processing unit is also adapted to transform the calculated 2D pixel coordinate data by solving for rotation and translation based on the 2D pixel coordinate data.
6. The head-mounted display device according to any one of claims 1 to 5, in, At least one target in the scene is a tool used by or near the HMD wearer, and at least one of the one or more additional geometric datasets includes a three-dimensional model representing the tool; and / or The objective in the scene is a marker defining a location within the scene, and at least one of the one or more additional geometric datasets includes a 3D model representing the marker.
7. The head-mounted display device according to claim 6, wherein, The scene includes at least two targets, and the data processing unit is further programmed to generate display data representing a path between the two targets in the scene when the 3D guided display data is generated.
8. The head-mounted display device according to any one of claims 1 to 7, further comprising a switchable illumination source operatively connected to a power component for power supply, the switchable illumination source being configured to excite the optical contrast agent in the scene.
9. The head-mounted display device according to any one of claims 1 to 8, wherein, The data processing component further includes a graphics processing unit (GPU) programmed to generate 3D guided display data based on the 3D guided data, referencing the one or more additional geometric datasets; and The data processing component is also adapted to output the 3D guided display data to the graphical user interface.
10. The head-mounted display device according to any one of claims 1 to 9, wherein, The data processing component is also adapted to The mismatch between the generated 3D guided display data and the wearer's eyes is determined based on distance measurements and the position of the wearer's eyes, and Based on the determined mismatch, the position of the generated guide display data in the graphical user interface is adjusted; and optionally... The distance measurement is performed based on stereo image data, or using an optional distance sensor of the HMD device.
11. The head-mounted display device according to any one of claims 1 to 10, wherein, The imaging component also generates eye pixel data representing the corresponding eye of the HMD wearer during use. The HMD device also includes... At least the second parallel processing module, which is configured with multiple data processing threads corresponding to eye pixels, is suitable for - The eye pixel data is filtered using a predetermined value to separate the first eye pixel data from the second eye pixel data, wherein the first eye pixel data represents at least a portion of the eye of the wearer or each wearer. Calculate two-dimensional (2D) eye pixel coordinate data from the separated first eye pixel data; as well as The data processing component is also adapted to - The 2D eye pixel coordinate data received from the or each of the second data parallel processing modules is transformed into three-dimensional (3D) eye coordinate data representing the wearer's eye relative to the focal point of the device. Transform the 3D coordinate data by referring to the 3D eye coordinate data, and The 3D guidance data is generated based on the transformed 3D coordinate data.
12. The head-mounted display device according to claim 11, wherein, The data processing component is also adapted to When the guided display data is generated, the 2D eye coordinate data is set as the fixation point; as well as The generated guided display data is output to the graphical user interface as display data recessed according to the gaze point.
13. The head-mounted display device according to any one of claims 1 to 12, wherein, The parallel processing module is selected from the group consisting of Field Programmable Gate Array ('FPGA'), Graphics Processing Unit ('GPU'), Video Processing Unit ('VPU'), Application-Specific Integrated Circuit ('ASIC'), Image Signal Processor ('ISP'), and Digital Signal Processor ('DSP'); alternatively The data processing component is selected from a group including a hybrid programmable parallel central processing unit and a configurable processor.
14. An image-based guidance system, comprising: At least one detectable target, one or more portions of which are configured with an optical contrast agent; as well as Head-mounted display (HMD) devices, including Imaging components that generate pixel data representing the scene during use; The display component outputs a graphical user interface during use; A power unit, a data storage unit for storing geometric datasets, and a data processing unit operatively interfaced with the imaging unit and the display unit, wherein the data processing unit includes - At least one parallel processing module, configured with multiple data processing threads corresponding to each pixel, is suitable for - Pixel data is filtered using a predetermined value to separate first pixel data from second pixel data, wherein the first pixel data represents one or more portions of the detectable target, and Calculate two-dimensional (2D) pixel coordinate data from the separated first pixel data; as well as The data processing component is also adapted to - By referencing a first geometric dataset representing at least one coordinate system originating from the HMD device, 2D pixel coordinate data is transformed into three-dimensional (3D) coordinate data representing one or more parts of the detectable target. 3D guidance data is generated based on the 3D coordinate data by referencing one or more other geometric datasets, where each geometric dataset represents a corresponding target in the scene. The 3D guidance data is output to the graphical user interface.
15. The system according to claim 14, wherein, The optical contrast agent is an active agent that emits light waves.
16. The system according to claim 14, wherein, The optical contrast agent is a passive agent, and the system further includes an illumination source configured to excite the optical contrast agent.
17. The system according to claim 16, wherein, The HMD device includes the lighting source.
18. The system according to any one of claims 14 to 17, wherein, Each of one or more parts of the detectable target is a marker having a predetermined relative geometric relationship with it.
19. The system according to claim 18, wherein, At least one detectable target is a tool used by or near the HMD wearer, and at least one of the one or more additional geometric datasets includes a three-dimensional model representing the tool; and / or At least one of the detectable targets is a marker defining its location in the scene, and at least one of the one or more additional geometric datasets includes a 3D model representing the marker.
20. The system according to claim 19, wherein, The mark is a matrix barcode, one or more of which are configured with the optical contrast agent.
21. The system according to claim 19 or 20, wherein, The scene includes at least two detectable targets, and the data processing unit is further programmed to generate display data representing a path between the two detectable targets in the scene when the 3D guided display data is generated.
22. A method for guiding a detectable target using a head-mounted display (HMD) device, comprising the following steps: The imaging sensor of the HMD device generates pixel data of a scene, wherein the detectable target is in the scene. The HMD device utilizes at least one parallel processing module, wherein the module is configured with multiple pixel-specific data processing threads. Pixel data is filtered using a predetermined value to separate first pixel data from second pixel data, wherein the first pixel data represents one or more portions of the detectable target configured with an optical contrast agent. Calculate two-dimensional (2D) pixel coordinate data from the separated first pixel data; as well as Using at least one additional processing unit of the HMD device, By referencing a first geometric dataset representing at least one coordinate system originating from the HMD device, 2D pixel coordinate data is transformed into three-dimensional (3D) coordinate data representing one or more portions of the detectable target. 3D guided display data is generated based on the 3D coordinate data by referencing one or more other geometric datasets, each geometric dataset representing a corresponding detectable target in the scene; as well as The 3D guided display data is output to a graphical user interface on at least one display of the HMD device.
23. The method according to claim 22, wherein, The transformation step further includes performing triangulation on the 2D pixel coordinate data with reference to the first geometric dataset.
24. The method according to claim 22, wherein, The transformation step also includes solving for rotation and translation based on the 2D pixel coordinate data.
Citation Information
Patent Citations
Augmented reality headset for medical imaging
GB2588774A