Method for controlling a display, computer program and mixed reality display device

By generating 3D point clouds of treatment subjects and using semantic segmentation and ICP algorithms, the problem of inaccurate registration between virtual information and real objects was solved, achieving high-precision medical imaging data overlay and reducing medical risks.

CN114270408BActive Publication Date: 2026-03-17APCOLLE LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-09
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

In medical applications, existing technologies struggle to achieve precise registration between virtual information and real-world objects, leading to increased illusions and medical risks in mixed reality devices.

Method used

By generating 3D target point clouds and source point clouds of the treatment object, and using semantic segmentation technology to determine the segmentation mask, the transformation matrix is ​​calculated to accurately align the medical imaging data with the real-world view, and the Iterative Closest Point (ICP) algorithm and Convolutional Neural Network (CNN) are used for precise registration.

Benefits of technology

It achieves high-precision alignment between medical imaging data and the treatment subject, reducing medical risks and improving the visualization effect and user experience of mixed reality devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114270408B_ABST
    Figure CN114270408B_ABST
Patent Text Reader

Abstract

A method for controlling a display of a mixed reality display device, wherein a source point cloud and a target point cloud representing a surface of a subject of treatment are generated from image data and medical imaging data of the subject of treatment. A plurality of segmentation masks are determined in the point clouds by applying semantic segmentation. A transformation between the source point cloud and the target point cloud is determined using the segmentation masks, and at least a portion of the medical imaging data is overlaid on the subject of treatment using the determined transformation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This invention relates to a method for controlling a display of a mixed reality device.

[0002] Furthermore, the present invention relates to a computer program having program code methods suitable for performing such methods.

[0003] In addition, the present invention relates to a mixed reality display device having such a computer program.

[0004] Generally speaking, this invention relates to the field of visualization of virtual information integrated with the real environment. Virtual information is overlaid on real-world objects on a display device's screen. This field is commonly referred to as "mixed reality."

[0005] The term "virtual continuum" extends from purely real environments to purely virtual environments, encompassing augmented reality and augmented virtuality between these extremes. The term "mixed reality" is generally defined as "anywhere between the extremes of the virtual continuum," meaning that mixed reality generally includes a fully virtual continuum excluding purely real and purely virtual environments. In the context of this application, the term "mixed reality" may specifically refer to "augmented reality."

[0006] Mixed reality technology holds particular promise for medical applications, such as those used in surgical procedures or other medical treatments. For example, medical imaging data (CT images, MRI images, etc.) visualizing the anatomy and / or physiological processes of a human or animal body can be overlaid onto a real-world view of the body using mixed reality display devices. In this way, for instance, surgeons can gain support during surgery by virtually placing this medical imaging data directly onto the treatment object—that is, onto the patient's body or a part of the patient's body.

[0007] One of the most significant challenges in mixed reality is the registration problem—the proper alignment of objects in the real world with those in the virtual world relative to each other. Without precise registration, the illusion of the two worlds coexisting is compromised. More seriously, in medical applications, inaccurate registration can lead to risks that affect medical success and even patient health. Therefore, it is crucial that visualized virtual information, such as medical imaging data, be precisely matched in position, size, and viewpoint to the real world—the object of treatment.

[0008] From EP 2 874 556 B1, augmented reality-based methods and corresponding systems are known, enabling instrument guidance in surgery and other interventional procedures. To this end, an interventional path used in the interventional procedure is obtained, wherein the path is planned based on 3D image data of the patient's interior, and camera images of the patient's exterior are acquired during the interventional procedure. A spatial correspondence is established between the camera images and the 3D image data, and a view of the interventional path corresponding to the camera images is calculated. Finally, the view of the interventional path is combined with the camera images to obtain a composite image displayed on a monitor.

[0009] The purpose of this invention is to provide an improved technique for visualizing virtual information in medical applications, which enables improved alignment between virtual and real objects.

[0010] The object of the present invention is achieved by a method for controlling a display of a mixed reality display device having the features of claim 1.

[0011] According to the present invention, the method includes at least the following steps:

[0012] a) Provide an image dataset containing multiple images of the treatment subject, wherein the treatment subject is the patient's body or a part of the patient's body, and the images depict the treatment subject from different perspectives.

[0013] b) Generate a 3D target point cloud from the image dataset, wherein the target point cloud includes multiple points defined in a three-dimensional coordinate system, and these points represent the surface of the treatment object.

[0014] c) Determine multiple semantic segmentation masks in the target point cloud by applying semantic segmentation.

[0015] d) Provide a medical imaging dataset that includes medical imaging data of the patients being treated.

[0016] e) Generate a 3D source point cloud from a medical imaging dataset, wherein the source point cloud comprises multiple points defined in a three-dimensional coordinate system, and these points also represent the surface of the object being treated.

[0017] f) Determine multiple semantic segmentation masks in the source point cloud by applying semantic segmentation.

[0018] g) Use the segmentation mask of the source point cloud and the segmentation mask of the target point cloud to determine the transformation between the source point cloud and the target point cloud, and

[0019] h) Visualize at least a portion of medical imaging data on a display, wherein the medical imaging data is overlaid on and aligned with the treatment object using a transformation between the source point cloud and the target point cloud.

[0020] The steps of this method need not be performed in a specified order, and therefore do not limit the invention; that is, the alphabetical order of the letters does not imply a specific order of steps a) to h). For example, steps a) to c) can certainly be performed after steps d) to f), or some steps of the method can be performed in parallel.

[0021] This invention proposes a method for controlling the display of a mixed reality device to overlay medical imaging data onto a treatment subject. Therefore, this invention provides a method for controlling the display of a mixed reality device to visualize medical imaging data on a treatment subject.

[0022] In the context of this application, the term "treatment object" refers to the patient's body or a part of the patient's body. The patient can be a person or an animal, that is, the treatment object can be a human or animal body or a part of a human or animal body.

[0023] The terms "mixed reality display device" and "mixed reality device" are used interchangeably. The term "computer" is used in its broadest sense, referring to any processing device that can be instructed to perform sequences of arithmetic and / or logical operations.

[0024] The term "2D" refers to two-dimensional coordinates. The term "3D" refers to three-dimensional coordinates. The term "4D" refers to four-dimensional coordinates.

[0025] In addition to a display, a mixed reality device may include a computer and memory. A mixed reality device may also include multiple computers. Furthermore, a mixed reality device may include cameras, particularly a 3D camera system, and / or multiple sensors, particularly at least one depth sensor, such as a time-of-flight depth sensor. A mixed reality display device may also include a positioning system and / or an inertial measurement unit.

[0026] In step a), an image dataset comprising multiple images of the treatment subject is provided, wherein the images depict the treatment subject from different perspectives. These images represent real-world views of the treatment subject. They can be generated, for example, by a camera and / or depth sensor, particularly by a camera and / or depth sensor of a mixed reality device.

[0027] In step d), a medical imaging dataset is provided, comprising medical imaging data of the patient being treated. This medical imaging data represents virtual information to be visualized on a display. The medical imaging data may include, for example, cross-sectional images of the patient being treated. For example, medical imaging methods such as magnetic resonance imaging (MRI) can be used to generate the medical imaging data. The medical imaging data may be generated before and / or during medical treatment, such as before and / or during surgery.

[0028] In steps b) and e), a three-dimensional point cloud representing the treatment object is generated. This can be achieved using 3D reconstruction methods.

[0029] In steps c) and f), semantic segmentation is applied to determine multiple segmentation masks in the target point cloud and multiple segmentation masks in the source point cloud, respectively. Semantic segmentation (also known as semantic image segmentation) can be defined as the task of clustering image parts that belong to the same object class together. In the context of this application, an object class can be, for example, a specific part of the anatomical structure of a treatment object. For example, if the treatment object is a human head, simple object classes could include nose, ears, mouth, eyes, eyebrows, etc. The term "semantic segmentation mask" refers to a part (or fragment) of an image that has been determined to belong to the same object class using semantic segmentation. Semantic segmentation can be performed on 2D data, i.e., pixel-based, or on 3D data, i.e., voxel-based.

[0030] In step g), a transformation between the source and target point clouds is determined using segmentation masks of the source and target point clouds. In this step, the transformation can be determined such that when applied to one of the point clouds, the points of the two point clouds are aligned with each other. The transformation can include translation and / or rotation. For example, the transformation can be in the form of a transformation matrix, particularly a 4×4 transformation matrix. The determined transformation can, in particular, approximately transform the source point cloud into the target point cloud, or approximately transform the target point cloud into the source point cloud.

[0031] Besides the segmentation mask, other parameters can also be used as inputs to determine the transformation between point clouds. Specifically, the transformation between the source and target point clouds can be determined using the segmentation masks of the source and target point clouds, as well as the coordinates of points in the source and target point clouds.

[0032] In step h), the medical imaging data is visualized on a display, wherein the medical imaging data is overlaid on and aligned with a real-world view of the treatment subject using the transformation determined in step g). This creates a virtual fusion of the visualized medical imaging data and the real-world view of the treatment subject.

[0033] This invention is based on the discovery that registration problems can be solved more accurately and reliably by using semantic segmentation. This is achieved by determining the transformation between the source and target point clouds based on segmentation masks of the source and target point clouds.

[0034] As an example, the optimization function used to determine the transformation between the source and target point clouds can be designed to facilitate precise matching of the transformations of corresponding semantic segmentation masks in the two point clouds, i.e., precise matching of the transformations of semantic segmentation masks with the same and / or similar object classes (e.g., nose, ears, eyes, eyebrows). This can be achieved, for example, by using a four-dimensional optimization algorithm to determine the transformation between the source and target point clouds, where, for each point in each point cloud, the object class of the corresponding semantic segmentation mask (e.g., nose, ears, mouth) is interpreted as the fourth dimension of the point (in addition to the point's 3D coordinates). For example, a 4D variant of the Iterative Closest Point (ICP) algorithm can be used for this purpose.

[0035] This invention enables highly accurate and reliable alignment of virtual information from medical imaging data with a real-world view of the patient. Therefore, an improved virtual fusion of visualized medical imaging data and a real-world view of the patient can be created, while avoiding risks that could compromise medical success and patient health.

[0036] According to an advantageous embodiment of the invention, it is proposed to design the display as an optically transparent display, particularly an optically transparent head-mounted display.

[0037] These embodiments of the present invention provide users, such as surgeons, with the advantage of a realistic and intuitive perception of their environment.

[0038] According to another advantageous embodiment of the invention, a mixed reality display device is proposed to include a head-mounted mixed reality display device and / or mixed reality smart glasses, or to consist of a head-mounted mixed reality display device and / or mixed reality smart glasses. The mixed reality display device may, for example, include a Microsoft HoloLens device or a Microsoft HoloLens 2 device or a similar device, or to consist of a Microsoft HoloLens device or a Microsoft HoloLens 2 device or a similar device.

[0039] These embodiments of the present invention offer the advantages of ease of use and provide powerful hardware for visualizing virtual information to users. In particular, many head-mounted mixed reality display devices and mixed reality smart glasses include, in addition to the display, powerful and versatile integrated hardware components, including high-performance processors and memory, 3D camera systems and time-of-flight depth sensors, positioning systems and / or inertial measurement units.

[0040] According to another advantageous embodiment of the invention, the mixed reality display device may include an additional external computer, such as an external server, connected to the display and adapted to perform at least a portion of the method proposed according to the invention. The mixed reality display device may, for example, include a head-mounted mixed reality display device and / or mixed reality smart glasses, and an additional computer, such as an external server, connected to the head-mounted mixed reality display device and / or mixed reality smart glasses via wired or wireless connections, respectively. The external server may be designed as a cloud server.

[0041] This implementation offers the advantage of additional computing power for complex and computationally intensive operations that are particularly common in related fields of computer vision and computer graphics.

[0042] According to another advantageous embodiment of the invention, all components of the mixed reality display device can be integrated into a head-mounted mixed reality display device and / or mixed reality smart glasses. This provides the advantage of a compact and therefore highly portable mixed reality display device.

[0043] According to another advantageous embodiment of the invention, it is proposed to use at least one of the following medical imaging methods to generate medical imaging data: magnetic resonance imaging (MRI), computed tomography (CT), cone-beam computed tomography (CBCT), digital volumetric tomography (DVT), intraoperative images with fluoroscopy, X-ray, radiography, ultrasound, endoscopy and / or nuclear medicine imaging.

[0044] This offers the following advantages: during medical treatments such as surgery, the results of powerful and versatile modern medical imaging methods can be advantageously utilized through mixed reality visualization. This enables, for example, visualization of a patient's internal organs and / or tumors and / or other internal defects in the patient's body.

[0045] According to another advantageous embodiment of the invention, a convolutional neural network configured for semantic segmentation is proposed to determine the semantic segmentation mask in the target point cloud and / or the semantic segmentation mask in the source point cloud.

[0046] Significant progress has been made in semantic segmentation using convolutional neural networks (CNNs) in recent years. Properly designed and trained CNNs enable reliable, accurate, and fast semantic segmentation of 2D and even 3D image data. Therefore, embodiments of the present invention using CNNs for semantic segmentation offer the advantage of leveraging the power of CNNs to improve mixed reality visualization. For example, the U-Net CNN architecture can be used for this purpose; that is, the convolutional neural network configured for semantic segmentation can be designed as U-NET CNN.

[0047] A CNN can be trained on semantic segmentation of the body or a specific body part using an appropriate training dataset, which includes semantic segmentation masks labeled with their respective object classes (e.g., nose, ears, eyes, and eyebrows in the case of a training set for the human head).

[0048] According to another advantageous embodiment of the invention, step c) includes the following steps:

[0049] - By applying semantic segmentation to images in an image dataset, specifically using a convolutional neural network configured for semantic segmentation, multiple semantic segmentation masks are determined in the images of the image dataset.

[0050] - Use the semantic segmentation mask in the image from the image dataset to determine the semantic segmentation mask in the target point cloud.

[0051] According to this implementation, a method for determining semantic segmentation masks in images within an image dataset is proposed. The images can be specifically designed as 2D images, particularly 2D RGB images. Based on these semantic segmentation masks in the images, 3D semantic segmentation masks can be determined in a target point cloud. For example, each point in the 3D target point cloud can be projected onto a 2D image to determine a semantic segmentation mask in the 2D image corresponding to the corresponding point in the 3D target point cloud.

[0052] This implementation offers the following advantages: it facilitates semantic segmentation and enables the determination of semantic segmentation masks using sophisticated available methods for semantic segmentation in 2D images, particularly 2D RGB images. For example, powerful convolutional neural networks and corresponding training datasets can be used for semantic segmentation in 2D RGB images. Leveraging their potential enables fast, accurate, and reliable semantic segmentation of 2D images. These results can be transferred to 3D point clouds for use in the claimed method, i.e., for determining the precise transformation between the source and target point clouds.

[0053] According to another advantageous embodiment of the invention, the image dataset comprises multiple visual and / or depth images of the treatment subject, and in particular, a 3D target point cloud is generated from the visual and / or depth images using photogrammetric and / or depth fusion methods.

[0054] Visual images can be generated using cameras, particularly 3D camera systems. These visual images can be designed as, for example, RGB images and / or grayscale images. Depth images can be generated using 3D scanners and / or depth sensors, particularly 3D laser scanners and / or time-of-flight depth sensors. Depth images can be designed as depth maps. A combination of visual and depth images can be designed as RGB-D image data. Therefore, images in an image dataset can be designed as RGB-D images.

[0055] 3D reconstruction methods, particularly active and / or passive 3D reconstruction methods, can be used to generate 3D target point clouds from visual and / or depth images.

[0056] 3D target point clouds can be generated from visual images using photogrammetry. Specifically, 3D target point clouds can be generated from visual images using Structure of Motion (SfM) and / or Multi-View Stereo (MVS) processes. For example, the COLMAP 3D reconstruction pipeline can be used for this purpose.

[0057] Deep fusion can be used to generate a 3D target point cloud from a depth image using 3D reconstruction from multiple depth images. Deep fusion can be based on a truncated signed distance function (TSDF). For example, a point cloud library (PCL) can be used to generate depth images using deep fusion. Kinect Fusion, for instance, can be used as a deep fusion method. Specifically, a Kinect Fusion implementation included in a PCL can be used for this purpose, such as KinFu.

[0058] As described above, these embodiments of the invention, including generating 3D target point clouds from visual images and / or depth images, offer the advantage of enabling accurate and detailed reconstruction of the surface of the treatment object.

[0059] According to another advantageous embodiment of the invention, step c) includes the following steps:

[0060] - By applying semantic segmentation to visual and / or depth images of an image dataset, specifically using a convolutional neural network configured for semantic segmentation, multiple semantic segmentation masks are determined in the visual and / or depth images of the image dataset.

[0061] - Use semantic segmentation masks in visual and / or depth images from an image dataset to determine the semantic segmentation mask in the target point cloud.

[0062] This implementation offers the advantages of facilitating semantic segmentation and enabling the use of sophisticated and readily available semantic segmentation methods to determine semantic segmentation masks in 2D and / or depth images. The semantic segmentation masks obtained from the visual and / or depth images can be transferred to a 3D point cloud for use in the claimed method, specifically for determining the precise transformation between the source and target point clouds. This allows for rapid, accurate, and detailed reconstruction of the surface of the treatment object.

[0063] According to another advantageous embodiment of the invention, the image dataset comprises multiple visual and depth images of the treatment subject, and step b) includes the following steps:

[0064] - Generate the first 3D point cloud from visual images in an image dataset.

[0065] - Generate a second 3D point cloud from depth images in an image dataset, and

[0066] - Use the first 3D point cloud and the second 3D point cloud, especially by merging the first 3D point cloud and the second 3D point cloud, to generate a 3D target point cloud.

[0067] This implementation offers the following advantages: it can improve the accuracy and completeness of 3D reconstruction, and therefore the accuracy and completeness of the resulting target point cloud. This is achieved by merging a first 3D point cloud, which may be an RGB-based point cloud, with a second 3D point cloud, which may be a depth-based point cloud.

[0068] According to another advantageous embodiment of the invention, step f) includes:

[0069] - By applying semantic segmentation to medical imaging data in a medical imaging dataset, specifically using a convolutional neural network configured for semantic segmentation, multiple semantic segmentation masks are determined in the medical imaging data of the medical imaging dataset.

[0070] - Use the semantic segmentation mask in the medical imaging data of the medical imaging dataset to determine the semantic segmentation mask in the source point cloud.

[0071] According to this implementation, a method for determining semantic segmentation masks in medical imaging data within a medical imaging dataset is proposed. The medical imaging data may include 2D medical images, particularly 2D cross-sectional images, and semantic segmentation masks can be determined within these 2D medical images. Based on these semantic segmentation masks in the medical imaging data, a 3D semantic segmentation mask can be determined in a source point cloud. For example, each point in the 3D source point cloud can be projected onto a 2D medical image to determine a semantic segmentation mask in the 2D medical image corresponding to the corresponding point in the 3D source point cloud.

[0072] This implementation offers the advantages of facilitating semantic segmentation and enabling the use of sophisticated and readily available medical semantic segmentation methods to determine semantic segmentation masks in 2D images. For example, powerful convolutional neural networks and corresponding training datasets can be used for medical semantic segmentation in 2D medical images. Examples include the U-Net convolutional neural network architecture. Utilizing the potential of these CNNs enables fast and accurate semantic segmentation of 2D medical images. These results can be transferred to 3D point clouds for use in the claimed method, i.e., for determining the precise transformation between the source and target point clouds.

[0073] Medical imaging data may also include, for example, 3D medical imaging data of a 3D medical imaging model of a treatment subject. 3D medical imaging data can be reconstructed from multiple 2D medical images of the treatment subject, particularly 2D cross-sectional images. These 2D medical images can be generated using medical imaging methods. In this implementation where the medical imaging data also includes 3D medical imaging data, a semantic segmentation mask can also be determined in the 3D medical imaging data. 3D semantic segmentation methods, particularly those based on convolutional neural networks configured for 3D semantic segmentation, can be used for this purpose. Based on the semantic segmentation mask in the 3D medical imaging data, a 3D semantic segmentation mask can be determined in the source point cloud.

[0074] According to another advantageous embodiment of the present invention, it is proposed that:

[0075] Step c) includes: determining a semantic segmentation mask in the target point cloud by directly applying semantic segmentation to the target point cloud, specifically using a convolutional neural network configured for semantic segmentation, and / or

[0076] Step f) includes: determining a semantic segmentation mask in the source point cloud by directly applying semantic segmentation to the source point cloud, in particular by using a convolutional neural network configured for semantic segmentation.

[0077] According to another advantageous embodiment of the invention, the transformation between the source point cloud and the target point cloud in step g) is determined by an Iterative Closest Point (ICP) algorithm, which uses the coordinates of the points in the source point cloud and the segmentation mask of the source point cloud, as well as the coordinates of the points in the target point cloud and the segmentation mask of the target point cloud.

[0078] According to this implementation, an Iterative Closest Point (ICP) algorithm is proposed to determine the transformation between a source point cloud and a target point cloud. Specifically, a 4D variant of the ICP algorithm can be used for this purpose. The coordinates of points in the source and target point clouds, along with a segmentation mask, are used as input to the ICP algorithm. This can be achieved, for example, by using a 4D variant of the ICP algorithm to determine the transformation between the source and target point clouds, where, for each point in the corresponding point cloud, the object class of the corresponding semantic segmentation mask (e.g., nose, ear, mouth) is interpreted as the fourth dimension of the point (in addition to the point's 3D coordinates). As described above, by including a semantic segmentation mask in the ICP algorithm, the ICP algorithm considers not only the spatial information of the points in the point cloud (i.e., their coordinates) but also their semantics, i.e., their meaning.

[0079] Semantic segmentation masks can also be used to determine the initial estimate of the transformation between the source point cloud and the target point cloud in the ICP algorithm, i.e., to determine the initial alignment.

[0080] The advantages provided by the above implementation are: improved guidance for the ICP algorithm to find the optimal transformation between the source and target point clouds. For example, by using semantic segmentation masks to determine the transformation, the ICP algorithm can avoid finding local optima as solutions. Therefore, the accuracy of the transformation can be improved, thereby improving the alignment of medical imaging data with the treatment object.

[0081] According to another advantageous embodiment of the present invention, it is proposed that:

[0082] Step a) includes: using semantic segmentation, specifically using a convolutional neural network configured for semantic segmentation, to remove background and / or other irrelevant parts from the image of the treatment subject, and / or

[0083] or

[0084] - Step d) includes: using semantic segmentation, in particular using a convolutional neural network configured for semantic segmentation, to remove background and / or other irrelevant parts from the medical imaging data of the treatment subject.

[0085] Irrelevant parts can be portions of images or medical imaging data that are irrelevant to the purpose of medical imaging and / or irrelevant to overlaying medical imaging data onto the treatment subject, such as hair on a human head.

[0086] This implementation offers the following advantages: it facilitates the generation of point clouds, the determination of semantic segmentation masks, and the determination of transformations between point clouds.

[0087] According to another advantageous embodiment of the invention, a medical imaging dataset is proposed that includes a 3D medical imaging model of a treatment subject, wherein the 3D medical imaging model is reconstructed from multiple 2D cross-sectional images of the treatment subject generated by a medical imaging method.

[0088] This implementation offers the following advantages: it facilitates the generation of 3D source point clouds while improving the accuracy of the obtained 3D source point clouds.

[0089] According to another advantageous embodiment of the invention, it is proposed that in step a), an image dataset including images of the treatment object is created by means of a camera of a mixed reality display device, particularly a 3D camera system and / or by means of a depth sensor of a mixed reality display device, particularly by means of a time-of-flight depth sensor.

[0090] In this way, the image dataset required to generate the target point cloud can be created and provided for the purposes of this invention in a very simple and user-friendly manner. For example, using a camera and / or depth sensor of a mixed reality display device, images of the image dataset can be created by automatically scanning the treatment subject. In particular, if the mixed reality display device includes or consists of a head-mounted mixed reality display device and / or mixed reality smart glasses, the camera and / or depth sensor can be integrated into the head-mounted mixed reality display device and / or mixed reality smart glasses, respectively. In this case, the user can simply point their head at the part of the treatment subject to be scanned.

[0091] According to another advantageous embodiment of the present invention, it is proposed that:

[0092] - In step a), when images are created using the camera of the mixed reality display device, the camera position is determined in a three-dimensional coordinate system for each image and the camera position is stored as a 3D camera position, and in step b), a target point cloud is generated using the 3D camera position, and / or

[0093] or

[0094] - In step a), when an image is created using the camera of the mixed reality display device, the camera orientation is determined in a three-dimensional coordinate system for each image and the camera orientation is stored as a 3D camera orientation, and in step b), the target point cloud is generated using the 3D camera orientation.

[0095] This implementation offers the advantage of enabling the generation of target point clouds from image datasets with particularly high accuracy and reliability.

[0096] According to another advantageous embodiment of the invention, it is proposed that any position, any orientation, and any transformation determined in steps a) to h) is determined without markers and / or using Simultaneous Localization and Mapping (SLAM).

[0097] This implementation offers the following advantages: it is particularly comfortable for the user because there is no need to provide optical markers or other reference markers.

[0098] The object of the present invention is also achieved by a computer program having program code units adapted to perform the method described above when the computer program is executed on a computer.

[0099] The object of the present invention is also achieved by a mixed reality display device having a display, a computer and a memory, wherein the aforementioned computer program is stored in the memory and the computer is adapted to execute the computer program.

[0100] Computer programs can be designed as distributed computer programs. Computers and memory can be designed as distributed computers and distributed memories, respectively; that is, the computers suitable for executing the computer program can include two or more computers. Distributed memory can include multiple memories, each of which can store at least a portion of the distributed computer program, and each of the two or more computers can be adapted to execute a portion of the distributed computer program.

[0101] Furthermore, the mixed reality display device may have an interface suitable for receiving medical imaging data from an external source. This interface may be adapted to connect, for example, to medical imaging equipment, such as magnetic resonance imaging (MRI) equipment and / or computed tomography (CT) equipment, and / or to a memory storing the medical imaging data. This may include wired and / or wireless connections.

[0102] As described above, a mixed reality display device may include a head-mounted mixed reality display device and / or mixed reality smart glasses, or may consist of a head-mounted mixed reality display device and / or mixed reality smart glasses. A mixed reality display device may, for example, include a Microsoft HoloLens device or a Microsoft HoloLens 2 device or a similar device, or may consist of a Microsoft HoloLens device or a Microsoft HoloLens 2 device or a similar device.

[0103] As described above, a mixed reality display device may include an additional external computer, such as an external server, connected to the display and adapted to perform at least a portion of the methods proposed according to the invention. The mixed reality display device may, for example, include a head-mounted mixed reality display device and / or mixed reality smart glasses, and an additional computer, such as an external server, connected to the head-mounted mixed reality display device and / or mixed reality smart glasses via wired and / or wireless connections, respectively. The external server may be designed as a cloud server.

[0104] As mentioned above, all components of a mixed reality display device can be integrated into a head-mounted mixed reality display device and / or mixed reality smart glasses.

[0105] The invention will be explained in more detail below using exemplary embodiments schematically illustrated in the accompanying drawings. The drawings are shown below:

[0106] Figure 1 - A schematic diagram of a mixed reality display device according to the present invention;

[0107] Figure 2 - A schematic diagram of a method for controlling a mixed reality display device according to the present invention;

[0108] Figure 3- A schematic diagram of a 3D target point cloud with a semantic segmentation mask;

[0109] Figure 4 - A schematic diagram of a 3D source point cloud with a semantic segmentation mask.

[0110] Figure 1 A schematic diagram of a mixed reality display device 1 is shown, which includes a head-mounted mixed reality display device in the form of mixed reality smart glasses 5. In this example embodiment, the mixed reality smart glasses 5 is of the Microsoft HoloLens type. The smart glasses 5 has a memory 13a and a computer 11a connected to the memory 13a, wherein the computer 11a includes several processing units, namely a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and an HPU (Holographic Processing Unit).

[0111] Furthermore, the smart glasses 5 of the mixed reality display device 1 have a camera 9 in the form of a 3D camera system. The camera 9 is adapted to create visual images of the treatment subject 15 from different perspectives. Additionally, the smart glasses 5 include multiple sensors 7, including a time-of-flight depth sensor adapted to create depth images of the treatment subject 15 from different perspectives. In this example embodiment, the treatment subject 15 is the head of a human patient 17.

[0112] Additionally, the smart glasses 5 of the mixed reality display device 1 have a display 3, which in this example embodiment is designed as an optical see-through head-mounted display. The see-through display 3 is adapted to visualize virtual information, in this example embodiment, medical imaging data, by overlaying virtual information onto the real view of the treatment subject 15.

[0113] also, Figure 1 The mixed reality display device 1 shown includes a server 21, which is connected to smart glasses 5 via a wireless and / or wired connection. The server 21 includes a memory 13b and a computer 11b connected to the memory 13b, wherein the computer 11b includes a CPU and a GPU. The server 21 has an interface 19 adapted to receive medical imaging data from an external source, i.e., from the memory storing the medical imaging data, via a wired and / or wireless connection.

[0114] Figure 2 A schematic diagram of an example method for controlling a mixed reality display device according to the present invention is shown.

[0115] In step a), an image dataset including multiple images of the treatment subject 15 is provided, wherein the images depict the treatment subject from different perspectives. In this example embodiment, the image dataset includes multiple visual images in the form of 2D RGB images and multiple depth images in the form of depth maps. The visual images are created by the camera 9 of the mixed reality display device 1, and the depth images are created by the depth sensor 7 of the mixed reality display device 1 (see...). Figure 1 In other words, this means creating an image dataset containing RGB-D image data. For this purpose, a user wearing smart glasses 5, such as a surgeon, can simply point their head at the treatment subject 15 from different perspectives, and the treatment subject 15 is automatically scanned from these different perspectives using camera 9 and depth sensor 7. When creating images, the position and orientation of camera 9 are determined in a three-dimensional coordinate system for each image, and these are stored separately as 3D camera position and 3D camera orientation. Simultaneous localization and mapping (SLAM) is used for this. The 3D camera position and 3D camera orientation of each image can be referred to as external camera parameters. The image dataset including the external camera parameters is then transferred from smart glasses 5 to server 21.

[0116] In step b), server 21 generates a 3D target point cloud from the image dataset, wherein the target point cloud includes multiple points defined in a three-dimensional coordinate system, and these points represent the surface of the treatment object 15. For this purpose, external camera parameters, namely 3D camera position and 3D camera orientation, are used.

[0117] A 3D target point cloud is generated from visual and depth images. To this end, a photogrammetric method is used, specifically a COLMAP 3D reconstruction pipeline that includes a Structure of Motion (SfM) process and a Multi-View Stereo (MVS) process, to generate a first 3D point cloud from visual images in an image dataset. Furthermore, depth fusion is used, i.e., 3D reconstruction from multiple depth images, to generate a second 3D point cloud from depth images in the image dataset. For this purpose, a Point Cloud Library (PCL) is used, specifically KinectFusion implemented within the PCL. Finally, the first and second 3D point clouds are merged to generate the 3D target point cloud.

[0118] In step c), multiple semantic segmentation masks in the target point cloud are determined by applying semantic segmentation.

[0119] First, semantic segmentation masks are determined in the 2D visual images of the image dataset by applying semantic segmentation to these 2D RGB images. Each semantic segmentation mask defines an object class for each pixel of each 2D visual image, where the object class may include object classes such as "Other," "Blank," or similar object classes for image regions that cannot be matched in other ways. To determine the semantic segmentation masks in the 2D visual images, a convolutional neural network (CNN) configured for semantic segmentation, specifically a CNN based on the U-Net CNN architecture (U-Net CNN), is used. The U-Net CNN is trained using an appropriate training dataset for semantic segmentation of the treatment object 15, specifically for the semantic segmentation of the human head. This appropriate training dataset includes semantic segmentation masks labeled with their respective object classes (e.g., nose, ears, eyes, eyebrows).

[0120] Secondly, semantic segmentation masks in the 3D target point cloud are determined using previously determined semantic segmentation masks in the 2D visual images of the image dataset. To do this, each point in the 3D target point cloud is projected onto multiple 2D RGB images to determine the semantic segmentation mask in the corresponding 2D image for each point in the 3D target point cloud.

[0121] Figure 3 The results of determining semantic segmentation masks in the target point cloud are illustrated schematically. The figure shows a 3D target point cloud 23 generated from an image dataset. Several semantic segmentation masks 25a, 25b, and 25c have been determined in the 3D target point cloud 23. Two semantic segmentation masks 25a represent the patient's right and left eyes, respectively. Another semantic segmentation mask 25b represents the patient's nose, and another semantic segmentation mask 25c represents the patient's mouth.

[0122] Now return to the reference. Figure 2 In step d), a medical imaging dataset including medical imaging data of the treatment subject is provided. In this example embodiment, the medical imaging data includes multiple 2D cross-sectional images of the treatment subject 15 created by magnetic resonance imaging (MRI) prior to surgery. These 2D MRI images in DICOM data format are received by server 21 via server interface 19. Server 21 reconstructs a 3D medical imaging model of the treatment subject 15 from the multiple 2D MRI images of the treatment subject 15.

[0123] Medical imaging data may also include metadata, for example, stored as tags in DICOM data in the form of attributes. Examples of such metadata include information about the slice thickness of a cross-sectional image and / or information about pixel spacing. This metadata can be used to reconstruct 3D imaging models and / or to generate 3D source point clouds.

[0124] In step e), a 3D source point cloud is generated from a medical imaging dataset. The source point cloud includes multiple points defined in a three-dimensional coordinate system, and these points also represent the surface of the treatment object 15. For this purpose, a point cloud library (PCL) is used.

[0125] In step f), multiple semantic segmentation masks are determined in the source point cloud by applying semantic segmentation.

[0126] First, semantic segmentation masks are determined in the 2D MRI images of medical imaging datasets by applying semantic segmentation to these datasets. Each semantic segmentation mask defines an object class for each pixel of each 2D MRI image, where the object class may include object classes such as "Other," "Blank," or similar object classes for image regions that cannot be matched in other ways. To determine the semantic segmentation masks in the 2D MRI images, a convolutional neural network (CNN) configured for semantic segmentation, specifically a CNN based on the U-Net CNN architecture (U-Net CNN), is used. The U-Net CNN is trained using an appropriate training dataset for semantic segmentation of treatment subject 15, specifically for the semantic segmentation of the human head. This appropriate training dataset includes semantic segmentation masks labeled with their respective object classes (e.g., nose, ears, eyes, eyebrows).

[0127] Secondly, semantic segmentation masks in the 3D source point cloud are determined using previously determined semantic segmentation masks in 2D MRI images from a medical imaging dataset.

[0128] Figure 4 The results of determining semantic segmentation masks in the source point cloud are illustrated schematically. The figure shows a 3D source point cloud 27 generated from a medical imaging dataset. Several semantic segmentation masks 29a, 29b, and 29c have been determined in the 3D source point cloud 27. Two semantic segmentation masks 29a represent the patient's right and left eyes, respectively. Another semantic segmentation mask 29b represents the patient's nose, and another semantic segmentation mask 29c represents the patient's mouth.

[0129] Now return to the reference. Figure 2 In step g), the segmentation masks of the source point cloud and the target point cloud are used to determine the transformation between the source and target point clouds. In this example implementation, a 4×4 transformation matrix including translation and rotation is determined, and when this 4×4 transformation matrix is ​​applied to the source point cloud, the points of the source point cloud are aligned with the points of the target point cloud. Therefore, the determined 4×4 transformation matrix is ​​used for the purpose of transforming between different viewpoints (positions and orientations) of the acquired source and target point clouds.

[0130] In this example implementation, a 4D variant of the Iterative Closest Point (ICP) algorithm is used to determine the 4×4 transformation matrix. The coordinates of points in the source and target point clouds and their segmentation masks are used as input to the algorithm, where, for each point in the respective point cloud, the object class of the corresponding semantic segmentation mask (e.g., nose, ear, mouth) is interpreted as a fourth dimension of the point (in addition to the point's 3D coordinates). The optimization function for the 4D ICP is designed to support exact matching of transformations between the corresponding semantic segmentation masks (nose and nose, ear and ear, mouth and mouth, etc.) in the two point clouds.

[0131] The 4×4 transformation matrix and medical imaging data are transmitted from server 21 to mixed reality smart glasses 5.

[0132] In step h), at least a portion of the medical imaging data, namely MRI imaging data, is visualized on the optical transillumination display 3 of the smart glasses 5 (see...). Figure 1 In this process, medical imaging data is overlaid on the real-world view of the treatment subject 15 and aligned with the treatment subject 15 using a transformation between the source point cloud and the target point cloud, i.e., using the 4×4 transformation matrix determined in step g).

[0133] In this way, surgeons using the mixed reality display device 1 can be supported during surgery by virtually visualizing the anatomical structure of the treatment subject 15 as shown in MRI images, which is precisely aligned with the real-world view of the treatment subject 15.

[0134] In this example implementation, steps a) and h) are performed by the mixed reality smart glasses 5, while steps b) through g) are performed by the server 21. In other implementations, additional or all steps of the method may be performed by the mixed reality smart glasses 5.

[0135] List of reference numerals

[0136] 1 Mixed Reality Display Device

[0137] 3. Monitors

[0138] 5 Mixed Reality Smart Glasses

[0139] 7 sensors

[0140] 9 cameras

[0141] 11a, 11b Computers

[0142] 13a and 13b memory

[0143] 15. Treatment subjects

[0144] 17 patients

[0145] 19 Interfaces

[0146] 21 servers

[0147] 23 Target Point Cloud

[0148] Segmentation masks for target point clouds 25a, 25b, and 25c

[0149] 27 Source Cloud

[0150] Segmentation masks for source point clouds 29a, 29b, and 29c

Claims

1. A method for controlling a display (3) of a mixed reality display device (1), the method comprising at least the following steps: a) providing an image data set comprising a plurality of images of a treatment object (15), wherein the treatment object (15) is a patient body or a part of the patient body and the images depict the treatment object (15) from different perspectives, b) generating a 3D target point cloud (23) from the image data set, wherein the target point cloud (23) comprises a plurality of points defined in a three-dimensional coordinate system and the points represent a surface of the treatment object, c) determining a plurality of semantic segmentation masks (25a, 25b, 25c) in the target point cloud (23) by applying semantic segmentation, d) providing a medical imaging data set comprising medical imaging data of the treatment object (15), e) generating a 3D source point cloud (27) from the medical imaging data set, wherein the source point cloud (27) comprises a plurality of points defined in a three-dimensional coordinate system and the points also represent a surface of the treatment object, f) determining a plurality of semantic segmentation masks (29a, 29b, 29c) in the source point cloud (27) by applying semantic segmentation, g) determining a transformation between the source point cloud (27) and the target point cloud (23) using the semantic segmentation masks (29a, 29b, 29c) of the source point cloud (27) and the semantic segmentation masks (25a, 25b, 25c) of the target point cloud (23), and h) visualizing at least a portion of the medical imaging data on the display (3), wherein the medical imaging data is overlaid on and aligned with the treatment object (15) using the transformation between the source point cloud (27) and the target point cloud (23), wherein step c) comprises: - determining a plurality of semantic segmentation masks in the images of the image data set by applying semantic segmentation to the images of the image data set, and - determining the semantic segmentation masks (25a, 25b, 25c) in the target point cloud (23) using the semantic segmentation masks in the images of the image data set, and / or step f) comprises: - determining a plurality of semantic segmentation masks in the medical imaging data of the medical imaging data set by applying semantic segmentation to the medical imaging data of the medical imaging data set, and - determining the semantic segmentation masks (29a, 29b, 29c) in the source point cloud (27) using the semantic segmentation masks in the medical imaging data of the medical imaging data set, wherein each semantic segmentation mask refers to segments of the images that have been determined to belong to the same object class using semantic segmentation, and wherein the semantic segmentation masks are used to determine an initial estimate of the transformation of an iterative closest point, ICP, algorithm. The display (3) is designed as an optical see-through display (3), and / or the mixed reality display device (1) comprises a head-mounted mixed reality display device (1). ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ 2. The method of claim 1, wherein, ​ 3. The method according to claim 1 or 2, characterized in that, The medical imaging data is generated using at least one of the following medical imaging methods: magnetic resonance imaging, MRI, computed tomography, CT, radiography, ultrasonography, endoscopy, and / or nuclear medicine imaging.

4. The method according to claim 1 or 2, characterized in that, determining semantic segmentation masks (25a, 25b, 25c) in the target point cloud (23) and / or semantic segmentation masks (29a, 29b, 29c) in the source point cloud (27) using a convolutional neural network configured for semantic segmentation.

5. The method according to claim 1 or 2, characterized in that, In step c), determining semantic segmentation masks in images of the image dataset using a convolutional neural network configured for semantic segmentation.

6. The method of claim 1 or 2, wherein, The image dataset comprises a plurality of visual images and / or depth images of the therapy object (15), and the 3D target point cloud (23) is generated from the visual images and / or the depth images.

7. The method of claim 6, wherein, Step c) comprises the following steps: - determining a plurality of semantic segmentation masks in visual images and / or depth images of the image dataset by applying semantic segmentation to the visual images and / or the depth images of the image dataset, and - determining semantic segmentation masks (25a, 25b, 25c) in the target point cloud (23) using the semantic segmentation masks in the visual images and / or the depth images of the image dataset.

8. The method of claim 1 or 2, wherein, The image dataset comprises a plurality of visual images and depth images of the therapy object (15), and step b) comprises the following steps: - generating a first 3D point cloud from the visual images of the image dataset, - generating a second 3D point cloud from the depth images of the image dataset, and - generating the 3D target point cloud (23) using the first 3D point cloud and the second 3D point cloud.

9. The method of claim 1 or 2, wherein, In step f), determining a plurality of semantic segmentation masks in the medical imaging data of the medical imaging dataset using a convolutional neural network configured for semantic segmentation.

10. The method according to claim 1 or 2, characterized in that: - step c) comprises determining semantic segmentation masks (25a, 25b, 25c) in the target point cloud (23) by directly applying semantic segmentation to the target point cloud (23), and / or - step f) comprises determining semantic segmentation masks (29a, 29b, 29c) in the source point cloud (27) by directly applying semantic segmentation to the source point cloud (27).

11. The method of claim 1 or 2, wherein, The transformation between the source point cloud (27) and the target point cloud (23) is determined by an iterative closest point, ICP, algorithm using coordinates of points of the source point cloud (27) and semantic segmentation masks (29a, 29b, 29c) of the source point cloud (27) and coordinates of points of the target point cloud (23) and semantic segmentation masks (25a, 25b, 25c) of the target point cloud (23).

12. The method according to claim 1 or 2, characterized in that: - step a) comprises removing background and / or other irrelevant parts from the images of the therapy object (15) using semantic segmentation, and / or - step a) comprises removing background and / or other irrelevant parts from the images of the therapy object (15) using semantic segmentation, and / or - step d) comprises removing background and / or other irrelevant parts from the medical imaging data of the therapy subject (15) using semantic segmentation.

13. The method of claim 1 or 2, wherein, The medical imaging data set comprises a 3D medical imaging model of the therapy subject (15), wherein the 3D medical imaging model is reconstructed from a plurality of 2D cross-sectional images of the therapy subject (15) generated by a medical imaging method.

14. The method of claim 1 or 2, wherein, In step a), the image data set comprising images of the therapy subject (15) is created by a camera (9) of the mixed reality display device (1) and / or by a depth sensor (7) of the mixed reality display device (1).

15. The method according to claim 14, wherein: - in step a), when the images are created by a camera (9) of the mixed reality display device (1), a position of the camera (9) is determined in a three-dimensional coordinate system for each image and stored as a 3D camera position, and in step b), the target point cloud (23) is generated using the 3D camera positions, and / or - in step a), when the images are created by a camera (9) of the mixed reality display device (1), an orientation of the camera (9) is determined in a three-dimensional coordinate system for each image and stored as a 3D camera orientation, and in step b), the target point cloud (23) is generated using the 3D camera orientations.

16. The method of claim 1 or 2, wherein, Any position, any orientation and any transformation determined in steps a) to h) is markerless determined and / or determined using simultaneous localization and mapping, SLAM.

17. The method of claim 2, wherein, The optical see-through display (3) is an optical see-through head-mounted display (3), and / or the head-mounted mixed reality display device (1) is a mixed reality smart glass (5).

18. The method of claim 6, wherein, The 3D target point cloud (23) is generated from the visual images and / or the depth images using photogrammetry methods and / or depth fusion methods.

19. The method of claim 7, wherein, A plurality of semantic segmentation masks in the visual images and / or the depth images of the image data set is determined using a convolutional neural network configured for semantic segmentation.

20. The method of claim 8, wherein, The 3D target point cloud (23) is generated by merging the first 3D point cloud and the second 3D point cloud.

21. The method according to claim 10, wherein: - step c) comprises determining a semantic segmentation mask (25a, 25b, 25c) in the target point cloud (23) using a convolutional neural network configured for semantic segmentation, and / or - step f) comprises determining a semantic segmentation mask (29a, 29b, 29c) in the source point cloud (27) using a convolutional neural network configured for semantic segmentation.

22. The method according to claim 12, wherein: - step a) comprises removing background and / or other irrelevant parts from the images of the therapy subject (15) using a convolutional neural network configured for semantic segmentation, and / or - step d) comprises removing background and / or other irrelevant parts from the medical imaging data of the subject (15) using a convolutional neural network configured for semantic segmentation.

23. The method of claim 14, wherein, The camera (9) is a 3D camera system and / or the depth sensor (7) is a time-of-flight depth sensor (7).

24. Computer program product having program code means adapted to perform the method according to any one of the preceding claims when executed on a computer (11a, 11b).

25. A mixed reality display device (1) having a display (3), a computer (11a, 11b) and a memory (13a, 13b), wherein, A computer program product is stored in the memory (13a, 13b) and performs the method according to any one of claims 1 to 23 when executed on the computer (11a, 11b).

Citation Information

Patent Citations

  • Augmented reality imaging system for surgical instrument guidance

    EP2874556B1

  • Augmenting real-time views of a patient with three-dimensional data

    US9892564B1