Robust processing of multimodal point cloud data of one or more maxillofacial structures
The deep neural network-based method for registering and fusing multimodal maxillofacial point cloud data addresses alignment and fusion challenges, ensuring high accuracy and resolution by determining geometric features and using clustering and stitching algorithms, enhancing dental treatment planning precision.
Patent Information
- Application Number
- PCT/EP2025/061605
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-26
- Filing Date
- 2025-04-28
- Publication Date
- 2025-10-30
AI Technical Summary
Existing methods for registering and fusing multimodal point cloud data of maxillofacial structures, such as CBCT and IOS data, face challenges in achieving robust, accurate, and efficient alignment and fusion, particularly due to sensitivity to partial field of views, low accuracy, and loss of resolution in iterative approaches like ICP and Euclidean distance-based mesh separation.
A method utilizing a deep neural network to determine geometric features for registration, followed by a clustering algorithm to establish correspondences and a transformation estimation algorithm for precise alignment, and a stitching algorithm for fusion, ensuring high accuracy and maintaining mesh integrity without iterative schemes.
The method provides robust, accurate, and efficient registration and fusion of point cloud modalities, maintaining high anatomical feature frequencies and resolution, outperforming conventional methods by achieving precise alignment and seamless fusion of multimodal data.
Smart Images

Figure EP2025061605_30102025_PF_FP_ABST
Abstract
Description
[0001] Robust processing of multimodal point cloud data of one or more maxillofacial structures
[0002] Technical field
[0003] The embodiments relate to robust processing, such as registration and / or fusion, of multimodal point cloud data representing one or more maxillofacial structures, and, in particular, though not exclusively, to methods and systems for robust registration and / or fusion of multimodal point cloud data representing one or more maxillofacial structures and a computer program product for executing such methods.
[0004] Background
[0005] Acquiring an accurate digital representation of the maxillofacial region of a patient is important for 3D computer-assisted dental applications such as dental implant planning, orthodontic treatment planning and orthognathic surgery. The creation of a digital 3D model of (part of) a dentition or virtual patient anatomy, relies on the acquisition of multiple data types acquired using different data acquisition systems resulting in different data modalities representing the maxillofacial regions (or a part thereof) of a patient.
[0006] Exemplary image acquisition protocols include 2D and 3D cone-beam computed tomography (CBCT) imaging, optical (OS) or intra-oral scanning (IOS), 2D panoramic x-ray images and 3D magnetic resonance imaging (MRI) images. Each of these modalities have known limitations. For example, CBCT imaging is a low-radiation acquisition system that has inconsistent image value with Hounsfield Units, poor contrast, lower resolution (0.15-0.3 mm) and is sensitive to metal. In contrast, IOS offers high resolution (5- 15 micron) of the patient anatomy but does not contain typically any information beyond the visible surface, in particular no information on tooth roots, jawbone, or nerves.
[0007] Combing different point cloud modalities such as CBCT and IOS point clouds of a maxillofacial structure or two IOS point clouds representing different parts of a maxillofacial structure can provide enhanced maxillofacial models. For example, as described in US2014 / 0169648 combining IOS and CBCT may enrich the anatomical digital model of the patient, providing more information to the clinical team to plan and perform the treatment more precisely through higher image quality. However, building a digital model based on different modalities requires accurate alignment (also known as registration) and fusion of the data, which are both non-trivial problems. Chung et al. describes in their article Automatic registration between dental cone-beam CT and scanned surface via deep pose regression neural networks and clustered similarities, IEEE Transactions on Medical Imaging Vol: 39, Issue: 12, December 2020 a way to address this problem by exploring deep pose regression neural networks applied in a reduced 2D domain to register IOS and CBCT 3D data. Their approach relies on a two-step registration, first a coarse alignment is performed by projecting both the CBCT and IOS scans to 2D views to extract and align the front teeth point and a vector pointing to the posterior direction following the teeth arch, second multiple point clusters are defined on the surfaces of the two modalities (segmented bone on CBCT) and used to optimized local transformations. A subsequent refined registration is performed iteratively until the best clusters are used to estimate the final transformation parameters. This approach shows improved alignment accuracy compared to the well-known conventional iterative closest points (ICP) methods but is not robust with respect to partial field of views and still report low accuracy with regards to sub millimeter registration requirements for image-guided dental treatment.
[0008] Fusing registered CBCT data representing a complete tooth-bone structures with high-resolution IOS data representing tooth crowns could provide a complete crown-root structure with the accuracy required for clinical applications. Liu. et al. described in their article Deep learning-enabled 3D multimodal fusion of cone-beam CT and intraoral mesh scans for clinically applicable tooth-bone reconstruction, an automatic multimodal framework for reconstructing 3D tooth-bone structures using CBCT and IOS. Qian et al. proposes in their paper An automatic tooth reconstruction method based on multimodal data a tooth reconstruction method based on the CBCT data and a corresponding laser-scanned crown mesh.
[0009] The primary challenges in fusion include preserving the topology and accuracy of tooth models from different modalities, avoiding modifications to the existing models, accurately splitting and merging tooth models from various modalities, creating smooth transitions between modalities, ensuring the mesh integrity, and producing aesthetically pleasing results. Both Liu et al. and Qian et al. employ a method based on a Euclidean distance for mesh separation that proves inaccurate because it requires setting a high threshold to ensure proper separation of the crown and root in the CBCT segmentations. This leads to a large gap between the IOS mesh and the CBCT root that requires an additional remeshing step, which removes the original resolution and details of the IOS modality. This holds significant importance for the IOS scan and it represents the foremost requirement emphasized by the clinical experts.
[0010] Hence, from the above it follows there is a need in the art for improved methods and systems for robust registration and / or fusion of point cloud modalities of a maxillofacial structure.
[0011] Summary As will be appreciated by one skilled in the art, aspects of the present invention may be embodied as a system, method or computer program product. Accordingly, aspects of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, micro-code, etc.) or an embodiment combining software and hardware aspects that may all generally be referred to herein as a "circuit," "module" or "system." Functions described in this disclosure may be implemented as an algorithm executed by a microprocessor of a computer. Furthermore, aspects of the present invention may take the form of a computer program product embodied in one or more computer readable medium(s) having computer readable program code embodied, e.g., stored, thereon.
[0012] Any combination of one or more computer readable medium(s) may be utilized. The computer readable medium may be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium would include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer readable storage medium may be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0013] A computer readable signal medium may include a propagated data signal with computer readable program code embodied therein, for example, in baseband or as part of a carrier wave. Such a propagated signal may take any of a variety of forms, including, but not limited to, electro-magnetic, optical, or any suitable combination thereof. A computer readable signal medium may be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0014] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber, cable, RF, etc., or any suitable combination of the foregoing. Computer program code for carrying out operations for aspects of the present invention may be written in any combination of one or more programming languages, including an object-oriented programming language such as Java(TM), Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0015] Aspects of the present invention are described below with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor, in particular a microprocessor or central processing unit (CPU), of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer, other programmable data processing apparatus, or other devices create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0016] These computer program instructions may also be stored in a computer readable medium that can direct a computer, other programmable data processing apparatus, or other devices to function in a particular manner, such that the instructions stored in the computer readable medium produce an article of manufacture including instructions which implement the function / act specified in the flowchart and / or block diagram block or blocks.
[0017] The computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. Additionally, the Instructions may be executed by any type of processors, including but not limited to one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FP- GAs), or other equivalent integrated or discrete logic circuitry. The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.
[0018] In an aspect, the embodiments in this disclosure may relate to a method of processing point cloud modalities of one or more maxillofacial structures wherein the method may comprise: providing a first point cloud modality of a first maxillofacial structure to an input of a deep neural network, which is trained to determine first geometric features for registration of the first point cloud modality; providing a second point cloud modality of a second maxillofacial structure to the input of the deep neural network to determine second geometric features for registration of the second point cloud modality; using a clustering algorithm to determine a set of correspondences between the first point cloud modality and the second point cloud modality based on the first and second geometric features respectively; and, using a transformation estimation algorithm and the set of correspondences to determine information about a coordinate transformation for registering the first point cloud modality with the second point cloud modality.
[0019] Hence, point cloud modalities may be registered based on geometric features of the point cloud modalities which are determined by a trained deep neural network and which are relevant for registration of the point cloud modalities. The method provides a robust, accurate and efficient way of registering point cloud modalities, which can be used for registering partially overlapping point clouds and which do not need an iterative scheme as known from the prior art.
[0020] In an embodiment, the method may further include: registering the first point cloud modality with the second point cloud modality based on the coordinate transformation.
[0021] In an embodiment, the set of correspondences may include point-to-point correspondences between subsets of points in the first and second point cloud modality, the point-to-point correspondences including pairs of points C(j,j) wherein a point i in the first point cloud modality has a corresponding point j in the second point cloud modality.
[0022] In an embodiment, the set of correspondences may include soft point-to-point correspondences between subsets of points in the first and second point cloud modality, the point-to-point correspondences including pairs of points wherein a point i in the first point cloud modality has a correspondence within a predefined distance of a corresponding point j in a the second point cloud modality.
[0023] In an embodiment, the deep neural network may be trained to determine for each point of the first and second point cloud modality a plurality of first and second feature vectors respectively, wherein a feature vector represent geometric features for a point of a point cloud modality.
[0024] In an embodiment, the geometric features may include local geometric features associated with an individual tooth in the maxillofacial structure, preferably the local geometric features including information about the orientation, position, shape and / or type of an individual tooth; and / or, wherein the global geometric features are associated with two or more teeth in the maxillofacial structure, preferably the global geometric features including about the dentition position (quadrant), orientation and / or neighbors.
[0025] In an embodiment, points of the first and / or second point cloud modality may be associated with one or more pre-defined attributes, preferably the one or more predefined attributes , for example a tooth type according to the FDI dental numbering system, a local curvature, a point normal, an implant, a bracket or an artificial crown.
[0026] In an embodiment, the first cloud modality may be generated by a first data acquisition system and the second point cloud modality is generated by a second data acquisition system that is different from the first data acquisition system, preferably the first and second data acquisition system being at one of: an inter-oral scanner (IOS), an X-ray scanner, a computed tomography (CT) system or a magnetic resonance imager (MRI) system.
[0027] In an embodiment the first point cloud modality may represent a first part of a dentition of a patient and the second point cloud modality may represent a second part of the dentition that partially overlaps the first part of the dentition.
[0028] In an embodiment, the first point cloud modality may represent a part of a dentition of a first patient and the second point cloud modality may represent a part of the dentition of a second patient, wherein the set of correspondences may be based on soft point-to-point correspondences.
[0029] In an embodiment, the deep neural network may be configured to receive all points of a cloud modality and, optionally, one or more pre-defined attributes associated with each point of the point cloud modality at its input; and / or, wherein the deep neural network is configured to output a dense set of geometrical features, preferably the dense set of geometrical features, including for a feature vector for each point of the point cloud modality.
[0030] In an embodiment, the deep neural network may be a fully convolutional and wherein the fully convolutional network is configured to receive all points of a cloud modality and, optionally, one or more pre-defined attributes associated with each point of the point cloud modality at its input and to apply generalized sparse convolution computations on each point of the point cloud modality, and, optionally, the one or more pre-defined attributes associated with each point of the point cloud modality.
[0031] In an embodiment, the first and second point cloud modalities may represent a first and second modality mesh respectively.
[0032] In an embodiment, the method may further comprise fusing the first modality with a separated part of the second modality, the fusing including: determining a separated part of the second modality mesh, the determining including removing portions of the second modality mesh that at least partially overlap with the first modality mesh by projecting surface normals of the second modality mesh in the direction of the first modality mesh; determining intersection of the projected surface normal with the first modality mesh; and, using a stitching algorithm to fuse the first modality mesh with the separated part of the second modality mesh.
[0033] In an embodiment, the removing of portions of the second modality may further comprise: projecting surface normals of the second modality mesh onto the surface of the first modality mesh in a direction along the surface normals, each of the surface normals being associated with a part of the surface of the second modality mesh; determining if the projection of one of the surface normals intersects the surface of the first modality mesh; and, removing the part of the surface of the second modality mesh that is associated with the projected surface normal, if it is determined that the projected surface normal intersects the surface of the first modality mesh.
[0034] The fusion method maintains the topology of the modality with the highest accuracy and / or highest anatomical feature frequencies. Moreover it is very fast compared to approaches such as Poisson reconstruction as used in the prior art.
[0035] In a further aspect, the embodiments may relate to a method of fusing at least two point cloud modalities of one or more maxillofacial structures wherein the method may comprise: registering a first point cloud modality with a second point cloud modality, the first and second point cloud modalities representing a first and second modality mesh respectively; determining a separated part of the second modality mesh, the determining including removing portions of the second modality mesh that at least partially overlap with the first modality mesh by projecting surface normals of the second modality mesh in the direction of the first modality mesh and determining intersection of the projected surface normal with the first modality mesh; and, using a stitching algorithm to fuse the first modality mesh with the separated part of the second modality mesh.
[0036] In an embodiment the first and second point cloud modalities may be registered using the method as described with reference to the embodiments in this disclosure.
[0037] In an embodiment, the stitching algorithm may be configured to: meshing edges of the first modality mesh and the separated part of the second modality; and, using one or more filtering operations applied to the seam region and to the resulting fused structure.
[0038] In another aspect, the embodiments in this disclosure may relate to a method of training a deep neural network to determine geometric features for registration of point cloud modalities of a maxillofacial structure, the method comprising: providing a first point cloud modality of a first maxillofacial structure and, optionally one or more pre-defined attributes, to a data augmentation unit, which is using transformations to augment the first point cloud modality into a first augmented first point cloud modality; providing a second point cloud modality of a second maxillofacial structure and, optionally, one or more pre-defined attributes, to the data augmentation unit to augment the second point cloud modality into a second augmented first point cloud modality; using a deep neural network to determine first and second geometrical features associated with the augmented first and second point cloud modality respectively; and, using the first and second geometrical features and predetermined point-to-point correspondences between points of the first point cloud modality and points of the second point cloud modality, to optimize parameters of the deep neural network based on a loss function to train the deep neural network to generate geometrical features for registration of point cloud modalities.
[0039] In an embodiment, the optimization of the network parameters may be based on a contrastive learning scheme in which the loss function is used to minimize a distance between the first and second geometrical features that correspond with point-to-point correspondences and to maximize a distance between the first and second geometrical features that do not correspond to the point-to-point correspondences.
[0040] In an embodiment, the transformations of the data augmentation unit may include one or more of: random cropping, translations, rotations, scaling, noise, tooth removal, preferably the tooth removal being based on the predefined-features of the first point cloud modality and / or the predefined-features of the second point cloud modality. In yet another aspect, the embodiments may relate to a system for processing point cloud modalities of a maxillofacial structure comprising: a computer readable storage medium having at least part of a program embodied therewith; and, a computer readable storage medium having computer readable program code embodied therewith, and a processor, preferably a microprocessor, coupled to the computer readable storage medium, wherein responsive to executing the computer readable program code, the processor is configured to perform executable operations comprising: providing a first point cloud modality of a maxillofacial structure to an input of a deep neural network, which is trained to determine first geometric features for registration of the first point cloud modality; providing a second point cloud modality of at least part of the maxillofacial structure that at least partially overlaps the first point cloud modality to an input of the deep neural network to determine second geometric features for registration of the second point cloud modality; using a clustering algorithm to determine a set of correspondences between the first point cloud modality and the second point cloud modality based on the first and second geometric features respectively; and, using a transformation estimation algorithm to determine information about a coordinate transformation for registering the first point cloud modality with the second point cloud modality based on the set of correspondences.
[0041] In a further embodiment, the embodiments may relate to a system for fusing at least two point cloud modalities of one or more maxillofacial structures comprising: a computer readable storage medium having at least part of a program embodied therewith; and, a computer readable storage medium having computer readable program code embodied therewith, and a processor, preferably a microprocessor, coupled to the computer readable storage medium, wherein responsive to executing the computer readable program code, the processor is configured to perform executable operations comprising: registering a first point cloud modality with a second point cloud modality, the first and second point cloud modalities representing a first and second modality mesh respectively, determining a separated part of the second modality mesh, the determining including removing portions of the second modality mesh that at least partially overlap with the first modality mesh by projecting surface normals of the second modality mesh in the direction of the first modality mesh and determining intersection of the projected surface normal with the first modality mesh; and, using a stitching algorithm to fuse the first modality mesh with the separated part of the second modality mesh.
[0042] The invention may also relate to a computer program product comprising software code portions configured for, when run in the memory of a computer, executing the method steps according to any of process steps described above. The invention will be further illustrated with reference to the attached drawings, which schematically will show embodiments according to the invention. It will be understood that the invention is not in any way restricted to these specific embodiments.
[0043] Brief description of the drawings
[0044] Fig. 1A and 1B illustrate registration examples of registered point cloud modalities of a maxillofacial structure using a conventional registration method.
[0045] Fig. 2 illustrates a system for automatic registering point cloud modalities of a maxillofacial structure according to an embodiment;
[0046] Fig. 3 depicts a flow chart of a method of automatic registering point cloud modalities of a maxillofacial structure according to an embodiment;
[0047] Fig. 4 illustrates a method for training a system to automatically determine geometrical features for registering point cloud modalities of a maxillofacial structure according to an embodiment;
[0048] Fig. 5 depicts a flow chart of training a model for generating geometrical features that are relevant for registering point cloud modalities according to an embodiment;
[0049] Fig. 6A-6C depict an example of a registration of two different point cloud modalities of a maxillofacial structure using a registration method according to the embodiments in this disclosure;
[0050] Fig. 7A-7D depict geometrical features associated with different point cloud modalities according to an embodiment;
[0051] Fig. 8A depicts an example of registered point cloud modalities of a maxillofacial structure using a registration method according to an embodiment;
[0052] Fig. 8B depicts static occlusion correspondences on two shifted point cloud to highlight the correspondence points;
[0053] Fig. 8C depicts an example a first lower IOS point cloud and a second, upper IOS point aligned on a CBCT scan of a dentition according to an embodiment;
[0054] Fig. 8D depicts an example of registered point cloud modalities of a maxillofacial structure for a static occlusion using a registration method according to an embodiment;
[0055] Fig. 9 depicts a flow chart of a method for automatic fusion of point cloud modalities of a maxillofacial structure according to an embodiment;
[0056] Fig. 10A and 10B depict an example of a fusion of two registered point cloud modalities of a maxillofacial structure using a fusion method according to an embodiment; Fig. 10C depicts an example of the method used to separate parts of the registered mesh modalities of a maxillofacial structure using the normal projection method according to an embodiment;
[0057] Fig. 10D depicts an example of the method used to remove portions of a second modality mesh by projecting surface normals onto a first modality mesh.
[0058] Description of the embodiments
[0059] Conventional registration methods that rely on an iterative approach, such as the well-known iterative closest point (ICP) method, may converge into a local minimum. Moreover, these methods are sensitive to different data modalities (e.g. different field of views, modalities, missing objects, artefacts, etc.) which often only partially overlap. Fig. 1A and 1B illustrate examples of registration of point cloud modalities representing a maxillofacial structure using a conventional registration method, in this case the well-known iterative closest point (ICP) method. Fig.lA depicts the registration between CBCT data and IOS data, showing bad posterior alignment on the upper jaw (102). Fig. 1B depicts an example of a registration between CBCT data and IOS data showing bad vertical alignment and posterior rotation on the upper jaw (103).
[0060] The embodiments in this disclosure aim to address these problems by providing methods and systems for automatically determining a transformation for accurate and robust registration (alignment) of at least two point cloud modalities of a maxillofacial structure (e.g. a point cloud generated by computed tomography and a point cloud generated by an optical scanner). The determination of the transformation parameters (which may represent a rigid, affine, or non-linear transformation) may include determining geometrical features for points of a point cloud modality using a trained 3D artificial neural network and determining point-to-point correspondences between points of the point cloud modalities based on the geometrical features. In some embodiment, points of the point cloud modalities may also be associated with pre-defined attributes, including but not limited to a tooth type (for example a tooth type according to the FDI dental numbering system (based on the World Dental Federation notation)), a local curvature, a point normal (direction associated with each point), other labels to indicate an implant, a bracket or an artificial crown. The pre-defined attributes may be use for training the 3D ANN to generate accurate 3D geometries features.
[0061] Fig. 2 illustrates a system for automatic registering different point cloud modalities of a digital maxillofacial structure according to an embodiment. In particular, the figure illustrates a system 202 configured to estimate transformation parameters for aligning two different 3D data modalities, e.g. a point cloud representation of a CBCT image and an IOS point cloud of a digital maxillofacial structure to the same canonical space.
[0062] The system may include at least one trained 3D artificially neural network 206 (ANN) including at least one input to receive input data representing a digital maxillofacial structure. Here, the ANN may be trained to receive multimodal data including different point cloud modalities of the digital maxillofacial structure generated by different data acquisition systems, e.g. intra-oral scanner, a CBCT scanner, etc. Input data may be represented as point cloud, e.g. a triangulated point cloud including coordinates of points in a space and associated attributes. In an embodiment, a point may be associated with additional predefined attributes e.g. color, intensity, an FDI value of the FDI dental numbering system, point normal, etc. This allows representing the anatomy of a maxillofacial structure with high precision compared to a conventional voxel presentation. The different digital representations may be based on different coordinate system associated with different data acquisition systems.
[0063] The ANN may be trained to determine 3D geometric features for a multimodal point cloud representation of a digital maxillofacial structure that is offered to the input of ANN. In an embodiment, the ANN may be trained to extract local geometrical features, e.g. geometrical features associated with individual tooth, and global geometric features, e.g. features associated with a set of tooth of the maxillofacial structure.
[0064] To handle large points cloud data sets, in which information is sparsely distributed (i.e. only at the points), a point cloud may be represented by a sparse tensor including a set of coordinates and associated features and convolutions may be computed using the generalized sparse convolution as described in the articles by Choy et al, Minkowski convolutional neural networks, and Fully convolutional geometric features, both published in CVPR 2019, which are hereby incorporated by the reference in this disclosure.
[0065] Hence, in an embodiment, the ANN may be implemented as a 3D fully convolutional network for sparse point clouds in which a point cloud modality that is provided to the input of the fully convolutional network may be represented by a sparse tensor and convolutions executed by the fully convolutional network may be computed using the generalized sparse convolution. Such 3D fully convolutional network may be configured to receive the entire set of points of a point cloud modality at its input and to apply generalized sparse convolutions on the set of points of a point cloud, without the need for cropping or the like. Such network can accurately learn both local and global 3D geometric features of a point cloud modality, as they produce a dense output unlike the non-fully convolutional networks, which is important for accurate alignment of point cloud modalities of a maxillofacial structure. The fully convolutional neural network may be trained to determine fully convolutional geometric features of a point cloud modality in which each point of the point cloud modality may be associated with a feature vector. The fully-convolutional features are locally correlated and are not considered independent and identically distributed. In an embodiment, the ANN may be implemented as Res-U-net, i.e. a U-net with residual connections as for example described in Choy et al, or potentially as other neural network that outputs dense features. For example, Zheng Qin et al. describe in their article Geometric Transformer for Fast and Robust Point Cloud Registration a KPConvFPN backbone to extract lower resolution geometric features and then leverage transformers to encode the global structure of the geometry. In another approach, Liang Pan et al. describe in their article PointAtrousGraph: Deep Hierarchical Encoder-Decoder with Point Atrous Convolution for Unorganized 3D Points point atrous convolution layers to encode both local and global features of a point-cloud.
[0066] The 3D ANN may be trained to learn relevant 3D geometric features for a point cloud modality of a maxillofacial structure points of the point cloud modality may be associated with additional pre-defined attributes 208 and 220. In an embodiment, the FDI dental numbering may be obtained using a semantic segmentation network (not shown) that is configured to receive raw point cloud data of a maxillofacial structure at its input and to output a label points in the point cloud representing a tooth with one or more attributes, e.g. a tooth numbering such as the FDI tooth number.
[0067] Here, relevant 3D geometric features may refer to features of a maxillofacial structure that are relevant for estimating a point-to-point correspondence between two point cloud modalities. In an embodiment, the ANN may be trained based on a contrastive learning scheme. In such scheme, geometrical features of a maxillofacial structure of known point-to- point correspondences (defining a ground-truth) are minimized during the training of the ANN, while geometrical features estimated for points that do not belong to known, groundtruth correspondences are maximized during training. This way, for each point of a point cloud modality (or at least a substantial part thereof), the ANN may determine a plurality of 3D geometrical features. These features may be represented as a vector, wherein each element of the vector may represent a value associated with a certain geometrical feature that is associated with a geometric property on the local-global continuum. For instance, one feature may represent the closeness to a cusp (a local feature), while another feature may represent distance from the left upper wisdom tooth (a global feature).
[0068] The 3D ANN may determine first 3D geometric features 210i for a first point cloud modality 208i of the digital maxillofacial structure and second 3D geometric features 2102 a second modality 2082Of the digital maxillofacial structure. The maxillofacial structure may relate to a raw CBCT or IOS scan or segmented data, e.g. a tooth, obtained by segmenting a CBCT or IOS scan. Examples of 3D geometrical features generated by the trained ANN are shown in Fig. 7A-7D. In particular, Fig. 7B shows information about a first geometric feature associated with a first point cloud modality indicating for each point of the point cloud a value representing a local orientation per tooth, (here the local orientation may be encoded in a grayscale value). Hence, this 3D geometric feature defines for each point of the point cloud that is provided to the input of the ANN, a value representing a local orientation. In a similar way, Fig. 7A shows information about a second geometric feature indicating for each point of the point cloud a global dentition quadrant or orientation (e.g. posterior / inferior direction). In a similar way, Fig. 7D and 7C illustrate geometrical features for a second point cloud modality. Many different types of geometric features may be defined by the learnt fully-convolutional features of the 3D ANN so that each point may be associated with different values (different grayscale values), each value representing a particular geometric feature.
[0069] A clustering algorithm 212 may be configured to receive the first and second 3D geometrical features. In some embodiment, the clustering algorithm may also be configured to receive pre-defined attributes 220i, 2 associated with the first and second point cloud modalities. The first and second 3D geometrical features and, in some embodiments, the pre-defined attributes, may be used by the clustering algorithm to determine point-to- point correspondences 214 between points of the two different point cloud modalities. In an embodiment, a nearest neighbors algorithm, such a k-nearest neighbors algorithm, may be used to group 3D geometrical features from the two modalities by comparing features of the first modality with features of the second modality using a metric, such as the squared L2 norm.
[0070] Hence, each point of the first and second point cloud modality may be associated with a plurality of 3D geometrical feature values, where each 3D geometrical feature value may represent an indication how relevant the point is for this particular feature. Based on the 3D geometrical feature values and, in some embodiment, the pre-defined attributes, the clustering algorithm may determine for a point in the first point cloud modality, a corresponding point in the second cloud modality using a minimum distance metric. This way, multiple point-to-point correspondences may be determined by the clustering algorithm between points of the first and second point cloud modality.
[0071] A transformation estimation algorithm 218 may be used to determine transformation parameters for registering the two data sets. The transformation parameters may include a transformation matrix, which may be estimated using a linear system that penalizes the distances between the point cloud correspondences. A transformation matrix may be optimized by minimization of the global distances of point-to-point correspondences that were determined by the clustering algorithm.
[0072] In an embodiment, the transformation estimation optimization step as used in registration algorithm as described in the article by Zhou et al., "Fast global registration." Computer Vision-ECCV 2016, may be used. For example, Zhou et al. proposes to optimize the transformation T such that the distances between the estimated correspondences K={(p, q)} are minimized, while spurious correspondences are seamlessly ignored: which represents the optimal transformation T that reduces the distances between the estimated correspondences (p, q). The term p( ) represents a robust penalty to reduce the impact of spurious correspondences in the transformation estimation.
[0073] It is submitted that that the embodiments are not limited to the implementation as described with reference to Fig. 2. For example, different neural network architectures which are capable of determining 3D geometrical features for point cloud modalities may be used. Different clustering algorithms may be used to group point coordinates between different modalities based on a similarity between high-dimensional features (e.g. ball query, Hungarian method). Further, different transformation matrix estimation algorithms may be used that can compute a transformation matrix based on a set of corresponding points. Such algorithms may include the Umeyama method.. As will be shown hereunder in more detail, embodiments provide fast, robust, and accurate registration of point cloud modalities, which does not require an iterative scheme relying on metric computation during inference.
[0074] Fig. 3 depicts a flowchart of a method of automatic registering point cloud modalities of a maxillofacial structures according to an embodiment. In particular, the figure illustrates a registration method for different point cloud modalities that is associated with the same or at least partially the same maxillofacial structure. The method may include providing a first point cloud modality of a maxillofacial structure to an input of a deep neural network, which is trained to determine first geometric features associated with the first point cloud modality (step 302). Here, the first point cloud modality may be associated with one or more first attributes, such as a tooth type (for example a tooth type according to the FDI dental numbering system based on the World Dental Federation notation), a local curvature, a point normal that indicates the orientation of the surface at each point other labels to indicate an implant, a bracket or an artificial crown. As described with reference to Fig. 2, in some embodiments, the deep neural network may be a fully convolutional network that can process sparse point cloud modalities, wherein a point cloud modality that is provided to the input of the fully convolutional network may be represented by a sparse tensor and wherein convolutions executed by the fully convolutional network may be computed using the generalized sparse convolution. The fully convolutional neural network may be trained to determine fully convolutional geometric features of a point cloud modality in which each point of the point cloud modality may be associated with a feature vector. The 3D fully convolutional network may be configured to receive the entire set of points of the first point cloud modality at its input and to apply generalized sparse convolution on the set of points of the first point cloud, without the need for cropping or the like. This way, the network can accurately determine both local and global 3D geometric features of the point cloud modality.
[0075] In a similar way, a second point cloud modality of at least part of the maxillofacial structure may be provided to an input of the deep neural network to determine second geometric features associated with the second point cloud modality (step 304), wherein the second point modality may be associated with one or more second attributes. Then, a clustering algorithm may be used to determine a set of correspondences, preferably point-to-point correspondences, between the first and the second point cloud modality based on the first and second geometric features respectively and, optionally, based on the predefined attributes associated with the first and second point cloud (step 306). Finally, a transformation estimation algorithm may be used to determine information about a coordinate transformation (step 308) for registering the first point cloud modality with the second point cloud modality based on the set of correspondences.
[0076] The registration scheme provides robust and accurate registration of two point cloud modalities. It allows the use of point clouds of the raw data without any preprocessing, e.g. segmentation of the structures of interest. The registered point clouds may be used in a subsequent fusion process to form a fused maxillofacial structure which matches the anatomy of the maxillofacial structure with high precision compared to prior art scheme.
[0077] Fig. 4 depicts a system for training an artificial neural network (ANN) to automatically determine geometric features which are relevant for registration of point cloud modalities representing a maxillofacial structure according to an embodiment. As shown in the figure, the system may include a data generation module 402 and a model optimization module 403. The data generation module may include one or more data storages for storing different point cloud modalities of different maxillofacial structures and may also include additional pre-defined attributes associated with each point in the point cloud modalities. For example, first and second point cloud modalities 404i,2may represent partly overlapping 3D point clouds of a maxillofacial structure that may correspond to the same patient or to different patients. Here, partially overlapping point clouds may relate to point clouds from different sources or viewpoints that exhibit some degree of spatial coincidence but do not completely align with each other. Partial overlap of point clouds of a maxillofacial structure may occur due to factors such as different viewpoints, occlusions, variations in geometry, or the presence of objects with similar shapes. For example, in dentistry, the alignment of the bite (static occlusion) can be estimated by utilizing two point clouds representing the mandible and maxilla which represent two partially overlapping point clouds.
[0078] Further, the data storage may include information about point-to-point correspondences 410 of points of different point cloud modalities, including the first and second point cloud modalities 404I,2 and their associated pre-defined attributes. The predefined attributes may be obtained using a semantic segmentation network that uses as input the raw point cloud data and outputs a label for each point in the point cloud corresponding to one or more attributes that are relevant for registration, e.g. a tooth type, such as an FDI number. Additionally, and / or alternatively, the one or more attributes may be determined in advance as hand-crafted features. These point-to-point correspondences may be determined upfront using a suitable correspondence algorithm.
[0079] In an embodiment, point-to-point correspondences may define pairs of points C(i ') indicating that point i in the first point cloud modality has a corresponding point j in the second point cloud modality. In another embodiment, point-to-point correspondence may be also defined as soft correspondence, when a point I in the first point cloud modality has a correspondence within a predefined distance of corresponding point j in a the second point cloud modality. The point-to-point correspondences may define a ground truth for correspondences between at least two point cloud modalities of a maxillofacial structure that the system needs to determine during training (as explained with reference to Fig. 4). A correspondence pair is logically associated with the same geometrical feature of the maxillofacial structure. In some embodiments, for some partially overlapping maxillofacial structures, the correspondence pair may be associated with geometrical features that have soft point-to-point correspondences. Hence, a plurality of correspondence pairs may be used to t the ANN to determine 3D geometrical features that are relevant for registration of points clouds representing a maxillofacial structure.
[0080] The data generation module may further include data augmentation unit 406I,2 for generating one or more augmented versions of a point cloud modality based on random, but known transformations, which may include cropping, translations, and / or rotations, tooth removal based on the FDI numbering. This way, based on the first and second point cloud modalities (having a known point-to-point correspondence) multiple augmented versions 408I,2 of the point cloud modalities may be generated and be used by the model optimization module to train an artificial neural network (ANN) 412. When training the ANN based on the augmented versions of the first and second point cloud modalities, the trained ANN will be more robust against variations in the input data. A sufficiently large set of training data may be generated by repeating this process for different point cloud modalities of different maxillofacial structures.
[0081] The pairs of digital models and the point-to-point correspondence may be used by the model optimization module 403 to train the ANN to determine 3D geometric features of a sparse point cloud modality representing a maxillofacial structure that is offered to the input of the ANN. As described with reference to Fig. 2, the ANN may be implemented as a fully convolutional network which is capable of efficiently processing sparse point cloud modalities and output dense geometrical features. In particular, the fully convolutional network may be configured to receive the full point cloud, e.g. in the form of a spare tensor, at its input and to apply 3D convolutions, e.g. generalized sparse convolutions, directly onto the point cloud modality that is offered to the input of the ANN. This way, the ANN may be trained to determine local and global 3D geometrical features for points of the point cloud modality.
[0082] In an embodiment, a contrastive learning method may be used to train the ANN to determine local and global 3D geometrical features for determining point-to-point correspondences. The objective of the contrastive learning scheme is encoding features of the input data, i.e. a point cloud modality, in a latent space wherein features are defined similar if they belong to known correspondences, while dissimilar features, defined as not being part of the ground-truth correspondences are penalized in the latent space by maximizing the distance between them. A loss function may be defined to achieve this objective. In an embodiment, the contrastive loss function is implemented using a method, inspired by Choy et al. in Fully convolutional geometric features, published in CVPR 2019 and Vassileios Balntas et al. in Learning local feature descriptors with triplets and shallow convolutional neural networks published in British Machine Vision Conference 2016.
[0083] Hence, during training a first point cloud modality 408i may be provided to the input of the ANN for determining first geometric features 414i and a second point cloud modality 408a may be provided to the input of the ANN to determine second geometric features 414s at its output. The ANN estimates for each point of a point cloud modality a plurality of feature vectors. Each feature vector represent 3D geometric features for a point of the point cloud modality. The ANN may be trained based on a loss function, which may include a positive loss Plossdefined according to the following expression:
[0084] Pioss = d^ Flj) (1) where d represents the £2 norm between the geometrical features estimated by the ANN on the basis of the first and second point cloud modality (F0 and Fl) for all features (FOj, Fly) extracted over the ground-truth correspondence (i,j). The loss function may also include a negative loss Nlosswhich may be defined as:
[0085] Nloss= max(0,M — d F0k, Fl^ (2) where d represents the £2 norm between the estimated features but for a subset of V correspondences (k, I) that do not have ground-truth correspondences. Here, M represents the margin distance to enforce feature dissimilarity between points without ground-truth correspondences. A final loss value may be defined as the summation of the positive and weighted negative losses: loss Pioss d- ^loss (3)
[0086] This bidirectional (i.e. correspondences are the same regardless of the registration directions) contrastive loss 416 may then be used in a backward gradient optimization scheme 418, computing gradients 420 of the loss function with respect to the weights of the ANN for modifying 422 the weights of the ANN until the loss is minimized.
[0087] Fig. 5 depicts a flow chart of a method for training a model for generating 3D geometrical features that are relevant for registering point cloud modalities according to an embodiment. The deep neural network may be trained to determine 3D geometrical features of a point cloud modality representing a maxillofacial structure that is augmented using different transformations, which may include cropping, translations, and / or rotations and / or modifications such as tooth removal based on the FDI numbering (step 502). Preferably, the transformation may be randomly selected from a known set of transformation.
[0088] In a similar way, a second point cloud modality of at least part of the maxillofacial structure that is augmented using different (randomly selected) transformations may be provided to an input of the deep neural network to determine second geometric features associated with the second point cloud modality (step 504). The deep neural network may be used to extract geometrical features that are relevant for registration of the the augmented first and the second point cloud modalities (step 506). An optimization algorithm and predetermined point-to-point correspondences between points of the first point cloud modality and the second point cloud modality may be used ot modify the parameters of the deep neural network so that the deep neural network is trained to generate geometrical features for registration of point cloud modalities (step 508). In an embodiment, a contrastive learning scheme may be used to train the deep neural network.
[0089] Fig. 6A-6C depict an example of a registration of two different point cloud modalities of a maxillofacial structure using a registration method according to the embodiments in this disclosure. In particular, Fig. 6A illustrates IOS data of crowns of a dentition representing a first point cloud modality 602 of a maxillofacial structure and Fig. 6B depicts CBCT data of teeth of the dentition representing a second point cloud modality 604 of the maxillofacial structure. These point cloud modalities are not aligned due to the different acquisition protocols that are used to acquire the point cloud modalities. Fig. 6C depicts the registered first and second point cloud modalities based on a registration method as described with reference to the embodiments in this application. The figure shows that accurate registration can be achieved between the different point cloud modalities.
[0090] Fig. 8A-8D illustrate examples of registering point cloud modalities using the registration schemes as described with reference to the embodiments in this disclosure. Fig. 8A depicts the registration of IOS point clouds onto a CBCT maxillofacial structure of a dentition. As shown in the picture, the IOS point clouds including points representing crowns and gingiva are registered with the CBCT point cloud representing teeth and maxillofacial bones. Fig. 8B depicts static occlusion correspondences on two shifted point cloud (both IOS data) to highlight the correspondence points. In particular, this figures depicts a soft point-to- point correspondences between two IOS point clouds corresponding to mandible and maxilla. The two point clouds are expected to be partially overlapping and the soft point-to- point correspondences are used in the registration process to estimate the bite of the patient between upper and lower dentition anatomies. The thin lines between the teeth of the mandible and corresponding teeth of the maxilla illustrate one-to-one soft correspondences between the two point clouds. Fig. 8C depicts an example a first lower IOS point cloud and a second, upper IOS point aligned on a CBCT scan of a dentition according to an embodiment. As shown in the figure, the IOS data also include points representing the gingiva. Although the gingiva is not a structure present in the CBCT scans, the gingiva is correctly propagated in the IOS point clouds to the CBCT anatomy. Fig. 8D depicts an example of the registration of two partially overlapping IOS point clouds representing mandible and maxilla to estimate the static occlusion of the patient. Fig. 9 depicts a system for fusing two point cloud modalities representing a maxillofacial structure according to an embodiment. As shown in the figure, the system may include a registration module 922, a separation module 924 and a stitching module 926. The registration module may include one or more data storages 902I,2 for storing different point cloud modalities of different maxillofacial structures. Additionally, it may also include one or more data storages 920I,2 for storing pre-defined attributes associated with each point in the point cloud modalities. For example, first and second point cloud modalities 902I,2 may represent partly overlapping 3D representations of a maxillofacial structure that may correspond to the same patient or to different patients. In an embodiment, the first point cloud modality may represent an IOS point cloud representing teeth of the dentition and the second point cloud modality may represent a CBCT tooth or a CBCT maxillofacial structure comprising teeth of a dentition . In other embodiment, the point cloud modalities may include two associated IOS point clouds representing different parts of a maxillofacial structure.
[0091] The registration module may include a registration algorithm 904 which is configured to register point cloud modalities, e.g. first and second point cloud modalities 902I,2. Different registrations algorithms may be used to achieve registration (alignment) of two point clouds. In an embodiment, a point cloud registration scheme as described with reference to Fig. 2 and 3 may be used. Registered point cloud modalities, e.g. first and second registered point cloud modalities 906I,2. and their corresponding meshes and surface normals 910I,2.
[0092] The point cloud registration module may also include a meshing module 908 for transforming a point cloud into a mesh if a mesh is not available for the input point clouds. Here, a mesh representation of a point cloud may define a collection of vertices, edges and / or surfaces that define the shape of the maxillofacial structure. Known mesh representations may be used including but not limited a polygon mesh, a face-vertex mesh, winged-edge mesh, half-edge mesh, quad-edge meshes or corner-table mesh. The meshing module may further include one or more surface reconstruction algorithm based on, for example, Delaunay triangulations, alpha shapes, Voronoi diagrams or Poisson reconstruction for obtaining mesh modalities with smooth surfaces. The different mesh modalities, e.g. a first and second modality mesh, may be stored on a data storage 910I,2.
[0093] After registration relevant parts of two associated modality meshes may be fused into an enhanced mesh representation of a maxillofacial structure. For example, a part of a second mesh modality, e.g. a crown of a CBCT mesh of a tooth, may be replaced by the first modality mesh, e.g. an IOS mesh of the crown. In that case, the part of the second modality mesh should be removed. For example, the part of a CBCT mesh that represents a crown of the tooth should be removed so that it can be replaced by an IOS mesh of the same crown. This removal should be performed in such a way that the edges of the meshes that need to be fused together can be accurately attached to each other without substantially changing the shape of the IOS mesh of the crown, i.e. without changing the clinical information that is embedded in the shape of the IOS mesh of the crown.
[0094] To that end, the first modality mesh, e.g. a IOS crown mesh, that is registered with the second modality mesh, e.g. a CBCT tooth mesh, may serve as inputs to the separation module 924. The separation module may employ a collision algorithm 912, which may be configured to determine surface normals at vertex position of the second modality mesh. For example, in an embodiment, for each vertex, edge and / or face, of the mesh a surface normal may be determined. The surface normals of the second modality mesh may be projected in positive and / or negative direction onto the first modality mesh. Surface parts (defined by vertices, edges and faces) of the second modality mesh for which the surface normals of the second modality mesh intersect with the first modality mesh may be removed. The use of the direction of the surface normals for determining which parts of the CBCT tooth mesh need to be removed may be regarded as a raytracing process using the direction of the surface normals.
[0095] This process may be repeated until all surface parts of the second modality mesh that have a normal that is intersecting with the surface of the first modality are removed. This way, the separation module removes part of the second modality mesh surface that overlaps with the first modality mesh. The technique of using surface normals to separate a part of a mesh into a separated mesh part without resampling and / or change position of points so that separated mesh part can be accurately merged with another mesh. The separation module 924 may further include a data storage for storing the separated mesh parts 914.
[0096] A stitching module 926 may be employed to fuse a first modality mesh with a separated mesh part of the second modality. The stitching module may be configured to execute steps including cleaning and repairing 916 the mesh of the first modality and the mesh of the separated part of the second modality around the region where the two meshes are fused. This region may be referred to as the seam region, representing a transitional area where the meshes of the first modality and the separated part of the second modality are fused. The cleaning and repairing may include techniques such as described in the article by Liepa, Filling holes with meshes, Eurographics Symposium on Geometry Processing(2003).
[0097] In embodiment, the seam region may be triangulated (meshed) 918 to ensure a smooth meshing between the two modalities. For example, the case of fusing the IOS crown mesh with a CBCT root mesh boundaries of the crown from the IOS mesh and the root from the CBCT mesh may be triangulated. In an embodiment, a graph algorithm may be used that defines triangles and / or polygons from the mesh boundaries. Polygons may be further split into triangles. The output for this process is a mesh including the IOS crown, CBCT root and the seam region connecting the IOS crown and CBCT root.
[0098] In a further embodiment, one or more filtering steps 920 may be applied to the seam region and the resulting fused structure of the first and second modalities. For example, in an embodiment, a smoothing technique, such as Laplacian smoothing, may be used to reshape the seam region to fit the correct topology of the tooth. In particular, Laplacian smoothing method may be used to further smooth the seam mesh. In an embodiment, during smoothing some point displacement may be allowed to ensure topology consistency with the context. The context from the seam point of view is the topology of the IOS crown and CBCT root meshes. The Laplacian regularization ensures that the seam mesh is deformed and smoothed in a way that follows the context topology.
[0099] The fusion method as described with reference to Fig. 9 provides many advantageous compared to known fusion methods, including but not limited to keeping the topology of the modality with the highest accuracy and / or highest anatomical feature frequencies. The fusion scheme is fast contrary to approaches such as Poisson reconstruction as is often used in the prior art.
[0100] The extended seam region may include more or less margin for each modality, e.g. 0.25 and 2.0 mm for the IOS and CBCT margins, respectively. This allows for user interactivity to change the algorithm to allow for more or less aesthetics and it allows for flexibility and context in the seam region to be regularized to follow the topology of the boundary points, i.e. points outside the margin from the crown and root meshes.
[0101] While embodiments described with reference to Fig. 9 are illustrated based on IOS and CBCT mesh modalities, the embodiments are not limited thereto. The method may be used for any modality that can be used to transfer context, e.g. additional clinical information, from one modality to another.
[0102] Fig. 10 depicts an example of the fusion results following the process explained in Fig.9 when using a first point cloud modality corresponding to an IOS crown 1002 and a second point modality corresponding to a CBCT tooth 1004. In this example, the collision algorithm applied to IOS crown and CBCT tooth is shown in Fig. 10A, 1006. The projected rays represent the surface normals of the CBCT tooth that may collide with the registered IOS mesh. The separated CBCT root mesh 1008 is obtained by removing the CBCT mesh surfaces where the CBCT surface normals collide with the IOS mesh. The seam region between the separated CBCT tooth and IOS crown is obtained via the stitching algorithm 1010 and the final fused tooth 1012 is shown in Fig. 10B. Fig. 10C depicts the collision algorithm applied to a first modality mesh that corresponds to an IOS crown (dark grey) and the second modality mesh that corresponds to the CBCT tooth mesh (light grey) that were registered in advance using a registration method. The rays correspond to the projected surface normals of the CBCT tooth mesh.
[0103] Fig. 10D depicts an example of the method used to remove portions of the second modality mesh 1016 by projecting surface normals 1020, 10181,2 of the second modality mesh onto the surface of the first modality mesh 1014 in both positive and negative directions. Each of the surface normals is associated with a part of the surface of the second modality mesh. The method determines if the projection of each surface normal intersects the surface of the first modality mesh; in this example normals 10181,2 intersect the first modality surface whereas normal 1020 does not. The method further comprises removing the part of the second modality mesh 1022 that is associated with the projected surface normals that intersects the first modality mesh.
[0104] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (I C) or a set of ICs (e.g., a chip set). Various components, modules, or units are described in this disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but do not necessarily require realization by different hardware units. Rather, as described above, various units may be combined in a codec hardware unit or provided by a collection of interoperative hardware units, including one or more processors as described above, in conjunction with suitable software and / or firmware.
[0105] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0106] The corresponding structures, materials, acts, and equivalents of all means or step plus function elements in the claims below are intended to include any structure, material, or act for performing the function in combination with other claimed elements as specifically claimed. The description of the present invention has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the invention in the form disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the invention. The embodiment was chosen and described in order to best explain the principles of the invention and the practical application, and to enable others of ordinary skill in the art to understand the invention for various embodiments with various modifications as are suited to the particular use contemplated.
Claims
CLAIMS1. Method of processing point cloud modalities of one or more maxillofacial structures including: providing a first point cloud modality of a first maxillofacial structure to an input of a deep neural network, which is trained to determine first geometric features for registration of the first point cloud modality; providing a second point cloud modality of a second maxillofacial structure to the input of the deep neural network to determine second geometric features for registration of the second point cloud modality; using a clustering algorithm to determine a set of correspondences between points in the first point cloud modality and point in the second point cloud modality based on the first and second geometric features respectively; using a transformation estimation algorithm and the set of correspondences to determine information about a coordinate transformation for registering the first point cloud modality with the second point cloud modality; and, optionally, registering the first point cloud modality with the second point cloud modality based on the coordinate transformation.
2. Method according to claim 1 wherein the set of correspondences define point-to-point correspondences between subsets of points in the first and second point cloud modality, the point-to-point correspondences including pairs of points C(j,j) wherein a point i in the first point cloud modality has a corresponding point j in the second point cloud modality; or, wherein the set of correspondences define soft point-to-point correspondences between subsets of points in the first and second point cloud modality, the point-to-point correspondences including pairs of pointswherein a point i in the first point cloud modality has a correspondence within a predefined distance of corresponding point j in a the second point cloud modality.
3. Method according to claims 1 or 2 wherein the deep neural network is trained to determine for each point of the first and second point cloud modality a plurality of first and second feature vectors respectively, wherein a feature vector represent geometric features for a point of a point cloud modality; and / or, wherein the geometric features include local geometric features associated with an individual tooth in the maxillofacial structure, preferably the local geometric features including information about the orientation, positionand / or type of an individual tooth; and / or, wherein the global geometric features are associated with two or more teeth in the maxillofacial structure, preferably the global geometric features including about the dentition position (quadrant), orientation and / or neighbors.
4. Method according to any of claims 1-3 wherein points of the first and / or second point cloud modality are associated with one or more pre-defined attributes, preferably the one or more pre-defined attributes including, a tooth type tooth type, for example a tooth type according to the FDI dental numbering system, a local curvature, a point normal, an implant, a bracket or an artificial crown.
5. Method according to any of claims 1-4 wherein the point cloud modality is generated by a first data acquisition system and the second point cloud modality is generated by a second data acquisition system that is different from the first data acquisition system, preferably the first and second data acquisition system being at one of: an inter-oral scanner (IOS), an X-ray scanner, a computed tomography (CT) system or a magnetic resonance imager (MRI) system.
6. Method according to any of claims 1-5 wherein the first point cloud modality represents a first part of a dentition of a patient and the second point cloud modality represents a second part of the dentition that partially overlaps the first part of the dentition; or, wherein the first point cloud modality represents a part of a dentition of a first patient and the second point cloud modality represents a part of the dentition of a second patient, wherein the set of correspondences are based on soft point-to-point correspondences.
7. Method according to any of claims 1-6 wherein the deep neural network is configured to receive all points of a cloud modality and, optionally, one or more pre-defined attributes associated with each point of the point cloud modality at its input; and / or, wherein the deep neural network is configured to output a dense set of geometrical features, preferably the dense set of geometrical features, including for a feature vector for each point of the point cloud modality; and / or, wherein the deep neural network is a fully convolutional and wherein the fully convolutional network is configured to receive all points of a cloud modality and, optionally, one or more pre-defined attributes associated with each point of the point cloud modality at its input and to apply generalized sparse convolution computations on each point of the point cloud modality, and, optionally, the one or more pre-defined attributes associated with each point of the point cloud modality.
8. Method according to any of claims 1-7 wherein the first and second point cloud modalities represent a first and second modality mesh respectively, the method further comprising: fusing the first modality with a separated part of the second modality, the fusing including: determining a separated part of the second modality mesh, the determining including removing portions of the second modality mesh that at least partially overlap with the first modality mesh by projecting surface normals of the second modality mesh in the direction of the first modality mesh and determining intersection of the projected surface normal with the first modality mesh; and, using a stitching algorithm to fuse the first modality mesh with the separated part of the second modality mesh; wherein, optionally, the removing of portions of the second modality further comprises: projecting surface normals of the second modality mesh onto the surface of the first modality mesh in a direction along the surface normals, each of the surface normals being associated with a part of the surface of the second modality mesh; determining if the projection of one of the surface normals intersects the surface of the first modality mesh; and, removing the part of the surface of the second modality mesh that is associated with the projected surface normal, if it is determined that the projected surface normal intersects the surface of the first modality mesh.
9. Method of fusing at least two point cloud modalities of one or more maxillofacial structures comprising: registering a first point cloud modality with a second point cloud modality, the first and second point cloud modalities representing a first and second modality mesh respectively, preferably the first and second point cloud modalities being registered using the method according to any of claims 1-8; determining a separated part of the second modality mesh, the determining including removing portions of the second modality mesh that at least partially overlap with the first modality mesh by projecting surface normals of the second modality mesh in the direction of the first modality mesh and determining intersection of the projected surface normal with the first modality mesh; and,using a stitching algorithm to fuse the first modality mesh with the separated part of the second modality mesh.
10. Method according to claims 9 wherein the removing of portions of the second modality further comprises: projecting surface normals of the second modality mesh onto the surface of the first modality mesh in a direction along the surface normals, each of the surface normals being associated with a part of the surface of the second modality mesh; determining if the projection of one of the surface normals intersects the surface of the first modality mesh; and, removing the part of the surface of the second modality mesh that is associated with the projected surface normal, if it is determined that the projected surface normal intersects the surface of the first modality mesh.
11. Method according to the claims 9 or 10 wherein the stitching algorithm comprises: meshing edges of the first modality mesh and the separated part of the second modality; and using one or more filtering operations applied to the seam region and to the resulting fused structure.
12. Method of training a deep neural network to determine geometric features for registration of point cloud modalities of a maxillofacial structure, the method comprising: providing a first point cloud modality of a first maxillofacial structure and, optionally one or more pre-defined attributes, to a data augmentation unit, which is using transformations to augment the first point cloud modality into a first augmented first point cloud modality; providing a second point cloud modality of a second maxillofacial structure and, optionally, one or more pre-defined attributes, to the data augmentation unit to augment the second point cloud modality into a second augmented first point cloud modality; using a deep neural network to determine first and second geometrical features associated with the augmented first and second point cloud modality respectively; using the first and second geometrical features and predetermined point-to- point correspondences between points of the first point cloud modality and points of the second point cloud modality, to optimize parameters of the deep neural network based on aloss function to train the deep neural network to generate geometrical features for registration of point cloud modalities.
13. Method according to claim 12, wherein the optimization of the network parameters is based on a contrastive learning scheme in which the loss function is used to minimize a distance between the first and second geometrical features that correspond with point-to-point correspondences and to maximize a distance between the first and second geometrical features that do not correspond to the point-to-point correspondences; and / or, wherein the transformations include one or more of: random cropping, translations, rotations, tooth removal, preferably the tooth removal being based on the predefined-features of the first point cloud modality and / or the predefined-features of the second point cloud modality.
14. A system for processing point cloud modalities of a maxillofacial structure comprising: a computer readable storage medium having at least part of a program embodied therewith; and, a computer readable storage medium having computer readable program code embodied therewith, and a processor, preferably a microprocessor, coupled to the computer readable storage medium, wherein responsive to executing the computer readable program code, the processor is configured to perform executable operations comprising: providing a first point cloud modality of a maxillofacial structure to an input of a deep neural network, which is trained to determine first geometric features for registration of the first point cloud modality; providing a second point cloud modality of at least part of the maxillofacial structure that at least partially overlaps the first point cloud modality to an input of the deep neural network to determine second geometric features for registration of the second point cloud modality; using a clustering algorithm to determine a set of correspondences between points of the first point cloud modality and points of the second point cloud modality based on the first and second geometric features respectively; and, using a transformation estimation algorithm to determine information about a coordinate transformation for registering the first point cloud modality with the second point cloud modality based on the set of correspondences.
15. A system for fusing at least two point cloud modalities of one or more maxillofacial structures comprising: a computer readable storage medium having at least part of a program embodied therewith; and, a computer readable storage medium having computer readable program code embodied therewith, and a processor, preferably a microprocessor, coupled to the computer readable storage medium, wherein responsive to executing the computer readable program code, the processor is configured to perform executable operations comprising: registering a first point cloud modality with a second point cloud modality, the first and second point cloud modalities representing a first and second modality mesh respectively, preferably the first and second point cloud modalities being registered using the method according to any of claims 1-11; determining a separated part of the second modality mesh, the determining including removing portions of the second modality mesh that at least partially overlap with the first modality mesh by projecting surface normals of the second modality mesh in the direction of the first modality mesh and determining intersection of the projected surface normal with the first modality mesh; and, using a stitching algorithm to fuse the first modality mesh with the separated part of the second modality mesh.
16. Computer program product comprising software code portions configured for, when run in the memory of a computer, executing the method steps according to any of 1-13.
Citation Information
Patent Citations
Integration of intra-oral imagery and volumetric imagery
US20140169648A1
Automatic complete tooth reconstruction method based on multi-modal data registration
CN112927358A
Method for generating digital data set representing target tooth arrangement for orthodontic treatment
US20230334771A1
Cited By
IOS image and CBCT image registration method based on particle swarm and single-tooth optimization
CN121544676A
Oral cavity scanning data occlusion contact point analysis method adopting deep learning
CN121812179A