Automatically determine the canonical pose of 3D objects and automatically superimpose 3D objects using deep learning
Through deep learning training of 3D deep neural networks, the regular posture of 3D dental structures is automatically determined and converted into regular coordinate systems, which solves the problem of accurate superposition of 3D dental models under low radiation conditions, and realizes high-precision automatic superposition and classification of dental structures.
Patent Information
- Application Number
- CN201980057637.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-07-03
- Filing Date
- 2019-07-03
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2039-07-03
AI Technical Summary
The prior art is difficult to generate high-precision 3D dental models under low radiation conditions, especially to deal with the superposition of 3D data sets of different modes, which leads to the inability of models to design accurate dental support templates in dental applications, and existing deep learning methods cannot effectively handle the large changes in 3D data sets of different modes.
Using deep learning methods, the regular pose of the 3D object is automatically determined by training the 3D deep neural network and converting it into a regular coordinate system to realize the automatic superposition of different 3D data sets, including using a convolutional neural network to process voxelized 3D data and determining the conversion parameters to convert the data set into a regular representation.
It realizes automatic superposition of high-precision 3D objects without manual intervention, improves the accuracy of dental structure segmentation and classification, ensures the accuracy and consistency of dental models, and is suitable for a variety of dental applications.
Smart Images

Figure CN112639880B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to automatically determining the canonical pose of 3D objects (e.g., 3D dental structures) and automatically superimposing 3D objects using deep learning; in particular, but not exclusively, to methods and systems for automatically determining the canonical pose of 3D objects and methods and systems for automatically superimposing 3D objects and computer program products enabling a computer system to perform these methods. Background Art
[0002] An accurate 3D model of a patient's dentition and jaws (maxilla and mandible) is crucial for 3D computer-assisted dental applications (orthodontic treatment planning, dental implant planning, orthognathic surgical planning (jaw surgery), etc.). The formation of such a 3D model is based on 3D image data of the patient, typically based on 3D computed tomography (CT) data representing a 3D object such as the dento-maxillofacial complex or other body part. A CT scan typically produces a voxel representation representing (a portion of) the 3D object, wherein each voxel is associated with an intensity value (typically the radio density of the scanned volume). In medical applications (e.g. dental applications), CT scans are typically acquired using cone-beam CT (CBCT) due to the lower radiation dose to the patient, lower purchase price of the equipment and greater ease of use compared to fan-beam CT.
[0003] However, CBCT technology is sensitive to artifacts (especially in the presence of metal), and there is no industry standard for converting the sensor output of a CBCT scanner into a radiation value representing the radio density in Hounsfield units (HU) in the scanned volume. Moreover, the use of low doses provides relatively poor contrast, making it difficult to distinguish structures of similar density in a 3D object (e.g., the dentofacial complex). These problems may lead to differences in 3D models derived from such voxel representations using, for example, thresholding techniques. Therefore, 3D models derived from voxel representations of CBCT data are not suitable, or at least less suitable, for designing precisely fitting tooth-supported templates, such as are used in, for example, orthodontics (clear aligner treatment), jaw surgery (orthognathic surgery), implant surgery (dental implantology), cosmetic dentistry (crowns, bridges), etc.
[0004] To address this problem, optical scanning data can be used to supplement, enhance and / or (partially) replace the voxel representation of the CBCT dataset or the 3D model derived from such voxel representation. Optical scanning data (e.g., intraoral scanning (IOS) data) are generated by performing a surface scan (e.g., laser or structured light) of the tooth surface (typically the crown and surrounding gingival surface) derived from a plaster model (or impression) of the patient's dentition, or by generating intraoral scanning (IOS) data of the patient's dentition. Compared to (CB)CT data, the advantages are the absence of radiation during data acquisition and higher spatial resolution. The typical accuracy of optical (extraoral) scanning and IOS is approximately in the range of 5 to 10 microns and 25 to 75 microns, respectively. However, the scan results cannot distinguish between the tooth (crown) area and the gingival area. In addition, information outside the visible surface cannot be captured, in particular information about the tooth root, jaw bone, nerves, etc. cannot be obtained. Intraoral scans can be supplemented with generalized model data, e.g., derived from a database of crown shapes with corresponding root shapes, to estimate the underlying structure, but such generalizations fail to take into account information about the actual 3D shape of the desired volume. Therefore, such model-based estimates are inherently imprecise.
[0005] More generally, when processing 3D image data, for example, to generate an accurate 3D model, to repair missing data in a 3D dataset, and to analyze and evaluate, for example, (potential) treatment effects / outcomes or for the purpose of disease progression analysis, it is advantageous or even necessary to combine 3D image datasets from different sources. This may mean aligning one or more voxelized 3D datasets (e.g., CBCT datasets) and / or one or more point clouds or 3D surface mesh datasets (e.g., IOS datasets) of the same 3D object (e.g., the same dental structure or bone structure), and merging the aligned sets into one dataset that can be used to determine an accurate 3D model or to perform an analysis of the dental structure. The process of aligning different image datasets is called image superposition or image registration. Therefore, the problem of superposition or registration involves finding a one-to-one mapping between one or more coordinate systems so that corresponding features of the models (e.g., 3D dental structures) in different coordinate systems are mapped to each other. The merging of aligned datasets into one dataset representing the dental structure is often called fusion.
[0006] In CT and CBCT imaging, known 3D superimposition techniques include point-based or landmark-based superimposition, surface-based or contour-based superimposition, and voxel-based superimposition. Examples of such techniques are described in the following articles: GKANTIDIS, N et al., Evaluation of 3-dimensional superimposition techniques on various skeletal structures of the head using surface models, PLoS One 2015, Vol. 10, No. 2; and JODA T et al., Systematic literature review of digital 3D superimposition techniques to create virtual dental patients, Int J Oral Maxillofac Implants, March-April 2015, Vol. 30, No. 2. Typically, these techniques require manual intervention, such as manual input.
[0007] The accuracy of point-based and surface-based overlay techniques depends on the accuracy of landmark identification and 3D surface model, respectively. This can be particularly problematic in the presence of artifacts and low-contrast areas. When matching different data sets, it can be challenging to identify corresponding landmarks with sufficient accuracy. Point-based matching algorithms (e.g., iterative closest point (ICP)) typically require user interaction to provide an initial state that is already relatively closely aligned. Voxel-based overlay may overcome some of the limitations of landmark-based and surface-based overlay techniques. This technique uses 3D volume information stored as voxel representations. The similarity between the 3D data to be overlaid can be inferred from the horizontal intensity of the voxels in the corresponding reference structures. This technique is particularly challenging when combining low-contrast, non-standardized CBCT data from different sources or combining data from different imaging modalities (e.g., CT and MRI, or CBCT and binary 3D image data (which may or may not be derived from a surface mesh that encloses a volume)). Additional difficulties may arise when the data sets only partially overlap. Existing voxel-based overlay methods are typically computationally expensive.
[0008] The large size of 3D data sets and the fact that clinical implementation requires very strict accuracy standards make it difficult to utilize traditional image superposition methods on high-dimensional medical images. With the recent development of deep learning, some efforts have been made to apply deep learning to the field of image registration. In one method, deep learning is used to estimate a similarity metric, which is then used to drive an iterative optimization scheme. This is reported, for example, by Simonovsky et al. in A Deep Metric for Multimodal Registration (deep metric for multimodal registration), MICCAI 2016 (Springer, Cham), pages 10-18, where the problem proposed is a classification task in which a CNN is set to distinguish between the alignment and misalignment of two superimposed image blocks. In another method, a deep regression (neural) network is used to predict the conversion parameters between images. For example, EP3121789 describes a method in which a deep neural network is used to directly predict the parameters of the conversion between a 3D CT image and a 2D X-ray image. Similarly, in the article “Non-rigid image registration using fully convolutional networks with deep self-supervision” by Li et al., dated September 3, 2017, a trained neural network receives two images and computes the deformations dx, dy, dx for each pixel, which is used to register one image to the other. This approach requires two input images that already have a certain degree of similarity and therefore cannot handle 3D datasets of different modalities. Therefore, the problem of registering 3D datasets of different modalities (pose, data type, coordinate system, etc.) of a specific 3D object is not addressed in the prior art.
[0009] The large variations in these 3D datasets that overlay systems should be able to handle (large variations in data format / modality, coordinate system, position and orientation of 3D objects, quality of image data, different amounts of overlap between existing structures, etc.) make the problem of accurate automatic overlay of 3D objects (e.g., overlay of 3D dental structures without any human intervention) a significant issue. Known overlay systems are unable to handle these issues in a reliable and robust manner. More generally, the large variations in 3D datasets of different modalities pose problems for accurate processing by deep neural network systems. This is not only a problem for accurate registration, but also for accurate segmentation and / or classification by deep neural networks.
[0010] Therefore, there is a need in the art for a method that can fully automatically, timely, and robustly superimpose 3D objects (e.g., 3D dental and maxillofacial structures, 3D datasets). More specifically, there is a need in the art for a solution in which, for example, a dental professional can obtain the superimposition results required for any of a variety of purposes without requiring additional knowledge or interaction from the professional, and with known accuracy and timeliness. Summary of the Invention
[0011] As will be appreciated by those skilled in the art, aspects of the present invention may be embodied as a system, method, or computer program product. Thus, aspects of the present invention may take the form of a complete hardware embodiment, a complete software embodiment (including firmware, resident software, microcode, etc.), or an embodiment incorporating software and hardware aspects (which may generally all be referred to herein as "circuits," "modules," or "systems"). The functions described in this disclosure may be implemented as algorithms executed by a microprocessor of a computer. Additionally, aspects of the present invention may take the form of a computer program product embodied in one or more computer-readable media having a computer-readable program code embodied thereon, such as stored thereon.
[0012] Any combination of one or more computer-readable media can be utilized. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of computer-readable storage media will include the following: an electrical connection with one or more wires, a portable computer floppy disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device or any suitable combination of the foregoing. In the context of this article, a computer-readable storage medium can be any tangible medium that can include or store a program for use by an instruction execution system, device or apparatus or used in conjunction with an instruction execution system, device or apparatus.
[0013] A computer-readable signal medium may include, for example, a propagated data signal in baseband or as part of a carrier wave having computer-readable program code embodied therein. Such a propagated signal may take any of a variety of forms, including but not limited to electromagnetic, optical, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium that is not a computer-readable storage medium and that can communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, device, or apparatus.
[0014] The program code embodied in the computer-readable medium can use any suitable medium to transmit, including but not limited to wireless, wired, optical fiber, cable, RF etc., or above-mentioned any suitable combination.Can write the computer program code for realizing the operation of aspect of the present invention with any combination of one or more programming languages, described programming languages comprise the function or object-oriented programming languages such as Java (TM), Scala, C++, Python etc., and the routine procedural programming language such as " C " programming language or similar programming language.Program code can be performed on user's computer completely, partly performs on user's computer, performs as independent software package, partly performs on user's computer and partly performs on remote computer, or performs completely on remote computer, server or virtual server.In the latter case, remote computer can be connected to user's computer by any type of network (comprising local area network (LAN) or wide area network (WAN)), or can be connected (for example, using Internet service provider to pass through the Internet) with external computer.
[0015] Aspects of the present invention are described below with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to embodiments of the present invention. It will be understood that each frame of the flowchart and / or block diagram and the combination of frames in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer or other programmable data processing device, particularly a microprocessor or central processing unit (CPU) or a graphics processing unit (GPU) to produce a machine so that the instructions are executed via the processor of the computer, other programmable data processing devices or other devices to create a component for implementing the function / action specified in one or more frames of the flowchart and / or block diagram.
[0016] These computer program instructions may also be stored in a computer-readable medium, which may direct a computer, other programmable data processing device, or other apparatus to operate in a specific manner so that the instructions stored in the computer-readable medium produce an article of manufacture, which includes instructions for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0017] The computer program instructions may also be loaded onto a computer, other programmable data processing device, or other apparatus to cause a series of operational steps to be performed on the computer, other programmable device, or other apparatus to produce a computer-implemented process, such that the instructions executing on the computer or other programmable device provide a process for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0018] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, segment or part of a code, which includes one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative embodiments, the functions specified in the box may not occur in the order specified in the figure. For example, depending on the functions involved, the two boxes shown in succession can actually be executed substantially simultaneously, or these boxes can sometimes be executed in reverse order. It should also be noted that each box in the block diagram and / or flow chart and the combination of the boxes in the block diagram and / or flow chart can be realized by a system based on dedicated hardware or a combination of dedicated hardware and computer instructions that performs a specified function or action.
[0019] In this application, "image" may refer to a data set that includes information in two, three or more spatial dimensions. "3D image data" may refer to any kind of three-dimensional data set (e.g., voxel intensity values, surface mesh definitions, etc.). Adjusting the coordinate system of one or more images to conform to the coordinate system of another reference image that includes (part of) the same structure is also referred to as data or image registration, matching, superposition or alignment. Unless the context indicates otherwise, these terms (and other words derived from these terms) can be used interchangeably. In this context, "(set of) transformation parameters" is a general term for information about how to rotate, translate and / or scale a data set in order to superimpose it on another data set or to represent the data set in an alternative coordinate system; it can be represented by a single matrix or by a collection of matrices, vectors and / or scalars, for example.
[0020] In one aspect, the present invention relates to a computer-implemented method for automatically determining a canonical pose of a 3D object in a 3D dataset. The method may include: a processor of a computer providing data points of one or more blocks of the 3D dataset to an input of a first 3D deep neural network, the first 3D deep neural network being trained to generate canonical pose information associated with a canonical coordinate system, the canonical coordinate system being defined relative to a position of a portion of the 3D object; the processor receiving the canonical pose information from an output of the first 3D deep neural network, the canonical pose information including, for each data point of the one or more blocks, a prediction of a position of the data point in the canonical coordinate system, the position being defined by the canonical coordinates; the processor using the canonical coordinates to determine an orientation of axes of the canonical coordinate system relative to axes and an origin of the first 3D coordinate system, a position of the origin of the canonical coordinate system, and / or a scaling of the axes of the canonical coordinate system; using the orientation and position to determine transformation parameters for transforming coordinates of the first coordinate system into canonical coordinates; and the processor determining a canonical representation of the 3D object, the determining comprising applying the transformation parameters to the coordinates of the data points of the 3D dataset.
[0021] In an embodiment, the 3D object may be a 3D dental structure. In an embodiment, the data points of the 3D dataset may represent voxels. In another embodiment, the data points of the 3D dataset may define points of a point cloud or points and normals of a 3D surface mesh.
[0022] In an embodiment, the first 3D deep neural network may be configured as a convolutional deep neural network configured to process voxelized 3D data.
[0023] In another embodiment, the first 3D deep neural network may be implemented as a deep multi-layer perceptron (MLP) based network capable of processing points of a 3D point cloud or a 3D surface mesh.
[0024] In an embodiment, the transformation parameters may include rotation, translation and / or scaling parameters.
[0025] In an embodiment, the canonical representation of the 3D object may be a canonical voxel representation or a canonical 3D mesh representation of the 3D object.
[0026] In an embodiment, the canonical pose information may include one or more voxel maps that are used to associate voxels of the voxel representation with predictions of the voxel's position in the canonical coordinate system.
[0027] In an embodiment, the one or more voxel maps may include a first 3D voxel map that associates (links) voxels with predictions of first canonical coordinates x' of a canonical coordinate system, a second 3D voxel map that associates (links) voxels with predictions of second canonical coordinates y' of a canonical coordinate system, and a third 3D voxel map that associates (links) voxels with predictions of third canonical coordinates z' of a canonical coordinate system.
[0028] In an embodiment, determining the orientation of an axis of a canonical coordinate system may further comprise determining a local gradient in canonical coordinates of a 3D voxel map in one or more 3D voxel maps for a voxel represented by the voxel, the local gradient representing a vector in a space defined by the first coordinate system, wherein the orientation of the vector represents a prediction of the orientation of the canonical axis and / or wherein the length of the vector defines a scaling factor associated with the canonical axis.
[0029] Thus, the method allows for the automatic determination of a canonical representation of a 3D object (e.g., a 3D dental structure). The method can be used to convert different 3D data modalities of a 3D object into a canonical pose of the 3D object, which can be used in a superposition process of different 3D datasets. Alternatively and / or additionally, the method can be used as a pre-processing step before providing the 3D dataset to a 3D input of one or more 3D deep neural networks, which are configured to segment the 3D objects (e.g., 3D dental structures) and / or determine a classification of the segmented 3D objects (e.g., teeth). Such a pre-processing step greatly improves the accuracy of 3D object segmentation and classification, because if the pose of the 3D object represented by the 3D dataset input to the system deviates too much from the normalized pose (especially with respect to orientation), the accuracy of such a trained neural network may be affected.
[0030] In another aspect, the present invention may be directed to a computer-implemented method for automatically superimposing a first 3D object (e.g., a first 3D dental structure) represented by (at least) a first 3D dataset and a second 3D object (e.g., a second 3D dental structure) represented by a second 3D dataset. In an embodiment, the first 3D object and the second 3D object are 3D dental structures of the same person. In an embodiment, the method may include: a processor of a computer providing one or more first blocks of voxels representing a first voxel of a first 3D object associated with a first coordinate system and one or more second blocks of voxels representing a second voxel of a second 3D object associated with a second coordinate system to an input of a first 3D deep neural network, the first 3D deep neural network being trained to generate canonical pose information associated with a canonical coordinate system, the canonical coordinate system being defined relative to a position of a portion of the 3D object; the processor receiving the first canonical pose information and the second canonical pose information from an output of the 3D deep neural network, the first canonical pose information including, for each voxel of the one or more first blocks, a prediction of a first position of the voxel in the canonical coordinate system; the second canonical pose information including, for each voxel of the one or more second blocks, a prediction of a second position of the voxel in the canonical coordinate system, the first position and the second canonical pose information being received from the output of the 3D deep neural network. The two positions are respectively defined by the first canonical coordinate and the second canonical coordinate; the processor uses the first canonical posture information to determine the first orientation of the axis and the first position of the origin of the axis in the first coordinate system, and uses the second canonical posture information to determine the second orientation of the axis of the canonical coordinate system and the second position of the origin of the axis in the second coordinate system; the processor uses the first orientation and the first position to determine the first transformation parameter (preferably the first rotation, translation and / or scaling parameter) for converting the coordinates of the first coordinate system into the coordinates of the canonical coordinate system; and uses the second orientation and the second position to determine the second transformation parameter (preferably the second rotation, translation and / or scaling parameter) for converting the coordinates of the second coordinate system into the canonical coordinate; and the processor determines the superposition of the first 3D object and the second 3D object, which determination includes using the first transformation parameter and the second transformation parameter to form a first canonical representation of the first 3D object and a second canonical representation of the second 3D object, respectively.
[0031] Thus, typically two or more different 3D datasets of a 3D object of the same patient (e.g., a 3D dental structure (representing a dento-maxillofacial complex)) can be superimposed by converting the coordinates associated with the 3D datasets into coordinates of a canonical coordinate system. In a typical example, the different 3D datasets can be different scans of (a portion of) the patient's dentition. Typically, the different 3D datasets at least partially overlap, i.e., the two datasets have at least a portion of the object elements in common (e.g., a tooth or crown in the case of a 3D dental structure). A 3D deep neural network (e.g., a 3D convolutional neural network) can determine the canonical pose of at least a portion of the 3D object. In some embodiments, the 3D deep neural network processes the coordinates block by block to meet computational constraints. The computer can then apply additional processing to derive the relative position (direction and scale) of the canonical origin and canonical axes for each voxel that is provided to the 3D convolutional network. Subsequently, transformation parameters for superimposing or aligning the two image datasets can be derived, and the transformation parameters can be used to align the 3D image datasets. The first 3D image dataset may be converted to align with the second 3D image dataset, or the second 3D image dataset may be converted to align with the first 3D image dataset, or both 3D image datasets may be converted so that they are aligned in a third orientation different from either received orientation.
[0032] 3D deep neural networks can be trained to be very robust to variations in the 3D dataset because the 3D deep neural networks are trained on a large number of typical 3D dentofacial structures, where these structures exhibit large spatial variations (translation, rotation and / or scaling). Problems related to the (limited) memory size of the 3D deep neural network can be solved by training the deep neural network based on subsamples (blocks) of the voxel representation. To this end, the voxel representation can be divided into voxel blocks of a predetermined size before being provided to the input of the 3D deep neural network. An additional advantage of using blocks is that the network can determine the canonical pose of even a limited amount of data (e.g., a few teeth instead of the entire dentition). Due to the fact that the canonical coordinate system is defined relative to a known (predetermined) standard of the object (e.g., the dentofacial structure), the obtained first canonical 3D dataset and the second canonical 3D dataset are aligned, where the accuracy may depend on the training time, the training sample variation and / or the available blocks for each received dataset.
[0033] In the present disclosure, a canonical pose defines a pose that includes position, orientation, and scale, and is a pose defined by assigning a predetermined position and / or orientation to a (preferably) reliably and unambiguously identifiable portion of a 3D object (e.g., a dental arch in the case of a 3D dental structure). In a similar manner, a canonical coordinate system can be defined by assigning an origin to a reliably and unambiguously identifiable position relative to the identifiable portion of the 3D object, thereby defining coordinate axes in a consistent manner, e.g., the x-axis along the tangent to the point of maximum curvature of the dental arch. Such a canonical coordinate system can define a standardized, unambiguous, predetermined coordinate system that is consistent across a certain type of 3D object data (e.g., all maxillofacial image data can be defined relative to a typical position, orientation, and scale of a reliably identifiable maxillofacial structure). Its function is to ensure that 3D objects in different image datasets are in the same relative position, orientation, and scale. Such a function can be used in a variety of applications where 3D data such as voxel representations, point clouds, or 3D meshes are processed by trained neural networks. The embodiments of the present disclosure utilize the insight that if two or more 3D objects are transformed into a canonical coordinate system, the 3D objects are also aligned with each other. Furthermore, they utilize the insight that the canonical pose of the dentofacial structure can be automatically determined by utilizing a 3D deep neural network (preferably a 3D convolutional neural network), thereby avoiding the need for human interaction.
[0034] In an embodiment, a first canonical representation of a first 3D object and a second canonical representation of a second 3D object, preferably the first 3D surface mesh and the second 3D surface mesh can be 3D surface meshes, and determining the superposition further comprises: segmenting the first canonical representation of the first 3D object into at least one 3D surface mesh of a 3D object element of the first 3D object (e.g., a first 3D dental object element), and segmenting the second canonical representation of the second 3D object into at least one 3D surface mesh of a second 3D object element of the second 3D object (e.g., a second 3D dental element); selecting at least three first non-collinear keypoints and at least three second non-collinear keypoints of the first 3D surface mesh and the second 3D surface mesh, keypoints (points of interest on surfaces of the 3D surface meshes); and aligning the first 3D dental element and the second 3D dental element based on the first and second first non-collinear keypoints and the first and second second non-collinear keypoints. In an embodiment, the keypoints can define local and / or global maxima or minima of a surface curvature of the first surface mesh.
[0035] In an embodiment, the first canonical representation and the second canonical representation of the first 3D object and the second 3D object may be voxel representations. In an embodiment, determining the superposition may further include: providing at least a portion of the first canonical voxel representation of the first 3D object and at least a portion of the second canonical voxel representation of the second 3D object to an input of a second 3D deep neural network, the second 3D deep neural network being trained to determine transformation parameters (preferably, rotation, translation and / or scaling parameters) for aligning the first canonical voxel representation and the second canonical voxel representation; aligning the first canonical representation of the first 3D object and the second canonical representation of the second 3D object based on the transformation parameters provided by the second 3D deep neural network output.
[0036] In an embodiment, determining the superposition may further include: the processor determining a volume of overlap between the canonical representation of the first 3D object and the canonical representation of the second 3D object.
[0037] In an embodiment, determining the overlay may further include: the processor determining a first volume of interest, the first volume of interest including first voxels of the first canonical representation in the overlapping volume; and determining a second volume of interest, the second volume of interest including second voxels of the second canonical representation in the overlapping volume.
[0038] In an embodiment, the method may further include: the processor providing first voxels contained in a first volume of interest (VOI) to an input of a third 3D deep neural network, the third 3D deep neural network being trained to classify and segment the voxels; and the processor receiving an activation value for each first voxel in the first volume of interest and / or an activation value for each second voxel in the second volume of interest from an output of the third 3D deep neural network, wherein the activation value of the voxel represents a probability that the voxel belongs to a predetermined 3D object element (e.g., a 3D dental element of a 3D dental structure (e.g., a tooth)); and the processor determining first and second voxel representations of the first and second 3D object elements in the first and second VOIs, respectively, using the activation values.
[0039] In an embodiment, the processor may determine first and second 3D surface meshes of the first and second 3D object elements using first and second voxel representations of the first and second 3D object elements.
[0040] In an embodiment, the method may further include: the processor selecting at least three first non-collinear key points and at least three second non-collinear key points of the first 3D surface mesh and the second 3D surface mesh, the key points preferably defining local and / or global maxima or minima in the surface curvature of the first surface mesh; and the processor preferably using an iterative closest point algorithm to align the first 3D object element and the second 3D object element based on the first and second first non-collinear key points and the first and second second non-collinear key points.
[0041] In an embodiment, the method may further include: the processor providing a first voxel representation of the first 3D dental element and a second voxel representation of the second 3D dental element to a fourth 3D deep neural network, the fourth 3D deep neural network being trained to generate an activation value for each of a plurality of candidate structure labels, the activation value associated with the candidate label representing a probability that the voxel representation received by the input of the fourth 3D deep neural network represents the structure type indicated by the candidate structure label; the processor receiving a plurality of first activation values and a second activation value from the output of the fourth 3D deep neural network, selecting a first structure label having a highest activation value among the first plurality of activation values, and selecting a second structure label having a highest activation value among the second plurality of activation values, and assigning the first structure label and the second structure label to the first 3D surface mesh and the second 3D surface mesh, respectively.
[0042] In an embodiment, the method may further include: the processor selecting at least three first non-collinear key points and at least three second non-collinear key points of the first 3D surface mesh and the second 3D surface mesh, the key points preferably defining local and / or global maxima or minima in the surface curvature of the first surface mesh; the processor marking the first key points and the second key points based on the first structure label assigned to the first 3D surface mesh and the second structure label assigned to the second 3D surface mesh, respectively; and the processor preferably using an iterative closest point algorithm to align the first 3D dental element and the second 3D dental element based on the first key points and the second key points of the first 3D surface mesh and the second 3D surface mesh, respectively, and the first structure label and the second structure label.
[0043] In another aspect, the present invention may relate to a computer-implemented method for training a 3D deep neural network to automatically determine a canonical pose of a 3D object (e.g., a 3D dental structure) represented by a 3D dataset. In embodiments, the method may include: receiving training data and associated target data, the training data comprising a voxel representation of the 3D object, the target data comprising canonical coordinate values of a canonical coordinate system for each voxel of the voxel representation, wherein the canonical coordinate system is a predetermined coordinate system defined relative to the position of a portion of the 3D dental structure; selecting one or more voxel blocks (one or more subsamples) of the voxel representation of a predetermined size and applying a random 3D rotation to the subsamples and applying the same rotation to the target data; providing the one or more blocks to an input of the 3D deep neural network, and the 3D deep neural network predicting a canonical coordinate of the canonical coordinate system for each voxel of the one or more blocks; and optimizing network parameter values of the 3D deep neural network by minimizing a loss function representing the deviation between the coordinate values predicted by the 3D deep neural network and the (appropriately transformed) canonical coordinates associated with the target data.
[0044] In another aspect, the present invention may relate to a computer system adapted to automatically determine a canonical pose of a 3D object (e.g., a 3D dental structure) represented by a 3D dataset, the computer system comprising: a computer-readable storage medium containing computer-readable program code, the program code comprising at least one trained 3D deep neural network; and at least one processor, preferably a microprocessor, coupled to the computer-readable storage medium, wherein, in response to executing the computer-readable program code, the at least one processor is configured to perform executable operations comprising: providing one or more voxel blocks of a voxel representation of the 3D object associated with a first coordinate system to an input of a first 3D deep neural network, the first 3D deep neural network being trained to generate canonical pose information associated with a canonical coordinate system, the canonical coordinate system The method comprises the steps of: receiving canonical pose information from an output of a first 3D deep neural network, the canonical pose information comprising, for each voxel of one or more blocks, a prediction of the position of the voxel in the canonical coordinate system, the position being defined by the canonical coordinates; determining, using the canonical coordinates, an orientation of axes of the canonical coordinate system and a position of the origin of the canonical coordinate system relative to axes and an origin of the first 3D coordinate system, and determining, using the orientation and the position, transformation parameters (preferably rotation, translation and / or scaling parameters) for transforming coordinates of the first coordinate system into canonical coordinates; and determining a canonical representation of the 3D object (preferably a canonical voxel representation or a canonical 3D mesh representation), the determining comprising applying the transformation parameters to coordinates of voxels of the voxel representation or a 3D dataset used to determine the voxel representation.
[0045] In yet another aspect, the present invention may relate to a computer system adapted to automatically superimpose a first 3D object (e.g., a first 3D dental structure) represented by a first 3D data set and a second 3D object (a second 3D dental structure) represented by a second 3D data set, the computer system comprising: a computer-readable storage medium containing computer-readable program code, the program code comprising at least one trained 3D deep neural network; and at least one processor, preferably a microprocessor, coupled to the computer-readable storage medium, wherein, in response to executing the computer-readable program code, the at least one processor is configured to perform executable operations comprising: providing one or more first voxel blocks of a first voxel representation of the first 3D object associated with a first coordinate system and one or more second voxel blocks of a second voxel representation of the second 3D object associated with a second coordinate system to an input of the 3D deep neural network; the 3D deep neural network being trained to generate canonical pose information associated with a canonical coordinate system defined relative to a position of a portion of the 3D object; receiving the first canonical pose information and the second canonical pose information from an output of the 3D deep neural network, the first The canonical pose information comprises, for each voxel of the one or more first blocks, a prediction of a first position of the voxel in the canonical coordinate system; the second canonical pose information comprises, for each voxel of the one or more second blocks, a prediction of a second position of the voxel in the canonical coordinate system, the first position and the second position being defined by the first canonical coordinate and the second canonical coordinate, respectively; determining a first orientation of an axis and a first position of an origin of the axis in the first coordinate system using the first canonical pose information, and determining a second orientation of the axis and a second position of the origin of the axis in the second coordinate system using the second canonical pose information; determining first transformation parameters (preferably first rotation, translation and / or scaling parameters) for transforming coordinates of the first coordinate system into coordinates of the canonical coordinate system using the first orientation and the first position; and determining second transformation parameters (preferably second rotation, translation and / or scaling parameters) for transforming coordinates of the second coordinate system into canonical coordinates using the second orientation and the second position; and determining a superposition of the first 3D object and the second 3D object, the determining comprising forming a first canonical representation of the first 3D object and a second canonical representation of the second 3D object using the first transformation parameters and the second transformation parameters, respectively.
[0046] In an embodiment, at least one of the first voxel representation and the second voxel representation may comprise (CB)CT data, wherein the voxel values represent radiodensity.
[0047] In an embodiment, at least one of the first voxel representation and the second voxel representation may comprise voxelized surface data or volume data obtained from a surface, preferably structured light or laser surface scanning data, more preferably intraoral scanner (IOS) data.
[0048] In another aspect, the present invention may also relate to a computer program product comprising software code portions configured to, when run in a memory of a computer, perform the method steps according to any of the above-described process steps.
[0049] The present invention will be further described with reference to the accompanying drawings, which schematically show embodiments according to the present invention. It will be understood that the present invention is not limited in any way to these specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 Depicts a schematic overview of a computer system for superimposing dental and maxillofacial 3D image data using deep learning according to an embodiment of the present invention;
[0051] Figure 2 depicts a schematic diagram of a system for determining a canonical pose of a 3D dental structure according to an embodiment of the present invention;
[0052] Figures 3A-3D depicts a schematic diagram illustrating a method of determining a canonical pose of a 3D dental structure according to an embodiment of the present invention;
[0053] Figure 4A -C shows the training and prediction data used by the system components according to an embodiment of the present invention;
[0054] Figure 5 Depicting an example of a 3D deep neural network architecture for generating canonical coordinates according to an embodiment of the present invention;
[0055] Figure 6 Depicts a schematic overview of system components for segmenting dento-maxillofacial 3D image data according to an embodiment of the present invention;
[0056] Figure 7A and 7B An example of a 3D deep neural network architecture for segmenting dentofacial 3D image data according to an embodiment of the present invention is depicted;
[0057] Figure 8 Depicts a schematic overview of system components for classification of dento-maxillofacial 3D image data according to an embodiment of the present invention;
[0058] Figure 9 An example of a 3D deep neural network architecture for classification of dental and maxillofacial 3D image data according to an embodiment of the present invention is depicted;
[0059] Figure 10A and 10B An example of generated keypoints is shown;
[0060] Figure 11depicts a schematic overview of system components for directly determining transformation parameters for superimposing voxel representations according to an embodiment of the present invention;
[0061] Figure 12A and 12B depicts received and transformed data employed within and derived from system components for directly generating transformation parameters in accordance with an embodiment of the present invention;
[0062] Figure 13 Depicts an example of a 3D deep neural network architecture for system components for directly deriving transformation parameters according to an embodiment of the present invention;
[0063] Figure 14 depicts a flow chart of system logic for selecting / determining conversion parameters to apply according to an embodiment of the present invention;
[0064] Figure 15A and 15B depicts conversion results on two exemplary 3D dentofacial image datasets according to various methods according to various embodiments of the present invention; and
[0065] Figure 16 is a block diagram illustrating an example data processing system that may be used to execute the methods and software products described in this disclosure. DETAILED DESCRIPTION
[0066] In the present disclosure, embodiments of computer systems and computer-implemented methods are described that use a 3D deep neural network to perform fully automatic, timely, accurate, and robust superposition of different 3D datasets representing 3D objects (e.g., 3D maxillofacial structures derived from a maxillofacial complex). The method and system enable superposition of at least two 3D datasets using a 3D deep neural network that is trained to determine a canonical pose for each of the two 3D datasets. The output of the trained neural network is used to determine transformation parameters that are used to determine a superimposed canonical 3D dataset, wherein the canonical 3D dataset represents a canonical representation of the 3D object (e.g., a maxillofacial structure). Other 3D deep learning networks and / or superposition schemes can be used to further improve the accuracy of the superposition. The system and method will be described in more detail below.
[0067] Figure 1A high-level schematic diagram depicts a computer system for automatically superimposing image data representing a 3D object, in this example, a 3D dental-maxillofacial complex, using deep learning, according to an embodiment of the present invention. A computer system 102 may include at least two inputs for receiving at least two 3D datasets, for example, a first dataset 106 and a second dataset 108. The first dataset 106 includes a first 3D object, for example, a first 3D dental structure, associated with a first coordinate system, and the second dataset 108 includes a second 3D object, for example, a second 3D dental structure, associated with a second coordinate system. The 3D datasets may represent the first 3D dental structure and the second 3D dental structure originating from the 3D dental-maxillofacial complex 104 (preferably from the same patient). The first and second 3D objects may have at least a common portion, for example, a common tooth portion in the case of a 3D dental structure. The 3D datasets may be generated by different scanners (e.g., different (CB) CT scanners and / or different optical scanners). Such scanning devices may include cone-beam CT scanners, fan-beam CT scanners, optical scanners such as intraoral scanners, and the like.
[0068] In the case of a CBCT scanner, the 3D dataset may include a voxel representation of the x-ray data generated by the CBCT scanner. The voxel representation may have a predetermined format, such as the DICOM format or a derivative thereof. The voxel representation defines a 3D voxel space of a predetermined size, such as a 400×400×400 voxel space, where each voxel is associated with a specific volume, and the position of the voxel in the voxel space may be defined based on a predetermined coordinate system.
[0069] Alternatively, in the case of an optical scanner, the 3D dataset may include surface mesh data, such as a set of points or vertices in 3D space connected by edges defining a set of surfaces, which in turn define a surface in the 3D space. The 3D dataset may also include point cloud data representing points in the 3D space defined by a 3D coordinate system. In an embodiment, the 3D dataset representing the surface mesh may be generated using an intraoral scanner, wherein the 3D dataset may have a predetermined format, such as an STL format or a derivative thereof. Also in this case, the 3D surface mesh representation defines a 3D space of a predetermined size, wherein the positions of the points and / or vertices are based on a predetermined coordinate system (which is different from the coordinate system used for other 3D datasets).
[0070] In some embodiments, the 3D surface mesh of the 3D dental structure can be segmented into separate segmented (i.e., separate) 3D dental elements, such as a crown and a surface belonging to the gums. Segmenting a 3D surface mesh into separate 3D surface meshes is a technique known in the art, as described, for example, in WU K et al., Tooth segmentation on dental meshes using morphologic skeleton, Comput Graph, February 2014, Vol. 38, pp. 199-211.
[0071] The 3D datasets may be generated (approximately) simultaneously or at different time points (pre- and post-operative scanning using the same or different scanning systems), wherein the representation of the 3D dental and maxillofacial complex may be defined based on a 3D coordinate system defined by image processing software, such that the orientation and / or proportions of the 3D dental and maxillofacial structures in the 3D dental and maxillofacial complexes of different 3D sets may vary significantly. The 3D dental and maxillofacial complex may include 3D dental and maxillofacial structures, referred to as 3D dental structures, such as jaws, teeth, gums, etc.
[0072] The large diversity of 3D data sets that can be provided as input to computer systems (in terms of data format / modality, coordinate system, position and orientation of 3D objects, quality of image data, different amounts of overlap between existing structures, etc.) makes the problem of accurate automatic superposition of 3D dental structures (i.e., superposition of 3D dental structures without any human intervention) a significant issue. Known superposition systems are unable to handle these problems in a reliable and robust manner.
[0073] To solve this problem, Figure 1The system may include a first trained 3D deep neural network 112 configured to receive voxel representations of different 3D datasets of a 3D dental and maxillofacial complex (preferably from a single patient). The 3D deep neural network is trained to determine a canonical pose of a 3D dental structure in a canonical coordinate system within the 3D dental and maxillofacial complex, wherein the canonical coordinate system defines a coordinate system relative to a location on a common dental and maxillofacial structure (e.g., a location on a dental arch). The 3D deep neural network may be configured to determine first transformation parameters 114 (in terms of translation, rotation, and / or scaling) for the voxel representations of the 3D datasets encoded in a memory of the 3D deep neural network. The first transformation parameters are determined based on translation, orientation, and / or scaling information of typical dental and maxillofacial features encoded in the 3D deep neural network and may be used to transform coordinates of a first coordinate system based on the first 3D dataset and coordinates of a second coordinate system based on the second 3D dataset into coordinates based on the canonical coordinate system. The first and second 3D datasets thus obtained represent the first and second 3D dental structures superimposed in the canonical coordinate system.
[0074] In the case where the first 3D dataset and / or the second 3D dataset is optical scan data, such data can be pre-processed before being provided as input to the first 3D deep neural network. Pre-processing can include converting the 3D scan data (e.g., a 3D mesh) into a voxel representation so that it can be processed by the 3D deep neural network. For example, the 3D surface mesh can be voxelized, for example, in such a way that the 3D voxel space represents at least the same real-world volume included in the 3D surface mesh data. For example, such a voxelized 3D surface mesh can have a binary voxel representation having a first default voxel value (e.g., "0") and a second voxel value (e.g., "1"), where the surface of the mesh data does not coincide with the representative voxel and the mesh data does coincide with the representative voxel. When the received 3D surface mesh defines an "open" 3D surface structure, the structure can be "closed" using an additional surface. Voxelization can be implemented as described above, where voxels located within a closed volume can also have a second value (e.g., "1"). In this way, a voxel representation of the volume is formed. The resolution (size of the voxels) may be chosen appropriately to produce accurate results throughout the system while still adhering to requirements that take into account, for example, available storage and processing.
[0075] In an embodiment, a 3D deep neural network can be used that is capable of determining the canonical pose of optical scan data (3D point cloud) directly based on the point cloud data. An example of such a network is a deep neural network based on a multilayer perceptron (MLP). MPL deep neural network architectures include PointNet (Qi, CR et al.: Pointnet: Deep learning on pointsets for 3d classication and segmentation. Proc. Computer Vision and Pattern Recognition (CVPR), IEEE 1(2), 4(2017)) or PointCNN (Li et al., "PointCNN: convolution onχ-transformed points", arXiv:1801.07791v5, November 5, 2018, to be published in Neural Information Processing Systems (NIPS) 2018). These MLP deep neural networks are capable of processing the points of the point cloud directly. As described in the present application, such a neural network can be trained to determine canonical pose information based on optical scan data. In practice, this would result in the ability to omit such a voxelization step as a pre-processing step, leading to faster processing and the ability to achieve higher accuracy results depending on the granularity of the point cloud data.
[0076] A further pre-processing step may include dividing the first 3D dataset and the second 3D dataset into blocks of a predetermined size. The block size may depend on the size of the 3D input space of the first 3D deep neural network and the memory space of the 3D deep neural network.
[0077] In an embodiment, a computer may determine the superimposed canonical first and second datasets by determining first transformation parameters for a first 3D dataset and first transformation parameters for a second 3D dataset and applying the thus determined transformation parameters to the first and second 3D datasets. A 3D deep neural network can be trained to be highly robust to large variations in the 3D dataset because the 3D deep neural network is trained on a large number of typical 3D dental and maxillofacial structures, wherein these structures exhibit large spatial variations (translation, rotation, and / or scaling). Issues associated with the (limited) memory size of the 3D deep neural network can be addressed by training the deep neural network based on subsamples (blocks) of the voxel representation. To this end, the voxel representation is first divided into blocks of a predetermined size before being provided to the input of the 3D deep neural network. Due to the fact that the canonical coordinate system is defined relative to a known (predetermined) standard of the dental and maxillofacial structures, the obtained first and second canonical 3D datasets are aligned, wherein the accuracy may depend on the training time, the variation of the training samples, and / or the available blocks of each received dataset. In addition, as described in more detail below, a specific network architecture can be used to encode a large amount of 3D image information taking into account spatial variations.
[0078] In some cases, it may be advantageous to further refine the accuracy of the superposition of the canonical 3D datasets. Therefore, in some embodiments, a further improvement of the superposition can be obtained using (partially overlapping) canonical voxel representations 118 of the first 3D dataset and the second 3D dataset and evaluating the superposition of the canonical voxel representations using a further second 3D deep learning network. In these embodiments, the computer can determine the overlap between the volumes defined by the 3D dental structures represented by the superimposed canonical first and second datasets. Here, the overlap can be defined as the volume within the space defined by the canonical coordinate system that is common to the 3D dental structures of the first and second datasets. The overlap can be used to select a volume of interest (VOI) in the canonical voxel representations of the first and second 3D datasets. In this way, a first VOI of the canonical voxel representation of the first 3D dataset and a second VOI of the canonical voxel representation of the second 3D dataset can be selected for input to a second 3D deep neural network 120, which is configured to determine second transformation parameters 122. The 3D deep neural network can be referred to as a direct conversion deep neural network because the neural network generates the conversion parameters in response to providing a canonical voxel representation to the input of the neural network. Applying the second conversion parameters to each of the first canonical 3D dataset and the second canonical 3D dataset (obtained based on the first conversion parameters) can further improve the accuracy of the superposition 116.
[0079] Alternatively and / or additionally, in some embodiments, further improvements in the overlay can be achieved by using canonical voxel representations of the first and second 3D datasets and evaluating the overlay of the canonical voxel representations based on an analytical overlay algorithm. In particular, in this embodiment, a canonical voxel representation of the first and second 3D data can be determined 124. Also in this case, the overlap between the volumes defined by the 3D dental structures represented by the overlaid canonical first and second datasets can be used to determine one or more first VOIs of the canonical voxel representation of the first 3D dataset and one or more second VOIs of the canonical voxel representation of the second 3D dataset. The one or more first VOIs and the one or more second VOIs can be provided as inputs to a third 3D deep neural network 126. The deep neural network is configured to classify the voxels of the VOIs of the voxel representation of the 3D dental structures and form voxel representations of different segmented 3D dental elements (e.g., teeth, jaw, gums, etc.). Furthermore, in some embodiments, a post-processing step can be applied in which a segmented 3D model of the segmented 3D dental elements is generated based on the classified voxels of the segmented 3D dental structures. Additionally, in some embodiments, an additional fourth 3D deep neural network may be used to label the segmented 3D voxel representations of dental elements according to known classification methods, e.g., to uniquely and consistently identify individual teeth.
[0080] The segmentation and classification processes can benefit from information derived from the first 3D deep neural network. In particular, determining and applying the first set of initial transformation parameters by the first 3D deep neural network can result in a canonical voxel representation of the 3D dataset, which allows for more accurate segmentation and / or classification results, as the accuracy of 3D deep neural networks used for segmentation and / or classification is relatively sensitive to large rotational variations of the 3D input data.
[0081] In addition, as described above, the third 3D deep neural network can use the overlap amount to determine in which volume of the space defined by the canonical coordinate system the same overlapping structures (e.g., 3D dental elements) of the first and second 3D datasets exist. The identification of the volume (VOI) that includes the overlapping structures in the first and second 3D datasets can be used to determine so-called key points. Key points are used to mark the same (overlapping) structures in two different datasets. Thus, a set of key points identifies the exact 3D positions of multiple points in the first 3D dataset that are linked to a set of associated key points in the second 3D dataset. The distance minimization algorithm can use the key points to calculate appropriate third transformation parameters 130 for accurately superimposing the first and second 3D datasets.
[0082] In an embodiment, the computer may use the superimposed canonical first and second 3D datasets (determined based on the first transformation parameters and optionally based on the second and / or third transformation parameters) to create a single fused 3D dataset 132 in a predetermined data format. Fusion of 3D datasets is known in the art, see for example the article by JUNG W et al., Combining volumetric dental CT and optical scan data for teeth modeling, Comput Aided Des, October 2015, Vol. 67-68, pp. 24-37.
[0083] Figure 2 A schematic diagram of a system for determining a canonical pose of a 3D dental structure in a canonical coordinate system according to an embodiment of the present invention is depicted. The system 200 includes at least one 3D deep neural network 222 having an input and an output. The system may include a training module 201 for training the 3D deep neural network based on a training set 212. Additionally, the system may include an inference module 203 configured to receive a 3D dataset representing a 3D object in a coordinate system and determine transformation parameters for transforming coordinates of voxels in the 3D dataset into canonical coordinates of a canonical coordinate system encoded in the 3D neural network during training.
[0084] The network can be trained based on a training set 212 comprising 3D image samples and their associated canonical coordinates. The training data can comprise a 3D dataset, for example, voxel intensity values, such as radio density in the case of (CB)CT data, or binary values, such as in the case of voxelized surface scan data. Canonical coordinate data, which can be represented as an (x, y, z) vector for each input voxel, can be used as target data.
[0085] A canonical coordinate system can be selected that is appropriate for a class of 3D objects (e.g., 3D dental structures). In an embodiment, in the case of 3D dental structures, the canonical coordinate system can be determined to have an origin (0,0,0) at a consistent point (inter-patient and intra-patient). Hereinafter, when referring to "real-world coordinates," this is considered to have axis directions relative to the patient's perspective, where the patient is standing upright, "lowest-highest" means "up-down" from the patient's perspective, "front-back" means "front-back" from the patient's perspective, and "left-right" means "left-right" from the patient's perspective. "Real-world" is intended to refer to the environment from which the information (e.g., a 3D dataset) originates. Such a consistent point can, for example, be the lowest point (in real-world coordinates) where the two most anteriorly positioned teeth (FDI system indices 11 and 21) are still in contact or would be in contact (e.g., if either of these teeth were missing). Taking into account the axis directions, the real-world directions of up-down, left-right, and front-back (as seen by the patient) can be defined, respectively, and encoded as x, y, and z values ranging from low to high values. To scale to real-world dimensions, various methods can be used, as long as it is done consistently across all training data, since the same scaling will be the output of the 3D deep learning network. For example, a value of 1 coordinate unit per 1 mm of real-world translation can be used.
[0086] To achieve a 3D deep neural network that is robust to variations in data and / or data modalities, a variety of training samples 212 can be generated based on an initial training set 202 comprising a 3D dataset (e.g., a voxel representation of a 3D dental structure) and associated canonical coordinate data. To this end, the training module can include one or more modules for preprocessing the training data. In one embodiment, to comply with the processing and storage requirements of the 3D deep neural network 222, a reduction module 204 can be configured to reduce the 3D dataset into a reduced 3D dataset and associated canonical coordinates of a predetermined resolution. This reduction operation results in a smaller 3D image dataset, for example, reducing the voxel resolution to 1 mm in each direction. In another embodiment, a transformation module 206 can be configured to generate different variations of a 3D dataset by applying random rotations to the (reduced) 3D data and associated canonical coordinates. Note that this operation can be performed for any available patient, effectively providing a data pool from which potential training samples can be drawn, with multiple patient datasets and multiple rotations of each dataset.
[0087] In a further embodiment, the training module may include a partitioning module 208 for partitioning the (reduced) 3D dataset and associated canonical coordinates into blocks (3D image samples), wherein each block has a predetermined size and is a subset of the total volume of the 3D dataset. For example, the 3D dataset provided as input to the training module may include a volume of 400×400×400 voxels, wherein each voxel has a dimension of 0.2 mm in each orthogonal direction. The 3D dataset may be reduced to a reduced 3D dataset having a volume of, for example, 80×80×80 voxels with a dimension of 1 mm in each direction. The partitioning module may then partition the reduced 3D dataset into 3D data blocks of a predetermined size (e.g., 24×24×24 voxels with a dimension of 1 mm in each direction). These blocks may be used to train a 3D deep neural network using canonical coordinates as targets. In an embodiment, the partitioning module may include a random selector for randomly selecting blocks that form the training set 212 for the 3D deep neural network 222.
[0088] Note that such a 3D deep learning network will inherently train on both varying rotations (from 206 ) and translations (from the random selection 208 ). Alternatively, in another embodiment, the samples may be presented at multiple scales that may be generated from 204 .
[0089] By means of a suitably trained 3D deep learning network 222, new 3D image data 214 (having arbitrary positions and orientations) can be presented as input to the system and appropriately processed, similar to the training 3D image data, more specifically, the reduced dataset is divided into image patches 218 of a predetermined size using a predetermined scaling factor 216, and the 3D image patches are rendered as requested by the 3D deep neural network 220. By rendering the image patches covering the entire space of the received 3D image data at least once, canonical coordinates can be predicted by the 3D deep neural network for each (down-sampled) voxel in the 3D image dataset.
[0090] This predicted data can be further processed 224 to generate a general set of transformation parameters that define how the received data can be transformed to align it as closely as possible with its canonical pose. This processing will be described and illustrated in more detail below. Note that with sufficient training samples from a relatively large real-world 3D space, the canonical pose can be determined for data received from a smaller volume (assuming it is representatively included in the training data). Note that in practice, the resolution of the input data may be around 1.25 mm. The predictions of the 3D deep neural network 222 can be generated in floating point values.
[0091] Figures 3A-3DDepicted is a schematic diagram illustrating a method of determining a canonical pose of a 3D object, such as a 3D dental structure, according to an embodiment of the present invention. Figure 3A A voxel representation 300 of a 3D object (e.g., a dental object such as a tooth) is schematically depicted. Voxels can be associated with intensity values (e.g., radiodensity obtained from a (CB)CT scan). Alternatively, voxels can be associated with binary values. In that case, the voxel representation can be a binary voxel representation of a voxelized surface or a volume derived from a voxelized surface obtained from a structured light scan or a laser surface scan. The 3D object can have specific features identifying the top (e.g., crown), bottom (e.g., root), front, back, and left and right portions. The voxel representation is associated with a first (orthogonal) coordinate system (x, y, z) 302, such as the coordinate system used by scanning software to represent scanned data in 3D space. These coordinates are provided, for example, as (meta)data in a DICOM image file. The 3D object can have a certain orientation, position, and size in the 3D space defined by the first coordinate system. Note, however, that such a coordinate system may not yet correspond to a system that can be defined relative to the object, illustrated here by "left," "right," "front," "back," "bottom," and "top." Using the trained 3D deep neural network, the 3D object can be (spatially) "normalized" (i.e., reoriented, repositioned, and scaled) 308 and defined based on a (orthogonal) canonical coordinate system. In the canonical coordinate system (x', y', z') 306, the normalized 3D object 305 can have a canonical pose, wherein certain features of the 3D object can be aligned with the axes of the canonical coordinate system. Thus, the system can receive a voxel representation of a 3D dental structure having a certain orientation, position, and size in a 3D space defined by a coordinate system defined by a scanning system, and determine a canonical voxel representation of the 3D object, wherein the 3D object is defined in the canonical coordinate system, the size of the object is scaled in the canonical coordinate system, and in the canonical coordinate system, certain features of the 3D dental structure are aligned with the axes of the canonical coordinate system.
[0092] Figure 3BA 3D deep neural network 318 is depicted that can be trained to receive voxels of a voxel representation 310 of a 3D object, where the voxels can have a certain position defined by a coordinate system 302 (x, y, z). The 3D deep neural network can be configured to generate so-called canonical pose information 303 associated with the voxel representation. The canonical pose information can include a prediction of the coordinates (x', y', z') in a space defined by a canonical coordinate system for each voxel 304 (x, y, z) of the voxel representation. The canonical coordinate system can be defined relative to typical positions, orientations, and proportions of reliably identifiable dentofacial structures (e.g., features of a dental arch). The information required to derive such a canonical coordinate system can be encoded in the 3D deep neural network during the training phase of the network. In this way, the canonical pose information can be used to place 3D data of different types and / or modalities representing the same dentofacial structure in the same relative position, orientation, and proportion.
[0093] Thus, for each input voxel 304, three corresponding output values 314, 324, 334 are generated by the 3D deep neural network, comprising predictions of the values of the x', y', and z' coordinates of the input voxel in the canonical coordinate system, respectively. In an embodiment, the canonical pose information may include three 3D voxel maps 312, 322, 332, wherein each 3D voxel map links the voxel represented at the input of the 3D neural network to the canonical coordinates.
[0094] Before providing the voxel representation to the input of the 3D deep neural network, the voxel representation can be divided into a set of voxel blocks (represented here by 316, hereinafter referred to as "blocks"), where the size of the voxel block matches the size of the input space of the 3D deep neural network. The block size may depend on the data storage capacity of the 3D deep neural network. Therefore, the 3D deep neural network can process the voxels in each block of the voxel representation and generate canonical pose information for the voxels of each block, that is, a prediction of the coordinates (x', y', z') of the canonical coordinate system for each voxel in the block. In an embodiment, the 3D deep neural network can generate three voxel maps 312, 322, and 332, the first voxel map 312 including the corresponding x' coordinate for each voxel in the block provided to the input of the 3D deep neural network; the second voxel map 322 including the y' coordinate for each voxel in the block; and the third voxel map 332 including the z' coordinate for each voxel in the block.
[0095] Figure 3CA voxel representation of a 3D object 300 is schematically shown, which is provided as an input to a 3D deep neural network and is defined based on a first coordinate system (x, y, z) 302 (e.g., the coordinate system used by the image processing software of a scanner used to generate the 3D image). These coordinates, or information used to determine these coordinates, can be included as metadata in a data file, such as a DICOM file. Based on the canonical pose information generated by the 3D deep neural network, a prediction of the canonical pose of the 3D object in the canonical coordinate system can be generated. Thus, the canonical pose information 350 can link the position (x, y, z) of each voxel in the first coordinate system to its position (x', y', z') in the canonical coordinate system. This information can be used to determine a transformation 360 that allows the system to transform the 3D object defined in the first coordinate system to its canonical pose 362 defined in the canonical coordinate system.
[0096] The pose information can be used to determine the orientation and scale factor associated with the axes of the canonical coordinate system (the canonical axes). Here, the orientation can be the orientation of the canonical axes in the space defined by the first coordinate system. The pose information can also be used to determine the position of the origin of the canonical coordinate system.
[0097] The orientation of the canonical axis can be determined based on the (local) gradient in one or more voxels in a 3D voxel map determined by a 3D deep neural network. For example, a local gradient can be determined for each or at least multiple voxels of the first 3D voxel map associated with the x' component of the canonical coordinates. The local gradient can be represented as a 3D vector in the x, y, z space defined by the first coordinate system. The direction of the vector represents a prediction of the orientation of the canonical x' axis at the location of the voxel. In addition, the length of the vector represents a prediction of the scale factor associated with the canonical x' axis. In an embodiment, the prediction of the orientation and scale factor associated with the canonical x' axis can be determined based on the x' value of the first 3D voxel map. For example, a statistically representative metric of the prediction of the voxels of the first 3D voxel map, such as a median or average gradient, can be determined. In an embodiment, the x' value of the first 3D voxel map can be preprocessed, for example, smoothed and / or filtered. For example, in an embodiment, a median filter can be used to remove (local) outliers. In the same manner, the orientation and scale factor prediction of the canonical y' axis can be determined based on the y' values in the second 3D voxel map, and the orientation and scale factor prediction of the canonical z' axis can be determined based on the z' values in the third 3D voxel map. The predicted orientations of the canonical x', y', z' axes can be post-processed to ensure that these axes are orthogonal or even orthonormal. Various known schemes (e.g., the Gram-Schmidt process) can be used to achieve this goal. The rotation and scaling parameters can be obtained by comparing the received coordinate system 302 with the coordinate system derived from the prediction.
[0098] The location of the origin of the canonical coordinate system (in terms of the translation vector in the space of the first coordinate system) can be obtained by determining the prediction of the canonical coordinates of the center of the voxel representation provided as input to the 3D deep learning network. These coordinates can be determined based on, for example, the average or median of the predicted x' value of the first 3D voxel map, the y' value of the second 3D voxel map, and the z' value of the third 3D voxel map. The translation vector can be determined based on the predicted canonical coordinates (xo', yo', zo') of the center of the block and the coordinates of the center of the block based on the first coordinate system, for example using a simple subtraction. Alternatively, the origin of the canonical coordinate system can be determined by aggregating multiple predictions of such blocks, the latter being effectively treated as canonical coordinates determined for the same size space of the received voxel representation. The above process can be repeated for each or at least most blocks of the 3D dataset. The information determined for each block (orientation, scale, and origin of the canonical coordinate system) can be used to obtain an average value that provides an accurate prediction.
[0099] therefore, Figure 2 The system and method depicted in FIG3 provides an efficient method for determining the canonical pose of a 3D dental structure. Figure 3D As shown, these methods include a first step 380 in which a processor of a computer provides a voxel representation of a 3D dental structure associated with a first coordinate system to an input of a 3D deep neural network, the neural network being configured to generate canonical pose information associated with a second canonical coordinate system. Thereafter, in step 382, the processor may receive the canonical pose information from the output of the 3D deep neural network, wherein, for each voxel of the voxel representation, the canonical pose information includes a prediction of the canonical coordinates of the voxel. Subsequently, the processor may perform a processing step 384 in which the canonical pose information is used to determine the orientation (and, if applicable, scaling) of the axes of the canonical coordinate system (e.g., by determining a vector representing the local gradient of the voxel position) and the position of the origin of the canonical coordinate system (e.g., by determining a vector representing the average (x', y', z') and thereby determining the average 3D distance to the canonical origin), and wherein the orientation and position (and, if applicable, scaling) are then used to determine transformation parameters used to transform coordinates of the first 3D coordinate system into coordinates of the second canonical coordinate system. Thereafter, in step 386, the processor determines the canonical pose of the 3D dental structure in the space represented by the second canonical coordinate system by applying the transformation parameters to the received 3D dataset. If the 3D dataset is represented by voxels, the parameters may be applied to the voxels. Alternatively, if the 3D dataset is represented by a mesh, the parameters may be applied to the coordinates of the mesh.
[0100] In this way, a canonical representation of a 3D object (e.g., a 3D dental structure) can be achieved. The method can be used to convert different 3D data modalities associated with the 3D object into a canonical pose of the 3D object, which can be used in a superposition process of different 3D datasets. Alternatively and / or additionally, the method can be used as a pre-processing step before providing the 3D dataset as a 3D input to one or more 3D deep neural networks, which are configured to segment the 3D object and (optionally) determine a classification of the segmented parts of the 3D object. Such a pre-processing step substantially improves the accuracy of the segmentation and classification of the 3D object (and / or reduces the training time or storage requirements for the same accuracy of such a 3D deep neural network), because if the pose of the 3D object represented by the 3D dataset input to the system deviates too much from the normalized pose (especially with respect to orientation), the accuracy of such a trained neural network will be affected.
[0101] Figure 4A -C shows the reference Figures 3A-3D Graphical representation of the training objectives and results used in the described method. Figure 4A Depicts three slices of the 3D dataset 400 1-3 , in this example, a CBCT scan of a 3D dental structure and the associated slices of a 3D voxel map of x', y' and z' coordinates that can be used to train a 3D deep neural network. These 3D voxel maps include the expected predictions of the canonical x' coordinate 4021, the canonical y' coordinate 4022 and the canonical z' coordinate 4023. The grayscale values visualize the gradient of the (encoded) values of the coordinates according to the canonical coordinate system. The coordinates (x, y, z) indicate the position of the voxels of the 3D dental structure based on the coordinate system associated with the CBCT scan. The axes of the visualization (including their direction) are indicated in the upper left of each picture. It is also worth noting that the grayscale values of the displayed gradients have been appropriately scaled so that the gradients are not distorted in the image. Figure 4A -C, the same values have the same grayscale value. This allows for better visual comparison of what is effectively a translation towards the canonical coordinate system, as encoded (for training) or predicted. Finally, note that all visualizations are 2D representations of a single intermediate "slice" (effectively a pixel of the 2D image data) and the associated voxel map, which is cut out of the actual 3D dataset, as indicated by the slice number visible in the upper left of each illustration.
[0102] For the purpose of training the system, e.g. Figure 4BAs shown, a 3D dataset representing a 3D dental structure can be assigned to a canonical coordinate system. In the case of these illustrations, for the illustration showing the gradient, the black value is -40.0 mm and the white value is +40 mm, effectively with the center of the patient scan as the origin (0,0,0). This data (3D image data and representation of the canonical system) has been appropriately scaled, as will be obtained by processor 204. This data can then be rotated (e.g., using linear or other interpolation methods) to produce the 3D data shown in the illustration at 406. Depending on the exact method used to perform this rotation, the size of the image space may also be expanded to include all voxels of the received 3D image dataset, which is not the case in these illustrations.
[0103] Figure 4B Training data 408 is shown, which may be obtained (performed by processor 208) from a random selection of appropriately sized blocks 412 (in this case, a subsample having dimensions of 24x24x24 voxels) from the randomly rotated input voxel representation 406. Note that when all three encoding directions of the canonical coordinate system are visualized in the same yz view (the middle yz slice of the 3D cube of voxels), as is done at 408, the gradient directions are visible, which effectively encode (in the case of this 2D visualization) the directions of the 2D components (in the yz plane) of the 3D direction vectors that encode the directions of the axes of the canonical coordinate system. Likewise, the value of each voxel effectively encodes the x', y', and z' coordinates of the voxel according to the canonical coordinate system. Note that when processing, for example, an entire dataset of 3D predictions, the 3D vector for each axis can be determined in terms of the canonical coordinate system. The selection of subsamples for training can be done in such a way that the selected smaller sized samples only include voxels that are part of the received 3D image dataset (i.e., not including "empty" patches of voxels along edges due to applied rotations, as can be seen from the illustration).
[0104] Figure 4CThe new input 416 (which may be obtained from the processor 216) after resizing is shown. For illustrative purposes, the input has been arbitrarily rotated. For the purposes of illustrations 418 and 420, only the xy views (slices) have been visualized. Set 418 shows the slices of predicted canonical coordinates x', y', and z'. It can be seen that the received image data has been divided into blocks (or subsamples), coordinate predictions performed on the blocks, and these predicted blocks are placed back into the total received 3D image set space (as can be seen from the square-like structures visible in the image, indicating the size of the blocks). Note that this effectively illustrates that the prediction of the parameters for the rotation, translation, and optional scaling used for the transformation to the canonical coordinate system can be performed on 3D image data of size 30×30×30 mm. The figure also illustrates that the trained network is again relatively robust to patches of "empty" data, which are generated from the rotations adopted for this visualization. (i.e., in the case of the network trained for the purposes of these illustrations, the "empty" data is assigned a value of always 0)
[0105] Shown at 420 are the encoded coordinate values that would become if 416 were training data, or the desired "target" values that should be produced from the 3D deep neural network. As can be seen, the general value of the gradient (indicating the distance of each voxel from the origin) and the general direction are very similar. If the real-world data were rotated, the "empty" patches seen in the illustration would not be present. The system can perform operations such as the following within the processor 224: 3D average filtering 418 of the predicted coordinate data, removal of outliers, and / or other methods for smoothing the resulting values. A representative measure of the predicted coordinate values (e.g., the mean or median) can be used to determine the position of the center of the voxel representation 416 relative to the canonical coordinate system. The translation can be determined based on the difference between the position of the center of the voxel representation 416 relative to the received coordinate system and the position relative to the canonical coordinate system.
[0106] The 3D gradient derivation algorithm (which is relatively computationally inexpensive) can generate three additional value cubes in each "axis cube of values", effectively generating three components of a vector describing the axis direction for each "axis cube". This can generate directional 3D vectors for the directions of all x-, y-, and z-axes for each voxel. A representative metric (e.g., mean or median) can be determined for these vectors for the desired coordinate axis. If applicable, these vectors for each axis can be converted to their equivalent unit vectors. In addition, the system can ensure that these three vectors are converted to the closest perfect orthogonal set of their three, so that the sum of the angular distances between the first set of vectors for each predicted axis and the resulting orthogonal set is minimized.
[0107] From these three (effective) predicted directions of the canonical axes, appropriate transformation parameters can be calculated that take into account the rotation of the received 3D image dataset towards the canonical orientation as part of the canonical pose. Subsequently, the system can determine what the average distance to the canonical origin is for each axis of the received 3D image dataset. From this, the transformation parameters for translating the received 3D image dataset can be calculated, effectively determining where the canonical origin should be within the received 3D image dataset (or its position relative to the received 3D image dataset), or conversely, where the canonical position should be in the coordinate system.
[0108] In another embodiment, a 3D deep neural network can be trained on varying scales, and the magnitude of the gradient of the resulting predictions can be used to determine the scale of the received 3D image dataset. This can be used to calculate transformation parameters toward the desired scaling of the received data.
[0109] Figure 5 An example of a 3D deep neural network architecture for determining canonical coordinates according to an embodiment of the present invention is depicted. The 3D deep neural network can have an architecture similar to a 3D U-net, which is actually a 3D implementation of a 2D U-net, as is known in the art.
[0110] The network can be implemented using various 3D neural network layers, such as (expanded) convolutional layers (3D CNN), 3D max pooling layers, 3D deconvolutional layers (3D de-CNN), and densely connected layers. These layers can use various activation functions, such as linear, tanh, ReLU, PreLU, Sigmoid, etc. The number of filters, filter sizes, and subsampling parameters of the 3D CNN and de-CNN layers can vary. The parameter initialization methods of the 3D CNN and de-CNN layers and the densely connected layers can vary. Dropout layers and / or batch normalization can be used throughout the architecture.
[0111] Following the 3D U-net architecture, during training, various filters in the 3D CNN and 3D de-CNN layers learn to encode meaningful features that contribute to prediction accuracy. During training, the matching set 502 of 3D image data and the encoded matching canonical coordinates 540 are used to optimize the prediction from the matching set of 3D image data toward the encoded matching canonical coordinates. A loss function can be used as the metric to be minimized. This optimization can be facilitated by using optimizers such as SGD and Adam.
[0112] Such an architecture can be used at various resolution scales, effectively downscaling 506, 510, 514 the results 504, 508, 512 from a previous set of 3D CNN layers through max pooling layers or (dilated and / or subsampled) convolutional layers. The term "meaningful features" refers to (continuous) derivatives of information relevant to determining the target output value, but are also encoded by 3D de-CNN layers, which effectively upscale while applying filters. By combining 520, 526, 532 the data generated from such 3D de-CNN layers 518, 524, 534 with data from the "previous" 3D CNN layer operating at the same resolution (512 to 520, 508 to 526, and 504 to 532), highly accurate predictions can be achieved. Additional 3D CNN layers 522, 528, 534 can be used throughout the upscaling path. Based on the incoming filter results of the 3D CNN layer 534, additional logic can be encoded within the parameters of the network by using densely connected layers that distill, for example, the logic of each voxel.
[0113] Input samples, when used for inference, and having been trained with encoded internal parameters in a manner such that validation produces sufficiently accurate results, can be presented and the 3D deep learning network can produce predicted canonical coordinates 542 for each voxel.
[0114] Figure 6 A schematic overview of system components for segmenting dento-maxillofacial 3D image data according to an embodiment of the present invention is depicted. Methods and systems for automatic segmentation based on deep learning are described in the following European Patent Application: No. 17179185.8 (entitled Classification and 3D modelling of 3D dento-maxillofacial structures using deep learning methods), which is incorporated herein by reference.
[0115] In particular, the computer system 602 can be configured to receive a 3D image data stack 604 of a dento-maxillofacial structure. The structure can include, for example, jaw structure, tooth structure, and neural structure. The 3D image data can include voxels, i.e., 3D spatial elements associated with voxel values, which represent radiation intensity or density values, such as grayscale values or color values. Preferably, the 3D image data stack can include CBCT image data according to a predetermined format (e.g., an image format or a derivative thereof).
[0116] In CBCT scans in particular, radiodensity (measured in Hounsfield units (HU)) is imprecise because different areas in the scan appear at different grayscale values depending on their relative location within the scanned organ. HU measurements from the same anatomical region by both CBCT and medical-grade CT scanners are not identical and are therefore unreliable for determining bone density of location-specific radiographic landmarks.
[0117] Furthermore, dental CBCT systems do not employ a standardized system for scaling the grayscale values representing reconstructed density values. Consequently, these values are arbitrary, making it impossible to assess bone quality. Without this standardization, grayscale interpretation is difficult, and it is impossible to compare values produced by different machines.
[0118] Teeth and jaw structures have similar densities, making it difficult for computers to distinguish between voxels belonging to teeth and those belonging to the jaw. Furthermore, CBCT systems are very sensitive to so-called beam hardening, which creates dark streaks between two highly attenuating objects (such as metal or bone) surrounded by brighter streaks.
[0119] For the reasons stated above, and as will be described in more detail below, for a superimposed system, utilizing the Figure 6 The described system components are particularly advantageous.
[0120] The system component may include a segmentation preprocessor 606 for preprocessing the 3D image data before it is fed to the input of a first 3D deep neural network 612, which is trained to produce a 3D set of classified voxels as output 614. Such preprocessing may, for example, include normalizing the voxel values to a range more favorable to the neural network. As will be described in more detail below, the 3D deep neural network can be trained according to a predetermined training regimen so that the trained neural network can accurately classify voxels in the 3D image data stack into different classes of voxels (e.g., voxels associated with teeth, jaws, and / or neural tissue). The 3D deep neural network may include a plurality of connected 3D convolutional neural network (3D CNN) layers.
[0121] The computer system may also include a segmentation post-processor 616 for accurately reconstructing 3D models of different parts of the dento-maxillofacial structure (e.g., teeth, jaws, and nerves) using voxels classified by a 3D deep neural network. The classified voxels 614 may include a collection of voxels representing, for example, all voxels classified as belonging to teeth, jaws, or nerve structures. It may be beneficial to create 3D data for these types of structures in such a way that each tooth and / or jaw (e.g., maxillary, mandibular) is represented by a separate 3D model. This can be achieved by volume reconstruction 620. For the case of separating the voxel collections belonging to each tooth, this can be achieved by (a combination of) 3D binary erosion, 3D marker creation, and 3D watershedding. For the combination separated into maxillary and mandibular parts, the distance from the origin along the upper and lower (real-world coordinate system) axes can be found at which the sum of the voxels in the plane perpendicular to that direction is at a minimum compared to other intersecting planes along the same axis. Using this distance, the maxillary and mandibular parts can be divided. In another embodiment, the jaws can be automatically segmented by a deep network by classifying corresponding voxels into separate jaw categories. The remaining portion of the classified voxels (e.g., voxels classified by the 3D deep neural network as belonging to nerves) can be post-processed using an interpolation function 618 and stored as 3D neural data 622. After segmentation, the 3D data of the various portions of the maxillofacial structure are post-processed, and the nerve, jaw, and tooth data 622-626 can be combined and formatted in separate 3D models 628 that accurately represent the maxillofacial structure in the input 3D image data fed to the computer system. Note that the classified voxels 614 as well as the 3D model 628 are defined in the same coordinate system as the input data 604.
[0122] In order to make the 3D deep neural network robust to variations present in, for example, current CBCT scan data, a module 638 may be used to train the 3D deep neural network to utilize 3D models of portions of the maxillofacial structure represented by the 3D image data. The 3D training data 630 may be correctly aligned to the CBCT image presented at 604 for which the associated target output is known (e.g., 3D CT image data of the maxillofacial structure and an associated 3D segmentation representation of the maxillofacial structure). Conventional 3D training data may be obtained by manually segmenting the input data, which may represent a significant amount of work. Additionally, manual segmentation results in low repeatability and consistency of the input data to be used.
[0123] To address this issue, in an embodiment, optically generated training data 630, i.e., an accurate 3D model of (part of) the maxillofacial structure, may be used instead of or in addition to the manually segmented training data. The maxillofacial structure used to generate the training data may be scanned using a 3D optical scanner. Such optical 3D scanners are known in the art and may be used to generate high-quality 3D jaw and tooth surface data. The 3D surface data may include a 3D surface mesh 632, which may be filled (determining which specific voxels are part of the volume enclosed by the mesh) and used by a voxel classifier 634. In this way, the voxel classifier is able to generate highly accurate classified voxels 636 for training. In addition, as described above, the training module may also use manually classified training voxels to train the network. The training module may use the classified training voxels as targets and the associated CT training data as input.
[0124] Figure 7A and 7B Depicted are examples of 3D deep neural network architectures for segmenting dentofacial 3D image data according to various embodiments of the present invention. Figure 7A As shown, the network can be implemented using a 3D convolutional neural network (3D CNN). The convolution layer can use an activation function associated with the neurons in the layer, such as a sigmoid function, a tanh function, a relu function, a softmax function, etc. Multiple 3D convolutional layers can be used, wherein slight changes in the number of layers and their defining parameters (e.g., different activation functions, number and size of kernels) and additional functional layers (e.g., dropout layers and / or batch normalization) can be used in the implementation without losing the essence of the 3D deep neural network design.
[0125] The network can include multiple convolution paths, in this example, three convolution paths: a first convolution path associated with a first set of 3D convolution layers 704, a second convolution path associated with a second set of 3D convolution layers 706, and a third set of 3D convolution layers 708. A computer performing data processing can provide a 3D dataset 702, such as CT image data, to the inputs of the convolution paths. The 3D dataset can be a voxel representation of a 3D dental structure.
[0126] exist Figure 7BThe functionality of the different paths is shown in more detail in . As shown in the figure, the voxels represented by the voxels can be provided to the input of the 3D deep neural network. The voxels represented by the voxels can define a predetermined volume, which can be referred to as the image volume 7014. The computer can divide the image volume into a first voxel block and provide the first block to the input of the first path. The 3D convolution layer of the first path 7031 can perform a 3D convolution operation on the first voxel block 7011. During processing, the output of a 3D convolution layer of the first path is the input of a subsequent 3D convolution layer in the first path. In this way, each 3D convolution layer can generate a 3D feature map representing a portion of the first pixel block provided to the input of the first path. Therefore, a 3D convolution layer configured to generate such a feature map can be referred to as a 3D CNN feature layer.
[0127] like Figure 7B As shown, the convolutional layer of the second path 7032 can be configured to process a second voxel block 7012 of voxel representation, wherein the second voxel block represents a downsampled version of the associated first voxel block, and wherein the first voxel block and the second voxel block have the same center origin. The representation volume of the second block is larger than the volume of the first block. In addition, the second voxel block represents a downsampled version of the associated first voxel block. The downsampling factor can be any suitable value. In an embodiment, the downsampling factor can be selected between 20 and 2, preferably between 5 and 3.
[0128] The first path 7031 can define a first set of 3D CNN feature layers (e.g., 5-20 layers) that are configured to process input data (e.g., a first voxel block at a predetermined position in the image volume) at a voxel resolution of the target (i.e., the voxels of the classified image volume). The second path can define a second set of 3D CNN feature layers (e.g., 5-20 layers) that are configured to process a second voxel block, wherein each block in the second voxel block 7012 has the same center point as its associated block from the first voxel block 7011. In addition, the voxels of the second block are processed at a resolution lower than the resolution of 7011. Therefore, the second voxel block represents a larger volume in real-world dimensions than the first block. In this way, the second 3D CNN feature layer processes the voxels to generate a 3D feature map that includes information about the direct neighbors of the associated voxels processed by the first 3D CNN feature layer. In this way, the second path enables the 3D deep neural network to determine contextual information, i.e., information about the context (e.g., its surroundings) of a voxel of 3D image data presented to the input of the 3D deep neural network.
[0129] In a similar manner, the third path 7033 can be utilized to determine additional contextual information for the first voxel block 7013. Thus, the third path can include a third set of 3D CNN feature layers (5-20 layers) configured to process a third voxel block, wherein each block in the third voxel block 7013 has the same center point as its associated blocks from the first voxel block 7011 and the second voxel block 7013. In addition, the voxels of the third block are processed at a resolution lower than the resolution of the first voxel block and the second voxel block. The downsampling factor can again be set to an appropriate value. In an embodiment, the downsampling factor can be selected between 20 and 3, preferably between 16 and 9.
[0130] By using three or more paths, 3D image data (input data) and contextual information about voxels of the 3D image data can be processed in parallel. Contextual information is important for classifying dentofacial structures, which often include densely packed, indistinguishable dental structures.
[0131] The outputs of the respective sets of 3D CNN feature layers are then combined and fed to the input of a set of fully connected 3D CNN layers 410, which are trained to derive the expected classification of voxels 412, which are provided at the input of the neural network and processed by the 3D CNN feature layers.
[0132] A collection of 3D CNN feature layers can be trained (via their learnable parameters) to derive and pass on the best useful information that can be determined from its specific input, with the fully connected layers encoding parameters that will determine how the information from the three previous paths should be combined to provide the best classified voxels 712. Here, the output of the fully connected layers (the last layer) can provide multiple activations for each voxel. Such voxel activations can represent a probability metric (prediction) that defines the probability that the voxel belongs to one of a plurality of classes (e.g., dental structure classes, such as teeth, jaws, and / or neural structures). For each voxel, the voxel activations associated with different dental structures can be thresholded to obtain a classified voxel. Thereafter, the classified voxels belonging to the different dental structure classes can be presented in image space 714. Thus, the output of the 3D deep neural network is a classified voxel in image space corresponding to the image space of the voxel at the input.
[0133] Please note that although Figure 6The segmentation 3D deep neural network described in conjunction with FIG7 may be inherently invariant to translations in the 3D image data space, but it may be advantageous to employ information from the processor 114 to apply an initial pre-alignment step 124 to at least adjust for rotations (albeit relatively roughly) to obtain a canonical pose. By having real-world orthogonal directions (e.g., patient up-down, left-right, and front-to-back) present in the 3D image data used in predefined canonical directions (e.g., internally (in the 3D dataset) representing up-down in z, left-right in x, and front-to-back in y, respectively), the memory bandwidth required for the segmentation 3D deep neural network can be reduced, training time can be reduced, and segmentation accuracy can be improved. This can be accomplished by specifically training the data and performing inference on the data (prediction of non-training samples) using the 3D dataset pre-aligned to account for the canonical rotation.
[0134] Figure 8 A schematic overview of system components for classifying 3D dental and maxillofacial image data according to an embodiment of the present invention is depicted. Methods and systems for automatic classification based on deep learning are described in the following European patent application: No. 17194460.6, entitled Automated classification and taxonomy of 3D teeth data using deep learning methods, which is incorporated herein by reference. System 800 may include two different processors, a first training module 802 for performing a process 826 of training a 3D deep neural network, and a second classification module 814 for performing a classification process based on new input data 816.
[0135] like Figure 8 As shown, the training module may include one or more repositories or databases 806, 812 of data sources intended for use in training. Such repositories may be obtained via an input 804 configured to receive input data (e.g., comprising 3D image data of the dentition), which may be stored in various formats along with corresponding desired labels. More specifically, at least a first repository or database 806 may be used to store 3D image data of the dentition and associated labels of the teeth within the dentition, which may be used by a computer system 808 configured to segment and extract volumes of interest 810 representing individual teeth that may be used for training. In the case of voxel (e.g., (CB)CT) data, such a system may be as described with respect to Figure 67 , or, alternatively, such a system may be, for example, a 3D surface mesh of individual tooth crowns, such as may be segmented from a 3D surface mesh comprising teeth and gums (e.g., IOS data). Similarly, a second repository or database 812 may be used to store 3D data in other formats, such as 3D surface meshes generated by optical scanning and labels of individual teeth that may be used during network training.
[0136] The 3D training data may be pre-processed 826 into a 3D voxel representation (voxelization) optimized for the 3D deep neural network 828. The training process may end at this stage, as the 3D deep neural network processor 826 may only need to train on a sample of individual teeth. In an embodiment, 3D dental data, such as a 3D surface mesh, may also be determined based on segmented 3D image data derived from a properly labeled complete dentition scan (808 to 812).
[0137] When using the classification module 800 for classifying (portions of) a new dentition 816, a variety of data formats may still be employed when converting the physical dentition into a 3D representation optimized for the 3D deep neural network 828. As described above, the classification system may, for example, utilize the 3D image data 106, 108 of the dentition and use a computer system 820 (which is 602) configured to segment and extract a volume of interest 822 (which is 626) including each tooth, similar to the training processor 808. Alternatively, another representation may be used, such as a surface mesh 824 of each tooth generated from an optical scan. Note again that the complete dentition data may be used to extract other 3D representations besides the volume of interest (820 to 824).
[0138] The data may be pre-processed 826 into the format required by the 3D deep neural network 828. Note that in the context of the entire overlay system, in the case where the received 3D image dataset is, for example, (CB)CT data, the classification of the segmented data by the network 828 may be performed directly using (a subset of) the data generated in the volume reconstruction 620. In the case where the received 3D image dataset is, for example, IOS data, the classification performed by 828 may be performed directly on (a subset of) the data after 3D surface mesh segmentation and voxelization of the crown.
[0139] The output of the 3D deep neural network can be fed into a post-classification processing step 830, which is designed to take advantage of knowledge about the dentition (e.g., the fact that each individual tooth index can only appear once in a single dentition) to ensure accuracy of the classification on the set of labels applied to the teeth of the dentition. This may result in the system outputting a tooth label for each identified individual tooth object. In an embodiment, to increase future accuracy after additional training of the 3D deep neural network, the correct labels can be fed back into the training data.
[0140] Figure 9 An example of a 3D deep neural network architecture for classification of maxillofacial 3D image data according to an embodiment of the present invention is depicted. The network can be implemented using 3D convolutional layers (3D CNN). Convolution can use activation functions. Multiple 3D convolutional layers 904-908 can be used, wherein slight changes in the number of layers and their defining parameters (e.g., different activation functions, number of kernels, use and size of subsampling) as well as additional functional layers (e.g., dropout layers and / or batch normalization layers) can be used in the implementation without losing the essence of the 3D deep neural network design.
[0141] In part to reduce the dimensionality of the internal representation of the data within the 3D deep neural network, a 3D max pooling layer 910 may be employed. At this point in the network, the internal representation may be passed to a densely connected layer 912, which serves as an intermediary for converting the representation in 3D space into activations of potential labels, particularly tooth type labels.
[0142] The final or output layer 914 may have the same dimension as the desired number of encoded labels and may be used to determine an activation value (similar to a prediction) for each potential label 918 .
[0143] The network can be trained using a dataset that has a pre-processed dataset 902 of 3D data as input to the 3D CNN layer, i.e., a 3D voxel representation of teeth. For each sample (which is a 3D representation of a single tooth), the matching representation of the correct label 916 can be used to determine the loss between the expected output and the actual output 914. This loss can be used as a metric to adjust parameters within the layers of the 3D deep neural network during training. An optimizer function can be used during training to aid in the efficiency of the training effort. The network can be trained for any number of iterations until the internal parameters result in the desired accuracy of the result. When properly trained, unlabeled samples may be presented as input and the 3D deep neural network can be used to derive predictions for each potential label.
[0144] Thus, when a 3D deep neural network is trained to classify a 3D data sample of teeth into one of a plurality of tooth types (e.g., 32 tooth types in the case of healthy dentition of an adult), the output of the neural network will be an activation value and an associated potential tooth type label. The potential tooth type label with the highest activation value can indicate to the classification system that the 3D data sample of teeth is most likely to represent the type of tooth indicated by the label. The potential tooth type label with the lowest or relatively low activation value can indicate to the classification system that the 3D data sample of teeth is least likely to represent the type of tooth indicated by such label.
[0145] Note that it may be necessary to train individual specific network models (same architecture with different final parameters after specific training) based on the type of input volume (e.g., the input voxel representation is the complete tooth volume, or the input voxel representation represents only the tooth crown).
[0146] It should also be noted that although Figure 8 and Figure 9 The described classification 3D deep neural network (as in the case of the segmentation 3D deep neural network) may be inherently invariant to translations in the 3D image data space, but it may be advantageous to use information from the processor 114 to apply an initial pre-alignment step 124 to at least adjust for rotations (albeit relatively roughly) to obtain a canonical pose. By having real-world orthogonal orientations (e.g., patient up-down, left-right, and front-to-back) present in the 3D image data used in predefined canonical orientations (e.g., internal (3D dataset) representations of z-up-down, x-left-right, and y-front-to-back, respectively), the memory bandwidth required for the classification 3D deep neural network can be reduced, training time can be reduced, and classification accuracy can be improved. This can be accomplished by specifically training and performing inference on the data using the 3D dataset pre-aligned to account for the canonical rotation.
[0147] Figure 10A and Figure 10B Examples of generated keypoints in two exemplary 3D dentofacial datasets, including and excluding classification information, are shown. Based at least on 3D image data (surface volumes) defining structures representing individual teeth or crowns, and based on, for example, in the case of (CB)CT data, on Figure 67 , or in the case of e.g. IOS data, by employing a more general determination of the surface mesh of the individual crowns, key points characterizing the surface can be determined. In practice, this can be viewed as a reduction step for reducing all available points in the surface mesh to a set of the most relevant (most significant) points. This reduction is beneficial because it reduces processing time and storage requirements. Additionally, methods for determining such points can be chosen that are expected to produce approximately the same set of points even if the input for their generation is a slightly different (set of) 3D surface meshes (still representing the same structure). Known methods in the art for determining keypoints from surface meshes typically include determining local or global surface descriptors (or features) that can be handcrafted (manually created) and / or machine-learned and optimized for repeatability across (slightly varying) input surface meshes, and can be optimized for performance (speed of determining salient points or keypoints), for example, as taught by TONIONI A et al., Learning to detect good 3D keypoints, Int J Comput Vis. 2018, Vol. 126, pp. 1-20. Examples of such features are local and global minima or maxima of surface curvature.
[0148] Figure 10A and Figure 10B Figure 1 shows a computer rendering of two received 3D image datasets, including the edges and vertices of the mesh defining the surface, thus showing the points defining the surface. The top four objects are individually processed and segmented crowns derived from the intraoral scan. The bottom four objects are based on the reference Figure 6 and 7 of the individual teeth obtained from CBCT scans. These two sets of four teeth were obtained from the same patient at approximately the same time. They have been processed by a processor (as described above with reference to FIG3, FIG4 and FIG5). Figure 5 As described in more detail, the teeth are roughly pre-aligned for the processor 114 that determines the canonical pose as described above. Based on the information after 114, the overlapping volume is determined and the 3D structure is segmented into separate surface meshes representing each tooth. Figure 10B In the case of Figure 8 and Figure 9 The described method performs a classification of 3D image data of individual teeth.
[0149] In particular, Figure 10AIn Figure 1, the points have been visualized using labels in the format P[number of received dataset]-[number of point]; the number of points has been reduced for visualization purposes. As can be seen, after keypoint generation, each received 3D image dataset has its own set of keypoints based on salient features of the volume, where identical points along the surface are marked with keypoints (albeit arbitrarily numbered). Note that it would be possible to group these points for each individual tooth in the original 3D dataset, but this would not provide any additional benefit, as it would be impossible to identify (the same) individual teeth in different 3D datasets.
[0150] exist Figure 10B In
[15] , using information from the additional classification step, the format of the labels has been visualized as P[number of received dataset]-[index of identified tooth]-[number of point]. For the same real-world tooth, the index of the identified tooth is the same in both received datasets. It should be noted that within each subgroup of individual teeth, the numbering of key points remains arbitrary.
[0151] It is worth noting that 3D surface mesh data (as well as point cloud data or sets of keypoints) are usually stored in the format of orthogonal x, y and z coordinates using floating point numbers. This opens up the possibility of highly accurate determination of the positions of keypoints and, therefore, the possibility of highly accurate alignment results with determined transformation parameters based on methods such as minimizing the calculated distance between clouds of such keypoints, as may be the case when employing, for example, an iterative closest point method.
[0152] like Figure 10B As shown, the additional information regarding which keypoints belong to which teeth (and the matching representations of the same teeth in other received 3D image datasets) can be particularly useful for enabling a more accurate determination of the alignment transformation parameters. For example, in cases where no initial pre-alignment has been performed, the average coordinates of each tooth can be used to determine the pre-alignment of one received 3D image dataset with another (in effect, roughly orienting the matching teeth as close to each other as possible). In other cases, it may be advantageous to first determine a set of transformation parameters (one set for each matching tooth between the two received 3D image datasets) and to determine the final transformation parameters based on the average of this set. This can be particularly beneficial in cases where the overlap volume between the two received 3D image datasets has not yet been (appropriately sufficiently) determined.
[0153] Note that in order to determine the alignment transformation parameters, at least three non-collinear points need to be determined.
[0154] Figure 11A schematic overview of system components for directly determining transformation parameters for superimposing voxel representations according to an embodiment of the present invention is depicted. System 1100 can be used to directly predict transformation parameters, such as applicable 3D rotations, 3D translations, and 3D scalings that define how one received 3D image dataset is aligned with another. Training data 1102 and inferred data 1116 can consist of 3D image data, such as voxel intensity values (e.g., radio density in the case of (CB)CT data) or binary values (e.g., in the case of voxelized surface scan data). The intensity values can be binarized by thresholding, such as by setting all voxel values above a value of, for example, 500 HU to 1 in the case of (CB)CT data, and the remaining voxels to 0. In particular, for the purpose of generating training data, this threshold can be randomly selected between the samples to be generated, for example, in the range of 400 to 800 HU.
[0155] The system can be used to predict parameters based on 3D image data with different modalities between two received 3D image datasets. Different sources of information, including information that accounts for different structures, can be trained by the same network. For example, when matching (CB)CT information with IOS information, the surface of a crown can be distinguished from both received datasets, while, for example, the gingiva would only be distinguishable in the IOS data, and, for example, the root would only be distinguishable in the CB)CT data.
[0156] During the training process, the internal parameters of the 3D deep neural network 1114 can be optimized towards providing results with sufficiently high accuracy for the network. This can be achieved by taking a set of 3D image datasets 1102 that may have varying modalities but that do include at least partial volumetric overlap of real-world structures. For the purpose of training such a network, it is desirable that the two input sets are aligned or overlapped with each other 1104. If this is not already the case in the data 1102, then the network can be trained, for example, according to the Figure 6-1 This can be done manually or automatically using the information from the method described in 0. The accuracy of the training data stacking may affect the accuracy of the output data.
[0157] Advantageously (for reasons of accuracy, memory bandwidth requirements, and potential processing speed) the data presented to the network 1114 for training comprises the same real-world structure and is scaled to the same real-world resolution in the voxel representation that will be provided to the 3D deep neural network 1114. If sufficient overlap does not already exist in the received datasets, this can be done manually or automatically based on the overlap region determined in the canonical coordinate system (e.g., according to the method described with respect to FIG3 ) 1106. If the input datasets have different resolutions, as known from metadata in the received data or derived, for example, by the method described with respect to FIG3 , it may be beneficial to rescale the high-resolution data to the resolution of the low-resolution data 1108.
[0158] Note that for the purpose of generating a large number of training samples from the same set of received 3D image data sets, selection of overlapping regions 1106 can be exploited in such a way that not only the largest overlapping volume of interest (VOI) is selected, but also smaller volumes within such largest overlapping volume are selected, thereby effectively "zooming in" on a subset of matching structural data in 3D.
[0159] For the purpose of generating a large number of training samples, random translation, rotation, and / or scaling transformations may be applied 1110, effectively misaligning any alignment present in the processed data until it reaches the processor 1110. For the purpose of using this introduced misalignment as a training target for the predicted transformation, this introduced misalignment may be passed to the 3D deep neural network 1114 in the form of applicable transformation parameters. The rotation and / or translation of the pre-processed dataset samples, or optionally the voxel representation of both samples, may be performed, for example, by rotation methods known in the art with linear (or other) interpolation.
[0160] A large number of samples obtained from pre-processing various sets of 3D image datasets including similar structures can be saved in a database (or memory) 1112 so that training 1114 of the network can be performed on multiple samples.
[0161] In another embodiment, a separate 3D deep neural network with a similar architecture can be trained for specific conditions (e.g., matching of a specific image modality, including real-world structures, and / or specific size scaling of voxel representations). This may produce results with potentially higher accuracy for specific situations while still adhering to hardware requirements, such as available system memory, processing speed, etc.
[0162] Once sufficiently trained 1114, "new data" 1116 can be presented for prediction or inference. This new data can be of the same type as described above, taking into account voxel representations of potentially different image modalities, etc. The canonical pose of the dental structures in the input dataset can be determined by a first 3D deep neural network 1118, and then a subset of data representing overlapping VOIs in a canonical coordinate system is selected 1120, for example, by the method described with reference to FIG3 . If the input datasets have different resolutions, as may be known from metadata in the received data, or derived, for example, by the method described with reference to FIG3 , a rescaling of the high-resolution data to the resolution of the low-resolution data can be performed 1122. This results in both datasets being pre-processed for receipt by the 3D deep learning network 1114. Note that the pre-alignment and selection of overlapping VOIs is expected to be less precise than the method described herein, and in this respect, this method may be considered an improvement over the higher-precision method described with reference to FIG3 .
[0163] Subsequently, the trained 3D deep neural network can process the preprocessed data 1114 and output as a prediction the transformation parameters 1126 for superimposing sample 1 and sample 2. The set of such parameters can, for example, include a vector of 6 values, the first 3 values encoding the applicable rotation to be performed in sequence along the three orthogonal axes of the received coordinate system for the data sample to be transformed (e.g., sample 2), and the last three values being the applicable translation, which can be positive and / or negative, to align or superimpose, for example, sample 2 to sample 1.
[0164] In another embodiment, these parameters may be trained in the form of, for example, rotation and / or translation matrices and / or translation matrices, achieving the same desired alignment or overlay result.
[0165] Note that the predicted transformation parameters for the received samples may not yet produce parameters for alignment or superposition of the original received 3D image dataset in the event that 1118, 1120, and / or 1122 have been employed. In such a case, the system 1100 may utilize processor 1128 to take into account information regarding any pre-processed transformations according to these three pre-processors (i.e., "stacking" any previous transformations with the predicted transformations for the samples), producing transformation parameters as outputs of 1128 that may be applicable to the received 3D image dataset.
[0166] Note that the inference utilization of the system can be considered to be relatively computationally inexpensive and therefore relatively fast. The accuracy of the system can be significantly higher when the pre-alignment and selection steps 1118 and 1120 are employed. The system can be highly robust to different imaging modalities and can operate at a variety of resolutions (voxel size in the received voxel representation), with a voxel resolution of 0.5-1 mm being employed depending on the amount of overlap between structures. However, it may not be as accurate in cases where there is insufficient overlap. The elements that make up the various transformation parameter sets can be in the form of floating point values.
[0167] Figure 12A and Figure 12B 12 shows diagrams of received data employed within a system component for directly deriving transformation parameters and transformed data resulting from the system component, according to an embodiment of the present invention. More specifically, they are visualizations of two received 3D image datasets (1202 and 1204). The visualizations are computer renderings of the 3D image datasets in their voxel representations.
[0168] In these particular visualizations, it can be seen that the voxel size used is 1 mm in either orthogonal direction. While the image 1202 received by the system component originates from CBCT data, for the purposes of this visualization, it is displayed as a 3D volume derived by thresholding the CBCT data above 500 Hounsfield units. 1204 is a voxelized representation of the IOS of the same patient, and the two received 3D image datasets were acquired at approximately the same time.
[0169] from Figure 12B As can be seen in FIG, by applying transformation parameters obtained from the system components, by means of 3D rotation and 3D translation, they have been aligned or superimposed 1204. In the case of this example, the received 3D image datasets already have the same scaling.
[0170] Figure 13 An example of a 3D deep neural network architecture for system components for directly deriving transformation parameters according to an embodiment of the present invention is depicted. Received (pre-processed) 3D image data, i.e., two voxel representations 1302, 1304 that match the voxel space of the input to the 3D deep neural network, can be passed through and processed by various layers 1306-1320 in the network. The first layer of the network can include multiple 3D convolutional layers 1306-1314.
[0171] After the data passes through the convolutional layers, the internal representation can be passed to a series of densely connected layers 1316–1318, which infer the rotation and translation distances between the 3D data.
[0172] Variations in the number of layers and their defining parameters (e.g., different activation functions, number of kernels, use and size of subsampling) as well as additional functional layers (e.g., dropout layers and / or batch normalization layers) can be used in the implementation without losing the essence of the 3D deep neural network design.
[0173] The final layer or output layer 1320 may represent the predictions of translation across three axes and rotation along three axes that should be applied to the data to obtain the correct overlay of the received 3D image dataset.
[0174] The training data may include as inputs 1302, 1304 a set of two voxel representations whose translations and rotations are known. For each data set of voxel representations to be processed, a random translation and rotation may be applied to either one, and the total translation and rotation difference may be used to determine the loss between the desired output 1322 and the actual output 1320. This loss may be used during training as a metric for adjusting parameters within each layer of the 3D deep neural network. Such a loss may be calculated so as to derive the most accurate predictions from the 3D deep learning network. An optimizer function may be used during training to aid in the efficiency of the training effort. The network may be trained for any number of iterations until the internal parameters result in a result of the desired accuracy. When properly trained, for example, two different voxel representations of maxillofacial structure may be presented as inputs, and the 3D deep neural network may be used to derive predictions 1324 of the translation and rotation required to accurately superimpose the set of inputs.
[0175] These layers can use various activation functions, such as linear, tanh, ReLU, PreLU, Sigmoid, etc. The number of filters, filter size, and subsampling parameters of the 3D CNN layers can vary. The parameter initialization method of these and the densely connected layers can also vary.
[0176] Figure 14 A flow chart of system logic for selecting / determining transformation parameters to be applied is depicted in accordance with an embodiment of the present invention. Note that this is an exemplary arrangement of system logic in accordance with various embodiments of the present invention as described above. For the purposes of the flow chart, the two input data sets are represented as having been appropriately voxelized. The two input data sets may be received at step 1402, at which point a first set of transformation parameters to a canonical pose may be determined. In an exemplary embodiment, this step is robust to large variations in the transformation parameters to be applied for alignment or overlay purposes. Accuracy may be low, and the resolution of the voxel representation of the received image data may be approximately 1 mm in any orthogonal direction.
[0177] Based on the information from 1402, a pre-alignment 1404 and a determination of sufficient overlap 1406 can be performed. Note that in embodiments, this step can perform two determinations of sufficient overlap, one for each optional subsequent method to be performed (starting at 1410 and 1416, respectively). The system can choose not to perform either or both of the methods starting at 1410, 1416 if the amount of overlap is insufficient according to a threshold value, which can be determined experimentally and can then be programmatically verified. That is, this can be considered a determination by the system that the conversion parameters generated by 1426 will not be improved due to the unfeasible results generated by one or both of these additional methods.
[0178] In the case of sufficient overlap, the direct derivation method can begin execution at step 1410, which is expected to produce more accurate results while being robust to different image modalities within the received 3D image dataset, especially if pre-alignment 1404 and VOI selection 1408 have already been performed. Note that applicable information after the previous transformation (potentially from 1404 and 1408) can be relayed for use in determining transformation parameters 1412 after the direct derivation method. The pre-processing 1410 employed in this method is expected to produce a voxel representation with a voxel resolution of 0.5-1.0 mm.
[0179] A sanity check 1414 may still be performed on the feasibility of the result. This may be done by comparing the parameters generated by 1402 with the parameters generated by 1414 and / or 1424. If the degree of deviation is too great, the system may, for example, choose not to relay the parameters to 1426, or 1426 may assign a weight of 0 to the resulting transformed parameters.
[0180] After determining applicable overlap, the system can employ a segmentation-based approach beginning at step 1416. The two received 3D image data sets can be automatically segmented 1416 using the 3D deep neural network-based approach described above, or other methods known in the art (as may be the case with IOS data). Note that in the latter case, such segmentation of the crown can be performed on the received 3D image data in the form of surface mesh data.
[0181] Classification 1418 may be performed on the (segmented) structural data and the resulting information may be relayed to a keypoint generation step 1420. It is expected that the ability to include identification of the same teeth in different received datasets will result in greater robustness to potential variations in the amount of overlap and data quality of the received datasets.
[0182] The resulting cloud of selected (sparse, closely matching) keypoints can be used to determine applicable transformation parameters for alignment or superposition at step 1422. Note again that any previous transformations, potentially produced by 1404, 1408, can be taken into account by 1422 to determine the set of transformation parameters used to set the transformation parameters.
[0183] A plausibility check 1424 for the method can again be performed, for example, by checking for deviations from the parameters generated at 1414 and / or 1402. In the event of a large discrepancy, the system can choose not to relay the parameters to 1426. Alternatively, 1426 can assign a weight of 0 to the resulting set of transformation parameters. An unworkable result may be the result of inaccurate data received, such as artifacts in the CBCT data, incorrect surface representation from IOS data, etc.
[0184] The point data of the surface mesh is stored in floating point precision, which makes it possible to produce highly accurate results. This method can be considered the most accurate within the system, while at the same time the least robust. However, it can be considered much more robust than current methods in the art due to the inclusion of determining pre-alignment, overlap, and segmentation of individual structures, as well as classification.
[0185] The transformation parameters can be represented internally in a variety of ways, for example, as three vectors describing three values for a rotation in sequence, three values for a translation to the origin, and / or three values for determining an applicable scaling factor, all of which have positive and / or negative values belonging to specific axes in an orthogonal 3D coordinate system. Alternatively, any combination of matrices known from linear algebra can be used, more specifically, rotations, translations, scalings, and / or combinations thereof that can be determined in an (affine) transformation matrix can be used.
[0186] A priori knowledge of considerations such as accuracy, robustness, etc. may be used, for example, to determine weights of importance for any / all transformation parameters received by 1426. Thus, step 1426 may programmatically combine parameters received from various methods to produce the most accurate desired transformation parameters for alignment or overlay.
[0187] Note that, depending on the desired results from such a system, the transformation parameters may be such that set 2 is matched to set 1, set 1 is matched to set 2, and / or the two sets are superimposed in an alternative (desired) coordinate system.
[0188] Figure 15A and Figure 15B Depicts the conversion results on two exemplary received data sets according to various embodiments of the present invention. More specifically, Figure 15A and Figure 15BComputer renderings of two 3D image datasets 1502 and 1504 are shown. These 3D image datasets are derived from a CBCT scanner and an intraoral scanner, respectively. Figure 14 With the described system setup, the system determined sufficient overlap and executed all three methods for generating transformation parameters, and employed pre-alignment according to the canonical pose method.
[0189] For the purpose of this visualization, the 3D CBCT image data are rendered with the aid of a surface mesh generated for each tooth structure resulting from the segmentation method. Figure 15A , the image data is shown in the received orientation, and it can be seen that the scaling between the two 3D image datasets is the same (e.g., 1 mm in real-world dimensions is equal to one unit value in each orthogonal axis of the two received datasets). It can also be seen that 1502 and 1504 are largely misaligned, taking into account rotation and translation.
[0190] The most accurate set of transformation parameters was determined to be the one produced by the segmentation and classification methods, matching and minimizing the distance between the keypoints generated for the two identified segmented (and labeled) teeth (crowns in the case of IOS data), so in the case of this example, no part of the applied transformation was a direct result of the other two methods. However, the transformation parameters from the regular pose method were used as preprocessing for the segmentation and classification based methods.
[0191] Figure 15B 1504 is shown as the transformation parameters determined by the system for applying to the 10S data, the system having been configured to determine and apply the transformation parameters that align one received 3D image dataset with another. Note that although overlap only exists for the image volumes defining tooth indices 41, 31, 32, 33, 34, 35, and 36 (as can be determined from the FDI representation), the final alignment or superposition step based on the application of the determined transformation parameters is automatically performed with a high degree of accuracy.
[0192] For example, for teeth that have overlap in the surface data (refer to Figure 1 132), the aligned or superimposed data as shown can be further fused or merged. In the case of visualized data, particularly showing the results of the segmentation step of generating a complete tooth including precise roots from CBCT data, combined with more precise information about the crown from IOS data, merging the surface of the IOS crown fused to the CBCT root can be very beneficial, for example, in the fields of implantology or orthodontics as previously described. Such merging methods are known in the art and can greatly benefit from the precise alignment produced by the system as described.
[0193] The method described above can provide the most accurate results for available overlays while being robust to the significant variability in input data conditions. This variability takes into account varying but potentially large "misalignments" between received 3D image datasets, different image modalities, and robustness against potential low data quality (e.g., misinterpreted surfaces, artifacts in CBCT data, etc.). The system can be fully automated and can deliver the most accurate alignment or overlay results in a timely manner. It should be noted that for any implementation of a 3D deep learning network, the expected results and robustness improve with longer training / utilization periods of more (variable) training data.
[0194] Although the examples in the figures are described with reference to 3D dental structures, it is clear that the embodiments of the present application can generally be used to automatically determine (thus, without any human intervention) the canonical pose of a 3D object in 3D datasets of different modalities. In addition, the embodiments of the present application can be used to automatically superimpose a first 3D object with a second 3D object, wherein the first 3D object and the second 3D object can be represented by 3D datasets of different modalities.
[0195] Figure 16 16 is a block diagram illustrating an exemplary data processing system that can be used as described in this disclosure. Data processing system 1600 may include at least one processor 1602 coupled to a memory element 1604 via a system bus 1606. In this way, the data processing system can store program code in the memory element 1604. In addition, processor 1602 can execute program code accessed from the memory element 1604 via the system bus 1606. In one aspect, the data processing system can be implemented as a computer suitable for storing and / or executing program code. However, it should be understood that data processing system 1600 can be implemented in the form of any system including a processor and a memory that can perform the functions described in this specification.
[0196] Memory element 1604 may include one or more physical memory devices, such as, for example, local memory 1608 and one or more mass storage devices 1610. Local memory may refer to random access memory or other non-persistent memory devices typically used during the actual execution of program code. Mass storage devices may be implemented as hard drives or other persistent data storage devices. Processing system 1600 may also include one or more cache memories (not shown) that provide temporary storage of at least some program code to reduce the number of times program code must be retrieved from mass storage devices 1610 during execution.
[0197] Input / output (I / O) devices, depicted as input device 1612 and output device 1614, may optionally be coupled to the data processing system. Examples of input devices may include, but are not limited to, a keyboard, a pointing device such as a mouse, and the like. Examples of output devices may include, but are not limited to, a monitor or display, speakers, and the like. The input and / or output devices may be coupled to the data processing system directly or through an intervening I / O controller. A network adapter 1616 may also be coupled to the data processing system to enable it to be coupled to other systems, computer systems, remote network devices, and / or remote storage devices through an intervening private or public network. A network adapter may include: a data receiver for receiving data transmitted to the data by the system, device, and / or network; and a data transmitter for transmitting data to the system, device, and / or network. Modems, cable modems, and Ethernet cards are examples of different types of network adapters that may be used with data processing system 1650.
[0198] like Figure 16 As shown, memory element 1604 can store application programs 1618. It should be understood that data processing system 1600 can also execute an operating system (not shown) that can facilitate the execution of application programs. Application programs implemented in the form of executable program code can be executed by data processing system 1600, for example, by processor 1602. In response to executing the application programs, the data processing system can be configured to perform one or more operations described in further detail herein.
[0199] In one aspect, for example, data processing system 1600 may represent a client data processing system. In this case, application 1618 may represent a client application that, when executed, configures data processing system 1600 to perform the various functions described herein with reference to a "client." Examples of a client may include, but are not limited to, a personal computer, a portable computer, a mobile phone, and the like.
[0200] The terms used herein are for the purpose of describing specific embodiments only and are not intended to limit the present invention. As used herein, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that when the terms "comprises" and / or "comprising" are used in this specification, they indicate the presence of the features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or groups thereof.
[0201] All parts or steps in the following claims plus the corresponding structure, materials, functions and equivalents of functional elements are intended to include any structure, material or function for performing the function in combination with other claimed elements for which protection is specifically claimed. The description of the present invention has been given for the purpose of illustration and description, but is not intended to be exhaustive or to limit the invention to the disclosed form. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the invention. The embodiments are chosen and described in order to best explain the principles of the invention and practical applications, and to enable others of ordinary skill in the art to understand the various embodiments of the invention and various modifications suitable for the intended specific use.
Claims
1. A computer-implemented method for automatically determining a canonical pose of a 3D dental structure represented by data points of a 3D data set, the method comprising: providing one or more blocks of data points of the 3D dataset associated with a first coordinate system to an input of a first 3D deep neural network, the first 3D deep neural network being trained to generate canonical pose information associated with a canonical coordinate system defined relative to a position of a portion of the 3D dental structure and associated with a canonical representation of the 3D dental structure, wherein one or more predetermined features of the canonical representation of the 3D dental structure are aligned with axes of the canonical coordinate system; receiving canonical pose information from an output of the first 3D deep neural network, the canonical pose information comprising, for each data point of the one or more blocks, a prediction of a position of the data point in the canonical coordinate system, the position of the data point being defined by canonical coordinates; using the canonical coordinates to determine an orientation and scaling of axes of the canonical coordinate system and a position of the origin of the canonical coordinate system relative to axes and an origin of the first coordinate system, and using the orientation and the position to determine transformation parameters for transforming coordinates of the first coordinate system to canonical coordinates; as well as A canonical representation of the 3D dental structure is determined, the determining comprising applying the transformation parameters to coordinates of data points of the 3D dataset.
2. The method according to claim 1, wherein The canonical pose information includes one or more voxel maps, and the one or more voxel maps are used to link the voxels represented by the voxel to the prediction of the position of the voxel in the canonical coordinate system, wherein the one or more voxel maps include a first 3D voxel map, a second 3D voxel map and a third 3D voxel map, the first 3D voxel map links the voxels to the prediction of the first x′ coordinate of the canonical coordinate system, the second 3D voxel map links the voxels to the prediction of the second y′ coordinate of the canonical coordinate system, and the third 3D voxel map links the voxels to the prediction of the third z′ coordinate of the canonical coordinate system.
3. The method according to claim 2, wherein: Determining the orientation of the axes of the canonical coordinate system further comprises: For a voxel of the voxel representation, a local gradient in canonical coordinates of a voxel map in the one or more voxel maps is determined, the local gradient representing a vector in a space defined by the first coordinate system, wherein the orientation of the vector represents a prediction of the orientation of a canonical axis and / or wherein the length of the vector defines a scaling factor associated with the canonical axis.
4. A computer-implemented method for automatically superimposing a first 3D dental structure represented by a first 3D dataset and a second 3D dental structure represented by a second 3D dataset, the method comprising: providing one or more first patches of voxels of a first voxel representation of a first 3D dental structure associated with a first coordinate system and one or more second patches of voxels of a second voxel representation of a second 3D dental structure associated with a second coordinate system to an input of a first 3D deep neural network trained to generate canonical pose information associated with a canonical coordinate system defined relative to positions of portions of the first 3D dental structure and the second 3D dental structure and associated with a canonical representation of the 3D dental structure, wherein one or more predetermined features of the canonical representation of the 3D dental structure are aligned with axes of the canonical coordinate system; Receiving first canonical pose information and second canonical pose information from an output of the 3D deep neural network, the first canonical pose information comprising, for each voxel of the one or more first blocks, a prediction of a first position of the voxel in the canonical coordinate system; and the second canonical pose information comprising, for each voxel of the one or more second blocks, a prediction of a second position of the voxel in the canonical coordinate system, the first position and the second position being defined by the first canonical coordinate and the second canonical coordinate, respectively; determining a first orientation and scale of the axes and a first position of the origins of the axes in the first coordinate system using the first canonical pose information, and determining a second orientation and scale of the axes of the canonical coordinate system and a second position of the origins of the axes in the second coordinate system using the second canonical pose information; determining first conversion parameters for converting coordinates of the first coordinate system to coordinates of the canonical coordinate system using the first orientation, scale, and first position; and determining second conversion parameters for converting coordinates of the second coordinate system to canonical coordinates using the second orientation, scale, and second position; and A superposition of the first 3D dental structure and the second 3D dental structure is determined, the determining comprising forming a first canonical representation of the first 3D dental structure and a second canonical representation of the second 3D dental structure using the first transformation parameter and the second transformation parameter, respectively.
5. The method according to claim 4, wherein The first canonical representation of the first 3D dental structure and the second canonical representation of the second 3D dental structure are 3D surface meshes, and determining the superposition further comprises: segmenting a first canonical representation of the first 3D dental structure into at least one 3D surface mesh of at least one 3D dental element of the first 3D dental structure, and segmenting a second canonical representation of the second 3D dental structure into at least one 3D surface mesh of at least one second 3D dental element of the second 3D dental structure; selecting at least three first non-collinear keypoints of the first 3D surface mesh and at least three second non-collinear keypoints of the second 3D surface mesh, the keypoints defining local and / or global maxima or minima in the surface curvature of the first 3D surface mesh; and The first 3D dental element and the second 3D dental element are aligned based on the first and second first non-collinear keypoints and the first and second second non-collinear keypoints.
6. The method according to claim 4, wherein: The first canonical representation of the first 3D dental structure and the second canonical representation of the second 3D dental structure are voxel representations, and determining the superposition further comprises: providing at least a portion of a first canonical voxel representation of the first 3D dental structure and at least a portion of a second canonical voxel representation of the second 3D dental structure to inputs of a second 3D deep neural network trained to determine transformation parameters for aligning the first canonical voxel representation and the second canonical voxel representation; and The first canonical representation of the first 3D dental structure and the second canonical representation of the second 3D dental structure are aligned based on the transformation parameters provided by the output of the second 3D deep neural network.
7. The method according to claim 4, wherein: Determining the overlay also includes: determining a volume of overlap between the canonical representation of the first 3D dental structure and the canonical representation of the second 3D dental structure; and A first volume of interest is determined, the first volume of interest including first voxels of the first canonical representation in the overlapping volume; and a second volume of interest is determined, the second volume of interest including second voxels of the second canonical representation in the overlapping volume.
8. The method according to claim 7, further comprising: providing a first voxel contained in the first volume of interest (VOI) to an input of a third 3D deep neural network trained to classify and segment the voxels; as well as receiving, from an output of the third 3D deep neural network, an activation value for each first voxel in the first volume of interest and / or an activation value for each second voxel in the second volume of interest, wherein the activation value of a voxel represents a probability that the voxel belongs to a tooth of the 3D dental structure; as well as A first voxel representation of a first 3D dental element in the first VOI and a second voxel representation of a second 3D dental element in the second VOI are determined using the activation values, respectively.
9. The method according to claim 8, further comprising: determining a first 3D surface mesh of the first 3D dental element and a second 3D surface mesh of the second 3D dental element using a first voxel representation of the first 3D dental element and a second voxel representation of the second 3D dental element; selecting at least three first non-collinear keypoints of the first 3D surface mesh and at least three second non-collinear keypoints of the second 3D surface mesh, the keypoints defining local and / or global maxima or minima in a surface curvature of the first 3D surface mesh; as well as The first 3D dental element and the second 3D dental element are aligned based on the first and second non-collinear keypoints.
10. The method according to claim 8, further comprising: determining a first 3D surface mesh of the first 3D dental element and a second 3D surface mesh of the second 3D dental element using a first voxel representation of the first 3D dental element and a second voxel representation of the second 3D dental element; providing a first voxel representation of the first 3D dental element and a second voxel representation of the second 3D dental element to a fourth 3D deep neural network, the fourth 3D deep neural network being trained to generate an activation value for each of a plurality of candidate structure labels, the activation value associated with the candidate label representing a probability that the voxel representation received as an input to the fourth 3D deep neural network represents a structure type indicated by the candidate structure label; The method further comprises receiving a plurality of first activation values and a plurality of second activation values from an output of the fourth 3D deep neural network, selecting a first structure label having a highest activation value among the first plurality of activation values, selecting a second structure label having a highest activation value among the second plurality of activation values, and assigning the first structure label and the second structure label to the first 3D surface mesh and the second 3D surface mesh, respectively.
11. The method according to claim 10, further comprising: selecting at least three first non-collinear keypoints of the first 3D surface mesh and at least three second non-collinear keypoints of the second 3D surface mesh, the keypoints defining local and / or global maxima or minima in a surface curvature of the first 3D surface mesh; labeling a first keypoint and a second keypoint based on a first structure tag assigned to the first 3D surface mesh and a second structure tag assigned to the second 3D surface mesh, respectively; as well as The first and second 3D dental elements are aligned based on the first and second keypoints and the first structure tags of the first 3D surface mesh and the second structure tags of the second 3D surface mesh, respectively.
12. A computer-implemented method for training a 3D deep neural network to automatically determine a canonical pose of a 3D dental structure represented by a 3D dataset, comprising: receiving training data and associated target data, the training data comprising a voxel representation of a 3D dental structure, the target data comprising canonical coordinate values of a canonical coordinate system for each voxel of the voxel representation, wherein the canonical coordinate system is a predetermined coordinate system defined relative to a position of a portion of the 3D dental structure, the canonical coordinate system being associated with the canonical representation of the 3D dental structure, wherein one or more predetermined features of the canonical representation of the 3D dental structure are aligned with axes of the canonical coordinate system; selecting one or more blocks of voxels, the one or more blocks representing one or more subsamples of a voxel representation of a predetermined size, and applying a random 3D rotation to the one or more subsamples; applying the same rotation to the target data; providing the one or more blocks to an input of a 3D deep neural network, and the 3D deep neural network predicting a canonical coordinate of a canonical coordinate system for each voxel of the one or more blocks; and The values of the network parameters of the 3D deep neural network are optimized by minimizing a loss function, wherein the loss function represents the deviation between the coordinate values predicted by the 3D deep neural network and the canonical coordinates associated with the target data.
13. A computer system adapted to automatically determine a canonical pose of a 3D dental structure represented by a 3D dataset, comprising: A computer-readable storage medium containing computer-readable program code, wherein the program code includes at least one trained 3D deep neural network, and at least one processor coupled to the computer-readable storage medium, wherein, in response to executing the computer-readable program code, the at least one processor is configured to perform executable operations comprising: providing one or more blocks of voxels of a voxel representation of a 3D dental structure associated with a first coordinate system to an input of a first 3D deep neural network, the first 3D deep neural network being trained to generate canonical pose information associated with a canonical coordinate system defined relative to a position of a portion of the 3D dental structure and associated with the canonical representation of the 3D dental structure, wherein one or more predetermined features of the canonical representation of the 3D dental structure are aligned with axes of the canonical coordinate system; receiving canonical pose information from an output of the first 3D deep neural network, the canonical pose information comprising, for each voxel of the one or more blocks, a prediction of a position of the voxel in the canonical coordinate system, the position being defined by canonical coordinates; using the canonical coordinates to determine the orientation and scale of the axes of the canonical coordinate system and the position of the origin of the canonical coordinate system relative to the axes and origin of the first coordinate system, and using the orientation, scale and position to determine transformation parameters for transforming coordinates of the first coordinate system to canonical coordinates; and A canonical representation of the 3D dental structure is determined, the determining comprising applying the transformation parameters to coordinates of voxels of the voxel representation or to a 3D dataset used to determine the voxel representation.
14. A computer system adapted to automatically superimpose a first 3D dental structure represented by a first 3D dataset and a second 3D dental structure represented by a second 3D dataset, the computer system comprising: A computer-readable storage medium containing computer-readable program code, wherein the program code includes at least one trained 3D deep neural network, and at least one processor coupled to the computer-readable storage medium, wherein, in response to executing the computer-readable program code, the at least one processor is configured to perform executable operations comprising: providing one or more first patches of voxels of a first voxel representation of a first 3D dental structure associated with a first coordinate system and one or more second patches of voxels of a second voxel representation of a second 3D dental structure associated with a second coordinate system to an input of a 3D deep neural network; the 3D deep neural network being trained to generate canonical pose information associated with a canonical coordinate system defined relative to positions of portions of the first and second 3D dental structures and associated with a canonical representation of the 3D dental structure, wherein one or more predetermined features of the canonical representation of the 3D dental structure are aligned with axes of the canonical coordinate system; Receiving first canonical pose information and second canonical pose information from an output of the 3D deep neural network, the first canonical pose information comprising, for each voxel of the one or more first blocks, a prediction of a first position of the voxel in the canonical coordinate system; and the second canonical pose information comprising, for each voxel of the one or more second blocks, a prediction of a second position of the voxel in the canonical coordinate system, the first position and the second position being defined by the first canonical coordinate and the second canonical coordinate, respectively; determining a first orientation and scale of the axes and a first position of the origins of the axes in the first coordinate system using the first canonical pose information, and determining a second orientation and scale of the axes of the canonical coordinate system and a second position of the origins of the axes in the second coordinate system using the second canonical pose information; determining first conversion parameters for converting coordinates of the first coordinate system into coordinates of the canonical coordinate system using the first orientation and the first position; and determining second conversion parameters for converting coordinates of the second coordinate system into canonical coordinates using the second orientation and the second position; and A superposition of the first 3D dental structure and the second 3D dental structure is determined, the determining comprising forming a first canonical representation of the first 3D dental structure and a second canonical representation of the second 3D dental structure using the first transformation parameter and the second transformation parameter, respectively.
15. A computer program product comprising software code portions configured to, when run in a memory of a computer, perform the method steps according to any one of claims 1 to 11.
Citation Information
Patent Citations
Method and system for convolutional neural network regression based 2d / 3d image registration
EP3121789A1
Integration of intra-oral imagery and volumetric imagery
EP2742857A1