Determining the spatial relationship between maxillary and mandibular teeth
The method uses 3D models and 2D images to optimize alignment between maxillary and mandibular teeth, addressing inaccuracies in traditional and digital methods, ensuring precise and natural alignment for dental prostheses.
Patent Information
- Application Number
- JP2022560417
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-12-09
- Filing Date
- 2020-12-08
- Publication Date
- 2025-10-02
- Estimated Expiration
- 2040-12-08
AI Technical Summary
Traditional methods for determining the spatial relationship between maxillary and mandibular teeth are error-prone and costly, and digital methods using 3D scanning disrupt natural jaw movements, leading to inaccurate results.
A computer-implemented method using 3D models of maxillary and mandibular teeth, combined with multiple 2D images, to determine spatial relationships through iterative alignment optimization based on cost scores derived from image features and optical flow, without requiring expensive equipment or intraoral placement.
Provides a cost-effective and accurate means to determine both static and dynamic relationships between maxillary and mandibular teeth, ensuring natural jaw movements and precise alignment for dental prostheses.
Smart Images

Figure 0007748107000001 
Figure 0007748107000002 
Figure 0007748107000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a computer-implemented method for determining spatial relationships between maxillary teeth (upper teeth) and mandibular teeth (lower teeth), and a system for determining spatial relationships between maxillary teeth (upper teeth) and mandibular teeth (lower teeth). [Background technology]
[0002] In dental treatment, it is often important to record the relationship between the maxillary (upper) and mandibular (lower) teeth. This can be a static relationship, such as maximum interdental position or other interdental positions. It is also important to record how the teeth function in movement, such as during typical chewing and other jaw movements. For example, knowledge of the relationship between the maxillary and mandibular teeth and their relative movement is helpful in fabricating dental prostheses such as crowns and bridges. This knowledge allows dental technicians to ensure that the prosthesis feels natural in the mouth and is not subject to undue stress. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Hartley, R. and Zisserman, A., 2003.Multiple view geometry in computer vision. Cambridge university press [Non-patent document 2] Fischler, MA and Bolles, RC, 1981, Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography, Communications of the ACM, 24(6), pp. 381-395 [Non-patent document 3] Walls, AWG, Wassell, RW and Steele, JG, 1991. A comparison of two methods for locating the intercuspal position (ICP) while mounting casts on an articulator. Journal of oral rehabilitation, 18(1), pp.43-48 [Non-patent document 4] Ullman, S., 1976. The Interpretation of Structure from Motion (No. AI-M-476). Massachusetts Inst of Tech Cambridge Artificial Intelligence Lab Summary of the Invention [Problem to be solved by the invention]
[0004] Traditional methods for determining these relationships can be error-prone and costly. For example, determining the static relationship between maxillary and mandibular teeth can be accomplished by clamping together two cast stone models, one from an impression of the maxillary teeth and the other from the mandibular teeth. However, this relies on the dental technician to properly position the models, and the clamping process can introduce alignment errors, such as an open bite. Similar problems can also occur when using a dental articulator.
[0005] The latest digital methods require expensive 3D scanning equipment and may not be as accurate as traditional methods, partly because dynamic motion recording requires placing instruments inside the patient's mouth, which disrupts proprioceptive feedback and results in unnatural jaw movements.
[0006] It is an object of the present disclosure to overcome the above-mentioned difficulties, as well as other difficulties that will be apparent to those skilled in the art in light of the description herein. It is a further object of the present disclosure to provide a cost-effective and accurate means for determining the static or dynamic relationship between maxillary and mandibular teeth. [Means for solving the problem]
[0007] According to the present invention there is provided an apparatus and method as defined in the accompanying claims. Other features of the invention will become apparent from the dependent claims and the description that follows.
[0008] According to a first aspect of the present disclosure, there is provided a computer-implemented method comprising: receiving a 3D model of the patient's maxillary teeth and a 3D model of the patient's mandibular teeth; receiving a plurality of 2D images (two-dimensional images), each 2D image representing at least a portion of the patient's maxillary and mandibular teeth; Determining the spatial relationship between the patient's maxillary and mandibular teeth based on the 2D image.
[0009] Determining the spatial relationship between the patient's maxillary teeth and mandibular teeth based on the 2D image can include determining an optimal alignment of the 2D image with one of the 3D models of the maxillary teeth and the 3D models of the mandibular teeth, and determining an optimal alignment between one of the 3D models of the maxillary teeth and the 3D models of the mandibular teeth and the other of the patient's 3D models of the maxillary teeth and the 3D models of the mandibular teeth.
[0010] Determining the optimal alignment may include, for each 2D image, rendering a 3D scene based on a current estimate of the camera pose at which the 2D image was captured and the spatial relationship between the 3D models of the patient's maxillary and mandibular teeth, extracting a 2D rendering from the rendered 3D scene based on the current estimate of the camera pose, and comparing the 2D rendering with the 2D image to determine a cost score indicative of the level of difference between the 2D rendering and the 2D image. Determining the optimal alignment may include iteratively obtaining the optimal alignment across all of the 2D images using an optimizer, preferably a non-linear optimizer.
[0011] The cost score may comprise a mutual information score calculated between the 2D rendering and the 2D image. The cost score may comprise a similarity score between corresponding image features extracted from the 2D rendering and the 2D image. The image feature may be a corner feature. The image feature may be an edge. The image feature may be an image gradient feature. The similarity score may be one of Euclidean distance, random sampling, or one-to-one. Before calculating the mutual information score, the rendering and the 2D image may be pre-processed. Convolution or a filter may be applied to the rendering and the 2D image.
[0012] The cost score may comprise a closest match score between the 2D rendering and the 2D image, which may be based on a comparison of 2D features from the 2D rendering and 2D features from the 2D image.
[0013] The cost score may comprise a 2D-3D-2D-2D cost (2d-3d-2d-2d cost) calculated by extracting 2D image features from the rendered 3D scene, reprojecting the extracted features onto the rendered 3D scene, extracting image features from the 2D images, reprojecting the extracted features from the 2D images onto the rendered 3D scene, and calculating a similarity measure indicating the difference between the reprojected features into 3D space.
[0014] The cost score may comprise an optical flow cost calculated by tracking pixels between successive images of the plurality of 2D images. The optical flow cost may be based on dense optical flow. The tracked pixels may exclude pixels determined to be occluded.
[0015] Determining the cost score may include determining a plurality of different cost scores based on different extracted features and / or similarity measures. The cost scores used in each iteration by the optimizer may be different. The cost score may alternate between a first selection of cost scores and a second selection of cost scores, where one of the first selection and the second selection comprises a 2D-3D-2D-2D cost and the other does not.
[0016] Each of the multiple 2D images can show the maxillary and mandibular teeth in substantially the same static alignment. The plurality of 2D images may each comprise at least a portion of the patient's upper and lower teeth. The plurality of 2D images may be captured while moving a camera around the patient's head.
[0017] Each of the plurality of 2D images can comprise at least a portion of a dental model of the patient's maxillary teeth and a dental model of the patient's mandibular teeth in an occluded state. The plurality of 2D images can be captured while moving a camera around the model of the patient's maxillary teeth and the model of the patient's mandibular teeth held in occlusion. The model of the maxillary teeth can comprise markers. The model of the mandibular teeth can comprise markers. Determining an optimal alignment of the 2D image with one of the 3D model of the maxillary teeth and the 3D model of the mandibular teeth can be based on markers placed on the dental models.
[0018] The method may include determining a first spatial relationship between the patient's maxillary and mandibular teeth in a first position, determining a second spatial relationship between the patient's maxillary and mandibular teeth in a second position, and determining a spatial transformation between the first and second spatial relationships. Determining the spatial transformation may include determining a lateral hinge axis. The first position may be a closed position. The second position may be an open position.
[0019] The plurality of 2D images may comprise a plurality of 2D image sets, each of which may comprise multiple simultaneously captured images of the patient's face from different viewpoints. The plurality of 2D images may comprise video of the maxillary and mandibular teeth in motion, captured simultaneously from different viewpoints. The method may include determining a spatial relationship based on each of the 2D image sets. The method may include determining, in each of the 2D image sets, an area of the 3D model of the maxillary teeth that is in contact with the 3D model of the mandibular teeth, and displaying the determined area on the 3D model of the maxillary or mandibular teeth.
[0020] According to a second aspect of the present disclosure, there is provided a system comprising a processor and a memory, the memory storing instructions that, when executed by the processor, cause the system to perform any of the methods defined herein.
[0021] According to a further aspect of the present invention, there is provided a tangible, non-transitory (non-transient) computer-readable storage medium having recorded thereon instructions which, when executed by a computing device, cause the computing device to configure as defined herein and / or perform any of the methods defined herein.
[0022] According to a further aspect of the present invention there is provided a computer program product comprising instructions which, when executed by a computer, cause the computer to carry out any of the methods described herein.
[0023] For a better understanding of the present invention, and to show how examples of the same may be carried into effect, reference will now be made, by way of example only, to the accompanying schematic drawings in which: [Brief explanation of the drawings]
[0024] [Figure 1]1 is a schematic flow chart of a first exemplary method for determining the spatial relationship between a patient's maxillary and mandibular teeth. [Figure 2] 1 is a schematic perspective view illustrating an exemplary method of capturing a 2D image comprising at least a portion of a patient's maxillary and mandibular teeth. [Figure 3] 3 is a schematic flow chart illustrating the exemplary method of FIGS. 1 and 2 in further detail; [Figure 4] 4 is a simplified flowchart illustrating the exemplary method of FIGS. 1-3 in further detail. [Figure 5] FIG. 5 is a schematic diagram illustrating the exemplary method of FIGS. 1-4 in further detail. [Figure 6] 1 is a schematic perspective view illustrating an exemplary method of capturing a 2D image comprising at least a portion of a patient's maxillary and mandibular teeth. [Figure 7] 4 is a schematic flow chart of an exemplary method for calculating a lateral horizontal axis. [Figure 8A] Schematic diagram showing the lateral horizontal axis of the patient. [Figure 8B] Schematic diagram showing the lateral horizontal axis of the patient. [Figure 9] 1 is a schematic diagram illustrating a method for capturing a 2D image comprising at least a portion of a patient's maxillary and mandibular teeth. [Figure 10] 10 is a schematic flow chart of a second exemplary method for determining the spatial relationship between a patient's maxillary and mandibular teeth. [Figure 11A] 10 is an exemplary GUI showing contacts between a patient's maxillary and mandibular teeth. [Figure 11B] 10 is an exemplary GUI showing contacts between a patient's maxillary and mandibular teeth. [Figure 12] 1 is a schematic block diagram of an exemplary system for determining the spatial relationship between a patient's maxillary and mandibular teeth. DETAILED DESCRIPTION OF THE INVENTION
[0025] In the drawings, corresponding reference characters indicate corresponding components. Those skilled in the art will appreciate that elements in the figures are illustrated for simplicity and clarity and have not necessarily been drawn to scale. For example, the dimensions of some elements in the figures may be exaggerated relative to other elements to improve understanding of the various illustrative embodiments. Also, common but well-understood elements that may be useful or necessary in commercially feasible examples are often not depicted so as not to overly obstruct the view of these various illustrative examples.
[0026] In summary, embodiments of the present disclosure provide a means for determining spatial relationships between maxillary and mandibular teeth based on multiple 2D images. Each image comprises at least a portion of the maxillary and mandibular teeth. The multiple 2D images may be used to align a 3D model of the maxillary teeth with a 3D model of the mandibular teeth. In some examples, the multiple 2D images comprise video captured by a camera moving around the patient's face or comprise cast stone models of the maxillary and mandibular teeth held in occlusion. In other examples, the multiple 2D images comprise multiple simultaneously captured video of the patient's face, each simultaneously captured from a different perspective.
[0027] 1 is a diagram illustrating an example of a method for determining the spatial relationship between maxillary and mandibular teeth. In the example of FIG. 1, the spatial relationship is a static spatial relationship, such as a maximum interdental position. In block S11, a 3D model of the patient's maxillary teeth is received. In block S12, a 3D model of the patient's mandibular teeth is received.
[0028] The 3D models of the upper and lower teeth can be obtained using a suitable 3D dental scanner, such as the dental scanner described in the applicant's pending UK patent application GB1913469.1. The dental scanner may be used to scan impressions taken of the upper and lower teeth, stone models cast from the impressions, or a combination of the impressions and stone models.
[0029] In a further example, the 3D model may be obtained by scanning the stone model or by other commercially available scanners in the form of intraoral scanners suitable for placement in the patient's mouth.
[0030] Each 3D model may take the form of a data file in STL or PLY format, or any other data format suitable for storing 3D models. At block S13, a plurality of 2D images are captured, each of which includes at least a portion of the maxillary teeth and a portion of the mandibular teeth held in a desired static spatial relationship.
[0031] FIG. 2 illustrates an example method for capturing multiple 2D images. As shown in FIG. 2, the multiple 2D images may be captured by a single camera 101, which may be, for example, the camera of a smartphone 100. The camera 101 may be placed in a video capture mode to capture images at a predetermined frame rate. The camera 101 is then moved in an arc A1 around the face of a patient P, who is holding the upper and lower teeth U1 and L1 in a desired static spatial relationship. The upper and lower teeth U1 and L1 are at least partially visible to the camera because the patient P is holding their lips apart or they are retracted by other standard dental means. Thus, multiple 2D images are captured, each comprising at least a portion of the patient's upper and lower teeth.
[0032] In one embodiment, the camera 101 is calibrated before capturing 2D images. In particular, the camera 101 may undergo a calibration process to determine or estimate parameters of the camera's 101 lens and image sensor, which can be used to correct for phenomena such as lens distortion and barrelling (also known as radial distortion) and enable accurate 3D scene reconstruction. The calibration process may also determine the focal length and optical center of the camera 101. The calibration process may include capturing images of an object with known dimensions and shape and estimating the lens and image sensor parameters based thereon. An exemplary method for calibrating a camera may be as described in Non-Patent Document 1, the contents of which are incorporated herein by reference.
[0033] Returning to FIG. 1 , in block S14, the spatial alignment of the maxillary and mandibular teeth is determined. The 2D images are used to align 3D models of the maxillary and mandibular teeth, such that the spatial alignment of the models corresponds to the relative alignment of the models as shown in the captured images. The method for aligning the maxillary and mandibular teeth using the 2D images is described in detail below with reference to FIGS. 3-5.
[0034] Figure 3 shows in more detail the processing of block S14 of Figure 1. This processing estimates the pose of the upper and lower teeth in each 2D image captured by the camera 101, as well as the camera pose (i.e., the position of the camera 101).
[0035] In a first step S31, the captured images are aligned with the maxillary teeth to determine the camera pose relative to the maxillary teeth. In one example, an initial guess for the camera pose of a first one of the images captured by camera 101 is received. This may be done via user input, such as by a user dragging a model via a user interface to align it with the first captured image. In another example, a standardized animation protocol may be used. For example, the user capturing the images may be instructed to start the animation with camera 101 pointed at a certain portion of the patient P's face, so that the approximate pose of the first image is known.
[0036] In the second step S32, a transformation T is performed to align the mandibular teeth L1 with the maxillary teeth U1. lower is simultaneously optimized across all captured 2D images. Thus, a determination of the alignment of the mandibular teeth L1 and the maxillary teeth U1 is made. In some examples, the camera pose and the transformation T lower are then simultaneously optimized iteratively.
[0037] Next, a method for aligning a 3D model to a 2D image will be described with reference to FIGS. In block S41, a 3D rendering of the scene S is performed based on the current estimate of the camera pose C and the model (upper teeth U1) pose, resulting in a 2D rendering image (R) of the current estimate of the scene as seen from the estimate of the camera position (C).
[0038] In block S42, the rendered image R is compared to the 2D captured image (I) to score the current estimate of the model pose and camera pose. The higher the similarity between the captured image I and the rendered image R, the higher the likelihood that the estimate is correct.
[0039] A similarity metric may be calculated to assess the similarity between the rendered image R and the captured image I. The similarity metric may output a cost score, with a higher cost score indicating a higher dissimilarity between the rendered image R and the captured image I.
[0040] In one example, a mutual information score (MI score) is calculated between the rendered image R and the captured image I. The mutual information MI may be calculated directly based on the rendered image R and the captured image I. However, in a further example, to improve the effectiveness of the mutual information MI score, the rendered image R and the captured image I may be processed in different ways before calculating the mutual information MI. For example, a filter such as a Sobel gradient process may be applied to the rendered image R and the captured image I. Alternatively, another machine-learned or engineered image convolution may be applied to the rendered image R and the captured image I.
[0041] In one example, feature extraction is performed on each of the rendered image R and the captured image I to extract salient features (e.g., regions or patches) from the rendered image R and the captured image I. For example, an edge extraction method is applied to extract edge points or regions from each of the rendered image R and the captured image I. In another example, a corner extraction method is applied to extract corner points or regions from each of the rendered image R and the captured image I. For example, image gradient features may be extracted. Alternatively, other machine-learned salient features, such as those trained by a convolutional neural network, may be extracted.
[0042] Corresponding salient features extracted from the rendered image R and the captured image I are then matched to determine the similarity between them. For example, the matching may employ a similarity measure such as the Euclidean distance between the extracted points. The similarity measure may employ a random sampling of extracted points as discussed in Non-Patent Document 2, or a measure such as one-to-one.
[0043] In a further example, the cost, referred to herein as the 2D-3D-2D-2D cost, is calculated as follows: First, 2D features (edges / corners / machine-learned or human-designed convolutions) are extracted from the 2D image (R) of the current 3D rendering guess, resulting in point set A2. Then, these 2D features are reprojected onto the 3D rendering to obtain a 3D point set (point set A3).
[0044] In the 2D camera image (captured image I), corresponding 2D features (point set B2) are extracted. Then, correspondence between B2 and A2 is assigned, typically using a nearest-point search, sometimes augmented with a nearest-neighbor similarity factor. The equivalent 3D correspondence (A3) is then calculated, resulting in a set of 2D-3D correspondences B2-A3. The problem here is to minimize the 2D reprojection error of point A3 (called A2'), which is done by optimizing the camera pose or model pose so that B2-A2' is minimized.
[0045] Conceptually, this score involves marking edge points (or other salient points) on a 3D model, finding the nearest edge points on the 2D camera image, and then moving the model so that a reprojection of these 3D edge points matches the 2D image points.
[0046] The whole process is repeated iteratively, as the edge points marked on the 3D model and the corresponding 2D points in the camera image become more and more likely to truly correspond with each iteration.
[0047] Further costs are extracted from the rendered image R (which may also have undergone preprocessing such as silhouette rendering) and from the camera image. Once the closest match between the two images is found, it is used as the current cost score and fed directly back into the optimizer, allowing for small adjustments to the model pose (or camera pose) and recomputing the rendering. This method differs from the 2D-3D-2D-2D cost in that the correspondences are updated at each iteration, and the correspondence distance is used as a score to guide the optimizer. In contrast, the 2D-3D-2D-2D cost is computed by iterating to minimize the current set of correspondences (i.e., the inner loop), then recomputing the rendering to obtain new correspondences and iterating again (the outer loop).
[0048] Another cost score may be calculated using optical flow, a technique for tracking pixels between successive frames in a sequence. As the frames are recorded sequentially, if the tracked position of a pixel in the rendered image R differs significantly from the tracked position of the pixel in the captured image I, this indicates that the model and camera pose estimation are inaccurate.
[0049] In one embodiment, optical flow prediction is performed forward, i.e., based on previous frames of the captured video. In one embodiment, optical flow prediction is performed backward, i.e., the order of frames in the video is reversed so that optical flow predictions based on subsequent frames can be calculated.
[0050] In one example, a dense optical flow technique is employed. Dense optical flow may be calculated for a subset of pixels in the image. For example, the pixels may be limited to pixels that are assumed to lie in a plane substantially orthogonal to a vector extending from the optical center of the camera based on the pixel's location. In one example, a pixel may be determined to be occluded by the face of patient P (e.g., by the nose) based on a face model fitted to the face. Such a pixel may then no longer be tracked by the optical flow.
[0051] In one example, multiple of the above methods are applied to determine multiple different cost scores based on different extracted features and / or similarity measures. The processing of blocks S41 and S42 is performed on a plurality of captured images. For example, the processing may be performed on all of the captured images. However, in other examples, a subset of the captured images may be selected. For example, a subsample of the captured images may be selected. The subsample may be a regular subsample, such as every other captured image, every third captured image, or every tenth captured image.
[0052] Therefore, based on the model-estimated position and the camera-estimated position, multiple cost scores are derived that reflect the degree of similarity between each captured image and its corresponding rendered image. In one example, the difference in estimated camera pose between consecutive frames may be calculated as an additional cost score. In other words, the distance between the estimated camera pose of the current image and the previously captured frame may be calculated, and / or the distance between the estimated camera pose of the current image and the subsequently captured frame may be calculated. Because images are captured consecutively, a large difference in camera pose between consecutive frames indicates an inaccurate estimation.
[0053] In block S43, the cost score for each image is fed to a nonlinear optimizer, which repeats the processes of blocks S41 and S42, adjusting the camera pose and model pose at each iteration. A simultaneous optimization solution is obtained for each captured image. The iterations stop when a convergence threshold is reached. In one example, the nonlinear optimizer employed is the Levenberg-Marquardt algorithm. In other examples, other algorithms may be employed from software libraries such as Eigen (http: / / eigen.tuxfamily.org / ), Ceres (http: / / ceres-solver.org / ), or nlopt (https: / / nlopt.readthedocs.io).
[0054] In a further example, the cost score used in each iteration may be different. In other words, a selection of the scores described above is calculated in each iteration, and the selection may be different for each iteration. For example, the method may alternate between a first selection of the scores and a second selection of the scores. In one example, one of the first and second selections comprises a 2D-3D-2D-2D cost, while the other does not. Using different cost metrics in the above method may result in a more robust solution.
[0055] FIG. 6 illustrates a further embodiment of a method for determining the spatial relationship between maxillary and mandibular teeth. FIG. 6 substantially corresponds to the method described above with reference to FIGS. 1 to 5. However, instead of passing a camera 101 around the face of patient P, the camera 101 is moved around a model of maxillary teeth U2 and a model of mandibular teeth L2. The models U2 and L2 are cast from dental impressions (impressions) of the patient P's maxillary teeth U1 and mandibular teeth L1. Furthermore, the models U2 and L2 are held in occlusion by a dental technician's hand H. This method takes advantage of the fact that manual occlusion of interdental positions by an experienced practitioner is relatively accurate (Non-Patent Document 3). Furthermore, the teeth of the models U2 and L2 are not obscured by the lips, which may facilitate accurate alignment.
[0056] In one embodiment, one or both of the dental models (U2, L2) may be provided with markers. The markers may take the form of, for example, QR codes or colored marks. The markers do not have to form part of the scanned 3D model. Thus, the markers can be applied to the dental models (U2, L2) after they have been scanned. The markers are then used to determine the camera pose relative to the model using standard "structure-from-motion" techniques. For example, the technique may be that described in [4]. This technique can therefore replace the first step S31 described above.
[0057] FIG. 7 shows an example of a method for determining a patient's transverse horizontal axis (THA). This is shown in FIGS. 8A and 8B, which show the patient's skull 10 in a jaw-closed position and an jaw-open position, respectively. The transverse horizontal axis THA 11 is the hinge axis about which the mandible may rotate during pure rotational opening and closing of the mouth. In FIG. 8B, an appliance 13 has been inserted between the maxillary teeth U1 and mandibular teeth L1 to statically hold them apart. The appliance 13 may be an anterior fixture or a deployment device.
[0058] In block S71 of Figure 7, the spatial relationship between the maxillary teeth U1 and the mandibular teeth L1 is determined with the jaws 12 in a substantially closed position, as shown in Figure 8A. In block S72, the spatial relationship between the maxillary teeth U1 and the mandibular teeth L1 is determined with the jaws 12 in an open position (maximum 20 mm open), as shown in Figure 8B, for example. In each case, the mandible may be centrically oriented.
[0059] In block S73, a lateral horizontal axis THA is determined based on the different positions of the mandibular teeth L1 relative to the maxillary teeth U1 in the two spatial alignments. In one example, the lateral horizontal axis THA is determined by aligning multiple points (e.g., three points) on both scans of the maxillary tooth U1 to create a common coordinate system. A 3x4 transformation matrix is then calculated that represents the movement of multiple points on the scan of the mandibular tooth L1 from a closed position to an open position. The upper left 3x3 submatrix of the transformation matrix can then be used to calculate the orientation of the rotation axis and the degree of rotation about that axis.
[0060] There are an infinite number of spatial locations for this axis, and each requires a different translation vector after rotation. To find the location of the transverse horizontal axis THA, a point is traced from the start position to the end position. A solution is then found that constrains the translation vector to a direction consistent with the calculated axis. This solution can be found using closed-form mathematics (e.g., singular value decomposition) or other open-form optimization methods such as gradient descent. This yields a vector corresponding to the point on the axis and the magnitude of the translation vector along the axis. The latter vector should ideally be zero.
[0061] In a further example, the spatial alignment of the maxillary and mandibular teeth can be determined for multiple positions, for example, by using jigs or deployments (deprogrammers) of different thicknesses to hold the jaws at different degrees of opening, so that multiple spatial alignments can be used to calculate the hinge axis.
[0062] In another embodiment, the spatial relationship between the maxillary tooth U1 and the mandibular tooth L1 is determined in a protrusive or lateral static registration. From such a registration, anterior or lateral guidance can be inferred. Therefore, dynamic jaw movement estimations can be made based on determining the spatial alignment of the teeth in various static positions.
[0063] 9, another embodiment of the present disclosure is illustrated. In this embodiment, two cameras 200A, 200B are employed. In contrast to the above-described embodiment, the cameras 200A, 200B are provided in fixed positions. That is, the cameras 200A, 200B are fixed relative to each other and do not move during image capture. For example, the cameras 200A, 200B may be supported in fixed positions by a suitable stand, jig, or other support mechanism (not shown).
[0064] Cameras 200A and 200B are positioned so that each camera has a different viewpoint and therefore captures images of different parts of the mouth. The line of sight of each camera 200 is represented by dotted lines 201A and 201B. For example, in FIG. 9, camera 200A is positioned to capture the right side of patient P's face. Camera 200B is positioned to capture the left side of patient P's face.
[0065] The cameras 200A and 200B are configured to operate synchronously, i.e., the cameras 200A and 200B are capable of capturing images simultaneously at specific time indexes, thus providing two simultaneously captured perspectives of the patient P's mouth.
[0066] FIG. 10 illustrates how cameras 200A, 200B can be used to determine the spatial relationship between, for example, upper and lower teeth U1 and L1. In block S1001, cameras 200A and 200B capture synchronized images. In block S1002, the spatial alignment of maxillary tooth U1 and mandibular tooth L1 is calculated for a time index based on the captured images at that time index.
[0067] 1-6, it is assumed that the maxillary teeth U1 and mandibular teeth L1 remain in the same spatial alignment in all images captured by camera 101. Thus, an optimal alignment solution is found based on multiple images captured by camera 101 as it passes around the patient's face or by a model held in occlusion.
[0068] In contrast, in the method of Figures 9-10, the simultaneous capture of each set of synchronously captured images by cameras 200A and 200B necessarily results in the maxillary tooth U1 and mandibular tooth L2 appearing in the same spatial alignment. Therefore, the method described above in connection with Figures 1-6 may be applied to each set of synchronously captured images to determine alignment for each time index.
[0069] This allows for dynamic determination of the spatial alignment between the maxillary tooth U1 and the mandibular tooth L2. Obtaining and combining two (or more) sets of costs per frame can improve the robustness and convergence properties of the optimization. Therefore, only one snapshot of the teeth (i.e., from two or more simultaneous images) can robustly align the 3D scan. Therefore, by capturing a video of the maxillary tooth U1 and mandibular tooth L2 moving (e.g., while chewing), a record of the interaction between the maxillary tooth U1 and the mandibular tooth L2 can be derived. Using this record, dental technicians can ensure that prosthetic devices, such as crowns, are designed to harmonize with the patient's chewing movements.
[0070] In one example, contact between maxillary tooth U1 and mandibular tooth L2 at each time index may be displayed to a user via a GUI. For example, FIG. 11A shows a GUI 300 displaying a 3D model 301 of mandibular tooth L1. A region 302 of mandibular tooth L1 that is in contact with maxillary tooth U1 at a particular time index T1 is highlighted. For example, region 302 may be displayed in a different color on the GUI 300. FIG. 11B shows the same 3D model 301 at a different time index T2, where there is now more contact between mandibular tooth L1 and maxillary tooth U1. The GUI 300 may sequentially display the contact between maxillary tooth U1 and mandibular tooth L1 at each time index, thereby displaying an animated video showing the contact between maxillary tooth U1 and mandibular tooth L1 throughout the captured image. Therefore, a dental technician can easily understand the contact between the maxillary tooth U1 / mandibular tooth L1 during chewing, for example, and can assist the user in designing an appropriate prosthesis. Furthermore, the data can be imported into a dental CAD package to enable diagnosis and prosthesis design that is in tune with the patient's jaw movements.
[0071] In the above example, two cameras 200A, 200B are used, but in other examples, more cameras may be positioned around the mouth of patient P and configured to capture images synchronously. In such examples, the method outlined above in relation to Figures 1-6 may be applied to all images captured by the multiple cameras at a particular time index.
[0072] FIG. 12 illustrates an exemplary system 400 for determining the spatial relationship between maxillary and mandibular teeth. The system 400 includes a controller 410. The controller may include a processor or other computing element, such as a field programmable gate array (FPGA), a graphics processing unit (GPU), or the like. In some examples, the controller 410 includes multiple computing elements. The system also includes a memory 420. The memory 420 may include any suitable storage device for temporarily or permanently storing any information necessary for the operation of the system 400. The memory 420 may include random access memory (RAM), read-only memory (ROM), flash storage, a disk drive, or any other type of suitable storage medium. The memory 420 stores instructions that, when executed by the controller 410, cause the system 400 to perform any of the methods described herein. In a further example, the system 400 includes a user interface 430, which may include a display and input means, such as a mouse and keyboard or a touchscreen. The user interface 430 may be configured to display the GUI 300. In a further example, system 400 may be configured with multiple computing devices, each with a controller and memory, connected via a network.
[0073] Various modifications can be made to the examples discussed herein within the scope of the present invention. For example, it will be understood that the order in which the 3D model and 2D images are received may be changed. For example, the 2D images may be captured before scanning the 3D model. The above example discusses determining the positions of the maxillary teeth relative to the camera and aligning the mandibular teeth to the determined positions of the maxillary teeth. However, it will be understood that the positions of the mandibular teeth may instead be determined relative to the camera before aligning the maxillary teeth with the mandibular teeth.
[0074] Although reference is made herein to the camera 101 of the smartphone 100, it will be appreciated that any suitable mobile camera may be employed. Advantageously, the above-described systems and methods provide a way to determine the alignment of upper and lower teeth that utilizes prior knowledge of tooth shape obtained through 3D scanning. This allows spatial relationships to be determined based on video captured by a camera, such as those found on smartphones. Therefore, no expensive specialized equipment beyond that used by general dentists is required. Furthermore, the above-described methods and systems advantageously do not require intraoral placement of equipment, thereby allowing the patient to move their jaws naturally.
[0075] At least some of the embodiments described herein may be constructed, in part or in whole, using dedicated, special-purpose hardware. As used herein, terms such as "component," "module," or "unit" may comprise, but are not limited to, hardware devices, such as circuits in the form of discrete or integrated components, field programmable gate arrays (FPGAs), or application-specific integrated circuits (ASICs), that perform particular tasks or provide related functionality. In some examples, the described elements may be configured to reside on tangible, persistent, addressable storage media and configured to execute on one or more processors. These functional elements may, in some examples, comprise components, processes, functions, attributes, procedures, subroutines, program code segments, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables, such as software components, object-oriented software components, class components, and task components, by way of example. While exemplary embodiments are described with reference to components, modules, and units discussed herein, such functional elements may be combined into fewer elements or separated into additional elements. While various combinations of optional features have been described herein, it will be understood that the described features may be combined in any suitable combination. In particular, features of any one embodiment may be combined with features of any other embodiment, as appropriate, except where such combinations are mutually exclusive. Throughout this specification, the terms "comprising" or "consisting of" mean comprising the specified component(s), but not excluding the presence of other components.
[0076] Attention is directed to all articles and documents related to this application that have been filed contemporaneously or previously hereto and are published herewith, and the contents of all such articles and documents are incorporated herein by reference.
[0077] All of the features disclosed in this specification (including the accompanying claims, abstract, and drawings), and / or all of the steps of any method or process so disclosed, may be combined in any combination, except combinations in which at least some of such features and / or steps are mutually exclusive.
[0078] Each feature disclosed in this specification (including the accompanying claims, abstract, and drawings), unless expressly stated otherwise, may be replaced by alternative features serving the same, equivalent, or similar purpose. Thus, unless expressly stated otherwise, each feature disclosed is only an example of a generic series of equivalent or similar features.
[0079] The invention is not limited to the details of the foregoing embodiments, and extends to any novel or novel combination of features disclosed in this specification (including the accompanying claims, abstract and drawings), or any novel or novel combination of method or process steps so disclosed.
Claims
1. 1. A computer-implemented method, the method comprising: receiving a 3D model of a patient's maxillary teeth and a 3D model of the patient's mandibular teeth; receiving a plurality of 2D images, each 2D image representing a position of the patient's maxillary teeth and mandibular teeth; performing a current estimation of a camera pose of a camera that captured one of the plurality of 2D images by aligning the 2D image of the maxillary teeth to a 3D model of the maxillary teeth using the set of 2D-3D correspondences obtained using feature-based recognition; determining a spatial relationship between the patient's maxillary teeth and mandibular teeth based on the 2D images; It is equipped with the plurality of 2D images comprising a plurality of 2D image sets; each said 2D image set comprising a plurality of simultaneously captured images of the patient's face from a plurality of different viewpoints; and The method further comprises determining the spatial relationship based on each of the sets of 2D images; determining the spatial relationship between the patient's maxillary teeth and mandibular teeth based on the 2D images, determining an optimal alignment of the 2D image with one of the 3D model of the maxillary teeth and the 3D model of the mandibular teeth; determining an optimal alignment between the one of the 3D model of the maxillary teeth and the 3D model of the mandibular teeth and the other of the 3D model of the maxillary teeth and the 3D model of the mandibular teeth of the patient; It is equipped with The step of determining the optimal alignment comprises, for each of the 2D images: rendering a 3D scene based on a current estimate of the camera pose at which the 2D image was captured and based on the spatial relationship between the 3D model of the maxillary teeth and the 3D model of the mandibular teeth of the patient; extracting a 2D rendering from the rendered 3D scene based on the current estimate of the camera pose; and comparing the 2D rendering to the 2D image to determine a cost score indicative of a level of difference between the 2D rendering and the 2D image; It is equipped with The cost score comprises a 2D-3D-2D-2D cost, the 2D-3D-2D-2D cost being: extracting 2D image features from the rendered 3D scene; reprojecting the extracted 2D image features onto the rendered 3D scene; extracting image features from the 2D image; reprojecting the image features extracted from the 2D image onto the rendered 3D scene; computing a similarity measure indicative of the differences between the reprojected features in 3D space; is calculated by method.
2. The method further comprises iteratively obtaining the optimal alignment across all of the 2D images using a non-linear optimizer. The method of claim 1.
3. the cost score comprises a mutual information score calculated between the 2D rendering and the 2D image.
3. The method according to claim 1 or 2.
4. the cost score comprises a similarity score between corresponding image features extracted from the 2D rendering and the 2D image. The method according to any one of claims 1 to 3.
5. the image feature is a corner feature, an edge feature, or an image gradient feature; The method of claim 4.
6. the similarity score is one of Euclidean distance, random sampling, or one-to-one; 6. The method according to claim 4 or 5.
7. the cost score comprises an optical flow cost calculated by tracking pixels between successive images of the plurality of 2D images. The method according to any one of claims 1 to 6.
8. The method further comprises determining a plurality of different cost scores based on at least one of the different extracted features and the similarity measure. The method according to any one of claims 1 to 7.
9. The cost scores used by the nonlinear optimizer in each iteration are different. The method of claim 2.
10. each of the plurality of 2D images showing the maxillary teeth and the mandibular teeth in substantially the same static alignment; The method according to any one of claims 1 to 9.
11. each of the plurality of 2D images comprising the maxillary teeth and the mandibular teeth of the patient; The method according to any one of claims 1 to 10.
12. each of the plurality of 2D images comprises a dental model of the patient's upper teeth and a dental model of the patient's lower teeth in occlusion; The method according to any one of claims 1 to 10.
13. The method further comprises determining an optimal alignment of the 2D image with one of the 3D model of the maxillary teeth and the 3D model of the mandibular teeth based on markers placed on a corresponding dental model. The method of claim 12.
14. The method further comprises: determining a first spatial relationship between the patient's maxillary teeth and mandibular teeth in a first position; determining a second spatial relationship between the patient's maxillary teeth and mandibular teeth in a second position; determining a spatial transformation between the first spatial relationship and the second spatial relationship; The method according to any one of claims 1 to 13, comprising:
15. determining the spatial transformation comprises determining a lateral hinge axis, the lateral hinge axis being the hinge axis about which the mandible may rotate during pure rotational opening and closing of the mouth; 15. The method of claim 14.
16. The method further comprises: determining, in each of the 2D image sets, areas of the 3D model of the maxillary teeth that are in contact with the 3D model of the mandibular teeth; displaying the region on the 3D model of the maxillary teeth or the mandibular teeth; The method according to any one of claims 1 to 15, comprising:
17. 1. A system comprising a processor and a memory, The memory stores computer readable instructions that, when executed by the processor, cause the system to perform the method of any one of claims 1 to 16. system.
18. - recorded with instructions which, when executed by a computing device, cause said computing device to perform the method of any one of claims 1 to 16; A tangible, non-transitory computer-readable storage medium.
Citation Information
Patent Citations
Method for matching a three-dimensional model of a patient's dentition to an image of the patient's face recorded by a camera - Patent Application 20070122999
JP2021514232A
Generating a virtual depiction of an orthodontic treatment of a patient
US20180263731A1
Method for aligning a three-dimensional model of a dentition of a patient to an image of the face of the patient
WO2019162164A1