Orthodontic treatment simulated with real-time augmented visualization
The integration of computer vision and augmented reality allows for real-time simulation of orthodontic treatments, addressing the lack of medical applications for augmented reality filters and enhancing treatment planning and visualization.
Patent Information
- Application Number
- JP2020563722
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-02-25
- Filing Date
- 2019-05-02
- Publication Date
- 2025-09-25
- Estimated Expiration
- 2039-05-02
AI Technical Summary
Augmented reality face filters are primarily used for entertainment purposes and lack applications in medical or dental fields, specifically for simulating orthodontic treatments.
A system and method using computer vision and augmented reality to overlay virtual orthodontic appliances and anatomical modifications onto a user's real-world view, providing real-time simulations of orthodontic treatments and treatment outcomes.
Enables patients and professionals to visualize and plan orthodontic treatments in real-time, allowing for informed decision-making and treatment planning without the need for initial professional visits.
Smart Images

Figure 0007744132000001 
Figure 0007744132000002 
Figure 0007744132000003
Abstract
Description
[Background technology]
[0001] The use of augmented reality face filters, such as in SNAPCHAT or INSTAGRAM products, is becoming increasingly popular. However, today these filters are primarily used in popular culture for entertainment purposes only, such as to visually apply depictions of hair, makeup, glasses, beards, and hats to a subject's face or head in real time in video augmentation. There is a need to extend this technology to use in medical or dental applications. Summary of the Invention
[0002] A method for simulating orthodontic treatment according to one embodiment of the present invention includes receiving an electronic image of a user's face and identifying an area of interest in the image that includes the user's teeth. Virtual orthodontic appliances are placed on the user's teeth in the image or detected appliances are removed, and the image of the user with the virtual orthodontic appliances or without the detected appliances is displayed on an electronic display device. The method is performed in real time to provide the user with an augmented image simulating treatment as or shortly after the image of the user is received.
[0003] A system for simulating orthodontic treatment according to one embodiment of the present invention includes a camera providing electronic digital images or video, an electronic display device, and a processor. The processor is configured to receive electronic images of a user's face from the camera, identify regions of interest in the images that include the user's teeth, place virtual orthodontic appliances on the user's teeth in the images or remove detected appliances, and display the images of the user with or without the virtual orthodontic appliances on the electronic display device. The processor operates in real time to provide the user with augmented images simulating treatment as or shortly after the images of the user are received from the camera.
[0004] Another method for simulating orthodontic treatment according to an embodiment of the present invention includes receiving an electronic image of a user's face and acquiring an electronic image or model of the user's facial anatomy. The method also includes identifying regions of interest in the image that include the user's teeth and placing virtual orthodontic appliances on the user's teeth in the image. The image of the user with the virtual orthodontic appliances and the image or model of the user's facial anatomy is displayed on an electronic display device. The method is performed in real time to provide the user with an augmented image simulating treatment as or shortly after the image of the user is received.
[0005] Another system for simulating orthodontic treatment according to an embodiment of the present invention includes a camera providing electronic digital images or video, an electronic display device, and a processor. The processor is configured to receive an electronic image of a user's face, acquire an electronic image or model of the user's facial anatomy, identify regions of interest in the image including the user's teeth, position virtual orthodontic appliances on the user's teeth in the image, and display the image of the user with the virtual orthodontic appliances and the image or model of the user's facial anatomy on the electronic display device. The processor operates in real time to provide the user with augmented images simulating treatment as the user's image is received from the camera or shortly thereafter. [Brief explanation of the drawings]
[0006] The accompanying drawings, which are incorporated in and constitute a part of this specification, and together with the description, serve to explain the advantages and principles of the present invention. [Figure 1] FIG. 1 is a diagram of a system for simulating orthodontic treatment. [Figure 2] 1 is a flowchart of a method for simulating orthodontic treatment. [Figure 3] 1 is a representation of an image showing an example of a detected face. [Figure 4]1 is a representation of an image showing facial landmarking. [Figure 5] 1 is a representation of an image showing facial pose estimation. [Figure 6] The cropped region of interest is shown. [Figure 7A] 13 shows adaptive thresholding applied to the inverse channel. [Figure 7B] 13 shows adaptive thresholding applied to the inverse channel. [Figure 7C] 13 shows adaptive thresholding applied to the inverse channel. [Figure 8A] 10 shows adaptive thresholding applied to the luminance channel. [Figure 8B] 10 shows adaptive thresholding applied to the luminance channel. [Figure 8C] 10 shows adaptive thresholding applied to the luminance channel. [Figure 9A] The opening and closing of the region of interest are shown. [Figure 9B] The opening and closing of the region of interest are shown. [Figure 10] Indicates the detected edges. [Figure 11] Shows closed contours with detected edges. [Figure 12] A virtual rendering of the treatment is shown. [Figure 13] Shows expansion of treatment. [Figure 14A] 10 is a user interface showing treatment extension within a single image display. [Figure 14B] 1 is a user interface showing a 3D model augmented with a 3D orthosis. [Figure 14C] 1 is a user interface showing a 2D image augmented with appliances and anatomical structures. [Figure 14D] 1 is a user interface showing a 3D model augmented with 3D orthotics and anatomical structures. DETAILED DESCRIPTION OF THE INVENTION
[0007] Embodiments of the invention include using computer vision and augmented reality techniques to overlay computer-generated imagery onto a user's view of the real world to provide one or a combination of the following views: 1. Adjustment of orthodontic appliances such as lingual brackets, labial brackets, aligners, braces, or retainers. 2. Addition or alteration of dental restorations such as crowns, bridges, inlays, onlays, veneers, dentures, or gums. 3. Modification or alteration of natural tooth anatomy, such as by using tooth whitening agents, reducing or strengthening cusp tips or incisal edges, or predicting the appearance of teeth after eruption (possibly including ectopic eruption) in children. 4. Modification of craniofacial structure by oral or maxillofacial surgery such as dentoalveolar surgery, dental and maxillofacial implants, cosmetic surgery, and mandibular surgery. 5. Addition or modification of maxillofacial, ocular, or craniofacial prostheses. 6. The results of orthodontic and / or restorative treatment to include the planned position and shape of the dental anatomy, i.e., teeth and gingiva. 7. The predicted consequences of not receiving orthodontic and / or restorative treatment, possibly indicating adverse consequences such as malocclusion, tooth wear, gum recession, periodontal disease, bone loss, caries, or tooth loss. 8. Modification of facial soft tissues, such as by botulinum toxin injection or more invasive cosmetic and plastic surgery, such as, for example, rhinoplasty, chin enhancement, cheek enhancement, wrinkle reduction, eyelid lift, neck lift, brow lift, cleft palate repair, burn repair, and scar repair. 9. Variable opacity overlay of staged treatment to compare scheduled and actual progress of treatment.
[0008] This technology can be used both as a tool for treatment planning by medical / dental professionals and for case presentation to patients. Other intended uses include exploration of treatment options by patients who want to visualize the appearance of their face during or after treatment, perhaps toggling or transitioning between before and after appearances for comparison. In some cases, patients may be able to consider treatment scenarios using only superficial image data without first visiting a medical or dental professional. In other cases, augmentation may require additional data, such as x-ray or cone beam computed tomography (CBCT) scan data, bite alignment data, three-dimensional (3D) oral scan data, virtual joint data, or calibrated measurements, to generate a treatment plan that more accurately depicts the patient after treatment. Therefore, additional time and effort may be required to organize and consolidate datasets, explore treatment options, perform measurements and analysis, and plan treatment steps before a reliable depiction of the treatment outcome can be presented to the patient. In some cases, two or more treatment scenarios may be generated and presented to the patient, thereby giving the patient a somewhat realistic expectation of treatment outcome. This can be useful in assisting the patient (or physician) in deciding on a course of action. In some cases, the cost of treatment may be weighted against the quality of the outcome. For example, a patient may forgo an expensive treatment for a slight improvement in appearance over a less expensive treatment. Similar rules may apply to the duration of treatment if time is treated like a cost.
[0009] Systems and methods FIG. 1 is a diagram of a system 10 for simulating orthodontic treatment. The system 10 includes a camera 12, a display device 14, an output device 18, a storage device 19, an input device 20, and a processor 16 electronically connected thereto. The camera 12 may be implemented, for example, by a digital image sensor providing electronic digital images or video. The display device 14 may be implemented by any electronic display, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a light emitting diode (LED) display, or an organic light emitting diode (OLED) display. The output device 18 may be implemented by any device for outputting information, such as a speaker, an LED indicator, or an auxiliary LCD display module. The input device 20 may be implemented by any device for inputting information or commands, such as a keyboard, a microphone, a cursor control device, or a touchscreen. The storage device 19 may be implemented by any electronic device for storing information, such as a hard drive, a solid-state drive, or a flash drive. Storage device 19 may optionally be located external to system 10 and accessed via a network connection. System 10 may be implemented, for example, by a desktop, notebook, or tablet computer, or a mobile phone.
[0010] 2 is a flowchart of a method 22 for simulating orthodontic treatment. The method may be implemented in software or firmware for execution by a processor, such as processor 16. The method is used to visualize the orthodontic treatment simulation by augmentation, i.e., by virtual overlay on captured real-world imagery, possibly involving real-time manipulation of an estimated 3D representation of a face captured in one or more video frames.
[0011] The following steps of method 22, described in more detail below, namely, face detection (step 24), full face landmarking (step 26), facial pose estimation (step 28), cropping the region of interest using facial landmarks (step 30), optional 3D representation of the face or region of interest (step 32), simulation (step 34), optional addition of an image or 3D representation of the face (step 35), and visualization (step 36), are used to generate an exemplary extended orthodontic appliance simulation starting from a single video frame (image).
[0012] Face Detection(24) This step uses a pre-trained model, such as a Viola-Jones model, a Local Binary Pattern Histogram (LBPH) cascade, or a model trained by deep learning, to find faces in video frames from camera 12. Once a face is detected, a facial context may be created to continue tracking the face across frames. This facial context may include key features used for reprojection or previous frame data, such as a point cloud obtained across previous frames. FIG. 3 is a representation of an image showing an example of a detected face received from camera 12, with a box in the image indicating the detection area. The image in FIG. 3 can optionally be an image of a user with orthodontic appliances, to virtually remove appliances rather than virtually add them.
[0013] Full face landmarking(26) Landmarks are interest points that describe important facial structures and can be found using shape prediction models, which may use SVM (support vector machine), PCA (principal component analysis), or other machine learning algorithms such as deep learning and random forests. These landmarks are used to crop the face or region of interest from the background, and in this example, will be used to roughly estimate the facial pose using a generalized 3D landmark model. They can also be used in the following steps to morph the generalized face mesh to more accurately represent the face detected in the video. Figure 4 is a representation of an image showing facial landmarking, where the detected area of the image in Figure 3 includes dots or other indicia for landmarking.
[0014] Facial Pose Estimation(28) Using the detected facial landmarks and a generalized 3D landmark model, a camera extrinsic estimation model (transformation) can be obtained by solving the general Perspective N-Point (PNP) problem. The PNP problem solution assumes that the detected landmarks are not affected or distorted by camera lens distortion (radial / tangential lens distortion). Instead, to obtain a more accurate camera extrinsic model, Structure from Motion (SfM) techniques can be applied to the static (non-moving) parts of the face once the following feature matching steps have been completed for multiple consecutive frames. Knowing the camera extrinsic features provides a good 3D point representation of each landmark, which can be used to help point out the pose and for more accurate augmented reality. This allows for more accurate cropping of the region of interest for augmented reality, but is not a required step. Figure 5 shows an image showing the face pose estimation as indicated by the two boxes on the image in Figure 3.
[0015] Region of interest cropping using facial landmarks (30) This example would crop the mouth region, with the inner edge of the lips representing the boundary of the cropped region. This could be the entire face or another area of the face if the simulation includes a larger or different area of the face, such as the chin. Figure 6 shows the cropped region of interest from the images of Figures 3-5.
[0016] 3D display of face or region of interest (32) This is an optional step and is only required for simulations that require a 3D model of the face or facial region, such as structural manipulation or enhanced precision enhancement. Either of the two methods or a combination of both can be used.
[0017] Method 1. The first method that can be employed is to generate a 3D mesh across multiple video frames using reprojection and mesh reconstruction. Reprojection is the process of finding the depth of features in one image and turning the features into a 3D point cloud. Mesh reconstruction is a technique used to generate a triangular mesh from that point cloud. This assumes two or more images of the same face, taken from slightly different angles and positions. Below are the steps for Method 1: i. Feature / keypoint detection. These are points (corners) on the image that stand out from other pixels in the image and are likely to be present in another image of the same scene at a different angle. ii. Feature Filtering. This is an optional step that may need to be performed if the list of detected features is too large. There are various techniques to prune the list of features to obtain a smaller list of the strongest features (most salient features). iii. Feature Matching. Each frame would have gone through the steps above and would have had its own list of features. This is where features from one image match features in the other image. Not all features will match, and those that do not match are discarded. Once features match, the list of features in the second image would be sorted in the same order as the features from the first image. iv. Rectification. The features in each image are adjusted using the camera extrinsic model so that they are row-aligned in the same plane. This means that a feature from one image is in the same x-axis row as the same feature in another image. v. Calculate the disparity map. The disparity map is a list of the disparities (distances) on the x-axis between matching points. The distances are measured in pixels. vi. Triangulation. Using the camera geometry, the method can calculate the epipolar shape and the fundamental and fundamental matrices needed to calculate the depth of each feature in an image. vii. Mesh Reconstruction. Turn the 3D point cloud into a 3D mesh for the simulation step. This may also include finding texture coordinates by ray tracing based on camera parameters (e.g., camera position).
[0018] Method 2. This method involves a parameterization of the face or region of interest (e.g., a NURBS surface or Beazier surface) or a generalized polygonal mesh (i.e., triangles or quadrilaterals), which is morphed, expanded, or stretched to best fit either the facial point cloud (see Method 1) or a previously obtained set of landmarks. When using a generalized mesh of the face or region of interest, vertices in the mesh can be assigned weights relative to the landmarks. When landmarks are found in the image, the vertices of the generalized mesh will be attracted to those landmarks based on the given vertex weights and facial pose. Alternatively, the surface topology can be predefined only abstractly, such as by graph-theoretic relationships between landmarks or between regions outlined by lines between landmarks. Once 3D points are obtained through feature recognition (landmark identification) and photogrammetric triangulation, a NURBS surface or mesh can be generated so that the landmarks coincide with the corners of the surface patches. Of course, additional points may necessarily be captured in the regions between the landmarks, and these points can serve to more precisely define the parameters of the surface, i.e., the control points and polynomial coefficients in the case of NURBS, or the intermediate vertices in the case of tessellated mesh surfaces. Distinguishing features of the face can serve as landmarks or control points for the mesh or parametric model. Therefore, increasing the resolution of the video image to the point where these fine features become visible and recognizable can improve the accuracy of the model.
[0019] Simulation (34) Once the above steps are complete, a virtual treatment can be applied, either by augmentation or 3D shape manipulation. This may include precisely identifying the location or area where the simulation will be applied. In this example, augmentation is performed to estimate the location on the region of interest where the appliance should be placed, and then the rendered orthodontic appliance is overlaid on the region of interest, or the detected appliance is virtually removed from the region of interest. For treatments involving 3D shape manipulation, additional steps may need to be performed, such as filling holes in the image caused by morphing the 3D shape. The rendered orthodontic appliance can be represented by a 2D image or a 3D model.
[0020] In this example, various image processing techniques are used on the region of interest to segment and identify the teeth, including region of interest enlargement, noise reduction, application of segmentation algorithms such as mean shift segmentation, histogram equalization, adaptive thresholding for specific channels of multiple color spaces, erosion and dilation (opening / closing), edge detection, and contour searching.
[0021] The region of interest is first expanded and transformed into a different color space, and an adaptive threshold is applied to the channel that best distinguishes teeth from non-teeth. Figures 7A-7C show adaptive thresholding applied to the inverted green-red channel of the region of interest in lab color space, i.e., the inverted channel "a" (green-red) in lab color space (Figure 7A), the mask after adaptive thresholding has been applied (Figure 7B), and the color image of the region of interest with the mask applied (Figure 7C). Figures 8A-8C show adaptive thresholding applied to the luminance channel of the region of interest in YUV color space, i.e., the channel "Y" (luminance) in YUV color space (Figure 8A), the mask after adaptive thresholding combined with the mask in Lab color space has been applied (Figure 8B), and the final color image of the region of interest with the combined mask applied (Figure 8C).
[0022] Once a mask has been created by adaptive thresholding of the color space channels as described above, opening (erosion followed by dilation) and closing (dilation followed by erosion) can be used to further segment the teeth and remove noise. Figures 9A and 9B show the opening and closing of the region of interest mask after color space thresholding, respectively.
[0023] Edges are detected from the mask using techniques such as Canny Edge Detection. Figure 10 shows the edges detected by canny edge detection.
[0024] A contour can be generated from the detected edges, allowing individual teeth or groups of teeth to be analyzed or treated as a single object. Figure 11 shows a closed contour derived from the detected edges and overlaid on a color region of interest.
[0025] From the analysis of the contours and the approximate facial pose, the treatment is rendered with the correct orientation and scale and directly overlaid on the region of interest. The original image can be analyzed to determine an approximate lighting model to apply to the rendered treatment, so that the treatment extension will fit more naturally into the scene. Figure 12 shows a virtual rendering of the treatment, which in this example is the brackets and archwire. Optionally, the virtual rendering can include removing detected appliances such as brackets and archwires.
[0026] The rendered treatment is extended to the region of interest. Post-processing, such as Gaussian blurring, can be performed to blend the extended treatment with the original image and make the final image look more natural, as if the extended treatment were part of the original image. FIG. 13 shows the extension of the treatment, with virtual brackets and archwires overlaid on the teeth in the region of interest. Optionally, the extension can include removing detected appliances and displaying the result. The extension can be achieved by directly aligning the tooth endpoints with the mouth endpoints, without explicitly segmenting the individual teeth and contours. However, approximately locating the center of each tooth will provide a more accurate / realistic result.
[0027] Add an image or 3D model (35) Another approach to augmenting facial images with appliances, restorations, or modified anatomical structures involves another aspect of 3D modeling. The above approach uses 3D modeling to determine the position, orientation, and scale of the face, which can be mathematically described by a 3D affine transformation in a virtual universe. Corresponding transformations are then applied to the dental appliances to position them on the patient's teeth in a somewhat realistic position and orientation, but 2D image analysis may ultimately be used to find the facial axis (FA) points of the teeth where brackets will be placed.
[0028] In another approach, a video camera may be used as a kind of 3D scanner to capture multiple 2D images of a person's face from different perspectives (i.e., viewpoints and viewpoint vectors, which, together with the image plane, form a series of separate view frustums). Using 3D photogrammetry, a 3D model of a person's face (and head) may be generated, initially in the form of a point cloud and later in the form of a triangular mesh. Accompanying these techniques is UV mapping (or texture mapping), which applies color values of pixels in the 2D image to 3D vertices or triangles in the mesh. Thus, a realistic-looking and reasonably accurate 3D model of a person's face may be generated in a virtual universe. An image of the 3D model may then be rendered onto a 2D plane positioned somewhere in the virtual universe according to the view frustum. Suitable rendering methods include polygonal rendering or ray tracing. The color values from the UV map may be applied in rendering to produce a more realistic color representation of the face (as opposed to a monochromatic mesh model, which is shaded only by the orientation of the triangle relative to both the viewpoint and each light source in the scene).
[0029] At this point, the 3D model of the face can be combined or integrated (or simply registered) with other 3D models or images, optionally obtained from different scanning sources, such as an intraoral scanner, a cone-beam CT (CBCT or 3D X-ray) scanner, an MRI scanner, a 3D stereoscopic camera, etc. Such data will likely have been previously obtained and used in treatment planning. Such other models or images can be stored in and retrieved from data storage 19, for example.
[0030] In some cases, video images serve as the only source of superficial 3D images, and therefore may capture facial soft tissues, making any dental anatomy visible. In other cases, soft tissue data may be captured in CBCT scans (without color information) or 3D volumetric images. Nevertheless, soft tissue data is an important element in modeling certain treatment plans, such as those involving mandibular surgery, occlusion class correction (with anterior-posterior tooth movement), significant changes to dental protrusion, and cosmetic and plastic surgery. Thus, embodiments of the present invention can provide patients with visual feedback regarding the expected appearance resulting from treatment, either during or after treatment, or both during and after treatment.
[0031] Capturing multiple 3D data sources of a patient's craniofacial anatomy may enable a physician to carefully devise one or more treatment plans based on the reality of the structural anatomy. These plans may be devised on the physician's own schedule, without the patient's presence, perhaps in collaboration with other physicians, technicians, or service providers. Prosthetics or prosthetics may be created and applied to the patient's natural anatomy, which may be augmented or modified as part of the treatment plan. 3D virtual models of such prosthetics and prosthetics may be created, for example, by a third-party laboratory or service provider and combined with datasets provided by the physician. The resulting modified anatomy may be generated by the physician, by a third party, or as a collaborative effort using remote software collaboration tools. The treatment plan may be presented to the patient for review, possible modifications, and approval. In some scenarios, the physician may be the reviewer, and a third-party laboratory or manufacturer may be the source of the proposed treatment plan.
[0032] Enough frames are captured using the video camera to generate a 3D model of the patient's face sufficient to align their own face with the 3D model containing the plan when presenting the treatment plan to the patient, which may be based on other data sources that share at least some anatomical features with the newly captured data. A complete rescan of the patient's face would not be required, since a full scan has already been performed previously to facilitate the creation of a detailed treatment plan. Thus, with only limited movement relative to the video camera, a partial 3D model of the patient's current anatomy can be created, optionally cropping out and removing portions of the anatomy altered by the plan to best fit the planning model. Because the view frustum of each video frame is known via 3D photogrammetry, the same view frustum can be used to represent a virtual camera, and a 2D rendering of the 3D model can be created in the camera's view plane. This 2D rendering can then be presented to the patient (or other observer) in real time as the patient moves relative to the camera. In effect, the physical and virtual worlds are kept synchronized with each other.
[0033] There are options when determining what to render and present to the patient. For example, if the plan is presented only in real time with the patient present, there may be no need to store a texture map in the planning model. In this scenario, alignment between individual video frames and the planning model can be performed without the use of a texture map, and the color values of each pixel in the video image presented to the patient can simply be derived from the values captured by the video camera. If appliances are attached to the patient's teeth (or other parts of the face), the appliances would be rendered in the virtual world, and their 2D representations would overlay the corresponding areas of the video frame, effectively hiding the underlying anatomical structures. The same method can be used if the anatomical structures are altered, except that the pixels of these anatomical structures may need to be colored using previously captured values UV-mapped onto the mesh of the planning model. Alternatively, color values can be obtained from the current video frame but then transformed to a different position as determined by a morph to the 3D anatomical structures according to the treatment plan. For example, a mandibular surgery may prescribe advancing the mandible by several millimeters. The color values of pixels used to render the patient's soft tissues (skin and lips) affected by the advancement will likely not be different as a result; only their position in the virtual world changes, thus altering the rendering. Therefore, the affected pixels in the 2D video image may simply be translated in the image plane according to a 3D affine transformation projected onto the view plane. In yet another scenario, once alignment of the physical video camera with the virtual camera is established, the entire 3D virtual scene may be rendered in real time and presented to the patient in synchronization with head movements.
[0034] In some of the above scenarios, techniques can be employed to continuously expand or update the planning model while capturing video for the purpose of presenting the treatment plan to the patient. In some scenarios, a 3D model of the patient may be generated using only minimal video coverage, making it likely that holes, islands, or noisy areas exist in the mesh that makes up the model. Even if the primary purpose of capture is to align the physical and virtual worlds and render an existing planning model for the patient, as new video images are received, the new images can be used to expand or update the existing model by filling holes, removing or joining islands, and reducing noise. In this way, the quality of the planning model can be improved in real time (or after a short processing delay). In some instances, the dimensional accuracy of the model can be improved by adjusting the position and orientation of smaller patches grouped together with accumulated error by a large, monolithic patch of scan data. This technique can also be used to improve the realism of avatar-style renderings by updating texture maps for current skin and lighting conditions. The texture maps used for the avatar would otherwise come from a previous scan captured earlier, possibly with different settings.
[0035] When the image or model is added, the image or 3D representation of the user's face can optionally be aligned with the image or model of the user's facial anatomy, e.g., using PCA or other techniques. Once aligned, the images or models are synchronized so that manipulation of one image or representation causes a corresponding manipulation of the other image or representation. For example, if the user's facial image is rotated, the aligned image or model of the user's facial anatomy is correspondingly rotated.
[0036] Visualization(36) This step includes overlaying the region of interest onto the original frame (FIG. 3) and displaying the final 2D image to the user (38). FIG. 14A shows the results of the treatment expansion in a single image display, for example, displayed to the user in a user interface by display device 14. Optionally, the single image in FIG. 14A can display the results of virtually removing a detected appliance worn by the user.
[0037] If optional step 32 is used to generate a 3D representation of the user's face, such 3D representation may be augmented and displayed by virtual therapy (40) for display on display device 14, as represented by the user interface of Figure 14B. If optional step 35 is used to generate additional 2D images or 3D representations, such additional 2D images or 3D representations may be augmented and displayed by virtual therapy (42) for display on display device 14, as represented by the user interfaces of Figures 14C and 14D.
[0038] Method 22 can be performed by processor 16 in real time as the image of the user is detected by camera 12 or shortly thereafter, such that the augmentation is shown to the user on display device 14. The term "real time" means a rate of at least 1 frame per second, and more preferably a rate of 15 to 60 frames per second. In addition to the above-described embodiments, the following aspects will be noted. (Appendix 1) 1. A method for simulating orthodontic treatment, comprising: receiving an electronic image of a user's face; identifying an area of interest in the image that includes the user's teeth; placing virtual orthodontic appliances on the user's teeth in the image; displaying an image of the user with the virtual orthodontic appliance on an electronic display device; The method, wherein the receiving, identifying, arranging, and displaying steps occur in real time. (Appendix 2) 2. The method of claim 1, further comprising detecting the electronic image of the user's face within a video frame. (Appendix 3) 2. The method of claim 1, further comprising applying landmarks to the electronic image of the user's face. (Appendix 4) 4. The method of claim 3, further comprising using the landmarks to estimate a facial pose of the user in the electronic image. (Appendix 5) 4. The method of claim 3, further comprising cropping the identified region of interest using the landmarks. (Appendix 6) 2. The method of claim 1, wherein the placing step includes placing virtual brackets and archwires on the user's teeth in the image. (Appendix 7) 10. The method of claim 1, further comprising generating a three-dimensional electronic representation of the user's face and displaying the three-dimensional representation of the user's face with the virtual orthodontic appliances on an electronic display device. (Appendix 8) 1. A system for simulating orthodontic treatment, comprising: a camera for providing electronic digital images or video; an electronic display device; a processor electronically connected to the camera and the electronic display device, receiving an electronic image of a user's face from the camera; identifying a region of interest in the image that includes the user's teeth; placing virtual orthodontic appliances on the user's teeth in the image; and displaying an image of the user with the virtual orthodontic appliance on an electronic display device; and a processor, wherein the receiving, identifying, locating, and displaying occur in real time. (Appendix 9) 9. The system of claim 8, wherein the processor is further configured to detect the electronic image of the user's face in video frames from the camera. (Appendix 10) 9. The system of claim 8, wherein the processor is further configured to apply landmarks to the electronic image of the user's face. (Appendix 11) 11. The system of claim 10, wherein the processor is further configured to use the landmarks to estimate a facial pose of the user in the electronic image. (Appendix 12) 11. The system of claim 10, wherein the processor is further configured to crop the identified region of interest using the landmarks. (Appendix 13) 9. The system of claim 8, wherein the processor is further configured to place virtual brackets and archwires on the user's teeth in the image. (Appendix 14) 9. The system of claim 8, wherein the processor is further configured to generate a three-dimensional electronic representation of the user's face and display the three-dimensional representation of the user's face with the virtual orthodontic appliances on an electronic display device. (Appendix 15) 1. A method for simulating orthodontic treatment, comprising: receiving an electronic image of a user's face; identifying an area of interest in the image that includes the user's teeth; removing the detected appliance; and displaying an image of the user without the detected appliance on an electronic display device; A method wherein the receiving, identifying, placing or removing, and displaying steps occur in real time. (Appendix 16) 1. A system for simulating orthodontic treatment, comprising: a camera for providing electronic digital images or video; an electronic display device; a processor electronically connected to the camera and the electronic display device, receiving an electronic image of a user's face from the camera; identifying a region of interest in the image that includes the user's teeth; removing the detected appliance; and and displaying the detected image of the user without the appliance on an electronic display device; and a processor, wherein the receiving, identifying, locating or removing, and displaying occur in real time. (Appendix 17) 1. A method for simulating orthodontic treatment, comprising: receiving an electronic image of a user's face; acquiring an electronic image or model of a user's facial anatomy; identifying an area of interest in the image that includes the user's teeth; placing virtual orthodontic appliances on the user's teeth in the image; displaying an image of the user along with the virtual orthodontic appliance and the image or model of the user's facial anatomy on an electronic display device; The method, wherein the receiving, identifying, arranging, and displaying steps occur in real time. (Appendix 18) 18. The method of claim 17, wherein the acquiring step includes acquiring an X-ray image. (Appendix 19) 18. The method of claim 17, wherein the acquiring step includes acquiring a cone beam computed tomography (CBCT) image. (Appendix 20) 18. The method of claim 17, further comprising aligning the image of the user with an image or model of the user's facial anatomy. (Appendix 21) 18. The method of claim 17, further comprising generating a three-dimensional electronic representation of the user's face; and displaying the three-dimensional representation of the user's face along with the virtual orthodontic appliances and the image or model of the user's facial anatomy on an electronic display device. (Appendix 22) 22. The method of claim 21, further comprising aligning the three-dimensional representation with an image or model of the user's facial anatomy. (Appendix 23) 1. A system for simulating orthodontic treatment, comprising: a camera for providing electronic digital images or video; an electronic display device; a processor electronically connected to the camera and the electronic display device, receiving an electronic image of a user's face; obtaining an electronic image or model of a user's facial anatomy; identifying a region of interest in the image that includes the user's teeth; placing virtual orthodontic appliances on the user's teeth in the image; displaying an image of the user along with the virtual orthodontic appliance and the image or model of the user's facial anatomy on an electronic display device; and a processor, wherein the receiving, identifying, locating, and displaying occur in real time. (Appendix 24) 24. The system of claim 23, wherein the processor is further configured to acquire an X-ray image of the user's facial anatomy. (Appendix 25) 24. The system of claim 23, wherein the processor is further configured to acquire a cone beam computed tomography (CBCT) image of the user's facial anatomy. (Appendix 26) 24. The system of claim 23, wherein the processor is further configured to align the image of the user with an image or model of the user's facial anatomy. (Appendix 27) 24. The system of claim 23, wherein the processor is further configured to generate a three-dimensional electronic representation of the user's face and display the three-dimensional representation of the user's face along with the virtual orthodontic appliances and the image or model of the user's facial anatomy on an electronic display device. (Appendix 28) 28. The system of claim 27, wherein the processor is further configured to align the three-dimensional representation with an image or model of the user's facial anatomy.
Claims
1. 1. A method for simulating orthodontic treatment in real time, comprising: receiving an electronic image; detecting facial images in the electronic image using a trained deep learning model; landmarking the facial image using a trained machine learning model; estimating a facial pose associated with the facial image based on the landmarking; identifying a region of interest in the facial image based on the landmarking and the estimated facial pose, such that the region of interest includes a representation of one or more teeth; virtually overlaying an image representing the orthodontic appliance onto the region of interest to form an orthodontic treatment simulated image; and outputting the orthodontic treatment simulated image for display via an electronic display device.
2. 1. A method for simulating an orthodontic treatment outcome in real time, comprising: receiving an electronic image; detecting facial images in the electronic image using a trained deep learning model; landmarking the facial image using a trained machine learning model; estimating a facial pose associated with the facial image based on the landmarking; identifying a region of interest in the facial image based on the landmarking and the estimated facial pose, such that the region of interest includes a representation of one or more teeth and orthodontic appliances on the teeth; removing the representation of the orthodontic appliance from the facial image to form a simulated post-orthodontic treatment image; and outputting the orthodontic treatment simulated image for display via an electronic display device.
3. 1. A method for simulating orthodontic treatment in real time, comprising: receiving an electronic image; receiving an electronic image or model of an anatomical structure corresponding to said electronic image; detecting facial images in the electronic image using a trained deep learning model; landmarking the facial image using a trained machine learning model; estimating a facial pose associated with the facial image based on the landmarking; identifying a region of interest in the facial image based on the landmarking and the estimated facial pose, such that the region of interest includes a representation of one or more teeth; virtually overlaying an image representing an orthodontic appliance onto the region of interest to form an orthodontic treatment simulated image; and outputting the orthodontic treatment simulated image, including an electronic image or model of the anatomical structure, for display via an electronic display device.
Citation Information
Patent Citations
Displaying method for moving-body
JP1992336048A
An interactive orthodontic care system based on intraoral scans of teeth
JP2004504077A
Method, system, and computer program product to execute digital orthodontics at one or more sites
JP2014091047A