3D facial reconstruction and visualization in dental treatment planning

Capturing the three-dimensional representation of the patient's face through a mobile device and integrating it with the three-dimensional intraoral mesh of the teeth solves the problem of insufficient precision in dental treatment planning in the prior art, and achieving a more accurate and authentic treatment plan.

CN120153401APending Publication Date: 2025-06-13ALIGN TECHNOLOGY INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380076945.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-08-29
Filing Date
2023-08-30
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The failure of existing dental treatment planning tools to effectively consider the relationship between tooth position and orientation and patient facial features, resulting in inaccurate and authentic enough in treatment planning.

Method used

By capturing a three-dimensional representation of the patient's face from multiple angles using a mobile device and integrating it with a three-dimensional intraoral mesh of the patient's teeth, a high-resolution combined representation is generated for visualizing dental treatment plans.

Benefits of technology

The accurate and true results of the dental treatment plan are achieved, allowing better consideration of facial relationships, and improving treatment effects and patient acceptance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120153401A_ABST
    Figure CN120153401A_ABST
Patent Text Reader

Abstract

The present disclosure is directed to capturing a three-dimensional representation of a patient's face from multiple angles to integrate (fuse) it with at least one three-dimensional intraoral mesh of the patient's teeth to visualize a dental treatment plan. In one aspect, a method includes capturing media of a patient's face from a plurality of angles using at least one device; converting the media into a three-dimensional representation of the patient's face; and transmitting the three-dimensional representation of the patient's face to one or more processing components to integrate it with the three-dimensional representation of the patient's intraoral scan to visualize the patient's dental treatment plan.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The subject matter of the present disclosure generally relates to the field of dental treatment planning, and more particularly to capturing three-dimensional representations of a patient's face from multiple angles using a mobile device to provide a complete visual display of a dental treatment plan to be implemented. Background Art

[0002] Dental treatment procedures typically involve repositioning a patient's teeth into a desired alignment in order to design and implement a treatment plan. To achieve these goals, various office imaging tools and devices are utilized. Existing tools and systems can provide three-dimensional representations of tooth alignment that a doctor can utilize to determine how treatment will affect tooth positioning and shape. Existing tools may not take into account the facial relationships between the position and orientation of the teeth and the shape and position of the patient's facial features.

[0003] The current treatment planning process may consider the facial relationships between the two-dimensional position and orientation of the teeth and the shape and position of the patient's facial features. Summary of the Invention

[0004] This Summary of the Invention is provided to introduce some concepts in a simplified form that will be further described in the Detailed Description section below. This Summary of the Invention is not intended to identify key features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.

[0005] The present disclosure provides tools for capturing three-dimensional representations of a patient's face from multiple angles (e.g., using a mobile device). The three-dimensional representation can then be integrated (fused) with at least one three-dimensional intraoral mesh of the patient's teeth. The resulting high-resolution combined representation can then be integrated with one or more three-dimensional treatment planning products, which can provide accurate and realistic results of the patient's dental treatment plan.

[0006] In one aspect, a method includes: capturing media of a patient's face from multiple angles using at least one device; converting the media into a three-dimensional representation of the patient's face; and transmitting the three-dimensional representation of the patient's face to one or more cloud-based processing components for integration with a three-dimensional representation of an intraoral scan of the patient to visualize the patient's dental treatment plan.

[0007] In another aspect, the at least one device is a mobile device having a built-in camera.

[0008] In another aspect, the media includes a plurality of two-dimensional images of the patient's face captured by moving the mobile device around the patient's face.

[0009] In another aspect, the media includes a video of the patient's face.

[0010] In another aspect, a multi-camera system is used to capture video, the multi-camera system being configured to capture images of a patient's face from multiple angles simultaneously, and the at least one device is the multi-camera system.

[0011] In another aspect, a three-dimensional intraoral scan is used as a reference for scaling a three-dimensional representation of the patient's face.

[0012] In another aspect, media is captured via an application installed on the at least one device, the application having access to the built-in camera of the at least one device.

[0013] In another aspect, the application provides real-time guidance for moving the at least one device around the patient's face to optimize the captured media.

[0014] In another aspect, the media includes a plurality of two-dimensional images of the patient's face, and converting the media into a three-dimensional representation of the patient's face includes: determining a plurality of camera poses, where each camera pose in the plurality of camera poses is determined for one of the plurality of two-dimensional images; generating a plurality of depth maps, where each depth map in the plurality of depth maps is generated for the two-dimensional image at least partially based on the camera pose for the two-dimensional image; and generating a three-dimensional representation of the patient's face based on combined information from the plurality of depth maps.

[0015] In one aspect, a method for generating a three-dimensional representation of a patient's face for dental treatment planning includes: receiving, at a processing component communicatively coupled to the at least one device, a three-dimensional representation of the patient's face based on media of the patient's face captured by the at least one device from multiple angles; and integrating, at the processing component, the three-dimensional representation of the patient's face with a three-dimensional representation of a patient intraoral scan to visualize the patient's dental treatment plan.

[0016] In another aspect, integrating the three-dimensional representation of the patient's face with the three-dimensional representation of the patient intraoral scan includes: identifying the inner mouth region of the patient in the three-dimensional representation of the patient's face; removing the inner mouth region from the three-dimensional representation of the patient's face; determining a scaling rigid relative transformation from the three-dimensional representation of the patient's face to the three-dimensional representation of the patient intraoral scan; replacing the removed inner mouth region from the three-dimensional representation with the three-dimensional representation of the intraoral scan; and aligning the three-dimensional representation of the patient's face with the three-dimensional representation of the patient intraoral scan based on the scaling rigid relative transformation.

[0017] In another aspect, a machine learning-based model is used to identify the inner mouth region in two-dimensional space.

[0018] In another aspect, the three-dimensional intraoral scan provides visualization of the patient's teeth after completion of the dental treatment plan.

[0019] In one aspect, a system includes at least one device; and a processing component communicatively coupled to the at least one device. The at least one device is configured to: capture media of a patient's face from multiple angles, convert the media into a three-dimensional representation of the patient's face, and transmit the three-dimensional representation of the patient's face to the processing component. The processing component is configured to integrate the three-dimensional representation of the patient's face with a three-dimensional representation of an intraoral scan of the patient to visualize a dental treatment plan for the patient.

[0020] In another aspect, the processing component is further configured to receive the three-dimensional representation of the intraoral scan of the patient from a dental scanner.

[0021] In another aspect, the processing component is further configured to integrate with one or more dental treatment planning applications to utilize the visualization of the dental treatment plan.

[0022] In another aspect, the at least one device is a mobile device having a built-in camera.

[0023] In another aspect, the media includes multiple two-dimensional images of the patient's face captured by moving the mobile device around the patient's face.

[0024] In another aspect, the media includes a video of the patient's face.

[0025] In another aspect, a multi-camera system is used to capture the video, the multi-camera system being configured to simultaneously capture images of the patient's face from multiple angles, wherein the at least one device is the multi-camera system.

[0026] In another aspect, the result of integrating the three-dimensional image of the patient's face with the three-dimensional image of the intraoral scan of the patient may allow an operator of the at least one device to visualize the movement of facial tissue, lips, and facial expressions after the dental treatment plan is applied to the patient's teeth.

[0027] In another aspect, a mobile application on the at least one device is configured to provide real-time guidance for moving the mobile device around the patient's face to optimize the captured media.

[0028] In another aspect, the processing component is configured to integrate the three-dimensional representation of the patient's face with the three-dimensional representation of the intraoral scan of the patient by: identifying the intraoral region of the patient in the three-dimensional representation of the patient's face; removing the intraoral region from the three-dimensional representation of the patient's face; determining a scaled rigid relative transformation of the three-dimensional representation of the patient's face to the three-dimensional representation of the intraoral scan; replacing the removed intraoral region from the three-dimensional representation with the three-dimensional representation of the intraoral scan; and aligning the three-dimensional representation of the patient's face with the three-dimensional representation of the intraoral scan based on the scaled rigid relative transformation.

[0029] In another aspect, a machine learning-based model is used to identify the inner region of the mouth in a two-dimensional space.

[0030] In another aspect, the processing component is configured to output the final result of integrating the three-dimensional representation of the patient's face with the three-dimensional representation of the intraoral scan of the patient to one or more dental planning tools for further analysis.

[0031] In another aspect, a system includes at least one media capture device; and a processing component (e.g., a cloud-based component) communicatively coupled to the at least one media capture device. The at least one media capture device is configured to: capture media of the patient's face from at least one angle, convert the media into a three-dimensional representation of the patient's face, and transmit the three-dimensional representation of the patient's face to the processing component. The processing component is configured to integrate the three-dimensional representation of the patient's face with the three-dimensional representation of the intraoral scan of the patient to visualize the patient's dental treatment plan.

[0032] In another aspect, the media includes a plurality of two-dimensional images of the patient's face, and in order to convert the media into a three-dimensional representation of the patient's face, the device is configured to: determine a plurality of camera poses, where each camera pose in the plurality of camera poses is determined for a two-dimensional image in the plurality of two-dimensional images; generate a plurality of depth maps, where each depth map in the plurality of depth maps is generated for the two-dimensional image at least partially based on the camera pose for the two-dimensional image in the plurality of two-dimensional images; and generate a three-dimensional representation of the patient's face based on combining information from the plurality of depth maps.

[0033] In one aspect, one or more non-transitory computer-readable media include computer-readable instructions that, when executed by one or more processors of a computing device, cause the computing device to receive a three-dimensional representation of a patient's face from at least one device based on media of the patient's face captured from multiple angles by the at least one device; and integrate the three-dimensional representation of the patient's face with the three-dimensional representation of the intraoral scan of the patient at the computing device to visualize the patient's dental treatment plan. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings, in which:

[0035] Figure 1A An example setup for obtaining a three-dimensional representation of a patient's face in accordance with some aspects of the present disclosure is shown;

[0036] Figure 1BShows another example setup for obtaining a three-dimensional representation of a patient's face in accordance with some aspects of the present disclosure;

[0037] Figure 2 Shows, in accordance with some aspects of the present disclosure, stored in Figure 1A a mobile device or Figure 1B an example application on a processing component for capturing media representing a patient's face;

[0038] Figure 3 Shows another example screenshot of an application for capturing media representing a patient's face as described in accordance with some aspects of the present disclosure; Figure 2 ;

[0039] Figure 4 Shows a three-dimensional representation of a patient's face based on media captured by an application using Figure 2 and Figure 3 in accordance with some aspects of the present disclosure;

[0040] Figure 5 Shows an example of a system for integrating a locally generated three-dimensional representation of a patient's face with one or more dental treatment plans in accordance with some aspects of the present disclosure;

[0041] Figure 6 Is an example method for capturing a three-dimensional representation of a patient's face in accordance with some aspects of the present disclosure;

[0042] Figure 7 Is an example method for post-processing a three-dimensional representation of a patient's face and fusing it with a three-dimensional dental treatment plan in accordance with some aspects of the present disclosure;

[0043] Figure 8A Shows an example output of generating a two-dimensional contour, projecting it onto a surface mesh, and generating a three-dimensional watertight mesh in accordance with some aspects of the present disclosure;

[0044] Figure 8B Shows an example output of removing the intraoral region from a three-dimensional representation using a three-dimensional watertight mesh in accordance with some aspects of the present disclosure;

[0045] Figure 9A Shows an example output of a scaled rigid alignment of a three-dimensional representation of a patient's face with a three-dimensional dental treatment plan after removing the intraoral region in accordance with some aspects of the present disclosure;

[0046] Figure 9B Shows an image processing pipeline for generating a 3D image of a face and merging it with one or more dental arch 3D models in accordance with an embodiment of the present disclosure;

[0047] Figure 10 Shows an example neural network that can be used to segment a three - dimensional representation of a patient's face to remove the intra - oral region according to some aspects of the present disclosure;

[0048] Figure 11 Shows an example computing system according to some aspects of the present disclosure.

[0049] Figure 12A Shows an orthodontic repositioning system including a plurality of appliances according to an embodiment of the present disclosure;

[0050] Figure 12B Shows a method of orthodontic treatment using a plurality of appliances according to an embodiment;

[0051] Figure 13 Shows a method for designing an orthodontic appliance to be produced by direct manufacturing or indirect manufacturing according to an embodiment; and

[0052] Figure 14 Shows a method for digitally planning orthodontic treatment and / or the design or manufacture of an appliance according to an embodiment. Detailed Description

[0053] A better understanding of the features and advantages of the present disclosure will be obtained by referring to the following detailed description and the accompanying drawings that set forth illustrative embodiments in which the principles of the embodiments of the present disclosure are utilized.

[0054] Although the detailed description contains many details, these details should not be construed as limiting the scope of the present disclosure, but should only be construed as illustrating different examples and aspects of the present disclosure. It should be understood that the scope of the present disclosure includes other embodiments not discussed in detail above. Various other modifications, changes, and variations that are obvious to those skilled in the art can be made to the arrangements, operations, and details of the methods, systems, and devices of the present disclosure provided herein without departing from the spirit and scope of the invention described herein.

[0055] As used herein, the terms “dental appliance” and “tooth - receiving appliance” are considered synonymous. As used herein, “dental positioning appliance” or “orthodontic appliance” can be considered synonymous and can include any dental appliance configured to change the position of a patient's teeth according to a plan (e.g., an orthodontic treatment plan). As used herein, “dental positioning appliance” or “orthodontic appliance” can include a set of dental appliances configured to progressively change the position of a patient's teeth over time. As described herein, dental positioning appliances and / or orthodontic appliances can include polymeric appliances configured to move a patient's teeth according to an orthodontic treatment plan.

[0056] As used herein, the term "and / or" may be used as a functional word to indicate that two words or expressions should be used together or separately. For example, the phrase "A and / or B" includes A alone, B alone, and A and B together. Depending on the context, the term "or" need not exclude one of multiple words / expressions. For example, the phrase "A or B" need not exclude A and B together.

[0057] As used herein, the terms "torque" and "moment" are considered synonymous.

[0058] As used herein, "moment" may encompass a force acting on an object (e.g., a tooth) at a distance from the center of resistance. For example, the moment can be calculated using the vector cross product of a vector force applied at a position corresponding to the displacement vector from the center of resistance. The moment can include a vector pointing in a certain direction. For example, a moment opposite to another moment can encompass one moment vector oriented towards one side of the object (e.g., a tooth) and another moment vector oriented towards the opposite side of the object (e.g., a tooth). Any discussion herein regarding applying a force to a patient's tooth equally applies to applying a moment to the tooth, and vice versa.

[0059] As used herein, "a plurality of teeth" may encompass two or more teeth. The plurality of teeth may, but need not, include adjacent teeth. In some embodiments, one or more posterior teeth include one or more of molars, premolars, or canines, while one or more anterior teeth include one or more of central incisors, lateral incisors, canines, first bicuspids, or second bicuspids.

[0060] The embodiments disclosed herein may be well suited for moving one or more teeth in a first set of one or more teeth or moving one or more teeth in a second set of one or more teeth, and combinations thereof.

[0061] The exemplary embodiments disclosed herein may be well suited for combination with one or more commercially available tooth movement assemblies (e.g., attachments and polymeric housing appliances (e.g., orthodontic appliances)). In some embodiments, the appliance and one or more attachments are configured to move one or more teeth along a tooth movement vector that includes six degrees of freedom, three of which are rotational and three of which are translational.

[0062] The present disclosure provides orthodontic appliances and related systems, methods, and devices. Tooth repositioning can be accomplished by using a series of removable elastomeric positioning appliances (e.g., available from the assignee of the present disclosure, Align Technology, Inc.) implemented by a system). Such an appliance can have a thin elastic material housing that generally conforms to the patient's teeth but is slightly misaligned with the initial or immediately preceding tooth configuration. Placing the appliance on the teeth applies a controlled force at specific locations to gradually move the teeth into a new configuration. This process is repeated with successive appliances that include the new configuration to ultimately move the teeth into the final desired configuration through a series of intermediate configurations or alignment patterns. The repositioning of the teeth can be achieved by other series of removable orthodontic and / or dental appliances, including polymer housing appliances.

[0063] Although appliances including polymer housing appliances are referenced herein, the embodiments disclosed herein are well-suited for use with many appliances that accommodate teeth (e.g., appliances that do not have one or more of polymers or a housing). The appliances can be manufactured using one or more of many materials (e.g., such as metals, glass, reinforcing fibers, carbon fibers, composite materials, reinforced composite materials, aluminum, biomaterials, and combinations thereof, etc.). The appliances can be formed in many ways, e.g., such as by thermoforming or direct manufacturing as described herein. Alternatively or in combination, the appliances can be manufactured using machining, e.g., appliances manufactured from a block of material using computer numerical control machining. Additionally, although orthodontic appliances are referenced herein, at least some of the techniques described herein can be applied to restorative dental appliances and / or other dental appliances, including but not limited to dental crowns, veneers, teeth whitening appliances, teeth protective appliances, etc.

[0064] In designing a virtual representation of a living being, the "uncanny valley" can relate to the degree of correspondence between the similarity of a virtual object to a living being and the emotional response to the virtual object. The concept of the uncanny valley may mean that humanoid virtual objects (robots, 3D computer animations, lifelike dolls, etc.) that look almost but not exactly like real humans can cause observers to have feelings of eeriness, horror, and revulsion that are both familiar and strange. There is a risk that virtual objects that look "almost" human can cause observers to have cold, eerie, and / or other non-emotional feelings.

[0065] In the context of treatment planning, the uncanny valley problem can lead to people having a negative reaction to their humanoid representation. For example, after undergoing an orthodontic treatment plan, people may face a strange, robot-like or non-humanoid view of themselves when viewing their 3D virtual representation.

[0066] Systems and methods that reduce the uncanny valley reaction and consider the relationship between the position and orientation of the teeth and the shape and position of the patient's facial features can help improve the effectiveness and acceptance of orthodontic treatment, particularly the effectiveness and acceptance of orthodontic treatment involving virtual representations and / or virtual 3D models of living beings before, during, and / or after the application of an orthodontic treatment plan.

[0067] The present disclosure provides tools for capturing a three-dimensional representation of a patient's face from multiple angles (e.g., using a mobile device). The three-dimensional representation can then be integrated (fused) with at least one three-dimensional intraoral mesh of the patient's teeth. The resulting high-resolution combined representation can then be integrated with one or more three-dimensional treatment planning products, which can provide accurate and realistic results for the patient's dental treatment plan.

[0068] Figure 1A An example setup for obtaining a three-dimensional representation of a patient's face in accordance with some aspects of the present disclosure is shown. Setup 100 can be located at a physical location (e.g., a dentist's office, the patient's home, or any other location). Mobile device 102 can be used to obtain media of patient 104's face. Mobile device 102 can be any known or to-be-developed consumer electronic device equipped with a media capture component (e.g., a built-in media capture component such as a camera for capturing two-dimensional photos and / or videos). Such a mobile device 102 can also be capable of having one or more applications installed thereon, as will be described below, which can be used to convert the captured media into a three-dimensional representation of patient 104's face on mobile device 102 itself in a very short time. Mobile device 102 can also be equipped with one or more electronic components that enable mobile device 102 to establish wired and / or wireless communication with one or more remote components and / or the Internet. One or more remote components can be cloud-accessible or cloud-based processing components and / or servers, such as server 106 (the use of which in the context of the present disclosure will be further described with reference to subsequent figures below). Cloud-based processing refers to processing that relies on the processing capabilities of remote servers (typically located in a data center) rather than using a local server or a personal device's computing. Cloud-based processing generally relies on virtual machines (VMs) that simulate physical computers, enabling users to run applications on cloud servers as if the applications were running on a local machine. Other options for cloud-based processing are to use containers, which are lightweight alternatives to VMs that encapsulate an application with its dependencies, libraries, and binaries in a single package. This ensures that the application will run consistently in different environments. Cloud-based processing can also utilize serverless computing, where the cloud provider dynamically manages the allocation of machine resources.

[0069] Non-limiting examples of mobile device 102 include, but are not limited to, smartphones (e.g., Apple iPhone, Samsung Galaxy, etc.), tablets (e.g., Apple iPad, Microsoft Tablet), and / or any other handheld electronic device having the above capabilities.

[0070] As shown in FIG. 1, the mobile device 102 can move around to capture media of the face of the patient 104 from different angles. In the non-limiting example of FIG. 1, three different positions (1), (2), and (3) are shown, each corresponding to a different angle at which the mobile device 102 can capture media representative of the face of the patient 104. However, the present disclosure is not limited thereto, and more or less media of the face of the patient 104 can be captured at different angles. As described above, the captured media can include two-dimensional photos (e.g., images) and / or videos. Throughout the present disclosure, a photo can be referred to as one or more still (static) images taken using one or more cameras (e.g., the built-in camera of the mobile device 102 or a number of interconnected cameras). In such a photo, static changes in the facial expression of the patient 104 (e.g., changes in tissue movement and joints) can be captured. In the present disclosure, (one or more) videos can refer to a series of still images and / or moving images of the patient 104 taken simultaneously using one or more cameras (which can be communicatively coupled to the mobile device 102 and / or the server 106) so that dynamic changes in the facial expression of the patient 104 (e.g., changes in tissue movement and joints) can be captured.

[0071] In one example, when the mobile device 102 moves between positions (1), (2), and (3), the movement of the mobile device 102 can be continuous (without pauses), or the movement of the mobile device 102 can involve a brief pause of the mobile device 102 at each angle to capture media.

[0072] Figure 1B Another example setup for obtaining a three-dimensional representation of a patient's face in accordance with some aspects of the present disclosure is shown. Figure 1B The example setup 150 is different from Figure 1A the example setup 100 in that instead of moving a mobile device around the face of the patient 104 to capture media of the face of the patient 104 from different angles, a group of fixed cameras 152, 154, and 156 (a multi-camera system) can be used. Although Figure 1B only three cameras are shown, the present disclosure is not limited thereto, and the number of cameras can be two or more.

[0073] In one example, each of the cameras 152, 154, and 156 can be in pairs (e.g., paired cameras 152, paired cameras 154, and paired cameras 156). The paired cameras can be used to capture stereoscopic images and / or three-dimensional images, videos, and / or effects.

[0074] Cameras 152, 154, and 156 can be used to capture media of patient 104's face in video form (as described below). Cameras 152, 154, and 156 can be communicatively coupled to a processing component (e.g., desktop 158) or can be directly connected to server 106. Then, generation of a three-dimensional representation of patient 104's face can occur on desktop 158 or server 106.

[0075] Figure 2 A user interface of an example application stored on mobile device 102 of FIG. 1 for capturing media representing a patient's face is shown, in accordance with some aspects of the present disclosure. Figure 2 The split screen 200 includes a snapshot of application 202 (downloadable and executable on mobile device 102 of FIG. 1). A user of mobile device 102 can select application 202 from among a plurality of available applications. A plurality of options can be displayed on application 202 (e.g., option 204 for selecting one of the available cameras on mobile device 102 (e.g., front camera or rear camera), reset button 206 for restarting the media capture process when applicable, scan button 208 for initiating the media capture process, hamburger menu 210 for accessing various other options of application 202 (e.g., settings, privacy notifications, media capture settings, previous recordings, etc.)). In one example, the scan button 208 can be changed to a stop button such that after the scan button 208 is selected, the media capture process stops when the button is selected.

[0076] Once the scan button 208 is selected, application 202 can begin capturing media of patient 104's face with identified contours 212.

[0077] The split screen 200 also shows a snapshot 214 of media captured using mobile device 102 and application 202 representing patient 104's face, as described above with reference to FIG. 1.

[0078] Figure 3 Shows a reference, in accordance with some aspects of the present disclosure Figure 2 Another example screenshot of an application for capturing media representing a patient's face, described. The example screenshot 300 is related to Figure 2is the same as screenshot 202, except that example screenshot 300 shows additional features of application 202. This additional feature is real-time media capture guidance 302. When the user of mobile device 102 moves around the patient's face to capture media of the face, real-time media capture guidance 302 can guide the user on how best to move mobile device 102 (e.g., left, right, up, and / or down) to optimize the captured media. In one example, the direction of the arrows in real-time media capture guidance 302 can change according to the recommended direction of movement (e.g., if the user is recommended to move mobile device 102 up, the arrow can point up, if the user is recommended to move mobile device 102 left, the arrow can point left, etc.).

[0079] Once media of the patient's face has been captured using mobile device 102 and application 202 (e.g., a face reconstruction application that can be Bellus3d or ARKit), a three-dimensional representation of the patient's face can be constructed and displayed within application 202. This conversion of two-dimensional photos / images and / or videos to a three-dimensional representation of the face is performed locally on mobile device 102 without using any external resources or computing power (e.g., server 106). This conversion can produce a high-resolution three-dimensional representation in a relatively short period of time (e.g., seconds or minutes). The conversion of images and videos to a three-dimensional representation can be performed according to any known or to-be-developed signal and image processing techniques for generating a three-dimensional representation of an object.

[0080] Figure 4 shows a three-dimensional representation of a patient's face based on media captured using the following aspects of the present disclosure Figure 2 and Figure 3 of an application. Snapshot 400 shows a three-dimensional representation 402 of the face of patient 104, the media of the face of which patient 104 was obtained using application 202 installed on mobile device 102, as described above with reference to FIGS. 1 to Figure 3 as described. In one example, the three-dimensional representation 402 can be interactive such that the user of mobile device 102 can move around (e.g., turn up, down, left, and right, and / or rotate around) the three-dimensional representation 402, zoom in and out, and view the three-dimensional representation 402 from different angles (e.g., by rotating the three-dimensional representation) using an input (e.g., a touch input or any other form of input).

[0081] Thereafter, the three-dimensional representation 402 can be transmitted from mobile device 102 to server 106 for further processing and integration with one or more dental treatment plans, as will be further described below, in order to provide a complete (overall) and realistic representation to the doctor and / or patient of what the patient will look like once the dental treatment plan is completed.

[0082] In one or more aspects, when the captured media is video from different angles, the overall representation will enable doctors and patients to see the effects of various facial movements once a dental treatment plan is implemented (e.g., visualizing what the patient's smile will look like after implementing the dental treatment plan, how the movement of the lips, mouth, tongue, and other facial tissues, as well as elements such as the eyes and cheeks, will affect the patient's appearance once the dental treatment plan is applied, etc.).

[0083] Figure 5 An example of a system for integrating a locally generated three-dimensional representation of a patient's face with one or more dental treatment plans is shown in accordance with some aspects of the present disclosure. Figure 5 System 500 shows a plurality of engines 502 - 508. Each engine can correspond to a set of computer-readable instructions that are stored on one or more memories and executed by one or more processors to implement a set of functions for fusing a three-dimensional representation of a patient's face with one or more dental treatment plans (e.g., with a three-dimensional model of the patient's upper and / or lower dental arches) in order to produce a visualization of the patient's final facial appearance after implementing the dental treatment plan. Figure 5 Non-limiting examples of the agents or engines shown in can be implemented on one or more servers (e.g., server 106 of FIG. 1). Although shown as separate logical engines in Figure 5 a single processor can implement the functions of two or more engines that form system 500.

[0084] System 500 can include a three-dimensional face engine 502, a dental treatment plan engine 504, a data fusion engine 506, and an output engine 508. In one example, the three-dimensional face engine 502 can receive a three-dimensional representation 402 of the face of patient 104 generated on mobile device 102 from mobile device 102. Then, the three-dimensional face engine 502 can perform a series of image segmentation and processing steps to remove the intraoral section of patient 104 from the three-dimensional representation 402 in order to replace it with a three-dimensional representation of the dental treatment plan provided by the dental treatment plan engine 504. The three-dimensional face engine 502 can perform any other functions related to modifying and / or enhancing the three-dimensional representation 402. This process will be described further below with reference to Figure 7 and.

[0085] The dental treatment plan engine 504 can interface with one or more existing three-dimensional dental planning software and / or tools (including but not limited to software and tools developed by Align Technology, Inc. of San Jose, California). Such non-limiting tools can include dental or intraoral scanners developed and sold by Align Technology, Inc. of San Jose, California, and Applications. Although these are referred to as exemplary three-dimensional dental planning tools, the present disclosure is not limited to these, and any other known or to-be-developed three-dimensional dental planning tools can be integrated / communicatively coupled with the dental treatment planning engine 504 to receive a three-dimensional dental treatment plan to be fused with the three-dimensional representation 402.

[0086] The fusion engine 506 can perform multiple image processing techniques to fuse (combine / integrate) the three-dimensional representation 402 (after removing the intraoral region of the three-dimensional representation 402) with the three-dimensional dental treatment plan received from the dental treatment planning engine 504. This fusion process will be described below with reference to Figure 7 Further description. The three-dimensional dental treatment plan can also refer to and / or can include one or more three-dimensional intraoral scans of the patient, and the one or more three-dimensional intraoral scans can be obtained using, for example, Tools.

[0087] Once the fusion process is complete, the output engine 508 can output the final three-dimensional representation 402, which is modified to include the three-dimensional dental treatment plan (e.g., one or more three-dimensional models of the patient's dental arches in one or more stages of orthodontic treatment), which will allow the dentist, healthcare professional, and / or patient to visualize how the dental treatment plan will look and feel on their facial appearance once implemented.

[0088] Figure 6 is an example method of capturing a three-dimensional representation of a patient's face according to some aspects of the present disclosure. The process will be described from the perspective of the mobile device 102 in FIG. 1 Figure 6 However, it should be noted that the mobile device 102 can have one or more memories storing computer-readable instructions of the application 202, and when these computer-readable instructions are executed by one or more processors on the mobile device 102, the mobile device 102 is enabled to implement Figure 6 Steps. When describing Figure 6 Reference may be made to one or more of FIGS. 1 to Figure 5 One or more of them.

[0089] At step 600, the mobile device 102 can activate the application 202. This activation can be triggered by the user of the mobile device 102 opening the application 202 and / or selecting the scan button 208 as described above with reference to Figure 2 Described.

[0090] At step 602, the mobile device 102 can capture media of the patient 104 from multiple angles (e.g., the positions (1), (2), and (3) described above with reference to FIG. 1). As described, the captured media can be one or more still images / photos of the patient 104's head and face, or can be a video of the patient 104's head and face. Throughout this disclosure, any reference to the patient's face can also include the patient's head. The media can be captured by the built-in camera of the mobile device 102 and / or an external media capture device (physically or communicatively) coupled to the mobile device 102.

[0091] As described above, the application 202 can provide real-time media capture guidance 302, which can direct the user of the mobile device 102 to move the mobile device 102 in different directions as the mobile device 102 moves around the patient 104 to optimize the captured media, with the aim of converting the captured media into a three-dimensional representation 402.

[0092] At step 604, the mobile device 102 can convert the media captured at step 602 into a three-dimensional representation 402 of the patient 104's head and face, as described above with reference to Figure 3 and Figure 4 described.

[0093] At step 606, the mobile device 102 can transmit the three-dimensional representation 402 to one or more servers 106 for further processing and fusion with a dental treatment plan (one or more three-dimensional intraoral scans of the patient 104), which will be described below with reference to Figure 7 further described.

[0094] Figure 7 is an example method for post-processing a three-dimensional representation of a patient's face and fusing it with a three-dimensional dental treatment plan (e.g., one or more three-dimensional models of the patient's upper and / or lower dental arches at one or more stages of treatment). The process will be described from the perspective of one of the servers 106 in FIG. 1. However, it should be noted that when more than one server 106 is utilized, any one or more of the servers 106 can perform Figure 7 the steps. Additionally, it should be noted that such a server can have one or more memories in which computer-readable instructions are stored, corresponding to the logic associated with each of the engines 502, 504, 506, and 508 described above with reference to Figure 7 which, when executed by one or more processors, enable the server 106 to implement Figure 5 the steps. When describing Figure 7 reference can be made to one or more of FIGS. 1 to Figure 7 in. Figure 6 one or more.

[0095] At step 700, the server 106 (e.g., a cloud-based processing component) can receive a three-dimensional representation 402 of the face of patient 104 from the mobile device 102.

[0096] At step 710, the server 106 can implement logic associated with the three-dimensional face engine 502 to remove at least one element or component from the three-dimensional representation 402. In one example, such an element or component can be the intraoral region of patient 104. In a non-limiting example of removing the intraoral region of patient 104 from the three-dimensional representation 402, the following steps can be taken.

[0097] First, the two-dimensional position of the intraoral contour can be generated or determined. In one example, one or more trained neural networks (machine learning processes) can be applied to generate or determine the contour. For example, a three-dimensional representation of the face of patient 104 on a plane from a predefined camera projection matrix can be generated. Then an image with the three-dimensional representation of the projection of the patient's face can be obtained. Using this image, a two-dimensional intraoral segmentation network is executed to define the region corresponding to the intraoral part of the patient. Then, image processing techniques can be applied to generate contour points defining the intraoral region. Then, with the projection matrix known, the intraoral contour points can be projected onto the three-dimensional representation, resulting in a three-dimensional representation of the patient's face with the contour of the patient's intraoral part embedded. This will be further described below with reference to Figure 11 This is further described.

[0098] Details of training one or more neural networks to generate the intraoral contour will be described below with reference to Figure 11 This is described.

[0099] Thereafter, the generated intraoral contour is projected onto a surface mesh to generate a three-dimensional watertight mesh.

[0100] Figure 8A An example output showing the generation of a two-dimensional contour, projecting it onto a surface mesh, and generating a three-dimensional mesh (e.g., a watertight mesh) according to some aspects of the present disclosure is shown. Example 800 includes a snapshot 802 that shows the above three-dimensional representation 402 and an example intraoral region 804 to be removed.

[0101] Example output 806 shows the result of applying a trained neural network to generate a two-dimensional intraoral contour 808, which is then projected onto a surface mesh 810.

[0102] Example output 812 shows the generated three-dimensional watertight mesh described above.

[0103] Returning to the reference Figure 7, once a three-dimensional watertight mesh is generated, a three-dimensional geometry processing algorithm can be applied to the three-dimensional representation 402 cut using the three-dimensional watertight mesh. Any known or to-be-developed three-dimensional geometry processing algorithm can be applied to cut the three-dimensional representation 402 using the above three-dimensional watertight mesh.

[0104] Figure 8B An example output of removing the oral internal region from a three-dimensional representation using a three-dimensional watertight mesh is shown according to some aspects of the present disclosure. In Figure 8B Example 820 of, outputs 822, 824, and 826 show the process of cutting the three-dimensional representation 402 using the three-dimensional watertight mesh, and output 828 shows the final result of the three-dimensional representation 402 with the oral internal region removed.

[0105] Returning to reference Figure 7 , after removing the oral internal region of the three-dimensional representation 402 at step 710, at step 720, the server 106 can receive a three-dimensional dental treatment plan. As described above, the dental treatment planning engine 504 can integrate (communicatively couple to) one or more dental planning tools to receive a three-dimensional representation of the dental treatment plan for the patient 104. Such a three-dimensional representation can visualize the appearance and feel of the teeth of the patient 104 after implementing the dental treatment plan.

[0106] At step 730, the server 106 can integrate the three-dimensional representation 402 with the three-dimensional dental treatment plan received at step 720 after removing the oral internal region. The integration process can be as follows. In one example, the server 106 can perform a scaled rigid alignment of the three-dimensional representation 402 with the three-dimensional dental treatment plan received at step 730 (e.g., one or more 3D models of the dental arches of the dental treatment plan). The scaled rigid alignment can be based on several poses of the patient 104 captured by the mobile device 102, the above two-dimensional machine learning segmentation, and the segmented intraoral scan. The server 106 can access several poses of the patient 104 (e.g., the originally captured media), because the mobile device 102 can also transmit the originally captured media to the server 106 together with the three-dimensional representation 402.

[0107] In one example, such a scaled rigid alignment can include the server 106 determining (computing) a scaled rigid relative transformation between the three-dimensional representation 402 and one or more 3D models of the dental arches of the patient 104 and / or one or more three-dimensional intraoral scans of the patient 104 using any known or to-be-developed method.

[0108] Thereafter, the server 106 places the three-dimensional intraoral scan or the 3D model (or a part thereof) of the upper dental arch and / or the lower dental arch in the removed oral internal region within the three-dimensional representation 402 (i.e., in Figure 8Bwithin the output 828. Once placed therein, the server 106 can perform a scaled rigid alignment process to align the three-dimensional intraoral scan or one or more 3D models of the upper dental arch and / or the lower dental arch within the three-dimensional representation 402, and scale the three-dimensional representation 402 such that the relative sizes of the patient 104's face and the three-dimensional intraoral scan are proportional to each other and do not appear unrealistic.

[0109] Figure 9A An example output of a scaled rigid alignment of a three-dimensional representation of a patient's face after removal of the intraoral region with a three-dimensional dental treatment plan (e.g., one or more 3D models of the patient's dental arches at a treatment stage) is shown, in accordance with some aspects of the present disclosure. Example 900 of FIG. 9 includes several outputs, which include a three-dimensional representation 828 of patient 104 with the intraoral region removed, as described above with reference to Figure 8B as described. Then, based on several poses of patient 104 (captured in output 902), the two-dimensional machine learning-based segmentation of the intraoral region as described above with reference to FIG. 8A (shown as output 904 in FIG. 9), and the segmentation of the three-dimensional intraoral scan or the 3D model of dental arch 906 in FIG. 9, the three-dimensional representation 828 can be aligned with the three-dimensional dental plan (e.g., the 3D models of the upper dental arch and / or the lower dental arch at the treatment stage). The resulting integration and alignment are shown as output 910 in FIG. 9. Output 910 includes three example snapshots 910-1, 910-2, and 910-3, which show step-by-step adjustments to obtain the alignment of the three-dimensional intraoral scan or the 3D model of dental arch 906 with the three-dimensional representation 402 of patient 104's face.

[0110] Returning to Figure 7 , after the integration of the 3D model of the dental arch or the three-dimensional intraoral scan with the three-dimensional representation of the patient 104's face is completed, at step 740, the server 106 can output a final result that provides a visualization of the patient 104's face with the dental treatment plan applied.

[0111] As described above, the server 106 can be coupled to one or more of the dental planning tools described above. Thus, the final result can be output to any one of such planning tools for further utilization by dentists, dental hygienists, dental technicians, and / or patient 104 to determine the adequacy of the proposed dental treatment plan and make any modifications and / or adjustments thereto if needed.

[0112] Figure 9BAn image processing pipeline 220 for generating a 3D image (e.g., a three-dimensional representation) of a face and merging it with one or more models of a dental arch 3D (e.g., a three-dimensional representation of an intraoral scan of a patient) according to an embodiment of the present disclosure is shown. The image processing pipeline 220 can be conceptually divided into a 2D-to-3D transformation pipeline 922 and a 3D image fusion pipeline (not labeled). Figure 9B is described with reference to an image, but is equally applicable to video (e.g., frames of a video).

[0113] The 2D-3D transformation pipeline 922 can include a camera pose determiner 926, which can perform tracking of the camera pose on the input image 924. The camera pose determiner 926 can receive a plurality of input 2D images 924 generated from different views corresponding to different camera positions and / or orientations (referred to as camera poses). The camera pose determiner 926 can estimate the camera pose 928 for each input image. Each camera pose 928 can include, for example, an x, y, z position and / or an x, y, z orientation (e.g., an angle relative to the global x, y, z axes). Each camera pose 928 can represent the position and orientation in space of the camera that generated the input image 924. The camera pose 928 can be determined based on: identifying 2D features in the input image 924, matching 2D features between images, and determining key frame pairs from the 2D feature matches. The adjustment between images can be determined from the 2D feature map, and the bundle adjustment can be used to estimate the camera pose, thereby ultimately estimating the 3D position of the 3D features. The camera pose determiner 926 can generate one or more camera matrices 928, which include the camera poses 928 for the plurality of input images 924 and camera parameters (e.g., extrinsic and / or intrinsic camera parameters). Intrinsic camera parameters are internal camera parameters that do not change regardless of the scene being captured. Intrinsic camera parameters are related to the internal characteristics of the camera. Examples of intrinsic parameters include the focal length (the distance between the projection center of the camera (usually the lens center) and the image plane (sensor)), the optical center or principal point (the point on the image plane where rays parallel to the optical axis converge), the skew value (describing the angle between pixel axes), and / or lens distortion parameters. Extrinsic parameters are parameters that describe the position and orientation of the camera in the world coordinate system. Extrinsic parameters are external to the camera and define its position relative to the world frame of reference. Extrinsic parameters can include, for example, a rotation matrix (e.g., a 3x3 matrix that captures the orientation of the camera in the world and relates the coordinates of a point in the camera system to its coordinates in the world system) and / or a translation vector (e.g., a 3x1 matrix that represents the position of the camera origin in world coordinates).

[0114] The depth map generator 932 receives the camera pose 928 and / or the camera matrix 930 from the camera pose determiner 926 and uses this information to estimate a depth map 934 for each input image 924. A depth map is a 2D representation of the depth information of a scene. For each pixel in the depth map, the value (usually a grayscale value) represents the distance between the viewpoint (usually the camera or viewer) and the corresponding point in the real-world scene. The estimated depth maps 934 of two or more images can be fused to generate a fused or stereo depth map 938. When two cameras (or two images generated by the same camera from different viewpoints) capture the same scene from slightly different viewpoints, the difference in the apparent position of an object in the two images (known as parallax) can be used to calculate the depth of the object. A stereo depth map can be based on an image pair and can be more accurate than a depth map generated from a single image.

[0115] The 3D image generator 940 receives the stereo depth map 938 from the depth map generator 932 and uses this information to generate a 3D mesh 942 by combining the point clouds from the depth map 934 and / or the stereo depth map 938 and the camera pose 938 associated with these depth maps or stereo depth maps. The 3D mesh 942 lacks color or texture information but provides a 3D model or image generated from the depth information contained in the stereo depth map 938. Then, the 3D image generator 940 applies texturing based on the color information contained in the input image 924 to the generated 3D mesh 942 and generates a 3D image 946 or model of the face that combines information from multiple 2D input images 924.

[0116] In some cases, the oral region in the 3D image 946 may not have sufficient quality for clinical purposes. Teeth may be highly reflective, which may result in inaccurate geometry of the internal oral region.

[0117] Once a 3D image 946 of a patient's face has been generated, it can be beneficial to merge the 3D image of the face with a 3D model of the patient's dental arch, as described elsewhere herein. In one embodiment, an image segmenter 950 receives one or more input images 924 and segments at least the oral region of the one or more input images 924 to generate a segmented image 952 that includes segmentation information of the intraoral region (e.g., where each tooth in the image has been identified and / or labeled). Image segmentation is a computer vision task in which an image is divided into multiple segments or regions, each of which represents a different object or part of an object. For example, one or more trained machine learning models (e.g., one or more neural networks) can be used to perform the segmentation. One or more neural networks can receive a 2D image of the face as input and can output segmentation information that segments the 2D image into an oral region and a non-oral region and optionally further segments the oral region into individual teeth and / or gums. This can be performed for each input image 924 or only for a subset of the input images 924. Depending on the architecture, the output of the machine learning model may be a label for each pixel (semantic segmentation) or a label for each different object instance (instance segmentation).

[0118] The segmented image 952 can be projected into 3D using an appropriate camera pose 928, depth map 934, and / or stereo depth map 938. Based on the projection of the segmented image 952 into 3D, the segmentation information of the intraoral region can be determined in the 3D image 946. An image updater 954 can update the 3D image by removing the data of the intraoral region based on the 3D intraoral segmentation information (e.g., including 3D tooth segmentation). An updated 3D image 960 can be generated in which the data of the intraoral region has been removed or deleted.

[0119] One or more 3D models 962 of the patient's dental arch may have been generated based on an intraoral scan (e.g., a 3D model of the patient's upper dental arch and a 3D model of the patient's lower dental arch). In an embodiment, a dental arch segmenter 964 segments the 3D model of the dental arch to generate a segmented 3D model 966. For example, one or more trained machine learning models (e.g., one or more neural networks) can be used to perform the segmentation. One or more neural networks can receive the 3D model of the dental arch and / or projections of the 3D model of the dental arch on one or more 2D planes as input and can output segmentation information that segments the 3D model into individual teeth and / or gums. This can be performed for both the upper dental arch 3D model and the lower dental arch 3D model.

[0120] In an embodiment, the segmented image 952, the input image 924, the camera matrix 930, and / or the stereo depth map 938 are input into the dental arch to 3D image aligner 970. Additionally, one or more segmented 3D models 966 can also be input into the dental arch to 3D image aligner 970. The dental arch to 3D image aligner 970 can perform multi-view facial alignment, where two or more of the segmented 2D images 952 are aligned (e.g., registered) with the (one or more) segmented 3D models 966. For example, for each segmented image 952, one or more transforms can be calculated, which can adjust the size (e.g., scale), x, y, and / or z positions, and / or rotation about the x, y, and / or z axes to register the segmented image 952 (e.g., the segmented oral region of the segmented image) to the segmented 3D model. Based on the registration / alignment of the multiple segmented images 952 to the segmented 3D model, a minimization or optimization problem can be solved to determine a single set of transform parameters that can be applied to the updated 3D image 960 to register the updated 3D image 960 with the segmented 3D model 966 (or the 3D model 962 of the dental arch). The single set of transform parameters can provide a rigid transform of the teeth relative to the camera pose of each corresponding image, which can be applied to the corresponding image to align the corresponding image with the 3D model. The rigid transform can be used to match all the different images with the 3D model of the dental arch.

[0121] Once the rigid transform is determined, the dental arch to 3D image aligner 970 can apply the determined rigid transform to incorporate the 3D model 962 of the dental arch (or the segmented 3D model 966) into the updated 3D image 960. Since the oral interior region is removed in the updated 3D image 960, the oral interior region will be filled with information from the 3D model 962 of the dental arch. The dental arch to 3D image aligner 970 can output the fused 3D image and 3D model 972 of the dental arch, which can be output to a display.

[0122] Thus, from the multi-view Figure 2 D images of the patient's face, the processing logic generates a 3D representation of the patient's facial structure and texture based on the estimation of the camera pose (e.g., the camera matrix), the depth map, and / or other information. Additionally, the processing logic registers the patient's 3D intraoral scan data (e.g., the 3D models of the upper dental arch and / or the lower dental arch) to the same multi-view Figure 2 D image set. The rigid transform is estimated, which is related to the estimated camera matrix. This allows the processing logic to estimate the rigid transform for aligning the 3D face with the 3D intraoral scan data (e.g., the 3D models of the upper dental arch and / or the lower dental arch).

[0123] Figure 10 An example neural network is shown that can be used to segment a three-dimensional representation of a patient's face to remove the oral interior region according to some aspects of the present disclosure.

[0124] The architecture 1000 includes a neural network 1010 defined by an example neural network description 1001 in a rendering engine model (neural controller) 1030. The neural network 1010 may represent a neural network implementation of a rendering engine for rendering media data. The neural network description 1001 may include a complete specification of the neural network 1010. For example, the neural network description 1001 may include a description or specification of the neural network 1010 (e.g., layers, layer interconnections, number of nodes in each layer, etc.); input and output descriptions indicating how the input and output are formed or processed; indications of activation functions, operations or filters in the neural network, etc.; neural network parameters, e.g., weights, biases, etc.; and so on.

[0125] In this example, the neural network 1010 may include an input layer 1002 that includes input data, such as an image of the face of patient 104 or a two-dimensional projection of a three-dimensional representation 402 of the face of patient 104.

[0126] The neural network 1010 may include hidden layers 1004A through 1004N (collectively referred to hereinafter as "1004"). The hidden layers 1004 may include n hidden layers, where n is an integer greater than or equal to 1. The number of hidden layers may include as many layers as required for a desired processing result and / or rendering intent. The neural network 1010 also includes an output layer 1006 that provides an output (e.g., an identification of an intraoral region within an image of the face of patient 104 or a two-dimensional projection of a three-dimensional representation 402 of the face of patient 104) resulting from the processing performed by the hidden layers 1004.

[0127] The neural network 1010 in this example may be a multi-layer neural network of interconnected nodes. Each node may represent a piece of information. The information associated with a node may be shared between different layers, and each layer retains the information as it is processed. In some cases, the neural network 1010 may include a feedforward neural network, in which case there are no feedback connections where the output of the neural network is fed back into itself. In other cases, the neural network 1010 may include a recurrent neural network, which may have a loop that allows information to be carried across nodes when reading in the input.

[0128] Information can be exchanged between nodes through node-to-node interconnections between the various layers. Nodes in the input layer 1002 can activate a group of nodes in the first hidden layer 1004A. For example, as shown, each of the input nodes in the input layer 1002 is connected to each of the nodes in the first hidden layer 1004A. Nodes in the hidden layer 1004A can transform the information of each input node by applying an activation function to the information. Then, the transformed information can be passed to nodes in the next hidden layer (e.g., 1004B) and can activate the nodes in the next hidden layer, which can perform their own specified functions. Example functions include convolution functions, upsampling functions, data transformation functions, pooling functions, and / or any other suitable functions. Then, the output of the hidden layer (e.g., 1004B) can activate nodes in the next hidden layer (e.g., 1004N), and so on. The output of the last hidden layer can activate one or more nodes in the output layer 1006, where the output is provided. In some cases, although the nodes (e.g., nodes 1008A, 1008B, 1008C) in the neural network 1010 are shown as having multiple output lines, the nodes have a single output, and all the lines shown as output from the node represent the same output value.

[0129] In some cases, each node or the interconnections between nodes can have weights, which are a set of parameters obtained from training the neural network 1010. For example, the interconnections between nodes can represent a piece of information learned about the interconnected nodes. The interconnections can have numerical weights that can be adjusted (e.g., based on a training data set), allowing the neural network 1010 to adapt to the input and be able to learn as more data is processed.

[0130] The neural network 1010 can be pre-trained to process the features of data from the input layer 1002 using different hidden layers 1004 in order to provide an output through the output layer 1006. In an example where the neural network 1010 is used to identify the oral internal region in the three-dimensional representation 402, the neural network 1010 can be trained using training data including various two-dimensional images with the identified oral internal regions annotated.

[0131] In some cases, the neural network 1010 can use a training process called backpropagation to adjust the weights of the nodes. Backpropagation can include a forward pass, a loss function, a backward pass, and weight updates. The forward pass, loss function, backward pass, and parameter updates are performed for one training iteration. This process can be repeated for a certain number of iterations for each training media data set until the weights of the layers are adjusted accurately.

[0132] For the first training iteration of neural network 1010, the output can include values that do not give a preference for any particular class since the weights are randomly selected during initialization. For example, if the output is a vector with probabilities of objects including different products and / or different users, the probability values for each of the different products and / or users can be equal or at least very similar (e.g., for ten possible products or users, each class can have a probability value of 0.1). With the initial weights, neural network 1010 cannot identify low-level features and thus cannot accurately determine what the classification of the object might be. A loss function can be used to analyze the error in the output. Any suitable loss function definition can be used.

[0133] For the first training dataset (e.g., an image), the loss (or error) will be high because the actual values will be different from the predicted output. The goal of training is to minimize the amount of loss such that the predicted output matches the target or ideal output. Neural network 1010 can perform backpropagation by determining which inputs (weights) contribute the most to the loss of neural network 1010 and can adjust the weights such that the loss decreases and ultimately is minimized.

[0134] The derivative of the loss with respect to the weights can be calculated to determine the weights that contribute the most to the loss of neural network 1010. After the derivative is calculated, weight updates can be performed by updating the weights of the filters. For example, the weights can be updated such that they change in the opposite direction of the gradient. The learning rate can be set to any suitable value, where a high learning rate includes larger weight updates and a lower value indicates smaller weight updates.

[0135] Neural network 1010 can include any suitable neural network or deep learning network. One example includes a convolutional neural network (CNN) that includes an input layer and an output layer with multiple hidden layers between the input layer and the output layer. The hidden layers of the CNN include a series of convolutional layers, non-linear layers, pooling layers (for downsampling), and fully connected layers. In other examples, neural network 1010 can represent any other neural network or deep learning network, such as an autoencoder, a deep belief network (DBN), a recurrent neural network (RNN), etc.

[0136] Figure 11 An example computing system is shown in accordance with some aspects of the present disclosure. Figure 11 The example computing system 1100 can be used to implement the above with reference to FIGS. 1 to Figure 10Any of the described components, including the mobile device 102, the server 106, etc. The computing system 1100 may include components that communicate electrically with each other using the connection 1105. The connection 1105 may be a physical connection via a bus or a direct connection to the processor 1110, such as in a chipset architecture. The connection 1105 may also be a virtual connection, a network connection, or a logical connection.

[0137] In some embodiments, the computing system 1100 is a distributed system, where the functions described in this disclosure may be distributed across data centers, multiple data centers, peer-to-peer networks, etc. In some embodiments, one or more of the described system components represent many such components, each component performing some or all of the functions described for the component. In some embodiments, the components may be physical devices or virtual devices.

[0138] The example system 1100 includes at least one processing unit (CPU or processor) 1110 and the connection 1105 that couples various system components, including the system memory 1115, such as read-only memory (ROM) 1120 and random access memory (RAM) 1125, to the processor 1110. The computing system 1100 may include a cache of high-speed memory 1112 that is directly connected to, in proximity to, or integrated as part of the processor 1110.

[0139] The processor 1110 may include any general-purpose processor and dedicated processors configured to control hardware services or software services for the processor 1110 (e.g., services 1132, 1134, and 1136 stored in the storage device 1130) and where software instructions are incorporated into the actual processor design. The processor 1110 may essentially be a fully self-contained computing system, containing multiple cores or processors, buses, memory controllers, caches, etc. The multi-core processor may be symmetric or asymmetric.

[0140] To enable user interaction, computing system 1100 includes an input device 1145, which can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice, etc. Computing system 1100 can also include an output device 1135, which can be one or more of a number of output mechanisms known to those skilled in the art. In some cases, a multimodal system can enable a user to provide multiple types of input / output to communicate with computing system 1100. Computing system 1100 can include a communication interface 1140, which generally can govern and manage user input and system output. There are no limitations on operating on any particular hardware arrangement, and thus the basic features here can be easily replaced as improved hardware or firmware arrangements are developed.

[0141] The storage device 1130 can be a non-volatile memory device and can be a hard disk or other type of computer-readable medium that can store data accessible by a computer, such as a magnetic tape cartridge, a flash memory card, a solid-state memory device, a digital versatile disc, a cassette tape, random access memory (RAM), read-only memory (ROM), and / or some combination of these devices.

[0142] The storage device 1130 can include software services, servers, services, etc., which, when the code defining such software is executed by the processor 1110, cause the system to perform functions. In some embodiments, the hardware services that perform a particular function can include software components stored in a computer-readable medium and the hardware components required to perform that function (e.g., processor 1110, connection 1105, output device 1135, etc.).

[0143] Figure 12AA tooth repositioning system 1210 is shown that includes a plurality of appliances 1212, 1214, 1216. The appliances 1212, 1214, 1216 can be designed based on the generation of a series of dental arch 3D models, which can be generated by performing an intraoral scan of a patient's oral cavity and registering and stitching between a plurality of intraoral scans generated during the intraoral scan process. Any appliance described herein can be designed and / or provided as part of a plurality of appliances in a set for use in a tooth repositioning system and can be designed in accordance with an orthodontic treatment plan generated according to embodiments of the present disclosure. Each appliance can be configured such that the tooth receiving cavity has a geometry corresponding to an intermediate or final tooth arrangement intended for the appliance. By placing a series of progressive position adjustment appliances on a patient's teeth, the patient's teeth can be gradually repositioned from an initial tooth alignment to a target tooth alignment. For example, the tooth repositioning system 1210 can include a first appliance 1212 corresponding to an initial tooth alignment, one or more intermediate appliances 1214 corresponding to one or more intermediate alignments, and a final appliance 1216 corresponding to a target alignment. The target tooth alignment can be a planned final tooth alignment selected for the patient's teeth at the end of all planned orthodontic treatment, optionally using a trained machine learning model to output the target tooth alignment. Alternatively, the target alignment can be one of some intermediate alignments of the patient's teeth during orthodontic treatment, which can include various different treatment scenarios, including but not limited to cases where surgery is recommended, cases where interproximal reduction (IPR) is appropriate, cases where progress checks are scheduled, cases where anchorage placement is optimal, cases where palatal expansion is desired, cases involving restorative dentistry (e.g., inlays, onlays, crowns, bridges, implants, veneers, etc.), and the like. Thus, it can be understood that the target tooth alignment can be any planned resultant alignment for the patient's teeth following one or more progressive repositioning phases. Similarly, the initial tooth alignment can be any initial alignment of the patient's teeth followed by one or more progressive repositioning phases.

[0144] In some embodiments, the appliances 1212, 1214, 1216 (or portions thereof) can be produced using indirect manufacturing techniques, such as by thermoforming on a positive or negative mold. Indirect manufacturing of orthodontic appliances can involve the following: producing a positive or negative mold of the patient's dentition in a target alignment (e.g., by rapid prototyping manufacturing, milling, etc.), and thermoforming one or more sheets on the mold to produce the appliance shell.

[0145] In an example of indirect manufacturing, a mold of a patient's dental arch can be fabricated from the digital model of the dental arch generated by the trained machine learning model as described above, and a shell can be formed on the mold (e.g., by thermoforming a polymer sheet on the mold of the dental arch and then trimming the thermoformed polymer sheet). Fabrication of the mold can be performed by a rapid prototyping machine (e.g., a stereolithography (SLA) 3D printer). After the digital models of the appliances 1212, 1214, 1216 have been processed by the processing logic of the computing device, the rapid prototyping machine can receive the digital model of the mold of the dental arch and / or the digital models of the appliances 1212, 1214, 1216. The processing logic can include hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executed by a processing device), firmware, or a combination thereof.

[0146] To fabricate the mold, the shape of the patient's dental arch at a treatment stage is determined based on a treatment plan. In an orthodontics example, a treatment plan can be generated based on an intraoral scan of the dental arch to be modeled. An intraoral scan of the patient's dental arch can be performed to generate a three-dimensional (3D) virtual model of the patient's dental arch (the mold). For example, a full scan of the patient's mandibular and / or maxillary arch can be performed to generate its 3D virtual model. The intraoral scan can be performed by creating multiple overlapping intraoral images or scans from different scan sites and then stitching the intraoral images or scans together to provide a composite 3D virtual model. In other applications, a virtual 3D model can also be generated based on a scan of the object to be modeled or based on the use of computer-aided drafting techniques (e.g., for designing a virtual 3D mold). Alternatively, an initial negative mold (e.g., a dental impression, etc.) can be generated from the actual object to be modeled. The negative mold can then be scanned to determine the shape of the positive mold to be produced.

[0147] Once the virtual 3D model of the patient's dental arch is generated, a dental practitioner can determine the desired treatment outcome including the final position and orientation of the patient's teeth. The processing logic can then determine a plurality of treatment stages for progressing the teeth from the starting position and orientation to the target final position and orientation. By calculating the progression of tooth movement throughout the orthodontic treatment from the initial tooth placement and orientation to the final corrected tooth placement and orientation, the shapes of the final virtual 3D model and each intermediate virtual 3D model can be determined. For each treatment stage, a separate virtual 3D model of the patient's dental arch at that treatment stage can be generated. The original virtual 3D model, the final virtual 3D model, and each intermediate virtual 3D model are unique and customized for the patient.

[0148] Thus, multiple different virtual 3D models (digital designs) of an arch can be generated for an individual patient. A first virtual 3D model can be a unique model of the patient's arch and / or teeth as they currently present, and a final virtual 3D model can be a model of the patient's arch and / or teeth after the correction of one or more teeth and / or jaws. Multiple intermediate virtual 3D models can be modeled, each being progressively different from the previous virtual 3D model.

[0149] Each virtual 3D model of the patient's arch can be used to generate a uniquely customized physical mold of the arch at a particular treatment stage. The shape of the mold can be at least partially based on the shape of the virtual 3D model at that treatment stage. The virtual 3D model can be represented in a file such as a computer-aided drafting (CAD) file or a 3D printable file such as a stereolithography (STL) file. The virtual 3D model of the mold can be sent to a third party (e.g., a clinician's office, a laboratory, a manufacturing facility, or other entity). The virtual 3D model can include instructions that will control a manufacturing system or device to produce a mold with a specified geometry.

[0150] A clinician's office, a laboratory, a manufacturing facility, or other entity can receive the virtual 3D model of the mold, i.e., the digital model created as described above. The entity can input the digital model into a 3D printer. 3D printing includes any layer-based additive manufacturing process. 3D printing can be achieved using an additive process in which successive layers of material are formed in a prescribed shape. 3D printing can be performed using extrusion deposition, granular material binding, lamination, photopolymerization, continuous liquid interface production (CLIP), or other techniques. 3D printing can also be achieved using a subtractive process (e.g., milling).

[0151] In some cases, stereolithography (SLA) (also known as optical fabrication solid imaging) is used to manufacture an SLA mold. In SLA, a mold is manufactured by successively printing thin layers of a photocurable material (e.g., a polymeric resin) on top of one another. A platform sits in a bath of liquid photopolymer or resin, just below the surface of the bath. A light source (e.g., an ultraviolet laser) traces a pattern on the platform, thereby curing the photopolymer that the light source is directed at to form the first layer of the mold. The platform is progressively lowered, and the light source traces a new pattern on the platform to form another layer of the mold with each progressive lowering. The process is repeated until the mold is fully manufactured. Once all the layers of the mold are formed, the mold can be cleaned and cured.

[0152] Materials such as polyesters, copolyesters, polycarbonates, thermally polymerized polyurethanes, polypropylenes, polyethylenes, polypropylene and polyethylene copolymers, acrylics, cyclic block copolymers, polyetheretherketones, polyamides, polyethylene terephthalates, polybutylene terephthalates, polyetherimides, polyethersulfones, polytrimethylene terephthalates, styrene block copolymers (SBCs), silicone rubbers, elastomer alloys, thermally polymerized elastomers (TPEs), thermally polymerized vulcanized rubbers (TPVs) elastomers, polyurethane elastomers, block copolymer elastomers, polyolefin blend elastomers, thermally polymerized copolyester elastomers, thermally polymerized polyamide elastomers, or combinations thereof can be used to directly form a mold. The materials used to fabricate the mold can be provided in an uncured form (e.g., as a liquid, resin, powder, etc.) and can be cured (e.g., by photopolymerization, photocuring, gas curing, laser curing, crosslinking, etc.). The properties of the material before curing may be different from the properties of the material after curing.

[0153] An appliance can be formed according to each mold and, when applied to a patient's teeth, can provide a force to move the patient's teeth as prescribed by a treatment plan. The shape of each appliance is unique and customized for a specific patient and a specific treatment phase. In an example, appliances 1212, 1214, and 1216 can be pressure-formed or thermoformed on a mold. Each mold can be used to fabricate an appliance that will apply a force to the patient's teeth at a specific phase of orthodontic treatment. Each of the appliances 1212, 1214, and 1216 has a tooth-receiving cavity that receives the teeth and elastically repositions the teeth according to a specific treatment phase.

[0154] In one embodiment, a sheet of material is press-formed or thermoformed on a mold. The sheet of material can be, for example, a polymer sheet (e.g., an elastomeric thermopolymer, a polymer material sheet, etc.). To thermoform a housing on a mold, the sheet of material can be heated to a temperature at which the sheet of material becomes soft. Pressure can be applied to the sheet of material simultaneously to form the now-soft sheet of material around the mold. Once the sheet of material cools, it will have a shape that conforms to the mold. In one embodiment, a release agent (e.g., a non-stick material) is applied to the mold before forming the housing. This can facilitate subsequent removal of the mold from the housing. A force can be applied to lift the appliance from the mold. In some cases, the removal force may cause cracking, warping, or deformation. Thus, the embodiments disclosed herein can determine, prior to fabrication, where possible one or more damage points may occur in the digital design of the appliance and can perform corrective measures.

[0155] Additional information can be added to the appliance. The additional information can be any information related to the appliance. Examples of such additional information include part number identifiers, patient names, patient identifiers, case numbers, sequence identifiers (e.g., indicating which appliance in a treatment sequence a particular liner is), manufacturing dates, clinician names, logos, etc. For example, after identifying possible damage points in the digital design of the appliance, indicators can be inserted into the digital design of the appliance. In some embodiments, the indicator can indicate a recommended location for gripping the polymeric appliance to prevent the damage points from manifesting during gripping.

[0156] After forming the appliance on the mold for a treatment phase, the appliance is removed from the mold (e.g., automatically removed from the mold), and then the appliance is trimmed along a cutting line (also referred to as a trimming line). The determination of the (one or more) cutting lines can be based on a virtual 3D model of the dental arch for a particular treatment phase, on a virtual 3D model of the appliance to be formed on the dental arch, or on a combination of the virtual 3D model of the dental arch and the virtual 3D model of the appliance. The location and shape of the cutting line are important for the functionality of the appliance (e.g., the ability of the appliance to apply desired forces to the patient's teeth) as well as the fit and comfort of the appliance. For shells such as orthodontic appliances, orthodontic retainers, and orthodontic splints, the trimming of the shell plays an important role in the efficacy of the shell for its intended purpose (e.g., aligning, retaining, or positioning one or more teeth of the patient) and the fit of the shell on the patient's dental arch. For example, if too much trimming is done on the shell, the shell may lose rigidity and the ability of the shell to apply forces to the patient's teeth may be impaired. If too much trimming is done on the shell, the shell may become weak at that location and may become a damage point when the patient removes the shell from their teeth or when the shell is removed from the mold. In some embodiments, as one of the corrective measures taken when identifying possible damage points in the digital design of the appliance, the cutting line can be modified in the digital design of the appliance.

[0157] On the other hand, if too little trimming is done on the shell, some parts of the shell may impinge on the patient's gums and cause discomfort, swelling, and / or other dental problems. Additionally, if too little trimming is done on the shell at a location, the shell may be too rigid at that location. In some embodiments, the cutting line can be a straight line passing through the appliance at, below, or above the gingival line. In some embodiments, the cutting line can be a gingival cutting line, which represents the interface between the appliance and the patient's gums. In such embodiments, the cutting line controls the distance between the edge of the appliance and the patient's gingival line or gum surface.

[0158] Each patient has a unique dental arch with unique gingiva. Thus, the shape and location of the cutting line can be unique and customized for each patient and each treatment stage. For example, the cutting line is customized to follow the gingival line (also known as the gum line). In some embodiments, the cutting line can deviate from the gingival line in some areas and be on the gingival line in other areas. For example, in some cases, it may be desirable for the cutting line to deviate from the gingival line (e.g., not touch the gingiva), where in the interdental areas between teeth, the housing will contact the teeth and be on the gingival line (e.g., touch the gingiva). Therefore, it is important to trim the housing along the predetermined cutting line.

[0159] Figure 12BA method 1250 of orthodontic treatment using multiple appliances according to an embodiment is shown. Method 1250 can be practiced using any appliance or group of appliances described herein. In block 1260, a first orthodontic appliance is applied to a patient's teeth to reposition the teeth from a first tooth alignment to a second tooth alignment. In block 1270, a second orthodontic appliance is applied to the patient's teeth to reposition the teeth from the second tooth alignment to a third tooth alignment. Method 1250 can be repeated using any suitable number and combination of sequential appliances as needed to progressively reposition the patient's teeth from an initial alignment to a target alignment. These appliances can be generated all at once in the same phase or in groups or batches (e.g., at the start of a treatment phase), or the appliances can be manufactured one at a time, and the patient can wear each appliance until the pressure of each appliance on the teeth is no longer felt, or until the maximum amount of tooth movement presented at that given phase has been achieved. Multiple different appliances (e.g., a set of appliances) can be designed and even manufactured before the patient wears any one of the multiple appliances. After the appliance has been worn for an appropriate period of time, the patient can replace the current appliance with the next appliance in the series until no appliances remain. The appliances are generally not fixed to the teeth, and the patient can place and replace the appliances at any time during the treatment (e.g., patient-removable appliances). The final appliance or appliances in the series can have one or more geometries selected for overcorrecting the tooth alignment. For example, one or more appliances can have a geometry that (if fully realized) would move individual teeth beyond what has been selected as the "final" tooth alignment. This overcorrection may be desirable to counteract potential regression after the repositioning method has ended (e.g., allowing individual teeth to move back towards their pre-correction positions). Overcorrection may also be beneficial for accelerating the correction speed (e.g., an appliance with a geometry positioned beyond the desired intermediate or final position may move individual teeth towards that position at a greater rate). In this case, the use of the appliance can be terminated before the teeth reach the position defined by the appliance. Additionally, overcorrection can be deliberately applied to compensate for any inaccuracies or limitations of the appliance.

[0160] Figure 13 A method 1300 of designing an orthodontic appliance to be produced by direct or indirect manufacturing according to an embodiment is shown. Method 1300 can be applied to any embodiment of the orthodontic appliances described herein, and one or more trained machine learning models can be used to perform method 1300 in an embodiment. Some or all of the blocks of method 1300 can be performed by any suitable data processing system or device (e.g., one or more processors configured with suitable instructions).

[0161] At block 1305, a target alignment of one or more teeth of a patient can be determined. The target alignment of the teeth (e.g., the desired and expected end result of orthodontic treatment) can be received from a clinician in the form of a prescription, can be calculated based on fundamental orthodontic principles, can be inferred computationally based on a clinical prescription, and / or can be generated by a trained machine learning model (e.g., a treatment plan generator). By specifying the desired final positions of the teeth and a digital representation of the teeth themselves, the final position and surface geometry of each tooth can be specified to form a complete model of the tooth alignment at the desired end of treatment.

[0162] In block 1310, a movement path for moving one or more teeth from an initial alignment to the target alignment is determined. The initial alignment can be determined, for example, from a mold or scan of the patient's teeth or oral tissues using techniques such as wax bite registration, direct contact scanning, x-ray imaging, tomography, ultrasound imaging, and other techniques for obtaining information about the position and structure of the teeth, jaws, gums, and other orthodontically relevant tissues. A digital data set representing the initial (e.g., pre-treatment) alignment of the patient's teeth and other tissues can be obtained from the data acquired, e.g., a 3D model of one or more dental arches of the patient. Optionally, the initial digital data set is processed to segment the tissue components from each other. For example, a data structure digitally representing the individual crowns can be produced. Advantageously, a digital model of the entire tooth can be produced, optionally including measured or inferred hidden surfaces and root structures as well as surrounding bone and soft tissue.

[0163] Having both the initial and target positions of each tooth, a movement path can be defined for the movement of each tooth. Determining the movement path of one or more teeth can include identifying a plurality of progressive alignments of one or more teeth for implementing the movement path. In some embodiments, the movement path implements one or more force systems on one or more teeth (e.g., as described below). In some embodiments, the movement path is configured to move the teeth in the fastest way with the least amount of back-and-forth movement to bring the teeth from their initial positions to their desired target positions. Optionally, the tooth path can be segmented, and the segments can be calculated such that the movement of each tooth within a segment remains within threshold limits of linear and rotational translation. In this way, the endpoints of each path segment can constitute clinically viable repositionings, and the set of segment endpoints can constitute a clinically viable sequence of tooth positions such that moving from one point to the next in the sequence does not cause the teeth to collide.

[0164] In some embodiments, a force system for generating movement of one or more teeth along a movement path is determined. The force system can include one or more forces and / or one or more torques. Different force systems can result in different types of tooth movement, such as tipping, translation, rotation, extrusion, intrusion, root movement, etc. Biomechanical principles, modeling techniques, force calculation / measurement techniques, etc. (including knowledge and methods commonly used in orthodontics) can be used to determine the appropriate force system to be applied to the teeth to achieve tooth movement. When determining the force system to be applied, sources can be considered, including literature, force systems determined through experiments or virtual modeling, computer-based modeling, clinical experience, minimization of unwanted forces, etc.

[0165] Determination of the force system can include constraints on the allowable forces, such as the allowable direction and magnitude and the desired movement caused by the applied forces. For example, when fabricating a palatal expander, different patients may require different movement strategies. For example, the amount of force required to separate the palate can depend on the patient's age, as very young patients may not have fully formed sutures. Thus, in adolescent patients with an incompletely closed palatal suture and others, palatal expansion can be accomplished with a lower force magnitude. Slower palatal movement can also help growing bone to fill the expanded suture. For other patients, more rapid expansion may be required, which can be achieved by applying greater forces. The structure and materials of the appliance can be selected according to these requirements as needed; for example, by selecting a palatal expander capable of applying large forces to split the palatal suture and / or cause rapid expansion of the palate. Subsequent appliance phases can be designed to apply different amounts of force, such as first applying large forces to disrupt the suture and then applying smaller forces to maintain suture separation or gradually expand the palate and / or dental arch.

[0166] Determination of the force system can also include modeling the patient's facial structure, such as the skeletal structure of the jaws and palate. For example, scan data of the palate and dental arch (such as X-ray data or 3D optical scan data) can be used to determine parameters of the skeletal and muscular systems of the patient's mouth in order to determine the forces sufficient to provide the desired expansion of the palate and / or dental arch. In some embodiments, the thickness and / or density of the palatal mid-suture can be considered. In other embodiments, the treating professional can select an appropriate treatment based on the patient's physiological characteristics. For example, the characteristics of the palate can also be estimated based on factors such as the patient's age. For example, younger adolescent patients generally require smaller forces to expand the suture than older patients because the suture is not fully formed.

[0167] In block 1330, the design of one or more dental appliances shaped to achieve a movement path is determined. In one embodiment, one or more dental appliances are shaped to move one or more teeth toward a corresponding progressive alignment. The determination of one or more dental appliances or orthodontic appliances, the geometry, material composition, and / or properties of the appliances can be performed using a treatment or force application simulation environment. The simulation environment can include, for example, a computer modeling system, a biomechanical system, or an instrument, etc. Optionally, a digital model of the appliance and / or teeth can be generated, such as a finite element model, a 3D virtual model of the dental arch, etc. A finite element model can be created using computer program application software available from various vendors. To create a solid geometry model, a computer-aided engineering (CAE) or computer-aided design (CAD) program can be used, such as the software product available from Autodesk, Inc. of San Rafael, California. To create and analyze a finite element model, program products from multiple vendors can be used, including the finite element analysis package from ANSYS, Inc. of Canonsburg, Pennsylvania, and the SIMULIA (Abaqus) software product from Dassault Systèmes of Waltham, Massachusetts.

[0168] In block 1340, instructions for manufacturing one or more dental appliances are determined or identified. In some embodiments, the instructions identify one or more geometries of one or more dental appliances. In some embodiments, the instructions identify the slices for each layer used to fabricate one or more dental appliances using a 3D printer. In some embodiments, the instructions identify one or more geometries of a mold that can be used to indirectly fabricate one or more dental appliances (e.g., by thermoforming a plastic sheet over a 3D printed mold). The dental appliances can include one or more of an aligner (e.g., an orthodontic aligner), a retainer, a progressive palatal expander, an attachment template, etc.

[0169] The instructions can be configured to control a manufacturing system or device to produce an orthodontic appliance having a specified orthodontic appliance. In some embodiments, the instructions are configured for direct manufacturing (e.g., stereolithography, selective laser sintering, fused deposition modeling, 3D printing, continuous direct manufacturing, multi-material direct manufacturing, etc.) to fabricate an orthodontic appliance using various methods presented herein. In alternative embodiments, the instructions can be configured for indirect manufacturing of the appliance, such as by 3D printing a mold and thermoforming a plastic sheet over the mold.

[0170] Method 1300 can include additional blocks: 1) intraorally scan the patient's upper dental arch and palate to generate three-dimensional data of the palate and upper dental arch; 2) determine the three-dimensional shape profile of the appliance to provide clearance and tooth engagement structures as described herein.

[0171] Although the above blocks illustrate a method 1300 for designing an orthodontic appliance according to some embodiments, those skilled in the art will recognize some variations based on the teachings described herein. Some blocks may include sub-blocks. Certain blocks may generally be repeated as needed. One or more blocks of method 1300 may be performed using any suitable manufacturing system or device, such as the embodiments described herein. Some blocks may be optional, and the order of the blocks may be changed as needed.

[0172] Figure 14 A method 1400 for digitally planning orthodontic treatment and / or designing or manufacturing an appliance according to an embodiment is shown. Method 1400 may be applied to any treatment procedure described herein and may be performed by any suitable data processing system.

[0173] In block 1410, a digital representation of a patient's teeth is received. The digital representation may include surface topography data of the patient's oral cavity (including teeth, gingival tissue, etc.). The surface topography data may be generated by directly scanning the oral cavity, a physical model (positive or negative mold) of the oral cavity, or an impression of the oral cavity using a suitable scanning device (e.g., a hand-held scanner, a desktop scanner, etc.).

[0174] In block 1420, one or more treatment phases are generated based on the digital representation of the teeth. Each treatment phase may include a 3D model of the dental arch generated in that treatment phase. The treatment phases may be progressive repositioning phases of an orthodontic treatment process, which are designed to move one or more of the patient's teeth from an initial tooth alignment to a target alignment. For example, the treatment phases may be generated by determining the initial tooth alignment indicated by the digital representation, determining the target tooth alignment, and determining the movement paths of one or more teeth in the initial alignment required to achieve the target tooth alignment. The movement paths may be optimized based on minimizing the total distance of movement, preventing collisions between teeth, avoiding more difficult-to-achieve tooth movements, or any other suitable criteria.

[0175] In block 1430, at least one orthodontic appliance is manufactured based on the generated treatment phases. For example, a set of appliances may be manufactured, each appliance being shaped according to the tooth alignment specified by one of the treatment phases, such that the appliances may be sequentially worn by the patient to progressively reposition the teeth from the initial alignment to the target alignment. The set of appliances may include one or more of the orthodontic appliances described herein. The manufacturing of the appliances may involve creating a digital model of the appliances to be used as an input to a computer-controlled manufacturing system. As needed, direct manufacturing methods, indirect manufacturing methods, or a combination thereof may be used to form the appliances.

[0176] In some cases, staging of the various alignments or treatment phases may not be necessary for the design and / or manufacturing of the appliances. As Figure 14As shown by the dashed lines in, the design and / or manufacture of an orthodontic appliance and possibly a specific orthodontic treatment can include using a representation of a patient's teeth (e.g., receiving a digital representation of the patient's teeth at block 1410), and then designing and / or manufacturing the orthodontic appliance based on the representation of the patient's teeth in the alignment represented by the received representation.

[0177] In some cases, for clarity of illustration, the present technology may be presented as including various functional blocks, which include functional blocks containing devices, device components, steps or routines in a method embodied in software or a combination of hardware and software.

[0178] Any step, operation, function or process described herein may be performed or implemented by hardware and software services or a combination of services alone or in combination with other devices. In some embodiments, a service may be software residing in the memory of a client device and / or one or more servers of a content management system, and when a processor executes the software associated with the service, the service performs one or more functions. In some embodiments, a service is a program or a collection of programs that perform a specific function. In some embodiments, a service may be regarded as a server. The memory may be a non-transitory computer-readable medium.

[0179] In some embodiments, a computer-readable storage device, medium, and memory may include a cable or wireless signal containing a bitstream, etc. However, when mentioned, non-transitory computer-readable storage media explicitly exclude media such as energy, carrier signals, electromagnetic waves, and signals themselves.

[0180] The method according to the above example can be implemented using computer-executable instructions stored in a computer-readable medium or otherwise obtainable from a computer-readable medium. Such instructions may include, for example, instructions and data that cause or otherwise configure a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform a certain function or group of functions. Portions of the computer resources used may be accessed via a network. The computer-executable instructions may be, for example, binaries, intermediate format instructions (such as assembly language), firmware, or source code. Examples of computer-readable media that may be used to store instructions, the information used, and / or the information created during the method according to the described example include magnetic or optical disks, solid-state memory devices, flash memory, USB devices equipped with non-volatile memory, networked storage devices, and the like.

[0181] Apparatuses for implementing the methods according to these disclosures may include hardware, firmware, and / or software, and may take any of a variety of form factors. Typical examples of such form factors include servers, laptop computers, smart phones, small personal computers, personal digital assistants, etc. The functions described herein may also be embodied in peripheral devices or add-on cards. As another example, such functions may also be implemented on a circuit board between different chips or in different processes executed in a single device.

[0182] Instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are means for providing the functions described in these disclosures.

[0183] Although this disclosure is described with respect to a specific application for dental treatment planning for humans, this disclosure is not limited thereto. The techniques described herein can equally be applied to any other medical application and planning. For example, the techniques described can be used for treatment planning for any other feature or element of the human body (e.g., eyes, nose, other facial elements, hands, legs, hips, feet, etc.). Additionally, the techniques described herein can equally be used for the purpose of treatment planning for animals and their body features, just as for humans.

[0184] Although various examples and other information are used to explain aspects within the scope of the appended claims, no limitation to the claims should be implied based on specific features or arrangements in such examples, because a person of ordinary skill in the art will be able to use these examples to arrive at various implementations. Additionally, although some subject matter may have been described in language specific to examples of structural features and / or method steps, it should be understood that the subject matter defined in the appended claims need not be limited to these described features or acts. For example, such functions may be distributed or performed differently in components other than those identified herein. Rather, the described features and steps are disclosed as examples of components of systems and methods within the scope of the appended claims.

[0185] Claim language reciting "at least one" and / or "one or more" in a set, or other language indicating that one member or more members (in any combination) of the set satisfy the claim. For example, claim language reciting "at least one of A and B" or "at least one of A or B" indicates A, B, or A and B. In another example, claim language reciting "at least one of A, B, and C" or "at least one of A, B, or C" indicates A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language of "at least one" and / or "one or more" in a set does not limit the set to the items listed in the set. For example, claim language reciting "at least one of A and B" or "at least one of A or B" can indicate A, B, or A and B, and can additionally include items not listed in the set of A and B.

Claims

1. A method, comprising: using at least one device to capture media of a patient's face from multiple angles; converting the media into a three-dimensional representation of the patient's face, the conversion comprising: determining a plurality of camera poses, wherein each camera pose in the plurality of camera poses is determined for a two-dimensional image of the media; generating a plurality of depth maps, wherein each depth map in the plurality of depth maps is generated for the two-dimensional image based at least in part on the camera pose of the two-dimensional image of the media; and generating the three-dimensional representation of the patient's face based on combining information from the plurality of depth maps; and transmitting the three-dimensional representation of the patient's face to one or more processing components for integration with a three-dimensional representation of an intraoral scan of the patient to visualize the patient's dental treatment plan.

2. The method according to claim 1, wherein, the at least one device is a mobile device having a built-in camera.

3. The method according to claim 2, wherein, the media comprises a plurality of two-dimensional images of the patient's face captured by moving the mobile device around the patient's face.

4. The method according to claim 1, wherein, the media comprises a video of the patient's face.

5. The method according to claim 4, wherein, the video is captured using a multi-camera system configured to simultaneously capture images of the patient's face from multiple angles, and the at least one device is the multi-camera system.

6. The method according to claim 1, wherein, the three-dimensional representation of the intraoral scan of the patient is used as a reference for scaling the three-dimensional representation of the patient's face.

7. The method according to claim 1, wherein, the media is captured via an application installed on the at least one device, and the application has access to the built-in camera of the at least one device.

8. The method according to claim 7, wherein, the application provides real-time guidance for moving the at least one device around the patient's face to optimize the captured media.

9. A method for generating a three-dimensional representation of a patient's face for dental treatment planning, the method comprising: receiving, at a processing component communicatively coupled to at least one device, the three-dimensional representation of the patient's face based on media of the patient's face captured by the at least one device from multiple angles; and integrating, at the processing component, the three-dimensional representation of the patient's face with a three-dimensional representation of an intraoral scan of the patient to visualize the patient's dental treatment plan.

10. The method according to claim 9, wherein, integrating the three-dimensional representation of the patient's face with the three-dimensional representation of the intraoral scan of the patient comprises: identifying an intraoral region of the patient in the three-dimensional representation of the patient's face; removing the intraoral region from the three-dimensional representation of the patient's face; determining a rigid relative transformation for scaling the three-dimensional representation of the patient's face to the three-dimensional representation of the intraoral scan of the patient; replacing the intraoral region removed from the three-dimensional representation with the three-dimensional representation of the intraoral scan of the patient; and Align the three-dimensional representation of the patient's face with the three-dimensional representation of the intraoral scan of the patient based on the scaled rigid relative transformation.

11. The method according to claim 10, wherein, using a machine learning-based model to identify the intraoral region in a two-dimensional space.

12. The method according to claim 9, wherein, the three-dimensional representation of the patient's intraoral scan provides a visualization of the patient's teeth after completion of the dental treatment plan.

13. A system, comprising: at least one device; and a processing component communicatively coupled to the at least one device, wherein the at least one device is configured to: capture media of the patient's face from multiple angles, convert the media into a three-dimensional representation of the patient's face, and transmit the three-dimensional representation of the patient's face to the processing component; and the processing component is configured to: integrate the three-dimensional representation of the patient's face with the three-dimensional representation of the patient's intraoral scan to visualize the patient's dental treatment plan.

14. The system according to claim 13, wherein, the processing component is further configured to receive the three-dimensional representation of the patient's intraoral scan from a dental scanner.

15. The system according to claim 13, wherein, the processing component is further configured to integrate with one or more dental treatment planning applications to use the visualization of the dental treatment plan.

16. The system according to claim 13, wherein, the at least one device is a mobile device with a built-in camera.

17. The system according to claim 16, wherein, the media includes a plurality of two-dimensional images of the patient's face captured by moving the mobile device around the patient's face.

18. The system according to claim 13, wherein, the media includes a video of the patient's face.

19. The system according to claim 18, wherein, the video is captured using a multi-camera system configured to simultaneously capture images of the patient's face from multiple angles, and the at least one device is the multi-camera system.

20. The system according to claim 18, wherein, the result of integrating the three-dimensional image of the patient's face with the three-dimensional image of the patient's intraoral scan allows the operator of the at least one device to visualize the movement of facial tissues, lips, and facial expressions after the dental treatment plan is applied to the patient's teeth.

21. The system according to claim 13, wherein, a mobile application on the at least one device is configured to provide real-time guidance for moving the at least one device around the patient's face to optimize the captured media.

22. The system according to claim 13, wherein, the processing component is configured to integrate the three-dimensional representation of the patient's face with the three-dimensional representation of the patient's intraoral scan by: identifying the intraoral region of the patient in the three-dimensional representation of the patient's face; removing the intraoral region from the three-dimensional representation of the patient's face; Determine a scaled rigid relative transformation of a three-dimensional representation of the patient's face to a three-dimensional representation of the patient's intraoral scan; Replace the intraoral region removed from the three-dimensional representation with the three-dimensional representation of the patient's intraoral scan; and Align the three-dimensional representation of the patient's face with the three-dimensional representation of the patient's intraoral scan based on the scaled rigid relative transformation.

23. The system according to claim 22, wherein, the intraoral region is identified in a two-dimensional space using a machine learning-based model.

24. The system according to claim 13, wherein, the processing component is configured to output a final result of integrating the three-dimensional representation of the patient's face with the three-dimensional representation of the patient's intraoral scan to one or more dental planning tools for further analysis.

25. The system according to claim 13, wherein, the media includes a plurality of two-dimensional images of the patient's face, and wherein, in order to convert the media into a three-dimensional representation of the patient's face, the at least one device is configured to: Determine a plurality of camera poses, wherein each camera pose in the plurality of camera poses is determined for a two-dimensional image in the plurality of two-dimensional images; Generate a plurality of depth maps, wherein each depth map in the plurality of depth maps is generated for the two-dimensional image based at least in part on the camera pose for the two-dimensional image in the plurality of two-dimensional images; and Generate a three-dimensional representation of the patient's face based on combining information from the plurality of depth maps.

26. A system, comprising: At least one media capture device; and A cloud-based processing component communicatively coupled to the at least one media capture device, wherein the at least one media capture device is configured to: Capture media of the patient's face from at least one angle, Convert the media into a three-dimensional representation of the patient's face, and Transmit the three-dimensional representation of the patient's face to the cloud-based processing component; and the cloud-based processing component is configured to: Integrate the three-dimensional representation of the patient's face with a three-dimensional representation of a patient's intraoral scan to visualize a dental treatment plan for the patient.