Dental treatment video

By aligning and adjusting the dental images, a synthetic video of the dental treatment process was generated, and a clear display of treatment progress was achieved.

CN120476429APending Publication Date: 2025-08-12ALIGN TECHNOLOGY INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380090225.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-10-27
Filing Date
2023-10-30
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The prior art cannot effectively generate continuous images of dental treatment processes, and it is difficult for patients to understand and evaluate changes before and after dental treatment.

Method used

By receiving and aligning multiple teeth images, color adjustments and transformations are used to generate synthetic images and videos to show the intermediate state of the dental treatment process.

Benefits of technology

Clearly demonstrate changes in dental treatment, helping doctors and patients understand treatment progress and outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120476429A_ABST
    Figure CN120476429A_ABST
Patent Text Reader

Abstract

A method and / or system generates a video of a tooth of an individual over time. In one example, images including teeth of an individual are received, where the images are arranged in a sequence and each image is associated with a different stage of a dental treatment. One or more of the images are modified and / or replaced to align the images with each other. One or more composite images are generated, where each composite image is generated based on a sequence image pair in the sequence and is an intermediate image comprising an intermediate state of the tooth between a first state of a first image of the sequence image pair and a second state of a second image of the sequence image pair. A video comprising the image and the one or more composite images is then generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of dentistry, and in particular to a system and method for generating a video of a dental treatment result from captured images and / or simulated images. Background Art

[0002] Generating images showing how a patient's teeth will appear after treatment would be helpful for dentists and patients considering orthodontic and / or other dental treatments. Furthermore, for patients considering treatment, it would be helpful to view images of other patients at different stages of treatment. Furthermore, for patients already undergoing treatment, it would be helpful to view images of their treatment at different stages. However, existing techniques, at best, generate a series of disconnected images that are difficult for patients to interpret and glean information from. Summary of the Invention

[0003] Various examples of embodiments of the present disclosure are provided. These examples should not be construed as limiting, but are for illustrative purposes only.

[0004] In a first example embodiment, a method includes: receiving a plurality of images comprising teeth of an individual, wherein the plurality of images are arranged in a sequence and each of the plurality of images is associated with a different stage of dental treatment; performing at least one of modifying one or more of the plurality of images or replacing one or more of the plurality of images so that the plurality of images are aligned with each other; generating one or more composite images, wherein each of the one or more composite images is generated based on a sequence image pair (a pair of sequence images) in a sequence and is an intermediate image comprising an intermediate state of the tooth between a first state of a first image of the sequence image pair and a second state of a second image of the sequence image pair; and generating a video comprising the plurality of images and the one or more composite images.

[0005] A second example embodiment may further extend the first example embodiment. In the second example embodiment, modifying one or more images of the plurality of images includes modifying a color of the plurality of images such that the color remains consistent across the plurality of images.

[0006] A third example embodiment may further expand upon the first example embodiment or the second example embodiment. In the third example embodiment, modifying the colors of the plurality of images includes inputting the plurality of images into a trained machine learning model, wherein the trained machine learning model outputs a color modification for one or more of the plurality of images.

[0007] A fourth exemplary embodiment may further extend the third exemplary embodiment. In the fourth exemplary embodiment, the trained machine learning model includes a convolutional neural network that performs one or more wavelet transforms.

[0008] The fifth example embodiment may further extend any one of the first to fourth example embodiments. In the fifth example embodiment, modifying an image in the one or more images includes performing at least one of a translation, rotation, or scaling change on one or more points of the image.

[0009] A sixth example embodiment may further extend the fifth example embodiment. In the sixth example embodiment, the method further includes: detecting a plurality of features common to at least some of the plurality of images; and determining, for a sequence image pair in the plurality of images in the sequence and one or more features of the plurality of features, an affine transformation of the feature between a first image and a second image of the sequence image pair, wherein applying the affine transformation to at least one of the first image or the second image results in at least one of a translation, rotation, or scaling change of one or more points of the image.

[0010] A seventh exemplary embodiment may further extend the sixth exemplary embodiment. In the seventh exemplary embodiment, detecting a plurality of features of an image includes inputting the image into a trained machine learning model, wherein the trained machine learning model outputs a location of each of the plurality of features in the image.

[0011] The eighth exemplary embodiment may further expand the sixth exemplary embodiment or the seventh exemplary embodiment. In the eighth exemplary embodiment, the plurality of features include one or more of teeth.

[0012] A ninth exemplary embodiment may further extend any one of the sixth to eighth exemplary embodiments. In the ninth exemplary embodiment, the plurality of images include facial images of an individual, wherein teeth of the individual are visible in the plurality of facial images, and wherein the plurality of features include one or more facial features.

[0013] The tenth exemplary embodiment may further extend any one of the fifth to ninth exemplary embodiments. In the tenth exemplary embodiment, replacing one or more images in the plurality of images includes: generating, for an image in the one or more images, a replacement image having: a) teeth corresponding to the teeth in the image; and b) one or more features different from the one or more features in the image and similar to the one or more features in an additional image in the plurality of images, wherein the replacement image is used to replace the image.

[0014] The eleventh exemplary embodiment may further expand upon the tenth exemplary embodiment. In the eleventh exemplary embodiment, generating the replacement image includes processing the image and the additional image using a trained machine learning model, wherein the trained machine learning model outputs the replacement image.

[0015] The twelfth exemplary embodiment can further extend the eleventh exemplary embodiment. In the twelfth exemplary embodiment, the trained machine learning model is a generative model.

[0016] The thirteenth exemplary embodiment may further extend any one of the tenth to twelfth exemplary embodiments. In the thirteenth exemplary embodiment, the one or more features of the image include a first camera viewpoint, and wherein the one or more features of the additional image include a second camera viewpoint.

[0017] The fourteenth exemplary embodiment may further extend the tenth exemplary embodiment to the thirteenth exemplary embodiment. In the fourteenth exemplary embodiment, the one or more features of the image include at least one of the following: a first facial expression, a first jaw position, a first upper and lower jaw relationship, a first color, a first lighting condition, tooth occlusion, a tooth accessory, a first hairstyle, or a first clothing item; and the one or more features of the additional image include at least one of the following: a second facial expression, a second jaw position, a second upper and lower jaw relationship, a second color, a second lighting condition, no tooth occlusion, no tooth accessory, a second hairstyle, or a second clothing item.

[0018] The fifteenth exemplary embodiment may further extend any one of the first to fourteenth exemplary embodiments. In the fifteenth exemplary embodiment, replacing one or more images from the plurality of images includes: generating, for an image from the one or more images, a replacement image having: a) teeth corresponding to the teeth in the image; and b) one or more features different from the one or more features in the image, wherein the replacement image is used to replace the image.

[0019] A sixteenth exemplary embodiment may further extend the fifteenth exemplary embodiment. In the sixteenth exemplary embodiment, generating a replacement image includes: receiving an input selecting one or more target features; and processing the image and the input using a trained machine learning model, wherein the trained machine learning model outputs a replacement image having one or more features corresponding to the one or more target features.

[0020] The seventeenth exemplary embodiment may further extend any one of the first to sixteenth exemplary embodiments. In the seventeenth exemplary embodiment, generating a composite image in the one or more composite images includes: determining, for a sequence image pair in a plurality of images in a sequence, an optical flow between a first image and a second image in the sequence image pair; and generating the composite image based on the optical flow.

[0021] The eighteenth exemplary embodiment may further extend any one of the first to seventeenth exemplary embodiments. In the eighteenth exemplary embodiment, generating a composite image in one or more composite images includes: inputting a sequence image pair from a plurality of images in a sequence into a trained machine learning model, wherein the trained machine learning model outputs the composite image.

[0022] The nineteenth exemplary embodiment may further extend the eighteenth exemplary embodiment. In the nineteenth exemplary embodiment, the trained machine learning model includes a generative model.

[0023] The twentieth example implementation may further extend the nineteenth example implementation.In the twentieth example implementation, one or more layers of the generative model determine optical flow between pairs of sequential images, and wherein the generative model uses the optical flow to generate the synthetic image.

[0024] A twenty-first example implementation may further extend any one of the first to twentieth example implementations. In the twenty-first example implementation, generating a composite image in one or more composite images includes: transforming a first image and a second image in a sequence into a feature space; determining an optical flow between the first image and the second image in the feature space; and generating a composite image using the optical flow in the feature space, the composite image being an intermediate image between the first image and the second image.

[0025] The twenty-second exemplary embodiment may further extend any one of the first to twenty-first exemplary embodiments. In the twenty-second exemplary embodiment, generating one or more composite images includes: generating a first composite image based on a first image and a second image in a sequence, the first composite image being an intermediate image between the first image and the second image; and generating a second composite image based on the first image and the first composite image, the second composite image being an intermediate image between the first image and the first composite image.

[0026] The twenty-third exemplary embodiment may further extend the twenty-second exemplary embodiment. In the twenty-third exemplary embodiment, the method further comprises: determining a similarity score between the first image and the first composite image; and generating a second composite image in response to determining that the similarity score does not satisfy a similarity threshold.

[0027] The twenty-fourth exemplary embodiment may further extend any one of the first to twenty-third exemplary embodiments. In the twenty-fourth exemplary embodiment, a non-transitory computer-readable medium includes instructions that, when executed by a processing device, cause the processing device to perform the method of any one of the first to twenty-third exemplary embodiments.

[0028] The twenty-fifth exemplary embodiment may further extend any one of the first to twenty-third exemplary embodiments. In the twenty-fifth exemplary embodiment, a system includes: a processing device; and a memory for storing instructions that, when executed by the processing device, cause the processing device to perform the method of any one of the first to twenty-third exemplary embodiments.

[0029] In a twenty-sixth example embodiment, a method includes: receiving an image including a current state of an individual's teeth; receiving or generating a treatment plan, the treatment plan including a three-dimensional (3D) model of a future state of the teeth during a treatment phase; generating a first composite image including the future state of the teeth during a treatment phase based on the received image and the 3D model of the future state of the teeth during a treatment phase; generating one or more additional composite images, the one or more additional composite images being intermediate images between the received image and the first composite image; and generating a video including the received image, the one or more additional composite images, and the first composite image.

[0030] The twenty-seventh exemplary embodiment may further extend the twenty-sixth exemplary embodiment. In the twenty-seventh exemplary embodiment, the received image has a facial image of the individual, wherein teeth of the individual are visible in the received facial image, and wherein the first composite image and the one or more additional composite images have facial images.

[0031] The twenty-eighth exemplary embodiment may further extend the twenty-sixth exemplary embodiment or the twenty-seventh exemplary embodiment. In the twenty-eighth exemplary embodiment, the treatment plan further includes a second 3D model of a second future state of the teeth at a second treatment stage, and the method further includes: generating a second composite image based on the received image and the second 3D model of the second future state of the teeth at the second treatment stage, the second composite image including the second future state of the teeth at the second treatment stage; and generating one or more additional composite images, the one or more additional composite images being intermediate images between the first composite image and the second composite image; wherein the video further includes the one or more additional composite images and the second composite image.

[0032] The twenty-ninth exemplary embodiment may further extend the twenty-sixth exemplary embodiment to the twenty-eighth exemplary embodiment. In the twenty-ninth exemplary embodiment, generating one or more additional composite images includes: determining an optical flow between the received image and the first composite image; and generating one or more additional composite images based on the optical flow.

[0033] The 30th exemplary embodiment may further extend the 26th to 29th exemplary embodiments. In the 30th exemplary embodiment, generating one or more additional composite images includes: inputting the received image and the first composite image into a trained machine learning model, wherein the trained machine learning model outputs the one or more additional composite images.

[0034] The thirty-first exemplary embodiment may further extend the thirtieth exemplary embodiment. In the thirty-first exemplary embodiment, the trained machine learning model includes a generative model.

[0035] The thirty-second example embodiment may further extend the thirty-first example embodiment. In the thirty-second example embodiment, one or more layers of the generative model determine an optical flow between the received image and the first composite image, and wherein the generative model uses the optical flow to generate one or more additional composite images.

[0036] The thirty-third exemplary embodiment may further extend any one of the twenty-sixth to thirty-second exemplary embodiments. In the thirty-third exemplary embodiment, generating one or more additional synthetic images includes: transforming the received image and the first synthetic image into a feature space; determining an optical flow between the received image and the first synthetic image in the feature space; and generating one or more additional synthetic images using the optical flow in the feature space, the one or more additional synthetic images being intermediate images between the received image and the first synthetic image.

[0037] The thirty-fourth exemplary embodiment may further extend any one of the twenty-sixth to thirty-third exemplary embodiments. In the thirty-fourth exemplary embodiment, generating one or more additional composite images includes: generating a second composite image that is an intermediate image between the received image and the first composite image; and generating a third composite image based on the received image and the second composite image, the third composite image being an intermediate image between the received image and the second composite image.

[0038] The thirty-fifth exemplary embodiment may further extend the thirty-fourth exemplary embodiment. In the thirty-fifth exemplary embodiment, the method further comprises: determining a similarity score between the received image and the second composite image; and generating a third composite image in response to determining that the similarity score does not meet the similarity threshold.

[0039] The thirty-sixth exemplary embodiment may further extend any one of the twenty-sixth exemplary embodiment to the thirty-fifth exemplary embodiment. In the thirty-sixth exemplary embodiment, a non-transitory computer-readable medium includes instructions that, when executed by a processing device, cause the processing device to perform the method of any one of the twenty-sixth exemplary embodiment to the thirty-sixth exemplary embodiment.

[0040] The thirty-seventh exemplary embodiment may further extend any one of the twenty-sixth to thirty-fifth exemplary embodiments. In the thirty-seventh exemplary embodiment, a system includes: a processing device; and a memory for storing instructions that, when executed by the processing device, cause the processing device to perform the method of any one of the twenty-sixth to thirty-sixth exemplary embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Embodiments of the disclosure are illustrated by way of example and not limitation in the figures of the accompanying drawings.

[0042] Figure 1 One embodiment of a system for treatment planning and / or smile video generation according to an embodiment is shown.

[0043] Figure 2 The model training workflow and model application workflow of the smile processing module according to an embodiment of the present disclosure are shown.

[0044] Figures 3A to 3E A flow chart illustrating a method of generating a video of a dental treatment result according to an embodiment is shown.

[0045] Figure 4A An example input image input into an image replacement module for image replacement operation and an output image generated by the image replacement module according to an embodiment are shown.

[0046] Figure 4B An input image into a color conversion model and an example output image generated by the color conversion model are shown according to an embodiment.

[0047] Figure 4C An example input image into a landmark detection model and the output of the landmark detection model are shown according to an embodiment.

[0048] Figure 4D An example input image with identified landmarks input into an affine transformation model and a modified image to which one or more determined affine transformations have been applied are shown in accordance with an embodiment.

[0049] Figure 4EAn example input image into an image generation model and an output image of the image generation model according to an embodiment are shown.

[0050] Figures 5A to 5C Various stages of a recursive composite image generation process for creating a video (eg, a video of a dental treatment over time) in an embodiment are shown, according to an embodiment.

[0051] Figures 5D to 5F Various stages of a recursive composite image generation process for creating a video (eg, a video of a dental treatment over time) in an embodiment are shown, according to an embodiment.

[0052] Figures 6A to 6B A flow chart illustrating a method of generating a video of a dental treatment over time after treatment has begun, according to an embodiment.

[0053] Figure 7 A flow chart illustrating a method of generating a video of a dental treatment over time before the treatment begins, according to an embodiment.

[0054] Figure 8 A flow chart illustrating a method for generating a simulated image of a dental treatment result according to an embodiment is shown.

[0055] Figure 9 Also shown is a flow chart of a method for generating a simulated image of a dental treatment result according to an embodiment.

[0056] Figure 10 A block diagram of an example computing device is shown, in accordance with an embodiment of the present disclosure. DETAILED DESCRIPTION

[0057] Described herein are methods and systems for generating videos of an individual's face, smile, and / or dentition over time, in accordance with embodiments of the present disclosure. In embodiments, the methods and systems described herein can convert one or more images or simulated images into a short video. These methods can be applied to demonstrate changes in dental health over time, the progress of dental treatments (e.g., orthodontic treatment and / or prosthodontic treatment) over time, and the like. The methods and systems described herein can be used for past treatments, ongoing treatments, and / or future treatments. In an example, a treatment video can be generated that shows before and after visualizations of treatments such as orthodontic treatment, restorative treatment, etc., to demonstrate how tooth shape, position, and / or orientation has changed or will change over time.

[0058] In an embodiment, an image processing pipeline is applied to images to fully automatically transform these images into short videos. A machine learning model (e.g., a neural network) can be constructed to perform operations such as keypoint detection, segmentation, style transfer, image generation, and / or frame interpolation in the image processing pipeline. Frame interpolation can be performed using a learning-based hybrid data-driven approach that estimates the movement between images even for irregular input data, so that the output can be combined into a visually smooth animated image. Frame interpolation can also be performed in a manner that can handle disocclusion, which is common in open bite images.

[0059] In an embodiment, the actual treatment images may be captured at irregular time intervals. Therefore, in some embodiments, motion estimation techniques are used to control the generation of intermediate frames, thereby obtaining a visually smooth output video.

[0060] The techniques described herein for generating video descriptions showing changes in dentition over time are also applicable to many other areas. For example, the techniques described herein for generating video descriptions showing changes in dentition over time can be used to generate occlusal views showing a person's dental arch over time, the visual effects of restorative dental treatment on tooth shape, the visual effects of removing accessories (e.g., accessories for orthodontic treatment), videos showing changes in a person's face and / or body over time (e.g., showing the effects of aging, which can take into account features such as changes in wrinkle progression), videos showing changes in plant conditions over time, videos showing changes in geographic locations over time, videos showing changes in houses or other buildings over time, and so on. Therefore, it should be understood that the examples described in connection with teeth, dentition, smiles, etc. are also applicable to any other type of object, person, organism, place, etc. whose condition or state may change over time. Therefore, in an embodiment, the techniques set forth herein can be used to generate videos showing, for example, any type of object, person, organism, place, etc. changing over time.

[0061] In the field of dentistry, a doctor, technician, or patient may periodically generate images of their smile, teeth, etc. The doctor, technician, patient, etc. may then view these periodically generated images sequentially when trying to determine how the patient's dentition has changed over time. However, it may be difficult to assess how a patient's dentition has changed (or may not have changed) over time simply by viewing a series of images. Each image will typically have different lighting conditions, be taken from a different angle and / or camera perspective, have a different zoom setting, be a different size, be a different color, show a different facial expression, show a different number of one or more teeth, and so on. These differences can make it challenging to determine which aspects of a patient's dentition have changed and which aspects have not changed over time.

[0062] Thus, in embodiments, systems and / or methods operate on discrete images to align them spatially (e.g., by scaling, reorienting, translating, etc.) and color (e.g., by color balancing, color matching, etc.) between the images. In embodiments, systems and / or methods process images to determine whether any images fail to meet one or more quality criteria (e.g., alignment criteria). For example, images captured from different camera perspectives may contain different data relative to the imaged face. After aligning the images and generating a video from the aligned images, the video may exhibit noticeable changes and / or movement between the images that should not exist. For example, if an image is captured from a different camera perspective than that used for other images, the image may not show teeth that are shown in the other images. After aligning the image with the other images, the image may still lack information about teeth that were not captured in the image. Consequently, the missing teeth in the image will appear as movement or change between the video frame based on the image and the video frames based on the other images. Thus, the one or more quality criteria may include a camera perspective criterion. An image whose camera perspective differs from the camera perspective of the other images by more than a threshold may fail to meet the camera perspective criterion.

[0063] In some embodiments, a replacement image is generated for one or more images that do not meet one or more quality standards. The replacement image can be generated by a trained machine learning model (e.g., a generative model). In an embodiment, a generative model such as StyleGAN is used to generate the replacement image. The generative model can receive as input the image to be replaced and style information for generating the replacement image. The style information can be received in the form of an additional image whose style information is to be copied and / or input indicating one or more style selections (e.g., values indicating lighting information, color information, pose information, camera perspective, etc.). The replacement image can include some details from the image to be replaced and some details based on the input style information (e.g., details of the additional input image). The original image that does not meet the quality standards can then be replaced with a synthetically generated image that should meet one or more quality standards.

[0064] The system and / or method can generate additional composite images, essentially interpolated images, that illustrate what the dentition might have looked like between the times the images were actually taken. The composite images are generated in a manner that aligns them color-wise and spatially with the captured images. The modified and composite images are then used to generate a video, where each image can be a frame of the video. The video can then be presented to a physician, patient, or the like to clearly and smoothly demonstrate the progression of the patient's dentition over time. If the images are captured at the beginning of dental treatment and then at various stages of dental treatment, the video demonstrates the progression of dental treatment over time. If the images are captured without dental treatment (e.g., as a physician would naturally do at each patient visit), the video can be used to demonstrate the progression of one or more dental conditions (e.g., tooth wear, gum recession, gum swelling, caries, discoloration, etc.). Such videos serve as a powerful tool, enabling physicians to clearly demonstrate the progression of dental problems to patients and can help persuade them to undergo treatment. Furthermore, treatment videos can be used to demonstrate the transformation of the patient's dentition during and after treatment. Such videos of past patients can also be shown to potential patients to demonstrate examples of treatment.

[0065] In some embodiments, the system and / or method generates a synthetic (simulated) image of the patient's future dentition based on current or past images of the patient's dentition and information in the treatment plan. Using the three-dimensional (3D) model in the treatment plan and the current or past images of the patient's dentition, accurate images of the patient's dentition at one or more future treatment stages can be generated. These synthetic images can then be used with the current or past images to generate additional synthetic images that are essentially interpolated images showing what the dentition might look like between the current or past image and the synthetic image associated with a particular treatment stage. A video is then generated using the current or past images and the synthetic images, where each image can be a frame of the video. The video can then be presented to a physician, patient, etc. to clearly and smoothly show the possible progression of the patient's dentition over time if they receive dental treatment.

[0066] A consumer smile simulation is a simulated image or video generated for a consumer (e.g., a patient) that shows how their smile will look after receiving a certain type of dental treatment (e.g., orthodontic treatment, etc.). A clinical smile simulation is a simulated image or video generated for a dental professional (e.g., an orthodontist, a dentist, etc.) to evaluate how a patient's smile will look after receiving a certain type of dental treatment. For both consumer smile simulations and clinical smile simulations, the goal is to generate a realistic photo-rendering of the patient at mid-treatment or post-treatment, which can be used by patients, potential patients, and / or dentists to view the results of the treatment. For both use cases, the general process of generating a simulated image showing a mid-treatment or post-treatment smile includes: taking a photograph of the patient's current smile; simulating or generating a treatment plan for the patient, which indicates the position and orientation of the teeth and gums at mid-treatment and / or post-treatment; and converting the data from the treatment plan back into a new simulated image showing the mid-treatment and / or post-treatment smile. Embodiments generate smile video simulations by further generating interpolated images (composite images showing intermediate states between treatment stages) and then aggregating all images together into a video.

[0067] Figure 1An embodiment of a treatment planning and / or smile video generation system 100 is shown. In one embodiment, the system 100 includes a computing device 105 and a data store 110. Additionally, the system 100 may include or be connected to an image capture device, such as a camera and / or an intraoral scanner. The computing device 105 may include a physical machine and / or a virtual machine hosted by the physical machine. The physical machine may be a rack-mounted server, a desktop computer, or other computing device. The physical machine may include a processing device, memory, auxiliary storage, one or more input devices (e.g., such as a keyboard, mouse, tablet, speakers, etc.), one or more output devices (e.g., a display, printer, etc.), and / or other hardware components. In one embodiment, the computing device 105 includes one or more virtual machines, which may be managed and provided by a cloud provider system. Each virtual machine provided by the cloud service provider may be hosted on one or more physical machines. The computing device 105 may be connected to the data store 110 directly or via a network. The network may be a local area network (LAN), a public wide area network (WAN) (e.g., the Internet), a private WAN (e.g., an intranet), or a combination thereof.

[0068] Data storage 110 may be an internal data storage, or may be an external data storage connected directly to or via a network to computing device 105. Examples of network data storage include storage area networks (SANs), network attached storage (NASs), and storage services provided by cloud provider systems. Data storage 110 may include one or more file systems, one or more databases, and / or other data storage devices.

[0069] Computing device 105 may receive one or more images from an image capture device, from multiple image capture devices, from data store 110, and / or from other computing devices. For example, the image capture device may be or include a charge-coupled device (CCD) sensor and / or a complementary metal oxide semiconductor (CMOS) sensor. For example, the image capture device may include a mobile device (e.g., a mobile phone) that may belong to the patient. Alternatively or additionally, the image capture device may include, for example, a doctor's camera. The image capture device may provide images or videos to computing device 105 for processing. For example, the image capture device may provide images to computing device 105, which may analyze the images to identify the patient's mouth, the patient's face, the patient's dental arches, etc. In some embodiments, the images captured by the image capture device may be stored in data store 110 as captured images 135. For example, captured images 135 may be stored in data store 110 as a patient history record, or used by computing device 105 to analyze the patient and / or generate simulated post-treatment images and / or videos (e.g., smile videos, dental stage progression videos, etc.). The image capture device may transmit discrete images and / or video to computing device 105, and computing device 105 may store captured images 135 in data store 110. In some embodiments, captured images 135 include two-dimensional data.

[0070] In some embodiments, the image capture device is a device located in a doctor's office. In some embodiments, the image capture device is a device of the patient. For example, the patient can use a webcam, mobile phone, tablet, laptop, digital camera, etc. to take one or more photos of their teeth, smile, and / or face. The patient can then send these photos to computing device 105, which can then be stored as captured images 135 in data storage 110.

[0071] In an embodiment, the computing device 105 includes a smile manipulation module 108 and a treatment planning module 120. The treatment planning module 120 is responsible for generating a treatment plan 158 that includes treatment outcomes for the patient. In an embodiment, the treatment plan may be stored in the data store 110. The treatment plan 158 may include and / or be based on one or more 2D images and / or intraoral scans of the patient's dental arch. For example, the treatment planning module 120 may receive a 3D intraoral scan of the patient's dental arch that is obtained based on an intraoral scan performed using an intraoral scanner. An example of an intraoral scanner is a 3D intraoral scanner manufactured by Align Technology, Inc. Intraoral Digital Scanner. Another example of an intraoral scanner is described in U.S. Publication No. 2019 / 0388193, filed June 19, 2019, which is incorporated herein by reference.

[0072] During an intraoral scan, the intraoral scanning application receives and processes intraoral scan data (e.g., intraoral scans) and, based on such processing, generates a 3D surface of a scanned area of the oral cavity (e.g., a dental site). To generate the 3D surface, the intraoral scanning application may register and "stitch" or merge together the intraoral scans generated during the intraoral scan in real time or near real time as the scan progresses. Once the scan is complete, the intraoral scanning application may then register and stitch or merge the intraoral scans together again using a more accurate and resource-intensive sequence of operations. In one embodiment, performing the registration includes capturing 3D data for various points of the surface in multiple scans (views from a camera) and registering the scans by calculating a transformation between the scans. The 3D data may be projected into 3D space for transformation and stitching. The scans may be integrated into a common reference frame by applying an appropriate transformation to the points of each registered scan and projecting each scan into 3D space.

[0073] In one embodiment, registration is performed on adjacent or overlapping intraoral scans (e.g., each consecutive frame of an intraoral video). A registration algorithm is performed to register two or more adjacent intraoral scans and / or to register the intraoral scans with a 3D surface that has been generated, which primarily involves determining a transformation that aligns one scan with another scan and / or with a 3D surface. Registration can involve identifying multiple points in each scan (e.g., a point cloud) of a scan pair (or a scan and a 3D model), performing a surface fit on the points, and matching the points of the two scans (or the scan and the 3D surface) using a local search around the points. For example, the intraoral scanning application can match points of one scan with the nearest interpolated points on the surface of another image and iteratively minimize the distance between the matched points. Other registration techniques can also be used. The intraoral scanning application can repeat the registration and stitching for all scans in a series of intraoral scans and update the 3D surface as scans are received.

[0074] The treatment planning module 120 can perform treatment planning in an automated manner and / or based on input from a user (e.g., a dental technician). The treatment planning module 120 can receive and / or store a 3D model 150 of the patient's current dental arch and can then determine the current position and orientation of the patient's teeth based on the virtual 3D model 150, and determine a target final position and orientation of the patient's teeth represented as a treatment outcome (e.g., a final treatment stage). The treatment planning module 120 can then generate a virtual 3D model 150 showing the patient's dental arch at the end of treatment, as well as one or more virtual 3D models 150 showing the patient's dental arch at various intermediate stages of treatment. Alternatively or additionally, the treatment planning module 120 can generate one or more 3D images and / or 2D images showing the patient's dental arch, teeth, smile, etc. at various stages of treatment.

[0075] As a non-limiting example, treatment results can be the result of various dental procedures. Such dental procedures can be roughly divided into oral restorative (dental restoration) procedures and orthodontic procedures, which are then further divided into the specific forms of these procedures. In addition, dental procedures can include the identification and treatment of gum disease, sleep apnea, and intraoral conditions. The term oral restorative procedure refers in particular to any procedure involving the oral cavity and for designing, manufacturing, or installing a dental prosthesis or its real or virtual model at a dental site in the oral cavity, or for designing and preparing a dental site to receive such a prosthesis. For example, a prosthesis can include any restoration, such as an implant, crown, veneer, inlay, onlay, and bridge, as well as any other artificial partial or complete denture. The term orthodontic procedure refers in particular to any procedure involving the oral cavity and for designing, manufacturing, or installing an orthodontic element or its real or virtual model at a dental site in the oral cavity, or for designing and preparing a dental site to receive such an orthodontic element. These elements can be appliances, including but not limited to brackets and archwires, retainers, clear aligners, or functional appliances. Any treatment results described herein or any update to treatment results can be based on these orthodontic and / or dental procedures. Examples of orthodontic treatments include treatments to reposition teeth, treatments such as mandibular advancement to manipulate the mandible, palatal expansion to widen the upper and / or lower palate, and the like. For example, updates on treatment results can be generated by performing one or more procedures on one or more portions of a patient's dental arch or mouth in interaction with a user. The AR systems described herein can assist in planning these orthodontic and / or dental procedures.

[0076] A treatment plan for producing a specific treatment outcome can first be generated by generating an intraoral scan of the patient's mouth. A virtual 3D model 150 of the patient's upper and / or lower dental arch can be generated based on the intraoral scan. The dentist or technician can then determine the desired final position and orientation of the patient's teeth on the upper and lower dental arches, and the desired final position and orientation of the patient's bite. This information can be used to generate a virtual 3D model 150 of the patient's upper and / or lower dental arches after orthodontic treatment and / or restorative treatment. This data can be used to create an orthodontic treatment plan, a restorative treatment plan (e.g., a dental prosthesis treatment plan), and / or a combination thereof. An orthodontic treatment plan can include a series of orthodontic treatment stages. Each orthodontic treatment stage can adjust the patient's dentition by a specified amount and can be associated with a 3D model 150 of the patient's dental arch showing the patient's dentition for that treatment stage.

[0077] In some embodiments, the treatment planning module 120 may receive or generate one or more virtual 3D models 150, virtual 2D models, 3D images, 2D images (e.g., captured images 135 and / or simulated images 145), or other treatment outcome models and / or images.

[0078] In some embodiments, the smile processing module 180 includes a historical smile processing module 155 and / or a future smile processing module 160. The historical smile processing module 155 can operate on received images to align them, generate replacement images, perform interpolation to generate intermediate simulated images (also known as composite images), and / or generate a video (e.g., smile video 140) based on the received images, replacement images, and / or simulated images. In some embodiments, the smile processing module 180 generates one or more replacement images 136 for one or more captured images 135, as discussed in more detail below. The composite images and simulated images can be images generated by a machine learning model (e.g., not by an image sensor).

[0079] The future smile processing module 160 can use seed images of the patient's dentition and data from the patient's treatment plan to generate simulated images of the patient's dentition during and / or after future treatment phases, and generate a video (e.g., smile video 140) based on the seed images and the simulated images of the future treatment phases. In some cases, the patient can begin treatment, and during treatment, a video of the patient's smile can be generated using a combination of the historical smile processing module 155 and the future smile processing module 160. For example, the historical smile processing module 155 can use images taken before the start of treatment and during past treatment phases to generate a first frame of video, while the future smile processing module 160 can use the latest images of the current state of the patient's dentition and data from the treatment plan (e.g., a 3D model of the future treatment phases) to generate simulated images that show how the patient's dentition will appear during and between future treatment phases. The images generated by the historical smile processing module 155 and the images generated by the future smile processing module 160 can then be combined into a single video (e.g., smile video 140) that shows the patient's dentition's progress to date and the projected future progress of the patient's dentition toward the final treatment goal.

[0080] In an embodiment, the historical smile processing module 155 performs a series of operations to align captured images of the patient's dentition, optionally replace one or more of the captured images, and then interpolate additional simulated images to ultimately generate a smile video 140 that shows the progression of the patient's dental condition over time. Such operations can be categorized at a high level into color conversion operations, landmark detection operations, physical alignment operations, image replacement operations, and interpolation operations (also referred to as image generation operations). Color conversion operations can include modifying the color of one or more of the images to achieve color matching and / or alignment between the images. Landmark detection operations can include identifying common features in each image, such as the patient's front teeth (e.g., the first six teeth), the patient's nose, eyes, ears, etc. Alignment operations can include determining an affine transformation between images (e.g., between sequential images), and then applying the affine transformation by performing translation, rotation, and / or scaling adjustments on one or more of the images to align the images in space, achieve the same size, and so on. Image replacement operations can include determining that one or more images do not meet one or more quality criteria, generating a synthetic replacement image for the one or more images, and replacing the one or more images with the corresponding one or more replacement images. Then, one or more interpolation operations can be performed (optionally, recursively) to generate intermediate images showing the state of the patient's dentition between the captured images. Finally, the modified image and all of the generated images can be added as frames to a smile video 140 showing the progression of the patient's smile (e.g., dental arch, dentition, teeth, etc.) over time. Figures 3A to 3E and Figures 4A to 4E, a possible sequence of operations performed by the historical smile processing module 155 to generate the smile video 140 is shown in FIG.

[0081] In some embodiments, the sequence of operations performed by the historical smile processing module 155 includes one or more of the above operations and one or more additional operations, such as distortion correction operations, blur correction operations, etc. In one embodiment, in order to perform distortion correction, landmarks are detected from a generated 3D model of the patient's dental arch. The detected 3D landmarks can then be used to correct distortions caused by the camera that generated one or more of the historical images. For example, the detected landmarks can be used to remove image distortion and artificially increase the focal length. In some cases, the distortion correction can be performed using the intraoral scan used to generate the 3D model. In an embodiment, a rigid fitting step can be performed in which the intraoral scan is fitted to each image separately, thereby providing an estimate of the camera position and focal length for each image. Each image can then be warped so that the camera position and focal length of all images are consistent to ultimately remove the distortion in the image.

[0082] In one embodiment, one or more image selection operations are performed on the images. A patient may submit a blurry, low-resolution image captured using a front-facing camera (e.g., a front-facing camera of a mobile device). One or more detectors and / or heuristic algorithms may be used to select a subset of images, which are then used to generate the final video. The heuristic algorithm / detector may analyze the images and may include criteria or rules that the images should meet for video generation. Examples of criteria include: an image with an open bite, a patient not wearing braces in the image, a patient's face angle with the camera within a target range (e.g., a camera viewing angle within a target range), and so on. In one embodiment, one or more deblurring (also known as blur removal) and / or image upscaling operations may be performed as part of the sequence of operations. For images that do not meet the image selection criteria, rather than excluding these images, processing logic may perform one or more operations to improve these images so that they meet the one or more criteria after modification. This may include deblurring and / or upscaling operations, in which a trained machine learning model transforms a "bad photo" into a clear, sharp, high-resolution image. In one embodiment, rather than discarding images that do not meet one or more criteria or rules, processing logic selects those images for replacement and generates composite replacement images to replace those images, as described in greater detail below.

[0083] Various operations may be performed using and / or with the assistance of one or more trained machine learning models, such as color conversion operations, landmark detection operations, image generation operations, deblurring operations, magnification operations, image selection operations, distortion correction operations, image replacement operations, and the like. Figure 2 2 shows a model training workflow 205 and a model application workflow 217 of a smile processing module according to an embodiment of the present disclosure. In an embodiment, the model training workflow 205 can be executed at a server and the trained model can be provided to another computing device (e.g., Figure 1 The smile processing module on the computing device 105) can execute the model application workflow 217. The model training workflow 205 and the model application workflow 217 can be executed by processing logic executed by a processor of the computing device. For example, one or more of these workflows 205, 217 can be implemented by one or more machine learning models implemented in the smile processing module 108 or in Figure 10 The illustrated computing device 1000 may be implemented by other software and / or firmware executing on the processing device.

[0084] The model training workflow 205 is used to train one or more machine learning models (e.g., deep learning models) to perform one or more tasks such as classification, image generation, landmark detection, color conversion, segmentation, detection, and recognition on images of smiles, teeth, dentition, faces, etc. The model application workflow 217 is used to apply one or more trained machine learning models to perform tasks such as classification, image generation, landmark detection, color conversion, segmentation, detection, and recognition on images of smiles, teeth, dentition, faces, etc.

[0085] Many different machine learning outputs are described herein. A specific number of machine learning models and a specific arrangement of machine learning models are described and illustrated. However, it should be understood that the number and type of machine learning models used, as well as the arrangement of such machine learning models, can be modified to achieve the same or similar end results. Therefore, the arrangements of machine learning models described and illustrated are merely examples and should not be construed as limiting. Furthermore, the embodiments discussed with reference to machine learning models can also be implemented using a conventional rule-based engine.

[0086] In an embodiment, one or more machine learning models are trained to perform one or more of the following tasks. Each task can be performed by a separate machine learning model. Alternatively, a single machine learning model can perform each task or a subset of tasks. Additionally or alternatively, different machine learning models can be trained to perform different combinations of tasks. In an example, one or several machine learning models can be trained, where the trained ML model is a single shared neural network with multiple shared layers and multiple higher-level different output layers, where each output layer outputs a different prediction, classification, identification, etc. The one or more trained machine learning models can be trained to perform the following tasks:

[0087] I) Color Modification / Conversion - This can include modifying the color, white balance, brightness, etc. of one or more images so that the color, white balance, brightness, etc. are uniform or approximately uniform across images. In some embodiments, the color modification / conversion is performed using a trained machine learning model that performs a wavelet transform. Examples of such machine learning models include style transfer machine learning models, such as a whitening and color transform (WCT) model. In embodiments, an example of a WCT model that can be used is described in Photorealistic Style Transfer via Wavelet Transforms, by Jaejun Yoo et al., September 29, 2019, which is incorporated herein by reference in its entirety.

[0088] II) Dental Object Segmentation - This can include point-level classification (e.g., pixel-level classification or voxel-level classification) of different types of dental objects from the image. For example, different types of dental objects can include: teeth, gums, palate, prepared teeth, unprepared dental restorations, implants, brackets, dental attachments, soft tissue, retraction cords (dental archwires), blood, saliva, etc. In some embodiments, the dentition image is segmented into individual teeth and optionally into gums.

[0089] III) Landmark Detection - This can include identifying landmarks in the image. In embodiments, the landmarks can be features of a particular type, such as the center of a tooth. In some embodiments, landmark detection is performed after dental object segmentation. In some embodiments, dental object segmentation and landmark detection are performed together by a single machine learning model. In one embodiment, landmark detection is performed using one or more stacked hourglass networks. An example of a model that can be used to perform landmark detection is a convolutional neural network that includes multiple stacked hourglass models, as described in Stacked Hourglass Networks for Human Pose Estimation by Alejandro Newell et al., July 26, 2016, which is incorporated herein by reference in its entirety.

[0090] IV) Image Generation / Interpolation - This can include generating (e.g., interpolating) simulated images that show how teeth, gums, etc. would look between existing images. Such images can be realistic images. In some embodiments, a generative model (e.g., a generative adversarial network (GAN), an encoder / decoder model, a diffusion model, a variational autoencoder (VAE), a neural radiance field (NeRF), etc.) is used to generate the intermediate simulated images. In one embodiment, a generative model is used that determines features of two input images in a feature space, determines an optical flow between the features of the two images in the feature space, and then uses the optical flow and one or both of the images to generate the simulated image. In one embodiment, a trained machine learning model that determines frame interpolation for large motion is used, such as described in Fitsum Reda et al., FILM: Frame Interpolation for Large Motion, Proceedings of the European Conference on Computer Vision (ECC) (2022), which is incorporated herein by reference in its entirety.

[0091] V) Image Generation - This can include generating an estimated image (e.g., a 2D image) of the expected appearance of the patient's teeth at a future stage of treatment (e.g., an intermediate stage of treatment and / or after treatment completion). Such an image can be a realistic image. In an embodiment, a generative model (e.g., such as a GAN, an encoder / decoder model, etc.) operates on image features extracted from the current image and a 2D projection of a 3D model of the future state of the patient's dental arch to generate a simulated image.

[0092] VI) Image Replacement - This can include a form of image generation where a synthetic replacement image is generated to replace an image that does not meet one or more quality criteria (e.g., one or more alignment criteria). In one embodiment, a Style Generative Adversarial Network (StyleGAN) is used to generate one or more replacement images. The StyleGAN can receive an image to be replaced and a style input, which can be another image and / or other style information (e.g., indicating target lighting conditions, camera viewpoint, pose information, etc.).

[0093] VII) Optical flow determination - This can include using a trained machine learning model to predict or estimate the optical flow between images. Such a trained machine learning model can be used to perform any of the optical flow determinations described herein.

[0094] One type of machine learning model that can be used to perform some or all of the above tasks is an artificial neural network, such as a deep neural network. An artificial neural network typically includes a feature representation component with a classifier or regression layer that maps features to a desired output space. For example, a convolutional neural network (CNN) hosts multiple layers of convolutional filters. Pooling is performed and nonlinear problems are solved at the lower layers. A multi-layer perceptron is typically attached on top of the lower layers to map the top-level features extracted by the convolutional layers to decisions (e.g., classification outputs). Deep learning is a class of machine learning algorithms that uses cascaded multi-layer nonlinear processing units for feature extraction and transformation. Each successive layer uses the output of the previous layer as input. Deep neural networks can learn in a supervised (e.g., classification) and / or unsupervised (e.g., pattern analysis) manner. A deep neural network includes a hierarchy of layers, in which different layers learn different representation levels corresponding to different levels of abstraction. In deep learning, each level learns to transform its input data into a slightly more abstract and complex representation. For example, in an image recognition application, the raw input might be a matrix of pixels; the first representation layer might abstract the pixels and encode edges; the second layer might combine and encode the arrangement of edges; the third layer might encode higher-level shapes (e.g., teeth, lips, gums, etc.); and the fourth layer might identify the scanned person. Notably, the deep learning process can autonomously learn which features are best placed in which level. The "depth" in "deep learning" refers to the number of layers through which the data is transformed. More precisely, deep learning systems have a significant credit assignment path (CAP) depth. CAP is the chain of transformations from input to output. CAP describes the possible causal relationships between input and output. For feedforward neural networks, the CAP depth can be the depth of the network or the number of hidden layers plus one. For recurrent neural networks (where signals can propagate through layers more than once), the CAP depth can be infinite.

[0095] In one embodiment, a deep learning model that performs whitening and color conversion (WCT) is used, for example, in the color conversion module 267. The model can be trained to perform realistic style transfer between images to be merged to form a video. The model can restore the structural information of a given content image while accurately stylizing the image (e.g., based on a second input image). In one embodiment, the model performs a wavelet-corrected transform based on whitening and color conversion (WCT). WCT can use any style for style transfer by directly matching the association between content and style in the Visual Geometry Group (VGG) feature domain. The model can project content features into a feature space by computing a singular value decomposition (SVD). The final stylized image can be obtained by feeding the converted features into a decoder. In an embodiment, a multi-stage stylization framework is employed that applies WCT to multiple encoder-decoder pairs.

[0096] In one embodiment, a pose estimation model is used to perform landmark detection, essentially detecting the pose of the patient's face and / or teeth in an image, such as for landmark detection module 270. In one embodiment, the pose estimation model is a convolutional neural network comprising multiple hourglass neural network modules stacked end-to-end. This allows for repeated bottom-up and top-down reasoning across scales.

[0097] In one embodiment, a generative model is used for one or more machine learning models. The generative model can be a generative adversarial network (GAN), an encoder / decoder model, a diffusion model, a variational autoencoder (VAE), a neural radiance field (NeRF), or other types of generative models. For example, the generative model can be used in the image generation module 274.

[0098] GAN is a type of artificial intelligence system that uses two artificial neural networks that compete with each other in a zero-sum game framework. GAN includes a first artificial neural network that generates candidate pairs and a second artificial neural network that evaluates the generated candidate pairs. GAN learns to map from a latent space to a specific data distribution of interest (a data distribution that varies from a photo to an input image that is indistinguishable to the human eye), while the discrimination network distinguishes between instances from a training data set and candidates generated by the generator. The training goal of the generative network is to improve the error rate of the discrimination network (for example, to deceive the discrimination network by generating new synthetic instances that appear to be from a training data set). The generative network and the discrimination network are trained together, and the generative network learns to generate images that are increasingly difficult for the discrimination network to distinguish from real images (from the training data set), while the discrimination network simultaneously learns to be able to better distinguish synthetic images from images from the training data set. Once the two networks of the GAN reach equilibrium, they are trained. GAN may include a generator network that generates artificial intraoral images and a discriminator network that attempts to distinguish real images from artificial intraoral images. In an embodiment, the discriminator network may be MobileNet.

[0099] In one embodiment, a generative model is used that is trained to perform frame interpolation—synthesizing intermediate images between pairs of input frames or images. The generative model receives a pair of input images and generates an intermediate image that can be placed between them in a video, for example, for frame rate amplification. In one embodiment, the generative model has three main stages: a shared feature extraction stage, a scale-independent motion estimation stage, an estimation stage, and a fusion stage that outputs the resulting color image. In one embodiment, the motion estimation stage is capable of handling temporally irregular input data streams. Feature extraction may include determining groups of features for each input image in a feature space, while scale-independent motion estimation may include determining the optical flow between features of the two images in the feature space. The optical flow and data from one or both images are then used to generate the intermediate image in the fusion stage. The generative model is capable of robustly tracking features without artifacts due to large motion. In one embodiment, the generative model can handle occlusion. Furthermore, compared to traditional image interpolation techniques, the generative model can provide improved image clarity. In one embodiment, the generative model generates the simulated image recursively. The number of recursions may not be fixed but may be determined based on a metric calculated from the image.

[0100] In one embodiment, one or more machine learning models are conditional generative adversarial (cGAN) networks, such as pix2pix. These networks not only learn a mapping from an input image to an output image, but also learn a loss function used to train this mapping. GANs are generative models that learn a mapping from a random noise vector z to an output image y, G:z→y. In contrast, conditional GANs learn a mapping from an observed image x and a random noise vector z to y, G:{x,z}→y. The generator G is trained to produce outputs that are indistinguishable from “real” images by an adversarially trained discriminator D, which is trained to detect the generator’s “fakes” as well as possible. In embodiments, the generator may comprise a U-net or encoder-decoder architecture. In embodiments, the discriminator may comprise a MobileNet architecture. An example of a cGAN machine learning architecture that can be used is the pix2pix architecture described in “Image-to-image translation with conditional adversarial networks” by Isola, Phillip, et al., arXiv preprint (2017).

[0101] In one embodiment, one or more machine learning models used to generate replacement images are StyleGANs. StyleGANs are an extension of the GAN architecture to control the decoupled style characteristics of generated images. In at least one embodiment, the generative network is a generative adversarial network (GAN), which includes a generator model and a discriminator model, wherein the generator model includes using a mapping network to map points in a latent space to an intermediate latent space, including using the intermediate latent space to control the style at each point in the generator model, and using introduced noise as a source of variation at one or more points in the generator model. The resulting generator model is not only capable of generating impressively realistic, high-quality synthetic images, but also controls the style of generated images at different levels of detail by varying the style vector and noise. Each style vector can correspond to a parameter or feature of clinical information, or a parameter or feature of non-clinical information. For example, in an embodiment, there may be a style vector for camera perspective or facial pose, a style vector for lighting, a style vector for patient clothing, a style vector for accessories, a style vector for facial expression, and so on. In at least one embodiment, the generator starts with a learned constant input and adjusts the "style" of the image at each convolutional layer based on a latent code, directly controlling the strength of image features at different scales.

[0102] In at least one embodiment, the StyleGAN generator uses two sources of randomness to generate synthetic images: an independent mapping network and a noise layer, and a starting point from the latent space. The output from the mapping network is a vector defining the style, which is integrated at each point in the generator model via a layer called adaptive instance normalization. The style of the generated image is controlled using this style vector. In at least one embodiment, random variation is introduced by noise added at each point in the generator model. Noise can be added to entire feature maps that allow the model to interpret style in a fine-grained, pixel-by-pixel manner. This block-by-block combination of style vectors and noise allows each block to localize both the interpretation of style and random variation to a given level of detail.

[0103] Training a neural network can be accomplished in a supervised learning fashion, which involves feeding a training dataset consisting of labeled inputs through the network, observing its output, defining an error (by measuring the difference between the output and the labeled value), and using techniques such as deep gradient descent and backpropagation to adjust the network's weights across all its layers and nodes to minimize the error. In many applications, repeating this process across many labeled inputs in the training dataset yields a network that can produce correct outputs when presented with inputs different from those present in the training dataset. In high-dimensional settings, such as large images, this generalization can be achieved when a sufficiently large and diverse training dataset is available.

[0104] For the model training workflow 205, a training dataset comprising hundreds, thousands, tens of thousands, hundreds of thousands, or more images should be used to form the training dataset. In an embodiment, up to millions of images of patient dentitions that have undergone restorative and / or orthodontic procedures can be used to form the training dataset, wherein each case can include various labels of one or more types of useful information. For example, each case can include data showing a 3D model of one or more dental sites, an intraoral scan, a height map, a color image at various stages of treatment, a NIRI image, etc., data showing pixel-wise segmentation of the data (e.g., a 3D model, an intraoral scan, a height map, a color image, a NIRI image, etc.) into various dental categories (e.g., teeth, dental restoration objects, gums, active tissue, palate, etc.), data showing one or more specified classifications of the data (e.g., teeth, gums, palate, nose, eyes, etc.), data associated with different style vectors, and the like. This data can be processed to generate one or more training datasets 236 for training one or more machine learning models. For example, a machine learning model can be trained to modify image color, perform landmark detection, perform segmentation, perform image interpolation, generate replacement images, and the like. Once trained, this trained machine learning model can be added to a smile processing module, e.g. Figure 1 The smile processing module 108.

[0105] In one embodiment, generating one or more training datasets 236 includes collecting one or more images 210 with labels. The labels used may depend on the training objectives of a particular machine learning model. For example, to train a machine learning model to classify dental sites and ultimately perform landmark detection (e.g., landmark detection module 270), the training dataset 236 may include pixel-level labels for various types of dental sites, such as teeth, gums, etc.

[0106] The processing logic may collect a training data set 236 comprising images having one or more associated labels. In an embodiment, one or more images, scans, surfaces and / or models in the training data set 236 and optionally associated labels may be resized. For example, a machine learning model may be used for images having a specific pixel size range, and if one or more images fall outside of these pixel size ranges, the size of one or more images may be resized. For example, methods such as nearest neighbor interpolation or box sampling may be used to resize the images. The training data set may additionally or alternatively be augmented. Training of large-scale neural networks typically uses tens of thousands of images, which are not readily available in many real-world applications. Data augmentation may be used to artificially increase the effective sample size. Common techniques include random rotation, shifting, shearing, flipping, etc. of existing images to increase the sample size.

[0107] To implement training, processing logic inputs the training dataset(s) 236 into one or more untrained machine learning models. Before inputting the first input into the machine learning model, the machine learning model may be initialized. Processing logic trains the untrained machine learning model(s) based on the training dataset(s) to generate one or more trained machine learning models that perform various operations as described above.

[0108] Training can be performed by inputting one or more images into the machine learning model one at a time. Each input can include data from an image in a training dataset. The machine learning model processes the input to generate an output. An artificial neural network includes an input layer, which consists of the values in the data points (e.g., the intensity values and / or height values of pixels in a height map). The next layer is called a hidden layer, and each node in the hidden layer receives one or more input values. Each node includes parameters (e.g., weights) that are applied to the input values. Therefore, each node essentially inputs the input values into a multivariate function (e.g., a nonlinear mathematical transformation) to produce an output value. The next layer can be another hidden layer or an output layer. In either case, the nodes in the next layer receive the output values from the nodes in the previous layer, and each node applies weights to these values and then generates its own output value. This can be performed at each layer. The final layer is the output layer, which contains a node for each category, prediction, and / or output that the machine learning model can generate. For example, for an artificial neural network trained to perform dental site classification, there may be a first category (teeth), a second category (gums), and / or one or more additional dental categories. Furthermore, the class, prediction, etc. can be determined for each pixel in the image / scan / surface, the class, prediction, etc. can be determined for the entire image / scan / surface, or the class, prediction, etc. can be determined for each region or group of pixels in the image / scan / surface. For pixel-level segmentation, a final layer applies, for each pixel in the image, a probability that the pixel belongs to a first class, a probability that the pixel belongs to a second class, and / or one or more additional probabilities that the pixel belongs to other classes.

[0109] Thus, the output may include one or more predictions and / or one or more probability maps. For example, for each pixel in the input image / scan / surface, the output probability map may include a first probability that the pixel belongs to a first dental class, a second probability that the pixel belongs to a second dental class, etc. For example, the probability map may include the probability that the pixel belongs to a dental class representing a tooth, gum, or restoration.

[0110] The processing logic can then compare the generated probability map and / or other output with the known probability map or prediction and / or label included in the training data item. The processing logic determines an error (i.e., a classification error) based on the difference between the output probability map and / or (one or more) labels and the provided probability map and / or (one or more) labels. The processing logic adjusts the weights of one or more nodes in the machine learning model based on the error. An error term or increment can be determined for each node in the artificial neural network. Based on the error, the artificial neural network adjusts one or more of its parameters for one or more of the nodes of the artificial neural network (the weights of one or more inputs of the node). The parameters can be updated in a backpropagation manner so that the nodes at the highest layer are updated first, then the nodes at the next layer are updated, and so on. The artificial neural network includes multiple layers of "neurons", where each layer receives as input values from the neurons at the previous layer. The parameters of each neuron include weights associated with the values received from each neuron at the previous layer. Therefore, adjusting the parameters can include adjusting the weights of each input assigned to one or more neurons at one or more layers in the artificial neural network.

[0111] Once the model parameters are optimized, model validation can be performed to determine whether the model has improved and to determine the current accuracy of the deep learning model. After one or more rounds of training, processing logic can determine whether a stopping criterion has been met. The stopping criterion can be a target level of accuracy, a target number of images processed from the training dataset, a target change in a parameter over one or more previous data points, a combination thereof, and / or other criteria. In one embodiment, the stopping criterion is met when at least a minimum number of data points have been processed and at least a threshold accuracy has been achieved. For example, the threshold accuracy can be 70%, 80%, or 90% accuracy. In one embodiment, the stopping criterion is met if the accuracy of the machine learning model has stopped improving. If the stopping criterion is not met, further training is performed. If the stopping criterion has been met, training can be completed. Once the machine learning model is trained, a retained portion of the training dataset can be used to test the model.

[0112] Once the one or more trained ML models 238 are generated, the one or more trained ML models 238 may be stored in the model storage 245 and may be added to the smile processing module or other applications (e.g., Figure 1 Then, in an embodiment, the smile processing module 108 may use one or more trained ML models 238 and additional processing logic to generate a smile video.

[0113] In one embodiment, the model application workflow 217 includes one or more trained machine learning models and / or other logic / modules arranged in a pipeline for generating a video of a face, smile, and / or teeth. For the model application workflow 217, according to one embodiment, one or more captured images 135 are received. For example, the captured images 135 may include images of the current state of a patient's dentition, images of the patient's dentition at various states during a dental treatment, and so on.

[0114] The captured image 135 can be input into the image evaluation module 271. The image evaluation module 271 can process the image to determine whether the image meets one or more quality criteria. For example, the image evaluation module 271 can determine whether the input image shows a sufficient number of teeth (e.g., whether the patient is smiling in the input image), whether the input image is too blurry, whether the camera angle of view and / or facial pose of the input image is within a threshold range, etc. For images that meet the quality criteria, these images can be input into the color conversion module 267. For images that do not meet the quality criteria, at least some of these images can be input into the image replacement module 273. Alternatively, images that do not meet the quality criteria can be discarded. In some embodiments, the image evaluation module 271 and / or the image replacement module 273 are not used.

[0115] In some embodiments, the image evaluation module 271 may include a segmentation module that segments the image. In some embodiments, the segmentation information of the image may be used to determine whether the image meets one or more quality criteria.

[0116] The image replacement module 273 may receive one or more captured images 135 to be replaced and / or one or more additional captured images 135, and in embodiments, copy style information from these images. The image replacement module 273 may be a StyleGAN trained to generate a synthetic replacement image of an input image of a face, smile, teeth, etc.

[0117] The application program (e.g., the smile processing module 108) that applies the model application workflow 217 may include a user interface (e.g., such as a graphical user interface (GUI)). The user interface may provide options for selecting one or more characteristics of one or more images input to the image replacement module 273 to be included in the generated composite replacement image. For example, via the user interface, a user may select to retain facial hair, hair color, accessories worn, posture, facial expression, lighting conditions, age, weight, missing accessories, etc. from one or more input images. In some embodiments, the user interface receives the image and / or second segmentation information from the image evaluation module 271. The image and / or segmentation information may be output to a display. The user may then select (e.g., click) an area in one or more of the images using, for example, a mouse pointer or a touch screen. In embodiments, the selected area in the image may be retained in the generated composite image. In some embodiments, the user interface provides a menu (e.g., a drop-down menu) that provides different features of the clinical and / or non-clinical information from the first input image and / or features of the clinical and / or non-clinical information from the second input image. The menu may include a generic graphic of the corresponding features, may not include a graphic of the corresponding features, or may include a custom graphic of the corresponding features determined based on the image and / or segmentation information. The user may select from the menu the features to be retained from one or more of the images.

[0118] In addition, the user interface may provide an option for the user to select values for one or more characteristics (e.g., tooth color, lighting, pose, etc.) to be included in the generated composite image. The selected characteristic values may not correspond to characteristic values from any input image. For example, the user interface may include a slider associated with one or more characteristics. The user may move the slider position of such a slider to a desired position associated with a particular characteristic value for that characteristic.

[0119] Image replacement module 273 receives one or more input images and, optionally, segmentation information for these images and / or selected features to be retained from one or more of these images and / or values for one or more features to be applied to the generated image. Image replacement module 273 receives these inputs and processes them to generate a composite replacement image for one or more images that do not meet image quality criteria. The composite image may retain clinical information from the image to be replaced and may include additional stylistic information based on a second input image (e.g., one that meets quality criteria). In an example, most captured images 135 may be captured from the same camera perspective (e.g., this may mean that the imaged patient's face has the same pose in each of these images). However, one or more of captured images 135 may be captured from a different camera perspective (e.g., the imaged patient may have a different pose in one or more images). These images, which have a different camera perspective and / or pose than the majority, may not meet quality criteria and may be marked for replacement. Additionally, image evaluation module 271 may identify a "best" image from the captured images 135. This "best" image may have optimal lighting conditions, minimal blur, a desired camera perspective, and so on. In an embodiment, the selected "best" image may be used as a reference image. The reference image may be input into the image replacement module 273 along with the image to be replaced. The image replacement module 273 may use style information from the reference image when generating the replacement image.

[0120] Image replacement module 273 can generate a new synthetic replacement image that retains first selected or predetermined information (e.g., clinical and / or non-clinical information, first features, dental data, etc.) from the image being replaced (e.g., information about the teeth from the image being replaced) and retains information from the reference image (e.g., one or more features in the reference image that differ from features in the image being replaced). In some embodiments, image replacement module 273 generates a replacement image that includes the teeth of the replacement image, the patient's face of the replacement image, etc., but generates the replacement image from the camera perspective and / or facial pose provided in the reference image. In some embodiments, the replacement image may have attachments on one or more teeth, may have a tongue, finger, etc. obscuring one or more teeth, may have poor lighting conditions, etc. In the reference image, the teeth may not be obscured by the tongue, finger, etc. Additionally or alternatively, in the reference image, the teeth may not have attachments, may have optimal lighting conditions, etc. The generated replacement image may have the same lighting as presented in the reference image, lack attachments, be free of objects obscuring the teeth, etc., but may have the teeth, face, etc. of the image being replaced. Similarly, a patient may have a first expression in the image to be replaced, but a second expression in the reference image. The generated replacement image may include the patient's teeth from the image to be replaced, but may display the second expression from the reference image. Similarly, the generated replacement image may include color information from the reference image, which may differ from the color information in the image to be replaced. Similarly, the generated replacement image may include hairstyle, clothing, etc. from the reference image, which may differ from the hairstyle, clothing, etc. in the image to be replaced.

[0121] For each image to be replaced, the image replacement module 273 may receive the image to be replaced and a reference image and / or style information to be applied. The generated replacement image may be output to the color conversion module 267 for further processing in the image processing pipeline. Alternatively, in some embodiments, the replacement image may be input to the image generation module 274 without further processing.

[0122] Figure 3A A flow chart 302 of a method of generating a video of a dental treatment result according to an embodiment is shown, wherein an image replacement operation 303 is performed. In an embodiment, the image replacement operation may be performed by the image replacement module 273.

[0123] Figure 4AInput images 405 and 410 are shown as input to the image replacement module 273 for image replacement operation 303, along with an output image 458 generated by the image replacement module 273, according to an embodiment. In one embodiment, the image replacement module comprises a StyleGAN model or other generative module trained to receive two input images, where the first input image (e.g., image 405) is the image to be replaced, and the second input image (e.g., image 410) is an image used to determine the style information and / or latent space to be applied to the first input image. In one embodiment, the image replacement module 273 is trained to consistently apply the same style conditions to all input images. In such an embodiment, the image replacement module 273 may receive a single input image at a time and may output a replacement image for that input image. In some embodiments, the image replacement module 273 generates replacement images for images that do not meet one or more quality criteria. In some embodiments, the image replacement module 273 generates replacement images for all images. In such an embodiment, image quality analysis may not be performed on the images before generating the replacement images.

[0124] In some embodiments, the output of the image replacement operation 303 is provided directly to the image generation and / or interpolation operation 330, bypassing the color conversion operation 305, the landmark detection operation 310, and / or the image alignment operation 315. For example, the generated replacement image may be generated in a manner that already has color information that matches the color information of the other images, has an image alignment that matches the image alignment of the other images, etc. Therefore, for these images, one or more of the color conversion operation 305, the landmark detection operation 310, and / or the image alignment operation 315 may be skipped. Alternatively, in some embodiments, the replacement image is processed using the color conversion operation 305, the landmark detection operation 310, and / or the image alignment operation 315 before being processed by the image generation and / or interpolation operation 330.

[0125] return Figure 2The captured image 135 (e.g., an image that satisfies the image quality assessment) and / or the replacement image output by the image replacement module 273 can be input to the color conversion module 267, which can include a trained neural network. The trained neural network of the color conversion module 267 can be trained to adjust the color, lighting, white balance, etc. of the input image. For an example in which there are multiple captured images 135 of a patient's teeth over time (e.g., at different stages of treatment), the color conversion module 267 can make each of these images have uniform color, shading, white balance, etc. across these images. The machine learning model of the color conversion module 267 can output an updated or modified version of each of the input images, or output instructions on how to modify the input images, which can be applied by further logic of the color conversion module 267 to generate a color-balanced, modified image.

[0126] Figure 3B A flow chart 302 of a method of generating a video of a dental treatment result is shown, wherein a color conversion operation 305 is performed, in accordance with an embodiment. In an embodiment, the color conversion operation may be performed by the color conversion module 267.

[0127] Figure 4B 3. Input images 405 and 410 are shown as input to a color conversion module 267 (e.g., including a color conversion model) for color conversion operation 305, and an output image 430 generated by the color conversion module 267, according to an embodiment. In one embodiment, the color conversion model is trained to receive two input images, where a first input image (e.g., image 405) is the image to be modified and a second input image (e.g., image 410) is the image whose colorization and / or lighting conditions are applied to the first input image. In one embodiment, the color conversion model is trained to always apply the same colorization and / or lighting conditions (e.g., the same style) to all input images. In such an embodiment, the color conversion module 267 can receive a single input image at a time and can output a modified version of the input image in which the colors have been modified.

[0128] return Figure 2 After the image is modified to have consistent color and / or lighting (e.g., a consistent style), the modified image can be provided to the landmark detection module 270, which can include one or more trained machine learning models. Alternatively, the unmodified image can be provided to the landmark detection module before or after the color conversion module 267 operates on the image. In some embodiments, the landmark detection module 270 and the color conversion module 267 process one or more captured images in parallel.

[0129] The trained neural network of the landmark detection module 270 can be trained to perform segmentation of an input image. The trained neural network can segment an image of a face, smile, mouth, etc. into different dental objects, such as individual teeth and / or gums. The neural network can identify multiple teeth in the image and assign a different object identifier to each of the identified teeth. In some embodiments, the neural network estimates a tooth number for each of the identified teeth (e.g., according to the Universal Dental Numbering System, according to Palmer notation, according to FDI World Dental Federation notation, etc.). The trained neural network or other trained neural networks can perform landmark detection on the image. In some embodiments, the trained neural network (or other trained neural network) can use tooth segmentation information to perform landmark detection. In some embodiments, segmentation is omitted and landmark detection is performed without first performing segmentation of the input image. In one embodiment, landmark detection includes identifying features or groups of features (e.g., landmarks) in each input image. In one embodiment, the identified landmarks are one or more teeth, the center of one or more teeth, an eye, a nose, etc. In some embodiments, the identified landmarks can be features common to some or all of the input captured images 135. The landmark detection module 270 may output information regarding the location (eg, coordinates) of each of a plurality of different features or landmarks in the input image. In an embodiment, the set of landmarks may indicate the pose (eg, position, orientation, etc.) of the dental arch.

[0130] Figure 3C It shows the embodiment Figure 3B Flowchart 302 of a method for generating a video of a dental treatment result is shown in FIG, wherein landmark detection operation 310 is highlighted, indicating that landmark detection can be performed after color conversion. Alternatively, landmark detection can be performed before and / or during color conversion.

[0131] Figure 4C An input image 420 is shown as input to the landmark detection module 270 (e.g., a landmark detection model) and an output 430 of the landmark detection module 270, according to an embodiment. In one embodiment, the landmark detection module 270 outputs the location of one or more teeth in the input image 420. For example, the landmark detection module 270 may output the center location of each of the one or more teeth (e.g., the first six or eight teeth) in the input image 420. In one embodiment, the input image 420 is a color-balanced image (e.g., the modified image generated in the color conversion operation 305).

[0132] return Figure 2, after landmark detection has been performed on the image, the image and landmark information may be provided to a physical alignment module 272. In one embodiment, the physical alignment module 272 comprises one or more trained machine learning models that are trained to receive a plurality of input images and generate modified versions of the input images, wherein the output image is approximately physically aligned with each of the other input images. In one embodiment, the physical alignment module 272 does not include any machine learning models. In one embodiment, the physical alignment module operates on pairs of input images (e.g., pairs of temporally consecutive images). The physical alignment module 272 may compute an affine transformation between each pair of input images. This may include estimating a scale, rotation, and / or translation between a first set of 2D points of a first image in the input image pair and a second set of 2D points of a second image in the input image pair. In an embodiment, the affine transformation may be computed using a least squares technique to compute a transformation matrix. For two images having given point sets p1 and p2 (where each point set comprises i points), an affine transformation S, R, T may be selected such that ∑||S·R·p 1,i +Tp 2,i || 2 Minimize, where T represents the translation amount, S represents the scaling matrix, and R represents the rotation matrix. The calculated affine transformation can then be applied to one or more of the input images (e.g., by performing an affine warp) to approximately align the two images spatially.

[0133] Figure 3D A flow chart 302 is shown of a method of generating a video of a dental treatment result, with the image alignment operation 315 highlighted, according to an embodiment.

[0134] Figure 4D Example input images 430, 435 with identified landmarks are shown, according to an embodiment, which are input to a physical alignment module 272 (e.g., which can calculate an affine transformation matrix and perform affine warping on one or more of the input images using the affine transformation matrix), which outputs one or more modified images 445 that are aligned with each other.

[0135] return Figure 2Once the images have been image registered, they can have substantially uniform coloring and lighting, and can also have substantially uniform spatial registration, rotation, camera view, pose, size, etc. Furthermore, the generated replacement images output by the image replacement module 273 can be generated to have spatial registration, rotation, size, camera view, pose, etc. that align with the spatial registration, rotation, size, camera view, pose, etc. of the other images. This allows the images to be presented sequentially one after another (e.g., as video frames) without inter-frame jittering behavior, sudden changes in position, color, etc. However, depending on the frequency at which the captured images 135 are generated, there may be significant frame motion between subsequent images / frames, which may still result in some jitter and / or abrupt transitions between one or more images. Therefore, in an embodiment, the image generation module 274 receives the physically aligned images and processes them to generate additional simulated images, which are interpolated images that represent intermediate versions of the patient's dentition between the captured images 135. The image generation module 274 can operate on a pair of sequential temporal images (e.g., a first temporal image and a second temporal image) and generate one or more simulated images depicting an intermediate state between the first temporal image and the second temporal image. This can be performed for each pair of sequential images.

[0136] Figure 3E A flow chart 302 is shown of a method of generating a video of a dental treatment result, with the image generation operation 320 highlighted, according to an embodiment.

[0137] Figure 4E Input images 445, 435 to the image generation module 274, and one or more output images 455 of the image generation module 274 are shown according to an embodiment. In embodiments, a variety of different techniques may be applied to generate simulated images, which may or may not rely on a trained machine learning model.

[0138] In one embodiment, the image generation module 274 calculates the optical flow between the input images 445 and 435 and generates a simulated image 455 based on the optical flow. In one embodiment, an intermediate image is generated using a generative model. The generative model can receive two input images and can generate an output image that shows an intermediate state between the two input images. In one embodiment, the generative model is a generative model that includes one or more layers that determine the features of each input image in a feature space and one or more layers that calculate the optical flow of the features in the feature space. For each pair of points between the first feature and the second feature, the optical flow may include a vector indicating the direction and magnitude of movement of the feature between the images. The generative model can then use the optical flow in the feature space to generate a simulated image that is an interpolation between the two input images. Such a generated image may be more accurate than a simulated image generated using a simple generative model or simple optical flow.

[0139] return Figure 2 , the image generation performed by the image generation module 274 can be performed recursively. For example, a first simulated image can be interpolated between a first image and a second image, and then a second simulated image can be interpolated between the first image and the first interpolated image, and / or a third simulated image can be interpolated between the first simulated image and the second image. This recursion can be performed until one or more stopping criteria are met. For example, a simulated image can be generated until the similarity score of the newly generated image and the input image used to generate the new simulated image exceeds a similarity threshold. In another example, a motion score can be calculated using key points (e.g., landmarks) between the two images, and if the motion score is below a threshold, additional simulated images may not be generated. In addition, the motion score can include details from the treatment plan, such as the projected 2D displacement and other indicators defined in the treatment plan. In one embodiment, the processing logic determines the motion score between the two images based on the amount of motion between the key points in the two images, and determines the number of recursions to be performed based on the motion score.

[0140] Figures 5A to 5C Various stages of a recursive composite image generation process for creating a video (eg, a video of a dental treatment over time) in an embodiment are shown, according to an embodiment. Figure 5A The generation of simulated image 515 is shown, which is an interpolation of received image 505 and received image 510 . Figure 5B5C shows a first recursion of the image generation process, where simulated image 520 is an interpolation of received image 505 and simulated image 515, and simulated image 525 is an interpolation of simulated image 515 and received image 510. 5C shows a second recursion of the image generation process, where simulated image 530 is an interpolation of received image 505 and simulated image 520, simulated image 535 is an interpolation of simulated image 530 and simulated image 515, simulated image 540 is an interpolation of simulated image 515 and simulated image 525, and simulated image 545 is an interpolation of simulated image 525 and received image 510. The recursion may continue until a certain stopping criterion is met.

[0141] The amount of movement of keypoints between different image pairs may be different, and / or more time may have passed between the capture of some image pairs than between other image pairs. As a result, more recursions of the image generation process may have been performed between some images than between others.

[0142] Figures 5D to 5F 1 shows various stages of a recursive composite image generation process for creating a video (e.g., a video of a dental treatment over time) in accordance with an embodiment. Figure 5D As shown, three images are captured at different times, including received image 555, received image 560, and received image 565. As an example, the similarity between received image 560 and received image 565 may be much higher than the similarity between received image 555 and received image 560. Figure 5E As shown, the image generation process can be performed to generate simulated image 570 as an interpolation between received image 555 and received image 560, and to generate simulated image 575 as an interpolation between received image 560 and received image 565. After the image generation process, simulated image 575 may be sufficiently similar to received image 560 and / or received image 565 that no additional simulated images between received image 560 and received image 565 are generated. However, received image 555 and received image 560 may be sufficiently different from simulated image 570 that a certain similarity criterion is not met. Therefore, the image generation process can be recursively performed to generate additional simulated image 580 as an interpolation between received image 555 and simulated image 570, and to generate additional simulated image 585 as an interpolation between simulated image 570 and received image 560. After the recursion, there may be sufficient similarity between the sequential images (eg, between the received image 555 and the simulated image 580 , between the simulated image 580 and the simulated image 570 , etc.) to abandon additional recursion of the image generation operation.

[0143] return Figure 2 Once all simulated images have been generated, the simulated images and the physically aligned (and color-mixed) captured images can be input into the video generation module 276. The video generation module 276 can then generate a video (e.g., smile video 278) including each received image, where each image can be a frame of the video. The images can be arranged sequentially from the earliest image to the latest image, with the simulated images interposed between the captured images. A viewer (e.g., a doctor and / or patient) can then view the smile video 278 to understand the progression of the state of their dentition over time (e.g., to understand the progression of dental treatment through one or more treatment phases).

[0144] In some embodiments, the patient and / or doctor may wish to view a video of the dental treatment before it begins and / or during an intermediate stage of the dental treatment. In such cases, it may be useful to generate a smile video that projects the state of the patient's dentition into the future. If the dental treatment has not yet begun, a smile video 278 may be generated using a single captured image of the current state of the patient's dentition and a treatment plan. If the treatment has already begun, the portion of the smile video from the start of treatment to the present may be generated using captured images of the various stages of treatment to date, and the portion of the smile video that estimates how the patient's dentition will look at various future stages of treatment may be generated using the latest images of the patient's dentition and the treatment plan. If the treatment has not yet begun, the operations of the color conversion module 267, the landmark detection module 270, and the physical alignment module 272 may be omitted.

[0145] To generate a smile video containing simulated future images of an individual's dentition, the image generation module 280, the image generation module 274, and the video generation module 276 can perform one or more operations. The image generation module 280 can receive one or more captured images 135 of a current state or a most recent state of the individual's dentition. In addition, the image generation module 280 can receive a treatment plan 158 and / or components of a treatment plan, such as a 3D model of the patient's current dentition and / or a 3D model that predicts a future state of the patient's dentition at a future stage of treatment. This can include a 3D model of the patient's teeth in a final state and / or 3D models of the patient's teeth at various intermediate stages of dental treatment (e.g., orthodontic treatment and / or restorative dental treatment). The image generation module 280 can then generate a simulated image of the patient's dentition at each treatment stage with an associated 3D model.

[0146] In one embodiment, to generate a simulated post-treatment image (a simulated image at a treatment stage), the image generation module 280 generates a color map based on the captured image 135. This may include determining one or more fuzzy functions based on the captured image 135. This may include setting the functions and then solving the one or more fuzzy functions using data from the initial pre-treatment captured image 135. In some embodiments, a first set of fuzzy functions is generated (e.g., set and then solved) for a first region depicting teeth in the captured image 135, and a second set of fuzzy functions is generated for a second region depicting gingiva in the captured image 135. Once the fuzzy functions are generated, a color map may be generated using these fuzzy functions. An abstract representation (e.g., a color map) and image data (e.g., a sketch depicting the outlines of the teeth and gingiva at a post-treatment or intermediate treatment stage obtained from a 3D model of the dental arch at the treatment stage (e.g., a 3D mesh in a treatment plan), and / or a normal map depicting normals of the surface of the 3D model) may be input into a generative model, which then uses this information to generate a post-treatment image of the patient's face and / or teeth.

[0147] In an embodiment, the blur function for the teeth and / or gums is a global blur function that is a parametric function. Examples of parametric functions that can be used include polynomial functions (e.g., such as biquadratic functions, etc.), trigonometric functions, exponential functions, fractional powers, etc. In one embodiment, a set of parametric functions is generated that will serve as a global blur mechanism for the patient. The parametric function can be a unique function generated for a specific patient based on an image of that patient's smile. Using parametric blurring, a set of functions can be generated (one for each color channel of interest), where each function provides the intensity I of a given color channel c at a given pixel position x, y according to the following equation:

[0148] I c (x,y)=f(x,y) (1)

[0149] Various parameter functions can be used for f. In one embodiment, a parameter function is used, where the parameter function can be expressed as:

[0150]

[0151] In one embodiment, a biquadratic function is used. The biquadratic function can be expressed as:

[0152] I c (x,y)=w0+w1x+w2y+w3xy+w4x 2 +w5y 2 (3)

[0153] Where w0, w1, ..., w5 are the weights (parameters) of each term of the biquadratic function, x is a variable representing the position on the x-axis, and y is a variable representing the position on the y-axis (e.g., the x and y coordinates of the pixel position, respectively).

[0154] Parametric functions (e.g., biquadratic functions) can be solved using linear regression (e.g., multiple linear regression). Some example techniques that can be used to perform linear regression include ordinary least squares, generalized least squares, iteratively reweighted least squares, instrumental variable regression, optimal instrument regression, total least squares regression, maximum likelihood estimation, rigid regression, least absolute deviation regression, adaptive estimation, Bayesian linear regression, and the like.

[0155] To solve the parametric function, a mask M of points may be used to indicate the pixel locations in the initial image that should be used to solve the parametric function. For example, if the parametric function is used to blur teeth, the mask M may specify some or all pixel locations in the image that represent teeth. Alternatively, if the parametric function is used to blur gums, the mask M may specify some or all pixel locations in the image that represent gums.

[0156] In the example, for any initial image and mask M of points, the biquadratic weights w0, w1, ..., w5 can be found by solving the least squares problem:

[0157] Aw T =b (4)

[0158] in:

[0159] w=[w0,w1,w2,w3,w4,w5] (5)

[0160]

[0161] By constructing blur functions (e.g., parametric blur functions) separately for the tooth region and the gum region, a group of color channels can be constructed that avoids any dark and bright spot patterns that may be present in the initial image due to shadows (e.g., due to one or more tooth depressions).

[0162] In an embodiment, the blur function of the gums is a local blur function, such as a Gaussian blur function. In an embodiment, the Gaussian blur function has a larger radius (e.g., a radius of at least 5, 10, 20, 40, or 50 pixels). Gaussian blur can be applied to the entire mouth area of the initial image to produce color information. Gaussian blurring of an image involves convolving the image with a two-dimensional convolution kernel and producing grouped results. The Gaussian kernel is parameterized by a kernel width σ specified in pixels. If the kernel width is the same in the x and y dimensions, the Gaussian kernel is typically a matrix of size 6σ+1, where the center pixel is the focus of the convolution and all pixels can be indexed by their distance from the center in the x and y dimensions. The value of each point in the kernel is given as follows:

[0163]

[0164] When the kernel width is different in the x and y dimensions, the kernel value is specified as:

[0165]

[0166] In some embodiments, a neural network (e.g., a generative network, a generative adversarial network (GAN), a conditional GAN, or a picture-to-picture GAN) can be used to generate a simulated image of a smile with teeth in a final or intermediate treatment position. The neural network can integrate data from a 3D model of the upper and / or lower dental arches associated with the treatment stage with a blurred version of the captured image 135 (which is a color image). A blurred color image (e.g., a color map) of the patient's smile can be generated by applying one or more generated blur functions to the data from the 3D model. The data can be received as 3D data or 2D data (e.g., a 2D view of a 3D virtual model of the patient's dental arch). The neural network can use the input data to generate a simulated post-treatment image that matches the colors, tones, shades, etc. from the blurred color image with the shape and contours of the teeth and gums from the post-treatment image data (e.g., data from the 3D model).

[0167] After training is complete, the neural network receives input for generating realistic renderings of the patient's teeth in clinical final positions and / or intermediate positions. To provide color information to the generative model, a blurred color image (e.g., a color map) representing a group of color channels is provided, along with a post-treatment or mid-treatment sketch of the patient's teeth and / or gums. In some embodiments, a normal map comprising normals of the surface of the post-treatment 3D model can also be generated and provided to the trained generative model. The color channels are based on an initial photograph and contain information about the color and lighting of the teeth and gums in the initial image. In some embodiments, to avoid suboptimal results from the generative model, no structural information (e.g., tooth position, shape, etc.) is retained in the blurred color image.

[0168] As discussed above, the input may include a color map of the patient's teeth and gums, an image (e.g., a sketch or outline) of the patient's teeth and / or gums in a clinical target position (e.g., a 2D rendering of a 3D model of the patient's teeth in the clinical target position), and / or a normal map of the treated teeth in the clinical target position, etc. For example, the clinical target position may have been determined according to the treatment plan 158.

[0169] The neural network uses the input and the set of trained model parameters to render a realistic image of the patient's teeth in the target position. The realistic image can then be integrated into the mouth opening of the captured image 135, and an alpha channel blur can be applied. In an embodiment, the image generation module 280 performs the operations set forth in U.S. application No. 16 / 041,613, filed on July 20, 2018, to generate the simulated image. U.S. application No. 16 / 041,613, filed on July 20, 2018, is incorporated herein in its entirety. In an embodiment, the image generation module 280 performs the operations set forth in U.S. application No. 16 / 579,673, filed on September 23, 2019, to generate the simulated image. U.S. application No. 16 / 579,673, filed on September 23, 2019, is incorporated herein in its entirety.

[0170] Once the image generation module 280 generates one or more simulated images, the simulated images and the one or more captured images 135 may be input into the image generation module 274. The image generation module may then perform the operations described above to generate additional simulated images that are interpolated images between the captured images and / or the simulated images generated by the image generation module 280.

[0171] The video generation module 276 can use data from the image generation module 280 and / or the image generation module 274 to generate a video showing a future stage of dental treatment. The video generation module 276 can merge the video showing the future stage of dental treatment with another video showing a previous stage of dental treatment (e.g., a video generated based on operations performed by the color conversion module 267, the landmark detection module 270, the physical alignment module 272, and the image generation module 274). In some embodiments, a first visualization or other indicator is used to indicate which images or frames in the video represent a past state of the patient's dentition, and a second visualization or other indicator is used to indicate which images or frames in the video represent a predicted future state of the patient's dentition. For example, different boundaries can be used to distinguish between past images and future images.

[0172] The following Figures 6A to 9 Methods associated with generating a simulated video of a patient's smile according to embodiments of the present disclosure are described. Figures 6A to 9 The methods described in the can be performed by processing logic that can include hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions running on a processing device), or a combination thereof. Various embodiments can be referred to by reference to Figure 1 The computing device 105 and / or Figure 10 The computing device 1000 shown executes.

[0173] Figures 6A to 6B A flow chart of a method 600 for generating a video of a dental treatment over time after the treatment has begun is shown, according to an embodiment. At block 610 of method 600, processing logic receives multiple images of the face, teeth, and / or mouth of an individual (e.g., a patient or person). These images may be images of the patient's smile and may show the patient's teeth, gums, lips, etc. In some embodiments, these images include the patient's entire head and / or shoulders. In some embodiments, these images include only the patient's face. In some embodiments, these images include the patient's teeth and do not show the patient's face, or only show a small area of the patient's face.

[0174] At block 612, processing logic may modify one or more of the images to align the images. In one embodiment, at block 614, processing logic modifies the colors of the images so that the colors are consistent between the images (e.g., to align the colors of the images). At block 616, processing logic determines features common to some or all of the images (e.g., by performing segmentation and / or landmark detection). At block 618, processing logic determines an affine transformation (or alternatively, another type of transformation) between a pair of images (e.g., between sequential images). At block 620, processing logic applies the corresponding affine transformation (or other transformation) to one or more of the images to achieve a translation of the image, a rotation of the image, and / or a scale change of the image.

[0175] At block 622, processing logic may replace one or more images with one or more replacement images that can be aligned with other images. In some embodiments, at block 624, the received images are evaluated to determine whether the images meet one or more quality criteria (e.g., alignment criteria). Quality criteria may include blur criteria, camera angle of view and / or facial pose criteria, lighting criteria, facial expression criteria, exposed teeth criteria, etc. In embodiments, images that do not meet one or more quality criteria may be replaced. Alternatively, all images may be replaced. Alternatively, images that do not meet one or more image quality criteria may be discarded.

[0176] In one embodiment, at block 626, one or more images that meet one or more image quality (e.g., alignment) criteria are determined. In embodiments, the images so determined may be used as reference images. In some embodiments, one or more target features (e.g., stylistic characteristics) may be selected and provided to processing logic. In some embodiments, the processing logic automatically determines the one or more target features (e.g., stylistic characteristics) based on the evaluation results of the received images. At block 628, the processing logic may optionally segment the one or more images and / or one or more additional images using a segmenter (e.g., a machine learning model trained to perform image segmentation on facial images). In embodiments, the segmenter may segment the input image into teeth, gums, lips, facial features, etc.

[0177] At block 630, processing logic may generate a replacement image for each of the one or more images that do not meet one or more image quality criteria. In one embodiment, at block 632, the image to be replaced and the reference image are input into a trained generative model. In some embodiments, segmentation information for both images is also input into the trained generative model. The trained generative model may then use features of the image to be replaced (e.g., teeth, etc.) and stylistic information from the reference image to generate the replacement image. For example, the replacement image may have teeth from the image to be replaced, but in the pose provided by the reference image. The replacement image may include synthetically generated data for one or more teeth not visible in the image to be replaced. Furthermore, the replacement image may include facial expressions from the reference image, lighting from the reference image, no occluding objects in the reference image, no dental accessories in the reference image, etc., even though the image to be replaced may have a different facial expression, different lighting, may include occluding objects in front of one or more teeth, may include dental accessories, etc. In one embodiment, at block 634, the image to be replaced is input into the generative model along with one or more selected features (e.g., selected stylistic characteristics). In some embodiments, a replacement image may be generated using selected features (eg, stylistic characteristics).

[0178] At box 640, processing logic generates one or more simulated or synthetic images. In one embodiment, at box 645, processing logic determines the optical flow between each pair of sequential images, and then, at box 650, uses the optical flow to generate a synthetic image that shows intermediate states between the pairs of sequential images. In one embodiment, at box 655, processing logic inputs the image pairs into a trained machine learning model (e.g., a generative model), which outputs a synthetic image for each pair of input images. In one embodiment, the generative model includes a layer that generates groups of features in a feature space for each image in the image pair, then determines the optical flow between the groups of features in the feature space, and generates a synthetic image using the optical flow in the feature space. In one embodiment, processing logic transforms each pair of images into a feature space at box 660, determines the optical flow between each pair of images in the feature space at box 665, and generates a synthetic image for each pair of images based on the optical flow in the feature space at box 670.

[0179] In one embodiment, at block 675, processing logic determines a similarity score and / or a motion score for each pair of sequential images (which may include received images, replacement images, and / or simulated images). Then, at block 680, processing logic may determine whether the similarity score and / or the motion score meet a stopping criterion. If any image pair does not meet the stopping criterion, the method returns to block 640 to generate one or more additional simulated images. If all image pairs meet one or more stopping criteria, the method continues to block 685. At block 685, processing logic generates a video (e.g., a smile video) that includes the received image and the generated composite image.

[0180] Figure 7 A flow chart of a method 700 for generating a video of a dental treatment over time before the treatment begins, according to an embodiment, is shown. At block 710 of method 700, processing logic receives one or more images of the face and / or mouth of an individual (e.g., a patient or person). These images may be images of the patient's smile and may show the patient's teeth, gums, lips, etc.

[0181] At box 715, processing logic receives or generates a treatment plan that includes a 3D model of one or more future states of the patient's teeth at one or more treatment stages. In one embodiment, the treatment plan is a detailed and clinically accurate treatment plan generated based on a 3D model of the patient's dental arch (which is generated based on an intraoral scan of the dental arch). Such a treatment plan may include a 3D model of the dental arch at multiple treatment stages. In one embodiment, the treatment plan is a simplified treatment plan that includes a rough 3D model of the final target state of the patient's dental arch and is generated based on one or more 2D images of the patient's current dentition (e.g., an image of the patient's current smile). In box 720, processing logic generates one or more composite images based on the received images and the 3D model, the one or more composite images including the future states of the teeth at one or more treatment stages. This may include projecting the 3D model onto a 2D plane to generate a 2D projection or sketch of the 3D model, and then processing the 2D sketch and a blurred version of the received image into a generated model that outputs a simulated image.

[0182] At box 740, processing logic generates one or more additional simulated or synthetic images. In one embodiment, at box 745, processing logic determines the optical flow between each pair of sequential images, and then, at box 750, uses the optical flow to generate a synthetic image that shows intermediate states between the pairs of sequential images. In one embodiment, at box 755, processing logic inputs the image pairs into a trained machine learning model (e.g., a generator of a GAN), which outputs a synthetic image for each pair of input images. In one embodiment, the GAN includes a layer that generates groups of features in a feature space for each image in the image pair, then determines the optical flow between the groups of features in the feature space, and generates a synthetic image using the optical flow in the feature space. In one embodiment, processing logic transforms each pair of images into a feature space at box 760, determines the optical flow between each pair of images in the feature space at box 765, and generates a synthetic image for each pair of images based on the optical flow in the feature space at box 770.

[0183] In one embodiment, at block 775, processing logic determines a similarity score and / or a motion score for each pair of sequential images. Then, at block 780, processing logic may determine whether the similarity score and / or the motion score meet a stopping criterion. If the stopping criterion is not met for any image pair, the method returns to block 740 to generate one or more additional simulated images. If all image pairs meet one or more stopping criteria, the method continues to block 785. At block 785, processing logic generates a video (e.g., a smile video) that includes the received images and the generated composite image.

[0184] In an embodiment, the operations of method 600 may be combined with the operations of method 700 to generate a video showing the patient's smile at previous treatment sessions and predicted future treatment sessions.

[0185] Figure 8 A flow chart illustrating a method 800 for generating a simulated image of a dental treatment result, according to an embodiment, is shown. In one embodiment, method 800 is performed at block 720 of method 700. Processing logic may receive a first image of a patient's face and / or mouth. The image may be an image of the patient smiling with their mouth open to illustrate the patient's teeth and gums. In one embodiment, the first image may be a two-dimensional (2D) color image.

[0186] At block 815, processing logic determines a first region from the first image that includes representations of teeth. The first region may include a first set of pixel locations (e.g., x and y coordinates of the pixel locations) in the first image. In some embodiments, the first region may be determined using a first mask of the first image, wherein the first mask identifies a first set of pixel locations in the first region that are associated with teeth. Additionally, the first mask may identify pixel locations in a second region of the first image that are associated with gums.

[0187] In one embodiment, at box 820, processing logic generates a first mask for the first image. The first mask can be generated based on user input identifying the first area and / or the second area. For example, the user can trace the outline of the teeth and the outline of the gums in the first image, and the first mask is generated based on the traced outlines. In one embodiment, the first mask is automatically generated using one or more trained neural networks (e.g., such as deep neural networks, etc.). For example, the first neural network can process the first image to determine a bounding box around the teeth and gums. The image data within the bounding box can then be processed using a second trained neural network and / or one or more image processing algorithms to identify the gums and / or teeth within the bounding box. This data can then be used to automatically generate the first mask without user input.

[0188] At block 825, processing logic generates a first parametric function for the first color channel based on the intensity of the first color channel at the pixel location in the first set of pixel locations identified in the first mask. Processing logic may also generate a second parametric function for the second color channel, a third parametric function for the third color channel, and / or one or more additional parametric functions for the additional color channels (applicable to color spaces with more than three channels). Any color space may be used for the color channels associated with the parametric functions. For example, a red-blue-green color space may be used, where a first parametric function may be generated for the red channel, a second parametric function may be generated for the blue channel, and a third parametric function may be generated for the green channel. A non-exhaustive list of other example color spaces that may be used includes a hue, saturation, value (HSV) color space, a hue, saturation, lightness (HSL) color space, a YUV color space, a LAB color space, and a cyan, magenta, yellow, black (CMYK) color space.

[0189] The parameter function generated at block 825 is a global fuzzy function that can be used to generate a fuzzy representation of the tooth. Any type of polynomial function can be used for the global fuzzy function. Some examples of polynomial functions that can be used include first-order polynomial functions, second-order polynomial functions, third-order polynomial functions, fourth-order polynomial functions, and the like. Other types of parameter functions that can be used include trigonometric functions, exponential functions, fractional powers, and the like. The parameter function can be a smooth function that varies in the x-direction and / or in the y-direction. For example, the parameter function can vary only in the x-direction, only in the y-direction, or in both the x-direction and the y-direction. The parameter function is a global function that contains some local information. In one embodiment, the parameter function is a biquadratic function (e.g., such as that described in Equation 3 above). In one embodiment, one or more of the parameter functions is a biquadratic function that lacks cross terms (e.g., Equation 3 above does not have the w3xy term). In other embodiments, the parameter function can be, for example, a linear polynomial function, a bilinear polynomial function, and the like.

[0190] Each parameter function may initially be set with unsolved weights (e.g., unsolved values for w0, w1, w2, w3, w4, and w5 of Equation 3 above). Processing logic may then perform linear regression to solve for the weight values (also referred to as parameters) using the intensity values at the pixel locations indicated by the mask. In one embodiment, a least squares method is applied to solve for the weights (e.g., as set forth in Equations 4-7 above).

[0191] A similar process as set forth above can also be used to generate a set of blur functions for the gums. Alternatively, a Gaussian blur function can be used for the gums (eg, as set forth in equations 8-9 above).

[0192] At block 830, processing logic receives image data and / or generates image data including a new mouth outline based on the treatment plan. The image data may be a 2D sketch of the dentition at the mid- or post-treatment stage, a projection of a 3D virtual model of the dental arch in a 2D plane, or other image data. In some embodiments, the 3D virtual model may be oriented such that mapping the 3D virtual model into a 2D plane produces a simulated 2D sketch of the teeth and gums from the same perspective as that from which the first image was taken. The 3D virtual model may be included in the treatment plan and may represent the final or intermediate shape of the patient's upper and / or lower dental arches after treatment is completed. Alternatively or additionally, the treatment plan may include one or more 2D sketches of the dentition at the mid- or post-treatment stage, regardless of whether the 3D virtual model of the dental arch is included. Alternatively or additionally, the one or more 2D sketches may be generated based on a 3D template. The image data may be a line drawing that includes the outline of the teeth and gums but lacks color data for one or more regions (e.g., regions associated with the teeth). In one embodiment, generating the image data includes projecting the 3D virtual model of the upper and / or lower dental arches into a 2D plane.

[0193] In one embodiment, generating image data includes inferring a possible 3D structure from the first image, matching the 3D structure to a template of a dental arch (e.g., a template having an ideal tooth arrangement), and then projecting the template into 2D. The 3D template can be selected from a set of available 3D templates, and the 3D template can be a template having a dental arch that most closely matches the dental arch in the first image. In some embodiments, the 3D template can be oriented so that mapping the 3D template into a 2D plane produces a 2D sketch of the teeth and gums from the same perspective as that from which the first image was taken.

[0194] At block 835, processing logic determines a second region of the image data that includes teeth. The second region that includes teeth may include a second set of pixel positions of the teeth that differ from the first set of pixel positions. For example, a treatment plan may require repositioning one or more of the patient's teeth. The first image may show the teeth in their initial position and / or orientation (e.g., possibly including a malocclusion), while the image data may show the teeth in their final position and / or orientation (e.g., a previous malocclusion may have been treated).

[0195] In one embodiment, at block 840, processing logic generates a second mask for the image data. Processing logic may also generate another gum mask for the image data. The second mask may identify a second set of pixel positions associated with the new position and / or orientation of the teeth. Another gum mask may indicate the pixel positions of the upper and / or lower gums after treatment. The second mask (and optionally, other masks) may be generated in the same manner as discussed above with respect to the first mask. In some embodiments, the 3D virtual model or 3D template includes information identifying the teeth and gums. In such embodiments, the second mask and / or other masks may be generated based on the information identifying the teeth and / or gums in the virtual 3D model or 3D template.

[0196] At block 845, processing logic generates a blurred color representation of the teeth by applying a first parametric function to a second set of pixel locations of the teeth identified in the second mask. This may include applying multiple different parametric functions to the pixel locations in the image data specified in the second mask. For example, a first parametric function for a first color channel may be applied to determine the intensity or value of the first color channel for each pixel location associated with the teeth, a second parametric function for a second color channel may be applied to determine the intensity or value of the second color channel for each pixel location associated with the teeth, and a third parametric function for a third color channel may be applied to determine the intensity or value of the third color channel for each pixel location associated with the teeth. The blurred color representation of the teeth may then include three different color values for each pixel location associated with the teeth in the image data, one for each color channel. A similar process may also be performed for the gums by applying one or more blur functions to pixel locations associated with the gums. Thus, a single blurred color image may be generated that includes blurred color representations of the teeth and blurred color representations of the gums, where different blur functions were used to generate the blurred color data for the teeth and gums.

[0197] At block 850, a new image is generated based on the image data (e.g., a sketch including outlines of teeth and gums) and the blurred color image (e.g., which may include blurred color representations of the teeth and, optionally, the gums). The shape of the teeth in the new simulated image may be based on the image data, while the color of the teeth (and, optionally, the gums) may be based on the blurred color image including blurred color representations of the teeth and / or gums. In one embodiment, the new image is generated by inputting the image data and the blurred color image into an artificial neural network that has been trained to generate an image based on the input line drawing (sketch) and the input blurred color image. In one embodiment, the artificial neural network is a GAN. In one embodiment, the GAN is a picture-to-picture GAN.

[0198] Figure 9Also shown is a flow chart of a method 900 for generating a simulated image of a dental treatment result, according to an embodiment. In one embodiment, method 900 is performed at block 720 of method 700. At block 910 of method 900, processing logic receives a first image of a patient's face and / or mouth. The image may be an image of the patient smiling with their mouth open to showcase the patient's teeth and gums. In one embodiment, the first image may be a two-dimensional (2D) color image.

[0199] At block 915, processing logic may determine a first region from the first image that includes a representation of teeth. The first region may include a first set of pixel locations (e.g., x and y coordinates of the pixel locations) in the first image. At block 920, processing logic may determine a second region from the first image that includes a representation of gums. The second region may include a second set of pixel locations in the first image.

[0200] In some embodiments, a first region may be determined using a first mask from the first image, wherein the first mask identifies a first set of pixel locations in the first region associated with teeth. A second region may be determined using a second mask from the first image, wherein the second mask identifies a second set of pixel locations in the second region associated with gums. In one embodiment, a single mask identifies the first region associated with teeth and the second region associated with gums. In one embodiment, processing logic generates the first mask and / or the second mask as described with reference to block 820 of method 800.

[0201] At block 925, processing logic generates a first parametric function for the first color channel based on the intensities (or other values) of the color channels at the pixel locations in the first set of pixel locations identified in the first mask. At block 930, processing logic may also generate a second parametric function for the second color channel, a third parametric function for the third color channel, and / or one or more additional parametric functions for the additional color channels (applicable to color spaces with more than three channels). Any color space may be used for the color channels associated with the parametric functions. For example, a red-blue-green color space may be used, where a first parametric function may be generated for the red channel, a second parametric function may be generated for the blue channel, and a third parametric function may be generated for the green channel. A non-exhaustive list of other example color spaces that may be used includes a hue, saturation, value (HSV) color space, a hue, saturation, lightness (HSL) color space, a YUV color space, a LAB color space, and a cyan, magenta, yellow, black (CMYK) color space.

[0202] The parametric functions generated at blocks 925 and 930 are global fuzzy functions that can be used to generate fuzzy representations of teeth.Any of the types of parametric functions described above can be used as global fuzzy functions.

[0203] At block 935, a blur function for the gums can be generated based on the first image. In one embodiment, a set of parametric functions is generated for the gums in the same manner as described above for the teeth. For example, a mask identifying pixel locations associated with the gums can be used to identify pixel locations for solving for weights of one or more parametric functions. Alternatively, a Gaussian blur function (e.g., as described above in Equations 8-9) can be applied to the gums using the pixel locations associated with the gums.

[0204] At box 940, processing logic receives image data and / or generates image data including a new mouth outline based on the treatment plan. The image data can be a projection of a 3D virtual model of the dental arch in a 2D plane. In some embodiments, the 3D virtual model can be oriented so that mapping the 3D virtual model into the 2D plane produces a simulated 2D sketch of the teeth and gums from the same perspective as the first image was taken. The 3D virtual model can be included in the treatment plan and can represent the final shape or intermediate shape of the patient's upper and / or lower dental arch after treatment is completed. The image data can be a line drawing that includes the outline of the teeth and gums but lacks color data. In one embodiment, generating the image data includes projecting the 3D virtual model of the upper and / or lower dental arch into a 2D plane.

[0205] At block 945, processing logic determines a third region of the image data that includes teeth. The third region that includes teeth may include a second set of pixel locations of the teeth that is different from the first set of pixel locations. In one embodiment, processing logic generates a second mask for the image data, and the second mask is used to determine the third region.

[0206] At block 950, processing logic generates a fuzzy color representation of the tooth by applying the first parameter function and, optionally, the second parameter function, the third parameter function, and / or the fourth parameter function to a second set of pixel positions of the tooth associated with the third region. The fuzzy color representation of the tooth can then include three (or four) different color values for each pixel position associated with the tooth in the image data, one for each color channel. The third region of the image data can have more or fewer pixels than the first region of the first image. The parameter function works equally well regardless of whether the third region has fewer, the same, or a greater number of pixels.

[0207] At block 955, processing logic may generate a blurred color representation of the gums by applying one or more blur functions for the gums to pixel locations associated with the gums in the image data. The blurred color representations of the teeth may be combined with the blurred color representation of the gums to generate a single blurred color image.

[0208] At block 960, a new image is generated based on the image data (e.g., the sketch including outlines of the teeth and gums) and the blurred color image (e.g., which may include blurred color representations of the teeth and, optionally, the gums). The shape of the teeth in the new simulated image may be based on the image data, while the color of the teeth (and, optionally, the gums) may be based on the blurred color image including blurred color representations of the teeth and / or gums. In one embodiment, the new image is generated by inputting the image data and the blurred color image into an artificial neural network that has been trained to generate an image based on the input line drawing (sketch) and the input blurred color image. In one embodiment, the artificial neural network is a GAN. In one embodiment, the GAN is a picture-to-picture GAN.

[0209] In some cases, the lower gums may not be visible in the first image, but may be visible in the new simulated image generated at block 960. In such cases, the parametric function generated to use the color data for the blurred gums may result in inaccurate coloring of the lower gums. In such cases, one or more Gaussian blur functions may be generated for the gums at block 935.

[0210] Figure 10 A schematic diagram of a machine in the example form of a computing device 1000 is shown, in which a set of instructions can be executed to cause the machine to perform any one or more of the methodologies discussed herein. In alternative embodiments, the machine can be connected (e.g., networked) to other machines in a local area network (LAN), an intranet, an extranet, or the Internet. The machine can operate in the capacity of a server or a client machine in a client-server network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine can be a personal computer (PC), a tablet computer, a set-top box (STB), a personal digital assistant (PDA), a cellular phone, a web appliance, a server, a network router, a switch or a bridge, or any machine capable of executing a set of instructions (sequentially or otherwise) specifying the actions to be taken by the machine. Furthermore, while a single machine is shown, the term "machine" should also be taken to include any collection of machines (e.g., computers) that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein. In one embodiment, the computing device 1000 corresponds to Figure 1 The computing device 105 in.

[0211] The example computing device 1000 includes a processing device 1002, a main memory 1004 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM), etc.), a static memory 1006 (e.g., flash memory, static random access memory (SRAM), etc.), and a secondary memory (e.g., data storage device 1028), which communicate with each other via a bus 1008.

[0212] Processing device 1002 represents one or more general-purpose processors such as microprocessors, central processing units, and the like. More specifically, processing device 1002 may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor implementing other instruction sets, or a processor implementing a combination of instruction sets. Processing device 1002 may also be one or more special-purpose processing devices, such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), a network processor, and the like. Processing device 1002 is configured to execute processing logic (instructions 1026) for performing the operations and steps discussed herein.

[0213] The computing device 1000 may also include a network interface device 1022 for communicating with a network 1064. The computing device 1000 may also include a video display unit 1010 (e.g., a liquid crystal display (LCD) or a cathode ray tube (CRT)), an alphanumeric input device 1012 (e.g., a keyboard), a cursor control device 1014 (e.g., a mouse), and a signal generating device 1020 (e.g., a speaker).

[0214] The data storage device 1028 may include a machine-readable storage medium (or, more specifically, a non-transitory computer-readable storage medium) 1024 on which is stored one or more sets of instructions 1026 embodying any one or more of the methods or functions described herein, such as instructions for the smile processing module 108. A non-transitory storage medium refers to a storage medium other than a carrier wave. During execution of the instructions 1026 by the computing device 1000, the instructions 1026 may also reside, completely or at least partially, within the main memory 1004 and / or within the processing device 1002, with the main memory 1004 and the processing device 1002 also constituting computer-readable storage media.

[0215] The computer-readable storage medium 1024 may also be used to store the smile processing module 108. The computer-readable storage medium 1024 may also store a software library containing methods for the smile processing module 108. Although the computer-readable storage medium 1024 is shown as a single medium in the example embodiment, the term "computer-readable storage medium" should be taken to include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more sets of instructions. The term "computer-readable storage medium" should also be taken to include any medium, other than a carrier wave, that is capable of storing or encoding a set of instructions for execution by a machine and causing the machine to perform any one or more of the methods of the present disclosure. Thus, the term "computer-readable storage medium" should be taken to include, but not be limited to, solid-state memories, and optical and magnetic media.

[0216] It should be understood that the above description is intended to be illustrative and not restrictive. After reading and understanding the above description, many other embodiments will be apparent. Although the embodiments of the present disclosure have been described with reference to specific example embodiments, it should be appreciated that the present disclosure is not limited to the described embodiments, but can be practiced through modification and alteration within the spirit and scope of the appended claims. Therefore, the description and drawings should be regarded as illustrative and not restrictive. Therefore, the scope of the present disclosure should be determined with reference to the appended claims and the full scope of equivalents to such claims.

Claims

1. A method comprising: receiving a plurality of images comprising a tooth of an individual, wherein the plurality of images are arranged in a sequence and each of the plurality of images is associated with a different stage of treatment of the tooth; performing at least one of modifying one or more images in the plurality of images or replacing one or more images in the plurality of images so that the plurality of images are aligned with one another; generating one or more composite images, wherein each of the one or more composite images is generated based on a pair of sequential images in the sequence and is an intermediate image comprising an intermediate state of the tooth between a first state of a first image of the sequential image pair and a second state of a second image of the sequential image pair; and A video is generated that includes the plurality of images and the one or more composite images.

2. The method according to claim 1, wherein Modifying one or more images of the plurality of images includes modifying a color of the plurality of images so that the color remains consistent across the plurality of images.

3. The method according to claim 2, wherein: Modifying the color of the plurality of images includes inputting the plurality of images into a trained machine learning model, wherein the trained machine learning model outputs a color modification for one or more of the plurality of images.

4. The method according to claim 3, wherein: The trained machine learning model includes a convolutional neural network that performs one or more wavelet transforms.

5. The method according to claim 1, wherein Modifying an image of the one or more images includes performing at least one of a translation, rotation, or scaling change on one or more points of the image.

6. The method according to claim 5, further comprising: detecting a plurality of features common to at least some of the plurality of images; as well as For a sequence image pair in the plurality of images in the sequence and one or more features of the plurality of features, determine an affine transformation of the feature between a first image and a second image of the sequence image pair, wherein applying the affine transformation to at least one of the first image or the second image results in at least one of a translation, rotation, or scaling change of one or more points of the images.

7. The method according to claim 6, wherein: Detecting the plurality of features of the image comprises inputting the image into a trained machine learning model, wherein the trained machine learning model outputs a location of each of the plurality of features in the image.

8. The method according to claim 6, wherein: The plurality of features includes one or more of the teeth.

9. The method according to claim 6, wherein: The plurality of images are of the individual's face, wherein teeth of the individual are visible in a plurality of the facial images, and wherein the plurality of features include one or more facial features.

10. The method according to claim 1, wherein Replacing one or more images of the plurality of images includes: For an image in the one or more images, generating a replacement image having: a) teeth corresponding to the teeth in the image, and b) one or more features that are different from the one or more features in the image and similar to the one or more features in an additional image in the plurality of images, wherein the replacement image is used to replace the image.

11. The method according to claim 10, wherein: Generating the replacement image includes: The image and the additional image are processed using a trained machine learning model, wherein the trained machine learning model outputs the replacement image.

12. The method according to claim 11, wherein The trained machine learning model is a generative model.

13. The method according to claim 10, wherein: The one or more characteristics of the image include a first camera perspective, and wherein the one or more characteristics of the additional image include a second camera perspective.

14. The method according to claim 10, wherein: The one or more features of the image include at least one of: a first facial expression, a first jaw position, a first upper and lower jaw relationship, a first color, a first lighting condition, tooth occlusion, a tooth attachment, a first hairstyle, or a first clothing item; and The one or more features of the additional image include at least one of the following: a second facial expression, a second jaw position, a second upper and lower jaw relationship, a second color, a second lighting condition, no tooth occlusion, no dental attachments, a second hairstyle, or a second clothing item.

15. The method according to claim 1, wherein Replacing one or more images of the plurality of images includes: For an image in the one or more images, a replacement image is generated, the replacement image having: a) teeth corresponding to the teeth in the image; and b) one or more features different from the one or more features in the image, the replacement image being used to replace the image.

16. The method according to claim 15, wherein Generating the replacement image includes: receiving input selecting one or more target features; and The image and the input are processed using a trained machine learning model, wherein the trained machine learning model outputs the replacement image having the one or more features corresponding to the one or more target features.

17. The method according to claim 1, wherein Generating a composite image of the one or more composite images includes: determining, for a sequence image pair in the plurality of images in the sequence, an optical flow between a first image and a second image in the sequence image pair; and The composite image is generated based on the optical flow.

18. The method according to claim 1, wherein Generating a composite image of the one or more composite images includes: Inputting sequence image pairs from the plurality of images in the sequence into a trained machine learning model, wherein the trained machine learning model outputs the composite image.

19. The method according to claim 18, wherein The trained machine learning model comprises a generative model.

20. The method according to claim 19, wherein One or more layers of the generative model determine optical flow between pairs of the sequence images, and wherein the generative model uses the optical flow to generate the composite image.

21. The method according to claim 1, wherein Generating a composite image of the one or more composite images includes: transforming a first image and a second image in the sequence into a feature space; determining an optical flow between the first image and the second image in the feature space; and The composite image is generated using the optical flow in the feature space, the composite image being an intermediate image between the first image and the second image.

22. The method according to claim 1, wherein Generating the one or more composite images comprises: generating a first composite image based on a first image and a second image in the sequence, the first composite image being an intermediate image between the first image and the second image; and Based on the first image and the first synthesized image, a second synthesized image is generated, the second synthesized image being an intermediate image between the first image and the first synthesized image.

23. The method according to claim 22, further comprising: determining a similarity score between the first image and the first composite image; as well as In response to determining that the similarity score does not satisfy a similarity threshold, the second composite image is generated.

24. A non-transitory computer-readable medium comprising instructions which, when executed by a processing device, cause the processing device to perform the method according to any one of claims 1 to 23.

25. A system comprising: processing device; as well as A memory for storing instructions which, when executed by the processing means, cause the processing means to perform the method according to any one of claims 1 to 23.

26. A method comprising: receiving an image including a current state of an individual's teeth; receiving or generating a treatment plan comprising a three-dimensional (3D) model of a future state of the tooth during a treatment phase; generating a first composite image including the future state of the teeth at the treatment stage based on the received image and the 3D model of the future state of the teeth at the treatment stage; generating one or more additional composite images, the one or more additional composite images being intermediate images between the received image and the first composite image; as well as A video is generated that includes the received image, the one or more additional composite images, and the first composite image.

27. The method according to claim 26, wherein The received image has a facial image of the individual, wherein teeth of the individual are visible in the received facial image, and wherein the first composite image and the one or more additional composite images have facial images.

28. The method according to claim 26, wherein The treatment plan further includes a second 3D model of a second future state of the tooth at a second treatment stage, the method further comprising: generating a second composite image including the second future state of the teeth at the second treatment stage based on the received image and the second 3D model of the second future state of the teeth at the second treatment stage; and generating one or more additional composite images, the one or more additional composite images being intermediate images between the first composite image and the second composite image; The video further includes the one or more additional composite images and the second composite image.

29. The method according to claim 26, wherein Generating the one or more additional composite images comprises: determining an optical flow between the received image and the first composite image; and The one or more additional composite images are generated based on the optical flow.

30. The method of claim 26, wherein: Generating the one or more additional composite images comprises: The received image and the first composite image are input into a trained machine learning model, wherein the trained machine learning model outputs the one or more additional composite images.

31. The method according to claim 30, wherein The trained machine learning model comprises a generative model.

32. The method according to claim 31, wherein One or more layers of the generative model determine an optical flow between the received image and the first composite image, and wherein the generative model uses the optical flow to generate the one or more additional composite images.

33. The method of claim 26, wherein: Generating the one or more additional composite images comprises: transforming the received image and the first composite image into a feature space; determining an optical flow in the feature space between the received image and the first composite image; and The one or more additional composite images are generated using the optical flow in the feature space, the one or more additional composite images being intermediate images between the received image and the first composite image.

34. The method of claim 26, wherein: Generating the one or more additional composite images comprises: generating a second composite image, the second composite image being an intermediate image between the received image and the first composite image; and A third composite image is generated based on the received image and the second composite image, the third composite image being an intermediate image between the received image and the second composite image.

35. The method of claim 34, further comprising: determining a similarity score between the received image and the second composite image; as well as In response to determining that the similarity score does not satisfy a similarity threshold, the third composite image is generated.

36. A non-transitory computer-readable medium comprising instructions that, when executed by a processing device, cause the processing device to perform the method according to any one of claims 26 to 35.

37. A system comprising: processing device; as well as A memory for storing instructions which, when executed by the processing means, cause the processing means to perform a method according to any one of claims 26 to 35.

Citation Information

Patent Citations

  • Parametric blurring of colors for teeth in generated images

    US10835349B2

  • Generic framework for blurring of colors for teeth in generated images using height map

    US20200105028A1