Intraoral 3D scanner employing multiple small cameras and multiple small pattern projectors

The use of multiple cameras and projectors with diffraction-based optical elements and advanced calibration techniques addresses the challenges of intraoral scanning, enabling high-contrast, high-resolution 3D imaging without opaque powders, enhancing scanning efficiency and user convenience.

JP7848379B2Active Publication Date: 2026-04-20ALIGN TECHNOLOGY INC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
ALIGN TECHNOLOGY INC
Filing Date
2025-04-02
Publication Date
2026-04-20

AI Technical Summary

Technical Problem

Conventional intraoral scanners using structured light three-dimensional imaging face challenges with high reflectivity and translucency of teeth, leading to reduced contrast of the structured light pattern, and the correspondence problem between projected and captured light patterns, necessitating the use of opaque powders for enhancement.

Method used

Employing multiple small cameras and projectors with diffraction and/or refraction pattern-generating optical elements, such as laser diodes, to project unencoded discrete light spots, and using calibration algorithms to triangulate and reconstruct 3D images without the need for opaque powders, combined with inertial measurement for scanner motion estimation and neural networks for improved accuracy and efficiency.

Benefits of technology

Enhances image capture and accuracy in intraoral scanning by maintaining pattern contrast and resolution, reducing the need for additional coatings, and providing efficient, cost-effective, and user-friendly scanning solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007848379000002
    Figure 0007848379000002
  • Figure 0007848379000003
    Figure 0007848379000003
  • Figure 0007848379000004
    Figure 0007848379000004
Patent Text Reader

Abstract

To provide a method for generating a digital three-dimensional image and making a three-dimensional image of the oral cavity using structured light illumination.SOLUTION: The method for generating a three-dimensional image includes a step 62 for driving a structured light projector(s) to project a light pattern onto a three-dimensional surface in the oral cavity, a step 64 for driving a camera(s) to capture an image including at least one of spots, and a step for using a processor to compare a series of images captured by each camera and determine which parts of the projection pattern can be tracked across the images. A 3D model of the 3D surface in the oral cavity is constructed at least partially on the basis of the comparison of the series of images.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to three-dimensional imaging, and more particularly to intraoral three-dimensional imaging using structured light illumination.

Background Art

[0002] Three-dimensional surfaces within a subject's oral cavity, such as dental impressions of teeth and gums, are used to plan dental procedures. Conventional dental impressions are made using a dental impression tray filled with an impression material such as PVS or alginate, and the subject bites down on it. The impression material then hardens, forming a negative imprint of the teeth and gums from which a three-dimensional model of the teeth and gums can be formed.

[0003] Digital dental impressions generate a three-dimensional digital model of the three-dimensional surface within a subject's oral cavity using an intraoral scan. Digital intraoral scanners often use structured light three-dimensional imaging. The surface of a subject's teeth can be highly reflective and somewhat translucent, which can reduce the contrast of the structured light pattern reflected from the teeth. Therefore, when using a digital intraoral scanner that utilizes structured light three-dimensional imaging, in order to improve the capture of the intraoral scan and to enhance the available level of contrast of the structured light pattern, for example, the subject's teeth are often coated with an opaque powder prior to scanning to make the surface a scattering surface. Intraoral scanners that utilize structured light three-dimensional imaging have advanced to some extent, but further advantages may be possible.

Summary of the Invention

[0004] The use of structured light 3D imaging can lead to a "correspondence problem," where it is necessary to determine the correspondence between points in a structured light pattern and points seen by a camera inspecting the pattern. One technique to address this problem is based on projecting an "encoded" light pattern and imaging the illuminated scene from one or more viewpoints. By encoding the emitted light pattern, parts of the light pattern become unique and identifiable when captured by the camera system. Because the pattern is encoded, the correspondence between points in the image and points in the projected pattern is found more easily. The decoded points are triangulated, and 3D information is reconstructed.

[0005] Examples of applications of the present invention include systems and methods relating to a three-dimensional intraoral scanning apparatus that includes one or more cameras and one or more pattern projectors. For example, a particular application of the present invention may relate to an intraoral scanning apparatus having multiple cameras and multiple pattern projectors.

[0006] Further applications of the present invention include methods and systems for decoding structured light patterns.

[0007] Another application of the present invention may relate to a system and method for three-dimensional intraoral scanning utilizing unencoded structured light patterns.

[0008] For example, in some specific applications of the present invention, an intraoral scanning device is provided, which includes an elongated handheld wand having a probe at its distal end. During scanning, the probe may be configured to enter the oral cavity of a subject. One or more miniature structured light projectors and one or more miniature cameras are coupled to a rigid structure located within the distal end of the probe. Each structured light projector transmits light using a light source such as a laser diode. In some applications, the structured light projectors have an illumination field of at least 45 degrees. Optionally, the illumination field may be less than 120 degrees. Each structured light projector may further include a pattern-generating optical element. The pattern-generating optical element may generate a light pattern by utilizing diffraction and / or refraction. In some applications, the light pattern may be a distribution of discrete, unconnected light spots. Optionally, the light pattern maintains a distribution of discrete, unconnected spots in all planes located between 1 mm and 30 mm from the pattern-generating optical element when a light source (e.g., a laser diode) is activated and transmits light through the pattern-generating optical element. In some applications, the pattern-generating optical element of each structured light projector may have an optical throughput efficiency, i.e., the proportion of light falling into the pattern generator that enters the pattern, of at least 80%, for example, at least 90%. Each camera includes a camera sensor and an objective optical system including one or more lenses.

[0009] Laser diode light sources and diffraction and / or refraction pattern-generating optical elements can offer specific advantages in some applications. For example, the use of laser diodes and diffraction and / or refraction pattern-generating optical elements can help maintain energy-efficient structured light projectors by preventing the probe from overheating during use. Furthermore, such components may help reduce costs because they do not require active cooling within the probe. For example, current laser diodes may consume less than 0.6 watts of power while transmitting continuously at high brightness (in contrast to, for example, current light-emitting diodes (LEDs)). When pulsed, as in some applications of the present invention, these current laser diodes may consume even less power; for example, when pulsed at a 10% duty cycle, the laser diode may consume less than 0.06 watts (however, in some applications, the laser diode may consume at least 0.2 watts while transmitting continuously at high brightness, and even less power when pulsed; for example, when pulsed at a 10% duty cycle, the laser diode may consume at least 0.02 watts). Furthermore, the diffracting and / or refraction pattern generating optical element may be configured to utilize almost all, if not all, of the transmitted light (in contrast to, for example, a mask that prevents some of the light rays from hitting the object).

[0010] In particular, diffraction and / or refraction-based pattern-generating optical elements generate patterns by diffraction, refraction, or interference of light, or any combination thereof, rather than by modulation of light as performed with transparent or transmissive masks. In some applications, this can be advantageous because the optical throughput efficiency (the proportion of light entering the pattern from the light falling into the pattern generator) is nearly 100%, for example, at least 80%, or for example, at least 90%, regardless of the pattern's "area-based duty cycle." In contrast, the optical throughput efficiency of transparent or transmissive mask pattern-generating optical elements is directly related to the "area-based duty cycle." For example, if the desired "area-based duty cycle" is 100:1, the throughput efficiency of a mask-based pattern-generating optical element is 1%, while the efficiency of a diffraction and / or refraction-based pattern-generating optical element remains nearly 100%. Furthermore, the focusing efficiency of a laser is more than 10 times higher than that of an LED with the same total light output, because lasers inherently have a smaller emission area and divergence angle, resulting in a brighter output illuminance per unit area. Higher efficiency in lasers and diffraction / / or refraction pattern generators enables thermally efficient configurations, limiting excessive probe heating during use and potentially eliminating or limiting the need for active cooling within the probe, thereby reducing costs. In some applications, laser diodes or DOEs are particularly preferred, but these are not essential, either alone or in combination. Other light sources such as LEDs, and pattern generation elements including transparent or transmissive masks, can be used in other applications, with or without active cooling.

[0011] In several applications, to improve image capture of intraoral scenes under structured light illumination without using contrast-enhancing means such as coating teeth with opaque powders, the inventors have found that light patterns, such as a distribution of discrete, unconnected light spots (rather than lines), may offer an improved balance in increasing pattern contrast while maintaining a useful amount of information. Generally, denser structured light patterns may provide more surface sampling and higher resolution, potentially allowing for better stitching together of each surface obtained from multiple image frames. However, overly dense structured light patterns can complicate the matching problem because there are more spots to solve. Furthermore, dense structured light patterns may lead to a decrease in pattern contrast due to the increased light in the system, which can be caused by a combination of (a) stray light reflected off the slightly glossy surface of the tooth and picked up by the camera, and (b) percolation, i.e., some of the light incident on the tooth reflects along multiple paths within the tooth and then moves away from the tooth in many different directions. As will be further explained below, a method and system are provided to solve the problem of correspondences presented by the distribution of discrete, unconnected light spots. In some applications, the discrete, unconnected light spots from each projector do not need to be encoded.

[0012] In some applications, the field of view of each camera may be at least 45 degrees, for example, at least 80 degrees, for example, 85 degrees. Optionally, the field of view of each camera may be less than 120 degrees, for example, less than 90 degrees. In some applications, one or more cameras have a fisheye lens or other optical element that provides a field of view of up to 180 degrees.

[0013] In any case, the fields of view of the various cameras may or may not be the same. Similarly, the focal lengths of the various cameras may or may not be the same. As used herein, the term “field of view” of each camera means the diagonal field of view of each camera. Furthermore, each camera may be configured to focus at an objective focal plane located between 1 mm and 30 mm, for example, at least 5 mm and / or less than 11 mm, for example, between 9 mm and 10 mm, from the lens furthest from each camera sensor. Similarly, in some applications, the illumination field of each structured light projector may be at least 45 degrees and optionally less than 120 degrees. The inventors have recognized that a larger field of view achieved by combining the respective fields of view of all cameras may improve accuracy because it reduces the amount of image stitching error, especially in edentulous regions where the gingival surface is smooth and there may be fewer clear, high-resolution 3D features. Having a wide field of view allows large, smooth features such as the overall curve of the teeth to be visible in each image frame, thus improving the accuracy of stitching together the surfaces obtained from multiple image frames.

[0014] Several application examples provide methods for generating digital three-dimensional images of the oral cavity surface. It should be noted that the term "three-dimensional image" as used in this application is based on a three-dimensional model, such as a point cloud, from which a three-dimensional image of the oral cavity surface is constructed. The resulting image is typically displayed on a two-dimensional screen, but contains data related to the three-dimensional structure of the scanned object, and can therefore typically be manipulated to display the scanned object from different viewpoints or perspectives. Furthermore, data from the three-dimensional image may be used to create a physical three-dimensional model of the scanned object.

[0015] For example, one or more structured light projectors may be driven to project a pattern of light onto the oral surface, such as a distribution of discrete, unconnected light spots, a pattern of intersecting lines (e.g., a grid), a checkerboard pattern, or other patterns, and one or more cameras may be driven to capture images of the projection. The images captured by each camera may include a portion of the projection pattern (e.g., at least one of the spots). In some embodiments, one or more structured light projectors project a spatially fixed pattern to one or more cameras.

[0016] Each camera includes a camera sensor having a pixel array, where each pixel has a corresponding ray in 3D space originating from the pixel facing the object being imaged, and each point along a particular one of these rays, when imaged on the sensor, descends onto the corresponding pixel on the sensor. In this specification, including in the claims, the term “camera ray” is used in this regard. Similarly, each projection spot from each projector has a corresponding projector ray. Each projector ray corresponds to a path of each pixel on at least one camera sensor; that is, when the camera views a feature or portion (e.g., a spot) of a pattern projected by a particular projector ray, that feature or portion (e.g., a spot) of the pattern is necessarily detected by a pixel on a particular path of the pixel corresponding to that particular projector ray. (a) the camera rays corresponding to each pixel on the camera sensor of each camera, and (b) the projector rays corresponding to each feature or portion (e.g., light spot) of the pattern projected from each projector may be stored during the calibration process as described below.

[0017] With respect to camera rays, in some applications, instead of storing individual values ​​for each camera ray corresponding to each pixel on the camera sensor of each camera, a smaller set of calibration values ​​that can be used to represent each camera ray is stored. For example, to define a camera ray, parameter values ​​for a parameterized camera calibration function that takes a given 3D position in space and translates it to a given pixel in the 2D pixel array of the camera sensor may be stored.

[0018] With respect to projector rays, (a) in some applications, an indexed list containing the value for each projector ray is stored, and (b) alternatively, in some applications, a smaller set of calibration values ​​that can be used to represent each projector ray is stored. For example, parameter values ​​of a parameterized projector calibration model that defines each projector ray of a given projector may be stored.

[0019] Based on stored calibration values, the processor may execute a corresponding algorithm to determine the three-dimensional location of each part of the projection light pattern feature on the surface (e.g., projection spots). For a given projector ray, the processor "sees" the corresponding camera sensor path on one of the cameras. Detected spots and other features along that camera sensor path have camera rays that intersect with the given projector ray. The intersection defines a three-dimensional point in space. The processor then searches the camera sensor paths on the other cameras corresponding to the given projector ray and determines how many of the other cameras have detected pattern features (e.g., spots) that intersect with those three-dimensional points in space on their respective camera sensor paths corresponding to the given projector ray. As used herein, when two or more cameras detect a portion or feature (e.g., a spot) of a pattern at the same point in three-dimensional space where their respective camera rays intersect with a given projector ray, the cameras are considered to "agree" that the portion or feature (e.g., a spot) is located at that three-dimensional point. This process is repeated for additional features (e.g., spots) along the camera sensor path, and the feature (e.g., spot) that the most cameras "agree" on is identified as the feature (e.g., spot) projected onto the surface from the given projector ray. Thus, the three-dimensional position on the surface is calculated for that feature (e.g., the spot) of the pattern.

[0020] In some embodiments, once a position on the surface is determined for a particular feature of the pattern (e.g., a specific spot), the projector ray that projected that feature (e.g., the spot), and all camera rays corresponding to that feature (e.g., the spot), may be excluded from consideration, and the corresponding algorithm is run again for the next projector ray.

[0021] Further applications of the present invention involve projecting a structured light pattern (e.g., parallel lines, grid, checkerboard, unconnected and / or uniform spots, random spot patterns, etc.) onto an intraoral object, capturing at least a portion of the structured light pattern projected onto the intraoral object, and tracking the captured portion of the structured light pattern across a series of images to scan the intraoral object. In some embodiments, tracking a captured portion of the structured light pattern across a series of images may help improve scanning speed and / or accuracy.

[0022] In a more specific example relating to a structured optical scanner using the aforementioned projection pattern (e.g., of unconnected spots), a processor may be used to compare a series of images captured by each camera (e.g., multiple consecutive images) to determine which features of the projection pattern (e.g., which projection spots) can be tracked across the series of images (e.g., multiple consecutive images). The inventors have noticed that the movement of a particular detected feature or spot can be tracked across multiple images in a series of images (e.g., in consecutive image frames). Thus, a correspondence solved for a particular spot in any of the images or frames in which the feature or spot is tracked provides a solution for the correspondence for that feature or spot in all of the images or frames in which the feature or spot is tracked. Since a detected feature or spot that can be tracked across multiple images is a feature or spot generated by the same particular projector ray, the trajectory of the tracked feature or spot will follow a particular camera sensor path corresponding to that particular projector ray.

[0023] In some applications, instead of tracking detected features or spots in a two-dimensional image, the length of each projector ray can be tracked in three-dimensional space. The length of a projector ray is defined as the distance between the origin of the ray, i.e., the light source, and the three-dimensional position where the ray intersects the oral cavity surface. As will be further explained below, tracking the length of a particular projector ray over time can help resolve ambiguities in correspondences. While some examples herein describe the above concepts of spot and ray tracking in relation to scanners that project unconnected spots, this is illustrative and not limiting, and it should be understood that the tracking technique is equally applicable to scanners that project other patterns (e.g., parallel lines, grids, checkerboards, unconnected and / or uniform spots, random spot patterns, etc.) onto oral cavity objects.

[0024] In some embodiments, for the purpose of objective scanning, it may be desirable to estimate the position of the scanner relative to the object to be scanned, i.e., the three-dimensional intraoral surface, during scanning, and in some embodiments, continuous estimation during scanning is desirable. In some applications of the present invention, the inventors have developed a method to address times when sufficient visual tracking of the scanner's motion is unavailable by combining visual tracking of the scanner's motion with inertial measurement of the scanner's motion. Using accumulated data of the motion of the intraoral scanner relative to the intraoral surface (visual tracking) and the motion of the intraoral scanner relative to the fixed coordinate system (inertial measurement), a predictive model of the motion of the intraoral surface relative to the fixed coordinate system can be constructed (further described below). If sufficient visual tracking is unavailable, the processor may calculate the estimated position of the intraoral scanner relative to the intraoral surface by adding (for example, subtracting in some embodiments) a prediction of the motion of the intraoral surface relative to the fixed coordinate system from the inertial measurement of the motion of the intraoral scanner relative to the fixed coordinate system (further described below). It should be understood that the scanner position estimation concepts described herein can be used with intraoral scanners regardless of the scanning technique employed (e.g., parallel confocal scanning, focal scanning, wavefront scanning, stereo vision, structured light, triangulation, light field, and / or combinations thereof). Therefore, while the concept of structured light described herein is discussed in relation to this, this is illustrative and not limiting.

[0025] In some embodiments of the structured light scanner described herein, the stored calibration values ​​may represent (a) camera rays corresponding to each pixel on the camera sensor of each camera, and (b) projector rays corresponding to each projected feature (e.g., light spot) from each structured light projector, where each projector ray corresponds to the path of each pixel on at least one of the camera sensors. However, over time, at least one of the cameras and / or at least one of the projectors may move (e.g., by rotation or translation), the optical system of at least one of the cameras and / or at least one of the projectors may be changed, or the wavelength of the laser may be changed, and as a result, the stored calibration values ​​may no longer accurately correspond to the camera rays and projector rays.

[0026] If, for any given projector ray, the processor collects data including the calculated 3D position of each of multiple detected features (e.g., spots) from that projector ray detected at different points in time, and superimposes them onto a single image, then all of those features (e.g., spots) should descend onto the camera sensor path of the pixel corresponding to that projector ray. If the camera or projector calibration is changed for any reason, then the features (e.g., spots) detected from that particular projector ray may not appear to descend onto the camera sensor path of the pixel expected according to the stored calibration value, but rather appear to be located on the camera sensor path of the newly updated pixel. If the calibration of the camera(s) and / or projector(s) is changed, the processor may reduce the difference between the updated pixel path and the original pixel path from the calibration data by (i) changing the stored calibration values ​​(e.g., stored parameter values ​​of a parameterized camera calibration model, e.g., a function) that represent the camera rays corresponding to each pixel on the camera sensor of one or more cameras, and / or (ii) changing the stored calibration values ​​(e.g., stored values ​​in an indexed list of projector rays, or stored parameter values ​​of a parameterized projector calibration model) that represent the projector rays r corresponding to each feature (e.g., a light spot) projected from one or more projectors.

[0027] The current calibration evaluation may be performed automatically periodically (e.g., every scan, every 10 scans, monthly, every few months, etc.) or in response to a particular criterion being met (e.g., in response to a threshold number of scans having been performed). As a result of the evaluation, the system may determine whether the calibration state is accurate or inaccurate. In one embodiment, as a result of the evaluation, the system determines whether the calibration is drifting. For example, if the previous calibration still maintains sufficient accuracy to produce high-quality scans, but the detected trend continues, the system may deviate and potentially be unable to produce accurate scans in the future. In one embodiment, the system determines a drift rate and projects that drift rate into the future to determine a predicted date and time when the calibration will become inaccurate. In one embodiment, automatic or manual calibration can be scheduled for that future date and time. In one example, the processing logic evaluates the calibration state over time (e.g., by comparing the calibration states at multiple different points in time) and determines a drift rate from such a comparison. From the drift rate, the processing logic can predict the timing at which calibration should be performed based on trend data.

[0028] Conventional intraoral scanners require the user to manually re-calibrate according to a set schedule (e.g., every six months). Conventional intraoral scanners do not have a function to monitor or evaluate the current status of calibration (e.g., a function to determine whether re-calibration should be performed). Furthermore, calibration of conventional intraoral scanners is performed manually using a special calibration target. Calibration of conventional intraoral scanners takes time and is inconvenient for the user. Thus, the dynamic calibration performed in certain embodiments described herein can enhance user convenience and be performed in less time compared to calibration of conventional intraoral scanners.

[0029] In some applications, when the calibration of the camera(s) and / or projector(s) is changed, the processor need not perform a recalibration; rather, it may simply determine that at least some of the stored calibration values for the camera(s) and / or projector(s) are incorrect. For example, based on the determination that the stored calibration values are incorrect, the user may be prompted to return the intraoral scanner to the manufacturer for maintenance and / or recalibration, or to request a new scanner.

[0030] Visual tracking of the motion of the intraoral scanner relative to the object being scanned may be obtained by stitching together each surface or point cloud obtained from adjacent image frames. As described herein, in some applications, illuminating the oral cavity under near-infrared (NIR) light can increase the number of visible features that can be used to stitch together each surface or point cloud obtained from adjacent image frames. In particular, since NIR light can penetrate teeth, images captured under NIR light contain features inside the teeth, such as cracks inside the teeth, in contrast to two-dimensional color images captured under broadband illumination in which only features appearing on the surface of the teeth are visible. These additional subsurface features may be used to stitch together each surface or point cloud obtained from adjacent image frames.

[0031] In some applications, the processor may use two-dimensional images (e.g., two-dimensional color images and / or two-dimensional monochromatic NIR images) in the 2D-3D surface reconstruction of the three-dimensional surface of the oral cavity. As described below, using two-dimensional images (e.g., two-dimensional color images and / or two-dimensional monochromatic NIR images) can significantly improve the resolution and speed of the three-dimensional reconstruction. Therefore, as described herein, in some applications, it is useful to enhance the three-dimensional reconstruction of the three-dimensional surface of the oral cavity by reconstructing from two-dimensional images (e.g., two-dimensional color images and / or two-dimensional monochromatic NIR images). In some applications, the processor calculates the three-dimensional position of multiple points on the three-dimensional surface of the oral cavity using, for example, the corresponding algorithm described herein, and calculates the three-dimensional structure of the three-dimensional surface of the oral cavity based on multiple two-dimensional images (e.g., two-dimensional color images and / or two-dimensional monochromatic NIR images) and the calculated three-dimensional positions on the oral cavity surface.

[0032] According to some applications of the present invention, the calculation of the 3D structure is performed by a neural network. The processor inputs (a) multiple 2D images (e.g., 2D color images) of the 3D surface of the oral cavity and (b) the calculated 3D positions of multiple points on the 3D surface of the oral cavity to the neural network, and the neural network determines and returns the respective estimated maps (e.g., depth map, normal map, and / or curvature map of the 3D surface of the oral cavity captured in each of the 2D images (e.g., 2D color images and / or 2D monochromatic NIR images)).

[0033] The inventors have noticed that when intraoral scanners are commercially produced, small manufacturing deviations may exist, which may (a) cause the calibration of the camera(s) and / or projector(s) on each commercially produced intraoral scanner to differ slightly from the calibration of the training stage camera(s) and / or projector(s), and / or (b) cause the illumination relationship between the camera(s) and projector(s) of a commercially produced intraoral scanner to differ slightly from the illumination relationship between the training stage camera(s) and projector(s) used to train a neural network. Other manufacturing deviations of the camera(s) and / or projector(s) may exist. Several applications of the present invention provide a method for using a processor to overcome manufacturing deviations of the camera(s) and / or projector(s) of an intraoral scanner and reduce the difference between the estimated map and the true structure of the intraoral 3D surface.

[0034] According to several applications of the present invention, one method by which manufacturing tolerances can be overcome is to modify, for example, crop and morph images from an intraoral scanner in the field to obtain modified images that match the field of view of a set of reference cameras used to train a neural network. The neural network is trained on images received from the set of reference cameras, and then the field images are modified so that the neural network receives these images as if they were captured by the reference cameras. The three-dimensional structure of the intraoral 3D surface is then calculated based on multiple modified two-dimensional images of the intraoral 3D surface, for example, the neural network determines the respective estimated map of the intraoral 3D surface captured in each of the multiple modified two-dimensional images.

[0035] In several applications of the present invention, a neural network determines an estimated depth map for each 3D surface of the oral cavity captured in each 2D image, and these depth maps are stitched together to obtain the 3D structure of the oral cavity surface. However, inconsistencies can sometimes occur between the estimated depth maps. The inventors have found it advantageous for the neural network to also determine an estimated confidence map for each estimated depth map determined by the neural network, with each confidence map indicating the confidence level for each region of the respective estimated depth map. Accordingly, a method for inputting multiple 2D images of the 3D surface of the oral cavity into a first neural network module and a second neural network module is provided herein. The first neural network module determines an estimated depth map for each 3D surface of the oral cavity captured in each 2D image. The second neural network module determines an estimated confidence map corresponding to each estimated depth map. Each confidence map indicates the confidence level for each region of the respective estimated depth map.

[0036] According to some applications of the present invention, a neural network is trained using (a) two-dimensional images of a training stage 3D surface, such as a model surface and / or an intraoral surface, and (b) a corresponding true output map of the training stage 3D surface calculated based on a structured optical image of the training stage 3D surface. For each two-dimensional image, the neural network estimates a predicted map of the intraoral 3D surface captured in each two-dimensional image, and then compares each predicted image with the corresponding true map of the intraoral 3D surface. Based on the difference between each predicted map and the corresponding true map, the neural network is optimized to better estimate subsequent predicted maps.

[0037] In some applications, when the 3D surface of the oral cavity is used to train a neural network, moving tissues, such as the subject's tongue, lips, and / or cheeks, may obstruct parts of the 3D surface from the field of view of one or more cameras. To prevent the neural network from "learning" based on images of moving tissues (as opposed to fixed tissues of the scanned 3D surface of the oral cavity), for 2D images in which moving tissues are identified, the images may be processed to exclude at least parts of the moving tissues before the 2D images are input to the neural network.

[0038] In some applications, to prevent cross-contamination between patients, a disposable sleeve is placed on the distal end of the intraoral scanner, for example, on the probe, before the probe is placed in the patient's mouth. As further described herein, due to the relative positional relationship between the structured light projector in the probe and the adjacent camera, some of the projected structured light pattern may be reflected from the sleeve and reach the camera sensor of the adjacent camera. As further described herein, due to the polarization of the laser light of the structured light projector, the laser may be rotated around its own optical axis so that the polarization angle of the laser light relative to the sleeve is determined to reduce the degree of reflection.

[0039] In several applications of the present invention, a simultaneous localization and mapping (SLAM) algorithm is used to track the motion of a handheld wand and generate a three-dimensional image. SLAM can generally be performed using two or more cameras that view the same image but from slightly different angles. However, due to the positioning of the camera 24 within the probe 28 and the proximity positioning of the probe 28 to the object being scanned, i.e., the three-dimensional surface of the oral cavity, the two or more cameras within the probe often do not view approximately the same image. Additional challenges can be encountered when using SLAM algorithms when scanning the three-dimensional surface of the oral cavity, as described below. The inventors have invented several methods to overcome these challenges in order to track the motion of a handheld wand and generate a three-dimensional image of the three-dimensional surface of the oral cavity using SLAM, as further described herein.

[0040] In some applications of the present invention, when scanning a three-dimensional intraoral surface using a handheld wand, and a structured light projector projects a distribution of features (e.g., a distribution of spots) onto the intraoral surface, some of the features (e.g., spots) may land on moving tissue (e.g., the patient's tongue). To improve the accuracy of the three-dimensional reconstruction algorithm, features (e.g., spots) that land on moving tissue should generally not be relied upon for the reconstruction of the three-dimensional intraoral surface. As described herein, whether a feature (e.g., spot) is projected onto moving tissue or stable tissue in the oral cavity may be determined on unstructured light image frames (e.g., broad-spectrum light) scattered within the structured light image frame. A confidence grading system may be used to assign confidence grades based on the determination of whether the detected features (e.g., spots) are projected onto fixed tissue or moving tissue. Based on the confidence grades for each of the multiple features (e.g., spots), the processor may run a three-dimensional reconstruction algorithm using the detected features (e.g., spots).

[0041] In one method for generating a digital three-dimensional image as defined in the present invention, the method includes the step of driving each optical projector of one or more structured optical projectors to project a pattern onto a three-dimensional surface of the oral cavity. The method further includes the step of driving each camera of one or more cameras to capture a plurality of images, each image including at least a portion of the projection pattern. The method further includes the step of using the processor to compare a series of images captured by the one or more cameras, determining which portions of the projection pattern can be tracked across the series of images based on the comparison of the series of images, and constructing a three-dimensional model of the three-dimensional surface of the oral cavity, at least in part based on the comparison of the series of images. In one embodiment, the method further includes the step of solving a correspondence algorithm for the tracked portion of the projection pattern in at least one image of the series of images, and using the solved correspondence algorithm to solve the correspondence algorithm for the tracked portion of the projection pattern in at least one image of the series of images, for example, in images of a series for which the correspondence algorithm has not been solved, and constructing a three-dimensional model using the solution of the correspondence algorithm. In one embodiment, the method further includes the step of solving a correspondence algorithm for the tracking portion of the projection pattern based on a portion of the tracking position of the tracking portion in each image throughout the entire series of images, and constructing a three-dimensional model using the solution of the correspondence algorithm.

[0042] In one embodiment of the method, the projection pattern includes a plurality of projection light spots, and the portion of the projection pattern corresponds to the projection spots of the plurality of projection light spots. In a further embodiment, the step of using the processor to compare a series of images based on stored calibration values ​​indicating (a) camera rays corresponding to each pixel on the camera sensor of each camera of the one or more cameras, and (b) projector rays corresponding to each projection light spot from each structured light projector of the one or more structured light projectors, wherein each projector ray corresponds to a projector ray corresponding to the path of each pixel on at least one of the camera sensors, and determining which portion of the projection pattern can be tracked, includes determining which of the projection spots s can be tracked across the series of images, each tracked spot s moves along the path of the pixel corresponding to the respective projector ray r.

[0043] In a further embodiment of the method, the step of using the processor further includes using the processor to determine a plurality of possible paths p of a pixel on one of the cameras for each tracking spot s, where each path p corresponds to a plurality of possible projector rays r. In a further embodiment, the step of using the processor further includes using the processor to perform a plurality of operations for each of the possible projector rays r by executing a corresponding algorithm. The plurality of operations include determining how many other cameras have detected each spot q corresponding to each camera ray that intersects the projector ray r with the camera ray of one of the cameras corresponding to the tracking spot s on the path p1 of their respective pixels corresponding to the projector ray r. The operations further include determining a given projector ray r1 from which the most other cameras detected each spot q. The operations further include determining the projector ray r1 as the specific projector ray r that generated the tracking spot s.

[0044] In a further embodiment of the method, the method includes using the processor to (a) execute a corresponding algorithm to calculate the respective three-dimensional positions of a plurality of detection spots on the three-dimensional surface of the oral cavity captured in a series of images, and (b) in at least one of the series of images, identifying that a detection spot originates from a particular projector ray r by identifying that the detection spot is a tracking spot s that moves along the path of a pixel corresponding to the particular projector ray r.

[0045] In a further embodiment of the method, the method includes using the processor to (a) perform a corresponding algorithm to calculate the three-dimensional position of each of a plurality of detection spots on the three-dimensional surface of the oral cavity that have been captured in the series of images, and (b) exclude from consideration as points on the three-dimensional surface of the oral cavity any spots that are not identified as points on the three-dimensional surface of the oral cavity occult algorithm to (a) perform a corresponding algorithm to calculate the three-dimensional position of each of a plurality of detection spots on the oral cavity that have been captured in the series of images.

[0046] In a further embodiment of the method, the method includes using the processor to (a) perform the corresponding algorithm to calculate the three-dimensional position of each of a plurality of detection spots on the three-dimensional surface of the oral cavity captured in the series of images, and (b) for detection spots that have been identified as originating from two different projector rays r based on the three-dimensional position calculated by the corresponding algorithm, the step of identifying the spot as originating from one of the two different projector rays r by identifying the detection spot as a tracking spot s that moves along one of the two different projector rays r.

[0047] In a further embodiment of the method, the method includes using the processor to (a) perform the corresponding algorithm to calculate the three-dimensional position of each of a plurality of detection spots on the three-dimensional surface of the oral cavity captured in a series of images, and (b) identify vulnerable spots whose three-dimensional position was not calculated by the corresponding algorithm as projection spots from a particular projector ray r by identifying them as tracking spots s moving along the path of pixels corresponding to a particular projector ray r.

[0048] In a further embodiment of the method, the method includes the step of using the processor to calculate the respective three-dimensional position on the oral cavity three-dimensional surface at the intersection of the projector ray r and the respective camera ray corresponding to the tracked spot s in each of the series of images in which the spot s is tracked.

[0049] In a further embodiment of the method, the three-dimensional model is constructed using a corresponding algorithm, which uses at least partially the portion of the projection pattern that is determined to be traceable across the series of images.

[0050] In a further embodiment of the method, the method includes using the processor to (a) determine parameters of a traceable portion of a projection pattern in at least two adjacent images from the series of images, wherein the parameters are selected from the group consisting of the size of the portion, the shape of the portion, the orientation of the portion, the intensity of the portion, and the signal-to-noise ratio (SNR) of the portion; and (b) predict the parameters of the traceable portion of the projection pattern in later images based on the parameters of the traceable portion of the projection pattern in the at least two adjacent images.

[0051] In a further embodiment of the method, the step of using the processor further includes using the processor to search for a portion of the projection pattern that substantially has the predicted parameters in the subsequent image, based on the predicted parameters of the tracking portion of the projection pattern.

[0052] In a further embodiment of the method, the parameter is the shape of the portion of the projection pattern, and the step of using the processor further includes using the processor to determine a search space in the next image for searching the tracking portion of the projection pattern, based on the predicted shape of the tracking portion of the projection pattern.

[0053] In a further embodiment of the method, the step of determining the search space using the processor includes the step of determining the search space in the next image for searching the tracking portion of the projection pattern using the processor, wherein the search space has a size and aspect ratio based on the size and aspect ratio of the predicted shape of the tracking portion of the projection pattern.

[0054] In a further embodiment of the method, the parameter is the shape of the portion of the projection pattern, and the step of using the processor further includes using the processor to (a) determine a velocity vector of the tracking portion of the projection pattern based on the direction and distance the tracking portion of the projection pattern has moved between the at least two adjacent images from the series of images; (b) predict the shape of the tracking portion of the projection pattern in a later image in response to the shape of the tracking portion of the projection pattern in at least one of the at least two adjacent images; and (c) determine a search space for searching the tracking portion of the projection pattern in the later image in response to a combination of (i) determining the velocity vector of the tracking portion of the projection pattern and (ii) predicting the shape of the tracking portion of the projection pattern.

[0055] In a further embodiment of the method, the parameter is the shape of the portion of the projection pattern, and the step of using the processor further includes using the processor to (a) determine the velocity vector of the tracking portion of the projection pattern based on the direction and distance the tracking portion of the projection pattern has moved between at least two adjacent images from the series of images; (b) predict the shape of the tracking portion of the projection pattern in a later image in response to the determination of the velocity vector of the tracking portion of the projection pattern; and (c) determine a search space for searching the tracking portion of the projection pattern in the later image in response to a combination of (i) the determination of the velocity vector of the tracking portion of the projection pattern and (ii) the predicted shape of the tracking portion of the projection pattern.

[0056] In a further embodiment of the method, the step of using the processor further includes using the processor to predict the shape of the tracking portion of the projection pattern in a later image in response to a combination of (i) determining the velocity vector of the tracking portion of the projection pattern and (ii) the shape of the tracking portion of the projection pattern in at least one of the two adjacent images.

[0057] In a further embodiment of the method, the step of using the processor further includes using the processor to (a) determine a velocity vector of the tracking portion of the projection pattern based on the direction and distance it has moved between two consecutive images in the series of images, and (b) determine a search space in a later image for searching the tracking portion of the projection pattern in response to the determination of the velocity vector of the tracking portion of the projection pattern.

[0058] In one embodiment of a second method for generating a digital three-dimensional image as defined herein, the method includes the steps of driving each of one or more structured light projectors to project a pattern of light along a plurality of projector rays onto a three-dimensional surface in the oral cavity, and driving each of one or more cameras to capture a plurality of images, each image comprising at least a portion of the projection pattern, and each of the one or more cameras comprising a camera sensor including an array of pixels. The second method further includes the steps of: using the processor to execute a corresponding algorithm to calculate the respective three-dimensional position of a plurality of detected features of the projection pattern on the oral cavity three-dimensional surface for each of the plurality of images; estimating a three-dimensional surface based on the at least three features using data corresponding to the respective three-dimensional positions of at least three features, each of which features corresponds to a respective projector ray r of a plurality of projector rays; estimating the three-dimensional position in the intersection space between the projector ray r1 and the estimated three-dimensional surface for projector rays r1 of a plurality of projector rays for which the three-dimensional position of the feature corresponding to the projector ray r1 has not been calculated; and using the three-dimensional position in the estimated space to identify a search space in the pixel array of at least one camera for searching for the feature corresponding to the projector ray r1.

[0059] In a further embodiment of the second method, the corresponding algorithm is performed based on stored calibration values ​​that indicate (a) camera rays corresponding to each pixel on the camera sensor of each of the one or more cameras, and (b) projector rays corresponding to each feature of the projection pattern from each of the one or more structured light projectors, wherein each projector ray corresponds to a projector ray corresponding to the path of each pixel on at least one of the camera sensors. Furthermore, the search space in the data includes a search space defined by one or more thresholds.

[0060] In a further embodiment of the second method, the processor sets a threshold such that detected features below the threshold are not considered by the corresponding algorithm, and in order to search for features corresponding to the projector ray r1 in the identified search space, the processor lowers the threshold to consider features that were not considered by the corresponding algorithm. In some embodiments, the threshold is an intensity threshold.

[0061] In a further embodiment of the second method, the light pattern includes a distribution of discrete spots, and each feature includes a spot from the distribution of discrete spots.

[0062] In a further embodiment of the second method, the data corresponding to the three-dimensional positions of each of the at least three features includes the step of using data corresponding to the three-dimensional positions of each of the at least three features that are all captured in one of the plurality of images.

[0063] In a further embodiment of the second method, the second method further includes a step of refining the estimation of the three-dimensional surface using data corresponding to the three-dimensional position of at least one additional feature of the projection pattern, wherein the at least one additional feature has a three-dimensional position calculated based on another image among the plurality of images. In a further embodiment of the second method, the step of refining the estimation of the three-dimensional surface includes a step of refining the estimation of the three-dimensional surface such that all of the at least three features and the at least one additional feature are located on the estimated three-dimensional surface.

[0064] In a further embodiment of the second method, the step of using data corresponding to the respective three-dimensional positions of at least three features includes the step of using data corresponding to at least three features captured in each of the plurality of images.

[0065] In one embodiment of a third method for generating a digital three-dimensional image, the third method includes the steps of driving each of one or more structured light projectors to project a pattern of light onto a three-dimensional surface inside the mouth, and driving each of a plurality of cameras to capture an image, wherein the image includes at least a portion of the projected pattern, and each of the plurality of cameras includes a camera sensor including an array of pixels. The third method further includes the steps of: using a processor to execute a corresponding algorithm to calculate the three-dimensional position of each of several features of the projection pattern on the three-dimensional surface of the oral cavity; using data from a first camera among the plurality of cameras to identify candidate three-dimensional positions of a given feature of a projection pattern that corresponds to or is otherwise associated with one or more specific projector rays r, wherein data from a second camera among the plurality of cameras is not used to identify the candidate three-dimensional positions; using the candidate three-dimensional positions seen by the first camera to identify a search space on the pixel array of the second camera for searching for features of the projection pattern from the projector rays r; and, if the features of the projection pattern from the projector rays r are identified in the search space, using data from the second camera to refine the candidate three-dimensional positions of the features of the projection pattern.

[0066] In a further embodiment of the third method, to identify candidate three-dimensional positions of a given spot corresponding to a particular projector ray r, the processor uses data from at least two of the cameras, data from another camera other than one of the at least two cameras is not used to identify the candidate three-dimensional position, and to identify the search space, the processor uses candidate three-dimensional positions seen by at least one of the at least two cameras.

[0067] In a further embodiment of the third method, the light pattern includes a distribution of discrete, unconnected light spots, and the features of the projection pattern include projection spots from the unconnected light spots.

[0068] In a further embodiment of the third method, the processor uses stored calibration values ​​that indicate (a) camera rays corresponding to each pixel on the camera sensor of each of the plurality of cameras, and (b) projector rays corresponding to each feature of the projection light pattern from each of the one or more structured light projectors, wherein each projector ray is a projector ray corresponding to the path of each pixel of at least one of the camera sensors.

[0069] A fourth method for generating a digital three-dimensional image as described herein, the fourth method comprising the steps of driving one or more structured light projectors to project a pattern of light onto a three-dimensional surface of the oral cavity, and driving one or more cameras to capture an image, the image comprising at least a portion of the pattern. The fourth method further comprises the steps of using the processor to execute a corresponding algorithm to calculate the respective three-dimensional positions of a plurality of features of the pattern on the three-dimensional surface of the oral cavity captured in a series of images, identifying the calculated three-dimensional positions of the detected features of the captured pattern to be associated with one or more specific projector rays r in at least a subset of the series of images, and evaluating the length associated with the one or more projector rays r in each image of the subset of images based on the three-dimensional positions of the detected features corresponding to the one or more projector rays r in the subset of images.

[0070] In the fourth method, the processor may further be used to calculate the estimated length of the one or more projector rays r in at least one image of a series of images in which the three-dimensional positions of features projected from the one or more projector rays have not been determined.

[0071] In one embodiment of the fourth method, each of the one or more cameras comprises a camera sensor including a pixel array, and the calculation of the respective three-dimensional positions of multiple features of a pattern on the three-dimensional surface of the oral cavity, and the identification that the calculated three-dimensional positions of the detected features of the pattern correspond to a particular projector ray r, is performed based on stored calibration values ​​indicating (i) a camera ray corresponding to a pixel of each of the one or more camera sensors, and (ii) a projector ray corresponding to each feature of the projection light pattern from each of the one or more projectors, wherein each projector ray corresponds to a projector ray corresponding to the path of each pixel on at least one of the camera sensors.

[0072] In a further embodiment of the fourth method, the step of using the processor further includes using the processor to calculate an estimated length of the projector ray r in at least one of the series of images in which the three-dimensional position of a feature projected from the projector ray r has not been determined, and determining a one-dimensional search space for searching for the feature projected from the projector ray r in at least one of the series of images based on the estimated length of the projector ray r in at least one of the series of images, wherein the one-dimensional search space is along the path of each pixel corresponding to the projector ray r.

[0073] In a further embodiment of the fourth method, the step of using the processor further includes using the processor to calculate an estimated length of the projector ray r in at least one image of the series of images in which the three-dimensional position of a feature projected from the projector ray r has not been determined; and, based on the estimated length of the projector ray r in at least one image of the series of images, determining a one-dimensional search space in the pixel array of each of a plurality of cameras that search for a spot projected from the projector ray r for each pixel array, wherein the one-dimensional search space is along the path of each pixel corresponding to the ray r.

[0074] In a further embodiment of the fourth method, the step of using the processor to determine a one-dimensional search space in each pixel array of a plurality of cameras includes the step of using the processor to determine a one-dimensional search space in each pixel array of all cameras that search for features projected from a projector ray r.

[0075] In a further embodiment of the fourth method, the step of using the processor further includes using the processor to identify a plurality of candidate three-dimensional positions of features projected from the projector ray r in each of at least one image of a set of images that are not in the subset of the image, based on the corresponding algorithm, and calculating an estimated length of the projector ray r in at least one image of the set of images in which the plurality of candidate three-dimensional positions of features projected from the projector ray r have been identified.

[0076] In a further embodiment of the fourth method, the step of using the processor further includes determining whether the projected feature is the correct three-dimensional position by using the processor to determine which of the plurality of candidate three-dimensional positions corresponds to the estimated length of the projector ray r in at least one of the series of images.

[0077] In a further embodiment of the fourth method, the step of using the processor further includes using the processor to determine a one-dimensional search space in at least one of the series of images for searching for features projected from the projector ray r, based on the estimated length of the projector ray r in at least one of the series of images; and determining which of the three-dimensional positions of the plurality of candidates for the projected feature is the correct three-dimensional position of the projected feature generated by the projector ray r, by determining which of the three-dimensional positions of the plurality of candidates corresponds to the feature generated by the projector ray r detected in the one-dimensional search space.

[0078] In a further embodiment of the fourth method, the step of using the processor includes using the processor to define a curve based on the evaluated length of a projector ray r in each image of a subset of the images, and removing from consideration a detection feature identified as originating from a projector ray r if the three-dimensional position of the projected feature corresponds to a length of the projector ray r that is at least a threshold distance away from the defined curve, as a point on the three-dimensional surface of the oral cavity.

[0079] In a further embodiment of the fourth method, the pattern comprises a plurality of spots, and each of the plurality of features of the pattern comprises one of the plurality of spots.

[0080] A fifth method for generating a digital three-dimensional image as defined herein, the method comprising the steps of driving each of one or more structured light projectors to project a pattern of light onto an intraoral three-dimensional surface along a plurality of projector rays, and driving each of one or more cameras to capture a plurality of images, each image comprising at least a portion of the projection pattern, and each of the one or more cameras comprising a camera sensor having an array of pixels. The method further includes the steps of: using the processor to execute a corresponding algorithm to calculate the respective three-dimensional positions of multiple detected features of the projection pattern on the oral cavity three-dimensional surface for each of the multiple images; using data corresponding to the respective three-dimensional positions of at least three of the detected features to estimate a three-dimensional surface based on at least three features, where each feature corresponds to a respective projector ray r among the multiple projector rays; for a projector ray r1 from the multiple projector rays, for which multiple candidate three-dimensional positions of features corresponding to projector ray r1 have been calculated, to estimate the three-dimensional position in the intersection space between projector ray r1 and the estimated three-dimensional surface; and using the estimated three-dimensional position in the intersection space of projector ray r1 to select which of the multiple candidate three-dimensional positions is the correct three-dimensional position of the feature corresponding to that projector ray r1.

[0081] In a further embodiment of the fifth method, the corresponding algorithm is performed based on stored calibration values ​​that indicate (a) camera rays corresponding to each pixel on the camera sensor of each of the one or more cameras, and (b) projector rays corresponding to each feature of the projection pattern from each of the one or more structured light projectors, wherein each projector ray corresponds to a projector ray corresponding to each path on at least one pixel of the camera sensor. Furthermore, the search space in the data includes a search space defined by one or more thresholds.

[0082] In a further embodiment of the fifth method, the light pattern comprises a distribution of discrete spots, and each feature comprises one spot from the distribution of discrete spots.

[0083] In a further embodiment of the fifth method, the data corresponding to the three-dimensional positions of each of the at least three features includes the step of using data corresponding to the three-dimensional positions of each of the at least three features that are all captured in one of the plurality of images.

[0084] In a further embodiment of the fifth method, the fifth method further includes a step of refining the estimation of the three-dimensional surface using data corresponding to the three-dimensional position of at least one additional feature of the projection pattern, wherein the at least one additional feature has a three-dimensional position calculated based on another image of a plurality of images. In a further embodiment of the fifth method, the step of refining the estimation of the three-dimensional surface includes a step of refining the estimation of the three-dimensional surface such that all of the at least three features and the at least one additional feature are located on the estimated three-dimensional surface.

[0085] In a further embodiment of the fifth method, the step of using data corresponding to the three-dimensional positions of at least three features includes the step of using data corresponding to at least three features captured in each of the plurality of images.

[0086] One method for tracking the motion of an intraoral scanner as defined herein includes the steps of: measuring the motion of the intraoral scanner relative to an intraoral surface being scanned using at least one camera coupled to the intraoral scanner; and measuring the motion of the intraoral scanner relative to an intraoral surface being scanned relative to a fixed coordinate system using at least one inertial measuring unit (IMU) coupled to the intraoral scanner. The method further includes the steps of using the processor to calculate the motion of the intraoral surface relative to the fixed coordinate system based on (a) the motion of the intraoral scanner relative to the intraoral surface and (b) the motion of the intraoral scanner relative to the fixed coordinate system; constructing a predictive model for the motion of the intraoral surface relative to the fixed coordinate system based on accumulated data of the motion of the intraoral surface relative to the fixed coordinate system; and further calculating the estimated position of the intraoral scanner relative to the intraoral surface based on (a) a prediction of the motion of the intraoral surface relative to the fixed coordinate system (derived based on the motion predictive model) and (b) the motion of the intraoral scanner relative to the fixed coordinate system (measured by the IMU). In a further embodiment of the motion tracking method, the method further includes the steps of determining whether the measurement of the motion of the intraoral scanner relative to the intraoral surface using the at least one camera is being obstructed, and calculating the estimated position of the intraoral scanner relative to the intraoral surface in response to the determination that the measurement of the motion is being obstructed. In a further embodiment of the motion tracking method, the motion calculation is performed by calculating the difference between (a) the motion of the intraoral scanner relative to the intraoral surface and (b) the motion of the intraoral scanner relative to a fixed coordinate system.

[0087] One method for determining whether calibration data for an intraoral scanner is inaccurate, as defined herein, includes the steps of: driving each light source of one or more light sources to project light onto a three-dimensional intraoral surface; and driving each camera of one or more cameras to capture a plurality of images of the three-dimensional intraoral surface. The method further includes the steps of: using a processor to execute a corresponding algorithm based on stored calibration data relating to one or more light sources and one or more cameras to calculate the respective three-dimensional positions of a plurality of features of the projected light on the three-dimensional intraoral surface; collecting data at a plurality of points in time, the data including the calculated respective three-dimensional positions of the plurality of features on the three-dimensional intraoral surface; and determining, based on the collected data, that at least a portion of the stored calibration data is inaccurate.

[0088] In a further embodiment of the method, the one or more light sources are one or more structured light projectors, the method comprises the steps of driving each of the one or more structured light projectors to project a pattern of light onto a three-dimensional surface in the oral cavity, and driving each of the one or more cameras to capture a plurality of images of the three-dimensional surface in the oral cavity, each image comprising at least a portion of the projection pattern, the camera of the one or more cameras comprising a camera sensor comprising an array of pixels, the stored calibration data comprising stored calibration values ​​indicating (a) camera rays corresponding to each pixel on the camera sensor of each of the cameras of the one or more cameras, and (b) projector rays corresponding to each feature of the light pattern projected from each of the one or more structured light projectors, wherein each projector ray corresponds to a path p of each pixel on at least one of the camera sensors.

[0089] In a further embodiment of the method, the step of determining that at least a portion of the stored calibration data is inaccurate includes: using the processor to define a pixel update path p' for each camera sensor such that, for each projector ray r, based on the collected data, all of the calculated three-dimensional positions corresponding to features generated by the projector ray r correspond to positions along the respective update path p' of the pixels for each camera sensor; comparing the respective update path p' of the pixels for each camera sensor corresponding to that projector ray r from the stored calibration values ​​to the pixel path p; and determining that at least some of the stored calibration values ​​are inaccurate in response that the update path p' for at least one camera sensor s is different from the pixel path p corresponding to that projector ray r from the stored calibration values.

[0090] One method of recalibration as defined herein includes the steps of driving each of one or more light sources to project light onto a three-dimensional intraoral surface, and driving each of one or more cameras to capture a plurality of images of the three-dimensional intraoral surface. The method includes the steps of using a processor to execute a corresponding algorithm based on stored calibration data of the one or more light sources and the one or more cameras to calculate the respective three-dimensional positions of a plurality of features of the projected light on the three-dimensional intraoral surface, collecting data at a plurality of points in time, the data including the calculated respective three-dimensional positions of the plurality of features on the three-dimensional intraoral surface, and using the collected data to recalibrate the stored calibration data.

[0091] In a further embodiment of the recalibration method, the one or more light sources are one or more structured light projectors, and the method includes the steps of driving each of the one or more structured light projectors to project a pattern of light onto a three-dimensional surface of the oral cavity, and driving each of the one or more cameras to capture a plurality of images of the three-dimensional surface of the oral cavity, each image including at least a portion of the projection pattern, and each of the one or more cameras is equipped with a camera sensor including an array of pixels. The processor performs a plurality of operations using the stored calibration data, the stored calibration data including stored calibration values, which include (a) camera rays corresponding to each pixel on the camera sensor of each of the cameras, and (b) projector rays corresponding to each feature of the light pattern projected from each of the one or more structured light projectors, wherein each projector ray corresponds to a path p of each pixel on at least one of the camera sensors. The operation includes the step of executing a corresponding algorithm to calculate the respective three-dimensional positions of multiple features of the projection pattern on the three-dimensional surface of the oral cavity. The operation further includes the step of collecting data at multiple points in time, the data including the calculated respective three-dimensional positions of the multiple features on the three-dimensional surface of the oral cavity. The operation further includes the step of defining a pixel update path p' for each of the camera sensors, based on the collected data, such that for each projector ray r, all of the calculated three-dimensional positions corresponding to features generated by the projector ray r correspond to positions along the respective pixel update path p' for each of the camera sensors. The operation further includes the step of recalibrating the stored calibration values ​​using the update path p'.

[0092] In a further embodiment of the recalibration method, the processor performs additional operations to recalibrate the stored calibration values. The additional operations include comparing each update pass p' of a pixel with the path p of the pixel corresponding to its projector ray r on each camera sensor from the stored calibration values. The additional operations further include, for at least one camera sensor s, if the update pass p' of the pixel corresponding to the projector ray r differs from the path p of the pixel corresponding to the projector ray r from the stored calibration values, reducing the difference between the update pass p' of the pixel corresponding to each projector ray r and the path p of the pixel corresponding to each projector ray r from the stored calibration values.

[0093] In a further embodiment of the recalibration method, the stored calibration data to be changed includes stored calibration values ​​indicating camera rays corresponding to each pixel on the camera sensor of each of the one or more cameras. Furthermore, the step of changing the stored calibration data includes changing one or more parameters of a parameterized camera calibration function that defines camera rays corresponding to each pixel on at least one camera sensor s in order to reduce the difference between (i) the calculated three-dimensional positions of the plurality of features of the projection pattern on the three-dimensional surface of the oral cavity and (ii) stored calibration values ​​indicating the camera rays corresponding to each pixel on the camera sensor that one of the plurality of features should be detected.

[0094] In a further embodiment of the recalibration method, the stored calibration data to be changed includes stored calibration values ​​indicating projector rays corresponding to each of the plurality of features from each of the one or more structured light projectors, and the step of changing the stored calibration data includes (i) an indexed list assigning each projector ray r to a pixel path p, or (ii) changing one or more parameters of a parameterized projector calibration model that defines each projector ray r.

[0095] In a further embodiment of the recalibration method, the step of modifying the stored calibration data includes modifying the indexed list by reassigning each projector ray r based on the respective update path p' of the pixel corresponding to each projector ray r.

[0096] In a further embodiment of the recalibration method, the step of changing the stored data includes changing (i) a stored calibration value representing a camera ray corresponding to each pixel on the camera sensor s of each of the one or more cameras, and (ii) a stored calibration value representing a projector ray r corresponding to each feature of a plurality of features from each of the one or more structured light projectors.

[0097] In a further embodiment of the recalibration method, the step of changing the stored calibration value includes the step of iteratively changing the stored calibration value.

[0098] In a further embodiment of the recalibration method, the method further includes the step of driving each camera of one or more cameras to capture a plurality of images of a calibration object having predetermined parameters. The recalibration method further includes the steps of using the processor to perform a triangulation algorithm to calculate the respective parameters of the calibration object based on the captured images, and performing an optimization algorithm to (a)(i) reduce the difference between the update path p' of the pixels corresponding to the projector rays r and (ii) the path p of the pixels corresponding to the projector rays r from the stored calibration values, and (b) using the respective parameters of the calibration object calculated based on the captured images.

[0099] In a further embodiment of the recalibration method, the object to be calibrated is a three-dimensional object of known shape, and the step of driving each of the one or more cameras to capture a plurality of images of the object to be calibrated includes the step of driving each of the one or more cameras to capture an image of the three-dimensional object to be calibrated, wherein the predetermined parameter of the object to be calibrated is the dimensions of the three-dimensional object to be calibrated. In a further embodiment, with respect to the processor using the calculated respective parameters of the object to be calibrated to execute an optimization algorithm, the processor further uses collected data including the calculated respective three-dimensional positions of the plurality of features on the three-dimensional surface of the oral cavity.

[0100] In a further embodiment of the recalibration method, the object to be calibrated is a two-dimensional object having visually distinguishable features, and the step of driving each camera of one or more cameras to capture a plurality of images of the object to be calibrated includes the step of driving each camera of one or more cameras to capture an image of the two-dimensional object to be calibrated, wherein a predetermined parameter of the two-dimensional object to be calibrated is the respective distance between each of the visually distinguishable features. In a further embodiment, with respect to the processor using the respective calculated parameters of the object to be calibrated to execute the optimization algorithm, the processor further uses acquired data including the respective calculated three-dimensional positions of the plurality of features on the three-dimensional surface of the oral cavity.

[0101] In a further embodiment of the recalibration method, the step of driving each camera of the one or more cameras to capture an image of the two-dimensional object to be calibrated includes the step of driving each camera of the one or more cameras to capture a plurality of images of the two-dimensional object to be calibrated from a plurality of different viewpoints of the two-dimensional object to be calibrated.

[0102] In one embodiment of an intraoral scanning device, the device includes an elongated handheld wand, the elongated handheld wand comprising a probe at its distal end, one or more illumination sources coupled to the probe, one or more near-infrared (NIR) light sources coupled to the probe, and one or more cameras coupled to the probe, configured to (a) capture an image using light from the one or more illumination sources, and (b) capture an image using NIR light from the NIR light sources. The device further includes a processor configured to execute a navigation algorithm to determine the position of the elongated handheld wand as it moves in space, the input to the navigation algorithm comprising (a) an image captured using light from the one or more illumination sources, and (b) an image captured using NIR light.

[0103] In a further embodiment of the intraoral scanning device, the one or more illumination sources include one or more structured light sources.

[0104] In a further embodiment of the intraoral scanning device, the one or more illumination sources include one or more non-coherent light sources.

[0105] A method for tracking the motion of an intraoral scanner includes the steps of: illuminating a three-dimensional intraoral surface using one or more illumination sources coupled to the intraoral scanner; driving each of the one or more NIR light sources coupled to the intraoral scanner to emit NIR light onto the three-dimensional intraoral surface; and using one or more cameras coupled to the intraoral scanner, (a) capturing a first set of images using the light from the one or more illumination sources; and (b) capturing a second set of images using the NIR light. The method further includes the step of using the processor to execute a navigation algorithm to track the motion of the intraoral scanner relative to the three-dimensional intraoral surface using (a) the first set of images captured using the light from the one or more illumination sources and (b) the second set of images captured using the NIR light.

[0106] In one embodiment of the motion tracking method, the step of using one or more light sources includes the step of illuminating the three-dimensional surface inside the oral cavity.

[0107] In one embodiment of the motion tracking method, the step of using one or more illumination sources includes the step of using one or more non-coherent light sources.

[0108] One embodiment of a sixth method for calculating the three-dimensional structure of an intraoral three-dimensional surface includes the steps of: driving one or more structured light projectors to project a structured light pattern onto the intraoral three-dimensional surface; driving one or more cameras to capture a plurality of structured light images, each structured light image including at least a portion of the structured light pattern; driving one or more unstructured light projectors to project unstructured light onto the intraoral three-dimensional surface; and driving one or more cameras to capture a plurality of two-dimensional images of the intraoral three-dimensional surface. The sixth method further includes the steps of using the processor to calculate the respective three-dimensional positions of a plurality of points on the intraoral three-dimensional surface captured in the plurality of structured light images; and calculating the three-dimensional structure of the intraoral three-dimensional surface, constrained by some or all of the calculated three-dimensional positions of the plurality of points, based on the plurality of two-dimensional images of the intraoral three-dimensional surface.

[0109] In some embodiments of the sixth method, the unstructured light is uncoherent light, and the plurality of two-dimensional images include a plurality of color two-dimensional images.

[0110] In some embodiments of the sixth method, the unstructured light is near-infrared (NIR) light, and the plurality of two-dimensional images comprises a plurality of monochromatic NIR images.

[0111] In a further embodiment of the sixth method, the step of driving one or more structured light projectors includes driving one or more structured light projectors to project a distribution of discrete, unconnected light spots, respectively.

[0112] In a further embodiment of the sixth method, the step of calculating the three-dimensional structure includes inputting a plurality of two-dimensional images of the oral cavity three-dimensional surface into a neural network, and the neural network determining an estimated map of the oral cavity three-dimensional surface captured in each of the two-dimensional images.

[0113] In a further embodiment of the sixth method, the sixth method further includes the step of inputting the calculated three-dimensional positions of a plurality of points on the three-dimensional surface of the oral cavity into the neural network.

[0114] In a further embodiment of the sixth method, the sixth method further includes the step of using the processor to stitch together the respective maps to obtain a three-dimensional structure of the oral cavity three-dimensional surface.

[0115] In a further embodiment of the sixth method, the sixth method further includes coordinating the capture of the structured light image and the capture of the two-dimensional image to generate an alternating sequence in which one or more image frames of unstructured light are interspersed with one or more image frames of structured light.

[0116] In a further embodiment of the sixth method, the determination step includes determining, by the neural network, the estimated depth map of each of the three-dimensional surfaces of the oral cavity captured in each two-dimensional image. In one embodiment, the processor is used to stitch together the respective estimated depth maps to obtain the three-dimensional structure of the three-dimensional surface of the oral cavity.

[0117] In one embodiment, (a) the processor generates a point cloud corresponding to the calculated 3D position of each of the plurality of points on the 3D surface of the oral cavity captured in each structured optical image, and the method further includes the step of using the processor to concatenate the respective estimated depth maps to the respective point clouds. In one embodiment, the method further includes the step of using the neural network to determine the respective estimated normal map of the 3D surface of the oral cavity captured in each of the 2D images.

[0118] In a further embodiment of the sixth method, the determination step includes determining the respective estimated normal maps of the oral cavity 3D surfaces captured in each of the 2D images by the neural network. In one embodiment, the processor is used to stitch together the respective estimated normal maps to obtain the 3D structure of the oral cavity 3D surfaces.

[0119] In one embodiment, the method further includes the step of interpolating the three-dimensional positions on the oral cavity three-dimensional surface between the calculated three-dimensional positions of the plurality of points on the oral cavity three-dimensional surface captured in the plurality of structured light images, based on the respective estimated normal maps of the oral cavity three-dimensional surface captured in each of the two-dimensional images.

[0120] In one embodiment, the method further includes coordinating the capture of the structured light image and the capture of the two-dimensional image to generate an alternating sequence in which one or more unstructured light image frames are interspersed with one or more structured light image frames. The method further includes using the processor to (a) generate a point cloud corresponding to the calculated respective three-dimensional positions of the plurality of points on the oral cavity three-dimensional surface captured in each structured light image frame, and (b) stitching together the point clouds for at least a subset of points in each point cloud using the normal to the surface at each point in the subset of points as stitching input, wherein for a given point cloud, the normal to the surface at at least one point in the subset of points is obtained from the respective estimated normal maps of the oral cavity three-dimensional surface captured in adjacent unstructured light image frames.

[0121] In one embodiment, the method further includes compensating for the motion of the intraoral scanner between a structured optical image frame and an adjacent unstructured optical image frame by using the processor to estimate the motion of the intraoral scanner based on a previous image frame.

[0122] In a further embodiment of the sixth method, the determining step includes determining the curvature of the oral cavity 3D surface captured in each of the 2D images by the neural network. In one embodiment, the determining step includes determining the respective estimated curvature map of the oral cavity 3D surface captured in each of the 2D images by the neural network.

[0123] In one embodiment, the method further includes the steps of using the processor to evaluate the curvature of the oral cavity 3D surface captured in each of the 2D images, and interpolating the 3D position of the oral cavity 3D surface between the calculated 3D positions of the plurality of points on the oral cavity 3D surface captured in the plurality of structured light images, based on the evaluated curvature of the oral cavity 3D surface captured in each of the 2D images.

[0124] In a further embodiment of the sixth method, the sixth method further includes coordinating the capture of the structured light image and the capture of the two-dimensional image to generate an alternating sequence in which one or more unstructured light image frames are interspersed within one or more structured light image frames.

[0125] In a further embodiment of the sixth method, the sixth method includes the step of driving one or more cameras to capture a plurality of structured light images, which includes the step of driving each of two or more cameras to capture each of the plurality of structured light images, and the step of driving one or more cameras to capture a plurality of two-dimensional images includes the step of driving each of the two or more cameras to capture each of the plurality of two-dimensional images.

[0126] In one embodiment, the step of driving the two or more cameras includes driving each of the two or more cameras in a given image frame to simultaneously capture each 2D image of each portion of the three-dimensional surface of the oral cavity. The step of inputting to the neural network includes inputting all of the 2D images in a given image frame as a single input to the neural network, wherein each of the 2D images has a field of view that overlaps with at least one other 2D image from among the 2D images. The step of determining by the neural network includes determining an estimated depth map of the three-dimensional surface of the oral cavity by combining each portion of the three-dimensional surface of the oral cavity for a given image frame.

[0127] In one embodiment, the step of driving two or more cameras to capture a plurality of structured light images includes the step of driving each camera of three or more cameras to capture each plurality of structured light images, and the step of driving two or more cameras to capture a plurality of two-dimensional images includes the step of driving each camera of three or more cameras to capture each plurality of two-dimensional images. In a given image frame, each camera of the three or more cameras is driven to simultaneously capture each two-dimensional image of each portion of the three-dimensional surface of the oral cavity. The step of inputting to a neural network includes, for a given image frame, inputting a subset of each two-dimensional image as a single input to the neural network, wherein the subset includes at least two images of each two-dimensional image, and each subset of each two-dimensional image has a field of view that overlaps with at least one other subset of each two-dimensional image. The step determined by the neural network includes determining an estimated depth map of the oral cavity 3D surface by combining the respective portions of the oral cavity 3D surface captured in each subset of the 2D images for a given image frame.

[0128] In one embodiment, the step of driving the two or more cameras includes driving each of the two or more cameras in a given image frame to simultaneously capture each 2D image of each portion of the three-dimensional surface inside the oral cavity, and the step of inputting to the neural network includes inputting each of the 2D images in the given image frame as separate inputs to the neural network. The step of determining by the neural network includes determining, for a given image frame, the estimated depth map of each portion of the three-dimensional surface inside the oral cavity captured in each of the 2D images captured in the given image frame.

[0129] In one embodiment, the method further includes the step of using the processor to merge the respective depth maps to obtain a composite estimated depth map of the oral cavity 3D surface captured in a given image frame. In one embodiment, the method further includes the step of training the neural network, wherein each input to the neural network during training includes an image captured by only one camera.

[0130] In one embodiment, the method further includes the step of determining an estimated confidence map corresponding to each of the estimated depth maps by the neural network, the confidence map indicating the confidence level for each region of the estimated depth map. In a further embodiment, the step of merging the estimated depth maps includes merging at least two estimated depth maps based on the confidence levels for each of the corresponding regions, as indicated by the confidence maps for each of the at least two estimated depth maps, in response to the processor determining inconsistencies between each of the corresponding regions in at least two of the estimated depth maps.

[0131] In a further embodiment of the sixth method, the step of driving one or more cameras includes driving one or more cameras of an intraoral scanner, and the method further includes training the neural network using training stage images captured by a plurality of training stage handheld wands. Each training stage handheld wand includes one or more reference cameras, and each of the one or more cameras of the intraoral scanner corresponds to one of the reference cameras of each of the one or more reference cameras of the each training stage handheld wand.

[0132] In a further embodiment of the sixth method, the step of driving one or more structured light projectors includes driving one or more structured light projectors of the intraoral scanner, the step of driving one or more unstructured light projectors includes driving one or more unstructured light projectors of the intraoral scanner, and the step of driving one or more cameras includes driving one or more cameras of the intraoral scanner. The neural network is initially trained using training stage images captured by one or more training stage cameras of the training stage handheld wand, and each of the one or more cameras of the intraoral scanner corresponds to each of the one or more training stage cameras. The method then includes the steps of: driving (i) one or more structured light projectors of the intraoral scanner and (ii) one or more unstructured light projectors of the intraoral scanner during a plurality of refining stage scans; driving one or more cameras of the intraoral scanner during the refining stage scans to capture (a) a plurality of refining stage structured light images and (b) a plurality of refining stage 2D images; calculating the 3D structure of the intraoral 3D surface based on the plurality of refining stage structured light images; and refining the training of the neural network for the intraoral scanner using (a) the plurality of refining stage 2D images captured during the refining stage scans and (b) the calculated 3D structure of the intraoral 3D surface calculated based on the plurality of refining stage structured light images.

[0133] In one embodiment, the neural network includes a plurality of layers, and the step of refining the training of the neural network includes the step of restricting a subset of the layers.

[0134] In one embodiment, the method further includes the step of selecting from a plurality of scans which of the plurality of scans to use as the smelting stage scan, based on the quality level of each scan.

[0135] In one embodiment, the method further includes the step of using the calculated three-dimensional structure of the intraoral three-dimensional surface, calculated based on the plurality of refinement stage structured optical images during the refining stage scan, as the final resulting three-dimensional structure of the intraoral three-dimensional surface for the user of the intraoral scanner.

[0136] In a further embodiment of the sixth method, the step of driving one or more cameras includes the step of driving one or more cameras of an intraoral scanner, wherein each of the one or more cameras of the intraoral scanner corresponds to one of the one or more reference cameras. The method further includes the step of using the processor to crop and morph at least one 2D image of the intraoral 3D surface from each camera c of one or more cameras of the intraoral scanner to obtain a plurality of cropped and morphed 2D images, each cropped and morphed image corresponding to a cropped and morphed field of view of camera c, the cropped and morphed field of view of camera c matching the cropped field of view of the corresponding reference camera, the step of inputting the plurality of 2D images into the neural network includes the step of inputting the plurality of cropped and morphed 2D images of the intraoral 3D surface into the neural network, the step of determining includes the step of determining by the neural network each estimated map of the intraoral 3D surface captured in each of the cropped and morphed 2D images, the neural network being trained using training stage images corresponding to the cropped field of view of each of the one or more reference cameras.

[0137] In a further embodiment, the steps of cropping and morphing include the processor using (a) stored calibration values ​​indicating camera rays corresponding to each pixel on the camera sensor of each camera of one or more cameras, and (b) reference calibration values ​​indicating (i) camera rays corresponding to each pixel on the reference camera sensor of each reference camera of one or more reference cameras, and (ii) the cropped field of view of each reference camera of the one or more reference cameras.

[0138] In a further embodiment, the cropped field of view of each of the one or more reference cameras is 85-97% of the total field of view of each of the one or more reference cameras.

[0139] In a further embodiment, the step of using the processor further includes, for each camera c, performing reverse morphing on each estimated map of the oral cavity 3D surface captured in each of the cropped and morphed 2D images to obtain each non-morphed estimated map of the oral cavity surface as seen in each of at least one 2D image from camera c before morphing.

[0140] In a further embodiment, the unstructured light is uncoherent light, and the plurality of two-dimensional images include a plurality of two-dimensional color images.

[0141] In a further embodiment, the unstructured light is near-infrared (NIR) light, and the plurality of two-dimensional images include a plurality of monochromatic NIR images.

[0142] In a further embodiment, the step of using the processor further includes, for each camera c, performing reverse morphing on each estimated map of the oral cavity 3D surface captured in each cropped and morphed 2D image to obtain each non-morphed estimated map of the oral cavity surface as seen in each of at least one 2D image from camera c before morphing.

[0143] It should be noted that all embodiments of the sixth method described above, relating to depth maps, normal maps, curvature maps, and their use, can be implemented with necessary modifications based on field-cropped and morphed runtime images.

[0144] In one embodiment, the step of driving one or more structured light projectors includes the step of driving one or more structured light projectors of the intraoral scanner, and the step of driving one or more unstructured light projectors includes the step of driving one or more unstructured light projectors of the intraoral scanner. The method further includes the step of driving one or more structured light projectors of the intraoral scanner and one or more unstructured light projectors of the intraoral scanner during a plurality of refining stage scans, after the neural network has been trained using training stage images corresponding to cropped fields of view of each of the one or more reference cameras. One or more cameras of the intraoral scanner are driven to capture (a) a plurality of refining stage structured light images and (b) a plurality of refining stage 2D images during a plurality of refining stage structured light scans. The three-dimensional structure of the three-dimensional surface of the oral cavity is calculated based on a plurality of refinement stage structured optical images, and the training of the neural network is refined for the intraoral scanner using (a) the plurality of refinement stage two-dimensional images captured during the refinement stage scan, and (b) the calculated three-dimensional structure of the three-dimensional surface of the oral cavity calculated based on the plurality of refinement stage structured optical images. In a further embodiment, the neural network comprises a plurality of layers, and the step of refining the training of the neural network comprises the step of constraining a subset of the layers.

[0145] In one embodiment, the step of driving one or more structured light projectors includes driving one or more structured light projectors of the intraoral scanner, the step of driving one or more unstructured light projectors includes driving one or more unstructured light projectors of the intraoral scanner, and the step of determining includes determining the respective estimated depth map of the intraoral 3D surface captured in each of the cropped and morphed 2D images by the neural network. The method further includes the steps of (a) calculating the three-dimensional structure of the oral cavity three-dimensional surface based on the calculated three-dimensional positions of a plurality of points on the oral cavity three-dimensional surface captured in a plurality of structured light images; (b) calculating the three-dimensional structure of the oral cavity three-dimensional surface based on the estimated depth maps of the oral cavity three-dimensional surface captured in each of the cropped and morphed two-dimensional images; and (c) comparing the three-dimensional structure of the oral cavity three-dimensional surface calculated based on (i) the calculated three-dimensional positions of the plurality of points on the oral cavity three-dimensional surface and (ii) the three-dimensional structure of the oral cavity three-dimensional surface calculated based on the estimated depth maps of the oral cavity three-dimensional surface. In response to determining a discrepancy between (i) and (ii), the method includes the steps of: driving (A) one or more structured light projectors of the intraoral scanner and (B) one or more unstructured light projectors of the intraoral scanner during a plurality of refining stage scans; driving one or more cameras of the intraoral scanner during a plurality of refining stage scans to capture (a) a plurality of refining stage structured light images and (b) a plurality of refining stage 2D images; calculating the 3D structure of the intraoral 3D surface based on the plurality of refining stage structured light images; and refining the training of a neural network for the intraoral scanner using (a) the plurality of 2D images captured during the refining stage scans and (b) the calculated 3D structure of the intraoral 3D surface calculated based on the plurality of refining stage structured light images.In a further embodiment, the neural network comprises multiple layers, and the step of refining the training of the neural network includes the step of restricting a subset of the layers.

[0146] In a further embodiment of the sixth method, the sixth method further includes the step of training the neural network, the training of which includes the steps of: driving one or more training stage structured light projectors to project a training stage structured light pattern onto the three-dimensional surface of the training stage; driving one or more training stage cameras to capture a plurality of structured light images, each image including at least a portion of the training stage structured light pattern; driving one or more training stage unstructured light projectors to project unstructured light onto the three-dimensional surface of the training stage; driving one or more training stage cameras to capture a plurality of two-dimensional images of the three-dimensional surface of the training stage using illumination from the training stage unstructured light projectors; coordinating the capture of the structured light images and the capture of the two-dimensional images to generate an alternating sequence in which one or more image frames of the two-dimensional images are interspersed within image frames of one or more structured light images; inputting the plurality of two-dimensional images into the neural network; and having the neural network process each two-dimensional image Steps include: estimating an estimated map of the captured training stage 3D surface; inputting a plurality of 3D reconstructions of the training stage 3D surface into a neural network based on a structured optical image of the training stage 3D surface, wherein the 3D reconstruction includes the calculated 3D positions of a plurality of points on the training stage 3D surface; interpolating the position of one or more training stage cameras relative to the training stage 3D surface in each 2D image frame based on the calculated 3D positions of a plurality of points on the training stage 3D surface calculated based on the structured optical image frames before and after each 2D image frame; projecting the 3D reconstructions onto the respective fields of view of the one or more training stage cameras; calculating a true map of the training stage 3D surface as seen in each 2D image, constrained by the calculated 3D positions of the plurality of points, based on the projection; and comparing each estimated depth map of the training stage 3D surface with the corresponding true map of the training stage 3D surface, based on the difference between each estimated map and the corresponding true map.This includes a step of optimizing the neural network to better estimate the subsequent estimated map.

[0147] In a further embodiment, the training includes initial training of the neural network, the step of driving one or more structured light projectors includes driving one or more structured light projectors of the intraoral scanner, the step of driving one or more unstructured light projectors includes driving one or more unstructured light projectors of the intraoral scanner, and the step of driving one or more cameras includes driving one or more cameras of the intraoral scanner. The method further includes, after initial training of the neural network, the steps of: driving (i) one or more structured light projectors of the intraoral scanner and (ii) one or more unstructured light projectors of the intraoral scanner during a plurality of refinement-stage structured light scans; driving one or more cameras of the intraoral scanner during the refinement-stage structured light scans to capture (a) a plurality of refinement-stage structured light images and (b) a plurality of refinement-stage 2D images; calculating the 3D structure of the intraoral 3D surface based on the plurality of refinement-stage structured light images; and refining the training of the neural network using (a) the plurality of 2D images captured during the refinement-stage scans and (b) the 3D structure of the intraoral 3D surface calculated based on the plurality of refinement-stage structured light images. In a further embodiment, the neural network comprises a plurality of layers, and the step of refining the training of the neural network comprises the step of restricting a subset of the layers.

[0148] In a further embodiment of the sixth method, the step of driving one or more structured light projectors to project a training stage structured light pattern includes the step of driving one or more structured light projectors to project a distribution of discrete, unconnected light spots onto the three-dimensional surface of the training stage.

[0149] In a further embodiment of the sixth method, the step of driving one or more training stage cameras includes the step of driving at least two training stage cameras.

[0150] In a further embodiment of the sixth method, the unstructured light includes broadband spectral light.

[0151] In one embodiment of an intraoral scanning device, the device includes an elongated handheld wand, the elongated handheld wand having a probe at its distal end configured to be detachably disposed within a sleeve. The device further includes at least one structured light projector coupled to the probe, the at least one structured light projector comprising (a) a laser configured to emit polarized laser light, and (b) a pattern generating optical element configured to generate a pattern of light when the laser is activated and transmits light through the pattern generating optical element. The device further includes a camera coupled to the probe, the camera comprising a camera sensor. The probe is configured so that light enters and exits the probe through the sleeve. Furthermore, the laser is positioned at a distance from the camera such that when the probe is placed within the sleeve, a portion of the pattern of light is reflected from the sleeve and reaches the camera sensor. Furthermore, the laser is positioned at an angle of rotation relative to its own optical axis such that, due to the polarization of the pattern of light, the degree of reflection of portions of the pattern of light by the sleeve is less than threshold reflection for all possible angles of rotation of the laser with respect to its optical axis.

[0152] In a further embodiment of the intraoral scanning device, the threshold is 70% of the maximum reflection for all possible rotation angles of the laser with respect to its optical axis.

[0153] In a further embodiment of the intraoral scanning device, the laser is positioned at an angle of rotation relative to its own optical axis such that the degree of reflection by the sleeve of a portion of the light pattern is less than 60% of the maximum reflection for all possible angles of rotation of the laser relative to its optical axis, due to the polarization of the light pattern.

[0154] In a further embodiment of the intraoral scanning device, the laser is positioned at an angle of rotation relative to its own optical axis such that, due to the polarization of the light pattern, the degree of reflection of portions of the light pattern by the sleeve is 15% to 60% of the maximum reflection for all possible rotation angles of the laser with respect to its optical axis.

[0155] In a further embodiment of the intraoral scanning device, when the elongated handheld wand is positioned within the sleeve, the distance between the structured light projector and the camera is 1 to 6 times the distance between the structured light projector and the sleeve.

[0156] In a further embodiment of the intraoral scanning device, the at least one structured light projector has an illumination field of at least 30 degrees, and the camera has a field of view of at least 30 degrees.

[0157] A seventh method for generating a three-dimensional image using an intraoral scanner, the seventh method includes capturing a plurality of images of the intraoral three-dimensional surface using at least two cameras rigidly connected to the intraoral scanner such that the fields of view of each camera have non-overlapping portions. The seventh method further includes using the processor to perform a simultaneous localization and mapping (SLAM) algorithm using the captured images from each camera for the non-overlapping portions of each field of view, wherein the localization of each camera is solved on the basis that the motion of each camera is the same as the motion of all the other cameras.

[0158] In a further embodiment of the seventh method, the fields of view of the first camera and the second camera also have overlapping portions. Furthermore, the capturing step includes capturing multiple images of the intraoral 3D surface such that features of the intraoral 3D surface in the overlapping portion of each field of view appear in the images captured by the first and second cameras. Furthermore, the processor-using step includes executing a SLAM algorithm using the intraoral 3D surface features appearing in the images of at least two cameras.

[0159] An eighth method for generating a three-dimensional image using an intraoral scanner, the eighth method comprising: driving one or more structured light projectors to project a pattern of structured light onto an intraoral three-dimensional surface; driving one or more cameras to capture a plurality of structured light images, each structured light image including at least a portion of the structured light pattern; driving one or more unstructured light projectors to project unstructured light onto an intraoral three-dimensional surface; driving at least one camera to capture a two-dimensional image of the intraoral three-dimensional surface using illumination from the unstructured light projectors; and coordinating the capture of structured light and unstructured light to generate an alternating sequence in which one or more structured light image frames are interspersed with one or more unstructured light image frames. The eighth method further comprises using the processor to calculate the three-dimensional position of each of a plurality of points on the intraoral three-dimensional surface captured in one or more image frames of structured light. The eighth method further includes the step of using the processor to interpolate the motion of at least one camera between a first unstructured light image frame and a second unstructured light image frame based on the calculated three-dimensional positions of a plurality of points in each structured light image frame before and after the unstructured light image frame. The eighth method further includes the steps of (a) using the features of the intraoral three-dimensional surface captured in the first and second unstructured light image frames by the at least one camera, and (b) executing a simultaneous localization and mapping (SLAM) algorithm constrained by the interpolated motion of the camera between the first unstructured light image frame and the second unstructured light image frame.

[0160] In one embodiment of the eighth method, the step of driving one or more structured light projectors to project the structured light pattern includes the step of driving one or more structured light projectors to project a distribution of discrete, unconnected light spots onto the three-dimensional surface of the oral cavity. In one embodiment, the unstructured light includes broadband spectral light, and the two-dimensional image includes a two-dimensional color image. In one embodiment, the unstructured light includes near-infrared (NIR) light, and the two-dimensional image includes a two-dimensional monochromatic NIR image.

[0161] In one embodiment of a ninth method for generating a three-dimensional image using an intraoral scanner, the ninth method includes the steps of: driving one or more structured light projectors to project a pattern of structured light onto an intraoral three-dimensional surface; driving one or more cameras to capture a plurality of structured light images, each of which structured light images includes at least a portion of the structured light pattern; driving one or more unstructured light projectors to project unstructured light onto an intraoral three-dimensional surface; driving one or more cameras to capture a two-dimensional image of the intraoral three-dimensional surface using illumination from the unstructured light projectors; and coordinating the capture of structured light and unstructured light to generate an alternating sequence in which one or more unstructured light image frames are interspersed within one or more structured light image frames. The ninth method further includes the steps of using the processor to calculate the three-dimensional position of a feature on a three-dimensional intraoral surface based on a structured light image frame, the feature also captured in a first unstructured light image frame and a second unstructured light image frame; further including the steps of calculating the motion of one or more cameras between the first unstructured light image frame and the second unstructured light image frame based on the calculated three-dimensional position of the feature; and executing a simultaneous localization and mapping (SLAM) algorithm using (i) the feature on the three-dimensional intraoral surface that was captured in the first and second unstructured light image frames by the one or more cameras but whose three-dimensional position was not calculated based on the structured light image frame; and (ii) the calculated motion of the cameras between the first and second unstructured light image frames. In one embodiment, the step of driving one or more structured light projectors to project the structured light pattern includes the step of driving one or more structured light projectors to project a distribution of discrete, unconnected light spots onto the three-dimensional surface of the oral cavity. In one embodiment, the unstructured light includes broadband spectral light, and the two-dimensional image includes a two-dimensional color image. In one embodiment, the unstructured light includes near-infrared (NIR) light, and the two-dimensional image includes a two-dimensional monochromatic NIR image.

[0162] A method for calculating the three-dimensional structure of a three-dimensional surface within the oral cavity of a subject, comprising the steps of: (a) driving one or more structured light projectors to project a pattern of structured light onto the three-dimensional surface within the oral cavity, the pattern comprising multiple features; (b) driving one or more cameras to capture multiple structured light images, each structured light image comprising at least one feature of the structured light pattern; (c) driving one or more unstructured light projectors to project unstructured light onto the three-dimensional surface within the oral cavity; (d) driving at least one camera to capture a two-dimensional image of the three-dimensional surface within the oral cavity using illumination from one or more unstructured light projectors; and (e) coordinating the capture of structured light and unstructured light to generate an alternating sequence in which one or more unstructured light image frames are interspersed within one or more structured light image frames. The method further includes using a processor to (a) determine, based on a two-dimensional image, whether one or more features of a plurality of features of a structured light pattern were projected onto moving tissue or stable tissue in the oral cavity; (b) assign a confidence grade to each of the one or more features based on the determination, assigning high confidence to fixed tissue and low confidence to moving tissue; and (c) run a three-dimensional reconstruction algorithm using one or more features based on the confidence grade for each of the one or more features. In one embodiment, the unstructured light includes broadband spectral light, and the two-dimensional image is a two-dimensional color image. In one embodiment, the unstructured light includes near-infrared (NIR) light, and the two-dimensional image is a two-dimensional monochromatic NIR image. In one embodiment, the plurality of features include a plurality of spots, and the step of driving one or more structured light projectors to project the structured light pattern includes driving one or more structured light projectors to project a distribution of discrete, unconnected light spots onto the three-dimensional surface of the oral cavity. In one embodiment, the step of executing the three-dimensional reconstruction algorithm is performed using only a subset of features, the subset consisting of features that have been assigned a confidence grade above a fixed tissue threshold.In one embodiment, the step of executing the three-dimensional reconstruction algorithm includes (a) assigning a weight to each feature based on the confidence grade assigned to that feature, and (b) using the respective weights of each feature in the three-dimensional reconstruction algorithm.

[0163] In one embodiment of a tenth method for calculating the three-dimensional structure of an intraoral three-dimensional surface, the tenth method includes the steps of driving one or more light sources of an intraoral scanner to project light onto the intraoral three-dimensional surface, and driving two or more cameras of the intraoral scanner to capture a plurality of two-dimensional images of the intraoral three-dimensional surface, each of which corresponds to one of two or more reference cameras. The method includes the step of using a processor to modify at least one of the two-dimensional images from each of the two or more cameras c of the intraoral scanner to obtain a plurality of modified two-dimensional images, each of which corresponds to the modified field of view of camera c, and the modified field of view of camera c matches the modified field of view of the corresponding camera among the reference cameras, and further includes the step of calculating the three-dimensional structure of the intraoral three-dimensional surface based on the plurality of modified two-dimensional images of the intraoral three-dimensional surface.

[0164] In one embodiment of the 11th method for calculating the three-dimensional structure of a three-dimensional surface in the oral cavity, the 11th method includes the steps of driving one or more light sources of an intraoral scanner to project light onto the three-dimensional surface in the oral cavity, and driving two or more cameras of the intraoral scanner to capture a plurality of two-dimensional images of the three-dimensional surface in the oral cavity, wherein each of the one or more cameras of the intraoral scanner corresponds to one of the two or more reference cameras. The method includes the step of using a processor to crop and morph at least one of the two-dimensional images from each of the two or more cameras c of the intraoral scanner to obtain a plurality of cropped and morphed two-dimensional images, wherein each cropped and morphed image corresponds to the cropped and morphed field of view of camera c, and the cropped and morphed field of view of camera c coincides with the cropped field of view of the corresponding camera among the reference cameras. The three-dimensional structure of the oral cavity's three-dimensional surface is calculated by inputting multiple cropped and morphed two-dimensional images of the oral cavity's three-dimensional surface into a neural network, and the neural network determines the respective estimated map of the oral cavity's three-dimensional surface captured in each of the multiple cropped and morphed two-dimensional images, the neural network being trained using training stage images corresponding to the cropped field of view of each of the one or more reference cameras.

[0165] In a further embodiment of the 11th method, the light is non-coherent light, and the plurality of two-dimensional images include a plurality of two-dimensional color images.

[0166] In a further embodiment of the 11th method, the light is near-infrared (NIR) light, and the plurality of two-dimensional images include a plurality of monochromatic NIR images.

[0167] In a further embodiment of the 11th method, the light is broadband spectral light, and the plurality of two-dimensional images include a plurality of two-dimensional color images.

[0168] In a further embodiment of the 11th method, the cropping and morphing step includes the step of the processor using (a) stored calibration values ​​indicating camera rays corresponding to each pixel on the camera sensor of each of the one or more cameras c, and (b) reference calibration values ​​indicating (i) camera rays corresponding to each pixel on the reference camera sensor of each of the one or more reference cameras, and (ii) the cropped field of view of each of the one or more reference cameras.

[0169] In a further embodiment of the 11th method, the cropped field of view of each of the one or more reference cameras is 85-97% of the total field of view of each of the one or more reference cameras.

[0170] In a further embodiment, the step of using the processor further includes, for each camera c, performing reverse morphing on each estimated map of the oral cavity 3D surface captured in each cropped and morphed 2D image to obtain each non-morphed estimated map of the oral cavity surface as seen in each of at least one 2D image from camera c before morphing.

[0171] It should be noted that all the above embodiments of the sixth method relating to depth maps, normal maps, curvature maps, and their use may be performed with necessary modifications from the eleventh method based on field-cropped and morphed runtime images.

[0172] It should also be noted that all the above embodiments of the sixth method relating to structured light may be performed with necessary modifications in the context of the eleventh method and a cropped and morphed runtime two-dimensional image.

[0173] In one embodiment of a 12th method for calculating the three-dimensional structure of a three-dimensional surface in the oral cavity, the 12th method includes the steps of driving one or more optical projectors to project light onto the three-dimensional surface in the oral cavity, and driving one or more cameras to capture a plurality of two-dimensional images of the three-dimensional surface in the oral cavity. The method includes the steps of using a processor to input the plurality of two-dimensional images of the three-dimensional surface in the oral cavity into a first neural network module and a second neural network module, the first neural network module to determine the respective estimated depth map of the three-dimensional surface in the oral cavity captured in each of the two-dimensional images, and the second neural network module to determine the respective estimated confidence map corresponding to the respective estimated depth map, the respective confidence map indicating the confidence level for each region of the respective estimated depth map.

[0174] In one embodiment of the twelfth method, the first neural network module and the second neural network module are separate modules of the same neural network.

[0175] In one embodiment of the twelfth method, the first and second neural network modules are not separate modules of the same neural network.

[0176] In a further embodiment of the 12th method, the method further includes the step of training a second neural network module to determine each estimated confidence map corresponding to each estimated depth map determined by the first neural network module, the determining step of initial training the first neural network module to determine each estimated depth map using a plurality of depth training stage 2D images, then (i) input a plurality of confidence training stage 2D images of the training stage 3D surface into the first neural network module, and (ii) the first neural network module determines the estimated depth of each training stage 3D surface captured in each confidence training stage 2D image The process is carried out in the following steps: (iii) determine the map, calculate the difference between each estimated depth map and the corresponding true depth map to obtain the respective target confidence map corresponding to each estimated depth map determined by the first neural network module, (iv) input multiple confidence training stage 2D images into the second neural network module, (v) have the second neural network module estimate each estimated confidence map showing the confidence level for each region of each estimated depth map, and (vi) compare each estimated confidence map with the corresponding target confidence map, and based on this comparison, optimize the second neural network module to better estimate subsequent estimated confidence maps.

[0177] In one embodiment, the plurality of confidence training stage 2D images are not the same as the plurality of depth training stage 2D images.

[0178] In one embodiment, the plurality of confidence training stage 2D images are the same as the plurality of depth training stage 2D images.

[0179] In a further embodiment of the 12th method, (a) the step of driving one or more cameras to capture a plurality of two-dimensional images includes, in a given image frame, driving each of the two or more cameras to simultaneously capture each two-dimensional image of each portion of the three-dimensional surface of the oral cavity, and (b) the step of inputting the plurality of two-dimensional images of the three-dimensional surface of the oral cavity into a first neural network module and a second neural network module, wherein, for a given image frame, each of the two-dimensional images is input to the first neural network module and the second neural network module separately The method includes the step of inputting as input, (c) a step determined by a first neural network module, which includes determining the respective estimated depth map of each portion of the oral cavity 3D surface captured in each of the 2D images captured in the given image frame for a given image frame, and (d) a step determined by a second neural network module, which includes determining the respective estimated confidence map corresponding to the respective estimated depth map of each portion of the oral cavity 3D surface captured in each of the 2D images captured in the given image frame for a given image frame. The method further includes the step of using the processor to merge the respective estimated depth maps to obtain a composite estimated depth map of the oral cavity 3D surface captured in the given image frame. In response to determining inconsistencies between corresponding regions in at least two of the estimated depth maps, the processor merges the at least two estimated depth maps based on the confidence levels of the corresponding regions, as indicated by the respective confidence maps for each of the at least two estimated depth maps.

[0180] In one embodiment of a thirteenth method for calculating the three-dimensional structure of an intraoral three-dimensional surface, the thirteenth method includes the steps of: driving one or more light sources of an intraoral scanner to project light onto an intraoral three-dimensional surface; and driving one or more cameras of the intraoral scanner to capture a plurality of two-dimensional images of the intraoral three-dimensional surface. The method further includes (a) using a processor to determine, by a neural network, an estimated map of the intraoral three-dimensional surface captured in each of the two-dimensional images; and (b) using a processor to overcome manufacturing deviations of one or more cameras of the intraoral scanner and reduce the difference between the estimated map and the true structure of the intraoral three-dimensional surface.

[0181] In one embodiment of the 13th method, the step of overcoming manufacturing deviations of one or more cameras includes the step of overcoming manufacturing deviations of one or more cameras from a reference set of one or more cameras.

[0182] In one embodiment of the 13th method, the intraoral scanner is one of a plurality of manufactured intraoral scanners, each manufactured intraoral scanner includes one or more sets of cameras, and the step of overcoming manufacturing deviations of one or more cameras of the intraoral scanner includes overcoming manufacturing deviations of one or more cameras from one or more sets of cameras of at least one other of the plurality of manufactured intraoral scanners.

[0183] In a further embodiment of the 13th method, the step of driving one or more cameras includes driving two or more cameras of the intraoral scanner to capture a plurality of two-dimensional images of the intraoral three-dimensional surface, each of which two or more cameras of the intraoral scanner corresponds to one reference camera of two or more reference cameras, and the neural network is trained using training stage images captured by the two or more reference cameras. The step of overcoming manufacturing deviation includes overcoming manufacturing deviation of the two or more cameras of the intraoral scanner, which is done by using the processor to (a) modify at least one of the two-dimensional images from each camera c of the two or more cameras of the intraoral scanner to obtain a plurality of modified two-dimensional images, each modified image corresponding to a modified field of view of camera c, and the modified field of view of camera c coincides with the modified field of view of the corresponding reference camera of the reference cameras, and (b) the neural network determines an estimated map of each of the intraoral three-dimensional surfaces based on the plurality of modified two-dimensional images of the intraoral three-dimensional surface.

[0184] In a further embodiment, the light is non-coherent light, and the plurality of two-dimensional images include a plurality of two-dimensional color images.

[0185] In a further embodiment, the light is near-infrared (NIR) light, and the plurality of two-dimensional images include a plurality of monochromatic NIR images.

[0186] In a further embodiment, the light is broadband spectral light, and the plurality of two-dimensional images include a plurality of two-dimensional color images.

[0187] In a further embodiment, the modification step includes cropping and morphing at least one of the two-dimensional images from camera c to obtain a plurality of cropped and morphed two-dimensional images, each cropped and morphed image corresponding to a cropped and morphed field of view of camera c, and the cropped and morphed field of view of camera c coincides with the cropped field of view of a corresponding reference camera of the reference camera.

[0188] In a further embodiment, the steps of cropping and morphing include the processor using (a) stored calibration values ​​indicating camera rays corresponding to each pixel on the camera sensor of each of the one or more cameras c, and (b) reference calibration values ​​indicating (i) camera rays corresponding to each pixel on the reference camera sensor of each of the one or more reference cameras, and (ii) the cropped field of view of each of the one or more reference cameras.

[0189] In a further embodiment, the cropped field of view of each of the one or more reference cameras is 85-97% of the total field of view of each of the one or more reference cameras.

[0190] In a further embodiment, the processor further includes performing reverse morphing on each estimated map of the oral cavity 3D surface captured in each cropped and morphed 2D image for each camera c to obtain each non-morphed estimated map of the oral cavity surface as seen in each of at least one 2D image from camera c before morphing.

[0191] In a further embodiment of the 13th method, the step of overcoming manufacturing deviations of the one or more cameras of the intraoral scanner includes the step of training the neural network using training stage images captured by a plurality of training stage intraoral scanners. Each training stage intraoral scanner includes one or more reference cameras, each of the one or more cameras of the intraoral scanner corresponds to one of the one or more reference cameras on each training stage intraoral scanner, and the manufacturing deviation of one or more cameras is the manufacturing deviation of one or more cameras from the corresponding one or more reference cameras.

[0192] In a further embodiment of the 13th method, the step of driving one or more cameras includes driving two or more cameras of the intraoral scanner to capture a plurality of two-dimensional images of the three-dimensional surface of the oral cavity, respectively. The step of overcoming manufacturing deviations includes overcoming manufacturing deviations between two or more cameras of the intraoral scanner, which is done by: training the neural network using training stage images each captured by only one camera; driving two or more cameras of the intraoral scanner in a given image frame to simultaneously capture two-dimensional images of each portion of the three-dimensional surface of the oral cavity; inputting each of the two-dimensional images as separate inputs to the neural network for a given image frame; the neural network determining each estimated depth map of each portion of the three-dimensional surface of the oral cavity captured in each of the two-dimensional images captured in the given image frame; and using the processor to merge the respective estimated depth maps to obtain a composite estimated depth map of the three-dimensional surface of the oral cavity captured in the given image frame.

[0193] In a further embodiment, the determination step further includes determining an estimated confidence map corresponding to each estimated depth map by the neural network, where each confidence map represents the confidence level for each region of the respective estimated depth map.

[0194] In a further embodiment, the step of merging the respective estimated depth maps includes, in response to the processor determining inconsistencies between the respective corresponding regions in at least two of the estimated depth maps, merging the at least two estimated depth maps based on the confidence levels of the respective corresponding regions as indicated by the respective confidence maps for each of the at least two estimated depth maps.

[0195] In a further embodiment of the 13th method, the step of overcoming manufacturing deviations of one or more cameras of the intraoral scanner includes (a) initially training the neural network using training stage images captured by one or more training stage cameras of one or more training stage handheld wands, each of which camera of the intraoral scanner corresponds to each of the one or more training stage cameras of each of the one or more training stage handheld wands; and (b) subsequently driving the intraoral scanner to perform a plurality of refining stage scans of the intraoral three-dimensional surface and refining the training of the neural network for the intraoral scanner using the refining stage scans of the intraoral three-dimensional surface.

[0196] In a further embodiment, the neural network comprises multiple layers, and the step of refining the training of the neural network includes the step of restricting a subset of the layers.

[0197] In a further embodiment, the method further includes the step of selecting from a plurality of scans which of the plurality of scans to use as the smelting stage scan, based on the quality level of each scan.

[0198] In a further embodiment, the step of driving the intraoral scanner to perform a plurality of refining stage scans includes, during the plurality of refining stage scans, (i) driving one or more structured light projectors of the intraoral scanner to project a pattern of structured light onto the intraoral three-dimensional surface; (ii) driving one or more unstructured light projectors of the intraoral scanner to project unstructured light onto the intraoral three-dimensional surface; driving one or more cameras of the intraoral scanner to capture, during the refining stage scan, (a) a plurality of refining stage structured light images using illumination from the structured light projectors and (b) a plurality of refining stage two-dimensional images using illumination from the unstructured light projectors; and calculating the three-dimensional structure of the intraoral three-dimensional surface based on the plurality of refining stage structured light images.

[0199] In a further embodiment, the step of refining the training of the neural network includes refining the training of the neural network for the intraoral scanner using (a) a plurality of two-dimensional refining stage images captured during a refining stage scan, and (b) a calculated three-dimensional structure of the intraoral three-dimensional surface calculated based on the plurality of refining stage structured optical images.

[0200] In a further embodiment, the method further includes using the calculated three-dimensional structure of the intraoral three-dimensional surface, calculated based on the plurality of refinement stage structured optical images during the refinement stage scan, as the final resulting three-dimensional structure of the intraoral three-dimensional surface for the user of the intraoral scanner.

[0201] In one embodiment of a 14th method for training a neural network for use with an intraoral scanner, the 14th method includes the steps of: inputting a plurality of 2D images of a three-dimensional intraoral surface into a neural network; the neural network estimating an estimated map of the three-dimensional intraoral surface captured in each of the 2D images; calculating a true map of the three-dimensional intraoral surface as seen in each of the 2D images based on a plurality of structured light images of the three-dimensional intraoral surface; comparing each estimated map of the three-dimensional intraoral surface with a corresponding true map of the three-dimensional intraoral surface; optimizing the neural network based on the difference between each estimated map and the corresponding true map to better estimate subsequent estimated maps, and, for 2D images in which moving tissue has been identified, processing the images to exclude at least a portion of the moving tissue before inputting the 2D images into the neural network.

[0202] In a further embodiment of the 14th method, the method includes the steps of: driving one or more structured light projectors to project a structured light pattern onto a three-dimensional intraoral surface; driving one or more cameras to capture a plurality of structured light images, each image comprising at least a portion of the structured light pattern; driving one or more unstructured light projectors to project unstructured light onto a three-dimensional intraoral surface; driving one or more cameras to capture a plurality of two-dimensional images of the three-dimensional intraoral surface using illumination from the unstructured light projectors; and coordinating the capture of the structured light images and the capture of the two-dimensional images to generate an alternating sequence in which one or more image frames of two-dimensional images are interspersed within image frames of one or more structured light images. Furthermore, the step of calculating a true map of the oral cavity 3D surface seen in each of the 2D images includes: inputting a plurality of 3D reconstructions of the oral cavity 3D surface into the neural network based on a structured optical image of the oral cavity 3D surface, wherein the 3D reconstruction includes the calculated 3D positions of a plurality of points on the oral cavity 3D surface; interpolating the position of one or more cameras relative to the oral cavity 3D surface for each 2D image frame based on the calculated 3D positions of a plurality of points on the oral cavity 3D surface calculated based on the structured optical image frames before and after each 2D image frame; and projecting the 3D reconstructions onto the respective fields of view of the one or more cameras, and calculating a true map of the oral cavity 3D surface seen in each 2D image, constrained by the calculated 3D positions of the plurality of points, based on the projection.

[0203] Furthermore, several applications of the present invention provide a method for generating digital 3D images, and the method is The steps include driving each of one or more structured light projectors to project a light pattern (e.g., a distribution of discrete, unconnected light spots) onto the three-dimensional surface of the oral cavity, Steps include: driving each camera of one or more cameras to capture multiple images, each image including at least a portion of the projection pattern, and each camera of the one or more cameras including a camera sensor including an array of pixels; The process includes the step of using a processor to compare multiple consecutive images captured by each camera and determining a portion of the captured projection pattern (e.g., projection spots s) that can be tracked across multiple images.

[0204] In some applications, the projection pattern is a distribution of unconnected light spots, and the processor may make a determination based on stored calibration values ​​indicating (a) camera rays corresponding to each pixel on the camera sensor of each of the one or more cameras, and (b) projector rays corresponding to each projection light spot of each of the one or more projectors. In some embodiments, each projector ray corresponds to the path of each pixel on at least one of the camera sensors. Furthermore, in some embodiments, the processor may determine which projection spots s can be tracked across multiple images, with each tracked spot s moving along the path of the pixel corresponding to each projector ray r.

[0205] In some applications, the step of using the processor further includes using the processor to calculate the respective three-dimensional position on the intraoral three-dimensional surface at the intersection of the projector ray r corresponding to the tracked spot s in each of a plurality of consecutive images in which the spot s is tracked, and the respective camera ray.

[0206] In some application examples, the step of using the processor further involves using the processor, (a) A step of determining the parameters of a tracking spot in at least two adjacent images from consecutive images, wherein the parameters include one or more of the size of the spot, the shape of the spot, the orientation of the spot, the intensity of the spot, and the signal-to-noise ratio (SNR) of the spot, (b) The step of predicting the parameters of the tracking spot in a later image based on the parameters of the tracking spot in at least two adjacent images.

[0207] In some applications, the step of using the processor further includes using the processor to search for spots substantially having the predicted parameters of the tracked spots in subsequent images, based on the predicted parameters of the tracked spots.

[0208] In some applications, the selected parameter is the shape of the spot, and the step of using the processor further includes using the processor to determine the search space in the next image for searching the tracking spot, based on the predicted shape of the tracking spot.

[0209] In some applications, the step of determining the search space using the processor includes the step of determining the search space in the next image in which the tracking spot is searched using the processor, wherein the search space has a size and aspect ratio based on the size and / or aspect ratio of the predicted shape of the tracking spot.

[0210] In some application examples, the selected parameter is the shape of the spot, and the step of using the processor further involves using the processor, (a) A step of determining the velocity vector of the tracking spot based on the direction and distance the tracking spot has moved between two adjacent images from consecutive images, (b) A step of predicting the shape of the tracking spot in a later image in response to the shape of the tracking spot in at least one of the two adjacent images, (c) The step of determining the search space in a subsequent image in which the tracking spot is searched, in response to the combination of (i) determining the velocity vector of the tracking spot and (ii) the predicted shape of the tracking spot.

[0211] In some application examples, the selected parameter is the shape of the spot, and the step of using the processor is to use the processor, (a) A step of determining the velocity vector of the tracking spot based on the direction and distance the tracking spot has moved between two adjacent images from the consecutive images, (b) A step of predicting the shape of the tracking spot in a later image in response to the determination of the velocity vector of the tracking spot, (c) The step of determining a search space in a subsequent image for searching for the tracking spot in response to a combination of (i) determining the velocity vector of the tracking spot and (ii) the predicted shape of the tracking spot.

[0212] In some applications, the step of using the processor includes using the processor to predict the shape of the tracking spot in a later image in response to a combination of (i) determining the velocity vector of the tracking spot and (ii) the shape of the tracking spot in at least one of two adjacent images.

[0213] In some application examples, the step of using the processor is to use the processor, (a) A step of determining the velocity vector of the tracking spot based on the direction and distance the tracking spot moved between two consecutive images, (b) The step of determining the search space in a subsequent image in which the tracking spot is searched, in response to determining the velocity vector of the tracking spot.

[0214] In some applications, the step of using the processor further includes determining, for each tracking spot s, a plurality of possible paths p of a pixel on one of the cameras, where each path p corresponds to a plurality of possible projector rays r.

[0215] In some application examples, the step of using the processor further involves using the processor to execute the corresponding algorithm. (a) For each possible projector ray r, Identify how many other cameras have detected each spot q corresponding to each camera beam that intersects with the projector beam r and the camera beam of a given camera corresponding to the tracking spot s, on the path p1 of each of their pixels corresponding to the projector beam r. (b) Identify a given projector ray r1 in which the most other cameras detected their respective spots q, (c) Including identifying the projector ray r1 as the specific projector ray r that generated the tracking spot s.

[0216] In some application examples, the step of using the processor is to use the processor, (a) A step of executing a corresponding algorithm to calculate the respective three-dimensional positions of multiple detection spots on the oral cavity three-dimensional surface captured in multiple consecutive images, (b) The step of identifying the detected spot as originating from the particular projector ray r by identifying the detected spot as a tracking spot s moving along the path of a pixel corresponding to the particular projector ray r in at least one of the plurality of consecutive images.

[0217] In some application examples, the step of using the processor further involves using the processor, (a) A step of executing a corresponding algorithm to calculate the respective three-dimensional positions of multiple detection spots on the oral cavity three-dimensional surface captured in multiple consecutive images, (b) The step of (i) excluding from consideration as points on the oral surface any spot that is identified as originating from a specific projector ray r based on its three-dimensional position calculated by the corresponding algorithm, and (ii) is not identified as a tracking spot s moving along the path of a pixel corresponding to the specific projector ray r.

[0218] In some application examples, the step of using the processor further involves using the processor, (a) A step of executing a corresponding algorithm to calculate the respective three-dimensional positions of multiple detection spots on the oral cavity three-dimensional surface captured in multiple consecutive images, (b) With respect to a detection spot that has been identified as originating from two different projector rays r based on the three-dimensional position calculated by the corresponding algorithm, the step of identifying the detection spot as a tracking spot s that moves along one of the two different projector rays r is included.

[0219] In some application examples, the step of using the processor further involves using the processor, (a) Execute the corresponding algorithm to calculate the three-dimensional position of each of the multiple detection spots on the three-dimensional surface of the oral cavity captured in multiple consecutive images, (b) Identifying a vulnerable spot whose three-dimensional position was not calculated by the corresponding algorithm as a projection spot from a specific projector ray r, thereby identifying the vulnerable spot as a tracking spot s that moves along the path of a pixel corresponding to a specific projector ray r.

[0220] A method for generating digital three-dimensional images is further provided according to several application examples of the present invention, and the method is The steps include driving each of one or more structured light projectors to project a light pattern (e.g., a distribution of discrete, unconnected light spots) onto the three-dimensional surface of the oral cavity, The steps include driving each camera of one or more cameras to capture an image, the image including at least a portion of the projection pattern, and each camera of one or more cameras including a camera sensor including an array of pixels, Using the processor, (a) A step of executing a corresponding algorithm to calculate the three-dimensional position of each portion of the detected pattern on the three-dimensional surface of the oral cavity that is captured in multiple consecutive images, (b) The steps of determining the calculated three-dimensional position of a portion of a detected pattern in at least a subset of multiple consecutive images, as corresponding to a specific projector ray r, (c) The step of calculating the length of the projector ray r in each image of the subset of the image based on the three-dimensional position of the detection portion of the pattern corresponding to the projector ray r in the subset of the image.

[0221] In some embodiments, the light pattern may be a distribution of unconnected spots. In some embodiments, the processor may perform steps (a) to (c) based on stored calibration values ​​indicating (i) camera rays corresponding to each pixel on the camera sensor of each of the one or more cameras, and (ii) projector rays corresponding to each one projected light spot from each of the one or more projectors. In some embodiments, each projector ray corresponds to the path of each pixel on at least one of the camera sensors.

[0222] In some applications, the step of using the processor further includes the step of using the processor to calculate an estimated length of the projector ray r in at least one of a plurality of consecutive images in which the three-dimensional position of the projection spot from the projector ray r was not determined in step (b).

[0223] In some applications, the step of using the processor further includes determining a one-dimensional search space in at least one of the images, which is used to search for a projection spot from the projector ray r based on the estimated length of the projector ray r in at least one of the images, wherein the one-dimensional search space is along the path of each pixel corresponding to the projector ray r.

[0224] In some applications, the step of using the processor further includes determining a one-dimensional search space in each pixel array of multiple cameras, which is used to search for a projection spot from a projector ray r based on the estimated length of the projector ray r in at least one of the multiple images, wherein the one-dimensional search space is along the path of each pixel corresponding to the ray r.

[0225] In some applications, the step of using a processor to determine a one-dimensional search space in each pixel array of multiple cameras includes the step of using a processor to determine a one-dimensional search space for searching for projection spots from projector rays r in each pixel array of all cameras.

[0226] In some applications, the step of using the processor further includes the step of calculating the estimated length of the projector ray r in at least one of a plurality of sequential images in which a plurality of candidate three-dimensional positions of the projection spot from the projector ray r are identified in step (b).

[0227] In some applications, the step of using the processor further includes determining whether the projection spot is in the correct three-dimensional position by determining which of a plurality of candidate three-dimensional positions corresponds to the estimated length of the projector ray r in at least one of the plurality of images.

[0228] In some application examples, the step of using the processor involves using the processor to estimate the length of the projector ray r in at least one of the multiple images, (a) A step of determining a one-dimensional search space in at least one of multiple images to search for a projection spot from a projector beam r, (b) The step of determining which of the plurality of candidate three-dimensional positions corresponds to the spot generated by the projector ray r detected in the one-dimensional search space, thereby determining which of the plurality of candidate three-dimensional positions of the projection spot is the correct three-dimensional position of the projection spot generated by the projector ray r.

[0229] In some application examples, the step of using the processor further involves using the processor, (i) A step of defining a curve based on the length of the projector ray r in each image of the subset of the image, (ii) If the three-dimensional position of the projection spot corresponds to the length of the projector ray r, which is at least a threshold distance from the defined curve, the detection spot identified in step (b) as originating from the projector ray r is excluded from consideration as a point on the oral surface.

[0230] A method for generating a digital three-dimensional image is further provided according to some applications of the present invention, and the method is The steps include driving each of the structured light projectors of one or more structured light projectors to project a distribution of discrete, unconnected light spots onto the three-dimensional surface of the oral cavity, Steps include: driving each camera of one or more cameras to capture an image, the image including at least one of the spots, and each camera of the one or more cameras including a camera sensor including an array of pixels; Based on stored calibration values ​​indicating (a) camera rays corresponding to each pixel on the camera sensor of each of the one or more cameras, and (b) projector rays corresponding to each projection light spot from each of the one or more projectors, wherein each projector ray corresponds to a projector ray corresponding to each path on at least one pixel of the camera sensor, Using the processor, (a) A step of executing a corresponding algorithm to calculate the respective three-dimensional positions of multiple projection spots on the three-dimensional surface of the oral cavity, (b) Using data from at least two of the cameras, a candidate three-dimensional position of a given spot corresponding to a specific projector ray r is identified, and substantially no data from another camera is used to identify the candidate three-dimensional position; (c) Using candidate three-dimensional positions seen by at least one of the two cameras, the step of identifying a search space on another one of the camera's pixel arrays for searching for a spot from the projector ray r, (d) If a spot from the projector ray r is identified in the search space, the step of using data from another camera to refine a candidate for the three-dimensional position of the spot.

[0231] A method for generating a digital three-dimensional image is further provided according to some applications of the present invention, and the method is The steps include driving each of the structured light projectors of one or more structured light projectors to project a distribution of discrete, unconnected light spots onto the three-dimensional surface of the oral cavity, Steps include: driving each camera of one or more cameras to capture multiple images, each image containing at least one spot, and each camera of one or more cameras containing a camera sensor including a pixel array, Based on stored calibration values ​​indicating (a) camera rays corresponding to each pixel on the camera sensor of each of the one or more cameras, and (b) projector rays corresponding to each projection light spot from each of the one or more projectors, wherein each projector ray corresponds to a projector ray corresponding to the path of each pixel on at least one of the camera sensors of the camera sensor, Using the processor, (a) A step of executing a corresponding algorithm to calculate the respective 3D positions of multiple detection spots on the 3D surface of the oral cavity for each of the multiple images, (b) Using data corresponding to the respective three-dimensional positions of at least three spots, each spot corresponding to a respective projector ray r, estimate the three-dimensional surface on which all at least three spots are located. (c) For projector rays r1 whose 3D position corresponding to the spot was not calculated in step (a), the step of estimating the 3D position in the intersection space between the projector ray r1 and the estimated surface, (d) The step of identifying a search space in the pixel array of at least one camera, using a three-dimensional position in the estimated space to search for a spot corresponding to a projector ray r1.

[0232] In some applications, the step of using data corresponding to the three-dimensional position of each of the at least three spots includes the step of using data corresponding to the three-dimensional position of each of the at least three spots that are all captured in one of the multiple images.

[0233] In some applications, the method further includes refining the estimation of the 3D surface using data corresponding to the 3D location of at least one additional spot, the at least one additional spot having a 3D location calculated based on another image from a plurality of images such that all of at least three spots and at least one additional spot lie on the 3D surface.

[0234] In some applications, the step of using data corresponding to the three-dimensional position of each of the at least three spots includes the step of using data corresponding to at least three spots, where each spot is captured in one image of a plurality of images.

[0235] According to some applications of the present invention, a method for tracking the motion of an intraoral scanner is further provided, and the method is (A) A step of measuring the motion of the intraoral scanner relative to the intraoral surface being scanned using at least one camera coupled to the intraoral scanner, (B) A step of measuring the motion of the intraoral scanner relative to a fixed coordinate system using at least one inertial measuring unit (IMU) coupled to the intraoral scanner, (C) Using the processor, (i) a step of calculating the motion of the oral cavity surface relative to a fixed coordinate system by (a) subtracting the motion of the oral cavity scanner relative to the oral cavity surface from (b) the motion of the oral cavity scanner relative to a fixed coordinate system, (ii) A step of constructing a predictive model of the motion of the oral cavity surface relative to a fixed coordinate system based on accumulated data of the motion of the oral cavity surface relative to the coordinate system, (iii) The step of calculating the estimated position of the intraoral scanner relative to the intraoral surface by (a) a prediction of the motion of the intraoral surface relative to the coordinate system, derived based on a prediction model, and (b) the motion of the intraoral scanner relative to the coordinate system, measured by an IMU, by subtracting this from the motion of the intraoral scanner relative to the coordinate system, which is measured by an IMU.

[0236] In some applications, the method further includes the steps of determining whether the measurement of the motion of the intraoral scanner relative to the oral surface is being obstructed using at least one camera, and calculating the estimated position of the intraoral scanner relative to the oral surface in response to the determination that the measurement of the motion is being obstructed.

[0237] Methods are provided according to several applications of the present invention, and said methods The steps include driving each of the structured light projectors of one or more structured light projectors to project a distribution of discrete, unconnected light spots onto the three-dimensional surface of the oral cavity, The steps include driving each camera of one or more cameras to capture multiple images, each image including at least one spot, and each camera of the one or more cameras including a camera sensor including a pixel array, Based on stored calibration values ​​indicating (a) camera rays corresponding to each pixel on the camera sensor of each of the one or more cameras, and (b) projector rays corresponding to each projection light spot from each of the one or more projectors, wherein each projector ray corresponds to each path p on at least one pixel of the camera sensor, Using the processor, (a) A step of executing a corresponding algorithm to calculate the respective three-dimensional positions of multiple projection spots on the three-dimensional surface of the oral cavity, (b) Collect data at multiple points in time, the data including the calculated three-dimensional position of each of the detected spots on the oral surface, (c) For each projector ray r, based on the collected data, define a pixel update path p' for each camera sensor such that all of the calculated 3D positions corresponding to the spots generated by the projector ray r correspond to positions along the respective pixel update path p' for each camera sensor. (d) The step of comparing each update path p' of a pixel with the path p of the pixel corresponding to the projector ray r of each camera sensor from the stored calibration value, (e) For at least one camera sensor s, if the update path p' of the pixel corresponding to the projector ray r is different from the path p of the pixel corresponding to the projector ray r from the stored calibration value, The difference between the update path p' of the pixel corresponding to each projector ray r and the path p of each pixel corresponding to each projector ray r from the stored calibration value is reduced, and this is then corrected. (i) Stored calibration values ​​indicating camera rays corresponding to each pixel on the camera sensor s of each of the one or more cameras, and (ii) Stored calibration values ​​indicating the projector ray r corresponding to each projector spot from each of the one or more projectors, The process includes the step of changing stored calibration data selected from a group consisting of the following:

[0238] In some application examples, the selected stored calibration data includes stored calibration values ​​that represent the camera rays corresponding to each pixel on the camera sensor of one or more cameras. The step of changing the stored calibration data involves changing one or more parameters of a parameterized camera calibration function that defines the camera ray corresponding to each pixel on at least one camera sensor s, (i) The calculated three-dimensional positions of each of the multiple detection spots on the oral cavity surface, (ii) A stored calibration value indicating the respective camera ray corresponding to each pixel on the camera sensor where each of the plurality of detection spots should have been detected, This includes a step to reduce the difference between them.

[0239] In some application examples, the selected stored calibration data includes stored calibration values ​​indicating projector rays corresponding to each projector light spot from each of the projectors of one or more projectors, and the step of changing the stored calibration data is: (i) An indexed list assigning each projector ray r to a pixel path p, or (ii) One or more parameters of the parameterized projector calibration model that defines each projector ray r, This includes a step of changing [something].

[0240] In some applications, the step of changing the stored calibration data includes changing the indexed list by reassigning each projector ray r based on the respective update path p' of the pixel corresponding to each projector ray r.

[0241] In some application examples, the step of changing the stored calibration data is: (i) Stored calibration values ​​indicating camera rays corresponding to each pixel on the camera sensor s of each of the one or more cameras, (ii) Stored calibration values ​​indicating the projector ray r corresponding to each projected light spot from each of the one or more projectors, This includes a step of changing [something].

[0242] In some applications, the step of changing the stored calibration value includes the step of iteratively changing the stored calibration value.

[0243] In some application examples, the method is further described as follows: The steps include driving each camera of one or more cameras to capture multiple images of a calibration target having predetermined parameters, Using the processor, The steps include: executing a triangulation algorithm to calculate the parameters of each object to be calibrated based on the captured image; Execute the optimization algorithm, (a)(i) reduce the difference between the update path p' of the pixel corresponding to the projector ray r and (ii) the path p of the pixel corresponding to the projector ray r from the stored calibration value, and (b) a step performed using each parameter of the calibration object calculated based on the captured image.

[0244] In some applications, the object to be calibrated is a three-dimensional object of known shape, the step of driving each camera of one or more cameras to capture an image of the three-dimensional object to be calibrated includes the step of driving each camera of one or more cameras to capture an image of the three-dimensional object to be calibrated, and the predetermined parameter of the object to be calibrated is the dimensions of the three-dimensional object to be calibrated.

[0245] In some applications, the object to be calibrated is a two-dimensional object having visually distinguishable features, and the step of driving each camera of one or more cameras to capture multiple images of the object to be calibrated includes the step of driving each camera of one or more cameras to capture an image of the two-dimensional object to be calibrated, wherein a predetermined parameter of the two-dimensional object to be calibrated is the respective distance between each of the visually distinguishable features.

[0246] According to some applications of the present invention, a method for calculating the three-dimensional structure of the three-dimensional surface of the oral cavity is further provided, and the method is The steps include scanning the oral cavity surface, The steps include driving one or more uniform light projectors to project broadband spectral light onto the three-dimensional surface inside the oral cavity, The steps include driving the camera to capture multiple 2D color images of the 3D surface inside the oral cavity, Using the processor, A step of calculating the three-dimensional positions of multiple points on the three-dimensional surface of the oral cavity based on an oral cavity surface scan, The method includes the step of calculating the three-dimensional structure of the oral cavity three-dimensional surface, constrained by the three-dimensional positions of the plurality of points, based on a plurality of two-dimensional color images of the oral cavity three-dimensional surface.

[0247] In some embodiments, the oral cavity surface is scanned by driving one or more structured light projectors to project a structured light pattern onto the three-dimensional oral cavity surface. One or more cameras are driven to capture multiple structured light images, each image containing at least a portion of the structured light pattern.

[0248] In some applications, the step of driving one or more structured light projectors includes driving one or more structured light projectors to project a distribution of discrete, unconnected light spots.

[0249] In some application examples, the step of calculating the 3D structure is, (a) Inputting a plurality of two-dimensional color images of the three-dimensional surface inside the oral cavity, and (b) the calculated three-dimensional positions of a plurality of points on the three-dimensional surface inside the oral cavity into the neural network, The method includes the step of determining a predicted depth map of each of the three-dimensional surfaces of the oral cavity captured in each of the two-dimensional color images using the neural network.

[0250] In some applications, the method further includes the step of using the processor to stitch together the respective depth maps to obtain a three-dimensional structure of the three-dimensional surface of the oral cavity.

[0251] In some applications, the method further includes coordinating the capture of the structured optical image and the capture of the two-dimensional color image to generate an alternating sequence in which one or more broadband spectral optical image frames are interspersed within one or more structured optical image frames.

[0252] In some application examples, The step of driving one or more cameras to capture multiple structured light images includes the step of driving each of two or more cameras to capture multiple structured light images, The step of driving the camera to capture the plurality of two-dimensional color images includes the step of driving each of the two or more cameras to capture the plurality of two-dimensional color images.

[0253] In some application examples, The step determined by the neural network includes determining, for a given image frame, the predicted depth map of each portion of the three-dimensional surface inside the oral cavity captured in a two-dimensional color image by each of the two or more cameras, The method further includes the step of using the processor to stitch together the respective depth maps to obtain a predicted depth map of the oral cavity three-dimensional surface captured in the given image frame.

[0254] In some applications, the method further includes the step of training the neural network, and the training is (a) The steps of driving one or more structured light projectors to project a training stage structured light pattern onto the three-dimensional surface of the training stage, (b) A step of driving one or more training stage cameras to capture multiple structured light images, each image including at least a portion of the training stage structured light pattern, (c) The step of driving one or more training stage uniform light projectors to project broadband spectral light onto the three-dimensional surface of the training stage, (d) The steps of driving one or more training stage cameras to capture multiple two-dimensional color images of the three-dimensional surface of the training stage using illumination from the training stage uniform light projector, (e) The steps of coordinating the capture of the structured light image and the capture of the two-dimensional color image to generate an alternating sequence in which one or more image frames of two-dimensional color images are scattered within one or more image frames of structured light images, (f) The step of inputting multiple 2D color images into the neural network, (g) Based on the structured optical image of the 3D surface of the training stage, a plurality of 3D reconstructions of the 3D surface of the training stage are input to the neural network, wherein the 3D reconstruction includes the calculated 3D positions of a plurality of points on the 3D surface of the training stage, (h) Interpolating the position of one or more training stage cameras relative to the training stage 3D surface for each 2D color image frame based on the calculated 3D positions of multiple points on the training stage surface calculated based on the structured light image frames before and after each 2D color image frame, (i) projecting the three-dimensional reconstruction onto the respective fields of view of one or more training stage cameras, and estimating a predicted depth map of the training stage three-dimensional surface as seen in each two-dimensional color image, constrained by the calculated three-dimensional positions of the plurality of points, based on the projection; (j) A step of comparing each predicted depth map of the 3D surface of the training stage with the corresponding true depth map of the 3D surface of the training stage, (k) A step of optimizing the neural network based on the difference between each predicted depth map and the corresponding true depth map to better estimate subsequent predicted depth maps, Includes.

[0255] In some applications, the step of driving one or more structured light projectors to project the training stage structured light pattern includes the step of driving one or more structured light projectors to project a distribution of discrete, unconnected light spots onto the three-dimensional surface of the training stage.

[0256] In some application examples, the step of driving one or more training stage cameras includes the step of driving at least two training stage cameras.

[0257] According to some applications of the present invention, an intraoral scanning device is further provided, and the device, A long, slender handheld wand, the long, slender handheld wand including a probe at its distal end, One or more illumination sources coupled to the probe, One or more near-infrared (NIR) light sources coupled to the probe, Coupled to the probe, one or more cameras configured to (a) capture an image using light from one or more illumination sources, and (b) capture an image using NIR light from the NIR light source, A processor configured to execute a navigation algorithm to determine the position of the handheld wand as it moves in space, wherein the input to the navigation algorithm is (a) an image captured using light from one or more illumination sources, and (b) an image captured using NIR light.

[0258] In some applications, the one or more illumination light sources are one or more structured light sources.

[0259] In some applications, the one or more illumination light sources are one or more uniform light sources.

[0260] According to some applications of the present invention, a method for tracking the motion of an intraoral scanner is further provided. The method includes: Illuminating an intraoral three-dimensional surface using one or more illumination sources coupled to the intraoral scanner; Using one or more near-infrared (NIR) light sources coupled to the intraoral scanner to drive each NIR light source of the one or more NIR light sources to emit NIR light onto the intraoral three-dimensional surface; Using one or more cameras coupled to the intraoral scanner to (a) capture a plurality of images using light from the one or more illumination light sources and (b) capture a plurality of images using the NIR light; Using a processor to Execute a navigation algorithm to track the motion of the intraoral scanner relative to the intraoral three-dimensional surface using (a) the images captured using light from the one or more illumination light sources and (b) the images captured using the NIR light. And including.

[0261] In some applications, the step of using one or more illumination light sources includes illuminating the intraoral three-dimensional surface using one or more structured light sources.

[0262] In some applications, the step of using one or more illumination sources includes using one or more uniform light sources.

[0263] According to some applications of the present invention, an intraoral scanning device for use with a sleeve is further provided. The device includes: An elongated handheld wand, comprising a probe at its distal end, configured to be detachably disposed within the sleeve, A structured light projector coupled to the probe, comprising (a) an illumination field of at least 30 degrees, (b) a laser configured to emit polarized laser light, and (c) a pattern generating optical element configured to generate a pattern of light when a laser diode is activated and transmits light through the pattern generating optical element, The probe is coupled to at least one camera, which includes a camera sensor, The probe is configured such that light enters and exits the probe through the sleeve. The laser is positioned at a distance from the camera such that when the probe is placed inside the sleeve, a portion of the light pattern is reflected from the sleeve and reaches the camera sensor. The laser is positioned at a rotation angle relative to its own optical axis such that, due to the polarization of the light pattern, the degree of reflection by the sleeve of a portion of the light pattern is less than 70% of the maximum reflection of all possible rotation angles of the laser relative to its optical axis.

[0264] In some applications, the distance between the structured light projector and the camera is 1 to 6 times the distance between the structured light projector and the sleeve, when the handheld wand is positioned within the sleeve.

[0265] In some application examples, each of the at least one cameras has a field of view of at least 30 degrees.

[0266] In some applications, the laser is positioned at a rotation angle relative to its own optical axis such that the degree of reflection by the sleeve of a portion of the light pattern is less than 60% of the maximum reflection of all possible rotation angles of the laser relative to its optical axis, due to the polarization of the light pattern.

[0267] In some applications, the laser is positioned at a rotation angle relative to its own optical axis such that, due to the polarization of the light pattern, the degree of reflection by the sleeve of a portion of the light pattern is 15% to 60% of the maximum reflection for all possible rotation angles of the laser with respect to its optical axis.

[0268] A method for generating a three-dimensional image using an intraoral scanner is further provided according to some applications of the present invention, and the method is (A) Using at least two cameras rigidly connected to an intraoral scanner such that the fields of view of each camera have non-overlapping portions, A step of capturing multiple images of the three-dimensional surface inside the oral cavity, (B) Using a processor, For the non-overlapping portions of each camera's field of view, a simultaneous localization and mapping (SLAM) algorithm is performed using images captured from each camera. The localization of each camera is solved based on the assumption that the motion of each camera is the same as the motion of all other cameras.

[0269] In some application examples, The fields of view of the first camera and the second camera among the aforementioned cameras also overlap, The capturing step includes capturing a plurality of images of the three-dimensional surface of the oral cavity such that the features of the three-dimensional surface of the oral cavity in the overlapping portion of each of the fields of view appear in the images captured by the first and second cameras. The step of using the processor includes the step of executing a SLAM algorithm using the features of the three-dimensional intraoral surface that appear in the images of the at least two cameras.

[0270] A method for generating a three-dimensional image using an intraoral scanner is further provided according to several applications of the present invention, and the method is The steps include driving one or more structured light projectors to project a structured light pattern onto a three-dimensional surface inside the oral cavity, Steps include: driving one or more cameras to capture multiple structured light images, each image including at least a portion of the structured light pattern; The steps include driving one or more uniform light projectors to project broadband spectral light onto the three-dimensional surface inside the oral cavity, The steps include driving at least one camera to capture a two-dimensional color image of the three-dimensional surface of the oral cavity using illumination from the uniform light projector, The steps include: adjusting the capture of the structured light and the capture of the broadband spectral light to generate an alternating sequence in which one or more broadband spectral light image frames are scattered within one or more structured light image frames; Using the processor, The steps include calculating the three-dimensional position of each of the multiple points on the three-dimensional surface of the oral cavity that are captured in the multiple structured optical image frames, The steps include interpolating the motion of at least one camera between a first broadband spectral image frame and a second broadband spectral image frame based on the calculated three-dimensional positions of a plurality of points in the structured optical image frames before and after the broadband spectral optical image frame, (a) using features of the intraoral three-dimensional surface captured in the first and second broadband spectral image frames by the at least one camera, and (b) performing a simultaneous localization and mapping (SLAM) algorithm constrained by the interpolated motion of the camera between the first and second broadband spectral image frames.

[0271] In some applications, the step of driving one or more structured light projectors to project the structured light pattern includes the step of driving one or more structured light projectors to project a distribution of discrete, unconnected light spots onto the three-dimensional surface of the oral cavity.

[0272] A method for generating a three-dimensional image using an intraoral scanner is further provided according to some applications of the present invention, and the method is Driving one or more structured light projectors to project a structured light pattern onto a three-dimensional surface within the oral cavity; Driving one or more cameras to capture a plurality of structured light images, each image including at least a portion of the structured light pattern; Driving one or more uniform light projectors to project broadband spectral light onto the three-dimensional surface within the oral cavity; Driving the one or more cameras to capture a two-dimensional color image of the three-dimensional surface within the oral cavity using illumination from the uniform light projector; Adjusting the capture of the structured light and the capture of the broadband spectral light to generate an alternating sequence in which one or more broadband spectral light image frames are interspersed among one or more structured light image frames; Using a processor, (a) calculating the three-dimensional position of features on the three-dimensional surface within the oral cavity based on the structured light image frames, the features also being captured in a first image frame of the broadband spectral light and a second image frame of the broadband spectral light; (b) calculating the motion of at least one camera between the first image frame of the broadband spectral light and the second image frame of the broadband spectral light based on the calculated three-dimensional position of the features; (c) performing a simultaneous localization and mapping (SLAM) algorithm using (i) features of the three-dimensional surface within the oral cavity for which a three-dimensional position was not calculated based on the structured light image frames captured in the first and second broadband spectral light image frames by the at least one camera, and (ii) the calculated motion of the camera between the first and second broadband spectral light image frames.

[0273] In some applications, the step of driving the one or more structured light projectors to project the structured light pattern includes driving the one or more structured light projectors to project respective distributions of discrete, unconnected light spots onto the three-dimensional surface within the oral cavity.

[0274] According to some applications of the present invention, a method for calculating the three-dimensional structure of the three-dimensional surface inside the oral cavity of a subject is further provided, and the method is The steps include driving one or more structured light projectors to project a structured light pattern spot onto the three-dimensional surface of the oral cavity, Steps include driving one or more cameras to capture multiple structured light images, each image including at least one of the spots, The steps include driving one or more uniform light projectors to project broadband spectral light onto the three-dimensional surface inside the oral cavity, The steps include driving at least one camera to capture a two-dimensional color image of the three-dimensional surface of the oral cavity using illumination from the uniform light projector, The steps include: adjusting the capture of the structured light and the capture of the broadband spectral light to generate an alternating sequence in which one or more broadband spectral light image frames are scattered within one or more structured light image frames; Using the processor, The steps include: determining, based on the two-dimensional color image, whether the spot was projected onto moving tissue or stable tissue in the oral cavity for each of the multiple spots; and, based on the determination, assigning a confidence grade to each of the detected multiple spots, assigning high confidence to fixed tissue and low confidence to moving tissue. The process includes the step of running a 3D reconstruction algorithm using the detection spots based on the confidence grade for each of the plurality of detection spots.

[0275] In some applications, the step of driving one or more structured light projectors to project the structured light pattern includes the step of driving one or more structured light projectors to project a distribution of discrete, unconnected light spots onto the three-dimensional surface of the oral cavity.

[0276] In some applications, the step of performing the three-dimensional reconstruction algorithm includes performing the three-dimensional reconstruction algorithm using only a subset of the detected spots, the subset consisting of spots that have been assigned a confidence grade above the fixed tissue threshold.

[0277] In some applications, the steps of executing the three-dimensional reconstruction algorithm include (a) assigning a weight to each spot based on the respective confidence level assigned to that spot, and (b) using the respective weights of each spot in the three-dimensional reconstruction algorithm.

[0278] This invention will be better understood from the following detailed description of its applications, accompanied by the drawings. [Brief explanation of the drawing]

[0279] [Figure 1] This is a schematic diagram of a handheld wand having multiple structured light projectors and cameras arranged within the probe at the distal end of the handheld wand, according to some application examples of the present invention. [Figure 2A] ~ [Figure 2B] These are schematic diagrams illustrating positioning configurations for a camera and a structured light projector, respectively, according to several application examples of the present invention. [Figure 2C] This chart illustrates several different configurations of the positions of the structured light projector and camera within the probe, based on several application examples of the present invention. [Figure 2D] ~ [Figure 2E] These isometric views show specific configurations of the positions of the structured light projector and camera within the probe, according to several application examples of the present invention, from two different viewpoints. [Figure 3] This is a schematic diagram of a structured light projector based on several application examples of the present invention. [Figure 4]This is a schematic diagram of a structured light projector that projects a distribution of discrete, unconnected light spots onto the focal planes of multiple objects, according to several application examples of the present invention. [Figure 5A] ~ [Figure 5B] This is a schematic diagram of a structured light projector, including a beam shaping optical element and an additional optical element positioned between the beam shaping optical element and the pattern generating optical element, according to several application examples of the present invention. [Figure 6A] ~ [Figure 6B] This is a schematic diagram of a structured light projector that projects discrete, unconnected spots and a camera sensor that detects the spots, according to several application examples of the present invention. [Figure 7] This flowchart outlines a method for generating digital 3D images using several application examples of the present invention. [Figure 8] This flowchart outlines how to perform a specific step in the method shown in Figure 7, based on several applications of the present invention. [Figure 9] ~ [Figure 12] This is a schematic diagram illustrating a simplified example of the steps in Figure 8, based on several applications of the present invention. [Figure 13] This flowchart outlines further steps in a method for generating digital 3D images, based on several applications of the present invention. [Figure 14] ~ [Figure 17] This is a schematic diagram illustrating a simplified example of the steps in Figure 13, based on several applications of the present invention. [Figure 18] This is a schematic diagram of a probe including a diffuse reflector, based on several application examples of the present invention. [Figure 19A] ~ [Figure 19B] This diagram shows a structured light projector and a schematic cross-section of a light beam transmitted by a laser diode, together illustrating pattern-generating optical elements arranged in the beam's optical path according to several application examples of the present invention. [Figure 20A] ~ [Figure 20E] This is a schematic diagram showing a microlens array used as a pattern-generating optical element in a structured light projector, according to several application examples of the present invention. [Figure 21A] ~ [Figure 21C] This is a schematic diagram of a composite two-dimensional diffraction periodic structure used as a pattern-generating optical element in a structured light projector, based on several application examples of the present invention. [Figure 22A] ~ [Figure 22B] This is a schematic diagram showing a single optical element having an aspherical first surface and a planar second surface opposite the first surface, and a structured light projector including the optical element, according to several application examples of the present invention. [Figure 23A] ~ [Figure 23B] This is a schematic diagram of an axicon lens and a structured light projector including an axicon lens, according to several application examples of the present invention. [Figure 24A] ~ [Figure 24B] This is a schematic diagram showing an optical element having an aspherical surface on a first side and a flat surface on a second side opposite to the first side, and a structured light projector including the optical element, according to several application examples of the present invention. [Figure 25] This is a schematic diagram of a single optical element in a structured light projector, based on several application examples of the present invention. [Figure 26A] ~ [Figure 26B] This is a schematic diagram of a structured light projector equipped with two or more laser diodes, according to some applications of the present invention. [Figure 27A] ~ [Figure 27B] This is a schematic diagram illustrating different methods of combining laser diodes of different wavelengths, based on several applications of the present invention. [Figure 28] This flowchart outlines the steps of a "spot tracking" method based on several applications of the present invention. [Figure 29]This invention provides several simplified examples of detection spots in some applications, and schematic diagrams illustrating how a processor can determine which set of detection spots can be considered tracked. [Figure 30] This flowchart outlines a method for determining tracking spots based on several applications of the present invention. [Figure 31] ~ [Figure 32] This flowchart outlines a method for detecting tracking spots in subsequent images, based on several applications of the present invention. [Figure 33] This schematic diagram illustrates an example of how spot tracking can be useful in identifying a detection spot that is projected from a specific projector beam, based on several applications of the present invention. [Figure 34A] ~ [Figure 34B] This is a simplified schematic diagram of a camera sensor showing two detection spots, based on several application examples of the present invention. [Figure 35] ~ [Figure 36] This flowchart outlines several applications of the present invention, illustrating each method in which spot tracking may be used. [Figure 37A] ~ [Figure 37B] This is a schematic diagram illustrating the points used in three-dimensional reconstruction before and after the processor implements spot tracking, according to several application examples of the present invention. [Figure 38] This flowchart outlines the steps of a method for generating digital 3D images, hereafter referred to as "ray tracing," based on several applications of the present invention. [Figure 39] This graph shows the time-series tracking of the length of a projector beam in several application examples of the present invention, and illustrates a specific simplified diagram of a camera sensor corresponding to a particular image frame. [Figure 40A] ~ [Figure 40B] This graph shows experimental sets of data before and after ray tracing, illustrating several applications of the present invention. [Figure 41]This is a schematic diagram of a projector that projects a spot using multiple camera sensors, according to several application examples of the present invention. [Figure 42A] ~ [Figure 42B] This flowchart outlines a method for generating 3D images using several application examples of the present invention. [Figure 43A] ~ [Figure 43B] This flowchart outlines various methods for tracking the motion of an intraoral scanner, based on several applications of the present invention. [Figure 44A] ~ [Figure 45] This schematic diagram shows a simplified camera image in which multiple detection spots from a single projector beam, all captured at different times, are superimposed on the same image, according to several application examples of the present invention. [Figure 46A] ~ [Figure 46D] This figure shows a simplified scenario in which a processor identifies the projector (and not any of the cameras) as the object that needs to be recalibrated, based on several applications of the present invention. [Figure 47A] ~ [Figure 47B] This figure shows a simplified scenario in which a processor identifies the camera (and not any of the projectors) as the component that needs to be recalibrated, based on several applications of the present invention. [Figure 48A] ~ [Figure 48B] This figure shows a simplified scenario in which the processor cannot reasonably assume that a shift occurred in the camera only or in the projector only, according to some applications of the present invention. [Figure 49A] ~ [Figure 49B] These are schematic diagrams showing three-dimensional and two-dimensional calibration targets, respectively, according to several application examples of the present invention. [Figure 50] This flowchart shows a method for tracking the motion of a handheld wand, according to several application examples of the present invention. [Figure 51A] ~ [Figure 51F]This flowchart illustrates a method for calculating the three-dimensional structure of the three-dimensional surface of the oral cavity, based on several application examples of the present invention. [Figure 51G] ~ [Figure 51I] This is a schematic diagram illustrating different combinations of inputs to a neural network, based on several application examples of the present invention. [Figure 52A] This flowchart shows a method for training a neural network using several application examples of the present invention. [Figure 52B] This is a block diagram of neural network training using several application examples of the present invention. [Figure 52C] This flowchart illustrates how a neural network outputs a depth map and a corresponding confidence map, based on several application examples of the present invention. [Figure 52D] ~ [Figure 52F] This is a schematic diagram illustrating the training of a neural network that outputs a depth map and a corresponding confidence map, based on several application examples of the present invention. [Figure 52G] This flowchart illustrates how confidence maps can be used in several applications of the present invention. [Figure 53A] This is a schematic diagram of a disposable sleeve, which is placed on the distal end of an intraoral scanner before the probe is placed in the patient's mouth, in order to prevent cross-contamination between patients, according to some applications of the present invention. [Figure 53B] This graph shows the reflectance of polarized laser light according to the Fresnel equation and its use in several applications of the present invention. [Figure 54A] This flowchart illustrates a method for generating a 3D image using a handheld wand, based on some applications of the present invention. [Figure 54B] This is a schematic diagram showing the positional relationship between a projector and a camera in several application examples of the present invention. [Figure 55]This flowchart shows a method for generating a 3D image using a handheld wand, according to several applications of the present invention. [Figure 56A] This flowchart shows a method for generating a 3D image using a handheld wand, according to several applications of the present invention. [Figure 56B] This is a schematic diagram illustrating two image frames of unstructured light and two features of a three-dimensional intraoral surface, based on several application examples of the present invention. [Figure 57] This flowchart shows a method for calculating the three-dimensional structure of the three-dimensional surface inside the oral cavity of a subject, based on several applications of the present invention. [Figure 58] ~ [Figure 59B] This is a schematic diagram of a neural network based on several application examples of the present invention. [Figure 60] This figure shows one embodiment of a system for performing an intraoral scan and generating a virtual 3D model of the dental arch. [Figure 61] This shows a block diagram of an exemplary computing device according to an embodiment of the present disclosure. [Figure 62] This flowchart shows a method for overcoming manufacturing discrepancies between intraoral scanners, based on several applications of the present invention. [Figure 63] This schematic diagram illustrates another method for overcoming manufacturing deviations between intraoral scanners, based on several applications of the present invention. [Figure 64] This flowchart shows a method for testing whether the cropping and morphing of each runtime image accurately accounts for possible manufacturing deviations relative to a given intraoral scanner, and if not, for refining the training of a neural network based on local refining stage scans relative to that given intraoral scanner, according to several applications of the present invention. [Figure 65] This flowchart shows a method for overcoming manufacturing discrepancies between intraoral scanners, based on several applications of the present invention. [Figure 66]This flowchart shows a method for training a neural network using several application examples of the present invention. [Modes for carrying out the invention]

[0280] Next, refer to Figure 1, a schematic diagram of an elongated handheld wand 20 for intraoral scanning according to several applications of the present invention. Multiple structured light projectors 22 and multiple cameras 24 are coupled to a rigid structure 26 located at the distal end 30 of the handheld wand within the probe 28. In some applications, during intraoral scanning, the probe 28 enters the oral cavity of the subject.

[0281] In some applications, the structured light projector 22 is positioned within the probe 28 so that each structured light projector 22 faces an object 32 outside the handheld wand 20, which is positioned within its field of view, in contrast to positioning the structured light projector at the proximal end of the handheld wand to illuminate the object by reflection of light from a mirror and subsequent reflection of light onto the object. Similarly, in some applications, the camera 24 is positioned within the probe 28 so that each camera 24 faces an object 32 outside the handheld wand 20, which is positioned within its field of view, in contrast to positioning the camera at the proximal end of the handheld wand to view the object by reflection of light from a mirror onto the camera. Such positioning of the projectors and cameras within the probe 28 allows the scanner to have a large overall field of view while maintaining a low-profile probe.

[0282] In some applications, the height H1 of the probe 28 is less than 15 mm, and the height H1 of the probe 28 is measured from the bottom surface 176 (sensing surface) where reflected light from the object 32 being scanned enters the probe 28 to the top surface 178 opposite the bottom surface 176. In some applications, the height H1 is between 10 and 15 mm.

[0283] In some applications, each camera 24 has a large field of view β (beta) of at least 45 degrees, for example, at least 70 degrees, for example, at least 80 degrees, for example, 85 degrees. In some applications, the field of view may be less than 120 degrees, for example, less than 100 degrees, for example, less than 90 degrees. Experiments conducted by the inventors have shown that a field of view β (beta) of each camera between 80 and 90 degrees is particularly useful as it provides a good balance between pixel size, field of view and camera overlap, optical quality, and cost. Camera 24 may include a camera sensor 58 and an objective optical system 60 including one or more lenses. To enable close-focus imaging, camera 24 may focus on an object focal plane 50 located between 1 mm and 30 mm from the lens furthest from the camera sensor, for example, between 4 mm and 24 mm, for example, between 5 mm and 11 mm, for example, between 9 mm and 10 mm. Experiments conducted by the inventors revealed that it is particularly useful for the object focal plane 50 to be located between 5mm and 11mm from the lens furthest from the camera sensor, because scanning teeth is easy at this distance and the depth of focus on most tooth surfaces is high. In some applications, the camera 24 can capture images at a frame rate of at least 30 frames / second, for example, at least 75 frames / second, for example, at least 100 frames / second. In some applications, the frame rate may be less than 200 frames per second.

[0284] As explained above, the large field of view achieved by combining the individual fields of view of all cameras can improve accuracy by reducing the amount of image stitching error, especially in edentulous areas where the gingival surface is smooth and there may be few clear, high-resolution 3D features. With a large field of view, large, smooth features such as the overall curve of the tooth become visible in each image frame, improving the accuracy of stitching together the surfaces obtained from such multiple image frames.

[0285] Similarly, each structured light projector 22 may have a large illumination field α (alpha) of at least 45 degrees, for example, at least 70 degrees. In some applications, the illumination field α (alpha) may be less than 120 degrees, for example, less than 100 degrees. Further features of the structured light projector 22 are described below.

[0286] In some applications, to improve image acquisition, each camera 24 has multiple discrete preset focal positions, at which point the camera focuses on the respective object's focal plane 50. Each camera 24 may include an autofocus actuator to select a focal position from the discrete preset focal positions to improve a given image acquisition. Additionally or alternatively, each camera 24 includes an optical aperture phase mask that extends the depth of field of the camera so that the image formed by each camera maintains focus over all object distances between 1mm and 30mm from the lens furthest from the camera sensor, for example, between 4mm and 24mm, for example between 5mm and 11mm, for example between 9mm and 10mm.

[0287] In some applications, the structured light projector 22 and camera 24 are coupled to a rigid structure 26 in close proximity and / or alternately such that (a) most of the field of view of each camera overlaps with the field of view of an adjacent camera, and (b) most of the field of view of each camera overlaps with the illumination field of an adjacent projector. Optionally, at least 20%, e.g., at least 50%, e.g., at least 75% of the projected light pattern is present in at least one field of view of the camera at the object focal plane 50 located at least 4 mm away from the lens furthest from the camera sensor. Due to the different possible configurations of the projectors and cameras, some parts of the projected pattern may never be seen within the field of view of any camera, and some parts of the projected pattern may be obscured from the field of view by the object 32 as the scanner moves around during scanning.

[0288] The rigid structure 26 may be a non-flexible structure to which the structured optical projector 22 and camera 24 are coupled in order to provide structural stability to the optical system within the probe 28. Coupling all projectors and cameras to a common rigid structure helps maintain the geometric integrity of the optical system of each structured optical projector 22 and camera 24 under changing ambient conditions, such as mechanical stress that may be induced by the subject's mouth. Furthermore, the rigid structure 26 helps maintain the stable structural integrity and relative positioning of the structured optical projector 22 and camera 24. As will be further described below, controlling the temperature of the rigid structure 26 may help maintain the geometric integrity of the optical system over a wide range of ambient temperatures when the probe 28 enters and exits the subject's oral cavity or when the subject breathes during scanning.

[0289] Next, we refer to Figures 2A and 2B, schematic diagrams of the respective positioning configurations of the camera 24 and structured light projector 22 in some application examples of the present invention. In some application examples, the camera 24 and structured light projector 22 are positioned so that they do not all face the same direction in order to improve the overall field of view and illumination field of the intraoral scanner. In some application examples, such as those shown in Figure 2A, multiple cameras 24 are coupled to a rigid structure 26 such that the angle θ (theta) between the two respective optical axes 46 of at least two cameras 24 is 90 degrees or less, for example, 35 degrees or less. Similarly, in some application examples, such as those shown in Figure 2B, multiple structured light projectors 22 are coupled to a rigid structure 26 such that the angle φ (phi) between the two respective optical axes 48 of at least two structured light projectors 22 is 90 degrees or less, for example, 35 degrees or less.

[0290] Next, we refer to Figure 2C, a chart illustrating several different configurations of the positions of the structured light projector 22 and camera 24 within the probe 28, according to several applications of the present invention. The structured light projector 22 is represented by a circle in Figure 2C, and the camera 24 is represented by a rectangle in Figure 2C. Note that the rectangle is used to represent the camera because typically the field of view β (beta) of each camera sensor 58 and each camera 24 has an aspect ratio of 1:2. Column (a) in Figure 2C shows bird's-eye views of various configurations of the structured light projector 22 and camera 24. The labeled x-axis in the first row of Column (a) corresponds to the central longitudinal axis of the probe 28. Column (b) shows side views of the camera 24 from various configurations, viewed from a line of sight that is coaxial with the central longitudinal axis of the probe 28. As shown in Figure 2A, Column (b) in Figure 2C shows cameras 24 positioned so that their optical axes 46 are at angles of 90 degrees or less, e.g., 35 degrees or less, relative to each other. Column (c) shows side views of the camera 24 in various configurations, viewed from a line of sight perpendicular to the central longitudinal axis of the probe 28.

[0291] Typically, the farthest (towards the positive x-direction in Figure 2C) and nearest (towards the negative x-direction in Figure 2C) cameras 24 are positioned such that their optical axes 46 are rotated slightly inward relative to the next nearest camera 24, for example, by an angle of less than 90 degrees, for example, less than 35 degrees. More centrally located cameras 24, i.e., cameras 24 that are neither the farthest nor the nearest camera 24, are positioned to face directly outward from the probe, and their optical axes 46 are substantially perpendicular to the central longitudinal axis of the probe 28. Note that in row (xi), the projector 22 is positioned at the farthest point of the probe 28, and therefore its optical axis 48 is pointed inward, so that more spots 33 projected from that particular projector 22 are seen by more cameras 24.

[0292] Typically, the number of structured light projectors 22 in the probe 28 may range from, for example, two as shown in row (iv) of Figure 2C to six as shown in row (xii). Typically, the number of cameras 24 in the probe 28 may range from, for example, four as shown in rows (iv) and (v) to seven as shown in row (ix). Note that the various configurations shown in Figure 2C are illustrative and not limiting, and the scope of the invention includes additional configurations not shown. For example, the scope of the invention includes more than five projectors 22 arranged in the probe 28 and more than seven cameras arranged in the probe 28.

[0293] Next, refer to Figures 2D-E, which are isometric views of specific configurations relating to the positions of the structured light projectors 22 and cameras 24 within the probe 28, shown from two different respective viewpoints, according to several applications of the present invention. Figure 2D is shown from the same bird's-eye view viewpoint as column (a) of Figure 2C. In some applications, six cameras 24 are arranged at equal intervals within the probe 28, with three cameras on each side of the probe 28, and five structured light projectors 22 are arranged within the central probe 28 along the central longitudinal axis of the probe 28 (shown by the dashed line 29).

[0294] In some applications, the camera 24 and structured light projector 22 are all coupled to a flexible printed circuit board (PCB) to accommodate the angular positioning of the camera 24 and structured light projector 22 within the probe 28. This angular positioning of the camera 24 and structured light projector 22 is shown in Figure 2E. The most distal (i.e., towards the positive x-direction) camera 24 and structured light projector 22 are positioned so that their respective optical axes are tilted backward toward the handheld wand 20 at an angle of, for example, 45 degrees or less, or 35 degrees or less. This allows, for example, the most distal camera to capture the posterior wall of the molars in the oral cavity. The most proximal (i.e., towards the negative x-direction) camera 24 and structured light projector 22 are positioned so that their respective optical axes are tilted forward toward the distal end of the probe 28 at an angle of, for example, 45 degrees or less, or 35 degrees or less, to obtain improved overlap of the fields of view of the camera 24. Furthermore, all of the structured light projectors 22 are positioned so that their respective optical axes are tilted toward the center of the probe 28, thereby improving the overlap of the illumination fields of each structured light projector 22. The inventors found that by arranging almost all of the structured light projectors 22 in a single row, they could be more easily connected to the same flexible PCB.

[0295] Furthermore, multiple uniform light projectors 118, multiple near-infrared (NIR) light projectors 292, and diffractive optical elements (DOEs) 39 placed on each structured light projector 22 are shown in Figures 2D to 2E and will be described further below.

[0296] Next, we refer to Figure 3, a schematic diagram of a structured light projector 22 according to several application examples of the present invention. In some application examples, the structured light projector 22 includes a laser diode 36, a beam-shaping optical element 40, and a pattern-generating optical element 38 that generates a distribution of discrete, unconnected light spots 34 (see Figure 4 for further explanation). In some application examples, the structured light projector 22 may be configured such that when the laser diode 36 transmits light through the pattern-generating optical element 38, it generates a distribution of discrete, unconnected light spots 34 in all planes located between 1 mm and 30 mm, for example, between 4 mm and 24 mm, from the pattern-generating optical element 38. In some application examples, the distribution of discrete, unconnected light spots 34 is focused in one plane located between 1 mm and 30 mm, for example, between 4 mm and 24 mm, but discrete, unconnected light spots still exist in all other planes located between 1 mm and 30 mm, for example, between 4 mm and 24 mm. While the use of laser diodes was described above, please understand that this is an illustrative and non-limiting example. Other light sources may be used in other applications. Furthermore, while the projection of a pattern of discrete, unconnected light spots was described, please understand that this is an illustrative and non-limiting example. Other patterns or arrays of light, including but not limited to lines, grids, checkerboards, and other arrays, may be used in other applications. In some applications, the light pattern projected by the structured light projector is spatially fixed to one or more cameras.

[0297] This specification describes embodiments with reference to discrete light spots and with reference to performing operations using or based on spots. Examples of such operations include solving a corresponding algorithm to determine the location of the light spots, tracking the light spots, mapping projector rays to the light spots, identifying vulnerable light spots, and generating a three-dimensional model based on the location of the spots. It should be understood that such operations and other operations described with reference to spots also work for other features of other projection light patterns. Thus, the considerations herein with reference to spots also apply to other features of projection light patterns.

[0298] The pattern-generating optical element 38 may be configured to have an optical throughput efficiency of at least 80%, for example, at least 90% (i.e., the proportion of light that enters the pattern out of the total light that falls on the pattern-generating optical element 38).

[0299] In some applications, each laser diode 36 of each structured light projector 22 transmits light at different wavelengths; that is, each laser diode 36 of at least two structured light projectors 22 transmits light at two different wavelengths. In some applications, each laser diode 36 of at least three structured light projectors 22 transmits light at three different wavelengths. For example, red, blue, and green laser diodes may be used. In some applications, each laser diode 36 of at least two structured light projectors 22 transmits light at two different wavelengths. For example, in some applications, there are six structured light projectors 22 arranged in a probe 28, three of which include a blue laser diode and three of which include a green laser diode.

[0300] Next, refer to Figure 4, a schematic diagram of a structured light projector 22 that projects a distribution of discrete, unconnected light spots onto multiple object focal planes, according to several applications of the present invention. The object to be scanned 32 may be one or more teeth or other intraoral objects / tissues in the mouth of a subject. The somewhat translucent and glossy properties of teeth can affect the contrast of the projected structured light pattern. For example, (a) some of the light hitting the teeth may scatter to other areas in the intraoral scene, causing some stray light, and (b) some of the light may penetrate the teeth and then exit the teeth at any other point. Therefore, we have found that a sparse distribution 34 of discrete, unconnected light spots can provide an improved balance in reducing the amount of projected light while maintaining a useful amount of information, in order to improve image capture of intraoral scenes under structured light illumination without using contrast-enhancing means such as coating the teeth with opaque powder. The sparseness of the distribution 34 can be characterized by the following ratios: (a) The illuminated region on the orthogonal plane 44 in the illumination field α (alpha), i.e., the sum of the areas of all projection spots 33 on the orthogonal plane 44 in the illumination field α (alpha), (b) The ratio of the illuminated area α to the unilluminated area on the orthogonal plane 44. In some applications, the density ratio may be at least 1:150 and / or less than 1:16 (e.g., at least 1:64 and / or less than 1:36).

[0301] In some applications, each structured light projector 22 projects at least 400 discrete, unconnected spots 33 onto the three-dimensional surface of the oral cavity during scanning. In some applications, each structured light projector 22 projects fewer than 3,000 discrete, unconnected spots 33 onto the oral cavity surface during scanning. In order to reconstruct the three-dimensional surface from the projected sparse distribution 34, the correspondence between each projected spot 33 (or other features of the projection pattern) and the spots (or other features) detected by the camera 24 must be determined, as further described below with reference to Figures 7-19.

[0302] In some applications, the pattern-generating optical element 38 is a diffractive optical element (DOE) 39 (Figure 3), and when a laser diode 36 transmits light to an object 32 through the DOE 39, it generates a distribution 34 of discrete, unconnected light spots 33. As used throughout this specification, including the claims, a light spot is defined as a small region of light having any shape. In some applications, each DOE 39 of different structured light projectors 22 generates spots of different shapes; that is, all spots 33 generated by a particular DOE 39 have the same shape, and the shape of a spot 33 generated by at least one DOE 39 is different from the shape of a spot 33 generated by at least one other DOE 39. For example, some DOE 39 may generate circular spots 33 (as shown in Figure 4), some DOE 39 may generate square spots, and some DOE 39 may generate elliptical spots. Optionally, some DOE39s may generate linear patterns that are connected or disconnected.

[0303] Next, refer to Figures 5A-B, schematic diagrams of structured light projectors 22 according to several applications of the present invention, which include a beam shaping optical element 40 and an additional optical element, e.g., a DOE 39, positioned between the beam shaping optical element 40 and the pattern generating optical element 38. Optionally, the beam shaping optical element 40 is a collimating lens 130. The collimating lens 130 may be configured to have a focal length of less than 2 mm. Optionally, the focal length may be at least 1.2 mm. In some applications, an additional optical element 42 positioned between the beam shaping optical element 40 and the pattern generating optical element 38, e.g., a DOE 39, generates a Bessel beam when the laser diode 36 transmits light through the optical element 42. In some applications, a Bessel beam is transmitted through a wide range of orthogonal planes 44 (e.g., between 1 mm and 30 mm from DOE 39, e.g., each orthogonal plane located between 4 mm and 24 mm from DOE 39, etc.) such that all discrete, unconnected optical spots 33 maintain small diameters (e.g., less than 0.06 mm, e.g., less than 0.04 mm, e.g., less than 0.02 mm). In the context of this patent application, the diameter of the spots 33 is defined by the full width at half maximum (FWHM) of the spot intensity.

[0304] Despite the above explanation that all spots are smaller than 0.06 mm, some spots with diameters near the upper end of these ranges (e.g., slightly smaller than 0.06 mm, or 0.02 mm) that are close to the edge of the illumination field of the projector 22 may elongate when they intersect a geometric plane orthogonal to DOE 39. In such cases, it is useful to measure their diameters when they intersect the inner surface of a geometric sphere centered at DOE 39 and having a radius of 1 mm to 30 mm corresponding to the distance of each orthogonal plane located between 1 mm and 30 mm from DOE 39. As used throughout this application, including in the claims, the term “geometric” is considered to relate to a theoretical geometric construct (such as a plane or sphere) and not to any physical device.

[0305] In some applications, when a Bessel beam is transmitted through DOE39, spots 33 with a diameter greater than 0.06 mm are generated in addition to spots with a diameter of less than 0.06 mm.

[0306] In some applications, the optical element 42 is an axicon lens 45, shown in Figure 5A and further described below with reference to Figures 23A-B. Alternatively, the optical element 42 may be an annular aperture ring 47, as shown in Figure 5B. Maintaining a small spot diameter improves 3D resolution and accuracy across the entire depth of field. Without the optical element 42, such as the axicon lens 45 or annular aperture ring 47, the size of the spot 33 may change, for example, become larger, as it moves further away from the best focal plane due to diffraction and defocusing.

[0307] Next, refer to Figures 6A-B, schematic diagrams of a structured light projector 22 that projects discrete, unconnected spots 33 and a camera sensor 58 that detects the spots 33', according to some applications of the present invention. In some applications, a method is provided for determining the correspondence between projected spots 33 on the oral cavity surface and detected spots 33' on each camera sensor 58. As previously mentioned, this method also applies to the step of determining the correspondence between other projected features on the oral cavity surface and detected features on each camera sensor. Once the correspondence is determined, a three-dimensional image of the surface is reconstructed. Each camera sensor 58 has a pixel array, and for each of them there exists a corresponding camera ray 86. Similarly, for each projected spot 33 from each projector 22, there exists a corresponding projector ray 88. Each projector ray 88 corresponds to a path 92 of each pixel on at least one camera sensor of the camera sensor 58. Therefore, when the camera views a spot 33' projected by a particular projector ray 88, that spot 33' will inevitably be detected by pixels on a specific path 92 of the pixel corresponding to that particular projector ray 88. Referring particularly to Figure 6B, the correspondence between each projector ray 88 and each camera sensor path 92 is shown. Projector ray 88' corresponds to camera sensor path 92', projector ray 88'' corresponds to camera sensor path 92'', and projector ray 88'' corresponds to camera sensor path 92''. For example, if a particular projector ray 88 projects a spot into a dusty space, it will illuminate lines of dust in the air. The lines of dust detected by the camera sensor 58 will follow the same path on the camera sensor 58 as the camera sensor path 92 corresponding to the particular projector ray 88.

[0308] During the calibration process, calibration values ​​are stored based on camera rays 86 corresponding to pixels on each camera sensor 58 of camera 24, and projector rays 88 corresponding to the projection spots 33 (or other features) of light from each structured light projector 22. For example, calibration values ​​may be stored for (a) multiple camera rays 86 corresponding to each of multiple pixels on each camera sensor 58 of camera 24, and (b) multiple projector rays 88 corresponding to each of multiple projection spots 33 of light from each structured light projector 22. When used throughout this application, including the claims, the stored calibration values ​​representing camera rays corresponding to each pixel on each camera sensor mean (a) the values ​​given to each camera ray, or (b) the parameterized camera calibration model, e.g., the parameter values ​​of a function. When used throughout the entire application, including the claims, the stored calibration values ​​indicating the projector rays corresponding to each projected spot (or other projected feature) of light from each structured light projector mean (a) values ​​given to each projector ray, for example, in an indexed list, or (b) parameterized projector calibration models, for example, parameter values ​​of a function.

[0309] As an example, the following calibration process may be used. A high-precision dot target, for example, a black dot on a white background, is illuminated from below, and images of the target are captured with all cameras. Next, the dot target is moved perpendicular to the camera, i.e., along the z-axis, to the target surface. For all dots at all positions in the z-axis direction, the dot center is calculated, and a 3D grid of dots is created in space. Then, using distortion and the camera pinhole model, the pixel coordinates of each 3D position of each dot center are determined, thereby defining the camera ray for each pixel as the ray starting from that pixel and moving toward the corresponding dot center in the 3D grid. Camera rays corresponding to pixels between grid points can be interpolated. The above camera calibration procedure is repeated for all respective wavelengths of each laser diode 36 so that the camera ray 86 corresponding to each pixel on each camera sensor 58 for each wavelength is included in the stored calibration value. Alternatively, the stored calibration value is the parameter value of the distortion and camera pinhole model, indicating the value of the camera ray 86 corresponding to each pixel on each camera sensor 58 for each wavelength.

[0310] After camera 24 is calibrated and all camera ray values ​​86 are stored, the structured light projector 22 may be calibrated as follows: A flat, featureless target is used, and the structured light projector 22 is turned on one at a time. Each spot (or other feature) is positioned on at least one camera sensor 58. Since camera 24 is now calibrated, the 3D spot position of each spot (or other feature) is calculated by triangulation based on images of the spot (or other feature) from multiple different cameras. The above process is repeated with featureless targets positioned at multiple different z-axis positions. Each projected spot (or other feature) on the featureless target will define a projector ray in space originating from the projector.

[0311] Next, refer to Figure 7, a flowchart outlining a method for generating a digital three-dimensional image according to several applications of the present invention. In steps 62 and 64 of the method outlined in Figure 7, each structured light projector 22 is driven to project a pattern of light (e.g., a distribution 34 of discrete, unconnected light spots 33) onto the three-dimensional surface of the oral cavity, and each camera 24 is driven to capture an image containing at least a portion of the pattern (e.g., one of the spots 33). Based on stored calibration values ​​indicating (a) camera rays 86 corresponding to each pixel on the camera sensor 58 of each camera 24 and (b) projector rays 88 corresponding to each projected spot 33 of light from each structured light projector 22, the corresponding algorithm is executed in step 66 using a processor 96 (Figure 1), which is further described below with reference to Figures 8-12. In some embodiments, the processor 96 is a processor located in an elongated handheld wand 20. In some embodiments, the processor 96 is located in a computing device as described below with reference to Figures 60-61, which may be operably connected to an elongated handheld wand 20 (e.g., via a wired or wireless connection). In some embodiments, multiple processors are used, in which case one or more processors may be located in the elongated handheld wand and / or one or more processors may be located in the computing device. Once the correspondence is resolved, the three-dimensional position on the oral cavity surface is calculated in step 68 and used to generate a digital three-dimensional image of the oral cavity surface. Furthermore, capturing the oral cavity scene using multiple cameras 24 results in an improvement in signal-to-noise ratio in capture that is the square root of the number of cameras.

[0312] Next, refer to Figure 8, a flowchart illustrating the corresponding algorithm for step 66 in Figure 7, based on several application examples of the present invention. Based on the stored calibration values, all projector rays 88 and all camera rays 86 corresponding to all detection spots 33' are mapped (step 70), and all intersections 98 between at least one camera ray 86 and at least one projector ray 88 are identified (step 72). Figures 9 and 10 are schematic diagrams showing simplified examples of steps 70 and 72 in Figure 8, respectively. As shown in Figure 9, three projector rays 88 are mapped along with eight camera rays 86 corresponding to a total of eight detection spots 33' on the camera sensor 58 of the camera 24. As shown in Figure 10, 16 intersections 98 are identified.

[0313] In steps 74 and 76 of Figure 7, the processor 96 determines the correspondence between the projection spots 33 and the detection spots 33' in order to determine the three-dimensional position of each projection spot 33 on the surface. Figure 11 is a schematic diagram illustrating steps 74 and 76 of Figure 8 using the simplified example described in the preceding paragraph. For a given projector ray i, the processor 96 "sees" the corresponding camera sensor path 90 on one of the camera sensors 58 of the camera 24. Each detection spot j along the camera sensor path 90 will have a camera ray 86 that intersects with the given projector ray i at an intersection point 98. The intersection point 98 defines a three-dimensional point in space. The processor 96 then "looks" at the camera sensor path 90' corresponding to a given projector ray i on each camera sensor 58' of the other cameras 24 and determines how many of the other cameras 24 similarly detected each spot k where the camera ray 86' intersects with the same three-dimensional point in space defined by the intersection point 98 on their respective camera sensor paths 90' corresponding to the given projector ray i. This process is repeated for all detected spots j along the camera sensor path 90, and the spot j that the most cameras 24 "agree" on is identified as the spot 33 (Figure 12) projected onto the surface from the given projector ray i. That is, the projector ray i is identified as the specific projector ray 88 that produced the detected spot j where the most of the other cameras detected each spot k. The three-dimensional position on the surface is calculated for that spot 33. The same process may be performed to calculate the three-dimensional positions on the surface of other features of the projection pattern.

[0314] In one example, as shown in Figure 11, all four cameras detect the respective spots where their respective camera rays intersect with the projector ray i at intersection point 98 on their respective camera sensor paths corresponding to the projector ray i, and intersection point 98 is defined as the intersection of the camera ray 86 corresponding to the detected spot j and the projector ray i. Therefore, it can be said that all four cameras "agree" that the spot 33 projected by the projector ray i is located at intersection point 98. However, when this process is repeated for the next spot j', none of the other cameras detect the respective spots where their respective camera rays intersect with the projector ray i at intersection 98' on their respective camera sensor paths corresponding to the projector ray i. Intersection 98' is defined as the intersection of camera ray 86" (corresponding to detected spot j') and projector ray i. Therefore, only one camera "agrees" that the spot 33 (or other feature) projected by the projector ray i is present at intersection 98', while all four cameras "agree" that the spot 33 (or other feature) projected by the projector ray i is present at intersection 98. Thus, the projector ray i is identified as the specific projector ray 88 that generated the detected spot j by projecting the spot 33 (or other feature) onto the surface at intersection 98 (Figure 12). As shown in step 78 of Figure 8, and also as shown in Figure 12, the three-dimensional position 35 on the oral cavity surface is calculated at the intersection point 98.

[0315] Next, refer to Figure 13, a flowchart outlining further steps of the correspondence algorithm according to some application examples of the present invention. Once the position 35 on the surface is determined, the projector ray i that projected spot j, as well as all camera rays 86 and 86' corresponding to spot j and each of the spots k, are excluded from consideration (step 80), and the correspondence algorithm is run again for the next projector ray i (step 82). Figure 14 depicts the simplified example described above after excluding the specific projector ray i that projected spot 33 onto position 35. As in step 82 of the flowchart in Figure 13, the correspondence algorithm is then run again for the next projector ray i. As shown in Figure 14, the remaining data shows that the three cameras "agree" that spot 33 at the intersection 98 of the camera rays 86 corresponds to the detected spot j and the projector ray i, and the intersection 98 is defined by the intersection of the camera ray 86 corresponding to the detected spot j and the projector ray i. Therefore, as shown in Figure 15, the three-dimensional position 37 is calculated at the intersection point 98.

[0316] As shown in Figure 16, once the three-dimensional position 37 on the surface is determined, the projector ray i that projected spot j, as well as all camera rays 86 and 86' corresponding to spot j and each of the spots k, are removed from consideration. The remaining data indicates the spot 33 projected by the projector ray i at intersection 98, and the three-dimensional position 41 on the surface is calculated at intersection 98. As shown in Figure 17, according to the simplified example, the three projected spots 33 of the three projector rays 88 of the structured light projector 22 are now located at three-dimensional positions 35, 37, and 41 on the surface. In some applications, each structured light projector 22 projects 400 to 3000 spots 33. Once the correspondences for all projector rays 88 are solved, a reconstruction algorithm can be used to reconstruct a digital image of the surface using the calculated three-dimensional positions of the projected spots 33.

[0317] Next, refer to Figure 28, a flowchart outlining the steps of a method hereinafter referred to as "spot tracking" for generating digital three-dimensional images, according to several applications of the present invention. Although this method is called "spot tracking," it can also be applied to tracking other types of features of a projection pattern. Due to the motion of the handheld intraoral scanner relative to the intraoral surface during scanning, the projection points move across the intraoral surface. The inventors have noticed that if the movement of a particular detected spot (or other feature) can be tracked in a series of image frames, then a correspondence solved for that particular spot (or other feature) in any of the frames in which the spot (or other feature) is tracked will give a solution for the correspondence for that spot (or other feature) in all of the frames in which the spot (or other feature) is tracked. That is, if the processor 96 understands that a given detection spot 33' was projected by a given projector ray 88 in a given frame, and if the processor 96 determines through spot tracking that the detection spot 33' in the next image is the same spot, the processor 96 automatically determines that it has generated a tracking spot in the next image in which the same projector ray 88 was detected.

[0318] Since the detection spot 33', which can be tracked across consecutive images, is generated by the same specific projector ray, the trajectory of the tracking spot will follow a specific camera sensor path 90 corresponding to that specific projector ray 88. When the correspondence is resolved for the detection spot 33' at one point along the specific camera sensor path 90, the three-dimensional position on the surface for all points along the camera sensor path 90 where the spot 33' was detected can be calculated. That is, the processor can calculate the three-dimensional position on the three-dimensional surface of the oral cavity at the intersection of the specific projector ray 88 that generated the detection spot 33' and the respective camera rays 86 corresponding to the tracking spot, for each of the multiple consecutive images in which the spot 33' was tracked. This can be particularly useful in situations where a specific detection spot is only visible by one camera (or a few cameras) in a particular image frame. If the specific detection spot is seen by other cameras 24 in previous consecutive image frames, and a correspondence is established for that specific detection spot in those previous image frames, then even in image frames where the specific detection spot is seen by only one camera 24, the processor knows which projector ray 88 generated the spot and can determine the three-dimensional position of the spot on the three-dimensional surface of the oral cavity.

[0319] For example, hard-to-reach areas within the oral cavity can be imaged by only a single camera 24. In this case, if the detected spot 33' on the camera sensor 58 of the single camera 24 can be tracked through a series of previous images, the three-dimensional position of the spot on the surface (even if it was only seen by the single camera 24) can be calculated based on the information obtained from the tracking, namely (a) which camera sensor path the spot is moving along, and (b) which projector ray generated the tracked spot 33'.

[0320] In step 180 of the method outlined in Figure 28, each structured light projector 22 is driven to project a pattern of light onto the three-dimensional surface of the oral cavity, which in one embodiment is a distribution 34 of discrete, unconnected light spots 33; and in step 182, each camera 24 is driven to capture an image that includes at least a portion of the projected pattern (e.g., at least one of the spots 33). In one embodiment, based on stored calibration values ​​indicating (a) camera rays 86 corresponding to each pixel on the camera sensor 58 of each camera 24 and (b) projector rays 88 corresponding to each projection spot 33 of light from each structured light projector 22, step 184 uses a processor 96 to compare a series of images (e.g., multiple consecutive images) captured by each camera 24 to determine which features of the projection pattern (e.g., which of the projection spots 33) can be tracked across multiple images, where each tracked feature (e.g., spot 33s') moves along the path p of the pixel corresponding to a specific camera sensor path 90 corresponding to each projector ray, e.g., projector ray 88. In some applications, step 186, the processor 96 calculates the three-dimensional position of each tracked feature (e.g., spot 33s') on the three-dimensional surface of the oral cavity in a series of images (e.g., in each of the consecutive images).

[0321] In one embodiment, each of the structured light projectors of one or more structured light projectors is driven to project a pattern onto the three-dimensional surface of the oral cavity. Furthermore, each of the cameras of one or more cameras is driven to capture multiple images, each image including at least a portion of the projection pattern. The projection pattern may include multiple projection light spots, and portions of the projection pattern may correspond to projection spots of the multiple projection light spots. The processor 96 then compares a series of images captured by one or more cameras, determines which portions of the projection pattern can be tracked across the series of images based on the comparison of the series of images, and constructs a three-dimensional model of the three-dimensional surface of the oral cavity, at least partially based on the comparison of the series of images. In one embodiment, the processor solves a correspondence algorithm for the tracking portion of the projection pattern in at least one image of the series of images, uses the solved correspondence algorithm to address the tracking portion of the projection pattern in at least one image of the series of images, for example, solves a correspondence algorithm for the tracking portion of the projection pattern in an image of a series of images for which the correspondence algorithm has not been solved, and constructs a three-dimensional model using the solution of the correspondence algorithm. In one embodiment, the processor solves a correspondence algorithm for the tracking portion of a projection pattern based on the position of the tracking portion in each image of the entire series of images, and constructs a three-dimensional model using the solution of the correspondence algorithm. In one embodiment, the processor compares a series of images based on stored calibration values ​​indicating (a) camera rays corresponding to each pixel on the camera sensor of each camera of one or more cameras, and (b) projector rays corresponding to each projection light spot from each structured light projector of one or more structured light projectors, where each projector ray corresponds to a projector ray corresponding to the path of each pixel on at least one camera sensor of the camera sensor, and each tracking spot s moves along the path of the pixel corresponding to the respective projector ray r.

[0322] Next, we refer to Figure 29, a schematic diagram illustrating several detection spots 33' (specifically 33a' and 33b') in some applications of the present invention, and how the processor 96 may determine which set of detection spots 33' can be considered tracked (step 184 in Figure 28). Spot 33a' is the detection spot 33' in the previous image, and spot 33b' is the detection spot 33' in the current image. The processor 96 performs a search within the search radius to detect possible matches between spots 33a' and 33b' that are considered to be the same spot 33' tracked between the two images.

[0323] The inventors have identified three typical factors that can influence how much the spot moves between frames. 1. The distance a spot travels between frames is typically inversely proportional to the camera's frame rate. That is, if the camera's frame rate is very fast, the spot will appear to travel a relatively small distance between consecutive frames, while if the camera's frame rate is slow, the spot will appear to travel a greater distance between consecutive frames. How fast the wand moves relative to the oral cavity surface also affects how far the spot travels between frames; that is, if the wand moves faster, the spot will appear to travel a greater distance between consecutive frames. 2. The distance the spot travels between frames typically depends on the degree of inclination of the oral surface being scanned, which is relative to the projector 22 and / or camera 24. If the spot is projected onto an inclined surface and the spot is moving in the direction of the inclination, the corresponding detection spot on the camera sensor will move faster, and therefore the tracking spot will travel further between consecutive sets of frames. 3. From the camera's field of view, how far the spot moves between frames typically depends on the distance between the scanning surface and the projector. When the surface is close to the projector, even small movement of the projector will cause a large movement of the tracking spot on the camera 24's sensor 58 between consecutive frames. In contrast, when the surface is far away, the same movement of the projector will cause less movement of the tracking spot on the camera 24's sensor 58 between consecutive frames. For example, if the surface is approaching infinity from the projector, the projector's movement will cause almost zero movement of the tracking spot on the camera 24's sensor 58.

[0324] In some applications, the processor 96 performs the search within a fixed search radius of at least 3 pixels and / or less than 10 pixels (e.g., 5 pixels). In some applications, the processor 96 calculates the search radius considering parameters such as the level of spot position error which may be determined during calibration. For example, the search radius is 2 * (Spot position error) or 3 * It may also be defined as (spot position error).

[0325] In the simplified example shown in Figure 29, spots 33a' and 33b' in each set 112 are considered to be close enough to each other that they are considered to be the same projection spot 33 moving through the two images. That is, for each set 112, the two detected spots 33a' and 33b' are considered to be the same tracking spot 33s'. In contrast, spots 33a' and 33b' in set 114 are too far apart to be considered tracked. Spots 33a' and 33b' in set 116 are close enough, but are not considered tracked because there are two or more matches. As will be further explained below, in set 116, a pair of tracking spots exists, and continuing to analyze more images may help determine which spots are actually tracking spots.

[0326] In one embodiment, to generate a digital three-dimensional image, an intraoral scanner drives each of one or more structured light projectors to project a pattern of light onto the intraoral three-dimensional surface. The intraoral scanner further drives each of a plurality of cameras to capture an image, the image including at least a portion of the projected pattern, and each of the plurality of cameras includes a camera sensor including a pixel array. The intraoral scanner further uses a processor to execute a corresponding algorithm to calculate the respective three-dimensional positions on the intraoral three-dimensional surface of multiple features of the projected pattern. The processor uses data from a first camera among the plurality of cameras, e.g., data from at least two cameras, to identify candidate three-dimensional positions of a given feature of the projected pattern corresponding to a particular projector ray r, while data from a second camera among the plurality of cameras, e.g., another camera other than one of the at least two cameras, is not used to identify its candidate three-dimensional position. The processor further uses the candidate three-dimensional positions seen by the first camera to identify a search space on the pixel array of the second camera for searching for features of the projected pattern from the projector ray r. If a feature of the projection pattern from a projector ray r is identified within the search space, the processor uses data from a second camera to refine candidate three-dimensional locations of the projection pattern feature. In one embodiment, the light pattern includes a distribution of discrete, unconnected light spots, and the projection pattern feature includes projection spots from the unconnected light spots. In one embodiment, the processor uses stored calibration values ​​indicating (a) camera rays corresponding to each pixel on the camera sensor of each camera of a plurality of cameras, and (b) projector rays corresponding to each feature of the projection light pattern from each structured light projector of one or more structured light projectors, where each projector ray corresponds to the path of each pixel on at least one camera sensor of the camera sensor.

[0327] Next, refer to Figure 30, a flowchart outlining a method for determining tracking features (e.g., tracking spot 33s') according to several applications of the present invention. Although Figure 30 discusses tracking spots with reference, the same applies to other types of tracking features. In some applications, in addition to searching for tracking spots by monitoring the proximity of spots in consecutive images, the processor 96 may also search for tracking spots 33s' based on the parameters(s) of the detected spot 33', which will be referred to below as "parametric tracking". The processor 96 determines the parameters of the detected spot 33' in the first of the consecutive images (step 188) and in adjacent images. The processor 96 then uses the determined parameters of the detected spot 33' in the two adjacent images to predict the same parameters of the spot in subsequent images, e.g., the next image (and subsequent images) (step 190). The processor 96 searches for spots in subsequent images, e.g., the next image, that substantially have the predicted parameters (step 192). For example, two specific detection spots 33a' and 33b' may be determined to originate from the same projector ray in two adjacent frames, either through the correspondence algorithm described above or through proximity tracking as described in the previous two paragraphs. Once the processor 96 understands that the detection spots 33a' and 33b' in two adjacent frames were caused by the same projector ray 88, the processor 96 can determine the parameters of the spots and predict the parameters of the spots in a later image, for example, the next image, based on the parameters of the spots in the two adjacent images.

[0328] In some applications, the spot parameters are the spot size, spot shape, and / or the spot aspect ratio, spot orientation, spot intensity, and / or spot signal-to-noise ratio (SNR). For example, if the determined parameter is the shape of the tracking spot 33s', the processor 96 predicts the shape of the tracking spot 33s' in a later image, e.g., the next image, and based on the predicted shape of the tracking spot 33s' in the later image, determines a search space in the later image for searching for the tracking spot 33s', e.g., a search space having a size and aspect ratio (e.g., within twice) based on the size and aspect ratio of the predicted shape of the tracking spot 33s'. In some applications, the spot shape may refer to the aspect ratio of an elliptical spot.

[0329] Refer again to Figure 29. In some applications, parametric tracking can help resolve ambiguities, such as those shown in set 116 of spots in Figure 29. As described above, spots 33a' and 33b' in set 116 are close enough to be considered a tracking spot, but two or more matches are found. Then, based on parametric tracking, processor 96 may be able to determine which spot in set 116 is actually the tracking spot.

[0330] Next, refer to Figure 31, a flowchart outlining a method for detecting tracking spot 33s' in a later image according to several applications of the present invention. Figure 31 also applies to detecting other tracking features in a later image. In some applications, based on the direction and distance the tracking spot 33s' has traveled between two images (e.g., between two consecutive images), the processor 96 determines the velocity vector of the tracking spot 33s' (step 194). The processor 96 then uses the velocity vector to determine the search space in a later image, e.g., the next image, for searching for the tracking spot 33s' (step 196).

[0331] In some applications, the search space in later images may be determined by estimating the new location of tracking spot 33s' using a predictive filter, such as a Kalman filter.

[0332] Next, refer to Figure 32, a flowchart outlining a method for detecting a tracking spot 33s' in a later image according to several applications of the present invention. Figure 32 also applies to detecting other tracking features in a later image. The determination of a velocity vector for the tracking spot 33s' may be used to assist in determining the search space for finding the tracking spot; that is, if the spot is moving faster, it has traveled further between consecutive frames, and therefore the processor 96 can set a larger search space for finding the tracking spot. We have noticed that the shape of the tracking spot 33s' and the direction in which it is moving may indicate the velocity of the spot. For example, in one application, if a spot projected as a circle appears elliptical, it is likely that the spot is descending onto a plane inclined with respect to the projector 22 and / or camera 24. Similarly, as described above, a spot moving along an inclined surface in the direction of the inclination moves faster than when it is not moving in the direction of the inclination, and the steeper the inclination of the surface, the faster the spot will move. Furthermore, if a spot appears stretched into an ellipse due to an inclined surface, it is highly likely that it appears stretched in the direction of the inclination, i.e., the major axis of the ellipse appears to be in the direction of the inclination. Therefore, movement of an elliptical spot along its major axis indicates that the spot is moving up or down and inclined, and thus indicates that it is moving faster than if the elliptical spot were moving along its minor axis (it may indicate that it is projected onto an inclined surface but is not moving in the direction of the inclination).

[0333] Therefore, in some applications, after determining the shape of the tracking spot 33s' (step 198), the processor 96 can determine the velocity vector of the tracking spot 33s' based on the direction and distance it has moved between two consecutive images (step 200). The processor 96 may then use the determined velocity vector and / or the shape of the tracking spot 33s' to predict the shape of the tracking spot 33s' in a later image, e.g., the next image (step 202). Following the prediction of the shape of the tracking spot 33s', the processor 96 may use the combination of the velocity vector and the predicted shape of the tracking spot 33s' to determine the search space in a later image, e.g., the next image, for searching for the tracking spot 33s'. Referring again to the above example of an elliptical spot, if the shape of the spot is determined to be elliptical and the spot is determined to be moving along its major axis, a larger search space will be specified compared to the case where an elliptical spot is moving along its minor axis.

[0334] Next, we refer to Figure 33, a schematic diagram illustrating an example in which spot tracking helps identify a detected spot 33' as projected from a specific projector ray 88, according to some applications of the present invention. This is useful when the correspondence algorithm does not provide a solution for a specific detected spot 33', for example, when the detected spot 33' in a particular frame is seen by only one camera 24. In such a case, once the detected spot 33' is identified as a tracking spot 33s' moving along a specific camera sensor path 90 of a pixel corresponding to a specific projector ray 88, it can be assumed that the specific projector ray 88 projected the spot, as described above. Based on the solved correspondence for the tracking spot 33s' in the previous frame, the correspondence can be solved for a frame in which only one camera detected the spot 33'. Therefore, in some application examples, after executing a correspondence algorithm such as the correspondence algorithm described above in relation to Figures 7-17, if it is determined that the detected spot 33' is a tracking spot 33s' that moves along a specific path 90 on the camera sensor 58 corresponding to a specific projector ray 88, it can be assumed that the specific projector ray 88 generated the detected spot 33'. That is, the processor 96 can identify the detected spot 33' as originating from a specific projector ray 88 by identifying it as a tracking spot 33s' that moves along a path 90 of the pixels of the camera sensor 58 corresponding to the specific projector ray 88.

[0335] In the example shown in Figure 33, in two consecutive frames captured at time 1 and time 2, each of the two camera sensors 58 detected a projected spot 33' from the projector ray 88. The correspondence algorithm described above determined that the projector ray 88 generated the detected spot 33' in frame 1 and frame 2, and solved the correspondence between the detected spot 33' in frame 1 and frame 2. However, in the third frame, only one camera detects the spot 33'. In this example, it is assumed that the correspondence algorithm could not solve the correspondence for the detected spot 33' in frame 3. However, the processor 96 determines that the detected spot 33' in frame 3 is a tracking spot 33s' that moves along the same projector ray 88 that generated the spot 33' in frame 1 and frame 2. Therefore, the processor 96 identifies the detected spot 33' in frame 3 as being generated by the projector ray 88.

[0336] Next, we refer to Figures 34A-B, which are simplified schematic diagrams of a camera sensor 58 showing two detection spots 33c' and 33d' according to some applications of the present invention. In Figure 34A, there is ambiguity regarding which projector ray generated each detection spot, i.e., which pixel of the camera sensor path 90 the detection spot corresponds to. In some applications, the processor 96 can resolve these ambiguities using spot tracking. Thus, in some applications, if, after executing a correspondence algorithm (such as the correspondence algorithm described above with respect to Figures 7-17), the detection spot 33' is identified as originating from two different candidate projector rays 88 and 88' based on the three-dimensional position calculated by the correspondence algorithm, the processor 96 may identify that the detection spot 33' originates from only one of the two different candidate projector rays 88 and 88' by identifying it as a tracking spot 33s' moving along either path 90 or path 90'.

[0337] An example of this type of ambiguity is represented by the detection spot 33c' in Figure 34A. Spot 33c' is located at the intersection of two different paths 90c and 90c'. The corresponding algorithm may know that such a detection spot 33c' is generated by both the projector ray 88 corresponding to path 90c and the projector ray 88' corresponding to path 90c'. Another type of this type of ambiguity is represented by the detection spot 33d' in Figure 34A. Spot 33d' is very close to two different paths 90d and 90d', but is not at the intersection between paths 90 and 90d. Due to signal noise, it may have been unclear in the corresponding algorithm whether spot 33d' was generated by the projector ray 88 corresponding to path 90d or by the projector ray 88' corresponding to path 90d'.

[0338] As shown in Figure 34B, the processor 96 may identify which projector ray generated each of the spots 33c' and 33d' by identifying spots 33c' and 33d' as tracking spots 33s' that move along a specific path. Detected spot 33c' is identified as a tracking spot 33s' that moves along path 90c', and therefore detected spot 33c' is identified as being generated by projector ray 88' corresponding to path 90c'. Detected spot 33d' is identified as a tracking spot 33s' that moves along path 90d', and therefore detected spot 33d' is identified as being generated by projector ray 88' corresponding to path 90d'.

[0339] Next, refer to Figure 35, a flowchart outlining additional or alternative methods in which spot tracking may be used, according to some applications of the present invention. The concepts shown in Figure 35 also apply to methods using other feature tracking. In some applications, the processor 96 may use spot tracking to exclude falsely detected spots 33' from being considered as points on the three-dimensional surface of the oral cavity. In some applications, after executing a correspondence algorithm such as the correspondence algorithm described above with reference to Figures 7-17 (step 206), the processor 96 may identify the detected spots 33' as originating from a specific projector ray 88 based on the correspondence algorithm (step 208). Step 206 typically follows step 186 of the method outlined in the flowchart of Figure 28. Furthermore, the processor 96 may identify a series of detected spots across multiple consecutive images, all of which are traced spots 33s' moving along a path 90 of pixels corresponding to the same specific projector ray 88. If the detection spot 33' is one of the tracking spots 33s', as shown by the determination diamond 210, then the detection spot 33' may be considered a point on the three-dimensional surface of the oral cavity (step 212). However, if the detection spot 33' is not identified as a tracking spot 33s' moving along the path 90 of the pixel corresponding to that particular projector ray 88, then the detection spot 33' may be assumed to be a false detection spot and may be excluded from being considered a point on the three-dimensional surface of the oral cavity (step 214).

[0340] Next, refer to Figure 36, a flowchart outlining additional or alternative methods in which spot tracking may be used, according to some applications of the present invention. The concepts shown in Figure 36 also apply to methods using other feature tracking. To reduce the occurrence of numerous false detections of spots on camera 24, processor 96 may set an intensity threshold so that any detected spots 33' below the threshold are not included as candidate spots in the corresponding algorithm. However, this may also result in false detections, i.e., spots that may have provided useful information, not being considered because they are weak spots (have an intensity below the threshold). For example, in hard-to-capture areas of an intraoral scene, some of the projected spots 33 may appear below the intensity threshold. Therefore, in some applications, after running a corresponding algorithm such as the corresponding algorithm described herein with reference to Figures 7-17 (step 216), processor 96 may identify weak spots 33' whose 3D position was not calculated by the corresponding algorithm (step 218) by, for example, lowering the intensity threshold and considering spots 33' that were not considered by the corresponding algorithm. If, as shown by the determination diamond 220, the vulnerable spot 33' is identified as a tracking spot 33s' that moves along the path 90 of the pixel corresponding to a particular projector ray 88, then the vulnerable spot 33' is identified as being projected from that particular projector ray 88 and is considered a point on the intraoral three-dimensional surface (step 222). If the vulnerable spot is not identified as a tracking spot 33s', then the vulnerable spot is excluded from being considered as a point on the intraoral three-dimensional surface (step 224).

[0341] In some applications, for a tracking spot 33s', the processor 96 may determine multiple possible camera sensor paths 90 of the pixels that the tracking spot 33s' is moving through, with each of the multiple possible projector rays 88 corresponding to a set of multiple possible projector rays 88. For example, the multiple projector rays 88 may closely correspond to the paths 90 of pixels on the camera sensor 58 of a given camera. The processor 96 may execute a corresponding algorithm to determine which of the possible projector rays 88 generated the tracking spot 33s' in order to calculate its three-dimensional position on the surface relative to each of its positions.

[0342] For a given camera sensor 58, for each of the multiple possible projector rays 88, the three-dimensional point in space lies at the intersection of each possible projector ray 88 and the camera ray corresponding to the tracking spot 33s' detected by the given camera sensor 58. For each possible projector ray 88, the processor 96 considers the camera sensor paths 90 corresponding to the possible projector ray 88 in each of the other camera sensors 58 and identifies how many other camera sensors 58 whose camera rays intersect with that three-dimensional point in space have similarly detected the spot 33' on their respective camera sensor paths 90 corresponding to that possible projector ray 88, i.e., how many other cameras agree that the tracking spot 33s' is projected by that projector ray 88. This process is repeated for all possible projector rays 88 corresponding to the tracking spot 33s'. The projector ray 88 that is most likely to be agreed upon by the majority of other cameras is determined to be the specific projector ray 88 that generated the tracking spot 33s'. Once the specific projector ray 88 for the tracking spot 33s' is determined, the camera sensor path 90 through which the spot is moving is identified, and for each of the consecutive images in which the spot 33s' is tracked, the three-dimensional position on the surface at the intersection of the specific projector ray 88 corresponding to the tracking spot 33s' and the respective camera ray is calculated.

[0343] Next, refer to Figures 37A and 37B, schematic diagrams showing points used for 3D reconstruction before and after the processor 96 performs spot tracking, according to several application examples of the present invention. The size of each data point represents how many cameras were used to solve for that point; that is, the larger the point, the more cameras that viewed that spot. Note that the size of the data points in the figure is used only to distinguish how many cameras viewed a given point, and does not represent the size of the projected spot on the surface. In Figure 37A, there are many small points located in the periphery that do not appear to be points on the oral cavity surface. These small points refer to detection spots that were viewed by a very small number of cameras, for example, only one, yet were assigned a 3D position in space based on the correspondence algorithm. After executing the correspondence algorithm, the processor 96 performs spot tracking and therefore may determine that these light points in the periphery are actually false positives (by determining that they are not tracking spots). Therefore, as shown in Figure 37B, after spot tracking, most of the spots that spot tracking determined to be false positives are removed from consideration as points on the oral cavity surface.

[0344] Next, refer to Figure 38, a flowchart outlining the steps of a method for generating a digital three-dimensional image, referred to herein as “ray tracing,” according to some applications of the present invention. In some applications, instead of or in addition to tracing detection spots and / or other features in a two-dimensional image (as described above), the length of each projector ray 88 may be traced in three-dimensional space. The length of the projector ray 88 is defined as the distance between the origin of the projector ray 88, i.e., the light source, and the three-dimensional position where the projector ray 88 intersects the oral cavity surface.

[0345] In step 226 of the method outlined in Figure 38, each structured light projector 22 is driven to project a pattern of light, such as a distribution 34 of discrete, unconnected light spots 33, onto the three-dimensional surface of the oral cavity. In step 228, each camera 24 is driven to capture multiple images, each image containing at least one feature of the projection pattern (e.g., at least one of the spots 33). Although this method has been described with reference to spots, it also works with other types of features. In step 230, the processor 96 is used to execute a correspondence algorithm, such as the correspondence algorithm described above with reference to Figures 7-17, based on stored calibration values ​​indicating (a) camera rays 86 corresponding to each pixel on the camera sensor 58 of each camera 24, and (b) projector rays 88 corresponding to each projected spot 33 of light from each structured light projector 22. As a result of the correspondence algorithm, each resolved projector ray 88 in each image frame yields a reconstructed three-dimensional point in space, which then defines the length of the resolved projector ray 88 in that frame.

[0346] Accordingly, in step 232, in at least a subset of the captured images, for example, in a series of images or a number of consecutive images, the processor 96 identifies that the calculated 3D position of the detected spot 33' (calculated from the corresponding algorithm) corresponds to a particular projector ray 88. In step 234, based on each 3D position corresponding to the projector ray 88 in the subset of images, the processor 96 evaluates, for example, calculates, the length of the projector ray 88 in each image of the subset of images. Because the camera 24 captures images at a relatively high frame rate, for example, about 100 Hz, the geometric shape of the spot seen by each camera does not change significantly between frames. Accordingly, if the evaluated, for example, calculated, length of the projector ray 88 is tracked and plotted against time, the data points will follow a relatively smooth curve, although some discontinuities may occur as will be discussed further below. Thus, the length of the projector ray over time forms a relatively smooth univariate function against time. As described above, the detection spot 33' corresponding to the projector ray 88 across multiple consecutive images will appear to move along a one-dimensional line which is the path 90 of the camera sensor pixels corresponding to the projector ray 88.

[0347] In one embodiment, a method for generating a digital three-dimensional image includes the steps of driving each of one or more structured light projectors to project a pattern onto a three-dimensional surface of the oral cavity, and driving each of one or more cameras to capture an image, the image including at least a portion of the pattern. The method further includes the step of using the processor to execute a corresponding algorithm to calculate the respective three-dimensional positions of a plurality of features of the pattern on the three-dimensional surface of the oral cavity captured in a series of images. The processor further identifies the calculated three-dimensional positions of the detected features of the captured pattern in at least a subset of the series of images as corresponding to one or more specific projector rays r. Based on the three-dimensional positions of the detected features corresponding to the one or more projector rays r in the subset of images, the processor evaluates, for example, calculates the length associated with one or more projector rays r in each image of the subset of images. In one embodiment, the processor calculates the estimated length of one or more projector rays r in at least one of a series of images in which the three-dimensional position of features projected from one or more projector rays r has not been identified. In one embodiment, each of the one or more cameras comprises a camera sensor including a pixel array, and the calculation of the three-dimensional position of each of a plurality of features of a pattern on the three-dimensional surface of the oral cavity, and the identification that the calculated three-dimensional position of the detected features of the pattern corresponds to a particular projector ray r, is performed based on stored calibration values ​​indicating (i) a camera ray corresponding to each pixel of the camera sensor of each of the one or more cameras, and (ii) a projector ray corresponding to each feature of the projected light pattern from each of the one or more projectors, wherein each projector ray corresponds to the path of each pixel on at least one of the camera sensors. In one embodiment, the pattern comprises a plurality of spots, and each of the plurality of features of the pattern comprises one of the plurality of spots.

[0348] Next, refer to Figure 39, a graph showing the time-series tracking of the length of a projector ray 88 and a specific simplified diagram of the camera sensor 58 corresponding to a particular image frame, according to several application examples of the present invention. The inventors have realized several use cases of the ray tracking described above. In some application examples, from a plurality of consecutive images, there may be at least one image in which the three-dimensional position of the projection spot 33 (or other feature) from a particular projector ray 88 was not identified in step 232 of the method shown in Figure 38. For example, the projection spot (or other feature) in a particular frame may be below an intensity threshold and not considered by the corresponding algorithm, or there may be a false detection of the spot in a particular frame. However, due to the time-series tracking of the length of the projector ray, the processor 96 may calculate the estimated length of the particular projector ray 88 in that image.

[0349] For example, in the illustrative graph shown in Figure 39, for scan frame s1 captured at time t1, the three-dimensional position of the projection spot 33 is not determined based on the corresponding algorithm, and therefore, as illustrated by the dashed circle 236, there are no data points corresponding to the length of the projector ray 88 for scan frame s1. However, because the length of the ray is tracked through multiple consecutive images, the estimated length L1 of the projector ray 88 can be calculated for scan frame s1, for example, by interpolation. As described above, all projection points by a particular projector ray 88 appear on a specific path 90 of the pixels of the camera sensor 58 corresponding to that particular projector ray 88. Therefore, for scan frame s1 in which the three-dimensional position of the spot 33 corresponding to a particular projector ray 88 was not determined in step 232, the processor 96 may determine a one-dimensional search space 238 within scan frame s1 for searching for the projection spot from that particular projector ray 88. The one-dimensional search space 238 is along the path 90 of each pixel corresponding to a particular projector ray 88. This is in contrast to the spot tracking algorithm described above, in which the processor 96 searches two-dimensionally within the image for spots that are sufficiently close to each other from one frame to the next, because each of the image frames is considered to be a tracking spot generated by the same projector ray.

[0350] In some applications, based on the estimated length L1 of the projector ray 88 in at least one of several images, the processor 96 may determine a one-dimensional search space in each pixel array of multiple cameras 24, e.g., all cameras 24, e.g., camera sensor 58. For each of the respective pixel arrays, the one-dimensional search space is along the path 90 of each pixel corresponding to the projector ray 88 in that particular pixel array, e.g., camera sensor 58. The length L1 of the projector ray 88 corresponds to a three-dimensional point in space, which corresponds to a two-dimensional position on the camera sensor 58. All other cameras 24 also have their respective two-dimensional positions on the camera sensor 58 corresponding to the same three-dimensional point in space. Thus, the length of the projector ray in a particular frame may be used to define a one-dimensional search space for multiple camera sensors 58, e.g., all camera sensors 58, for that particular frame.

[0351] In some applications, there is a possibility that a false detection of the projection spot 33 (or other feature) may occur, as opposed to a false detection where the expected spot (or other feature) was not detected, in which case at least one of several consecutive images may exist in which multiple candidate 3D positions were calculated for the projection spot 33 (or other feature) from a particular projector ray 88. For example, in the exemplary graph shown in Figure 39, based on the corresponding algorithm, for a scan frame s2 captured at time t2, it is calculated that two candidate detection spots 33' and 33'' both originate from a specific projector ray 88, and thus two candidate 3D positions of the projection spot 33 are calculated. The processor 96 then calculates two candidate lengths of the projector ray 88 corresponding to each of the candidate 3D positions for that frame. For example, the processor 96 calculates that the candidate length of the projector ray 88 for the candidate detection spot 33' is L2 (represented by data point 242 in Figure 39), and the candidate length of the projector ray 88 for the candidate detection spot 33'' is L3 (represented by data point 244 in Figure 39). Because the length of the projector ray 88 is tracked across multiple consecutive images, when the ray length data for scan frame s2 are added together, it becomes clear which of the candidate lengths, L2 or L3, is the estimated length of the projector ray 88 for scan frame s2. In this way, the processor 96 can determine the correct three-dimensional position of the projection spot 33 by determining which of the multiple candidate three-dimensional positions of the projection spot 33 corresponds to the estimated length of the projector ray 88 for that image.

[0352] Based on at least one of the multiple images, for example, the estimated length of the projector ray 88 in scan frame s2, the processor 96 may determine a one-dimensional search space 246 in scan frame s2. The processor 96 may then determine which of the multiple candidate three-dimensional positions of the projection spot 33' is the correct three-dimensional position of the projection spot 33 generated by the projector ray 88 by determining which of the multiple candidate three-dimensional positions of the projection spot 33' corresponds to the spot 33' generated by the projector ray 88 and is detected in the one-dimensional search space 246. Before additional information is provided by ray tracing, the camera sensor 58 for scan frame s2 should have shown both candidate detection spots 33' and 33'' on the path 90 of pixels corresponding to the projector ray 88. The processor 96 calculates the estimated length of the projector ray 88 based on the length of the ray being traced across multiple consecutive images, thereby enabling the processor 96 to determine the one-dimensional search space 246 and determine that candidate detection spot 33' is indeed the correct spot. Candidate detection spot 33'' is then excluded from consideration as a point on the three-dimensional intraoral surface.

[0353] In some applications, the processor 96 may define a curve 248 based on the evaluated, for example, calculated, length of the projector ray 88 in each image of a subset of images, such as a plurality of consecutive images. In our assumptions, it can be reasonably assumed that any detected point corresponding to the length of the projector ray r that is at least a threshold distance away from the defined curve 248, based on the corresponding algorithm, can be considered a false detection and excluded from being considered a point on the three-dimensional oral surface.

[0354] Next, we refer to Figures 40A and 40B, which are graphs showing experimental sets of data before and after ray tracing in several application examples of the present invention. In Figure 40A, the length of a particular projector ray 88 is plotted for all spots and / or other features calculated by the corresponding algorithm. Before ray tracing is applied, the data includes ray lengths corresponding to spots and / or other features that appear far from the general curve defined by the projector ray 88. Figure 40B shows the appearance of the data after ray tracing has been applied and used to determine which falsely detected spots and / or other features should be removed from being considered as points on the three-dimensional intraoral surface.

[0355] Next, we refer to Figure 41, a schematic diagram showing multiple camera sensors and a projector that projects a spot or other feature, according to some application examples of the present invention. In some application examples, after executing the correspondence algorithms described above, such as the correspondence algorithms, with reference to Figures 7 to 17, the processor 96 can determine candidate three-dimensional positions of the projection spot 33 (or other feature) with a certain degree of certainty, depending on how many cameras 24 detected a given projection spot 33 (or other feature) on their respective pixel arrays (i.e., camera sensors 58). In Figure 41, the candidate three-dimensional positions of the projection spot 33 (or other feature) are indicated by dashed circles 250.

[0356] The more cameras 24 that have viewed the projection spot 33 (or other feature), the higher the certainty of the candidate 3D position 250. Therefore, using data from at least two of the cameras 24, the processor 96 can identify candidate 3D positions 250 for a given spot 33 (or other feature) corresponding to a particular projector ray 88. Assuming that the identification of the candidate 3D position 250 was substantially determined without using data from at least another camera 24', there may be some error in the candidate 3D position 250, and if the processor 96 had data from another camera 24', the candidate 3D position 250 could be refined.

[0357] Thus, assuming that at least two cameras 24 have seen the projection spot 33 (or other feature) after obtaining correspondence, the processor 96 now knows (a) which projector ray 88 generated the projection spot 33 (or other feature), and (b) the candidate 3D position 250 of the spot (or other feature). Combining (a) and (b), the processor 96 can determine a one-dimensional search space 252 for the pixel array of another camera 24', i.e., the camera sensor 58', to search for the spot (or other feature) from the projector ray 88. The one-dimensional search space 252 is along the path 90 of the pixels on the camera sensor 58' of the other camera 24', and may be along a specific segment of the path 90 corresponding to the candidate 3D position 250. If a spot 33' (or other feature) originating from the projector ray 88, for example, a false detection spot 33' that was not considered by the corresponding algorithm (for example, due to an intensity below a threshold), is identified in the one-dimensional search space 252, the candidate three-dimensional position 250 of the spot 33 (or other feature) may be refined using the data from the other camera 24' obtained in this instance to obtain a refined three-dimensional position 254.

[0358] Next, we refer to Figures 42A-B, which show flowcharts outlining methods for generating 3D images according to several applications of the present invention. In some applications, once 3D positions are determined for at least three projection spots 33 (or other features) from three different projector rays 88, a 3D surface can be estimated such that all three of the determined 3D positions lie on the estimated surface. For another projector ray 88' whose 3D position was not determined, for example, because the intensity of the projection spot 33 was below a threshold, a candidate 3D position may be calculated at the intersection of the other projector ray 88' and the estimated 3D surface. As described above with reference to Figure 41, the processor 96 may determine a one-dimensional search space in at least one camera sensor 58 for searching for detected spots 33' from other projector rays 88' using a combination of (a) knowing a specific projector ray, i.e., knowing the path 90 on the sensor to be viewed, and (b) knowing a candidate 3D position.

[0359] Accordingly, in step 256 of the method outlined in Figures 42A-B, each structured light projector 22 is driven to project a structured light pattern, e.g., a distribution 34 of discrete, unconnected light spots 33 onto the three-dimensional surface of the oral cavity, and in step 258, each camera 24 is driven to capture multiple images, each image containing at least one of the spots. In step 260, the processor 96 is used to execute a correspondence algorithm (e.g., the correspondence algorithm described above with reference to Figures 7-17) based on stored calibration values ​​indicating (a) camera rays 86 corresponding to each pixel on the camera sensor 58 of each camera 24 and (b) projector rays 88 corresponding to each projected spot 33 of light from each structured light projector 22, to calculate the respective three-dimensional positions of the multiple detected spots 33' on the three-dimensional surface of the oral cavity for each of the multiple images. In step 262, using data corresponding to the three 3D positions of at least three detection spots 33' that correspond to each projector ray 88, the processor 96 estimates the 3D surface on which all at least three detection spots 33' are located. The processor 96 then considers another projector ray 88'. As shown by the determination hexagon 266, if the correspondence algorithm has already identified the 3D position of one detection spot 33' from the other projector ray 88' in step 260, that detection spot 33' from the other projector ray 88' is considered to be a point on the oral surface at that 3D position (step 268). However, for another projector ray 88' for which the 3D position of the spot 33 corresponding to the other projector ray 88' was not calculated in step 260, the processor 96 may estimate the 3D position in the intersection space of the other projector ray 88' and the estimated 3D surface (step 270). In step 272, the processor 96 uses the three-dimensional position in the estimated space to identify a search space (e.g., a one-dimensional search space) for searching for detection spots 33' corresponding to other projector rays 88' in the pixel array (e.g., camera sensor 58) of at least one camera 24.

[0360] As described above, in order to reduce the occurrence of the camera 24 detecting many false detection spots, the processor 96 may set a threshold, for example, an intensity threshold, and any detection feature below the threshold, for example, spot 33', will not be considered by the corresponding algorithm. Therefore, for example, because the intensity of the detected spot 33' is below the threshold, the 3D position of spot 33 corresponding to the other projector ray 88' may not have been calculated in step 260. In step 272, in order to search for features, for example, the detected spot 33', the processor may lower the threshold in order to consider features that were not initially considered by the corresponding algorithm.

[0361] When used in the entire application, including the claims, if a search space is identified for searching for detected features, for example, detected spot 33', it may be as follows: (a) A falsely mis-detected spot (for example, a spot below a threshold that was not initially considered by the corresponding algorithm, or a spot blocked by moving tissue), in which case the processor 96 may lower the threshold to re-search that particular region, i.e., the identified search space, for the detected spot 33', or (b) A false positive spot, i.e., a number of candidate 3D locations for a spot identified by the corresponding algorithm, in which case the processor 96 may determine which spot is the correct spot based on its specific region for the detected spot 33', i.e., a re-search of the identified search space.

[0362] In some applications, in step 260, the corresponding algorithm may identify multiple candidate 3D positions for the detection spot 33' from the projector ray 88'. For the projector ray 88' for which multiple candidate 3D positions have been identified for the detection spot 33', as shown by the determination hexagon 267, the processor 96 may estimate the 3D position in the intersection space between the projector ray 88' and the estimated 3D surface (step 269). In step 271, the processor 96 selects which of the candidate 3D positions for the detection spot 33' from the projector ray 88' is the correct position based on the 3D position of the intersection point between the projector ray 88' and the estimated 3D surface.

[0363] In some applications, in step 262, the processor 96 uses data corresponding to the respective 3D positions of at least three detection spots 33' captured in one of the multiple images. Furthermore, after the 3D surface has been estimated, the estimation may be refined by adding data points from subsequent images, i.e., by using data corresponding to the 3D position of at least one additional spot whose 3D position has been calculated based on another image of the multiple images, so that all spots (the three used in the original estimation plus at least one additional spot) lie on the refined estimated 3D surface. In some applications, in step 262, the processor 96 uses data corresponding to the respective 3D positions of at lea...

Claims

1. Intraoral scanner and One or more structured light projectors, each structured light projector configured to project a light pattern onto the oral cavity surface, Two or more cameras configured to capture one or more sets of images, wherein each set of images includes at least one image from each of the two or more cameras, and each image includes at least a portion of the projected light pattern, One or more processors, By solving the correspondence problem within one of the sets of images, a point in three-dimensional (3D) space is determined based on the correspondence between a captured feature in the set of images and at least a portion of the projected features of the projected pattern, and the point in 3D space forms the solution to the correspondence problem. One or more processors configured to calibrate the intraoral scanner by performing changes to stored calibration data associated with the intraoral scanner, wherein the changes are based on the solution to the corresponding problem, An intraoral scanning system equipped with [specific features / features].

2. The intraoral scanning system according to claim 1, wherein one or more processors are further configured to calibrate the intraoral scanner during scanning of the intraoral surface.

3. The one or more processors described above are: The camera beam corresponding to the captured features and the projector beam corresponding to the projected features are projected together into the 3D space. The intraoral scanning system according to claim 1, further configured to minimize the distance between one or more camera beams and a given related projector beam from among the projector beams by changing the stored calibration data.

4. The stored calibration data includes the stored calibration model, and changing the stored calibration data is, The intraoral scanning system according to claim 3, comprising changing one or more parameters in the stored calibration model to reduce the difference between (i) the update path of pixels corresponding to one or more projector rays of one or more structured light projectors and (ii) the path of stored pixels corresponding to one or more projector rays of one or more structured light projectors from the stored calibration data.

5. The intraoral scanner comprises a wand having a probe at the distal end of the wand, the one or more structured light projectors comprises at least two structured light projectors within the probe, and the two or more cameras comprises at least eight cameras within the probe. The intraoral scanning system according to claim 1, wherein each of the one or more structured light projectors is configured to project a spatially fixed pattern onto the two or more cameras.

6. The intraoral scanning system according to claim 1, wherein the projected light pattern includes a checkerboard pattern of light.

7. The intraoral scanning system according to claim 1, wherein the intraoral scanner is calibrated without using a calibration target.

8. The intraoral scanning system according to claim 1, wherein the projected light pattern includes a plurality of projection light spots, and the portion of the projected light pattern includes one of the plurality of projection light spots.

9. The aforementioned light pattern is defined by multiple projector rays, The stored calibration data associates the camera rays corresponding to pixels on each of the two or more camera sensors with the projector rays among the plurality of projector rays. Each of the plurality of projector rays corresponds to the path of each pixel on the camera sensor of each of the two or more cameras. In order to solve the aforementioned correspondence problem, one or more processors For each projector ray i, for each detected feature j on the camera sensor path corresponding to the projector ray i, the number of other cameras that detected each feature k corresponding to each camera ray where the projector ray i and the camera ray corresponding to the detected feature j intersect on each camera sensor path corresponding to the projector ray i is identified, and the projector ray i is identified as the specific projector ray that produced the detected feature j, where the most other cameras detected each feature k. The intraoral scanning system according to claim 1, which executes a corresponding algorithm for calculating the respective three-dimensional positions on the intraoral surface at the intersection of the projector beam i and the respective camera beams corresponding to the detected feature j and the respective detected feature k.

10. The one or more processors described above are: Based on the data accumulated by the intraoral scanner, the accuracy of the stored calibration data is determined. The intraoral scanning system according to claim 1, further configured to automatically calibrate the intraoral scanner in response to a determination that the accuracy falls below an accuracy threshold.

11. The one or more processors described above are: Determine whether the inaccuracy in the stored calibration data is due to one of the two or more cameras, or to one of the one or more structured light projectors. In response to the determination that the inaccuracy in the stored calibration data is due to the camera, the camera is recalibrated. The intraoral scanning system according to claim 1, further configured to recalibrate the structured light projector in response to the determination that the inaccuracy of the stored calibration data is attributable to the structured light projector.

12. The structured light projector is recalibrated by performing recalibration of the projector beam associated with the structured light projector based on updates to the calibration data of the projector beam in the parameterized projector calibration model. The intraoral scanning system according to claim 11, wherein the camera is recalibrated by redefining a parameterized calibration function that acquires a given 3D position in space and converts the given 3D position to a given pixel in a two-dimensional (2D) pixel array of the camera sensor of the camera.

13. The stored calibration data includes stored calibration values ​​indicating (a) camera rays corresponding to each pixel on each camera sensor of the two or more cameras, and (b) projector rays corresponding to each feature of the light pattern projected from each of the one or more structured light projectors, wherein each projector ray r corresponds to the bus p of each pixel on at least one of the camera sensors, and the one or more processors The intraoral scanning system according to claim 1, further configured such that for each projector ray r, a path p' of updated pixels is defined on each of the camera sensors, so that all of the points in 3D space corresponding to the features of the light pattern generated by the projector ray r correspond to positions along the updated pixel path p' for each of the camera sensors, and the updated path p' is used to recalibrate the stored calibration data.

14. The one or more processors described above are: The calibration evaluation values ​​of the intraoral scanner are determined, and each of these calibration evaluation values ​​represents the amount of deviation of the intraoral scanner from its calibrated state over a certain period of time. The intraoral scanning system according to claim 1, further configured to determine at least one of the drift or drift rate of the calibrated state of the intraoral scanner based on the calibration evaluation value.

15. The oral cavity surface includes a calibration target having known parameters, and the one or more processors further, The intraoral scanning system according to claim 1, further configured to perform a triangulation algorithm to calculate the parameters of each of the calibration objects based on the one or more captured images.

16. The intraoral scanning system according to claim 1, wherein the intraoral scanning system needs to be recalibrated multiple times during an intraoral scanning session.

17. A method for automatically recalibrating an intraoral scanner, The steps include driving one or more light sources to project light onto the three-dimensional surface inside the oral cavity, The steps include driving one or more cameras to capture multiple images of the three-dimensional surface inside the oral cavity, Based on the stored calibration data of the one or more light sources and the one or more cameras, Execute the corresponding algorithm to calculate the three-dimensional position of each of the multiple features of the projected light on the three-dimensional surface of the oral cavity. The collection of data at one or more points in time, wherein the data includes the calculated three-dimensional position of each of the plurality of features on the three-dimensional surface of the oral cavity, and The steps include performing the following: using the collected data, recalibrating the stored calibration data; A method that includes this.

18. The method according to claim 17, wherein the stored calibration data is recalibrated during scanning of the three-dimensional surface of the oral cavity.

19. The method according to claim 17, wherein the stored calibration data is recalibrated multiple times during an intraoral scan session.

20. The steps include projecting together a camera ray corresponding to the captured features of the projected light and a projector ray corresponding to the projected features of the projected light into a three-dimensional space, The steps include minimizing the distance between one or more camera rays and a given related projector ray from among the projector rays by changing the stored calibration data, The method according to claim 17, further comprising:

21. The stored calibration data includes the stored calibration model, and changing the stored calibration data is, The method according to claim 20, comprising varying one or more parameters in the stored calibration model to reduce the difference between (i) the updated pixel path corresponding to the given projector ray to a specific light source among the one or more light sources and (ii) the stored pixel path corresponding to the given projector ray to the specific light source from the stored calibration data.

22. The method according to claim 17, wherein the intraoral scanner is calibrated without using a calibration target.

23. The method according to claim 17, wherein the aforementioned multiple features include multiple light spots.

24. A non-temporary computer-readable medium that, when executed by a processing unit, While light is projected onto the three-dimensional surface of the oral cavity by one or more light sources of the intraoral scanner, multiple images of the three-dimensional surface of the oral cavity captured by one or more cameras of the intraoral scanner are received, and Based on the stored calibration data of the one or more light sources and the one or more cameras of the intraoral scanner, Execute the corresponding algorithm to calculate the three-dimensional position of each of the multiple features of the projected light on the three-dimensional surface of the oral cavity. The collection of data at one or more points in time, wherein the data includes the calculated three-dimensional positions of the plurality of features on the three-dimensional surface of the oral cavity, and The command includes causing the processing device to perform an operation that includes using the collected data to recalibrate the stored calibration data, Non-temporary computer-readable media.

Citation Information

Patent Citations

  • Device for determining three-dimensional coordinate of object, tooth in particular

    JP2010069301A

  • Apparatus and method for non-contact detection of three-dimensional contours

    JP2010507079A

  • Shape measuring instrument and shape measuring method

    JP2011242178A

  • Intraoral three-dimensional measuring device, intraoral three-dimensional measuring method, and method for displaying intraoral three-dimensional measuring result

    JP2017020930A

  • JPP7661250B