Intraoral 3D scanner using multiple small cameras and multiple small pattern projectors

The system with multiple cameras and projectors using laser diodes and diffraction/refraction optics addresses contrast and correspondence issues in intraoral scanning, achieving accurate and efficient 3D imaging without opaque powders and automated calibration.

JP2025106354AActive Publication Date: 2025-07-15ALIGN TECHNOLOGY INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025060895
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-06-23
Filing Date
2025-04-02
Publication Date
2025-07-15
Estimated Expiration
2040-06-24

AI Technical Summary

Technical Problem

Conventional digital intraoral scanners face challenges in capturing high-reflectivity and translucent dental surfaces due to reduced contrast of structured light patterns, necessitating the use of opaque powders, and struggle with the correspondence problem in structured light imaging, which complicates the stitching of multiple image frames.

Method used

Employing a system with multiple small cameras and projectors that utilize laser diodes and diffraction/refraction pattern generators to project discrete, unconnected light spots, coupled with advanced image processing algorithms to track and triangulate these spots across images, and integrate visual and inertial tracking for precise scanning.

Benefits of technology

Enhances scanning accuracy and efficiency by maintaining high contrast without powders, reducing heating, and automating calibration, while improving image stitching and surface reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025106354000001_ABST
    Figure 2025106354000001_ABST
Patent Text Reader

Abstract

To provide a method for generating a digital three-dimensional image and making a three-dimensional image of the oral cavity using structured light illumination.SOLUTION: The method for generating a three-dimensional image includes a step 62 for driving a structured light projector(s) to project a light pattern onto a three-dimensional surface in the oral cavity, a step 64 for driving a camera(s) to capture an image including at least one of spots, and a step for using a processor to compare a series of images captured by each camera and determine which parts of the projection pattern can be tracked across the images. A 3D model of the 3D surface in the oral cavity is constructed at least partially on the basis of the comparison of the series of images.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to three-dimensional imaging, and more particularly to three-dimensional intraoral imaging using structured light illumination.

Background Art

[0002] Three-dimensional intraoral surfaces of a subject, such as dental impressions of teeth and gums, are used for planning dental treatments. Conventional dental impressions are made using a dental impression tray filled with an impression material such as PVS or alginate, and the subject bites on it. The impression material then hardens to become a negative imprint of the teeth and gums, from which a three-dimensional model of the teeth and gums can be formed.

[0003] Digital dental impressions generate a three-dimensional digital model of the three-dimensional intraoral surface of a subject using intraoral scanning. Digital intraoral scanners often use structured light three-dimensional imaging. The surface of a subject's teeth may have a high reflectivity and be somewhat translucent, which may reduce the contrast of the structured light pattern reflected from the teeth. Therefore, when using a digital intraoral scanner that utilizes structured light three-dimensional imaging, in order to improve the capture of intraoral scans and to promote the available level of contrast of the structured light pattern, for example, the subject's teeth are often coated with an opaque powder before scanning in order to make the surface a scattering surface. Intraoral scanners that utilize structured light three-dimensional imaging have advanced to a certain extent, but there may still be further advantages.

Summary of the Invention

[0004] The use of structured light 3D imaging can lead to the so-called "correspondence problem" which requires determining the correspondence between points in the structured light pattern and the points seen by the camera viewing the pattern. One technique for addressing this problem is based on projecting a "coded" light pattern and imaging the illuminated scene from one or more viewpoints. By coding the emitted light pattern, a part of the light pattern becomes unique and identifiable when captured by the camera system. Since the pattern is coded, the correspondence between the points in the image and the points in the projection pattern can be found more easily. The decoded points are triangulated and the 3D information is restored.

[0005] Application examples of the present invention include systems and methods related to a three-dimensional intraoral scanning device including one or more cameras and one or more pattern projectors. For example, certain application examples of the present invention may relate to an intraoral scanning device having a plurality of cameras and a plurality of pattern projectors.

[0006] Further application examples of the present invention include methods and systems for decoding a structured light pattern.

[0007] Yet another application example of the present invention may relate to systems and methods for three-dimensional intraoral scanning using an unencoded structured light pattern.

[0008] For example, in some particular applications of the present invention, an intraoral scanning device is provided, which includes an elongated hand-held wand having a probe at its distal end. During scanning, the probe may be configured to enter the subject's oral cavity. One or more small structured light projectors and one or more small cameras are coupled to a rigid structure disposed within the distal end of the probe. Each structured light projector uses a light source such as a laser diode to transmit light. In some applications, the structured light projector has an illumination field of at least 45 degrees. Optionally, the illumination field may be less than 120 degrees. Each structured light projector may further include a pattern generation optical element. The pattern generation optical element may utilize diffraction and / or refraction to generate a light pattern. In some applications, the light pattern may be a distribution of discrete, unconnected light spots. Optionally, when the light source (e.g., a laser diode) is activated and transmits light through the pattern generation optical element, the light pattern maintains a distribution of discrete, unconnected spots in all planes located between 1 mm and 30 mm from the pattern generation optical element. In some applications, the pattern generation optical element of each structured light projector may have a light throughput efficiency, i.e., the ratio of the light that enters the pattern out of the light that falls on the pattern generator, of at least 80%, for example, at least 90%. Each camera includes a camera sensor and an objective optical system including one or more lenses.

[0009] A laser diode light source and a diffraction and / or refraction pattern generating optical element can provide certain advantages in some applications. For example, the use of a laser diode, a diffraction and / or refraction pattern generating optical element can help maintain an energy-efficient structured light projector to prevent the probe from being heated during use. Further, such components may help reduce costs as they do not require active cooling within the probe. For example, current laser diodes can transmit continuously at high brightness while consuming less than 0.6 watts of power (in contrast to current light-emitting diodes (LEDs), for example). When pulsed according to some applications of the present invention, these current laser diodes may use even less power. For example, when pulsed at a 10% duty cycle, the laser diode may use less than 0.06 watts (however, in some applications, the laser diode may use at least 0.2 watts while transmitting continuously at high brightness and may use even less power when pulsed. For example, when pulsed at a 10% duty cycle, the laser diode may use at least 0.02 watts). Further, the diffraction and / or refraction pattern generating optical element may be configured to utilize most, if not all, of the transmitted light (in contrast to a mask that stops some of the light rays from hitting the object, for example).

[0010] In particular, a pattern generation optical element based on diffraction and / or refraction generates a pattern by diffraction, refraction, or interference of light, or any combination thereof, rather than modulation of light as performed with a transparent mask or a transmissive mask. In some applications, this can be advantageous because the light throughput efficiency (the ratio of the light that enters the pattern to the light that falls on the pattern generator), regardless of the "area-based duty cycle" of the pattern, is nearly 100%, for example at least 80%, for example at least 90%. In contrast, the light throughput efficiency of a pattern generation optical element with a transparent mask or a transmissive mask is directly related to the "area-based duty cycle". For example, when the desired "area-based duty cycle" is 100:1, the throughput efficiency of a mask-based pattern generation optical element is 1%, while the efficiency of a diffraction and / or refraction-based pattern generation optical element remains nearly 100%. Further, the light collection efficiency of a laser is more than 10 times higher than that of an LED with the same total light output. This is because the laser has a smaller emission area and divergence angle in essence, resulting in a brighter output illuminance per unit area. The high efficiency of lasers and diffraction and / or refraction pattern generators enables a configuration with high thermal efficiency, limits the large heating of the probe during use, and potentially eliminates or limits the need for active cooling within the probe, thereby reducing costs. In some applications, laser diodes and DOEs may be particularly preferred, but these are by no means essential, either alone or in combination. Other light sources such as LEDs and pattern generation elements including transparent masks or transmissive masks can be used in other applications, regardless of the presence or absence of active cooling.

[0011] In some applications, in order to improve the image capture of the intraoral scene under structured light illumination without using contrast enhancement means such as coating teeth with opaque powder, the inventors have found that light patterns such as the distribution of discrete, unconnected light spots (e.g., not lines) may offer an improved balance with respect to increasing the contrast of the pattern while maintaining a useful amount of information. Generally, a higher density structured light pattern can provide a greater amount of surface sampling and higher resolution, and may be able to better stitch together each surface obtained from multiple image frames. However, with a structured light pattern that is too dense, the number of spots for which correspondence problems need to be solved increases, which can make the correspondence problems more complex. Further, a high density structured light pattern may lead to a decrease in pattern contrast due to increasing the light in the system, which can be caused by a combination of (a) stray light reflected from the somewhat glossy surface of the teeth and picked up by the camera, and (b) percolation, i.e., a portion of the light incident on the teeth is reflected along multiple paths within the teeth and then exits the teeth in many different directions. As will be further described below, a method and system are provided for solving the correspondence problem presented by the distribution of discrete, unconnected light spots. In some applications, the discrete, unconnected light spots from each projector may not be encoded.

[0012] In some applications, the field of view of each camera may be at least 45 degrees, for example, at least 80 degrees, for example, even 85 degrees. Optionally, the field of view of each camera may be less than 120 degrees, for example, less than 90 degrees. In some applications, one or more of the cameras have a fish-eye lens or other optical element that provides a field of view of up to 180 degrees.

[0013] In any case, the fields of view of the various cameras may or may not be the same. Similarly, the focal lengths of the various cameras may or may not be the same. As used herein, the term "field of view" of each camera means the diagonal field of view of each camera. Further, each camera may be configured to focus on an object focal plane located between 1 mm and 30 mm, for example, at least 5 mm and / or less than 11 mm, for example, between 9 mm and 10 mm, from the lens farthest from each camera sensor. Similarly, in some applications, each illumination field of the structured light projector may be at least 45 degrees and optionally less than 120 degrees. The inventors have recognized that the large field of view achieved by combining the fields of view of all the cameras may improve the accuracy because the amount of image stitching error is reduced, particularly in the edentulous region where the surface of the gingiva is smooth and there may be few clear high-resolution 3D features. The presence of a wide field of view enables large smooth features such as the overall curve of the teeth to appear in each image frame, thus improving the stitching accuracy of the respective surfaces obtained from multiple image frames.

[0014] In some applications, a method for generating a digital three-dimensional image of the oral cavity surface is provided. It should be noted that the expression "three-dimensional image" used in this application is based on a three-dimensional model, for example, a point cloud, from which an image of the three-dimensional oral cavity surface is constructed. The resulting image is generally displayed on a two-dimensional screen, but contains data related to the three-dimensional structure of the scanned object, and thus can typically be manipulated to display the scanned object from different viewpoints and perspectives. Further, a physical three-dimensional model of the scanned object may be created using data from the three-dimensional image.

[0015] For example, one or more structured light projectors may be driven to project a pattern of light, such as a distribution of discrete, unconnected light spots, an intersecting line pattern (e.g., a grid), a checkerboard pattern, or other pattern, onto the inner surface of the oral cavity, and one or more cameras may be driven to capture an image of the projection. The image captured by each camera may include a portion of the projection pattern (e.g., at least one of the spots). In some embodiments, the one or more structured light projectors project a spatially fixed pattern relative to the one or more cameras.

[0016] Each camera includes a camera sensor having a pixel array, and for each pixel, there is a corresponding light ray in 3D space that emanates from the pixel in the direction facing the object being imaged, and each point along a particular one of these light rays, when imaged on the sensor, falls on a corresponding respective pixel on the sensor. Throughout this specification, including the claims, the term used in this regard is "camera ray". Similarly, for each projection spot from each projector, there is a corresponding projector ray. Each projector ray corresponds to the path of each pixel on at least one camera sensor, i.e., when the camera views a feature or portion (e.g., a spot) of the pattern projected by a particular projector ray, that feature or portion (e.g., the spot) of the pattern is necessarily detected by the pixels on a particular path of the pixels corresponding to that particular projector ray. The values of (a) the camera rays corresponding to each pixel on the camera sensor of each camera and (b) the projector rays corresponding to each feature or portion (e.g., a light spot) of the pattern projected from each projector may be stored during a calibration process as described below.

[0017] Regarding camera rays, in some applications, instead of storing individual values for each camera ray corresponding to each pixel on the camera sensor of each camera, a set of smaller calibration values that can be used to indicate each camera ray is stored. For example, to define a camera ray, parameter values for a parameterized camera calibration function that takes a given 3D position in space and converts it to a given pixel in the 2D pixel array of the camera sensor may be stored.

[0018] Regarding projector rays, (a) in some applications, an indexed list containing the values of each projector ray is stored, and (b) alternatively, in some applications, a set of smaller calibration values that can be used to indicate each projector ray is stored. For example, parameter values for a parameterized projector calibration model that defines each projector ray of a given projector may be stored.

[0019] Based on the stored calibration values, the processor may execute a corresponding algorithm to identify the three-dimensional position of each part (e.g., projection spot) of the features of the projected light pattern on the surface. For a given projector ray, the processor "looks" at the corresponding camera sensor path on one of the cameras. Detection spots and other features along that camera sensor path have camera rays that intersect that given projector ray. That intersection point defines a three-dimensional point in space. Next, the processor searches through the camera sensor paths corresponding to that given projector ray on the other cameras and identifies how many other cameras detected a pattern feature (e.g., spot) whose camera ray intersects that three-dimensional point in space on their respective camera sensor paths corresponding to that given projector ray. As used herein throughout this application, when two or more cameras detect a portion or feature (e.g., spot) of a pattern where their respective camera rays intersect a given projector ray at the same point in three-dimensional space, the cameras are considered to "agree" that the portion or feature (e.g., spot) is located at that three-dimensional point. This process is repeated for additional features (e.g., spots) along the camera sensor paths, and the feature (e.g., spot) that the most cameras "agree" on is identified as the feature (e.g., spot) projected from the given projector ray onto the surface. In this way, the three-dimensional position on the surface is calculated for that feature of the pattern (e.g., that spot).

[0020] In some embodiments, once the position on the surface has been determined for a particular feature (e.g., a particular spot) of the pattern, the projector ray that projected that feature (e.g., spot) and all camera rays corresponding to that feature (e.g., spot) may be excluded from consideration, and the corresponding algorithm may be run again for the next projector ray.

[0021] Further application examples of the present invention involve projecting a structured light pattern (e.g., parallel lines, grids, checkerboards, unconnected and / or uniform spots, random spot patterns, etc.) onto an intraoral object, capturing at least a portion of the structured light pattern projected onto the intraoral object, and tracking a portion of the captured structured light pattern across consecutive images to scan the intraoral object. In some embodiments, tracking a portion of the structured light pattern captured across consecutive images can help improve the scanning speed and / or accuracy.

[0022] In a more specific example related to a structured light scanner using the above-described projection patterns (e.g., of unconnected spots), a processor may be used to compare a series of images (e.g., a plurality of consecutive images) captured by each camera to determine which features of the projection pattern (e.g., which of the projection spots) can be tracked across the series of images (e.g., across a plurality of consecutive images). The inventors have noticed that the detected movement of a particular feature or spot can be tracked across multiple images (e.g., consecutive image frames) in a series of images. Thus, the correspondence solved for that particular spot in any of the images or frames where the feature or spot was tracked provides the solution of the correspondence for that feature or spot in all the images or frames where the feature or spot was tracked. Since the detected features or spots that can be tracked across multiple images are features or spots generated by the same particular projector ray, the trajectory of the tracked feature or spot will be along the particular camera sensor path corresponding to that particular projector ray.

[0023] In some applications, instead of or in addition to tracking detected features or spots within a 2D image, the length of each projector ray can be tracked in 3D space. The length of a projector ray is defined as the distance between the origin of the projector ray, i.e., the light source, and the 3D position where the projector ray intersects the oral cavity surface. As will be further explained below, tracking the length of a particular projector ray over time can help resolve ambiguity in the correspondence. In some examples herein, the above concepts of spot and ray tracking are described with respect to a scanner that projects unconnected spots, but this is exemplary and in no way limiting, and it should be understood that the tracking techniques are equally applicable to scanners that project other patterns (e.g., parallel lines, grids, checkerboards, unconnected and / or uniform spots, random spot patterns, etc.) onto objects within the oral cavity.

[0024] In some embodiments, for the purpose of object scanning, it may be desirable to estimate the position of the scanner with respect to the object to be scanned, i.e., the three-dimensional intraoral surface, during scanning, and in certain embodiments, continuous estimation during scanning is desirable. Through several application examples of the present invention, the inventors have developed a method to combine visual tracking of the scanner's motion with inertial measurement of the scanner's motion to address situations where sufficient visual tracking may not be obtained. Using the accumulated data of the motion of the intraoral scanner with respect to the intraoral surface (visual tracking) and the motion of the intraoral scanner with respect to the fixed coordinate system (inertial measurement), a predictive model of the motion of the intraoral surface with respect to the fixed coordinate system can be constructed (described further below). When sufficient visual tracking is not available, the processor may calculate the estimated position of the intraoral scanner with respect to the intraoral surface by taking into account (e.g., subtracting in some embodiments) the prediction of the motion of the intraoral surface with respect to the fixed coordinate system from the inertial measurement of the motion of the intraoral scanner with respect to the fixed coordinate system (described further below). It should be understood that the scanner position estimation concept described herein can be used with intraoral scanners regardless of the scanning technology employed (e.g., parallel confocal scanning, focal scanning, wavefront scanning, stereovision, structured light, triangulation, light field, and / or combinations thereof). Thus, although discussed in relation to the concept of structured light described herein, this is exemplary and in no way limiting.

[0025] In some embodiments of the structured light scanner described herein, the stored calibration values may indicate (a) camera rays corresponding to each pixel on the camera sensor of each camera, and (b) projector rays corresponding to each projected feature (e.g., light spot) from each structured light projector, where each projector ray corresponds to the path of each pixel on at least one of the camera sensors of the camera sensor. However, over time, at least one of the cameras and / or at least one of the projectors may move (e.g., by rotation or translation), the optical system of at least one of the cameras and / or at least one of the projectors may be changed, and the wavelength of the laser may be changed, resulting in the possibility that the stored calibration values may not accurately correspond to the camera rays and projector rays.

[0026] For any given projector ray, the processor collects data including the calculated three-dimensional positions on the oral cavity surface of a plurality of detected features (e.g., spots) from that projector ray detected at respective different times. When these are overlaid on one image, all of those features (e.g., spots) should fall on the camera sensor path of the pixel corresponding to that projector ray. If the calibration of the camera or projector is changed by something, the features (e.g., spots) detected from that particular projector ray may appear not to fall on the camera sensor path of the pixel expected according to the stored calibration values, but rather may appear to be located on the camera sensor path of the newly updated pixel. When the calibration of the camera(s) and / or projector(s) is changed, the processor may reduce the difference between the path of the updated pixel and the path of the original pixel from the calibration data by (i) changing the stored calibration values (e.g., the stored parameter values of a parameterized camera calibration model, e.g., a function) indicating the camera rays corresponding to each pixel on the camera sensor of one or more cameras, and / or (ii) changing the stored calibration values (e.g., the values stored in an indexed list of projector rays, or the stored parameter values of a parameterized projector calibration model) indicating the projector ray r corresponding to each feature (e.g., light spot) projected from each projector of one or more projectors.

[0027] The current calibration evaluation may be performed automatically on a regular basis (e.g., every scan, every 10 scans, monthly, every few months, etc.) or in response to a particular criterion being met (e.g., in response to a threshold number of scans having been performed). As a result of the evaluation, the system may determine whether the calibration state is accurate or inaccurate. In one embodiment, as a result of the evaluation, the system determines whether the calibration is drifting. For example, the previous calibration still maintains sufficient accuracy to produce high-quality scans, but if the detected trend continues, the system may deviate and potentially be unable to produce accurate scans in the future. In one embodiment, the system determines a drift rate and projects that drift rate into the future to determine a predicted date and time when the calibration will become inaccurate. In one embodiment, automatic or manual calibration can be scheduled for that future date and time. In one example, the processing logic evaluates the calibration state over time (e.g., by comparing the calibration states at multiple different points in time) and determines a drift rate from such comparison. From the drift rate, the processing logic can predict the timing at which calibration should be performed based on trend data.

[0028] Conventional intraoral scanners require the user to manually re-calibrate according to a set schedule (e.g., every six months). Conventional intraoral scanners do not have a function to monitor or evaluate the current calibration status (e.g., a function to determine whether re-calibration should be performed). Furthermore, calibration of conventional intraoral scanners is performed manually using special calibration targets. Calibration of conventional intraoral scanners is time-consuming and inconvenient for the user. Thus, the dynamic calibration performed in the particular embodiments described herein can enhance user convenience and be performed in less time compared to calibration of conventional intraoral scanners.

[0029] In some applications, if the calibration of the camera(s) and / or projector(s) is changed, the processor does not need to perform re-calibration. Instead, it may simply determine that at least some of the stored calibration values of the camera(s) and / or projector(s) are incorrect. For example, based on the determination that the stored calibration values are incorrect, the user may be prompted to return the intraoral scanner to the manufacturer for maintenance and / or re-calibration, or to request a new scanner.

[0030] Visual tracking of the motion of the intraoral scanner relative to the object being scanned may be obtained by stitching together each surface or point cloud obtained from adjacent image frames. As described herein, in some applications, illuminating the oral cavity under near-infrared (NIR) light can increase the number of visible features that can be used to stitch together each surface or point cloud obtained from adjacent image frames. In particular, since NIR light can penetrate teeth, the images captured under NIR light contain features inside the teeth, such as cracks inside the teeth, in contrast to two-dimensional color images taken under broadband illumination where only features appearing on the surface of the teeth are visible. These additional subsurface features may be used to stitch together each surface or point cloud obtained from adjacent image frames.

[0031] In some application examples, the processor may use a two-dimensional image (e.g., a two-dimensional color image and / or a two-dimensional monochromatic NIR image) in the 2D-3D surface reconstruction of the three-dimensional surface in the oral cavity. As described below, by using a two-dimensional image (e.g., a two-dimensional color image and / or a two-dimensional monochromatic NIR image), the resolution and speed of the three-dimensional reconstruction can be significantly improved. Therefore, as described in this specification, in some application examples, it is useful to enhance the three-dimensional reconstruction of the three-dimensional surface in the oral cavity by three-dimensional reconstruction from a two-dimensional image (e.g., a two-dimensional color image and / or a two-dimensional monochromatic NIR image). In some application examples, the processor calculates the three-dimensional position of each of a plurality of points on the three-dimensional surface in the oral cavity, for example, using the corresponding algorithm described in this specification, and calculates the three-dimensional structure of the three-dimensional surface in the oral cavity based on the plurality of two-dimensional images (e.g., a two-dimensional color image and / or a two-dimensional monochromatic NIR image) and the calculated three-dimensional positions on the oral cavity surface.

[0032] According to some application examples of the present invention, the calculation of the three-dimensional structure is performed by a neural network. The processor inputs into the neural network (a) a plurality of two-dimensional images (e.g., two-dimensional color images) of the three-dimensional surface in the oral cavity and (b) the calculated three-dimensional positions of a plurality of points on the three-dimensional surface in the oral cavity, and the neural network determines and returns each estimated map (e.g., a depth map, a normal map, and / or a curvature map of the three-dimensional surface in the oral cavity captured in each of the two-dimensional images (e.g., two-dimensional color images and / or two-dimensional monochromatic NIR images)).

[0033] The inventors noticed that when an intraoral scanner is commercially produced, small manufacturing deviations may exist, and the manufacturing deviations (a) slightly differ from the calibration of the camera(s) and / or projector(s) on each commercially produced intraoral scanner to the calibration of the training stage camera(s) and / or projector(s), and / or (b) slightly differ from the lighting relationship between the camera(s) and projector(s) of the commercially produced intraoral scanner to the lighting relationship between the training stage camera(s) and projector(s) used to train the neural network. Other manufacturing deviations of the camera and / or projector may exist. According to some applications of the present invention, a method is provided for using a processor to overcome the manufacturing deviations of the camera(s) and / or projector(s) of the intraoral scanner and reduce the difference between the estimated map and the true structure of the intraoral three-dimensional surface.

[0034] According to some applications of the present invention, one way the manufacturing tolerances can be overcome is to modify, e.g., crop and morph, the images from the in-field intraoral scanner so as to obtain modified images that match the field of view of the set of reference cameras used to train the neural network. The neural network is trained based on the images received from the set of reference cameras, and then the in-field images are modified as if the neural network had received these images as if they had been captured by the reference cameras. Then, the three-dimensional structure of the intraoral three-dimensional surface is calculated based on the plurality of modified two-dimensional images of the intraoral three-dimensional surface. For example, the neural network determines an estimated map of each of the intraoral three-dimensional surfaces captured in each of the plurality of modified two-dimensional images.

[0035] According to some application examples of the present invention, a neural network determines an estimated depth map for each of the intraoral three-dimensional surfaces captured in each two-dimensional image, and the depth maps are stitched together to obtain the three-dimensional structure of the intraoral surface. However, there may sometimes be contradictions between the estimated depth maps. The inventors have found it advantageous that for each estimated depth map determined by the neural network, the neural network also determines an estimated confidence map, and each confidence map indicates the confidence for each region of each estimated depth map. Accordingly, a method is provided herein for inputting a plurality of two-dimensional images of an intraoral three-dimensional surface into a first neural network module and a second neural network module. The first neural network module determines an estimated depth map for each of the intraoral three-dimensional surfaces captured in each two-dimensional image. The second neural network module determines an estimated confidence map corresponding to each estimated depth map. Each confidence map indicates the confidence for each region of each estimated depth map.

[0036] According to some application examples of the present invention, a neural network is trained using (a) two-dimensional images of a training stage three-dimensional surface, such as a model surface and / or an intraoral surface, and (b) a corresponding true output map of the training stage three-dimensional surface calculated based on a structured light image of the training stage three-dimensional surface. For each two-dimensional image, the neural network estimates an estimated map of the intraoral three-dimensional surface captured in each two-dimensional image, and then compares each estimated image with the corresponding true map of the intraoral three-dimensional surface. Based on the difference between each estimated map and the corresponding true map, the neural network is optimized to better estimate subsequent estimated maps.

[0037] In some applications, when the three-dimensional oral surface is used for training a neural network, moving tissues, such as the subject's tongue, lips, and / or cheeks, may block a portion of the three-dimensional oral surface from the field of view of one or more cameras. To avoid the neural network "learning" based on images of the moving tissues (as opposed to the stationary tissues of the three-dimensional oral surface being scanned), for a two-dimensional image in which the moving tissues are identified, the image may be processed to exclude at least a portion of the moving tissues before inputting the two-dimensional image into the neural network.

[0038] In some applications, to prevent cross-contamination between patients, a disposable sleeve is placed over the distal end of the intraoral scanner, such as over the probe, before the probe is placed in the patient's mouth. As further described herein, due to the relative positional relationship between the structured light projector within the probe and the adjacent camera, a portion of the projected structured light pattern may be reflected from the sleeve and reach the camera sensor of the adjacent camera. As further described herein, due to the polarization of the laser light of the structured light projector, the laser may be rotated about its own optical axis such that the polarization angle of the laser light with respect to the sleeve results in a reduced degree of reflection.

[0039] In some application examples of the present invention, a simultaneous localization and mapping (SLAM) algorithm is used to track the motion of a handheld wand and generate a three-dimensional image. SLAM can generally be performed using two or more cameras that view the same image but from slightly different angles. However, due to the positioning of the cameras 24 within the probe 28 and the proximity positioning of the probe 28 with respect to the scan object, i.e., the three-dimensional surface within the oral cavity, it is often the case that two or more cameras within the probe do not view substantially the same image. As will be described below, additional challenges may be encountered with respect to utilizing the SLAM algorithm when scanning the three-dimensional surface within the oral cavity. The inventors have invented several ways to overcome these challenges to track the motion of a handheld wand using SLAM and generate a three-dimensional image of the three-dimensional surface within the oral cavity, as further described herein.

[0040] In some application examples of the present invention, when scanning a three-dimensional surface within the oral cavity using a handheld wand, when a structured light projector projects a distribution of their features (e.g., a distribution of spots) onto the oral cavity surface, some of the features (e.g., spots) may land on a moving tissue (e.g., the patient's tongue). To improve the accuracy of the three-dimensional reconstruction algorithm, features (e.g., spots) that fall on a moving tissue generally should not be relied upon for the reconstruction of the three-dimensional surface within the oral cavity. As described herein, whether a feature (e.g., a spot) is projected onto a moving tissue or a stable tissue within the oral cavity may be determined on an unstructured light image frame (e.g., which may be broad-spectrum light) scattered within the structured light image frame. A reliability grading system may be used to assign a reliability grade based on a determination of whether a detected feature (e.g., a spot) is projected onto a fixed tissue or a moving tissue. Based on the reliability grades for each of a plurality of features (e.g., spots), the processor may execute a three-dimensional reconstruction algorithm using the detected features (e.g., spots).

[0041] In one method of generating a digital three-dimensional image as defined by the present invention, the method includes driving each light projector of one or more structured light projectors to project a pattern onto a three-dimensional surface within the oral cavity. The method further includes driving each camera of one or more cameras to capture a plurality of images, each image including at least a portion of the projected pattern. The method further includes using the processor to compare a series of images captured by the one or more cameras, and based on the comparison of the series of images, determining which portions of the projected pattern can be tracked across the series of images, and constructing a three-dimensional model of the three-dimensional surface within the oral cavity based at least in part on the comparison of the series of images. In one embodiment, the method further includes solving a correspondence algorithm for the tracked portion of the projected pattern in at least one image of the series of images, and using the solved correspondence algorithm to address the tracked portion of the projected pattern in at least one image of the series of images, for example, solving a correspondence algorithm for the tracked portion of the projected pattern in an image of a series of images in which the correspondence algorithm has not been solved, and constructing a three-dimensional model using the solution of the correspondence algorithm. In one embodiment, the method further includes solving a correspondence algorithm for the tracked portion of the projected pattern based on a portion of the tracked position of the tracked portion in each image throughout the series of images, and constructing a three-dimensional model using the solution of the correspondence algorithm.

[0042] In one embodiment of the method, the projection pattern includes a plurality of projection light spots, and the portion of the projection pattern corresponds to the projection spots of the plurality of projection light spots. In a further embodiment, using the processor, (a) camera rays corresponding to each pixel on the camera sensor of each of the one or more cameras, and (b) projector rays corresponding to each one projection light spot from each of the one or more structured light projectors, each projector ray corresponding to the path of each pixel on at least one of the camera sensors of the camera sensors, and comparing a series of images based on the stored calibration values indicating the same, and determining which portion of the projection pattern can be tracked includes determining which of the projection spots s can be tracked across a series of images, and each tracking spot s moves along the path of the pixel corresponding to each respective projector ray r.

[0043] In a further embodiment of the method, the step of using the processor further includes using the processor to determine, for each tracking spot s, a plurality of possible paths p of pixels on a given one of the cameras, each path p corresponding to a respective plurality of possible projector rays r. In a further embodiment, the step of using the processor further includes using the processor to execute a corresponding algorithm to perform a plurality of operations for each of the possible projector rays r. The plurality of operations includes identifying, for each camera ray corresponding to each spot q that intersects the projector ray r and the camera ray of a given one of the cameras corresponding to the tracking spot s, how many other cameras detected it on their respective paths p1 of pixels corresponding to the projector ray r. The operations further include identifying a given projector ray r1 where the most other cameras detected each spot q. The operations further include identifying the projector ray r1 as the particular projector ray r that generated the tracking spot s.

[0044] In a further embodiment of the method, the method uses the processor to: (a) execute a corresponding algorithm to calculate the three-dimensional position of each of a plurality of detection spots on the intraoral three-dimensional surface captured in a series of images; and (b) in at least one image of the series of images, identify that a detection spot is a tracking spot s that moves along a path of pixels corresponding to a specific projector ray r by identifying that the detection spot is derived from the specific projector ray r.

[0045] In a further embodiment of the method, the method uses the processor to: (a) execute a corresponding algorithm to calculate the three-dimensional position of each of a plurality of detection spots on the intraoral three-dimensional surface captured in the series of images; and (b) exclude from consideration as points on the intraoral three-dimensional surface spots that are (i) identified as being from a specific projector ray r based on the three-dimensional positions calculated by the corresponding algorithm and (ii) not identified as tracking spots s that move along a path of pixels corresponding to the specific projector ray r.

[0046] In a further embodiment of the method, the method uses the processor to: (a) execute the corresponding algorithm to calculate the three-dimensional position of each of a plurality of detection spots on the intraoral three-dimensional surface captured in the series of images; and (b) for detection spots identified as being derived from two different projector rays r based on the three-dimensional positions calculated by the corresponding algorithm, identify the spot as being derived from one of the two different projector rays r by identifying that the detection spot is a tracking spot s that moves along one of the two different projector rays r.

[0047] In a further embodiment of the method, the method uses the processor to: (a) execute the corresponding algorithm to calculate the three-dimensional position of each of a plurality of detection spots on the three-dimensional surface within the oral cavity captured in a series of images; and (b) identify, by identifying a vulnerable spot not calculated for its three-dimensional position by the corresponding algorithm as a tracking spot s moving along a path of pixels corresponding to a specific projector ray r, that it is a projection spot from the specific projector ray r.

[0048] In a further embodiment of the method, the method uses the processor to calculate the three-dimensional position of each on the three-dimensional surface within the oral cavity at the intersection of the projector ray r and each of the camera rays corresponding to the tracking spot s in each of the series of images in which the spot s was tracked.

[0049] In a further embodiment of the method, the three-dimensional model is constructed using a corresponding algorithm that at least partially uses portions of the projection pattern determined to be trackable across the series of images.

[0050] In a further embodiment of the method, the method uses the processor to: (a) determine parameters of a trackable portion of the projection pattern in at least two adjacent images from the series of images, the parameters being selected from the group consisting of the size of the portion, the shape of the portion, the orientation of the portion, the intensity of the portion, and the signal-to-noise ratio (SNR) of the portion; and (b) predict parameters of the trackable portion of the projection pattern in a later image based on the parameters of the trackable portion of the projection pattern in the at least two adjacent images.

[0051] In a further embodiment of the method, the step of using the processor further includes using the processor to search for a portion of the projection pattern in the subsequent image that substantially has the prediction parameters based on the prediction parameters of the tracking portion of the projection pattern.

[0052] In a further embodiment of the method, the parameter is the shape of the portion of the projection pattern, and the step of using the processor further includes using the processor to determine a search space in a next image for searching for the tracking portion of the projection pattern based on the predicted shape of the tracking portion of the projection pattern.

[0053] In a further embodiment of the method, the step of determining the search space using the processor includes using the processor to determine a search space in the next image for searching for the tracking portion of the projection pattern, and the search space has a size and aspect ratio based on the size and aspect ratio of the predicted shape of the tracking portion of the projection pattern.

[0054] In a further embodiment of the method, the parameter is the shape of the portion of the projection pattern, and the step of using the processor further includes using the processor to: (a) determine a velocity vector of the tracking portion of the projection pattern based on a direction and distance that the tracking portion of the projection pattern has moved between at least two adjacent images from the series of images; (b) predict a shape of the tracking portion of the projection pattern in a subsequent image in response to the shape of the tracking portion of the projection pattern in at least one of the at least two adjacent images; and (c) determine a search space for searching for the tracking portion of the projection pattern in the subsequent image in response to a combination of (i) the determination of the velocity vector of the tracking portion of the projection pattern and (ii) the predicted shape of the tracking portion of the projection pattern.

[0055] In a further embodiment of the method, the parameter is the shape of the portion of the projection pattern, and the step of using the processor further comprises using the processor to: (a) determine a velocity vector of the tracking portion of the projection pattern based on the direction and distance that the tracking portion of the projection pattern has moved between at least two adjacent images from the series of images; (b) predict the shape of the tracking portion of the projection pattern in a subsequent image in response to determining the velocity vector of the tracking portion of the projection pattern; and (c) determine a search space for searching for the tracking portion of the projection pattern in the subsequent image in response to a combination of (i) determining the velocity vector of the tracking portion of the projection pattern and (ii) the predicted shape of the tracking portion of the projection pattern.

[0056] In a further embodiment of the method, the step of using the processor further comprises using the processor to predict the shape of the tracking portion of the projection pattern in a subsequent image in response to a combination of (i) determining the velocity vector of the tracking portion of the projection pattern and (ii) the shape of the tracking portion of the projection pattern in at least one of the two adjacent images.

[0057] In a further embodiment of the method, the step of using the processor further comprises using the processor to: (a) determine a velocity vector of the tracking portion of the projection pattern based on the direction and distance that the tracking portion of the projection pattern has moved between two consecutive images in the series of images; and (b) determine a search space for searching for the tracking portion of the projection pattern in a subsequent image in response to determining the velocity vector of the tracking portion of the projection pattern.

[0058] In one embodiment of a second method of generating a digital three-dimensional image as defined herein, the method includes driving each structured light projector of one or more structured light projectors to project a pattern of light along a plurality of projector rays onto an intraoral three-dimensional surface, and driving each camera of one or more cameras to capture a plurality of images, each image including at least a portion of the projected pattern, and each camera of the one or more cameras including a camera sensor including an array of pixels. The second method further includes using the processor to execute a corresponding algorithm to calculate, for each of the plurality of images, the respective three-dimensional positions on the intraoral three-dimensional surface of a plurality of detected features of the projected pattern, and estimating a three-dimensional surface based on at least three of the features using data corresponding to the respective three-dimensional positions of at least three of the features, each feature corresponding to a respective projector ray r of the plurality of projector rays, and estimating a three-dimensional position in the intersection space between a projector ray r1 of the plurality of projector rays for which the three-dimensional position of the feature corresponding to the projector ray r1 has not been calculated and the estimated three-dimensional surface, and identifying a search space in the pixel array of at least one of the cameras for searching for the feature corresponding to the projector ray r1 using the estimated three-dimensional position within the estimated space.

[0059] In a further embodiment of the second method, the corresponding algorithm is executed based on stored calibration values indicating (a) camera rays corresponding to each pixel on the camera sensor of each of the one or more cameras, and (b) projector rays corresponding to each feature of the projected pattern from each of the one or more structured light projectors, each projector ray corresponding to the path of each pixel on at least one of the camera sensors of the camera sensors. Further, the search space in the data includes a search space defined by one or more thresholds.

[0060] In a further embodiment of the second method, the processor sets a threshold such that detected features below the threshold are not considered by the corresponding algorithm, and in order to search for features corresponding to the projector ray r1 in the specified search space, the processor lowers the threshold in order to consider features that were not considered by the corresponding algorithm. In some embodiments, the threshold is an intensity threshold.

[0061] In a further embodiment of the second method, the pattern of light includes a distribution of discrete spots, and each feature includes a spot from the distribution of discrete spots.

[0062] In a further embodiment of the second method, the step of using data corresponding to the three-dimensional position of each of the at least three features includes using data corresponding to the three-dimensional position of at least three features all captured in one of the plurality of images.

[0063] In a further embodiment of the second method, the second method further includes the step of refining the estimation of the three-dimensional surface using data corresponding to the three-dimensional position of at least one additional feature of the projection pattern, the at least one additional feature having a three-dimensional position calculated based on another one of the plurality of images. In a further embodiment of the second method, the step of refining the estimation of the three-dimensional surface includes the step of refining the estimation of the three-dimensional surface such that all of the at least three features and the at least one additional feature are located on the estimated three-dimensional surface.

[0064] In a further embodiment of the second method, the step of using data corresponding to the three-dimensional position of each of the at least three features includes using data corresponding to at least three features captured in each of the plurality of images.

[0065] In one embodiment of a third method for generating a digital three-dimensional image, the third method includes driving each structured light projector of one or more structured light projectors to project a pattern of light onto a three-dimensional surface within the oral cavity, and driving each camera of a plurality of cameras to capture an image, the image including at least a portion of the projected pattern, and each camera of the plurality of cameras including a camera sensor including an array of pixels. The third method further includes using a processor to execute a corresponding algorithm to calculate the respective three-dimensional positions on the three-dimensional surface within the oral cavity of a plurality of features of the projected pattern, using data from a first camera of the plurality of cameras to identify candidate three-dimensional positions of a given feature of the projected pattern corresponding to or otherwise associated with one or more specific projector rays (plurality) r, data from a second camera of the plurality of cameras not being used to identify those candidate three-dimensional positions, identifying a search space on the pixel array of the second camera to search for features of the projected pattern from the projector ray r using the candidate three-dimensional positions seen by the first camera, and refining the candidate three-dimensional positions of the features of the projected pattern using data from the second camera when the features of the projected pattern from the projector ray r are identified within the search space.

[0066] In a further embodiment of the third method, to identify candidate three-dimensional positions of a given spot corresponding to a specific projector ray r, the processor uses data from at least two of the cameras, data from another camera other than one of the at least two cameras not being used to identify those candidate three-dimensional positions, and to identify the search space, the processor uses the candidate three-dimensional positions seen by at least one of the at least two cameras.

[0067] In a further embodiment of the third method, the pattern of light includes a distribution of discrete non-connected light spots, and the features of the projected pattern include projected spots from the non-connected light spots.

[0068] In a further embodiment of the third method, the processor uses stored calibration values indicating (a) camera rays corresponding to each pixel on the camera sensor of each of the plurality of cameras, and (b) projector rays corresponding to each feature of the projection light pattern from each of the one or more structured light projectors, wherein each projector ray corresponds to the path of each pixel of at least one of the camera sensors.

[0069] In a fourth method of generating a digital three-dimensional image described herein, the fourth method includes driving each of one or more structured light projectors to project a pattern of light onto an intraoral three-dimensional surface, and driving each of one or more cameras to capture an image, the image including at least a portion of the pattern. The fourth method further includes using the processor to execute a corresponding algorithm to calculate the three-dimensional position of each of a plurality of features of the pattern on the intraoral three-dimensional surface captured in a series of images, identifying the calculated three-dimensional position of the detected features of the imaged pattern as being related to one or more specific projector rays r in at least a subset of the series of images, and evaluating the length related to the one or more projector rays r in each image of the subset of images based on the three-dimensional position of the detected features corresponding to the one or more projector rays r in the subset of the images.

[0070] In the fourth method, the processor may further be used to calculate an estimated length of the one or more projector rays r in at least one image of a series of images in which the three-dimensional position of a feature projected from the one or more projector rays is not identified.

[0071] In one embodiment of the fourth method, each of the one or more cameras comprises a camera sensor including a pixel array, and the calculation of the three-dimensional position of each of a plurality of features of the pattern on the intraoral three-dimensional surface and the identification that the calculated three-dimensional positions of the detected features of the pattern correspond to a specific projector ray r are performed based on stored calibration values indicating (i) camera rays corresponding to the pixels of the camera sensor of each of the one or more camera sensors, and (ii) projector rays corresponding to each of the features of the projected light pattern from each of the one or more projectors, each projector ray corresponding to the path of a respective pixel on at least one of the camera sensors.

[0072] In a further embodiment of the fourth method, the step of using the processor further comprises using the processor to calculate an estimated length of the projector ray r in at least one image of the series of images in which the three-dimensional position of the feature projected from the projector ray r is not identified, and based on the estimated length of the projector ray r in at least one image of the series of images, determining, in at least one image of the series of images, a one-dimensional search space for searching for the feature projected from the projector ray r, the one-dimensional search space being along the path of the respective pixels corresponding to the projector ray r.

[0073] In a further embodiment of the fourth method, the step of using the processor further comprises using the processor to calculate an estimated length of the projector ray r in at least one image of the series of images in which the three-dimensional position of the feature projected from the projector ray r is not identified, and based on the estimated length of the projector ray r in the at least one image of the series of images, for each pixel array, determining a one-dimensional search space within the pixel array of each of a plurality of cameras for searching for a spot projected from the projector ray r, the one-dimensional search space being along the path of the respective pixels corresponding to the ray r.

[0074] In a further embodiment of the fourth method, the step of using the processor to determine a one-dimensional search space in the pixel array of each of the plurality of cameras includes the step of using the processor to determine a one-dimensional search space in the pixel array of each of all the cameras that search for features projected from the projector ray r.

[0075] In a further embodiment of the fourth method, the step of using the processor further includes the steps of using the processor to identify, based on the corresponding algorithm, a plurality of candidate 3D positions of the features projected from the projector ray r in at least one image of a series of images that are not in the subset of the images, and calculating an estimated length of the projector ray r within at least one image of the series of images in which a plurality of candidate 3D positions of the features projected from the projector ray r are identified.

[0076] In a further embodiment of the fourth method, the step of using the processor further includes the step of using the processor to determine whether the correct 3D position of the projected feature is by determining which of the plurality of candidate 3D positions corresponds to the estimated length of the projector ray r in at least one image of the series of images.

[0077] In a further embodiment of the fourth method, the step of using the processor further includes the steps of using the processor to determine a one-dimensional search space in at least one image of the series of images that search for features projected from the projector ray r based on the estimated length of the projector ray r in at least one image of the series of images, and determining which of the plurality of candidate 3D positions corresponds to the feature generated by the projector ray r detected within the one-dimensional search space, thereby determining which of the plurality of candidate 3D positions of the projected feature is the correct 3D position of the projected feature generated by the projector ray r.

[0078] In a further embodiment of the fourth method, the step of using the processor comprises using the processor to define a curve based on the evaluated lengths of the projector rays r in each image of the subset of images, and removing from consideration as points on the intraoral three-dimensional surface a detected feature identified as originating from a projector ray r when the three-dimensional position of the projected feature is at least a threshold distance away from the defined curve.

[0079] In a further embodiment of the fourth method, the pattern includes a plurality of spots, and each of the plurality of features of the pattern includes one of the plurality of spots.

[0080] In a fifth method of generating a digital three-dimensional image as defined herein, the method includes driving each structured light projector of one or more structured light projectors to project a pattern of light onto the intraoral three-dimensional surface along a plurality of projector rays, and driving each camera of one or more cameras to capture a plurality of images, each image including at least a portion of the projected pattern, and each camera of the one or more cameras including a camera sensor comprising an array of pixels. The method further includes using the processor to execute a corresponding algorithm to calculate for each of the plurality of images the respective three-dimensional positions on the intraoral three-dimensional surface of a plurality of detected features of the projected pattern, estimating a three-dimensional surface based on at least three features using data corresponding to the respective three-dimensional positions of at least three of the detected features, each feature corresponding to a respective projector ray r of the plurality of projector rays, estimating for the projector ray r1 a three-dimensional position within the intersection space between the projector ray r1 and the estimated three-dimensional surface for a plurality of candidate three-dimensional positions of the feature corresponding to the projector ray r1 that have been calculated for the plurality of projector rays r1, and selecting, using the estimated three-dimensional position within the intersection space of the projector ray r1, which of the plurality of candidate three-dimensional positions is the correct three-dimensional position of the feature corresponding to that projector r1.

[0081] In a further embodiment of the fifth method, the corresponding algorithm is executed based on stored calibration values indicating (a) camera rays corresponding to each pixel on the camera sensor of each of the one or more cameras, and (b) projector rays corresponding to each feature of the projection pattern from each of the one or more structured light projectors, wherein each projector ray corresponds to a respective path on at least one pixel of the camera sensor. Further, the search space in the data includes a search space defined by one or more thresholds.

[0082] In a further embodiment of the fifth method, the pattern of light includes a distribution of discrete spots, and each feature includes one spot from the distribution of discrete spots.

[0083] In a further embodiment of the fifth method, the step of using data corresponding to the three-dimensional positions of each of the at least three features includes using data corresponding to the three-dimensional positions of at least three features all captured in one of the plurality of images.

[0084] In a further embodiment of the fifth method, the fifth method further includes refining the estimation of the three-dimensional surface using data corresponding to the three-dimensional position of at least one additional feature of the projection pattern, wherein the at least one additional feature has a three-dimensional position calculated based on another one of the plurality of images. In a further embodiment of the fifth method, the step of refining the estimation of the three-dimensional surface includes refining the estimation of the three-dimensional surface such that all of the at least three features and the at least one additional feature are located on the estimated three-dimensional surface.

[0085] In a further embodiment of the fifth method, the step of using data corresponding to the three-dimensional positions of each of the at least three features includes using data corresponding to at least three features respectively captured in each one of the plurality of images.

[0086] In one method of tracking the motion of an intraoral scanner as defined herein, the method includes using at least one camera coupled to the intraoral scanner to measure the motion of the intraoral scanner relative to the intraoral surface being scanned, and using at least one inertial measurement unit (IMU) coupled to the intraoral scanner to measure the motion of the intraoral scanner relative to the intraoral surface being scanned in a fixed coordinate system. The method further includes using the processor to calculate the motion of the intraoral surface relative to the fixed coordinate system based on (a) the motion of the intraoral scanner relative to the intraoral surface and (b) the motion of the intraoral scanner relative to the fixed coordinate system, constructing a prediction model regarding the motion of the intraoral surface relative to the fixed coordinate system based on the accumulated data of the motion of the intraoral surface relative to the fixed coordinate system, and further calculating the estimated position of the intraoral scanner relative to the intraoral surface based on (a) the prediction of the motion of the intraoral surface relative to the fixed coordinate system (derived based on the motion prediction model) and (b) the motion of the intraoral scanner relative to the fixed coordinate system (measured by the IMU). In a further embodiment of the method of tracking motion, the method further includes determining whether the measurement of the motion of the intraoral scanner relative to the intraoral surface using the at least one camera is inhibited, and in response to determining that the measurement of the motion is inhibited, calculating the estimated position of the intraoral scanner relative to the intraoral surface. In a further embodiment of the method of tracking motion, the calculation of the motion is performed by calculating the difference between (a) the motion of the intraoral scanner relative to the intraoral surface and (b) the motion of the intraoral scanner relative to the fixed coordinate system.

[0087] One method of determining whether calibration data of an intraoral scanner, as defined herein, is inaccurate includes driving each of one or more light sources to project light onto an intraoral three-dimensional surface, and driving each of one or more cameras to capture a plurality of images of the intraoral three-dimensional surface. The method further includes using a processor to execute a corresponding algorithm based on stored calibration data regarding the one or more light sources and the one or more cameras to calculate, for each of a plurality of features of the projected light, a respective three-dimensional position on the intraoral three-dimensional surface, collecting data at a plurality of time points, the data including the respective calculated three-dimensional positions on the intraoral three-dimensional surface of the plurality of features, and based on the collected data, determining that at least a portion of the stored calibration data is inaccurate.

[0088] In a further embodiment of the method, the one or more light sources are one or more structured light projectors, and the method includes driving each of the one or more structured light projectors to project a pattern of light onto the intraoral three-dimensional surface, and driving each of the one or more cameras to capture a plurality of images of the intraoral three-dimensional surface, each image including at least a portion of the projected pattern, each of the one or more cameras comprising a camera sensor including an array of pixels, and the stored calibration data including stored calibration values indicative of (a) camera rays corresponding to each pixel on the camera sensor of each of the one or more cameras, and (b) projector rays corresponding to each feature of the pattern of light projected from each of the one or more structured light projectors, each projector ray corresponding to a path p of a respective pixel on at least one of the camera sensors.

[0089] In a further embodiment of the method, the step of determining that at least a portion of the stored calibration data is inaccurate includes, using the processor, for each projector ray r, based on the collected data, defining for each camera sensor a pixel update path p' such that all of the calculated 3D positions corresponding to the features generated by the projector ray r are along the positions along the respective pixel update path p' of each pixel of the camera sensor; comparing the pixel update path p' of each pixel with the pixel path p corresponding to that projector ray r of each camera sensor from the stored calibration values; and in response to the update path p' for at least one camera sensor s being different from the pixel path p corresponding to that projector ray r from the stored calibration values, determining that at least some of the stored calibration values are inaccurate.

[0090] One method of recalibration as defined herein includes driving each of one or more light sources to project light onto an intraoral three-dimensional surface and driving each of one or more cameras to capture a plurality of images of the intraoral three-dimensional surface. The method includes using a processor to execute a corresponding algorithm to calculate, based on the stored calibration data of the one or more light sources and the one or more cameras, the respective three-dimensional positions on the intraoral three-dimensional surface of a plurality of features of the projected light; collecting data at a plurality of time points, the data including the respective calculated three-dimensional positions on the intraoral three-dimensional surface of the plurality of features; and using the collected data to perform recalibration of the stored calibration data.

[0091] In a further embodiment of the recalibration method, the one or more light sources are one or more structured light projectors, and the method includes driving each of the one or more structured light projectors to project a pattern of light onto the three-dimensional surface within the oral cavity, and driving each of the one or more cameras to capture a plurality of images of the three-dimensional surface within the oral cavity, each image including at least a portion of the projected pattern, and each of the one or more cameras comprising a camera sensor including an array of pixels. The processor executes a plurality of operations using the stored calibration data, the stored calibration data including stored calibration values indicating (a) camera rays corresponding to each pixel on the camera sensor of each of the one or more cameras, and (b) projector rays corresponding to each feature of the light pattern projected from each of the one or more structured light projectors, each projector ray corresponding to a path p of each pixel on at least one of the camera sensors. The operations include executing a corresponding algorithm to calculate the respective three-dimensional positions of a plurality of features of the projected pattern on the three-dimensional surface within the oral cavity. The operations further include collecting data at a plurality of time points, the data including the respective calculated three-dimensional positions of the plurality of features on the three-dimensional surface within the oral cavity. The operations further include, for each projector ray r, based on the collected data, defining an updated path p' of pixels for each of the camera sensors such that all of the calculated three-dimensional positions corresponding to the features generated by the projector ray r correspond to positions along the updated path p' of each pixel for each of the camera sensors. The operations further include recalibrating the stored calibration values using the updated path p'.

[0092] In a further embodiment of the recalibration method, to recalibrate the stored calibration values, the processor performs additional operations. The additional operations include the step of comparing each updated path p' of a pixel with the path p of the pixel on its projector ray r on each camera sensor from the stored calibration values. The additional operations further include, for at least one camera sensor s, when the updated path p' of the pixel corresponding to the projector ray r is different from the path p of the pixel corresponding to the projector ray r from the stored calibration values, (i) the stored calibration values indicating the camera rays corresponding to each pixel on the camera sensor s of each of the one or more cameras, and (ii) the stored calibration values indicating the projector rays r corresponding to each feature projected from each of the one or more structured light projectors, by changing the stored calibration data selected from the group consisting of, reducing the difference between the updated path p' of the pixel corresponding to each projector ray r and the path p of the pixel corresponding to each projector ray r from the stored calibration values.

[0093] In a further embodiment of the recalibration method, the stored calibration data to be changed includes the stored calibration values indicating the camera rays corresponding to each pixel on the camera sensor s of each of the one or more cameras. Further, the step of changing the stored calibration data includes changing one or more parameters of the parameterized camera calibration function defining the camera rays corresponding to each pixel on at least one camera sensor s to reduce the difference between (i) the calculated respective three-dimensional positions on the intraoral three-dimensional surface of the plurality of features of the projection pattern and (ii) the stored calibration values indicating the respective camera rays corresponding to each pixel on each camera sensor on which each of the plurality of features is to be detected.

[0094] In a further embodiment of the recalibration method, the stored calibration data to be changed includes stored calibration values indicating projector light rays corresponding to each feature of the plurality of features from each structured light projector of the one or more structured light projectors, and the step of changing the stored calibration data includes (i) an indexed list that assigns each projector light ray r to a pixel path p, or (ii) changing one or more parameters of a parameterized projector calibration model that defines each projector light ray r.

[0095] In a further embodiment of the recalibration method, the step of changing the stored calibration data includes changing the indexed list by reassigning each projector light ray r based on the respective updated path p' of the pixels corresponding to each projector light ray r.

[0096] In a further embodiment of the recalibration method, the step of changing the stored data includes changing (i) stored calibration values indicating camera light rays corresponding to each pixel on the camera sensor s of each of the one or more cameras, and (ii) stored calibration values indicating projector light rays r corresponding to each feature of the plurality of features from each structured light projector of the one or more structured light projectors.

[0097] In a further embodiment of the recalibration method, the step of changing the stored calibration values includes repeatedly changing the stored calibration values.

[0098] In a further embodiment of the recalibration method, the method further includes driving each of the one or more cameras to capture a plurality of images of a calibration object having predetermined parameters. The recalibration method further includes using the processor to execute a triangulation algorithm to calculate respective parameters of the calibration object based on the captured images, and executing an optimization algorithm to (a) reduce the difference between (i) an updated path p' of a pixel corresponding to the projector ray r and (ii) a path p of the pixel corresponding to the projector ray r from the stored calibration value, and (b) using respective parameters of the calibration object calculated based on the captured images.

[0099] In a further embodiment of the recalibration method, the calibration object is a three-dimensional calibration object of a known shape, and the step of driving each of the one or more cameras to capture a plurality of images of the calibration object includes driving each of the one or more cameras to capture an image of the three-dimensional calibration object, and the predetermined parameters of the calibration object are the dimensions of the three-dimensional calibration object. In a further embodiment, with respect to the use of the calculated respective parameters of the calibration object by the processor to execute the optimization algorithm, the processor further uses the collected data including the calculated respective three-dimensional positions of the plurality of features on the intraoral three-dimensional surface.

[0100] In a further embodiment of the re - calibration method, the object to be calibrated is a two - dimensional calibration object having visually distinguishable features. The step of driving each of the one or more cameras to capture a plurality of images of the calibration object includes the step of driving each of the one or more cameras to capture an image of the two - dimensional calibration object. The predetermined parameters of the two - dimensional calibration object are the respective distances between the respective visually distinguishable features. In a further embodiment, with respect to the processor using the calculated respective parameters of the calibration object to execute the optimization algorithm, the processor further uses the collected data including the calculated respective three - dimensional positions of the plurality of features on the three - dimensional surface within the oral cavity.

[0101] In a further embodiment of the re - calibration method, the step of driving each of the one or more cameras to capture an image of the two - dimensional calibration object includes the step of driving each of the one or more cameras to capture a plurality of images of the two - dimensional calibration object from a plurality of different viewpoints with respect to the two - dimensional calibration object.

[0102] In one embodiment of an intra - oral scanning device, the device includes an elongated hand - held wand. The elongated hand - held wand includes a probe at its distal end, one or more illumination sources coupled to the probe, one or more near - infrared (NIR) light sources coupled to the probe, and one or more cameras coupled to the probe and configured to (a) capture an image using light from the one or more illumination sources and (b) capture an image using NIR light from the NIR light sources. The device further includes a processor configured to execute a navigation algorithm to determine the position of the elongated hand - held wand as the elongated hand - held wand moves within the space. The input to the navigation algorithm includes (a) the images captured using light from the one or more illumination sources and (b) the images captured using the NIR light.

[0103] In a further embodiment of the intraoral scanning device, the one or more illumination sources include one or more structured light sources.

[0104] In a further embodiment of the intraoral scanning device, the one or more illumination sources include one or more non - coherent light sources.

[0105] A method for tracking the motion of an intraoral scanner includes illuminating an intraoral three - dimensional surface using one or more illumination sources coupled to the intraoral scanner; driving each NIR light source of the one or more NIR light sources coupled to the intraoral scanner to emit NIR light onto the intraoral three - dimensional surface; using one or more cameras coupled to the intraoral scanner to (a) capture a first plurality of images using light from the one or more illumination light sources and (b) capture a second plurality of images using the NIR light. The method further includes using the processor to execute a navigation algorithm to track the motion of the intraoral scanner relative to the intraoral three - dimensional surface using (a) the first plurality of images captured using light from the one or more illumination sources and (b) the second plurality of images captured using the NIR light.

[0106] In one embodiment of the method for tracking motion, the step of using the one or more illumination sources includes illuminating the intraoral three - dimensional surface.

[0107] In one embodiment of the method for tracking motion, the step of using the one or more illumination sources includes using one or more non - coherent light sources.

[0108] One embodiment of a sixth method for calculating the three-dimensional structure of an intraoral three-dimensional surface includes driving one or more structured light projectors to project a structured light pattern onto the intraoral three-dimensional surface; driving one or more cameras to capture a plurality of structured light images, each structured light image including at least a portion of the structured light pattern; driving one or more unstructured light projectors to project unstructured light onto the intraoral three-dimensional surface; and driving the one or more cameras to capture a plurality of two-dimensional images of the intraoral three-dimensional surface. The sixth method further includes using the processor to calculate the three-dimensional position of each of a plurality of points on the intraoral three-dimensional surface captured in the plurality of structured light images, and calculating the three-dimensional structure of the intraoral three-dimensional surface constrained by some or all of the calculated three-dimensional positions of the plurality of points based on the plurality of two-dimensional images of the intraoral three-dimensional surface.

[0109] In some embodiments of the sixth method, the unstructured light is incoherent light and the plurality of two-dimensional images include a plurality of color two-dimensional images.

[0110] In some embodiments of the sixth method, the unstructured light is near-infrared (NIR) light and the plurality of two-dimensional images include a plurality of monochromatic NIR images.

[0111] In a further embodiment of the sixth method, the step of driving the one or more structured light projectors includes driving the one or more structured light projectors to project respective distributions of discrete non-connected light spots.

[0112] In a further embodiment of the sixth method, the step of calculating the three-dimensional structure includes inputting the plurality of two-dimensional images of the intraoral three-dimensional surface into a neural network, and determining, by the neural network, an estimated map of the intraoral three-dimensional surface captured in each of the two-dimensional images.

[0113] In a further embodiment of the sixth method, the sixth method further includes inputting the calculated three-dimensional positions of a plurality of points on the intraoral three-dimensional surface into the neural network.

[0114] In a further embodiment of the sixth method, the sixth method further includes using the processor to stitch together each map to obtain the three-dimensional structure of the intraoral three-dimensional surface.

[0115] In a further embodiment of the sixth method, the sixth method further includes adjusting the capture of the structured light image and the capture of the two-dimensional image to generate an alternating sequence in which one or more image frames of unstructured light are interspersed among one or more image frames of structured light.

[0116] In a further embodiment of the sixth method, the determining step includes determining, by the neural network, each estimated depth map of the intraoral three-dimensional surface captured in each two-dimensional image. In one embodiment, the processor is used to stitch together each of the estimated depth maps to obtain the three-dimensional structure of the intraoral three-dimensional surface.

[0117] In one embodiment, (a) the processor generates each point cloud corresponding to the calculated three-dimensional position of each of the plurality of points on the intraoral three-dimensional surface captured in each structured light image, and the method further includes using the processor to stitch together each of the estimated depth maps with each of the point clouds. In one embodiment, the method further includes determining, by the neural network, each estimated normal map of the intraoral three-dimensional surface captured in each of the two-dimensional images.

[0118] In a further embodiment of the sixth method, the determining step includes determining, by the neural network, an estimated normal map for each of the intraoral three-dimensional surfaces captured in each of the two-dimensional images. In one embodiment, the processor is used to stitch together the respective estimated normal maps to obtain the three-dimensional structure of the intraoral three-dimensional surface.

[0119] In one embodiment, the method further includes interpolating three-dimensional positions on the intraoral three-dimensional surface between the calculated respective three-dimensional positions of the plurality of points on the intraoral three-dimensional surface captured in the plurality of structured light images based on the respective estimated normal maps for each of the intraoral three-dimensional surfaces captured in each of the two-dimensional images.

[0120] In one embodiment, the method further includes adjusting the capture of the structured light images and the capture of the two-dimensional images to produce an alternating sequence in which one or more unstructured light image frames are interspersed among one or more structured light image frames. The method further uses the processor to: (a) generate respective point clouds corresponding to the calculated respective three-dimensional positions of the plurality of points on the intraoral three-dimensional surface captured in each image frame of the structured light; and (b) stitch together the respective point clouds for at least a subset of the plurality of points of each point cloud, using the surface normals at each point of the subset of points as inputs for the stitching, wherein for a given point cloud, the surface normals at at least one point of the subset of points are obtained from the respective estimated normal maps of the intraoral three-dimensional surface captured in adjacent unstructured light image frames.

[0121] In one embodiment, the method further includes compensating for the motion of the intraoral scanner between a structured light image frame and an adjacent unstructured light image frame by estimating the motion of the intraoral scanner based on a previous image frame using the processor.

[0122] In a further embodiment of the sixth method, the determining step includes determining, by the neural network, the curvature of the intraoral three-dimensional surface captured in each of the two-dimensional images. In one embodiment, the determining step comprises determining, by the neural network, an estimated curvature map for each of the intraoral three-dimensional surfaces captured in each of the two-dimensional images.

[0123] In one embodiment, the method further comprises using the processor to evaluate the curvature of the intraoral three-dimensional surface captured in each of the two-dimensional images and interpolating the three-dimensional position of the intraoral three-dimensional surface between the calculated respective three-dimensional positions of the plurality of points of the intraoral three-dimensional surface captured in the plurality of structured light images based on the evaluated curvature of the intraoral three-dimensional surface captured in each of the two-dimensional images.

[0124] In a further embodiment of the sixth method, the sixth method further comprises adjusting the capture of the structured light images and the capture of the two-dimensional images to produce an alternating sequence in which one or more unstructured light image frames are interspersed among one or more structured light image frames.

[0125] In a further embodiment of the sixth method, the sixth method includes driving the one or more cameras to capture a plurality of structured light images, which includes driving each of the two or more cameras to capture respective pluralities of structured light images, and driving the one or more cameras to capture a plurality of two-dimensional images includes driving each of the two or more cameras to capture respective pluralities of two-dimensional images.

[0126] In one embodiment, the step of driving the two or more cameras includes driving each of the two or more cameras in a given image frame to simultaneously capture respective two-dimensional images of respective portions of the three-dimensional surface within the oral cavity. The step of inputting into the neural network includes, in a given image frame, inputting all of the respective two-dimensional images as a single input into the neural network, and each of the respective two-dimensional images has a field of view that overlaps with at least one other two-dimensional image among the respective two-dimensional images. The step of determining by the neural network includes, for a given image frame, determining an estimated depth map of the three-dimensional surface within the oral cavity by combining respective portions of the three-dimensional surface within the oral cavity.

[0127] In one embodiment, the step of driving the two or more cameras to capture a plurality of structured light images includes driving each of the three or more cameras to capture respective pluralities of structured light images, and the step of driving the two or more cameras to capture a plurality of two-dimensional images includes driving each of the three or more cameras to capture respective pluralities of two-dimensional images. In a given image frame, each of the three or more cameras is driven to simultaneously capture respective two-dimensional images of respective portions of the three-dimensional surface within the oral cavity. The step of inputting into the neural network includes, for a given image frame, inputting a subset of the respective two-dimensional images as a single input into the neural network, the subset includes at least two images of the respective two-dimensional images, and each of the subsets of the respective two-dimensional images has a field of view that overlaps with at least one other of the subsets of the respective two-dimensional images. The step of determining by the neural network includes, for a given image frame, determining an estimated depth map of the three-dimensional surface within the oral cavity by combining respective portions of the three-dimensional surface within the oral cavity captured by the subsets of the respective two-dimensional images.

[0128] In one embodiment, the step of driving the two or more cameras includes, in a given image frame, driving each of the two or more cameras to simultaneously capture respective two-dimensional images of respective portions of the three-dimensional surface within the oral cavity, and the step of inputting into the neural network includes, in a given image frame, inputting each of the respective two-dimensional images as a separate input into the neural network. The step of determining by the neural network includes, for a given image frame, determining respective estimated depth maps of respective portions of the three-dimensional surface within the oral cavity captured in each of the respective two-dimensional images captured in the given image frame.

[0129] In one embodiment, the method further includes using the processor to merge the respective depth maps to obtain a composite estimated depth map of the three-dimensional surface within the oral cavity captured in the given image frame. In one embodiment, the method further includes the step of training the neural network, and each input to the neural network during the training includes an image captured by only one camera.

[0130] In one embodiment, the method further includes the step of determining, by the neural network, respective estimated confidence maps corresponding to the respective estimated depth maps, and each confidence map indicates a confidence level for each region of the respective estimated depth map. In a further embodiment, the step of merging the respective estimated depth maps includes using the processor to merge at least two of the estimated depth maps based on the confidence levels of the corresponding respective regions indicated by the respective confidence maps for at least two of the estimated depth maps in response to determining a conflict between corresponding respective regions in at least two of the estimated depth maps.

[0131] In a further embodiment of the sixth method, the step of driving the one or more cameras includes driving one or more cameras of an intraoral scanner, and the method further includes training the neural network using training stage images captured by a plurality of training stage handheld wands. Each training stage handheld wand includes one or more reference cameras, and each of the one or more cameras of the intraoral scanner corresponds to one of the one or more reference cameras of each of the training stage handheld wands.

[0132] In a further embodiment of the sixth method, the step of driving the one or more structured light projectors includes driving one or more structured light projectors of the intraoral scanner, the step of driving the one or more unstructured light projectors includes driving one or more unstructured light projectors of the intraoral scanner, and the step of driving the one or more cameras includes driving one or more cameras of the intraoral scanner. The neural network is initially trained using training stage images captured by one or more training stage cameras of the training stage handheld wand, and each of the one or more cameras of the intraoral scanner corresponds to one of each of the one or more training stage cameras. Thereafter, the method includes, during a plurality of refinement stage scans, driving (i) the one or more structured light projectors of the intraoral scanner and (ii) the one or more unstructured light projectors of the intraoral scanner, and during the refinement stage scans, driving the one or more cameras of the intraoral scanner to capture (a) a plurality of refinement stage structured light images and (b) a plurality of refinement stage two-dimensional images, calculating the three-dimensional structure of the intraoral three-dimensional surface based on the plurality of refinement stage structured light images, and refining the training of the neural network for the intraoral scanner using (a) the plurality of refinement stage two-dimensional images captured during the refinement stage scans and (b) the calculated three-dimensional structure of the intraoral three-dimensional surface calculated based on the plurality of refinement stage structured light images.

[0133] In one embodiment, the neural network includes a plurality of layers, and the step of refining the training of the neural network includes the step of constraining a subset of the layers.

[0134] In one embodiment, the method further includes, from a plurality of scans, selecting which of the plurality of scans to use as a refinement stage scan based on the quality level of each scan.

[0135] In one embodiment, the method further includes, during the refinement stage scan, using the calculated three-dimensional structure of the intraoral three-dimensional surface calculated based on the plurality of refinement stage structured light images as the final result three-dimensional structure of the intraoral three-dimensional surface for the user of the intraoral scanner.

[0136] In a further embodiment of the sixth method, the step of driving the one or more cameras includes driving one or more cameras of an intraoral scanner, and each camera of the one or more cameras of the intraoral scanner corresponds to each one of one or more reference cameras. The method further includes, for each camera c of the one or more cameras of the intraoral scanner, using the processor to crop and morph at least one two-dimensional image of the intraoral three-dimensional surface from the camera c to obtain a plurality of cropped and morphed two-dimensional images, each cropped and morphed image corresponding to the cropped and morphed field of view of the camera c, the cropped and morphed field of view of the camera c being identical to the cropped field of view of the corresponding reference camera, the step of inputting the plurality of two-dimensional images into the neural network including inputting the plurality of cropped and morphed two-dimensional images of the intraoral three-dimensional surface into the neural network, and further, the step of determining includes determining, by the neural network, an estimated map of each of the intraoral three-dimensional surfaces captured in each of the cropped and morphed two-dimensional images, the neural network being trained using training stage images corresponding to the cropped fields of view of each of the one or more reference cameras.

[0137] In a further embodiment, the step of cropping and morphing includes the processor using (a) stored calibration values indicative of camera rays corresponding to each pixel on the camera sensor of each camera of the one or more cameras, and (b) (i) camera rays corresponding to each pixel on the reference camera sensor of each of the one or more reference cameras, and (ii) reference calibration values indicative of the cropped fields of view of each of the one or more reference cameras.

[0138] In a further embodiment, the cropped field of view of each of the one or more reference cameras is 85-97% of the full field of view of each of the one or more reference cameras.

[0139] In a further embodiment, the step of using the processor further includes, for each camera c, performing reverse morphing on each of the estimated maps of the respective intraoral three-dimensional surfaces captured in each of the cropped and morphed two-dimensional images to obtain a non-morphed estimated map of each of the intraoral surfaces as seen in at least one two-dimensional image from camera c prior to morphing.

[0140] In a further embodiment, the unstructured light is incoherent light, and the plurality of two-dimensional images includes a plurality of two-dimensional color images.

[0141] In a further embodiment, the unstructured light is near-infrared (NIR) light, and the plurality of two-dimensional images includes a plurality of monochromatic NIR images.

[0142] In a further embodiment, the step of using the processor further includes, for each camera c, performing reverse morphing on each of the estimated maps of the respective intraoral three-dimensional surfaces captured in each of the cropped and morphed two-dimensional images to obtain a non-morphed estimated map of each of the intraoral surfaces as seen in at least one two-dimensional image from camera c prior to morphing.

[0143] Note that all of the above-described embodiments of the sixth method related to depth maps, normal maps, curvature maps, and their use can be implemented with the necessary modifications based on the on-site cropped and morphed runtime images.

[0144] In one embodiment, the step of driving the one or more structured light projectors includes the step of driving one or more structured light projectors of the intraoral scanner, and the step of driving the one or more unstructured light projectors includes the step of driving one or more unstructured light projectors of the intraoral scanner. The method further includes driving one or more structured light projectors of the intraoral scanner and one or more unstructured light projectors of the intraoral scanner during a plurality of refinement stage scans after the neural network has been trained using training stage images corresponding to the cropped fields of view of each of the one or more reference cameras. One or more cameras of the intraoral scanner are driven to capture (a) a plurality of refinement stage structured light images and (b) a plurality of refinement stage two-dimensional images during structured light scans of the plurality of refinement stages. The three-dimensional structure of the intraoral three-dimensional surface is calculated based on the plurality of refinement stage structured light images, and the training of the neural network is refined for the intraoral scanner using (a) the plurality of refinement stage two-dimensional images captured during the refinement stage scans and (b) the calculated three-dimensional structure of the intraoral three-dimensional surface calculated based on the plurality of refinement stage structured light images. In a further embodiment, the neural network includes a plurality of layers, and the step of refining the training of the neural network includes the step of constraining a subset of the layers.

[0145] In one embodiment, the step of driving the one or more structured light projectors includes driving one or more structured light projectors of the intraoral scanner, the step of driving the one or more unstructured light projectors includes driving one or more unstructured light projectors of the intraoral scanner, and the step of determining includes determining, by the neural network, an estimated depth map for each of the intraoral three-dimensional surfaces captured in each of the cropped and morphed two-dimensional images. The method further includes (a) calculating a three-dimensional structure of the intraoral three-dimensional surface based on calculated respective three-dimensional positions of a plurality of points on the intraoral three-dimensional surface captured in a plurality of structured light images; (b) calculating a three-dimensional structure of the intraoral three-dimensional surface based on an estimated depth map for each of the intraoral three-dimensional surfaces captured in each of the cropped and morphed two-dimensional images; and (c) (i) the three-dimensional structure of the intraoral three-dimensional surface calculated based on the calculated respective three-dimensional positions of the plurality of points on the intraoral three-dimensional surface; and (ii) the three-dimensional structure of the intraoral three-dimensional surface calculated based on an estimated depth map for each of the intraoral three-dimensional surfaces, comparing the two. In response to determining a mismatch between (i) and (ii), the method includes, during a plurality of refinement stage scans, driving (A) the one or more structured light projectors of the intraoral scanner and (B) the one or more unstructured light projectors of the intraoral scanner; during a plurality of refinement stage scans, driving one or more cameras of the intraoral scanner to capture (a) a plurality of refinement stage structured light images and (b) a plurality of refinement stage two-dimensional images; calculating a three-dimensional structure of the intraoral three-dimensional surface based on the plurality of refinement stage structured light images; and refining the training of a neural network for the intraoral scanner using (a) the plurality of two-dimensional images captured during the refinement stage scans and (b) the calculated three-dimensional structure of the intraoral three-dimensional surface calculated based on the plurality of refinement stage structured light images.In a further embodiment, the neural network includes a plurality of layers, and the step of refining the training of the neural network includes the step of constraining a subset of the layers.

[0146] In a further embodiment of the sixth method, the sixth method further includes the step of training the neural network, and the training includes driving one or more training stage structured light projectors to project a training stage structured light pattern onto a training stage three-dimensional surface; driving one or more training stage cameras to capture a plurality of structured light images, each image including at least a portion of the training stage structured light pattern; driving one or more training stage unstructured light projectors to project unstructured light onto the training stage three-dimensional surface; driving the one or more training stage cameras to capture a plurality of two-dimensional images of the training stage three-dimensional surface using illumination from the training stage unstructured light projectors; adjusting the capture of the structured light images and the capture of the two-dimensional images to generate an alternating sequence in which the image frames of one or more two-dimensional images are interspersed among the image frames of one or more structured light images; inputting a plurality of two-dimensional images into the neural network; estimating, by the neural network, an estimated map of the training stage three-dimensional surface captured in each two-dimensional image; inputting, into the neural network, a plurality of three-dimensional reconstructions of the training stage three-dimensional surface based on the structured light images of the training stage three-dimensional surface, the three-dimensional reconstructions including the calculated three-dimensional positions of a plurality of points on the training stage three-dimensional surface; interpolating the positions of the one or more training stage cameras with respect to the training stage three-dimensional surface for each two-dimensional image frame based on the calculated three-dimensional positions of a plurality of points on the training stage three-dimensional surface calculated based on each of the structured light image frames before and after each two-dimensional image frame; projecting the three-dimensional reconstruction into the field of view of each of the one or more training stage cameras and calculating, based on the projection, a true map of the training stage three-dimensional surface seen in each two-dimensional image constrained by the calculated three-dimensional positions of the plurality of points; comparing each estimated depth map of the training stage three-dimensional surface with the corresponding true map of the training stage three-dimensional surface and, based on the difference between each estimated map and the corresponding true map,Optimizing the neural network to better estimate subsequent estimated maps, and the like.

[0147] In a further embodiment, the training includes initial training of the neural network, the step of driving the one or more structured light projectors includes driving one or more structured light projectors of the intraoral scanner, the step of driving the one or more unstructured light projectors includes driving one or more unstructured light projectors of the intraoral scanner, and the step of driving one or more cameras includes driving one or more cameras of the intraoral scanner. The method further includes, after the initial training of the neural network, during a plurality of refinement stage structured light scans, (i) driving the one or more structured light projectors of the intraoral scanner and (ii) driving the one or more unstructured light projectors of the intraoral scanner, and during the refinement stage structured light scans, driving one or more cameras of the intraoral scanner to capture (a) a plurality of refinement stage structured light images and (b) a plurality of refinement stage two-dimensional images, calculating a three-dimensional structure of the intraoral three-dimensional surface based on the plurality of refinement stage structured light images, and refining the training of the neural network using (a) the plurality of two-dimensional images captured during the refinement stage scans and (b) the three-dimensional structure of the intraoral three-dimensional surface calculated based on the plurality of refinement stage structured light images. In a further embodiment, the neural network includes a plurality of layers, and the step of refining the training of the neural network includes constraining a subset of the layers.

[0148] In a further embodiment of the sixth method, the step of driving the one or more structured light projectors to project a training stage structured light pattern includes driving the one or more structured light projectors to project a distribution of discrete unconnected light spots onto the training stage three-dimensional surface, respectively.

[0149] In a further embodiment of the sixth method, the step of driving one or more training stage cameras includes the step of driving at least two training stage cameras.

[0150] In a further embodiment of the sixth method, the unstructured light includes broadband spectral light.

[0151] In one embodiment of an intraoral scanning device, the device includes an elongated hand-held wand, the elongated hand-held wand including, at its distal end, a probe configured to be removably disposed within a sleeve. The device further includes at least one structured light projector coupled to the probe, the at least one structured light projector comprising: (a) a laser configured to emit polarized laser light; and (b) a pattern generation optical element configured to generate a pattern of light when the laser is actuated to transmit light through the pattern generation optical element. The device further includes a camera coupled to the probe, the camera including a camera sensor. The probe is configured such that light passes into and out of the probe through the sleeve. Further, the laser is positioned at a distance from the camera such that when the probe is disposed within the sleeve, a portion of the pattern of light is reflected from the sleeve and reaches the camera sensor. Further, the laser is positioned at an angle of rotation with respect to its own optical axis such that, due to the polarization of the pattern of light, the degree of reflection of the portion of the pattern of light by the sleeve is less than a threshold reflection for all possible angles of rotation of the laser with respect to its optical axis.

[0152] In a further embodiment of the intraoral scanning device, the threshold is 70% of the maximum reflection for all possible angles of rotation of the laser with respect to its optical axis.

[0153] In a further embodiment of the intraoral scanning device, the laser is positioned by an angle of rotation with respect to its own optical axis such that, due to the polarization of the light pattern, the degree of reflection of the portion of the light pattern by the sleeve is less than 60% of the maximum reflection for all possible angles of rotation of the laser with respect to its optical axis.

[0154] In a further embodiment of the intraoral scanning device, the laser is positioned by an angle of rotation with respect to its own optical axis such that, due to the polarization of the light pattern, the degree of reflection of the portion of the light pattern by the sleeve is between 15% and 60% of the maximum reflection for all possible angles of rotation of the laser with respect to its optical axis.

[0155] In a further embodiment of the intraoral scanning device, when the elongated handheld wand is disposed within the sleeve, the distance between the structured light projector and the camera is 1 to 6 times the distance between the structured light projector and the sleeve.

[0156] In a further embodiment of the intraoral scanning device, the at least one structured light projector has an illumination field of at least 30 degrees, and the camera has a field of view of at least 30 degrees.

[0157] In a seventh method of generating a three-dimensional image using an intraoral scanner, the seventh method includes using at least two cameras rigidly attached to the intraoral scanner such that the fields of view of the respective cameras have non-overlapping portions to capture a plurality of images of the intraoral three-dimensional surface. The seventh method further includes using the processor to perform a simultaneous localization and mapping (SLAM) algorithm using the captured images from the respective cameras for the non-overlapping portions of the respective fields of view, wherein the localization of each camera is solved based on the motion of each camera being the same as the motion of all other cameras.

[0158] In a further embodiment of the seventh method, the fields of view of the first camera and the second camera among the cameras also have an overlapping portion. Further, the capturing step includes capturing a plurality of images of the intraoral three-dimensional surface such that the features of the intraoral three-dimensional surface in the overlapping portion of the respective fields of view appear in the images captured by the first and second cameras. Further, the step of using the processor includes executing a SLAM algorithm using the features of the intraoral three-dimensional surface appearing in the images of at least two cameras.

[0159] In an eighth method of generating a three-dimensional image using an intraoral scanner, the eighth method includes driving one or more structured light projectors to project a pattern of structured light onto an intraoral three-dimensional surface; driving one or more cameras to capture a plurality of structured light images, each structured light image including at least a portion of the structured light pattern; driving one or more unstructured light projectors to project unstructured light onto the intraoral three-dimensional surface; driving at least one camera to capture a two-dimensional image of the intraoral three-dimensional surface using illumination from the unstructured light projector; and adjusting the capture of the structured light and the capture of the unstructured light to generate an alternating sequence in which one or more unstructured light image frames are interspersed among one or more structured light image frames. The eighth method further includes using the processor to calculate the three-dimensional position of each of a plurality of points on the intraoral three-dimensional surface captured in one or more image frames of the structured light. The eighth method further includes using the processor to interpolate the motion of at least one camera between a first unstructured light image frame and a second unstructured light image frame based on the calculated three-dimensional positions of the plurality of points in each of the structured light image frames before and after the unstructured light image frame. The eighth method further includes executing a simultaneous localization and mapping (SLAM) algorithm that (a) uses features of the intraoral three-dimensional surface captured in the first and second unstructured light image frames by the at least one camera and (b) is constrained by the interpolated motion of the camera between the first unstructured light image frame and the second unstructured light image frame.

[0160] In one embodiment of the eighth method, the step of driving the one or more structured light projectors to project the structured light pattern includes driving the one or more structured light projectors to project respective distributions of discrete, unconnected light spots on the three-dimensional surface within the oral cavity. In one embodiment, the unstructured light includes broadband spectral light and the two-dimensional image includes a two-dimensional color image. In one embodiment, the unstructured light includes near-infrared (NIR) light and the two-dimensional image includes a two-dimensional monochromatic NIR image.

[0161] In one embodiment of a ninth method of generating a three-dimensional image using an intraoral scanner, the ninth method includes driving one or more structured light projectors to project a pattern of structured light onto an intraoral three-dimensional surface; driving one or more cameras to capture a plurality of structured light images, each structured light image including at least a portion of the structured light pattern; driving one or more unstructured light projectors to project unstructured light onto the intraoral three-dimensional surface; driving one or more cameras to capture a two-dimensional image of the intraoral three-dimensional surface using illumination from the unstructured light projector; and adjusting the capture of the structured light and the capture of the unstructured light to generate an alternating sequence in which one or more unstructured light image frames are interspersed among one or more structured light image frames. The ninth method further includes using the processor to calculate the three-dimensional position of features on the intraoral three-dimensional surface based on the structured light image frames, the features also being captured in a first image frame of the unstructured light and a second image frame of the unstructured light, and further calculating the motion of one or more cameras between the first image frame of the unstructured light and the second image frame of the unstructured light based on the calculated three-dimensional position of the features, and executing a simultaneous localization and mapping (SLAM) algorithm using (i) features of the intraoral three-dimensional surface for which the three-dimensional position was not calculated based on the structured light image frames captured in the first and second unstructured light image frames by the one or more cameras, and (ii) the calculated motion of the camera between the first and second unstructured light image frames. In one embodiment, the step of driving one or more structured light projectors to project the structured light pattern includes driving the one or more structured light projectors to project respective distributions of discrete, unconnected light spots onto the intraoral three-dimensional surface. In one embodiment, the unstructured light includes broadband spectral light and the two-dimensional image includes a two-dimensional color image. In one embodiment, the unstructured light includes near-infrared (NIR) light and the two-dimensional image includes a two-dimensional monochromatic NIR image.

[0162] In a method of calculating the three-dimensional structure of the three-dimensional surface within the oral cavity of a subject, the method comprises: (a) driving one or more structured light projectors to project a pattern of structured light onto the three-dimensional surface within the oral cavity, the pattern including a plurality of features; (b) driving one or more cameras to capture a plurality of structured light images, each structured light image including at least one of the features of the structured light pattern; (c) driving one or more unstructured light projectors to project unstructured light onto the three-dimensional surface within the oral cavity; (d) driving at least one camera to capture a two-dimensional image of the three-dimensional surface within the oral cavity using illumination from the one or more unstructured light projectors; and (e) adjusting the capture of structured light and the capture of unstructured light to produce an alternating sequence in which one or more unstructured light image frames are interspersed among one or more structured light image frames. The method further comprises using a processor to: (a) for one or more of the plurality of features of the structured light pattern, determine, based on the two-dimensional image, whether the feature was projected onto a moving tissue within the oral cavity or onto a stable tissue; (b) based on the determination, assign a respective confidence grade to each of the one or more features, with a high confidence grade assigned to fixed tissue and a low confidence grade assigned to moving tissue; and (c) based on the confidence grade for each of the one or more features, execute a three-dimensional reconstruction algorithm using the one or more features. In one embodiment, the unstructured light includes broadband spectral light and the two-dimensional image is a two-dimensional color image. In one embodiment, the unstructured light includes near-infrared (NIR) light and the two-dimensional image is a two-dimensional monochromatic NIR image. In one embodiment, the plurality of features includes a plurality of spots, and the step of driving the one or more structured light projectors to project the structured light pattern includes driving the one or more structured light projectors to project, onto the three-dimensional surface within the oral cavity, respective distributions of discrete, unconnected light spots. In one embodiment, the step of executing the three-dimensional reconstruction algorithm is performed using only a subset of the plurality of features, the subset consisting of features assigned a confidence grade above a fixed tissue threshold.In one embodiment, the step of executing the three-dimensional reconstruction algorithm includes: (a) for each feature, assigning a weight to the feature based on each reliability grade assigned to the feature; and (b) using the weight of each feature in the three-dimensional reconstruction algorithm.

[0163] In one embodiment of a tenth method for calculating the three-dimensional structure of an intraoral three-dimensional surface, the tenth method includes driving one or more light sources of an intraoral scanner to project light onto the intraoral three-dimensional surface, and driving two or more cameras of the intraoral scanner to respectively capture a plurality of two-dimensional images of the intraoral three-dimensional surface, wherein each of the two or more cameras of the intraoral scanner corresponds to one of the two or more reference cameras. The method includes using a processor to, for each camera c of the two or more cameras of the intraoral scanner, modifying at least one of the two-dimensional images from camera c to obtain a plurality of modified two-dimensional images, each modified image corresponding to a modified field of view of camera c, the modified field of view of camera c being identical to the modified field of view of the corresponding camera among the reference cameras, and further includes calculating the three-dimensional structure of the intraoral three-dimensional surface based on the plurality of modified two-dimensional images of the intraoral three-dimensional surface.

[0164] In an embodiment of an eleventh method for calculating the three-dimensional structure of an intraoral three-dimensional surface, the eleventh method includes driving one or more light sources of an intraoral scanner to project light onto the intraoral three-dimensional surface, and driving two or more cameras of the intraoral scanner to respectively capture a plurality of two-dimensional images of the intraoral three-dimensional surface, and each of the one or more cameras of the intraoral scanner corresponds to one of the two or more reference cameras. The method includes using a processor to crop and morph at least one of the two-dimensional images from each camera c of the two or more cameras of the intraoral scanner to obtain a plurality of cropped and morphed two-dimensional images, each cropped and morphed image corresponding to the cropped and morphed field of view of camera c, and the cropped and morphed field of view of camera c matching the cropped field of view of the corresponding camera among the reference cameras. The three-dimensional structure of the intraoral three-dimensional surface is calculated based on the plurality of cropped and morphed two-dimensional images of the intraoral three-dimensional surface by inputting the plurality of cropped and morphed two-dimensional images of the intraoral three-dimensional surface into a neural network, and the neural network determining an estimated map of each of the intraoral three-dimensional surfaces captured in each of the plurality of cropped and morphed two-dimensional images, and the neural network being trained using training stage images corresponding to the cropped fields of view of each of the one or more reference cameras.

[0165] In a further embodiment of the eleventh method, the light is non-coherent light, and the plurality of two-dimensional images include a plurality of two-dimensional color images.

[0166] In a further embodiment of the eleventh method, the light is near-infrared (NIR) light, and the plurality of two-dimensional images include a plurality of monochromatic NIR images.

[0167] In a further embodiment of the eleventh method, the light is broadband spectral light, and the plurality of two-dimensional images include a plurality of two-dimensional color images.

[0168] In a further embodiment of the eleventh method, the cropping and morphing step comprises the processor using (a) stored calibration values indicative of camera rays corresponding to each pixel on the camera sensor of each camera c of the one or more cameras c, and (b) (i) camera rays corresponding to each pixel on the reference camera sensor of each reference camera of the one or more reference cameras, and (ii) reference calibration values indicative of the cropped fields of view of each of the one or more reference cameras.

[0169] In a further embodiment of the eleventh method, the cropped field of view of each of the one or more reference cameras is 85% to 97% of the full field of view of each of the one or more reference cameras.

[0170] In a further embodiment, the step of using the processor further comprises, for each camera c, performing reverse morphing on each of the estimated maps of the respective three-dimensional intraoral surfaces captured in each of the cropped and morphed two-dimensional images to obtain a non-morphed estimated map of each of the intraoral surfaces as seen in at least one two-dimensional image from camera c prior to morphing.

[0171] It should be noted that all of the above-described embodiments of the sixth method related to depth maps, normal maps, curvature maps, and their use may be performed with the necessary modifications from the eleventh method, based on runtime images that are cropped and morphed on-site.

[0172] It should also be noted that all of the above-described embodiments of the sixth method related to structured light may be performed with the necessary modifications in the context of the eleventh method and the cropped and morphed runtime two-dimensional images.

[0173] In one embodiment of the twelfth method for calculating the three-dimensional structure of the three-dimensional surface in the oral cavity, the twelfth method includes driving one or more light projectors to project light onto the three-dimensional surface in the oral cavity and driving one or more cameras to capture a plurality of two-dimensional images of the three-dimensional surface in the oral cavity. The method includes using a processor to input the plurality of two-dimensional images of the three-dimensional surface in the oral cavity into a first neural network module and a second neural network module, determining, by the first neural network module, respective estimated depth maps of the three-dimensional surface in the oral cavity captured in each of the two-dimensional images, and determining, by the second neural network module, respective estimated confidence maps corresponding to the respective estimated depth maps, wherein each confidence map indicates the confidence for each region of the respective estimated depth map.

[0174] In one embodiment of the twelfth method, the first neural network module and the second neural network module are separate modules of the same neural network.

[0175] In one embodiment of the twelfth method, each of the first and second neural network modules is not a separate module of the same neural network.

[0176] In a further embodiment of the twelfth method, the method further includes training a second neural network module to determine respective estimated confidence maps corresponding to the respective estimated depth maps determined by the first neural network module, and the determining step includes initially training the first neural network module to determine respective estimated depth maps using a plurality of depth training stage two-dimensional images, and then, (i) inputting a plurality of confidence training stage two-dimensional images of the training stage three-dimensional surface into the first neural network module, (ii) determining, by the first neural network module, respective estimated depth maps of the respective training stage three-dimensional surfaces captured in each of the confidence training stage two-dimensional images, (iii) calculating the difference between each of the estimated depth maps and the corresponding respective true depth maps to obtain respective target confidence maps corresponding to the respective estimated depth maps determined by the first neural network module, (iv) inputting the plurality of confidence training stage two-dimensional images into the second neural network module, (v) estimating, by the second neural network module, respective estimated confidence maps indicating the confidence of each region of the respective estimated depth maps, and (vi) comparing each of the estimated confidence maps with the corresponding target confidence map and optimizing the second neural network module based on the comparison to better estimate subsequent estimated confidence maps, which is performed by the steps.

[0177] In one embodiment, the plurality of confidence training stage two-dimensional images are not the same as the plurality of depth training stage two-dimensional images.

[0178] In one embodiment, the plurality of confidence training stage two-dimensional images are the same as the plurality of depth training stage two-dimensional images.

[0179] In a further embodiment of the twelfth method, (a) the step of driving the one or more cameras to capture a plurality of two-dimensional images includes, in a given image frame, driving each of the two or more cameras to simultaneously capture a respective two-dimensional image of each respective part of the three-dimensional surface within the oral cavity; (b) the step of inputting the plurality of two-dimensional images of the three-dimensional surface within the oral cavity into a first neural network module and a second neural network module includes, for a given image frame, inputting each image of the respective two-dimensional images as a separate input into the first neural network module and the second neural network module; (c) the step of determining by the first neural network module includes, for a given image frame, determining an estimated depth map of each respective part of the three-dimensional surface within the oral cavity captured in each of the respective two-dimensional images captured in the given image frame; (d) the step of determining by the second neural network module includes, for a given image frame, determining an estimated confidence map corresponding to each estimated depth map of each respective part of the three-dimensional surface within the oral cavity captured in each of the respective two-dimensional images captured in the given image frame. The method further includes using the processor to merge the respective estimated depth maps to obtain a composite estimated depth map of the three-dimensional surface within the oral cavity captured in the given image frame. In response to determining a conflict between corresponding respective regions in at least two of the estimated depth maps, the processor merges the at least two estimated depth maps based on the confidence of the corresponding respective regions indicated by the respective confidence maps for each of the at least two estimated depth maps.

[0180] In one embodiment of a thirteenth method for calculating the three-dimensional structure of an intraoral three-dimensional surface, the thirteenth method includes driving one or more light sources of an intraoral scanner to project light onto the intraoral three-dimensional surface, and driving one or more cameras of the intraoral scanner to capture a plurality of two-dimensional images of the intraoral three-dimensional surface. The method further includes (a) using a processor to determine, by a neural network, an estimated map of each of the intraoral three-dimensional surfaces captured in each of the two-dimensional images, and (b) using a processor to overcome manufacturing deviations of one or more cameras of the intraoral scanner and reduce the difference between the estimated map and the true structure of the intraoral three-dimensional surface.

[0181] In one embodiment of the thirteenth method, the step of overcoming manufacturing deviations of one or more cameras includes overcoming manufacturing deviations of one or more cameras from a reference set of one or more cameras.

[0182] In one embodiment of the thirteenth method, the intraoral scanner is one of a plurality of manufactured intraoral scanners, each manufactured intraoral scanner includes a set of one or more cameras, and the step of overcoming manufacturing deviations of one or more cameras of the intraoral scanner includes overcoming manufacturing deviations of one or more cameras from a set of one or more cameras of at least one other of the plurality of manufactured intraoral scanners.

[0183] In a further embodiment of the 13th method, the step of driving one or more cameras includes driving two or more cameras of the intraoral scanner to respectively capture a plurality of two-dimensional images of the intraoral three-dimensional surface, and two or more cameras of the intraoral scanner respectively correspond to each one of two or more reference cameras, and the neural network is trained using training stage images captured by the two or more reference cameras. The step of overcoming manufacturing deviations includes the step of overcoming manufacturing deviations of two or more cameras of the intraoral scanner, and doing so using the processor, (a) for each camera c of the two or more cameras of the intraoral scanner, modifying at least one two-dimensional image from camera c to obtain a plurality of modified two-dimensional images, each modified image corresponding to a modified field of view of camera c, and the modified field of view of camera c matching the modified field of view of the corresponding one of the reference cameras, and (b) determining, by the neural network, an estimated map of each of the intraoral three-dimensional surfaces based on the plurality of modified two-dimensional images of the intraoral three-dimensional surface.

[0184] In a further embodiment, the light is non-coherent light, and the plurality of two-dimensional images include a plurality of two-dimensional color images.

[0185] In a further embodiment, the light is near-infrared (NIR) light, and the plurality of two-dimensional images include a plurality of monochromatic NIR images.

[0186] In a further embodiment, the light is broadband spectral light, and the plurality of two-dimensional images include a plurality of two-dimensional color images.

[0187] In a further embodiment, the modifying step includes cropping and morphing at least one of the two-dimensional images from camera c to obtain a plurality of cropped and morphed two-dimensional images, each cropped and morphed image corresponding to a cropped and morphed field of view of camera c, and the cropped and morphed field of view of camera c being identical to the cropped field of view of a corresponding one of the reference cameras of the reference camera.

[0188] In a further embodiment, the cropping and morphing step includes the processor using (a) stored calibration values indicative of camera rays corresponding to each pixel on the camera sensor of each camera c of the one or more cameras c, and (b) (i) camera rays corresponding to each pixel on the reference camera sensor of each reference camera of the one or more reference cameras, and (ii) reference calibration values indicative of the cropped fields of view of each of the one or more reference cameras.

[0189] In a further embodiment, the cropped field of view of each of the one or more reference cameras is 85-97% of the respective full field of view of each of the one or more reference cameras.

[0190] In a further embodiment, the processor further performs reverse morphing on each of the estimated maps of the intraoral three-dimensional surface captured in each of the cropped and morphed two-dimensional images for each camera c to obtain a non-morphed estimated map of each of the intraoral surfaces seen in each of at least one of the two-dimensional images from camera c prior to morphing.

[0191] In a further embodiment of the 13th method, the step of overcoming the manufacturing deviation of the one or more cameras of the intraoral scanner includes training the neural network using training stage images captured by a plurality of training stage intraoral scanners. Each training stage intraoral scanner includes one or more reference cameras, each of the one or more cameras of the intraoral scanner corresponding to each one of the one or more reference cameras on each training stage intraoral scanner, and the manufacturing deviation of the one or more cameras being the manufacturing deviation of the one or more cameras from the corresponding one or more reference cameras.

[0192] In a further embodiment of the 13th method, the step of driving the one or more cameras includes driving two or more cameras of the intraoral scanner to respectively capture a plurality of two-dimensional images of the intraoral three-dimensional surface. The step of overcoming the manufacturing deviation includes overcoming the manufacturing deviation of two or more cameras of the intraoral scanner, which is done by training the neural network using training stage images each captured by only one camera, driving two or more cameras of the intraoral scanner to simultaneously image each respective part of the intraoral three-dimensional surface in a given image frame, inputting each of the respective two-dimensional images as separate inputs into the neural network for the given image frame, determining, by the neural network, an estimated depth map for each respective part of the intraoral three-dimensional surface captured in each respective two-dimensional image captured in the given image frame, and using the processor to merge the respective estimated depth maps to obtain a composite estimated depth map of the intraoral three-dimensional surface captured in the given image frame.

[0193] In a further embodiment, the determining step further includes determining, by the neural network, an estimated confidence map corresponding to each estimated depth map, each confidence map indicating the confidence for each region of each estimated depth map.

[0194] In a further embodiment, the step of merging each estimated depth map comprises, in response to determining, using the processor, a conflict between corresponding respective regions in at least two of the estimated depth maps, merging at least two estimated depth maps based on the reliability of the corresponding respective regions indicated by respective reliability maps for each of the at least two estimated depth maps.

[0195] In a further embodiment of the thirteenth method, the step of overcoming manufacturing deviations of one or more cameras of the intraoral scanner comprises: (a) initially training the neural network using training stage images captured by one or more training stage cameras of one or more training stage hand-held wands, each of the one or more cameras of the intraoral scanner corresponding to a respective one of the one or more training stage cameras of each of the one or more training stage hand-held wands; and (b) subsequently driving the intraoral scanner to perform a plurality of refinement stage scans of the intraoral three-dimensional surface and refining the training of the neural network for the intraoral scanner using the refinement stage scans of the intraoral three-dimensional surface.

[0196] In a further embodiment, the neural network includes a plurality of layers, and the step of refining the training of the neural network comprises constraining a subset of the layers.

[0197] In a further embodiment, the method further comprises selecting, from the plurality of scans, which of the plurality of scans to use as refinement stage scans based on the quality level of each scan.

[0198] In a further embodiment, the step of driving the intraoral scanner to perform a plurality of refinement stage scans includes, during the plurality of refinement stage scans, (i) driving one or more structured light projectors of the intraoral scanner to project a pattern of structured light onto the intraoral three-dimensional surface; (ii) driving one or more unstructured light projectors of the intraoral scanner to project unstructured light onto the intraoral three-dimensional surface; and driving one or more cameras of the intraoral scanner to capture, during the refinement stage scan, (a) a plurality of refinement stage structured light images using the illumination from the structured light projector and (b) a plurality of refinement stage two-dimensional images using the illumination from the unstructured light projector; and calculating a three-dimensional structure of the intraoral three-dimensional surface based on the plurality of refinement stage structured light images.

[0199] In a further embodiment, the step of refining the training of the neural network includes refining the training of the neural net for the intraoral scanner using (a) a plurality of refinement stage two-dimensional images captured during the refinement stage scan and (b) the calculated three-dimensional structure of the intraoral three-dimensional surface calculated based on the plurality of refinement stage structured light images.

[0200] In a further embodiment, the method further includes using, during the refinement stage scan, the calculated three-dimensional structure of the intraoral three-dimensional surface calculated based on the plurality of refinement stage structured light images as the final result three-dimensional structure of the intraoral three-dimensional surface for the user of the intraoral scanner.

[0201] In one embodiment of a fourteenth method of training a neural network for use with an intraoral scanner, the fourteenth method includes inputting a plurality of two-dimensional images of an intraoral three-dimensional surface into the neural network; estimating, by the neural network, an estimated map of the intraoral three-dimensional surface captured in each of the two-dimensional images; calculating, based on a plurality of structured light images of the intraoral three-dimensional surface, a true map of the intraoral three-dimensional surface seen in each of the two-dimensional images; comparing each estimated map of the intraoral three-dimensional surface with the corresponding true map of the intraoral three-dimensional surface; optimizing the neural network based on a difference between each estimated map and the corresponding true map to better estimate subsequent estimated maps; and, for two-dimensional images in which a moving tissue is identified, processing the images to exclude at least a portion of the moving tissue prior to inputting the two-dimensional images into the neural network.

[0202] In a further embodiment of the 14th method, the method includes driving one or more structured light projectors to project a structured light pattern onto the intraoral three-dimensional surface; driving one or more cameras to capture a plurality of structured light images, each image including at least a portion of the structured light pattern; driving one or more unstructured light projectors to project unstructured light onto the intraoral three-dimensional surface; driving the one or more cameras to capture a plurality of two-dimensional images of the intraoral three-dimensional surface using the illumination from the unstructured light projector; adjusting the capture of the structured light images and the capture of the two-dimensional images to generate an alternating sequence in which the image frames of the one or more two-dimensional images are interspersed among the image frames of the one or more structured light images. Further, the step of calculating the true map of the intraoral three-dimensional surface seen in each of the two-dimensional images includes inputting a plurality of three-dimensional reconstructions of the intraoral three-dimensional surface into the neural network based on the structured light images of the intraoral three-dimensional surface, the three-dimensional reconstructions including the calculated three-dimensional positions of a plurality of points on the intraoral three-dimensional surface; interpolating the positions of the one or more cameras relative to the intraoral three-dimensional surface for each two-dimensional image frame based on the calculated three-dimensional positions of a plurality of points on the intraoral three-dimensional surface calculated based on the structured light image frames before and after each two-dimensional image frame; projecting the three-dimensional reconstructions into the respective fields of view of the one or more cameras and calculating the true map of the intraoral three-dimensional surface seen in each of the plurality of two-dimensional images constrained by the calculated three-dimensional positions of the plurality of points based on the projection.

[0203] Furthermore, according to some applications of the present invention, a method for generating a digital three-dimensional image is provided, the method comprising: driving each structured light projector of one or more structured light projectors to project a pattern of light (e.g., a distribution of discrete non-connected light spots) onto the intraoral three-dimensional surface; Drive each of the one or more cameras to capture a plurality of images, each image including at least a portion of the projection pattern, and each of the one or more cameras including a camera sensor including an array of pixels, a step; Using a processor, comparing a plurality of consecutive images captured by each camera and determining a portion of the captured projection pattern (e.g., projection spot s) that can be tracked across the plurality of images.

[0204] In some applications, the projection pattern is a distribution of unconnected light spots, and the processor may make a determination based on stored calibration values indicating (a) camera rays corresponding to each pixel on the camera sensor of each of the one or more cameras, and (b) projector rays corresponding to each projection light spot of each of the one or more projectors. In some embodiments, each projector ray corresponds to the path of each pixel on at least one of the camera sensors. Further, in some embodiments, the processor may determine which projection spots s can be tracked across the plurality of images, and each tracked spot s moves along the path of the pixel corresponding to each projector ray r.

[0205] In some applications, the step of using the processor further includes using the processor to calculate, at the intersection of the projector ray r corresponding to the tracked spot s and each camera ray in each of the plurality of consecutive images in which the spot s is tracked, each three-dimensional position on the three-dimensional surface within the oral cavity.

[0206] In some applications, the step of using the processor further includes using the processor to (a) determining parameters of the tracked spots in at least two adjacent images from consecutive images, the parameters including one or more of the size of the spot, the shape of the spot, the orientation of the spot, the intensity of the spot, and the signal-to-noise ratio (SNR) of the spot; (b) including the step of predicting the parameters of the tracking spot in a subsequent image based on the parameters of the tracking spot in the at least two adjacent images.

[0207] In some application examples, the step of using the processor further includes the step of using the processor to search for a spot having substantially the predicted parameters in a subsequent image based on the predicted parameters of the tracking spot.

[0208] In some application examples, the selected parameter is the shape of the spot, and the step of using the processor further includes the step of using the processor to determine a search space in a next image for searching for the tracking spot based on the predicted shape of the tracking spot.

[0209] In some application examples, the step of determining the search space using the processor includes the step of using the processor to determine a search space in a next image for searching for the tracking spot, and the search space has a size and aspect ratio based on the size and / or aspect ratio of the predicted shape of the tracking spot.

[0210] In some application examples, the selected parameter is the shape of the spot, and the step of using the processor further includes using the processor to (a) determining a velocity vector of the tracking spot based on a direction and distance by which the tracking spot has moved between two adjacent images from consecutive images; (b) predicting a shape of the tracking spot in a subsequent image in response to the shape of the tracking spot in at least one of the two adjacent images; (c) (i) determining a velocity vector of the tracking spot and (ii) determining a search space in a subsequent image for searching for the tracking spot in response to a combination of the determined velocity vector of the tracking spot and the predicted shape of the tracking spot.

[0211] In some application examples, the selected parameter is the shape of the spot, and the step of using the processor is to use the processor to (a) determining a velocity vector of the tracking spot based on a direction and a distance by which the tracking spot has moved between two adjacent images from the series of images; (b) predicting a shape of the tracking spot in a subsequent image in response to determining the velocity vector of the tracking spot; and (c) determining a search space within a subsequent image for searching for the tracking spot in response to a combination of (i) determining the velocity vector of the tracking spot and (ii) the predicted shape of the tracking spot.

[0212] In some application examples, the step of using the processor includes using the processor to predict a shape of the tracking spot in a subsequent image in response to a combination of (i) determining a velocity vector of the tracking spot and (ii) the shape of the tracking spot in at least one of two adjacent images.

[0213] In some application examples, the step of using the processor is to use the processor to (a) determining a velocity vector of the tracking spot based on a direction and a distance by which the tracking spot has moved between two consecutive images; and (b) determining a search space within a subsequent image for searching for the tracking spot in response to determining the velocity vector of the tracking spot.

[0214] In some application examples, the step of using the processor further includes, for each tracking spot s, determining a plurality of possible paths p of pixels on a given one of the cameras of the camera, where the paths p each correspond to a plurality of possible projector rays r.

[0215] In some application examples, the step of using the processor further includes using the processor to execute a corresponding algorithm to (a) for each possible projector ray r, determine, for each camera ray corresponding to each camera ray that intersects the projector ray r and the camera ray of a given one of the cameras corresponding to the tracking spot s, how many other cameras detected it on the path p1 of each pixel of its own corresponding to the projector ray r, (b) identify a given projector ray r1 where the most other cameras detected the respective spots q, (c) identify the projector ray r1 as the specific projector ray r that generated the tracking spot s.

[0216] In some application examples, the step of using the processor includes using the processor to (a) execute a corresponding algorithm to calculate the three-dimensional position of each of a plurality of detection spots on the three-dimensional surface inside the oral cavity captured in a plurality of consecutive images, and (b) in at least one of the plurality of consecutive images, identify the detection spot as a tracking spot s that moves along the path of the pixel corresponding to the specific projector ray r, thereby identifying that the detection spot is derived from the specific projector ray r.

[0217] In some application examples, the step of using the processor further includes using the processor to (a) execute a corresponding algorithm to calculate the three-dimensional position of each of a plurality of detection spots on the three-dimensional surface inside the oral cavity captured in a plurality of consecutive images, and (b) exclude from consideration as points on the inner surface of the oral cavity spots that are (i) identified as being derived from a specific projector ray r based on the three-dimensional positions calculated by the corresponding algorithm and (ii) not identified as a tracking spot s that moves along the path of the pixel corresponding to the specific projector ray r.

[0218] In some application examples, the step of using the processor further includes using the processor to (a) execute a corresponding algorithm to calculate the three-dimensional position of each of a plurality of detection spots on the intraoral three-dimensional surface captured in a plurality of consecutive images; and (b) identifying the detection spots as tracking spots s that move along one of the two different projector rays r for the detection spots identified as being derived from two different projector rays r based on the three-dimensional positions calculated by the corresponding algorithm, thereby identifying that the detection spots are derived from the two different projector rays r.

[0219] In some application examples, the step of using the processor further includes using the processor to (a) execute a corresponding algorithm to calculate the three-dimensional position of each of a plurality of detection spots on the intraoral three-dimensional surface captured in a plurality of consecutive images, and (b) identifying the vulnerable spots, for which the three-dimensional positions have not been calculated by the corresponding algorithm, as tracking spots s that move along the path of the pixels corresponding to a specific projector ray r by identifying them as projection spots from a specific projector ray r.

[0220] According to some application examples of the present invention, a method for generating a digital three-dimensional image is further provided. The method includes driving each structured light projector of one or more structured light projectors to project a pattern of light (e.g., a distribution of discrete non-connected light spots) onto the intraoral three-dimensional surface; and driving each camera of one or more cameras to capture an image, the image including at least a portion of the projected pattern, and each camera of the one or more cameras including a camera sensor including an array of pixels; and using a processor to (a) Execute a corresponding algorithm to calculate the 3D position of each part of the detected pattern on the 3D surface within the oral cavity captured in a plurality of consecutive images; (b) In at least a subset of a plurality of consecutive images, identify the calculated 3D positions of the parts of the detected pattern corresponding to a specific projector ray r; (c) Calculate the length of the projector ray r in each image of the subset of images based on the 3D positions of the detected parts of the pattern corresponding to the projector ray r in the subset of images.

[0221] In some embodiments, the pattern of light may be a distribution of unconnected spots. In some embodiments, the processor may execute steps (a) - (c) based on stored calibration values indicating (i) camera rays corresponding to each pixel on the camera sensor of each of the one or more cameras, and (ii) projector rays corresponding to each projected light spot from each of the one or more projectors. In some embodiments, each projector ray corresponds to the path of each pixel on at least one of the camera sensors.

[0222] In some applications, the step of using the processor further includes calculating an estimated length of the projector ray r using the processor in at least one of a plurality of consecutive images in which the 3D position of the projection spot from the projector ray r was not identified in step (b).

[0223] In some applications, the step of using the processor further includes determining a one - dimensional search space in at least one of a plurality of images for searching for the projection spot from the projector ray r based on the estimated length of the projector ray r in at least one of the plurality of images using the processor, and the one - dimensional search space follows the path of each pixel corresponding to the projector ray r.

[0224] In some application examples, the step of using the processor further includes using the processor to search for a projection spot from the projector light ray r based on the estimated length of the projector light ray r in at least one of a plurality of images, and determining a one-dimensional search space in each pixel array of a plurality of cameras, where the one-dimensional search space is along the path of each pixel corresponding to the light ray r.

[0225] In some application examples, the step of using the processor to determine a one-dimensional search space in each pixel array of a plurality of cameras includes using the processor to determine a one-dimensional search space for searching for a projection spot from the projector light ray r in each pixel array of all the cameras.

[0226]

[0225] In some application examples, the step of using the processor further includes calculating the estimated length of the projector light ray r in at least one of a plurality of consecutive images in which a plurality of candidate three-dimensional positions of the projection spot from the projector light ray r are identified in step (b).

[0227] In some application examples, the step of using the processor further includes determining whether the correct three-dimensional position of the projection spot is determined by determining which of the plurality of candidate three-dimensional positions corresponds to the estimated length of the projector light ray r in at least one of a plurality of images.

[0228] In some application examples, the step of using the processor includes using the processor to determine, based on the estimated length of the projector light ray r in at least one of a plurality of images, (a) determining a one-dimensional search space in at least one of a plurality of images for searching for a projection spot from the projector light ray r; (b) Determining which of the plurality of candidate three-dimensional positions corresponds to the spot generated by the projector light ray r detected within the one-dimensional search space, thereby determining which of the plurality of candidate three-dimensional positions of the projection spot is the correct three-dimensional position of the projection spot generated by the projector light ray r, including steps.

[0229] In some application examples, the step of using the processor further includes using the processor to (i) defining a curve based on the length of the projector light ray r in each image of the subset of the images; and (ii) when the three-dimensional position of the projection spot corresponds to a length of the projector light ray r that is at least a threshold distance away from the defined curve, excluding the detection spot identified as being derived from the projector light ray r in step (b) from consideration as a point on the oral cavity inner surface, including steps.

[0230] According to some application examples of the present invention, a method for generating a digital three-dimensional image is further provided, the method including driving each structured light projector of one or more structured light projectors to project a distribution of discrete, non-connected light spots onto the three-dimensional surface within the oral cavity; driving each camera of one or more cameras to capture an image, the image including at least one of the spots, and each camera of the one or more cameras including a camera sensor including an array of pixels; based on stored calibration values indicating (a) camera rays corresponding to each pixel on the camera sensor of each camera of the one or more cameras, and (b) projector rays corresponding to each one projection light spot from each projector of the one or more projectors, each projector ray corresponding to a respective path on at least one pixel of the camera sensor, using a processor to (a) Executing a corresponding algorithm to calculate the respective three-dimensional positions of the plurality of projection spots on the intraoral three-dimensional surface; (b) Using data from at least two of the cameras to identify candidate three-dimensional positions of a given spot corresponding to a specific projector ray r, and substantially not using data from another one of the cameras to identify the candidate three-dimensional positions; (c) Using candidate three-dimensional positions seen by at least one of the two cameras to identify a search space on another one of the pixel arrays of the camera for searching for spots from the projector ray r; (d) When a spot from the projector ray r is identified within the search space, using data from the other one of the cameras to refine a candidate for the three-dimensional position of the spot.

[0231] According to some application examples of the present invention, a method for generating a digital three-dimensional image is further provided, the method comprising: Driving each structured light projector of one or more structured light projectors to project a distribution of discrete, non-connected light spots on the intraoral three-dimensional surface; Driving each camera of one or more cameras to capture a plurality of images, each image including at least one of the spots, and each camera of the one or more cameras including a camera sensor including a pixel array; Based on stored calibration values indicating (a) camera rays corresponding to each pixel on the camera sensor of each of the one or more cameras, and (b) projector rays corresponding to each one projection light spot from each of the one or more projectors, wherein each projector ray corresponds to the path of each pixel on at least one of the camera sensors of the camera sensor; Using a processor to: (a) Executing a corresponding algorithm to calculate, for each of the plurality of images, the respective three-dimensional positions of the plurality of detected spots on the intraoral three-dimensional surface; (b) Using data corresponding to the three-dimensional positions of at least three spots, each spot corresponding to a respective projector ray r, to estimate a three-dimensional surface on which all of the at least three spots are located; (c) For a projector ray r1 for which the three-dimensional position of the spot corresponding to that projector ray r1 was not calculated in step (a), estimating a three-dimensional position on the intersection space between the projector ray r1 and the estimated surface; (d) Using the three-dimensional positions within the estimation space to identify a search space within the pixel array of at least one camera for searching for the spot corresponding to the projector ray r1.

[0232] In some applications, the step of using data corresponding to the three-dimensional positions of each of the at least three spots includes using data corresponding to the three-dimensional positions of at least three spots all captured in one of the plurality of images.

[0233] In some applications, the method further includes refining the estimation of the three-dimensional surface using data corresponding to the three-dimensional positions of at least one additional spot, the at least one additional spot having a three-dimensional position calculated based on another one of the plurality of images such that all of the at least three spots and the at least one additional spot are on the three-dimensional surface.

[0234] In some applications, the step of using data corresponding to the three-dimensional positions of each of the at least three spots includes using data corresponding to at least three spots, each spot being captured in each one of the plurality of images.

[0235] In accordance with some applications of the present invention, a method for tracking the motion of an intraoral scanner is further provided, the method comprising: (A) Using at least one camera coupled to the intraoral scanner to measure the motion of the intraoral scanner with respect to the intraoral surface being scanned; (B) Measuring the motion of the intraoral scanner relative to a fixed coordinate system using at least one inertial measurement unit (IMU) coupled to the intraoral scanner; (C) Using a processor, (i) (a) Calculating the motion of the intraoral surface relative to the fixed coordinate system by subtracting (b) the motion of the intraoral scanner relative to the intraoral surface from (b) the motion of the intraoral scanner relative to the fixed coordinate system; (ii) Constructing a prediction model of the motion of the intraoral surface relative to the fixed coordinate system based on the accumulated data of the motion of the intraoral surface relative to the coordinate system; (iii) Calculating an estimated position of the intraoral scanner relative to the intraoral surface by subtracting (a) the prediction of the motion of the intraoral surface relative to the coordinate system derived based on the prediction model from (b) the motion of the intraoral scanner relative to the coordinate system measured by the IMU.

[0236] In some applications, the method further includes determining whether measurement of the motion of the intraoral scanner relative to the intraoral surface is inhibited using at least one camera, and calculating an estimated position of the intraoral scanner relative to the intraoral surface in response to determining that measurement of the motion is inhibited.

[0237] According to some applications of the present invention, a method is provided, the method comprising: Driving each structured light projector of one or more structured light projectors to project a distribution of discrete, unconnected light spots on an intraoral three-dimensional surface; Driving each camera of one or more cameras to capture a plurality of images, each image including at least one of the spots, and each camera of the one or more cameras including a camera sensor including a pixel array; (a) The camera rays corresponding to each pixel on the camera sensor of each of the one or more cameras, and (b) the projector rays corresponding to each one projection light spot from each of the one or more projectors, wherein each projector ray corresponds to a respective path p on at least one pixel of the camera sensor, based on the stored calibration values indicating the same. Using a processor, (a) Executing a corresponding algorithm to calculate the respective three-dimensional positions of the plurality of projection spots on the three-dimensional intraoral surface; (b) Collecting data at a plurality of time points, the data including the calculated respective three-dimensional positions of the detected plurality of spots on the intraoral surface; (c) For each projector ray r, based on the collected data, defining an updated path p' of pixels for each camera sensor such that all of the calculated three-dimensional positions corresponding to the spots generated by the projector ray r correspond to positions along the respective updated paths p' of the pixels for each camera sensor; (d) Comparing each updated path p' of the pixels with the path p of the pixels corresponding to the projector ray r of that camera sensor from the stored calibration values; (e) For at least one camera sensor s, when the updated path p' of the pixels corresponding to the projector ray r is different from the path p of the pixels corresponding to the projector ray r from the stored calibration values, Reducing the difference between the updated path p' of the pixels corresponding to each projector ray r and the respective path p of the pixels corresponding to each projector ray r from the stored calibration values, and doing so (i) By changing the stored calibration data selected from the group consisting of the stored calibration values indicating the camera rays corresponding to each pixel on the camera sensor s of each of the one or more cameras, and (ii) The stored calibration values indicating the projector rays r corresponding to each one projection light spot from each of the one or more projectors. This includes the step of performing by changing the stored calibration data so selected.

[0238] In some applications, the selected stored calibration data includes stored calibration values indicating camera rays corresponding to each pixel on the camera sensor of each of one or more cameras, The step of varying the stored calibration data varies one or more parameters of a parameterized camera calibration function that defines a camera ray corresponding to each pixel on at least one camera sensor s, (i) the respective calculated three-dimensional positions on the oral surface of the plurality of detection spots, and (ii) the stored calibration values indicating the respective camera rays corresponding to each pixel on the camera sensor on which each of the plurality of detection spots was to be detected, including the step of reducing the difference between.

[0239] In some applications, the selected stored calibration data includes stored calibration values indicating projector rays corresponding to each one projection light spot from each of one or more projectors, and the step of varying the stored calibration data is (i) an indexed list that assigns each projector ray r to a pixel path p, or (ii) one or more parameters of a parameterized projector calibration model that defines each projector ray r, including the step of varying.

[0240] In some applications, the step of varying the stored calibration data includes varying the indexed list by reassigning each projector ray r based on the respective updated path p' of the pixel corresponding to each projector ray r.

[0241] In some applications, the step of varying the stored calibration data is (i) the stored calibration values indicating the camera rays corresponding to each pixel on the camera sensor s of each of the one or more cameras, (ii) the stored calibration value indicating the projector ray r corresponding to each one projection light spot from each projector of one or more projectors, and includes steps of changing

[0242] In some application examples, the step of changing the stored calibration value includes the step of repeatedly changing the stored calibration value.

[0243] In some application examples, the method further includes steps of driving each camera of one or more cameras to capture a plurality of images of a calibration object having predetermined parameters, and using a processor, executing a triangulation algorithm to calculate respective parameters of the calibration object based on the captured images, and executing an optimization algorithm to (a) (i) reduce the difference between the updated path p' of the pixel corresponding to the projector ray r and (ii) the path p of the pixel corresponding to the projector ray r from the stored calibration value, and make it (b) perform using each parameter of the calibration object calculated based on the captured images.

[0244] In some application examples, the calibration object is a three-dimensional calibration object with a known shape, and the step of driving each camera of the one or more cameras to capture an image of the three-dimensional calibration object includes the step of driving each camera of the one or more cameras to capture an image of the three-dimensional calibration object, and the predetermined parameters of the calibration object are the dimensions of the three-dimensional calibration object.

[0245] In some application examples, the object to be calibrated is a two-dimensional calibration object having visually distinguishable features, and the step of driving each of the one or more cameras to capture a plurality of images of the calibration object includes driving each of the one or more cameras to capture an image of the two-dimensional calibration object, and the predetermined parameters of the two-dimensional calibration object are the respective distances between the respective visually distinguishable features.

[0246] According to some applications of the present invention, a method for calculating the three-dimensional structure of an intraoral three-dimensional surface is further provided, the method comprising: scanning the intraoral surface; driving one or more uniform light projectors to project broadband spectral light onto the intraoral three-dimensional surface; driving a camera to capture a plurality of two-dimensional color images of the intraoral three-dimensional surface; using a processor, calculating the three-dimensional positions of a plurality of points on the intraoral three-dimensional surface based on the intraoral surface scan; calculating the three-dimensional structure of the intraoral three-dimensional surface constrained by the three-dimensional positions of the plurality of points based on the plurality of two-dimensional color images of the intraoral three-dimensional surface.

[0247] In some embodiments, the intraoral surface is scanned by driving one or more structured light projectors to project a structured light pattern onto the intraoral three-dimensional surface, driving one or more cameras to capture a plurality of structured light images, each image including at least a portion of the structured light pattern.

[0248] In some application examples, the step of driving the one or more structured light projectors includes driving the one or more structured light projectors to project respective distributions of discrete non-connected light spots.

[0249] In some application examples, the step of calculating the three-dimensional structure includes (a) The plurality of two-dimensional color images of the intraoral three-dimensional surface, and (b) the step of inputting the calculated three-dimensional positions of the plurality of points on the intraoral three-dimensional surface into the neural network; Determining, by the neural network, a respective predicted depth map of the intraoral three-dimensional surface captured in each of the two-dimensional color images, including.

[0250] In some applications, the method further includes the step of using the processor to stitch together the respective depth maps to obtain a three-dimensional structure of the intraoral three-dimensional surface.

[0251] In some applications, the method further includes the step of adjusting the capture of the structured light image and the capture of the two-dimensional color image to generate an alternating sequence in which one or more broadband spectral light image frames are interspersed in one or more structured light image frames.

[0252] In some applications, The step of driving one or more cameras to capture a plurality of structured light images includes the step of driving each of the two or more cameras to capture a plurality of structured light images, The step of driving the cameras to capture the plurality of two-dimensional color images includes the step of driving each of the two or more cameras to capture the plurality of two-dimensional color images.

[0253] In some applications, The step of determining by the neural network includes the step of determining, for a given image frame, a respective predicted depth map of the portion of the intraoral three-dimensional surface captured in the two-dimensional color image by each of the two or more cameras, The method further includes the step of using the processor to stitch together the respective depth maps to obtain a predicted depth map of the intraoral three-dimensional surface captured in the given image frame.

[0254] In some applications, the method further includes the step of training the neural network, and the training includes (a) driving one or more structured light projectors to project a training stage structured light pattern onto a training stage three-dimensional surface; (b) driving one or more training stage cameras to capture a plurality of structured light images, each image including at least a portion of the training stage structured light pattern; (c) driving one or more training stage uniform light projectors to project broadband spectral light onto the training stage three-dimensional surface; (d) driving one or more training stage cameras to capture a plurality of two-dimensional color images of the training stage three-dimensional surface using illumination from the training stage uniform light projector; (e) coordinating the capture of the structured light images and the capture of the two-dimensional color images to produce an alternating sequence in which image frames of one or more two-dimensional color images are interspersed among image frames of one or more structured light images; (f) inputting the plurality of two-dimensional color images into the neural network; (g) based on the structured light images of the training stage three-dimensional surface, inputting into the neural network respective multiple three-dimensional reconstructions of the training stage three-dimensional surface, the three-dimensional reconstructions including the calculated three-dimensional positions of multiple points on the training stage three-dimensional surface; (h) interpolating the positions of the one or more training stage cameras relative to the training stage three-dimensional surface for each two-dimensional color image frame based on the calculated three-dimensional positions of multiple points on the training stage surface calculated based on the respective structured light image frames before and after each two-dimensional color image frame; (i) projecting the three-dimensional reconstructions into the respective fields of view of the one or more training stage cameras and, based on the projection, estimating a predicted depth map of the training stage three-dimensional surface as seen in each two-dimensional color image, constrained by the calculated three-dimensional positions of the multiple points; (j) Comparing each predicted depth map of the training stage 3D surface with the corresponding true depth map of the training stage 3D surface; (k) Optimizing the neural network based on the difference between each predicted depth map and the corresponding true depth map to better estimate subsequent predicted depth maps; comprising.

[0255] In some applications, the step of driving the one or more structured light projectors to project the training stage structured light pattern includes driving the one or more structured light projectors to project, respectively, a distribution of discrete and unconnected light spots on the training stage 3D surface.

[0256] In some applications, the step of driving one or more training stage cameras includes driving at least two training stage cameras.

[0257] According to some applications of the present invention, an intraoral scanning device is further provided, the device comprising: an elongated hand-held wand including a probe at its distal end; one or more illumination sources coupled to the probe; one or more near-infrared (NIR) light sources coupled to the probe; one or more cameras coupled to the probe and configured to (a) capture an image using light from the one or more illumination light sources and (b) capture an image using NIR light from the NIR light sources; a processor configured to execute a navigation algorithm to determine the position of the hand-held wand as the hand-held wand moves within a space, wherein inputs to the navigation algorithm are (a) an image captured using light from the one or more illumination light sources and (b) an image captured using the NIR light.

[0258] In some application examples, the one or more illumination light sources are one or more structured light sources.

[0259] In some application examples, the one or more illumination light sources are one or more uniform light sources.

[0260] According to some application examples of the present invention, a method for tracking the motion of an intraoral scanner is further provided, the method comprising: Illuminating the intraoral three-dimensional surface using one or more illumination sources coupled to the intraoral scanner; Using one or more near-infrared (NIR) light sources coupled to the intraoral scanner to drive each of the one or more NIR light sources to emit NIR light onto the intraoral three-dimensional surface; Using one or more cameras coupled to the intraoral scanner to (a) capture a plurality of images using the light from the one or more illumination light sources and (b) capture a plurality of images using the NIR light; Using a processor to Execute a navigation algorithm to track the motion of the intraoral scanner relative to the intraoral three-dimensional surface using (a) the images captured using the light from the one or more illumination light sources and (b) the images captured using the NIR light; Including.

[0261] In some application examples, the step of using one or more illumination light sources includes illuminating the intraoral three-dimensional surface using one or more structured light sources.

[0262] In some application examples, the step of using one or more illumination sources includes using one or more uniform light sources.

[0263] According to some application examples of the present invention, an intraoral scanning device for use with a sleeve is further provided, the device comprising: An elongated hand-held wand, comprising an elongated hand-held wand having a probe at its distal end configured to be removably disposed within the sleeve, At least one structured light projector coupled to the probe, the structured light projector including: (a) a laser having an illumination field of at least 30 degrees, (b) a laser configured to emit polarized laser light, and (c) a pattern generation optical element configured to generate a pattern of light when the laser diode is activated to transmit light through the pattern generation optical element. At least one camera coupled to the probe and including a camera sensor. The probe is configured such that light enters and exits the probe through the sleeve. The laser is disposed at a distance from the camera such that when the probe is disposed within the sleeve, a portion of the pattern of light is reflected from the sleeve and reaches the camera sensor. The laser is positioned at an angle of rotation with respect to its own optical axis such that, due to the polarization of the pattern of light, the degree of reflection of the portion of the pattern of light by the sleeve is less than 70% of the maximum reflection for all possible angles of rotation of the laser with respect to its optical axis.

[0264] In some applications, the distance between the structured light projector and the camera is 1 to 6 times the distance between the structured light projector and the sleeve when the hand-held wand is disposed within the sleeve.

[0265] In some applications, each camera of the at least one camera has a field of view of at least 30 degrees.

[0266] In some applications, the laser is positioned at an angle of rotation with respect to its own optical axis such that, due to the polarization of the pattern of light, the degree of reflection of the portion of the pattern of light by the sleeve is less than 60% of the maximum reflection for all possible angles of rotation of the laser with respect to its optical axis.

[0267] In some applications, the laser is positioned at an angle of rotation with respect to its own optical axis such that, due to the polarization of the pattern of light, the degree of reflection by the sleeve of the portion of the pattern of light is between 15% and 60% of maximum reflection for all possible angles of rotation of the laser with respect to its optical axis.

[0268] In accordance with some applications of the present invention, a method of generating a three-dimensional image using an intraoral scanner is further provided, the method comprising: (A) using at least two cameras rigidly attached to the intraoral scanner such that the respective fields of view of the cameras have non-overlapping portions, capturing a plurality of images of the intraoral three-dimensional surface; (B) using a processor, for each non-overlapping portion of the respective fields of view, performing a simultaneous localization and mapping (SLAM) algorithm using the captured images from each camera, wherein the localization of each camera is solved based on the motion of each camera being the same as the motion of all the other cameras.

[0269] In some applications, the respective fields of view of the first camera among the cameras and the second camera among the cameras also have an overlapping portion, the capturing step includes capturing a plurality of images of the intraoral three-dimensional surface such that features of the intraoral three-dimensional surface in the overlapping portion of the respective fields of view appear in the images captured by the first and second cameras. The step of using the processor includes performing a SLAM algorithm using the features of the intraoral three-dimensional surface that appear in the images of the at least two cameras.

[0270] In accordance with some applications of the present invention, a method of generating a three-dimensional image using an intraoral scanner is further provided, the method comprising: driving one or more structured light projectors to project a structured light pattern onto the intraoral three-dimensional surface; Driving one or more cameras to capture a plurality of structured light images, each image including at least a portion of the structured light pattern; Driving one or more uniform light projectors to project broadband spectral light onto the three-dimensional surface within the oral cavity; Driving at least one camera to capture a two-dimensional color image of the three-dimensional surface within the oral cavity using illumination from the uniform light projector; Adjusting the capture of the structured light and the capture of the broadband spectral light to generate an alternating sequence in which one or more broadband spectral light image frames are interspersed among one or more structured light image frames; Using a processor, Calculating the three-dimensional position of each of a plurality of points on the three-dimensional surface within the oral cavity captured in the plurality of structured light image frames; Interpolating the motion of the at least one camera between a first broadband spectral light image frame and a second broadband spectral light image frame based on the calculated three-dimensional positions of the plurality of points in each of the structured light image frames before and after the broadband spectral light image frames; Performing a simultaneous localization and mapping (SLAM) algorithm that is (a) constrained by the motion of the camera interpolated between the first broadband spectral light image frame and the second broadband spectral light image frame and (b) uses features of the three-dimensional surface within the oral cavity captured in the first and second broadband spectral light image frames by the at least one camera.

[0271] In some applications, the step of driving the one or more structured light projectors to project the structured light pattern includes driving the one or more structured light projectors to project respective distributions of discrete, unconnected light spots onto the three-dimensional surface within the oral cavity.

[0272] According to some applications of the present invention, there is further provided a method of generating a three-dimensional image using an intraoral scanner, the method comprising: Driving one or more structured light projectors to project a structured light pattern onto the three-dimensional surface within the oral cavity; Driving one or more cameras to capture a plurality of structured light images, each image including at least a portion of the structured light pattern; Driving one or more uniform light projectors to project broadband spectral light onto the three-dimensional surface within the oral cavity; Driving the one or more cameras to capture a two-dimensional color image of the three-dimensional surface within the oral cavity using the illumination from the uniform light projector; Adjusting the capture of the structured light and the capture of the broadband spectral light to generate an alternating sequence in which one or more broadband spectral light image frames are interspersed among one or more structured light image frames; Using a processor, (a) calculating the three-dimensional positions of features on the three-dimensional surface within the oral cavity based on the structured light image frames, the features also being captured in a first image frame of the broadband spectral light and a second image frame of the broadband spectral light; (b) calculating the motion of at least one camera between the first image frame of the broadband spectral light and the second image frame of the broadband spectral light based on the calculated three-dimensional positions of the features; (c) performing a simultaneous localization and mapping (SLAM) algorithm using (i) features of the three-dimensional surface within the oral cavity for which three-dimensional positions were not calculated based on the structured light image frames captured in the first and second broadband spectral light image frames by the at least one camera, and (ii) the calculated motion of the camera between the first and second broadband spectral light image frames.

[0273] In some applications, the step of driving the one or more structured light projectors to project the structured light pattern includes driving the one or more structured light projectors to project respective distributions of discrete, unconnected light spots onto the three-dimensional surface within the oral cavity.

[0274] According to some applications of the present invention, a method for calculating the three-dimensional structure of the intraoral three-dimensional surface in a subject is further provided, the method comprising: driving one or more structured light projectors to project structured light pattern spots onto the intraoral three-dimensional surface; driving one or more cameras to capture a plurality of structured light images, each image including at least one of the spots; driving one or more uniform light projectors to project broadband spectral light onto the intraoral three-dimensional surface; driving at least one camera to capture a two-dimensional color image of the intraoral three-dimensional surface using illumination from the uniform light projector; adjusting the capture of the structured light and the capture of the broadband spectral light to generate an alternating sequence in which one or more broadband spectral light image frames are interspersed among one or more structured light image frames; using a processor, judging, for each of the plurality of spots, whether the spot is projected onto a moving tissue or a stable tissue in the oral cavity based on the two-dimensional color image, and assigning a respective confidence grade to each of the detected plurality of spots based on the judgment, with a high confidence assigned to fixed tissue and a low confidence assigned to moving tissue; executing a three-dimensional reconstruction algorithm using the detected spots based on the confidence grade for each of the plurality of detected spots.

[0275] In some application examples, the step of driving the one or more structured light projectors to project the structured light pattern includes driving the one or more structured light projectors to project respective distributions of discrete, unconnected light spots onto the intraoral three-dimensional surface.

[0276] In some application examples, the step of executing the three-dimensional reconstruction algorithm includes the step of executing the three-dimensional reconstruction algorithm using only a subset of the detected spots, and the subset consists of spots assigned a confidence grade exceeding a fixed tissue threshold.

[0277] In some application examples, the step of executing the three-dimensional reconstruction algorithm includes: (a) for each spot, assigning a weight to the spot based on the respective confidence assigned to the spot; and (b) using the respective weights of each spot in the three-dimensional reconstruction algorithm.

[0278] The present invention will be more fully understood from the following detailed description of its application examples taken in conjunction with the drawings.

Brief Description of the Drawings

[0279]

Figure 1

Figure 2A

Figure 2B

Figure 2C

Figure 2D

Figure 2E

Figure 3

Figure 4

Figure 5A

Figure 5B

Figure 6A

Figure 6B

Figure 7

Figure 8

Figure 9

Figure 12

Figure 13

Figure 14

Figure 17

Figure 18

Figure 19A

Figure 19B

Figure 20A

Figure 20E

Figure 21A

Figure 21C

Figure 22A

Figure 22B

Figure 23A

Figure 23B

Figure 24A

Figure 24B

Figure 25

Figure 26A

Figure 26B

Figure 27A

Figure 27B

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34A

Figure 34B

Figure 35

Figure 36

Figure 37A

Figure 37B

Figure 38

Figure 39

Figure 40A

Figure 40B

Figure 41

Figure 42A

Figure 42B

Figure 43A

Figure 43B

Figure 44A

Figure 45

Figure 46A

Figure 46D

Figure 47A

Figure 47B

Figure 48A

Figure 48B

Figure 49A

Figure 49B

Figure 50

Figure 51A

Figure 51F

Figure 51G

Figure 51I

Figure 52A

Figure 52B

Figure 52C

Figure 52D

Figure 52F

Figure 52G

Figure 53A

Figure 53B

Figure 54A

Figure 54B

Figure 55

Figure 56A

Figure 56B

Figure 57

Figure 58

Figure 59B

Figure 60

Figure 61

Figure 62

Figure 63

Figure 64

Figure 65

Figure 66

Best Mode for Carrying Out the Invention

[0280] Next, refer to FIG. 1, which is a schematic diagram of an elongated hand-held wand 20 for intraoral scanning according to some application examples of the present invention. A plurality of structured light projectors 22 and a plurality of cameras 24 are coupled to a rigid structure 26 disposed at the distal end 30 of the hand-held wand within the probe 28. In some application examples, during intraoral scanning, the probe 28 enters the subject's oral cavity.

[0281] In some application examples, the structured light projector 22 is positioned within the probe 28 such that each structured light projector 22 faces an object 32 external to the hand-held wand 20 within its illumination field, as opposed to positioning the structured light projector at the proximal end of the hand-held wand and illuminating the object by reflection of light from a mirror and subsequent reflection of the light onto the object. Similarly, in some application examples, the camera 24 is disposed within the probe 28 such that each camera 24 faces an object 32 external to the hand-held wand 20 within its field of view, as opposed to positioning the camera at the proximal end of the hand-held wand and viewing the object by reflection of light from a mirror to the camera. Such positioning of the projectors and cameras within the probe 28 allows the scanner to have a relatively large overall field of view while maintaining a low-profile probe.

[0282] In some application examples, the height H1 of the probe 28 is less than 15 mm, and the height H1 of the probe 28 is measured from the lower surface 176 (sensing surface) where the reflected light from the object 32 being scanned enters the probe 28 to the upper surface 178 on the opposite side of the lower surface 176. In some application examples, the height H1 is between 10 and 15 mm.

[0283] In some applications, each of the cameras 24 has a large field of view β (beta) of at least 45 degrees, for example, at least 70 degrees, for example, at least 80 degrees, for example, 85 degrees. In some applications, the field of view may be less than 120 degrees, for example, less than 100 degrees, for example, less than 90 degrees. In experiments conducted by the inventors, it has been found that a field of view β (beta) of each camera being between 80 degrees and 90 degrees is particularly useful as it provides a good balance among pixel size, field of view and camera overlap, optical quality, and cost. The camera 24 may include an image sensor 58 and an objective optical system 60 including one or more lenses. To enable close-focus imaging, the camera 24 may be focused on an object focal plane 50 located between 1 mm and 30 mm, for example, between 4 mm and 24 mm, for example, between 5 mm and 11 mm, for example, between 9 mm and 10 mm from the lens farthest from the camera sensor. In experiments conducted by the inventors, it has been found that it is particularly useful for the object focal plane 50 to be located between 5 mm and 11 mm from the lens farthest from the camera sensor, because it is easy to scan teeth at this distance and the focus degree of most tooth surfaces is high. In some applications, the camera 24 can capture images at a frame rate of at least 30 frames per second, for example, at least 75 frames per second, for example, at least 100 frames per second. In some applications, the frame rate may be less than 200 frames per second.

[0284] As described above, the large field of view achieved by combining the fields of view of each of all the cameras can improve the accuracy by reducing the amount of image stitching error, particularly in edentulous regions where the gingival surface is smooth and there may be few distinct high-resolution 3D features. With a large field of view, large smooth features such as the overall curve of the teeth appear in each image frame, and the stitching accuracy of each surface obtained from such a plurality of image frames is improved.

[0285] Similarly, each structured light projector 22 may have a large illumination field α (alpha) of at least 45 degrees, for example at least 70 degrees. In some applications, the illumination field α (alpha) may be less than 120 degrees, for example less than 100 degrees. Further features of the structured light projector 22 will be described below.

[0286] In some applications, to improve image capture, each camera 24 has a plurality of discrete preset focus positions, and at each focus position, the camera is focused on the respective object focal plane 50. Each of the cameras 24 may include an autofocus actuator that selects a focus position from the discrete preset focus positions to improve a given image capture. Additionally or alternatively, each camera 24 includes an optical aperture phase mask that extends the depth of focus of the camera so that the image formed by each camera remains in focus over all object distances where the image is between 1 mm and 30 mm, for example between 4 mm and 24 mm, for example between 5 mm and 11 mm, for example between 9 mm and 10 mm, from the lens farthest from the camera sensor.

[0287] In some applications, the structured light projectors 22 and cameras 24 are coupled to the rigid structure 26 closely and / or alternately such that (a) most of the field of view of each camera overlaps with the field of view of an adjacent camera, and (b) most of the field of view of each camera overlaps with the illumination field of an adjacent projector. Optionally, at least 20%, for example at least 50%, for example at least 75% of the projected light pattern is present in at least one field of view of the camera at the object focal plane 50 that is at least 4 mm away from the lens farthest from the camera sensor. Due to the different possible configurations of the projectors and cameras, a portion of the projected pattern may never be seen within the field of view of any camera, and a portion of the projected pattern may be blocked from the field of view by the object 32 as the scanner moves around during scanning.

[0288] The rigid structure 26 may be a non-flexible structure to which the structured light projector 22 and the camera 24 are coupled to provide structural stability to the optical system within the probe 28. Coupling all the projectors and all the cameras to a common rigid structure serves to maintain the geometric integrity of the optical systems of each structured light projector 22 and each camera 24 under changing ambient conditions, such as mechanical stresses that may be induced by the subject's mouth. Further, the rigid structure 26 helps to maintain the stable structural integrity and the positioning of the structured light projector 22 and the camera 24 relative to each other. As will be further described below, controlling the temperature of the rigid structure 26 can help to maintain the geometric integrity of the optical system over a wide range of ambient temperatures when the probe 28 enters and exits the subject's oral cavity or when the subject breathes during the scan.

[0289] Next, reference is made to FIGS. 2A - B, which are schematic diagrams of the respective positioning configurations of the camera 24 and the structured light projector 22 according to some application examples of the present invention. In some application examples, the camera 24 and the structured light projector 22 are positioned so as not to face in the same direction as each other in order to improve the overall field of view and illumination field of the intraoral scanner. In some application examples as shown in FIG. 2A, a plurality of cameras 24 are coupled to the rigid structure 26 such that the angle θ (theta) between the two respective optical axes 46 of at least two cameras 24 is 90 degrees or less, for example, 35 degrees or less. Similarly, in some application examples as shown in FIG. 2B, a plurality of structured light projectors 22 are coupled to the rigid structure 26 such that the angle φ (phi) between the two respective optical axes 48 of at least two structured light projectors 22 is 90 degrees or less, for example 35 degrees or less.

[0290] Next, refer to FIG. 2C, which is a chart depicting a plurality of different configurations regarding the positions of the structured light projector 22 and the camera 24 within the probe 28, according to several application examples of the present invention. The structured light projector 22 is represented by a circle in FIG. 2C, and the camera 24 is represented by a rectangle in FIG. 2C. Note that typically, since the field of view β (beta) of each camera sensor 58 and each camera 24 has an aspect ratio of 1:2, a rectangle is used to represent the camera. Column (a) of FIG. 2C shows a bird's-eye view of various configurations of the structured light projector 22 and the camera 24. The labeled x-axis in the first row of column (a) corresponds to the central longitudinal axis of the probe 28. Column (b) shows side views of the camera 24 from various configurations, as viewed along a line of sight that is coaxial with the central longitudinal axis of the probe 28. Similar to that shown in FIG. 2A, column (b) of FIG. 2C shows cameras 24 arranged to have optical axes 46 at an angle of 90 degrees or less, for example 35 degrees or less, relative to each other. Column (c) shows side views of the camera 24 from various configurations, as viewed along a line of sight perpendicular to the central longitudinal axis of the probe 28.

[0291] Typically, the most distal (towards the positive x-direction in FIG. 2C) and the most proximal (towards the negative x-direction in FIG. 2C) cameras 24 are positioned such that their optical axes 46 are rotated slightly inward, for example 90 degrees or less, for example 35 degrees or less, relative to the next-nearest camera 24. The more centrally located camera(s) 24, i.e., the camera 24 that is neither the most distal camera 24 nor the most proximal camera 24, are arranged to directly face the outside of the probe, and their optical axes 46 are substantially perpendicular to the central longitudinal axis of the probe 28. Note that in row (xi), the projector 22 is arranged at the most distal position of the probe 28, and thus the optical axis 48 of that projector 22 faces inward, such that more spots 33 projected from that particular projector 22 can be seen by more cameras 24.

[0292] Typically, the number of structured light projectors 22 within the probe 28 may range from two, as shown in row (iv) of FIG. 2C for example, to six, as shown in row (xii) for example. Typically, the number of cameras 24 of the probe 28 may range from four, as shown in rows (iv) and (v) for example, to seven, as shown in row (ix) for example. It should be noted that the various configurations shown in FIG. 2C are illustrative and not limiting, and the scope of the present invention includes additional configurations not shown. For example, the scope of the present invention includes more than five projectors 22 arranged in the probe 28 and more than seven cameras arranged in the probe 28.

[0293] Next, refer to FIGS. 2D - E, which are isometric views of a particular configuration regarding the positions of the structured light projectors 22 and cameras 24 within the probe 28, shown from two different respective viewpoints according to some application examples of the present invention. FIG. 2D is shown from the same bird's-eye view perspective as that of column (a) of FIG. 2C. In some application examples, six cameras 24 are arranged equidistantly within the probe 28, with three cameras on each side of the probe 28, and five structured light projectors 22 are arranged within the central probe 28 along the central longitudinal axis of the probe 28 (illustrated by the dashed line 29).

[0294] In some applications, the camera 24 and the structured light projector 22 are all coupled to a flexible printed circuit board (PCB) to accommodate the angular positioning of the camera 24 and the structured light projector 22 within the probe 28. This angular positioning of the camera 24 and the structured light projector 22 is shown in FIG. 2E. The most distal (i.e., in the positive x - direction) camera 24 and structured light projector 22 are positioned such that their respective optical axes are tilted rearward at an angle of, for example, 45 degrees or less, for example 35 degrees or less, towards the hand - held wand 20. Thereby, for example, the most distal camera can capture the posterior wall of the posterior molars in the oral cavity. The most proximal (i.e., in the negative x - direction) camera 24 and structured light projector 22 are positioned such that their respective optical axes are tilted forward at an angle of, for example, 45 degrees or less, for example 35 degrees or less, towards the distal end of the probe 28 to obtain an improved overlap of the respective fields of view of the camera 24. Further, all of the structured light projectors 22 are arranged such that their respective optical axes are tilted towards the center of the probe 28, thereby improving the overlap of the respective illumination fields of the structured light projectors 22. The inventors have realized that by arranging substantially all of the structured light projectors 22 in a row, they can be more easily connected to the same flexible PCB.

[0295] Further, a plurality of uniform light projectors 118, a plurality of near - infrared (NIR) light projectors 292, and diffractive optical elements (DOE) 39 disposed on each structured light projector 22 are shown in FIGS. 2D - E and will be further described below.

[0296] Next, refer to FIG. 3, which is a schematic diagram of the structured light projector 22 according to some application examples of the present invention. In some application examples, the structured light projector 22 includes a laser diode 36, a beam shaping optical element 40, and a pattern generation optical element 38 that generates a distribution 34 of discrete and unconnected light spots (further described below with reference to FIG. 4). In some application examples, the structured light projector 22 is configured to generate a distribution 34 of discrete and unconnected light spots on all planes located between 1 mm and 30 mm, for example, between 4 mm and 24 mm, when the laser diode 36 transmits light through the pattern generation optical element 38. In some application examples, the distribution 34 of discrete and unconnected light spots is in focus on one plane located between 1 mm and 30 mm, for example, between 4 mm and 24 mm, but there are still discrete and unconnected light spots on all other planes located between 1 mm and 30 mm, for example, between 4 mm and 24 mm. Although the use of a laser diode has been described above, it should be understood that this is an exemplary and non-limiting application example. Other light sources may be used in other application examples. Further, although the projection of a pattern of discrete and unconnected light spots has been described, it should be understood that this is an exemplary and non-limiting application example. In other application examples, other patterns or arrays of light, including but not limited to lines, gratings, checkerboards, and other arrays, may be used. In some application examples, the light pattern projected by the structured light projector is spatially fixed with respect to one or more cameras.

[0297] In this specification, embodiments are described with reference to discrete light spots and with reference to performing operations that use or are based on the spots. Examples of such operations include steps of solving a corresponding algorithm to determine the position of a light spot, tracking a light spot, mapping a projector beam to the light spot, identifying a vulnerable light spot, and generating a three-dimensional model based on the position of the spot. It should be understood that such operations and other operations described with reference to spots also function for other features of other projection light patterns. Thus, the considerations in this specification with reference to spots also apply to other features of the projection light pattern.

[0298] The pattern generation optical element 38 may be configured to have a light throughput efficiency of at least 80%, for example at least 90% (i.e., the ratio of the light entering the pattern to the total light incident on the pattern generation optical element 38).

[0299] In some applications, each laser diode 36 of each structured light projector 22 transmits light at a different wavelength, i.e., each laser diode 36 of at least two structured light projectors 22 transmits light at two different wavelengths. In some applications, each laser diode 36 of at least three structured light projectors 22 transmits light at three different wavelengths. For example, red, blue, and green laser diodes may be used. In some applications, each laser diode 36 of at least two structured light projectors 22 transmits light at two different wavelengths. For example, in some applications, there are six structured light projectors 22 disposed within the probe 28, three of which include blue laser diodes and three of which include green laser diodes.

[0300] Next, refer to FIG. 4, which is a schematic diagram of a structured light projector 22 that projects a distribution of discrete, unconnected light spots onto multiple object focal planes according to some application examples of the present invention. The object 32 to be scanned may be one or more teeth or other intraoral objects / tissues in the subject's mouth. The somewhat translucent and glossy properties of the teeth can affect the contrast of the projected structured light pattern. For example, (a) part of the light hitting the teeth may scatter into other regions within the intraoral scene, potentially causing a certain degree of stray light, and (b) part of the light may pass through the teeth and then exit the teeth at any other point. Therefore, the inventors have found that, without using contrast enhancement means such as coating the teeth with an opaque powder, a sparse distribution 34 of discrete, unconnected light spots can provide an improved balance regarding reducing the amount of projected light while maintaining a useful amount of information for improving image capture of the intraoral scene under structured light illumination. The sparsity of the distribution 34 can be characterized by the following ratio. (a) The ratio of, the total area of the illuminated regions on the orthogonal plane 44 in the illumination field α (alpha), i.e., the sum of the areas of all the projection spots 33 on the orthogonal plane 44 in the illumination field α (alpha),

[0301] In some application examples, each structured light projector 22 projects at least 400 discrete, unconnected spots 33 onto the intraoral three-dimensional surface during scanning. In some application examples, each structured light projector 22 projects less than 3000 discrete, unconnected spots 33 onto the intraoral surface during scanning. To reconstruct the three-dimensional surface from the projected sparse distribution 34, as will be further described below with reference to FIGS. 7 - 19, the correspondence between each projection spot 33 (or other features of the projection pattern) and the detected spots (or other features) by the camera 24 must be determined.

[0302] In some applications, the pattern generation optical element 38 is a diffractive optical element (DOE) 39 (FIG. 3), and when the laser diode 36 transmits light through the DOE 39 to the object 32, it generates a distribution 34 of discrete, unconnected light spots 33. Throughout this specification, including the claims, a light spot is defined as a small region of light having any shape. In some applications, each DOE 39 of the different structured light projectors 22 generates spots having different respective shapes, i.e., all spots 33 generated by a particular DOE 39 have the same shape, and the shape of the spots 33 generated by at least one DOE 39 is different from the shape of the spots 33 generated by at least one other DOE 39. By way of example, some DOE 39 may generate circular spots 33 (as shown in FIG. 4), some DOE 39 may generate square spots, and some DOE 39 may generate elliptical spots. Optionally, some DOE 39 may generate linear patterns that are connected or unconnected.

[0303] Next, refer to FIGS. 5A - B, which are schematic diagrams of structured light projector 22 according to some applications of the present invention, including beam shaping optical element 40 and additional optical elements, such as DOE 39, disposed between beam shaping optical element 40 and pattern generation optical element 38. Optionally, beam shaping optical element 40 is collimating lens 130. Collimating lens 130 may be configured to have a focal length of less than 2 mm. Optionally, the focal length may be at least at least 1.2 mm. In some applications, additional optical element 42 disposed between beam shaping optical element 40 and pattern generation optical element 38, such as DOE 39, generates a Bessel beam when laser diode 36 transmits light through optical element 42. In some applications, the Bessel beam passes through a wide range of orthogonal planes 44 (e.g., each orthogonal plane located between 1 mm and 30 mm from DOE 39, e.g., each orthogonal plane located between 4 mm and 24 mm from DOE 39, etc.) such that all discrete unconnected light spots 33 maintain a small diameter (e.g., less than 0.06 mm, e.g., less than 0.04 mm, e.g., less than 0.02 mm). The diameter of spot 33 is defined by the full width at half maximum (FWHM) of the intensity of the spot in the context of this patent application.

[0304] Despite the above description that all spots are smaller than 0.06 mm, some spots having a diameter (e.g., only slightly smaller than 0.06 mm, or 0.02 mm) near the upper end of these ranges, also near the edge of the illumination field of projector 22, may elongate when intersecting a geometric plane orthogonal to DOE 39. In such cases, it is useful to measure their diameters when intersecting the inner surface of a geometric sphere having a radius of 1 mm or more and 30 mm or less corresponding to the distance of each orthogonal plane located 1 mm or more and 30 mm or less from DOE 39 centered on DOE 39. Throughout this application, including the claims, the term "geometric" is considered to relate to theoretical geometric constructs (such as planes or spheres, etc.) and is not part of any physical device.

[0305] In some application examples, when the Bessel beam is transmitted through DOE39, in addition to the spot having a diameter less than 0.06 mm, a spot 33 having a diameter exceeding 0.06 mm is generated.

[0306] In some application examples, the optical element 42 is an axicon lens 45 shown in FIG. 5A and further described below with reference to FIGS. 23A - B. Alternatively, the optical element 42 may be an annular aperture ring 47 as shown in FIG. 5B. By maintaining the small diameter of the spot, the 3D resolution and accuracy are improved throughout the depth of focus. Without the optical element 42, for example, the axicon lens 45 or the annular aperture ring 47, the size of the spot 33 may change, for example, increase, as it moves further away from the best focus plane due to diffraction and defocus.

[0307] Next, refer to FIGS. 6A - B, which are schematic diagrams of a structured light projector 22 that projects discrete, unconnected spots 33 according to several application examples of the present invention, and a camera sensor 58 that detects spots 33'. In several application examples, a method is provided for determining the correspondence between the projection spots 33 on the oral cavity inner surface and the detection spots 33' on each camera sensor 58. As described above, this method is also applicable to the step of determining the correspondence between other projected features on the oral cavity inner surface and the detected features on each camera sensor. Once the correspondence is determined, a three - dimensional image of the surface is reconstructed. Each camera sensor 58 has a pixel array, and for each of them, there is a corresponding camera ray 86. Similarly, for each projection spot 33 from each projector 22, there is a corresponding projector ray 88. Each projector ray 88 corresponds to the path 92 of each pixel on at least one of the camera sensors 58. Therefore, when a camera views a spot 33' projected by a specific projector ray 88, that spot 33' will necessarily be detected by the pixels on a specific path 92 of the pixels corresponding to that specific projector ray 88. Referring particularly to FIG. 6B, the correspondence between each projector ray 88 and each camera sensor path 92 is shown. Projector ray 88' corresponds to camera sensor path 92', projector ray 88'' corresponds to camera sensor path 92'', and projector ray 88'' corresponds to camera sensor path 92''. For example, when a specific projector ray 88 projects a spot into a dusty space, the lines of dust in the air will be illuminated. The lines of dust detected by the camera sensor 58 will follow the same path on the camera sensor 58 as the camera sensor path 92 corresponding to the specific projector ray 88.

[0308] During the calibration process, calibration values are stored based on camera rays 86 corresponding to pixels on the camera sensor 58 of each camera of the cameras 24, and projector rays 88 corresponding to the projection spots 33 (or other features) of the light from each structured light projector 22. For example, the calibration values may be stored for (a) a plurality of camera rays 86 corresponding to respective ones of the plurality of pixels on each camera sensor 58 of the cameras 24, and (b) a plurality of projector rays 88 corresponding to respective ones of the plurality of projection spots 33 of the light from each structured light projector 22. As used throughout this application, including the claims, a stored calibration value indicative of a camera ray corresponding to each pixel on the camera sensor of each camera means (a) a value given to each camera ray, or (b) a parameterized camera calibration model, e.g., a parameter value of a function. As used throughout this application, including the claims, a stored calibration value indicative of a projector ray corresponding to each projection spot (or other projected feature) of the light from each structured light projector means (a) a value given to each projector ray, e.g., in an indexed list, or (b) a parameterized projector calibration model, e.g., a parameter value of a function.

[0309] As an example, the following calibration process may be used. A high-precision dot target, e.g., black dots on a white background, is illuminated from below, and images of the target are taken with all cameras. Next, the dot target is moved vertically towards the cameras, i.e., along the z-axis, up to the target plane. For all dots at all positions in the z-axis direction, the dot centers are calculated, and a three-dimensional grid of dots is created in space. Then, using distortion and the camera pinhole model, the pixel coordinates of each three-dimensional position of each dot center are determined, thereby defining, for each pixel, a camera ray in the direction from that pixel towards the corresponding dot center of the three-dimensional grid. The camera rays corresponding to pixels between grid points may be interpolated. The above-described camera calibration procedure is repeated for all respective wavelengths of each laser diode 36 such that the camera rays 86 corresponding to each pixel on each camera sensor 58 for each respective wavelength are included in the stored calibration values. Alternatively, the stored calibration values are parameter values of the distortion and camera pinhole model, indicating the values of the camera rays 86 corresponding to each pixel on each camera sensor 58 for each respective wavelength.

[0310] After the camera 24 is calibrated and the values of all camera rays 86 are stored, the structured light projector 22 may be calibrated as follows. A target without flat features is used, and the structured light projector 22 is turned on one at a time. Each spot (or other feature) is placed on at least one camera sensor 58. Since the camera 24 is now calibrated, the three-dimensional spot position of each spot (or other feature) is calculated by triangulation based on the images of the spot (or other feature) in a plurality of different cameras. The above-described process is repeated with the featureless target placed at a plurality of different z-axis positions. Each projected spot (or other feature) on the featureless target will define a projector ray in space starting from the projector.

[0311] Next, refer to FIG. 7, which is a flowchart showing an overview of a method for generating a digital three-dimensional image according to some application examples of the present invention. In steps 62 and 64 of the method outlined in FIG. 7, each structured light projector 22 is driven to project a pattern of light (e.g., a distribution 34 of discrete non-connected light spots 33) onto the three-dimensional surface within the oral cavity, and each camera 24 is driven to capture an image including at least a portion of the pattern (e.g., one of the spots 33). Based on stored calibration values indicating (a) each camera ray 86 corresponding to each pixel on the camera sensor 58 of each camera 24 and (b) each projector ray 88 corresponding to each projection spot 33 of light from each structured light projector 22, a corresponding algorithm is executed in step 66 using a processor 96 (FIG. 1), which will be further described below with reference to FIGS. 8 - 12. In some embodiments, the processor 96 is a processor disposed in the elongated handheld wand 20. In some embodiments, the processor 96 is disposed in a computing device as described below with reference to FIGS. 60 - 61, and this device may be operably connected to the elongated handheld wand 20 (e.g., via a wired or wireless connection). In some embodiments, multiple processors are used, in which case one or more processors may be disposed in the elongated handheld wand and / or one or more processors may be disposed in the computing device. Once the correspondence is resolved, the three-dimensional positions on the oral cavity surface are calculated in step 68 and used to generate a digital three-dimensional image of the oral cavity surface. Further, by using multiple cameras 24 to capture the oral cavity scene, an improvement in signal-to-noise in the capture by a factor of the square root of the number of cameras is achieved.

[0312] Next, refer to FIG. 8, which is a flowchart showing an overview of the corresponding algorithm for step 66 of FIG. 7 according to some application examples of the present invention. Based on the stored calibration values, all projector rays 88 and all camera rays 86 corresponding to all detection spots 33' are mapped (step 70), and all intersections 98 (FIG. 10) between at least one camera ray 86 and at least one projector ray 88 are identified (step 72). FIGS. 9 and 10 are schematic diagrams showing simplified examples of step 70 and step 72 of FIG. 8, respectively. As shown in FIG. 9, three projector rays 88 are mapped together with eight camera rays 86 corresponding to a total of eight detection spots 33' on the camera sensor 58 of the camera 24. As shown in FIG. 10, 16 intersections 98 are identified.

[0313] In steps 74 and 76 of FIG. 7, the processor 96 determines the correspondence between the projection spot 33 and the detection spot 33' in order to identify the three-dimensional position of each projection spot 33 on the surface. FIG. 11 is a schematic diagram depicting steps 74 and 76 of FIG. 8 using the simplified example described in the immediately preceding paragraph. For a given projector ray i, the processor 96 "looks at" the corresponding camera sensor path 90 on one camera sensor 58 of the camera 24. Each detection spot j along the camera sensor path 90 will have a camera ray 86 that intersects the given projector ray i at the intersection point 98. The intersection point 98 defines a three-dimensional point in space. The processor 96 then "looks at" the camera sensor paths 90' corresponding to the given projector ray i on the respective camera sensors 58' of the other cameras 24, and identifies how many other cameras 24 have similarly detected the respective spots k where their camera rays 86' intersect that same three-dimensional point in space defined by the intersection point 98. This process is repeated for all detection spots j along the camera sensor path 90, and the spot j that the most cameras 24 "agree" on is identified as the spot 33 (FIG. 12) projected onto the surface from the given projector ray i. That is, the projector ray i is identified as the specific projector ray 88 that generated the detection spot j where the most other cameras detected the respective spots k. The three-dimensional position on the surface is calculated for that spot 33. The same process may be performed to calculate the three-dimensional positions on the surface of other features of the projection pattern.

[0314] In one example, as shown in FIG. 11, all four cameras detect, on their respective camera sensor paths corresponding to the projector ray i, the respective spots where their respective camera rays intersect the projector ray i at the intersection point 98, and the intersection point 98 is defined as the intersection of the camera ray 86 corresponding to the detected spot j and the projector ray i. Thus, it can be said that all four cameras "agree" that the spot 33 projected by the projector ray i exists at the intersection point 98. However, when this process is repeated for the next spot j', none of the other cameras detect, on their respective camera sensor paths corresponding to the projector ray i, the respective spots where their respective camera rays intersect the projector ray i at the intersection point 98', and the intersection point 98' is defined as the intersection of the camera ray 86" (corresponding to the detected spot j') and the projector ray i. Thus, it can be said that only one camera "agrees" that the spot 33 (or other feature) projected by the projector ray i exists at the intersection point 98', while four cameras "agree" that the spot 33 (or other feature) projected by the projector ray i exists at the intersection point 98. Thus, the projector ray i is identified as the specific projector ray 88 that generated the detected spot j by projecting the spot 33 (or other feature) onto the surface at the intersection point 98 (FIG. 12). As in step 78 of FIG. 8 and as shown in FIG. 12, the three-dimensional position 35 on the oral cavity inner surface is calculated at the intersection point 98.

[0315] Next, refer to FIG. 13, which is a flowchart showing an overview of further steps of the corresponding algorithm according to some application examples of the present invention. When the position 35 on the surface is determined, the projector ray i that projects the spot j, as well as all the camera rays 86 and 86' corresponding to the spot j and each spot k, are excluded from consideration (step 80), and the corresponding algorithm is executed again for the next projector ray i (step 82). FIG. 14 depicts the above-described simplified example after excluding a specific projector ray i that projects the spot 33 at the position 35. As in step 82 of the flowchart of FIG. 13, the corresponding algorithm is then executed again for the next projector ray i. As shown in FIG. 14, the remaining data indicates that the three cameras are "in agreement" in that the spot 33 at the intersection 98 of the camera ray 86 corresponds to the detected spot j and the projector ray i, and the intersection 98 is defined by the intersection of the camera ray 86 corresponding to the detected spot j and the projector ray i. Therefore, as shown in FIG. 15, the three-dimensional position 37 is calculated at the intersection 98.

[0316] As shown in FIG. 16, when the three-dimensional position 37 on the surface is determined, the projector ray i that projects the spot j, as well as all the camera rays 86 and 86' corresponding to the spot j and each spot k, are removed from consideration. The remaining data shows the spot 33 projected by the projector ray i at the intersection 98, and the three-dimensional position 41 on the surface is calculated at the intersection 98. As shown in FIG. 17, according to the simplified example, the three projection spots 33 of the three projector rays 88 of the structured light projector 22 are now arranged at the three-dimensional positions 35, 37, and 41 on the surface. In some application examples, each structured light projector 22 projects 400 to 3000 spots 33. When the correspondence is solved for all the projector rays 88, a digital image of the surface can be reconstructed using the calculated three-dimensional positions of the projection spots 33 with a reconstruction algorithm.

[0317] Next, refer to FIG. 28, which is a flowchart showing an overview of the steps of a method hereinafter referred to as "spot tracking" for generating a digital three-dimensional image according to some application examples of the present invention. Although this method is called "spot tracking", this method can be similarly applied to track other types of features of the projection pattern. Due to the motion of the handheld intraoral scanner with respect to the intraoral surface during scanning, the projection points move across the intraoral surface. The inventors have noticed that if the movement of a specific detection spot (or other feature) can be tracked in consecutive image frames, the correspondence solved for that specific spot (or other feature) in any of the frames in which the spot (or other feature) was tracked gives the solution to the correspondence for that spot (or other feature) in all the frames in which the spot (or other feature) was tracked. That is, in a certain frame, when the processor 96 interprets that a given detection spot 33' was projected by a given projector ray 88, and the processor 96 determines by spot tracking that the detection spot 33' in the next image is the same spot, the processor 96 automatically determines that the same projector ray 88 was detected in the next image to generate the tracked spot.

[0318] The detection spot 33' that can be tracked across consecutive images is generated by the same specific projector ray, so the trajectory of the tracking spot will follow a specific camera sensor path 90 corresponding to that specific projector ray 88. When the correspondence is solved for a detection spot 33' at a point along the specific camera sensor path 90, the three-dimensional positions on the surface for all points along the camera sensor path 90 where that spot 33' was detected can be calculated. That is, the processor can calculate the respective three-dimensional positions on the three-dimensional surface inside the oral cavity at the intersection of the specific projector ray 88 that generated the detection spot 33' and each camera ray 86 corresponding to the tracking spot in each of the plurality of consecutive images in which the spot 33' was tracked. This can be particularly useful for situations where a specific detection spot is only seen by one camera (or a small number of cameras) in a specific image frame. If that specific detection spot was seen by other cameras 24 in previous consecutive image frames and the correspondence was solved for the specific detection spot in those previous image frames, even in an image frame where the specific detection spot is only seen by one camera 24, the processor knows which projector ray 88 generated that spot, and can determine the three-dimensional position of the spot on the three-dimensional surface inside the oral cavity.

[0319] For example, areas in the oral cavity that are difficult to reach can be imaged by only a single camera 24. In this case, if the detection spot 33' on the camera sensor 58 of the single camera 24 can be tracked through a plurality of previous consecutive images, based on the information obtained from the tracking, i.e., (a) along which camera sensor path the spot is moving, and (b) which projector ray generated the tracking spot 33', the three-dimensional position on the surface for the spot (even if it is only seen by a single camera 24) can be calculated.

[0320] In step 180 of the method outlined in FIG. 28, each structured light projector 22 is driven to project a pattern of light, which in one embodiment is a distribution 34 of discrete non - connected light spots 33, onto the three - dimensional surface within the oral cavity. In step 182, each camera 24 is driven to capture an image that includes at least a portion of the projected pattern (e.g., at least one of the spots 33). In one embodiment, in step 184, based on stored calibration values indicating (a) each camera ray 86 corresponding to each pixel on the camera sensor 58 of each camera 24 and (b) each projector ray 88 corresponding to each projection spot 33 of light from each structured light projector 22, the processor 96 is used to compare a series of images (e.g., a plurality of consecutive images) captured by each camera 24 and determine which features of the projected pattern (e.g., which of the projection spots 33) can be tracked across the plurality of images. Each feature to be tracked (e.g., spot 33s') moves along a path p of pixels corresponding to a particular camera sensor path 90 corresponding to a respective projector ray, e.g., projector ray 88. In some applications, in step 186, the processor 96 calculates the respective three - dimensional positions on the three - dimensional surface within the oral cavity of the features (e.g., spots 33s') tracked in the series of images (e.g., in each of the consecutive images).

[0321] In one embodiment, each structured light projector of one or more structured light projectors is driven to project a pattern onto the three-dimensional surface within the oral cavity. Further, each camera of one or more cameras is driven to capture a plurality of images, each image including at least a portion of the projection pattern. The projection pattern may include a plurality of projection light spots, and the portion of the projection pattern may correspond to the projection spots of the plurality of projection light spots. The processor 96 then compares a series of images captured by the one or more cameras and determines, based on the comparison of the series of images, which portions of the projection pattern can be tracked across the series of images, and constructs a three-dimensional model of the three-dimensional surface within the oral cavity based at least in part on the comparison of the series of images. In one embodiment, the processor solves a correspondence algorithm for the tracked portion of the projection pattern in at least one image of the series of images and uses the solved correspondence algorithm to handle the tracked portion of the projection pattern in at least one image of the series of images. For example, in an image of the series of images in which the correspondence algorithm has not been solved, the correspondence algorithm for the tracked portion of the projection pattern is solved, and the solution of the correspondence algorithm is used to construct the three-dimensional model. In one embodiment, the processor solves a correspondence algorithm for the tracked portion of the projection pattern based on the position of the tracked portion in each image of the entire series of images and uses the solution of the correspondence algorithm to construct the three-dimensional model. In one embodiment, the processor compares a series of images based on stored calibration values indicating (a) camera rays corresponding to each pixel on the camera sensor of each camera of the one or more cameras and (b) projector rays corresponding to each one projection light spot from each structured light projector of the one or more structured light projectors, each projector ray corresponding to the path of a respective pixel on at least one camera sensor of the camera sensor, and each tracked spot s moves along the path of the pixel corresponding to the respective projector ray r.

[0322] Next, refer to FIG. 29, which depicts some detection spots 33' (specifically, 33a' and 33b') according to some application examples of the present invention, and describes how the processor 96 can determine which set of detection spots 33' can be considered to have been tracked (step 184 in FIG. 28). Spot 33a' is the detection spot 33' in the previous image, and spot 33b' is the detection spot 33' in the current image. The processor 96 performs a search within the search radius to detect possible matches between spot 33a' and spot 33b' that are considered to be the same spot 33' tracked between the two images.

[0323] The inventors have identified three typical factors that can affect how far a spot moves between frames. 1. How far a spot moves between frames typically varies inversely with the frame rate of the camera. That is, when the frame rate of the camera is very fast, the spot appears to move a relatively small distance between successive sets of frames, and when the frame rate of the camera is slow, the spot appears to move farther between successive sets of frames. How fast the wand is moving relative to the inner surface of the oral cavity also affects how far the spot moves between frames, that is, when the movement of the wand is fast, the spot appears to move farther between successive sets of frames. 2. How far a spot moves between frames typically varies depending on the degree of inclination of the inner surface of the oral cavity being scanned, where the degree of inclination is with respect to the projector 22 and / or the camera 24. When the spot is projected onto an inclined surface and the spot is moving in the direction of the inclination, the corresponding detection spot on the camera sensor moves faster, and thus the tracking spot moves farther between successive sets of frames. 3. As seen from the camera's field of view, how far the spot moves between frames usually varies depending on the distance between the scan surface and the projector. When the surface is close to the projector, even a small movement of the projector can cause a large movement of the tracking spot on the sensor 58 of the camera 24 between successive sets of frames. In contrast, when the surface is far away, the same movement of the projector causes less movement of the tracking spot on the sensor 58 of the camera 24 between successive sets of frames. For example, if the surface is approaching an infinite distance from the projector, the movement of the projector will cause almost zero movement of the tracking spot on the sensor 58 of the camera 24.

[0324] In some applications, the processor 96 performs the search within a fixed search radius of at least 3 pixels and / or less than 10 pixels (e.g., 5 pixels). In some applications, the processor 96 calculates the search radius taking into account parameters such as the level of spot position error that can be determined during calibration. For example, the search radius may be defined as 2 * (spot position error) or 3 * (spot position error).

[0325] In the simplified example shown in FIG. 29, the spots 33a' and 33b' in each set 112 are considered to be close enough to each other such that they are regarded as the same projected spot 33 moving through the two images. That is, for each set 112, the two detected spots 33a' and 33b' are considered to be the same tracking spot 33s'. In contrast, the spots 33a' and 33b' in set 114 are too far apart to be considered tracked. The spots 33a' and 33b' in set 116 are close enough, but since more than two matches are seen, they are not considered tracked. As will be further explained below, in set 116, there are sets of tracking spots, and continuing to analyze more images can help determine which spots are actually the tracking spots.

[0326] In one embodiment, to generate a digital three-dimensional image, an intraoral scanner drives each structured light projector of one or more structured light projectors to project a pattern of light onto an intraoral three-dimensional surface. The intraoral scanner further drives each of a plurality of cameras to capture an image, the image including at least a portion of the projection pattern, and each camera of the plurality of cameras includes a camera sensor including a pixel array. The intraoral scanner further uses a processor to execute a corresponding algorithm to calculate the respective three-dimensional positions on the intraoral three-dimensional surface of a plurality of features of the projection pattern. The processor uses data from a first camera of the plurality of cameras, e.g., data from at least two cameras, to identify a candidate three-dimensional position of a given feature of the projection pattern corresponding to a particular projector ray r, and data from a second camera of the plurality of cameras, e.g., another camera that is not one of at least two cameras, is not used to identify that candidate three-dimensional position. The processor further uses the candidate three-dimensional position seen by the first camera to identify a search space on the pixel array of the second camera for searching for the feature of the projection pattern from projector ray r. When the feature of the projection pattern from projector ray r is identified within that search space, the processor uses data from the second camera to refine the candidate three-dimensional position of the feature of the projection pattern. In one embodiment, the pattern of light includes a distribution of discrete non-connected light spots, and the features of the projection pattern include projection spots from non-connected light spots. In one embodiment, the processor uses stored calibration values indicating (a) camera rays corresponding to each pixel on the camera sensor of each camera of the plurality of cameras, and (b) projector rays corresponding to each feature of the features of the projection pattern from each structured light projector of one or more structured light projectors, each projector ray corresponding to the path of a respective pixel on at least one camera sensor of the camera sensor.

[0327] Next, refer to FIG. 30, which is a flowchart showing an overview of a method for determining tracking features (e.g., tracking spot 33s', etc.) according to some application examples of the present invention. Although FIG. 30 discusses with reference to the tracking spot, it is equally applicable to other types of tracking features. In some application examples, in addition to searching for the tracking spot by monitoring the proximity of spots in consecutive images, the processor 96 may search for the tracking spot 33s' based on the parameter(s) of the detected spot 33', which is hereinafter referred to as "parametric tracking". The processor 96 determines the parameters of the detected spot 33' in the first one of the consecutive images (step 188) and in the adjacent image. Next, the processor 96 uses the determined parameters of the detected spot 33' in the two adjacent images to predict the same parameters of the spot in a later image, e.g., the next image (and subsequent images) (step 190). The processor 96 searches for a spot having substantially the predicted parameters in a later image, e.g., the next image (step 192). For example, two specific detected spots 33a' and 33b' can be determined to be derived from the same projector ray in two adjacent frames via the corresponding algorithm as described above, or via proximity tracking as described in the previous two paragraphs. When the processor 96 ascertains that the detected spots 33a' and 33b' in two adjacent frames are generated by the same projector ray 88, the processor 96 can determine the parameters of the spot and predict the parameters of the spot in a later image, e.g., the next image, based on the parameters of the spot in the two adjacent images.

[0328] In some application examples, the parameters of the spot are the size of the spot, the shape of the spot, for example, the aspect ratio of the spot, the orientation of the spot, the intensity of the spot, and / or the signal-to-noise ratio (SNR) of the spot. For example, if the determined parameter is the shape of the tracking spot 33s', the processor 96 predicts the shape of the tracking spot 33s' in a later image, for example, the next image, and based on the predicted shape of the tracking spot 33s' in the later image, determines a search space in the later image for searching for the tracking spot 33s', for example, a search space having a size and aspect ratio (e.g., within a factor of 2) based on the size and aspect ratio of the predicted shape of the tracking spot 33s'. In some application examples, the shape of the spot may refer to the aspect ratio of an elliptical spot.

[0329] Referring again to FIG. 29. In some application examples, parametric tracking may help resolve ambiguities as shown in the set of spots 116 in FIG. 29. As described above, the spots 33a' and 33b' in the set 116 are close enough to be considered tracking spots, but more than one match is found. Then, based on parametric tracking, the processor 96 may be able to determine which spot in the set 116 is actually the tracking spot.

[0330] Next, refer to FIG. 31, which is a flowchart showing an overview of a method for detecting the tracking spot 33s' in a later image according to some application examples of the present invention. FIG. 31 is also applicable when detecting other tracking features in a later image. In some application examples, based on the direction and distance that the tracking spot 33s' has moved between two images (e.g., between two consecutive images), the processor 96 determines a velocity vector of the tracking spot 33s' (step 194). Next, the processor 96 uses the velocity vector to determine a search space in a later image, for example, the next image, for searching for the tracking spot 33s' (step 196).

[0331] In some application examples, the search space in subsequent images may be determined by using a prediction filter, such as a Kalman filter, to estimate the new position of the tracking spot 33s'.

[0332] Next, refer to FIG. 32, which is a flowchart showing an overview of a method for detecting the tracking spot 33s' in subsequent images according to some application examples of the present invention. FIG. 32 is also applicable when detecting other tracking features in subsequent images. The determination of the velocity vector for the tracking spot 33s' may be used to assist in determining the search space for searching for the tracking spot. That is, if the spot is moving faster, it has moved farther between consecutive frames. Therefore, the processor 96 can set a larger search space for searching for the tracking spot. The inventors have noticed that the shape of the tracking spot 33s' and the direction in which it is moving can suggest the speed of the spot. For example, in one application example, when a spot projected as a circle appears elliptical, the spot is likely to be descending on a surface inclined with respect to the projector 22 and / or the camera 24. Similarly, as described above, a spot moving in the direction of inclination along an inclined surface moves faster than when it is not moving in the direction of inclination, and the steeper the inclination of the surface, the faster the spot moves. Further, when the spot appears stretched into an ellipse by the inclined surface, it is likely to appear stretched in the direction of inclination, that is, the major axis of the ellipse appears to be in the direction of inclination. Therefore, the movement of the elliptical spot along its major axis indicates that the spot is moving up or down and is inclined. Therefore, it indicates that it is faster than when the elliptical spot moves along its minor axis (although projected on the inclined surface, it can indicate that the spot is not moving in the direction of inclination).

[0333] Accordingly, in some applications, after determining the shape of the tracking spot 33s' (step 198), the processor 96 can determine the velocity vector of the tracking spot 33s' based on the direction and distance that the tracking spot 33s' has moved between two consecutive images (step 200). The processor 96 can then use the determined velocity vector and / or the shape of the tracking spot 33s' to predict the shape of the tracking spot 33s' in a later image, e.g., the next image (step 202). Following the prediction of the shape of the tracking spot 33s', the processor 96 can use the combination of the velocity vector and the predicted shape of the tracking spot 33s' to determine a search space in a later image, e.g., the next image, in which to search for the tracking spot 33s'. Referring again to the above example of an elliptical spot, if the shape of the spot is determined to be elliptical and the spot is determined to be moving along its major axis, a larger search space will be specified compared to the case where the elliptical spot is moving along its minor axis.

[0334] Next, refer to FIG. 33, which is a schematic diagram showing an example in which spot tracking according to some application examples of the present invention helps to identify the detected spot 33' as being projected from a specific projector ray 88. This is useful when the corresponding algorithm does not present a solution for a specific detected spot 33', for example, when the detected spot 33' in a specific frame is only seen by one camera 24. In such a case, when it is identified that the detected spot 33' is a tracking spot 33s' that moves along a specific camera sensor path 90 of pixels corresponding to a specific projector ray 88, it can be assumed, as described above, that the specific projector ray 88 projected the spot. Based on the resolved correspondence of the tracking spot 33s' in the previous frame, the correspondence can be resolved for the frame in which only one camera detected the spot 33'. Thus, in some application examples, after executing a correspondence algorithm such as the correspondence algorithm described above in relation to FIGS. 7 to 17, if it is determined that the detected spot 33' is a tracking spot 33s' that moves along a specific path 90 on the camera sensor 58 corresponding to a specific projector ray 88, it can be assumed that the specific projector ray 88 generated the detected spot 33'. That is, the processor 96 can identify that the detected spot 33' is derived from a specific projector ray 88 by identifying the detected spot 33' as a tracking spot 33s' that moves along the path 90 of the pixels of the camera sensor 58 corresponding to the specific projector ray 88.

[0335] In the example shown in FIG. 33, in two consecutive frames respectively photographed at time 1 and time 2, each of the two camera sensors 58 detected the projection spot 33' from the projector light beam 88. The above-mentioned correspondence algorithm determined that the projector light beam 88 generated the detection spots 33' in Frame 1 and Frame 2, and solved the correspondence relationship of the detection spots 33' in Frame 1 and Frame 2. However, in the third frame, only one camera detected the spot 33'. In this example, it is assumed that the correspondence algorithm could not solve the correspondence relationship for the detection spot 33' in Frame 3. However, the processor 96 determines that the detection spot 33' in Frame 3 is the tracking spot 33s' that moves along the same projector light beam 88 that generated the spots 33' in Frame 1 and Frame 2. Therefore, the processor 96 identifies the detection spot 33' in Frame 3 as being generated by the projector light beam 88.

[0336] Next, refer to FIGS. 34A - B, which are schematic diagrams of the camera sensor 58 showing two detection spots 33c' and 33d' according to some application examples of the present invention. In FIG. 34A, for each of the detection spots, there is ambiguity as to which projector light beam generated the spot, that is, which pixel of which camera sensor path 90 the detection spot corresponds to. In some application examples, the processor 96 can use spot tracking to resolve these ambiguities. Therefore, in some application examples, after executing the correspondence algorithm (such as the correspondence algorithm described above with respect to FIGS. 7 - 17), when the detection spot 33' is identified as being derived from two different candidate projector light beams 88 and 88' based on the three-dimensional position calculated by the correspondence algorithm, the processor 96 may identify that the detection spot 33' is the tracking spot 33s' that moves along either path 90 or path 90', thereby identifying that it is derived from only one of the two different candidate projector light beams 88 and 88'.

[0337] Examples of such ambiguities are represented by the detection spot 33c' in FIG. 34A. Spot 33c' is located at the intersection of two different paths 90c and 90c'. The corresponding algorithm may know that such a detection spot 33c' is generated by both the projector ray 88 corresponding to path 90c and the projector ray 88' corresponding to path 90c'. Another type of such ambiguity is represented by the detection spot 33d' in FIG. 34A. Spot 33d' is very close to two different paths 90d and 90d', but is not at the intersection between paths 90 and 90d. Due to signal noise, it may have been unclear in the corresponding algorithm whether spot 33d' was generated by the projector ray 88 corresponding to path 90d or the projector ray 88' corresponding to path 90d'.

[0338] As shown in FIG. 34B, the processor 96 may identify which projector ray generated each of the spots 33c' and 33d' by identifying spots 33c' and 33d' as tracking spots 33s' that each move along a particular one of the paths. Detection spot 33c' is identified as a tracking spot 33s' that moves along path 90c', and thus detection spot 33c' is identified as being generated by the projector ray 88' corresponding to path 90c'. Detection spot 33d' is identified as a tracking spot 33s' that moves along path 90d', and thus detection spot 33d' is identified as being generated by the projector ray 88' corresponding to path 90d'.

[0339] Next, refer to FIG. 35, which is a flowchart showing an overview of additional or alternative methods by which spot tracking can be used according to some application examples of the present invention. The concepts shown in FIG. 35 are also applicable to methods using other feature tracking. In some application examples, the processor 96 may be able to use spot tracking to exclude misdetected spots 33' from consideration as points on the intraoral three-dimensional surface. In some application examples, after executing a corresponding algorithm (step 206), such as the corresponding algorithm described above with reference to FIGS. 7 to 17, the processor 96 may identify the detected spot 33' as being derived from a specific projector ray 88 based on the corresponding algorithm (step 208). Step 206 typically occurs following step 186 of the method outlined in the flowchart of FIG. 28. Further, the processor 96 may identify a series of spots detected over a plurality of consecutive images that are tracking spots 33s' all moving along a path 90 of pixels corresponding to the same specific projector ray 88. As indicated by decision diamond 210, if the detected spot 33' is one of the tracking spots 33s', the detected spot 33' may be considered a point on the intraoral three-dimensional surface (step 212). However, if the detected spot 33' is not identified as being a tracking spot 33s' moving along the path 90 of pixels corresponding to that specific projector ray 88, the detected spot 33' may be assumed to be a misdetected spot, and the detected spot 33' may be excluded from consideration as a point on the intraoral three-dimensional surface (step 214).

[0340] Next, refer to FIG. 36, which is a flowchart showing an overview of additional or alternative methods by which spot tracking can be used according to some application examples of the present invention. The concepts shown in FIG. 36 are also applicable to methods using other feature tracking. To reduce the occurrence of detection of a large number of false detection spots in camera 24, processor 96 may set an intensity threshold, and any detection spot 33' below the threshold is not included as a candidate spot in the corresponding algorithm. However, this may also result in the situation that false detection spots, i.e., spots that may have provided useful information, are not considered because it is a vulnerable spot (having an intensity below the threshold). For example, in a difficult-to-capture area of an intraoral scene, a part of projection spot 33 may be displayed below the intensity threshold. Therefore, in some application examples, after executing a corresponding algorithm (step 216), such as the corresponding algorithm described herein with reference to FIGS. 7 to 17, processor 96 may identify vulnerable spots 33' whose three-dimensional positions were not calculated by the corresponding algorithm (step 218), for example, by lowering the intensity threshold and considering spots 33' that were not considered by the corresponding algorithm. As shown by decision diamond 220, if the vulnerable spot 33' is identified as a tracking spot 33s' that moves along path 90 of pixels corresponding to a specific projector ray 88, the vulnerable spot 33' is identified as being projected from that specific projector ray 88 and considered to be a point on the intraoral three-dimensional surface (step 222). If the vulnerable spot is not identified as a tracking spot 33s', the vulnerable spot is excluded from being considered as a point on the intraoral three-dimensional surface (step 224).

[0341] In some application examples, for the tracking spot 33s’, the processor 96 may determine a plurality of possible camera sensor paths 90 of the pixels where the tracking spot 33s’ is moving, and the plurality of paths 90 correspond to the plurality of possible projector light rays 88 respectively. For example, the plurality of projector light rays 88 may closely correspond to the paths 90 of the pixels on the camera sensor 58 of a given camera. The processor 96 may execute a corresponding algorithm to identify which of the possible projector light rays 88 generated the tracking spot 33s’ in order to calculate the three-dimensional position on the surface for each position of the tracking spot 33s’.

[0342] For each of the plurality of possible projector light rays 88 for a given camera sensor 58, the three-dimensional point in space exists at the intersection of each of the possible projector light rays 88 and the camera ray corresponding to the tracking spot 33s’ detected in the given camera sensor 58. For each of the possible projector light rays 88, the processor 96 considers the camera sensor paths 90 corresponding to the possible projector light rays 88 in each of the other camera sensors 58, and determines how many other camera sensors 58 whose camera rays intersect that three-dimensional point in space detected the spot 33’ on their respective camera sensor paths 90 corresponding to that possible projector light ray 88 as well, that is, how many other cameras agree that the tracking spot 33s’ is projected by that projector light ray 88. This process is repeated for all possible projector light rays 88 corresponding to the tracking spot 33s’. The possible projector light ray 88 that most other cameras agree on is determined to be the specific projector light ray 88 that generated the tracking spot 33s’. When the specific projector light ray 88 for the tracking spot 33s’ is determined, the camera sensor path 90 where the spot is moving is identified, and in each of the consecutive images where the spot 33s’ is tracked, the respective three-dimensional positions on the surface at the intersection of the specific projector light ray 88 corresponding to the tracking spot 33s’ and each camera ray are calculated.

[0343] Next, refer to FIGS. 37A - B, which are schematic diagrams showing points used for 3D reconstruction before and after the processor 96 performs spot tracking according to some application examples of the present invention. The size of each data point represents how many cameras were used to solve that point, that is, the larger the point, the more cameras saw that spot. Note that the size of the data point is used in the figure only to distinguish how many cameras saw any point and does not indicate the size of the projected spot on the surface. In FIG. 37A, there are many small points located at the periphery and not seemingly points on the oral cavity inner surface. These small points refer to detected spots whose 3D positions in space were assigned based on the corresponding algorithm, even though they were seen by a very small number of cameras, for example, only one. After executing the corresponding algorithm, the processor 96 may execute spot tracking and thus may determine that these light points at the periphery are actually misdetected points (by determining that they are not tracking spots). Therefore, as shown in FIG. 37B, after spot tracking, most of the spots determined to be misdetected spots by spot tracking are removed because they are considered to be points on the oral cavity inner surface.

[0344] Next, refer to FIG. 38, which is a flowchart showing an overview of the steps of a method for generating a digital 3D image, referred to herein as "ray tracing" according to some application examples of the present invention. In some application examples, instead of or additionally to tracking detected spots and / or other features in a 2D image (as described above), the length of each projector ray 88 can be tracked in 3D space. The length of the projector ray 88 is defined as the distance between the origin of the projector ray 88, i.e., the light source, and the 3D position where the projector ray 88 intersects the oral cavity inner surface.

[0345] In step 226 of the method outlined in FIG. 38, each structured light projector 22 is driven to project a pattern of light, such as a distribution 34 of discrete, unconnected light spots 33, onto the oral cavity 3D surface, and in step 228 each camera 24 is driven to capture a plurality of images, each image including at least one feature of the projected pattern (e.g., at least one of the spots 33). Although the method is described with reference to spots, other types of features will also function. In step 230, using processor 96, a correspondence algorithm, such as the correspondence algorithm described above with reference to FIGS. 7-17, is executed based on stored calibration values indicative of (a) camera rays 86 corresponding to each pixel on camera sensor 58 of each camera 24, and (b) projector rays 88 corresponding to each projection spot 33 of light from each structured light projector 22. As a result of the correspondence algorithm, each resolved projector ray 88 in each image frame yields a reconstructed spatial 3D point, which in turn defines the length of the resolved projector ray 88 in that frame.

[0346] Accordingly, at step 232, in at least a subset of the captured images, for example, in a series of images or a plurality of consecutive images, the processor 96 identifies that the calculated three-dimensional position of the detection spot 33' (calculated from the corresponding algorithm) corresponds to a particular projector ray 88. At step 234, based on each three-dimensional position corresponding to the projector ray 88 in the subset of images, the processor 96 evaluates, for example, calculates the length of the projector ray 88 in each image of the subset of images. Since the camera 24 captures images at a relatively high frame rate, for example, about 100 Hz, the geometric shape of the spots seen by each camera does not vary significantly between frames. Thus, when the evaluated, for example, calculated, length of the projector ray 88 is tracked and plotted against time, the data points will follow a relatively smooth curve, although some discontinuities may occur as further discussed below. Accordingly, the length of the projector ray over time forms a relatively smooth univariate function against time. As described above, the detection spots 33' corresponding to the projector ray 88 over a plurality of consecutive images will appear to move along a one-dimensional line, which is the pixel path 90 of the camera sensor corresponding to the projector ray 88.

[0347] In one embodiment, a method for generating a digital three-dimensional image includes driving each structured light projector of one or more structured light projectors to project a pattern onto a three-dimensional surface within the oral cavity, and driving each camera of one or more cameras to capture an image, the image including at least a portion of the pattern. The method further includes using the processor to execute a corresponding algorithm to calculate the three-dimensional position of each of a plurality of features of the pattern on the three-dimensional surface within the oral cavity captured in a series of images. The processor further identifies, in at least a subset of the series of images, the calculated three-dimensional positions of the detected features of the imaged pattern as corresponding to one or more specific projector rays r. Based on the three-dimensional positions of the detected features corresponding to the one or more projector rays r in the subset of the images, the processor evaluates, e.g., calculates, a length associated with the one or more projector rays r in each image of the subset of the images. In one embodiment, the processor calculates an estimated length of one or more projector rays r in at least one of the series of images in which the three-dimensional positions of the features projected from the one or more projector rays r are not identified. In one embodiment, each camera of the one or more cameras includes a camera sensor including a pixel array, and the calculation of the three-dimensional position of each of a plurality of features of the pattern on the three-dimensional surface within the oral cavity and the identification that the calculated three-dimensional positions of the detected features of the pattern correspond to a specific projector ray r are performed based on stored calibration values indicating (i) camera rays corresponding to each pixel of the camera sensor of each of the one or more cameras, and (ii) projector rays corresponding to each feature of the projected light pattern from each projector of the one or more projectors, each projector ray corresponding to a path of a respective pixel on at least one of the camera sensors. In one embodiment, the pattern includes a plurality of spots, and each of the plurality of features of the pattern includes one of the plurality of spots.

[0348] Next, refer to FIG. 39, which is a graph showing the tracking of the length of the projector light beam 88 over time and a specific schematic diagram of the camera sensor 58 corresponding to a specific image frame according to some application examples of the present invention. The inventors have realized a plurality of usage examples of the above-described light beam tracking. In some application examples, there may be at least one image in which the three-dimensional position of the projection spot 33 (or other feature) from a specific projector light beam 88 is not specified in step 232 of the method shown in FIG. 38 from a plurality of consecutive images. For example, the projection spot (or other feature) in a specific frame may be below the intensity threshold and may not be considered by the corresponding algorithm, or there may be a false detection of the spot in a specific frame. However, due to the tracking of the length of the projector light beam over time, the processor 96 may calculate the estimated length of the specific projector light beam 88 in that image.

[0349] For example, in the exemplary graph shown in FIG. 39, for the scan frame s1 captured at time t1, the three-dimensional position of the projection spot 33 is not determined based on the corresponding algorithm, and thus, as illustrated by the dashed circle 236, there is no data point corresponding to the length of the projector ray 88 for the scan frame s1. However, due to the length of the ray being tracked through a plurality of consecutive images, the estimated length L1 of the projector ray 88 can be calculated for the scan frame s1, for example, by interpolation. As described above, all projection points by a particular projector ray 88 appear on a particular path 90 of the pixels of the camera sensor 58 corresponding to that particular projector ray 88. Therefore, for the scan frame s1 in which the three-dimensional position of the spot 33 corresponding to a particular projector ray 88 is not determined in step 232, the processor 96 may determine a one-dimensional search space 238 within the scan frame s1 for searching for the projection spot from that particular projector ray 88. The one-dimensional search space 238 is along the path 90 of each pixel corresponding to the particular projector ray 88. This is in contrast to the spot tracking algorithm described above, in which the processor 96 searches two-dimensionally within the image for spots that are close enough to each other from one frame to the next, considering that they are tracking spots generated by the same projector ray in each of the image frames.

[0350] In some application examples, based on the estimated length L1 of the projector ray 88 in at least one of the plurality of images, the processor 96 may determine a one-dimensional search space in each of the plurality of cameras 24, for example, in the pixel arrays of all the cameras 24, for example, in the camera sensor 58. For each of the respective pixel arrays, the one-dimensional search space is along the respective pixel paths 90 corresponding to the projector ray 88 in that particular pixel array, for example, in the camera sensor 58. The length L1 of the projector ray 88 corresponds to a three-dimensional point in space, which corresponds to a two-dimensional position on the camera sensor 58. All the other cameras 24 also have respective two-dimensional positions on the camera sensor 58 corresponding to the same three-dimensional space point. Thus, using the length of the projector ray in a particular frame, for that particular frame, a one-dimensional search space may be defined in a plurality of camera sensors 58, for example, in all the camera sensors 58.

[0351] In some application examples, in contrast to a false detection where an expected spot (or other feature) was not detected, there may be at least one image among a plurality of consecutive images in which a plurality of candidate 3D positions were calculated for a projection spot 33 (or other feature) from a specific projector ray 88, that is, there may have been a false detection of the projection spot 33 (or other feature). For example, in the exemplary graph shown in FIG. 39, based on the corresponding algorithm, for the scan frame s2 taken at time t2, two candidate detection spots 33' and 33" were both calculated to be derived from a specific projector ray 88, and thus two candidate 3D positions of the projection spot 33 were calculated. The processor 96 calculated two candidate lengths of the projector ray 88 corresponding to each of the candidate 3D positions for that frame. For example, for the candidate detection spot 33', the processor 96 calculated that the candidate length of the projector ray 88 is L2 (represented by the data point 242 in FIG. 39), and for the candidate detection spot 33", the processor 96 calculated that the candidate length of the projector ray 88 is L3 (represented by the data point 244 in FIG. 39). Due to the length of the projector ray 88 being tracked across a plurality of consecutive images, when the ray length data of the scan frame s2 is added, it becomes clear which of the candidate lengths L2 or L3 is the estimated length of the projector ray 88 of the scan frame s2. In this way, the processor 96 can determine which of the correct 3D positions of the projection spot 33 is by determining which of the plurality of candidate 3D positions of the projection spot 33 corresponds to the estimated length of the projector ray 88 for that image.

[0352] Based on the estimated length of the projector ray 88 in at least one of a plurality of images, e.g., the scan frame s2, the processor 96 may determine a one-dimensional search space 246 in the scan frame s2. Thereafter, the processor 96 determines which of the plurality of candidate 3D positions of the projection spot 33' corresponds to the spot 33' generated by the projector ray 88 and is detected within the one-dimensional search space 246, thereby determining which of the plurality of candidate 3D positions of the projection spot 33 is the correct 3D position of the projection spot 33 generated by the projector ray 88. Before additional information is provided by ray tracing, the camera sensor 58 for the scan frame s2 should have shown both of the two candidate detection spots 33' and 33" on the path 90 of the pixels corresponding to the projector ray 88. The processor 96 calculates the estimated length of the projector ray 88 based on the length of the ray being tracked across a plurality of consecutive images, enabling the processor 96 to determine the one-dimensional search space 246 and determine that the candidate detection spot 33' is the actual correct spot. Then, the candidate detection spot 33" is excluded since it is considered to be a point on the 3D intraoral surface.

[0353] In some applications, the processor 96 may define a curve 248 based on an evaluated, e.g., calculated, length of the projector ray 88 in a subset of the images, e.g., each of a plurality of consecutive images. Under the assumption of the inventors, any detection point corresponding to a length of the projector ray r that is at least a threshold distance away from the defined curve 248 for a 3D position based on the corresponding algorithm is considered to be a false detection and can reasonably be assumed to be excluded since it is considered to be a point on the 3D intraoral surface.

[0354] Next, refer to FIGS. 40A - B, which are graphs showing experimental sets of data before and after ray tracing according to some application examples of the present invention. In FIG. 40A, the lengths of specific projector rays 88 are plotted for all spots and / or other features calculated by the corresponding algorithm. Before ray tracing is applied, the ray lengths corresponding to spots and / or other features that appear far from the general curve defined by the projector ray 88 are included in the data. FIG. 40B represents the aspect of the data after ray tracing is applied and used to determine which false - detected spots and / or other features should be removed as being considered points on the three - dimensional intra - oral surface.

[0355] Next, refer to FIG. 41, which is a schematic diagram showing a plurality of camera sensors and a projector that projects spots or other features according to some application examples of the present invention. In some application examples, after executing a corresponding algorithm such as the corresponding algorithm described above with reference to FIGS. 7 - 17, depending on how many cameras 24 detect a given projection spot 33 (or other feature) on their respective pixel arrays (i.e., camera sensors 58), the processor 96 can determine the candidate three - dimensional position of the projection spot 33 (or other feature) with a certain degree of certainty. The candidate three - dimensional position of the projection spot 33 (or other feature) in FIG. 41 is indicated by the dashed circle 250.

[0356] The higher the number of cameras 24 that view the projection spot 33 (or other feature), the higher the certainty of the three - dimensional position candidate 250. Thus, using data from at least two of the cameras 24, the processor 96 can identify the candidate three - dimensional position 250 of a given spot 33 (or other feature) corresponding to a specific projector ray 88. Assuming that the identification of the candidate three - dimensional position 250 is substantially determined without using data from at least another camera 24', there may be some error in the candidate three - dimensional position 250, and if the processor 96 has data from another camera 24', the candidate three - dimensional position 250 can be refined.

[0357] Thus, after obtaining the correspondence, assuming that at least two cameras 24 saw the projection spot 33 (or other feature), at this point the processor 96 knows (a) which projector ray 88 generated the projection spot 33 (or other feature), and (b) the candidate three-dimensional position 250 of the spot (or other feature). Combining (a) and (b), the processor 96 can determine a one-dimensional search space 252 of a pixel array of another camera 24', i.e., the camera sensor 58', to search for the spot (or other feature) from the projector ray 88. The one-dimensional search space 252 is along a path 90 of pixels on the camera sensor 58' of the other camera 24', and may be along a specific segment of the path 90 corresponding to the candidate three-dimensional position 250. If a spot 33' (or other feature) derived from the projector ray 88, e.g., a false detection spot 33' not considered by the correspondence algorithm (e.g., due to intensity below a threshold), is identified within the one-dimensional search space 252, the candidate three-dimensional position 250 of the spot 33 (or other feature) may be refined to a refined three-dimensional position 254 using the data of the other...

Claims

1. A method for generating a digital three-dimensional image, comprising: driving each structured light projector of one or more structured light projectors to project a pattern onto a three-dimensional surface within the oral cavity; driving each camera of one or more cameras to capture a plurality of images, each image including at least a portion of the projected pattern; using a processor to compare a series of images captured by the one or more cameras; determine, based on the comparison of the series of images, which portions of the projected pattern can be tracked across the series of images; construct a three-dimensional model of the three-dimensional surface within the oral cavity, based at least in part on the comparison of the series of images; A method comprising the above steps.

2. The step of using the processor further comprises using the processor to solve a corresponding algorithm for the tracked portion of the projected pattern in at least one image of the series of images; using the solved corresponding algorithm for the tracked portion in at least one image of the series of images to solve a corresponding algorithm for the tracked portion of the projected pattern in at least another image of the series of images, and constructing the three-dimensional model using the solution of the corresponding algorithm; The method according to claim 1, comprising the above steps.

3. The step of using the processor further comprises using the processor to solve a corresponding algorithm for the tracked portion of the projected pattern based on the positions of the tracked portions in each image throughout the series of images, and the step of constructing the three-dimensional model comprises constructing the three-dimensional model using the solution of the corresponding algorithm. The method according to claim 1, comprising the above steps.

4. The method according to claim 1, wherein the one or more structured light projectors project a spatially fixed pattern with respect to the one or more cameras.

5. The method according to any one of claims 1 to 4, wherein the projected pattern includes a plurality of projected light spots, and the portion of the projected pattern corresponds to one projected light spot s of the plurality of projected light spots.

6. The step of comparing the series of images using the processor comprises using the processor to comparing the series of images based on stored calibration values indicating (a) camera rays corresponding to respective pixels on the camera sensors of the one or more cameras and (b) projector rays corresponding to respective projection light spots of the projection light spots from the one or more structured light projectors, each projector ray corresponding to the path of a respective pixel on at least one of the camera sensors; The step of determining which portions of the projection pattern can be tracked includes determining which of the projection spots s can be tracked over the series of images, each tracking spot s moving along the path of a pixel corresponding to a respective projector ray r. The method according to claim 5.

7. The step of using the processor further includes using the processor to determine, for each tracking spot s, a path p of a plurality of possible pixels on a given one of the cameras, the path p corresponding to a respective plurality of possible projector rays r. The method according to claim 6.

8. The step of using the processor further includes using the processor to execute a corresponding algorithm to for each of the possible projector rays r identifying how many other cameras detected, on their respective paths p1 of pixels corresponding to the projector ray r, respective spots q corresponding to the camera rays that intersect the projector ray r and the camera rays of the given one of the cameras corresponding to the tracking spot s; identifying a given projector ray r1 where the most other cameras detected the respective spots q; identifying the projector ray r1 as the particular projector ray r that generated the tracking spot s. The method according to claim 7.

9. The step of using the processor further includes using the processor to execute a corresponding algorithm to calculate the three-dimensional position of each of the plurality of detection spots on the three-dimensional surface within the oral cavity captured in the series of images. The method according to claim 6, comprising the step of identifying, in at least one image of the series of images, that a detection spot is derived from a specific projector ray r by identifying the detection spot as a tracking spot s that moves along the path of the pixel corresponding to the specific projector ray r.

10. The step of using the processor further comprises using the processor to execute a corresponding algorithm to calculate the three-dimensional position of each of a plurality of detection spots on the three-dimensional surface within the oral cavity captured in the series of images; and excluding, as points on the three-dimensional surface within the oral cavity, spots that are identified as being derived from a specific projector ray r based on the three-dimensional positions calculated by the corresponding algorithm and that are not identified as tracking spots s that move along the path of the pixel corresponding to the specific projector ray r. The method according to claim 6.

11. The step of using the processor further comprises using the processor to execute a corresponding algorithm to calculate the three-dimensional position of each of a plurality of detection spots on the three-dimensional surface within the oral cavity captured in the series of images; and identifying, based on the three-dimensional positions calculated by the corresponding algorithm, for detection spots identified as being derived from two different projector rays r, the detection spots as tracking spots s that move along one of the two different projector rays r, thereby identifying that the detection spots are derived from one of the two different projector rays r. The method according to claim 6.

12. The step of using the processor further comprises using the processor to execute a corresponding algorithm to calculate the three-dimensional position of each of a plurality of detection spots on the three-dimensional surface within the oral cavity captured in the series of images; and identifying, as projection spots derived from a specific projector ray r, vulnerable spots for which the three-dimensional position was not calculated by the corresponding algorithm by identifying the vulnerable spots as tracking spots s that move along the path of the pixel corresponding to the specific projector ray r. The method according to claim 6.

13. The step of using the processor further includes using the processor to calculate, at the intersection of the projector ray r and each of the camera rays corresponding to the tracking spot s in each of the series of images in which the spot s is tracked, each three-dimensional position on the three-dimensional surface within the oral cavity. The method according to claim 6.

14. The three-dimensional model is constructed using a corresponding algorithm, and the corresponding algorithm uses at least in part the portion of the projection pattern determined to be trackable across the series of images. The method according to any one of claims 1 to 4.

15. The step of using the processor further includes using the processor to determine parameters of the tracked portion of the projection pattern in at least two adjacent images from the series of images, the parameters being selected from the group consisting of the size of the portion, the shape of the portion, the orientation of the portion, the intensity of the portion, and the signal-to-noise ratio (SNR) of the portion; predicting parameters of the tracked portion of the projection pattern in a subsequent image based on the parameters of the tracked portion of the projection pattern in the at least two adjacent images. The method according to any one of claims 1 to 4.

16. The step of using the processor further includes using the processor to search for a portion of the projection pattern having substantially the predicted parameters in the subsequent image based on the predicted parameters of the tracked portion of the projection pattern. The method according to claim 15.

17. The selected parameter is the shape of the portion of the projection pattern, and the step of using the processor further includes using the processor to determine a search space in a next image for searching for the tracked portion of the projection pattern based on the predicted shape of the tracked portion of the projection pattern. The method according to claim 15.

18. The step of determining the search space using the processor includes the step of determining, using the processor, a search space in a next image for searching a tracking portion of the projection pattern, the search space having a size and an aspect ratio based on a size and an aspect ratio of a predicted shape of the tracking portion of the projection pattern, the method according to claim 17.

19. The selected parameter is the shape of the portion of the projection pattern, and the step of using the processor includes using the processor to determine a velocity vector of the tracking portion of the projection pattern based on a direction and a distance by which the tracking portion of the projection pattern has moved between the at least two adjacent images from the series of images; predict a shape of the tracking portion of the projection pattern in a later image in response to the shape of the tracking portion of the projection pattern in at least one of the at least two adjacent images; The method according to claim 15, comprising determining a search space in the later image for searching the tracking portion of the projection pattern in response to a combination of (i) the determination of the velocity vector of the tracking portion of the projection pattern and (ii) the predicted shape of the tracking portion of the projection pattern.

20. The parameter is the shape of the portion of the projection pattern, and the step of using the processor further includes using the processor to determine a velocity vector of the tracking portion of the projection pattern based on a direction and a distance by which the tracking portion of the projection pattern has moved between the at least two adjacent images from the series of images; predict the shape of the tracking portion of the projection pattern in a later image in response to the determination of the velocity vector of the tracking portion of the projection pattern; The method according to claim 15, comprising determining a search space in the later image for searching the tracking portion of the projection pattern in response to a combination of (i) the determination of the velocity vector of the tracking portion of the projection pattern and (ii) the predicted shape of the tracking portion of the projection pattern.

21. The step of using the processor includes using the processor to The method according to claim 20, comprising the step of predicting the shape of the tracking portion of the projection pattern in the subsequent image in response to a combination of (i) the determination of the velocity vector of the tracking portion of the projection pattern and (ii) the shape of the tracking portion of the projection pattern in at least one of the two adjacent images. **Claim 22** The step of using the processor comprises using the processor to determine a velocity vector of the tracking portion of the projection pattern based on a direction and a distance by which the tracking portion of the projection pattern has moved between two consecutive images in the series of images; and determining a search space in a subsequent image for searching for the tracking portion of the projection pattern in response to the determination of the velocity vector of the tracking portion of the projection pattern. The method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Device for determining three-dimensional coordinate of object, tooth in particular

    JP2010069301A

  • Apparatus and method for non-contact detection of three-dimensional contours

    JP2010507079A

  • Shape measuring instrument and shape measuring method

    JP2011242178A

  • Intraoral three-dimensional measuring device, intraoral three-dimensional measuring method, and method for displaying intraoral three-dimensional measuring result

    JP2017020930A

  • An intraoral 3D scanner that uses multiple small cameras and multiple small pattern projectors

    JP7661250B2