Intraoral 3D scanner employing multiple miniature cameras and multiple miniature pattern projectors

By using multiple miniature cameras and pattern projectors in a dental digital scanner, combined with diffraction and refraction optical elements and calibration algorithms, the problem of capturing structured light patterns on highly reflective tooth surfaces has been solved, achieving efficient 3D imaging and accurate dental scanning.

CN114302672BActive Publication Date: 2026-01-13ALIGN TECHNOLOGY INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080060018.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-06-23
Filing Date
2020-06-24
Publication Date
2026-01-13
Estimated Expiration
2040-06-24

AI Technical Summary

Technical Problem

Existing dental digital scanners have difficulty effectively capturing structured light patterns on highly reflective and translucent tooth surfaces, resulting in reduced contrast and difficulties in correspondence, which affects the accuracy of 3D imaging.

Method used

By employing multiple miniature cameras and miniature pattern projectors, combined with diffraction and/or refraction pattern generating optical elements, discrete and unconnected light spot distributions are generated. Laser diodes are used as the light source, and calibration algorithms and neural network optimization processors are used to solve the correspondence problem, thereby improving the accuracy of image stitching and 3D reconstruction.

Benefits of technology

It improves the accuracy and scanning speed of three-dimensional imaging of tooth and gum surfaces, reduces heat buildup in the device, simplifies the calibration process, reduces the need for active cooling, and enhances the convenience and cost-effectiveness of the device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114302672B_ABST
    Figure CN114302672B_ABST
Patent Text Reader

Abstract

A method for generating a 3D image includes driving a structured light projector to project a light pattern on an intraoral 3D surface and driving a camera to capture images, each image including at least a portion of the projected pattern, each camera including an array of pixels. A processor compares a series of images captured by each camera and determines which portions of the projected pattern can be tracked across the images. The processor constructs a three-dimensional model of the intraoral three-dimensional surface based at least in part on the comparison of the series of images. Other embodiments are also described.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention generally relates to three-dimensional imaging, and more specifically to intraoral three-dimensional imaging using structured light illumination. Background Technology

[0002] Dental impressions of a subject's three-dimensional intraoral surfaces (e.g., teeth and gums) are used to plan dental procedures. Traditional dental impressions are made using dental impression trays filled with impression material (e.g., PVS or alginate) into which the subject bites. The impression material then cures to form a negative impression of the teeth and gums, from which a three-dimensional model of the teeth and gums can be formed.

[0003] Digital dental impressions utilize intraoral scanning to generate a three-dimensional digital model of the subject's intraoral surface. Digital intraoral scanners typically employ structured light 3D imaging. The surface of a subject's teeth can be highly reflective and somewhat translucent, which may reduce the contrast of the structured light pattern reflected from the teeth. Therefore, to improve the capture rate of intraoral scans, when using digital intraoral scanners utilizing structured light 3D imaging, the subject's teeth are often coated with an opaque powder before scanning to enhance the usable level of contrast in the structured light pattern, for example, to make the surface a scattering surface. While intraoral scanners utilizing structured light 3D imaging have made some progress, there may be additional advantages. Summary of the Invention

[0004] The use of structured light 3D imaging can lead to a "correspondence problem," where it's necessary to determine the correspondence between points in a structured light pattern and points seen by a camera observing that pattern. One technique for addressing this problem is to project an "encoded" light pattern and image the illuminated scene from one or more viewpoints. Encoding the emitted light pattern makes portions of the pattern unique and distinguishable when captured by the camera system. Because the pattern is encoded, it's easier to find the correspondence between points in the image and points in the projected pattern. Triangulation and 3D information recovery can then be performed on the decoded points.

[0005] Applications of this invention include systems and methods related to a three-dimensional intraoral scanning apparatus, which includes one or more cameras and one or more pattern projectors. For example, some applications of this invention may relate to an intraoral scanning apparatus having multiple cameras and multiple pattern projectors.

[0006] Other applications of the present invention include methods and systems for decoding structured light patterns.

[0007] Other applications of the present invention may relate to systems and methods for three-dimensional intraoral scanning utilizing non-coded structured light patterns.

[0008] For example, in certain applications of the invention, an apparatus for intraoral scanning is provided, comprising an elongated handheld rod with a probe at its distal end. During scanning, the probe can be configured to enter the oral cavity of a subject. One or more miniature structured light projectors and one or more miniature cameras are coupled to a rigid structure disposed within the distal end of the probe. Each structured light projector uses a light source such as a laser diode to transmit light. In some applications, the structured light projector may have an illumination field of at least 45 degrees. Optionally, the illumination field may be less than 120 degrees. Each structured light projector may also include a pattern-generating optics element. The pattern-generating optics element may utilize diffraction and / or refraction to generate a light pattern. In some applications, the light pattern may be a discrete distribution of unconnected light spots. Optionally, when the light source (e.g., a laser diode) is activated to emit light through the pattern-generating optics element, the light pattern maintains a discrete distribution of unconnected spots on all planes located between 1 mm and 30 mm from the pattern-generating optics element. In some applications, the pattern-generating optics of each structured light projector can have a light throughput efficiency of at least 80% (e.g., at least 90%), that is, the proportion of light falling on the pattern generator that enters the pattern. Each camera includes a camera sensor and an objective lens optics comprising one or more lenses.

[0009] Laser diode light sources and diffraction and / or refractive pattern generating optics can offer certain advantages in some applications. For example, using laser diodes and diffraction and / or refractive pattern generating optics can help maintain highly energy-efficient structured light projectors, thereby preventing the probe from overheating during use. Furthermore, such components can help reduce costs by making active cooling within the probe unnecessary. For example, today's laser diodes can use less than 0.6 watts of power while continuously emitting at high brightness (e.g., compared to today's light-emitting diodes (LEDs)). When pulsed excitation is performed according to some applications of the invention, these today's laser diodes can use even less power; for example, when pulsed excitation is performed at a 10% duty cycle, the laser diode can use less than 0.06 watts (but for some applications, the laser diode can use at least 0.2 watts while continuously transmitting at high brightness, and when pulsed excitation can be performed with even less power, for example, when pulsed excitation is performed at a 10% duty cycle, the laser diode can use at least 0.02 watts). Furthermore, diffraction and / or refraction pattern generating optical elements can be configured to utilize most (if not all) of the emitted light (e.g., as opposed to a mask that blocks some light from hitting an object).

[0010] Specifically, diffraction and / or refraction-based pattern-generating optics generate patterns through the diffraction, refraction, or interference of light, or any combination thereof, rather than through the modulation of light, as is done by transparency or transmission masks. In some applications, this can be advantageous because the light conversion efficiency (the ratio of light entering the pattern to light falling on the pattern generator) is close to 100%, for example at least 80%, or at least 90%, regardless of the pattern's area-based duty cycle. Conversely, the light conversion efficiency of transparency or transmission mask pattern-generating optics is directly related to the area-based duty cycle. For example, for a desired area-based duty cycle of 100:1, the conversion efficiency of a mask-based pattern generator would be 1%, while the efficiency of a diffraction and / or refraction-based pattern-generating optics remains close to 100%. Furthermore, due to the inherently smaller emission area and divergence angle of lasers, their light-gathering efficiency is at least 10 times higher than that of an LED with the same total light output, resulting in brighter output illumination per unit area. The high efficiency of lasers and diffraction and / or refraction pattern generators can help achieve highly thermally efficient configurations that limit significant probe heating during use, thereby reducing costs by potentially eliminating or limiting the need for active cooling within the probe. While laser diodes and DOEs may be particularly preferred in some applications, they are by no means essential, either individually or in combination. Other light sources (including LEDs) and pattern generating elements (including transparency masks and transmission masks) can be used in other applications with or without active cooling.

[0011] In some applications, to improve image capture of intraoral scenes under structured light illumination without using contrast enhancement methods (e.g., coating teeth with opaque powder), inventors have realized that light patterns (e.g., the distribution of discrete, unconnected light spots (e.g., opposite to lines)) can provide an improved balance between increasing pattern contrast and maintaining useful information. Generally, denser structured light patterns can provide more surface sampling, higher resolution, and better stitching of corresponding surfaces obtained from multiple image frames. However, overly dense structured light patterns can lead to more complex correspondence problems due to the larger number of spots for which they address the correspondence problem. Furthermore, denser structured light patterns may have lower pattern contrast due to the greater amount of light in the system, which may be caused by a combination of: (a) stray light reflected from the somewhat glossy surfaces of the teeth and potentially picked up by the camera, and (b) percolation, i.e., some light enters the tooth, reflects along multiple paths within the tooth, and then exits the tooth in many different directions. As further described below, methods and systems are provided for addressing the correspondence problem presented by the distribution of discrete, unconnected light spots. In some applications, discrete, unconnected light spots from each projector can be uncoded.

[0012] In some applications, the field of view of each camera can be at least 45 degrees, for example, at least 80 degrees, for example, 85 degrees. Optionally, the field of view of each camera can be less than 120 degrees, for example, less than 90 degrees. For some applications, one or more cameras have fisheye lenses or other optics that provide a field of view of up to 180 degrees.

[0013] In any case, the fields of view of each camera can be the same or different. Similarly, the focal lengths of each camera can be the same or different. As used herein, the term "field of view" for each camera refers to the diagonal field of view of each camera. Furthermore, each camera can be configured to focus on the object's focal plane at a distance between 1 mm and 30 mm from the lens, for example, at least 5 mm and / or less than 11 mm, for example, 9 mm to 10 mm, where the lens is furthest from the corresponding camera sensor. Similarly, in some applications, the illumination field of each structured light projector can be at least 45 degrees and optionally less than 120 degrees. The inventors have realized that a large field of view achieved by combining the respective fields of view of all cameras can improve accuracy due to a reduction in the amount of image stitching error, particularly in edentulous areas where high-resolution 3D features, such as smooth and sharp gingival surfaces, may be less abundant. A larger field of view allows large, smooth features (e.g., the overall curve of a tooth) to appear in each image frame, which improves the accuracy of stitching the corresponding surfaces obtained from multiple such image frames.

[0014] In some applications, a method is provided for generating digital three-dimensional images of the intraoral surface. Note that the phrase "three-dimensional image" as used in this application refers to an image of the three-dimensional intraoral surface constructed from a three-dimensional model (e.g., a point cloud). While the resulting image is generally displayed on a two-dimensional screen, it contains data related to the three-dimensional structure of the scanned object and can therefore often be manipulated to display the scanned object from different viewpoints and angles. Furthermore, data from the three-dimensional image can be used to create a physical three-dimensional model of the scanned object.

[0015] For example, one or more structured light projectors may be driven to project light patterns (e.g., a distribution of discrete, unconnected light spots, a pattern of intersecting lines (e.g., a grid), a checkerboard pattern, or some other pattern on an intraorbital surface), and one or more cameras may be driven to capture the projected image. The image captured by each camera may include a portion of the projected pattern (e.g., at least one spot). In some implementations, one or more structured light projectors project a pattern that is spatially fixed relative to one or more cameras.

[0016] Each camera includes a camera sensor with a pixel array. For each pixel, there exists a corresponding ray in 3D space originating from that pixel, directed toward the object being imaged. When each point along a particular ray of these rays is imaged on the sensor, it falls on its corresponding pixel on the sensor. As used throughout this application, including the claims, the term used here is “camera ray.” Similarly, for each projected spot from each projector, there exists a corresponding projector ray. Each projector ray corresponds to a corresponding path of a pixel on at least one camera sensor; that is, if the camera sees a feature or portion (e.g., a spot) of a pattern projected by a particular projector ray, then that feature or portion (e.g., a spot) of the pattern will necessarily be detected by a pixel on a particular pixel path corresponding to that particular projector ray. The (a) value of the camera ray corresponding to each pixel on the camera sensor of each camera and the (b) value of the projector ray corresponding to each projected feature or portion of a pattern (e.g., a spot) from each projector can be stored during the calibration process, as described below.

[0017] Regarding camera lighting, for some applications, instead of storing individual values ​​for each camera light corresponding to each pixel on each camera sensor, a smaller set of calibration values ​​can be stored that can be used to indicate each camera light. For example, parameter values ​​can be stored for a parameterized camera calibration function that takes a given three-dimensional position in space and translates it into a given pixel in the two-dimensional pixel array of the camera sensor to define the camera light.

[0018] Regarding projector rays, (a) for some applications, an indexed list containing values ​​for each projector ray is stored, and (b) alternatively, for some applications, a smaller set of calibration values ​​is stored, which can be used to indicate each projector ray. For example, parameter values ​​can be stored for a parametric projector calibration model that defines each projector ray for a given projector.

[0019] Based on stored calibration values, the processor can be used to run a correspondence algorithm to identify the three-dimensional location of each portion (e.g., a projected spot) of a feature of a projected light pattern on a surface. For a given projector ray, the processor “looks” at the corresponding camera sensor path on one of the cameras. Each detected spot or other feature along that camera sensor path will have a camera ray intersecting the given projector ray. This intersection defines a three-dimensional point in space. The processor then searches the camera sensor paths corresponding to the given projector ray on other cameras and identifies how many other cameras also detect a feature (e.g., a spot) of the pattern whose camera ray intersects the three-dimensional point in space on their corresponding camera sensor paths corresponding to the given projector ray. As used throughout this application, if two or more cameras detect a portion or feature (e.g., a spot) of the pattern whose corresponding camera ray intersects the given projector ray at the same three-dimensional point in space, then these cameras are considered to “agree” that the portion or feature (e.g., a spot) is located at that three-dimensional point. The process is repeated for additional features (e.g., spots) along the camera sensor path, and the feature (e.g., spot) that the most cameras “agree” with is identified as the feature (e.g., spot) being projected onto the surface from a given projector beam. Therefore, the three-dimensional position on the surface is calculated for this feature (e.g., the spot) of the pattern.

[0020] In some embodiments, once the position on the surface is determined for a specific feature of the pattern (e.g., a specific spot), the projector ray projecting that feature (e.g., the spot) and all camera rays corresponding to that feature (e.g., the spot) can be removed from consideration, and the correspondence algorithm is run again for the next projector ray.

[0021] Other applications of the invention relate to scanning intraoral objects by projecting a structured light pattern (e.g., parallel lines, grid, checkerboard, disjointed and / or uniform spots, random spot patterns, etc.) onto the intraoral object, capturing at least a portion of the structured light pattern projected onto the intraoral object, and tracking the captured portion of the structured light pattern across successive images. In some embodiments, tracking the captured portion of the structured light pattern across successive images can help improve scanning speed and / or accuracy.

[0022] In a more specific example related to the structured light scanner using the aforementioned projection pattern (e.g., a projection pattern of non-connected spots), the processor can be used to compare a series of images (e.g., multiple consecutive images) captured by each camera to determine which features (e.g., which projection spots) of the projection pattern can be tracked across that series of images (e.g., across multiple consecutive images). The inventors have realized that the motion of a specific detected feature or spot can be tracked across multiple images in a series of images (e.g., in consecutive image frames). Therefore, the correspondence solved for a specific point in any image or frame across which a tracked feature or spot is located provides a solution for the correspondence of features or points across all images or frames across which a tracked feature or point is located. Since the detected features or points that can be tracked across multiple images are features or points generated by the same specific projector ray, the trajectory of the tracked feature or point will follow a specific camera sensor path corresponding to that specific projector ray.

[0023] For some applications, as an alternative or addition to tracking features or points detected within a two-dimensional image, the length of each projector ray can be tracked in three-dimensional space. The length of the projector ray is defined as the distance between the origin of the projector ray (i.e., the light source) and the three-dimensional location where the projector ray intersects the intraoral surface. As further described below, tracking the length of a particular projector ray over time can help resolve correspondence uncertainties. Although the above concepts of spot and ray tracking are described in some examples of this paper concerning scanners projecting disjoint spots, it should be understood that this is exemplary and not limiting—the tracking technique is equally applicable to scanners that project other patterns (e.g., parallel lines, grids, checkerboards, disjoint and / or uniform dots, random dot patterns, etc.) onto intraoral objects.

[0024] In some embodiments, for the purpose of object scanning, it may be necessary to estimate the position of the scanner relative to the object being scanned (i.e., a three-dimensional intraoral surface) during scanning, and in some embodiments, estimation is required at all times during scanning. According to some applications of the invention, the inventors have developed a method that combines visual tracking of scanner motion with inertial measurement of scanner motion to accommodate situations where sufficient visual tracking may not be available. Accumulated data of the motion of the intraoral scanner relative to the intraoral surface (visual tracking) and the motion of the intraoral scanner relative to a fixed coordinate system (inertial measurement) can be used to build a predictive model of the motion of the intraoral surface relative to the fixed coordinate system (described further below). When sufficient visual tracking is unavailable, the processor can calculate the estimated position of the intraoral scanner relative to the intraoral surface by taking into account (e.g., subtracting, in some embodiments) the prediction of the motion of the intraoral surface relative to the fixed coordinate system from the inertial measurement of the intraoral scanner's motion relative to the fixed coordinate system (described further below). It should be understood that the scanner position estimation concept described herein can be used with intraoral scanners regardless of the scanning technique employed (e.g., parallel confocal scanning, focused scanning, wavefront scanning, stereo vision, structured light, triangulation, light field, and / or combinations thereof). Therefore, while the concept of structured light described herein is discussed, it is exemplary and in no way limiting.

[0025] In some embodiments of the structured light scanner described herein, the stored calibration values ​​may indicate (a) camera rays corresponding to each pixel on the camera sensor of each camera, and (b) projector rays corresponding to each projection feature (e.g., spot) from each structured light projector, whereby each projector ray corresponds to a corresponding pixel path on at least one camera sensor. However, over time, at least one camera and / or at least one projector may move (e.g., by rotation or translation), the optics of at least one camera and / or at least one projector may be changed, or the wavelength of the laser may be changed, causing the stored calibration values ​​to no longer accurately correspond to the camera rays and projector rays.

[0026] For any given projector ray, if the processor collects data including the calculated 3D positions of multiple detected features (e.g., specks) from that projector ray (detected at various time points) on the intraorific surface and overlays these features onto an image, then the features (e.g., specks) should all fall on the camera sensor pixel path corresponding to that projector ray. If something changes the calibration of the camera or projector, then depending on the stored calibration values, the detected features (e.g., specks) from that particular projector ray may appear not to fall on the expected camera sensor pixel path, but rather on the new, updated camera sensor pixel path. When the calibration of the camera and / or projector has been altered, the processor can reduce the difference between the updated pixel path from the calibration data and the original pixel path by changing (i) the stored calibration value of the camera ray corresponding to each pixel on the camera sensor of one or more cameras (e.g., the stored parameter value of the parameterized camera calibration model (e.g., a function)), and / or (ii) the stored calibration value of the projector ray r corresponding to each projection feature (e.g., a spot) from one or more projectors (e.g., a stored value in the index list of projector rays, or the stored parameter value of the parameterized projector calibration model).

[0027] Current calibration evaluation can be performed periodically (e.g., every scan, every 10 scans, monthly, every few months, etc.) or automatically in response to meeting certain criteria (e.g., in response to a threshold number of scans having been performed). As a result of the evaluation, the system can determine whether the calibration status is accurate or inaccurate. In one embodiment, as a result of the evaluation, the system determines whether the calibration has drifted. For example, a previous calibration may still be accurate enough to produce high-quality scans, but the system may have deviated such that if the detected trend continues, it will no longer be able to produce accurate scans in the future. In one embodiment, the system determines the drift rate and projects that drift rate into the future to determine the expected date / time when the calibration will no longer be accurate. In one embodiment, automatic or manual calibration can be scheduled for this future date / time. In the example, the processing logic evaluates the calibration status over time (e.g., by comparing the calibration status at multiple different points in time) and determines the drift rate from this comparison. Based on the drift rate, the processing logic can predict when calibration should be performed based on trend data.

[0028] Conventional intraoral scanners are manually recalibrated by the user according to a set schedule (e.g., every six months). Conventional intraoral scanners lack the ability to monitor or evaluate the current calibration status (e.g., determine whether recalibration should be performed). Furthermore, conventional intraoral scanner calibration is performed manually using specific calibration targets. Conventional intraoral scanner calibration is both time-consuming and inconvenient for the user. Therefore, dynamic calibration, performed in some embodiments described herein, provides increased convenience to the user and can be performed in a shorter time compared to conventional intraoral scanner calibration.

[0029] For some applications, if the camera and / or projector calibration has been altered, the processor may not perform recalibration, but instead determine that at least some of the stored calibration values ​​for the camera and / or projector are incorrect. For example, based on the determination that the stored calibration values ​​are incorrect, the user may be prompted to return the intraoral scanner to the manufacturer for maintenance and / or recalibration, or request a new scanner.

[0030] Visual tracking of motion of an intraoral scanner relative to the scanned object can be achieved by stitching together corresponding surfaces or point clouds obtained from adjacent image frames. As described herein, for some applications, illuminating the oral cavity under near-infrared (NIR) light can increase the number of visible features available for stitching together corresponding surfaces or point clouds obtained from adjacent image frames. In particular, NIR light penetrates teeth, allowing images captured under NIR light to include features inside the teeth, such as fissures, in contrast to two-dimensional color images taken under broad-spectrum illumination, where only features appearing on the surface of the teeth are visible. These additional sub-surface features can be used to stitch together corresponding surfaces or point clouds obtained from adjacent image frames.

[0031] For some applications, the processor can use two-dimensional images (e.g., two-dimensional color images and / or two-dimensional monochrome NIR images) in the 2D to 3D surface reconstruction of intraoral three-dimensional surfaces. As described below, using two-dimensional images (e.g., two-dimensional color images and / or two-dimensional monochrome NIR images) can significantly improve the resolution and speed of three-dimensional reconstruction. Therefore, as described herein, it is useful for some applications to enhance the three-dimensional reconstruction of intraoral three-dimensional surfaces using three-dimensional reconstruction from two-dimensional images (e.g., two-dimensional color images and / or two-dimensional monochrome NIR images). For some applications, the processor, for example, uses the correspondence algorithm described herein to calculate the corresponding three-dimensional positions of multiple points on the intraoral three-dimensional surface and calculates the three-dimensional structure of the intraoral three-dimensional surface based on multiple two-dimensional images (e.g., two-dimensional color images and / or two-dimensional monochrome NIR images) and the calculated three-dimensional positions on the intraoral surface.

[0032] According to some applications of the present invention, the calculation of the three-dimensional structure is performed by a neural network. The processor inputs (a) multiple two-dimensional images (e.g., two-dimensional color images) of the intraoral three-dimensional surface and (b) the calculated three-dimensional positions of multiple points on the intraoral three-dimensional surface, and the neural network determines and returns a corresponding estimated map (e.g., depth map, normal map, and / or curvature map) of the intraoral three-dimensional surface captured in each two-dimensional image (e.g., two-dimensional color image and / or two-dimensional monochrome NIR image).

[0033] The inventors have recognized that small manufacturing deviations may exist when intraoral scanners are commercially manufactured, resulting in (a) slight differences in the calibration of the camera and / or projector on each commercially manufactured intraoral scanner compared to the calibration during the training phase, and / or (b) slight differences in the illumination relationship between the camera and projector on each commercially manufactured intraoral scanner compared to the illumination relationship between the camera and projector during the training phase used to train the neural network. Other manufacturing deviations may also exist in the camera and / or projector. According to some applications of the invention, a method is provided in which a processor is used to overcome manufacturing deviations in the camera and / or projector of the intraoral scanner to reduce the difference between the estimated map of the intraoral three-dimensional surface and the actual structure.

[0034] According to some applications of the invention, one way to overcome manufacturing tolerances is by modifying (e.g., cropping and deforming) images from a field intraoral scanner to obtain modified images that match the field of view of a set of reference cameras used to train a neural network. The neural network is trained based on images received from the set of reference cameras, and then the field images are modified as if the neural network were receiving these images already captured by the reference cameras. The three-dimensional structure of the intraoral three-dimensional surface is then calculated based on multiple modified two-dimensional images of the intraoral three-dimensional surface; for example, the neural network determines a corresponding estimated map of the intraoral three-dimensional surface captured in each of the multiple modified two-dimensional images.

[0035] According to some applications of the present invention, a neural network determines a corresponding estimated depth map of the intraoral three-dimensional surface captured in each two-dimensional image, and the depth maps are stitched together to obtain the three-dimensional structure of the intraoral surface. However, inconsistencies may sometimes exist between the estimated depth maps. The inventors have realized that it would be advantageous if, for each estimated depth map determined by the neural network, the neural network also determined an estimated confidence map, each confidence map indicating the region-specific confidence level of the corresponding estimated depth map. Therefore, this paper provides a method for inputting multiple two-dimensional images of the intraoral three-dimensional surface into a first neural network module and a second neural network module. The first neural network module determines a corresponding estimated depth map of the intraoral three-dimensional surface captured in each two-dimensional image. The second neural network module determines a corresponding estimated confidence map corresponding to each estimated depth map. Each confidence map indicates the region-specific confidence level of the corresponding estimated depth map.

[0036] According to some applications of the present invention, two-dimensional images of the three-dimensional surface (e.g., model surface and / or intraoral surface) during the training phase are used, and (b) corresponding ground truth output images of the three-dimensional surface during the training phase are computed based on structured light images of the three-dimensional surface during the training phase. The neural network estimates an estimated image of the intraoral three-dimensional surface captured in each two-dimensional image, and then each estimated image is compared with a corresponding ground truth image of the intraoral three-dimensional surface. Based on the difference between each estimated image and its corresponding ground truth image, the neural network is optimized to better estimate subsequent estimated images.

[0037] In some applications, when the intraoral 3D surface is used to train a neural network, moving tissues (such as an object's tongue, lips, and / or cheeks) may obscure a portion of the intraoral 3D surface from the viewpoint of one or more cameras. To prevent the neural network from "learning" based on images of moving tissues (as opposed to fixed tissues on the scanned intraoral 3D surface), the images containing the moving tissues can be processed to exclude at least a portion of the moving tissues before being input into the neural network.

[0038] For some applications, a disposable cannula is placed on the distal end of an intraoral scanner, for example, before the probe is inserted into a patient's mouth, to prevent cross-contamination between patients. Due to the relative positioning of the structured light projector within the probe and the adjacent camera, as further described herein, a portion of the projected structured light pattern may be reflected from the cannula and reach the camera sensor of the adjacent camera. As further described herein, due to the polarization of the laser in the structured light projector, the laser can rotate about its own optical axis to find the polarization angle of the laser relative to the cannula, thereby reducing the degree of reflection.

[0039] According to some applications of the present invention, a Simultaneous Localization and Mapping (SLAM) algorithm is used to track the movement of a handheld stick and generate a 3D image. SLAM can be performed using two or more cameras that see substantially the same image, but from slightly different angles. However, due to the positioning of camera 24 within probe 28, and the close positioning of probe 28 to the object being scanned (i.e., the intraoral 3D surface), it is not typically the case that the two or more cameras in the probe see substantially the same image. As described below, additional challenges may arise when using SLAM algorithms when scanning the intraoral 3D surface. The inventors have devised numerous methods to overcome these challenges in order to utilize SLAM to track the movement of a handheld stick and generate a 3D image of the intraoral 3D surface, as further described herein.

[0040] According to some applications of the present invention, when a handheld probe is used to scan the intraoral 3D surface, some features (e.g., speckle distribution) may fall onto moving tissue (e.g., a patient's tongue) as a structured light projector projects its feature distribution (e.g., spot distribution) onto the intraoral surface. To improve the accuracy of 3D reconstruction algorithms, it is generally not advisable to rely on features (e.g., speckles) falling onto moving tissue to reconstruct the intraoral 3D surface. As described herein, features (e.g., speckles) can be determined on an image frame of unstructured light (e.g., which may be broad-spectrum light) scattered within a structured light image frame, whether they have been projected onto moving tissue or stable tissue within the oral cavity. A confidence grading system can be used to assign confidence levels based on the determination that detected features (e.g., speckles) have been projected onto fixed or moving tissue. Based on the confidence level of each of a plurality of features (e.g., speckles), a processor can run a 3D reconstruction algorithm using the detected features (e.g., speckles).

[0041] In a method for generating a digital 3D image as described herein, the method includes driving each of one or more structured light projectors to project a pattern onto a 3D surface within an oral cavity. The method also includes driving each of one or more cameras to capture a plurality of images, each image including at least a portion of the projected pattern. The method further includes using a processor to compare a series of images captured by the one or more cameras, determining, based on the comparison of the series of images, which portions of the projected pattern can be tracked across the series of images, and constructing a 3D model of the 3D surface within the oral cavity, at least partially based on the comparison of the series of images. In one implementation, the method further includes solving a correspondence algorithm for the tracked portion of the projected pattern in at least one of the series of images, and using the solved correspondence algorithm in at least one of the series of images to address the tracked portion of the projected pattern in images where the correspondence algorithm has not been solved, for example, solving the correspondence algorithm for the tracked portion of the projected pattern in images where the correspondence algorithm has not been solved, wherein the solution of the correspondence algorithm is used to construct the 3D model. In one implementation, the method further includes a correspondence algorithm for the tracked portion of the projection pattern based on the tracked portion's position in each image of the entire series of images, wherein the solution of the correspondence algorithm is used to construct a 3D model.

[0042] In one implementation of the method, the projection pattern comprises multiple projection spots, and portions of the projection pattern correspond to projection spots within the multiple projection spots. In another implementation, a processor is used to compare a series of images based on stored calibration values ​​indicating (a) camera rays corresponding to each pixel on a camera sensor of each of one or more cameras, and (b) projector rays corresponding to each projection spot from each of one or more structured light projectors, wherein each projector ray corresponds to a corresponding path of a pixel on at least one camera sensor, wherein determining which portions of the projection pattern can be tracked includes determining which projection spots s can be tracked across the series of images, and wherein each tracked spot s moves along a path corresponding to a pixel of a corresponding projector ray r.

[0043] In another implementation of the method, the processor further includes using the processor to determine, for each tracked spot s, a plurality of possible paths p for pixels on a given camera, each path p corresponding to a plurality of possible projector rays r. In another implementation, the processor further includes using the processor to run a correspondence algorithm to perform a plurality of operations for each possible projector ray r. The plurality of operations includes identifying how many other cameras have detected a corresponding spot q corresponding to a corresponding camera ray on a corresponding path p1 of a pixel corresponding to the projector ray r, the corresponding camera ray intersecting the projector ray r and the camera ray of the given camera corresponding to the tracked spot s. The operation also includes identifying a given projector ray r1 for which the maximum number of other cameras have detected the corresponding spot q. The operation also includes identifying the projector ray r1 as the specific projector ray r that generated the tracked spot s.

[0044] In another implementation of the method, the method includes using a processor (a) to run a correspondence algorithm to calculate the corresponding three-dimensional positions of a plurality of detected spots on a three-dimensional surface inside the mouth captured in a series of images, and (b) in at least one of the series of images, identifying the detected spots as originating from a specific projector ray r by identifying the detected spots as tracked spots s that move along a path corresponding to a pixel of a specific projector ray.

[0045] In another implementation of the method, the method includes using a processor to (a) run a correspondence algorithm to calculate the corresponding three-dimensional positions of a plurality of detected spots on a three-dimensional surface of the mouth captured in a series of images, and (b) remove the spots that are considered to be points on the three-dimensional surface of the mouth: the spots (i) are identified as originating from a specific projector ray r based on the three-dimensional position calculated by the correspondence algorithm, and (ii) are not identified as tracked spots s that move along a path corresponding to a pixel of the specific projector ray r.

[0046] In another implementation of the method, the method includes using a processor (a) to run a correspondence algorithm to calculate the corresponding three-dimensional positions of a plurality of detected spots on a three-dimensional surface inside the mouth captured in a series of images, and (b) for a detected spot that is identified as originating from two different projector rays r based on the three-dimensional position calculated by the correspondence algorithm, the detected spot is identified as originating from one of the two different projector rays r by identifying the detected spot as a tracked spot s moving along one of the two different projector rays r.

[0047] In another implementation of the method, the method includes using a processor (a) to run a correspondence algorithm to calculate the corresponding three-dimensional positions of multiple detected spots on a three-dimensional surface inside the mouth captured in a series of images, and (b) identifying the weak spot as a projected spot from a specific projector ray r by identifying the weak spot as a tracked spot s that moves along a path corresponding to a pixel of a specific projector ray r, wherein the three-dimensional position of the weak spot is not calculated by the correspondence algorithm.

[0048] In another implementation of the method, the method includes using a processor to calculate the corresponding three-dimensional position on an intraoral three-dimensional surface at the intersection of a projector ray r corresponding to the tracked spot s and a corresponding camera ray in each of a series of images across the tracked spot s.

[0049] In another implementation of this method, a correspondence algorithm is used to construct a 3D model, wherein the correspondence algorithm uses at least in part the projection pattern determined as a traceable portion across a series of images.

[0050] In another implementation of the method, the method includes using a processor (a) to determine parameters of the tracked portion of the projection pattern in at least two adjacent images from the series of images, the parameters being selected from the group consisting of: the size of the portion, the shape of the portion, the orientation of the portion, the intensity of the portion, and the signal-to-noise ratio (SNR) of the portion, and (b) to predict parameters of the tracked portion of the projection pattern in subsequent images based on the parameters of the tracked portion of the projection pattern in the at least two adjacent images.

[0051] In another implementation of the method, the processor is further used to search for portions of the projection pattern that substantially have the predicted parameters in subsequent images, based on the predicted parameters of the tracked portions of the projection pattern.

[0052] In another implementation of the method, the parameter is the shape of a portion of the projection pattern, and the use of a processor also includes using the processor to determine a search space in the next image in which to search for the tracked portion of the projection pattern, based on the predicted shape of the tracked portion of the projection pattern.

[0053] In another implementation of the method, using a processor to determine the search space includes using the processor to determine a search space in the next image in which the tracked portion of the projection pattern is to be searched, the search space having a size and aspect ratio based on the predicted shape of the tracked portion of the projection pattern.

[0054] In another implementation of the method, the parameter is the shape of a portion of the projection pattern, and the processor further includes using the processor to (a) determine the velocity vector of the tracked portion of the projection pattern based on the direction and distance the tracked portion of the projection pattern has moved between at least two adjacent images from the series of images, (b) predict the shape of the tracked portion of the projection pattern in a subsequent image in response to the shape of the tracked portion of the projection pattern in at least one of the at least two adjacent images, and (c) determine a search space in the subsequent image in response to (i) determining the velocity vector of the tracked portion of the projection pattern and combining it with (ii) the predicted shape of the tracked portion of the projection pattern to search for the tracked portion of the projection pattern.

[0055] In another implementation of the method, the parameter is the shape of a portion of the projection pattern, and the processor further includes (a) determining a velocity vector of the tracked portion of the projection pattern based on the direction and distance the tracked portion of the projection pattern has moved between at least two adjacent images from the series of images, (b) predicting the shape of the tracked portion of the projection pattern in subsequent images in response to determining the velocity vector of the tracked portion of the projection pattern, and (c) determining a search space in subsequent images in response to (i) determining the velocity vector of the tracked portion of the projection pattern and in combination with (ii) the predicted shape of the tracked portion of the projection pattern, for searching the tracked portion of the projection pattern.

[0056] In another implementation of the method, using a processor includes using the processor to predict the shape of the tracked portion of the projection pattern in a subsequent image in response to (i) determining the velocity vector of the tracked portion of the projection pattern and (ii) the shape of the tracked portion of the projection pattern in at least one of two adjacent images.

[0057] In another implementation of the method, the processor further includes using the processor to (a) determine the velocity vector of the tracked portion of the projection pattern based on the direction and distance that the tracked portion of the projection pattern has moved between at least two consecutive images in the series of images, and (b) in response to determining the velocity vector of the tracked portion of the projection pattern, determine a search space in subsequent images in which the tracked portion of the projection pattern is to be searched.

[0058] In one implementation of the second method for generating digital three-dimensional images described herein, the method includes driving each of one or more structured light projectors to project a light pattern onto an intraoral three-dimensional surface along a plurality of projector rays, and driving each of one or more cameras to capture a plurality of images, each image including at least a portion of the projected pattern, each of the one or more cameras including a camera sensor comprising a pixel array. The second method further includes using a processor to: run a correspondence algorithm to compute corresponding three-dimensional positions of a plurality of detected features of the projected pattern on the intraoral three-dimensional surface for each of the plurality of images; estimate the three-dimensional surface based on at least three features using data corresponding to the corresponding three-dimensional positions of at least three features, each feature corresponding to a corresponding projector ray r among the plurality of projector rays; for projector ray r1 among the plurality of projector rays, estimate the three-dimensional position in space of the intersection point of projector ray r1 with the estimated three-dimensional surface, wherein no three-dimensional position of a feature corresponding to projector ray r1 is computed for that projector ray r1; and identify a search space in the pixel array of at least one camera to search for features corresponding to projector ray r1 using the estimated three-dimensional positions in space.

[0059] In another implementation of the second method, a correspondence algorithm is run based on stored calibration values ​​indicating (a) camera rays corresponding to each pixel on a camera sensor of each of one or more cameras, and (b) projector rays corresponding to each feature of a projection pattern from each of one or more structured light projectors, whereby each projector ray corresponds to a corresponding path to a pixel on at least one camera sensor. Furthermore, the search space in the data includes a search space defined by one or more thresholds.

[0060] In another implementation of the second method, the processor sets a threshold such that the correspondence algorithm does not consider detected features below the threshold, and in order to search for features corresponding to the projector ray r1 in the search space, the processor lowers the threshold to consider features not considered by the correspondence algorithm. In some implementations, the threshold is an intensity threshold.

[0061] In another implementation of the second method, the light pattern comprises a distribution of discrete spots, and each feature comprises spots derived from the distribution of discrete spots.

[0062] In another implementation of the second method, the data corresponding to the corresponding three-dimensional positions of at least three features includes data using the corresponding three-dimensional positions of at least three features captured in one of a plurality of images.

[0063] In another implementation of the second method, the method further includes refining the estimate of the three-dimensional surface using data corresponding to the three-dimensional position of at least one additional feature of the projection pattern, the at least one additional feature having a three-dimensional position calculated based on another of a plurality of images. In yet another implementation of the second method, refining the estimate of the three-dimensional surface includes refining the estimate of the three-dimensional surface such that all of the at least three features and at least one additional feature lie on the estimated three-dimensional surface.

[0064] In another implementation of the second method, using data corresponding to the corresponding three-dimensional positions of at least three features includes using data corresponding to at least three features, each of which is captured in one of multiple images.

[0065] In one implementation of a third method for generating a digital 3D image, the method includes driving each of one or more structured light projectors to project a light pattern onto an intraoral 3D surface, and driving each of a plurality of cameras to capture an image including at least a portion of the projected pattern, each of the cameras including a camera sensor comprising a pixel array. The third method further includes using a processor to: run a correspondence algorithm to compute corresponding 3D positions of a plurality of features of the projected pattern on the intraoral 3D surface; using data from a first camera of the plurality of cameras to identify candidate 3D positions of a given feature of the projected pattern corresponding to one or more specific projector rays r, or otherwise associated with one or more specific projector rays r, wherein data from a second camera of the plurality of cameras is not used to identify the candidate 3D positions; using the candidate 3D positions seen by the first camera to identify a search space on the pixel array of the second camera, wherein features of the projected pattern from the projector rays r are to be searched in the search space; and if features of the projected pattern from the projector rays r are identified in the search space, refining the candidate 3D positions of the features of the projected pattern using data from the second camera.

[0066] In another implementation of the third method, in order to identify a candidate 3D location corresponding to a given spot of a particular projector ray r, the processor uses data from at least two cameras, wherein data from another camera that is not one of the at least two cameras is not used to identify the candidate 3D location, and in order to identify the search space, the processor uses candidate 3D locations seen by at least one of the at least two cameras.

[0067] In another implementation of the third method, the light pattern comprises a distribution of discrete, unconnected light spots, and wherein the projection pattern is characterized by projection spots from the unconnected light spots.

[0068] In another implementation of the third method, the processor uses stored calibration values ​​that indicate (a) camera rays corresponding to each pixel on the camera sensor of each of the plurality of cameras, and (b) projector rays corresponding to each feature of the projection light pattern from each of one or more structured light projectors, whereby each projector ray corresponds to a corresponding path to a pixel on at least one camera sensor.

[0069] In the fourth method for generating digital 3D images described herein, the method includes driving each of one or more structured light projectors to project a light pattern onto an intraoral 3D surface, and driving each of one or more cameras to capture an image including at least a portion of the pattern. The fourth method also includes using a processor to run a correspondence algorithm to compute corresponding 3D positions of multiple features of the pattern captured on the intraoral 3D surface in a series of images; identifying the computed 3D positions of detected features of the imaged pattern in a subset of at least one series of images as associated with one or more specific projector rays r; and evaluating the length associated with the one or more projector rays r in each image within the image subset based on the 3D positions of the detected features corresponding to the one or more projector rays r in the image subset.

[0070] In the fourth method, the processor can also be used to calculate the estimated length of one or more projector rays r in at least one image of a series of images, in which the three-dimensional positions of the projection features from one or more projector rays are not identified.

[0071] In one implementation of the fourth method, each of the one or more cameras includes a camera sensor comprising a pixel array, wherein calculating the corresponding three-dimensional positions of multiple features of the pattern on a three-dimensional surface within the mouth and identifying the calculated three-dimensional positions of the detected features of the pattern corresponding to a particular projector ray r is performed based on stored calibration values ​​indicating (i) camera rays corresponding to each pixel on the camera sensor of each of the one or more camera sensors, and (ii) projector rays corresponding to each feature of the projected light pattern from each of the one or more projectors, whereby each projector ray corresponds to a corresponding path of a pixel on at least one camera sensor.

[0072] In another implementation of the fourth method, the use of a processor further includes using the processor to calculate the estimated length of the projector ray r in at least one image of a series of images, in which the three-dimensional position of the projected features from the projector ray r is not identified, and based on the estimated length of the projector ray r in the at least one image of the series of images, determining a one-dimensional search space in the at least one image of the series of images, wherein the projected features from the projector ray r are to be searched in the one-dimensional search space along the corresponding path corresponding to the pixel of the projector ray r.

[0073] In another implementation of the fourth method, the use of a processor further includes using the processor to calculate the estimated length of the projector ray r in at least one image of a series of images, in which the three-dimensional position of the projected features from the projector ray r is not identified, and based on the estimated length of the projector ray r in at least one image of the series of images, determining a one-dimensional search space in the corresponding pixel arrays of a plurality of cameras, wherein the projected spots from the projector ray r are to be searched in the one-dimensional search space, and for each corresponding pixel array, the one-dimensional search space is along the corresponding path of the pixel corresponding to the ray r.

[0074] In another implementation of the fourth method, using a processor to determine a one-dimensional search space in the corresponding pixel arrays of multiple cameras includes using a processor to determine a one-dimensional search space in the corresponding pixel arrays of all cameras, wherein projection features from the projector ray r are to be searched in the one-dimensional search space.

[0075] In another implementation of the fourth method, the use of the processor further includes using the processor to identify more than one candidate three-dimensional position of the projection feature of the projector ray r in each of at least one image in a series of images that is not in the image subset, based on a correspondence algorithm, and to calculate the estimated length of the projector ray r in at least one image in the series of images, in which more than one candidate three-dimensional position of the projection feature of the projector ray r is identified.

[0076] In another implementation of the fourth method, the use of a processor further includes using the processor to determine which of the more than one candidate three-dimensional locations is the correct three-dimensional location of the projection feature by determining which of the more than one candidate three-dimensional locations corresponds to the estimated length of the projector ray r in at least one image in the series of images.

[0077] In another implementation of the fourth method, the processor further includes using the processor to: determine a one-dimensional search space in the at least one image in the series of images based on the estimated length of the projector ray r in the at least one image in the series of images, wherein a projection feature from the projector ray r is to be searched in the one-dimensional search space; and determine which of the more than one candidate three-dimensional positions of the projection feature corresponds to the feature generated by the projector ray r found in the one-dimensional search space.

[0078] In another implementation of the fourth method, the processor also includes using the processor to: evaluate the length-limited curve of the projector ray r in each image of the image subset; and remove the detected feature identified from the projector ray r from the points considered to be on the intraoral three-dimensional surface if the three-dimensional position of the projected feature corresponds to a length of at least a threshold distance from the projector ray r to the defined curve.

[0079] In another implementation of the fourth method, the pattern comprises multiple spots, and each of the multiple features of the pattern comprises spots among multiple spots.

[0080] In the fifth method for generating digital 3D images described herein, the method includes driving each of one or more structured light projectors to project a light pattern onto an intraoral 3D surface along a plurality of projector rays, and driving each of one or more cameras to capture a plurality of images, each image including at least a portion of the projected pattern, each of the one or more cameras including a camera sensor comprising a pixel array. The method further includes: running a correspondence algorithm using a processor to compute corresponding 3D positions of a plurality of detected features of the projected pattern for each of the plurality of images on the intraoral 3D surface; estimating the 3D surface based on at least three features using data corresponding to the corresponding 3D positions of at least three detected features, each feature corresponding to a corresponding projector ray r among the plurality of projector rays; estimating the 3D position in space of the intersection point of the projector ray r1 with the estimated 3D surface for the projector ray r1 among the plurality of projector rays, wherein more than one candidate 3D position of the feature corresponding to the projector ray r1 is computed for the projector ray r1; and selecting, using the estimated 3D position in space of the intersection point of the projector ray r1, which of the more than one candidate 3D position is the correct 3D position of the feature corresponding to the projector ray r1.

[0081] In another implementation of the fifth method, a correspondence algorithm is run based on stored calibration values ​​indicating (a) camera rays corresponding to each pixel on a camera sensor of each of one or more cameras, and (b) projector rays corresponding to each feature of a projection pattern from each of one or more structured light projectors, whereby each projector ray corresponds to a corresponding path to a pixel on at least one camera sensor. Furthermore, the search space in the data includes a search space defined by one or more thresholds.

[0082] In another implementation of the fifth method, the light pattern comprises a distribution of discrete spots, and each feature comprises spots derived from the distribution of discrete spots.

[0083] In another implementation of the fifth method, the data corresponding to the corresponding three-dimensional positions of at least three features includes data using the corresponding three-dimensional positions of at least three features captured in one of the multiple images.

[0084] In another implementation of the fifth method, the fifth method further includes refining the estimate of the three-dimensional surface using data corresponding to the three-dimensional position of at least one additional feature of the projection pattern, the at least one additional feature having a three-dimensional position calculated based on another of a plurality of images. In another implementation of the fifth method, refining the estimate of the three-dimensional surface includes refining the estimate of the three-dimensional surface such that all of the at least three features and at least one additional feature lie on the estimated three-dimensional surface.

[0085] In another implementation of the fifth method, using data corresponding to the corresponding three-dimensional positions of at least three features includes using data corresponding to at least three features, each of which is captured in one of multiple images.

[0086] In one method for tracking the motion of an intraoral scanner described herein, the method includes using at least one camera coupled to the intraoral scanner to measure the motion of the intraoral scanner relative to a scanned intraoral surface, and using at least one inertial measurement unit (IMU) coupled to the intraoral scanner to measure the motion of the intraoral scanner relative to the scanned intraoral surface, relative to a fixed coordinate system. The method also includes using a processor to calculate the motion of the intraoral surface relative to a fixed coordinate system based on (a) the motion of the intraoral scanner relative to the intraoral surface and (b) the motion of the intraoral scanner relative to the fixed coordinate system; establishing a predictive model of the motion of the intraoral surface relative to the fixed coordinate system based on accumulated data of the motion of the intraoral surface relative to the fixed coordinate system; and calculating an estimated position of the intraoral scanner relative to the intraoral surface based on (a) the predicted motion of the intraoral surface relative to the fixed coordinate system (derived from the motion prediction model) and (b) the motion of the intraoral scanner relative to the fixed coordinate system (measured by the IMU). In another implementation of the method for tracking motion, the method further includes determining whether to prohibit the use of at least one camera to measure the motion of the intraoral scanner relative to the intraoral surface, and calculating the estimated position of the intraoral scanner relative to the intraoral surface in response to determining that motion measurement is prohibited. In another implementation of the method for tracking motion, motion calculation is performed by calculating the difference between (a) the motion of the intraoral scanner relative to the intraoral surface and (b) the motion of the intraoral scanner relative to a fixed coordinate system.

[0087] This paper describes a method for determining whether calibration data of an intraoral scanner is incorrect. This method involves driving each of one or more light sources to project light onto a three-dimensional surface within the mouth, and driving each of one or more cameras to capture multiple images of the three-dimensional surface within the mouth. The method also includes, based on stored calibration data from the one or more light sources and the one or more cameras, using a processor: running a correspondence algorithm to calculate the corresponding three-dimensional positions of multiple features of the projected light on the three-dimensional surface within the mouth; collecting data at multiple time points, including the calculated corresponding three-dimensional positions of the multiple features on the three-dimensional surface within the mouth; and, based on the collected data, determining that at least some of the stored calibration data is incorrect.

[0088] In another implementation of the method, the one or more light sources are one or more structured light projectors, and the method includes driving each of the one or more structured light projectors to project a light pattern onto a three-dimensional surface within the mouth, driving each of the one or more cameras to capture multiple images of the three-dimensional surface within the mouth, each image including at least a portion of the projected pattern, wherein each of the one or more cameras includes a camera sensor comprising a pixel array, and stored calibration data includes stored calibration values ​​indicating (a) camera rays corresponding to each pixel on the camera sensor of each of the one or more cameras, and (b) projector rays corresponding to each feature of the projected light pattern from each of the one or more structured light projectors, whereby each projector ray corresponds to a corresponding path p of a pixel on at least one camera sensor.

[0089] In another implementation of the method, determining that at least some of the stored calibration data is incorrect includes using a processor: for each projector ray r, based on the collected data, defining an updated path p' for the pixels on each camera sensor such that all calculated 3D positions corresponding to the features generated by the projector ray r correspond to the positions along the corresponding updated path p' of the pixels on each camera sensor; comparing each updated path p' of the pixels with the path p of the pixels on each camera sensor corresponding to that projector ray r from the stored calibration values; and determining that at least some of the stored calibration values ​​are incorrect in response to an updated path p' of at least one camera sensor s being different from the path p of the pixels corresponding to that projector ray r from the stored calibration values.

[0090] One recalibration method described in this paper involves driving each of one or more light sources to project light onto a three-dimensional surface within the mouth, and driving each of one or more cameras to capture multiple images of the three-dimensional surface within the mouth. The method also includes, based on stored calibration data from the one or more light sources and the one or more cameras, using a processor to: run a correspondence algorithm to calculate the corresponding three-dimensional positions of multiple features of the projected light on the three-dimensional surface within the mouth; collect data at multiple time points, including the calculated corresponding three-dimensional positions of the multiple features on the three-dimensional surface within the mouth; and recalibrate the stored calibration data using the collected data.

[0091] In another implementation of the recalibration method, one or more light sources are one or more structured light projectors, and the method includes driving each of the one or more structured light projectors to project a light pattern onto an intraoral three-dimensional surface, driving each of the one or more cameras to capture multiple images of the intraoral three-dimensional surface, each image including at least a portion of the projected pattern, wherein each of the one or more cameras includes a camera sensor comprising a pixel array. A processor performs multiple operations using stored calibration data, including stored calibration values ​​indicating (a) camera rays corresponding to each pixel on the camera sensor of each of the one or more cameras, and (b) projector rays corresponding to each feature of the projected light pattern from each of the one or more structured light projectors, whereby each projector ray corresponds to a corresponding path p of a pixel on at least one camera sensor. The operations include running a correspondence algorithm to calculate the corresponding three-dimensional positions of the multiple features of the projected pattern on the intraoral three-dimensional surface. The operations also include collecting data at multiple time points, the data including the calculated corresponding three-dimensional positions of the multiple features on the intraoral three-dimensional surface. The operation also includes, for each projector ray r, defining an updated path p' for each pixel on each camera sensor based on the collected data, such that all computed 3D positions corresponding to features generated by the projector ray r correspond to positions along the corresponding updated path p' of the pixels on each camera sensor. The operation also includes recalibrating stored calibration values ​​using the updated path p'.

[0092] In another implementation of the recalibration method, the processor performs additional operations to recalibrate the stored calibration values. These additional operations include comparing each updated path p' of a pixel with the path p of the pixel corresponding to that projector ray r on each camera sensor from the stored calibration values. The additional operations also include reducing the difference between the updated path p' of the pixel corresponding to the projector ray r and the corresponding path p of the pixel corresponding to the projector ray r from the stored calibration values ​​if, for at least one camera sensor s, the updated path p' of the pixel corresponding to the projector ray r differs from the path p of the pixel corresponding to the projector ray r from the stored calibration values ​​by changing stored calibration data selected from the group consisting of: (i) stored calibration values ​​indicating the camera ray for each pixel on each of the one or more cameras, and (ii) stored calibration values ​​indicating the projector ray r for each projection feature from each of the one or more structured light projectors.

[0093] In another implementation of the recalibration method, the modified stored calibration data includes stored calibration values ​​indicating camera rays corresponding to each pixel on the camera sensor of one or more cameras. Furthermore, modifying the stored calibration data includes changing one or more parameters of a parameterized camera calibration function, which defines the camera rays corresponding to each pixel on at least one camera sensor, to reduce differences between: (i) the calculated corresponding three-dimensional positions of multiple features of the projected pattern on the three-dimensional surface within the mouth; and (ii) the stored calibration values ​​indicating the corresponding camera rays corresponding to each pixel on the camera sensor, at which a corresponding feature among multiple features should have been detected.

[0094] In another implementation of the recalibration method, the modified stored calibration data includes stored calibration values ​​indicating projector rays corresponding to each of a plurality of features from each of one or more structured light projectors, and the modified stored calibration data includes changing: (i) an index list of paths p to which each projector ray r is assigned to a pixel, or (ii) one or more parameters of a parameterized projector calibration model that defines each projector ray r.

[0095] In another implementation of the recalibration method, changing the stored calibration data involves modifying the index list by reallocating each projector ray r based on the corresponding updated path p' for each pixel of the projector ray r.

[0096] In another implementation of the recalibration method, changing the stored calibration data includes changing: (i) stored calibration values ​​indicating camera rays corresponding to each pixel on the camera sensors s of one or more cameras, and (ii) stored calibration values ​​indicating projector rays r corresponding to each of a plurality of features from one or more structured light projectors.

[0097] In another implementation of the recalibration method, changing the stored calibration value involves iteratively changing the stored calibration value.

[0098] In another implementation of the recalibration method, the method further includes driving each of one or more cameras to capture multiple images of a calibration object having predetermined parameters. The recalibration method also includes using a processor to: run a triangulation algorithm to calculate the corresponding parameters of the calibration object based on the captured images; and run an optimization algorithm to (b) reduce the difference between (i) the updated path p' corresponding to the pixel of the projector ray r and (ii) the path p corresponding to the pixel of the projector ray r from the stored calibration values ​​using (a) the calculated corresponding parameters of the calibration object based on the captured images.

[0099] In another implementation of the recalibration method, the calibration object is a three-dimensional calibration object of known shape, and wherein driving each of one or more cameras to capture multiple images of the calibration object includes driving each of one or more cameras to capture images of the three-dimensional calibration object, and wherein predetermined parameters of the calibration object are the dimensions of the three-dimensional calibration object. In another implementation, the processor uses calculated parameters of the calibration object to run an optimization algorithm, and the processor also uses collected data, including calculated three-dimensional positions of multiple features on the intraoral three-dimensional surface.

[0100] In another implementation of the recalibration method, the calibration object is a two-dimensional calibration object with visually distinguishable features, wherein driving each of one or more cameras to capture multiple images of the calibration object includes driving each of the one or more cameras to capture images of the two-dimensional calibration object, and wherein predetermined parameters of the two-dimensional calibration object are corresponding distances between corresponding visually distinguishable features. In another implementation, the processor uses the calculated corresponding parameters of the calibration object to run an optimization algorithm, and the processor also uses collected data, including the calculated three-dimensional positions of multiple features on a three-dimensional surface within the mouth.

[0101] In another implementation of the recalibration method, driving each of one or more cameras to capture an image of a two-dimensional calibration object includes driving each of one or more cameras to capture multiple images of the two-dimensional calibration object from multiple different viewpoints relative to the two-dimensional calibration object.

[0102] In one implementation of a device for intraoral scanning, the device includes an elongated handheld rod comprising a probe at its distal end, one or more illumination sources coupled to the probe, one or more near-infrared (NIR) light sources coupled to the probe, and one or more cameras coupled to the probe, configured to (a) capture images using light from the one or more illumination sources, and (b) capture images using NIR light from the NIR light sources. The device also includes a processor configured to run a navigation algorithm to determine the position of the elongated handheld rod as it moves through space, the inputs to which are (a) images captured using light from the one or more illumination sources, and (b) images captured using NIR light.

[0103] In another implementation of the device for intraoral scanning, one or more illumination sources include one or more structured light sources.

[0104] In another implementation of the device for intraoral scanning, one or more illumination sources include one or more incoherent light sources.

[0105] A method for tracking the motion of an intraoral scanner includes: illuminating an intraoral three-dimensional surface with one or more illumination sources coupled to the intraoral scanner; driving each of one or more NIR light sources coupled to the intraoral scanner to emit NIR light onto the intraoral three-dimensional surface; and using one or more cameras coupled to the intraoral scanner to (a) capture a first plurality of images using light from the one or more illumination sources, and (b) capture a second plurality of images using NIR light. The method further includes using a processor to run a navigation algorithm to track the motion of the intraoral scanner relative to the intraoral three-dimensional surface using (a) the first plurality of images captured using light from the one or more illumination sources, and (b) the second plurality of images captured using NIR light.

[0106] In one implementation of the method for tracking motion, one or more illumination sources are used to illuminate a three-dimensional surface within the aperture.

[0107] In one implementation of a method for tracking motion, using one or more illumination sources includes using one or more incoherent light sources.

[0108] One implementation of a sixth method for calculating the three-dimensional structure of an intraoral three-dimensional surface includes driving one or more structured light projectors to project a structured light pattern onto the intraoral three-dimensional surface, driving one or more cameras to capture multiple structured light images, each structured light image including at least a portion of the structured light pattern, driving one or more unstructured light projectors to project unstructured light onto the intraoral three-dimensional surface, and driving one or more cameras to capture multiple two-dimensional images of the intraoral three-dimensional surface. The sixth method also includes using a processor to calculate the corresponding three-dimensional positions of multiple points on the intraoral three-dimensional surface captured in the multiple structured light images, and calculating the three-dimensional structure of the intraoral three-dimensional surface based on the multiple two-dimensional images of the intraoral three-dimensional surface, which is constrained by some or all of the calculated three-dimensional positions of the multiple points.

[0109] In some implementations of the sixth method, the unstructured light is incoherent light, and the multiple two-dimensional images include multiple color two-dimensional images.

[0110] In some implementations of the sixth method, the unstructured light is near-infrared (NIR) light, and the multiple two-dimensional images include multiple monochromatic NIR images.

[0111] In another implementation of the sixth method, driving one or more structured light projectors includes driving one or more structured light projectors to project a distribution of discrete, unconnected light spots, respectively.

[0112] In another implementation of the sixth method, calculating the three-dimensional structure includes: inputting multiple two-dimensional images of the intraoral three-dimensional surface into a neural network; and having the neural network determine a corresponding estimated map of the intraoral three-dimensional surface captured in each two-dimensional image.

[0113] In another implementation of the sixth method, the sixth method also includes inputting the calculated three-dimensional positions of multiple points on the three-dimensional surface inside the mouth into the neural network.

[0114] In another implementation of the sixth method, the sixth method also includes using a processor to stitch the individual graphs together to obtain a three-dimensional structure of the intraoral three-dimensional surface.

[0115] In another implementation of the sixth method, the sixth method further includes adjusting the capture of the structured light image and the capture of the two-dimensional image to generate an alternating sequence of one or more structured light image frames interspersed with one or more non-structured light image frames.

[0116] In another implementation of the sixth method, determining includes using a neural network to identify corresponding estimated depth maps of the intraoral three-dimensional surfaces captured in each two-dimensional image. In one implementation, a processor is used to stitch the individual estimated depth maps together to obtain the three-dimensional structure of the intraoral three-dimensional surfaces.

[0117] In one implementation, (a) the processor generates a corresponding point cloud corresponding to the computed three-dimensional positions of multiple points on the intraoral three-dimensional surface captured in each structured light image, and the method further includes using the processor to stitch the respective estimated depth maps together with the corresponding point cloud. In another implementation, the method further includes determining, by a neural network, a corresponding estimated normal map of the intraoral three-dimensional surface captured in each two-dimensional image.

[0118] In another implementation of the sixth method, determining includes using a neural network to determine the corresponding estimated normal maps of the intraoral three-dimensional surfaces captured in each two-dimensional image. In one implementation, a processor is used to stitch the individual estimated normal maps together to obtain the three-dimensional structure of the intraoral three-dimensional surfaces.

[0119] In one implementation, the method further includes interpolating the three-dimensional positions on the intraoral three-dimensional surface between calculated corresponding three-dimensional positions of multiple points captured in multiple structured light images, based on the corresponding estimated normal map of the intraoral three-dimensional surface captured in each two-dimensional image.

[0120] In one implementation, the method further includes adjusting the capture of structured light images and the capture of two-dimensional images to produce an alternating sequence of one or more structured light image frames interspersed with one or more unstructured light image frames. The method includes using a processor to further: (a) generate corresponding point clouds corresponding to computed corresponding three-dimensional positions of a plurality of points on an intraoral three-dimensional surface captured in each structured light image frame; and (b) stitch the individual point clouds together for at least one subset of the plurality of points, using surface normals at each point of the subset as stitching input, wherein, for a given point cloud, the surface normals at at least one point of the subset are obtained from corresponding estimated normal maps of the intraoral three-dimensional surface captured in adjacent unstructured light image frames.

[0121] In one implementation, the method further includes using a processor to compensate for the intraoral scanner’s motion between structured light image frames and adjacent unstructured light image frames by estimating the intraoral scanner’s motion based on previous image frames.

[0122] In another implementation of the sixth method, the determination includes determining the curvature of the intraoral three-dimensional surface captured in each two-dimensional image by a neural network. In one implementation, the determination includes determining a corresponding estimated curvature map of the intraoral three-dimensional surface captured in each two-dimensional image by a neural network.

[0123] In one implementation, the method further includes using a processor to: evaluate the curvature of the intraoral three-dimensional surface captured in each two-dimensional image; and interpolate the calculated three-dimensional positions on the intraoral three-dimensional surface between corresponding three-dimensional positions of multiple points captured in multiple structured light images based on the evaluated curvature of the intraoral three-dimensional surface captured in each two-dimensional image.

[0124] In another implementation of the sixth method, the sixth method further includes adjusting the capture of the structured light image and the capture of the two-dimensional image to generate an alternating sequence of one or more structured light image frames interspersed with one or more non-structured light image frames.

[0125] In another implementation of the sixth method, the sixth method includes driving one or more cameras to capture a plurality of structured light images, which includes driving each of two or more cameras to capture a corresponding plurality of structured light images; and driving one or more cameras to capture a plurality of two-dimensional images, which includes driving each of two or more cameras to capture a corresponding plurality of two-dimensional images.

[0126] In one implementation, driving two or more cameras includes, in a given image frame, driving each of the two or more cameras to simultaneously capture a corresponding two-dimensional image of a corresponding portion of the intraoral three-dimensional surface. Inputting into the neural network includes, for a given image frame, feeding all the corresponding two-dimensional images as a single input to the neural network, wherein each of the corresponding two-dimensional images has a field of view overlapping with at least one other image of the corresponding two-dimensional surface. Determination by the neural network includes, for a given image frame, determining an estimated depth map of the intraoral three-dimensional surface, which combines the corresponding portions of the intraoral three-dimensional surface.

[0127] In one implementation, driving two or more cameras to capture multiple structured light images includes driving each of three or more cameras to capture corresponding multiple structured light images, and driving two or more cameras to capture multiple two-dimensional images includes driving each of three or more cameras to capture corresponding multiple two-dimensional images. In a given image frame, driving each of the three or more cameras simultaneously captures a corresponding two-dimensional image of a corresponding portion of the intraoral three-dimensional surface. Inputting to the neural network includes, for a given image frame, inputting a subset of the corresponding two-dimensional images as a single input to the neural network, wherein the subset includes at least two of the corresponding two-dimensional images, and each image in the subset of the corresponding two-dimensional images has a field of view overlapping with at least another image in the subset of the corresponding two-dimensional images. Determination by the neural network includes, for a given image frame, determining an estimated depth map of the intraoral three-dimensional surface, the depth map combining the corresponding portions of the intraoral three-dimensional surface captured in the subset of the corresponding two-dimensional images.

[0128] In one implementation, driving two or more cameras includes, in a given image frame, driving each of the two or more cameras to simultaneously capture a corresponding two-dimensional image of a corresponding portion of the intraoral three-dimensional surface, and inputting the image into a neural network includes, for a given image frame, feeding each of the corresponding two-dimensional images as a separate input into the neural network. Determination by the neural network includes, for a given image frame, determining a corresponding estimated depth map of each corresponding portion of the intraoral three-dimensional surface captured in each of the corresponding two-dimensional images captured in the given image frame.

[0129] In one implementation, the method further includes using a processor to merge the individual depth maps together to obtain an estimated depth map of the combination of intraoral three-dimensional surfaces captured in a given image frame. In another implementation, the method further includes training a neural network, wherein each input to the neural network during training comprises an image captured by only one camera.

[0130] In one implementation, the method further includes determining, by a neural network, a corresponding estimated confidence map for each estimated depth map, each confidence map indicating a confidence level for each region of the corresponding estimated depth map. In another implementation, merging the estimated depth maps includes using a processor, in response to determining inconsistencies between corresponding regions in at least two estimated depth maps, combining the confidence levels of each of the corresponding regions indicated by the corresponding confidence maps of each of the at least two estimated depth maps, and combining the at least two estimated depth maps.

[0131] In another implementation of the sixth method, driving one or more cameras includes driving one or more cameras of an intraoral scanner, and the method further includes training a neural network using images captured by a handheld stick during multiple training phases. Each training phase of the handheld stick includes one or more reference cameras, and each of the one or more cameras of the intraoral scanner corresponds to a corresponding one of the one or more reference cameras on the handheld stick during each training phase.

[0132] In another implementation of the sixth method, driving one or more structured light projectors includes driving one or more structured light projectors of an intraoral scanner, driving one or more unstructured light projectors includes driving one or more unstructured light projectors of an intraoral scanner, and driving one or more cameras includes driving one or more cameras of an intraoral scanner. The neural network is initially trained using images captured by one or more cameras of a training phase handheld stick, each of the one or more cameras of the intraoral scanner corresponding to a corresponding one of the one or more training phase cameras. The method then includes driving (i) one or more structured light projectors of an intraoral scanner and (ii) one or more unstructured light projectors of an intraoral scanner during scanning at multiple refinement stages; driving one or more cameras of the intraoral scanner during scanning at the refinement stages to capture (a) structured light images of multiple refinement stages and (b) two-dimensional images of multiple refinement stages; calculating a three-dimensional structure of an intraoral three-dimensional surface based on the structured light images of multiple refinement stages; and refining the training of a neural network for the intraoral scanner using (a) the two-dimensional images of multiple refinement stages captured during scanning at the refinement stages and (b) the calculated three-dimensional structure of the intraoral three-dimensional surface based on the structured light images of multiple refinement stages.

[0133] In one implementation, the neural network comprises multiple layers, and the training of the refined neural network includes a subset of the constraint layers.

[0134] In one implementation, the method further includes selecting which of the multiple scans to be used as scans in the refinement stage based on the quality level of each scan.

[0135] In one implementation, the method further includes, during the scanning phase of the refinement stage, using the calculated three-dimensional structure of the intraoral three-dimensional surface calculated based on the structured light images of multiple refinement stages as the final three-dimensional structure of the intraoral three-dimensional surface of the user of the intraoral scanner.

[0136] In another implementation of the sixth method, driving one or more cameras includes driving one or more cameras of an intraoral scanner, each of the one or more cameras of the intraoral scanner corresponding to a corresponding one of one or more reference cameras. The method also includes, using a processor: for each of the one or more cameras c of the intraoral scanner, cropping and deforming at least one of two-dimensional images of the intraoral three-dimensional surface from camera c to obtain a plurality of cropped and deformed two-dimensional images, each cropped and deformed image corresponding to a cropped and deformed field of view of camera c, the cropped and deformed field of view of camera c matching the cropped field of view of the corresponding reference camera; inputting the plurality of two-dimensional images into a neural network includes inputting the plurality of cropped and deformed two-dimensional images of the intraoral three-dimensional surface into the neural network; and determining a corresponding estimated map of the intraoral three-dimensional surface captured in each cropped and deformed two-dimensional image determined by the neural network, the neural network having been trained using images from a training phase corresponding to the cropped field of view of each of the one or more reference cameras.

[0137] In another implementation, the cropping and deformation steps include the processor using (a) stored calibration values ​​that indicate camera rays for each pixel on a camera sensor corresponding to each of the one or more cameras, and (b) reference calibration values ​​that indicate (i) camera rays for each pixel on a reference camera sensor corresponding to each of the one or more reference cameras, and (ii) the cropped field of view for each of the one or more reference cameras.

[0138] In another implementation, the cropped field of view of each of the one or more reference cameras is 85-97% of the corresponding full field of view of each of the one or more reference cameras.

[0139] In another implementation, the processor further includes, for each camera c, performing an inverse deformation of each corresponding estimated map of the intraoral three-dimensional surface captured in each cropped and deformed two-dimensional surface, to obtain a corresponding undeformed estimated map of the intraoral surface seen in each of at least one two-dimensional image from camera c before deformation.

[0140] In another implementation, unstructured light is incoherent light, and multiple two-dimensional images include multiple two-dimensional color images.

[0141] In another implementation, the unstructured light is near-infrared (NIR) light, and the multiple two-dimensional images comprise multiple monochromatic NIR images.

[0142] In another implementation, the processor further includes, for each camera c, performing an inverse deformation of each corresponding estimated map of the intraoral three-dimensional surface captured in each cropped and deformed two-dimensional surface, to obtain a corresponding undeformed estimated map of the intraoral surface seen in each of at least one two-dimensional image from camera c before deformation.

[0143] It should be noted that, with necessary modifications, all of the above implementations of the sixth method related to depth maps, normal maps, curvature maps, and their use can be performed based on a cropped and deformed runtime image from the field.

[0144] In one implementation, driving one or more structured light projectors includes driving one or more structured light projectors of an intraoral scanner, and driving one or more unstructured light projectors includes driving one or more unstructured light projectors of an intraoral scanner. The method further includes, after the neural network has been trained using images from training phases corresponding to cropped fields of view of each of one or more reference cameras: driving one or more structured light projectors of the intraoral scanner and one or more unstructured light projectors of the intraoral scanner during scanning at multiple refinement phases. The one or more cameras of the intraoral scanner are driven to capture (a) structured light images of multiple refinement phases and (b) two-dimensional images of multiple refinement phases during structured light scanning at multiple refinement phases. A three-dimensional structure of the intraoral three-dimensional surface is calculated based on the structured light images of the multiple refinement phases, and the training of the neural network for the intraoral scanner is refined using (a) the two-dimensional images of the multiple refinement phases captured during scanning at multiple refinement phases and (b) the calculated three-dimensional structure of the intraoral three-dimensional surface calculated based on the structured light images of the multiple refinement phases. In another implementation, the neural network includes multiple layers, and refining the training of the neural network includes a subset of constraint layers.

[0145] In one implementation, driving one or more structured light projectors includes driving one or more structured light projectors of an intraoral scanner, driving one or more unstructured light projectors includes driving one or more unstructured light projectors of an intraoral scanner, and determining a corresponding estimated depth map of the intraoral three-dimensional surface captured in each cropped and deformed two-dimensional image, determined by a neural network. The method further includes: (a) calculating a three-dimensional structure of the intraoral three-dimensional surface based on calculated corresponding three-dimensional positions of multiple points on the intraoral three-dimensional surface captured in multiple structured light images; (b) calculating a three-dimensional structure of the intraoral three-dimensional surface based on the corresponding estimated depth map of the intraoral three-dimensional surface captured in each cropped and deformed two-dimensional image; and (c) comparing (i) the three-dimensional structure of the intraoral three-dimensional surface calculated based on the calculated corresponding three-dimensional positions of multiple points on the intraoral three-dimensional surface with (ii) the three-dimensional structure of the intraoral three-dimensional surface calculated based on the corresponding estimated depth map of the intraoral three-dimensional surface. In response to determining the difference between (i) and (ii), the method includes: driving one or more structured light projectors of (A) an intraoral scanner and one or more unstructured light projectors of (B) an intraoral scanner during scanning of multiple refinement stages; driving one or more cameras of the intraoral scanner to capture (a) structured light images of the multiple refinement stages and (b) two-dimensional images of the multiple refinement stages during scanning of multiple refinement stages; calculating a three-dimensional structure of an intraoral three-dimensional surface based on the structured light images of the multiple refinement stages; and refining the training of a neural network for the intraoral scanner using (a) the multiple two-dimensional images captured during scanning of the refinement stages and (b) the calculated three-dimensional structure of the intraoral three-dimensional surface based on the structured light images of the multiple refinement stages. In another implementation, the neural network includes multiple layers, and refining the training of the neural network includes a subset of constraint layers.

[0146] In another implementation of the sixth method, the sixth method further includes training a neural network, the training comprising: driving one or more structured light projectors of the training phase to project a structured light pattern of the training phase onto a three-dimensional surface of the training phase; driving one or more cameras of the training phase to capture a plurality of structured light images, each image including at least a portion of the structured light pattern of the training phase; driving one or more unstructured light projectors of the training phase to project unstructured light onto the three-dimensional surface of the training phase; driving one or more cameras of the training phase to capture a plurality of two-dimensional images of the three-dimensional surface of the training phase using illumination from the unstructured light projectors of the training phase; adjusting the capture of structured light images and the capture of two-dimensional images to generate an alternating sequence of image frames of one or more structured light images interspersed with image frames of one or more two-dimensional images; inputting the plurality of two-dimensional images into the neural network; estimating by the neural network an estimated map of the three-dimensional surface of the training phase captured in each two-dimensional image; and based on the three-dimensional surface of the training phase... The neural network is fed a structured light image of a 3D surface during the training phase, along with multiple 3D reconstructions of the 3D surface, including the calculated 3D positions of multiple points on the 3D surface during the training phase. For each 2D image frame, the positions of one or more cameras relative to the 3D surface during the training phase are interpolated based on the calculated 3D positions of the multiple points on the 3D surface during the training phase, which are calculated based on corresponding structured light image frames before and after each 2D image frame. The 3D reconstructions are projected onto the corresponding field of view of each of the one or more cameras during the training phase, and based on the projection, a ground truth map of the 3D surface during the training phase, constrained by the calculated 3D positions of the multiple points, is calculated in each 2D image. Each estimated depth map of the 3D surface during the training phase is compared with the corresponding ground truth map of the 3D surface during the training phase. Based on the difference between each estimated map and the corresponding ground truth map, the neural network is optimized to better estimate subsequent estimated maps.

[0147] In another implementation, training includes initial training of a neural network, driving one or more structured light projectors (including driving one or more structured light projectors of an intraoral scanner), driving one or more unstructured light projectors (including driving one or more unstructured light projectors of an intraoral scanner), and driving one or more cameras (including driving one or more cameras of an intraoral scanner). The method further includes, after the initial training of the neural network: driving (i) one or more structured light projectors of the intraoral scanner and (ii) one or more unstructured light projectors of the intraoral scanner during structured light scanning at multiple refinement stages; driving one or more cameras of the intraoral scanner during structured light scanning at refinement stages to capture (a) structured light images at multiple refinement stages and (b) two-dimensional images at multiple refinement stages; calculating a three-dimensional structure of an intraoral three-dimensional surface based on the structured light images at multiple refinement stages; and refining the training of the neural network for the intraoral scanner using (a) the multiple two-dimensional images captured during the refinement stage scanning and (b) the calculated three-dimensional structure of the intraoral three-dimensional surface based on the structured light images at multiple refinement stages. In another implementation, the neural network comprises multiple layers, and the training of the refined neural network includes a subset of the constraint layers.

[0148] In another implementation of the sixth method, driving one or more structured light projectors to project a structured light pattern for the training phase includes driving one or more structured light projectors to project a distribution of discrete, unconnected light spots onto a three-dimensional surface during the training phase.

[0149] In another implementation of the sixth method, driving one or more training phases of the camera includes driving at least two training phases of the camera.

[0150] In another implementation of the sixth method, unstructured light includes broadband light.

[0151] In one implementation of a device for intraoral scanning, the device includes an elongated handheld rod with a probe at its distal end, the probe being configured to be removably disposed within a sleeve. The device also includes at least one structured light projector coupled to the probe, the at least one structured light projector (a) including a laser configured to emit polarized laser light, and (b) including a pattern-generating optics configured to generate a light pattern when the laser is activated to emit light that passes through the pattern-generating optics. The device also includes a camera coupled to the probe, the camera including a camera sensor. The probe is configured such that light exits through the sleeve and enters the probe. Additionally, the laser is positioned at a distance relative to the camera such that when the probe is disposed within the sleeve, a portion of the light pattern is reflected by the sleeve and reaches the camera sensor. Furthermore, the laser is positioned at a rotational angle relative to its own optical axis such that, due to the polarization of the light pattern, the degree of reflection of a portion of the light pattern by the sleeve is less than a threshold reflection for all possible rotational angles of the laser relative to its optical axis.

[0152] In another implementation of the device for intraoral scanning, the threshold is 70% of the maximum reflection for all possible rotation angles of the laser relative to its optical axis.

[0153] In another implementation of the device for intraoral scanning, the laser is positioned at a rotational angle relative to its own optical axis, such that due to the polarization of the light pattern, the sleeve reflects less than 60% of the maximum reflection for all possible rotational angles of the laser relative to its optical axis.

[0154] In another implementation of the device for intraoral scanning, the laser is positioned at a rotational angle relative to its own optical axis, such that due to the polarization of the light pattern, the sleeve reflects a portion of the light pattern to a degree of 15%–60% of the maximum reflection for all possible rotational angles of the laser relative to its optical axis.

[0155] In another implementation of the device for intraoral scanning, when the elongated handheld rod is positioned in the sleeve, the distance between the structured light projector and the camera is 1-6 times the distance between the structured light projector and the sleeve.

[0156] In another implementation of the device for intraoral scanning, at least one structured light projector has an illumination field of at least 30 degrees, and wherein the camera has a field of view of at least 30 degrees.

[0157] A seventh method for generating 3D images using an intraoral scanner includes capturing multiple images of an intraoral 3D surface using at least two cameras rigidly connected to the intraoral scanner, such that the respective fields of view of each camera have non-overlapping portions. The seventh method also includes using a processor to run a Simultaneous Localization and Mapping (SLAM) algorithm using the captured images from each camera for the non-overlapping portions of their respective fields of view, wherein the localization of each camera is solved based on the fact that the motion of each camera is the same as the motion of all other cameras.

[0158] In another implementation of the seventh method, the respective fields of view of the first and second cameras also overlap. Furthermore, capturing includes capturing multiple images of the intraoral 3D surface such that features of the intraoral 3D surface in the overlapping portion of the respective fields of view appear in the images captured by the first and second cameras. Additionally, using a processor includes running a SLAM algorithm using features of the intraoral 3D surface appearing in the images from at least two cameras.

[0159] In an eighth method for generating three-dimensional images using an intraoral scanner, the method includes driving one or more structured light projectors to project a structured light pattern onto an intraoral three-dimensional surface; driving one or more cameras to capture a plurality of structured light images, each structured light image including at least a portion of the structured light pattern; driving one or more unstructured light projectors to project unstructured light onto the intraoral three-dimensional surface; driving at least one camera to capture a two-dimensional image of the intraoral three-dimensional surface using illumination from the unstructured light projectors; and adjusting the capture of structured light and unstructured light to produce an alternating sequence of one or more structured light image frames interspersed with one or more unstructured light image frames. The eighth method also includes using a processor to calculate the corresponding three-dimensional positions of a plurality of points on the intraoral three-dimensional surface captured in the one or more structured light image frames. The eighth method further includes using a processor to interpolate the motion of at least one camera between a first unstructured light image frame and a second unstructured light image frame based on the calculated three-dimensional positions of the plurality of points in corresponding structured light image frames before and after the unstructured light image frames. The eighth method further includes (a) running a Simultaneous Localization and Mapping (SLAM) algorithm using features of the intraoral three-dimensional surface captured by at least one camera in a first unstructured light image frame and a second unstructured light image frame, the algorithm being (b) subject to motion constraints of the camera between the first unstructured light image frame and the second unstructured light image frame.

[0160] In one implementation of the eighth method, driving one or more structured light projectors to project a structured light pattern includes driving the one or more structured light projectors to each project a distribution of discrete, unconnected light spots onto a three-dimensional surface within the mouth. In one implementation, the unstructured light includes broad-spectrum light, and the two-dimensional image includes a two-dimensional color image. In one implementation, the unstructured light includes near-infrared (NIR) light, and the two-dimensional image includes a two-dimensional monochromatic NIR image.

[0161] In one implementation of a ninth method for generating three-dimensional images using an intraoral scanner, the ninth method includes driving one or more structured light projectors to project a structured light pattern onto an intraoral three-dimensional surface, driving one or more cameras to capture a plurality of structured light images, each structured light image including at least a portion of the structured light pattern, driving one or more unstructured light projectors to project unstructured light onto the intraoral three-dimensional surface, driving one or more cameras to capture two-dimensional images of the intraoral three-dimensional surface using illumination from the unstructured light projectors, and adjusting the capture of structured light and unstructured light to produce an alternating sequence of one or more structured light image frames interspersed with one or more unstructured light image frames. The ninth method further includes using a processor to calculate the three-dimensional position of a feature on a three-dimensional surface within the mouth based on a structured light image frame, the feature also being captured in a first and a second unstructured light image frame; calculating the motion of one or more cameras between the first and second unstructured light image frames based on the calculated three-dimensional position of the feature; and running a Simultaneous Localization and Mapping (SLAM) algorithm using (i) the feature of the three-dimensional surface within the mouth captured by one or more cameras in the first and second unstructured light image frames, without calculating the three-dimensional position of the feature based on the structured light image frame, and (ii) the calculated motion of the cameras between the first and second unstructured light image frames. In one implementation, driving one or more structured light projectors to project a structured light pattern includes driving one or more structured light projectors to each project a distribution of discrete, unconnected light spots onto the three-dimensional surface within the mouth. In one implementation, the unstructured light includes broad-spectrum light, and the two-dimensional image includes a two-dimensional color image. In one implementation, the unstructured light includes near-infrared (NIR) light, and the two-dimensional image includes a two-dimensional monochromatic NIR image.

[0162] In a method for calculating the three-dimensional structure of an intraoral three-dimensional surface within the oral cavity of an object, the method includes (a) driving one or more structured light projectors to project a structured light pattern onto the intraoral three-dimensional surface, the pattern including a plurality of features; (b) driving one or more cameras to capture a plurality of structured light images, each structured light image including at least one feature of the structured light pattern; (c) driving one or more unstructured light projectors to project unstructured light onto the intraoral three-dimensional surface; (d) driving at least one camera to capture a two-dimensional image of the intraoral three-dimensional surface using illumination from the one or more unstructured light projectors; and (e) adjusting the capture of structured light and unstructured light to produce an alternating sequence of one or more structured light image frames interspersed with one or more unstructured light image frames. The method further includes using a processor to (a) determine, based on a two-dimensional image, one or more features of a plurality of features for a structured light pattern, whether the feature is being projected onto moving or stable tissue within the oral cavity; (b) based on the determination, assign a corresponding confidence level to each of the one or more features, with high confidence for stable tissue and low confidence for moving tissue; and (c) based on the confidence level of each of the one or more features, run a three-dimensional reconstruction algorithm using the one or more features. In one implementation, the unstructured light includes broad-spectrum light, and the two-dimensional image is a two-dimensional color image. In one implementation, the unstructured light includes near-infrared (NIR) light, and the two-dimensional image is a two-dimensional monochromatic NIR image. In one implementation, the plurality of features includes a plurality of spots, and driving one or more structured light projectors to project the structured light pattern includes driving one or more structured light projectors to each project a distribution of discrete, unconnected light spots onto a three-dimensional surface within the oral cavity. In one implementation, the three-dimensional reconstruction algorithm is performed using only a subset of the plurality of features, which consists of features assigned confidence levels above a stable tissue threshold. In one implementation, running the 3D reconstruction algorithm includes (a) assigning weights to each feature based on the corresponding confidence level assigned to that feature, and (b) using the corresponding weights of each feature in the 3D reconstruction algorithm.

[0163] In one implementation of a tenth method for calculating the three-dimensional structure of an intraoral three-dimensional surface, the tenth method includes driving one or more light sources of an intraoral scanner to project light onto the intraoral three-dimensional surface, and driving two or more cameras of the intraoral scanner to capture multiple two-dimensional images of the intraoral three-dimensional surface, each of the two or more cameras of the intraoral scanner corresponding to a corresponding one of two or more reference cameras. The method includes, using a processor, modifying at least one two-dimensional image from camera c for each of the two or more cameras of the intraoral scanner to obtain multiple modified two-dimensional images, each modified image corresponding to a modified field of view of camera c, the modified field of view of camera c matching the modified field of view of the corresponding reference camera; and calculating the three-dimensional structure of the intraoral three-dimensional surface based on the multiple modified two-dimensional images of the intraoral three-dimensional surface.

[0164] In one implementation of an eleventh method for calculating the three-dimensional structure of an intraoral three-dimensional surface, the eleventh method includes driving one or more light sources of an intraoral scanner to project light onto the intraoral three-dimensional surface, and driving two or more cameras of the intraoral scanner to each capture multiple two-dimensional images of the intraoral three-dimensional surface, each of the one or more cameras of the intraoral scanner corresponding to a corresponding one of two or more reference cameras. The method includes, using a processor, cropping and deforming at least one two-dimensional image from each of the two or more cameras c of the intraoral scanner to obtain multiple cropped and deformed two-dimensional images, each cropped and deformed image corresponding to a cropped and deformed field of view of camera c, the cropped and deformed field of view of camera c being matched with the cropped field of view of the corresponding reference camera. The three-dimensional structure of the intraoral three-dimensional surface is calculated based on multiple cropped and deformed two-dimensional images of the intraoral three-dimensional surface by: inputting multiple cropped and deformed two-dimensional images of the intraoral three-dimensional surface into a neural network; and the neural network determining a corresponding estimated map of the intraoral three-dimensional surface captured in each of the multiple cropped and deformed two-dimensional images, the neural network having been trained with images from a training phase corresponding to cropped fields of view of each of one or more reference cameras.

[0165] In another implementation of the eleventh method, the light is incoherent, and the multiple two-dimensional images include multiple two-dimensional color images.

[0166] In another implementation of the eleventh method, the light is near-infrared (NIR) light, and the multiple two-dimensional images include multiple monochromatic NIR images.

[0167] In another implementation of the eleventh method, the light is broad-spectrum light, and the multiple two-dimensional images include multiple two-dimensional color images.

[0168] In another implementation of the eleventh method, the cropping and deformation steps include the processor using (a) stored calibration values ​​that indicate the camera light on each pixel of the camera sensor corresponding to each of the one or more cameras c, and (b) reference calibration values ​​that indicate (i) the camera light on each pixel of the reference camera sensor corresponding to each of the one or more reference cameras, and (ii) the cropped field of view of each of the one or more reference cameras.

[0169] In another implementation of the eleventh method, the cropped field of view of each of the one or more reference cameras is 85-97% of the corresponding full field of view of each of the one or more reference cameras.

[0170] In another implementation, the processor further includes, for each camera c, performing an inverse deformation of each corresponding estimated map of the intraoral three-dimensional surface captured in each cropped and deformed two-dimensional surface, to obtain a corresponding undeformed estimated map of the intraoral surface seen in each of at least one two-dimensional image from camera c before deformation.

[0171] It should be noted that, with necessary modifications, all of the above implementations of the sixth method related to depth maps, normal maps, curvature maps, and their use can be performed based on the clipped and deformed runtime image in the field from the eleventh method.

[0172] It should also be noted that, with the necessary modifications, all of the above implementations of the sixth method related to structured light can be performed in the context of the eleventh method and a clipped and deformed runtime 2D image.

[0173] In one implementation of a twelfth method for calculating the three-dimensional structure of an intraoral three-dimensional surface, the twelfth method includes driving one or more light projectors to project light onto the intraoral three-dimensional surface, and driving one or more cameras to capture multiple two-dimensional images of the intraoral three-dimensional surface. The method includes using a processor to input the multiple two-dimensional images of the intraoral three-dimensional surface into a first neural network module and a second neural network module; the first neural network module determining a corresponding estimated depth map of the intraoral three-dimensional surface captured in each two-dimensional image; and the second neural network module determining a corresponding estimated confidence map corresponding to each estimated depth map, each confidence map indicating the confidence level of each region of the corresponding estimated depth map.

[0174] In one implementation of the twelfth method, the first neural network module and the second neural network module are separate modules of the same neural network.

[0175] In one implementation of the twelfth method, each of the first neural network module and the second neural network module is not a separate module of the same neural network.

[0176] In another implementation of the twelfth method, the method further includes: training a second neural network module to determine a corresponding estimated confidence map corresponding to each estimated depth map determined by the first neural network module by initially training a first neural network to determine corresponding estimated depth maps using two-dimensional images from multiple depth training phases; and subsequently: (i) inputting two-dimensional images of the three-dimensional surface of the training phases from multiple confidence training phases into the first neural network module; and (ii) having the first neural network module determine the corresponding estimated depth map of the three-dimensional surface of the training phase captured in the two-dimensional images of each confidence training phase. iii) Calculate the difference between each estimated depth map and the corresponding ground truth depth maps to obtain a corresponding target confidence map for each estimated depth map determined by the first neural network module; (iv) input the two-dimensional images of multiple confidence training phases into the second neural network module; (v) estimate the corresponding estimated confidence maps by the second neural network module, which indicate the confidence level of each region of each corresponding estimated depth map; and (vi) compare each estimated confidence map with the corresponding target confidence map, and based on the comparison, optimize the second neural network module to better estimate subsequent estimated confidence maps.

[0177] In one implementation, the two-dimensional images trained on multiple confidence levels are different from the two-dimensional images trained on multiple depth levels.

[0178] In one implementation, the two-dimensional images for multiple confidence training phases are the same as the two-dimensional images for multiple depth training phases.

[0179] In another implementation of the twelfth method, (a) driving one or more cameras to capture multiple two-dimensional images includes, in a given image frame, driving each of two or more cameras to simultaneously capture a corresponding two-dimensional image of a corresponding portion of the intraoral three-dimensional surface; (b) inputting the multiple two-dimensional images of the intraoral three-dimensional surface to a first neural network module and a second neural network module includes, for a given image frame, inputting each of the corresponding two-dimensional images as a separate input to the first neural network module and the second neural network module; (c) determination by the first neural network module includes, for a given image frame, determining a corresponding estimated depth map for each corresponding portion of the intraoral three-dimensional surface captured in each of the corresponding two-dimensional images captured in the given image frame; and (d) determination by the second neural network module includes, for a given image frame, determining a corresponding estimated confidence map corresponding to each corresponding estimated depth map of each corresponding portion of the intraoral three-dimensional surface captured in each of the corresponding two-dimensional images captured in the given image frame. The method also includes using a processor to combine the respective estimated depth maps to obtain an estimated depth map of a combination of the intraoral three-dimensional surfaces captured in the given image frame. In response to determining inconsistencies between corresponding regions in at least two estimated depth maps, the processor merges the at least two estimated depth maps based on the confidence level of each of the corresponding regions indicated by the corresponding confidence map of each of the at least two estimated depth maps.

[0180] In one implementation of a thirteenth method for calculating the three-dimensional structure of an intraoral three-dimensional surface, the thirteenth method includes driving one or more light sources of an intraoral scanner to project light onto the intraoral three-dimensional surface, and driving one or more cameras of the intraoral scanner to capture multiple two-dimensional images of the intraoral three-dimensional surface. The method includes (a) using a processor, a neural network determines a corresponding estimated map of the intraoral three-dimensional surface captured in each two-dimensional image, and (b) using the processor, overcoming manufacturing biases of the one or more cameras of the intraoral scanner to reduce the discrepancy between the estimated map of the intraoral three-dimensional surface and the true structure.

[0181] In one implementation of the thirteenth method, overcoming manufacturing deviations of one or more cameras includes overcoming manufacturing deviations of one or more cameras relative to a reference set of one or more cameras.

[0182] In one implementation of the thirteenth method, the intraoral scanner is one of a plurality of manufactured intraoral scanners, each manufactured intraoral scanner including a set of one or more cameras, and overcoming manufacturing deviations of the one or more cameras of the intraoral scanner includes overcoming manufacturing deviations of the one or more cameras relative to the set of one or more cameras of at least another of the plurality of manufactured intraoral scanners.

[0183] In another implementation of the thirteenth method, driving one or more cameras includes driving two or more cameras of an intraoral scanner to each capture multiple two-dimensional images of the intraoral three-dimensional surface, each of the two or more cameras of the intraoral scanner corresponding to a corresponding one of two or more reference cameras, and the neural network has been trained using images from a training phase captured by the two or more reference cameras. Overcoming manufacturing bias includes overcoming manufacturing bias of the two or more cameras of the intraoral scanner by using a processor, (a) for each of the two or more cameras c of the intraoral scanner, modifying at least one two-dimensional image from camera c to obtain multiple modified two-dimensional images, each modified image corresponding to a modified field of view of camera c, the modified field of view of camera c matching the modified field of view of the corresponding reference camera; and (b) determining a corresponding estimated map of the intraoral three-dimensional surface by the neural network based on the multiple modified two-dimensional images of the intraoral three-dimensional surface.

[0184] In another implementation, the light is incoherent, and the multiple two-dimensional images include multiple two-dimensional color images.

[0185] In another implementation, the light is near-infrared (NIR) light, and the multiple two-dimensional images include multiple monochromatic NIR images.

[0186] In another implementation, the light is broad-spectrum light, and the multiple two-dimensional images include multiple two-dimensional color images.

[0187] In another implementation, the modification step includes cropping and deforming at least one two-dimensional image from camera c to obtain a plurality of cropped and deformed two-dimensional images, each cropped and deformed image corresponding to a cropped and deformed field of view of camera c, and the cropped and deformed field of view of camera c matching the cropped field of view of a corresponding reference camera.

[0188] In another implementation, the cropping and deformation steps include the processor using (a) stored calibration values ​​that indicate the camera light on each pixel of the camera sensor corresponding to each of the one or more cameras c, and (b) reference calibration values ​​that indicate (i) the camera light on each pixel of the reference camera sensor corresponding to each of the one or more reference cameras, and (ii) the cropped field of view of each of the one or more reference cameras.

[0189] In another implementation, the cropped field of view of each of the one or more reference cameras is 85-97% of the corresponding full field of view of each of the one or more reference cameras.

[0190] In another implementation, the processor further includes, for each camera c, performing an inverse deformation on each of the corresponding estimated maps of the intraoral three-dimensional surface captured in each cropped and deformed two-dimensional surface to obtain a corresponding undeformed estimated map of the intraoral surface seen in each of at least one two-dimensional image from camera c before deformation.

[0191] In another implementation of the thirteenth method, overcoming manufacturing bias of one or more cameras of the intraoral scanner includes training the neural network using images captured by the intraoral scanner during multiple training phases. Each training phase of the intraoral scanner includes one or more reference cameras, each of the one or more cameras of the intraoral scanner corresponding to a corresponding one of the one or more reference cameras on the intraoral scanner during each training phase, and the manufacturing bias of the one or more cameras is the manufacturing bias of the one or more cameras relative to the corresponding one or more reference cameras.

[0192] In another implementation of the thirteenth method, driving one or more cameras includes driving two or more cameras of an intraoral scanner to each capture multiple two-dimensional images of the intraoral three-dimensional surface. Overcoming manufacturing bias includes overcoming manufacturing bias of the two or more cameras of the intraoral scanner by: training a neural network using images from a training phase, each captured by only one camera; driving the two or more cameras of the intraoral scanner to simultaneously capture corresponding two-dimensional images of corresponding portions of the intraoral three-dimensional surface in a given image frame; for a given image frame, feeding each of the corresponding two-dimensional images as a separate input to the neural network; having the neural network determine a corresponding estimated depth map of each corresponding portion of the intraoral three-dimensional surface captured in each of the corresponding two-dimensional images captured in the given image frame; and using a processor, merging the respective estimated depth maps together to obtain an estimated depth map of the combined intraoral three-dimensional surface captured in the given image frame.

[0193] In another implementation, the determination also includes determining a corresponding estimated confidence map for each estimated depth map by a neural network, each confidence map indicating the confidence level of each region of the corresponding estimated depth map.

[0194] In another implementation, merging the estimated depth maps together includes using a processor to merge the at least two estimated depth maps based on the confidence level of each of the corresponding regions indicated by the corresponding confidence map of each of the at least two estimated depth maps in response to determining inconsistencies between corresponding regions in the at least two estimated depth maps.

[0195] In another implementation of the thirteenth method, overcoming manufacturing deviations of one or more cameras of the intraoral scanner includes: (a) initially training the neural network using images of one or more training phases captured by one or more training phase cameras of one or more training phase handheld rods, each of the one or more cameras of the intraoral scanner corresponding to a corresponding one of the one or more training phase cameras on each of the one or more training phase handheld rods; and (b) subsequently driving the intraoral scanner to perform multiple refinement phase scans of the intraoral three-dimensional surface, and refining the training of the neural network for the intraoral scanner using the refinement phase scans of the intraoral three-dimensional surface.

[0196] In another implementation, the neural network comprises multiple layers, and the training of the refined neural network includes a subset of the constraint layers.

[0197] In another implementation, the method further includes selecting which of the multiple scans to be used as refinement stage scans based on the quality level of each scan.

[0198] In another implementation, driving the intraoral scanner to perform multiple refinement stages of scanning includes: during the multiple refinement stages of scanning, driving (i) one or more structured light projectors of the intraoral scanner to project a structured light pattern onto a three-dimensional surface within the intraoral cavity, and (ii) one or more unstructured light projectors of the intraoral scanner to project unstructured light onto the three-dimensional surface within the intraoral cavity; driving one or more cameras of the intraoral scanner to (a) capture structured light images of the multiple refinement stages using illumination from the structured light projectors, and (b) capture two-dimensional images of the multiple refinement stages using illumination from the unstructured light projectors; and calculating a three-dimensional structure of the three-dimensional surface within the intraoral cavity based on the structured light images of the multiple refinement stages.

[0199] In another implementation, training the neural network involves refining the training of the neural network for the intraoral scanner using (a) two-dimensional images of multiple refinement stages captured during the scan in the refinement stage and (b) a computed three-dimensional structure of the intraoral three-dimensional surface computed based on the structured light images of the multiple refinement stages.

[0200] In another implementation, the method further includes, during the scanning phase of the refinement phase, using the calculated three-dimensional structure of the intraoral three-dimensional surface calculated based on the structured light images of multiple refinement phases as the final three-dimensional structure of the intraoral three-dimensional surface of the user of the intraoral scanner.

[0201] In one implementation of a fourteenth method for training a neural network used with an intraoral scanner, the fourteenth method includes inputting multiple two-dimensional images of an intraoral three-dimensional surface into the neural network; estimating an estimated map of the intraoral three-dimensional surface captured in each two-dimensional image by the neural network; calculating a true map of the intraoral three-dimensional surface seen in each two-dimensional image based on multiple structured light images of the intraoral three-dimensional surface; comparing each estimated map of the intraoral three-dimensional surface with a corresponding true map of the intraoral three-dimensional surface; and optimizing the neural network to better estimate subsequent estimated maps based on the differences between each estimated map and the corresponding true map, wherein, for a two-dimensional image in which mobile tissue is identified, the image is processed to exclude at least a portion of the mobile tissue before the two-dimensional image is input into the neural network.

[0202] In another implementation of the fourteenth method, the method further includes: driving one or more structured light projectors to project a structured light pattern onto a three-dimensional surface within the mouth; driving one or more cameras to capture a plurality of structured light images, each image including at least a portion of the structured light pattern; driving one or more unstructured light projectors to project unstructured light onto the three-dimensional surface within the mouth; driving one or more cameras to capture a plurality of two-dimensional images of the three-dimensional surface within the mouth using illumination from the unstructured light projectors; and adjusting the capture of the structured light images and the capture of the two-dimensional images to produce an alternating sequence of image frames of one or more structured light images interspersed with image frames of one or more two-dimensional images. Furthermore, calculating the true image of the intraoral three-dimensional surface seen in each two-dimensional image includes: inputting a corresponding multiple three-dimensional reconstructions of the intraoral three-dimensional surface into a neural network based on a structured light image of the intraoral three-dimensional surface, the three-dimensional reconstructions including the calculated three-dimensional positions of multiple points on the intraoral three-dimensional surface; for each two-dimensional image frame, interpolating the positions of one or more cameras relative to the intraoral three-dimensional surface based on the calculated three-dimensional positions of multiple points on the intraoral three-dimensional surface, which are calculated based on the corresponding structured light image frames before and after each two-dimensional image frame; and projecting the three-dimensional reconstructions onto the corresponding field of view of each of the one or more cameras, and based on the projection, calculating the true image of the intraoral three-dimensional surface seen in each two-dimensional image, which is constrained by the calculated three-dimensional positions of the multiple points.

[0203] According to some applications of the present invention, a method for generating digital three-dimensional images is additionally provided, the method comprising:

[0204] Drive each of one or more structured light projectors to project a light pattern (e.g., a distribution of discrete, unconnected light spots) onto a three-dimensional surface inside the mouth.

[0205] Each of one or more cameras is driven to capture multiple images, each image including at least a portion of a projected pattern; each of the one or more cameras includes a camera sensor comprising a pixel array; and

[0206] A processor is used to compare multiple consecutive images captured by each camera and to determine portions of the captured projection pattern (e.g., projection spots) that can be tracked across multiple images.

[0207] For some applications, the projection pattern is a distribution of disjoint light spots, and the processor can determine this based on stored calibration values ​​indicating (a) camera rays corresponding to each pixel on a camera sensor of one or more cameras, and (b) projector rays corresponding to each projected light spot from one or more projectors. In some embodiments, each projector ray corresponds to a corresponding path of a pixel on at least one camera sensor. Furthermore, in some embodiments, the processor can determine which projected spots s can be tracked across multiple images, each tracked spot s moving along a path corresponding to a pixel of a corresponding projector ray r.

[0208] For some applications, the use of the processor also includes using the processor to calculate the corresponding three-dimensional position on the intraorbital three-dimensional surface at the intersection of the projector ray r and the corresponding camera ray, the corresponding camera ray corresponding to the tracked spot s in each of a plurality of consecutive images, wherein the spot s is tracked across the plurality of consecutive images.

[0209] For some applications, using a processor also includes using the processor to:

[0210] (a) Determine parameters of the tracked blob in at least two adjacent images from consecutive images, the parameters including one or more of the following: blob size, blob shape, blob orientation, blob intensity, and blob signal-to-noise ratio (SNR), and

[0211] (b) Based on the parameters of the tracked blob in at least two adjacent images, predict the parameters of the tracked blob in subsequent images.

[0212] For some applications, the use of the processor also includes using predicted parameters based on the tracked blobs to search for blobs that essentially have predicted parameters in subsequent images.

[0213] For some applications, the selected parameter is the shape of the blob, and the use of the processor also includes using the processor to determine the search space in the next image where to search for the tracked blob based on the predicted shape of the tracked blob.

[0214] For some applications, using a processor to determine the search space includes using a processor to determine the search space in the next image in which to search for the tracked blob, the search space having the size and aspect ratio based on the predicted shape of the tracked blob.

[0215] For some applications, the selected parameter is the shape of the spot, and the use of the processor also includes using the processor to:

[0216] (a) Determine the velocity vector of the tracked blob based on the direction and distance the blob has moved between two adjacent images from consecutive images.

[0217] (b) In response to determining the velocity vector of the tracked blob, predict the shape of the tracked blob in subsequent images, and

[0218] (c) In response to (i) determining the velocity vector of the tracked blob and combining (ii) the predicted shape of the tracked blob, determine the search space in the subsequent image in which to search for the tracked blob.

[0219] For some applications, the selected parameter is the shape of the spot, and the use of the processor also includes using the processor to:

[0220] (a) Determine the velocity vector of the tracked blob based on the direction and distance the blob has moved between two adjacent images from consecutive images.

[0221] (b) In response to determining the velocity vector of the tracked blob, predict the shape of the tracked blob in subsequent images, and

[0222] (c) In response to (i) determining the velocity vector of the tracked blob and combining (ii) the predicted shape of the tracked blob, determine the search space in the subsequent image in which to search for the tracked blob.

[0223] For some applications, the processor includes, in response to (i) determining the velocity vector of the tracked blob and (ii) combining the shape of the tracked blob in at least one of two adjacent images, using the processor to predict the shape of the tracked blob in a subsequent image.

[0224] For some applications, using a processor also includes using the processor to:

[0225] (a) Determine the velocity vector of the tracked blob based on the direction and distance the blob has moved between two consecutive images, and

[0226] (b) In response to determining the velocity vector of the tracked blob, determine the search space in the subsequent image in which to search for the tracked blob.

[0227] For some applications, the use of a processor also includes using the processor to determine, for each tracked spot s, multiple possible paths p of pixels on a given camera, where path p corresponds to multiple possible projector rays r.

[0228] For some applications, using the processor also includes using the processor to run mapping algorithms:

[0229] (a) For each possible projector ray r:

[0230] The number of other cameras that detect the corresponding spot q corresponding to the corresponding camera ray on the corresponding path p1 of the pixel corresponding to the projector ray r, the corresponding camera ray intersecting with the projector ray r and the camera ray of a given camera corresponding to the tracked spot s;

[0231] (b) Identify a given projector ray r1, for which the maximum number of other cameras detect the corresponding spot q; and

[0232] (c) Identify the projector ray r1 as the specific projector ray r that produces the tracked spot s.

[0233] For some applications, using a processor also includes using the processor to:

[0234] (a) Run the correspondence algorithm to calculate the corresponding 3D positions of multiple detected spots on the intraoral 3D surface captured in multiple consecutive images.

[0235] (b) In at least one of the plurality of consecutive images, the detected spots are identified as spots that move along a path of a pixel corresponding to a particular projector ray r.

[0236] For some applications, using a processor also includes using the processor to:

[0237] (a) Run the correspondence algorithm to calculate the corresponding 3D positions of multiple detected spots on the intraoral 3D surface captured in multiple consecutive images, and

[0238] (b) Remove the following spots from the points considered to be on the intraoral three-dimensional surface: the spots (i) are identified as originating from a specific projector ray r based on the three-dimensional position calculated by the correspondence algorithm, and (ii) are not identified as tracked spots s that move along the path of the pixel corresponding to the specific projector ray r.

[0239] For some applications, using a processor also includes using the processor to:

[0240] (a) Run the correspondence algorithm to calculate the corresponding 3D positions of multiple detected spots on the intraoral 3D surface captured in multiple consecutive images, and

[0241] (b) For a detected spot whose three-dimensional position is calculated by the correspondence algorithm and is identified as coming from two different projector rays r, the detected spot is identified as coming from one of the two different projector rays r by identifying the detected spot as a tracked spot s moving along one of the two different projector rays r.

[0242] For some applications, using a processor also includes using the processor to:

[0243] (a) Run the correspondence algorithm to calculate the corresponding 3D positions of multiple detected spots on the intraoral 3D surface captured in multiple consecutive images, and

[0244] (b) The weak spot is identified as a tracked spot s that moves along a path corresponding to a pixel of a specific projector ray r, and the three-dimensional position of the weak spot is not calculated by the correspondence algorithm.

[0245] According to some applications of the present invention, a method for generating digital three-dimensional images is also provided, the method comprising:

[0246] Drive each of one or more structured light projectors to project a light pattern (e.g., a distribution of discrete, unconnected light spots) onto a three-dimensional surface inside the mouth.

[0247] Drive each of one or more cameras to capture an image that includes at least a portion of a projected pattern, each of the one or more cameras including a camera sensor that includes a pixel array;

[0248] Use the processor to:

[0249] (a) Run the correspondence algorithm to calculate the corresponding 3D position of the detected pattern portion on the intraoral 3D surface captured in multiple consecutive images.

[0250] (b) In at least one subset of multiple consecutive images, the calculated three-dimensional position of a portion of the detected pattern is identified as corresponding to a specific projector ray r, and

[0251] (c) Calculate the length of the projector ray r in each image of the image subset based on the three-dimensional position of the portion of the detected pattern corresponding to the projector ray r in the image subset.

[0252] In some embodiments, the light pattern may be a distribution of disjoint spots. In some embodiments, the processor may perform steps (a)-(c) based on stored calibration values ​​indicating (i) camera rays corresponding to each pixel on a camera sensor of each of one or more cameras, and (ii) projector rays corresponding to each projected light spot from each of one or more projectors. In some embodiments, each projector ray corresponds to a corresponding path of a pixel on at least one camera sensor.

[0253] For some applications, the use of a processor also includes using the processor to calculate the estimated length of the projector ray r in at least one of a plurality of consecutive images, in which the three-dimensional position of the projected spot from the projector ray r was not identified in step (b).

[0254] For some applications, the processor also includes using the processor to determine a one-dimensional search space in at least one of the plurality of images, based on the estimated length of the projector ray r in at least one of the plurality of images, in which to search for the projected spot from the projector ray r, the one-dimensional search space being along the corresponding path corresponding to the pixel of the projector ray r.

[0255] For some applications, the use of the processor also includes using the processor to determine a one-dimensional search space in the respective pixel arrays of the plurality of cameras, based on the estimated length of the projector ray r in at least one of the plurality of images, in which the projected spot from the projector ray r is to be searched, the one-dimensional search space being along the respective path of the pixel corresponding to the ray r for each respective pixel array.

[0256] For some applications, using a processor to determine a one-dimensional search space in the corresponding pixel arrays of multiple cameras includes using a processor to determine a one-dimensional search space in the corresponding pixel arrays of all cameras in which to search for the projection spot from the projector ray r.

[0257] For some applications, the use of a processor also includes using the processor to calculate the estimated length of the projector ray r in at least one of a plurality of consecutive images, in which more than one candidate three-dimensional position of the projected spot from the projector ray r is identified in step (b).

[0258] For some applications, the use of a processor also includes using the processor to determine which of more than one candidate 3D locations corresponds to the estimated length of the projector ray r in at least one of the multiple images to determine which of the more than one candidate 3D locations is the correct 3D location of the projected spot.

[0259] For some applications, the processor also includes an estimated length of the projector ray r in at least one of multiple images, using the processor to:

[0260] (a) Determine a one-dimensional search space in at least one of a plurality of images in which to search for a projection spot from the projector ray r, and

[0261] (b) The correct three-dimensional position of the projection spot is determined by identifying which of the more than one candidate three-dimensional positions corresponds to the spot generated by the projector ray r found in the one-dimensional search space.

[0262] For some applications, using a processor also includes using the processor to:

[0263] (i) A length-limited curve based on the projector ray r in each image of the image subset, and

[0264] (ii) If the three-dimensional position of the projected spot corresponds to the length of the projector ray r at least a threshold distance from the defined curve, the detected spot identified in step (b) as coming from the projector ray r is removed from the point considered to be on the three-dimensional surface inside the mouth.

[0265] According to some applications of the present invention, a method for generating digital three-dimensional images is also provided, the method comprising:

[0266] Drive each of one or more structured light projectors to project a discrete, unconnected distribution of light spots onto a three-dimensional surface inside the mouth.

[0267] Drive each of one or more cameras to capture an image including at least one blob, each of the one or more cameras including a camera sensor including a pixel array;

[0268] Based on stored calibration values, it indicates (a) camera rays corresponding to each pixel on the camera sensor of each of one or more cameras, and (b) projector rays corresponding to each projection spot from each of one or more projectors, whereby each projector ray corresponds to a corresponding path to a pixel on at least one camera sensor:

[0269] Using a processor:

[0270] (a) Run the correspondence algorithm to calculate the corresponding three-dimensional positions of multiple projected spots on the three-dimensional surface inside the mouth.

[0271] (b) Using data from at least two cameras, identify candidate 3D locations for a given spot corresponding to a specific projector ray r, and substantially without using data from another camera to identify the candidate 3D location.

[0272] (c) Using candidate 3D positions seen by at least one of the two cameras, identify the search space on the pixel array of the other camera in which to search for spots from the projector ray r, and

[0273] (d) If a spot from the projector ray r is identified in the search space, the candidate 3D position of the spot is refined using data from the other camera.

[0274] According to some applications of the present invention, a method for generating digital three-dimensional images is also provided, the method comprising:

[0275] Drive each of one or more structured light projectors to project a discrete, unconnected distribution of light spots onto a three-dimensional surface inside the mouth.

[0276] Drive each of one or more cameras to capture multiple images, each image including at least one blob, each of the one or more cameras including a camera sensor including a pixel array;

[0277] Based on stored calibration values, it indicates (a) camera rays corresponding to each pixel on the camera sensor of each of one or more cameras, and (b) projector rays corresponding to each projection spot from each of one or more projectors, whereby each projector ray corresponds to a corresponding path to a pixel on at least one camera sensor:

[0278] Using a processor:

[0279] (a) Run the correspondence algorithm to calculate the corresponding three-dimensional positions of the multiple detected spots on the three-dimensional surface inside the mouth for each of the multiple images.

[0280] (b) Using data corresponding to the respective 3D positions of at least three spots, each spot corresponding to a corresponding projector ray r, estimate the 3D surface on which all at least three spots reside.

[0281] (c) For projector ray r1 for which the three-dimensional position of the corresponding spot was not calculated in step (a), estimate the three-dimensional position in space of the intersection point of projector ray r1 and the estimated surface, and

[0282] (d) Using the estimated three-dimensional position in the space, identify the search space in the pixel array of at least one camera in which to search for the spot corresponding to the projector ray r1.

[0283] For some applications, using data corresponding to the corresponding 3D positions of at least three spots includes using data corresponding to the corresponding 3D positions of at least three spots captured in one of multiple images.

[0284] For some applications, the method also includes refining the estimate of a three-dimensional surface using data corresponding to the three-dimensional position of at least one additional spot, which has a calculated three-dimensional position based on another of multiple images, such that at least three spots and at least one additional spot are located on the three-dimensional surface.

[0285] For some applications, using data corresponding to the corresponding 3D positions of at least three spots includes using data corresponding to at least three spots, each spot being captured in one of multiple images.

[0286] According to some applications of the present invention, a method for tracking the motion of an intraoral scanner is also provided, the method comprising:

[0287] (A) Using at least one camera coupled to an intraoral scanner, measure the motion of the intraoral scanner relative to the intraoral surface being scanned.

[0288] (B) Using at least one inertial measurement unit (IMU) coupled to an intraoral scanner, the motion of the intraoral scanner relative to a fixed coordinate system is measured; and

[0289] (C) Using a processor:

[0290] (i) The motion of the intraoral surface relative to the fixed coordinate system is calculated by subtracting the motion of the intraoral scanner relative to the intraoral surface from the motion of the intraoral scanner relative to the fixed coordinate system in (b).

[0291] (ii) Based on the accumulated data of the motion of the intraoral surface relative to a fixed coordinate system, establish a predictive model for the motion of the intraoral surface relative to a fixed coordinate system, and

[0292] (iii) The estimated position of the intraoral scanner relative to the intraoral surface is calculated by subtracting the predicted position of the intraoral scanner relative to the coordinate system from the position of the intraoral scanner relative to the coordinate system measured by the IMU in (b).

[0293] For some applications, the method also includes determining whether to prohibit the use of at least one camera to measure the movement of the intraoral scanner relative to the intraoral surface, and in response to determining that the movement measurement is prohibited, calculating an estimated position of the intraoral scanner relative to the intraoral surface.

[0294] According to some applications of the present invention, a method is also provided, comprising:

[0295] Drive each of one or more structured light projectors to project a discrete, unconnected distribution of light spots onto a three-dimensional surface inside the mouth.

[0296] Drive each of one or more cameras to capture multiple images, each image including at least one blob, each of the one or more cameras including a camera sensor including a pixel array;

[0297] Based on stored calibration values, it indicates (a) camera rays corresponding to each pixel on the camera sensor of each of one or more cameras, and (b) projector rays corresponding to each projection spot from each of one or more projectors, whereby each projector ray corresponds to a corresponding path p of a pixel on at least one camera sensor:

[0298] Using a processor:

[0299] (a) Run the correspondence algorithm to calculate the corresponding three-dimensional positions of multiple projected spots on the three-dimensional surface inside the mouth.

[0300] (b) Data was collected at multiple time points, including the calculated three-dimensional locations of multiple detected spots on the three-dimensional surface of the oral cavity.

[0301] (c) For each projector ray r, based on the collected data, define an updated path p' for the pixels on each camera sensor such that all calculated 3D positions corresponding to the spots generated by the projector ray r correspond to the positions along the corresponding updated path p' of the pixels on each camera sensor.

[0302] (d) Compare each updated path p' of a pixel with the path p of the pixel corresponding to the projector ray r on each camera sensor.

[0303] (e) If, for at least one camera sensor s, the updated path p' of the pixel corresponding to the projector ray r is different from the path p of the pixel corresponding to the projector ray r from the stored calibration value:

[0304] By changing the stored calibration data selected from the following groups, the difference between the updated path p' corresponding to the pixel of each projector ray r and the corresponding path p from the stored calibration value for each projector ray r is reduced:

[0305] (i) Indicates the stored calibration value of the camera light for each pixel on the camera sensor s corresponding to each of one or more cameras, and

[0306] (ii) Indicates the stored calibration value of the projector ray r corresponding to each projection spot from one or more projectors.

[0307] For some applications:

[0308] The selected stored calibration data includes stored calibration values ​​that indicate the camera light intensity for each pixel on the camera sensor corresponding to each of one or more cameras.

[0309] Changing the stored calibration data involves altering one or more parameters of a parameterized camera calibration function, which is defined for camera light corresponding to each pixel on at least one camera sensor, to reduce differences between the following:

[0310] (i) The calculated three-dimensional locations of multiple detected spots on the intraoral surface, and

[0311] (ii) Stored calibration values ​​that indicate the corresponding camera light for each pixel on the camera sensor, at which one of a plurality of detected spots should be detected.

[0312] For some applications, the selected stored calibration data includes stored calibration values ​​that indicate the projector light corresponding to each projection spot from one or more projectors, and wherein changing the stored calibration data includes changing:

[0313] (i) Assign each projector ray r to a list of indices of pixel path p, or

[0314] (ii) Define one or more parameters of the parameterized projector calibration model for each projector ray r.

[0315] For some applications, changing the stored calibration data involves altering the index list by reassigning each projector ray r based on the corresponding updated path p' for each pixel.

[0316] For some applications, changing the stored calibration data includes changing:

[0317] (i) Stored calibration values, which indicate the camera light intensity corresponding to each pixel on the camera sensors of one or more cameras, and

[0318] (ii) Stored calibration values ​​that indicate the projector ray r corresponding to each projection spot from one or more projectors.

[0319] For some applications, changing the stored calibration value involves iteratively changing the stored calibration value.

[0320] For some applications, the method also includes:

[0321] Drive each of one or more cameras to capture multiple images of a calibrated object with predetermined parameters;

[0322] Using a processor:

[0323] Run a triangulation algorithm to calculate the corresponding parameters of the calibrated object based on the captured image; and

[0324] Run the optimization algorithm:

[0325] Using (b) the corresponding parameters calculated based on the captured image of the calibrated object,

[0326] (a) Reduce the difference between (i) the updated path p' of the pixel corresponding to the projector ray r and (ii) the path p of the pixel corresponding to the projector ray r from the stored calibration value.

[0327] For some applications, the calibration object is a three-dimensional calibration object of known shape, and wherein driving each of one or more cameras to capture multiple images of the calibration object includes driving each of one or more cameras to capture images of the three-dimensional calibration object, and the predetermined parameter of the calibration object is the size of the three-dimensional calibration object.

[0328] For some applications, the calibration object is a two-dimensional calibration object with visually distinguishable features. Driving each of one or more cameras to capture multiple images of the calibration object includes driving each of one or more cameras to capture images of the two-dimensional calibration object, and wherein predetermined parameters of the two-dimensional calibration object are corresponding distances between corresponding visually distinguishable features.

[0329] According to some applications of the present invention, a method for calculating the three-dimensional structure of a three-dimensional surface within an oral cavity is also provided, the method comprising:

[0330] Scan the inner surface of the port;

[0331] Drive one or more uniform light projectors to project a broad spectrum of light onto a three-dimensional surface inside the mouth;

[0332] Drive the camera to capture multiple two-dimensional color images of the three-dimensional surface within the mouth; and

[0333] Using a processor:

[0334] The three-dimensional positions of multiple points on the three-dimensional surface of the oral cavity are calculated based on intraoral surface scanning.

[0335] Based on multiple two-dimensional color images of the intraoral three-dimensional surface, the three-dimensional structure of the intraoral three-dimensional surface is calculated, which is constrained by the three-dimensional position of multiple points.

[0336] In some embodiments, the intraoral surface is scanned by driving one or more structured light projectors to project a structured light pattern onto the intraoral three-dimensional surface, and

[0337] Drive one or more cameras to capture multiple structured light images, each image including at least a portion of a structured light pattern.

[0338] For some applications, driving one or more structured light projectors involves driving one or more structured light projectors to project a distribution of discrete, unconnected light spots.

[0339] For some applications, calculating 3D structures includes:

[0340] (a) Multiple two-dimensional color images of the intraoral three-dimensional surface, and (b) the calculated three-dimensional positions of multiple points on the intraoral three-dimensional surface, are input into a neural network; and

[0341] The neural network determines the corresponding predicted depth map of the intraoral three-dimensional surface captured in each two-dimensional color image.

[0342] For some applications, the method also includes using a processor to stitch the individual depth maps together to obtain a three-dimensional structure of the intraoral three-dimensional surface.

[0343] For some applications, the method also includes adjusting the capture of structured light images and two-dimensional color images to produce an alternating sequence of one or more structured light image frames interspersed with one or more broadband light image frames.

[0344] For some applications:

[0345] Driving one or more cameras to capture multiple structured light images includes driving each of two or more cameras to capture multiple structured light images, and

[0346] Driving a camera to capture multiple two-dimensional color images includes driving each of two or more cameras to capture multiple two-dimensional color images.

[0347] For some applications:

[0348] The determination by the neural network includes, for a given image frame, determining a corresponding predicted depth map of a portion of the intraoral three-dimensional surface captured in a two-dimensional color image by each of two or more cameras, and

[0349] The method also includes using a processor to stitch the individual depth maps together to obtain a predicted depth map of the intraoral three-dimensional surface captured in a given image frame.

[0350] For some applications, the method also includes training a neural network, which includes:

[0351] (a) Drive one or more structured light projectors to project a structured light pattern of the training phase onto a three-dimensional surface of the training phase.

[0352] (b) Drive one or more cameras in the training phase to capture multiple structured light images, each image including at least a portion of the structured light pattern of the training phase;

[0353] (c) Drive one or more uniform light projectors in the training phase to project a broad spectrum of light onto the three-dimensional surface of the training phase.

[0354] (d) Drive one or more cameras in the training phase to capture multiple two-dimensional color images of the three-dimensional surface of the training phase using illumination from a uniform light projector in the training phase;

[0355] (e) Adjusting the capture of structured light images and two-dimensional color images to produce an alternating sequence of image frames of one or more structured light images and image frames of one or more two-dimensional color images;

[0356] (f) Input multiple two-dimensional color images into a neural network;

[0357] (g) Based on the structured light image of the three-dimensional surface during the training phase, input the corresponding multiple three-dimensional reconstructions of the three-dimensional surface during the training phase into the neural network. The three-dimensional reconstructions include the calculated three-dimensional positions of multiple points on the three-dimensional surface during the training phase.

[0358] (h) For each two-dimensional color image frame, interpolate the position of one or more cameras relative to the three-dimensional surface of the training phase based on the calculated three-dimensional position of multiple points on the three-dimensional surface of the training phase, which is calculated based on the corresponding structured light image frames before and after each two-dimensional color image frame.

[0359] (i) Project the 3D reconstruction onto the corresponding field of view of each of the cameras in one or more training phases, and based on the projection, estimate the predicted depth map of the 3D surface seen in each 2D color image of the training phase, which is constrained by the calculated 3D position of multiple points.

[0360] (j) Compare each predicted depth map of the 3D surface during the training phase with the corresponding true depth map of the 3D surface during the training phase; and

[0361] (k) Based on the difference between each predicted depth map and the corresponding true depth map, optimize the neural network to better estimate subsequent predicted depth maps.

[0362] For some applications, driving one or more structured light projectors to project structured light patterns during the training phase involves driving one or more structured light projectors to each project a distribution of discrete, unconnected light spots onto a three-dimensional surface during the training phase.

[0363] For some applications, driving a camera for one or more training phases includes driving a camera for at least two training phases.

[0364] According to some applications of the present invention, an apparatus for intraoral scanning is also provided, the apparatus comprising:

[0365] A slender handheld stick, including a probe at the distal end of the handheld stick;

[0366] One or more lighting sources are coupled to the probe;

[0367] One or more near-infrared (NIR) light sources are coupled to the probe;

[0368] One or more cameras, coupled to the probe, are configured to (a) capture images using light from one or more illumination sources, and (b) capture images using NIR light from an NIR light source; and

[0369] The processor is configured to run a navigation algorithm to determine the position of the handheld stick as it moves in space. The inputs to the navigation algorithm are (a) an image captured using light from one or more illumination sources and (b) an image captured using NIR light.

[0370] For some applications, one or more lighting sources are one or more structured light sources.

[0371] For some applications, one or more lighting sources are one or more uniform light sources.

[0372] According to some applications of the present invention, a method for tracking the motion of an intraoral scanner is also provided, the method comprising:

[0373] The three-dimensional surface inside the mouth is illuminated using one or more illumination sources coupled to the intraoral scanner;

[0374] Using one or more near-infrared (NIR) light sources coupled to an intraoral scanner, each of the one or more NIR light sources is driven to emit NIR light onto a three-dimensional surface inside the mouth;

[0375] Using one or more cameras coupled to an intraoral scanner, (a) multiple images are captured using light from one or more illumination sources, and (b) multiple images are captured using NIR light;

[0376] Using a processor:

[0377] A navigation algorithm is run to track the motion of the intraoral scanner relative to the intraoral three-dimensional surface using (a) images captured using light from one or more illumination sources and (b) images captured using NIR light.

[0378] For some applications, one or more lighting sources are used, including one or more structured light sources, to illuminate the three-dimensional surface inside the mouth.

[0379] For some applications, the use of one or more lighting sources includes the use of one or more uniform light sources.

[0380] According to some applications of the present invention, an apparatus for intraoral scanning used with a cannula is also provided, the apparatus comprising:

[0381] A slender handheld rod, including a probe at the distal end of the handheld rod, the probe being configured to be removably disposed within a sleeve;

[0382] At least one structured light projector coupled to the probe, the structured light projector (a) having an illumination field of at least 30 degrees, (b) including a laser configured to emit polarized laser light, and (c) including a pattern generating optics configured to generate a light pattern when a laser diode is activated to emit light passing through the pattern generating optics; and

[0383] At least one camera coupled to the probe, the camera including a camera sensor,

[0384] The probe is configured such that light passes through the sleeve and enters the probe.

[0385] The laser is positioned at a certain distance relative to the camera, so that when the probe is placed inside the sleeve, a portion of the light pattern is reflected by the sleeve and reaches the camera sensor.

[0386] The laser is positioned at a rotational angle relative to its own optical axis, such that due to the polarization of the light pattern, the sleeve reflects less than 70% of the maximum reflection for all possible rotational angles of the laser relative to its optical axis.

[0387] For some applications, when the handheld stick is set in the sleeve, the distance between the structured light projector and the camera is 1-6 times the distance between the structured light projector and the sleeve.

[0388] For some applications, each of at least one camera has a field of view of at least 30 degrees.

[0389] For some applications, the laser is positioned at a rotational angle relative to its own optical axis, such that due to the polarization of the light pattern, the sleeve reflects less than 60% of the maximum reflection for all possible rotational angles of the laser relative to its optical axis.

[0390] For some applications, the laser is positioned at a rotational angle relative to its own optical axis, such that due to the polarization of the light pattern, the sleeve reflects a portion of the light pattern to a degree that is 15%–60% of the maximum reflection for all possible rotational angles of the laser relative to its optical axis.

[0391] According to some applications of the present invention, a method for generating three-dimensional images using an intraoral scanner is also provided, the method comprising:

[0392] (A) Using at least two cameras rigidly connected to an intraoral scanner, such that the respective fields of view of each camera have non-overlapping portions:

[0393] Capture multiple images of the three-dimensional surface within the mouth; and

[0394] (B) Using a processor:

[0395] Simultaneous Localization and Mapping (SLAM) algorithm is run using images captured from each camera for the non-overlapping portion of the corresponding field of view. The localization of each camera is solved based on the fact that the motion of each camera is the same as the motion of every other camera.

[0396] For some applications:

[0397] The fields of view of the first and second cameras in the camera system also overlap.

[0398] The capture includes capturing multiple images of the intraoral three-dimensional surface, such that features of the intraoral three-dimensional surface in the overlapping portion of the respective fields of view appear in the images captured by the first and second cameras, and

[0399] The processor used includes running the SLAM algorithm using features of the intraoral 3D surface that appears in images from at least two cameras.

[0400] According to some applications of the present invention, a method for generating three-dimensional images using an intraoral scanner is also provided, the method comprising:

[0401] Drive one or more structured light projectors to project structured light patterns onto a three-dimensional surface inside the mouth;

[0402] Drive one or more cameras to capture multiple structured light images, each structured light image including at least a portion of a structured light pattern;

[0403] Drive one or more uniform light projectors to project a broad spectrum of light onto a three-dimensional surface inside the mouth;

[0404] Drive at least one camera to capture two-dimensional color images of the three-dimensional surface inside the mouth using illumination from a uniform light projector;

[0405] Adjusting the capture of structured light and broadband light to generate an alternating sequence of one or more structured light image frames interspersed with one or more broadband light image frames; and

[0406] Using a processor:

[0407] Calculate the corresponding 3D positions of multiple points on the intraoral 3D surface captured in one or more structured light image frames.

[0408] Based on the calculated 3D positions of multiple points in corresponding structured light image frames before and after the broadband light image frame, the motion of at least one camera is interpolated between the first and second broadband light image frames.

[0409] Simultaneous Localization and Mapping (SLAM) algorithm is run, which (a) uses features of the intraoral three-dimensional surface captured by at least one camera in a first and second broadband light image frame, and (b) is constrained by the interpolation motion of the camera between the first and second broadband light image frames.

[0410] For some applications, driving one or more structured light projectors to project structured light patterns includes driving one or more structured light projectors to each project a distribution of discrete, unconnected light spots onto a three-dimensional surface within the mouth.

[0411] According to some applications of the present invention, a method for generating three-dimensional images using an intraoral scanner is also provided, the method comprising:

[0412] Drive one or more structured light projectors to project structured light patterns onto a three-dimensional surface inside the mouth;

[0413] Drive one or more cameras to capture multiple structured light images, each structured light image including at least a portion of a structured light pattern;

[0414] Drive one or more uniform light projectors to project a broad spectrum of light onto a three-dimensional surface inside the mouth;

[0415] Drive one or more cameras to capture two-dimensional color images of the three-dimensional surface inside the mouth using illumination from a uniform light projector;

[0416] Adjusting the capture of structured light and broadband light to generate an alternating sequence of one or more structured light image frames interspersed with one or more broadband light image frames; and

[0417] Using a processor:

[0418] (a) The three-dimensional position of a feature on a three-dimensional surface within the mouth is calculated based on a structured light image frame, which is also captured in the first and second broadband light image frames.

[0419] (b) Based on the calculated 3D position of the features, calculate the motion of at least one camera between the first and second broadband light image frames, and

[0420] (c) Using (i) features of the intraoral three-dimensional surface captured by at least one camera in the first and second broadband light image frames, without calculating the three-dimensional position of the features based on the structured light image frames, and (ii) the calculated motion of the camera between the first and second broadband light image frames, run the Simultaneous Localization and Mapping (SLAM) algorithm.

[0421] For some applications, driving one or more structured light projectors to project structured light patterns includes driving one or more structured light projectors to each project a distribution of discrete, unconnected light spots onto a three-dimensional surface within the mouth.

[0422] According to some applications of the present invention, a method for calculating the three-dimensional structure of intraoral three-dimensional surfaces within the oral cavity of an object is also provided, the method comprising:

[0423] Drive one or more structured light projectors to project a structured light pattern of spots onto a three-dimensional surface inside the mouth;

[0424] Drive one or more cameras to capture multiple structured light images, each image including at least one spot;

[0425] Drive one or more uniform light projectors to project a broad spectrum of light onto a three-dimensional surface inside the mouth;

[0426] Drive at least one camera to capture two-dimensional color images of the three-dimensional surface inside the mouth using illumination from a uniform light projector;

[0427] Adjusting the capture of structured light and broadband light to generate an alternating sequence of one or more structured light image frames interspersed with one or more broadband light image frames; and

[0428] Using a processor:

[0429] Based on two-dimensional color images, determine whether, for each of a plurality of spots, the spot is being projected onto moving or stable tissue within the oral cavity.

[0430] Based on this determination, a corresponding confidence level is assigned to each of the multiple detected spots: high confidence for fixed tissue and low confidence for moving tissue.

[0431] A 3D reconstruction algorithm is run using the detected blobs based on the confidence level of each of the multiple detected blobs.

[0432] For some applications, driving one or more structured light projectors to project structured light patterns includes driving one or more structured light projectors to each project a distribution of discrete, unconnected light spots onto a three-dimensional surface within the mouth.

[0433] For some applications, running a 3D reconstruction algorithm involves running the algorithm using only a subset of the detected spots, which consists of spots assigned a confidence level above a fixed tissue threshold.

[0434] For some applications, running a 3D reconstruction algorithm includes (a) assigning weights to each blob based on the corresponding confidence level assigned to that blob, and (b) using the corresponding weights of each blob in the 3D reconstruction algorithm. Attached Figure Description

[0435] The invention will be more fully understood through the following detailed description of its application in conjunction with the accompanying drawings, wherein:

[0436] Figure 1 This is a schematic diagram of a handheld stick having multiple structured light projectors and cameras disposed in a probe at the distal end of the handheld stick according to some applications of the present invention.

[0437] Figure 2A -B are schematic diagrams of the positioning configuration of a camera and the positioning configuration of a structured light projector, respectively, for some applications according to the present invention.

[0438] Figure 2C It is a diagram depicting various different configurations of the structured light projector and camera in the probe for some applications according to the present invention;

[0439] Figure 2D -E is an isometric illustration of a specific configuration of the position of the structured light projector and camera in the probe, shown from two different corresponding perspectives, according to some applications of the present invention;

[0440] Figure 3 This is a schematic diagram of a structured light projector for some applications according to the present invention;

[0441] Figure 4 This is a schematic diagram of a structured light projector according to some applications of the present invention projecting the distribution of discrete, unconnected light spots onto the focal plane of multiple objects.

[0442] Figure 5A -B is a schematic diagram of a structured light projector according to some applications of the present invention, the structured light projector including a beam shaping optical element and an additional optical element disposed between the beam shaping optical element and a pattern generating optical element;

[0443] Figure 6A -B is a schematic diagram of a structured light projector projecting discrete, unconnected spots and a camera sensor detecting the spots in some applications according to the present invention.

[0444] Figure 7 This is a flowchart outlining a method for generating digital three-dimensional images according to some applications of the present invention;

[0445] Figure 8 This is an overview of some applications of the present invention for performing Figure 7 A flowchart of a specific step in a method;

[0446] Figure 9 , 10 11 and 12 are depictions of some applications according to the present invention. Figure 8 A simplified example of the steps is illustrated in the diagram.

[0447] Figure 13 This is a flowchart outlining other steps in a method for generating digital three-dimensional images according to some applications of the present invention;

[0448] Figure 14 , 15 16 and 17 are depictions of some applications according to the present invention. Figure 13 A simplified example of the steps is illustrated in the diagram.

[0449] Figure 18This is a schematic diagram of a probe including a diffuse reflector for some applications according to the present invention;

[0450] Figure 19A -B is a schematic diagram of a structured light projector for some applications of the present invention and a schematic diagram of the cross-section of a beam emitted by a laser diode, wherein the pattern generating optical element shown is disposed in the optical path of the beam;

[0451] Figure 20A -E is a schematic diagram of a microlens array used as a pattern generating optical element in a structured light projector according to some applications of the present invention;

[0452] Figure 21A -C is a schematic diagram of a composite two-dimensional diffraction periodic structure used as a pattern generating optical element in a structured light projector according to some applications of the present invention.

[0453] Figure 22A -B is a schematic diagram illustrating a single optical element having an aspherical first side and a planar second side opposite to the first side, and a schematic diagram of a structured light projector including the optical element, according to some applications of the present invention.

[0454] Figure 23A -B is a schematic diagram of an axial cone lens for some applications according to the present invention and a schematic diagram of a structured light projector including an axial cone lens;

[0455] Figure 24A -B is a schematic diagram illustrating an optical element having an aspherical surface on a first side and a planar surface on a second side opposite to the first side, according to some applications of the invention, and a schematic diagram of a structured light projector including the optical element;

[0456] Figure 25 This is a schematic diagram of a single optical element in a structured light projector according to some applications of the present invention;

[0457] Figure 26A -B is a schematic diagram of a structured light projector having more than one laser diode according to some applications of the present invention;

[0458] Figure 27A -B is a schematic diagram illustrating different ways of combining laser diodes of different wavelengths according to some applications of the present invention;

[0459] Figure 28 This is a flowchart outlining the steps of a "spot tracking" method according to some applications of the present invention;

[0460] Figure 29 This is a simplified example depicting some detected spots in some applications according to the present invention, and a schematic diagram of how a processor can determine which sets of detected spots can be considered to be tracked;

[0461] Figure 30 This is a flowchart outlining a method for determining a tracked spot according to some applications of the present invention;

[0462] Figures 31-32 This is a flowchart outlining various methods for finding tracked spots in subsequent images according to some applications of the present invention;

[0463] Figure 33 This is a schematic diagram illustrating an example of how spot tracking, according to some applications of the present invention, helps to identify detected spots as being projected from a particular projector beam;

[0464] Figure 34A -B is a simplified schematic diagram of a camera sensor according to some applications of the present invention, showing two detected spots;

[0465] Figures 35-36 This is a flowchart outlining corresponding methods that can be used with blob tracking for some applications according to the present invention;

[0466] Figure 37A -B is a schematic diagram illustrating points used for 3D reconstruction before and after the processor has implemented blob tracking, according to some applications of the present invention;

[0467] Figure 38 This is a flowchart outlining the steps of a method for generating digital three-dimensional images (hereinafter referred to as "ray tracing") according to some applications of the present invention;

[0468] Figure 39 This is a diagram illustrating the length of a projector ray tracked over time in some applications according to the present invention, and a specific simplified view of a camera sensor corresponding to a particular image frame;

[0469] Figure 40A -B is a graph showing experimental datasets before and after ray tracing for some applications according to the present invention;

[0470] Figure 41 This is a schematic diagram of a projector with multiple camera sensors and projection spots for some applications according to the present invention;

[0471] Figure 42A -B shows a flowchart outlining a method for generating three-dimensional images according to some applications of the present invention;

[0472] Figure 43A -B is a flowchart outlining various methods for tracking the movement of an intraoral scanner according to some applications of the present invention; and

[0473] Figure 44A ,44B 45 are schematic diagrams illustrating a simplified view of a camera image with multiple detected spots from a single projector beam, which are captured at corresponding times and superimposed on the same image, according to some applications of the present invention.

[0474] Figure 46A -D illustrates a simplified scenario where the processor identifier, according to some applications of the present invention, should be recalibrated for the projector (and not any camera);

[0475] Figure 47A -B illustrates a simplified scenario where, according to some applications of the present invention, the processor identifier should recalibrate the camera (and not any projector);

[0476] Figure 48A -B illustrates a simplified scenario in which the processor of some applications according to the present invention cannot reasonably assume that the offset occurs only in the camera or only in the projector;

[0477] Figure 49A -B are schematic diagrams of three-dimensional and two-dimensional calibrated objects according to some applications of the present invention;

[0478] Figure 50 This is a flowchart depicting a method for tracking the movement of a handheld stick according to some applications of the present invention;

[0479] Figure 51A -F is a flowchart depicting a method for calculating the three-dimensional structure of an intraoral three-dimensional surface according to some applications of the present invention;

[0480] Figure 51G -I is a schematic diagram graphically depicting different combinations of neural network inputs for some applications of the present invention;

[0481] Figure 52A This is a flowchart depicting a method for training a neural network according to some applications of the present invention;

[0482] Figure 52B This is a block diagram of training a neural network according to some applications of the present invention;

[0483] Figure 52C This is a flowchart depicting a method for outputting a depth map of a neural network and a corresponding confidence map in some applications according to the present invention;

[0484] Figure 52D -F is a schematic diagram depicting a trained neural network according to some applications of the present invention to output a depth map and a corresponding confidence map;

[0485] Figure 52GThis is a flowchart depicting how confidence graphs can be used in some applications according to the present invention;

[0486] Figure 53A This is a schematic diagram of a disposable cannula placed on the distal end of an intraoral scanner before the probe is placed inside the patient's mouth, in order to prevent cross-contamination between patients according to some applications of the present invention.

[0487] Figure 53B This is a graph illustrating the reflectivity of polarized lasers according to Fresnel equations and their use in some applications according to the present invention;

[0488] Figure 54A -B are flowcharts depicting methods for generating three-dimensional images using a handheld stick, and schematic diagrams illustrating the positioning of a projector and a camera, respectively, according to some applications of the present invention.

[0489] Figure 55 This is a flowchart describing a method for generating three-dimensional images using a handheld stick according to some applications of the present invention;

[0490] Figure 56A -B are flowcharts describing a method for generating three-dimensional images using a handheld stick, and schematic diagrams of two image frames of unstructured light and two features of the three-dimensional surface inside the mouth, respectively, according to some applications of the present invention.

[0491] Figure 57 This is a flowchart depicting a method for calculating the three-dimensional structure of an intraoral three-dimensional surface within the oral cavity of an object, according to some applications of the present invention;

[0492] Figure 58 and 59A -B is a schematic diagram of a neural network according to some applications of the present invention;

[0493] Figure 60 An embodiment of a system for performing intraoral scanning and generating a virtual 3D model of the dental arch is shown;

[0494] Figure 61 A block diagram of an example computing device according to an embodiment of the present disclosure is shown;

[0495] Figure 62 This is a flowchart depicting a method for overcoming manufacturing deviations between intraoral scanners according to some applications of the present invention;

[0496] Figure 63 This is a schematic diagram of another method for overcoming manufacturing deviations between intraoral scanners according to some applications of the present invention;

[0497] Figure 64This is a flowchart depicting a method for testing whether cropping and deformation of each runtime image accurately addresses possible manufacturing deviations of a given intraoral scanner, and if not, training a scan refinement neural network based on the local refinement stage of that given intraoral scanner, according to some applications of the present invention.

[0498] Figure 65 This is a flowchart depicting a method for overcoming manufacturing deviations between intraoral scanners according to some applications of the present invention; and

[0499] Figure 66 This is a flowchart depicting a method for training a neural network according to some applications of the present invention.

[0500] Specific implementation method

[0501] Now for reference Figure 1 This is a schematic diagram of an elongated handheld rod 20 for intraoral scanning according to some applications of the present invention. A plurality of structured light projectors 22 and a plurality of cameras 24 are coupled to a rigid structure 26 disposed within a probe 28 at the distal end 30 of the handheld rod. In some applications, the probe 28 enters the patient's oral cavity during intraoral scanning.

[0502] For some applications, the structured light projectors 22 are positioned within the probe 28 such that each structured light projector 22 faces an object 32 placed in the projector's illumination field outside the handheld stick 20, rather than positioning the structured light projector near the handheld stick and illuminating the object by reflecting light from a mirror and then onto the object. Similarly, for some applications, the cameras 24 are positioned within the probe 28 such that each camera 24 faces an object 32 placed in the camera's field of view outside the handheld stick 20, rather than positioning the camera near the handheld stick and observing the object by reflecting light from a mirror and entering the camera. This positioning of the projectors and cameras within the probe 28 allows the scanner to have a large overall field of view while maintaining a low-profile probe.

[0503] In some applications, the height H1 of the probe 28 is less than 15 mm. The height H1 of the probe 28 is measured from the lower surface 176 (sensing surface) to the upper surface 178 opposite to the lower surface 176. Reflected light from the scanned object 32 enters the probe 28 through the lower surface 176. In some applications, the height H1 is between 10 and 15 mm.

[0504] In some applications, each of the cameras 24 has a large field of view β(beta) of at least 45 degrees (e.g., at least 70 degrees, e.g., at least 80 degrees, e.g., 85 degrees). In some applications, the field of view may be less than 120 degrees, e.g., less than 100 degrees, e.g., less than 90 degrees. In experiments conducted by the inventors, a field of view β(beta) of 80 to 90 degrees for each camera was found to be particularly useful, as this provides a good balance between pixel size, field of view and camera overlap, optical quality, and cost. The camera 24 may include a camera sensor 58 and an objective optics 60 comprising one or more lenses. To achieve near-focus imaging, the camera 24 may focus on an object focal plane 50 located between 1 mm and 30 mm, e.g., between 4 mm and 24 mm, e.g., between 5 mm and 11 mm, e.g., between 9 mm and 10 mm, at the distance from the lens furthest from the camera sensor. In experiments conducted by the inventors, it was found particularly useful that the object's focal plane 50 is located between 5mm and 11mm from the lens furthest from the camera sensor, because teeth are easily scanned at this distance and because most tooth surfaces are well focused. In some applications, the camera 24 can capture images at a frame rate of at least 30 frames per second, for example, at least 75 frames per second, or at least 100 frames per second. In some applications, the frame rate can be lower than 200 frames per second.

[0505] As mentioned above, a large field of view achieved by combining the respective fields of view of all cameras can improve accuracy due to a reduction in the amount of image stitching error, especially in edentulous areas where high-resolution 3D features such as smooth and clear gingival surfaces may be less abundant. A larger field of view allows large, smooth features (e.g., the overall curve of the teeth) to appear in each image frame, which improves the accuracy of stitching the corresponding surfaces obtained from multiple such image frames.

[0506] Similarly, each of the structured light projectors 22 may have a large illumination field α (alph) of at least 45 degrees (e.g., at least 70 degrees). In some applications, the illumination field α (alph) may be less than 120 degrees, for example, less than 100 degrees. Other features of the structured light projector 22 are described below.

[0507] For some applications, to improve image capture, each camera 24 has multiple discrete preset focus positions, at each focus position the camera focuses on the corresponding object focal plane 50. Each camera 24 may include an autofocus actuator that selects a focus position from the discrete preset focus positions to improve a given image capture. Additionally or alternatively, each camera 24 includes an optical aperture phase mask that extends the depth of focus of the camera, such that the image formed by each camera remains in focus at all object distances between 1 mm and 30 mm (e.g., between 4 mm and 24 mm, e.g., between 5 mm and 11 mm, e.g., 9 mm–10 mm) from the lens furthest from the camera sensor.

[0508] In some applications, the structured light projector 22 and camera 24 are coupled to the rigid structure 26 in a closely packed and / or alternating manner, such that (a) a large portion of the field of view of each camera overlaps with the field of view of a neighboring camera, and (b) a large portion of the field of view of each camera overlaps with the illumination field of a neighboring projector. Optionally, at least 20% (e.g., at least 50%, e.g., at least 75%) of the projected light pattern is in the field of view of at least one camera at an object focal plane 50 located at least 4 mm from the lens furthest from the camera sensor. Due to the different possible configurations of the projector and camera, some projected patterns may never be seen in the field of view of any camera, and some projected patterns may be blocked from the field of view by the object 32 as the scanner moves back and forth during scanning.

[0509] The rigid structure 26 can be a non-flexible structure to which the structured light projector 22 and camera 24 are coupled to provide structural stability for the optics within the probe 28. Coupling all projectors and cameras to a common rigid structure helps maintain the geometrical integrity of the optics of each structured light projector 22 and each camera 24 under varying environmental conditions, such as mechanical stresses that may be caused by the object's mouth. Furthermore, the rigid structure 26 helps maintain stable structural integrity and the positioning of the structured light projectors 22 and cameras 24 relative to each other. As further described below, controlling the temperature of the rigid structure 26 can help maintain the geometrical integrity of the optics over a wide range of ambient temperatures when the probe 28 enters and exits the object's mouth or when the object breathes during scanning.

[0510] Now for reference Figure 2A-B, which are schematic diagrams of the positioning configuration of the camera 24 and the structured light projector 22, respectively, for some applications according to the present invention. For some applications, in order to improve the overall field of view and illumination field of the intraoral scanner, the camera 24 and the structured light projector 22 are positioned such that they do not both face the same direction. For some applications, for example... Figure 2A As shown, multiple cameras 24 are coupled to a rigid structure 26 such that the angle θ (theta) between two corresponding optical axes 46 of at least two cameras 24 is 90 degrees or less, for example, 35 degrees or less. Similarly, for some applications, such as Figure 2B As shown, a plurality of structured light projectors 22 are coupled to a rigid structure 26 such that the angle between two corresponding optical axes 48 of at least two structured light projectors 22 is such that... It is 90 degrees or less, for example, 35 degrees or less.

[0511] Now for reference Figure 2C This is a diagram depicting various different configurations of the structured light projector 22 and camera 24 within the probe 28 for some applications according to the present invention. The structured light projector 22... Figure 2C The center is represented by a circle, and the camera 24 is... Figure 2C The rectangle represents the camera. Note that the rectangle is used to represent the camera because each camera sensor 58 and field of view β (beta) of each camera 24 typically has an aspect ratio of 1:2. Figure 2C Column (a) shows a bird's-eye view of various configurations of the structured light projector 22 and camera 24. The x-axis marked in the first row of column (a) corresponds to the central longitudinal axis of the probe 28. Column (b) shows a side view of the various configurations of the camera 24 as viewed from a line of sight coaxial with the central longitudinal axis of the probe 28. Figure 2A Similar to the example shown, Figure 2C Column (b) shows the camera 24 positioned such that the optical axis 46 is at an angle of 90 degrees or less (e.g., 35 degrees or less) relative to each other. Column (c) shows a side view of the camera 24 in various configurations when viewed from a line of sight perpendicular to the central longitudinal axis of the probe 28.

[0512] Typically, the farthest end (facing) Figure 2C (in the positive x direction) and the nearest end (towards) Figure 2CCamera 24 (in the negative x direction) is positioned such that its optical axis 46 is slightly rotated inward relative to the next nearest camera 24, for example, at an angle of 90 degrees or less (e.g., 35 degrees or less). A more central camera 24 (i.e., neither the farthest nor the closest camera 24) is positioned such that it faces directly outward from the probe, with its optical axis 46 substantially perpendicular to the central longitudinal axis of the probe 28. It should be noted that in row (xi), projector 22 is located at the farthest point of the probe 28, and therefore the optical axis 48 of projector 22 points inward, thus allowing a greater number of spots 33 projected from that particular projector 22 to be seen by more cameras 24.

[0513] Typically, the number of structured light projectors 22 in probe 28 can range from two (e.g., as...). Figure 2C The number of cameras 24 in probe 28 can typically range from four (e.g., as shown in rows (iv) and (v)) to seven (e.g., as shown in row (ix)). Note that Figure 2C The various configurations shown are exemplary and not limiting, and therefore the scope of the invention includes additional configurations not shown. For example, the scope of the invention includes more than five projectors 22 and more than seven cameras located in probe 28.

[0514] Now for reference Figure 2D -E, which is an isometric illustration of the specific configuration of the positions of the structured light projector 22 and camera 24 in the probe 28, shown from two different corresponding perspectives according to some applications of the invention. Figure 2D From and Figure 2C The same bird's-eye view is shown in column (a). For some applications, there are six cameras 24 evenly spaced within the probe 28 (with three cameras on each side of the probe 28), and five structured light projectors 22 centrally positioned within the probe 28 along the central longitudinal axis of the probe 28 (shown by dashed line 29).

[0515] For some applications, both camera 24 and structured light projector 22 are coupled to a flexible printed circuit board (PCB) to accommodate their angular positioning within probe 28. This angular positioning of camera 24 and structured light projector 22... Figure 2EAs shown in the diagram, the farthest (i.e., facing the positive x-direction) camera 24 and structured light projector 22 are positioned such that their respective optical axes are tilted backward toward the handheld stick 20 at an angle of 45 degrees or less (e.g., 35 degrees or less). This allows the farthest camera to capture the posterior wall of the posterior molars within the oral cavity. The nearest (i.e., facing the negative x-direction) camera 24 and structured light projector 22 are positioned such that their respective optical axes are tilted forward toward the distal end of the probe 28 at an angle of 45 degrees or less (e.g., 35 degrees or less) to achieve improved overlap of the respective fields of view of the camera 24. Furthermore, all structured light projectors 22 are positioned such that their respective optical axes are tilted toward the center of the probe 28, which improves the overlap of the individual illumination fields of the structured light projectors 22. The inventors have realized that positioning the structured light projectors 22 substantially all in a single line allows for easier connection of them all to the same flexible PCB.

[0516] In addition Figure 2D Shown in E and further described below are a plurality of uniform light projectors 118, a plurality of near-infrared (NIR) light projectors 292, and a diffractive optical element (DOE) 39 disposed above each structured light projector 22.

[0517] Now for reference Figure 3 This is a schematic diagram of a structured light projector 22 according to some applications of the present invention. In some applications, the structured light projector 22 includes a laser diode 36, a beam shaping optics element 40, and a pattern generating optics element 38, which generates a discrete, unconnected distribution 34 of light spots (refer to below). Figure 4 (Further discussion). In some applications, the structured light projector 22 can be configured to generate a distribution 34 of discrete, unconnected light spots on all planes located between 1 mm and 30 mm (e.g., 4 mm to 24 mm) from the pattern-generating optics 38 when the laser diode 36 emits light. For some applications, the distribution 34 of discrete, unconnected light spots is focused on one plane located between 1 mm and 30 mm (e.g., 4 mm to 24 mm), while all other planes located between 1 mm and 30 mm (e.g., 4 mm to 24 mm) still contain discrete, unconnected light spots. Although the above description uses a laser diode, it should be understood that this is an exemplary and non-limiting application. Other light sources can be used in other applications. Furthermore, although described as projecting a pattern of discrete, unconnected light spots, it should be understood that this is an exemplary and non-limiting application. Other light patterns or arrays, including but not limited to lines, grids, checkerboards, and other arrays, can be used in other applications. In some applications, the light pattern projected by the structured light projector is spatially fixed relative to one or more cameras.

[0518] This document describes embodiments with reference to discrete light spots and operations performed using or based on these spots. Examples of these operations include solving correspondence algorithms to determine the location of the light spot, tracking the light spot, mapping projector rays to the light spot, identifying faint light spots, and generating 3D models based on the location of the spots. It should be understood that these and other operations described with reference to the spots are also applicable to other features of other projection light patterns. Therefore, the discussion of reference to spots in this document also applies to any other feature of projection light patterns.

[0519] The pattern generating optical element 38 can be configured to have a light conversion efficiency of at least 80% (e.g., at least 90%) (i.e., the proportion of light entering the pattern out of the total light falling on the pattern generating optical element 38).

[0520] For some applications, the corresponding laser diodes 36 of each structured light projector 22 emit light at different wavelengths; that is, the corresponding laser diodes 36 of at least two structured light projectors 22 emit light at two different wavelengths, respectively. For some applications, the corresponding laser diodes 36 of at least three structured light projectors 22 emit light at three different wavelengths, respectively. For example, red, blue, and green laser diodes can be used. For some applications, the corresponding laser diodes 36 of at least two structured light projectors 22 emit light at two different wavelengths, respectively. For example, in some applications, six structured light projectors 22 are arranged within the probe 28, three of which contain blue laser diodes and three of which contain green laser diodes.

[0521] Now for reference Figure 4 This is a schematic diagram of a structured light projector 22, according to some applications of the present invention, projecting a distribution of discrete, unconnected light spots onto the focal plane of multiple objects. The scanned objects 32 may be one or more teeth or other intraoral objects / tissues within the mouth. Some translucent and glossy properties of teeth may affect the contrast of the structured light pattern being projected. For example, (a) some light illuminating the teeth may scatter into other areas of the intraoral scene, resulting in a certain amount of stray light, and (b) some light may penetrate the teeth and subsequently exit from the teeth at any other point. Therefore, in order to improve image capture of intraoral scenes under structured light illumination without using contrast enhancement means (e.g., using opaque powder to coat the teeth), the inventors have realized that a sparse distribution 34 of discrete, unconnected light spots can provide an improved balance between reducing the amount of projected light and maintaining the amount of useful information. The sparsity of the distribution 34 can be characterized by the ratio of the following:

[0522] (a) The illumination area on the orthogonal plane 44 in the illumination field α(alpha), that is, the sum of the areas of all projected spots 33 on the orthogonal plane 44 in the illumination field α(alpha), and

[0523] (b) The non-illuminated area on the orthogonal plane 44 in the illumination field α (alpha). In some applications, the sparsity ratio may be at least 1:150 and / or less than 1:16 (e.g., at least 1:64 and / or less than 1:36).

[0524] In some applications, each structured light projector 22 projects at least 400 discrete, unconnected spots 33 onto the intraoral 3D surface during scanning. In other applications, each structured light projector 22 projects fewer than 3000 discrete, unconnected spots 33 onto the intraoral surface during scanning. To reconstruct the 3D surface from the sparse distribution 34 of the projections, the correspondence between the corresponding projected spots 33 (or other features of the projection pattern) and the spots (or other features) detected by the camera 24 must be determined, as referenced below. Figure 7-1 9. Further description.

[0525] For some applications, pattern generating optical element 38 is a diffractive optical element (DOE) 39. Figure 3 When the laser diode 36 emits light through the DOE 39 onto the object 32, the element generates a distribution 34 of discrete, unconnected spots 33. As used herein throughout this application, including the claims, the spot is defined as a small area of ​​light having any shape. For some applications, the corresponding DOE 39 of different structured light projectors 22 generates spots with different corresponding shapes; that is, each spot 33 generated by a particular DOE 39 has the same shape, while the shape of the spot 33 generated by at least one DOE 39 is different from the shape of the spot 33 generated by at least one other DOE 39. For example, some DOEs 39 may generate circular spots 33 (e.g., circular spots 33). Figure 4 As shown, some DOE 39s may generate square spots, while others may generate elliptical spots. Optionally, some DOE 39s may generate connected or disconnected line patterns.

[0526] Now for reference Figure 5A-B, which is a schematic diagram of a structured light projector 22 according to some applications of the present invention, the structured light projector 22 including a beam-shaping optics element 40 and an additional optics element disposed between the beam-shaping optics element 40 and a pattern-generating optics element 38 (e.g., DOE 39). Optionally, the beam-shaping optics element 40 is a collimating lens 130. The collimating lens 130 may be configured to have a focal length of less than 2 mm. Optionally, the focal length may be at least 1.2 mm. For some applications, when the laser diode 36 emits light through the optics element 42, the additional optics element 42 disposed between the beam-shaping optics element 40 and the pattern-generating optics element 38 (e.g., DOE 39) generates a Bessel beam. In some applications, the Bessel beam passes through the DOE 39 such that all discrete, unconnected spots 33 maintain small diameters (e.g., less than 0.06 mm, less than 0.04 mm, less than 0.02 mm) passing through a series of orthogonal planes 44 (e.g., each orthogonal plane located between 1 mm and 30 mm from the DOE 39, or between 4 mm and 24 mm from the DOE 39, etc.). In the context of this patent application, the diameter of the spot 33 is defined by the full width at half maximum (FWHM) of the spot intensity.

[0527] Although all the aforementioned spots are less than 0.06 mm, some points near the upper end of these ranges (e.g., only slightly less than 0.06 mm or 0.02 mm) and also near the edge of the illumination field of projector 22 may be elongated when they intersect a geometric plane orthogonal to DOE 39. In this case, it is useful to measure their diameters when they intersect the inner surface of a geometric sphere centered on DOE 39 and having a radius between 1 mm and 30 mm, corresponding to a distance from the corresponding orthogonal plane of DOE 39 between 1 mm and 30 mm. As used throughout this application, including the claims, the word “geometry” is considered to be related to theoretical geometric constructions (e.g., planes or spheres) and is not part of any physical device.

[0528] For some applications, when a Bessel beam passes through a DOE 39, in addition to spots with a diameter less than 0.06 mm, spots 33 with a diameter greater than 0.06 mm are also generated.

[0529] For some applications, optical element 42 is an axonoconical lens 45, for example in... Figure 5A As shown and referenced below Figure 23A -B further description. Alternatively, optical element 42 can be an annular aperture ring 47, for example... Figure 5BAs shown. Maintaining a small-diameter spot improves 3D resolution and accuracy throughout the depth of focus. Without optical elements 42 (e.g., axial-cone lens 45 or annular aperture ring 47), the spot size of one or more spots 33 may change, for example, become larger, as they move further away from the optimal focal plane due to diffraction and defocus.

[0530] Now for reference Figure 6A -B, which is a schematic diagram of a structured light projector 22 projecting discrete, unconnected spots 33 and a camera sensor 58 detecting spots 33' according to some applications of the present invention. For some applications, a method is provided for determining the correspondence between projected spots 33 on an intraoral surface and detected spots 33' on the corresponding camera sensor 58. As previously described, the method is also applicable to determining the correspondence between other projected features on an intraoral surface and detected features on the corresponding camera sensor. Once the correspondence is determined, a three-dimensional image of the surface is reconstructed. Each camera sensor 58 has a pixel array, and for each pixel, there is a corresponding camera ray 86. Similarly, for each projected spot 33 from each projector 22, there is a corresponding projector ray 88. Each projector ray 88 corresponds to a corresponding path 92 of a pixel on at least one camera sensor 58. Therefore, if the camera sees a spot 33' projected by a particular projector ray 88, then that spot 33' will necessarily be detected by a pixel on the specific pixel path 92 corresponding to that particular projector ray 88. See details Figure 6B The diagram illustrates the correspondence between each projector ray 88 and its corresponding camera sensor path 92. Projector ray 88' corresponds to camera sensor path 92', projector ray 88" corresponds to camera sensor path 92", and projector ray 88"' corresponds to camera sensor path 92"'. For example, if a particular projector ray 88 projects a spot onto a dusty space, it will illuminate a dust line in the air. The dust line detected by camera sensor 58 will follow the same path on camera sensor 58 as the camera sensor path 92 corresponding to that particular projector ray 88.

[0531] During calibration, calibration values ​​are stored based on camera rays 86 corresponding to pixels on camera sensors 58 of each camera 24 and projector rays 88 corresponding to projection spots 33 (or other features) from each structured projector 22. For example, calibration values ​​can be stored for (a) multiple camera rays 86 corresponding to corresponding multiple pixels on camera sensors 58 of each camera 24 and (b) multiple projector rays 88 corresponding to corresponding multiple projection spots 33 from each structured light projector 22. As used throughout this application, including the claims, indicating the stored calibration value corresponding to each pixel on the camera sensor of each camera means (a) a value assigned to each camera ray, or (b) a parameter value of a parameterized camera calibration model (e.g., a function). As used throughout this application, including the claims, indicating the stored calibration value corresponding to each projection spot (or other projection feature) from each structured light projector means (a) a value assigned to each projector ray, for example, in an indexed list, or (b) a parameter value of a parameterized projector calibration model (e.g., a function).

[0532] As an example, the following calibration process can be used. A high-accuracy point target (e.g., a black dot on a white background) is illuminated from below, and images of the target are captured using all cameras. The point target is then moved perpendicularly (i.e., along the z-axis) toward the camera onto the target plane. Point centers are calculated for all points at all corresponding z-axis positions to create a 3D grid of points in space. The pixel coordinates of each 3D position of the corresponding point center are then found using a distortion and camera pinhole model, thus defining the camera ray for each pixel as a ray originating from a pixel oriented toward the corresponding point center in the 3D grid. Interpolation can be performed on the camera rays corresponding to pixels between grid points. The above camera calibration process is repeated for all corresponding wavelengths of the respective laser diodes 36, such that the stored calibration values ​​include the camera ray 86 corresponding to each pixel on each camera sensor 58 for each wavelength. Alternatively, the stored calibration values ​​are parameter values ​​of the distortion and camera pinhole model, indicating the value of the camera ray 86 corresponding to each pixel on each camera sensor 58 for each wavelength.

[0533] After camera 24 has been calibrated and all camera ray values ​​86 have been stored, structured light projector 22 can be calibrated as follows: A flat, featureless target is used, and one structured light projector 22 is turned on at a time. Each spot (or feature) is located on at least one camera sensor 58. Since camera 24 is now calibrated, the three-dimensional spot position of each spot (or feature) is calculated by triangulation based on images of that spot (or feature) in multiple different cameras. The above process is repeated for featureless targets located at multiple different z-axis positions. Each projected spot (or feature) on the featureless target will define the projector rays originating from the projector in space.

[0534] Now for reference Figure 7 This is a flowchart outlining a method for generating digital three-dimensional images according to some applications of the present invention. (The flowchart is repeated in the original text.) Figure 7 In steps 62 and 64 of the outlined method, each structured light projector 22 is driven to project a light pattern (e.g., a distribution 34 of discrete, unconnected light spots 33) onto a three-dimensional surface within the mouth, and each camera 24 is driven to capture an image including at least a portion of the pattern (e.g., one of the spots 33). In step 66, based on stored calibration values, a processor 96 ( Figure 1 The corresponding relationship algorithm will be run, and the following text will refer to it. Figure 8-12 Further described, the stored calibration values ​​indicate (a) camera rays 86 corresponding to each pixel on the camera sensor 58 of each camera 24, and (b) projector rays 88 corresponding to each projection spot 33 from each structured light projector 22. In some embodiments, the processor 96 is a processor disposed in the elongated handheld stick 20. In some embodiments, the processor 96 is disposed in a computing device, for example, as referenced below. Figures 60-61 Described, the computing device can be operatively connected to the elongated handheld stick 20 (e.g., via a wired or wireless connection). In some embodiments, multiple processors are used, wherein one or more processors may be disposed in the elongated handheld stick and / or one or more processors may be disposed in the computing device. Once the correspondence is solved, in step 68, the three-dimensional position on the intraoral surface is calculated and used to generate a digital three-dimensional image of the intraoral surface. Furthermore, using multiple cameras 24 to capture the intraoral scene provides an improved signal-to-noise ratio in the capture by a factor of the square root of the number of cameras.

[0535] Now for reference Figure 8 This outlines some applications according to the invention. Figure 7The flowchart of the correspondence algorithm in step 66 is as follows. Based on the stored calibration values, all projector rays 88 and all camera rays 86 corresponding to all detected spots 33' are mapped (step 70), and all intersections 98 of at least one camera ray 86 and at least one projector ray 88 are identified (step 72). Figure 9 and Figure 10 They are Figure 8 A simplified example of steps 70 and 72 is illustrated in the diagram. Figure 9 As shown, the three projector rays 88, along with the eight camera rays 86 of a total of eight detected spots 33' on the camera sensor 58 corresponding to camera 24, are mapped. Figure 10 As shown, sixteen intersection points 98 are marked.

[0536] exist Figure 7 In steps 74 and 76, the processor 96 determines the correspondence between the projected spot 33 and the detected spot 33' in order to identify the three-dimensional position of each projected spot 33 on the surface. Figure 11 The simplified example described in the previous paragraph is illustrated below. Figure 8 A schematic diagram of steps 74 and 76. For a given projector ray i, processor 96 “looks” at the corresponding camera sensor path 90 on the camera sensor 58 of one of the cameras 24. Each detected spot j along the camera sensor path 90 will have a camera ray 86 that intersects the given projector ray i at the intersection 98. The intersection 98 defines a three-dimensional point in space. Then processor 96 “looks” at the corresponding camera sensor paths 90' of the other cameras 24 corresponding to the given projector ray i, and identifies how many other cameras 24 also detect corresponding spots k on their corresponding camera sensor paths 90' corresponding to the given projector ray i, where their camera rays 86' intersect the same three-dimensional point in the space defined by the intersection 98. This process is repeated for all detected spots j along the camera sensor path 90, and the spot j that the most cameras 24 “agree” to be is identified as spot 33. Figure 12 The spot 33 is projected onto the surface from a given projector ray i. That is, projector ray i is identified as the specific projector ray 88 that produces the detected spot j, for which at most a number of other cameras detect corresponding spots k. Therefore, the three-dimensional position on the surface is calculated for this spot 33. The same process can be performed to calculate the three-dimensional positions of other features of the projected pattern on the surface.

[0537] In the example, such as Figure 11As shown, all four cameras detect corresponding spots where their respective camera rays intersect with projector ray i at intersection point 98 on their respective camera sensor paths corresponding to projector ray i. This intersection point 98 is defined as the intersection point of camera ray j corresponding to the detected spot j and projector ray i. Therefore, all four cameras are said to "agree" that spot 33 projected by projector ray i exists at intersection point 98. However, when the process is repeated for the next spot j', no other camera detects the corresponding spot where its corresponding camera ray intersects with the projector ray i at intersection 98', defined as the intersection of camera ray 86 (corresponding to the detected spot j') and projector ray i on its corresponding camera sensor path corresponding to projector ray i. Therefore, only one camera is considered to "agree" to have spot 33 (or other features) projected by projector ray i at intersection 98', while four cameras are considered to have spot 33 (or other features) projected by projector ray i at intersection 98. Thus, projector ray i is identified by projecting spot 33 (or other features) onto the surface at intersection 98. Figure 12 ) and generate specific projector rays 88 for the detected spot j. According to Figure 8 Step 78 and as follows Figure 12 As shown, calculate the three-dimensional position 35 on the inner surface of the mouth at intersection 98.

[0538] Now for reference Figure 13 This is a flowchart outlining other steps in the correspondence algorithm according to some applications of the present invention. Once the position 35 on the surface is determined, the projector ray i of the projected spot j and all camera rays 86 and 86' corresponding to spot j and the corresponding spot k are removed from consideration (step 80), and the correspondence algorithm is run again for the next projector ray i (step 82). Figure 14 A simplified example as described above is depicted after removing a specific projector ray i from the projected spot 33 at position 35. According to Figure 13 Step 82 in the flowchart, and then the correspondence algorithm is run again for the next projector ray i. Figure 14 As shown, the remaining data indicates that the three cameras "agree" to have spot 33 at intersection 98, which is defined by the intersection of camera ray 86 corresponding to the detected spot j and projector ray i. Therefore, as... Figure 15 As shown, calculate the three-dimensional position 37 at intersection point 98.

[0539] like Figure 16As shown, once the three-dimensional position 37 on the surface is determined, the projector ray i projecting spot j, along with all camera rays 86 and 86' corresponding to spot j and the corresponding spot k, are removed from consideration again. The remaining data shows spot 33 projected by projector ray i at intersection 98, and the three-dimensional position 41 on the surface at intersection 98 is calculated. Figure 17 As shown, in a simplified example, the three projected spots 33 of the three projector rays 88 of the structured light projector 22 are now located on the surface at three-dimensional positions 35, 37, and 41. In some applications, each structured light projector 22 projects 400-3000 spots 33. Once the correspondences for all projector rays 88 have been solved, a reconstruction algorithm can be used to reconstruct a digital image of the surface using the calculated three-dimensional positions of the projected spots 33.

[0540] Now for reference Figure 28 This is a flowchart outlining the steps of a method for generating digital three-dimensional images (hereinafter referred to as "blob tracking") according to some applications of the present invention. Although the method is called "blob tracking," it can also be applied to tracking other types of features of a projected pattern. Projected blobs move on the intraoral surface due to the movement of the handheld intraoral scanner relative to the intraoral surface during scanning. The inventors have realized that if the movement of a particular detected blob (or other feature) can be tracked in consecutive image frames, the correspondence solved for that particular blob (or other feature) in any of the frames across which the blob (or other feature) is tracked provides a solution for the correspondence of that blob (or other feature) in all the frames across which the blob (or other feature) is tracked. That is, if in one frame the processor 96 solves that a given detected blob 33' is projected by a given projector ray 88, and the processor 96 has determined, due to blob tracking, that the detected blob 33' in the next image is the same blob, then the processor 96 automatically determines that the same projector ray 88 produces the detected tracked blob in the next image.

[0541] Since the detected spot 33', which can be tracked across consecutive images, is generated by the same specific projector ray, the trajectory of the tracked spot will follow a specific camera sensor path 90 corresponding to that specific projector ray 88. If a correspondence is solved for the detected spot 33' at a point along the specific camera sensor path 90, the three-dimensional position on the surface can be calculated for all points along the camera sensor path 90 that detected the spot 33'. That is, the processor can calculate the corresponding three-dimensional position on the intraoral three-dimensional surface at the intersection of the specific projector ray 88 that generated the detected spot 33' and the corresponding camera ray 86 corresponding to the tracked spot in each of multiple consecutive images across which it tracks the spot 33'. This can be particularly useful when only one camera (or a few cameras) sees a specific detected spot in a specific image frame. If other cameras 24 have seen the specific detected spot in previous consecutive image frames, and a correspondence has been solved for the specific detected spot in those previous image frames, then even in an image frame in which only one camera 24 sees the specific detected spot, the processor knows which projector ray 88 produced the spot and can determine the three-dimensional position of the spot on the three-dimensional surface inside the mouth.

[0542] For example, areas that are difficult to reach inside the mouth may be imaged by only a single camera 24. In this case, if a spot 33' detected on the camera sensor 58 of the single camera 24 can be tracked through multiple previous consecutive images, the three-dimensional position on the surface of that point (even if it is only seen by a single camera 24) can be calculated based on the information obtained from the tracking, i.e., (a) along which camera sensor path the spot moves and (b) which projector light produced the tracked spot 33'.

[0543] exist Figure 28In step 180 of the method outlined herein, each structured light projector 22 is driven to project a light pattern onto a three-dimensional surface within the mouth. In one embodiment, the light pattern is a distribution 34 of discrete, unconnected light spots 33. In step 182, each camera 24 is driven to capture an image including at least a portion of the projected pattern (e.g., at least one spot 33). In one embodiment, in step 184, based on stored calibration values, a processor 96 is used to compare a series of images (e.g., multiple consecutive images) captured by each camera 24 and determine which features of the projected pattern (e.g., which projected spots 33) can be tracked across the multiple images. The stored calibration values ​​indicate (a) camera rays 86 corresponding to each pixel on the camera sensor 58 of each camera 24, and (b) projector rays 88 corresponding to each projected spot 33 from each structured light projector 22. Each tracked feature (e.g., spot 33s') moves along a path p corresponding to the pixel of the respective projector ray (e.g., a specific camera sensor path 90 corresponding to the projector ray 88). For some applications, in step 186, the processor 96 calculates the corresponding three-dimensional position of the feature (e.g., spot 33s') being tracked in the series of images (e.g., in each of the consecutive images) on the three-dimensional surface inside the mouth.

[0544] In one embodiment, each of one or more structured light projectors is driven to project a pattern onto a three-dimensional surface within the mouth. Additionally, each of one or more cameras is driven to capture multiple images, each image including at least a portion of the projected pattern. The projected pattern may include multiple projected light spots, and a portion of the projected pattern may correspond to a projected spot of the multiple projected light spots. The processor 96 then compares a series of images captured by the one or more cameras, determines, based on the comparison of the series of images, which portions of the projected pattern can be tracked across the series of images, and constructs a three-dimensional model of the three-dimensional surface within the mouth, at least in part based on the comparison of the series of images. In one embodiment, the processor solves a correspondence algorithm for the tracked portion of the projected pattern in at least one of the series of images, and uses the correspondence algorithm solved in the at least one of the series of images to find the tracked portion of the projected pattern in images in the series of images where the correspondence algorithm has not been solved, for example, by solving the correspondence algorithm for the tracked portion of the projected pattern, wherein the solution of the correspondence algorithm is used to construct the three-dimensional model. In another embodiment, the processor solves a correspondence algorithm for the tracked portion of the projected pattern based on the position of the tracked portion in each image of the entire series of images, wherein the solution of the correspondence algorithm is used to construct the three-dimensional model. In one embodiment, the processor compares the series of images based on stored calibration values ​​indicating (a) camera rays corresponding to each pixel on a camera sensor of each of one or more cameras, and (b) projector rays corresponding to each of a projection spot from each of one or more structured light projectors, wherein each projector ray corresponds to a corresponding path of a pixel on at least one camera sensor, and wherein each tracked spot s moves along the path of a pixel corresponding to the corresponding projector ray r.

[0545] Now for reference Figure 29 This is a simplified example depicting some detected spots 33' (particularly 33a' and 33b') according to some applications of the invention, and how the processor 96 can determine which set of detected spots 33' can be considered to be tracked. Figure 28 A schematic diagram of step 184). Spot 33a' is spot 33' detected in the previous image, and spot 33b' is spot 33' detected in the current image. Processor 96 searches within the search radius to find possible matches between spot 33a' and spot 33b', which are considered to be the same spot 33' tracked between the two images.

[0546] The inventors have recognized that there are generally three factors that can affect how far a blob moves between frames:

[0547] 1. The distance a speckle travels between frames is generally inversely proportional to the camera's frame rate. That is, if the camera's frame rate is very fast, the speckle will appear to travel only a relatively small distance between pairs of consecutive frames, while if the camera's frame rate is slow, the speckle will appear to travel a greater distance between pairs of consecutive frames. The speed at which the rod moves relative to the inner surface of the mouth also affects the distance the speckle travels between frames; that is, when the rod moves faster, the speckle will appear to travel a greater distance between pairs of consecutive frames.

[0548] 2. The distance a spot travels between frames typically varies with the tilt of the scanned intraoral surface relative to the projector 22 and / or camera 24. If a spot is projected onto a tilted surface and moves along the tilt direction, the corresponding detected spot on the camera sensor will move faster, thus the tracked spot will travel a greater distance between pairs of consecutive frames.

[0549] 3. From the camera's perspective, how far a spot moves between frames typically varies with the distance between the scanned surface and the projector. When the surface is closer to the projector, even a small movement of the projector can cause a large movement of the tracked spot on sensor 58 of camera 24 between pairs of consecutive frames. Conversely, when the surface is farther away, the same movement of the projector will cause a smaller movement of the tracked spot on sensor 58 of camera 24 between pairs of consecutive frames. For example, if the distance between the surface and the projector is close to infinity, movement of the projector will cause almost zero movement of the tracked spot on sensor 58 of camera 24.

[0550] For some applications, processor 96 searches within a fixed search radius of at least three pixels and / or less than ten pixels (e.g., five pixels). For some applications, processor 96 considers parameters such as the blob location error level that can be determined during calibration when calculating the search radius. For example, the search radius may be limited to 2*(blob location error) or 3*(blob location error).

[0551] exist Figure 29In the simplified example shown, blobs 33a' and 33b' in each set 112 are considered close enough to be considered the same projected blobs 33 that moved through both images. That is, for each set 112, the two detected blobs 33a' and 33b' are considered the same tracked blobs 33s'. In contrast, blobs 33a' and 33b' in set 114 are too far apart to be considered tracked. Blobs 33a' and 33b' in set 116 are close enough, but more than one match was found, so they cannot be considered tracked. As further described below, there may be pairs of tracked blobs in set 116, and further analysis of more images can help determine which blobs are indeed tracked.

[0552] In one embodiment, to generate a digital 3D image, an intraoral scanner drives each of one or more structured light projectors to project a light pattern onto an intraoral 3D surface. The intraoral scanner also drives each of a plurality of cameras to capture an image including at least a portion of the projected pattern; each camera includes a camera sensor comprising a pixel array. The intraoral scanner also uses a processor to run a correspondence algorithm to compute corresponding 3D positions of a plurality of features of the projected pattern on the intraoral 3D surface. The processor uses data from a first camera (e.g., data from at least two of the plurality of cameras) to identify candidate 3D positions of a given feature of the projected pattern corresponding to a particular projector ray r, wherein data from a second camera (e.g., another camera that is not one of the at least two cameras) is not used to identify the candidate 3D position. The processor also uses the candidate 3D positions seen by the first camera to identify a search space on the pixel array of the second camera, in which features of the projected pattern from the projector ray r are to be searched. If a feature of the projected pattern from the projector ray r is identified within the search space, the processor uses data from the second camera to refine the candidate 3D positions of the projected pattern features. In one embodiment, the light pattern comprises a distribution of discrete, unconnected light spots, and wherein the projection pattern is characterized by projection spots from the unconnected light spots. In one embodiment, the processor uses stored calibration values ​​that indicate (a) camera rays corresponding to each pixel on a camera sensor of each of a plurality of cameras, and (b) projector rays corresponding to each feature of a projection light pattern from each of one or more structured light projectors, whereby each projector ray corresponds to a corresponding path to a pixel on at least one camera sensor.

[0553] Now for reference Figure 30It is a flowchart outlining a method for determining tracked features (e.g., such as tracked spots 33s') according to some applications of the present invention. Figure 30 This discussion is based on the tracked blob, but it also applies to other types of tracked features. For some applications, in addition to searching for tracked blobs by monitoring the proximity of blobs in consecutive images, processor 96 can also search for tracked blobs 33s' based on the parameters of detected blobs 33', referred to below as "parameter tracking". Processor 96 determines the parameters of the detected blobs 33' in the first image in consecutive images (step 188) and the adjacent images. Processor 96 then uses the determined parameters of the detected blobs 33' in the two adjacent images to predict the same parameters of the blobs in subsequent images (e.g., in the next image (and in subsequent images)) (step 190). Processor 96 searches for blobs with substantially the predicted parameters in the subsequent images (e.g., in the next image) (step 192). For example, either by the correspondence algorithm described above, or by proximity tracking as described in the previous two paragraphs, it can be determined that two specific detected blobs 33a' and 33b' in two adjacent frames both originate from the same projector ray. Once the processor 96 knows that the detected spots 33a' and 33b' in two adjacent frames were generated by the same projector light 88, the processor 96 can determine the parameters of the spots and, based on the parameters of the spots in the two adjacent images, predict the parameters of the spots in the subsequent image (e.g., the next image).

[0554] For some applications, the parameters of a blob are its size, shape (e.g., aspect ratio), orientation, intensity, and / or signal-to-noise ratio (SNR). For example, if the determined parameter is the shape of the tracked blob 33s', the processor 96 predicts the shape of the tracked blob 33s' in a subsequent image (e.g., the next image) and, based on the predicted shape of the tracked blob 33s' in the subsequent image, determines a search space in which to search for the tracked blob 33s', for example, a search space with size and aspect ratio based on the predicted shape of the tracked blob 33s' in the subsequent image (e.g., within a factor of two). For some applications, the shape of the blob may refer to the aspect ratio of an elliptical blob.

[0555] Refer again Figure 29 For some applications, parameter tracking can help resolve uncertainties, such as in... Figure 29The set of spots 116 is shown in the diagram. As mentioned above, spots 33a' and 33b' in set 116 are close enough to be considered as tracked spots, but more than one match is found. Based on parameter tracking, processor 96 is now able to determine which spots in set 116 are indeed tracked spots.

[0556] Now for reference Figure 31 It is a flowchart outlining a method for finding tracked spots 33s' in a subsequent image according to some applications of the present invention. Figure 31 This also applies to finding other tracked features in subsequent images. For some applications, based on the direction and distance the tracked spot 33s' has moved between two images (e.g., between two consecutive images), the processor 96 determines the velocity vector of the tracked spot 33s' (step 194). The processor 96 then uses the velocity vector to determine the search space in subsequent images (e.g., the next image) in which to search for the tracked spot 33s' (step 196).

[0557] For some applications, the new location of the tracked blob 33s' can be estimated by using a predictive filter (e.g., a Kalman filter) to determine the search space in the background image.

[0558] Now for reference Figure 32 It is a flowchart outlining a method for finding tracked spots 33s' in a subsequent image according to some applications of the present invention. Figure 32 This also applies to finding other tracked features in subsequent images. Determining the velocity vector of the tracked spot 33s' can also help determine the search space in which to search for the tracked spot; that is, if the spot moves faster, it will move further between consecutive frames, so the processor 96 can set a larger search space in which to search for the tracked spot. The inventors have realized that the shape of the tracked spot 33s' and the direction in which it is moving can indicate the speed of the spot. For example, for some applications, if a spot projected as a circle appears elliptical, the spot is likely to fall on an inclined surface relative to the projector 22 and / or camera 24. Similarly, as described above, a spot moving along an inclined surface in the inclined direction moves faster than when it is not moving in the inclined direction, and the steeper the inclination of the surface, the faster the spot moves. Furthermore, if a spot appears stretched into an ellipse due to the inclined surface, it may appear stretched in the inclined direction, i.e., the major axis of the ellipse is in the inclined direction. Therefore, an elliptical spot moving along its major axis indicates that the spot is moving upward or downward and tilting, and thus it moves faster than an elliptical spot moving along its minor axis (this may indicate that the spot is not moving in the tilting direction despite being projected onto a tilted surface).

[0559] Therefore, for some applications, after determining the shape of the tracked spot 33s' (step 198), the processor 96 can determine the velocity vector of the tracked spot 33s' based on the direction and distance of its movement between two consecutive images (step 200). The processor 96 can then use the determined velocity vector and / or shape of the tracked spot 33s' to predict its shape in a subsequent image (e.g., the next image) (step 202). After predicting the shape of the tracked spot 33s', the processor 96 can use the combination of the velocity vector and the predicted shape to determine the search space in the subsequent image (e.g., the next image) to which the tracked spot 33s' is to be searched. Referring again to the example of the elliptical spot above, a larger search space will be specified if the shape of the spot is determined to be elliptical and the spot is determined to be moving along its major axis, and vice versa if the elliptical spot is moving along its minor axis.

[0560] Now for reference Figure 33 This is a schematic diagram illustrating how speckle tracking according to some applications of the invention helps identify a detected speckle 33' as an example of being projected from a particular projector ray 88. This is useful for cases where the correspondence algorithm does not provide a solution for a particular detected speckle 33' (e.g., a speckle 33' detected in a particular frame is only seen by one camera 24). In this case, if the detected speckle 33' is identified as a tracked speckle 33s' moving along a particular camera sensor path 90 corresponding to a particular pixel of the particular projector ray 88, then, as described above, it can be assumed that the particular projector ray 88 projects the speckle. Based on the correspondence solved for the tracked speckles 33s' in previous frames, the correspondence can be solved for frames in which only one camera detects the speckle 33'. Therefore, for some applications, when running the correspondence algorithm (e.g., referred to above), Figure 7-17 After the correspondence algorithm described, if it is determined that the detected spot 33' is a tracked spot 33s' that moves along a specific path 90 on the camera sensor 58 corresponding to a specific projector ray 88, then it can be considered that the specific projector ray 88 generated the detected spot 33'. That is, the processor 96 can identify the detected spot 33' as originating from the specific projector ray 88 by identifying the detected spot 33' as a tracked spot 33s' that moves along the path 90 of the pixel of the camera sensor 58 corresponding to the specific projector ray 88.

[0561] exist Figure 33In the example shown, in two consecutive frames captured at time 1 and time 2 respectively, each of the two camera sensors 58 detects a spot 33' projected from the projector ray 88. The correspondence algorithm described above has solved the correspondence of the detected spot 33' in frames 1 and 2, thus determining that the projector ray 88 generated the detected spot 33' in frames 1 and 2. However, in the third frame, only one camera detects the spot 33'. For this example, assume that the correspondence algorithm cannot solve the correspondence of the spot 33' detected in frame 3. However, the processor 96 determines that the detected spot 33' in frame 3 is a tracked spot 33s' that moves along the same projector ray 88 that generated the spot 33' in frames 1 and 2. Therefore, the processor 96 identifies the spot 33' detected in frame 3 as generated by the projector ray 88.

[0562] Now for reference Figure 34A -B, which is a simplified schematic diagram of a camera sensor 58 according to some applications of the present invention, showing two detected spots 33c' and 33d'. Figure 34A In this context, for each detected blob, there is uncertainty regarding which projector light ray produced the blob, i.e., on which camera sensor path 90 of the pixel the detected blob falls. For some applications, processor 96 is capable of using blob tracking to resolve these uncertainties. Therefore, for some applications, when running correspondence algorithms (e.g., as referenced above)... Figure 7-17 After the correspondence algorithm described, if the detected spot 33' is identified as coming from two different candidate projector rays 88 and 88' based on the three-dimensional position calculated by the correspondence algorithm, then the processor 96 can identify the detected spot 33' as coming from only one of the two different candidate projector rays 88 and 88' by identifying that the detected spot 33' is a tracked spot 33s' moving along path 90 or path 90'.

[0563] Examples of this uncertainty include Figure 34A The detected spot 33c' is represented in the diagram. Spot 33c' is located at the intersection of two different paths 90c and 90c'. The correspondence algorithm may have found that this detected spot 33c' is generated by projector ray 88 corresponding to path 90c and projector ray 88' corresponding to path 90c'. Another type of uncertainty is caused by... Figure 34A The detected spot 33d' is represented in the image. Spot 33d' is very close to two different paths 90d and 90d', but not at the intersection between paths 90 and 90d. Due to noise in the signal, it may not be clear during the correspondence algorithm whether spot 33d' is generated by projector ray 88 corresponding to path 90d or by projector ray 88' corresponding to path 90d'.

[0564] like Figure 34B As shown, the processor 96 can identify which projector rays generate each of spots 33c' and 33d' by identifying spots 33c' and 33d' as tracked spots 33s' moving along a specific path. Detected spot 33c' is identified as a tracked spot 33s' moving along path 90c', and therefore detected spot 33c' is identified as generated by projector rays 88' corresponding to path 90c'. Detected spot 33d' is identified as a tracked spot 33s' moving along path 90d', and therefore detected spot 33d' is identified as generated by projector rays 88' corresponding to path 90d'.

[0565] Now for reference Figure 35 This is a flowchart outlining some applications of the invention that may use additional or alternative methods of speckle tracking. Figure 35 The concepts illustrated also apply to other feature-tracking methods. For some applications, processor 96 might be able to use blob tracking to remove erroneously detected blobs 33' from points considered to be on the intraoral 3D surface. For some applications, when running correspondence algorithms (e.g., see reference above), Figure 7-17 Following the described correspondence algorithm (step 206), the processor 96 can identify the detected spot 33' as originating from a specific projector ray 88 based on the correspondence algorithm (step 208). Step 206 typically occurs during... Figure 28 Following step 186 of the method outlined in the flowchart. Furthermore, processor 96 can identify a series of spots detected across multiple consecutive images, all of which are tracked spots 33s' moving along a path 90 corresponding to the pixel of the same specific projector ray 88. As shown in decision diamond 210, if the detected spot 33' is one of the tracked spots 33s', then the detected spot 33' can be considered a point on the intraoral three-dimensional surface (step 212). However, if the detected spot 33' is not identified as a tracked spot 33s' moving along a path 90 corresponding to the pixel of that specific projector ray 88, then the detected spot 33' can be considered a false positive detection of a spot and the detected spot 33' is removed from the list of points considered to be on the intraoral three-dimensional surface (step 214).

[0566] Now for reference Figure 36 This is a flowchart outlining some applications of the invention that may use additional or alternative methods of speckle tracking. Figure 36The concepts illustrated also apply to other feature-tracking methods. To reduce the occurrence of many false positives detected by camera 24, processor 96 can set an intensity threshold, so that any detected spots 33' below the threshold are not included as candidate spots in the correspondence algorithm. However, this can also lead to falsely mis-detected spots, that is, spots that could provide useful information may be overlooked because they are weak spots (with intensities below the threshold). For example, in areas where it is difficult to capture intraoral scenes, some projected spots 33 may be below the intensity threshold. Therefore, for some applications, when running the correspondence algorithm (e.g., as mentioned above), Figure 7-17 After the correspondence algorithm described (step 216), the processor 96 can, for example, identify weak spots 33' whose three-dimensional positions were not calculated by the correspondence algorithm by lowering the intensity threshold and considering spots 33' not considered by the correspondence algorithm (step 218). As shown in decision rhombus 220, if a weak spot 33' is identified as a tracked spot 33s' that moves along a path 90 corresponding to a pixel of a particular projector ray 88, then the weak spot 33' is identified as a point projected from that particular projector ray 88 and considered to be on the intraoral three-dimensional surface (step 222). If a weak spot is not identified as a tracked spot 33s', then the weak spot is removed from the list of points considered to be on the intraoral three-dimensional surface (step 224).

[0567] For some applications, for a tracked spot 33s', the processor 96 can determine multiple possible camera sensor pixel paths 90 along which the tracked spot 33s' moves, with multiple paths 90 corresponding to multiple possible projector rays 88. For example, there may be more than one projector ray 88 that closely corresponds to a path 90 on a pixel of a given camera sensor 58. The processor 96 can run a correspondence algorithm to identify which of the possible projector rays 88 produces the tracked spot 33s' in order to calculate a three-dimensional position on the surface for the corresponding position of the tracked spot 33s'.

[0568] For a given camera sensor 58, for each of a plurality of possible projector rays 88, a three-dimensional point in space exists at the intersection of each possible projector ray 88 and the camera ray corresponding to the tracked spot 33s' detected in the given camera sensor 58. For each possible projector ray 88, the processor 96 considers the camera sensor path 90 on each other camera sensor 58 corresponding to the possible projector ray 88 and identifies how many other camera sensors 58 also detect, on their respective camera sensor path 90 corresponding to that possible projector ray 88, the spot 33' whose camera ray intersects with the three-dimensional point in space, i.e., how many other cameras agree that the tracked spot 33s' is projected by that projector ray 88. This process is repeated for all possible projector rays 88 corresponding to the tracked spot 33s'. The possible projector ray 88 agreed upon by the most other cameras is determined as the specific projector ray 88 that generates the tracked spot 33s'. Once the specific projector ray 88 of the tracked spot 33s' is determined, the camera sensor path 90 along which the spot moves is known, and the corresponding three-dimensional position on the surface at the intersection of the specific projector ray 88 and the corresponding camera ray of the tracked spot 33s' in each consecutive frame is calculated, where the spot 33s' is tracked across these consecutive frames.

[0569] Now for reference Figure 37A -B is a schematic diagram illustrating points used for 3D reconstruction before and after speckle tracking has been implemented in processor 96, according to some applications of the invention. The size of each data point represents the number of cameras used to solve for that point; that is, the larger the point, the more cameras that see the speckle. It should be noted that the size of the data points in the figure is only used to distinguish how many cameras see any given point and does not represent the size of the projected speckle on the surface. Figure 37A Within the sample, there exist numerous smaller points that appear to be located on the periphery rather than on the inner surface of the sample. These smaller points refer to the detected blobs seen by a very small number (e.g., only one) of cameras, but are still assigned a three-dimensional position in space based on a correspondence algorithm. After running the correspondence algorithm, the processor 96 can perform blob tracking and thus determine that these few points on the periphery are actually false positives (by determining that they are not the tracked blobs). Therefore, as... Figure 37B As shown, after speckle tracking, most of the specks identified as false positives by speckle tracking have been removed from the points that were thought to be on the intraoral surface.

[0570] Now for reference Figure 38This is a flowchart outlining the steps of a method for generating digital three-dimensional images (hereinafter referred to as "ray tracing") according to some applications of the present invention. For some applications, as an alternative or additional scheme to tracing spots and / or other features detected within a two-dimensional image (as described above), the length of each projector ray 88 can be tracked in three-dimensional space. The length of the projector ray 88 is defined as the distance between the origin of the projector ray 88 (i.e., the light source) and the three-dimensional position where the projector ray 88 intersects with the inner surface of the aperture.

[0571] exist Figure 38 In step 226 of the method outlined above, each structured light projector 22 is driven to project a light pattern (e.g., a distribution 34 of discrete, unconnected light spots 33) onto a three-dimensional surface within the mouth, and in step 228, each camera 24 is driven to capture multiple images, each image including at least one feature of the projected pattern (e.g., at least one spot 33). This method is described with reference to spots, but is also applicable to other types of features. In step 230, based on stored calibration values, a correspondence algorithm (e.g., as described above) is run using processor 96. Figure 7-17 The correspondence algorithm described indicates (a) the camera ray 86 corresponding to each pixel on the camera sensor 58 of each camera 24, and (b) the projector ray 88 corresponding to each projection spot 33 from each structured light projector 22. As a result of the correspondence algorithm, each solved projector ray 88 in each image frame generates a reconstructed 3D point in space, which in turn defines the length of the solved projector ray 88 in that frame.

[0572] Therefore, in step 232, in at least a subset of the captured images (e.g., in a series of images or multiple consecutive images), processor 96 identifies the calculated 3D position (e.g., calculated from the correspondence algorithm) of the detected spot 33' as corresponding to a specific projector ray 88. In step 234, based on each 3D position corresponding to the projector ray 88 in the image subset, processor 96 evaluates (e.g., calculates) the length of the projector ray 88 in each image of the image subset. Since camera 24 captures images at a relatively high frame rate (e.g., approximately 100 Hz), the geometry of the spot seen by each camera does not change significantly between frames. Therefore, if the evaluated (e.g., calculated) length of the projector ray 88 is tracked and plotted relative to time, the data points will follow a relatively smooth curve, although some discontinuities may occur, as discussed further below. Thus, the length of the projector ray over time forms a relatively smooth univariate function relative to time. As described above, the detected spot 33' corresponding to the projector ray 88 on multiple consecutive images will appear to move along a one-dimensional line, which is the path 90 of the pixel in the camera sensor corresponding to the projector ray 88.

[0573] In one embodiment, a method for generating a digital three-dimensional image includes driving each of one or more structured light projectors to project a pattern onto an intraoral three-dimensional surface, and driving each of one or more cameras to capture an image including at least a portion of the pattern. The method also includes using a processor to run a correspondence algorithm to compute corresponding three-dimensional positions of multiple features of the pattern captured on the intraoral three-dimensional surface in a series of images. The processor also identifies the computed three-dimensional positions of detected features of the imaged pattern corresponding to one or more specific projector rays r in at least one subset of the image series. Based on the three-dimensional positions of the detected features corresponding to the one or more projector rays r in the image subset, the processor evaluates (e.g., computes) the length associated with the one or more projector rays r in each image of the image subset. In one embodiment, the processor computes an estimated length of the one or more projector rays r in at least one image of the image series, in which the three-dimensional positions of projected features from the one or more projector rays r are not identified. In one embodiment, each of the one or more cameras includes a camera sensor comprising a pixel array, wherein calculating the corresponding three-dimensional positions of a plurality of features of the pattern on a three-dimensional surface within the mouth and identifying the calculated three-dimensional positions of the detected features of the pattern corresponding to a particular projector ray r is performed based on stored calibration values ​​indicating (i) camera rays corresponding to each pixel on the camera sensor of each of the one or more cameras, and (ii) projector rays corresponding to each feature of the projected light pattern from each of the one or more projectors, whereby each projector ray corresponds to a corresponding path of a pixel on at least one camera sensor. In one embodiment, the pattern comprises a plurality of spots, and each of the plurality of features of the pattern comprises a spot among a plurality of spots.

[0574] Now for reference Figure 39 This is a diagram illustrating the length of a tracked projector ray 88 over time in some applications according to the invention, and a specific simplified view of a camera sensor 58 corresponding to a particular image frame. The inventors have implemented various uses of the aforementioned ray tracking. For some applications, there may be at least one image from multiple consecutive images in which no image is present, as in... Figure 38 In step 232 of the method shown, the three-dimensional position of the projected spot 33 (or other feature) from a specific projector ray 88 is identified. For example, a projected spot (or other feature) in a particular frame may be below an intensity threshold and not considered by the correspondence algorithm, or there may be false detection spots in a particular frame. However, since the length of the projector ray is tracked over time, the processor 96 can calculate the estimated length of the specific projector ray 88 in the image.

[0575] For example, in Figure 39 In the exemplary diagram shown, for scan frame s1 captured at time t1, the three-dimensional position of the projected spot 33 is not identified based on the correspondence algorithm. Therefore, as shown by the dashed circle 236, for scan frame s1, there is no data point corresponding to the length of the projector ray 88. However, since the length of the ray is tracked through multiple consecutive images, the estimated length L1 of the projector ray 88 for scan frame s1 can be calculated, for example, by interpolation. As described above, all points projected by a particular projector ray 88 appear on a specific path 90 in the camera sensor 58 corresponding to the pixel of that particular projector ray 88. Therefore, for scan frame s1 in which the three-dimensional position of the spot 33 corresponding to the particular projector ray 88 is not identified in step 232, the processor 96 can determine a one-dimensional search space 238 in scan frame s1 in which to search for the projected spot from that particular projector ray 88. The one-dimensional search space 238 is along the corresponding path 90 of the pixel corresponding to the particular projector ray 88. This contrasts with the speckle tracking algorithm described above, in which the processor 96 searches for specks that are close enough to each other in two dimensions within the image from one frame to the next, so that they are considered as tracked specks generated by the same projector light in each image frame.

[0576] For some applications, based on the estimated length Ll of the projector ray 88 in at least one of multiple images, the processor 96 can determine a one-dimensional search space in the corresponding pixel arrays (e.g., camera sensors 58) of the multiple cameras 24 (e.g., all cameras 24). For each corresponding pixel array, the one-dimensional search space is along a corresponding path 90 in that particular pixel array (e.g., camera sensor 58) corresponding to the pixel of the projector ray 88. The length Ll of the projector ray 88 corresponds to a three-dimensional point in space, which corresponds to a two-dimensional position on the camera sensor 58. All other cameras 24 also have corresponding two-dimensional positions on their camera sensors 58 corresponding to the same three-dimensional point in space. Therefore, the length of the projector ray in a particular frame can be used to define a one-dimensional search space in the multiple camera sensors 58 (e.g., all camera sensors 58) for that particular frame.

[0577] For some applications, contrary to false detections of not detecting the expected spot (or other feature), there may be at least one of multiple consecutive images in which more than one candidate 3D position is calculated for the projected spot 33 (or other feature) from a particular projector ray 88; that is, a false positive detection of the projected spot 33 (or other feature) occurs. For example, in Figure 39In the exemplary figure shown, for scan frame s2 captured at time t2, based on the correspondence algorithm, it is calculated that the two candidate detected spots 33' and 33" both originate from a specific projector ray 88. Therefore, two candidate three-dimensional positions of the projected spot 33 are calculated. The processor 96 calculates two candidate lengths of the projector ray 88 corresponding to each candidate three-dimensional position for this frame. For example, the processor 96 calculates that for the candidate detected spot 33', the candidate length of the projector ray 88 is L2 (from...). Figure 39 Data point 242 in the image represents the candidate length of the projector ray 88 for the candidate detected spot 33, which is L3 (as indicated by the data point 242 in the image). Figure 39 (Data point 244 is represented in the image). Since the length of the projector ray 88 is tracked on multiple consecutive images, when ray length data for scan frame s2 is added, it becomes obvious which candidate length L2 or L3 is the estimated length of the projector ray 88 for scan frame s2. Therefore, the processor 96 is able to determine which of the more than one candidate three-dimensional positions of the projection spot 33 is the correct three-dimensional position of the projection spot 33 by determining which of the candidate three-dimensional positions corresponds to the estimated length of the projector ray 88 for that image.

[0578] Based on the estimated length of the projector ray 88 in at least one of multiple images (e.g., in scan frame s2), the processor 96 can determine a one-dimensional search space 246 in scan frame s2. Subsequently, the processor 96 can determine which of the more than one candidate three-dimensional locations of the projected spot 33 is the correct three-dimensional location of the projected spot 33 generated by the projector ray 88 by determining which one corresponds to the spot 33' found within the one-dimensional search space 246. Before ray tracing provides additional information, the camera sensor 58 for scan frame s2 has shown two candidate spots 33' and 33'" both on the path 90 of the pixel corresponding to the projector ray 88. The processor 96 calculates the estimated length of the projector ray 88 based on the length of the ray traced over multiple consecutive images, thereby allowing the processor 96 to determine the one-dimensional search space 246 and to determine that the candidate detected spot 33' is indeed the correct spot. The candidate detected spot 33'" is then removed from the list of points considered to be on the three-dimensional inner surface of the aperture.

[0579] For some applications, processor 96 may define curve 248 based on the evaluated (e.g., calculated) length of projector ray 88 in each image of a subset of images (e.g., multiple consecutive images). The inventors assume that it is reasonable to consider any detected point whose three-dimensional position corresponds to the length of projector ray r at least a threshold distance from the defined curve 248, based on a correspondence algorithm, as a false positive detection and can be removed from points considered to be on the three-dimensional intraorbital surface.

[0580] Now for reference Figure 40A -B is a graph illustrating experimental datasets before and after ray tracing for some applications according to the present invention. Figure 40A In the data, a specific projector ray 88 length is plotted for each spot and / or other feature, calculated using a correspondence algorithm. Before applying ray tracing, the ray lengths corresponding to spots and / or other features that appear to deviate from the overall curve defined by the projector ray 88 are included in the data. Figure 40B This indicates how the data looks after ray tracing has been applied and used to determine which false positive spots and / or other features should be removed from points that are considered to be on the three-dimensional intraoral surface.

[0581] Now for reference Figure 41 This is a schematic diagram of a projector with multiple camera sensors and projection spots or other features for some applications according to the present invention. For some applications, a correspondence algorithm is run (e.g., referenced above). Figure 7-17 Following the described correspondence algorithm, the processor 96 can determine the candidate three-dimensional position of the projection spot 33 (or other feature) with a certain degree of determinism based on how many cameras 24 detect a given projection spot 33 (or other feature) on their respective pixel arrays (i.e., camera sensors 58). Figure 41 The candidate 3D positions of the projected spot 33 (or other features) are marked by dashed circles 250.

[0582] The more cameras 24 that see the projected spot 33 (or other feature), the higher the certainty of the candidate 3D location 250. Therefore, using data from at least two cameras 24, the processor 96 can identify candidate 3D locations 250 corresponding to a given spot 33 (or other feature) of a particular projector ray 88. Assuming that the identification of candidate 3D locations 250 is not determined substantially using data from at least another camera 24', there may be some error in the candidate 3D locations 250, and the processor 96 can refine the candidate 3D locations 250 if it has data from other cameras 24'.

[0583] Therefore, assuming that after the correspondence, at least two cameras 24 see the projected spot 33 (or other feature), the processor 96 knows (a) which projector ray 88 produced the projected spot 33 (or other feature) and (b) the candidate 3D location 250 of the spot (or other feature). Combining (a) and (b) allows the processor 96 to determine a one-dimensional search space 252 in the pixel array of the other camera 24' (i.e., camera sensor 58') to search for the spot (or other feature) from the projector ray 88. The one-dimensional search space 252 is along the path 90 of the pixels on the camera sensor 58' of the other camera 24', and can be along specific segments of the path 90 corresponding to the candidate 3D location 250. If a spot 33' (or other feature) from the projector ray 88 is identified within the one-dimensional search space 252 (e.g., a falsely detected spot 33' not considered by the correspondence algorithm (e.g., because it has a subthreshold intensity)), then, using the data now obtained from the other camera 24', the processor 96 can refine the candidate three-dimensional position 250 of the spot 33 (or other feature) to become the refined three-dimensional position 254.

[0584] Now for reference Figure 42A -B, which shows a flowchart outlining a method for generating a three-dimensional image according to some applications of the invention. For some applications, once the three-dimensional positions of at least three projection spots 33 (or other features) from three different projector rays 88 are identified, a three-dimensional surface can be estimated such that all three identified three-dimensional positions lie on the estimated surface. For another projector ray 88', whose three-dimensional position of projection spot 33 (or other feature) is not determined, for example, because projection spot 33 has a subthreshold intensity, candidate three-dimensional positions can be calculated at the intersection of the other projector ray 88' and the estimated three-dimensional surface. Similar to the reference above. Figure 41 As described, the processor 96 can use a combination of the following to determine a one-dimensional search space in at least one camera sensor 58, along which a detected spot 33' from another projector ray 88' is to be searched: (a) the specific projector ray is known, i.e., it is known which path 90 on the sensor is to be viewed, and (b) the candidate three-dimensional position is known.

[0585] Therefore, in Figure 42A In step 256 of the method outlined in -B, each structured light projector 22 is driven to project a structured light pattern (e.g., a distribution 34 of discrete, unconnected light spots 33) onto a three-dimensional surface within the mouth, and in step 258, each camera 24 is driven to capture multiple images, each image including at least one spot. In step 260, based on stored calibration values, a correspondence algorithm (e.g., as described above) is run using processor 96. Figure 7-17 The correspondence algorithm described is used to calculate the corresponding three-dimensional positions of multiple detected spots 33' on the intraoral three-dimensional surface for each of multiple images, the stored calibration values ​​indicating: (a) the camera ray 86 corresponding to each pixel on the camera sensor 58 of each camera 24, and (b) the projector ray 88 corresponding to each projected spot 33 from each structured light projector 22. In step 262, using data corresponding to the corresponding three-dimensional positions of at least three detected spots 33', each detected spot 33' corresponding to a corresponding projector ray 88, the processor 96 estimates the three-dimensional surface on which all at least three detected spots 33' are located. The processor 96 then considers another projector ray 88'. As shown in decision hexagon 266, if the correspondence algorithm has identified a three-dimensional position of a detected spot 33' from another projector ray 88' in step 260, then that detected spot 33' from the other projector ray 88' can be considered a point on the intraoral surface at that three-dimensional position (step 268). However, for another projector ray 88', for which the three-dimensional position of the spot 33 corresponding to the other projector ray 88' was not calculated in step 260, the processor 96 can estimate the three-dimensional position in space of the intersection of the other projector ray 88' and the estimated three-dimensional surface (step 270). In step 272, the processor 96 uses the estimated three-dimensional position in space to identify a search space (e.g., a one-dimensional search space) in the pixel array of at least one camera 24 (e.g., camera sensor 58), along which the detected spot 33' corresponding to the other projector ray 88' is to be searched.

[0586] As described above, to reduce the occurrence of many false positive spots detected by camera 24, processor 96 can set a threshold, such as an intensity threshold, so that any detected features below the threshold (e.g., spot 33') are not considered by the correspondence algorithm. Therefore, for example, since the detected spot 33' has a sub-threshold intensity, the three-dimensional position of spot 33 corresponding to another projector ray 88' may not be calculated in step 260. In step 272, in order to search for features (e.g., detected spot 33'), the processor can lower the threshold to consider features that were initially not considered by the correspondence algorithm.

[0587] As used throughout this application, including the claims, when identifying the search space in which to search for detected features (e.g., detected spot 33') may be as follows:

[0588] (a) False detections of blobs (e.g., subthreshold blobs initially not considered by the correspondence algorithm, or blobs occluded by moving tissue), in which case the processor 96 can lower the threshold to re-search for detected blobs 33' in that specific region (i.e., the identified search space), or

[0589] (b) False positive spots, i.e., more than one candidate 3D position of a spot is identified by the correspondence algorithm. In this case, the processor 96 can re-search the detected spots 33' in the specific region (i.e., the identified search space) to determine which spot is the correct spot.

[0590] For some applications, in step 260, the correspondence algorithm can identify more than one candidate 3D location for the detected spot 33' from the projector ray 88'. As shown in decision hexagon 267, for the projector ray 88', given that it has identified more than one candidate 3D location for the detected spot 33', the processor 96 can estimate the 3D location in space of the intersection of the projector ray 88' and the estimated 3D surface (step 269). In step 271, the processor 96 selects which of the candidate 3D locations for the detected spot 33' from the projector ray 88' is the correct location based on the 3D location of the intersection of the projector ray 88' and the estimated 3D surface.

[0591] For some applications, in step 262, the processor 96 uses data corresponding to the respective 3D positions of all at least three detected spots 33' captured in one of the multiple images. Furthermore, after estimating the 3D surface, the estimation can be refined by adding data points from subsequent images (i.e., using data corresponding to the 3D position of at least one additional spot, calculated based on another image in the multiple images), such that all spots (the three spots used in the original estimation and the at least one additional spot) lie on the refined estimated 3D surface. For some applications, in step 262, the processor 96 uses data corresponding to the respective 3D positions of at least three detected spots 33' each captured in a separate image (i.e., in one of the multiple images).

[0592] Note that the above discussion concerns false positives in projection spots. False positives could also be genuine false positives, where the spot corresponding to a particular projector beam is not actually projected onto the intraoral surface, for example, due to obstruction by moving tissue (e.g., a patient's tongue, a patient's cheek, or a doctor's finger).

[0593] In one embodiment, a method for generating a digital three-dimensional image includes driving each of one or more structured light projectors to project a light pattern onto an intraoral three-dimensional surface, and driving each of one or more cameras to capture a plurality of images, each image including at least a portion of the projected pattern, each of the one or more cameras including a camera sensor comprising a pixel array. The method further includes using a processor to run a correspondence algorithm to compute corresponding three-dimensional positions of a plurality of detected features of the projected pattern for each of the plurality of images on the intraoral three-dimensional surface. The processor uses data corresponding to the corresponding three-dimensional positions of at least three features to estimate the three-dimensional surface in which all at least three features reside, each feature corresponding to a respective projector ray r. For projector ray r1, where no three-dimensional position of a feature corresponding to projector ray r1 is computed, or more than one three-dimensional position of a feature corresponding to projector ray r1 is computed, the processor estimates the three-dimensional position in space of the intersection of projector ray r1 and the estimated three-dimensional surface. The processor then uses the estimated three-dimensional position in space to identify a search space in the pixel array of at least one camera, wherein a feature corresponding to projector ray r1 is to be searched in the search space. In one embodiment, a correspondence algorithm is run based on stored calibration values ​​indicating (a) camera rays corresponding to each pixel on a camera sensor of each of one or more cameras, and (b) projector rays corresponding to each feature of a projection pattern from each of one or more structured light projectors, whereby each projector ray corresponds to a corresponding path to a pixel on at least one camera sensor. In one embodiment, the search space in the data includes a search space defined by one or more thresholds. In one embodiment, the light pattern includes a distribution of discrete spots, and each feature includes spots from the distribution of discrete spots.

[0594] Now for reference Figure 43A-B, which is a flowchart outlining various methods for tracking the movement of an intraoral scanner (i.e., handheld stick 20) ​​according to some applications of the invention. For the purpose of object scanning, it is always necessary to estimate the position of the scanner relative to the object being scanned 32 (i.e., the intraoral three-dimensional surface) during scanning. Typically, the intraoral scanner can use at least one camera 24 coupled to the intraoral scanner to measure the movement of the intraoral scanner relative to the object being scanned 32 via visual tracking (step 274). Visual tracking of the movement of the intraoral scanner relative to the object being scanned 32 is obtained by stitching together the individual surfaces or point clouds obtained from adjacent image frames or by using a Simultaneous Localization and Mapping (SLAM) algorithm, which in turn provides information about how the intraoral scanner moves between one frame and the next. However, sometimes sufficient visual tracking of the movement of the intraoral scanner relative to the object 32 may not be possible during scanning, for example, in areas of the intraoral scene that are difficult to capture, or when moving tissue (e.g., such as a patient's tongue, a patient's cheek, or a doctor's finger) obstructs the camera. An inertial measurement unit (IMU) coupled to the intraoral scanner can measure the movement of the intraoral scanner relative to a fixed coordinate system. However, using the IMU alone is usually insufficient to determine the position of the intraoral scanner relative to the scanned object 32, since the object 32 is part of the object's own movable head.

[0595] Therefore, the inventors have developed a method that combines (a) visual tracking of scanner motion with (b) inertial measurement of scanner motion to (i) accommodate time when sufficient visual tracking is unavailable, and optionally (ii) when visual tracking is available, to help provide an initial guess of the movement of the intraoral scanner relative to object 32 from one frame to the next, so that only the refinement of the intraoral scanner position obtained from visual tracking is needed, thereby reducing stitching time. In step 274, at least one camera (e.g., camera 24) coupled to the intraoral scanner is used to measure (A) the motion of the intraoral scanner relative to the scanned intraoral surface. In step 276, at least one IMU coupled to the intraoral scanner is used to measure (B) the motion of the intraoral scanner relative to a fixed coordinate system (i.e., the Earth reference system). In step 278, a processor (e.g., processor 96) is used to calculate the motion of the intraoral surface relative to the fixed coordinate system by subtracting the motion of the intraoral scanner relative to the intraoral surface from the motion of the intraoral scanner relative to the fixed coordinate system (B). Alternatively, the motion of the intraoral surface relative to a fixed coordinate system can be calculated in other ways based on (A) the motion of the intraoral scanner relative to the intraoral surface and (B) the motion of the intraoral scanner relative to the fixed coordinate system. The motion of the intraoral surface can be calculate...

Claims

1. A method of generating a digital three-dimensional image, the method comprising: driving each of one or more structured light projectors of an intraoral scanner to project a pattern on an intraoral three-dimensional surface; driving each of one or more cameras of the intraoral scanner to capture a plurality of images, each image including at least a portion of the projected pattern; and using a processor to: compare a series of images captured by the one or more cameras at different points in time; based on the comparison of the series of images, determine which portions of the projected pattern are trackable across the series of images; solve a correspondence algorithm for the tracked portions of the projected pattern in at least one image of the series of images; and construct a three-dimensional model of the intraoral three-dimensional surface based at least in part on the comparison of the series of images, wherein the solution of the correspondence algorithm is used to construct the three-dimensional model.

2. The method of claim 1, wherein, using the processor further comprises using the processor to: use the solved correspondence algorithm for the tracked portions in the at least one image of the series of images to solve a correspondence algorithm for the tracked portions of the projected pattern in at least another image of the series of images.

3. The method of claim 1, wherein, using the processor further comprises using the processor to: solve the correspondence algorithm for the tracked portions based on the locations of the tracked portions of the projected pattern in each image of the series of images, wherein constructing the three-dimensional model comprises using the solution of the correspondence algorithm to construct the three-dimensional model.

4. The method of claim 1, wherein, the one or more structured light projectors project a pattern that is spatially fixed relative to the one or more cameras.

5. The method of any one of claims 1-4, wherein, the projected pattern includes a plurality of projected spots, and wherein the portions of the projected pattern correspond to projected spots s of the plurality of projected spots.

6. The method of claim 5, wherein, using the processor to compare the series of images comprises using the processor to compare the series of images based on stored calibration values indicating (a) camera rays corresponding to each pixel on a camera sensor of each of the one or more cameras, and (b) projector rays corresponding to each of the projected spots from each of the one or more structured light projectors, wherein each projector ray corresponds to a respective path of a pixel on at least one camera sensor, wherein determining which portions of the projected pattern are trackable across the series of images comprises determining which projected spots s are trackable across the series of images, and wherein each tracked spot s moves along a path of a pixel corresponding to a respective projector ray r.

7. The method of claim 6, wherein, using the processor further comprises using the processor to, for each tracked spot s, determine a plurality of possible paths p of a pixel on a given one of the cameras, a path p corresponding to a respective plurality of possible projector rays r.

8. The method of claim 7, wherein, using the processor further comprises using the processor to run a correspondence algorithm to: for each possible projector ray r: identifying how many other cameras detected a respective spot q corresponding to a respective camera ray that intersects the projector ray r and the camera ray of the given one of the cameras corresponding to the tracked spot s; identifying a given projector ray r1 for which a maximum number of other cameras detected a respective spot q; and identifying the projector ray r1 as the particular projector ray r that produced the tracked spot s.

9. The method of claim 6, wherein, using the processor further comprises using the processor to: run the correspondence algorithm to compute respective three-dimensional positions of a plurality of detected spots on the intraoral three-dimensional surface captured in the series of images, in at least one of the series of images, identify a detected spot as being from a particular projector ray r by identifying the detected spot as being a tracked spot s that moves along a path of pixels corresponding to the particular projector ray r.

10. The method of claim 6, wherein, the using processor further comprises using the processor to: run the correspondence algorithm to compute respective three-dimensional positions of a plurality of detected spots on the intraoral three-dimensional surface captured in the series of images, and remove from being considered points on the intraoral three-dimensional surface spots that are (i) identified as being from a particular projector ray r based on the three-dimensional positions computed by the correspondence algorithm and (ii) not identified as being tracked spots s that move along a path of pixels corresponding to the particular projector ray r.

11. The method of claim 6, wherein, using the processor further comprises using the processor to: run the correspondence algorithm to compute respective three-dimensional positions of a plurality of detected spots on the intraoral three-dimensional surface captured in the series of images, and for a detected spot that is identified as being from two different projector rays r based on the three-dimensional positions computed by the correspondence algorithm, identify the detected spot as being from one of the two different projector rays r by identifying the detected spot as being a tracked spot s that moves along the one of the two different projector rays r.

12. The method of claim 6, wherein, using the processor further comprises using the processor to: run the correspondence algorithm to compute respective three-dimensional positions of a plurality of detected spots on the intraoral three-dimensional surface captured in the series of images, and identify a faint spot as being a projected spot from a particular projector ray r by identifying the faint spot as being a tracked spot s that moves along a path of pixels corresponding to the particular projector ray r, a three-dimensional position of the faint spot not being computed by the correspondence algorithm. using the processor further comprises using the processor to: run the correspondence algorithm to compute respective three-dimensional positions of a plurality of detected spots on the intraoral three-dimensional surface captured in the series of images, and identify a faint spot as being a projected spot from a particular projector ray r by identifying the faint spot as being a tracked spot s that moves along a path of pixels corresponding to the particular projector ray r, a three-dimensional position of the faint spot not being computed by the correspondence algorithm.

13. The method of claim 6, wherein, Using the processor further includes using the processor to compute a respective three-dimensional position on the intra-oral three-dimensional surface at an intersection of a projector ray r and a corresponding camera ray corresponding to a tracked speckle s in each image of the series of images, wherein the speckle s is tracked across the series of images.

14. The method of any one of claims 1-4, wherein, Constructing the three-dimensional model using the correspondence algorithm, wherein the correspondence algorithm uses at least in part a portion of the projected pattern determined to be trackable across the series of images.

15. The method of any one of claims 1-4, wherein, Using the processor further includes using the processor to: determine parameters of the tracked portion of the projected pattern in at least two adjacent images from the series of images, the parameters selected from the group consisting of: a size of the portion, a shape of the portion, an orientation of the portion, an intensity of the portion, and a signal-to-noise ratio SNR of the portion; and predict parameters of the tracked portion of the projected pattern in a subsequent image based on the parameters of the tracked portion of the projected pattern in the at least two adjacent images.

16. The method of claim 15, wherein, Using the processor further includes using the processor to search for a portion of the projected pattern in the subsequent image that substantially has the predicted parameters based on the predicted parameters of the tracked portion of the projected pattern.

17. The method of claim 15, wherein, The selected parameter is a shape of a portion of the projected pattern, and wherein using the processor further includes using the processor to determine a search space in a next image in which to search for the tracked portion of the projected pattern based on a predicted shape of the tracked portion of the projected pattern.

18. The method of claim 17, wherein, Using the processor to determine the search space includes using the processor to determine a search space in the next image in which to search for the tracked portion of the projected pattern, the search space having a size and an aspect ratio based on a size and an aspect ratio of the predicted shape of the tracked portion of the projected pattern.

19. The method of claim 15, wherein, The selected parameter is a shape of a portion of the projected pattern, and wherein using the processor further includes using the processor to: determine a velocity vector of the tracked portion of the projected pattern based on a direction and a distance that the tracked portion of the projected pattern has moved between the at least two adjacent images from the series of images; predict a shape of the tracked portion of the projected pattern in a subsequent image in response to a shape of the tracked portion of the projected pattern in at least one of the at least two adjacent images; and determine a search space in the subsequent image in which to search for the tracked portion of the projected pattern in response to (i) determining the velocity vector of the tracked portion of the projected pattern in combination with (ii) the predicted shape of the tracked portion of the projected pattern.

20. The method of claim 15, wherein, The parameter is a shape of a portion of the projected pattern, and wherein using the processor further includes using the processor to: based on a direction and distance that the tracked portion of the projected pattern has moved between the at least two adjacent images from the series of images, determine a velocity vector for the tracked portion of the projected pattern; in response to determining the velocity vector for the tracked portion of the projected pattern, predict a shape of the tracked portion of the projected pattern in a later image; and in response to (i) determining the velocity vector for the tracked portion of the projected pattern, in combination with (ii) the predicted shape of the tracked portion of the projected pattern, determine a search space in the later image in which to search for the tracked portion of the projected pattern.

21. The method of claim 20, wherein, using the processor includes using the processor to predict a shape of the tracked portion of the projected pattern in a later image in response to (i) determining the velocity vector for the tracked portion of the projected pattern, in combination with (ii) a shape of the tracked portion of the projected pattern in at least one of the two adjacent images.

22. The method of any one of claims 1-4, wherein, using the processor also includes using the processor to: based on a direction and distance that the tracked portion of the projected pattern has moved between at least two consecutive images in the series of images, determine a velocity vector for the tracked portion of the projected pattern; and in response to determining the velocity vector for the tracked portion of the projected pattern, determine a search space in a later image in which to search for the tracked portion of the projected pattern.

23. An intraoral scanning system comprising: an intraoral scanner and a processor, wherein the intraoral scanning system is configured to perform the method of any one of claims 1-22.

Citation Information

Patent Citations

  • 3D photogrammetry using projected patterns

    US20080101688A1