Methods, systems, and computer-readable media for registering intraoral measurements
By using deep learning methods to automatically identify and correct registration errors in intraoral measurements, the registration problem caused by soft tissue deformation was solved, and high-quality 3D registration and cleaned 3D model generation were achieved.
Patent Information
- Application Number
- CN202080064381.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-24
- Filing Date
- 2020-09-22
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2040-09-22
AI Technical Summary
Existing technologies make it difficult to achieve accurate three-dimensional registration in intraoral measurements due to registration errors and interruptions caused by soft tissue deformation.
By employing deep learning methods, a deep neural network is trained to automatically identify registration error sources in individual images. The output labels and probability values are used for image segmentation, and registration is performed based on predetermined weights to eliminate or reduce registration errors and generate cleaned 3D models.
It improves the accuracy of intraoral measurements, reduces registration errors and interruptions, generates high-quality global 3D images, and improves the scanning process.
Smart Images

Figure CN114424246B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This patent application claims the benefit and priority of U.S. Application No. 16 / 580,084, filed September 24, 2019, which is incorporated herein by reference for all purposes. Invention Field
[0003] This application generally relates to a method, system, and computer-readable storage medium for registration in intraoral measurements, and more specifically, to a method, system, and computer-readable storage medium for semantically registering intraoral measurements using deep learning methods. Background Technology
[0004] Dentists can be trained to produce satisfactory acquisition results during scanning by using appropriate scanning techniques, such as keeping soft tissue outside the dental camera's field of view. Soft tissue may deform during scanning, resulting in multiple shapes for the same area, which can introduce errors and / or interruptions during registration.
[0005] Currently, feature-based techniques such as Fast Point Feature Histogram (FPFH) can be used to compute transformations that enable registration of scan / 3D measurements without prior knowledge of the relative orientation of the scans. However, for these techniques to be effective, it may be necessary to avoid scanning / 3D measurements in potentially deformable areas.
[0006] U.S. Patent No. 9,456,754 B2 discloses a method for recording multiple 3D images of a dental object, wherein each 3D image may include color data and 3D measurement data of the object's measurement surfaces, wherein a computer-aided recording algorithm is used to combine individual images into a whole image. This patent is incorporated herein by reference for all purposes as if it were fully disclosed herein.
[0007] U.S. Patent No. 7,698,068B2 discloses a method for providing data useful in oral-related procedures by: providing at least one digital entity representing the three-dimensional surface geometry and color of at least a portion within the oral cavity; and manipulating the entity to provide desired data from it. Typically, the digital entity includes surface geometry and color data associated with said portion within the oral cavity, and the color data includes actual or perceived visual characteristics, including hue, chroma, value, translucency, and reflectance.
[0008] WO2018219800A1 discloses a method and apparatus for generating and displaying a 3D representation of a portion of an intraoral scene, comprising determining 3D point cloud data representing a portion of the intraoral scene in a point cloud coordinate space; acquiring a color image of the same portion of the intraoral scene in a camera coordinate space; and marking color image elements within an image region representing the surface of the intraoral scene.
[0009] US Patent No. 9436868B2 discloses a method for rapid automatic object classification of a measured three-dimensional (3D) object scene. The target scene is illuminated with a light pattern, and an image sequence of the target scene illuminated by the pattern at different spatial phases is acquired.
[0010] U.S. Patent No. 9788917B2 discloses a method for employing artificial intelligence in automated orthodontic diagnosis and treatment planning. The method may include: providing an intraoral imager configured for patient operation; receiving patient data regarding the orthodontic condition; accessing a database containing or having access to information obtained from orthodontic treatment; generating an electronic model of the orthodontic condition; and instructing at least one computer program to analyze the patient data and identify at least one diagnostic and treatment plan for the orthodontic condition based on the information obtained from orthodontic treatment.
[0011] U.S. Patent Application Publication No. 20190026893A1 discloses a method for evaluating the shape of an orthodontic appliance, wherein an analysis image is submitted to a deep learning device to determine tooth attribute values associated with teeth represented on the analysis image, and / or at least one value of an image attribute associated with the analysis image.
[0012] PCT application PCT / EP2018 / 055145 discloses a method for constructing restorations, in which the condition of teeth is measured using a dental camera and a three-dimensional (3D) model of the tooth condition is generated. Computer-aided detection algorithms can then be applied to the 3D model of the tooth condition and automatically determine the type of restoration, the number of teeth, or the location of the restoration.
[0013] U.S. Patent Application Publication No. 20180028294A1 discloses an automated dental CAD method using deep learning. The method may include: receiving patient scan data representing at least a portion of a patient's dentition dataset; and using a trained deep neural network to identify one or more tooth features in the patient scan. Here, design automation can be performed after the complete scan is generated. However, this method does not improve the actual scanning process.
[0014] WO2018158411A1 discloses a method for constructing restorations, wherein the condition of teeth is measured using a dental camera and a 3D model of the tooth condition is generated. In this case, a computer-aided detection algorithm is applied to the 3D model of the tooth condition, wherein the type of restoration and / or at least the number of teeth and / or the location of the restoration to be inserted are automatically determined. Summary of the Invention
[0015] Methods, systems, and computer-readable storage media for semantically registering intraoral measurements using deep learning approaches can overcome existing limitations, as well as others, associated with the foregoing.
[0016] In one aspect of this invention, the present invention can provide a computer-implemented method for three-dimensional (3D) registration, the method comprising: receiving an individual image of a patient's dentition via one or more computing devices; automatically identifying sources of registration errors in the individual image using one or more output labels (e.g., output probability values of a trained deep neural network), wherein the output labels / probability values are obtained by segmenting the individual image into regions corresponding to one or more objects; wherein the individual image is a depth and / or corresponding color image; the method further comprising registering the individual images together based on one or more output labels such as probability values to form a registered 3D image with no or substantially no registration errors.
[0017] In another aspect of this paper, the computer-implemented method may further include one or more combinations of the following steps: (i) wherein registration is achieved by generating a point cloud from a depth image by projecting pixels of the depth image into space; assigning a color value and a label / probability value to each point in the point cloud, respectively, using a corresponding color image and an output label / probability value of a trained deep neural network; and discarding or partially including points in the point cloud using predetermined weights based on the assigned label / probability value, thereby eliminating or reducing the contribution of the discarded or partially included points to the registration; (ii) wherein the individual image is an individual three-dimensional optical image; (iii) wherein the individual image is received as a time series of images; (iv) wherein the individual image is received as a pair of color and depth images; (v) wherein the one or more object categories include hard gingiva, soft tissue gingiva, teeth, and dentition; and (vi) wherein the indication of the relevance of the identified registration error sources is based on the geometry around them. The shape, (vii) wherein the deep neural network is a network selected from the group consisting of: convolutional neural networks (CNN), fully convolutional neural networks (FCN), recurrent neural networks (RNN), and recurrent convolutional neural networks (Recurrent-CNN), (vii) further comprising: training the deep neural network using one or more computing devices and multiple individual training images to map one or more tissues in at least a portion of each training image to one or more labels / probability values, wherein the training is performed at the pixel level by classifying individual training images, pixels of individual training images, or superpixels of individual training images into one or more classes corresponding to semantic data types and / or erroneous data types, (viii) wherein the training images comprise a 3D mesh and a registered pair of depth and color images, (ix) wherein the 3D mesh is labeled, and a transformation function is used to transfer labels to the registered pairs of 3D and color images.
[0018] In another aspect of the invention, a non-transitory computer-readable storage medium storing a program, when executed by a computer system, causes the computer system to perform a process comprising the steps of: receiving an individual image of a patient's dentition via one or more computing devices; automatically identifying sources of registration errors in the individual image using one or more output probability values of a trained deep neural network, wherein the output probability values are obtained by segmenting the individual image into regions corresponding to one or more object categories; wherein the individual image is a depth and / or corresponding color image; the method further comprising registering the individual images together based on one or more output probability values to form a registered 3D image with no or substantially no registration errors.
[0019] Furthermore, a system for three-dimensional (3D) registration can be provided, the system including a processor configured to: receive individual images of a patient's dentition via one or more computing devices; automatically identify sources of registration errors in the individual images using one or more output probability values of a trained deep neural network, wherein the output probability values are obtained by segmenting the individual images into regions corresponding to one or more object categories; wherein the individual images are depth and / or corresponding color images; wherein the processor is configured to register the individual images together based on the one or more output probability values to form a registered 3D image with no or substantially no registration errors.
[0020] In another aspect of the invention, the deep neural network in the system is selected from the group consisting of convolutional neural networks (CNN), fully convolutional neural networks (FCN), recurrent neural networks (RNN), and recurrent convolutional neural networks (Recurrent-CNN). Attached Figure Description
[0021] The exemplary embodiments will be more fully understood from the following detailed description and accompanying drawings, in which like elements are designated by like reference numerals, and these drawings are given by way of illustration only and therefore do not limit the exemplary embodiments herein. In the drawings:
[0022] Figure 1 This is a schematic diagram of a cross-section of the oral cavity, showing the different surrounding geometries caused by soft tissue deformation.
[0023] Figure 2A This is a schematic diagram of a top view of the oral cavity, illustrating the scanning / recording of individual images of the patient's dentition.
[0024] Figure 2B This is a schematic diagram illustrating an example registration according to an embodiment of the present invention.
[0025] Figure 3A This is a high-level block diagram of a system according to an embodiment of the present invention.
[0026] Figure 3B An example training image is shown according to an embodiment of the present invention.
[0027] Figure 4 It is a perspective view of a global 3D image of a dental arch formed from individual images of soft tissue.
[0028] Figure 5 It is a perspective view of a corrected global 3D image of a dental arch formed from individual images of which soft tissue contributions have been removed or reduced in weight, according to an embodiment of the present invention.
[0029] Figure 6This is a high-level block diagram illustrating the structure of a deep neural network according to one embodiment.
[0030] Figure 7A This is a flowchart illustrating a method according to an embodiment of the present invention.
[0031] Figure 7B This is a flowchart illustrating a method according to an embodiment of the present invention.
[0032] Figure 8 This is a block diagram illustrating a training method according to an embodiment of the present invention.
[0033] Figure 9 This is a block diagram illustrating a computer system according to an exemplary embodiment of the present invention.
[0034] Different figures may have at least some of the same reference numerals to identify the same components; however, a detailed description of each such component may not be provided below for each figure. Invention Details
[0036] Based on the example aspects described herein, a method, system, and computer-readable storage medium may be provided for semantically segmenting and registering individual intraoral measurements using deep learning methods.
[0037] System for registration intra-oral measurements
[0038] Incorrect registration can hinder accurate 3D measurements of a patient's oral cavity. In intraoral measurements of the jawbone, a camera is used, producing single scans that capture only a subset of the entire jawbone. These scans can be registered together to form a complete model. The camera can be handheld, and the exact location of each individual scan is often unknown. Transformations are determined based on information from these individual scans (e.g., 3D data, color data) to bring the individual scans into a common reference frame (common 3D coordinate system). However, when the camera performs multiple single scans at high frequency, deformation / shape changes in the oral cavity can distort the registration because most registration processes are performed under rigidity assumptions. Therefore, only rigid components should be considered for registration.
[0039] Because scans are performed at different points in time, the geometry of certain tissues (especially the soft tissues of the oral cavity) may change between scans due to soft tissue deformation or the presence of moving foreign bodies. This can hinder registration that relies on matching 3D data (see...). Figure 1 This illustrates typical errors, such as those generated by techniques based on minimizing the sum of squared errors.
[0040] These improvements to the technique can be achieved by considering only the rigid portions used for registration and by discarding irrelevant (i.e., non-rigid) portions or by weighting their contribution to registration less, i.e., when rigid portions or hard tissue 12 (e.g., teeth) are considered for registration, the registration is robust (iii) and the surrounding geometries 13a, 13b of the rigid portions or hard tissue 12 are aligned, and vice versa (iv), as... Figure 1 As shown. Therefore, the term "rigid" can be used below to describe anatomical structures or parts thereof that are unlikely to deform during the time period of the scanning procedure. If sufficient force is applied to the gingiva near the teeth, it may deform, but this is generally not the case when scanning with an intraoral scanner. Therefore, it would be a reasonable assumption to consider it rigid. On the other hand, the inner side of the cheek may deform on a scan-by-scan basis and can therefore be considered soft tissue 15 / soft portion ( Figure 2A ).
[0041] Preferably, the system described herein can acquire images, such as individual three-dimensional optical images 2 ( Figure 2A Each three-dimensional optical image 2 preferably includes 3D measurement data and color data of the measurement surface of the tooth, and is preferably recorded sequentially in the oral cavity by direct intraoral scanning. For example, this may occur in a dental clinic or office and may be performed by a dentist or dental technician. Images can also be obtained indirectly through a sequence of stored images.
[0042] Using these images (preferably obtained as a time series), a computer-implemented system can automatically identify regions in the images that can be considered for registration. This can be done in real time. Of course, the images can also be individual two-dimensional (2D) images, RGB images, range images (2.5D), or 4-channel images (RGB-D), where depth and color may not be perfectly aligned, i.e., depth and color images may be obtained at different time periods.
[0043] During the scanning process, multiple individual images can be created, and then sequences of at least two individual images or multiple sequences of images can be combined to form a global 3D image. Figure 4 More specifically, such as Figure 2AAs shown, an individual 3D optical image 2 (shown in rectangular form) can be acquired by a scanner / dental camera 3, which can move relative to the object 1 along a measurement path 4 during measurement. In some embodiments, the measurement path can be arbitrary, i.e., measurements can be taken from different directions. The dental camera 3 can be a handheld camera, for example, that uses edge projection to measure the object 1. Other methods of 3D measurement will be recognized by those skilled in the art. A first overlapping region 5 between the first image 6 and the second image 7, shown by dashed lines, is examined using a computer to determine if recording conditions are met, and if so, the 3D optical images 2 can be combined / registered together to form a global 3D image 10.
[0044] Recording conditions can include sufficient size, sufficient waviness, sufficient roughness, and / or a sufficient number and arrangement of feature geometries. However, it can be difficult to program conventional computers to identify the sources of registration errors and how to prevent them. Manually programming the features of registration or segmentation methods to cover all possible scenarios can be tedious, especially given the high frequency of measurements. This is particularly true if the context of the entire image needs to be considered. Using machine learning methods (especially neural networks) and proper training data can solve the problem more effectively. On the other hand, neural networks can learn from a single scan / single 3D measurement to identify the sources of registration errors and semantically segment the data, and decide whether these areas of the oral cavity can be considered for registration. To this end, the labels for different objects / object categories segmented can be defined as including, but not limited to: (i) hard tissue (e.g., teeth, crowns, bridges, hard gingiva near teeth, and other tooth-like objects), (ii) soft tissue (e.g., tongue, cheek, soft gingiva, etc.), and (iii) disposable items for instruments / intraoral applications (e.g., mirrors, scanners, cotton rolls, brackets, etc.). Of course, other definitions can be added as appropriate, such as glare 21 ( Figure 2B (Caused by strong light). Segmentation can be based on color, and color can be interpreted in a context-aware manner, that is, the relevance of potential sources of registration error can be indicated based on the surrounding geometry 13a, 13b. Furthermore, segmentation can be based on a single pixel of color data or a larger area of an individual 3D optical image 2.
[0045] Since the crown, teeth, or hard gingiva near the teeth are rigid, registration errors can be eliminated or significantly reduced by considering a registration algorithm that takes into account proper segmentation. Furthermore, by removing clutter introduced by accessories such as cotton rolls, a cleaned 3D model can be generated, containing only data relevant to dental treatment.
[0046] Therefore, the system can train neural networks (such as deep neural networks) using multiple training datasets to automatically identify sources of registration errors in the 3D optical image 2 and prevent these sources from contributing to the registration, preferably in real time. This reduces or eliminates the propagation of errors to the global 3D image 10. Figure 4 ) incorrect registration ( Figure 1 ,iv), such as Figure 5 As shown in the corrected global 3D image 9, and / or due to fewer / no interruptions caused by misregistration, the scanning process can be improved.
[0047] This system can also semantically identify and label data (in a context-aware manner, i.e., context may be important for selecting an appropriate correction method. For example, gingiva close to the teeth can be considered hard tissue 12, while gingiva far from the teeth can be considered soft tissue 15).
[0048] Furthermore, the system can determine and / or apply corrective measures when a source of registration error is detected. For example, when the ratio of hard tissue 12 to soft tissue 15 is high in an individual's three-dimensional optical image 2, it may be advantageous to give the hard tissue 12 significantly more weight than the soft tissue 15, because deformation or movement of the patient's cheeks or lips may cause soft tissue deformation, resulting in erroneous recording, such as... Figure 1 As shown in (iv). However, if the proportion of hard tissue in the image is low, soft tissue 15 can be weighted more heavily to improve recording quality. As an example, Figure 2B An image is shown where two central teeth are missing, such that a first region with hard tissue 12 occupies approximately 10% of the total area of the individual 3D optical image 2, and a second region with soft tissue 15 occupies approximately 50% of the total image area. A remaining third region, which cannot be assigned to any tissue, occupies approximately 40%; the percentage of the first region with hard tissue 12 is below a predetermined threshold of 30%. Therefore, the first region with hard tissue 12 is weighted by a first weighting factor, and the second region with soft tissue is weighted by a second weighting factor, such that the second weighting factor increases as the first region decreases. For example, when the predetermined threshold is exceeded, the first weighting factor for the first region with hard tissue 12 can be 1.0, while the second variable weighting factor for the second region with soft tissue 15 can be 0.5, and can reach 1.0 as the first region decreases. The dependence of the second weighting factor on the first region can be defined according to any function such as an exponential or linear function, and the hard and soft tissues can be segmented using the neural network described herein.
[0049] Figure 3AA block diagram of a system 200 for identifying dental information from an individual three-dimensional optical image 2 of a patient's dentition, according to one embodiment, is shown. System 200 may include a dental camera 3, a training module 204, an image registration module 206, a computer system 100, and a database 202. In another embodiment, the database 202, the image registration module 206, and / or the training module 204 may be part of the computer system 100, and / or may be able to directly and / or indirectly adjust the parameters of the dental camera 3 based on a correction scheme. The computer system 100 may also include at least one computer processor 122, a user interface 126, and an input unit 130. The computer processor may receive various requests and may load appropriate instructions stored on a storage device into memory, and then execute the loaded instructions. The computer system 100 may also include a communication interface 146, which enables software and data to be transferred between the computer system 100 and external devices.
[0050] The computer system 100 can receive registration requests from external devices such as a dental camera 3 or from a user (not shown), and can load appropriate instructions for semantic registration. Preferably, the computer system can independently register images upon receiving a separate three-dimensional optical image 2 without waiting for a request.
[0051] In one embodiment, computer system 100 may use multiple training datasets (which may include, for example, multiple individual three-dimensional optical images 2) from database 202 to train one or more deep neural networks, and database 202 may be part of training module 204. Figure 3B(i-iv) illustrate example images used for training, including color images, depth images, color images mapped to depth images, and depth images mapped to color images. Mapping images may be advantageous for convolutional neural networks (CNNs) because CNNs operate on local neighborhoods. Therefore, the representation regions of RGB and depth images can be represented using the same 2D pixel coordinates. For example, mapping a depth image to an RGB image or vice versa means generating images such that pixels with the same 2D pixel coordinates in the RGB image and the generated image correspond to the same points in space. This typically involves the application of a pinhole camera model and adaptation to motion transformations (determined by a registration algorithm). The motion compensation step can be omitted, and the network can be expected to handle the resulting offset (which is expected to be small). In some embodiments, system 200 may include a neural network module (not shown) containing various deep learning neural networks, such as convolutional neural networks (CNNs), fully convolutional neural networks (FCNs), recurrent neural networks (RNNs), and recurrent convolutional neural network networks (recurrent CNNs). An example fully convolutional neural network is described in the publication "Fully Convolutional Networks for Semantic Segmentation" by Jonathan Long et al., dated March 8, 2015, which is incorporated herein by reference in its entirety as it is fully disclosed herein. Thus, a fully convolutional neural network (an efficient convolutional network architecture for per-pixel segmentation) can be trained to segment RGB(D) images or sequences of RGB(D) images using a recurrent model. Recurrent models can be used in contrast to simple feedforward networks. Therefore, the network can receive the output of a layer as input to the next forward computation, so its current activation can be considered dependent on the states of all previous inputs, thus enabling it to process sequences. Furthermore, an example recurrent CNN model is described in the publication "Recurrent Convolutional Neural Networks: A Better Model of Biological Object Recognition" by Courtney J. Spoerer et al., dated September 12, 2017, which is incorporated herein by reference in its entirety as it is fully disclosed herein.
[0052] The training dataset and / or input of a neural network can be preprocessed. For example, to process color data in conjunction with 3D measurements, calibration (e.g., determining the parameters of a camera model) can be applied to align the color image with a 3D surface. Furthermore, standard data augmentation procedures such as synthetic rotation and scaling can be applied to the training dataset and / or input.
[0053] Training module 204 can use a labeled training dataset to supervise the learning process of a deep neural network. Labels can be used to weight data points. Training module 204 can conversely use an unlabeled training dataset to train a generative deep neural network.
[0054] In an example embodiment, to train a deep neural network to detect sources of registration errors, a dataset of three-dimensional optical images of multiple real-life individuals with tissue types and object categories as described above can be used. In another example, to train a deep neural network to recognize semantic data (e.g., hard gingiva near teeth), multiple additional training datasets from real dental patients, comprising one or more hard gingival regions near teeth and one or more soft gingival regions away from teeth, are selected to form a set of training datasets. Database 202 can therefore contain different sets of training datasets, for example, one set for each object category and / or each semantic data type.
[0055] In some embodiments, training module 204 may pre-train one or more deep neural networks using a training dataset from database 202, enabling computer system 100 to easily use one or more pre-trained deep neural networks to detect sources of registration errors. It can then send information about the detected sources and / or the individual 3D optical image 2, preferably automatically and in real-time, to image registration module 206, where the sources of registration errors are considered before registration.
[0056] Database 202 can also store data related to the deep neural network and the identified sources, as well as the corresponding individual three-dimensional optical images 2. Furthermore, computer system 100 may have a display unit 126 and an input unit 130, through which users can perform functions such as submitting requests and receiving and viewing identified sources of registration errors during training.
[0057] In an example implementation of the training process, S600, as follows: Figure 8As shown, labels can be generated by collecting images representing real-world cases, step S602. These cases may involve meshes (e.g., 3D triangular meshes) and individual images (registration pairs of depth and color images). The meshes can be segmented by an expert capable of cutting out teeth from the meshes. The cut-out meshes can then be labeled, step S604. They can be labeled as teeth, crowns, implants, or other tooth-like objects (step S606), or as hard and soft gums (step S608). Furthermore, in step S610, outliers, such as 3D points in the optical path of camera 3 that do not belong to the oral cavity, can be labeled. The labeling of the meshes can be transferred to the pixels of individual images, thereby reducing the workload during training. In step S612, all final labels can be determined by combining information from steps S606, S608, and S610, step S612. Furthermore, knowing the transformation that aligns the individual images together (because they are already registered), these final labels can be transferred from the cut-out meshes to the individual images. Thus, many images can be labeled simultaneously through mesh cutting / slicing.
[0058] Other embodiments of system 200 may include different and / or additional components. Furthermore, functionality may be distributed among the components in a manner different from that described herein.
[0059] Figure 6 A block diagram illustrating the structure of a deep neural network 300 according to an embodiment of the present invention is shown. It may have several layers, including an input layer 302, one or more hidden layers 304, and an output layer 306. Each layer may consist of one or more nodes 308, represented by small circles. Information can flow from the input layer 302 to the output layer 306, i.e., from left to right, although in other embodiments it may be from right to left, or both. For example, a recurrent network might consider previously observed data when processing new data in sequence 8 (e.g., the current image could be segmented considering previous images), while a non-recurrent network processes new data in isolation.
[0060] Nodes 308 can have inputs and outputs, and the nodes 308 in the input layer can be passive, meaning they cannot modify the data. For example, nodes 308 in input layer 302 can each receive a single value (e.g., a pixel value) on their input and copy that value to their multiple outputs. Conversely, nodes in hidden layer 304 and output layer 306 can be active and therefore capable of modifying the data. In the example structure, each value from input layer 302 can be copied and sent to all hidden nodes. Values entering hidden nodes can be multiplied by weights, which can be a set of predetermined numbers associated with each hidden node. The weighted inputs can then be summed to produce a single number.
[0061] In an embodiment of the present invention, when detecting object categories, the deep neural network 300 can use pixels from an individual three-dimensional optical image 2 as input. The individual three-dimensional optical image 2 can be a color image and / or a depth image. Here, the number of nodes in the input layer 302 can be equal to the number of pixels in the individual three-dimensional optical image 2.
[0062] In one example embodiment, a single neural network can be used for all object categories, while in another embodiment, different networks can be used for different object categories. In another example, when detecting object categories (such as those caused by ambient light), the deep neural network 300 can classify / label individual 3D optical images 2 (rather than individual pixels). In a further embodiment, the image can be a subsampled input, for example, every 4 pixels.
[0063] In yet another embodiment, the deep neural network 300 may have multiple data points acquired by the dental camera 3 as input, such as color images, depth measurements, acceleration, and device parameters such as exposure time and aperture. The deep neural network can output labels, which may be, for example, probability vectors comprising one or more probability values for each pixel input belonging to a certain object category. For example, the output may contain a probability vector containing probability values, where the highest probability value may define the location of the hard tissue 12. The deep neural network can also output a graph (or map) with no probability label values. A deep neural network may be created for each category, although this may not be necessary.
[0064] Methods for registration in intraoral measurements
[0065] Already described Figure 3A System 200, now for reference Figure 7A This illustrates process S400 according to at least some example embodiments described herein.
[0066] Process S400 can begin by obtaining and labeling the region of interest with predetermined labels in the training dataset, step S402. For example, Figure 3B The sample soft tissue 415 on the sample image 413 shown in (i) can be labeled as soft tissue. Figure 3B The sample hard tissue 412 on the sample image 413 shown in (i) can be labeled as hard tissue. The labeling of training images can be done numerically, for example by setting points on the image corresponding to the point of interest.
[0067] Training data can be labeled to assign semantics to individual 3D optical images2. This can be done at the per-pixel level for color or depth information. Alternatively, a mesh of the complete 3D model can be cut to compute corresponding per-pixel labels for individual images. Furthermore, the mesh can be segmented, allowing the labeling process to be automated. These labels can distinguish between teeth, cheeks, lips, tongue, gums, fillings, and ceramics, without assigning labels to anything else. Data unrelated to registration may include cheeks, lips, tongue, glare, and unlabeled data.
[0068] The training data can also be labeled to assign sources of mislabeled registrations to individual 3D optical images 2. This can also be done at the pixel level, for example, for image or depth information. For example, training data can be labeled at the pixel level for hard tissue 12 and soft tissue 15, and / or disposable items for instruments / intraoral applications, etc.
[0069] Semantic tags can overlap with tags used to identify registration error sources, such as tags like “hard tissue + glare”, “soft tissue adjacent to hard tissue”, and “tongue + hard tissue”, which can be distinguished from other tags such as “cheek + glare”.
[0070] Using this set of labeled or classified images, a deep neural network 300 can be constructed and fed labeled images, allowing the network to "learn" from them, enabling the network to generate network wiring that can spontaneously segment new images.
[0071] As an alternative to segmentation that involves classification on a per-pixel basis, segmentation can involve classification and training at a level slightly above the per-pixel level (i.e., at each "superpixel" level, where a "superpixel" is a part of the image that is larger than the normal pixels of the image).
[0072] The instructions and algorithms of process S400 can be stored in the memory of computer system 100 and can be loaded and executed by processor 122 to train (step S404) one or more deep neural networks using a training dataset to detect one or more defects based on one or more output labels / probability values. For example, if one of the probability values corresponding to the probability vector of glare is 90%, the neural network can detect glare 21 as one of the registration error sources in the individual three-dimensional optical image 2.
[0073] Training can be performed once, multiple times, or intermittently. Training can also be semi-supervised or self-supervised. For example, after the initial training, the deep neural network can receive or acquire previously unseen images and outputs, and can provide corresponding feedback, allowing the network to preferably eventually operate autonomously to classify images without human assistance. Therefore, the deep neural network 300 can be trained such that when a sequence 8 of individual 3D optical images 2 is input into the deep neural network 300, the deep neural network can return a resulting label / probability vector for each image, the label / probability vector indicating the category to which a portion of the image belongs.
[0074] After training, the deep neural network can acquire or receive a sequence 8 of individual three-dimensional optical images from the dental camera 3 for real-time segmentation (step S406), and can detect sources of registration errors in the images (step 408). When the source is detected, the image registration module 206 can register the images together based on predetermined weights of the segmented segments by ensuring that the detected registration error source does not contribute to the registration process (step S410). Figure 7A Steps S406-S410 are also included Figure 7B The flowchart will be discussed below. Figure 7B .
[0075] Figure 7BThe illustration shows process S500, which can be a subset of process S400. Process S500 can begin at step S502, where a color image and / or a depth image are acquired from a dental camera 3. In step S504, the color and depth images are used as input to a trained deep neural network 300, and a corresponding output labeled image is obtained, which shows the probability of segmentation segments belonging to the object category. Using both types of images makes it easier to distinguish different labels. For example, one might expect that training a network to distinguish between soft and hard gums based on color is more difficult than using depth and color. Since the images can be “mapped,” it may not make much difference which image is labeled. In one embodiment, one is labeled / segmented, while the other can simply provide additional features to determine the segmentation. According to this embodiment, there may be a one-to-one correspondence between the resulting labeled image and the depth image, or between the resulting labeled image and the color image. The labeled image may have the same lateral resolution and channels for labeling as the depth / color image. In step S506, a point cloud can be generated from the depth image by projecting each pixel of the depth image into space. For each point in the point cloud, a color value from a color image and a probability vector from a labeled image can be assigned. In one embodiment, the labeled image, the point cloud, and the resulting mesh can all have labels, label probabilities, or probability vectors assigned to them. In step S508, more images and corresponding output labels are acquired, such that each incoming point cloud is registered to the already registered point cloud. Here, points with a high probability of being hard tissue (e.g., above a predetermined threshold or weighted as a function of probability in a predetermined manner) are discarded or given less weight compared to other points with a high probability of being hard tissue. In step S510, each point in the incoming point cloud is added to the corresponding mesh cell to average position, color, and probability. The transformation that aligns the individual images to each other can then be optimized by using predetermined weights for soft tissue 15, hard tissue 12, and / or any other object category. If the transformation changes, the entries in the mesh sampler can be updated accordingly. Of course, implementations can be made according to this specification. Figure 7B Different other embodiments.
[0076] Computer system for registration of intraoral measurements
[0077] Already described Figure 7A and 7B The process, now refer to Figure 9This document illustrates a block diagram of a computer system 100 that can be used according to at least some of the exemplary embodiments described herein. Although various embodiments of this exemplary computer system 100 are described herein, those skilled in the art will understand how to implement the invention using other computer systems and / or architectures after reading this specification.
[0078] Computer system 100 may include a training module 204, a database 202, and / or an image registration module 206, or may be separate from the training module 204, database 202, and / or image registration module 206. These modules may be implemented in hardware, firmware, and / or software. The computer system may also include at least one computer processor 122, a user interface 126, and an input unit 130. In one exemplary embodiment, the input unit 130 may be used by a dentist in conjunction with a display unit 128, such as a monitor, to send instructions or requests during training. In another exemplary embodiment herein, the input unit 130 is a finger or stylus to be used on a touchscreen interface (not shown). Alternatively, the input unit 130 may also be a gesture / voice recognition device, a trackball, a mouse, or other input device, such as a keyboard or stylus. In one example, the display unit 128, the input unit 130, and the computer processor 122 may collectively form the user interface 126.
[0079] Computer processor 122 may include, for example, a central processing unit, a multiplexing unit, an application-specific integrated circuit (“ASIC”), a field-programmable gate array (“FPGA”), etc. Processor 122 may be connected to communication infrastructure 124 (e.g., a communication bus or network). In one embodiment herein, processor 122 may receive a request for 3D measurement and may automatically detect sources of registration errors in an image, and use image registration module 206 to automatically register the image based on the detected registration error sources. Processor 122 may achieve this by loading corresponding instructions stored in a non-transitory storage device as computer-readable program instructions and executing the loaded instructions.
[0080] Computer system 100 may further include main memory 132, which may be random access memory (“RAM”), and may also include secondary memory 134. Secondary memory 134 may include, for example, hard disk drive 136 and / or removable storage drive 138. Removable storage drive 138 may read from and / or write to removable storage unit 140 in a well-known manner. Removable storage unit 140 may be, for example, a floppy disk, magnetic tape, optical disk, flash memory device, etc., which may be written to and read from by removable storage drive 138. Removable storage unit 140 may include a non-transitory computer-readable storage medium storing computer-executable software instructions and / or data.
[0081] In another alternative embodiment, auxiliary memory 134 may include other computer-readable media storing a computer-executable program or other instructions to be loaded into computer system 100. Such a device may include: a removable storage unit 144 and an interface 142 (e.g., a program cartridge and a cartridge interface); a removable memory chip (e.g., an erasable programmable read-only memory (“EPROM”) or a programmable read-only memory (“PROM”)) and an associated memory slot; and other removable storage units 144 and interfaces 142 that allow software and data to be transferred from removable storage unit 144 to other parts of computer system 100.
[0082] Computer system 100 may also include a communication interface 146 that enables the transfer of software and data between computer system 100 and external devices. Such an interface may include a modem, a network interface (e.g., an Ethernet card, a wireless interface, an Internet cloud delivery hosting service, etc.), a communication port (e.g., a Universal Serial Bus (“USB”) port or... Port), PCMCIA (Personal Computer Memory Card International Association) interface, Software and data transmitted via communication interface 146 may be in the form of signals, which may be electronic, electromagnetic, optical, or other types of signals that can be sent and / or received by communication interface 146. Signals may be provided to communication interface 146 via communication path 148 (e.g., a channel). Communication path 148 may carry signals and may be implemented using wires or cables, optical fibers, telephone lines, cellular links, radio frequency (“RF”) links, etc. Communication interface 146 may be used to transfer software or data or other information between computer system 100 and remote servers or cloud-based storage.
[0083] One or more computer programs or computer control logic may be stored in main memory 132 and / or auxiliary memory 134. The computer program may also be received via communication interface 146. The computer program may include computer-executable instructions that, when executed by computer processor 122, cause computer system 100 to perform the methods described herein.
[0084] In another embodiment, the software may be stored in a non-transitory computer-readable storage medium and loaded into the main memory 132 and / or secondary memory 134 of the computer system 100 using a removable storage drive 138, a hard disk drive 136, and / or a communication interface 146. When executed by the processor 122, the control logic (software) causes the computer system 100, more generally, for a system used to detect scanning interference, to perform all or some of the methods described herein.
[0085] Based on this specification, implementations of other hardware and software arrangements for performing the functions described herein will be readily apparent to those skilled in the art.
Claims
1. A computer-implemented method for three-dimensional (3D) registration, the method comprising: receiving, by one or more computing devices, individual images of a patient's dentition configured as a depth and a corresponding color image that are mapped together; providing pixels of the individual images as input to a trained deep neural network; automatically identifying a source of a registration error in the individual images using one or more output label values of the trained deep neural network, wherein the output label values are obtained by segmenting the individual images into regions corresponding to one or more object classes; registering the individual images together based on the one or more output label values to form a registered 3D image that is free of registration errors or substantially free of registration errors, wherein the individual images are depth and corresponding color images that are mapped together by mapping a depth image to a corresponding color image or mapping a corresponding color image to a depth image, wherein the method further comprises: generating a point cloud from the depth image by projecting pixels of the depth image into space; in response to the generating, assigning each point in the point cloud a color value and a label value using the corresponding color image and the output label values of the trained deep neural network, respectively; and based on the assigned label values, discarding or including one or more points in the point cloud using a predetermined weight, thereby eliminating or reducing a contribution of the discarded or partially included one or more points to the registration.
2. The method of claim 1, wherein, the individual images are individual three-dimensional optical images.
3. The method of claim 1, wherein, the individual images are received as a time sequence of images.
4. The method of claim 1, wherein, the one or more object classes include hard gingiva, soft tissue gingiva, tongue, cheek, teeth, and dentiform objects.
5. The method of claim 1, wherein a relevance of the identified source of a registration error is based on a geometry of its surroundings.
6. The method of claim 1, wherein, the deep neural network is a network selected from a group consisting of a convolutional neural network (CNN), a fully convolutional neural network (FCN), a recurrent neural network (RNN), and a recurrent convolutional neural network (Recurrent-CNN).
7. The method of claim 1, further comprising: training the deep neural network using the one or more computing devices and a plurality of individual training images to map one or more tissues in at least a portion of each training image to one or more label values, wherein the training is performed at a pixel level by classifying the individual training images, pixels of the individual training images, or superpixels of the individual training images into one or more classes corresponding to semantic data types and / or error data types.
8. The method of claim 7, wherein, the training images include pairs of 3D meshes and registered depth and color images.
9. The method of claim 8, wherein, the 3D meshes are labeled and the labels are transferred to the pairs of 3D and color images using a transformation function.
10. A non-transitory computer-readable storage medium storing a program which, when executed by a computer system, causes the computer system to perform processes comprising: receiving, by one or more computing devices, individual images of a patient's dentition configured as depth and corresponding color images that are mapped together; providing pixels of the individual images as input to a trained deep neural network; automatically identifying a source of registration error in the individual images using one or more output label values of the trained deep neural network, wherein the output label values are obtained by segmenting the individual images into regions corresponding to one or more object classes; wherein the individual images are depth and corresponding color images that are mapped together by mapping depth images to corresponding color images or mapping corresponding color images to depth images; the process includes registering the individual images together based on the one or more output label values to form a registered 3D image that is free of registration errors or substantially free of registration errors; and the process further includes: generating a point cloud from the depth images by projecting pixels of the depth images into space; in response to the generating, assigning each point in the point cloud a color value and a label value using the corresponding color image and the output label values of the trained deep neural network, respectively; and based on the assigned label values, discarding or including one or more points in the point cloud using a predetermined weight, thereby eliminating or reducing the contribution of the discarded or partially included one or more points to the registration.
11. A system for three-dimensional (3D) registration, comprising a processor configured to: receive, by one or more computing devices, individual images of a patient's dentition configured as depth and corresponding color images that are mapped together; provide pixels of the individual images as input to a trained deep neural network; automatically identify a source of registration error in the individual images using one or more output label values of the trained deep neural network, wherein the output label values are obtained by segmenting the individual images into regions corresponding to one or more object classes; wherein the individual images are depth and corresponding color images that are mapped together by mapping depth images to corresponding color images or mapping corresponding color images to depth images; wherein, the processor is configured to register the individual images together based on the one or more output label values to form a registered 3D image that is free of registration errors or substantially free of registration errors, wherein the processor is further configured to: generate a point cloud from the depth images by projecting pixels of the depth images into space; in response to the generating, assign each point in the point cloud a color value and a label value using the corresponding color image and the output label values of the trained deep neural network, respectively; and based on the assigned label values, discard or include one or more points in the point cloud using a predetermined weight, thereby eliminating or reducing the contribution of the discarded or partially included one or more points to the registration.
12. The system of claim 11, wherein, The deep neural network is a network selected from the group consisting of a convolutional neural network (CNN), a fully convolutional neural network (FCN), a recurrent neural network (RNN), and a recurrent-convolutional neural network (Recurrent-CNN). The deep neural network is a network selected from the group consisting of a convolutional neural network (CNN), a fully convolutional neural network (FCN), a recurrent neural network (RNN), and a recurrent-convolutional neural network (Recurrent-CNN). The deep neural network is a network selected from the group consisting
Citation Information
Patent Citations
Dental CAD automation using deep learning
US20180028294A1
Method for analyzing an image of a dental arch
US20190026893A1
Method for providing data associated with the intraoral cavity
US7698068B2
Object classification for measured three-dimensional object scenes
US9436868B2
Method for recording multiple three-dimensional images of a dental object
US9456754B2