Tooth identification method of multi-view tooth picture
Through a two-layer deep neural network architecture, combined with instance segmentation and multi-view integration network, the problem of inaccurate tooth recognition in multi-view dental photos, especially the misidentification of deciduous and permanent teeth, is solved, achieving higher recognition accuracy.
Patent Information
- Application Number
- CN202410406093.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-07
- Publication Date
- 2025-10-21
AI Technical Summary
Existing multi-view tooth photo recognition methods easily identify the same tooth as different tooth numbers from different perspectives of the same patient, especially when identifying deciduous teeth, which is not accurate enough.
Using a two-layer deep neural network architecture, the system first uses an instance segmentation network to identify tooth masks and predict tooth numbers in dental photos. Then, it uses a multi-view integration network based on an attention mechanism to integrate multi-view information, and solves the tooth number assignment problem through the Hungarian algorithm to ensure accurate recognition.
Improved tooth recognition accuracy for multi-view dental photos, corrected misidentification of deciduous and permanent teeth, and achieved higher recognition accuracy.
Smart Images

Figure CN120823418A_ABST
Abstract
Description
Technical Field
[0001] The present application generally relates to a method for tooth recognition from multi-view dental photographs. Background Art
[0002] With the rapid development of computer technology, dental treatment is increasingly relying on computer technology. In digital orthodontic treatment, tooth recognition from multi-view dental photos of patients is a very important step, and digital orthodontic treatment has extremely high requirements for the accuracy of tooth recognition.
[0003] The existing technical solution is to process dental photos taken from different perspectives separately. In some cases, the same tooth in dental photos taken from different perspectives of the same patient may be identified as different tooth numbers, especially for deciduous teeth.
[0004] In view of the above, it is necessary to provide a new tooth recognition method for tooth photos. Summary of the Invention
[0005] One aspect of the present application provides a computer-implemented tooth recognition method for multi-perspective dental photographs, comprising: obtaining first and second dental photographs of a patient, wherein the first dental photograph is a first-perspective dental photograph, which shows the labial and buccal surfaces of multiple teeth in the patient's first dentition, and the second dental photograph is a second-perspective dental photograph, which shows the occlusal surfaces of multiple teeth in the patient's first dentition, and the first dentition is the maxillary or mandibular dentition; using a trained first deep neural network to identify teeth in the first dental photograph, and obtain a first group of masks and a 1-1 group of predicted tooth numbers for the multiple teeth; using the first deep neural network to identify teeth in the second dental photograph, and obtain a second group of masks and a 2-1 group of predicted tooth numbers for the multiple teeth, wherein , the first deep neural network is an instance segmentation network; the first group of tooth features is extracted from the first tooth photo using the first group of masks, and the second group of tooth features is extracted from the second tooth photo using the second group of masks; and the trained second deep neural network is used to predict the 1-2 group of predicted tooth numbers of the multiple teeth in the first tooth photo and the 2-2 group of predicted tooth numbers of the multiple teeth in the second tooth photo based on the 1-1 group of predicted tooth numbers, the first group of tooth features, the 2-1 group of predicted tooth numbers and the second group of tooth features, wherein the second deep neural network is a multi-view integration network, which integrates the information of the first and second tooth photos in the process of predicting the 1-2 group and the 2-2 group of tooth numbers.
[0006] In some embodiments, the second deep neural network is a neural network based on an attention mechanism.
[0007] In some embodiments, the second deep neural network includes at least two fusion modules connected in series.
[0008] In some embodiments, each fusion module includes a first cross attention layer, a second crossattention layer, and a self attention layer, wherein the first cross attention layer updates a fusion token based on the tooth features of all input perspectives, the second cross attention layer updates the tooth features of all perspectives based on the fusion token updated by the first crossattention layer, and the self attention layer further updates the tooth features updated by the second cross attention layer, wherein the fusion token corresponds one-to-one to the tooth numbers of the multiple teeth, and the fusion token is a medium for integrating the information of the first and second tooth photos.
[0009] In some embodiments, the first set of dental features is extracted from the first dental photo by the first deep neural network based on the first set of masks, and the second set of dental features is extracted from the second dental photo by the first deep neural network based on the second set of masks.
[0010] In some embodiments, the method further includes: using the second deep neural network to predict the first group of tooth number probability distributions of the multiple teeth in the first dental photo and the second group of tooth number probability distributions of the multiple teeth in the second dental photo based on the 1-1 group of predicted tooth numbers, the first group of tooth features, the 2-1 group of predicted tooth numbers and the second group of tooth features; and using a method for solving an allocation problem or an optimal transmission problem to respectively calculate the 1-2 group and 2-2 group of predicted tooth numbers based on the first group and the second group of tooth number probability distributions.
[0011] In some embodiments, the predicted tooth numbers of Group 1-2 and Group 2-2 are calculated using the Hungarian algorithm based on the probability distribution of the tooth numbers of the first and second groups, respectively.
[0012] In some embodiments, the training data set for training the second deep neural network is obtained by the following method: obtaining multiple groups of dental photographs, wherein each group of dental photographs includes a patient's first-perspective dental photograph and a second-perspective dental photograph; performing the following operations on each of the dental photographs in each group: using the first deep neural network to identify each tooth in the dental photograph to obtain a group of tooth masks and a group of predicted tooth numbers; using the first deep neural network to extract a group of tooth features from the dental photograph based on the group of tooth masks; and obtaining the annotation of the actual tooth number of each tooth in the dental photograph, wherein the tooth features and predicted tooth numbers corresponding to each group of dental photographs serve as a group of inputs when training the second deep neural network, and the annotation of the actual tooth number corresponding to the group of dental photographs serves as the ground truth corresponding to the group of inputs.
[0013] Another aspect of the present application provides a computer-implemented tooth recognition method for multi-perspective dental photographs, comprising: obtaining dental photographs of a patient from N predetermined perspectives, which show the labial and buccal surfaces and occlusal surfaces of multiple teeth of the patient; using a trained first deep neural network to perform tooth recognition on the dental photographs from the N perspectives, respectively, to obtain corresponding N groups of masks and N groups of predicted tooth numbers, wherein N is a natural number greater than or equal to 2, and the first deep neural network is an instance segmentation network; using the N groups of masks to extract corresponding N groups of tooth features from the dental photographs from the N perspectives; and using a trained second deep neural network to update the N groups of predicted tooth numbers based on the N groups of predicted tooth numbers and the N groups of tooth features, wherein the second deep neural network is a multi-view integration network, which integrates the information of the dental photographs from the N perspectives in the process of updating the N groups of tooth numbers.
[0014] In some embodiments, the second deep neural network is a neural network based on an attention mechanism, which includes at least two fusion modules connected in series, each fusion module including a first cross attention layer, a second crossattention layer and a self attention layer, wherein the first cross attention layer updates a fusion token based on the tooth features of all input perspectives, the second cross attention layer updates the tooth features of all perspectives based on the fusion token updated by the first crossattention layer, and the self attention layer further updates the tooth features of all perspectives updated by the second cross attention layer, wherein the fusion token corresponds one-to-one to the tooth numbers of the multiple teeth, and the fusion token is a medium for integrating the information of the tooth photos of the N perspectives. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The above and other features of the present application will be further described below in conjunction with the accompanying drawings and detailed description thereof. It should be understood that these drawings only illustrate several exemplary embodiments of the present application and should not be considered to limit the scope of protection of the present application. Unless otherwise specified, the drawings are not necessarily to scale, and similar reference numerals represent similar components.
[0016] Figure 1 This is a schematic flow chart of a method for tooth recognition from multi-view dental photos in one embodiment of the present application;
[0017] Figures 2A-2E An example shows the results of the first deep neural network on tooth recognition of a patient's teeth from multiple perspectives;
[0018] Figure 3 Schematically illustrates the input and output of a second deep neural network in one embodiment;
[0019] Figure 4 is a schematic module diagram of the second deep neural network in one embodiment;
[0020] Figure 5 Schematically shows Figure 4 The structure of the subsequent processing module of the second deep neural network shown;
[0021] Figure 6 Schematically shows the structure of M fusion modules in one; and
[0022] Figures 7A-7E The second deep neural network is based on Figures 2A-2EThe teeth recognition results shown are as well as the final teeth recognition results predicted from the teeth features extracted from the patient's teeth photos from multiple viewing angles. DETAILED DESCRIPTION
[0023] The following detailed description refers to the drawings that form a part of this specification. The illustrative embodiments mentioned in the specification and drawings are for illustrative purposes only and are not intended to limit the scope of protection of this application. In light of this application, those skilled in the art will understand that many other embodiments can be adopted and various changes can be made to the described embodiments without departing from the subject matter and scope of protection of this application. It should be understood that the various aspects of the present application described and illustrated herein can be arranged, replaced, combined, separated and designed according to many different configurations, all of which are within the scope of protection of this application.
[0024] One aspect of the present application provides a computer-implemented tooth recognition method for multi-view dental photos, which can accurately recognize each tooth in each view dental photo based on multiple view dental photos.
[0025] Please refer to Figure 1 , which is a schematic flowchart of a computer-implemented tooth recognition method 100 for multi-view tooth photos in one embodiment of the present application.
[0026] In 101, a plurality of dental photos of a patient from different perspectives are obtained.
[0027] The dental photographs are two-dimensional, and each of the dental photographs includes a plurality of maxillary teeth and / or a plurality of mandibular teeth.
[0028] In one embodiment, the patient's teeth photos can be taken by non-dental professionals, such as the patient himself or his family or friends, etc. In this way, the patient does not need to return to the dental professional institution for this purpose, which reduces the burden on the patient.
[0029] Currently, there are devices that can allow non-dental professionals to take dental photos, for example, the dental photography device disclosed in Chinese Patent No. ZL202220617972.5 or Chinese Patent No. ZL202223452359.1.
[0030] In one embodiment, if only multiple teeth of one dentition (maxillary or mandibular) of the patient are to be identified, a photo of the teeth with the labial and buccal surfaces of the teeth visible and a photo of the teeth with the occlusal surfaces of the teeth visible are required.
[0031] In one embodiment, if all teeth of a patient need to be identified, then typically, dental photos of the patient from five viewing angles, namely left, middle, right, top, and bottom, are required to cover the labial and buccal surfaces and occlusal surfaces of all teeth.
[0032] It is understood that the number of dental photos required depends on whether the dental photos taken cover the labial and buccal surfaces and occlusal surfaces of the corresponding teeth. For example, if the frontal dental photos cover the labial and buccal surfaces of all teeth, then only three views of the patient's teeth are needed to identify all the patient's teeth: the middle, upper, and lower views.
[0033] The tooth recognition method for multi-view tooth photos of the present application is described in detail below using tooth photos from five viewpoints as an example.
[0034] In 103 , teeth in the plurality of dental photos are identified using the trained first deep neural network.
[0035] The first deep neural network is an instance segmentation network, for example, it can be a YOLO network or other applicable deep neural networks based on CNN (Convolutional Neural Network) or Transformer.
[0036] Corresponding to an input dental photo, the trained first deep neural network predicts the mask and tooth number of each tooth in the dental photo.
[0037] In one embodiment, only one first deep neural network can be trained to process dental photos from different perspectives. In another embodiment, a first deep neural network can also be trained for each perspective. The method of this application is described in detail below using the former as an example.
[0038] The training dataset used to train the first deep neural network is referred to as the first training dataset, which includes dental photos from various viewing angles labeled with masks and tooth numbers. In a preferred embodiment, the first training dataset includes multiple dental photos corresponding to each viewing angle.
[0039] The dental photos from the five perspectives are input into the first deep neural network one by one to obtain the predicted mask and tooth number of each tooth in these dental photos.
[0040] In one embodiment, each tooth may be numbered using the FDI tooth position representation method, and the numbering method is shown in Table 1 below.
[0041]
[0042]
[0043] Table 1
[0044] In some cases, some patients may have extra teeth, known as supernumerary teeth, in addition to their normal number of teeth. To identify these supernumerary teeth, in one embodiment, in addition to the 52 tooth numbers used in the FDI tooth position representation, four additional tooth numbers are assigned to these supernumerary teeth, for a total of 56 tooth numbers.
[0045] Please refer to Figures 2A-2E , respectively showing the results of the first deep neural network segmenting the dental photos of the five perspectives in an example, and the segmentation results of the dental photos corresponding to the Nth perspective are recorded as the N-1th group of predicted tooth numbers and the N-1th group of predicted masks.
[0046] On the one hand, the number of permanent teeth in the first training data set used to train the first deep neural network is much larger than that of deciduous teeth. On the other hand, to accurately identify a deciduous tooth, images of the deciduous tooth from different perspectives are generally required, namely, images of the lip and cheek side surfaces and images of the occlusal surface, while the first deep neural network only performs tooth recognition based on tooth photos from a single perspective. On yet another hand, to accurately identify a deciduous tooth, it is also necessary to refer to the conditions of other teeth, while the first deep neural network does not refer to the conditions of other teeth when identifying a tooth. Therefore, the first deep neural network is likely to mistakenly detect deciduous teeth as permanent teeth.
[0047] It can be noted that in Figures 2A-2E In the image, all teeth are detected as permanent teeth, but in fact some of these teeth are deciduous teeth. Therefore, the output of the first deep neural network needs to be further processed to correct these false detections.
[0048] In 105 , the tooth number of each tooth in the plurality of dental photos is updated based on the tooth number of each tooth in the plurality of dental photos predicted by the first deep neural network using the trained second deep neural network.
[0049] The second deep neural network is a multi-view aggregation network based on the attention mechanism, which can integrate information from multi-view dental photos to correctly identify permanent teeth and deciduous teeth.
[0050] Please refer to Figure 3 , is a schematic module diagram of using the second deep neural network to update the tooth number of each tooth in the multiple tooth photos in one embodiment.
[0051] The features of the teeth detected in each of the dental photos and the tooth number 201 are input into the second deep neural network 203.
[0052] In the following, the features extracted from the dental photo of the Nth perspective are recorded as the N-1th group of dental features, and the tooth number of each tooth in the dental photo of the Nth perspective identified by the first deep neural network is recorded as the N-1th group of tooth numbers.
[0053] In one embodiment, the first deep neural network can be used as the backbone of the second deep neural network to extract tooth features. In one embodiment, the specific operation of extracting features from one of the multiple tooth photos is as follows.
[0054] First, the dental photo is input into the first deep neural network to extract the feature map after 4 times downsampling.
[0055] Next, for each tooth detected in the dental photograph, the tooth mask is scaled according to the feature map, and the tooth feature is extracted from the feature map using the scaled mask. In one embodiment, the feature values of all pixels belonging to the tooth can be averaged or maximized to serve as the feature of the tooth.
[0056] The second deep neural network 203 outputs a tooth number probability distribution 205 of each tooth detected in the tooth photo. The tooth number probability distribution corresponding to the tooth photo at the Nth viewing angle is recorded as the N-1th group of tooth number probability distribution.
[0057] If the tooth number probability distribution is directly converted into tooth numbers, there is a small probability that duplicate tooth numbers may occur, i.e., two different teeth are assigned the same tooth number. In one embodiment, the Hungarian algorithm 207 can be used to solve this tooth number assignment problem and obtain the final tooth number 209 for each tooth detected in the dental photograph. The final set of tooth numbers corresponding to the dental photograph at the Nth angle is recorded as the N-2th set of tooth numbers.
[0058] In one embodiment, the cost of the Hungarian algorithm can be expressed by the following equation (1):
[0059] cost = 1 - tooth number probability distribution equation (1)
[0060] The tooth number probability distribution of each tooth photo is input into the Hungarian algorithm module to calculate the non-repeated predicted tooth number.
[0061] In light of this application, it can be understood that in addition to the Hungarian algorithm, other methods for solving allocation problems or optimal transmission problems can also be used, for example, the algorithm proposed in "On implementing 2D rectangular assignment algorithms" published in IEEE Transactions on Aerospace and Electronic Systems, 52(4):1679-1696, August 2016.
[0062] Please refer to Figure 4 , which is a schematic module diagram of the second deep neural network 203.
[0063] The tooth features extracted from the tooth photos of all angles are recorded as the first group of tooth features 2011, which can be spliced together by the tooth features extracted from each of the tooth photos, that is, spliced together by the tooth features from group 1-1 to group 5-1.
[0064] All the groups of tooth numbers identified by the first deep neural network in the dental photos are recorded as the first group of predicted tooth numbers 2013, which can be spliced together by the tooth numbers identified in each of the dental photos, that is, from the 1-1 group of tooth numbers to the 5-1 group of tooth numbers.
[0065] The tooth number encoding module 2031 encodes the first set of predicted tooth numbers 2013 using tooth number embedding. It is understood that tooth number encoding is not limited to tooth number embedding, and the first set of predicted tooth numbers 2013 can also be encoded using a neural network with a one-hot vector as input.
[0066] The addition module 2033 adds the first set of tooth features and the first set of encoded predicted tooth numbers to obtain a second set of tooth features, which is then output to M serially connected fusion modules 2035. Simultaneously, a first set of fused tokens is output to the M fusion modules 2035. The first set of fused tokens is obtained after training the second deep neural network 203. Each fused token corresponds to a tooth number, resulting in 56 fused tokens in this embodiment. In one embodiment, M is greater than or equal to 2.
[0067] The M fusion modules 2035 output the third group of tooth features 2051 and the second group of fusion tokens 2019, wherein the third group of tooth features 2051 has a group of tooth features corresponding to each perspective of the tooth photo, and the group of tooth features corresponding to the tooth photo of the Nth perspective is recorded as the N-3th group of tooth features.
[0068] Please refer to Figure 5, schematically showing the second deep neural network 203 in Figure 4 The structure of the subsequent processing module of the structure shown.
[0069] The N-3 group of tooth features 2051 are subjected to L2 norm normalization processing using the L2 norm normalization module 2037, and the second group of fusion tokens 2019 are subjected to L2 norm normalization processing using the L2 norm normalization module 2039. The two are then matrix multiplied using the matrix multiplication module 2041 to obtain a similarity matrix between the teeth in the N-th perspective tooth photo and the 56 tooth numbers. This is then subjected to a layer of softmax processing to obtain the N-1 group of tooth number probability distribution 209.
[0070] In light of this application, it can be understood that in addition to using matrix calculations to calculate the probability distribution of tooth numbers, deep neural networks can also be used directly to predict the probability distribution of tooth numbers.
[0071] Please refer to Figure 6 , schematically showing the structure of the first of the M fusion modules 2035.
[0072] The M fusion modules 2035 are connected in series, that is, the output of one fusion module is the input of the next fusion module.
[0073] In one embodiment, each fusion module includes two cross attention layers and one self-attention layer. In light of this application, it is understood that in addition to cross attention and self attention, various variants of them can also be used. Therefore, in this application, cross attention and self attention include them and their variants.
[0074] The second group of tooth features 2055 and the first group of fused tokens 2017 output by the addition module 2033 are input to the first cross attention layer 2035a of the first fusion module, which updates the first group of fused tokens 2017 to obtain the third group of fused tokens 2021.
[0075] The third set of fused tokens 2021 and the second set of tooth features 2055 are input into the second crossattention layer 2035a, which updates the second set of tooth features 2055 to obtain the fourth set of tooth features 2057. The fused tokens integrate information about teeth from different perspectives, and the second set of tooth features 2055 is updated based on the third set of fused tokens 2021.
[0076] The fourth group of tooth features 2057 is input into a self attention layer 2035d, which generates a fifth group of tooth features 2059 based on the fourth group of tooth features 2057. The fifth group of tooth features 2059 and the third group of fusion tokens 2021 will serve as inputs to the next fusion module.
[0077] It should be noted that, in order to simplify the description, in the above detailed description, the same reference numerals are used for a group of tooth features in all viewing angles and a group of tooth features in one viewing angle.
[0078] In light of this application, it can be understood that the structure of multi-perspective information exchange is not limited to fusion tokens and fusion modules, but can also be other network structures with information interaction functions, such as cross attention between two perspectives.
[0079] Please refer to Figures 7A-7E , showing that the second deep neural network is based on Figures 2A-2E The tooth recognition results shown (i.e., the tooth recognition results predicted by the first deep neural network based on the multiple tooth photos) and the tooth features extracted from the multiple tooth photos, and the final tooth recognition results predicted.
[0080] It can be noted that compared with the tooth recognition results predicted by the first deep neural network based on the multiple tooth photos, the tooth numbers of the maxillary dentition, except for teeth 16, 22, and 26, were corrected to the corresponding deciduous tooth numbers, and the tooth numbers of the mandibular dentition, except for teeth 36, 31, 41, 42, and 46, were corrected to the corresponding deciduous tooth numbers. The second deep neural network accurately distinguished and identified deciduous and permanent teeth.
[0081] The training data set used to train the second deep neural network is recorded as a second training data set. In one embodiment, the second training data set can be obtained by the following method.
[0082] A plurality of sets of dental photographs are obtained, each set including a plurality of dental photographs of a patient at predetermined viewing angles, for example, left, middle, right, top, and bottom viewing angles.
[0083] For each set of dental photos, the first deep neural network is used to identify the teeth in each dental photo to obtain a set of predicted tooth numbers and a set of tooth masks. The first deep neural network is then used as the backbone to extract a set of tooth features from the dental photos based on the set of tooth masks. The predicted tooth numbers and tooth features of the set of dental photos are used as the input for a training session. The tooth numbers correctly labeled in the set of dental photos are used as the ground truth for this training session. In this way, the second training data set is obtained. The trained second deep neural network is obtained by training the second deep neural network using the second training data set.
[0084] In one embodiment, the correctly marked tooth number can be obtained by manual marking.
[0085] Although various aspects and embodiments of the present application are disclosed herein, other aspects and embodiments of the present application will be readily apparent to those skilled in the art in light of this disclosure. The various aspects and embodiments disclosed herein are for illustrative purposes only and are not intended to be limiting. The scope and subject matter of this application are determined solely by the appended claims.
[0086] Similarly, various diagrams may illustrate exemplary architectures or other configurations of the disclosed methods and systems that aid in understanding the features and functionality that may be included in the disclosed methods and systems. The claimed content is not limited to the exemplary architectures or configurations shown, and the desired features may be implemented using a variety of alternative architectures and configurations. Furthermore, for flow charts, functional descriptions, and method claims, the order of blocks presented herein should not limit various embodiments to being implemented in the same order to perform the described functionality, unless the context clearly dictates otherwise.
[0087] Unless otherwise expressly stated, the terms and phrases used herein and their variations should be interpreted as open ended rather than restrictive. In some instances, the appearance of broad words and phrases such as "one or more," "at least," "but not limited to," or other similar terms should not be understood as intending or requiring a narrowing of the context in which such broad terms may not be used.
Claims
1. A computer-implemented method for tooth recognition from multi-view dental photographs, comprising: Obtaining first and second dental photographs of a patient, wherein the first dental photograph is a first-perspective dental photograph showing labial and buccal surfaces of a plurality of teeth in a first dentition of the patient, and the second dental photograph is a second-perspective dental photograph showing occlusal surfaces of a plurality of teeth in a first dentition of the patient, wherein the first dentition is the maxillary or mandibular dentition; Using a trained first deep neural network to identify teeth in the first dental photograph to obtain a first set of masks for the plurality of teeth and a 1-1 set of predicted tooth numbers; using the first deep neural network to identify teeth in the second dental photograph to obtain a second set of masks for the plurality of teeth and a 2-1 set of predicted tooth numbers, wherein the first deep neural network is an instance segmentation network; extracting a first set of tooth features from the first tooth photograph using the first set of masks, and extracting a second set of tooth features from the second tooth photograph using the second set of masks; and A trained second deep neural network is used to predict the 1-1 group of predicted tooth numbers, the first group of tooth features, the 2-1 group of predicted tooth numbers and the second group of tooth features to obtain the 1-2 group of predicted tooth numbers of the multiple teeth in the first tooth photo and the 2-2 group of predicted tooth numbers of the multiple teeth in the second tooth photo, wherein the second deep neural network is a multi-view integration network, which integrates the information of the first and second tooth photos in the process of predicting the 1-2 group and the 2-2 group of tooth numbers.
2. The method according to claim 1, wherein The second deep neural network is a neural network based on the attention mechanism.
3. The method according to claim 2, wherein The second deep neural network includes at least two fusion modules connected in series.
4. The method according to claim 3, wherein Each fusion module includes a first crossattention layer, a second cross attention layer and a self attention layer, wherein the first crossattention layer updates the fusion token based on the tooth features of all input perspectives, the second cross attention layer updates the tooth features of all perspectives based on the fusion token updated by the first cross attention layer, and the self attention layer further updates the tooth features updated by the second cross attention layer, wherein the fusion token corresponds one-to-one to the tooth numbers of the multiple teeth, and the fusion token is a medium for integrating the information of the first and second tooth photos.
5. The method according to claim 1, wherein The first set of tooth features is extracted from the first tooth photo by the first deep neural network based on the first set of masks, and the second set of tooth features is extracted from the second tooth photo by the first deep neural network based on the second set of masks.
6. The method according to claim 1, wherein It also includes: Using the second deep neural network, based on the 1-1 group of predicted tooth numbers, the first group of tooth features, the 2-1 group of predicted tooth numbers, and the second group of tooth features, predict a first group of tooth number probability distributions for the plurality of teeth in the first dental photograph and a second group of tooth number probability distributions for the plurality of teeth in the second dental photograph; and The predicted tooth numbers of group 1-2 and group 2-2 are calculated based on the tooth number probability distributions of the first and second groups respectively using a method for solving the allocation problem or the optimal transmission problem.
7. The method according to claim 6, wherein The predicted tooth numbers of group 1-2 and group 2-2 are calculated using the Hungarian algorithm based on the probability distribution of the tooth numbers of the first group and the second group, respectively.
8. The method according to claim 1, wherein The training dataset for training the second deep neural network is obtained by the following method: Acquire multiple groups of dental photographs, wherein each group of dental photographs includes a dental photograph of a patient taken from the first perspective and a dental photograph taken from the second perspective; For each of the sets of dental photos, the following operations are performed: using the first deep neural network to identify each tooth in the dental photo to obtain a set of tooth masks and a set of predicted tooth numbers; using the first deep neural network to extract a set of tooth features from the dental photo based on the set of tooth masks; and obtaining the actual tooth number of each tooth in the dental photo. Among them, the tooth features and predicted tooth numbers corresponding to each group of dental photos serve as a group of inputs for training the second deep neural network, and the annotations of the actual tooth numbers corresponding to the group of dental photos serve as the GroundTruth corresponding to the group of inputs.
9. A computer-implemented method for tooth recognition from multi-view dental photographs, comprising: Acquire dental photographs of a patient at N predetermined viewing angles, wherein the dental photographs show labial and buccal surfaces and occlusal surfaces of a plurality of teeth of the patient; Using a trained first deep neural network to perform tooth recognition on the tooth photos from the N perspectives, respectively, to obtain corresponding N groups of masks and N groups of predicted tooth numbers, where N is a natural number greater than or equal to 2, and the first deep neural network is an instance segmentation network; Extracting corresponding N groups of tooth features from the tooth photos of the N viewing angles using the N groups of masks respectively; and The N groups of predicted tooth numbers are updated based on the N groups of predicted tooth numbers and the N groups of tooth features using a trained second deep neural network, wherein the second deep neural network is a multi-view integration network, and in the process of updating the N groups of tooth numbers, it integrates the information of the tooth photos from the N perspectives.
10. The method according to claim 9, wherein The second deep neural network is a neural network based on the attention mechanism, which includes at least two fusion modules connected in series, each fusion module including a first crossattention layer, a second cross attention layer and a self attention layer, wherein the first crossattention layer updates the fusion token based on the tooth features of all input perspectives, the second cross attention layer updates the tooth features of all perspectives based on the fusion token updated by the first cross attention layer, and the self attention layer further updates the tooth features of all perspectives updated by the second cross attention layer, wherein the fusion token corresponds one-to-one to the tooth numbers of the multiple teeth, and the fusion token is a medium for integrating the information of the tooth photos of the N perspectives.
Citation Information
Patent Citations
Dental imaging device
CN217218989U
Dental imaging device
CN219126318U