METHOD FOR DETERMINING A TOOTH COLOR
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- IVOCLAR VIVADENT AG
- Filing Date
- 2022-02-16
- Publication Date
- 2026-04-23
AI Technical Summary
Existing methods for determining tooth color in real-world conditions suffer from inaccuracies due to deviations in lighting and angles, which are not reproducible, leading to a loss of laboratory accuracy.
An iterative learning algorithm using a high-quality imaging device captures images of sample teeth under various lighting conditions and angles, creating a database for a CNN to learn and assign corresponding colors, allowing for accurate tooth color determination on-site with a smartphone or professional camera.
The method significantly improves tooth color recognition accuracy by leveraging a comprehensive database of diverse lighting and angle conditions, ensuring precise color determination even in real-world environments.
Description
[0001] The invention relates to a method for determining a tooth color, according to the preamble of claim 1.
[0002] From EP 3 613 382 A1, it is known to use a color selection object that is held next to a tooth whose color is to be determined. A joint image of the tooth and the color selection object, which serves as an auxiliary object, is taken. Since the auxiliary object has a known tooth color, this makes it easier and more accurate to determine the tooth color.
[0003] CNNs trained to recognize tooth color are known from CN 111 652 839 A.
[0004] The aforementioned solution requires, where possible, reproducible lighting conditions. While the presence of the auxiliary body with the known tooth shade allows for the calibration or standardization of the captured tooth color, experience has shown that deviations still occur in practice, meaning that the accuracy of color determination achievable in the laboratory is not maintained in real-world situations.
[0005] In contrast, the invention is based on the objective of creating a method for determining a tooth color according to the preamble of claim 1 that ensures accurate color determinations in practice.
[0006] This problem is solved according to the invention by claim 1. Advantageous embodiments are described in the dependent claims.
[0007] According to the invention, an evaluation device is provided to first capture and evaluate images of sample teeth under different lighting conditions and shooting angles in an iterative learning algorithm and to learn an assignment of the respective captured images to the known applicable sample tooth color.
[0008] In particular, the acquisition and learning at different recording angles is important according to the invention and contributes significantly to improving the recognition capability of the evaluation device.
[0009] A special imaging device is used to capture images of sample teeth. This device can be of particularly high quality. For example, a professional SLR camera can be used to read out the raw data. The quality of the raw data is typically better than that of the camera's data converted to a standard format like JPG.
[0010] However, it is also possible to use a camera from an end device as the initial recording device, for example, that of a smartphone.
[0011] A recording device, e.g., another recording device, is used in an evaluation step. In this evaluation step, an image of a tooth whose color is to be determined is recorded together with the auxiliary body of the known color.
[0012] The evaluation device accesses the learned images and the image data captured with them.
[0013] The evaluation device preferably has a first part, which is used in the preparatory step, and a second part, which is then used in practice for the evaluation device.
[0014] The second part accesses the same data as the first part.
[0015] This makes it possible to use the evaluation device on site, for example in a dental practice, without having to forgo the data obtained in preparation.
[0016] To enable easy access to the data and insights gained, it is preferable that these are stored in a cloud or at least in an area that is both protected and practically accessible.
[0017] In practice, the second part of the evaluation device is then used, which accesses the same data as the first part.
[0018] This data is therefore always available to the evaluation device. However, this does not mean that full access to the data is required at all times in the method according to the invention.
[0019] It is preferable, however, that the evaluation device in the second part has a memory whose contents are periodically synchronized with the cloud.
[0020] What the second part of the evaluation devices then does is evaluate the images captured by the - further - recording device and assign them to sample tooth colors, based on the learned assignment to sample tooth colors, and then output the tooth color according to a common tooth key based on that again.
[0021] The sample tooth color does not need to exist in reality and be stored; a virtual creation of the sample tooth color, i.e., its definition in a specified color space, is sufficient. It is also possible to create a tooth color purely numerically in a virtual color space, preferably the RGB space, and use this as a reference. Such a color is referred to here as an RGB tooth color, whereby it is understood that this also includes colors created in other virtual tooth spaces.
[0022] There are several ways to create such a virtual tooth color: 1. A scan of a tooth is used to determine the corresponding RGB values; 2. The values of existing tooth libraries are used; or 3. Color measuring devices such as photospectrometers are used to define the colors numerically.
[0023] Surprisingly, the large number of recording situations in the preparatory step results in a significantly better recognition and determination of the actual color of the tooth to be determined.
[0024] Shooting situations include those where different light sources and brightness levels are used.
[0025] For example, the same sample teeth can be photographed using the light spectrum of a halogen lamp, an LED lamp, and daylight under sunny skies, as well as under cloudy skies. This can be done at three different brightness levels and at 5 to 15 different shooting angles in both the vertical and horizontal directions.
[0026] These preparations result in a profound database of, for example, 100 to 300 different recording situations for each sample tooth.
[0027] It is understood that the foregoing explanation is merely exemplary and, in particular, the invention is not limited to the number of recording situations in the preparatory step.
[0028] In an advantageous embodiment of the method according to the invention, it is provided that, in the preparatory step of the iteration, it is checked to what extent the result changes in each iteration. If the change in the result is less than a predetermined threshold, it is assumed that the desired accuracy has been achieved.
[0029] This design can also be modified so that the iteration is only terminated after the change threshold has been undercut several times.
[0030] Preferably, after completion of the preparatory step, all data obtained, which includes the mapping between the result of the preparatory step and the sample tooth colors, is transferred to a cloud where it can be accessed when necessary.
[0031] Before an evaluation step is performed by an end device, it carries out a data comparison to ensure that the determined data is stored locally on the end device to the extent necessary and, in particular, completely.
[0032] This data is regularly synchronized, so that changes are regularly transmitted to the end devices.
[0033] Therefore, when the terminal device is to perform the evaluation step, the current data of the evaluation device is always available, insofar as it was provided and made available in the preparatory step.
[0034] The first part of the evaluation device operates exclusively in the preparatory step, but remains available if further lighting conditions or shooting angles need to be recorded.
[0035] In cases where the relevant lighting conditions and shooting angles have been recorded, but further adjustments may be necessary, the evaluation device can also be transferred to the cloud as an executable or compileable program.
[0036] This solution has the advantage that a dentist can also use the evaluation device in its first part himself if necessary, provided he has the required equipment and can take his own special lighting conditions into account, provided that he has sample tooth colors available which he needs to carry out the first step.
[0037] The data he has collected, which is new with regard to the special lighting conditions, can then be made available to other dentists in the cloud if desired.
[0038] A similar approach is also possible if regional adjustment is desired: The light spectrum of daylight differs geographically between equatorial and polar regions, as the absorption bands of the atmosphere are significantly more pronounced in polar regions.
[0039] If an average daylight value is used as a basis in the first step, it is possible that a dentist in an equatorial region may conclude, based on the data supplied by the evaluation device, that the set daylight data requires correction.
[0040] He can then make the regionally adapted data available to other dentists in his area, for example in the cloud.
[0041] This is just one example of the methods preferred according to the invention for enabling data exchange of the provided data in the cloud. It is understood that data exchange for other reasons is also included according to the invention.
[0042] For the actual evaluation in the evaluation device, it is advantageous if reference points are attached to the auxiliary body. The reference points are selected so that they can be recognized by the evaluation device. If, for example, three or four reference points are provided, their arrangement in the captured image allows the angle at which the image was taken to be deduced.
[0043] According to the invention, it is advantageous if a conclusion or comparison with the data assigned to the relevant recording angle is drawn from this angle measurement.
[0044] In a further advantageous embodiment, the recorded image is divided into segments.
[0045] Segmentation has the advantage that areas can be hidden whose data recorded in the evaluation step suggest that reflections are present.
[0046] Surprisingly, this measure significantly increases the precision of the evaluations, especially in bright and therefore reflective environments.
[0047] Various parameters can be used for the evaluation performed by the evaluation device: For example, it is possible to normalize the data to the ambient brightness using an ambient light sensor. Typically, devices such as smartphones have an ambient light sensor that usually adjusts the display brightness.
[0048] In an advantageous embodiment of the invention, this ambient light sensor is used to perform the aforementioned normalization.
[0049] High-quality smartphones can also distinguish spectrally between artificial light and daylight. This distinction is also made via the built-in ambient light sensor. According to the invention, this differentiation result can also be preferably utilized by not providing the evaluation device with the irrelevant data from the outset.
[0050] Further advantages, details and features will become apparent from the following description of an embodiment of the method according to the invention with reference to the drawing.
[0051] They show: Fig. 1 a schematic flowchart of the preparatory step according to the invention; Fig. 2 a schematic flowchart of the image pipeline as a subroutine, which is used in both the preparatory step and the evaluation step; Fig. 3 a schematic flowchart of the evaluation step; Fig. 4 a learning algorithm according to the invention as a function with input and output; Fig. 5 a schematic representation of a CNN convolution layer in the algorithm; and Fig. 6 a schematic representation of a leg game for Max Pooling.
[0052] In Fig. 1 The preparatory step 12 of the inventive method for determining a tooth color is shown schematically. Preparatory step 12 is referred to here as training 10 and begins according to the flowchart of Fig. 1 with the start of training on the 10th.
[0053] First, images, which can also be referred to as "recordings", are taken in step 12. The recordings represent sample tooth colors and the corresponding teeth, designed as a pattern, are recorded, captured and evaluated under different lighting conditions and recording angles.
[0054] The recorded data is fed to an image pipeline 14 as a subprogram, the design of which consists of Fig. 2 as is evident.
[0055] After processing in the image pipeline 14, the prepared data are fed to a training algorithm 16. This performs the actual training, i.e., the optimized reproduction of the sample tooth colors under different lighting conditions and recording angles of the data.
[0056] Following this, step 18 checks whether the training was sufficient. If this was not the case, meaning greater accuracy is required, the process returns to block 12 and the acquired data goes through image pipeline 14 again.
[0057] If, however, the training is deemed sufficient, the "trained" data is saved in the cloud in step 20. This completes the training at block 22.
[0058] Out of Fig. 2 The individual steps of image pipeline 14 are shown. Image pipeline 14 is started in step 26. Image data 28 is available, which is stored in the image memory in step 28.
[0059] Image 30 is now available and its resolution and format are checked in step 32. If the resolution and format are insufficient, path 34 is followed; if both format and resolution are sufficient, execution continues with path 36.
[0060] Image 30 is then checked in path 36, in step 38, to determine the reference points of the reference object; these are recorded. Again, it is possible, in path 40, that no or insufficient reference points could be determined.
[0061] If, however, reference points were found, path 42 is executed further and the relevant information is extracted from image 30 in block 44. According to block 46, this includes color values extracted from the reference object.
[0062] The system checks whether a tooth segment exists in image 30. If so, path 48 is followed.
[0063] In parallel, color values are processed via path 50. These color values are fed to the algorithm in block 52, which then generates the color information for image 30 in path 54. The data is then transferred to the calling program in block 56, marking the end of the image pipeline subroutine.
[0064] According to path 48, the tooth segment data is processed further. Block 58 checks for reflections. If these exceed a threshold, path 60 is followed, which, like paths 34 and 40, ends with no result according to block 62.
[0065] However, if, according to path 64, the reflections are below a setpoint, dominant tooth colors are calculated using k-means clustering. This occurs in block 66.
[0066] This results in color values in path 68, which are then fed to algorithm 52.
[0067] Out of Fig. 3 The evaluation step is visible. This step is intended to be executed on the end device, for example, a smartphone. The first block, 70, represents the cloud, and the second block, 72, represents the smartphone.
[0068] In an advantageous configuration, the data of the image captured by the end user is forwarded to the cloud in block 74, and it enters the image pipeline in block 76 ( Fig. 2 This image pipeline is the same as image pipeline 14 in Fig. 1 .
[0069] After processing and evaluating the data, step 78 checks whether enough images are available. If so, path 80 is followed, and a classification algorithm is performed in block 82. The output of block 82 is thus classified colors at block 84. These are then routed via block 86 to the smartphone 72.
[0070] The smartphone receives the data in block 88 and the color classification is completed in block 90.
[0071] If, however, step 78 detects that there are not enough images, path 92 is followed. In this case, the color classification is started in block 94 by the smartphone, and a photo is taken at step 96.
[0072] The recording of the image or picture is thus triggered or initiated via path 92.
[0073] On the output side of the smartphone 72, an image is located in path 98. In block 100, this image is routed to the cloud via path 102 and loaded there, so that the execution in block 74 can start with this image.
[0074] The following is an example of a learning algorithm: Once the algorithm is fully trained, it can, as shown in the following, Fig. 4 It is evident that at the highest level, this can be seen as a function that assigns each input image a natural number (including 0) from 0 to N (number of classes). The output numbers represent the different classifications; thus, the upper limit of the numbers depends on the use case or the number of different objects to be classified. In the embodiment according to the invention, these are the 16 different colors of the tooth shade guide. (A1-D4)
[0075] Fig. 4 The algorithm is shown as a function with input and output. This algorithm belongs to a variation of neural networks called convolutional neural networks (CNNs). CNNs are neural networks primarily used to classify images (i.e., name what you see), group images by similarity (photo search), and recognize objects in scenes. For example, CNNs are used to identify faces, people, street signs, tumors, animals, and many other aspects of visual data.
[0076] Tests have shown that CNNs are particularly effective at image recognition and enable deep learning. The well-known deep convolutional architecture AlexNet (ImageNet competition 2012) can be used; at the time, applications for this architecture were discussed in the areas of self-driving cars, robotics, drones, security, and medical diagnostics.
[0077] The CNNs according to the invention, as used, do not perceive images like humans do. Rather, it is crucial how an image is fed to a CNN and processed by it.
[0078] CNNs perceive images as volumes, i.e., as three-dimensional objects, rather than as a flat canvas measured only by width and height. This is because digital color images have a red-blue-green (RGB) coding, where these three colors are mixed to create the color spectrum perceived by humans. A CNN interprets such images as three separate, stacked layers of color.
[0079] A CNN receives a standard color image as a rectangular box, whose width and height are measured by the number of pixels in those dimensions, and whose depth comprises three layers, one for each letter in RGB. These depth layers are called channels.
[0080] These numbers are the initial, raw, sensory features that are fed into the CNN, and the purpose of the CNN is to find out which of these numbers are significant signals that help it to classify images more accurately into certain classes.
[0081] CNNs consist roughly of three different layers through which input images are propagated sequentially via mathematical operations. The number, properties, and arrangement of the layers can be modified depending on the application to optimize the results. Fig. 5 This shows a possible architecture of a CNN. The following sections describe the different layers in more detail.
[0082] Fig 5 This shows a schematic representation of a CNN convolution layer. Instead of focusing on one pixel at a time, a CNN takes square ranges of pixels and processes them through a filter. This filter is also a square matrix, smaller than the image itself and the same size as the array. It is also called a kernel, and the filter's function is to find patterns in the pixels. This process is called convolution.
[0083] An example of a possible visualization of the folding process can be seen at the following link: https: / / cs231n.github.io / assets / conv-demo / index.html
[0084] The next layer in a CNN has three names: Max-Pooling, Downsampling, and Subsampling. Fig. 6 This shows an example of a max pooling layer. The inputs from the previous layer are fed into a downsampling layer, and as with convolutions, this method is applied patch by patch. In this case, max pooling simply takes the largest value from a field of an image (see Fig. 6 ), inserted into a new matrix alongside the maximum values from other fields, and the rest of the information contained in the activation cards was discarded.
[0085] This step results in the loss of a lot of information about lower values, which has spurred research into alternative methods. However, downsampling has the advantage of reducing storage and processing requirements precisely because information is lost. Dense Layer
[0086] Dense layers are "traditional" layers also used in classical neural networks. They consist of a variable number of neurons. Neurons in such layers have complete connections to all outputs in the previous layer, just like in normal neural networks. Their outputs can therefore be computed using matrix multiplication followed by a bias offset. For further explanation, please refer to the following link: www.a dventuresinmachinelearning.com / wp-content / uploads / 2020 / 02 / A-beginners-introduction-to-neural-networks-V3.pdf. Learning process
[0087] The learning process of CNNs is largely identical to that of classical neural networks. Therefore, please refer to the following link for full details: https: / / futurism.com / how-do-artificial-neural-networks-learn
[0088] While CNNs have so far been used exclusively for the detection of objects in images, according to the invention this architecture is used to classify the colors of objects with very high accuracy.
Claims
1. A method for determining a tooth colour, wherein - an evaluation device has an iterative learning algorithm using CNNs, which, in a preparatory step, based on at least one previously known sample tooth colour, possibly generated virtually in the RGB space, acquires and evaluates images under different lighting conditions and capture angles and learns to assign them to the applicable sample tooth colour, - wherein, in an evaluation step, - a capture device, in particular a camera of a terminal device, such as a smartphone, or a scanner is provided, with which an image of an auxiliary body with a previously known, possibly virtually generated, sample tooth colour is captured together with at least one tooth, - the capture device acquires at least two images of the combination of the tooth to be determined and the auxiliary body from different capture angles and feeds them to the evaluation device, and - the evaluation device, based on the learned assignment to the applicable, possibly virtually generated, sample tooth colour, evaluates the captured images and outputs the tooth colour of the tooth to be determined according to a reference value, such as a common tooth key, e.g. A1, B2, etc.
2. The method according to claim 1, characterized in that, in the preparatory step, the iteration of the learning algorithm is terminated when the accuracy of the assignment to the applicable sample tooth colour exceeds a threshold value, or the change of the data provided in the iteration with respect to the data of the previous iteration falls below a threshold value.
3. The method according to claim 1 or 2, characterized in that, at the end of the preparatory step, the evaluation device or its data is or are transferred to a cloud, and, in the evaluation step, the transfer of the at least two images to the cloud occurs.
4. The method according to one of the preceding claims, characterized in that the terminal device outputs the tooth colour of the tooth to be determined.
5. The method according to one of the preceding claims, characterized in that the auxiliary body has reference points and, in the evaluation step, the evaluation starts only if reference points are detected in the image to be evaluated and, in particular, requests another image in the case of an absence of reference points.
6. The method according to one of the preceding claims, characterized in that, in the evaluation step, the evaluation device divides the image to be evaluated into segments, in particular after detecting reference points.
7. The method according to claim 6, characterized in that, in the evaluation step, the evaluation device determines a segment from among the segments to be a tooth segment if the segment has a surface essentially in the shape of a tooth and with tooth-like minor differences in colour and brightness.
8. The method according to claim 6 or 7, characterized in that, in the evaluation step, the evaluation device determines a segment from among the segments to be an auxiliary body segment.
9. The method according to claim 6, 7 or 8, characterized in that, in the evaluation step, the evaluation device searches for reflections in a segment, in particular the tooth segment and / or the auxiliary body segment, and only continues the evaluation if only reflections below a predetermined threshold value are detected, and in particular requests another image if the reflections exceed the predetermined threshold value.
10. The method according to one of the preceding claims, characterized in that, in the evaluation step, the evaluation device determines dominant colours of a segment, in particular of the tooth segment, and carries out the evaluation based on the assignment learned, therefore in particular taking into account different lighting conditions, according to a method of the smallest colour distance.
11. The method according to one of the preceding claims, characterized in that, in the evaluation step, the evaluation device also carries out the evaluation based on a comparison of the colour and brightness values of the auxiliary body segment or segments and of the tooth segment.
12. The method according to one of the preceding claims, characterized in that the capture device, such as the camera or the terminal device, has an ambient light sensor, the output signal of which is fed to the evaluation device.
13. The method according to one of the preceding claims, characterized in that the evaluation device, after a predetermined number of evaluations, carries out the preparatory step again, taking into account the evaluations that have been carried out.
14. The method according to one of the preceding claims, characterized in that the evaluation device has a geometry detection unit which emits a warning signal based on an alignment or non-parallelism of segment boundaries, and in that the evaluation device displays the warning signal as an indication of a desired alignment of the terminal device to be changed on the screen thereof.
15. The method according to one of the preceding claims, characterized in that known tooth shapes occurring in practice for defining the segments and / or for improving the geometry detection are stored in the evaluation device and are compared with the detected shapes or segment boundaries.
16. The method according to one of the preceding claims, characterized in that the capture device is designed as a camera and the evaluation device comprises a smartphone app, said app performing at least part of the evaluation via the evaluation device.
17. The method according to one of the preceding claims, characterized in that the capture device is designed as a scanner, where said scanner integrates a computer, in particular a minicomputer such as a Raspberry Pi, which performs at least part of the evaluation via the evaluation device.
18. The method according to one of the preceding claims, wherein the mentioned sample tooth colours are each designed as RGB tooth colours.