Image enhancement using generative machine learning

By replacing the interfered image part with enhanced models and generative machine learning models in dental image processing, the false positive and false negative diagnosis problems of oral health assessment in the prior art are solved, and more efficient and accurate image processing and diagnosis are achieved.

CN120202487APending Publication Date: 2025-06-24KONINKLIJKE PHILIPS NV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202380078010.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-10
Filing Date
2023-11-01
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In the prior art, the results and efficiency of computerized dental feature recognition methods have not yet reached the optimal level when evaluating oral health, which can easily lead to false positive and false negative diagnosis, affecting patients, society and economy.

Method used

Using an enhanced model, by acquiring digital images and using a trained generative machine learning model, augmented digital images are generated, replacing the interfered image parts, especially in dental image processing, interference removal and image enhancement are achieved using machine learning technology.

Benefits of technology

It improves image clarity and diagnostic accuracy, reduces the risk of false positive and false negative diagnosis, reduces the workload of medical practitioners, and improves the efficiency of diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120202487A_ABST
    Figure CN120202487A_ABST
Patent Text Reader

Abstract

The invention improves image processing, in particular medical and / or dental image processing, by means of image enhancement techniques, while minimizing misdiagnosis risks. The method comprises the steps of: acquiring a digital image of at least a portion of the body of the person, in particular the oral cavity, wherein the digital image comprises at least one disturbed image portion; and generating an enhanced digital image based at least in part on the acquired digital image and using a generative machine learning model, wherein at least one disturbed image portion has been replaced by the artificially created image portion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to the field of digital image processing, and in particular to the computerized processing and analysis of medical images such as dental and oral scan images. Background Art

[0002] Digitalization is everywhere in today's world and has fundamentally changed the way we live and communicate. One area of application for digitalization is personal care or health devices. For example, users can use a smart toothbrush equipped with a camera that generates images of their oral cavity while the user is brushing their teeth. Companion apps on the smart toothbrush or the user's smart phone can indicate whether the brushing motion was successful or which teeth require more care, and these indications can be provided in real time during or after the brushing motion. Another example is a smart razor.

[0003] The collected image data can also provide additional advantages in the medical, nursing, or personal health fields. For example, the images can be sent to medical practitioners for remote diagnosis. However, when the images are generated by a smart toothbrush, for example, the images may include disturbed image portions caused by vibrations, the presence of toothpaste or saliva, and / or the presence of the toothbrush or its bristles. Remote diagnosis based on distorted images can impose a significant additional workload on medical practitioners and may even lead to inaccurate diagnoses, or in the worst case, even incorrect diagnoses.

[0004] Efforts have been made in the prior art to alleviate these problems. For example, WO20211 / 75713A1 discloses a computerized tooth feature recognition method for evaluating oral health. The disclosure specifically includes a computer program and a computer-based system for remotely evaluating the oral health status of a person by obtaining at least one digital image of the oral cavity of the person and additional non-image data containing the person's past information. The digital image is segmented using a statistical image segmentation algorithm to extract visible segments, and these visible image segments are further processed to predict any invisible segments.

[0005] However, the results and efficiency of prior art methods have not reached an optimal level. Therefore, further improvements and refinements are needed. Especially when evaluating oral health, the health diagnosis (performed manually and / or by computer) may be accidentally changed during the process. This ultimately leads to error-prone situations and may result in false positive and false negative diagnoses, which have an adverse impact on patients, society, and the economy. These effects are either not recognized, resulting in continued damage, or alternatively, they lead medical practitioners to reject any computer-aided image processing that would alter the original image, thus preventing the relevant patients from enjoying the associated benefits of modern imaging technology.

[0006] Therefore, the fundamental problem of the present invention is to further improve the computerized processing of digital images, thereby at least partially overcoming the above-mentioned disadvantages of the prior art.

[0007] US2022 / 012815 A1 describes a dental procedure where one or more dental images and documents are processed to extract data and label and / or measure dental anatomy or pathology.

[0008] US2021 / 059796 A1 describes how to label and / or correct the marginal lines and / or other features of a dental position. In one example, a three-dimensional model of the dental position is generated. Summary of the Invention

[0009] Currently, a solution to this problem has been designed for the subject matter of the independent claims. Therefore, an image processing method as defined in claim 1 is provided.

[0010] The method may include acquiring a digital image of at least a part of a human body. The digital image may include at least one disturbed image part, where the at least one disturbed image part is caused by vibration, and / or the presence of medical, care, or body substances such as water, liquid, toothpaste, or saliva, and / or the presence of medical, care, or personal health devices such as a toothbrush. The method may include: generating an enhanced digital image of at least a part of the human body at least partially based on the acquired digital image and using a trained generative machine learning model called an enhancement model, in which at least one disturbed image part has been replaced by an artificially created image part; where the part of the human body is the oral cavity of the human; and where the digital image is acquired by a toothbrush equipped with a camera.

[0011] The artificially created image part (preferably, each artificially created image part) may show a healthy part of the human body. The artificially created image part (preferably, each artificially created image part) may show a part of the human body without substantial health problems. The enhancement model may have been trained such that it only adds artificially created image parts without substantial health problems.

[0012] A non-limiting application of the method is the medical field, especially the field of dental care. In such a scenario, the method may be a dental image processing method. The part of the human body may be the oral cavity of the human. At least one disturbed image part may be caused by vibration, and / or the presence of medical, care, or body substances such as water, liquid, toothpaste, or saliva, and / or the presence of medical, care, or personal health devices such as a toothbrush.

[0013] The disturbed image portion can be of several different types. By way of example, they may include the result of vibrations that occur when taking a digital image. In the case of oral and dental health applications, they may include, for example, toothpaste or saliva, as well as other liquids or substances that a dentist may use or employ during patient examination and treatment and that may thus also appear during an oral scan. In other words, the disturbed image portion can be understood as a distorted image portion (caused by vibrations) and / or a blurred image portion (containing toothpaste or saliva, etc.). In some embodiments, the disturbed image portion can include a distorted image portion caused by vibrations that occur when capturing a digital image. Machine learning techniques (such as a trained machine learning model) can be used to achieve interference removal and replace (enhance) it with an artificially generated image portion. Such models have proven to be particularly suitable for achieving viable results. Enhancement of the image can facilitate and can improve subsequent diagnostic steps, whether these steps involve manual diagnosis and visual perception by a medical practitioner or whether alternatively or additionally they involve further automated data processing and analysis steps. Using the enhanced image for diagnosis, a medical practitioner can save more effort because the image is clearer and easier to understand.

[0014] If image enhancement is applied to the healthy body parts depicted in the image, especially if only the healthy parts are enhanced while keeping any unhealthy parts unchanged, then the risk of the diagnosis being altered is reduced. Conversely, the diagnosis may be altered, which would be detrimental to the quality and accuracy of the diagnosis. For example, there are no gum health problems, tooth health problems, or other substantial health problems in the healthy parts. As a result, with the present invention, the continuous diagnostic process is facilitated by the enhancement process. At the same time, the risk of adding artificial health problems that do not exist in the original digital image to the image and the risk of accidentally removing health problems that exist in the original digital image during the enhancement process are reduced. Therefore, the present invention has significant advantages over the prior art. For example, studies have shown that clear images require less manual processing time for classification and understanding.

[0015] It also helps to indicate hidden areas and / or project predictions onto the correct areas. For example, specific colors can be applied to these areas. In extreme cases, the health problems predicted for these areas can be added to the enhanced image.

[0016] To correctly train an enhancement model, a method for training a generative machine learning model for generating enhanced digital images of at least a part of a human body is provided. The method may include providing a training data set and / or training a generative machine learning model using the training data set. The training data set may include a plurality of pairs of digital images. Each pair may include a first digital image of at least a part of a human body. The first digital image may include at least one disturbed image portion, where at least one disturbed image portion is caused by vibration and / or the presence of a medical, care, or body substance such as water, liquid, toothpaste, or saliva, and / or the presence of a medical, care, or personal health device such as a toothbrush. Each pair may include a second digital image of at least a part of a human body, in which at least one disturbed image portion is missing or attenuated. The part of the human body is the oral cavity of the human; and, in this training method, the first digital image is acquired by a toothbrush equipped with a camera.

[0017] At least one disturbed image portion in the first digital image of a given pair may correspond to an image portion in the second digital image of the pair that shows a part (especially the oral cavity) of the human body with no substantial health problems (i.e., health problems that do not affect diagnostic or triage results).

[0018] "Attenuated" may also be understood as "reduced". In other words, the disturbed image portion is either missing in the second digital image or its pixel values are presented less in the second digital image (i.e., the intensity and / or amount are reduced).

[0019] The absence of substantial health problems in the relevant part of the training data can train the enhancement model in a desired way, that is, the trained model can avoid any unexpected changes to the image content related to diagnosis. In this way, the risks of false negatives and false positives can be significantly reduced when using the trained model for continuous (oral) health assessment.

[0020] For diagnosis or further processing, masked digital images may also be used and / or displayed, which allows for immediate identification of the parts that have been modified by the enhancement model. The masked digital images can be used to verify the enhancement process by adding additional checks to prevent errors. On the one hand, the method may include replacing only the disturbed image portions consisting of healthy parts, without replacing the disturbed image portions including unhealthy parts.

[0021] On the one hand, the healthy parts may not contain gum health problems, tooth health problems, or other substantial health problems.

[0022] On the one hand, the method can include using a segmentation machine learning model (referred to as a segmentation model) to segment the acquired digital image to produce an indication of at least one undisturbed image portion in the acquired digital image. The method can include comparing at least one undisturbed image portion in the acquired digital image with at least one corresponding portion in the generated enhanced digital image to produce a similarity score. As proposed herein, segmenting the original digital image can also calculate a similarity score, which allows verification and validation of the proposed method, particularly the above-mentioned enhancement method.

[0023] The data sets useful for various aspects of the present invention can be advantageously enhanced and / or extended, particularly by synthetic data generated artificially. As an important example, interference can be added to clean regions to artificially interfere with a certain segment or portion of the image. This data is very suitable for use as training data. In other words, the training data set can include at least one synthetic digital image in which at least one artificially disturbed image portion has been added to a digital image of at least one body part of a person (particularly the oral cavity). This method is very advantageous because it is difficult to obtain real data that includes both disturbed images and corresponding clean image versions in practice. By using the enhanced data set of the proposed synthetic data, the workload can be reduced. A machine learning model can be employed to generate the proposed synthetic data. Optionally, at least one tooth or gum problem (such as tartar, dental caries, dental plaque, and / or gum recession) can be artificially added to at least one synthetic digital image. This improves the stability of the model that is ultimately trained using the artificially generated data. Preferably, the added health problem intervals and the added interference intervals should not overlap. This avoids incorrect training of the receiving machine learning model.

[0024] According to various aspects of the present invention, the images can be used and / or displayed in various ways. For example, the acquired digital image and the generated enhanced digital image can be displayed on an electronic display device. In particular, if the similarity score reaches or exceeds a predetermined similarity threshold, the display of the enhanced image can depend on certain conditions. In this case, if the enhancement algorithm is in a state prone to errors, the practitioner is prevented from being faced with a potentially misleading enhanced image.

[0025] To further assist the diagnosing practitioner in understanding the enhancement process in depth, a masked digital image indicating at least one undisturbed image portion in the acquired digital image, that is, the image portion modified by the enhancement model, can also be displayed on the electronic display device.

[0026] Powerful security measures can be implemented to avoid misdiagnosis. For example, if the similarity score is below a predetermined similarity threshold, the acquired digital image (excluding the generated enhanced digital image) can be displayed on an electronic display device, optionally with a message that the digital image cannot be enhanced.

[0027] For image enhancement, exemplary enhancement models can include or can be convolutional neural networks (CNNs), deep CNNs, generative adversarial networks (GANs), and / or StyleGANs. Segmentation models can include or can be CNNs, deep CNNs, image segmentation models, and / or U-Nets. A suitable similarity score can be generated by a comparison model that includes or can be a CNN, deep CNN, siamese neural network, and / or deep siamese neural network. However, these machine learning implementations are merely exemplary, and any suitable machine learning model, algorithm, or technique can be adapted and adopted by a skilled practitioner according to the task at hand.

[0028] In particular, a validation dataset different from the training dataset can be used to validate and / or verify the performance of the machine learning model. This increases the value of the verification itself and can serve as a safeguard against erroneously verifying error-prone models.

[0029] The present invention also provides a computer program or a computer-readable medium storing a computer program. The computer program can include instructions that, when the program is executed in a computer and / or a computer network, cause the computer and / or the computer network to perform any method disclosed herein. The computer program can run on a computer, in a personal health device, in edge computing, or in the cloud.

[0030] The present invention also provides a machine learning model data structure. The data structure can embody a machine learning model, particularly a generative machine learning model, for generating an enhanced digital image of at least a part of a human body. The model can be trained using any method disclosed herein.

[0031] The present invention also provides a training, validation, verification, or test dataset for a machine learning model (particularly a generative machine learning model). The dataset can include a collection of any of the above digital images, such as the above first digital image and / or the above second digital image and / or the above synthetic digital image and / or the above generated enhanced digital image and / or the above masked digital image. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The present disclosure can be better understood with reference to the following drawings:

[0033] Figure 1 : A method for processing dental images according to an embodiment of the present invention.

[0034] Figures 2A - 2B : An exemplary digital image according to an embodiment of the present invention having a disturbed image portion caused by toothpaste / saliva ( Figure 2A ) and vibration ( Figure 2B ).

[0035] Figure 3 : An exemplary image pair that can be used as training data according to an embodiment of the present invention.

[0036] Figure 4 : An exemplary image pair having artificially generated data samples according to an embodiment of the present invention.

[0037] Figure 5 : A schematic overview of a model, training, and inference architecture according to an embodiment of the present invention.

[0038] Figure 6 : An exemplary graphical user interface of a smartphone application according to an embodiment of the present invention. Detailed Description

[0039] Preferred embodiments of the dental image processing method 100 will now be described with reference to Figure 1 . Although the embodiments are explained in the context of dental images (i.e., images showing at least a part of a human oral cavity), the basic concepts of the present invention are not limited to the dental field, but can also be applied to other types of medical images or general images where remote diagnosis and / or triage such as skin diagnosis, nail diagnosis are desired, images taken by users of home medical testers (such as COVID-19 tests or pregnancy tests), or remote injury diagnosis based on injury images taken by users.

[0040] The input to method 100 is a digital medical image 102. In the illustrated embodiment, the digital image 102 depicts a part of a human oral cavity. The digital image can be acquired by a camera device, for example, when a person brushes their teeth with a camera-equipped toothbrush or in a professional environment (e.g., during an intraoral scanning process).

[0041] The acquired digital image 102 includes interferences that are present in the disturbed image portion. FIG. 2a shows a magnified view of image 102, from which it can be seen that a part of a person's teeth is covered with a toothpaste and saliva mixture. FIG. 2b shows another example of digital image 102, in which there are interferences caused by vibration (e.g., by an electric toothbrush).

[0042] Returning to Figure 1 , image enhancement is performed in step 104. This image enhancement produces an enhanced digital image 106. As Figure 1As can be seen, the disturbed image portion in the enhanced image 106 has been replaced by the undisturbed portion, thereby producing an overall improved image 106.

[0043] In one embodiment, step 104 is performed by a generative machine learning model (also referred to herein as an enhancement model). Multiple implementations are possible. As an option, the baseline for such a model can be a deep convolutional neural network, which can be used for image-to-image translation, deformation, and / or style transfer applications. The inventors have performed tests using a generative adversarial network (GAN) design (such as a StyleGAN-like structure) and achieved good results, but the principles of the present invention are not limited to this particular implementation. In an example implementation, the generator attempts to generate a new enhanced image 106 based on the original distorted image 102 as input, replacing the disturbed portion with new portions (healthy portions, since in the preferred embodiment, all training data has healthy examples). The discriminator compares the newly generated image 106 with the original ground truth image.

[0044] Returning to Figure 1 , image 102 is also the input for the image segmentation in step 108. Steps 104 and 108 can be performed simultaneously, or can be performed in any suitable order. The image segmentation step 108 produces a segmented or masked image 114, in which the disturbed portion of the input image 102 has been masked.

[0045] In one embodiment, step 108 is performed by a segmentation machine learning model (also referred to herein as a segmentation model). As a non-limiting practical implementation, the model can be a U-Net. The segmentation model can be trained using at least partially synthetic data, where ground truth data is available, which can lead to good training results. In one example, the training data includes two different types of segments, namely clean segments and enhanced segments.

[0046] In one embodiment, the training data and test data are split into dedicated subsets for all models and for training, as they will later be used for verification and should not be affected by overlapping data during training.

[0047] Returning to Figure 1 , in step 110, image segmentation is applied to the enhanced image 106 to produce a masked enhanced image 112.

[0048] In step 116, the masked enhanced image 112 and the masked original image 114 are compared, and a similarity score is calculated. The similarity score can be calculated using various techniques, such as basic comparison, matrix section comparison, or more advanced machine learning solutions, such as a CNN-based deep siamese network.

[0049] The similarity score can be used to verify the results produced by the enhancement model. This ensures that the clean sections (i.e., the sections in the input image 102 that depict the healthy parts of the oral cavity of a person) are not overly altered. The similarity score should generally be high, or even very high. Otherwise, the enhancement model may not be able to process the provided image, and medical professionals should not use the resulting enhanced image alone for diagnosis, but rather use the original image or use both images. In such a case, it may not be possible to make a diagnosis based on these images. The risk of false positives can be verified based on the test dataset and shown to be lower than the typical human error for the perturbed images.

[0050] As an alternative to using similarity calculations during inference, a trained segmentation model can also be used during the training of the enhancement model to compare specific sections with the ground truth and penalize specific mismatches in the clean regions during training. In this case, the segmentation model should not be used during inference to perform the verification, as the corresponding similarity check has been used to train the algorithm.

[0051] Optionally, a dental problem detection algorithm can be used to detect features of "dental problem-like" that may be present in the generated region of the enhanced image 106.

[0052] Figure 3 Exemplary training image data available according to an embodiment of the present invention is shown. In Figure 3 the image in the lower half there are interferences. These interferences are not present in Figure 3 the image in the upper half. These two images together form an image pair and serve as a dataset for training data. In the context of the present invention, such data can be used as training data for, for example, training an enhancement model using any type of supervised learning method. In certain embodiments of the present invention, the training data is unique in that the parts of the image that are hidden or partially hidden due to damage, interference, or distortion should not contain any dental (or other medical) problems, as this may lead to a negative diagnosis. The model of the preferred embodiment of the present invention is trained using this training data to fill in the damaged, disrupted, or distorted parts, where due to the training data, the replaced image parts cannot have any dental problems.

[0053] The training data can be collected through specific camera settings, where the captured images may or may not have damage, interference, and / or distortion.

[0054] Although the model can be trained in various ways, certain embodiments of the present invention ensure that the two networks (the enhancement model and the segmentation model) do not have overlapping training data to achieve the verification function.

[0055] The performance of the trained model should achieve security enhancement without affecting diagnostic false positives. In scenarios where the intention is remote diagnosis or classification, the impact of false negatives is less. Since the model is trained only with valid images, where no dental problems are added in the disturbed sections and at the top, a comparison can be made between the original image and the enhanced image for verification. Additionally, in cases where the verification score is low, i.e., the model may have added or removed a health problem and thus should not be used for secure diagnosis, the original disturbed image is available to dental professionals, thereby reducing the risk of false positive diagnosis to a very low level.

[0056] Collecting a sufficient amount of the above-mentioned real-world training data may be difficult. Therefore, synthetic training data can be generated. Figure 4 An artificially generated data sample according to an embodiment of the present invention is shown. This figure can illustrate how (for example) machine learning models and techniques can be used to artificially add disturbances to image data. In this case, Figure 4 the image in the upper half of Figure 4 is from a real oral scan, while the image in the lower half of

[0057] is artificially disturbed. It can be seen that disturbances for implementing the presence of toothpaste are added to the image (see the circled area). Since images taken by intraoral scanners (IOS) are widely available and relatively easy to collect, this technique is particularly feasible. Additionally, information, data, and images of typical damages, disturbances, or distortions are also available and easy to generate. Combining the two (e.g., in a random manner) enables the generation of a large amount of training data. Figure 4 This generation of artificially generated image data depends, for example, on a parametric description of the nature of certain types of disturbances (e.g., toothpaste, vibrations). As a result, it may be possible to start with just one image (e.g.,

[0058] the upper image in

[0059] The IOS used as the base image can include a mixed image of real tooth IOS. Additionally, images obtained from tooth demonstration models can also be used.

[0060] To train the model to keep any dental problems in the undisturbed regions of the image, some embodiments may involve using images in a dataset that have dental problems that are visible and not hidden by the disturbance.

[0061] Known visible dental problems such as tartar, cavities, dental plaque, and / or gum recession can also be added artificially in a manner similar to that explained above in the context of adding disturbances. Such dental problems that affect diagnosis should only be added to the undisturbed regions of the image. For example, in cases that are difficult to classify, such as dental fillings that usually have a lot of similarity to toothpaste in the image, special attention may be required. Figure 4 The lower half of shows that the artificially created disturbance representing toothpaste can show similarity to the filling material.

[0062] The locations for the artificially added dental problems can be randomly selected (similar to the disturbances, as explained above). The dental problems can also be specifically added to common dental problem regions, which can improve the complexity, performance, and scale of the algorithm. As mentioned above, one goal of some embodiments of the present invention is to train the model to add and / or transfer only healthy medical (oral) image segments to the disturbed regions while leaving the undisturbed segments unaffected.

[0063] Additional data augmentation techniques can be employed in some embodiments of the present invention. For example, the images can be subtly adapted or changed, such as adjusting the zoom level, rotation, brightness, and / or color variations. In the case of 3D input data (such as 3DIOS), the camera angle can also be changed. These techniques can be effectively used to even further expand the training dataset without collecting real-life data and at the same time prevent adding dental problems in the distorted regions of the image.

[0064] Figure 5 Illustrates a schematic overview of a model architecture in which embodiments of the present invention can be implemented.

[0065] The principles and various features of the embodiments disclosed herein can be implemented locally, remotely, or distributively, for example, in personal health devices, smartphones, computers, and / or cloud computing environments. Real-time enhancement is possible, and thus implementation as a preview or on the diagnostic interface of a dental practitioner is also possible.

[0066] Figure 6 Illustrates an exemplary graphical user interface for an application for a smartphone or other electronic user device. The enhanced image can be used for diagnosis, as described above. By combining the enhanced image and the original image (after automatic verification) in one view (see Figure 6In the upper left and lower left figures), the effect of achieving twice the result with half the effort can still be achieved, and at the same time, the risk of misdiagnosis is even further reduced by presenting the original image. If the above automatic verification fails, only the original image can be presented, along with a message indicating that the image cannot be enhanced (see Figure 6 in the upper right figure). It is also feasible to present all three image versions (the original image, the enhanced image, and the masked image; see Figure 6 in the lower right figure).

[0067] Although some aspects are described in the context of a device, it is obvious that these aspects also represent a description of the corresponding process, where a block or device corresponds to a process step or the function of a process step. Similarly, the aspects described in the context of a process step also constitute a description of the corresponding module, element, or feature of the corresponding device.

[0068] Embodiments of the present invention may be implemented in a computer system. The computer system may be a local computing device (e.g., a personal computer, laptop, tablet, or phone) having one or more processors and one or more storage devices, or may be a distributed computing system (e.g., a cloud computing system having one or more processors or one or more storage devices distributed at different locations (e.g., distributed at local clients and / or one or more remote server farms and / or data centers)). The computer system may include any circuit or combination of circuits. In one embodiment, the computer system may include one or more processors, which may be of any type. As used herein, a "processor" may mean any type of computing circuit, such as but not limited to a microprocessor, a microcontroller, a complex instruction set microprocessor (CISC), a reduced instruction set microprocessor (RISC), a very long instruction word (VLIW; VLIW) microprocessor, a graphics processor, a digital signal processor (DSP), a multi-core processor, a field programmable gate array (FPGA), or any other type of processor or processing circuit. Other types of circuits that may be included in the computer system may include custom circuits, application specific integrated circuits (ASICs), etc., such as one or more circuits (e.g., communication circuits) for use in wireless devices (such as phones, tablets, laptops, two-way radios, and similar electronic systems). The computer system may include one or more storage devices, which may include one or more storage elements suitable for a particular application, such as main memory in the form of random access memory (RAM), one or more hard disk drives, and / or one or more drives for handling removable media (such as CDs, flash cards, DVDs), etc. The computer system may also include a display device, one or more speakers, and a keyboard and / or controller, which may include a mouse, a trackball, a touch screen, a voice recognition device, or any other device that allows a system user to input information to and receive information from the computer system.

[0069] Some or all of the method steps may be performed by (or using) a hardware device, such as a processor, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the critical method steps may be performed by such a device.

[0070] Depending on the specific implementation requirements, embodiments of the present invention may be implemented in hardware or software. The implementation may be achieved using a non-volatile storage medium (such as a digital storage medium, such as a floppy disk, DVD, Blu-ray disc, CD, ROM, PROM, EPROM, EEPROM, or flash memory), on which electronically readable control signals are stored that can interact (or are capable of interacting) with a programmable computer system to cause the corresponding processing to be performed. Thus, the digital storage medium may be computer-readable.

[0071] Some embodiments according to the present invention include a data carrier having electronically readable control signals that can interact with a programmable computer system to perform any of the methods described herein.

[0072] Generally, embodiments of the present invention can be implemented as a computer program product having program code that, when run on a computer, can effectively perform any method. For example, the program code can be stored on a machine-readable medium.

[0073] Other embodiments include a computer program stored on a machine-readable medium for performing any of the methods described herein.

[0074] In other words, example embodiments of the present invention thus include a computer program having program code that, when run on a computer, is for performing any of the methods described herein.

[0075] Thus, another embodiment of the present invention is a storage medium (or digital storage medium or computer-readable medium) that includes a computer program stored thereon for performing any of the methods described herein when executed by a processor. The data carrier, digital storage medium, or recording medium is generally tangible and / or non-transitory. Another embodiment of the present invention is an apparatus as described herein that includes a processor and a storage medium.

[0076] Thus, another embodiment of the present invention is a data stream or signal sequence representing a computer program for performing any of the methods described herein. For example, the data stream or signal sequence can be configured to be transmitted via a data communication link (such as via the Internet).

[0077] Another example embodiment includes a processing component, such as a computer or a programmable logic device, that is configured or adapted to perform any of the methods described herein.

[0078] Another example embodiment includes a computer on which a computer program for performing any of the methods described herein is installed.

[0079] Another embodiment according to the present invention includes a device or system configured to transmit (e.g., electronically or optically) a computer program for performing any of the methods described herein to a receiver. The receiver can be, for example, a computer, a mobile device, a storage device, etc. The device or system can include, for example, a file server for transmitting the computer program to the receiver.

[0080] In some embodiments, a programmable logic device (e.g., a field programmable gate array, FPGA) may be used to perform some or all of the functionality of the methods described herein. In some embodiments, the field programmable gate array may cooperate with a microprocessor to perform any of the processes described herein. Generally, these methods are preferably performed by any hardware device.

[0081] Embodiments may be based on the use of artificial intelligence, particularly machine learning models or machine learning algorithms. Machine learning may refer to algorithms and statistical models that a computer system can use to perform a specific task without using explicit instructions, rather than relying on models and inference. For example, machine learning may use a data transformation that can be inferred from the analysis of historical and / or training data, rather than using a rule-based data transformation. For example, a machine learning model or a machine learning algorithm may be used to analyze the content of an image. To enable a machine learning model to analyze the content of an image, training images may be used as input and training content information may be used as output to train the machine learning model. By training the machine learning model with a large number of training images and / or training sequences (e.g., words or sentences) and associated training content information (e.g., labels or annotations), the machine learning model "learns" to identify the content of the image such that the machine learning model can be used to identify image content not included in the training data. The same principle may also be applied to other types of sensor data: by training a machine learning model with training sensor data and expected outputs, the machine learning model "learns" the transformation between the sensor data and the output, which can be used to provide an output based on non-training sensor data provided to the machine learning model. The data provided (e.g., sensor data, metadata, and / or image data) may be preprocessed to obtain feature vectors that are used as input to the machine learning model.

[0082] A machine learning model can be trained using training input data. The above example uses a training method called supervised learning. In supervised learning, a machine learning model is trained using multiple training samples, where each sample can include multiple input data values and multiple expected output values, i.e., each training sample is associated with an expected output value. By specifying the training samples and the expected output values, the machine learning model can "learn" which output value to provide based on input samples similar to those provided during training. In addition to supervised learning, semi-supervised learning can also be used. In semi-supervised learning, some of the training samples lack expected output values. Supervised learning can be based on supervised learning algorithms (e.g., classification algorithms, regression algorithms, or similarity learning algorithms). When the output is restricted to a finite set of values (categorical variables), a classification algorithm can be used, i.e., the input is classified into one of the values in the finite set. When the output exhibits a certain numerical value (within a certain range), a regression algorithm can be used. Similarity learning algorithms are similar to classification algorithms and regression algorithms, but similarity learning is based on learning from examples using a similarity function that measures how similar or related two objects are. In addition to supervised learning or semi-supervised learning, unsupervised learning can also be used to train a machine learning model. In unsupervised learning, (only) input data can be provided, and then unsupervised learning algorithms can be used to find structures in the input data (e.g., by grouping or clustering the input data to find commonalities in the data). Clustering is the assignment of input data including multiple input values to subsets (clusters) such that the input values within the same cluster are similar according to one or more (predetermined) similarity criteria and dissimilar to the input values in other clusters.

[0083] Reinforcement learning is a third type of machine learning algorithm. In other words, reinforcement learning can be used to train a machine learning model. In reinforcement learning, one or more software actors (called "software agents") are trained to take actions in an environment. Based on the actions taken, a reward is calculated. Reinforcement learning is based on training one or more software agents to select actions such that the cumulative reward is increased, thus making the software agents better at the tasks they are given (an increase in reward is evidence).

[0084] In addition, some techniques can be applied to some of the machine learning algorithms. For example, feature learning can be used. In other words, a machine learning model can be trained at least in part using feature learning, and / or a machine learning algorithm can include a feature learning component. Feature learning algorithms (called representation learning algorithms) can retain the information in their input but transform it to make it useful, typically used as a preprocessing stage before performing classification or prediction. For example, feature learning can be based on principal component analysis or cluster analysis.

[0085] In some examples, anomaly detection (i.e., outlier detection) can be used, which aims to provide the identification of input values that are suspect because they are significantly different from most of the input and training data. In other words, the machine learning model can be trained at least in part using anomaly detection, and / or the machine learning algorithm can include an anomaly detection component.

[0086] In some examples, the machine learning algorithm can use a decision tree as a prediction model. In other words, the machine learning model can be based on a decision tree. In a decision tree, an observation about an item (e.g., a set of input values) can be represented by a branch of the decision tree, and the output value corresponding to that item can be represented by a leaf of the decision tree. Decision trees can support discrete values and continuous values as output values. When discrete values are used, the decision tree can be called a classification tree; when continuous values are used, the decision tree can be called a regression tree.

[0087] Association rules are another technique that can be used in machine learning algorithms. In other words, the machine learning model can be based on one or more association rules. Association rules are created by identifying relationships between variables given a large amount of data. Machine learning algorithms can identify and / or use one or more ratio rules, which represent knowledge inferred from the data. These rules can be used, for example, to store, manipulate, or apply knowledge.

[0088] Machine learning algorithms are generally based on machine learning models. In other words, the term "machine learning algorithm" can refer to a set of instructions that can be used to create, train, or use a machine learning model. The term "machine learning model" can refer to a data structure and / or set of rules that represents the learned knowledge (e.g., based on training performed by a machine learning algorithm). In an embodiment, using a machine learning algorithm may mean using the underlying machine learning model (or models). Using a machine learning model may mean that the machine learning model and / or the data structure / rule set that constitutes the machine learning model is trained by a machine learning algorithm.

[0089] For example, the machine learning model can be an artificial neural network (ANN, artificial neural network). An ANN is a system inspired by biological neural networks (such as those in the retina or the brain). An ANN includes multiple interconnected nodes and multiple connections (referred to as edges) between the nodes. There are generally three types of nodes: input nodes that receive input values, hidden nodes that are (only) connected to other nodes, and output nodes that provide output values. Each node can represent an artificial neuron. Each edge can send information from one node to another node. The output of a node can be defined as a (non-linear) function of its input (e.g., the sum of its inputs). The input of a node can be used in the function based on the "weights" of the edge or node that provides the input. The weights of the nodes and / or edges can be adjusted during the learning process. In other words, training an artificial neural network can include adjusting the weights of the nodes and / or edges, i.e., achieving a desired output for a given input.

[0090] Alternatively, the machine learning model can be a support vector machine, a random forest model, or a gradient boosting model. A support vector machine (i.e., a support vector network) is a supervised learning model, and its associated learning algorithm can be used to analyze data (e.g., for classification or regression analysis). A support vector machine can be trained by providing an input with multiple training input values (which belong to one of two categories). A support vector machine can be trained to assign new input values to one of the two categories. Alternatively, the machine learning model can be a Bayesian network, which is a probabilistic directed acyclic graph model. A Bayesian network can use a directed acyclic graph to represent a set of random variables and their conditional dependencies. Alternatively, the machine learning model can be based on a genetic algorithm, which is a search algorithm and heuristic technique that simulates the process of natural selection.

Claims

1. An image processing method (100), comprising the following steps: Obtaining a digital image (102) of at least a part of a person's body, wherein the digital image includes at least one disturbed image part, and wherein the at least one disturbed image part is caused by vibration, and / or the presence of a medical, care or body substance such as water, liquid, toothpaste or saliva, and / or the presence of a medical, care or personal health device such as a toothbrush; and Generating (104) an enhanced digital image (106) of at least a part of the person's body, at least partially based on the obtained digital image (102) and using a trained generative machine learning model called an enhancement model, in which the at least one disturbed image part has been replaced by an artificially created image part; wherein the part of the person's body is the oral cavity of the person; and wherein the digital image is obtained by a toothbrush equipped with a camera.

2. The method according to claim 1, wherein only the disturbed image parts including healthy parts are replaced, and the disturbed image parts including unhealthy parts are not replaced.

3. The method according to claim 2, wherein the healthy parts do not include gum health problems, or tooth health problems, or other substantial health problems.

4. The method according to claim 1, wherein the artificially created image part, preferably each artificially created image part, shows a part of the person's body without substantial health problems; and / or wherein the enhancement model has been trained such that it only adds artificially created image parts without substantial health problems.

5. The method according to claim 1 or 2, further comprising: Using a segmentation machine learning model called a segmentation model to segment (108) the obtained digital image (102) to produce an indication of at least one undisturbed image part (114) in the obtained digital image; and Comparing (116) the at least one undisturbed image part (114) in the obtained digital image with at least one corresponding part (112) in the generated enhanced digital image to produce a similarity score.

6. The method according to claim 5, further comprising: If the similarity score reaches or exceeds a predetermined similarity threshold, displaying the obtained digital image (102) and the generated enhanced digital image (106) on an electronic display device.

7. The method according to claim 6, further comprising: Displaying a masked digital image on the electronic display device, the masked digital image indicating the at least one undisturbed image part in the obtained digital image.

8. The method according to any one of the preceding claims, wherein the enhancement model comprises or is a convolutional neural network CNN, a deep CNN, a generative adversarial network GAN, and / or a StyleGAN; and / or wherein the segmentation model includes or is a CNN, a deep CNN, an image segmentation model, and / or a U-Net; and / or wherein the similarity score is generated by a comparison model, the comparison model including or being a CNN, a deep CNN, a siamese neural network, and / or a deep siamese neural network.

9. A method of training a generative machine learning model for generating an enhanced digital image (104) of at least a part of a human body for use in the method according to any one of claims 1 to 8, the method comprising: providing a training data set including a plurality of pairs of digital images, each pair including: a first digital image of at least a part of a human body, wherein the first digital image includes at least one disturbed image portion, the at least one disturbed image portion being caused by vibration and / or the presence of a medical, care or body substance such as water, liquid, toothpaste or saliva and / or the presence of a medical, care or personal health device such as a toothbrush; and a second digital image of the at least a part of the human body, in which the at least one disturbed image portion is missing or attenuated; and using the training data set to train the generative machine learning model; wherein the part of the human body is the oral cavity of the human; wherein the first digital image is acquired by a toothbrush equipped with a camera.

10. The method according to claim 9, wherein the training data set includes at least one synthetic digital image, in which at least one artificially disturbed image portion has been added to a digital image of at least a part of a human body; wherein optionally, at least one tooth or gum problem, such as tartar, dental caries, dental plaque and / or gum recession, has been artificially added to at least one of the synthetic digital images.

11. The method according to any one of claims 9 to 10, further comprising: validating the performance of the machine learning model, in particular using a validation data set different from the training data set.

12. A computer program or a computer-readable medium storing a computer program, the computer program including instructions which, when the program is executed on a computer and / or a computer network, cause the computer and / or the computer network to execute the method according to any one of claims 1 to 11.

13. A machine learning model data structure embodying a machine learning model, in particular a generative machine learning model, for generating an enhanced digital image of at least a part of a human body, the machine learning model data structure being trained using the method according to any one of claims 9 to 11.

14. A training, validation, verification, or test data set for a machine learning model, particularly for a generative machine learning model, the data set comprising a collection of the first digital image defined in claim 9, and / or the second digital image defined in claim 9, and / or the synthetic digital image defined in claim 10, and / or the generated enhanced digital image defined in claim 1, and / or the masked digital image defined in claim 7.

15. A data processing device or system comprising means for performing the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Automated detection, generation and / or correction of dental features in digital models

    US20210059796A1

  • Artificial Intelligence Architecture For Evaluating Dental Images And Documentation For Dental Procedures

    US20220012815A1

  • Dental feature identification for assessing oral health

    WO2021175713A1