Method for generating image of simulated development or treatment result of dental condition
By generating images that simulate the development of dental conditions or the outcome of treatment using deep neural networks, the problem of low communication efficiency between doctors and patients in dental diagnosis and treatment is solved, and clearer communication results are achieved.
Patent Information
- Application Number
- PCT/CN2025/089284
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-29
- Filing Date
- 2025-04-16
- Publication Date
- 2025-12-04
Smart Images

Figure CN2025089284_04122025_PF_FP_ABST
Abstract
Description
Methods for generating images that simulate the development of dental conditions or the outcome of treatment. Technical Field
[0001] This application generally relates to methods for generating images that simulate the development or treatment outcomes of a dental condition. Background Technology
[0002] In the field of dental care, dentists often need to diagnose a patient's dental condition based on images of the patient's oral cavity, such as panoramic radiographs and dental photographs. After diagnosis, the dentist will communicate with the patient about subsequent treatment methods. However, whether communication is based on verbal means or images, it can be affected by other factors, such as patient bias or misunderstanding of medical terminology, thus reducing the efficiency of the consultation.
[0003] In this context, having images simulating the development of a patient's dental condition or the outcome of treatment can effectively improve the efficiency of doctor-patient communication. Therefore, it is necessary to provide a method for generating images simulating the development of a dental condition or the outcome of treatment. Summary of the Invention
[0004] One aspect of this application provides a computer-executed method for generating images simulating the development or treatment outcome of a dental condition, comprising: acquiring a first image obtained by overlaying a mask onto a second image, the second image showing dental tissues of a patient; acquiring a simulation category representing the development or treatment outcome of a dental condition to be simulated; and generating a third image based on the first image and the simulation category using a trained deep neural network, the third image having a region corresponding to the mask representing the simulated development or treatment outcome of the dental condition.
[0005] In some implementations, the third image is substantially the same as the first image in the portion outside the masked area.
[0006] In some embodiments, the second image shows the current state of the patient's dental-related tissues and includes a dental condition area representing the portion of the patient's dental-related tissues that requires treatment, the mask covering the dental condition area.
[0007] In some implementations, when the simulation category is the outcome of the dental condition, the coverage of the mask is determined manually by a dental professional based on experience.
[0008] In some implementations, the deep neural network is an image completion network that can complete the image of the masked region based on the first image and the simulated category.
[0009] In some implementations, the deep neural network is a generative adversarial network.
[0010] In some implementations, the adversarial generative network includes an encoder and a decoder, and the method further includes: the encoder extracting features from the first image to obtain a feature map; the encoder sampling the distribution of the map in the first image and incorporating an embedding corresponding to the simulated category into the sampling result; and the decoder generating the third image based on the feature map and the sampling result.
[0011] In some implementations, the adversarial generative network further includes a discriminator, which, during the training of the adversarial generative network, calculates a multi-class cross-entropy based on the completed image generated by the adversarial generative network to determine whether the completed image belongs to a specified simulated class, wherein the completed image is the portion of the image generated by the adversarial generative network that corresponds to the mask.
[0012] In some implementations, the adversarial generative network (GCN) is trained using a training dataset comprising multiple sets of training data. Each set of training data includes a mask image, a completed image, and a simulated category. The mask image is an image of a patient's dental tissue overlaid with a mask, and the completed image is an image corresponding to the masked region. The training of the GCN includes: a first training step for training the GCN's ability to generate realistic images, in which both the mask image and the completed image from the training dataset are input into the GCN; and a second training step for training the GCN's ability to generate completed images conforming to a specified simulated category, in which only the mask image and the completed image from the training dataset are input into the GCN.
[0013] In some implementations, during the first training step, the encoder generates feature maps and sampling results based on the mask image and the completed image, respectively. The sampling results of the two are added to the feature map of the mask image and then input into the decoder to generate a complete image, wherein the complete image refers to an image of dental tissues without masking. Attached Figure Description
[0014] The above and other features of this application will be further described below with reference to the accompanying drawings and their detailed description. It should be understood that these drawings only illustrate several exemplary embodiments according to this application and should not be considered as limiting the scope of protection of this application. Unless otherwise specified, the drawings are not necessarily to scale, and similar reference numerals denote similar parts.
[0015] Figure 1 is a schematic flowchart of a method for generating images of simulated dental condition development or treatment outcomes in one embodiment of this application;
[0016] Figure 2 schematically illustrates the structure of the deep neural network used in this application in one embodiment;
[0017] Figure 3A shows an example of an image to be completed;
[0018] Figure 3B shows an example of a completed image generated using the method of this application, based on the image shown in Figure 3A, with dental caries as the simulated category.
[0019] Figure 3C shows an example of a completed image generated using the method of this application, based on the image shown in Figure 3A, with tooth filling as the simulation category.
[0020] Figure 4A shows the images to be completed in a set of training data;
[0021] Figure 4B shows the image of the completed portion from the same training data as the image shown in Figure 4A;
[0022] Figure 4C shows the complete image obtained by merging the images shown in Figure 4A and Figure 4B;
[0023] Figure 5A shows the image to be completed in another set of training data;
[0024] Figure 5B shows the image of the completed portion from the same training data as the image shown in Figure 5A; and
[0025] Figure 5C shows the complete image obtained by merging the images shown in Figure 5A and Figure 5B. Detailed Implementation
[0026] The following detailed description incorporates the accompanying drawings, which form part of this specification. The illustrative embodiments mentioned in the specification and drawings are for illustrative purposes only and are not intended to limit the scope of this application. Those skilled in the art will understand, based on the teachings of this application, that many other embodiments can be employed and various changes can be made to the described embodiments without departing from the spirit and scope of this application. It should be understood that the various aspects of this application illustrated herein can be arranged, substituted, combined, separated, and designed in many different configurations, all of which are within the scope of this application.
[0027] One aspect of this application provides a computer-executed method for generating images simulating the development or treatment outcome of a dental condition, which generates images simulating the development or treatment outcome of a dental condition based on an image of a patient's current dental condition, wherein the development of a dental condition refers to the degree to which the patient's dental condition will deteriorate in the future without treatment.
[0028] Please refer to Figure 1, which is a schematic flowchart of a computer-executed method 100 for generating images of simulated dental condition development or treatment outcomes in one embodiment of this application.
[0029] In 101, obtain the first image showing the patient's dental-related tissues along with a mask.
[0030] The first image shows the current state of the patient's dental-related tissues, including at least one area of dental condition. The dental condition area refers to the region of dental-related tissue in the first image that requires treatment or intervention. Dental-related tissues include, but are not limited to, teeth, alveolar bone, periodontal ligament, gingiva, and temporomandibular joint. Dental conditions include, but are not limited to, dental caries, periapical periodontitis, residual roots, and missing teeth.
[0031] The first image can be an image taken using X-rays, such as a panoramic or lateral view, or it can be an image taken using a conventional camera, such as an image taken using the dental imaging apparatus disclosed in Chinese Patent No. ZL202220617972.5 or Chinese Patent No. ZL202223452359.1. The type of the first image depends on the needs of the dental diagnosis.
[0032] In one embodiment, the mask may be a mask manually annotated on the first image by a dental professional, covering at least the area of the dental condition. In some cases, without treatment or intervention, the deterioration of the patient's dental condition may lead to an expansion of the dental condition area. To simulate this situation later, the dental professional may, based on experience, predict the outcome of the deterioration and appropriately expand the area covered by the mask.
[0033] In another embodiment, the mask may also be generated by a trained deep neural network based on the first image.
[0034] In 103, obtain the simulation category.
[0035] The simulation category refers to the category of the simulation result. For example, for the dental condition of tooth decay, if we want to simulate the result of tooth decay worsening without treatment, the simulation category can be defined as "tooth decay". Subsequently, the deep neural network will generate an image within a masked area based on the simulation category of "tooth decay" to simulate the result of tooth decay worsening without treatment. If we want to simulate the result of treating tooth decay with a filling, the simulation category can be defined as "filling". Subsequently, the deep neural network will generate an image within a masked area based on the simulation category of "filling" to simulate the result of treating tooth decay with a filling.
[0036] For some dental conditions, similarly, two simulation categories can be defined: one corresponding to the progression or deterioration of the condition without treatment or intervention, and the other corresponding to the outcome of treatment or intervention. For other dental conditions, only one simulation category can be defined, corresponding to the outcome of treatment or intervention. For example, for dental conditions involving missing teeth, only one simulation category can be defined, such as dental implants. Understandably, for some dental conditions, there may be different treatment methods; correspondingly, a simulation category can be defined for each treatment method.
[0037] The simulation category can be selected by dental professionals based on the patient's dental condition and needs (i.e., whether to simulate a worsening situation or the outcome of treatment).
[0038] In step 105, a second image is generated based on the first image, a mask, and a simulated category using a trained deep neural network.
[0039] First, the mask is superimposed on the first image to obtain a third image, which is the first image in which the masked area is not visible; this third image can also be called the mask image. Next, the deep neural network generates the second image based on the third image and the simulated category.
[0040] The second image is an image representing the simulation result. Compared with the first image, the two are basically the same in the parts outside the masked area, but different in the parts inside the masked area. The part of the second image within the masked area is the completion part generated by the deep neural network based on the third image and the simulation category.
[0041] The deep neural network is a neural network used to complete images. In one embodiment, it can be a Generative Adversarial Network (GAN).
[0042] Please refer to Figure 2, which schematically illustrates the structure of a deep neural network 200 in one embodiment of this application. The process of the deep neural network 200 generating the second image will be described in detail below with reference to Figure 2.
[0043] The deep neural network 200 mainly includes an encoder 201, a decoder 203, and a discriminator 205.
[0044] When the second image is generated based on the third image and the simulated category, image_m301 in Figure 2 is the third image. image_m represents an image showing the patient's dental tissues overlaid with a mask, wherein the masked area is not visible.
[0045] Encoder 201 extracts features from image_m 301 to obtain feature map_m 303, where feature map represents the feature map.
[0046] Encoder 201 randomly samples the distribution of the image_m 301 to obtain sample_m 305, where sample represents the sampling result. The corresponding embedding is obtained according to the simulated category, concatenated with sample_m 305, and then fused through convolution. In Figure 2, class embedding 307 represents the embedding corresponding to the simulated category.
[0047] Feature map_m 303 and sample_m 305 are added together and then input into decoder 205 to generate the second image. Image_gen 309 in Figure 2 represents the second image, which can also be called a complete image because it is unmasked.
[0048] Please refer to Figure 3A, which shows an example image_m, where the two black areas are masks.
[0049] When the simulation category is dental caries, the deep neural network 200 generates the image shown in Figure 3B based on the image shown in Figure 3A, wherein the darker shades in the image completion part (i.e. the part corresponding to the mask) represent the dental caries part.
[0050] When the simulation category is tooth filling, the deep neural network 200 generates the image shown in Figure 3C based on the image shown in Figure 3A, wherein the lighter shade in the image completion part represents the tooth filling part.
[0051] In one embodiment, the training dataset for training the deep neural network 200 may include multiple sets of training data. One set of training data may include an image_m, an image_c, and a simulated class. Here, image_c represents the image of the completed masked region, also known as the completed image. It can be understood that image_m and image_c can be generated based on a complete image showing the patient's dental tissue and a corresponding mask. For example, image_m is obtained by overlaying the mask onto the complete image, and image_c is obtained by using the mask to extract the masked region from the complete image.
[0052] In one embodiment, a set of training data can be used to perform a first training and a second training on the deep neural network 200. The first training trains the encoder and decoder to generate more realistic images, and the second training trains the deep neural network 200 to generate completed images that conform to a specified simulation category. In one embodiment, these two trainings can be performed in parallel. These two trainings are described in detail below.
[0053] During the first training, image_m 301 and image_c 311 are simultaneously input into deep neural network 200. For image_m 301, encoder 201 generates feature map_m 303 and sample_m 305, which are then added together. Correspondingly, for image_c 311, encoder 201 generates feature map_c 313 and randomly samples the distribution of the image_c 311 to obtain sample_c 315. Then, feature map_m 303 and sample_c 315 are added together. Since sample_c 315 contains no information other than the mask, only information about the part to be completed, feature map_m 303 is added to sample_c 315 instead of feature map_c 313 to sample_c 315. The results of the two additions are concatenated and input into decoder 203, which generates image_rec 315, the reconstructed complete image.
[0054] Discriminator 207 calculates L1 loss based on image_rec 315 and ground truth, and determines whether image_rec 315 is a real image based on the calculated L1 loss. The ground truth can be obtained by merging image_mask 301 and image_complement 311.
[0055] During the second training, image_mask 301 and the simulated class are input into the deep neural network 200 to generate image_gen 309.
[0056] On one hand, the discriminator 207 calculates the L1 loss based on image_gen 309 and the ground truth, and determines whether image_gen 309 is a real image based on the L1 loss. In one embodiment, the region corresponding to the mask in image_gen 309 can be extracted, and image_complement 311 can be used as the ground truth, and the L1 loss can be calculated based on both.
[0057] On the other hand, the discriminator 207 calculates a multi-class cross-entropy based on the region corresponding to the mask in image_gen 309, and determines whether the completed image belongs to the specified class accordingly.
[0058] Please refer to Figures 4A-4C, which show images from a set of training data. Figure 4A is the image mask, Figure 4B is the image complement, and Figure 4C is the complete image obtained by merging Figures 4A and 4B. The simulated category for this set of training data is dental caries, so the darker shaded areas in Figure 4B represent dental caries.
[0059] Please refer to Figures 5A-5C, which show images from another set of training data. Figure 5A is the image_mask, Figure 5B is the image_complement, and Figure 5C is the complete image obtained by merging Figures 5A and 5B. The simulation category of this set of training data is tooth filling, so the brighter part in Figure 5B represents the filling.
[0060] Although various aspects and embodiments of this application have been disclosed herein, other aspects and embodiments of this application will be apparent to those skilled in the art upon inspiration from this application. The various aspects and embodiments disclosed herein are for illustrative purposes only and not for limiting purposes. The scope and spirit of this application are determined solely by the appended claims.
[0061] Similarly, the diagrams may illustrate exemplary architectures or other configurations of the disclosed methods and systems, which aid in understanding the features and functions that may be included in the disclosed methods and systems. The claims are not limited to the exemplary architectures or configurations shown, and the desired features may be implemented with various alternative architectures and configurations. Furthermore, the order of the blocks given herein with respect to flowcharts, functional descriptions, and method claims should not be limited to various embodiments implemented in the same order to perform the said functions, unless explicitly indicated in the context.
[0062] Unless otherwise expressly stated, the terms and phrases used herein, and their variations thereof, should be interpreted as open-ended rather than restrictive. In some instances, the appearance of extended words and phrases such as “one or more,” “at least,” “but not limited to,” or other similar expressions should not be construed as an intention or necessity to indicate a narrower scope in examples where such extended expressions might not exist.
Claims
1. A computer-executed method for generating images simulating the development or treatment outcome of a dental condition, comprising: The first image is obtained by overlaying a mask onto a second image, which shows the patient's dental tissues. Obtain the simulation category, which represents the development or treatment outcome of the dental condition that needs to be simulated; as well as A third image is generated based on the first image and the simulated category using a trained deep neural network. The region of the third image corresponding to the mask is a simulated development or treatment outcome of the dental condition.
2. The method as described in claim 1, characterized in that, The third image is substantially the same as the first image in the portion outside the masked area.
3. The method as described in claim 1, characterized in that, The second image shows the current state of the patient's dental-related tissues and includes a dental condition area representing the portion of the patient's dental-related tissues that requires treatment, the mask covering the dental condition area.
4. The method as described in claim 3, characterized in that, When the simulation category represents the developmental outcome of the dental condition, the coverage of the mask is determined manually by dental professionals based on experience.
5. The method as described in claim 1, characterized in that, The deep neural network is an image completion network that can complete the image of the masked region based on the first image and the simulated category.
6. The method as described in claim 5, characterized in that, The deep neural network in question is a generative adversarial network.
7. The method as described in claim 6, characterized in that, The adversarial generative network includes an encoder and a decoder, and the method further includes: The encoder extracts features from the first image to obtain a feature map; The encoder samples the distribution of the graph in the first image and incorporates the embedding corresponding to the simulation category into the sampling result; and The decoder generates the third image based on the feature map and the sampling results.
8. The method as described in claim 7, characterized in that, The adversarial generative network further includes a discriminator. During the training of the adversarial generative network, the discriminator is used to calculate a multi-class cross-entropy based on the completed image generated by the adversarial generative network to determine whether the completed image belongs to a specified simulated class. The completed image is the part of the image generated by the adversarial generative network that corresponds to the mask.
9. The method as described in claim 8, characterized in that, The adversarial generative network is trained using a training dataset, which includes multiple sets of training data. Each set of training data includes a mask image, a completed image, and a simulated category. The mask image is an image of the patient's dental tissue overlaid with a mask, and the completed image is the image corresponding to the masked region. The training of the adversarial generative network includes: The first training step, used to train the adversarial generative network's ability to generate realistic images, involves inputting both the masked image and the completed image from a set of training data into the adversarial generative network; and The second training step is used to train the adversarial generative network to generate completed images that conform to a specified simulation category. In this training, only the masked image and the completed image in the set of training data are input into the adversarial generative network.
10. The method as described in claim 9, characterized in that, In the first training step, the encoder generates feature maps and sampling results based on the mask image and the completed image, respectively. The sampling results of the two are added to the feature map of the mask image and then input into the decoder to generate a complete image. The complete image refers to an image of dental tissues without masking.
Citation Information
Patent Citations
Method for generating dental images
CN114586069A
Method for generating tooth-exposed face image with aligned teeth of patient based on generative adversarial network
CN116630149A
Method and device for generating orthodontic effect preview image
CN117115352A
Domain specific image quality assessment
US20210174477A1