Focus feature-based dental fluorosis end-to-end diagnosis method
By constructing a lightweight convolutional neural network and model pruning technology based on MobileNet-V3, the problem of high information cleavage and computational complexity in fluorosis recognition is solved, and efficient and accurate fluorosis diagnosis is achieved, which is suitable for resource-constrained devices.
Patent Information
- Application Number
- CN202510774091.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art has problems such as grouping convolution information fragmentation, excessive parameter quantity and high computational complexity in the recognition of fluorosis, resulting in low accuracy and efficiency of diagnostic results.
Using an end-to-end diagnostic method based on lesion features, a lightweight convolutional neural network based on MobileNet-V3 is used to combine data augmentation and model pruning technology to build a fluorosis tooth recognition model, extract key features of tooth images through feature lesion module, and enhance the focusing ability of the model using mask generation and image reconstruction mechanism.
It realizes high-precision tooth fluorosis recognition, significantly reduces the computing volume and memory requirements, improves the accuracy and efficiency of diagnosis, is suitable for resource-constrained equipment, and provides reliable diagnostic auxiliary tools.
Smart Images

Figure CN120299686A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of dental disease diagnosis, and particularly relates to an end-to-end diagnosis method for dental fluorosis based on lesion characteristics. Background Art
[0002] Dental fluorosis, also known as enamel fluorosis or mottled enamel, is a dental hard tissue disease caused by excessive intake of fluoride by the human body during the tooth development and mineralization period. Its pathogenesis is mainly that excessive fluoride interferes with the normal function of ameloblasts during the enamel formation process, resulting in enamel hypoplasia and abnormal mineralization. This disease not only causes white spots, yellowish-brown or brown patches on the teeth in appearance, but also causes surface defects and unevenness of the teeth when severe, greatly affecting dental health.
[0003] From the aesthetic point of view, dental fluorosis has a negative impact on the social and mental health of patients. In today's society, people increasingly pay attention to their external image, and a set of white and neat teeth is often regarded as a symbol of health and confidence. However, due to the abnormal color and shape of their teeth, patients with dental fluorosis may have an inferiority complex in interpersonal communication, affecting their normal social activities and mental health. In terms of function, the hard tissue mineralization degree of the teeth of patients with dental fluorosis is relatively low, and the enamel structure is loose, which reduces the wear resistance and acid resistance of the teeth and makes them more vulnerable to wear and dental caries. Some studies have shown that the probability of dental caries in patients with dental fluorosis is several times higher than that of normal people.
[0004] Globally, the distribution of dental fluorosis shows obvious regional characteristics, mainly concentrated in some high-fluoride areas. In China, the prevalence of dental fluorosis is also relatively serious. According to the relevant survey data of the National Health Commission, there are dental fluorosis epidemic areas of varying degrees in many provinces across the country, and dental fluorosis seriously affects the oral health and quality of life of local residents.
[0005] The screening of dental fluorosis is an important part of the national prevention and control work of endemic fluorosis. How to complete the screening of dental fluorosis in an efficient and accurate manner is an urgent problem to be solved. At present, the identification of dental fluorosis in clinical practice mainly relies on the naked-eye observation and experience judgment of doctors. However, this is easily affected by the subjective factors of doctors, resulting in poor accuracy and consistency of the diagnosis results. In addition, the efficiency of traditional identification methods is relatively low. In large-scale oral disease screenings, such as community oral health surveys or school oral examinations, doctors need to observe and judge the teeth of patients one by one in detail, which requires a lot of time and effort.
[0006] In recent years, deep learning has made breakthrough progress in the field of image recognition, demonstrating many advantages. By training on a large amount of image data, it automatically extracts complex features in images, enabling accurate recognition and classification of new and unseen images. Currently, deep learning has been widely applied to tasks such as disease diagnosis, lesion detection and classification, and has achieved remarkable results.
[0007] The fluorosis tooth recognition methods adopted in the prior art have problems such as fragmented grouped convolution information, excessive number of parameters, and high computational complexity. Summary of the Invention
[0008] The purpose of the present invention is to provide an end-to-end diagnosis method for fluorosis teeth based on lesion features, aiming to solve the problems of fragmented grouped convolution information, excessive number of parameters, and high computational complexity in the prior art.
[0009] The present invention is implemented as follows. An end-to-end diagnosis method for fluorosis teeth based on lesion features, the method comprising: Collect oral cavity photos and annotate the oral cavity photos; Construct an oral cavity tooth detection model, process the oral cavity photos through the oral cavity tooth detection model, and output tooth images; Construct a dataset based on the tooth images, and perform data augmentation and expansion on the dataset; Construct a fluorosis tooth recognition model based on the MobileNet-V3 basic framework, perform pruning processing on it, perform integration processing on the model to obtain an integrated model, and perform fluorosis tooth recognition processing on the image to be recognized through the integrated model.
[0010] Preferably, in the step of collecting oral cavity photos and annotating the oral cavity photos, collect a preset number of oral cavity photos, where the oral cavity photos include normal tooth photos, fluorosis tooth photos, and suspicious tooth photos, and their ratio is a preset value. Extract the normal tooth photos and fluorosis tooth photos among them, construct a plurality of non-overlapping subsets, and set a test set.
[0011] Preferably, in the step of annotating the oral cavity photos, use the labelimg image annotation software to annotate the oral cavity tooth part, and generate a file for each picture storing the position information of the rectangular frame and the target category information.
[0012] Preferably, in the step of constructing the oral tooth detection model, a deep learning model framework is built based on the pytorch neural network library, and the oral tooth detection is realized based on the YOLOv11n basic model framework and trained. The oral tooth detection model consists of a backbone network Backbone, a neck network Neck, and a detection head Head. The backbone network Backbone extracts features from the image through downsampling of the convolutional layer, feature fusion of the C3k2 module, multi-scale pooling of the SPPF module, and feature enhancement of the C2PSA module, specifically including 6 convolutional layers Conv, 6 C3k2 modules, 1 SPPF module, and 1 C2PSA module.
[0013] Preferably, the C3k2 module is the core component of the Backbone, designed based on the CSP structure, and the input feature map is divided into two parts for separate processing and then merged. The C3k2 module contains two Bottleneck modules and two C3k modules. The SPPF module realizes the fusion of multi-scale features through spatial pyramid pooling operations.
[0014] Preferably, the neck network Neck consists of multiple upsampling Upsample, feature concatenation Concat, convolutional layer Conv, and C3k2 modules, which are used to fuse multi-scale features and enhance the feature expression ability. The low-resolution feature map is upsampled and concatenated with the corresponding scale feature map from the backbone network Backbone to form multi-scale feature fusion.
[0015] Preferably, in the step of constructing a dataset based on tooth images and performing data augmentation and expansion on the dataset, the images are unified to a preset size, randomly horizontally flipped, converted from the PIL format to PyTorch tensors, normalized for each channel of the image, and normalized using predefined mean and standard deviation. Horizontal flipping and vertical flipping are used to simulate images from different perspectives, and random rotation at different angles is used to simulate changes in the image in different directions. The number of images is increased to a specified number.
[0016] Preferably, in the step of constructing a dental fluorosis recognition model based on the MobileNet-V3 basic framework, three feature lesion modules of dynamic color basis decomposition, multi-scale texture extraction, and edge-sensitive convolution are defined. The dynamic color basis decomposition feature lesion module contains three learnable basis vectors of chalk spot, enamel pigmentation, and background suppression, and is used to generate dynamic mixing weights; the multi-scale texture extraction feature lesion module generates an LBP feature map reflecting the local texture features of the image by calculating the relative gray-scale relationship between each pixel in the image and its neighboring pixels; the edge-sensitive convolution module defines a horizontal edge detection convolution kernel of a Sobel operator and converts it into a torch.Tensor type. In the forward propagation process of the model, the input features of the three feature lesion modules are concatenated and fused, imported into the backbone network for simulation training, and these feature maps are input into the classifier to finally output the classification result.
[0017] Preferably, in the step of constructing a dental fluorosis recognition model based on the MobileNet-V3 basic framework, a branch is established to provide feedback through mask generation and image reconstruction tasks. A mask generation branch is defined, which consists of two convolutional layers and uses the Sigmoid function to output a mask with a range of [0, 1]. An image reconstruction branch is defined, which uses convolutional layers and deconvolutional layers to gradually upsample the feature map to the size of the original image and uses the Sigmoid function to limit the output within the range of [0, 1].
[0018] Preferably, in the integrated model, the Adam optimizer is adopted, and the StepLR learning rate scheduler is used. An L2 regularization term is introduced to prevent the model from overfitting after pruning; in terms of model performance evaluation, the k-fold cross-validation method is adopted.
[0019] The end-to-end dental fluorosis diagnosis method based on lesion features provided by the present invention predicts and classifies by selecting candidate regions. In the first stage, the oral teeth part in the face image is detected, and the detection accuracy reaches more than 90%. The tooth image in the detection frame is extracted for subsequent prediction and classification tasks, effectively reducing the interference of background noise, greatly improving the accuracy of diagnosis and recognition, realizing end-to-end design, lightweight MobileNet-V3 backbone network and model pruning operations, significantly reducing the calculation amount and memory, greatly improving the inference speed, and also facilitating easy deployment on resource-constrained devices, making up for the disadvantages of poor consistency, time-consuming and laborious in the current traditional visual inspection methods in clinical diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a flowchart of the end-to-end dental fluorosis diagnosis method based on lesion features provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0022] It can be understood that the terms "first", "second", etc. used in this application can be used in this document to describe various elements, but unless otherwise specified, these elements are not limited by these terms. These terms are only used to distinguish the first element from another element. For example, without departing from the scope of this application, the first xx script can be called the second xx script, and similarly, the second xx script can be called the first xx script.
[0023] Optimize the model feature extraction according to the characteristics of dental fluorosis lesions. According to the diagnostic hygienic industry standard for dental fluorosis in China (WS / T 10028—2024), the following design is carried out on the model to enable the automatic diagnosis of the deep learning model to conform to the diagnostic standard basis.
[0024] (1) Chalky changes: That is, part or all of the enamel on the tooth surface loses its luster, showing opaque chalky or rough stripe, spot, plaque-like patterns, or the entire tooth surface presents a pathological change like white marble. Based on this, we will design a multi-branch network starting from the characteristics of color transparency and texture, and then fuse the features.
[0025] Regarding the color transparency problem, we use the method of converting the color space to Lab, design a color-sensitive module, and enhance the differences in the Lab channels. In the Lab space, the L channel is the brightness, and the a and b channels are the color components. The white of normal teeth usually has luster, while the white of chalky spots is opaque and may be darker. Therefore, compared with normal teeth, chalky spots have different distributions on Lab, with higher L channel values but lower a and b channel values.
[0026] Regarding the texture problem. The surface of chalky spots is rough and granular, while normal teeth are smooth. Based on this, a multi-scale texture extraction module is designed, mainly used to capture the rough texture of chalky spots (i.e., chalky granularity). In terms of structural design, first, fine-grained LBP is adopted, that is, the local binary pattern is used to extract tiny spots to obtain the fine texture information in the image. At the same time, through wavelet high-frequency enhancement means, the Haar wavelet transform is used to extract the high-frequency subbands (HL / LH / HH) to enhance the granularity in the image. Finally, through the two-stream fusion method, the LBP features and wavelet features are spliced and then input into the lightweight convolutional layer to further integrate and optimize the texture features.
[0027] Jointly learn the above two branches.
[0028] Enamel discoloration: The enamel on the surface of dental fluorosis teeth shows varying degrees of color changes, such as light yellow, yellowish-brown, dark brown or black. To further enhance the extraction of key color features in dental images, a dynamic color basis decomposition function is introduced. This function aims to decompose the input image into basis vectors such as yellowish-brown, which is used to accurately characterize the enamel discoloration situation. Its implementation process is based on the Lab color space and is dynamically weighted and mixed through learnable basis vectors. In this process, the model can effectively enhance pathological color features, such as the unique yellowish-brown tone of dental fluorosis, while suppressing background interference, such as irrelevant color information like the red color of the gums. This operation provides a more targeted color information dimension for subsequent classification and diagnosis, helping to improve the model's analysis and recognition ability of color features related to dental fluorosis.
[0029] Enamel defect: To further enhance the response to enamel defect features, the second-layer convolutional kernel is initialized as a Sobel edge detection operator. Enamel defects usually have obvious edge features, and the Sobel operator can highlight these edges, enabling the model to focus on the edges of enamel defects during the early convolution process. Whether it is a small-scale local enamel defect or a large-scale enamel exfoliation, their edges can generate a stronger response under the action of the convolutional kernel initialized by the Sobel operator, providing key edge information for the subsequent model learning and judgment of enamel defects.
[0030] Build an end-to-end deep learning network model and achieve efficient image feature extraction and classification with a relatively small model size, providing strong support for deployment on mobile and embedded devices. Users only need to simply upload dental images to quickly obtain diagnostic results, providing an efficient and convenient innovative tool for oral health screening and diagnosis.
[0031] (1) Build an end-to-end model architecture. Compared with the model trained in stages, end-to-end training can automatically balance the relationship between various modules during the training process, directly adjust the parameters of the entire model according to the final output result, enabling the model to better learn the mapping relationship from input to output, and avoiding the performance degradation caused by the mismatch between sub-models. The automatic transfer and processing of data within the model make the entire system more integrated and efficient, and can be better applied to actual image analysis tasks. In addition, the end-to-end model design avoids the errors and low efficiency caused by manual operations in the traditional staged processing method.
[0032] (2)Select MobileNet-V3 as the basic architecture. MobileNet-V3 is a lightweight convolutional neural network structure designed for efficient operation in resource-constrained environments such as mobile devices. It mainly adopts depthwise separable convolutions, that is, splitting the standard convolution into a depthwise convolution and a pointwise convolution. The depthwise convolution performs convolution operations independently for each input channel, only considering the feature extraction in the spatial dimension; the pointwise convolution linearly combines the output channels of the depthwise convolution through 1×1 convolution kernels, integrates the features and changes the number of channels. This structure greatly reduces the number of model parameters and the computational amount, significantly reduces the memory occupation and operation time while maintaining a comparable classification accuracy, and greatly improves the inference speed.
[0033] (3)Model pruning, as a very crucial optimization method, aims to improve the model efficiency and reduce the computational cost by removing redundant parameters and connections in the model. After the model training is completed, not all parameters are equally important for the final classification decision. For example, in a convolutional neural network, the weights of some convolutional kernels may be close to zero, and these weights contribute very little to the output of the model and belong to the redundant part. In our model, in some convolutional layers for multi-scale texture modeling, some convolutional kernels always output similar and tiny values, which are not actually helpful for distinguishing normal and enamel texture features, so the corresponding parameters can be regarded as the objects for pruning. In addition, model pruning can also avoid overfitting of the model to a certain extent. Because after removing the redundant parameters, the model complexity decreases, and the sensitivity to the noise in the training data also decreases, thus improving the model generalization ability and enabling more accurate recognition of tooth images in different scenarios.
[0034] By establishing branches and using the mask generation and image reconstruction mechanisms, the model can enhance its ability to focus on key features, thereby improving the classification performance. For those background information or interference factors that are irrelevant to the diagnosis, the mask will block them, thus guiding the model to focus its attention on the truly key features. During the reconstruction process, the model will learn more refined feature representations of the key regions, further enhancing the understanding and expression ability of the key features.
[0035] The Grad-CAM++ heatmap and the recognition and diagnosis results of dental images are output together, providing strong assistance for doctors' diagnosis work. Grad-CAM++ is a visualization technique based on deep learning models, which intuitively presents the image regions that the model focuses on when judging dental fluorosis. For example, in the region of abnormal enamel coloring, the heatmap will be prominently displayed in warm colors such as red and yellow, indicating that the model considers this region to be of high importance for judging dental fluorosis. While the heat value corresponding to the normal enamel region is relatively low, and it may appear in cold colors such as blue and green. This enables doctors to refer to the heatmap simultaneously when viewing the diagnosis results given by the model, understand the basis for the model's judgment, which greatly improves the reliability and interpretability of the diagnosis.
[0036] In the case of a small volume of the image dataset, cross-validation is a very practical method for evaluating the performance of the model. It can make more full use of the limited data and reduce the evaluation bias caused by the randomness of data division. Applying cross-validation in this model can effectively avoid overfitting and underfitting problems, ensuring that the model can maintain stable performance and good generalization ability in different datasets and scenarios.
[0037] As Figure 1 shown, it is the flowchart of the end-to-end diagnosis method for dental fluorosis based on lesion characteristics provided by the embodiment of the present invention. The method includes: Collect oral photos and annotate the oral photos.
[0038] In this step, 300 oral photos are collected in total. Among them, there are 126 normal teeth, 116 dental fluorosis teeth, and 58 suspicious teeth. 242 photos of normal teeth and dental fluorosis teeth are randomly divided into 5 non-overlapping subsets, and the cross-validation method is adopted. Another 13 photos are used as the test set. 58 suspicious oral photos are used for branch mask generation and image reconstruction tasks. Use the labelimg image annotation software to annotate the oral teeth part. Frame the target with a rectangular box. At the same time, a file in xml or txt format storing information such as the position information of the rectangular box and the target category is generated for each picture.
[0039] Construct an oral tooth detection model, and process the oral photos through the oral tooth detection model to output tooth images.
[0040] In this step, a deep learning model framework is built based on the pytorch neural network library, and oral tooth detection is implemented based on the YOLOv11n basic model framework. The model is trained using the training set. The detection efficiency of the model reaches more than 90%. It consists of three major parts: the backbone network, the neck network, and the detection head. The model inputs a face image of 640×640 pixels.
[0041] The Backbone extracts image features through downsampling of convolutional layers, feature fusion of the C3k2 module, multi-scale pooling of the SPPF module, and feature enhancement of the C2PSA module, specifically including 6 convolutional layers (Conv), 6 C3k2 modules, 1 SPPF module, and 1 C2PSA module.
[0042] The convolutional layers are mainly used for downsampling and preliminary feature extraction, gradually reducing the resolution of the input image through 3x3 convolutions with a stride of 2 to generate multi-scale feature maps.
[0043] The C3k2 module is the core component of the Backbone, designed based on the CSP (Cross Stage Partial) structure. By splitting the input feature map into two parts, processing them separately and then merging, it significantly reduces the computational amount and enhances the feature fusion ability. Specifically, the implementation of the C3k2 module uses two Bottleneck modules and two C3k modules. The SPPF module realizes the fusion of multi-scale features through spatial pyramid pooling operations, capable of capturing context information at different scales. The C2PSA module may incorporate an attention mechanism, which helps the model focus on key regions in the image.
[0044] In the Neck part, the overall structure consists of multiple upsampling (Upsample), feature concatenation (Concat), convolutional layers (Conv), and C3k2 modules, which are used to fuse multi-scale features and enhance the feature expression ability. The low-resolution feature map is upsampled, and then concatenated (Concat) with the corresponding scale feature map from the Backbone to form multi-scale feature fusion. Then, the C3k2 module is used to further process the concatenated features. Then comes downsampling again. The Neck part aggregates features of different resolutions, realizes the efficient fusion of multi-scale features, and passes them to the head, providing rich feature information for object detection.
[0045] The Head performs forward pass processing on the input feature map and realizes the prediction of generating output bounding boxes and classes.
[0046] Extract the tooth images in the detection box for subsequent prediction and classification tasks.
[0047] Build a dataset based on the tooth images and perform data augmentation and expansion on the dataset.
[0048] In this step, the dataset is preprocessed, and the dataset is augmented by data augmentation methods (such as flipping the image up, down, left, and right, cropping, perspective transformation, scale transformation, rotation, etc.). For the dataset used to train the backbone network for the MobileNet-V3 binary classification task, which includes non-dental fluorosis images and dental fluorosis images, in order to adapt the image data to the input requirements of the deep learning model, we uniformly adjust the input image size to 224×224 pixels; randomly flip the image horizontally to enhance the diversity of the data and improve the generalization ability of the model; convert the image from the PIL format to a PyTorch tensor for subsequent numerical calculations; normalize each channel of the image using the predefined mean (mean=[0.485, 0.456, 0.406]) and standard deviation (std=[0.229, 0.224, 0.225]). This process helps to accelerate the convergence of the model and improve the stability of training, which is beneficial to enhancing the generalization of the model. In addition, for the dataset of suspected dental fluorosis used to train the mask generation and image reconstruction task branch, the image is flipped horizontally and vertically with a probability of 1.0 to increase the data diversity and simulate images from different perspectives; the image is randomly rotated by 90°, 180°, and 270° to simulate the changes of the image in different directions and enhance the robustness of the model to direction changes. Through the above operations, the dataset images are amplified to 200. The size, conversion to tensor, and normalization processing are the same as those of the binary classification dataset.
[0049] Taking MobileNet-V3 as the lightweight backbone, customized module enhancement is carried out for the three major pathological features (color, texture, geometry) of dental fluorosis. MobileNet-V3 is responsible for extracting general low-level features (such as edges, texture primitives) from the input image.
[0050] Define three feature lesion modules: dynamic color basis decomposition, multi-scale texture extraction (LBP+wavelet), and edge-sensitive convolution (Sobel initialization). Dynamic color basis decomposition emphasizes three learnable basis vectors: chalky spots, enamel staining, and background suppression, and then dynamically mixes the weights to generate. This module converts the input RGB image into Lab as the input.
[0051] The multi-scale texture extraction module integrates two techniques: Local Binary Pattern (LBP) and wavelet transform. This module takes the RGB-to-grayscale input, calculates the relative grayscale relationship between each pixel in the image and its neighboring pixels, and generates an LBP feature map that reflects the local texture features of the image. For wavelet transform, the input image is first transferred from the GPU to the CPU and converted into a NumPy array. The two-dimensional Haar wavelet transform is performed using the pywt.dwt2 function to obtain the high-frequency subband coefficients cH, cV, and cD in the horizontal, vertical, and diagonal directions. Then, the absolute values of these high-frequency subband coefficients are added together, converted back to a PyTorch tensor, and transferred back to the original device to obtain the wavelet features that reflect the high-frequency texture details of the image. Finally, the original input image x is concatenated with the calculated LBP features and wavelet features along the channel dimension, thus achieving the fusion of multi-scale texture features.
[0052] The edge-sensitive convolution module defines a horizontal edge detection convolution kernel of the Sobel operator and converts it to the torch.Tensor type. The Sobel operator used here can detect horizontal edges in the image. Then, the convolution kernel is copied three times to match the case where the number of input channels is 3. The defined Sobel convolution kernel is assigned to the weights of the convolutional layer. In this way, when the model is initialized, the convolutional kernel of the convolutional layer has the characteristics of the Sobel operator and can perform edge detection on the input image.
[0053] During the forward propagation of the model, first, the input features of the above three modules are concatenated, and the fused features are input into the backbone network for model training. Then, these feature maps are input into the classifier, and finally, the classification results are output.
[0054] Construct a dental fluorosis recognition model based on the MobileNet-V3 framework, perform pruning on it, and integrate the model to obtain an integrated model. Use the integrated model to perform dental fluorosis recognition on the image to be recognized.
[0055] In this step, the model is pruned. Before the pruning operation is executed, it will first check whether the model has been pruned. If not, the convolutional layer of the color branch will be pruned using the l1_unstructured method, and the weights will be trimmed according to the L1 norm. After that, global pruning will be performed on multiple modules. The weight parameters of these modules will be used as the parameter set to be pruned, and the prune.L1Unstructured method will be used to perform global unified pruning according to the specified pruning ratio. After pruning is completed, it will be marked as pruned to avoid repeated pruning.
[0056] Based on the MobileNet-V3 as the basic framework, a branch is established to generate feedback through mask generation and image reconstruction tasks to optimize the feature extraction part. Define the mask generation branch, which consists of two convolutional layers, and finally uses the Sigmoid function to output a mask with a range of [0, 1]. Define the image reconstruction branch, which uses convolutional layers and transposed convolutional layers to gradually upsample the feature map to the size of the original image, and finally uses the Sigmoid function to limit the output within the range of [0, 1]. The input image extracts features through the backbone network. The features are input into the mask generation branch to generate a mask and multiply it with the features to obtain the masked features. The masked features are input into the image reconstruction branch to reconstruct the original image. Finally, the features are mapped to the output space of binary classification through a fully connected layer.
[0057] Integrate the model to build an end-to-end complete architecture. Connect the input layer, YOLO segmentation module, and MobileNet-V3 classification module in sequence. When implementing the forward propagation of the model, data flows from the unified input layer into the YOLO segmentation module, and the segmented image is obtained after the segmentation operation. Subsequently, the segmented image is passed to the MobileNet-V3 classification module for classification, and finally the classification result is output.
[0058] In the embodiment of the present invention, the Adam optimizer is used, and the StepLR learning rate scheduler is used. The learning rate decays continuously as the training progresses, which can make the model converge more stably in the later stage of training. The specific scheduling strategy is that the initial learning rate is set to 0.001, and every 3 epochs, the learning rate is halved. For the classification task, calculate the cross-entropy loss between the model output and the true label. For the reconstruction task, calculate the mean squared error loss between the model output and the true value. The loss function is a function used to estimate the degree of difference between the predicted value and the actual value. It is a non-negative value function. When the value of the loss function is smaller, it means better robustness. Introduce the L2 regularization term to prevent the model from overfitting after pruning. Batchsize = 32, and the early stopping method is adopted. The data is trained iteratively. As the number of training times increases, the Loss value of the training dataset is continuously decreasing. As the range of the Loss value decrease stabilizes, the model training is successful.
[0059] In terms of model performance evaluation, the k-fold cross-validation method is adopted. The specific operation is to randomly divide the original dataset into 5 non-overlapping subsets. Each time, 4 of these subsets are selected as the training set, and the remaining 1 subset is used as the validation set. The model is trained on the training set and then evaluated on the validation set, and the evaluation metrics are recorded. This process is repeated 4 times, with a different subset used as the validation set each time. Finally, the results of the 4 evaluations are averaged to obtain the final performance evaluation of the model. The model performance evaluation metrics are as follows: AUC evaluates the classification ability of the model at different classification thresholds. By calculating the area under the curve of the true positive rate and the false positive rate, it measures the discrimination ability of the model between positive and negative examples; accuracy is used to calculate the proportion of samples with correct predictions in the total samples, intuitively reflecting the overall prediction accuracy of the model; precision focuses on evaluating the proportion of samples predicted as positive examples that are actually positive, reflecting the precision of the model's prediction of positive examples; recall measures the proportion of samples that are actually positive and are correctly predicted as positive, showing the coverage ability of the model for positive examples; the F1 value comprehensively considers precision and recall. Through the harmonic mean of the two, it comprehensively evaluates the performance of the model in predicting positive examples, and systematically examines the advantages and disadvantages of the model from multiple dimensions.
[0060] The above-described embodiments merely represent several implementation manners of the present invention. Their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.
[0061] The above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An end-to-end diagnosis method for dental fluorosis based on lesion characteristics, characterized in that The method includes: Collect oral cavity photos and annotate the oral cavity photos; Construct an oral cavity tooth detection model, process the oral cavity photos through the oral cavity tooth detection model, and output tooth images; Construct a dataset based on the tooth images, and perform data augmentation and expansion on the dataset; Construct a dental fluorosis recognition model with MobileNet-V3 as the basic framework, perform pruning on it, perform integration processing on the model to obtain an integrated model, and perform dental fluorosis recognition processing on the image to be recognized through the integrated model.
2. The end-to-end diagnosis method for dental fluorosis based on lesion characteristics according to claim 1, wherein In the step of collecting oral cavity photos and annotating the oral cavity photos, collect a preset number of oral cavity photos, where the oral cavity photos include normal tooth photos, dental fluorosis photos, and suspicious tooth photos, and their ratio is a preset value. Extract the normal tooth photos and dental fluorosis photos among them, construct multiple non-overlapping subsets, and set a test set.
3. The end-to-end diagnosis method for dental fluorosis based on lesion characteristics according to claim 1, wherein In the step of annotating the oral cavity photos, use the labelimg image annotation software to annotate the oral cavity tooth part, and generate a file for each picture storing the position information of the rectangular frame and the target category information.
4. The end-to-end diagnosis method for dental fluorosis based on lesion characteristics according to claim 1, characterized in that, In the step of constructing the oral cavity tooth detection model, build a deep learning model framework based on the pytorch neural network library, implement oral cavity tooth detection based on the YOLOv11n basic model framework, and perform training. The oral cavity tooth detection model consists of a backbone network Backbone, a neck network Neck, and a detection head Head. The backbone network Backbone extracts features from the image through downsampling of convolutional layers, feature fusion of C3k2 modules, multi-scale pooling of SPPF modules, and feature enhancement of C2PSA modules, specifically including 6 convolutional layers Conv, 6 C3k2 modules, 1 SPPF module, and 1 C2PSA module.
5. The end-to-end diagnosis method for dental fluorosis based on lesion characteristics according to claim 4, wherein The C3k2 module is the core component of the Backbone, designed based on the CSP structure, and merges the input feature map after processing the two parts separately. The C3k2 module contains two Bottleneck modules and two C3k modules. The SPPF module realizes the fusion of multi-scale features through spatial pyramid pooling operations.
6. The method for end-to-end diagnosis of dental fluorosis based on lesion characteristics according to claim 4, characterized in that The neck network Neck consists of multiple upsampling Upsample, feature concatenation Concat, convolutional layers Conv, and C3k2 modules, and is used to fuse multi-scale features and enhance the feature expression ability. The low-resolution feature map is upsampled through upsampling and concatenated with the corresponding scale feature map from the backbone network Backbone to form multi-scale feature fusion.
7. The end-to-end diagnosis method for dental fluorosis based on lesion characteristics according to claim 1, characterized in that, In the step of constructing a dataset based on the tooth images and performing data augmentation and expansion on the dataset, unify the images to a preset size, perform random horizontal flipping on the images, convert the images from the PIL format to PyTorch tensors, standardize each channel of the images, normalize them using predefined mean and standard deviation values, simulate images from different perspectives by adopting horizontal flipping and vertical flipping, simulate the changes of images in different directions by adopting random rotation at different angles, and expand the number of images to a specified number.
8. The end-to-end diagnosis method for dental fluorosis based on lesion characteristics according to claim 1, wherein In the steps of constructing a dental fluorosis recognition model based on the MobileNet-V3 basic framework, three feature lesion modules, namely dynamic color basis decomposition, multi-scale texture extraction, and edge-sensitive convolution, are defined. The dynamic color basis decomposition feature lesion module contains three learnable basis vectors: chalky spot, enamel pigmentation, and background suppression, and is used to generate dynamic mixing weights. The multi-scale texture extraction feature lesion module generates an LBP feature map reflecting the local texture features of the image by calculating the relative gray-scale relationship between each pixel in the image and its neighboring pixels. The edge-sensitive convolution module defines a horizontal edge detection convolution kernel of the Sobel operator and converts it into the torch.Tensor type. During the forward propagation process of the model, the input features of the three feature lesion modules are concatenated and fused, then imported into the backbone network for simulation training, and these feature maps are input into the classifier to finally output the classification result.
9. The end-to-end diagnosis method for dental fluorosis based on lesion characteristics according to claim 1, wherein In the steps of constructing a dental fluorosis recognition model based on the MobileNet-V3 basic framework, a branch is established to provide feedback through mask generation and image reconstruction tasks. The mask generation branch is defined, which consists of two convolutional layers and uses the Sigmoid function to output a mask with a range of [0, 1]. The image reconstruction branch is defined, which gradually upsamples the feature map to the size of the original image using convolutional layers and deconvolutional layers, and uses the Sigmoid function to limit the output within the range of [0, 1].
10. The end-to-end diagnosis method for dental fluorosis based on lesion characteristics according to claim 1, wherein, In the integrated model, the Adam optimizer is adopted, and the StepLR learning rate scheduler is used. The L2 regularization term is introduced to prevent the model from overfitting after pruning. In terms of model performance evaluation, the k-fold cross-validation method is adopted.
Citation Information
Patent Citations
Method for detecting and diagnosing coal-fired pollution type endemic fluorosis
CN111141897A
Micro-nucleus omics image detection method based on deep learning and image processing algorithm
CN113658174A
Oral cavity panoramic film permanent tooth segmentation and tooth position identification method based on deep learning
CN115240227A
Deep learning-based dental fluorosis image processing and identification method
CN117934926A
Dental diagnostics workflow and interface
WO2025072374A1