Automatic classification auxiliary decision-making system for caries
Through the improved Mask R-CNN and YOLO-Teeth network models, combining feature fusion and loss function optimization, the rapid and accurate diagnosis of multiple oral lesions is solved, and an automatic classification and assisted decision-making system for caries is provided to provide auxiliary diagnostic support for dentists.
Patent Information
- Application Number
- CN202510358998.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-04
AI Technical Summary
Traditional oral disease diagnosis methods are prone to misdiagnosis and missed diagnosis, and existing deep learning technologies cannot meet the needs of rapid and accurate diagnosis of a variety of oral lesions, especially the identification of complex diseases such as caries, periodontitis, apical inflammation, root bifurcation lesions.
The improved Mask R-CNN model was used for tooth segmentation and YOLO-Teeth network model to identify multiple oral lesions, combining Triplet attention mechanism, BiFPN module and MPDIoU loss function to improve the positioning accuracy and recognition accuracy of the lesion area.
It realizes accurate identification and classification display of caries, impaired teeth, periarthritis apical and bifurcation lesions, and provides dentists with auxiliary diagnostic means to meet clinical diagnosis needs.
Smart Images

Figure CN120259843A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and particularly relates to an automatic classification and auxiliary decision-making system for dental caries. Background Art
[0002] In recent years, oral and dental diseases have gradually shown a trend of getting younger and more prevalent. For example, children aged five or six are troubled by dental caries, and middle-aged and elderly people have problems with tooth loss. At the same time, oral and dental diseases can also induce a variety of high-risk diseases. The traditional method for diagnosing oral diseases is to use a tongue depressor and a flashlight to examine the lesions of the teeth inside the oral cavity, and make a diagnosis based on the doctor's clinical experience for the observed results. Due to the influence of the oral environment on the lighting conditions, the observed results deviate, which is prone to misdiagnosis and missed diagnosis, thus resulting in improper treatment.
[0003] With the development of deep learning technology, its application in the medical field has become more and more extensive, especially in image analysis. However, traditional deep learning-based tooth recognition only involves the classification of tooth lesions and does not involve the positioning problem of the lesion area, which is very limited in helping doctors make rapid and accurate diagnoses. At present, most studies focus on single dental diseases, while patients usually carry multiple oral lesions at the same time, such as dental caries, periodontitis, periapical periodontitis, furcation lesions, etc. The complexity of these multiple diseases makes the existing technology unable to meet the actual needs of clinical diagnosis. Summary of the Invention
[0004] The present invention provides an automatic classification and auxiliary decision-making system for dental caries. Through the Mask R-CNN model, the shape and position of each tooth are obtained. At the same time, through the YOLO-Teeth network model, dental caries, impacted teeth, periapical periodontitis and furcation lesions are identified. It can identify multiple oral lesions, improve the recognition accuracy, and finally classify, display and store the disease images to provide an auxiliary diagnosis means for dentists and meet the actual needs of clinical diagnosis.
[0005] The present invention provides an automatic classification and auxiliary decision-making system for dental caries, including a controller, an image acquisition module, an auxiliary display module, an image annotation module, an image segmentation module, a disease recognition module, a disease extraction module, and a classification storage module. The controller is respectively connected to the image acquisition module, the image annotation module, the image segmentation module, the disease recognition module and the disease extraction module. The image acquisition module is also connected to the auxiliary display module. The disease recognition module is respectively connected to the classification storage module and the disease extraction module;
[0006] The image acquisition module is used to take oral images by using an integrated circuit board in the shape of a tongue depressor composed of a camera and a light source, and perform real-time display through the auxiliary display module to obtain the original image;
[0007] The image annotation module is used to perform tooth annotation on the original image by using the pair tool;
[0008] The image segmentation module uses an improved Mask R-CNN to perform tooth recognition and segmentation on the annotated original image, obtains the shape and position of each tooth and superimposes them on the original image as the recognition image;
[0009] The disease recognition module uses YOLO-Teeth to recognize dental caries, periapical periodontitis, furcation lesion and impacted tooth on the recognition image to obtain a disease recognition image; wherein, the disease recognition image includes the disease area and its corresponding disease category;
[0010] The disease extraction module is used to crop the disease area to obtain a disease image and associate the disease image with its corresponding disease category;
[0011] The classification and storage module is used to classify the disease images according to the disease category and classify and store the disease images of the same category and their original images.
[0012] Furthermore, in the image segmentation module, the skip connection structure in the U-net model is integrated with multi-scale attention information to improve the Mask R-CNN segmentation branch, and an improved Mask R-CNN is obtained;
[0013] The improved Mask branch uses a convolutional layer for downsampling encoding and a transposed convolutional layer for upsampling. The formula for upsampling is:
[0014]
[0015] where k is the size of the transposed convolution kernel and f is the upsampling factor, i.e., the stride. During upsampling, shallow high-resolution features of different scales are input into the transposed convolutional layer through skip connections and the SE module;
[0016] The working process of the improved Mask R-CNN is as follows:
[0017] Input a 14×14 feature map, obtain an 8×8 feature map through 3 convolutional layers, and obtain a 28×28 feature map through a transposed convolutional layer; for the input of the transposed convolutional layer, the symmetric layers of the encoder network and the decoder network provide skip connections, and the result of each convolutional operation in the encoder network is concatenated with the result of upsampling in the decoder network after passing through the SE module, and finally a binary segmentation mask is generated through a sigmoid layer.
[0018] Furthermore, the feature fusion method of the skip connection is the concatenation of feature maps in the channel dimension, and the formula is:
[0019]
[0020] Among them, W(h, w, a) and V(h, w, b) come from feature maps of different layers respectively, F(h, w, c) is the feature map after concatenation, h and w are the length and width of the feature map, and a, b, and c are all the number of channels of the feature map.
[0021] Furthermore, the attention mechanism SE module includes compression and excitation operations, specifically:
[0022] The feature X is convolved to change its number of channels from C' to C, and the feature map U is passed to the compression operation. The compression operation uses global average pooling to compress each feature channel into a real number, expanding the receptive field to the global range. The formula for the compression calculation process is:
[0023]
[0024] where, u c is the feature map obtained after convolution, c is the number of channels of U, and H×W is the spatial dimension of U; the excitation operation captures the information of the compressed real number sequence, uses two fully connected layers to increase the non-linearity of the module, first reduces the dimension through the first fully connected layer, then activates through the rectified linear unit ReLU, then increases the dimension through the second fully connected layer, and finally activates through the sigmoid activation function. The whole process is:
[0025] s = σ[W2δ(W1z)]
[0026] where, δ is the non-linear activation function ReLU, W1 and W2 are the parameters of the two fully connected layers respectively, and σ is the sigmoid function; finally, the original features are weighted, and the original features are multiplied by the channel importance coefficients obtained by the excitation operation channel by channel to obtain the features with attention information: k = 1, 2, …, C.
[0027] Furthermore, in the disease recognition module, the YOLO-Teeth model includes an input unit, a backbone network unit, a neck unit, and a detector unit;
[0028] The backbone network unit uses the CSPDarknet53 network as the feature extraction network, and adds a Triplet attention mechanism module after its spatial pyramid pooling layer to improve the attention of the model to the lesion area;
[0029] The neck unit uses a weighted BiFPN to quickly capture information at different scales and perform efficient feature fusion, thereby optimizing the model's processing ability for complex data in the image;
[0030] The detector unit adopts the MPDIoU loss function to improve the localization accuracy of the model for the lesion area in the image.
[0031] Furthermore, the Triplet attention mechanism module captures the interaction between the spatial dimensions h, w and the channel dimension c of the input tensor, and uses three branches to capture the dependencies between the (c, h), (c, w) and (h, w) dimensions of the input tensor respectively, and finally weights and averages the output features of the three branches; where,
[0032] The first branch is the spatial attention calculation branch. First, perform the Z-Pool operation on the input vector F0(c×h×w) to convert it into a 2×h×w scale feature; subsequently, after the 7×7 convolution operation f 7×7 and batch normalization BN processing, generate the spatial attention weight M s (F0) through the Sigmoid activation function, and its expression is:
[0033] M s (F0) = σ{f 7×7 [P Avgpool (F0); P Maxpool (F0)]}
[0034] where: σ is the Sigmoid activation function, and the generated spatial attention weight M s (F0) is then applied to F0;
[0035] In the second branch, capture the interaction between the channel dimension c and the spatial dimension w. First, rotate the input vector F(c×h×w) counterclockwise by 90° around the w axis to obtain the F1(h×c×w) feature; subsequently, perform the Z-Pool operation in the h dimension to convert it into a 2×c×w scale feature; then, after f 7×7 and BN processing, generate the attention weight M s (F1) through the Sigmoid activation function, and its expression is:
[0036] M s (F1) = σ{f 7×7 [P Avgpool (F1); P Maxpool (F1)]}
[0037] The generated attention weight M s (F1) is then applied to F1, and finally rotated clockwise by 90° around the w axis to be converted into a c×h×w dimension feature;
[0038] In the third branch, the channel dimension c and the spatial dimension h capture dimensional interactions. Different from the second branch, the input vector F(c×h×w) is rotated counterclockwise by 90° around the h-axis to obtain the F2(h×c×w) feature; finally, it is rotated clockwise by 90° around the h-axis to become the c×h×w scale feature, and the intermediate process is the same; the attention weight M s (F2) is generated, and its expression is:
[0039] M s (F2) = σ{f 7×7 [P Avgpool (F2); [P Maxpool (F2)]}
[0040] Finally, the output features of the three branches are added and averaged.
[0041] Furthermore, the weighted BiFPN removes nodes with only one input, and when the original input and output nodes are at the same scale, an additional path is introduced. Each bidirectional path is regarded as a feature network and repeated multiple times to achieve more advanced feature fusion;
[0042] When fusing lesion features of different scales, additional weights are introduced for each lesion input feature to let the network learn the importance of each lesion input feature. BiFPN adopts a fast normalized weighted fusion method, and its formula is:
[0043]
[0044] where O is the weighted fusion quantity; ∈ is 0.0001; w is a learnable weight; I is the input feature; a ReLU activation function is added after each w i to ensure w i ≥0. At the same time, ∈ = 0.0001 is introduced to avoid numerical instability, and the finally obtained normalized weights are between 0 and 1.
[0045] Furthermore, the MPDIoU loss function minimizes the distance between the upper left and lower right points of the predicted lesion bounding box and the ground truth lesion bounding box, including the overlapping or non-overlapping regions, the distance between the center points, and the width and height deviations; its calculation formula is:
[0046]
[0047] where d1 is the distance between the upper left point of the predicted box and the upper left point of the ground truth box; d2 is the distance between the lower right point of the predicted box and the lower right point of the ground truth box; are the coordinates of the upper left and lower right points of the predicted box respectively; are the coordinates of the upper left and lower right points of the ground truth box respectively; R MPDIoU is the intersection over union ratio; A is the ground truth box; B is the predicted box; wN and h N are the width and height of the input image of the current network layer respectively; C a is the minimum enclosed area;
[0048] The MPDIoU loss function L MPDIoU has the following expression: L MPDIoU = 1 - R MPDIoU .
[0049] Furthermore, it also includes an image transmission module and a disease display module. The image transmission module is respectively connected to the controller and the disease display module. The controller transmits the disease recognition image, the original image, and the disease image through the image transmission module. The disease display module is used to receive the disease recognition image, the original image, and the disease image, and overlays and displays the disease image at the corresponding position of the original image according to the disease recognition image.
[0050] The beneficial effects of the present invention are as follows:
[0051] In the present invention, the teeth in the original image are segmented based on the Mask R-CNN model, and the morphology and position of each tooth on the original image are accurately obtained. Furthermore, caries, impacted teeth, periapical periodontitis, and furcation lesions are identified through the YOLO-Teeth network model, and multiple oral diseases can be identified. In the YOLO-Teeth network model, the Triplet attention mechanism module is introduced to enhance the feature extraction ability of the network, the BiFPN module is used to make the fusion of deep and shallow feature layers more sufficient, and the MPDIoU loss function is used to replace the CIoU loss function to improve the disease localization accuracy of the network. Finally, the disease images are classified, displayed, and stored, providing an auxiliary diagnosis means for dentists and meeting the actual needs of clinical diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 is a schematic structural diagram of the automatic classification and auxiliary decision-making system for dental caries of the present invention.
[0053] The realization, functional characteristics, and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0055] For example Figure 1As shown, the present invention provides an automatic classification and auxiliary decision-making system for dental caries, including a controller, an image acquisition module, an auxiliary display module, an image annotation module, an image segmentation module, a disease recognition module, a disease extraction module, and a classification storage module. The controller is respectively connected to the image acquisition module, the image annotation module, the image segmentation module, the disease recognition module, and the disease extraction module. The image acquisition module is also connected to the auxiliary display module. The disease recognition module is respectively connected to the classification storage module and the disease extraction module. Among them,
[0056] The image acquisition module is used to take oral images by using a tongue depressor-shaped integrated circuit board composed of a camera and a light source, and display them in real time through the auxiliary display module to obtain the original image. The image acquisition module consists of an OV2640 camera, an LED light source, and a membrane key to form a tongue depressor-shaped integrated circuit board. The camera is installed on the top of the circuit board. When it extends into the oral cavity to collect images, the LED lamp provides light, and then the membrane key is used to control the camera to collect image data. Among them, the camera collects oral data in real time at a speed of 30fps and a clarity of 300,000 pixels.
[0057] The image annotation module is used to label teeth in the original image by using the pair tool;
[0058] The image segmentation module uses the improved Mask R-CNN to identify and segment teeth in the annotated original image, obtains the shape and position of each tooth, and superimposes them on the original image as the recognition image;
[0059] The disease recognition module uses YOLO-Teeth to identify dental caries, periapical periodontitis, furcation lesions, and impacted teeth in the recognition image to obtain the disease recognition image. Among them, the disease recognition image includes the disease area and its corresponding disease category;
[0060] The disease extraction module is used to crop the disease area to obtain the disease image and associate the disease image with its corresponding disease category;
[0061] The classification storage module is used to classify the disease images according to the disease category, and classify and store the disease images and their original images containing the same category of diseases.
[0062] In one embodiment, an image transmission module and a disease display module are further included. The image transmission module is respectively connected to the controller and the disease display module. The controller transmits the disease recognition image, the original image, and the disease image through the image transmission module. The disease display module is used to receive the disease recognition image, the original image, and the disease image, and overlap and display the disease image at the corresponding position of the original image according to the disease recognition image.
[0063] In one embodiment, in the image segmentation module, the skip connection structure in the U-net model is incorporated into the multi-scale attention information to improve the Mask R-CNN segmentation branch, and an improved Mask R-CNN is obtained;
[0064] The improved Mask branch uses a convolutional layer for downsampling encoding and a transposed convolutional layer for upsampling. The formula for upsampling is as follows:
[0065]
[0066] where k is the size of the transposed convolution kernel, and f is the upsampling factor, i.e., the stride. During upsampling, high-resolution features of different scales at shallow layers are input into the transposed convolutional layer through skip connections and the SE module.
[0067] The improved Mask branch network is as follows: Layers 1, 2, and 3 are all convolutional layers with a convolution kernel size of 3 and a stride of 1. After each convolution, a batch normalization (BN) layer and a ReLU activation function are followed. Layers 4, 5, 6, and 7 are transposed convolutional layers, where the convolution kernel sizes of layers 4, 5, and 6 are 3 and the stride is 1, and the convolution kernel size of layer 7 is 2 and the stride is 2.
[0068] The working process of the improved Mask R-CNN is as follows:
[0069] Input a 14×14 feature map, obtain an 8×8 feature map through 3 convolutional layers, and obtain a 28×28 feature map through the transposed convolutional layer; for the input of the transposed convolutional layer, skip connections are provided by symmetric layers of the encoder network and the decoder network. The result of each convolutional operation of the encoder network is concatenated with the result of the upsampling of the decoder network after passing through the SE module, and finally a binary segmentation mask is generated through the sigmoid layer. In this way, the information contained in different-scale feature maps is fully utilized, the feature utilization rate is improved, the segmentation branch can obtain richer detailed features in a larger receptive field, and the fine-grained segmentation effect of the target is improved.
[0070] Skip connection:
[0071] Skip connections were first used in the Fully Convolution Network (FCN) for semantic segmentation. Subsequently, based on skip connections, the U-net architecture for semantic segmentation of medical images was proposed. The difference between the FCN and the U-net architecture is that the FCN uses summation operations for feature fusion, while the U-net concatenates features.
[0072] The feature fusion method of the skip connection is the concatenation of feature maps in the channel dimension, and the calculation formula is:
[0073]
[0074] Among them, W(h, w, a) and V(h, w, b) are feature maps from different layers respectively, and F(h, w, c) is the feature map after concatenation. h and w are the length and width of the feature map, and a, b, and c are the number of channels of the feature map. This skip connection structure combines the features in the low-level feature map, avoiding learning directly on the high-level feature map, so that the finally obtained feature map contains both high-level features and many low-level features, realizing the fusion of features at multiple scales.
[0075] SE module:
[0076] Although skip connections better fuse context semantic information and effectively extract more tooth detail information, the uneven brightness and low contrast in low-level features still interfere with the fine-grained segmentation of teeth. By introducing the attention mechanism SE (Squeeze and Excitation) module to capture high-level semantic information, each feature channel is weighted according to the value of the feature image, the weight of important features is increased, and the weight of unimportant features is reduced, thereby improving the effect of feature extraction and the segmentation accuracy of the model. The SE module includes compression and excitation operations, specifically:
[0077] Feature X passes through a convolution to change its number of channels from C′ to C, and the feature map U is passed to the compression operation. The compression operation uses global average pooling to compress each feature channel into a real number, expanding the receptive field to the global range. The compression calculation process formula is:
[0078]
[0079] Among them, u c is the feature map obtained after convolution, c is the number of channels of U, and H×W is the spatial dimension of U; the excitation operation captures the information of the compressed real number sequence, uses two fully connected layers to increase the non-linearity of the module, first reduces the dimension through the first fully connected layer, then activates through the rectified linear unit ReLU, then increases the dimension through the second fully connected layer, and finally activates through the sigmoid activation function. The whole process is:
[0080] s = σ[W2δ(W1z)]
[0081] Wherein, δ is the non-linear activation function ReLU, W1 and W2 are the parameters of two fully connected layers respectively, and σ is the sigmoid function; finally, the original features are weighted, and the channel importance coefficients obtained by the excitation operation are multiplied by the original features channel by channel to obtain the features with attention information: k = 1, 2, …, C.
[0082] In one embodiment, in the disease recognition module, the YOLO-Teeth model includes an input unit, a backbone network unit, a neck unit, and a detector unit;
[0083] Backbone network unit:
[0084] The backbone network unit uses the CSPDarknet53 network as the feature extraction network, and adds a Triplet attention mechanism module after its spatial pyramid pooling layer to improve the model's attention to the lesion area.
[0085] Among them, the Triplet attention mechanism can focus on the key information required for the current task, filter out the content related to the disease from numerous input information, reduce the attention to other information, and filter out irrelevant information. Assist the feature extraction network to more effectively focus on the areas related to the disease in dental radiographs when processing multiple targets, improve the relevance of lesion targets, reduce redundant information, and suppress background interference, thereby improving the disease recognition performance.
[0086] The Triplet attention mechanism module captures the interaction between the spatial dimensions h, w and the channel dimension c of the input tensor. The three branches are respectively used to capture the dependencies between the (c, h), (c, w) and (h, w) dimensions of the input tensor. The Triplet attention mechanism emphasizes the importance of calculating attention weights to provide rich feature representations, while capturing important information of cross-dimensional interactions, and finally weighted averaging the outputs of the three branches.
[0087] The first branch is the spatial attention calculation branch. First, perform the Z-Pool operation on the input vector F0(c×h×w) to convert it into a 2×h×w scale feature; then, after the 7×7 convolution operation f 7×7 and batch normalization BN processing, generate the spatial attention weight M s (F0) through the Sigmoid activation function. Its expression is:
[0088] M s (F0) = σ{f 7×7 [P Avgpool (F0); P Maxpool (F0)]}
[0089] where: σ is the Sigmoid activation function, and the generated spatial attention weight M s (F0) is then applied to F0;
[0090] In the second branch, the interaction between the channel dimension c and the spatial dimension w is captured. First, the input vector F(c×h×w) is rotated counterclockwise by 90° around the w-axis to obtain the F1(h×c×w) feature; subsequently, a Z-Pool operation is performed in the h dimension to convert it into a 2×c×w scale feature; then, after f 7×7 and BN processing, the attention weight M s (F1) is generated through the Sigmoid activation function, and its expression is:
[0091] M s (F1) = σ{f 7×7 [P Avgpool (F1); P Maxpool (F1)]}
[0092] The generated attention weight M s (F1) is then applied to F1, and finally rotated clockwise by 90° around the w-axis to be converted into a c×h×w dimension feature;
[0093] In the third branch, the interaction between the channel dimension c and the spatial dimension h is captured. The difference from the second branch is that the input vector F(c×h×w) is rotated counterclockwise by 90° around the h-axis to obtain the F2(h×c×w) feature; finally, it is rotated clockwise by 90° around the h-axis to become a c×h×w scale feature, and the intermediate process is the same; the attention weight M s (F2) is generated, and its expression is:
[0094] M s (F2) = σ{f 7×7 [P Avgpool (F2); P Maxpool (F2)]}
[0095] Finally, the output features of the three branches are added and averaged.
[0096] Z-Pool operation: By connecting the average pooling and the max pooling on this dimension, the Z-Pool layer can reduce the 0th dimension of the tensor to 2 dimensions. It helps to retain the rich features of the tensor and reduce the depth, facilitating subsequent calculations. The max pooling and average pooling operations are performed in the Z-Pool layer, and its formula is:
[0097] P z-Pool = [P Avgpool (x); P Maxpool (x)]
[0098]
[0099] In the formula: x is the input tensor; P Avgpool is the average pooling operation; P Maxpool is the max pooling operation; W is the width of the feature map; H is the height of the feature map; i, j are the number of input features.
[0100] Neck unit:
[0101] The neck unit adopts a weighted BiFPN to quickly capture information at different scales and perform efficient feature fusion, thereby optimizing the model's processing ability for complex data in images.
[0102] Among them, the weighted BiFPN improves the feature fusion mechanism and enhances the accuracy of disease recognition in the network. In fusing lesion features at different scales, BiFPN adopts a refined strategy. Instead of simply adding or concatenating features, it performs weight processing on input features at different scales. First, nodes with only one input are removed because they have no feature fusion process and contribute less to the entire network. Second, when the original input and output nodes are at the same scale, an additional path is introduced, which can better fuse lesion features without significantly increasing the computational cost. Finally, each bidirectional path is regarded as a feature network and repeated multiple times to achieve more advanced feature fusion.
[0103] When fusing lesion features at different scales, an additional weight is introduced for each lesion input feature to let the network learn the importance of each lesion input feature. BiFPN adopts a fast normalized weighted fusion method, and its formula is:
[0104]
[0105] where, O is the weighted fusion quantity; ∈ is 0.0001; w is a learnable weight; I is the input feature; after each w i add a ReLU activation function to ensure that w i ≥0. At the same time, ∈ = 0.0001 is introduced to avoid numerical instability, and the finally obtained normalized weight value is between 0 and 1.
[0106] Detector unit:
[0107] The detector unit adopts the MPDIoU loss function to improve the localization accuracy of the lesion area in the image by the model.
[0108] Among them, the MPDIoU loss function minimizes the distance between the upper left and lower right points of the predicted lesion bounding box and the ground truth lesion bounding box, comprehensively considering all relevant factors in the existing loss functions, including overlapping or non-overlapping regions, center point distance, width and height deviations; its calculation formula is:
[0109]
[0110]
[0111] Among them, d1 is the distance between the upper left point of the predicted box and the upper left point of the ground truth box; d2 is the distance between the lower right point of the predicted box and the lower right point of the ground truth box; are the coordinates of the upper left and lower right points of the predicted box respectively; are the coordinates of the upper left and lower right points of the ground truth box respectively; R MPDIoU is the intersection over union ratio; A is the ground truth box; B is the predicted box; w N and h N are the width and height of the input image of the current network layer respectively; C a is the minimum enclosing area;
[0112] The expression of the MPDIoU loss function L MPDIoU is: L MPDIoU = 1 - R MPDIoU .
[0113] In the present invention, the teeth in the original image are segmented based on the Mask R-CNN model, and the shape and position of each tooth on the original image are accurately obtained. Furthermore, caries, impacted teeth, periapical periodontitis and furcation lesions are identified through the YOLO-Teeth network model, and multiple oral lesions can be identified; in the YOLO-Teeth network model, the Triplet attention mechanism module is introduced to enhance the feature extraction ability of the network, the BiFPN module is used to make the fusion of deep and shallow feature layers more sufficient, and the MPDIoU loss function is used to replace the CIoU loss function to improve the disease localization accuracy of the network; finally, the disease images are classified, displayed and stored to provide an auxiliary diagnosis means for dentists and meet the actual needs of clinical diagnosis.
[0114] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, device, article or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, device, article or method including that element.
[0115] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.
Claims
1. An automatic classification and auxiliary decision-making system for dental caries, characterized in that, It includes a controller, an image acquisition module, an auxiliary display module, an image annotation module, an image segmentation module, a disease recognition module, a disease extraction module, and a classification storage module. The controller is respectively connected to the image acquisition module, the image annotation module, the image segmentation module, the disease recognition module, and the disease extraction module. The image acquisition module is also connected to the auxiliary display module. The disease recognition module is respectively connected to the classification storage module and the disease extraction module; The image acquisition module is used to capture oral images by using a spatula-shaped integrated circuit board composed of a camera and a light source, and perform real-time display through the auxiliary display module to obtain an original image; The image annotation module is used to perform tooth annotation on the original image by using the pair tool; The image segmentation module uses an improved Mask R-CNN to perform tooth recognition and segmentation on the annotated original image, obtains the shape and position of each tooth, and superimposes them on the original image as a recognition image; The disease recognition module uses YOLO-Teeth to recognize dental caries, periapical periodontitis, furcation lesions, and impacted teeth on the recognition image to obtain a disease recognition image. Among them, the disease recognition image includes the disease area and its corresponding disease category; The disease extraction module is used to crop the disease area to obtain a disease image and associate the disease image with its corresponding disease category; The classification storage module is used to classify the disease images according to the disease category, and classify and store the disease images of the same category and their original images; 2. The automatic caries classification and auxiliary decision-making system according to claim 1, characterized in that, In the image segmentation module, the skip connection structure in the U-net model is incorporated into the multi-scale attention information to improve the Mask R-CNN segmentation branch, resulting in an improved Mask R-CNN; The improved Mask branch uses a convolutional layer for downsampling encoding and a transposed convolutional layer for upsampling. The formula for upsampling is: where k is the size of the transposed convolution kernel, f is the upsampling factor, which is the stride. During upsampling, different-scale shallow high-resolution features are input into the transposed convolutional layer through skip connections and the SE module; The working process of the improved Mask R-CNN is as follows: Input a 14×14 feature map, obtain an 8×8 feature map through 3 convolutional layers, and obtain a 28×28 feature map through a transposed convolutional layer. For the input of the transposed convolutional layer, the symmetric layers of the encoder network and the decoder network provide skip connections. The result of each convolutional operation in the encoder network is concatenated with the result of the upsampling in the decoder network after passing through the SE module, and finally a binary segmentation mask is generated through a sigmoid layer.
3. The automatic caries classification assisted decision-making system according to claim 2, wherein The feature fusion method of the skip connection is the concatenation of feature maps in the channel dimension, and the formula is: where W(h, w, a) and V(h, w, b) are feature maps from different layers, F(h, w, c) is the feature map after concatenation, h and w are the length and width of the feature map, and a, b, and c are the number of channels of the feature map.
4. The automatic caries classification and auxiliary decision-making system according to claim 3, characterized in that, The attention mechanism SE module includes compression and excitation operations, specifically: Feature X is convolved to change its number of channels from C' to C, and the feature map U is passed to a compression operation. The compression operation uses global average pooling to compress each feature channel into a real number, expanding the receptive field to the global scope. The formula for the compression calculation process is: Among them, u c is the feature map obtained after convolution, c is the number of channels of U, and H×W is the spatial dimension of U; the excitation operation captures the information of the compressed real number sequence, and two fully connected layers are used to increase the non-linearity of the module. First, it is dimension-reduced through the first fully connected layer, then activated by the rectified linear unit ReLU, then dimension-increased through the second fully connected layer, and finally activated by the sigmoid activation function. The whole process is as follows: s = σ[W2δ(W1z)] Among them, δ is the non-linear activation function ReLU, W1 and W2 are the parameters of two fully connected layers respectively, and σ is the sigmoid function; finally, the original features are weighted by multiplying the original features channel by channel with the channel importance coefficients obtained by the excitation operation to obtain the features with attention information:
5. The automatic caries classification and auxiliary decision-making system according to claim 1, characterized in that In the disease recognition module, the YOLO-Teeth model includes an input unit, a backbone network unit, a neck unit, and a detector unit; The backbone network unit uses the CSPDarknet53 network as the feature extraction network, and a Triplet attention mechanism module is added after its spatial pyramid pooling layer to improve the model's attention to the lesion area; The neck unit uses a weighted BiFPN to quickly capture information at different scales and perform efficient feature fusion, thereby optimizing the model's processing ability for complex data in the image; The detector unit uses the MPDIoU loss function to improve the model's localization accuracy for the lesion area in the image.
6. The automatic caries classification and auxiliary decision-making system according to claim 5, characterized in that The Triplet attention mechanism module captures the interaction between the spatial dimensions h, w and the channel dimension c of the input tensor. The three branches are respectively used to capture the dependencies between the (c, h), (c, w) and (h, w) dimensions of the input tensor. Finally, the output features of the three branches are weighted and averaged; among them, The first branch is the spatial attention calculation branch. First, perform Z-Pool operation on the input vector F0 (c×h×w) to convert it into a 2×h×w scale feature; subsequently, after 7×7 convolution operation f 7×7 , batch normalization BN processing, and then generate the spatial attention weight M s (F0) through the Sigmoid activation function. Its expression is: M s (F0) = σ{f 7×7 [P Avgpool (F0); P Maxpool (F0)]} where: σ is the Sigmoid activation function, and the generated spatial attention weight M s (F0) is then applied to F0; In the second branch, the interaction between the channel dimension c and the spatial dimension w is captured. First, the input vector F(c×h×w) is rotated counterclockwise by 90° around the w-axis to obtain the F1(h×c×w) feature. Subsequently, a Z-Pool operation is performed in the h dimension to convert it into a 2×c×w scale feature. Then, after passing through f 7×7 and BN processing, the attention weight M s (F1) is generated through the Sigmoid activation function, and its expression is: M s (F1) = σ{f 7×7 [P Avgpool (F1); P Maxpool (F1)]} The generated attention weight M s (F1) is then applied to F1 and finally rotated 90° clockwise around the w-axis to be transformed into a c×h×w dimensional feature; In the third branch, the channel dimension c and the spatial dimension h capture dimensional interactions. Different from the second branch, the input vector F(c×h×w) is rotated counterclockwise by 90° around the h-axis to obtain the feature F2(h×c×w); finally, it is rotated clockwise by 90° around the h-axis to become a feature of the c×h×w scale. The intermediate process is the same; the attention weight M s (F2) is generated, and its expression is: M s (F2) = σ{f 7×7 [P Avgpool (F2); P Maxpool (F2)]} Finally, the output features of the three branches are added and averaged.
7. The automatic caries classification and auxiliary decision-making system according to claim 5, wherein The weighted BiFPN removes nodes with only one input, and when the original input and output nodes are at the same scale, an additional path is introduced. Each bidirectional path is regarded as a feature network and repeated multiple times to achieve more advanced feature fusion; When fusing lesion features of different scales, additional weights are introduced for each lesion input feature to let the network learn the importance of each lesion input feature. BiFPN adopts a fast normalized weighted fusion method, and its formula is: Among them, O is the weighted fusion amount; ∈ is 0.0001; w is a learnable weight; I is the input feature; after each w i add the ReLU activation function to ensure that w i ≥ 0. At the same time, introduce ∈ = 0.0001 to avoid numerical instability. The finally obtained normalized weights are between 0 and 1.
8. The automatic caries classification and auxiliary decision-making system according to claim 5, wherein The MPDIoU loss function minimizes the distance between the upper left and lower right points of the lesion prediction bounding box and the lesion ground truth bounding box, including overlapping or non-overlapping regions, center point distance, width and height deviations; its calculation formula is: Among them, d1 is the distance between the upper left point of the predicted box and the upper left point of the ground truth box; d2 is the distance between the lower right point of the predicted box and the lower right point of the ground truth box; are the coordinates of the upper left and lower right points of the predicted box respectively; are the coordinates of the upper left and lower right points of the ground truth box respectively; R MPDIoU is the intersection over union ratio; A is the ground truth box; B is the predicted box; w N and h N are the width and height of the input image of the current network layer respectively; C a is the minimum enclosed area; MPDIoU loss function L MPDIoU The expression of which is: L MPDIoU = 1 - R MPDIoU .
9. The automatic caries classification and auxiliary decision-making system according to claim 1, wherein It also includes an image transmission module and a disease display module. The image transmission module is respectively connected to the controller and the disease display module. The controller transmits the disease recognition image, the original image, and the disease image through the image transmission module. The disease display module is used to receive the disease recognition image, the original image, and the disease image, and overlap and display the disease image at the corresponding position of the original image according to the disease recognition image.