A tooth instance segmentation method based on nonlinear features

By constructing and training a dental instance segmentation network containing nonlinear feature extraction and weighting modules, the problem of insufficient efficiency and accuracy of dental image recognition is solved, and fast and accurate dental image recognition and diagnostic support is achieved.

CN119649035BActive Publication Date: 2025-05-23CHINA JILIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510170370.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-23
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

The prior art has shortcomings in the rapid and accurate recognition of dental images, especially in the case of insufficient supply of oral medical resources, and a method that can effectively improve the efficiency and accuracy of dental image recognition is needed.

Method used

A tooth instance segmentation method based on nonlinear features is adopted to build and train a tooth instance segmentation network by acquiring original tooth images, performing quality evaluation and denoising processing. The network includes an encoder module, a reverse mechanism module, a dense expansion module, a decoder module, a splicing module, an output module and a receptive field feature convolution kernel attention mixing module. These modules are used to extract and weight features, and finally output the tooth position represents the segmentation map.

Benefits of technology

It realizes fast and accurate recognition of dental images, improves the efficiency and accuracy of dental auxiliary diagnosis, ensures the accuracy of segmentation results, and provides reliable diagnostic support when oral medical resources are limited.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119649035B_ABST
    Figure CN119649035B_ABST
Patent Text Reader

Abstract

The invention discloses a tooth instance segmentation method based on nonlinear features, belonging to the field of image processing and deep learning technology. Firstly, oral X-rays of patients are collected and image quality is evaluated, unqualified images are denoised, mask images are made according to the FDI tooth position representation method, and a tooth instance segmentation network with nonlinear features is constructed and trained to perform instance segmentation of tooth images. The backbone network consists of an encoder module, a reverse mechanism module, a dense expansion module, a decoder module, a splicing module, an output module, and a receptive field feature convolution kernel attention hybrid module. The invention deeply mines deep feature information of teeth to generate more accurate instance segmentation results, overcomes the limitations of traditional methods when faced with complex tooth images and a wide variety of tooth types, and performs index evaluation on the training results after the training is completed. The segmentation results are output after meeting the set benchmark indicators, thereby achieving high-precision recognition and segmentation of teeth.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing and deep learning technology, and specifically relates to a tooth instance segmentation method based on nonlinear features. Background Art

[0002] With the development of computing technology, computer-aided diagnosis technology has been widely used in modern medical diagnosis to process various types of medical images with remarkable results. Medical images in computer-aided diagnosis technology can provide a large amount of medical information and provide doctors with a large amount of real and reliable reference basis in clinical diagnosis.

[0003] The dental industry as a whole has problems such as high incidence rate, low consultation rate, low shortage ratio of dentists, uneven medical level, etc. The treatment cycle is long and the patient experience needs to be improved urgently, but oral medical resources are seriously in short supply.

[0004] Deep learning is a hot field in machine learning research. It imitates the structure and function of human neural networks and trains models through large amounts of data to achieve automated image recognition and classification. Therefore, deep learning can be applied to auxiliary dental diagnosis, effectively improving doctors' efficiency and accuracy. Summary of the invention

[0005] In order to solve the shortcomings of the prior art and achieve the purpose of rapid and accurate recognition of tooth images, the present invention adopts the following technical solutions:

[0006] A tooth instance segmentation method based on nonlinear features comprises the following steps:

[0007] Step 1, obtaining the original tooth image;

[0008] Step 2: quality assessment of the original tooth image to remove noise;

[0009] Step 3, construct and train a tooth instance segmentation network, predict the tooth image, obtain a tooth instance segmentation map based on the FDI (Federation Dentaire Internationale, International Dental Federation) tooth position representation, and output it for visualization; the tooth instance segmentation network includes an encoder module, a reverse mechanism module, a dense expansion module, a decoder module, a splicing module, an output module, and a receptive field feature convolution kernel attention hybrid module; the tooth instance segmentation network has a multi-layer U-shaped structure, the features extracted by the encoder module of the current layer for the tooth image are respectively transmitted to the reverse mechanism module of the current layer and the encoder module of the next layer, after the reverse mechanism module of the current layer performs a reverse operation on the extracted features, the receptive field feature convolution kernel attention hybrid module of the current layer further extracts features based on the attention mechanism, the encoder module of the next layer continues to pass the extracted features to the current layer and the next layer, until the Kan-mamba module of the next layer extracts deep features, the dense expansion module uses a dense connection method to send the deep features to multiple expansion branches respectively, and then splices the features of the branch fusion output with the output of the receptive field feature convolution kernel attention hybrid module of the previous layer as the input of the decoder module of the previous layer, until the top decoder module, the tooth position representation segmentation map is output through the output module;

[0010] Step 4: Evaluate the segmentation accuracy of the tooth position representation segmentation map and output the instance segmentation result based on the accuracy threshold.

[0011] Furthermore, in the step 2, the peak signal-to-noise ratio of the original tooth image is calculated as the image quality evaluation value, and is judged by the set evaluation threshold. When the peak signal-to-noise ratio PSNR is less than the evaluation threshold η, the original tooth image is denoised using an improved wavelet threshold and then input into the tooth instance segmentation network. When the peak signal-to-noise ratio PSNR is greater than or equal to the evaluation threshold η, the original tooth image is directly input into the tooth instance segmentation network.

[0012] Furthermore, the calculation formula of the peak signal-to-noise ratio value is as follows:

[0013]

[0014] In the formula, max represents the maximum value of the image pixel, M and N represent the length and width of the image (in pixels), W and They represent the images before and after denoising respectively. The PSNR value is proportional to the image quality.

[0015] Furthermore, the original tooth image with a peak signal-to-noise ratio value PSNR less than the evaluation threshold η is decomposed into a low-frequency wavelet signal and a high-frequency wavelet signal through wavelet, the low-frequency signal retains the effective information of the image, and the high-frequency signal preserves the noise contained in the image and some detailed information of the image. The high-frequency signal is screened by selecting a wavelet threshold and using an improved threshold function, and the screened high-frequency information and low-frequency information are inversely transformed through wavelet reconstruction to obtain a denoised image; the function expression of the improved threshold function is:

[0016]

[0017] Where f(w) represents the optimized wavelet coefficient, w represents the initial input coefficient, t represents the wavelet threshold, and γ represents the adjustable parameter, through which the denoising effect of the image can be changed.

[0018] Furthermore, the encoder module includes a group of residually connected sub-modules for extracting features, and the encoder sub-module includes a convolutional layer, a batch normalization layer (BN layer), and a ReLU activation function; the decoder module includes a group of spliced ​​sub-modules, and the decoder sub-module includes a convolutional layer, a batch normalization layer (BN layer), and a ReLU activation function.

[0019] Furthermore, the reverse mechanism module includes a sigmoid function, an inverter and a multiplier. The output features of the encoder module are subjected to the sigmoid function and then reversely transformed through the inverter, and the obtained results are then multiplied element by element with the input features of the encoder module through the multiplier.

[0020] Furthermore, the dense expansion module sends the input features to multiple expansion convolution blocks in a densely connected manner to form multiple independent branches. After the expansion convolution processing, the dense independent branches are re-fused into one branch and output.

[0021] Furthermore, the receptive field feature convolution kernel attention hybrid module inputs the acquired features into parallel channel attention paths and spatial attention paths. The channel attention path captures higher-dimensional feature relationships through the finer kernel operations of the Kan layer to obtain a channel attention map; the spatial attention path obtains a converted feature map through a group convolution layer, normalization operation and ReLU activation operation, and tensor view conversion operation, and then performs average pooling and maximum pooling, Cov3×3 convolution and Sigmoid activation function on the converted features to obtain a spatial attention map; the spatial attention map, channel attention map and tensor view converted feature map are redistributed with spatial channel weights, and finally, the channel and spatial attention are multiplied and re-weighted to the input feature map to form a weighted feature map, and the weighted feature map is output through a convolution layer.

[0022] Furthermore, in the Kan-mamba module, the input features are sequentially passed through the Patch embedding block, the KAN block, the activation layer, the SSM state space model, the channel attention module, and the spatial attention module. Finally, the adder adds the input features, the features output by the activation layer, and the features output by the spatial attention module and outputs them.

[0023] Furthermore, the training of the tooth instance segmentation network adopts a batch strategy, and the loss function formula is as follows:

[0024] L total =L BA-Dice +L CrossEntropy

[0025]

[0026] Where, L BA-Dice represents the boundary-aware BA-Dice loss function, L CrossEntropy represents the cross entropy loss function, N is the number of batches, p i represents the predicted probability value, g i Represents the true value of each pixel i, weight W b It is used to determine the priority of boundary pixels and improve the accuracy of predicted boundaries. λ represents a set constant.

[0027] Input the prepared training data set into the network for training. If L total If the value continues to decrease, continue training until the final model is obtained after m iterations. total If the value tends to be relatively stable halfway, the iteration is stopped to obtain the final model.

[0028] The advantages and beneficial effects of the present invention are:

[0029] The present invention collects X-rays through oral X-ray equipment, calculates the peak signal-to-noise ratio of the collected X-rays; if the value of the peak signal-to-noise ratio is greater than or equal to the set threshold, the original image is directly input into the tooth instance segmentation network for the next step of processing; if the value of the peak signal-to-noise ratio is less than the set threshold, an improved wavelet threshold denoising algorithm is used to preprocess the X-ray image, and then a tooth mask label is made according to the FDI tooth position representation method, and then the X-ray and the made tooth mask image are input into the tooth instance segmentation network for the next step of processing. The tooth instance segmentation network is embedded with a receptive field feature convolution kernel attention hybrid module, which is used to sense the The tooth information is extracted by convolution of the receptive field features, and KAN is used to assign different weights to different channel features. Finally, the receptive field features and the kernel attention module are integrated to achieve multi-dimensional weighting of the features. A pooling layer is added before upsampling to strengthen supervised learning and improve the ability to learn features. A dense expansion module is embedded before the first decoder to enhance the extraction of global contextual information. A reverse mechanism module is embedded in the jump connection of each encoder to solve the problem of gradient vanishing or saturation caused by excessive forward superposition during training. After the segmentation result is output, it is evaluated and the final result is output after the evaluation is qualified to ensure the accuracy of the segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is a flowchart of a tooth instance segmentation method based on nonlinear features according to an embodiment of the present invention.

[0031] Figure 2 Schematic diagram of the dental X-ray image preprocessing process in an embodiment of the present invention.

[0032] Figure 3 Schematic diagram of the structure of the tooth instance segmentation network in the embodiment of the present invention.

[0033] Figure 4 It is a schematic diagram of the structure of the submodule in an embodiment of the present invention.

[0034] Figure 5 It is a schematic diagram of the connection structure between two reverse mechanism modules in an embodiment of the present invention.

[0035] Figure 6 It is a schematic diagram of the structure of a dense expansion module in an embodiment of the present invention.

[0036] Figure 7 It is a structural diagram of the receptive field feature convolution kernel attention hybrid module in an embodiment of the present invention.

[0037] Figure 8 Schematic diagram of the structure of the KAN layer in an embodiment of the present invention.

[0038] Fig. 9 Schematic diagram of the structure of the KAN-Mamba module in an embodiment of the present invention.

[0039] Fig.10 It is a visualization comparison diagram of instance segmentation in an embodiment of the present invention. DETAILED DESCRIPTION

[0040] The specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the present invention, and is not used to limit the present invention.

[0041] like Figure 1 As shown, a tooth instance segmentation method based on nonlinear features includes the following steps:

[0042] Step 1: Collect the patient's oral X-ray;

[0043] Step 2: Set the image quality assessment threshold and judge the obtained image quality assessment value. When the peak signal-to-noise ratio value PSNR is less than the threshold value η, use an improved wavelet threshold denoising algorithm to preprocess the original image and then input it into the tooth instance segmentation network. When the peak signal-to-noise ratio value PSNR is greater than or equal to the threshold value η, directly input the original image into the tooth instance segmentation network for the next step of processing.

[0044] Collect the patient's oral X-rays and use the peak signal-to-noise ratio (PSNR) method to evaluate the quality of the original image. The calculation formula is as follows:

[0045]

[0046] In the formula, max represents the maximum value of the image pixel, M and N represent the length and width of the image (in pixels), W and They represent the images before and after denoising respectively. The PSNR value is proportional to the image quality.

[0047] like Figure 2 As shown in the figure, the original X-ray film (including the noisy image) is decomposed by wavelet transform method, and it is decomposed into two wavelet signals, one of which is a low-frequency wavelet signal and the other is a high-frequency wavelet signal. The two signals contain different information. The low-frequency signal retains the effective information of the image, and the high-frequency signal preserves the noise contained in the image and some detailed information of the image. The appropriate threshold is selected and the improved threshold function is used to filter the high-frequency signal. The filtered high-frequency information and low-frequency information are inversely transformed through wavelet reconstruction to obtain the denoised image; the function expression of the improved threshold function is:

[0048]

[0049] In the formula, f(w) represents the optimized wavelet coefficient, w represents the initial input coefficient, t represents the threshold, and γ represents an adjustable parameter, by adjusting its size, the denoising effect of the image can be changed.

[0050] Step three, construct and train a tooth instance segmentation network, predict the preprocessed image, obtain a tooth instance segmentation map based on the FDI (Federation Dentaire Internationale, International Dental Federation) tooth position representation, and output it for visualization.

[0051] The construction and training process of the tooth instance segmentation network includes the following steps:

[0052] Step 3.1, data set preparation: annotate 1500 dental X-rays in the public data set Trends with dental instances, and divide the annotated data set into training data and test data in a ratio of 7:3. The specific example annotation method is to use the LabelImg tool to annotate all data using the FDI tooth position representation method, and divide the patient's dental panorama into four quadrants, namely upper right, upper left, lower left, and lower right (left and right are from the patient's perspective). Each tooth of the patient is represented by two Arabic numerals. The first digit represents the quadrant in which the tooth is located. In permanent teeth, each quadrant is represented as 1, 2, 3, and 4 respectively; the second digit represents the position of the tooth. The central incisor, lateral incisor, canine, first bicuspid, second bicuspid, first molar, second molar, and wisdom tooth are represented as 1, 2, 3, 4, 5, 6, 7, and 8 respectively. Therefore, the central incisor, lateral incisor, canine, first bicuspid, second bicuspid, first molar, second molar, and wisdom tooth in the upper right quadrant of the patient are labeled as 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 7 7, 18. The central incisor, lateral incisor, canine, first bicuspid, second bicuspid, first molar, second molar, and wisdom tooth in the upper left quadrant are marked as 21, 22, 23, 24, 25, 26, 27, and 28 respectively. The central incisor, lateral incisor, canine, first bicuspid, second bicuspid, first molar, second molar, and wisdom tooth in the lower left quadrant are marked as 31, 32, 33, 34, 35, 36, 37, and 38 respectively. The central incisor, lateral incisor, canine, first bicuspid, second bicuspid, first molar, second molar, and wisdom tooth in the lower right quadrant are marked as 41, 42, 43, 44, 45, 46, 47, and 48 respectively. If the second molar or wisdom tooth is missing in any quadrant, additional marking will be made; if other types of teeth except the second molar and wisdom tooth are missing in any quadrant, no additional marking will be made.

[0053] Step 3.2, construction of tooth instance segmentation network structure: design an end-to-end nonlinear feature-based instance segmentation network, such as Figure 3As shown in the figure, the network consists of 5 encoder modules, 4 reverse mechanism modules, 1 dense expansion module, 4 decoder modules, 4 splicing modules, 1 output module, and 4 receptive field feature convolution kernel attention hybrid modules.

[0054] The encoder module consists of four identical submodules with residual connections, which are used to extract features. Figure 4 As shown in the figure, the submodule contains 1 3×3 convolution layer, 1 batch normalization layer (BN layer), and 1 ReLU activation function. The specific calculation formula of the encoder module is as follows:

[0055]

[0056] In the formula, I represents the input feature, represents the output feature, and F(·) represents the feature extraction processing operation of the submodule.

[0057] like Figure 5 As shown, the reverse mechanism module consists of a sigmoid function, an inverter and a multiplier, and the calculation formula is shown as follows:

[0058]

[0059] W i =τ(S i )

[0060] R i =I⊙W i ,(i=1,2,3…)

[0061] In the formula, ⊙ represents the element-by-element multiplication symbol, i.e., the multiplier, τ(·) represents the reverse operator, i.e., the inverter, which subtracts the input from the all-1 matrix E, and W i Represents the weight of the reverse mechanism.

[0062] The dense dilation module consists of three dilated convolution blocks with different dilation rates, such as Figure 6 As shown in the figure, the expansion rates of the three dilated convolution blocks are 1, 2, and 3 respectively. The input features are deep features extracted by KAN-mamba. After the features are input into the dense convolution block, they are densely connected and sent to multiple dilated convolution blocks to form multiple independent branches. After the dilated convolution processing, these dense independent branches are re-fused into one branch, and finally the features are output. The output mathematical expression of each layer of dilated convolution is as follows:

[0063]

[0064] In the formula, K represents the convolution kernel, S l represents the dilation rate of the lth layer, represents the dilated convolution operation, [...] represents the concatenation operation, [yl-1 ,y l-2 ,…,y 0 ] represents the feature map generated by concatenating all the outputs of the previous layer and the input feature map.

[0065] like Figure 7 As shown in Figure 1, the receptive field feature convolution kernel attention hybrid module consists of two parallel attention paths:

[0066] 1. Channel attention path: The more refined kernel operation of the Kan layer is used to capture higher-dimensional feature relationships. A C×1×1 channel attention map is obtained by Kan. The KAN layer structure is as follows: Figure 8 shown.

[0067] 2. Spatial attention path: The input feature map passes through the group convolution layer to generate a CK2×H×W feature map, which is then processed by normalization and ReLU activation operations, and then converted to C×KH×KW through the tensor view conversion operation. The obtained feature map is then average pooled and max pooled to obtain a 2×KH×KW feature map, which is then activated by Cov3×3 and Sigmoid to obtain a 1×KH×KW spatial attention map.

[0068] 3. Reassign spatial channel weights to the spatial attention map, channel attention map, and feature map converted from tensor view. Finally, multiply the channel and spatial attention, re-weight them to the input feature map to form a weighted feature map, and output the weighted feature map through a convolution layer. The specific formula is as follows:

[0069]

[0070] A rf =Softmax(g 1×1 (AvgPool(X)))

[0071] F rf =ReLU(Norm(g k×k (X)))

[0072] F=(C att ·A rf )×F rf

[0073] F out =Conv 3×3,stride=K (F)

[0074] In the formula, X represents the input feature map; Φ k (X) represents the nonlinear feature transformation operation of the kth layer. The ° in the upper left corner represents the compound operation, that is, passing the input X to the function Φ 0 , whose output is then passed as input to Φ1 , and so on, until the Kth layer Φ K-1 ; φ k,q,p represents the (q,p)th sigmoid activation function of the kth layer; W k,q,p Represented as the weight matrix of the kth layer, responsible for controlling the transformation between input and output; X p Represented as a feature Figure X The pth channel of k,q,p represents the bias term of the kth layer, p represents the channel index of the input feature map, q represents the channel index of the output feature map, and n in Indicates the number of channels of the input feature map, n out Indicates the number of channels of the output feature map. Softmax represents the activation function, which converts a real vector into a probability distribution, in which the value of each element is between 0 and 1, and the sum of all elements is 1; C att Represented as the channel weight obtained by the KAN block; g 1×1 and g k×k Respectively represent 1×1 and k×k group convolution operations; A here rf Represented as the receptive field attention weight; F rf It is represented by the spatial features generated by RFA; F is the feature map after attention weighting; Norm represents the normalization operation and AvgPool represents the average pooling operation.

[0075] like Fig. 9 As shown in the figure, the Kan-mamba module consists of 2 Patch embedding blocks, 2 KAN blocks, 2 activation layers, an SSM state space model, a spatial attention module, and a channel attention module, which mainly consists of three parallel branches. The input data first passes through the Patch embedding block and then is processed using a single KAN block. The processed features are passed through the activation function and then fed into the SSM block. After the spatial transformation, the result is processed by the channel attention module and the spatial attention layer, and finally an adder is used to add the processing results of the three branches and output them. The specific calculation formula is as follows:

[0076]

[0077] X 1 =PatchEmbed(X in )

[0078] X 2 =KAN(X 1 )

[0079] X'=Ψ(X 2 )

[0080] X 3=SSM(X')

[0081] ChannelAttention(X 3 )=σ(W 2 ReLU(W 1 Pooling(X 4 )))

[0082] SpatialAttention(X c )=(Conv(concat(Pooling avg (X c ),Pooling max (X c ))))

[0083] X”=SpatialAttention(ChannelAttention(X 3 ))

[0084] In the formula, X', X", X in Represents the output results of three different branches, X c ChannelAttention(X) represents the actual input 4 ); PatchEmbed represents the input feature X in The blocks are divided and embedded to provide preliminary feature representation for the subsequent KAN blocks; Ψ() represents the activation feature operation, SSM() represents the state-space model transformation, and SpatialAttention (ChannelAttention()) represents the feature extraction operation after the channel attention and then the spatial attention module.

[0085] The decoder module is composed of two identical submodules, which include a 3×3 convolutional layer, a batch normalization layer (BN layer), and a ReLU activation function. The output of the decoder module is upsampled by 3×3 and then concatenated with the output of the reverse mechanism module to obtain the input feature map of the next decoder module.

[0086] The last layer of the network is the output module, which outputs the FDI tooth position representation segmentation map. The output module contains 1 1×1 convolution layer, 1 batch normalization layer (BN layer), and 1 ReLU activation function.

[0087] Step 3.3, training of tooth instance segmentation network: batch strategy is used for training, initialization values ​​are assigned to network parameters, batch size is set to 8, maximum number of iterations of the network is set to 500, optimizer is RMSProp, impulse is 0.9, weight decay is 1×10 -8, the learning rate is set to 1×10 -5 , the loss function is a loss function that combines boundary perception-Dice and binary cross entropy. The loss function formula is as follows:

[0088] L total =L BA-Dice +L CrossEntropy

[0089]

[0090] Where, L BA-Dice represents the boundary-aware BA-Dice loss function, L CrossEntropy represents the cross entropy loss function, N is the number of batches, p i represents the predicted probability value, g i Represents the true value of each pixel i, weight W b It is used to determine the priority of boundary pixels and improve the accuracy of predicted boundaries. λ is a constant set to 1.

[0091] Input the prepared training data set into the network for training. total If the value continues to decrease, continue training until the final model is obtained after m iterations. total If the value tends to be relatively stable halfway, the iteration is stopped to obtain the final model.

[0092] like Fig.10 As shown in the figure, after the tooth instance segmentation network training is completed, the predicted image is predicted, and the json file of the predicted tooth points is output, which is then mapped and visualized with the original image.

[0093] Step 4, input the FDI tooth position representation segmentation map in step 3 into the segmentation accuracy evaluation module, evaluate the segmentation accuracy of the instance segmentation result, and output the instance segmentation result when it meets the set threshold;

[0094] The segmentation accuracy evaluation module consists of the following indicators:

[0095] Instance-level DSC:

[0096] Definition: At the instance level, DSC is used to evaluate the similarity between each instance predicted by the model and the true instance. For each instance, DSC is twice the intersection area of ​​the predicted instance and the true instance, divided by the total area of ​​the union area of ​​the predicted instance and the true instance. The formula is

[0097]

[0098] Where A and B represent the predicted instance region and the real instance region, respectively. It is often used to evaluate the performance of segmentation algorithms in identifying and segmenting specific instances (such as tumors, organs, etc.).

[0099] The image-level DSC is the average of the instance-level DSC and is used to evaluate the quality of the entire image segmentation, including the overall performance of all instances, where N is the total number of samples.

[0100] Instance-level NSD:

[0101] Definition: At the instance level, NSD is used to evaluate the surface distance difference between the predicted instance and the true instance. Specifically, it is to calculate the average distance from the predicted surface to the true surface, and these distances usually need to be normalized.

[0102] Formula: NSD is usually the average distance between two surfaces. Normalization can make it consistent for instances of different sizes. i represents the true value, g i Represents the predicted value.

[0103]

[0104] Image-level NSD:

[0105] Definition: The image-level NSD is the average of the instance-level NSD, which is used to evaluate the overall situation of the surface distance between the predicted and true regions in the entire image.

[0106] MIoU:

[0107] Definition: MIoU is calculated by adding the IoU (Intersection over Union) of each instance and then averaging all instances. IoU is the ratio of the area of ​​the intersection of the predicted area (A) and the true area (B) to the area of ​​their union.

[0108] formula:

[0109]

[0110] Applications: Used to evaluate the accuracy of single instance segmentation, especially when dealing with multiple instances.

[0111] Identification Accuracy (IA):

[0112] Definition: Recognition accuracy is often used to evaluate the accuracy in classification tasks. It measures the model's ability to correctly classify each instance. For instance segmentation tasks, IA can be understood as the ratio of correctly identified instances to the total number of instances.

[0113] formula:

[0114]

[0115] The threshold value set in the present invention is obtained through multiple experiments, as shown in the benchmark row of Table 1. When the threshold value is greater than or equal to the threshold value, an accurate instance segmentation result can be obtained.

[0116] Table 1 Instance segmentation evaluation index table

[0117]

[0118] After passing the segmentation accuracy evaluation module, the tooth instance segmentation map is output after reaching the set threshold. It can be seen from Table 1 that compared with the U-Net, SegFormer, and Maskrcnn methods in the prior art, the evaluation indicators of the present invention are higher than the benchmark indicators.

[0119] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some or all of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A tooth instance segmentation method based on nonlinear features, characterized in that The steps include: Step 1, obtaining the original tooth image; Step 2: quality assessment of the original tooth image to remove noise; Step 3: construct and train a tooth instance segmentation network to predict the tooth image and obtain a tooth instance segmentation map based on tooth position representation; The tooth instance segmentation network includes an encoder module, a reverse mechanism module, a dense expansion module, a decoder module, a splicing module, an output module, and a receptive field feature convolution kernel attention hybrid module; The tooth instance segmentation network has a multi-layer U-shaped structure. The features extracted from the tooth image by the encoder module of the current layer are respectively passed to the reverse mechanism module of the current layer and the encoder module of the next layer. After the reverse mechanism module of the current layer performs a reverse operation on the extracted features, the features are further extracted based on the attention mechanism through the receptive field feature convolution kernel attention hybrid module of the current layer. The encoder module of the next layer continues to pass the extracted features to the current layer and the next layer until the Kan-mamba module of the next layer extracts deep features. The dense expansion module uses a dense connection method to send the deep features to multiple expansion branches respectively, and then splices the features of the branch fusion output with the output of the receptive field feature convolution kernel attention hybrid module of the previous layer as the input of the decoder module of the previous layer until the top decoder module, and then outputs the tooth position representation segmentation map through the output module; Step 4: Evaluate the segmentation accuracy of the tooth position representation segmentation map and output the instance segmentation result based on the accuracy threshold.

2. The tooth instance segmentation method based on nonlinear features according to claim 1, characterized in that: In the step 2, the peak signal-to-noise ratio value of the original tooth image is calculated as the image quality evaluation value, and is judged by the set evaluation threshold. When the peak signal-to-noise ratio value is less than the evaluation threshold, the original tooth image is denoised using the wavelet threshold and then input into the tooth instance segmentation network. When the peak signal-to-noise ratio value is greater than or equal to the evaluation threshold, the original tooth image is directly input into the tooth instance segmentation network.

3. The tooth instance segmentation method based on nonlinear features according to claim 2, characterized in that: The calculation formula of the peak signal-to-noise ratio value is as follows: In the formula, max represents the maximum value of the image pixel, M and N represent the length and width of the image respectively, and W and Respectively represent the images before and after denoising.

4. The tooth instance segmentation method based on nonlinear features according to claim 2, characterized in that: The original tooth image with a peak signal-to-noise ratio value less than the evaluation threshold is decomposed into a low-frequency wavelet signal and a high-frequency wavelet signal through wavelet decomposition. The low-frequency signal retains the effective information of the image, and the high-frequency signal preserves the noise contained in the image and some detailed information of the image. The high-frequency signal is screened by selecting a wavelet threshold and using an improved threshold function. The screened high-frequency information and low-frequency information are inversely transformed through wavelet reconstruction to obtain a denoised image; the function expression of the improved threshold function is: Where f(w) represents the optimized wavelet coefficient, w represents the initial input coefficient, t represents the wavelet threshold, and γ represents the adjustable parameter, through which the denoising effect of the image can be changed.

5. The tooth instance segmentation method based on nonlinear features according to claim 1, characterized in that: The encoder module includes a group of residually connected sub-modules for extracting features, and the encoder sub-module includes a convolution layer, a batch normalization layer, and an activation function; the decoder module includes a group of spliced ​​sub-modules, and the decoder sub-module includes a convolution layer, a batch normalization layer, and an activation function.

6. The tooth instance segmentation method based on nonlinear features according to claim 1, characterized in that: The reverse mechanism module includes a sigmoid function, an inverter and a multiplier. The output features of the encoder module are subjected to the sigmoid function and then reversely transformed through the inverter. The obtained result is then multiplied element by element with the input features of the encoder module through the multiplier.

7. The tooth instance segmentation method based on nonlinear features according to claim 1, characterized in that: The dense expansion module sends the input features to multiple expansion convolution blocks in a densely connected manner to form multiple independent branches. After the expansion convolution processing, the dense independent branches are re-fused into one branch and output.

8. The tooth instance segmentation method based on nonlinear features according to claim 1, characterized in that: The receptive field feature convolution kernel attention hybrid module inputs the acquired features into the parallel channel attention path and spatial attention path. The channel attention path captures the high-dimensional feature relationship through the Kan layer to obtain the channel attention map; the spatial attention path obtains the transformed feature map through the group convolution layer, normalization operation and activation operation processing, and tensor view conversion operation, and then performs average pooling and maximum pooling, convolution and activation function on the transformed features to obtain the spatial attention map; The spatial attention map, channel attention map, and feature map converted from the tensor view are reassigned spatial channel weights. Finally, the channel and spatial attention are multiplied and re-weighted onto the input feature map to form a weighted feature map.

9. The tooth instance segmentation method based on nonlinear features according to claim 1, characterized in that: In the Kan-mamba module, the input features are sequentially passed through the Patch embedding block, the KAN block, the activation layer, the SSM state space model, the channel attention module, and the spatial attention module. Finally, the adder adds the input features, the features output by the activation layer, and the features output by the spatial attention module and outputs them.

10. The tooth instance segmentation method based on nonlinear features according to claim 1, characterized in that: The training of the tooth instance segmentation network adopts a batch strategy, and the loss function formula is as follows: L total =L BA-Dice +L CrossEntropy Where, L BA-Dice represents the boundary-aware BA-Dice loss function, L CrossEntropy represents the cross entropy loss function, N is the number of batches, p i represents the predicted probability value, g i Represents the true value of each pixel i, weight W b Used to determine the priority of boundary pixels, λ represents a set constant.

Citation Information

Patent Citations

  • Tooth occlusion fin decayed tooth segmentation method based on improved U-Net network

    CN116205925A

  • Tooth instance segmentation method and system based on self-attention and receptive field adjustment

    CN116485809A