Oral cavity image segmentation method and system based on attention mechanism

By introducing channel attention modules and non-local modules into the oral image segmentation model, the attention mechanism is used to improve the feature capture and expression ability of the model, and the problems of low accuracy of oral CBCT images segmentation and insufficient generalization ability in the existing technology are solved, achieving higher segmentation accuracy and robustness.

CN120070460APending Publication Date: 2025-05-30BEIJING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510118334.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has problems with low accuracy and poor performance in oral CBCT imaging segmentation, and the model generalization ability is insufficient to cope with the diversity of anatomical characteristics and lesion status of different patients.

Method used

Using the segmentation method based on attention mechanism, the model's ability to capture key features of the image and express global features by introducing channel attention modules into the encoder of the segmentation model and non-local modules in the decoder.

Benefits of technology

Improves the accuracy and robustness of oral image segmentation, and enhances the adaptability and accuracy of the model when dealing with complex anatomical structures and diverse image features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070460A_ABST
    Figure CN120070460A_ABST
Patent Text Reader

Abstract

The invention discloses an oral image segmentation method and system based on an attention mechanism, and the method comprises the steps: constructing a segmentation model composed of an encoder block, a jump connection module and a decoder block, the encoder block comprises a down-sampling convolution and channel attention module, and the decoder block comprises an up-sampling convolution and non-local module; the jump connection module fuses the features extracted by the encoder block and the decoder block; when the model is trained, model parameters are adjusted based on a cross entropy loss function and a Dice loss function, and when the performance of the model is no longer improved or reaches a preset number of iterations, training is stopped, and a trained model is obtained; a to-be-segmented data set is established and input into the segmentation model, the to-be-segmented data set is subjected to encoder block convolution to obtain a deep feature map, a channel attention module processes the deep feature map to obtain a key feature map, and a decoder block non-local module processes the key feature map to obtain a global feature map; and outputting an oral cavity segmentation image based on the key feature map and the global feature map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image segmentation, and particularly relates to a multi-class segmentation method and system for oral CBCT images based on an attention mechanism. Background Art

[0002] Oral CBCT (Cone Beam Computed Tomography) images are a technology for three-dimensional imaging using cone beam X-rays, which can generate high-resolution three-dimensional images and provide detailed oral structure information. Its main features include low radiation dose, fast scanning, and high resolution, which reduce the risk of radiation exposure for patients and improve the comfort of scanning. Segmenting oral CBCT images, that is, by precisely separating different structures (such as teeth, jaws, and soft tissues), doctors can more clearly analyze the anatomical relationships inside the oral cavity, which not only improves the readability of the images but also enhances the accuracy of diagnosis and the pertinence of treatment.

[0003] Currently, significant progress has been made in medical image segmentation methods based on convolutional neural networks. Guan et al. (2022) proposed an automatic brain tumor MRI data segmentation framework, adding a channel attention (SE) module to each encoder to automatically enhance useful information and suppress useless information in the channels using channel relationships; adding an attention-guided filtering (AG) module to each decoder to guide edge information using the attention mechanism and remove the influence of irrelevant information such as noise; in addition, adding a skip connection module between the encoder and the decoder to fuse the feature maps at corresponding positions of the encoder and decoder. Chen et al. (2023) proposed a new semi-supervised multi-organ segmentation network of a teacher-student model. First, the labeled and unlabeled three-dimensional images are divided into many small cubes and randomly mixed, and the prior anatomical structures of different organs are used to let the unlabeled images learn the semantic information of the organs to guide data augmentation and reduce the mismatch between the labeled and unlabeled images in semi-supervised learning; then, the pseudo-labels predicted by the teacher network are fused with the local feature representations learned by the small cubes to improve the quality of the pseudo-labels.

[0004] In recent years, deep learning technology has made remarkable progress in medical image segmentation. However, for the segmentation task of oral CBCT images, there are still the following problems and challenges:

[0005] Although CBCT images can provide high-resolution three-dimensional data, they also have some inherent defects, such as image noise, artifacts, and insufficient contrast. These defects will significantly affect the performance of segmentation algorithms, especially in areas with blurred boundaries or overlapping structures. In addition, the anatomical structures inside the oral cavity are extremely complex, and the shapes and arrangements of teeth vary significantly, making the segmentation task more challenging. The diversity of anatomical features and lesion states among different patients may also lead to changes in image features, thus affecting the generalization ability of the model. For this reason, even under similar imaging conditions, accurately segmenting different oral-related structures remains a difficult task. This complexity not only increases the difficulty of segmentation but also poses higher requirements for the accuracy of clinical diagnosis.

[0006] There is still room for improvement in the segmentation accuracy of oral CBCT images. Existing algorithms perform well on specific datasets, but in clinical applications, due to the variation in image quality and the diversity of pathological features, the accuracy is often insufficient to meet clinical needs. Summary of the Invention

[0007] Aiming at the deficiencies in the existing technology, the present invention provides an oral image segmentation method and system based on an attention mechanism. By introducing a channel attention module and a non-local module into the segmentation model, the segmentation accuracy and robustness of oral images are improved, and the technical problem of low accuracy in oral image segmentation is solved.

[0008] The first object of the present invention is to provide an oral image segmentation method based on an attention mechanism, including:

[0009] Construct a segmentation model, which is a convolutional neural network composed of an encoder block, a skip connection module, and a decoder block. The encoder block includes a downsampling convolution and a channel attention module, the decoder block includes an upsampling convolution and a non-local module, and the skip connection module is used to fuse the features extracted by the encoder block and the decoder block;

[0010] Train the segmentation model with an oral image dataset containing original images and label data, continuously adjust the parameters of the segmentation model based on the cross-entropy loss function and the Dice loss function, and stop training when the performance of the segmentation model no longer improves or reaches a preset number of iterations to obtain a trained segmentation model for oral image segmentation;

[0011] Obtain oral image data and construct a dataset to be segmented based on the oral image data;

[0012] Input the dataset to be segmented into the segmentation model. After the downsampling convolution operation of the encoder block on the dataset to be segmented, a deep feature map is obtained. The channel attention module of the encoder block processes the deep feature map to obtain a key feature map. The non-local module of the decoder block processes the key feature map to obtain a global feature map. Based on the key feature map and the global feature map, an oral cavity segmentation image is output.

[0013] As a further improvement of the present invention, the channel attention module includes a global average pooling layer, a fully connected layer, and an output layer;

[0014] The channel attention module of the encoder block processes the deep feature map to obtain a key feature map, including:

[0015] Input the deep feature map into the global average pooling layer to obtain a channel feature map;

[0016] Input the channel feature map into the fully connected layer to obtain the weight of each channel;

[0017] Fuse the weight with the channel feature map to obtain a calibrated feature map, and the output layer outputs the key feature map based on the calibrated feature map.

[0018] As a further improvement of the present invention, the non-local module includes a query feature convolution layer, a key feature convolution layer, and a value feature convolution layer;

[0019] The non-local module of the decoder block processes the key feature map to obtain a global feature map, including:

[0020] Input the key feature map into the query feature convolution layer, the key feature convolution layer, and the value feature convolution layer respectively to correspondingly obtain query feature data, key feature data, and value feature data;

[0021] Obtain a similarity matrix based on the query feature data and the key feature data;

[0022] Perform weighted summation on the value feature data based on the similarity matrix to obtain a weighted feature map, and obtain a global feature map based on the weighted feature map and the key feature map.

[0023] As a further improvement of the present invention, obtaining the similarity matrix based on the query feature data and the key feature data includes:

[0024] Multiply the query feature data and the key feature data point by point to obtain an initial similarity matrix;

[0025] Perform normalization processing on the initial similarity matrix based on the Softmax function to obtain a similarity matrix.

[0026] As a further improvement of the present invention, obtaining the global feature map based on the weighted feature map and the key feature map includes: adding the weighted feature map and the key feature map based on residual connection to obtain the global feature map.

[0027] As a further improvement of the present invention, constructing the dataset to be segmented based on the oral image data includes:

[0028] Preprocessing the oral image data to obtain the preprocessed dataset, where the preprocessing includes normalization processing and denoising processing;

[0029] Performing data augmentation processing on the preprocessed dataset to construct the dataset to be segmented, where the data augmentation processing includes rotation, translation, and scaling.

[0030] As a further improvement of the present invention, each decoder block of the segmentation model is connected to the corresponding encoder block through a skip connection module.

[0031] The second object of the present invention is to provide an oral image segmentation system based on an attention mechanism, including:

[0032] A model construction unit for constructing a segmentation model, where the segmentation model is a convolutional neural network composed of an encoder block, a skip connection module, and a decoder block. The encoder block includes a downsampling convolution and a channel attention module, the decoder block includes an upsampling convolution and a non-local module, and the skip connection module is used to fuse the features extracted by the encoder block and the decoder block;

[0033] A model training unit for training the segmentation model with an oral image dataset containing original images and label data, continuously adjusting the parameters of the segmentation model based on the cross-entropy loss function and the Dice loss function, and stopping training when the performance of the segmentation model no longer improves or reaches a preset number of iterations to obtain a trained segmentation model for oral image segmentation;

[0034] A data acquisition unit for acquiring oral image data and constructing a dataset to be segmented based on the oral image data;

[0035] An image segmentation unit for inputting the dataset to be segmented into the segmentation model. After the downsampling convolution operation of the encoder block in the dataset to be segmented, a deep feature map is obtained. The channel attention module of the encoder block processes the deep feature map to obtain a key feature map. The non-local module of the decoder block processes the key feature map to obtain a global feature map, and an oral segmentation image is output based on the key feature map and the global feature map.

[0036] As a further improvement of the present invention, the data acquisition unit is further configured to: preprocess and perform data augmentation processing on the acquired oral image data to construct a dataset to be segmented.

[0037] As a further improvement of the present invention, the model construction unit is further configured to: when constructing the segmentation model, each decoder block is connected to the corresponding encoder block through a skip connection module.

[0038] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0039] Add a channel attention module to the encoder of the segmentation model to focus on and capture the key features of the image, so as to effectively improve the performance of the segmentation model when processing regions with blurred boundaries or overlapping structures. Furthermore, add a non-local module to the decoder of the segmentation model to utilize the context information within the global range to enhance the expression ability of the global feature map.

[0040] Based on the cross-entropy loss function and the Dice loss function, optimize and train the segmentation model, so that the trained segmentation model can effectively handle the subtle differences of different types of oral images, provide more accurate oral image segmentation results, and improve the robustness and accuracy of oral image segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is a flowchart of the oral image segmentation method;

[0042] Figure 2 is a network architecture diagram of the segmentation model;

[0043] Figure 3 is a network architecture diagram of the channel attention module;

[0044] Figure 4 is a network architecture diagram of the non-local module;

[0045] Figure 5 is a structural diagram of the oral image segmentation system. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0047] The following further describes the present invention in detail with reference to the accompanying drawings:

[0048] Please refer to Figure 1 , this embodiment provides an oral image segmentation method based on an attention mechanism, including:

[0049] S1. Construct a segmentation model, which is a convolutional neural network composed of an encoder block, a skip connection module, and a decoder block. The encoder block includes a downsampling convolution and a channel attention module. The decoder block includes an upsampling convolution and a non-local module. The skip connection module is used to fuse the features extracted by the encoder block and the decoder block.

[0050] The segmentation model can be a convolutional neural network with VNet as the backbone network for three-dimensional medical image segmentation. For the structure of the segmentation model, please refer to Figure 2 .

[0051] In this embodiment, VNet is selected as the backbone network. The network structure of VNet is mainly composed of an encoder and a decoder, and multi-layer three-dimensional convolutions and downsampling operations are used to extract features. In the encoder part, the input three-dimensional image gradually reduces the spatial dimension through a series of convolutional layers and max pooling layers, while extracting high-level features. A batch normalization and an activation function (such as the ReLU activation function) are connected behind each convolutional block to enhance the non-linear expression ability of the network. In the decoder part, VNet uses transposed convolutional layers for upsampling to gradually restore the spatial resolution of the image, and at the same time makes skip connections with the feature maps of the encoder to retain detailed information. This enables the initial segmentation model to better retain information during the training process, improving the learning ability and expression ability of the network.

[0052] The VNet convolutional neural network can make full use of the spatial characteristics of three-dimensional data and adapt to complex oral anatomical structures. Its three-dimensional convolutional structure enables it to make full use of spatial information when processing CBCT images, thereby improving the segmentation accuracy. Secondly, VNet adopts residual connections, which can alleviate the problem of gradient disappearance in the training of deep networks and enhance the training stability of the network. Especially when processing boundary-blurred or overlapping regions, it can effectively improve the robustness of segmentation. In addition, the VNet architecture has good flexibility and can be adjusted and extended according to specific needs. This characteristic enables it to adapt to different clinical application scenarios and data sets, enhancing the generality of the model. Therefore, choosing the VNet architecture as the backbone network in this application can well adapt to the characteristics of oral CBCT images and the requirements of the segmentation task.

[0053] In this embodiment, the VNet backbone network includes 5 encoder blocks, 5 decoder blocks, and a skip connection module to achieve efficient feature extraction and information reconstruction. Specifically, the VNet backbone network accepts an input with the shape of (batchsize, channels, depth, height, width). In the encoder part, the 5 encoder blocks are stacked layer by layer. Each encoder block consists of multiple convolutional layers, and 3D convolution is used to process three-dimensional data. First, the input feature map undergoes a series of convolutional operations, combined with batch normalization and the ReLU activation function, so as to extract deep features. Then, each encoder block performs downsampling through a max-pooling layer to gradually reduce the spatial dimension of the feature map, increase the receptive field, and extract more feature information.

[0054] In the decoder part, the network gradually restores the spatial resolution of the feature map through deconvolution operations. Each decoder block is connected to the corresponding encoder block through a skip connection, directly passing the high-resolution features extracted in the encoder to the decoder. This not only effectively preserves local detail information but also better integrates context information. In each decoder block, the feature map after deconvolution is concatenated with the feature map from the skip connection to enhance the feature expression ability. Subsequently, after being processed by convolutional layers, features are further extracted, and finally, an output feature map with the same number as the number of categories is obtained through the last convolutional layer.

[0055] Based on the VNet backbone network framework, the encoder is improved by introducing a channel attention module into the encoder. By introducing the channel attention module, the segmentation model can dynamically adjust the feature importance of each channel, thus paying more attention to key information more effectively and suppressing redundant or irrelevant features. In this embodiment, the channel attention module is integrated into the five encoders of the VNet backbone network respectively to enhance the expression ability and overall performance of the segmentation model.

[0056] Based on the VNet backbone network framework, the decoder is also improved by introducing a non-local module into the decoder. The 3D non-local module is added to the first two decoders of the VNet decoder to capture global context information. Different from traditional convolutional operations that only focus on local regions, the 3D non-local module considers the relationships between all positions in the input feature map, enhancing the comprehensiveness of feature representation.

[0057] S2. Use the oral imaging dataset containing the original images and label data to train the segmentation model. Based on the cross-entropy loss function and the Dice loss function, continuously adjust the parameters of the segmentation model. When the performance of the segmentation model no longer improves or reaches the preset number of iterations, stop training to obtain the trained segmentation model for oral imaging segmentation.

[0058] Specifically, during the training process, a strategy of combining the cross-entropy loss function and the Dice loss function, along with the consistency loss (weight 0.2), is adopted to evaluate the difference between the model output and the ground truth label. Meanwhile, the Adam optimizer is used, with the base learning rate set to 0.01, the maximum number of iterations set to 30000, and the batch size to 4 to ensure the efficiency of training. To improve the repeatability of training, the random seed can be set to 1337, and the consistency boost (boost period 40.0) and EMA decay rate (0.99) can be dynamically adjusted.

[0059] During the training process, the parameters of the overall network structure are jointly adjusted, including the slice size, the convolutional layer parameters in the module, the batch size, and the learning rate of the network. Through multiple experiments and optimizations, the optimal combination of each parameter is found, enabling the network to converge quickly and avoid overfitting while ensuring the segmentation accuracy. In addition, a strategy of combining the cross-entropy loss function and the Dice loss function is adopted, enabling the model to take into account the learning of local and global features during the training process, further improving the recognition ability of fine structures. This synergistic effect of multiple parameters not only improves the segmentation accuracy but also enhances the adaptability of the model, enabling it to perform well on different image qualities and different oral anatomical structures.

[0060] S3. Obtain oral image data and construct a dataset to be segmented based on the oral image data. The oral image data can be Cone beam CT (CBCT), and can include four categories of annotations: maxilla, mandible, teeth, and background.

[0061] The steps for constructing the dataset to be segmented include: normalizing and denoising the oral image data to obtain a preprocessed dataset, and improving the quality of the oral images through preprocessing;

[0062] Performing at least one of data augmentation processes such as rotation, translation, and scaling on the preprocessed dataset to construct the dataset to be segmented. Performing data augmentation can increase the diversity of training data and prevent the segmentation model from overfitting.

[0063] S4. Input the dataset to be segmented into the segmentation model. After the encoder block performs downsampling convolution operations on the dataset to be segmented, a deep feature map is obtained. The channel attention module of the encoder block processes the deep feature map to obtain a key feature map. The non-local module of the decoder block processes the key feature map to obtain a global feature map, and an oral segmentation image is output based on the key feature map and the global feature map.

[0064] The channel attention module can effectively identify and separate different structural features, such as tooth edges and alveolar bone contours, in the multi-class segmentation task of teeth and alveolar bone by adaptively adjusting the weights of each channel, enabling the target segmentation model to focus on more important features. Secondly, noise and irrelevant features often exist in CBCT images, which may affect the segmentation performance. The channel attention module can reduce the activation values of unimportant channels, helping the target segmentation model better ignore these interferences and thus improving the accuracy of the segmentation results. In addition, the introduction of the channel attention module enables the segmentation model to adaptively adjust the responses of different channels to meet the requirements of different segmentation tasks. For example, in some cases, the features of the alveolar bone may be more important, while in other cases, the features of the teeth may be more prominent, enabling the target segmentation model to flexibly handle complex oral structures.

[0065] To address the problem that traditional convolutional operations cannot effectively capture long-range relationships between different anatomical structures, a non-local module that can capture global dependencies between various positions in the feature map is added to the decoder, thereby improving the recognition ability of the segmentation model for different categories (such as different types of teeth and alveolar bone) and reducing confusion. Secondly, when dealing with similar structures, the non-local module enhances the grasp of details by integrating global information, thus improving the segmentation accuracy, especially between teeth and alveolar bone with similar shapes and boundaries. In addition, for complex anatomical structures such as teeth and alveolar bone, the local module can adaptively adjust the feature representation, enhancing the learning ability of the segmentation model for different anatomical features and enabling it to better handle multi-class segmentation tasks.

[0066] A channel attention module is added to the encoder of the segmentation model to focus on and capture the key features of the image, effectively improving the performance of the segmentation model when dealing with regions with blurred boundaries or overlapping structures. Furthermore, a non-local module is added to the decoder of the segmentation model to utilize the context information within the global scope to enhance the expressive ability of the global feature map. Based on the cross-entropy loss function and Dice loss function, the segmentation model is optimized and trained, enabling the trained segmentation model to effectively handle the subtle differences in different categories of oral images, providing more accurate oral image segmentation results and improving the robustness and accuracy of oral image segmentation.

[0067] Furthermore, the channel attention module includes a global average pooling layer, a fully connected layer, and an output layer. For the structure of the channel attention module, please refer to Figure 3 ;

[0068] The key feature map is obtained in the following way:

[0069] The deep feature map is input into the global average pooling layer to obtain the channel feature map;

[0070] Input the channel feature map into the fully connected layer to obtain the weight of each channel;

[0071] Fuse the weight with the channel feature map to obtain the calibrated feature map, and the output layer outputs the key feature map based on the calibrated feature map.

[0072] Specifically, this embodiment proposes a channel attention module to enhance the target segmentation model's ability to focus on key features. The channel attention module first receives a feature map x with a shape of (batchsize, channels, depth, height, width). Next, perform global average pooling operation on the input feature map x to compress the spatial information of each channel into a single value to represent the overall feature of each channel. Subsequently, perform fully connected layer processing on the feature map after global average pooling operation. The steps of fully connected layer processing are as follows: First, reduce the number of channels from channel to channel / / ratio to reduce the number of model parameters and computational complexity; then, introduce the ReLU non-linear activation function to enable the model to learn more complex feature representations; finally, restore the number of channels to the original channel number through the fully connected layer again to generate the weight of each channel. The output layer uses the Sigmoid function to map these weight values to the range of [0,1] to represent the importance of each channel. Finally, generate the calibrated feature map by multiplying the input feature map x by the weight y, thereby enhancing the network's response ability to important features, improving the performance of the model, effectively enhancing the network's learning ability, enabling it to better focus on key features in complex tasks, and thus improving the overall performance.

[0073] By focusing on important features, the channel attention module can effectively improve the model's performance when processing regions with blurred boundaries or overlapping structures. It not only improves the segmentation accuracy of different teeth and related structures, but also enhances the model's adaptability to anatomical feature variations of different patients, achieving more accurate and reliable segmentation results.

[0074] Furthermore, the non-local module includes a query feature convolutional layer, a key feature convolutional layer, and a value feature convolutional layer. For the structure of the non-local module, please refer to Figure 4 ;

[0075] The global feature map is obtained in the following way:

[0076] Input the key feature map into the query feature convolutional layer, the key feature convolutional layer, and the value feature convolutional layer respectively to obtain the query feature data, the key feature data, and the value feature data correspondingly;

[0077] Obtain the similarity matrix based on the query feature data and the key feature data;

[0078] Weighted sum is performed on the value feature data based on the similarity matrix to obtain a weighted feature map, and a global feature map is obtained based on the weighted feature map and the key feature map.

[0079] Specifically, this embodiment proposes a non-local module to enhance the ability of the target segmentation model to utilize global context information. The non-local module first receives a key feature map x with a shape of (batchsize, channels, depth, height, width). Next, the input key feature map x generates query (Q), key (K), and value (V) features through three independent convolutional layers, and a similarity matrix is obtained based on the query feature data and the key feature data. Subsequently, the generated similarity matrix is used to perform a weighted sum on the V feature to obtain a new feature representation. This process effectively integrates global information into local features by calculating the weighted values of each position in the global context.

[0080] Furthermore, the similarity matrix is obtained in the following manner:

[0081] The query feature data and the key feature data are multiplied pointwise to obtain an initial similarity matrix;

[0082] The initial similarity matrix is normalized based on the Softmax function to obtain the similarity matrix.

[0083] When calculating the feature similarity, the dot product method is used to calculate the similarity between Q and K to generate a similarity matrix. To ensure that the output of the similarity matrix is within the same range, the Softmax function is used to normalize the similarity matrix. The normalization process makes the sum of the elements in each row equal to 1, which can thus be interpreted as the relative importance between different features.

[0084] Furthermore, obtaining the global feature map based on the weighted feature map and the key feature map includes: adding the weighted feature map and the key feature map based on residual connection to obtain the global feature map. After passing through the convolutional layer, the weighted feature map is fused with the key feature map, and the two are added together in a residual connection manner, thereby retaining the information of the key feature map and enhancing its expression ability.

[0085] The effectiveness of the oral image segmentation method provided in this embodiment is verified through specific experiments:

[0086] Select 443 oral CBCT image data from the publicly available Tooth Fairy dataset of oral CBCT images for experiments, and use 90 subsets of it for verification. The dataset contains four categories of annotations: maxilla, mandible, teeth, and background. In the initial stage of the verification experiment, each image is preprocessed first, randomly cropped to 96×96×96 pixels, and the volume is divided into small cubes of 32×32×32. To improve the generalization ability of the target segmentation model, a 4-fold cross-validation method is adopted, and a labeled dataset is randomly selected for training.

[0087] The verification experiment is carried out on an RTX 3090 GPU, and the environment used is torch 2.0.1, CUDA 11.8, and Python 3.8.19. During the network training process, the initial learning rate is set to 0.01, and the learning rate is decayed to 0.1 every 12,000 iterations. The Stochastic Gradient Descent (SGD) optimizer is used, the batch size is set to 4, and the maximum number of iterations is 30,000, which ensures the effective training and performance evaluation of the target segmentation model.

[0088] The experimental results are shown in the following table:

[0089]

[0090]

[0091] The Dice Similarity Coefficient (DSC), 95% Hausdorff Distance (HD95), Normalized Surface Distance (NSD), and Average Surface Distance (ASD) are used as evaluation metrics in the experiment. ↑ indicates that the higher the value, the better, and ↓ indicates that the lower the value, the better. It can be seen from the experimental results that using VNet as the backbone network and adding a channel attention module and a non-local module respectively, all indicators have been improved.

[0092] By introducing a channel attention module, the importance of each channel can be automatically adjusted, thereby enhancing the selectivity for key features, especially in the recognition of tiny structures, significantly improving the segmentation accuracy. Secondly, the application of the non-local module can capture global context information, effectively understand the mutual relationship between anatomical structures, reduce misjudgment, and improve the consistency of segmentation.

[0093] In addition, the skip connections of the encoder-decoder structure achieve multi-scale feature extraction, ensuring the effective recognition ability of the target segmentation model at different resolutions. The optimized training based on the cross-entropy loss function and the Dice loss function further improves the robustness and adaptability of the model on different image qualities and oral anatomical structures.

[0094] Please refer to Figure 5, this embodiment provides an oral image segmentation system based on an attention mechanism, including:

[0095] A model construction unit for constructing a segmentation model, which is a convolutional neural network composed of an encoder block, a skip connection module, and a decoder block. The encoder block includes a downsampling convolution and a channel attention module, and the decoder block includes an upsampling convolution and a non-local module. Each decoder block is connected to the corresponding encoder block through a skip connection module, and the skip connection module is used to fuse the features extracted by the encoder block and the decoder block;

[0096] A model training unit for training the segmentation model with an oral image dataset containing original images and label data, continuously adjusting the parameters of the segmentation model based on the cross-entropy loss function and the Dice loss function, and stopping the training when the performance of the segmentation model no longer improves or reaches a preset number of iterations to obtain a trained segmentation model for oral image segmentation;

[0097] A data acquisition unit for acquiring oral image data, performing preprocessing and data augmentation on the acquired oral image data to construct a dataset to be segmented;

[0098] An image segmentation unit for inputting the dataset to be segmented into the segmentation model. After the downsampling convolution operation of the encoder block, a deep feature map is obtained. The channel attention module of the encoder block processes the deep feature map to obtain a key feature map. The non-local module of the decoder block processes the key feature map to obtain a global feature map, and an oral segmentation image is output based on the key feature map and the global feature map.

[0099] The above is only the preferred embodiment of the present invention and is not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An oral image segmentation method based on attention mechanism, characterized in that: include: Constructing a segmentation model, which is a convolutional neural network consisting of an encoder block, a skip connection module, and a decoder block, wherein the encoder block includes a downsampling convolution and a channel attention module, the decoder block includes an upsampling convolution and a non-local module, and the skip connection module is used to fuse the features extracted by the encoder block and the decoder block; The segmentation model is trained using an oral image dataset containing original images and label data, and the parameters of the segmentation model are continuously adjusted based on the cross entropy loss function and the Dice loss function. When the performance of the segmentation model no longer improves or reaches a preset number of iterations, the training is stopped to obtain a trained segmentation model for oral image segmentation; Acquire oral image data, and construct a data set to be segmented based on the oral image data; The data set to be segmented is input into the segmentation model. The data set to be segmented is downsampled and convolved by the encoder block to obtain a deep feature map. The channel attention module of the encoder block processes the deep feature map to obtain a key feature map. The non-local module of the decoder block processes the key feature map to obtain a global feature map. The oral segmentation image is output based on the key feature map and the global feature map.

2. The oral image segmentation method according to claim 1, characterized in that: The channel attention module includes a global average pooling layer, a fully connected layer and an output layer; The channel attention module of the encoder block processes the deep feature map to obtain the key feature map, including: Input the deep feature map into the global average pooling layer to obtain the channel feature map; Input the channel feature map into the fully connected layer to obtain the weight of each channel; The weights are fused with the channel feature map to obtain a calibration feature map, and the output layer outputs a key feature map based on the calibration feature map.

3. The oral image segmentation method according to claim 1, characterized in that: The non-local module includes a query feature convolution layer, a key feature convolution layer and a value feature convolution layer; The non-local module of the decoder block processes the key feature map to obtain a global feature map, including: The key feature graph is input into the query feature convolution layer, the key feature convolution layer and the value feature convolution layer respectively to obtain the query feature data, the key feature data and the value feature data correspondingly; A similarity matrix is ​​obtained based on the query feature data and the key feature data; The value feature data is weighted and summed based on the similarity matrix to obtain a weighted feature map, and a global feature map is obtained based on the weighted feature map and the key feature map.

4. The oral image segmentation method according to claim 3, characterized in that: The obtaining of a similarity matrix based on the query feature data and the key feature data includes: Multiply the query feature data and the key feature data to get the initial similarity matrix; The initial similarity matrix is ​​normalized based on the Softmax function to obtain the similarity matrix.

5. The oral image segmentation method according to claim 3, characterized in that: The obtaining of the global feature map based on the weighted feature map and the key feature map includes: adding the weighted feature map and the key feature map based on a residual connection to obtain the global feature map.

6. The oral image segmentation method according to claim 1, characterized in that: The step of constructing a data set to be segmented based on oral image data includes: Preprocessing the oral image data to obtain a preprocessed data set, wherein the preprocessing includes normalization processing and denoising processing; The preprocessed data set is subjected to data enhancement processing to construct a data set to be segmented, wherein the data enhancement processing includes rotation, translation, and scaling.

7. The oral image segmentation method according to claim 1, characterized in that: Each decoder block of the segmentation model is connected to the corresponding encoder block through a skip connection module.

8. An oral image segmentation system based on attention mechanism, characterized in that: include: A model construction unit, used to construct a segmentation model, wherein the segmentation model is a convolutional neural network composed of an encoder block, a skip connection module, and a decoder block, wherein the encoder block includes a downsampling convolution and a channel attention module, the decoder block includes an upsampling convolution and a non-local module, and the skip connection module is used to fuse the features extracted by the encoder block and the decoder block; A model training unit is used to train the segmentation model using an oral image data set containing original images and label data, and continuously adjust the parameters of the segmentation model based on a cross entropy loss function and a Dice loss function. When the performance of the segmentation model is no longer improved or reaches a preset number of iterations, the training is stopped to obtain a trained segmentation model for oral image segmentation; A data acquisition unit, used for acquiring oral image data and constructing a data set to be segmented based on the oral image data; The image segmentation unit is used to input the data set to be segmented into the segmentation model. The data set to be segmented is subjected to downsampling and convolution operations of the encoder block to obtain a deep feature map. The channel attention module of the encoder block processes the deep feature map to obtain a key feature map. The non-local module of the decoder block processes the key feature map to obtain a global feature map. The oral segmentation image is output based on the key feature map and the global feature map.

9. The oral image segmentation system according to claim 8, characterized in that: The data acquisition unit is further configured to: perform preprocessing and data enhancement processing on the acquired oral image data to construct a data set to be segmented.

10. The oral image segmentation system according to claim 8, characterized in that: The model construction unit is also configured to: when constructing the segmentation model, each decoder block is connected to the corresponding encoder block through a skip connection module.

Citation Information

Patent Citations

  • Oral cavity CBCT image tooth and soft tissue segmentation model method based on improved U-Net model

    CN117115132A

  • Depth visual perception method and system for building space analysis

    CN118967917A