Dental maxillofacial region CBCT automatic point fixing method and device based on neural network
By improving the encoder of the U-Net model, the 3D multi-scale feature global attention module was introduced, and the accuracy and applicability of key points positioning of soft and hard tissues in the CBCT image processing of toothmaxillary facial CBCT is solved, and high-precision detection of all ages and complex malformations is achieved.
Patent Information
- Application Number
- CN202510426466.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-22
AI Technical Summary
The existing U-Net model has clinical applicability limitations and insufficient multi-tissue detection accuracy in the processing of dental and maxillofacial CBCT images. Especially when covering all ages and complex mishap malformations, it is difficult to achieve high-precision positioning of key points of soft and hard tissues.
Improve the encoder of the U-Net model, introduce a 3D multi-scale feature global attention module, combine it with the 3D downsampling module, extract local features and capture global information, and improve the accuracy and robustness of key point detection through multi-scale global attention calculation and cross-attention operation.
High-precision positioning of key points of soft and hard tissue in CBCT images of toothma maxillofacial areas is achieved, which improves the robustness and generalization ability of the model and meets the needs of clinical application.
Smart Images

Figure CN120355669A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computational models, and particularly to an automatic tooth and jaw facial CBCT point positioning method and device based on a neural network. Background Art
[0002] In recent years, deep learning has been widely applied in the field of medical images, especially in medical image analysis, disease diagnosis, surgical assistance, etc. The proposal of U-Net has greatly promoted the development of image segmentation tasks. As a classic convolutional neural network architecture, U-Net can efficiently handle pixel-level segmentation problems in medical images with its unique encoder-decoder structure and skip connection method, and is particularly suitable for fine-grained structure recognition, such as tumor regions, vascular networks, etc. The U-Net network architecture has shown excellent recognition performance in the field of two-dimensional image processing, but its application in three-dimensional medical images (especially dental and maxillofacial cone beam computed tomography (CBCT)) still has significant limitations. Through a systematic analysis of existing CBCT automatic point positioning technologies, the following key problems need to be solved urgently: 1. Limited clinical applicability: Existing models are mostly developed based on regular dentition or a single type of malocclusion. In actual clinical practice, patients of different ages, dentition development stages, and malocclusion types have significant morphological differences in the hard and soft tissues of the tooth and jaw face. Currently, there is still a lack of a CBCT automatic point positioning solution that can cover all age groups, multiple dentition stages, and adapt to complex malocclusion types, which requires further optimization of the traditional U-Net model to improve the robustness and generalization ability of the model; 2. Insufficient accuracy in multi-tissue detection: When existing models simultaneously detect key points of hard and soft tissues, their spatial positioning accuracy has not yet reached the clinical practical requirements. Summary of the Invention
[0003] In order to solve the above problems, the purpose of the present invention is to provide an automatic tooth and jaw facial CBCT point positioning method based on a neural network for different malocclusion types of all age groups, which improves the encoder of the U-NET model, enables the network to better capture fine features in the image, and improves the accuracy and robustness of key point detection.
[0004] To achieve the above purpose, the present invention adopts the following technical solutions:
[0005] Technical Solution 1
[0006] An automatic tooth and jaw facial CBCT point positioning method based on a neural network, comprising the following steps:
[0007] Construct a U-Net model: It includes an encoder, a decoder, and a prediction head. The encoder contains a 3D downsampling module and a 3D multi-scale feature global attention module connected in sequence. The 3D downsampling module extracts local features, and the 3D multi-scale feature global attention module captures global information. The feature maps output by each 3D downsampling module and 3D multi-scale feature global attention module are transmitted in two paths. One path is output to the decoder, and the other path is transmitted to the next module connected to it. The decoder contains multiple 3D upsampling modules. Each 3D upsampling module receives one output of the encoder and the output of the previous module connected to it. The output of the decoder is connected to the prediction head, and the prediction head outputs a heat map containing the position information of the landmark points.
[0008] Construct a CBCT image dataset and annotate the soft tissues and hard tissues in the images, and divide the dataset into a training set and a test set.
[0009] Use the training set to train the model and use the test set to evaluate the overall localization performance of the model.
[0010] Detect the dental and maxillofacial CBCT images through the trained U-Net model to detect the positions of the soft tissues and hard tissues.
[0011] Preferably, the 3D multi-scale feature global attention module includes a dynamic position encoding module, a multi-scale global attention network, a multi-scale feature fusion network, and a 3D downsampling module connected in sequence.
[0012] Preferably, the multi-scale global attention network performs 3D normalization and multi-scale global attention calculation. The multi-scale global attention calculation process is as follows: The feature map X after 3D normalization is respectively input into multiple depthwise separable convolution modules to obtain feature maps with different scales. Each of the feature maps with different scales reduces the number of channels to 1 / 4 of the original through pointwise convolution, and each scale of the feature map respectively outputs a set of feature vectors, including the value-key vector (K) for cross-attention calculation, the query vector (q) for self-attention calculation, the key vector (k), and the value vector (v); the feature map X also obtains the feature vector Q through pointwise convolution. shared ; Perform a self-attention calculation on the query vector (q), the key vector (k), and the value vector (v) to output the feature vector V, and combine the feature vector V with the feature vector Q shared and the value-key vector (K) for cross-attention calculation to output the feature map M, and combine the feature map M with the feature vector Q sharedPerform splicing, add the result of splicing to the input feature map X element by element, and then output the multi-scale feature fusion network.
[0013] Preferably, the multi-scale feature fusion network includes a 3D batch normalization layer, a 3D convolutional layer F1, a 3D depth convolutional layer F2, and a GELU activation function. The core calculation process is defined as follows: X = X + F1(F2(GELU(F1(X)))); where X represents the output of the multi-scale global attention network.
[0014] Preferably, the 3D upsampling module performs upsampling by the method of trilinear interpolation. Through the feature fusion of multiple 3D upsampling modules, the spatial size of the feature maps output by each 3D upsampling module gradually expands, and at the same time, the number of channels gradually decreases, and finally aligns with the dimension of the original input feature.
[0015] The present invention also provides a dental and maxillofacial CBCT automatic positioning device based on a neural network. The technical solution is as follows:
[0016] Technical solution two
[0017] The dental and maxillofacial CBCT automatic positioning device based on a neural network includes
[0018] U-Net model construction module: including an encoder, a decoder, and a prediction head. The encoder includes a 3D downsampling module and a 3D multi-scale feature global attention module connected in sequence. The 3D downsampling module extracts local features, and the 3D multi-scale feature global attention module captures global information. The feature maps output by each 3D downsampling module and 3D multi-scale feature global attention module are transmitted in two paths, one of which is output to the decoder, and the other is transmitted to the next module connected thereto. The decoder includes multiple 3D upsampling modules. Each 3D upsampling module receives one output of the encoder and the output of the previous module connected thereto. The output of the decoder is connected to the prediction head, and the prediction head outputs a heat map containing the position information of the landmark points;
[0019] Dataset construction module: construct a CBCT image dataset and label the soft tissue and hard tissue in the image, and divide the dataset into a training set and a test set;
[0020] Model training and evaluation module: use the training set to train the model and use the test set to evaluate the overall positioning performance of the model;
[0021] Application module: input the dental and maxillofacial CBCT image into the trained U-Net model to detect the positions of the soft tissue and hard tissue.
[0022] The present invention has the following beneficial effects:
[0023] The automatic key-point positioning method and device for dental and maxillofacial CBCT based on neural network designed a 3D multi-scale feature global attention module, which was combined with the traditional down-sampling module and used as the encoder of the model to replace the traditional high-dimensional feature extraction process, effectively extracting the global information and local information of the image. The 3D multi-scale feature global attention module first performs multi-scale convolution to extract local feature information at different scales, and then further makes up for the lack of global feature information through the self-attention mechanism to obtain the global features and local features at different scales and the attention relationship weights at different scales. Finally, through the cross-attention operation with the feature map under conventional convolution, the fusion and alignment of information at different scales are realized, providing rich multi-scale information support for key-point positioning and improving the accuracy of key-point positioning. At the same time, feature extraction is performed through multi-scale global attention in the encoder stage, effectively improving the robustness of the model. Description of the Drawings
[0024] Figure 1 It is the structural block diagram of the U-Net model of the present invention;
[0025] Figure 2 It is the structural block diagram of the multi-scale feature global attention module of the present invention;
[0026] Figure 3 It is the structural block diagram of the 3D multi-scale global attention network of the present invention;
[0027] Figure 4 It is the schematic block diagram of the automatic key-point positioning device of the present invention. Detailed Embodiments
[0028] The following further describes the present invention in detail with reference to the drawings and specific embodiments:
[0029] Embodiment 1
[0030] The automatic key-point positioning method for dental and maxillofacial CBCT based on neural network includes the following steps:
[0031] Step 10: Construct a U-Net model, including an encoder, a decoder, and a prediction head. Please refer to Figure 1 , and specifically execute steps 11 - 18.
[0032] Step 11: Use a 1×1×1 convolutional layer (CL) to expand the number of channels. This process is defined as follows: X = CL(I), where X represents the output of CL and I represents the input image. Taking the input image I ∈ R H×W×D as an example, an additional channel dimension is first introduced, and then a 1×1×1 convolutional layer (CL) is applied to expand the number of channels.
[0033] Step 12: The encoder consists of two stages. The first stage is a traditional 3D downsampling module that extracts local features, and the second stage is a 3D multi-scale feature global attention module that captures global information.
[0034] In this embodiment, the first stage of the encoder has two layers, that is, it includes two 3D downsampling modules. Each 3D downsampling module contains two sets of convolutional layers (3D convolution + 3D batch normalization + ReLU activation function). The downsampling module reduces the image size by 1 / 2 and doubles the number of channels, gradually becoming an abstract feature map. The downsampling process can be defined as follows: X = Conv(MaxPool(X)); where Maxpool(·) represents the max pooling layer, and Conv(·) represents two 3D convolutional layers (CL). Each convolutional layer uses a 3×3×3 convolutional kernel with a stride of 2 and a padding of 1 to reduce the spatial dimension while maintaining local features.
[0035] Step 13: The second stage of the encoder has two layers, that is, it includes two 3D multi-scale feature global attention modules. Please refer to Figure 2 , each 3D multi-scale feature global attention module includes a dynamically positional encoding module, a multi-scale global attention network, a multi-scale feature fusion network, and a 3D downsampling module connected in sequence.
[0036] The dynamically positional encoding module (DPC) contains a 3D convolution with a size of 3x3 and a stride of 1, and uses depthwise separable convolution, that is, each input channel performs convolution operations independently. The dynamically positional encoding calculation process can be defined as follows: X = X + DPC(X), where X represents the output after dynamically positional encoding, and the output X overwrites the input X as the input for the next stage. The result of the convolution is added element-wise to the input X to form a residual connection.
[0037] The multi-scale global attention network contains a 3D normalization layer and a multi-scale global attention calculation operation. Please refer to Figure 3 , the multi-scale global attention calculation process is as follows: The feature map X after 3D normalization is respectively input into three depthwise separable convolution modules to obtain three feature maps with different scales, and then the number of channels is reduced to 1 / 4 of the original through pointwise convolution. Each of the feature maps with different scales passes through a 1×1×1 convolutional kernel with a stride of 1 for convolution to output a set of feature vectors respectively, including the value-key vector (K) for cross-attention calculation, the query vector (q) for self-attention calculation, the key vector (k), and the value vector (v). At the same time, the feature map X also obtains the feature vector Q through pointwise convolution shared . Perform a self-attention calculation on the query vector (q), the key vector (k), and the value vector (v) to output the feature vector V. The feature vector Qshared Perform cross-attention calculation with the value key vector (K) and the feature vector V, output the feature map M, and concatenate the feature map M with the feature vector Q shared Perform concatenation, add the result of concatenation to the input feature map X element-wise, and then output the multi-scale feature fusion network.
[0038] The multi-scale global attention calculation process can be defined as follows:
[0039] Q shared = PWC(X);
[0040] q 1,2,3 , k 1,2,3 , v 1,2,3 , K 1,2,3 = PWC(DWC(X));
[0041]
[0042] X = X + Concat(Q shared , M1, M2, M3);
[0043] Where PWC(·) represents pointwise convolution, using a 1×1×1 convolution kernel, with a stride of 1, and the number of channels is reduced to 1 / 4 of the original; DWC(·) represents depthwise separable convolution, using convolution kernels of 4, 8, and 12 respectively, with strides of 2, 4, and 6 respectively, and paddings of 1, 2, and 3 respectively, and the number of channels remains unchanged; the feature maps q 1,2,3 , k 1,2,3 , v 1,2,3 Perform self-attention calculation once to obtain the feature map V 1,2,3 , and concatenate the feature map V 1,2,3 with the feature map Q shared , K 1,2,3 to perform cross-attention calculation and output the feature map M 1,2,3 , and concatenate the feature map M 1,2,3 with Q shared to perform concatenation operation, and add the result of concatenation to the input X element-wise to form a residual connection, and the output X overwrites the input X as the input of the next stage. The subscripts 1, 2, 3 in q 1,2,3 , k 1,2,3 , v 1,2,3 , V 1,2,3 , K 1,2,3 , M 1,2,3 represent the relationship of OR, for example, the feature map K 1,2,3 represents the feature maps K1, K2, and K3.
[0044] Step 14. The multi-scale feature fusion network includes a 3D batch normalization layer, a 3D convolutional layer F1, a 3D depth convolutional layer F2, and a GELU activation function. The core calculation process is defined as follows:
[0045] X = X + F1(F2(GELU(F1(X)))); where X represents the output of the multi-scale global attention network, and the output X overwrites the input X as the input for the next stage;
[0046] Step 15. Input the feature map output by the multi-scale global attention network into a downsampling module. The output feature map is used as the feature extraction result of one of the 3D multi-scale feature global attention modules, achieving further spatial dimension reduction while extracting richer high-level features. The 3D multi-scale feature global attention module increases the diversity of the generated global features, thereby improving the sensitivity of the model to the key point positions.
[0047] Step 16. The feature maps output by each of the 3D downsampling modules and 3D multi-scale feature global attention modules are transmitted in two paths. One path is output to the decoder, and the other path is transmitted to the next module connected to it.
[0048] Step 17. The decoder includes multiple sequentially connected 3D upsampling modules. Each 3D upsampling module receives one output of the encoder and the output of the previous module connected to it. The feature maps of each encoder stage are reused first. The 3D upsampling module performs upsampling by trilinear interpolation. Through the feature fusion of multiple 3D upsampling modules, the spatial size of the feature maps output by each 3D upsampling module gradually expands, while the number of channels gradually decreases, and finally aligns with the dimension of the original input feature. The 3D upsampling module can be defined as follows: X = Conv(Concat(τ(X d ), X e , dim = c)); where X d and X e represent the decoder feature and the encoder feature map respectively. τ(·) represents trilinear interpolation, which magnifies the resolution to 2 times the original. X d and X e are concatenated along the channel dimension and the upsampling process is completed through a Conv(·).
[0049] Step 18: The output of the decoder is connected to the prediction head, and the prediction head outputs a heatmap containing the landmark position information. Specifically, in the prediction stage, through encoder and decoder feature extraction, we obtain feature maps containing local and global information, and the resolution of these feature maps is the same as that of the original input image. Then, using a 1×1×1 convolutional prediction head, these feature maps are converted into N heatmaps, and each heatmap corresponds to one of the N landmarks. For each heatmap, the position of the maximum value is determined, and its coordinates are considered as the final predicted key point positions.
[0050] Step 20: Construct a CBCT image dataset and label the soft and hard tissues in the images, and divide the dataset into a training set and a test set;
[0051] Step 30: Use the training set to train the model and use the test set to evaluate the overall localization performance of the model. To evaluate the overall localization performance on the test set, calculate for each landmark: 1) the mean radial error (MRE), which is the average Euclidean distance between the reference landmark and the predicted landmark; and 2) the successful detection rate (SDR), which is the proportion of landmarks with a radial error less than 2 mm, 3 mm, and 4 mm, and its calculation formula is as follows:
[0052]
[0053] where x k , y k , and z k represent the coordinates of the manually labeled landmarks along the x, y, and z axes respectively, while and represent the predicted landmark coordinates along the x, y, and z axes respectively. The calculation formula is as follows:
[0054]
[0055] where, d k represents the localization error of the k-th key point, which is usually used to calculate the Euclidean distance between the predicted landmark position and the manually calibrated landmark position. If d k the error is lower than the threshold σ (d k < σ), the indicator function is equal to 1, otherwise it is equal to 0, which means that the k-th landmark is considered to be successfully detected when the error is lower than the threshold.
[0056] The model of the present invention was trained on a server equipped with an Intel Xeon Gold 6325 central processing unit (2.90 GHz) and an NVIDIA A40 graphics processing unit (GPU), with 48 GB of memory. GPU acceleration was achieved using cuDNN version 8.9.0.2, and the PyTorch framework, version 2.2.2, was used. Table 4 shows the accuracy of the automatic detection of key points of this algorithm:
[0057] Table 4
[0058]
[0059] The overall mean absolute error of 43 landmark points was 0.71 mm on the x-axis, 0.67 mm on the y-axis, and 0.85 mm on the z-axis. The overall MRE of 43 landmark points was 1.76 ± 1.13 mm. Within the error ranges of 2 mm, 3 mm, and 4 mm of the manual markings, the SDRs were 60.16%, 91.05%, and 97.58% respectively. Among the 43 landmark points, 32 were hard tissue landmark points and 11 were soft tissue landmark points.
[0060] This method achieved the task of automatic landmarking of 3D CBCT cephalometry by a neural network. The average MRE of hard tissue landmark points was 1.73 mm, and the average MRE of soft tissue landmark points was 1.84 mm, both of which met the clinical accuracy criteria (a distance of no more than 2 mm between automatic and manual recognition is considered clinically accurate). The MREs of 15 hard tissue landmark points (25 / 32, 78.1%) and 7 soft tissue landmark points (8 / 11, 72.7%) were less than 2 mm.
[0061] Step 40: Detect the soft and hard tissues of the dental and maxillofacial CBCT images through the trained U-Net model.
[0062] The automatic tooth and jaw facial CBCT fixed-point method based on neural network of the present invention deeply improves the encoder in the U-Net model, designs a 3D multi-scale feature global attention module, and replaces the traditional high-dimensional feature extraction process. The 3D multi-scale feature global attention module first performs multi-scale convolution operations on the input feature map to effectively extract local feature information at different scales, and then further makes up for the lack of global feature information through the self-attention mechanism to ensure the complementarity of global and local features, and obtains global features and local features at different scales. Finally, by performing cross-attention operations on the feature map under conventional convolution, the fusion and alignment of different-scale information are realized. Performing the similarity calculation of the feature map in the cross-attention operation helps to integrate multi-scale information onto the feature map of the original size, thereby providing rich multi-scale information support for key-point positioning and improving the accuracy of key-point positioning. At the same time, the present invention performs feature extraction through multi-scale global attention in the encoder stage, improving the robustness of the model.
[0063] Based on the same inventive concept, the present application also provides a device corresponding to the method in Embodiment 1. For details, see Embodiment 2.
[0064] Embodiment 2
[0065] Please refer to Figure 4 , the automatic tooth and jaw facial CBCT fixed-point device based on neural network, includes
[0066] U-Net model construction module: including an encoder, a decoder and a prediction head. The encoder includes a 3D downsampling module and a 3D multi-scale feature global attention module connected in sequence. The 3D downsampling module extracts local features, and the 3D multi-scale feature global attention module captures global information. The feature maps output by each 3D downsampling module and 3D multi-scale feature global attention module are both transmitted in two paths, one path is output to the decoder, and the other path is transmitted to the next module connected thereto. The decoder includes a plurality of 3D upsampling modules. Each 3D upsampling module receives one output of the encoder and the output of the previous module connected thereto. The output of the decoder is connected to the prediction head, and the prediction head outputs a heat map including the position information of the landmark points.
[0067] The 3D multi-scale feature global attention module includes a dynamic position encoding module, a multi-scale global attention network, a multi-scale feature fusion network, and a 3D downsampling module connected in sequence. The multi-scale global attention network performs 3D normalization and multi-scale global attention calculation. The multi-scale global attention calculation process is as follows: The feature map X after 3D normalization is respectively input into three depthwise separable convolution modules to obtain three feature maps with different scales. Each of the feature maps with different scales reduces the number of channels to 1 / 4 of the original through pointwise convolution, and each scale of the feature map respectively outputs a set of feature vectors, including the value-key vector (K) for cross-attention calculation, the query vector (q) for self-attention calculation, the key vector (k), and the value vector (v); The feature map X also obtains the feature vector Q through pointwise convolution. shared Perform a self-attention calculation on the query vector (q), the key vector (k), and the value vector (v), output the feature vector V, and combine the feature vector V with the feature vector Q shared Perform cross-attention calculation with the value-key vector (K), output the feature map M, and combine the feature map M with the feature vector Q shared Perform concatenation, add the concatenated result to the input feature map X element-wise, and then output to the multi-scale feature fusion network.
[0068] The multi-scale feature fusion network includes a 3D batch normalization layer, a 3D convolutional layer F1, a 3D depth convolutional layer F2, and a GELU activation function. The core calculation process is defined as follows:
[0069] X = X + F1(F2(GELU(F1(X)))); where X represents the output of the multi-scale global attention network.
[0070] The 3D upsampling module performs upsampling by trilinear interpolation. Through the feature fusion of multiple 3D upsampling modules, the spatial size of the feature map output by each 3D upsampling module gradually expands, and at the same time, the number of channels gradually decreases, and finally aligns with the dimension of the original input feature.
[0071] Dataset construction module: Construct a CBCT image dataset, label the soft tissue and hard tissue in the image, and divide the dataset into a training set and a test set;
[0072] Model training and evaluation module: Use the training set to train the model and use the test set to evaluate the overall localization performance of the model;
[0073] Application module: Input the dental and maxillofacial CBCT graphics into the trained U-Net model to detect the positions of soft tissue and hard tissue.
[0074] Since the device introduced in the second embodiment of the present invention is the device adopted for implementing the method of the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of the device, so it will not be elaborated here. Any device adopted for the method of the first embodiment of the present invention falls within the scope of protection of the present invention.
[0075] The above is only a specific implementation manner of the present invention, and does not limit the patent scope of the present invention. Any equivalent structural transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. An automatic dental and maxillofacial CBCT point - setting method based on a neural network, characterized in that, It includes the following steps: Construct a U-Net model: It includes an encoder, a decoder, and a prediction head. The encoder contains a 3D downsampling module and a 3D multi-scale feature global attention module connected in sequence. The 3D downsampling module extracts local features, and the 3D multi-scale feature global attention module captures global information. The feature maps output by each 3D downsampling module and 3D multi-scale feature global attention module are transmitted in two paths. One path is output to the decoder, and the other path is transmitted to the next module connected to it. The decoder contains multiple 3D upsampling modules. Each 3D upsampling module receives one output of the encoder and the output of the previous module connected to it. The output of the decoder is connected to the prediction head, and the prediction head outputs a heat map containing the position information of the landmark points; Construct a CBCT image dataset and label the soft tissues and hard tissues in the images, and divide the dataset into a training set and a test set; Use the training set to train the model and use the test set to evaluate the overall localization performance of the model; Detect the dental and facial CBCT images through the trained U-Net model to detect the positions of the soft tissues and hard tissues.
2. The automatic tooth and jaw facial CBCT fixed-point method based on a neural network according to claim 1, characterized in that: The 3D multi-scale feature global attention module includes a dynamic position encoding module, a multi-scale global attention network, a multi-scale feature fusion network, and a 3D downsampling module connected in sequence.
3. The method for automatically determining points on dental and maxillofacial CBCT based on a neural network according to claim 2, wherein: The multi-scale global attention network performs 3D normalization and multi-scale global attention calculation. The multi-scale global attention calculation process is as follows: The feature map X after 3D normalization is respectively input into multiple depthwise separable convolution modules to obtain feature maps with different scales. Each of the feature maps with different scales reduces the number of channels to 1 / 4 of the original through pointwise convolution, and a set of feature vectors are respectively output for each scale of the feature map, including the value-key vector (K) for cross-attention calculation, the query vector (q) for self-attention calculation, the key vector (k), and the value vector (v); The feature map X also obtains the feature vector Q through pointwise convolution. shared ; Perform a self-attention calculation on the query vector (q), the key vector (k), and the value vector (v), and output the feature vector V. Combine the feature vector V with the feature vector Q shared and the value-key vector (K) for cross-attention calculation, and output the feature map M. Combine the feature map M with the feature vector Q shared for concatenation. The concatenated result is added element-wise to the input feature map X, and then output to the multi-scale feature fusion network.
4. The automatic tooth and jaw facial CBCT positioning method based on neural network according to claim 2 or 3, characterized in that: The multi-scale feature fusion network contains a 3D batch normalization layer, a 3D convolutional layer F1, a 3D depth convolutional layer F2, and a GELU activation function. The core calculation process is defined as follows: X = X + F1(F2(GELU(F1(X)))); where X represents the output of the multi-scale global attention network.
5. The automatic tooth and jaw facial CBCT fixed-point method based on a neural network according to claim 1, wherein: The 3D upsampling module performs upsampling by the method of trilinear interpolation. Through the feature fusion of multiple 3D upsampling modules, the spatial dimensions of the feature maps output by each 3D upsampling module are gradually expanded, and at the same time, the number of channels is gradually reduced, and finally aligned with the dimension of the original input feature.
6. The automatic tooth and jaw facial CBCT positioning device based on a neural network is characterized in that, Include U-Net model construction module: It includes an encoder, a decoder, and a prediction head. The encoder contains a 3D downsampling module and a 3D multi-scale feature global attention module connected in sequence. The 3D downsampling module extracts local features, and the 3D multi-scale feature global attention module captures global information. The feature maps output by each 3D downsampling module and 3D multi-scale feature global attention module are transmitted in two paths. One path is output to the decoder, and the other path is transmitted to the next module connected to it. The decoder contains multiple 3D upsampling modules. Each 3D upsampling module receives one output of the encoder and the output of the previous module connected to it. The output of the decoder is connected to the prediction head, and the prediction head outputs a heat map containing the position information of the landmark points; Dataset construction module: Construct a CBCT image dataset and label the soft tissues and hard tissues in the images, and divide the dataset into a training set and a test set; Model training and evaluation module: Use the training set to train the model and use the test set to evaluate the overall localization performance of the model; Application module: Input the dental and maxillofacial CBCT images into the trained U-Net model to detect the positions of soft tissues and hard tissues.
7. The automatic tooth and jaw facial CBCT fixed-point device based on a neural network according to claim 6, wherein, The 3D multi-scale feature global attention module includes a dynamically positional encoding module, a multi-scale global attention network, a multi-scale feature fusion network, and a 3D downsampling module connected in sequence.
8. The automatic tooth and jaw facial CBCT fixed-point device based on a neural network according to claim 7, characterized in that: The multi-scale global attention network performs 3D normalization and multi-scale global attention calculation. The multi-scale global attention calculation process is as follows: The feature map X after 3D normalization is respectively input into multiple depthwise separable convolution modules to obtain feature maps with different scales. Each of the feature maps with different scales reduces the number of channels to 1 / 4 of the original through pointwise convolution, and each scale of feature map respectively outputs a set of feature vectors, including the value-key vector (K) for cross-attention calculation, the query vector (q) for self-attention calculation, the key vector (k), and the value vector (v); the feature map X also obtains the feature vector Q through pointwise convolution. shared ; Perform a self-attention calculation on the query vector (q), the key vector (k), and the value vector (v), and output the feature vector V. Combine the feature vector V with the feature vector Q shared , the value-key vector (K) for cross-attention calculation, and output the feature map M. Combine the feature map M with the feature vector Q shared for splicing. The result after splicing is added element-wise to the input feature map X, and then output to the multi-scale feature fusion network.
9. The neural network-based automatic tooth and jaw facial CBCT point positioning device according to claim 7 or 8, characterized in that: The multi-scale feature fusion network includes a 3D batch normalization layer, a 3D convolutional layer F1, a 3D depth convolutional layer F2, and a GELU activation function. The core calculation process is defined as follows: X = X + F1(F2(GELU(F1(X)))); where X represents the output of the multi-scale global attention network.
10. The automatic tooth and jaw facial CBCT positioning device based on a neural network according to claim 6, characterized in that: The 3D upsampling module performs upsampling by trilinear interpolation. Through the feature fusion of multiple 3D upsampling modules, the spatial dimensions of the feature maps output by each 3D upsampling module gradually expand, while the number of channels gradually decreases, and finally aligns with the dimension of the original input features.