A tooth lesion detection method based on global feature optimization of transformer
Patent Information
- Application Number
- CN202411176167.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2024-08-12
- Filing Date
- 2024-08-26
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-08-26
AI Technical Summary
[0004]针对现有技术中的上述不足,本发明提供了一种基于全局特征优化Transformer的牙齿病变检测方法,用于解决现有牙齿检测方法中存在的处理小尺度病变时的不足,尤其是解决在磁共振成像图像的复杂背景下检测困难的问题
[0040]1.本发明所提出的一种基于全局特征优化Transformer的牙齿病变检测方法,通过引入全局驱动语义增强模块,构建改进的Transformer结构,用于牙齿病变检测,不仅能够提高小尺度病变的检测精度与效率,还丰富特征表示能力,减少了计算复杂度的影响,提高了牙齿病变诊断效率与准确性。
Smart Images

Figure CN119067950B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of dental image data processing technology, and specifically to a method for detecting dental lesions based on a global feature-optimized Transformer. Background Technology
[0002] Dental diseases, including caries, pulpitis, and periodontal disease, often exhibit diverse morphological and structural features. These diseases have significant diagnostic and therapeutic value in clinical practice. Accurate detection and diagnosis of dental diseases are crucial for developing effective treatment plans. Magnetic resonance imaging (MRI) technology is widely used in the field of medical imaging, particularly in the detection and diagnosis of dental diseases. MRI offers advantages such as non-invasiveness, high contrast, and multi-parameter imaging, providing detailed information about teeth and surrounding tissues, making it an important tool for detecting dental diseases. However, in MRI images, these lesions may show slight differences in grayscale values and texture compared to normal tooth tissue and adjacent anatomical structures. Furthermore, the complex morphology and dense arrangement of teeth make the precise localization and boundary determination of lesion areas even more challenging.
[0003] Due to the complexity of the aforementioned magnetic resonance imaging images and the diversity of tooth structures, accurate detection of dental lesions, especially small-sized lesions, has always been a challenge. Currently, computer-aided detection (CAD) systems have made some progress in medical imaging analysis, especially deep learning-based methods, which have shown significant advantages. However, existing deep learning-based tooth detection methods have the following problems: (1) insufficient detection performance for small-scale lesions; (2) limitations in global feature extraction; and (3) insufficient richness of feature representation. Summary of the Invention
[0004] To address the aforementioned shortcomings in existing technologies, this invention provides a method for detecting dental lesions based on a global feature-optimized Transformer, which solves the deficiencies in existing dental detection methods when handling small-scale lesions, and particularly addresses the difficulty of detection against the complex background of magnetic resonance imaging images.
[0005] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows:
[0006] A method for detecting dental lesions based on global feature optimization of Transformer includes the following steps:
[0007] S1. Obtain magnetic resonance imaging of the teeth and input them into the backbone network to extract multi-scale features;
[0008] S2. Introduce a globally driven semantic enhancement module to build an improved Transformer structure;
[0009] The improved Transformer structure includes an encoder, a globally driven semantic enhancement module, and a decoder connected in sequence.
[0010] A global-driven semantic enhancement module is used to capture global features at different scales and enhance the detection of small-sized dental lesions;
[0011] S3. Input the multi-scale features into the improved Transformer structure for training to obtain a trained improved Transformer structure for identifying dental lesions.
[0012] S4. Obtain new magnetic resonance imaging images of the teeth and input them into the trained and improved Transformer structure to obtain the detection results of dental lesions.
[0013] Furthermore, the multi-scale features in step S1 are:
[0014] {S3, S4, S5}
[0015] Where S3 represents the third scale feature, S4 represents the fourth scale feature, and S5 represents the fifth scale feature.
[0016] Furthermore, the global-driven semantic enhancement module includes a first convolutional layer module, a second convolutional layer module, a first weight extraction layer module, a second weight extraction layer module, a third weight extraction layer module, a first upsampling module, a second upsampling module, and a convolutional block module.
[0017] Furthermore, step S3 specifically includes:
[0018] S31. The S3 scale feature in the multi-scale feature is positionally encoded and added to the S3 scale feature. Then, it is input into the encoder of the improved Transformer structure for encoding operation to obtain the global feature F3.
[0019] S32. Input global feature F3, S4 scale feature and S5 scale feature from multi-scale features into global driving semantic enhancement module to perform multi-scale feature fusion to obtain multi-scale fused feature image;
[0020] S33. Input the multi-scale fused feature image into the decoder of the improved Transformer structure to classify and locate the region of interest of dental lesions, and obtain the trained improved Transformer structure for recognizing dental lesions.
[0021] Furthermore, step S32 specifically includes:
[0022] S321. Input the global feature F3 into the first convolutional layer module for size alignment to obtain a global feature F3 with the same size as the feature at scale S5. Then input it into the first weight extraction layer module for convolution activation to obtain the weight W. s5 At the same time, the weight W s5 After element-wise multiplication with the S5-scale feature, and then element-wise addition with the S5-scale feature, the first weighted feature S′5 is obtained, i.e.:
[0023]
[0024] Where conv represents the convolution operation, ↓ represents the down-sizing operation, and sigmoid represents the activation function;
[0025] S322. After upsampling the first weighted feature S′5 into the first upsampling module, it is concatenated with the S4 scale feature to obtain the concatenated feature S′4, i.e.:
[0026] S′4=C(↑S′5,S4)
[0027] Where C represents splicing, and ↑ represents upsampling;
[0028] S323. Input the global feature F3 into the second convolutional layer module for size alignment to obtain a global feature F3 with the same size as the feature at scale S4. Then, input it into the second weight extraction layer module for convolutional activation to obtain the weight W. s4 At the same time, the weight W s4 After element-wise multiplication with the concatenated feature S′4, and then element-wise addition with the concatenated feature S′4, the second weighted feature S″4 is obtained, that is:
[0029]
[0030] S324. After upsampling the second weighted feature S″4 into the second upsampling module, it is concatenated with the global feature F3 to obtain the concatenated global feature F′3. Simultaneously, the global feature F3 is input into the third weight extraction layer module for convolution activation to obtain the weight W. s3 The weight W s3 After element-wise multiplication with the concatenated global feature F′3, and then element-wise addition with the concatenated global feature F′3, the third weighted feature F″3 is obtained, i.e.:
[0031]
[0032] S325. Input the third weighted feature F″3 into the convolutional block module for feature integration and optimization to obtain a multi-scale fused feature image.
[0033] Furthermore, step S33 specifically includes:
[0034] S331. Encode the multi-scale fused feature image to generate input features containing location information;
[0035] S332. The decoder is embedded into the input multi-head self-attention mechanism to capture the correlation within the decoder embedding and generate new decoder embedding features.
[0036] S333. The input features containing location information and the new decoder embedded features interact through a multi-head cross-attention mechanism to obtain encoder information related to the current decoding position. The information is then sequentially input into the feedforward neural network, summation and normalization, and multilayer perceptron to perform multilayer fully connected operations, residual summation and normalization, and classification and localization operations, respectively, to generate the predicted classification confidence and the localized bounding box offset.
[0037] S334. Based on the predicted bounding box offset and the predefined anchor box, generate a new bounding box and obtain the localization and classification results of the region of interest for dental lesions.
[0038] S335. Input the localization results, classification results, and ground truth labels of the region of interest of the dental lesion into the loss function. Use the loss function to calculate the gradient and update the network parameters of the improved Transformer structure to obtain the trained improved Transformer structure.
[0039] The present invention has the following beneficial effects:
[0040] 1. The present invention proposes a dental lesion detection method based on global feature optimization Transformer. By introducing a globally driven semantic enhancement module, an improved Transformer structure is constructed for dental lesion detection. This method not only improves the detection accuracy and efficiency of small-scale lesions, but also enriches the feature representation capability, reduces the impact of computational complexity, and improves the efficiency and accuracy of dental lesion diagnosis.
[0041] 2. In order to help the encoder of the improved Transformer structure distinguish the positional information of the input features and thus better capture the relative relationships between features, positional encoding is added to the S3 scale features before encoding. Then, the positional encoding is added to the S3 scale features and input into the encoder to extract global features, effectively capturing the global context within the S3 scale features and providing a comprehensive understanding of the medical scene. Attached Figure Description
[0042] Figure 1 This is a schematic diagram of the process of a dental lesion detection method based on global feature optimization Transformer proposed in this invention;
[0043] Figure 2 This is a schematic diagram of the improved Transformer structure. Detailed Implementation
[0044] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.
[0045] like Figure 1 As shown, a method for detecting dental lesions based on global feature optimization of Transformer includes the following steps S1-S4:
[0046] S1. Obtain magnetic resonance imaging of the teeth and input them into the backbone network to extract multi-scale features.
[0047] Specifically, the multi-scale features in step S1 are:
[0048] {S3, S4, S5}
[0049] Where S3 represents the third scale feature, S4 represents the fourth scale feature, and S5 represents the fifth scale feature.
[0050] In this embodiment, magnetic resonance imaging (MRI) images of teeth are first acquired and input into a backbone network to extract multi-scale features, which are {S3, S4, S5}. The multi-scale features {S3, S4, S5} extracted from the backbone network differ in spatial resolution and semantic information. Specifically, S3 has higher spatial resolution and lower semantic information, making it suitable for detecting small objects and thus capturing more subtle details. S4 and S5, on the other hand, have higher semantic information but lower spatial resolution, which helps to provide rich semantic context for identifying and classifying regions of interest (ROIs).
[0051] S2. Introduce a global-driven semantic enhancement module to construct an improved Transformer structure. The improved Transformer structure includes an encoder, a global-driven semantic enhancement module, and a decoder connected in sequence. The global-driven semantic enhancement module is used to capture global features at different scales and enhance the detection of small-sized dental lesions.
[0052] Specifically, the global-driven semantic enhancement module includes a first convolutional layer module, a second convolutional layer module, a first weight extraction layer module, a second weight extraction layer module, a third weight extraction layer module, a first upsampling module, a second upsampling module, and a convolutional block module.
[0053] In this embodiment, the improved Transformer structure is as follows: Figure 2 As shown, Figure 2 The improved Transformer structure includes an encoder, a global-driven semantic enhancement module, and a decoder connected in sequence. The global-driven semantic enhancement module includes a first convolutional layer module, a second convolutional layer module, a first weight extraction layer module, a second weight extraction layer module, a third weight extraction layer module, a first upsampling module, a second upsampling module, and a convolutional block module. The working principle of the improved Transformer structure is as follows: S3 is positionally encoded and then summed element-wise with S3. This sum is then input into the encoder to extract global features F3. The global-driven semantic enhancement module integrates and optimizes global features F3, S4, and S5 to enhance the overall feature representation capability. At the same time, the details of S3 are combined with the semantic richness of S4 and S5 to obtain the entire global feature information (multi-scale fused feature image), which is then input into the decoder to generate classification and localization results. The module comprises three layers: a first convolutional layer and a second convolutional layer, which align the global feature F3 with the multi-scale features in terms of size; a first weight extraction layer, a second weight extraction layer, and a third weight extraction layer, which extract different weight information; a first upsampling layer and a second upsampling layer, which perform upsampling and downsampling operations, respectively; and a convolutional block module, which integrates and optimizes the features to generate a rich global feature representation, i.e., a multi-scale fused feature image. Therefore, the globally driven semantic enhancement module significantly improves the ability of the improved Transformer structure to detect small lesions by focusing on enriching the global contextual features with high-level semantic features, thus enriching the feature representation capabilities and reducing the impact of computational complexity.
[0054] S3. Input the multi-scale features into the improved Transformer structure for training to obtain a trained improved Transformer structure for identifying dental lesions.
[0055] Specifically, step S3 includes S31-S33:
[0056] S31. The S3 scale feature in the multi-scale feature is positionally encoded and then added to the S3 scale feature. The result is then input into the encoder of the improved Transformer structure for encoding to obtain the global feature F3.
[0057] In this embodiment, to help the encoder of the improved Transformer structure distinguish the positional information of the input features and thus better capture the relative relationships between features, positional encoding is added to the S3-scale features before encoding. This positional encoding is then added to the S3-scale features and input into the encoder to extract global features. The purpose of this design is to effectively capture the global context within the S3-scale features, providing a comprehensive understanding of the medical scenario.
[0058] S32. Input the global feature F3 and the S4 and S5 scale features from the multi-scale features into the global driving semantic enhancement module to perform multi-scale feature fusion and obtain a multi-scale fused feature image.
[0059] In this embodiment, the global-driven semantic enhancement module is used to integrate and optimize the global feature F3 and the S4 and S5 scale features in the multi-scale features to enhance the overall feature representation, so as to combine the details of the S3 scale features with the semantic richness of the S4 and S5 scale features, thereby realizing the detection of dental lesions.
[0060] Specifically, step S32 includes S321-S325:
[0061] S321. Input the global feature F3 into the first convolutional layer module for size alignment to obtain a global feature F3 with the same size as the feature at scale S5. Then input it into the first weight extraction layer module for convolution activation to obtain the weight W. s5 At the same time, the weight W s5 After element-wise multiplication with the S5-scale feature, and then element-wise addition with the S5-scale feature, the first weighted feature S′5 is obtained, i.e.:
[0062]
[0063] Here, conv represents the convolution operation, ↓ represents the downsizing operation, and sigmoid represents the activation function.
[0064] In this embodiment, the stride of the first convolutional layer module (conv4×4) is 4. This module's function is to match the size of the global feature F3 with the S5 scale feature. The first weight extraction layer module consists of a conv1×1 layer and a sigmoid activation function, and its function is to generate weights W specific to the global features. s5 And using weight W s5 Multiplying with the S5-scale feature to embed global contextual information into the S5-scale feature, and then obtaining the first weighted feature S′5 through residual connection, i.e., addition operation; this process embodies global feature-driven semantic enhancement.
[0065] S322. After upsampling the first weighted feature S′5 into the first upsampling module, it is concatenated with the S4 scale feature to obtain the concatenated feature S′4, i.e.:
[0066] S′4=C(↑S′5,S4)
[0067] Where C represents splicing and ↑ represents upsampling.
[0068] In this embodiment, the purpose of upsampling is to match the size of the first weighted feature S′5 with the size of the S4 scale feature.
[0069] S323. Input the global feature F3 into the second convolutional layer module for size alignment to obtain a global feature F3 with the same size as the feature at scale S4. Then, input it into the second weight extraction layer module for convolutional activation to obtain the weight W. s4 At the same time, the weight W s4 After element-wise multiplication with the concatenated feature S′4, and then element-wise addition with the concatenated feature S′4, the second weighted feature S″4 is obtained, that is:
[0070]
[0071] In this embodiment, the purpose of this process is to utilize the extracted weight W. s4 The concatenated feature S′4 is further enhanced. The second weight extraction layer module also consists of a conv1×1 layer and a sigmoid activation function.
[0072] S324. After upsampling the second weighted feature S″4 into the second upsampling module, it is concatenated with the global feature F3 to obtain the concatenated global feature F′3. Simultaneously, the global feature F3 is input into the third weight extraction layer module for convolution activation to obtain the weight W. s3 The weight W s3 After element-wise multiplication with the concatenated global feature F′3, and then element-wise addition with the concatenated global feature F′3, the third weighted feature F″3 is obtained, i.e.:
[0073]
[0074] In this embodiment, the purpose of upsampling is to align the size of the second weighted feature S″4 with that of F3; simultaneously, using the extracted weight W s3 The global feature F′3 is optimized to integrate global feature information. The third weight extraction layer module also consists of a conv1×1 layer and a sigmoid activation function.
[0075] S325. Input the third weighted feature F″3 into the convolutional block module for feature integration and optimization to obtain a multi-scale fused feature image.
[0076] In this embodiment, the convolutional block module integrates and optimizes F″3 to generate rich global feature representations, i.e., a multi-scale fused feature image. Therefore, the globally driven semantic enhancement module significantly improves the ability of the improved Transformer structure to detect small lesions by focusing on using high-level semantic features to enrich global contextual features. The convolutional block module consists of convolutional layers and a tanh activation function.
[0077] S33. Input the multi-scale fused feature image into the decoder of the improved Transformer structure to classify and locate the region of interest of dental lesions, and obtain the trained improved Transformer structure for recognizing dental lesions.
[0078] In this embodiment, the multi-scale fused feature image is input into the decoder of the improved Transformer structure. Through multi-layer processing, including positional encoding, decoder embedding, and anchor box definition, classification and localization results are finally generated. Positional encoding helps the decoder distinguish the positional information of the input features, thereby better capturing the relative relationships between features. Anchor boxes are a predefined set of bounding boxes that cover the image at different scales and aspect ratios, helping the encoder predict the location and size of the region of interest in dental lesions. Decoder embedding uses learnable positional encoding as input embedding and undergoes multiple layer processing, ultimately transforming it into the output classification result and bounding box offset. Finally, the predefined anchor boxes are updated using the learned bounding box offsets to obtain the final localization result, i.e., the trained improved Transformer structure, used for recognizing dental lesions.
[0079] Specifically, step S33 includes S331-S335:
[0080] S331. Perform position encoding on the multi-scale fused feature image to generate input features containing position information.
[0081] S332. The decoder is embedded into the input multi-head self-attention mechanism to capture the correlation within the decoder embedding and generate new decoder embedding features.
[0082] S333. The input features containing location information and the new decoder embedded features interact through a multi-head cross-attention mechanism to obtain encoder information related to the current decoding position. The information is then sequentially input into a feedforward neural network, summation and normalization, and a multilayer perceptron to perform multilayer fully connected operations, residual summation and normalization, and classification and localization operations, respectively, to generate the predicted classification confidence and the localized bounding box offset.
[0083] S334. Based on the predicted bounding box offset and the predefined anchor box, generate a new bounding box and obtain the localization and classification results of the region of interest for dental lesions.
[0084] S335. Input the localization results, classification results, and ground truth labels of the region of interest of the dental lesion into the loss function. Use the loss function to calculate the gradient and update the network parameters of the improved Transformer structure to obtain the trained improved Transformer structure.
[0085] In this embodiment, when training the improved Transformer structure, the Hungarian method is used to determine the best binary match between the predicted target and the real target. At the same time, focal loss is used to supervise the classification loss between the predicted and real targets, and L1 loss and GIoU loss are used to supervise the localization loss between the predicted and real targets. The above three losses constitute the loss function used for gradient calculation and backpropagation. The network parameters of the improved Transformer structure are updated using this loss function to obtain the trained improved Transformer structure.
[0086] S4. Obtain new magnetic resonance imaging images of the teeth and input them into the trained and improved Transformer structure to obtain the detection results of dental lesions.
[0087] In this embodiment, after the improved Transformer structure is trained, the newly acquired magnetic resonance imaging image of the tooth can be input into the trained improved Transformer structure to detect and identify the target of interest in dental lesions.
[0088] In summary, the proposed method for detecting dental lesions based on global feature optimization Transformer has the following advantages: (1) It improves the detection accuracy of small lesions; that is, it can effectively capture and integrate global features of different scales, thereby enhancing the detection capability of small lesions; at the same time, the encoder can establish long-distance dependencies of images and capture global context information, so that small lesions can be accurately identified and detected in complex backgrounds; (2) It enriches feature representation; that is, the global feature optimization module further enriches feature representation by integrating high-level semantic information; so that the improved Transformer structure can better understand and distinguish different types of lesions by combining global context features with high-level semantic features; this multi-scale feature enhancement method not only improves the detection accuracy, but also significantly enhances the detection accuracy. (3) Reduced computational complexity: Although the Transformer structure performs well in global feature extraction, its computational complexity increases with the square of the input data size, which is limited by hardware conditions. The improved Transformer structure proposed in this invention effectively alleviates the problem of the complexity of the traditional Transformer structure by combining the encoder with the global driving semantic enhancement module, so that the model can still achieve efficient detection of dental lesions under the condition of limited hardware resources. (4) Improved diagnostic efficiency and accuracy: The combination of the encoder and the global driving semantic enhancement module improves feature extraction and enhancement, significantly improves the overall performance of the improved Transformer structure, and makes the detection and segmentation of dental lesions more accurate. It can not only improve diagnostic efficiency, but also provide clinicians with more accurate diagnostic information to assist them in the early diagnosis and treatment of diseases. Therefore, this invention uses the improved Transformer structure to detect dental lesions, which significantly improves the accuracy and efficiency of dental lesion detection, overcomes the shortcomings of existing methods in the detection of small-sized lesions, and provides an innovative and effective solution for the field of medical image analysis.
[0089] Specific embodiments have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.
[0090] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.
Claims
1. A method for detecting dental lesions based on global feature optimization of Transformer, characterized in that, Includes the following steps: S1. Obtain magnetic resonance imaging of the teeth and input them into the backbone network to extract multi-scale features; S2. Introduce a globally driven semantic enhancement module to build an improved Transformer structure; The improved Transformer structure includes an encoder, a globally driven semantic enhancement module, and a decoder connected in sequence. A global-driven semantic enhancement module is used to capture global features at different scales and enhance the detection of small-sized dental lesions; The global-driven semantic enhancement module includes a first convolutional layer module, a second convolutional layer module, a first weight extraction layer module, a second weight extraction layer module, a third weight extraction layer module, a first upsampling module, a second upsampling module, and a convolutional block module. S3. Input the multi-scale features into the improved Transformer structure for training to obtain a trained improved Transformer structure, which is used to identify dental lesions, specifically: S31, Incorporating multi-scale features After the scale features are positionally encoded, and The scale features are summed and then input into an encoder with an improved Transformer structure for encoding to obtain global features. ; S32, Global Features Multi-scale features Scale characteristics and The scale features are input into the global-driven semantic enhancement module to perform multi-scale feature fusion, resulting in a multi-scale fused feature image, specifically: S321, Global Features The first convolutional layer module is input for size alignment, resulting in... Global features of the same scale The weights are then input into the first weight extraction layer module for convolution activation to obtain the weights. At the same time, the weights and After performing element-wise multiplication of the scale features, and with The scale features are summed element by element to obtain the first weighted feature. ,Right now: in, This represents the convolution operation. This indicates a downward adjustment operation. Indicates the activation function; S322, Apply the first weighted feature After the first upsampling module performs the upsampling operation, and... The scale features are concatenated to obtain the concatenated features. ,Right now: in, Indicates splicing, Indicates upsampling; S323, Global Features The second convolutional layer module is input for size alignment, resulting in... Global features of the same scale The weights are then input into the second weight extraction layer module for convolution activation to obtain the weights. At the same time, the weights splicing features After element-wise multiplication, and then with the splicing features By adding elements one by one, we obtain the second weighted feature. ,Right now: ; S324, Apply the second weighted feature After the second upsampling module performs the upsampling operation, it is compared with the global features. By splicing the features together, we can obtain the spliced global features. At the same time, global features The third weight extraction layer module is input to perform convolutional activation operations to obtain the weights. , weight With splicing global features After performing element-wise multiplication, it is combined with the concatenated global features. By adding elements one by one, we obtain the third weighted feature. ,Right now: ; S325, the third weighted feature The input convolutional block module is used for feature integration and optimization to obtain a multi-scale fused feature image; S33. Input the multi-scale fused feature image into the decoder of the improved Transformer structure to classify and locate the region of interest of dental lesions, and obtain the trained improved Transformer structure for recognizing dental lesions. S4. Obtain new magnetic resonance imaging images of the teeth and input them into the trained and improved Transformer structure to obtain the detection results of dental lesions.
2. The method for detecting dental lesions based on global feature optimization Transformer according to claim 1, characterized in that, The multi-scale features in step S1 are: in, This represents the third-scale feature. This represents the fourth scale feature. This represents the 5th scale feature.
3. The method for detecting dental lesions based on global feature optimization Transformer according to claim 1, characterized in that, Step S33 specifically includes: S331. Encode the multi-scale fused feature image to generate input features containing location information; S332. The decoder is embedded into the input multi-head self-attention mechanism to capture the correlation within the decoder embedding and generate new decoder embedding features. S333. The input features containing location information and the new decoder embedded features interact through a multi-head cross-attention mechanism to obtain encoder information related to the current decoding position. The information is then sequentially input into the feedforward neural network, summation and normalization, and multilayer perceptron to perform multilayer fully connected operations, residual summation and normalization, and classification and localization operations, respectively, to generate the predicted classification confidence and the localized bounding box offset. S334. Based on the predicted bounding box offset and the predefined anchor box, generate a new bounding box and obtain the localization and classification results of the region of interest for dental lesions. S335. Input the localization results, classification results, and ground truth labels of the region of interest of the dental lesion into the loss function. Use the loss function to calculate the gradient and update the network parameters of the improved Transformer structure to obtain the trained improved Transformer structure.
Citation Information
Patent Citations
Deep learning-based root tip X-ray film disease identification method and system
CN117132835A
KR20240107030A