A method and system for detecting landmarks on a lateral cephalogram
Patent Information
- Application Number
- CN202311632252.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-01
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-12-01
AI Technical Summary
[0005]1、固定权重分配:普通深度学习模型通常在处理输入数据时,所有的特征都被平等对待,没有考虑到不同特征的重要性差异
[0035]本发明提供的一种头颅侧位X光片地标点检测方法及系统,其构建了一种U型混合注意力网络模型,用于进行头颅侧位X光片地标点检测,U型混合注意力网络模型首先采用多层多头注意力像素卷积和多层下采样将输入图像投影到多个子空间,然后采用相对应的多层上采样将下采样之后的图像投射到分辨率更高的空间,下采样和上采样之间通过通道注意力机制融合。
Smart Images

Figure CN117635572B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical image processing technology, and in particular to a method and system for detecting landmarks on lateral cephalometric X-ray films. Background Technology
[0002] Currently, landmark detection on head X-rays is a crucial task in the field of medical imaging. It is applied in medical imaging diagnosis and treatment processes, including head morphology analysis, craniofacial deformity assessment, and orthodontic treatment planning. By automatically or semi-automatically detecting specific key points on head X-rays, such as the tip of the nose, the edge of the orbit, and the temporomandibular joint, quantitative morphological data and anatomical information can be provided, assisting physicians in making diagnostic and treatment decisions.
[0003] Existing head X-ray landmark detection technologies are primarily based on computer vision and image processing methods. Generally, the main steps of these technologies include the following: a. Preprocessing: Preprocessing the head X-ray, such as denoising, contrast enhancement, and image correction. This helps improve the accuracy and stability of landmark detection. b. Feature extraction: Extracting specific features or feature descriptors from the head X-ray to identify and locate key points. Commonly used feature extraction methods include edge detection, local feature extraction (such as SIFT, SURF, ORB, etc.), and deep learning methods (such as convolutional neural networks). c. Landmark localization: Using machine learning or optimization algorithms to locate landmarks on the head X-ray. Machine learning methods can be trained using labeled datasets, such as Support Vector Machines (SVM), Random Forest, and regression models. Optimization algorithms can iteratively adjust the location of landmarks based on specific constraints and objective functions, such as minimizing deformation energy or maximizing feature similarity. d. Post-processing: Post-processing the detected landmarks, such as filtering outliers, smoothing trajectories, and applying boundary constraints. This helps improve the accuracy and consistency of landmark detection.
[0004] Current deep learning models in the field of cephalometric analysis still have the following problems:
[0005] 1. Fixed weight allocation: Typical deep learning models treat all features equally when processing input data, without considering the differences in importance between different features. This may cause the model to fail to effectively focus on key information when processing complex data;
[0006] 2. Difficulty in capturing feature correlations: When dealing with highly correlated features, ordinary models struggle to effectively capture the correlations between these features, which may lead to information loss.
[0007] 3. Insufficient information focus: Ordinary models may not be able to effectively focus attention on the important parts of the input data, which may result in poor performance when processing long sequences of data or large images;
[0008] 4. Low computational efficiency: When dealing with a large amount of input data, ordinary models may need to process all features at the same time, which leads to a decrease in computational efficiency;
[0009] 5. Adversarial attacks: Ordinary models may be affected by adversarial attacks when dealing with some complex tasks because they do not have a clear mechanism to focus on the features that adversarial attacks may occur.
[0010] 6. Insufficient generalization ability: Ordinary models may lack generalization ability because they cannot effectively learn important information from the input data, making it difficult to adapt to different tasks and domains;
[0011] 7. Poor model interpretability: Ordinary models often have difficulty explaining their decision-making process and cannot provide a clear explanation of why certain features or parts are selected. Summary of the Invention
[0012] This invention provides a method and system for detecting landmarks on lateral cephalometric X-ray images, addressing the technical problem of overcoming the aforementioned issues with current deep learning models in the field of cephalometric analysis.
[0013] To solve the above technical problems, the present invention provides a method for detecting landmarks on a lateral cephalometric X-ray, comprising the following steps:
[0014] S1. Construct a U-shaped hybrid attention network model;
[0015] The U-shaped hybrid attention network model includes an input layer, a U-shaped feature extraction module, and an output layer;
[0016] The input layer is used to convert the number of channels of the input lateral cephalometric X-ray from 3 to C;
[0017] The U-shaped feature extraction module adopts a U-shaped network architecture, including an encoder and a decoder. Each layer of the encoder includes a multi-head attention convolution module and a downsampling module; each layer of the decoder includes an upsampling module and a concatenation module. The encoder and decoder are linked by a channel convolution module. The multi-head attention convolution module extracts the multi-head attention of its own input feature map, concatenates them, and then performs pixel-level convolution to obtain the output feature map. The downsampling module increases the number of channels in its own input feature map. The channel convolution module obtains the attention weights of its own input feature map. The upsampling module decreases the number of channels in its own input feature map. The concatenation module concatenates the attention weights of the corresponding layer with the feature map output by the upsampling module.
[0018] The output layer is used to find and mark the maximum value position in each channel of the output feature map of the U-shaped feature extraction module, and then convert the number of channels to 3 to obtain the corresponding cephalometric X-ray landmark image.
[0019] S2. Train the U-shaped hybrid attention network model;
[0020] S3. The trained U-shaped hybrid attention network model is used to detect landmarks on the lateral cephalometric X-ray images to be detected.
[0021] Specifically, the channel convolution module includes a channel splicing unit, a reshaping summation unit, an average pooling layer, a multilayer perceptron, a SoftMax layer, and a weight reshaping unit;
[0022] The channel stitching unit is used to stitch its own input feature map along the channel dimension and then input it into the reshaping summation unit;
[0023] The reshaping and summing unit is used to reshape its own input feature map into a five-dimensional tensor, which is represented as B×self.height×C×H×W, where H and W represent the height and width of the lateral cephalometric X-ray film input to the input layer, B represents the batch size, and self.height represents the number of input feature maps; the reshaping and summing unit is also used to sum the feature maps in the vertical dimension to obtain the total feature map;
[0024] The average pooling layer is used to perform adaptive max pooling on the total feature map, reducing the total feature map to a size of 1×1;
[0025] The multilayer perceptron is used to perform multilayer perceptron processing on the pooled total feature map, and the SoftMax layer is used to obtain attention weights.
[0026] The weight reshaping unit is used to reshape the attention weights acquired by the multilayer perceptron into a shape of B×self.height×C×1×1.
[0027] Specifically, the encoder has a 4-layer encoding structure. The 4-layer encoding structure has 4 downsampling modules that perform 4 downsampling operations. During each downsampling, the height and width of the image are halved, and the number of channels is doubled.
[0028] The decoder has a 4-layer decoding structure. The 4-layer decoding structure has 4 upsampling modules that perform 4 upsampling operations. Each time it is upsampled, the height and width of the image are doubled and the number of channels is halved.
[0029] Specifically, the multi-head attention convolution module includes a multi-head attention layer and a point convolution layer. The multi-head attention layer is placed before the point convolution layer. The multi-head attention layer is used to extract and concatenate multi-head attention, and the point convolution layer is used to perform point convolution on the output of the multi-head attention layer.
[0030] Specifically, the output layer includes a 1×1 convolutional layer, which is used to convert the number of channels of the multi-channel feature map output by the U-shaped feature extraction module back into the final required number of channels, 3.
[0031] Specifically, in step S2, the loss function used is the pixel-wise binary cross-entropy loss, expressed as:
[0032]
[0033] Among them, F i mha Y represents the i-th feature map output by the model. i Indicates with F i mha The corresponding real image, where y represents a landmark in the real image, f represents the probability of predicting the landmark from a pixel, and L i This represents the calculated loss.
[0034] The present invention also provides a landmark detection system for lateral cephalometric X-ray images, the key feature of which is: it includes an intelligent agent, on which the trained U-shaped hybrid attention network model is mounted.
[0035] This invention provides a method and system for detecting landmarks in lateral cephalometric X-ray images. It constructs a U-shaped hybrid attention network model for detecting landmarks in lateral cephalometric X-ray images. The U-shaped hybrid attention network model first uses multi-layer multi-head attention pixel convolution and multi-layer downsampling to project the input image into multiple subspaces. Then, it uses corresponding multi-layer upsampling to project the downsampled image into a higher resolution space. The downsampling and upsampling are fused through a channel attention mechanism.
[0036] The U-shaped hybrid attention network model proposed in this invention offers more interpretable weight allocation, taking into account the varying importance of different features in lateral cephalometric radiographs. The downsampling component in the hybrid attention mechanism more effectively captures the correlations between highly relevant features. The network exhibits stronger generalization performance, eliminating the need for fundamental modifications for different datasets, and consistently performs optimally on various datasets, making it adaptable to different tasks and domains. Furthermore, this invention employs channel attention for skip links, offering the following advantages compared to ordinary skip links:
[0037] Channel attention enhances feature selectivity: Channel attention allows the model to dynamically learn the weights of each encoder channel so that the encoder channels can be used more precisely and selectively in the decoder. This means that in the decoder, the model can focus more on the encoder feature channels that are most important to the current task, thereby improving the model's performance.
[0038] Adaptive Feature Fusion: Channel attention allows the decoder to adaptively select encoder feature channels to fuse, unlike traditional skip connections, which are typically hard-coded. This adaptive feature fusion can better adapt to different tasks and variations in input data, improving the model's generalization ability.
[0039] Reduced parameters and computational cost: Channel attention mechanisms typically require fewer parameters than hard-jump links and can be implemented without introducing additional convolutional layers. This reduces the computational cost of the model, especially for deep models, and helps accelerate the training and inference processes.
[0040] Suitable for multi-scale tasks: Channel attention can be more easily extended to multi-scale tasks because it can perform adaptive feature fusion between feature representations at different scales, thereby improving the model's sensitivity to information at different scales.
[0041] Adaptability and interpretability: Channel attention mechanisms can provide the model with adaptability, enabling it to automatically select the feature channels to focus on based on the input and task. At the same time, the channel attention weights are interpretable because they reflect the importance of each feature channel to the task, which is helpful for the interpretability of the model. Attached Figure Description
[0042] Figure 1 This is an architecture diagram of the U-shaped hybrid attention network model provided in an embodiment of the present invention;
[0043] Figure 2 This is a structural diagram of the multi-head attention convolution module provided in an embodiment of the present invention;
[0044] Figure 3 This is a structural diagram of the channel convolution module provided in an embodiment of the present invention;
[0045] Figure 4 This is a map showing the landmark detection results of an adult dataset and an exemplary adult dataset, provided by an embodiment of the present invention. Detailed Implementation
[0046] The embodiments of the present invention are described in detail below with reference to the accompanying drawings. The embodiments are given for illustrative purposes only and should not be construed as limiting the present invention. The accompanying drawings are for reference and illustration only and do not constitute a limitation on the scope of patent protection of the present invention, because many changes can be made to the present invention without departing from the spirit and scope of the present invention.
[0047] This invention provides a method for detecting landmarks on a lateral cephalometric X-ray, comprising the following steps:
[0048] S1. Construct a U-shaped hybrid attention network model;
[0049] S2. Train the U-shaped hybrid attention network model;
[0050] S3. The trained U-shaped hybrid attention network model is used to detect landmarks on the lateral cephalometric X-ray images to be detected.
[0051] like Figure 1 As shown, the U-shaped hybrid attention network model constructed in step S1 includes an input layer, a U-shaped feature extraction module, and an output layer.
[0052] The input layer is used to convert the number of channels of the input lateral cephalometric X-ray (512×512×3) from 3 to C(64).
[0053] The U-shaped feature extraction module adopts a U-shaped network architecture, including an encoder and a decoder. Each layer of the encoder includes a multi-head attention convolution module and a downsampling module; each layer of the decoder includes an upsampling module and a concatenation module. The encoder and decoder layers are skipped by channel convolution modules. The multi-head attention convolution module extracts the multi-head attention of its own input feature map, concatenates them, and then performs pixel-level convolution to obtain the output feature map. The downsampling module increases the number of channels of its own input feature map. The channel convolution module obtains the attention weights of its own input feature map. The upsampling module reduces the number of channels of its own input feature map. The concatenation module concatenates the attention weights of the corresponding layers with the feature map output by the upsampling module.
[0054] The output layer is used to find and mark the maximum value location in each channel of the output feature map of the U-shaped feature extraction module, and then convert the number of channels to 3 to obtain the corresponding cephalometric X-ray landmark image.
[0055] The encoder extracts local features from the input lateral cephalometric X-ray image. These local features primarily refer to the features projected onto various subspaces of the image matrix through a multi-head attention mechanism. Skip connections are a mechanism that connects features from different levels of the encoder to corresponding levels of the decoder. In architectures such as U-Net, each decoder level corresponds to an encoder level, and these connections allow the decoder to access feature representations at different levels within the encoder. This ensures that low-level feature information from the encoder is preserved during decoding. This helps overcome the vanishing gradient problem while allowing the decoder to access both low- and high-level feature representations from the encoder. The decoder utilizes channel attention to combine features previously ignored with the extracted local features and outputs a feature heatmap. Specifically, attention scores are calculated using a multi-head attention mechanism and merged using pixel-level convolutions.
[0056] Specifically, such as Figure 1 As shown, the output layer includes a 1×1 convolutional layer, which is used to convert the number of channels of the multi-channel feature map output by the U-shaped feature extraction module back into the final required number of channels, 3.
[0057] 1×1 convolutions can learn the correlations between different channels, achieving feature fusion. Compared to larger convolutional kernels, 1×1 convolutions require fewer computational resources, reducing computational burden while maintaining network performance.
[0058] like Figure 2 As shown, the encoder has a 4-layer coding structure (U-net has been tested and verified to have the best effect, so the default is a 4-layer structure). The 4-layer coding structure has 4 downsampling modules that perform 4 downsampling operations. In each downsampling operation, the height and width of the image are halved and the number of channels is doubled.
[0059] The decoder has a 4-layer decoding structure. The 4-layer decoding structure has 4 upsampling modules that perform 4 upsampling operations. Each time it is upsampled, the height and width of the image are doubled and the number of channels is halved.
[0060] like Figure 2 As shown, the multi-head attention convolution module uses a multi-head attention mechanism to calculate the multi-head attention score (different from ordinary convolution operation, which projects the data to multiple different subspaces, i.e., features are considered in the heads), and then uses pixel-level convolution to merge them. The multi-head attention convolution module includes a multi-head attention layer and a point convolution layer. The multi-head attention layer is placed before the point convolution layer. The multi-head attention layer first captures the patterns and relationships in the input feature map data, which improves the model's ability to recognize data features (spatial features and long-distance dependencies).
[0061] like Figure 3 As shown, the channel convolution module includes a channel splicing unit, a reshaping summation unit, an average pooling layer (AvgPool), a multilayer perceptron (MLP), a SoftMax layer, and a weight reshaping unit.
[0062] The channel stitching unit is used to stitch its own input feature map along the channel dimension and then input it into the reshaping summation unit;
[0063] The reshaping summation unit reshapes its input feature map into a five-dimensional tensor, represented as B×self.height×C×H×W, where H and W represent the height and width of the lateral cephalometric X-ray image input to the input layer, B represents the batch size (representing the number of data samples processed simultaneously in one forward propagation), self.height represents the number of input feature maps (used for feature fusion), and C represents the number of channels, representing the depth of the feature map. The reshaping summation unit is also used to sum the feature maps along their vertical dimension to obtain the total feature map.
[0064] The average pooling layer is used to perform adaptive max pooling on the total feature map, reducing the total feature map to a size of 1×1.
[0065] The multilayer perceptron is used to perform multilayer perceptron processing on the pooled total feature map, and the SoftMax layer is used to obtain attention weights.
[0066] The weight reshaping unit is used to reshape the attention weights acquired by the multilayer perceptron into a shape of B×self.height×C×1×1.
[0067] In step S2, the loss function used is the pixel-wise binary cross-entropy loss, expressed as:
[0068]
[0069] Among them, F i mha Y represents the i-th feature map output by the model. i Indicates with F i mha The corresponding real image, where y represents a landmark in the real image, f represents the probability of predicting the landmark from a pixel, and L i This indicates the calculated loss.
[0070] After training, the model can accept new input images and generate corresponding heatmaps F. mha In the reasoning phase, through F mha By finding the location of the maximum value on each channel, the model can determine the exact location of each landmark. This method allows the model to accurately identify and label landmarks in images, and is suitable for various scenarios requiring precise landmark identification, such as medical image analysis or geographic information systems.
[0071] To facilitate implementation, this invention also provides a cephalometric X-ray landmark detection system, which includes an agent equipped with a trained U-shaped hybrid attention network model. In practical applications, the cephalometric X-ray image to be detected is simply input into the U-shaped hybrid attention network model, and the corresponding annotation results are output.
[0072] This invention provides a method and system for detecting landmarks in lateral cephalometric X-ray images. It constructs a U-shaped hybrid attention network model for detecting landmarks in lateral cephalometric X-ray images. The U-shaped hybrid attention network model first uses multi-layer multi-head attention pixel convolution and multi-layer downsampling to project the input image into multiple subspaces. Then, it uses corresponding multi-layer upsampling to project the downsampled image into a higher resolution space. The downsampling and upsampling are fused through a channel attention mechanism.
[0073] This invention provides a method and system for detecting landmarks in lateral cephalometric X-ray images. It constructs a U-shaped hybrid attention network model for detecting landmarks in lateral cephalometric X-ray images. The U-shaped hybrid attention network model first uses multi-layer multi-head attention pixel convolution and multi-layer downsampling to project the input image into multiple subspaces. Then, it uses corresponding multi-layer upsampling to project the downsampled image into a higher resolution space. The downsampling and upsampling are fused through a channel attention mechanism.
[0074] The U-shaped hybrid attention network model proposed in this invention provides more interpretable weight allocation, taking into account the differences in the importance of different features in lateral cephalometric radiographs. The downsampling component in the hybrid attention mechanism can more effectively capture the correlation between highly relevant features. The network exhibits stronger generalization performance, eliminating the need for fundamental modifications for different datasets, and consistently performs optimally on various datasets, making it adaptable to different tasks and domains. Furthermore, this invention employs channel attention for skip links, offering the following advantages compared to ordinary skip links:
[0075] Channel attention enhances feature selectivity: Channel attention allows the model to dynamically learn the weights of each encoder channel so that the encoder channels can be used more precisely and selectively in the decoder. This means that in the decoder, the model can focus more on the encoder feature channels that are most important to the current task, thereby improving the model's performance.
[0076] Adaptive Feature Fusion: Channel attention allows the decoder to adaptively select encoder feature channels to fuse, which differs from traditional skip connections, which are usually hard-coded connections. This adaptive feature fusion can better adapt to different tasks and changes in input data, improving the model's generalization ability.
[0077] Reduced parameters and computational cost: Channel attention mechanisms typically require fewer parameters than hard-jump links and can be implemented without introducing additional convolutional layers, which reduces the computational cost of the model, especially for deep models, and helps to accelerate the training and inference process.
[0078] Suitable for multi-scale tasks: Channel attention can be more easily extended to multi-scale tasks because it can perform adaptive feature fusion between feature representations at different scales to improve the model’s sensitivity to information at different scales.
[0079] Adaptability and interpretability: Channel attention mechanisms can provide the model with adaptability, enabling it to automatically select the feature channels to focus on based on the input and task. At the same time, the channel attention weights are interpretable because they reflect the importance of each feature channel to the task, which is helpful for the interpretability of the model.
[0080] The present invention was compared with other models, and the results are shown in Table 1. In Table 1, rows 1 to 4 are ordinary algorithm iterations that discard the proposed algorithm, row 5 is Unet, row 6 is GU2net, and row 7 corresponds to the network model proposed in the present invention.
[0081] Table 1
[0082]
[0083] In Table 1, MRE stands for Mean Radial Error, representing the mean error between the predicted point and the true value, expressed in millimeters. The four columns on the right show the accuracy of the predicted points within the 2-4 mm range. Table 1 demonstrates the effectiveness of our proposed algorithm (highest accuracy in the 2-4 mm range) and the significant improvement in accuracy within the 2 mm range.
[0084] Figure 4 This image visualizes the prediction points of the proposed network model over 0-25 epochs on an adult dataset (first row) and a children's dataset (second row). The MRE value is shown in the upper right corner. Figure 4 As can be seen, the mean radial error decreases continuously with the increase of epochs, and it is effective not only on adult datasets but also on children's datasets, demonstrating the effectiveness and generalization of the network model proposed in this invention.
[0085] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.
Claims
1. A method for detecting landmarks on a lateral cephalometric X-ray film, characterized in that, Including the following steps: S1. Construct a U-shaped hybrid attention network model; The U-shaped hybrid attention network model includes an input layer, a U-shaped feature extraction module, and an output layer; The input layer is used to convert the number of channels of the input lateral cephalometric X-ray from 3 to C; The U-shaped feature extraction module adopts a U-shaped network architecture, including an encoder and a decoder. Each layer of the encoder includes a multi-head attention convolution module and a downsampling module; each layer of the decoder includes an upsampling module and a concatenation module. The encoder and decoder are linked by a channel convolution module. The multi-head attention convolution module extracts the multi-head attention of its own input feature map, concatenates them, and then performs pixel-level convolution to obtain the output feature map. The downsampling module increases the number of channels in its own input feature map. The channel convolution module obtains the attention weights of its own input feature map. The upsampling module decreases the number of channels in its own input feature map. The concatenation module concatenates the attention weights of the corresponding layer with the feature map output by the upsampling module. The output layer is used to find and mark the maximum value position in each channel of the output feature map of the U-shaped feature extraction module, and then convert the number of channels to 3 to obtain the corresponding cephalometric X-ray landmark image; the channel convolution module includes a channel stitching unit, a reshaping summation unit, an average pooling layer, a multilayer perceptron, a SoftMax layer and a weight reshaping unit. The channel stitching unit is used to stitch its own input feature map along the channel dimension and then input it into the reshaping summation unit; The reshaping and summing unit is used to reshape its own input feature map into a five-dimensional tensor, which is represented as B×self.height×C×H×W, where H and W represent the height and width of the lateral cephalometric X-ray film input to the input layer, B represents the batch size, and self.height represents the number of input feature maps; the reshaping and summing unit is also used to sum the feature maps in the vertical dimension to obtain the total feature map; The average pooling layer is used to perform adaptive max pooling on the total feature map, reducing the total feature map to a size of 1×1; The multilayer perceptron is used to perform multilayer perceptron processing on the pooled total feature map, and the SoftMax layer is used to obtain attention weights. The weight reshaping unit is used to reshape the attention weights acquired by the multilayer perceptron into a shape of B×self.height×C×1×1; The encoder has a 4-layer encoding structure. The 4-layer encoding structure has 4 downsampling modules that perform 4 downsampling operations. During each downsampling, the height and width of the image are halved, and the number of channels is doubled. The decoder has a 4-layer decoding structure. The 4-layer decoding structure has 4 upsampling modules that perform 4 upsampling operations. Each time it is upsampled, the height and width of the image are doubled and the number of channels is halved. The multi-head attention convolution module includes a multi-head attention layer and a point convolution layer. The multi-head attention layer is placed before the point convolution layer. The multi-head attention layer is used to extract and concatenate multi-head attention, and the point convolution layer is used to perform point convolution on the output of the multi-head attention layer. S2. Train the U-shaped hybrid attention network model; S3. The trained U-shaped hybrid attention network model is used to detect landmarks on the lateral cephalometric X-ray images to be detected.
2. The method for detecting landmarks on a lateral cephalometric X-ray film according to claim 1, characterized in that: The multilayer perceptron consists of two convolutional layers, which are used to perform convolution processing on the feature maps input from the front, i.e., feature extraction.
3. The method for detecting landmarks on a lateral cephalometric X-ray film according to claim 1, characterized in that: The output layer includes Convolutional layer, the The convolutional layer is used to convert the number of channels in the multi-channel feature map output by the U-shaped feature extraction module back into the final required number of channels, 3.
4. The method for detecting landmarks on a lateral cephalometric X-ray film according to claim 1, characterized in that, In step S2, the loss function used is the pixel-wise binary cross-entropy loss, expressed as: , in, This represents the i-th feature map output by the model. Indicates and The corresponding real-world image, where y represents a landmark in the real image, and f represents the probability of predicting the landmark from a pixel. This indicates the calculated loss.
5. A system for detecting landmarks on lateral cephalometric X-ray films, characterized in that: It includes an agent equipped with a U-shaped hybrid attention network model trained according to any one of claims 1 to 4.
Citation Information
Patent Citations
Skull side position film key point automatic detection system and method based on deep learning
CN114820517A
Blood vessel image segmentation method and system based on three-dimensional deep network
CN115546570A
Hybrid attention retinal vessel segmentation method based on residual U-shaped network
CN116363060A