Head shadow key point detection method and system based on interactive attention
Through the interactive attention-based head shadow key point detection method, the U-shaped structure of the encoder and decoder and the global and window self-attention modules are utilized to automatically identify the head shadow key points, which solves the problems of inaccurate and inconsistent detection results in traditional methods and achieves efficient and accurate head shadow detection.
Patent Information
- Application Number
- CN202510478498.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-09-12
AI Technical Summary
Traditional head shadow key point detection methods rely on manual identification, resulting in inaccurate and inconsistent detection results, especially among inexperienced clinicians.
A head shadow key point detection method based on interactive attention is adopted. Through the U-shaped structure of encoder and decoder, combined with the global attention aggregation module and the window self-attention module, a head shadow key point detection model is constructed. Multi-scale global attention and long-distance dependencies are used to automatically identify key points, and the model training process is optimized through the loss function.
It significantly improves the accuracy and consistency of head shadow detection, reduces the inconsistency of detection results caused by manual operation, improves computing efficiency and model generalization ability, is suitable for resource-constrained devices, and enhances the detection ability of different populations and pathological characteristics.
Smart Images

Figure CN120634944A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image detection, and in particular to a head shadow key point detection method and system based on interactive attention. Background Art
[0002] Medical imaging is dedicated to using various imaging technologies to acquire, process and display the internal structure and function of the human body. These medical images are one of the most important diagnostic and treatment methods in modern medicine, and can provide doctors with detailed information about diseases, abnormalities and organ status, thereby supporting diagnosis, treatment and monitoring of patients' disease progression.
[0003] Cephalometric landmark detection (measuring), first proposed by Broadbent and Hofrath in 1931, has become an essential step in the orthodontic diagnostic process and involves identifying and locating anatomical landmarks on lateral skull radiographs. These landmarks are then used to measure distances, angles, and ratios, providing insights into the structural characteristics and interrelationships of craniofacial soft and hard tissues.
[0004] However, landmark detection on lateral skull X-ray images is a significant challenge. Traditional methods for detecting key points in the cephalogram rely on manual identification and labeling of landmarks, a process that is not only time-consuming and labor-intensive, but also prone to errors due to subjective visual interpretation and differences in clinician experience levels, leading to inconsistent test results during repeated assessments. To improve detection accuracy and consistency, some clinicians have chosen computer-assisted semi-automatic cephalogram detection methods, in which clinicians manually identify relevant landmarks and software performs distance, angle, and ratio measurements. However, regardless of the method used, manual landmark identification is necessary and can lead to inaccuracies, especially among less experienced clinicians.
[0005] Therefore, how to provide a head shadow key point detection method and system based on interactive attention to improve the accuracy and consistency of head shadow detection has become a technical problem that needs to be solved urgently. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a head shadow key point detection method and system based on interactive attention, so as to improve the accuracy and consistency of head shadow detection.
[0007] In a first aspect, the present invention provides a head shadow key point detection method based on interactive attention, comprising the following steps:
[0008] Step S1: obtaining a large number of historical head images, preprocessing and annotating each of the historical head images to construct a data set, and dividing the data set into a training set, a test set, and a validation set;
[0009] Step S2: creating a head shadow key point detection model based on the encoder, decoder, and output module, and setting the loss function and hyperparameters of the head shadow key point detection model;
[0010] The encoder and decoder are jump-connected to form a U-shaped structure; the encoder is equipped with a global attention aggregation module and a grouped convolution module; the decoder is composed of pure linear layers and is equipped with a windowed self-attention module;
[0011] Step S3, training the head shadow key point detection model using the training set, testing the trained head shadow key point detection model using the test set, and verifying the head shadow key point detection model that passes the test using the validation set;
[0012] Step S4: deploying the verified head shadow key point detection model, and detecting the head shadow key points using the deployed head shadow key point detection model.
[0013] Furthermore, the step S1 is specifically as follows:
[0014] A large number of historical cephalograms of different ages, genders, malocclusion patterns, and skeletal patterns are obtained, and each of the historical cephalograms is preprocessed, including at least geometric transformation, grayscale transformation, noise reduction, and image enhancement. Two people respectively annotate each of the preprocessed historical cephalograms with a preset number of key points, and the average of the two people's annotation results is taken as the true label to complete the annotation. A data set is constructed based on the annotated historical cephalograms, and the data set is divided into a training set, a test set, and a validation set based on a preset ratio.
[0015] Furthermore, in step S2, the encoder is used to downsample the input head image to obtain a downsampled image; the decoder is used to upsample each downsampled image to obtain an original size image; the global attention aggregation module is used to extract multi-scale global attention during the downsampling process; the grouped convolution module is provided in the last layer of the encoder to improve computational efficiency; the windowed self-attention module is used to extract long-distance dependencies during the upsampling process; the output module is used to identify key points based on interactive attention including multi-scale global attention and long-distance dependencies, mark the key points on the original size image and output them as a heat map;
[0016] The loss function is constructed based on the weighted combination of key point position loss function, key point visibility loss function, bounding box loss function and classification loss function.
[0017] Furthermore, the step S3 is specifically as follows:
[0018] The head shadow key point detection model is trained using the training set, and during the training process, the head shadow key point detection model is continuously optimized, including at least hyperparameters of a learning rate, a regularization parameter, a network structure parameter, an activation function parameter, an optimizer parameter, and a random dropout rate, until a loss value of the loss function is less than a preset loss threshold;
[0019] The detection accuracy is calculated using the test set to test the trained head shadow key point detection model. If the test fails, the data set is expanded to continue training. If the test passes, then:
[0020] The confidence level is calculated using the validation set to validate the head shadow key point detection model that has passed the test. If the validation fails, the data set is expanded to continue training; if the validation passes, the training ends.
[0021] Furthermore, the step S4 is specifically as follows:
[0022] The verified head shadow key point detection model is pruned and knowledge distilled and then deployed to the medical terminal. The authentication mechanism of the head shadow key point detection model is set. After authentication based on the authentication mechanism, the input real-time head shadow image is input into the deployed head shadow key point detection model, and a heat map with key points marked is output to detect the head shadow key points. The head shadow key point detection model is continuously optimized based on the marking accuracy of the heat map.
[0023] In a second aspect, the present invention provides a head shadow key point detection system based on interactive attention, comprising the following modules:
[0024] A data set construction module is used to obtain a large number of historical head images, pre-process and annotate each of the historical head images to construct a data set, and divide the data set into a training set, a test set, and a validation set;
[0025] A head shadow key point detection model creation module is used to create a head shadow key point detection model based on the encoder, decoder and output module, and set the loss function and hyperparameters of the head shadow key point detection model;
[0026] The encoder and decoder are jump-connected to form a U-shaped structure; the encoder is equipped with a global attention aggregation module and a grouped convolution module; the decoder is composed of pure linear layers and is equipped with a windowed self-attention module;
[0027] a head shadow key point detection model training module, configured to train the head shadow key point detection model using the training set, test the trained head shadow key point detection model using the test set, and verify the head shadow key point detection model that passes the test using the verification set;
[0028] The head shadow key point detection module is used to deploy the verified head shadow key point detection model and detect the head shadow key points using the deployed head shadow key point detection model.
[0029] Furthermore, the dataset construction module is specifically used to:
[0030] A large number of historical cephalograms of different ages, genders, malocclusion patterns, and skeletal patterns are obtained, and each of the historical cephalograms is preprocessed, including at least geometric transformation, grayscale transformation, noise reduction, and image enhancement. Two people respectively annotate each of the preprocessed historical cephalograms with a preset number of key points, and the average of the two people's annotation results is taken as the true label to complete the annotation. A data set is constructed based on the annotated historical cephalograms, and the data set is divided into a training set, a test set, and a validation set based on a preset ratio.
[0031] Furthermore, in the head shadow key point detection model creation module, the encoder is used to downsample the input head shadow image to obtain a downsampled image; the decoder is used to upsample each downsampled image to obtain an original size image; the global attention aggregation module is used to extract multi-scale global attention during the downsampling process; the grouped convolution module is provided at the last layer of the encoder to improve computational efficiency; the window self-attention module is used to extract long-distance dependencies during the upsampling process; the output module is used to identify key points based on interactive attention including multi-scale global attention and long-distance dependencies, mark the key points on the original size image and output them as a heat map;
[0032] The loss function is constructed based on the weighted combination of key point position loss function, key point visibility loss function, bounding box loss function and classification loss function.
[0033] Furthermore, the head shadow key point detection model training module is specifically used to:
[0034] The head shadow key point detection model is trained using the training set, and during the training process, the head shadow key point detection model is continuously optimized, including at least hyperparameters of a learning rate, a regularization parameter, a network structure parameter, an activation function parameter, an optimizer parameter, and a random dropout rate, until a loss value of the loss function is less than a preset loss threshold;
[0035] The detection accuracy is calculated using the test set to test the trained head shadow key point detection model. If the test fails, the data set is expanded to continue training. If the test passes, then:
[0036] The confidence level is calculated using the validation set to validate the head shadow key point detection model that has passed the test. If the validation fails, the data set is expanded to continue training; if the validation passes, the training ends.
[0037] Furthermore, the head shadow key point detection module is specifically used to:
[0038] The verified head shadow key point detection model is pruned and knowledge distilled and then deployed to the medical terminal. The authentication mechanism of the head shadow key point detection model is set. After authentication based on the authentication mechanism, the input real-time head shadow image is input into the deployed head shadow key point detection model, and a heat map with key points marked is output to detect the head shadow key points. The head shadow key point detection model is continuously optimized based on the marking accuracy of the heat map.
[0039] The advantages of the present invention are:
[0040] 1. By obtaining a large number of historical head shadow images, pre-processing and annotating each historical head shadow image, a data set is constructed, and the data set is divided into a training set, a test set, and a validation set; then a head shadow key point detection model is created based on the encoder, decoder, and output module, and the loss function and hyperparameters of the head shadow key point detection model are set; the encoder and the decoder are jump-connected to form a U-shaped structure; the encoder is equipped with a global attention aggregation module and a group convolution module; the decoder is composed of a pure linear layer and is equipped with a window self-attention module; then the head shadow key point detection model is trained through the training set, the trained head shadow key point detection model is tested through the test set, and the head shadow key point detection model that has passed the test is verified through the validation set. Finally, the verified head shadow key point detection model is deployed, and the head shadow key point detection model is used for head shadow key point detection. Detection; that is, head shadow key point detection is performed through a pre-trained head shadow key point detection model created based on an encoder, a decoder and an output module. The encoder extracts multi-scale global attention through a global attention aggregation module during the downsampling process, and the decoder extracts long-distance dependencies through a window self-attention module during the upsampling process. The output module identifies key points based on the interactive attention including multi-scale global attention and long-distance dependencies, and marks the key points on the original size image output by the decoder and outputs them as a heat map. By fusing multi-scale global attention and long-distance dependencies, the feature extraction capability can be effectively improved, and global features and detail information can be effectively learned. Automatic detection through the head shadow key point detection model can avoid the problem of inconsistent detection results during repeated evaluations caused by traditional manual operations, and ultimately greatly improve the accuracy and consistency of head shadow detection.
[0041] 2. Two people respectively annotate a preset number of key points on each pre-processed historical head shadow image, and take the average of the two people's annotation results as the true label, which effectively improves the accuracy of the annotation and thus effectively improves the quality of the dataset, thereby greatly improving the training effect of the head shadow key point detection model.
[0042] 3. By setting the decoder to consist of pure linear layers, the receptive field of the head shadow key point detection model (neural network) is effectively expanded while ensuring the accuracy of head shadow key point detection.
[0043] 4. By introducing a grouped convolution module in the last layer of the encoder, the grouped convolution module divides the input channels into multiple groups and performs convolution operations in each group, thereby significantly reducing the amount of calculation and the number of parameters, thereby greatly improving the computational efficiency.
[0044] 5. By setting the loss function based on the weighted construction of key point position loss function, key point visibility loss function, bounding box loss function and classification loss function, the key point position loss function is used to measure the difference between the predicted key point position and the true position, the key point visibility loss function is used to measure the difference between the predicted key point visibility and the true visibility, the bounding box loss function is used to measure the difference between the predicted bounding box and the true bounding box, and the classification loss function is used to measure the difference between the predicted category and the true category. That is, the advantages of different loss functions are integrated, and the contributions between different losses are balanced by weights, which further improves the training effect of the head shadow key point detection model.
[0045] 6. The head shadow key point detection model is trained using the training set. During the training process, the hyperparameters including learning rate, regularization parameter, network structure parameter, activation function parameter, optimizer parameter and random dropout rate are continuously optimized until the loss value of the loss function is less than the preset loss threshold. The detection accuracy is calculated using the test set to test the trained head shadow key point detection model, and the confidence is calculated using the validation set to verify the head shadow key point detection model that has passed the test. After deployment and use, the head shadow key point detection model is continuously optimized based on the identification accuracy of the heat map. That is, the head shadow key point detection model is continuously optimized, tested and trained during the training process, and continuously optimized after it is put into use, thereby greatly improving the accuracy of head shadow detection.
[0046] 7. By performing pruning and knowledge distillation before deploying the head shadow key point detection model, the model size is effectively reduced, making it easier to deploy on resource-constrained devices, thereby greatly improving the scope of applicability.
[0047] 8. By setting up an authentication mechanism, the deployed head shadow key point detection model can be called for head shadow detection only after authentication through the authentication mechanism, preventing the head shadow key point detection model from being illegally called.
[0048] 9. By setting the dataset to include head images of different ages, genders, malocclusion patterns and skeletal patterns, the generalization ability of the head key point detection model for different populations and pathological characteristics is effectively improved.
[0049] 10. By setting the preprocessing of historical head images to include geometric transformation, grayscale transformation, noise reduction and image enhancement, the data robustness is enhanced and the risk of overfitting is reduced.
[0050] 11. A U-shaped skip connection structure is used to achieve multi-scale feature fusion, retaining underlying details and high-level semantic information, and improving the accuracy of key point positioning. By introducing a global attention aggregation module in the encoder, multi-scale global context is dynamically captured, enhancing the ability to model complex anatomical structures. The decoder uses a pure linear layer combined with a windowed self-attention module to extract long-distance dependencies while reducing computational complexity, balancing efficiency and performance.
[0051] 12. By continuously optimizing the hyperparameters including learning rate, regularization parameters, network structure parameters, activation function parameters, optimizer parameters and random dropout rate during the training process of the head shadow key point detection model, the model performance is adaptively adjusted to avoid local optimality. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0053] Figure 1 This is a flow chart of a head shadow key point detection method based on interactive attention of the present invention.
[0054] Figure 2 It is a structural schematic diagram of a head shadow key point detection system based on interactive attention of the present invention.
[0055] Figure 3 It is a schematic diagram of the process of detecting key points of head shadow of the present invention.
[0056] Figure 4 It is a structural diagram of the global attention aggregation module of the present invention.
[0057] Figure 5 Schematic diagram of multi-scale global attention extraction of the present invention.
[0058] Figure 6 Schematic diagram of the structure of the window self-attention module of the present invention.
[0059] Figure 7 It is a structural diagram of the convolutional feedforward neural network of the present invention. DETAILED DESCRIPTION
[0060] The technical solution in the embodiments of the present application has the following overall idea: head shadow key point detection is performed through a head shadow key point detection model created by an encoder, a decoder and an output module. The encoder extracts multi-scale global attention through a global attention aggregation module during the downsampling process. The decoder extracts long-distance dependencies through a window self-attention module during the upsampling process. The output module identifies key points based on interactive attention including multi-scale global attention and long-distance dependencies, marks the key points on the original-size image output by the decoder and outputs them as a heat map. By fusing multi-scale global attention and long-distance dependencies, the feature extraction capability can be effectively improved, and global features and detail information can be effectively learned. Automatic detection is performed through the head shadow key point detection model, which can avoid the problem of inconsistent detection results during repeated evaluations, thereby improving the accuracy and consistency of head shadow detection.
[0061] Please refer to Figures 1 to 7 As shown, a preferred embodiment of the head shadow key point detection method based on interactive attention of the present invention includes the following steps:
[0062] Step S1: obtaining a large number of historical head images, preprocessing and annotating each of the historical head images to construct a data set, and dividing the data set into a training set, a test set, and a validation set;
[0063] Step S2: creating a head shadow key point detection model based on the encoder, decoder, and output module, and setting the loss function and hyperparameters of the head shadow key point detection model;
[0064] The encoder and decoder are jump-connected to form a U-shaped structure; the encoder is equipped with a global attention aggregation module (GAA) and a grouped convolution module; the decoder is composed of pure linear layers and is equipped with a windowed self-attention module (WSA);
[0065] By setting the decoder to consist of pure linear layers, the receptive field of the head shadow key point detection model (neural network) is effectively expanded while ensuring the accuracy of head shadow key point detection.
[0066] By introducing a grouped convolution module in the last layer of the encoder, the grouped convolution module divides the input channels into multiple groups and performs convolution operations in each group, thereby significantly reducing the amount of calculation and the number of parameters, thereby greatly improving the computational efficiency.
[0067] Multi-scale feature fusion is achieved through a U-shaped jump connection structure, retaining underlying details and high-level semantic information, and improving the accuracy of key point positioning; by introducing a global attention aggregation module in the encoder, multi-scale global context is dynamically captured, enhancing the ability to model complex anatomical structures; through the decoder, a pure linear layer combined with a windowed self-attention module is used to extract long-distance dependencies while reducing computational complexity, balancing efficiency and performance.
[0068] Step S3, training the head shadow key point detection model using the training set, testing the trained head shadow key point detection model using the test set, and verifying the head shadow key point detection model that passes the test using the validation set;
[0069] Step S4: deploying the verified head shadow key point detection model, and detecting the head shadow key points using the deployed head shadow key point detection model.
[0070] The step S1 is specifically as follows:
[0071] A large number of historical cephalograms of different ages, genders, malocclusion patterns, and skeletal patterns are obtained, and each of the historical cephalograms is preprocessed, including at least geometric transformation, grayscale transformation, noise reduction, and image enhancement. Two people respectively annotate each of the preprocessed historical cephalograms with a preset number of key points, and the average of the two people's annotation results is taken as the true label to complete the annotation. A data set is constructed based on the annotated historical cephalograms, and the data set is divided into a training set, a test set, and a validation set based on a preset ratio.
[0072] Two people respectively annotate a preset number of key points on each preprocessed historical head shadow image, and take the average of the two people's annotation results as the true label, which effectively improves the accuracy of the annotation and thus effectively improves the quality of the dataset, thereby greatly improving the training effect of the head shadow key point detection model.
[0073] By setting the dataset to include cephalogram images of different ages, genders, malocclusion patterns and skeletal patterns, the generalization ability of the cephalogram key point detection model for different populations and pathological characteristics can be effectively improved.
[0074] By setting the preprocessing of historical head images to include geometric transformation, grayscale transformation, noise reduction and image enhancement, data robustness is enhanced and the risk of overfitting is reduced.
[0075] In step S2, the encoder is used to downsample the input head image to obtain a downsampled image; the decoder is used to upsample each downsampled image to obtain an original size image; the global attention aggregation module is used to extract multi-scale global attention during the downsampling process; the grouped convolution module is provided in the last layer of the encoder to improve computational efficiency; the windowed self-attention module is used to extract long-distance dependencies during the upsampling process; the output module is used to identify key points based on the interactive attention including multi-scale global attention and long-distance dependencies, mark the key points on the original size image and output them as a heat map;
[0076] That is, features are extracted from the head shadow image in layers, the image resolution is reduced at each layer, and global attention processing is performed on the image features of the downsampled 2 and 3 layers, and local attention processing is performed on the image features of the upsampled 3 and 4 layers. Features of multiple scales are fused through left and right staggered interactions, and then sent to the decoder of the same layer for upsampling and feature decoding. The position of each key point is predicted by generating a heat map.
[0077] The loss function is constructed based on the weighted combination of key point position loss function, key point visibility loss function, bounding box loss function and classification loss function.
[0078] By setting the loss function based on the weighted construction of key point position loss function, key point visibility loss function, bounding box loss function and classification loss function, the key point position loss function is used to measure the difference between the predicted key point position and the true position, the key point visibility loss function is used to measure the difference between the predicted key point visibility and the true visibility, the bounding box loss function is used to measure the difference between the predicted bounding box and the true bounding box, and the classification loss function is used to measure the difference between the predicted category and the true category. That is, the advantages of different loss functions are integrated, and the contributions between different losses are balanced by weights, which further improves the training effect of the head shadow key point detection model.
[0079] The global attention aggregation module performs dynamic position encoding of convolution on the input data (head image) to capture sequential dependencies and enhance feature expression capabilities. It then standardizes the data through layer normalization to improve the training stability of the model. It then performs feature extraction through local attention operations to enhance the expression capabilities of local information. Finally, it uses convolution transformation operations to further aggregate multi-scale features to improve the feature extraction and generalization capabilities of the model.
[0080] The calculation process of convolutional dynamic position encoding (CDPE: Convolutional Dynamic Position Encoding) is:
[0081] For the input labeled tensor Xin ∈R C×H×W , the convolutional dynamic position encoding (CDPE: Convolutional Dynamic Position Encoding) dynamically adds position information to all tags through a 3×3, stride 1, and fill 1 depth separation convolution calculation, aiming to enhance the spatial position information of the feature map, and then through the residual connection (X in +DWC(·)) alleviates the gradient vanishing problem and ultimately enhances the spatial position information while maintaining the size of the feature map, providing support for the precise positioning of key points. The calculation definition formula of CDPE is expressed as:
[0082] X=X in +DWC(Pad(X in , p);
[0083] Among them, X in ∈R C×H×W Represents the input image, R is a real number, C is the number of channels, H is the image height, and W is the image width; Pad represents padding; p represents the number of paddings; X represents the dynamic position encoding of the convolution.
[0084] Multi-Scale Global Attention (MSG) is a global feature extraction mechanism for head key point detection. It is particularly suitable for capturing features with relatively stable spatial distribution, such as the skull, and aims to achieve a balance between local details and global structural features. In the head key point detection task, it not only relies on the meticulous capture of local details, but also requires the model to have the ability to capture global features and provide rich contextual information to avoid errors caused by isolated analysis of local information. However, the global attention mechanism of traditional VIT incurs high computational costs. On large-scale datasets, VIT can automatically learn these features through a large amount of data, but on small-scale datasets, the model cannot fully learn the spatial structural information in the image, resulting in a decrease in the model's generalization ability. MSG achieves global feature extraction by introducing an aggregation-based global attention mechanism, making the model adaptable and stable on datasets of different scales.
[0085] MSG extracts input features by performing multi-scale convolution operations and combines them with a global attention mechanism to achieve multi-level feature fusion. Specifically, MSG uses convolution kernels of different sizes (such as 4×4, 8×8, and 12×12) to replace traditional convolution with a combination of depthwise convolution (DWC) and pointwise convolution (PWC) to obtain multi-scale features and effectively reduce the computational complexity of the global attention mechanism (GA). The number of output channels of PWC is set to one-quarter of the number of input feature channels. At the same time, MSC reduces the spatial size of the feature map through downsampling operations, sets a larger convolution kernel (greater than 3), and a stride of half the size of the convolution kernel, thereby alleviating the problem of loss of detail information that may result from the downsampling process. After MSC processing, each multi-scale feature map x i Calculate the global attention (GA), the process is defined as follows:
[0086]
[0087] in, Represents the output of the multi-scale feature map; represents the output of GA; Q(·), K(·) and V(·) are linear projection operations; W q , W k and W v represents the learning parameter; d is equal to C, which represents the channel dimension.
[0088] Since the stride of MSC is set to be large, some local details will be lost during the feature downsampling process. To alleviate this problem, a cross attention (CA) mechanism is introduced after each GA operation. The query of CA is to apply a regular convolution (kernel size and stride are both 1) to the input feature X, and the generated key and value are X respectively. i and GA(X i ). The CA process is defined as follows:
[0089]
[0090] Among them, Y i ∈R C×H×W , represents the output of CA; M∈R c×H×W , represents the query feature. After the output of CA, the spatial dimension is restored to its original size, effectively avoiding the over-compression caused by the feature map during the downsampling process.
[0091] In the channel dimension, the output of each channel attention (CA) is concatenated with the global features and input into the convolutional feedforward neural network (ConvFFN), further enabling the fusion and processing of multi-scale features. ConvFFN consists of two 1×1 convolutions, a 3×3 deep convolution, and a nonlinear activation function (GELU). It is worth noting that the two deep convolutions in CDPE and ConvFFN can effectively compensate for the model's shortcomings in learning local correlations. The process is defined as follows:
[0092]
[0093] After feature concatenation, the input feature map is represented as Y i ∈R C×H×W ; and Both represent a standard 1×1 convolution operation; D(·) represents a 3×3 depthwise convolution (DWC). By fusing multiple global features, convolutional feedforward neural networks avoid over-reliance on a single specific feature, further enabling efficient fusion and processing of multi-scale features. Furthermore, the deep convolution design in CDPE and ConvFFN effectively compensates for the model's shortcomings in learning local correlations, significantly enhancing its generalization and generalization capabilities, and enabling stronger global feature representation and robustness in complex scenarios.
[0094] As a hierarchical visual Transformer model, the Swin Transformer has demonstrated excellent performance in various image processing tasks. In particular, the introduction of the Window-based Self-Attention (WSA) mechanism effectively improves the ability to model local features. However, when applied to large-scale feature maps, this mechanism requires a large number of sliding window calculations, resulting in a significant increase in computational resource overhead. Given that key points are usually focused on local features in the target area, we introduced a lightweight WSA module. While strengthening the ability to express local features, it reduces the computational burden on large-resolution feature maps, thereby achieving a better balance between computational efficiency and feature expression capabilities.
[0095] As network depth increases, especially during downsampling in the final layer, the number of channels in the feature map reaches its maximum, and the number of convolution kernel parameters also reaches its maximum. Therefore, grouped convolution (optimal when grouped to 8) is employed in the final downsampling layer to reduce computational overhead. After downsampling, the model enters the upsampling phase, where layer-by-layer interpolation restores the image resolution and high-resolution features are transferred via skip connections. However, skip connections primarily focus on transferring global information, while local details are easily overlooked during feature fusion. Therefore, a windowed self-attention module is introduced during the upsampling phase. The core of the WSA module lies in the application of the windowed multi-head self-attention (W-MSA) mechanism. This module partitions the input feature map into multiple fixed-size sub-windows, each of which is a non-overlapping spatial region, within which self-attention is independently computed. In the experiments, each window covers a 21×21 feature region, and no sliding or overlapping window strategies are employed. W-MSA achieves efficient local feature extraction by performing self-attention within the partitioned fixed windows. The core of the window self-attention module lies in the application of the window-based multi-head self-attention (W-MSA) mechanism, which divides the input features (such as images or sequences) into multiple fixed-size windows. For example, an image with a resolution of H×W can be divided into h×w windows, each of which is processed independently. W-MSA achieves efficient local feature extraction by performing self-attention operations within the window. Compared with the traditional global self-attention mechanism, W-MSA retains the ability to capture fine-grained features without sacrificing accuracy, while avoiding the high computational burden brought by the global attention mechanism. After downsampling, the model enters the upsampling stage, restoring the image resolution through layer-by-layer interpolation, and introducing a window self-attention module to enhance the ability to accurately capture the locations of craniofacial key points. The introduction of this module in the upsampling stage can effectively improve the model's ability to capture long-range dependencies while ensuring refined detection of local features. The self-attention calculation definition of W-MSA is as follows:
[0096]
[0097] Where Q, K, and V represent the query matrix, key matrix, and value matrix, respectively; Qw, Kw, and Vw represent the query, key, and value sub-matrices extracted in each window, respectively; and N represents the total number of windows.
[0098] A windowed self-attention module is introduced during the upsampling phase to gradually enhance the ability to extract local information. By introducing global attention modules and windowed self-attention modules at different layers in the encoder-decoder, the modeling of global-local information interactions is significantly improved. This allows the capture of local details while maintaining the ability to model global features, ensuring accurate capture of long-range dependencies and spatial structure in the head shadow keypoint detection task.
[0099] The heat map is constructed based on the Gaussian kernel density estimation method. The formula of the Gaussian kernel function is:
[0100]
[0101] Where σ represents the peak width of the Gaussian function, which controls the diffusion range of each key point in the image. For each pixel position in the heat map, the distance from the key point is calculated and weighted summed using the Gaussian kernel. The resulting density value reflects the model's attention at that position, which is specifically defined as follows:
[0102]
[0103] Among them, (x i ,y i ) represents the coordinates of each keypoint; (x, y) represents the pixel location on the heatmap. This visualization method helps clinicians and researchers intuitively understand the anatomical regions that the model focuses on when detecting craniofacial keypoints.
[0104] The step S3 is specifically as follows:
[0105] The head shadow key point detection model is trained using the training set, and during the training process, the head shadow key point detection model is continuously optimized, including at least hyperparameters of a learning rate, a regularization parameter, a network structure parameter, an activation function parameter, an optimizer parameter, and a random dropout rate, until a loss value of the loss function is less than a preset loss threshold;
[0106] By continuously optimizing hyperparameters including learning rate, regularization parameter, network structure parameter, activation function parameter, optimizer parameter and random dropout rate during the training process of the head shadow key point detection model, the model performance is adaptively adjusted to avoid local optimality.
[0107] The detection accuracy is calculated using the test set to test the trained head shadow key point detection model. If the test fails, the data set is expanded to continue training. If the test passes, then:
[0108] The confidence level is calculated using the validation set to validate the head shadow key point detection model that has passed the test. If the validation fails, the data set is expanded to continue training; if the validation passes, the training ends.
[0109] The step S4 is specifically as follows:
[0110] The verified head shadow key point detection model is pruned and knowledge distilled and then deployed to the medical terminal. The authentication mechanism of the head shadow key point detection model is set. After authentication based on the authentication mechanism, the input real-time head shadow image is input into the deployed head shadow key point detection model, and a heat map with key points marked is output to detect the head shadow key points. The head shadow key point detection model is continuously optimized based on the marking accuracy of the heat map.
[0111] The head shadow key point detection model is trained using the training set. During the training process, hyperparameters including learning rate, regularization parameter, network structure parameter, activation function parameter, optimizer parameter and random dropout rate are continuously optimized until the loss value of the loss function is less than the preset loss threshold. The detection accuracy is calculated using the test set to test the trained head shadow key point detection model, and the confidence is calculated using the validation set to verify the head shadow key point detection model that has passed the test. After deployment and commissioning, the head shadow key point detection model is continuously optimized based on the identification accuracy of the heat map. That is, the head shadow key point detection model is continuously optimized, tested and trained during its training process, and continuously optimized after it is put into use, thereby greatly improving the accuracy of head shadow detection.
[0112] By performing pruning and knowledge distillation before deploying the head shadow key point detection model, the model size is effectively reduced, making it easier to deploy on resource-constrained devices, thereby greatly improving its scope of applicability.
[0113] By setting up an authentication mechanism, the deployed head shadow key point detection model can be called for head shadow detection only after passing the authentication mechanism, preventing the head shadow key point detection model from being illegally called.
[0114] A preferred embodiment of the head shadow key point detection system based on interactive attention of the present invention includes the following modules:
[0115] A data set construction module is used to obtain a large number of historical head images, pre-process and annotate each of the historical head images to construct a data set, and divide the data set into a training set, a test set, and a validation set;
[0116] A head shadow key point detection model creation module is used to create a head shadow key point detection model based on the encoder, decoder and output module, and set the loss function and hyperparameters of the head shadow key point detection model;
[0117] The encoder and decoder are jump-connected to form a U-shaped structure; the encoder is equipped with a global attention aggregation module (GAA) and a packet convolution module (WSA); the decoder is composed of pure linear layers and is equipped with a window self-attention module;
[0118] By setting the decoder to consist of pure linear layers, the receptive field of the head shadow key point detection model (neural network) is effectively expanded while ensuring the accuracy of head shadow key point detection.
[0119] By introducing a grouped convolution module in the last layer of the encoder, the grouped convolution module divides the input channels into multiple groups and performs convolution operations in each group, thereby significantly reducing the amount of calculation and the number of parameters, thereby greatly improving the computational efficiency.
[0120] Multi-scale feature fusion is achieved through a U-shaped jump connection structure, retaining underlying details and high-level semantic information, and improving the accuracy of key point positioning; by introducing a global attention aggregation module in the encoder, multi-scale global context is dynamically captured, enhancing the ability to model complex anatomical structures; through the decoder, a pure linear layer combined with a windowed self-attention module is used to extract long-distance dependencies while reducing computational complexity, balancing efficiency and performance.
[0121] a head shadow key point detection model training module, configured to train the head shadow key point detection model using the training set, test the trained head shadow key point detection model using the test set, and verify the head shadow key point detection model that passes the test using the verification set;
[0122] The head shadow key point detection module is used to deploy the verified head shadow key point detection model and detect the head shadow key points using the deployed head shadow key point detection model.
[0123] The dataset construction module is specifically used for:
[0124] A large number of historical cephalograms of different ages, genders, malocclusion patterns, and skeletal patterns are obtained, and each of the historical cephalograms is preprocessed, including at least geometric transformation, grayscale transformation, noise reduction, and image enhancement. Two people respectively annotate each of the preprocessed historical cephalograms with a preset number of key points, and the average of the two people's annotation results is taken as the true label to complete the annotation. A data set is constructed based on the annotated historical cephalograms, and the data set is divided into a training set, a test set, and a validation set based on a preset ratio.
[0125] Two people respectively annotate a preset number of key points on each preprocessed historical head shadow image, and take the average of the two people's annotation results as the true label, which effectively improves the accuracy of the annotation and thus effectively improves the quality of the dataset, thereby greatly improving the training effect of the head shadow key point detection model.
[0126] By setting the dataset to include cephalogram images of different ages, genders, malocclusion patterns and skeletal patterns, the generalization ability of the cephalogram key point detection model for different populations and pathological characteristics can be effectively improved.
[0127] By setting the preprocessing of historical head images to include geometric transformation, grayscale transformation, noise reduction and image enhancement, data robustness is enhanced and the risk of overfitting is reduced.
[0128] In the head shadow key point detection model creation module, the encoder is used to downsample the input head shadow image to obtain a downsampled image; the decoder is used to upsample each downsampled image to obtain an original size image; the global attention aggregation module is used to extract multi-scale global attention during the downsampling process; the grouped convolution module is provided in the last layer of the encoder to improve computational efficiency; the window self-attention module is used to extract long-distance dependencies during the upsampling process; the output module is used to identify key points based on interactive attention including multi-scale global attention and long-distance dependencies, mark the key points on the original size image and output them as a heat map;
[0129] That is, features are extracted from the head shadow image in layers, the image resolution is reduced at each layer, and global attention processing is performed on the image features of the downsampled 2 and 3 layers, and local attention processing is performed on the image features of the upsampled 3 and 4 layers. Features of multiple scales are fused through left and right staggered interactions, and then sent to the decoder of the same layer for upsampling and feature decoding. The position of each key point is predicted by generating a heat map.
[0130] The loss function is constructed based on the weighted combination of key point position loss function, key point visibility loss function, bounding box loss function and classification loss function.
[0131] By setting the loss function based on the weighted construction of key point position loss function, key point visibility loss function, bounding box loss function and classification loss function, the key point position loss function is used to measure the difference between the predicted key point position and the true position, the key point visibility loss function is used to measure the difference between the predicted key point visibility and the true visibility, the bounding box loss function is used to measure the difference between the predicted bounding box and the true bounding box, and the classification loss function is used to measure the difference between the predicted category and the true category. That is, the advantages of different loss functions are integrated, and the contributions between different losses are balanced by weights, which further improves the training effect of the head shadow key point detection model.
[0132] The global attention aggregation module performs dynamic position encoding of convolution on the input data (head image) to capture sequential dependencies and enhance feature expression capabilities. It then standardizes the data through layer normalization to improve the training stability of the model. It then performs feature extraction through local attention operations to enhance the expression capabilities of local information. Finally, it uses convolution transformation operations to further aggregate multi-scale features to improve the feature extraction and generalization capabilities of the model.
[0133] The calculation process of convolutional dynamic position encoding (CDPE: Convolutional Dynamic Position Encoding) is:
[0134] For the input labeled tensor X in ∈R C×H×W , the convolutional dynamic position encoding (CDPE: Convolutional Dynamic Position Encoding) dynamically adds position information to all tags through a 3×3, stride 1, and fill 1 depth separation convolution calculation, aiming to enhance the spatial position information of the feature map, and then through the residual connection (X in +DWC(·)) alleviates the gradient vanishing problem and ultimately enhances the spatial position information while maintaining the size of the feature map, providing support for the precise positioning of key points. The calculation definition formula of CDPE is expressed as:
[0135] X=X in +DWC(Pad(X in , p);
[0136] Among them, X in ∈R C×H×W Represents the input image, R is a real number, C is the number of channels, H is the image height, and W is the image width; Pad represents padding; p represents the number of paddings; X represents the dynamic position encoding of the convolution.
[0137] Multi-Scale Global Attention (MSG) is a global feature extraction mechanism for head key point detection. It is particularly suitable for capturing features with relatively stable spatial distribution, such as the skull, and aims to achieve a balance between local details and global structural features. In the head key point detection task, it not only relies on the meticulous capture of local details, but also requires the model to have the ability to capture global features and provide rich contextual information to avoid errors caused by isolated analysis of local information. However, the global attention mechanism of traditional VIT incurs high computational costs. On large-scale datasets, VIT can automatically learn these features through a large amount of data, but on small-scale datasets, the model cannot fully learn the spatial structural information in the image, resulting in a decrease in the model's generalization ability. MSG achieves global feature extraction by introducing an aggregation-based global attention mechanism, making the model adaptable and stable on datasets of different scales.
[0138] MSG extracts input features by performing multi-scale convolution operations and combines them with a global attention mechanism to achieve multi-level feature fusion. Specifically, MSG uses convolution kernels of different sizes (such as 4×4, 8×8, and 12×12) to replace traditional convolution with a combination of depthwise convolution (DWC) and pointwise convolution (PWC) to obtain multi-scale features and effectively reduce the computational complexity of the global attention mechanism (GA). The number of output channels of PWC is set to one-quarter of the number of input feature channels. At the same time, MSC reduces the spatial size of the feature map through downsampling operations, sets a larger convolution kernel (greater than 3), and a stride of half the size of the convolution kernel, thereby alleviating the problem of loss of detail information that may result from the downsampling process. After MSC processing, each multi-scale feature map x i Calculate the global attention (GA), the process is defined as follows:
[0139]
[0140] in, Represents the output of the multi-scale feature map; represents the output of GA; Q(·), K(·) and V(·) are linear projection operations; W q , W k and W v represents the learning parameter; d is equal to C, which represents the channel dimension.
[0141] Since the stride of MSC is set to be large, some local details will be lost during the feature downsampling process. To alleviate this problem, a cross attention (CA) mechanism is introduced after each GA operation. The query of CA is to apply a regular convolution (kernel size and stride are both 1) to the input feature X, and the generated key and value are X respectively. iand GA(X i ). The CA process is defined as follows:
[0142]
[0143] Among them, Y i ∈R C×H×W , represents the output of CA; M∈R c×H×W , represents the query feature. After the output of CA, the spatial dimension is restored to its original size, effectively avoiding the over-compression caused by the feature map during the downsampling process.
[0144] In the channel dimension, the output of each channel attention (CA) is concatenated with the global features and input into the convolutional feedforward neural network (ConvFFN), further enabling the fusion and processing of multi-scale features. ConvFFN consists of two 1×1 convolutions, a 3×3 deep convolution, and a nonlinear activation function (GELU). It is worth noting that the two deep convolutions in CDPE and ConvFFN can effectively compensate for the model's shortcomings in learning local correlations. The process is defined as follows:
[0145]
[0146] After feature concatenation, the input feature map is represented as Y i ∈R C×H×W ; and Both represent a standard 1×1 convolution operation; D(·) represents a 3×3 depthwise convolution (DWC). By fusing multiple global features, convolutional feedforward neural networks avoid over-reliance on a single specific feature, further enabling efficient fusion and processing of multi-scale features. Furthermore, the deep convolution design in CDPE and ConvFFN effectively compensates for the model's shortcomings in learning local correlations, significantly enhancing its generalization and generalization capabilities, and enabling stronger global feature representation and robustness in complex scenarios.
[0147] As a hierarchical visual Transformer model, the Swin Transformer has demonstrated excellent performance in various image processing tasks. In particular, the introduction of the Window-based Self-Attention (WSA) mechanism effectively improves the ability to model local features. However, when applied to large-scale feature maps, this mechanism requires a large number of sliding window calculations, resulting in a significant increase in computational resource overhead. Given that key points are usually focused on local features in the target area, we introduced a lightweight WSA module. While strengthening the ability to express local features, it reduces the computational burden on large-resolution feature maps, thereby achieving a better balance between computational efficiency and feature expression capabilities.
[0148] As network depth increases, especially during downsampling in the final layer, the number of channels in the feature map reaches its maximum, and the number of convolution kernel parameters also reaches its maximum. Therefore, grouped convolution (optimal when grouped to 8) is employed in the final downsampling layer to reduce computational overhead. After downsampling, the model enters the upsampling phase, where layer-by-layer interpolation restores the image resolution and high-resolution features are transferred via skip connections. However, skip connections primarily focus on transferring global information, while local details are easily overlooked during feature fusion. Therefore, a windowed self-attention module is introduced during the upsampling phase. The core of the WSA module lies in the application of the windowed multi-head self-attention (W-MSA) mechanism. This module partitions the input feature map into multiple fixed-size sub-windows, each of which is a non-overlapping spatial region, within which self-attention is independently computed. In the experiments, each window covers a 21×21 feature region, and no sliding or overlapping window strategies are employed. W-MSA achieves efficient local feature extraction by performing self-attention within the partitioned fixed windows. The core of the window self-attention module lies in the application of the window-based multi-head self-attention (W-MSA) mechanism, which divides the input features (such as images or sequences) into multiple fixed-size windows. For example, an image with a resolution of H×W can be divided into h×w windows, each of which is processed independently. W-MSA achieves efficient local feature extraction by performing self-attention operations within the window. Compared with the traditional global self-attention mechanism, W-MSA retains the ability to capture fine-grained features without sacrificing accuracy, while avoiding the high computational burden brought by the global attention mechanism. After downsampling, the model enters the upsampling stage, restoring the image resolution through layer-by-layer interpolation, and introducing a window self-attention module to enhance the ability to accurately capture the locations of craniofacial key points. The introduction of this module in the upsampling stage can effectively improve the model's ability to capture long-range dependencies while ensuring refined detection of local features. The self-attention calculation definition of W-MSA is as follows:
[0149]
[0150] Where Q, K, and V represent the query matrix, key matrix, and value matrix, respectively; Qw, Kw, and Vw represent the query, key, and value sub-matrices extracted in each window, respectively; and N represents the total number of windows.
[0151] A windowed self-attention module is introduced during the upsampling phase to gradually enhance the ability to extract local information. By introducing global attention modules and windowed self-attention modules at different layers in the encoder-decoder, the modeling of global-local information interactions is significantly improved. This allows the capture of local details while maintaining the ability to model global features, ensuring accurate capture of long-range dependencies and spatial structure in the head shadow keypoint detection task.
[0152] The heat map is constructed based on the Gaussian kernel density estimation method. The formula of the Gaussian kernel function is:
[0153]
[0154] Where σ represents the peak width of the Gaussian function, which controls the diffusion range of each key point in the image. For each pixel position in the heat map, the distance from the key point is calculated and weighted summed using the Gaussian kernel. The resulting density value reflects the model's attention at that position, which is specifically defined as follows:
[0155]
[0156] Among them, (x i ,y i ) represents the coordinates of each keypoint; (x, y) represents the pixel location on the heatmap. This visualization method helps clinicians and researchers intuitively understand the anatomical regions that the model focuses on when detecting craniofacial keypoints.
[0157] The head shadow key point detection model training module is specifically used for:
[0158] The head shadow key point detection model is trained using the training set, and during the training process, the head shadow key point detection model is continuously optimized, including at least hyperparameters of a learning rate, a regularization parameter, a network structure parameter, an activation function parameter, an optimizer parameter, and a random dropout rate, until a loss value of the loss function is less than a preset loss threshold;
[0159] By continuously optimizing hyperparameters including learning rate, regularization parameter, network structure parameter, activation function parameter, optimizer parameter and random dropout rate during the training process of the head shadow key point detection model, the model performance is adaptively adjusted to avoid local optimality.
[0160] The detection accuracy is calculated using the test set to test the trained head shadow key point detection model. If the test fails, the data set is expanded to continue training. If the test passes, then:
[0161] The confidence level is calculated using the validation set to validate the head shadow key point detection model that has passed the test. If the validation fails, the data set is expanded to continue training; if the validation passes, the training ends.
[0162] The head shadow key point detection module is specifically used for:
[0163] The verified head shadow key point detection model is pruned and knowledge distilled and then deployed to the medical terminal. The authentication mechanism of the head shadow key point detection model is set. After authentication based on the authentication mechanism, the input real-time head shadow image is input into the deployed head shadow key point detection model, and a heat map with key points marked is output to detect the head shadow key points. The head shadow key point detection model is continuously optimized based on the marking accuracy of the heat map.
[0164] The head shadow key point detection model is trained using the training set. During the training process, hyperparameters including learning rate, regularization parameter, network structure parameter, activation function parameter, optimizer parameter and random dropout rate are continuously optimized until the loss value of the loss function is less than the preset loss threshold. The detection accuracy is calculated using the test set to test the trained head shadow key point detection model, and the confidence is calculated using the validation set to verify the head shadow key point detection model that has passed the test. After deployment and commissioning, the head shadow key point detection model is continuously optimized based on the identification accuracy of the heat map. That is, the head shadow key point detection model is continuously optimized, tested and trained during its training process, and continuously optimized after it is put into use, thereby greatly improving the accuracy of head shadow detection.
[0165] By performing pruning and knowledge distillation before deploying the head shadow key point detection model, the model size is effectively reduced, making it easier to deploy on resource-constrained devices, thereby greatly improving its scope of applicability.
[0166] By setting up an authentication mechanism, the deployed head shadow key point detection model can be called for head shadow detection only after passing the authentication mechanism, preventing the head shadow key point detection model from being illegally called.
[0167] In summary, the advantages of the present invention are:
[0168] 1. By obtaining a large number of historical head shadow images, pre-processing and annotating each historical head shadow image, a data set is constructed, and the data set is divided into a training set, a test set, and a validation set; then a head shadow key point detection model is created based on the encoder, decoder, and output module, and the loss function and hyperparameters of the head shadow key point detection model are set; the encoder and the decoder are jump-connected to form a U-shaped structure; the encoder is equipped with a global attention aggregation module and a group convolution module; the decoder is composed of a pure linear layer and is equipped with a window self-attention module; then the head shadow key point detection model is trained through the training set, the trained head shadow key point detection model is tested through the test set, and the head shadow key point detection model that has passed the test is verified through the validation set. Finally, the verified head shadow key point detection model is deployed, and the head shadow key point detection model is used for head shadow key point detection. Detection; that is, head shadow key point detection is performed through a pre-trained head shadow key point detection model created based on an encoder, a decoder and an output module. The encoder extracts multi-scale global attention through a global attention aggregation module during the downsampling process, and the decoder extracts long-distance dependencies through a window self-attention module during the upsampling process. The output module identifies key points based on the interactive attention including multi-scale global attention and long-distance dependencies, and marks the key points on the original size image output by the decoder and outputs them as a heat map. By fusing multi-scale global attention and long-distance dependencies, the feature extraction capability can be effectively improved, and global features and detail information can be effectively learned. Automatic detection through the head shadow key point detection model can avoid the problem of inconsistent detection results during repeated evaluations caused by traditional manual operations, and ultimately greatly improve the accuracy and consistency of head shadow detection.
[0169] 2. Two people respectively annotate a preset number of key points on each pre-processed historical head shadow image, and take the average of the two people's annotation results as the true label, which effectively improves the accuracy of the annotation and thus effectively improves the quality of the dataset, thereby greatly improving the training effect of the head shadow key point detection model.
[0170] 3. By setting the decoder to consist of pure linear layers, the receptive field of the head shadow key point detection model (neural network) is effectively expanded while ensuring the accuracy of head shadow key point detection.
[0171] 4. By introducing a grouped convolution module in the last layer of the encoder, the grouped convolution module divides the input channels into multiple groups and performs convolution operations in each group, thereby significantly reducing the amount of calculation and the number of parameters, thereby greatly improving the computational efficiency.
[0172] 5. By setting the loss function based on the weighted construction of key point position loss function, key point visibility loss function, bounding box loss function and classification loss function, the key point position loss function is used to measure the difference between the predicted key point position and the true position, the key point visibility loss function is used to measure the difference between the predicted key point visibility and the true visibility, the bounding box loss function is used to measure the difference between the predicted bounding box and the true bounding box, and the classification loss function is used to measure the difference between the predicted category and the true category. That is, the advantages of different loss functions are integrated, and the contributions between different losses are balanced by weights, which further improves the training effect of the head shadow key point detection model.
[0173] 6. The head shadow key point detection model is trained using the training set. During the training process, the hyperparameters including learning rate, regularization parameter, network structure parameter, activation function parameter, optimizer parameter and random dropout rate are continuously optimized until the loss value of the loss function is less than the preset loss threshold. The detection accuracy is calculated using the test set to test the trained head shadow key point detection model, and the confidence is calculated using the validation set to verify the head shadow key point detection model that has passed the test. After deployment and use, the head shadow key point detection model is continuously optimized based on the identification accuracy of the heat map. That is, the head shadow key point detection model is continuously optimized, tested and trained during the training process, and continuously optimized after it is put into use, thereby greatly improving the accuracy of head shadow detection.
[0174] 7. By performing pruning and knowledge distillation before deploying the head shadow key point detection model, the model size is effectively reduced, making it easier to deploy on resource-constrained devices, thereby greatly improving the scope of applicability.
[0175] 8. By setting up an authentication mechanism, the deployed head shadow key point detection model can be called for head shadow detection only after authentication through the authentication mechanism, preventing the head shadow key point detection model from being illegally called.
[0176] 9. By setting the dataset to include head images of different ages, genders, malocclusion patterns and skeletal patterns, the generalization ability of the head key point detection model for different populations and pathological characteristics is effectively improved.
[0177] 10. By setting the preprocessing of historical head images to include geometric transformation, grayscale transformation, noise reduction and image enhancement, the data robustness is enhanced and the risk of overfitting is reduced.
[0178] 11. A U-shaped skip connection structure is used to achieve multi-scale feature fusion, retaining underlying details and high-level semantic information, and improving the accuracy of key point positioning. By introducing a global attention aggregation module in the encoder, multi-scale global context is dynamically captured, enhancing the ability to model complex anatomical structures. The decoder uses a pure linear layer combined with a windowed self-attention module to extract long-distance dependencies while reducing computational complexity, balancing efficiency and performance.
[0179] 12. By continuously optimizing the hyperparameters including learning rate, regularization parameters, network structure parameters, activation function parameters, optimizer parameters and random dropout rate during the training process of the head shadow key point detection model, the model performance is adaptively adjusted to avoid local optimality.
[0180] Although the specific embodiments of the present invention are described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and are not intended to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be included within the scope of protection of the claims of the present invention.
Claims
1. A head shadow key point detection method based on interactive attention, characterized by: The steps include: Step S1: obtaining a large number of historical head images, preprocessing and annotating each of the historical head images to construct a data set, and dividing the data set into a training set, a test set, and a validation set; Step S2: creating a head shadow key point detection model based on the encoder, decoder, and output module, and setting the loss function and hyperparameters of the head shadow key point detection model; The encoder and decoder are jump-connected to form a U-shaped structure; the encoder is equipped with a global attention aggregation module and a grouped convolution module; the decoder is composed of pure linear layers and is equipped with a windowed self-attention module; Step S3, training the head shadow key point detection model using the training set, testing the trained head shadow key point detection model using the test set, and verifying the head shadow key point detection model that passes the test using the validation set; Step S4: deploying the verified head shadow key point detection model, and detecting the head shadow key points using the deployed head shadow key point detection model.
2. The head shadow key point detection method based on interactive attention according to claim 1, characterized in that: The step S1 is specifically as follows: A large number of historical cephalograms of different ages, genders, malocclusion patterns, and skeletal patterns are obtained, and each of the historical cephalograms is preprocessed, including at least geometric transformation, grayscale transformation, noise reduction, and image enhancement. Two people respectively annotate each of the preprocessed historical cephalograms with a preset number of key points, and the average of the two people's annotation results is taken as the true label to complete the annotation. A data set is constructed based on the annotated historical cephalograms, and the data set is divided into a training set, a test set, and a validation set based on a preset ratio.
3. The head shadow key point detection method based on interactive attention according to claim 1, characterized in that: In step S2, the encoder is used to downsample the input head image to obtain a downsampled image; the decoder is used to upsample each downsampled image to obtain an original size image; the global attention aggregation module is used to extract multi-scale global attention during the downsampling process; the grouped convolution module is provided in the last layer of the encoder to improve computational efficiency; the windowed self-attention module is used to extract long-distance dependencies during the upsampling process; the output module is used to identify key points based on the interactive attention including multi-scale global attention and long-distance dependencies, mark the key points on the original size image and output them as a heat map; The loss function is constructed based on the weighted combination of key point position loss function, key point visibility loss function, bounding box loss function and classification loss function.
4. The head shadow key point detection method based on interactive attention according to claim 1, characterized in that: The step S3 is specifically as follows: The head shadow key point detection model is trained using the training set, and during the training process, the head shadow key point detection model is continuously optimized, including at least hyperparameters of a learning rate, a regularization parameter, a network structure parameter, an activation function parameter, an optimizer parameter, and a random dropout rate, until a loss value of the loss function is less than a preset loss threshold; The detection accuracy is calculated using the test set to test the trained head shadow key point detection model. If the test fails, the data set is expanded to continue training. If the test passes, then: The confidence level is calculated using the validation set to validate the head shadow key point detection model that has passed the test. If the validation fails, the data set is expanded to continue training; if the validation passes, the training ends.
5. The head shadow key point detection method based on interactive attention according to claim 1, characterized in that: The step S4 is specifically as follows: The verified head shadow key point detection model is pruned and knowledge distilled and then deployed to the medical terminal. The authentication mechanism of the head shadow key point detection model is set. After authentication based on the authentication mechanism, the input real-time head shadow image is input into the deployed head shadow key point detection model, and a heat map with key points marked is output to detect the head shadow key points. The head shadow key point detection model is continuously optimized based on the marking accuracy of the heat map.
6. A head shadow key point detection system based on interactive attention, characterized by: Includes the following modules: A data set construction module is used to obtain a large number of historical head images, pre-process and annotate each of the historical head images to construct a data set, and divide the data set into a training set, a test set, and a validation set; A head shadow key point detection model creation module is used to create a head shadow key point detection model based on the encoder, decoder and output module, and set the loss function and hyperparameters of the head shadow key point detection model; The encoder and decoder are jump-connected to form a U-shaped structure; the encoder is equipped with a global attention aggregation module and a grouped convolution module; the decoder is composed of pure linear layers and is equipped with a windowed self-attention module; a head shadow key point detection model training module, configured to train the head shadow key point detection model using the training set, test the trained head shadow key point detection model using the test set, and verify the head shadow key point detection model that passes the test using the verification set; The head shadow key point detection module is used to deploy the verified head shadow key point detection model and detect the head shadow key points using the deployed head shadow key point detection model.
7. The head shadow key point detection system based on interactive attention according to claim 6, characterized in that: The dataset construction module is specifically used for: A large number of historical cephalograms of different ages, genders, malocclusion patterns, and skeletal patterns are obtained, and each of the historical cephalograms is preprocessed, including at least geometric transformation, grayscale transformation, noise reduction, and image enhancement. Two people respectively annotate each of the preprocessed historical cephalograms with a preset number of key points, and the average of the two people's annotation results is taken as the true label to complete the annotation. A data set is constructed based on the annotated historical cephalograms, and the data set is divided into a training set, a test set, and a validation set based on a preset ratio.
8. The interactive attention-based head shadow key point detection system according to claim 6, characterized in that: In the head shadow key point detection model creation module, the encoder is used to downsample the input head shadow image to obtain a downsampled image; the decoder is used to upsample each downsampled image to obtain an original size image; the global attention aggregation module is used to extract multi-scale global attention during the downsampling process; the grouped convolution module is provided in the last layer of the encoder to improve computational efficiency; the window self-attention module is used to extract long-distance dependencies during the upsampling process; the output module is used to identify key points based on interactive attention including multi-scale global attention and long-distance dependencies, mark the key points on the original size image and output them as a heat map; The loss function is constructed based on the weighted combination of key point position loss function, key point visibility loss function, bounding box loss function and classification loss function.
9. The head shadow key point detection system based on interactive attention according to claim 6, characterized in that: The head shadow key point detection model training module is specifically used for: The head shadow key point detection model is trained using the training set, and during the training process, the head shadow key point detection model is continuously optimized, including at least hyperparameters of a learning rate, a regularization parameter, a network structure parameter, an activation function parameter, an optimizer parameter, and a random dropout rate, until a loss value of the loss function is less than a preset loss threshold; The detection accuracy is calculated using the test set to test the trained head shadow key point detection model. If the test fails, the data set is expanded to continue training. If the test passes, then: The confidence level is calculated using the validation set to validate the head shadow key point detection model that has passed the test. If the validation fails, the data set is expanded to continue training; if the validation passes, the training ends.
10. The head shadow key point detection system based on interactive attention according to claim 6, characterized in that: The head shadow key point detection module is specifically used for: The verified head shadow key point detection model is pruned and knowledge distilled and then deployed to the medical terminal. The authentication mechanism of the head shadow key point detection model is set. After authentication based on the authentication mechanism, the input real-time head shadow image is input into the deployed head shadow key point detection model, and a heat map with key points marked is output to detect the head shadow key points. The head shadow key point detection model is continuously optimized based on the marking accuracy of the heat map.