Automatic classification method of sagittal facial profile in children

By constructing a specialized neural network model and combining data annotation and transfer learning, the problem of sagittal facial pattern classification in children has been solved, achieving high-accuracy automatic diagnosis, especially in the identification of skeletal Class III.

CN116612314BActive Publication Date: 2025-11-04GUANGXI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310414928.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-18
Publication Date
2025-11-04
Estimated Expiration
2043-04-18

AI Technical Summary

Technical Problem

Existing technologies are difficult to effectively classify sagittal facial patterns in children, especially due to challenges such as difficulty in sample collection, large variability in children's facial patterns, the need for multi-source data analysis, and blurred boundaries between adjacent stages, resulting in high diagnostic errors.

Method used

A neural network model specifically designed for sagittal facial pattern classification in children was constructed. Diagnosis was performed using lateral or lateral cephalometric radiographs. The model was trained using the DenseNet121 model, incorporating data annotation, data augmentation, and transfer learning. Label distribution learning was employed to reduce label confusion in borderline cases.

Benefits of technology

It achieves automatic classification of children's sagittal skeletal patterns based on a single image input, with an accuracy rate of 80%, especially for the identification of skeletal type III, the accuracy rate reaches 85%, reducing diagnostic errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116612314B_ABST
    Figure CN116612314B_ABST
Patent Text Reader

Abstract

The present application provides a child sagittal facial pattern automatic classification method, comprising the following steps: step one: data collection; step two: data labeling, after manual measurement, classification is performed by three orthodontic experts with 20 years of work experience, and child 90 degree side view photos are classified and labeled according to corresponding child X-ray lateral cephalogram; step three: data processing and data enhancement, YOLOV5 is used to automatically extract the image area related to A point, N point and B point from the original image to reduce the interference from other anatomical structures, and then the extracted image is resized to 224*224 pixel size, and then a subset ANB angle-Subset is extracted from all data sets ANB angle-all, which only contains accurately classified samples; label distribution learning; the present application constructs a neural network model specially used for child sagittal facial pattern classification by using X-ray lateral cephalogram or only using side view photos, so that the child sagittal facial pattern can be obtained by only inputting the lateral cephalogram or the side view photo.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer intelligent classification, and in particular relates to an automatic classification method for sagittal facial features in children. Background Technology

[0002] In recent years, artificial intelligence technology based on convolutional neural networks (CNNs) has become an efficient and reliable tool for medical image diagnosis. In the field of orthodontics, many scholars have also attempted to apply neural networks to cephalometrics. In 2017, Ar1k et al. first used CNNs for automatic landmark detection and measurement in lateral cephalometric radiographs, and Yoon et al. used cascaded CNNs for landmark detection in cephalometric analysis. Deep learning algorithms have demonstrated excellent performance in cephalometric analysis, but these studies have focused on the automatic identification and detection of landmarks. Like traditional cephalometric methods, they still require landmark detection before measurement and diagnosis.

[0003] Considering the potential for high errors in traditional diagnostic methods that rely on landmark detection, in 2019, Yu et al. developed a fully automated, one-step, end-to-end deep learning system for skeletal classification and diagnosis. This system can automatically diagnose sagittal bone types without the need for landmark detection or measurement, achieving an accuracy rate exceeding 90%. In 2022, Gao et al. compared the performance of four different CNN algorithms for automatic sagittal bone classification on lateral cephalometric radiographs. In the same year, Kim et al. found that a DCNN-based AI model outperformed automatic tracking AI software in sagittal bone classification.

[0004] Compared to existing research, sagittal bone facial pattern classification faces three new challenges. First, due to the difficulty in collecting samples, there is no specific model developed for sagittal bone classification in children. Children, being in a period of rapid growth and development, exhibit significant variations in bone facial patterns, which differ greatly from adults and are more difficult to diagnose. Second, for children with sagittal problems, multi-source data analysis is required. Third, the ambiguous boundaries between adjacent stages cause uncertainty in labeling; children's sagittal bone classification may remain between two adjacent types, making it difficult even for human experts to distinguish the exact stage. Summary of the Invention

[0005] To address the problems in the background art, this invention provides an automatic classification method for sagittal facial features in children, constructs a neural network model specifically for sagittal facial feature classification in children, and performs diagnosis using only lateral cephalometric radiographs or only side photographs, laying the foundation for subsequent multimodal fusion model research.

[0006] An automatic classification method for sagittal facial features in children includes the following steps:

[0007] Step 1: Data Collection

[0008] Lateral cephalometric radiographs and 90° lateral cephalometric photographs of children were taken. The lateral cephalometric radiographs were all taken using a Mindray Hyperion X9 (Italy's Sefello Group), with original image resolutions of 2460×1950 or 1752×2108 pixels and a resolution of 0.1 mm / pixel. The 90° lateral cephalometric photographs were all taken using a Nikon D7200, with original image resolutions of 2510×2000 pixels and a resolution of 0.1 mm / pixel. All images were stored in high-resolution JPG format.

[0009] Step 2: Data Labeling

[0010] The standard for annotation is the A-Nasion-B point (ANB angle) and the Wits assessment method. The ANB angle is used to assess the sagittal relationship between the maxilla and mandible, and is measured by the angle formed by point A, the root of the nose, and point B. The WITS assessment is an analytical method for classifying sagittal skeletal relationships that describe the severity of anterior-posterior disharmony of the jaw. This method requires projecting points A and B perpendicularly onto the occlusal plane passing through the intersection of the maximal cusps.

[0011] Based on the normal average values ​​of the ANB angle and WITS for Chinese patients, all X-ray images were categorized into three classes: skeletal Class I (5° ≥ ANB angle ≥ 0° and 2 ≥ WITS ≥ -3), skeletal Class II (ANB angle > 5° and WITS > 2), and skeletal Class III (ANB angle < 0° and WITS < -3). After manual measurement, the images were classified by three orthodontic experts with 20 years of experience. If two experts had different opinions on an image, that image would be circulated among all experts for discussion. Children's 90° lateral profile photos were also categorized and labeled according to the corresponding X-ray images.

[0012] Step 3: Data Processing and Data Augmentation

[0013] YOLOv5 was used to automatically extract image regions involving points A, N, and B from the original images to reduce interference from other anatomical structures. The extracted images were then resized to 224×224 pixels. A subset, ANB-Subset, containing only samples that were accurately classified, was then extracted from all datasets (ANB angle-all).

[0014] To avoid overfitting the model on small datasets, the following data augmentation methods should be used randomly: random rotation, random scaling, random translation, and random changes in contrast and brightness. In each training cycle, the training set data has a 50% probability of being augmented.

[0015] Step 4: Label Distribution Learning

[0016] Define a set l = {1, 2, 3} to represent the three category labels for skeletal classification. Given an input image x, the one-hot label of x is defined as y (y ∈ l), and the label distribution is obtained by transforming y into d = {p1, p2, p3} through a function, where p i This indicates that x belongs to label l i The probability value. Transform y using a Gaussian function:

[0017]

[0018] Where σ is a hyperparameter that needs to be set, it determines the width of the Gaussian function curve; y represents the true label value of the input image x, l i This represents the true label value of the i-th category.

[0019] Considering that in reality, Class I skeletal bone is a transitional stage between Class II and Class III skeletal bone, and the distance between Class II and Class III skeletal bone is much greater, the above constraints were imposed on this situation.

[0020] Step 5: Model Structure and Training Details

[0021] Convolutional neural networks (CNNs) for compressing medical data typically employ transfer learning strategies, which utilize model parameters that have already been pre-trained on non-medical data. A representative CNN model was selected as the backbone network for this method: DenseNet121.

[0022] The training steps for the DenseNet121 model are as follows:

[0023] Step 1: Use the pre-trained weight parameters corresponding to the large-scale image dataset ImageNet as the initial weights of the model;

[0024] Step 2: Train and upgrade all layers of the CNN model using fine-tuning techniques;

[0025] Step 3: After initializing the parameters of the pre-trained model, the Stochastic Gradient Descent (SGD) optimizer was used to train each CNN model in this study for 200 epochs and 150 epochs for side photos. The hyperparameters of the neural network were adjusted multiple times based on the model performance on the validation set.

[0026] Step 4: A personalized combination of hyperparameters, including learning rate, batch size, momentum, and weight decay, was determined to maximize the capabilities of the CNN;

[0027] All training processes in this method were performed on a computer equipped with an NVIDIA GeForce RTX 3080 GPU.

[0028] Step Six: Model Performance Evaluation

[0029] Model performance is tested using the confusion matrix, classification accuracy (ACC), sensitivity (SN), specificity (SP), receiver operating characteristic (ROC) curve, and area under the curve (AUC). Accuracy is one of the most commonly used classification evaluation metrics. It is calculated by dividing the number of correctly classified samples by the total number of images.

[0030] The following algorithms are used to calculate accuracy, sensitivity, and specificity:

[0031] Accuracy (ACC) = (TP + TN) / (TP + FN + TN + FP) (100%)

[0032] Sensitivity (SN) = TP / (TP + FN) (100%)

[0033] Specificity (SP) = TN / (TN + FP) (100%)

[0034] Where TP represents true positive, TN represents true negative, FP represents false positive, and FN represents false negative. Calculate the ROC curve and AUC for each bone category.

[0035] This method uses two types of averages: micro-average and macro-average. AUC is an effective and comprehensive measure of the sensitivity and specificity of the inherent validity of a diagnostic test and the overall performance of the ROC curve.

[0036] For each method, perform 5-fold cross-validation: randomly divide the dataset into 5 parts according to the proportion of label categories. In each experiment, take 4 parts as the training set and the remaining part as the validation set, and calculate the mean and standard deviation of the 5 results.

[0037] Considering that label confusion mainly exists between borderline cases, the first-stage bias is defined as follows: it is acceptable to misclassify a borderline case as its adjacent category.

[0038] Beneficial effects:

[0039] 1. This method allows you to obtain the sagittal facial profile of a child simply by inputting a lateral or side view of the skull.

[0040] 2. This method uses a CNN model based on 90° side view photos for automatic bone classification. The results show that the CNN model has an accuracy of 80% and can classify sagittal bone types in children to a large extent. The highest accuracy is achieved for bone type III, reaching 85%. For more severe bone deformities, the bone type can be preliminarily judged based on the soft tissue profile alone. Attached Figure Description

[0041] Figure 1 This is a flowchart of the invention;

[0042] Figure 2 This describes the label distribution under case a and case σ.

[0043] Figure 3 This describes the label distribution under case b and under case σ.

[0044] Figure 4 This describes the label distribution under case c and under case σ.

[0045] Figure 5 The Epoch curves of the training and validation sets of the CNN model trained based on lateral cephalometric radiographs (a);

[0046] Figure 6 The Epoch curves of the training and validation sets of the CNN model trained based on the side view photo (b) are shown.

[0047] Figure 7 A is a visualization of class activation maps of images in Class I skeletal facial type classification;

[0048] Figure 8 This is a visualization of class activation maps of images in Class I skeletal facial type classification (B).

[0049] Figure 9 A is a visualization of class activation maps of images in Class II skeletal facial type classification;

[0050] Figure 10 This is a visualization of class activation maps of images in Class II skeletal facial type classification (B).

[0051] Figure 11 A is a visualization of class activation maps of images in Class III skeletal facial type classification;

[0052] Figure 12 This is a visualization of class activation maps of images in Class III skeletal facial type classification (B).

[0053] Figure 13 This is the confusion matrix diagram A, which is based on the ANB-all cephalometric radiograph dataset used to train a CNN model.

[0054] Figure 14 This is the confusion matrix B, which is based on the ANB-sub cephalometric radiograph dataset used to train a CNN model.

[0055] Figure 15 The confusion matrix C is based on the ANB-a ll cephalometric radiograph dataset and incorporates LDL to train a CNN model;

[0056] Figure 16 The confusion matrix D is a diagram of a CNN model trained on the ANB-all profile photo dataset.

[0057] Figure 17 E is the confusion matrix diagram E for training a CNN model based on the ANB-sub profile photo dataset;

[0058] Figure 18 F is the confusion matrix diagram F based on the ANB-a ll profile photo dataset and LDL-trained CNN model;

[0059] Figure 19 Figure A shows the ROC curve of a CNN model trained based on the ANB-all cephalometric radiograph dataset.

[0060] Figure 20 Figure B shows the ROC curve of a CNN model trained based on the ANB-sub cephalometric radiograph dataset.

[0061] Figure 21 The ROC curve C is based on the ANB-a ll cephalometric radiograph dataset and LDL is added to train the CNN model;

[0062] Figure 22 D is the ROC curve of a CNN model trained on the ANB-all profile photo dataset;

[0063] Figure 23 The graph E shows the ROC curve of a CNN model trained on the ANB-sub profile photo dataset.

[0064] Figure 24 F is the ROC curve of a CNN model trained using the ANB-a ll profile photo dataset and LDL. Detailed Implementation

[0065] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0066] according to Figure 1 As shown, the automatic classification method for sagittal facial features in children includes the following steps:

[0067] Step 1: Data Collection

[0068] Lateral cephalometric radiographs and 90° lateral cephalometric photographs of children were taken. The lateral cephalometric radiographs were all taken using a Mindray Hyperion X9 (Italy's Sefello Group), with original image resolutions of 2460×1950 or 1752×2108 pixels and a resolution of 0.1 mm / pixel. The 90° lateral cephalometric photographs were all taken using a Nikon D7200, with original image resolutions of 2510×2000 pixels and a resolution of 0.1 mm / pixel. All images were stored in high-resolution JPG format.

[0069] Step 2: Data Labeling

[0070] The standard for annotation is the A-Nasion-B point (ANB angle) and the Wits assessment method. The ANB angle refers to the sagittal relationship between the maxilla and mandible, measured by the angle formed by points A, the root of the nose, and B. The WITS assessment is an analytical method for classifying the sagittal skeletal relationship that describes the severity of anterior-posterior disharmony of the jaw. This method requires projecting points A and B perpendicularly onto the occlusal plane passing through the intersection of the maximal cusps.

[0071] Based on the normal average values ​​of the ANB angle and WITS for Chinese patients, all X-ray images were categorized into three classes: skeletal Class I (5° ≥ ANB angle ≥ 0° and 2 ≥ WITS ≥ -3), skeletal Class II (ANB angle > 5° and WITS > 2), and skeletal Class III (ANB angle < 0° and WITS < -3). After manual measurement, the images were classified by three orthodontic experts with 20 years of experience. If two experts had different opinions on an image, that image would be circulated among all experts for discussion. Children's 90° lateral profile photos were also categorized and labeled according to the corresponding X-ray images.

[0072] Step 3: Data Processing and Data Augmentation

[0073] YOLOv5 was used to automatically extract image regions involving points A, N, and B from the original images to reduce interference from other anatomical structures. The extracted images were then resized to 224×224 pixels. A subset, ANB-Subset, containing only samples that were accurately classified, was then extracted from all datasets (ANB angle-all).

[0074] To avoid overfitting the model on small datasets, the following data augmentation methods should be used randomly: random rotation, random scaling, random translation, and random changes in contrast and brightness. In each training cycle, the training set data has a 50% probability of being augmented.

[0075] Step 4: Label Distribution Learning

[0076] Define a set l = {1, 2, 3} to represent the three category labels for skeletal classification. Given an input image x, the one-hot label of x is defined as y (y ∈ l), and the label distribution is obtained by transforming y into d = {p1, p2, p3} through a function, where p i This indicates that x belongs to label l i The probability value. Transform y using a Gaussian function:

[0077]

[0078] Where σ is a hyperparameter that needs to be set, it determines the width of the Gaussian function curve; y represents the true label value of the input image x, l i This represents the true label value of the i-th category.

[0079] Considering that in reality, Class I skeletal bone is a transitional stage between Class II and Class III skeletal bone, and the distance between Class II and Class III skeletal bone is much greater, the above constraints were imposed on this situation.

[0080] All model training was performed on a computer equipped with a GeForce RTX 3080 GPU. The PyTorch architecture was used to train and fine-tune the models, and the original code was developed using the PyCharm development software.

[0081] Example

[0082] 1. Data Collection

[0083] The study included 797 males and 816 females, aged 4-14 years. Lateral cephalometric radiographs and 90° lateral cephalometric photographs were collected; non-standardized or low-resolution images were excluded. Lateral cephalometric radiographs were taken using a Mindray Hyperion X9 (Italy's Sefello Group), with original images at 2460×1950 or 1752×2108 pixels and a resolution of 0.1 mm / pixel. 90° lateral cephalometric photographs were taken using a Nikon D7200, with original images at 2510×2000 pixels and a resolution of 0.1 mm / pixel. All images were stored in high-resolution JPG format.

[0084] 2. Data labeling

[0085] The A-Nasion-B point (ANB angle) marking standard and the Wits assessment method are two commonly used methods for diagnosing sagittal skeletal relationships (see appendix). Figure 1The ANB angle refers to the sagittal relationship between the maxilla and mandible, measured by the angle formed by points A, the root of the nose, and B. The WITS assessment is an analytical method for classifying sagittal skeletal relationships that describe the severity of anterior-posterior disharmony of the jaw. This method requires projecting points A and B perpendicularly onto the occlusal plane passing through the intersection of the maximal cusps. The WITS assessment differs from the ANB angle measurement because it describes a basal plane relationship independent of the anterior cranial base angle.

[0086] This study categorized all X-ray images into three classes based on the normal mean values ​​of the ANB angle and WITS for Chinese patients: skeletal Class I (5° ≥ ANB angle ≥ 0° and 2 ≥ WITS ≥ -3), skeletal Class II (ANB angle > 5° and WITS > 2), and skeletal Class III (ANB angle < 0° and WITS < -3). After manual measurement, the images were classified by three orthodontic experts with nearly 20 years of experience. If two experts disagreed on an image, that image was shared with all experts for discussion. 90° lateral images were classified and labeled according to the corresponding X-ray images.

[0087] 3. Data processing and data augmentation

[0088] YOLOv5 was used to automatically extract image regions involving points A, N, and B (1500×800 pixels) from the original images (2460×1950 or 1752×2108 pixels) to reduce interference from other anatomical structures. The extracted images were then resized to 224×224 pixels. As mentioned earlier, even with careful annotation, label ambiguity is inevitable for samples near the boundaries of the two stages. We extracted a subset, ANB-Subset, from all datasets (ANB-all), which contains only 1400 samples that were accurately classified. This embodiment validates the effectiveness of the method on AP-all and AP-Subset.

[0089] To avoid overfitting the model on small datasets, this study also randomly employed the following data augmentation techniques: random rotation, random scaling, random translation, and random variations in contrast and brightness. During each training epoch, the training set data had a 50% probability of data augmentation, resulting in a maximum of 85,425 (1139 × 150 × 0.5) new data points after 150 training epochs. Detailed information about the dataset is shown in the table below.

[0090] 4. Label Distribution Learning

[0091] Table 1. Number of patient data points distributed across each bone category and descriptive statistics of the sample in this study.

[0092]

[0093]

[0094] As shown in Table 1, there are borderline types among the three categories of sagittal bony profile. These borderline types are real and do not represent decision boundaries generated by the model classification. For example, for a borderline type between bony type I and II, classifying it as either type I or type II is acceptable, and there is currently no universally accepted clinical guideline for classifying borderline types. Therefore, label confusion is unavoidable for borderline types. To address this issue, we use label distribution learning instead of one-hot encoding for borderline types. First, we define a set l = {1, 2, 3} to represent the three category labels for bony classification. Given an input image x, the one-hot label of x is defined as y (y ∈ l), and the label distribution is obtained by transforming y into d = {p1, p2, p3} through a function, where p i This indicates that x belongs to label l i The probability value. Here, the Gaussian function is used to transform y:

[0095]

[0096] Where σ is a hyperparameter that needs to be set, it determines the width of the Gaussian function curve; y represents the true label value of the input image x, l i This represents the true label value of the i-th category. Considering that in reality, skeletal category 1 is a transitional stage between skeletal categories 2 and 3, and the distance between skeletal categories 2 and 3 is greater, the above constraint was applied to this situation. Figure 2-4 The label distribution under different labels and σ is shown respectively.

[0097] 5. Model Structure and Training Details

[0098] Convolutional neural networks (CNNs) for compressing medical data often employ transfer learning strategies, using model parameters that have already been pre-trained on non-medical data. A representative CNN model was selected as the backbone network for this study: DenseNet121. Pre-trained weight parameters on the large-scale image dataset ImageNet were used as the initial weights for the model.

[0099] All layers of the CNN models were trained and upgraded using fine-tuning techniques. After initializing the parameters of the pre-trained models, each CNN model in this study was trained for 200 epochs (150 epochs for side views) using a stochastic gradient descent (SGD) optimizer. The hyperparameters of the neural networks were tuned multiple times based on their model performance on the validation set. Finally, a personalized combination of hyperparameters, including learning rate, batch size, momentum, and weight decay, was determined to maximize the capabilities of the CNN. All training in this study was performed on a computer equipped with an NVIDIA GeForce RTX 3080 GPU.

[0100] 6. Model performance evaluation:

[0101] according to Figure 13-24 The model performance was tested using the confusion matrix, classification accuracy (ACC), sensitivity (SN), specificity (SP), receiver operating characteristic (ROC) curve, and area under the curve (AUC). Accuracy is one of the most commonly used classification evaluation metrics. It is calculated by dividing the number of correctly classified samples by the total number of images.

[0102] The following algorithms are used to calculate accuracy, sensitivity, and specificity:

[0103] Accuracy (ACC) = (TP + TN) / (TP + FN + TN + FP) (100%)

[0104] Sensitivity (SN) = TP / (TP + FN) (100%)

[0105] Specificity (SP) = TN / (TN + FP) (100%)

[0106] TP represents true positive, TN represents true negative, FP represents false positive, and FN represents false negative. ROC curves and AUCs were calculated for each bone category. Two types of means were used in this study: micro-mean and macro-mean. AUC is a valid and comprehensive measure of the inherent validity of the diagnostic test and the sensitivity and specificity of the overall performance of the ROC curve.

[0107] This embodiment performs 5-fold cross-validation for each method: the dataset is randomly divided into 5 parts according to the proportion of label categories. In each experiment, 4 parts are used as the training set and the remaining part is used as the validation set. The mean and standard deviation of the 5 results are calculated.

[0108] Due to the ambiguity of labeling, some researchers argue that using one-stage bias for accuracy is reasonable. In this task, considering that label confusion mainly exists between borderline cases, we define one-stage bias as follows: misclassifying a borderline case into its adjacent category is acceptable. In this paper, we report normal accuracy and one-stage bias accuracy.

[0109] This embodiment creates a Class Activation Map (CAM) to better understand the model's learning style. The CAM visually highlights areas in lateral cephalometric radiographs and lateral images, which provide the most information in distinguishing skeletal classifications.

[0110] Figure 5-6 The results of the model-based ablation study are shown. Significant overfitting exists when drop_rate = 0.2. After incorporating label distribution learning, the accuracy of the CNN model improved by 1.0%, and overfitting was reduced.

[0111] The diagnostic model based on individual lateral photographs also showed an average clinical performance close to 80%. Similarly, after learning the label distribution for borderline cases, the overall accuracy improved by 0.6%, and overfitting was reduced.

[0112] Figure 7-12 The image displays class activation maps for images across different skeletal facial types. Red represents high attention, and blue represents low attention.

[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for automatic classification of children's sagittal facial profiles, characterized in that, Comprising the following steps: Step one: data collection Take X-ray lateral cephalograms of children and 90° lateral photographs of children; Step two: data labeling After artificial measurement of X-ray lateral cephalograms of children, three orthodontic experts with 20 years of experience classify the standard by labeling ANB angle and Wits evaluation method; Classify and label 90° lateral photographs of children according to the corresponding X-ray films; Step three: data processing and data enhancement Use YOLOV5 for data processing to accurately classify data samples; the data processing specifically uses YOLOV5 to automatically extract image regions involving points A, N and B from the original image to reduce interference from other anatomical structures, and then the extracted image is resized to 224x224 pixels, and a subset ANB angle-Subset is extracted from all data sets ANB angle-all, which only contains accurately classified samples; In order to avoid overfitting of the model on small data sets, the following data enhancement methods are randomly used: random rotation, random scaling, random translation and random changes in contrast and brightness, and in each training cycle, there is a 50% probability of data enhancement for the training set data; Step four: label distribution learning Definition of set l = {1,2,3}, representing three class labels of bone classification, given an input image x, the one-hot label of x is defined as y (y e l), and the label distribution is converted from y through a function to d = {p1, p2, p3}, where p i represents the probability value of x belonging to label l i , and the Gaussian function is used to convert y: where σ is a hyper-parameter to be set, which determines the width of the Gaussian function curve; y represents the true label value of the input picture x, l i represents the true label value of the i-th category. Step five: model structure and training details The representative CNN model used in this step is DenseNet121; Step six: model performance evaluation Confusion matrix, classification accuracy accuracy, ACC, sensitivity sensitivity, SN, specificity specificity, SP, receiver operating characteristic ROC curve, and area under the curve AUC are used to test the performance of the model; Considering that label confusion mainly exists between boundary cases, one-stage bias is defined as follows: it is acceptable to misclassify boundary cases as their adjacent categories.

2. The method for automatic classification of child sagittal facial profiles according to claim 1, characterized in that, In step one, X-ray lateral cephalograms of children are taken using Mindray Hyperion X9, the original image has a pixel of 2460x1950 or 1752x2108, and the resolution is 0.1mm / pixel, 90° lateral photographs of children are taken using Nikon D7200, the original image has a pixel of 2510x2000, and the resolution is 0.1mm / pixel, all images are stored in high-resolution JPG format.

3. The method for automatic classification of child sagittal facial profiles according to claim 1, characterized in that, In step two, the labeling standard is ANB angle and Wits evaluation method, ANB angle is used to evaluate the relationship between maxilla and mandible, which is measured by the angle formed by points A, nasion and B, WITS evaluation is an analysis method for classifying sagittal skeletal relationship describing the severity of jaw bone forward and backward imbalance, which requires vertical projection of points A and B to the occlusal plane through the maximum dental crossbite position; According to the Chinese normal mean of ANB angle and WITS, all X-ray films are labeled into three categories: skeletal class I is 5°≥ANB angle≥0° and 2≥WITS≥-3, skeletal class II is ANB angle>5° and WITS>2, and skeletal class III is ANB angle<0° and WITS<-3.

4. The method for automatic classification of child sagittal facial profiles according to claim 1, wherein, In step five, the DenseNet121 model training steps are as follows: Step 1: Use the corresponding pre-training weight parameters on the large-scale image dataset ImageNet as the initial weight of the model; Step 2: Use fine-tuning technology to train and upgrade all layers of the CNN model; Step 3: After initializing the parameters of the pre-trained model, use the stochastic gradient descent (SGD) optimizer to train each CNN model in this study for 200 Epochs, 150 Epochs for side photos, and the hyperparameters of the neural network are adjusted multiple times according to the model performance on the validation set; Step 4: A personalized combination of hyperparameters, including learning rate, batch size, momentum, and weight decay, was determined to maximize the capabilities of the CNN.

5. The method for automatic classification of child sagittal facial profiles according to claim 1, wherein, Step six is as follows: The following algorithm is used to calculate accuracy, sensitivity, and specificity: Accuracy (ACC) = (TP + TN) / (TP + FN + TN + FP) (100%) Sensitivity (SN) = TP / (TP + FN) (100%) Specificity (SP) = TN / (TN + FP) (100%) Where TP represents true positive, TN represents true negative, FP represents false positive, and FN represents false negative, the ROC curve and AUC for each bone class are calculated; This step uses two types of average: micro-average and macro-average, and AUC is an effective and comprehensive measure of sensitivity and specificity to evaluate the inherent effectiveness of diagnostic tests and the overall performance of ROC curves; 5-fold cross-validation is performed for each method: that is, the dataset is randomly divided into 5 parts according to the label category ratio, in each experiment, 4 parts are taken as the training set, and the remaining part is taken as the validation set, and the average value and standard deviation of the 5 results are calculated.

Citation Information

Patent Citations

  • Feature correction small sample learning labeling method and device and classification identification method

    CN114842283A

  • Artificial Intelligence Architecture For Identification Of Periodontal Features

    US20200364860A1