Facial contour guidance and multi-modal feature fusion orthodontic bone malformation screening method
The intelligent diagnostic method, which combines facial contour guidance with multimodal feature fusion, solves the problems of single modality and insufficient accuracy in the diagnosis of malocclusion in existing technologies. It achieves high-precision and low-cost screening of skeletal deformities and is suitable for radiation-free assessment of children and adolescents.
Patent Information
- Application Number
- CN202511035167.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-11-25
AI Technical Summary
Existing technologies for diagnosing malocclusion suffer from problems such as limited modality, simple structure, low accuracy, and poor adaptability, making it difficult to achieve early identification and low-cost screening, especially in children and adolescents where there is a lack of effective radiation-free assessment methods.
Employing a multi-stage training strategy that combines facial contour guidance with multimodal feature fusion, and through deep network structure design, this system integrates cephalometric radiographs and individual feature information to construct an intelligent diagnostic model that is trained on multiple modes and deployed on a single mode. This model includes facial contour feature extraction, cephalometric radiograph analysis, multimodal fusion, and physiological feature-assisted modeling, thereby achieving high-precision prediction of bony deformities.
It significantly improves the accuracy and generalization ability of skeletal deformity classification, reduces human intervention and radiation risks, and is suitable for early screening of sensitive populations such as children, adolescents and pregnant women. It has the advantages of low cost, convenient deployment and efficient screening.
Smart Images

Figure CN121011331A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of orthodontic skeletal deformity screening, and particularly relates to a facial contour guidance and multi-modal feature fusion orthodontic skeletal deformity screening method. BACKGROUND
[0002] Orthodontics is an important branch of oral medicine, mainly focusing on the evaluation and intervention of craniofacial skeletal structure, adjusting the relative spatial relationship between teeth, alveolar bone and maxilla and mandible to achieve the purpose of preventing, blocking and treating malocclusion. Malocclusion not only affects the appearance of the face and the function of mastication, but also can cause a series of oral and systemic health problems such as temporomandibular joint disorder, periodontal disease, language development disorder, and even adversely affect the psychological health and social life of patients.
[0003] Currently, the diagnosis and treatment of malocclusion still faces many challenges, including late discovery, low patient follow-up rate, and high cost of existing diagnostic methods and ionizing radiation hazards. In addition, as a chronic disease with long-term progression, malocclusion needs to rely on continuous monitoring and dynamic assessment. However, current commonly used radiological imaging methods such as lateral cephalogram and cone beam CT (CBCT) have important value in bone structure evaluation, but due to high cost, complex operation process and ionizing radiation risk, their application in routine screening and long-term follow-up is limited. In contrast, non-radiological methods such as intraoral direct vision examination, plaster or digital model analysis, and three-dimensional oral scanning have the advantage of safety, but lack of direct presentation of bone development status, making it difficult to fully assess the nature of the deformity and its development trend. Moreover, existing research on predicting bone morphology based on craniofacial soft tissue images is still limited, mainly focusing on using facial photos for face shape or skin disease diagnosis, and using intraoral photos to identify oral diseases. Such research mostly relies on single modal data input and uses end-to-end deep learning models for automatic classification or regression prediction. They generally have the following limitations:
[0004] 1. Insufficient model accuracy: Due to the lack of multi-modal information fusion, it is difficult to simulate the diagnostic process of clinicians who comprehensively judge multi-modal information, resulting in insufficient generalization ability and prediction accuracy of the model;
[0005] 2. Single algorithm structure: Existing research mostly uses traditional convolutional neural network architectures such as ResNet and VGG, lacking optimization design of deeper, multi-scale or attention mechanisms, limiting the model's ability to capture complex structural features;
[0006] 3. Incomplete diagnostic dimension: Most current research focuses on sagittal bone deformity classification, and there is still a lack of systematic exploration of automatic classification of vertical bone deformity, resulting in insufficient diagnostic coverage.
[0007] The prior art still has problems of single modality, simple structure, low precision, poor adaptability and the like in predicting skeletal deformities based on soft tissue images.
[0008] Therefore, there is an urgent need for a safe, economical and bone evaluation capable screening technology to improve the early identification rate of malocclusion deformities, assist the public in forming correct disease cognition, and thus realize early detection, early diagnosis and early treatment of malocclusion deformities, and reduce the treatment burden and additional risks brought by subsequent disease progression. SUMMARY
[0009] The present application aims to adopt a deeper network structure and a model with more scale feature extraction, design a multi-stage training strategy from coarse to fine, fuse the cephalogram and personal feature information, realize high-precision skeletal deformity prediction of multi-modal training and single-modal deployment, and realize efficient, low-cost and non-invasive intelligent diagnosis and screening.
[0010] The present application provides the following technical scheme: a facial contour guided and multi-modal feature fused orthodontic skeletal deformity screening method, comprising the following steps:
[0011] Step 1, sample data acquisition; obtaining facial frontal soft tissue images, facial lateral soft tissue images and cephalogram of the target sample;
[0012] Step 2, diagnosis label construction and image preprocessing; manually pointing each sample obtained in step 1 and obtaining a diagnosis classification label; marking the facial contour key points on the two-dimensional soft tissue photo, so that the contour marking points accurately fit the facial contour of the sample person, and obtaining an accurate soft tissue contour point image;
[0013] Step 3, constructing an automatic diagnosis classification model based on two-dimensional soft tissue images and deep learning; inputting the sample data into a convolutional neural network to train the model, constructing sagittal and vertical skeletal deformity automatic diagnosis classification models respectively, and generating a multi-stage neural network model with different parameters after training;
[0014] Step 4, inputting the sample to be screened into the multi-stage neural network model with different parameters trained in step 3 to obtain the prediction result of the skeletal deformity screening.
[0015] Preferably, in step 1, the samples obtained by sample data acquisition are randomly divided into a training set, a validation set and a test set in proportion; the training set is used for model training, the validation set is used for hyperparameter optimization, and the test set is used for final model performance evaluation.
[0016] Preferably, in step 2, manual positioning is adopted to mark the landmark points related to sagittal and vertical classification of bone anatomy, record the position information of the landmark points, and classify according to the diagnostic criteria.
[0017] More preferably, the landmark points include: Sn-GoGn angle landmark point, the triangular and landmark point, ANB angle landmark point and Wits value landmark point;
[0018] The classification includes: sagittal diagnostic classification of bone type I, II and III, and vertical diagnostic classification of normal angle, high angle and low angle.
[0019] Preferably, in step 3, the automatic diagnostic classification model comprises:
[0020] A facial contour feature extraction module is used to process the facial contour image by constructing a deep separable convolutional network; for a side photo, first, facial key anatomical landmark points (including frontal points, nasal root points, nasal tip points, subnasal points, lip points, chin points, etc.) are extracted to generate a facial contour map. This module is based on a deep convolutional structure design, combined with an attention mechanism to filter out key regional features closely related to skeletal deformities, providing efficient representation for subsequent fusion.
[0021] A cephalogram analysis module is used to process the cephalogram by using a multi-scale residual network architecture; this module designs a multi-scale structure and residual connection for the cephalogram, which can simultaneously extract local and global features of the image, and retain key structural information related to orthodontic diagnosis (such as the position relationship between the upper and lower jaws), and the results of this part are used as the "gold standard reference features" for skeletal classification.
[0022] A multi-modal fusion module is used to effectively fuse the contour map, cephalogram and facial photo information by establishing a multi-level feature alignment and weighting mechanism; to effectively fuse the contour map, cephalogram and facial photo information, this module establishes a multi-level feature alignment and weighting mechanism, and through the attention interaction between different modal features, the relevance and discriminability of feature expression are improved. The fusion process is adaptive, and the contribution weight of each modality can be dynamically adjusted according to the importance of the input features.
[0023] A physiological feature assisted modeling module is used to realize the fusion of auxiliary features and image features through a dual-channel architecture by using a gating attention mechanism; considering the influence of individual factors such as gender, age, height and weight on the craniofacial structure, an auxiliary branch network is introduced to process these non-image continuous and categorical variables, and the image features are fused. This module improves the recognition ability of individual differences through the gating mechanism, thereby improving the generalization performance and prediction accuracy of the overall model.
[0024] More preferably, the face contour feature extraction module training process specifically includes:
[0025] In the face contour feature extraction stage, first, a depth separable convolutional network is constructed for processing the face contour image; second, the anatomical landmark positioning point coordinates are identified, including the highest point of the forehead, the lowest point between the eyebrows, the most prominent point of the nose tip, the junction between the nasal columella and the upper lip, the most prominent point of the upper lip, the most prominent point of the lower lip, the most forward point of the soft tissue chin, the most lower point of the soft tissue chin, and the anterior edge point of the tragus; third, a 256-dimensional feature representation is output by the feature encoder, and a face contour image is generated.
[0026] More preferably, the cephalogram analysis module training process specifically includes:
[0027] A feature pyramid structure is adopted, the bottom layer uses a 7x7 large kernel convolution to capture global features, the middle layer extracts local features through three residual blocks, and the high layer uses a dilated convolution to expand the receptive field; each residual block contains two convolution layers and a shortcut connection, and uses group normalization GN and GeLU activation function; a spatial attention mechanism is provided at the top of the network to highlight key region features through channel weighting; local and global features of the cephalogram are extracted at the same time to retain key structural information related to orthodontic diagnosis.
[0028] More preferably, the multi-modal fusion module training process specifically includes:
[0029] The three-level attention fusion strategy: the primary fusion layer performs spatial alignment and splicing on the pixel-level features; the middle-level fusion layer implements channel attention reweighting, dynamically adjusts the contribution of each modal feature through the SE module; the high-level fusion layer adopts cross-attention mechanism to establish the correlation mapping between the face features and the contour features, and the X-ray image features.
[0030] The fusion weight is adjusted by a learnable temperature coefficient τ, which dynamically adjusts the contribution weight of each modal according to the importance of the input feature, and realizes dynamic feature fusion by using an attention-based soft assignment strategy.
[0031] More preferably, the physiological feature auxiliary modeling module training process specifically includes:
[0032] In the dual-path architecture, the continuous feature path uses a three-layer fully connected network to process age, gender, height, and weight information, each layer is followed by LayerNorm and Dropout; the categorical feature path uses an embedding layer to process gender information and then splices it with the continuous features; the auxiliary features and image features are fused through a gated attention mechanism, and the gating weight is generated by a sigmoid function.
[0033] Preferably, in step 4, the difference between the obtained diagnosis prediction result of the skeletal deformity and the actual gold standard is compared, the accuracy of the automatic diagnosis classification model is evaluated through the accuracy and the change of the loss function, and the model structure and the super parameter are optimized accordingly.
[0034] The beneficial effects of the present application are:
[0035] 1. The present application realizes an intelligent recognition framework of "multi-modal fusion in training phase and single-modal deployment in inference phase", and significantly improves the classification accuracy and generalization ability of the model for skeletal deformity by combining structure perception, semantic guidance and feature constraint mechanism; compared with the traditional process relying on artificial measurement and radiographic imaging, the present application greatly reduces the risk of artificial participation and radiation, has advantages of early identification, convenient deployment, low-cost screening, etc., and has wide application prospect and promotion value in the scenes of children screening, family monitoring and portable diagnosis and treatment, etc.
[0036] 2. The contour point guidance mechanism of the present application introduces facial contour key points into the skeletal deformity discrimination process, and establishes explicit connection for facial feature representation and deep skeletal relationship.
[0037] 3. The individual physiological feature constraint of the present application uses gender, age and other feature information to constrain the deep learning process, guides the network to mine more discriminative features in heterogeneous samples, and improves the accuracy and generalization ability of the model in actual application.
[0038] 4. The multi-modal fusion modeling of the present application fuses facial contour structure, gold standard information of cephalogram and individual physiological features for deep learning, constructs a multi-branch deep learning framework, cooperatively extracts image features of different scales and depths, breaks through the limitation of single modal for skeletal deformity recognition, and significantly improves the prediction accuracy; the vertical diagnosis classification is included, and the sagittal diagnosis classification is included to form a comprehensive output result.
[0039] 5. The present application can be de-artificialized and de-radiated, compared with the traditional method relying on radiographic image and doctor measurement, the present application significantly reduces the risk of artificial dependence and radiation exposure, is suitable for early screening and family self-monitoring of sensitive groups such as children, adolescents and pregnant women, and has good application prospect and transformation value. BRIEF DESCRIPTION OF DRAWINGS
[0040] Figure 1 is the schematic diagram of collecting two-dimensional soft tissue images of patients in the present application of facial contour guidance and multi-modal feature fusion of orthodontic skeletal deformity screening method;
[0041] Figure 2 is the sagittal classification index schematic diagram of the present application: (a) ANB angle; (b) Wits value;
[0042] Figure 3This is a schematic diagram of the vertical classification index of the present invention: (a) SN-GoGn angle; (b) Triangle sum;
[0043] Figure 4 This is a schematic diagram of the present invention, which shows the process of manually calibrating after marking the landmark to obtain a dotted image that conforms to the contour of facial soft tissue.
[0044] Figure 5 This is a schematic diagram of the overall design of the automatic diagnosis and classification model for orthodontic skeletal deformities based on two-dimensional plane and deep learning in this invention.
[0045] Figure 6 This is a schematic diagram of the algorithm framework for building an automatic diagnostic classification model for predicting bony deformities from two-dimensional facial photographs according to the present invention;
[0046] Figure 7 This is an application example diagram of the automatic diagnosis and classification applet for orthodontic skeletal deformities based on two-dimensional plane images according to the present invention;
[0047] Figure 8 This is the sagittal skeletal deformity prediction label map of the present invention;
[0048] Figure 9 This is the vertical skeletal deformity prediction label map of the present invention. Detailed Implementation
[0049] The related technologies of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0050] like Figures 1-9 As shown, this embodiment includes the following steps:
[0051] 1. Obtain frontal and lateral soft tissue images of the target subject's face and lateral cephalometric radiographs.
[0052] The sample data used in this implementation method were collected from the Dentofacial Development Management Center of a university dental hospital in China, from 2023 to 2025, with informed consent obtained and approved by the hospital's ethics committee (ethics approval number: 2023-XJKQIEC-043-001). All study subjects met the following inclusion criteria and were strictly screened according to the following exclusion criteria:
[0053] 1. History of systemic diseases, medication use, maxillofacial trauma, and surgery;
[0054] 2. Jaw cysts, tumors and other lesions;
[0055] 3. Large area of intraoral prosthesis or edentulous;
[0056] 4. There is obvious enlargement or reduction of the radiograph, there is shielding, and the image clarity is poor.
[0057] All study subjects were collected the following information:
[0058] 1. Two-dimensional soft tissue images (photos) of facial frontal and lateral position;
[0059] 2. Corresponding lateral cephalogram;
[0060] 3. Age, gender, height, weight information.
[0061] Canon (Japan, Model-No: DS126631) was used to take frontal and lateral soft tissue image photos, vertical resolution 72 dpi, horizontal resolution 72 dpi, bit depth 24, output format. JPG. (such as Figure 1 ) Sirona Dental Systems GmbH (Germany, Model-No: D3352) system was used to take lateral cephalogram, and JPG format lateral cephalogram was obtained, and the accurate actual gender, age, height, weight of the target object were recorded.
[0062] A total of 1761 sample data were collected, with an age range of 8-67 years old. In order to analyze the age distribution, the samples were divided into 6 age groups every 10 years (see Table 1). Each sample contains three types of image data: frontal facial photo, lateral facial photo and corresponding lateral cephalogram.
[0063] In order to construct a deep learning automatic diagnosis model, all samples were randomly divided into training set, validation set and test set according to the ratio of 7:2:1. Among them, the training set was used for model training, the validation set was used for hyperparameter optimization, and the test set was used for final model performance evaluation. This dataset provides a complete input basis for subsequent establishment of automatic diagnosis model of skeletal deformity based on two-dimensional soft tissue image. Table 1 shows the age and gender distribution of all samples.
[0064] Age (y) Male (n) Female (n) Total (n) 8-17 376 704 1080 18-27 95 495 590 28-37 28 51 79 38-47 3 5 8 48-57 1 1 2 58-67 1 1 2 Total (n) 504 1257 1761
[0065] Table 1
[0066] II. Diagnosis label construction and image preprocessing.
[0067] 1. Artificial fixed point and gold standard diagnosis label establishment:
[0068] The application adopts artificially labeled cephalogram diagnostic information as a gold standard to establish a high-accuracy deep learning model. In the data set, all cephalograms are manually labeled by Labelme 4.5.10 software, typical bone anatomical measurement indexes related to sagittal and vertical bone deformity judgment are selected, including ANB angle, Wits value, SN-GoGn angle and triangle and equal key points (see Figure 2 、 3 ). For samples with structural overlap, the midpoint of the connecting line of the left and right side key points is taken as the final labeling point coordinates after labeling.
[0069] The meanings of the sagittal and vertical classification indexes are as follows:
[0070] (1) Sagittal diagnostic classification:
[0071] 1) ANB angle: the angle formed by the upper alveolar point, the nasal root point and the lower alveolar point, that is, the difference between the SNA angle and the SNB angle, which reflects the mutual position relationship of the upper and lower jaw bones with the skull.
[0072] 2) Wits value: the distance between the two vertical foot points made from the upper and lower alveolar points A and B to the functional plane to reflect the mutual position relationship of the upper and lower jaw bones.
[0073] (2) Vertical diagnostic classification:
[0074] 1) Sn-GoGn angle: the intersection angle of the anterior cranial base plane (SN) and the connecting line (Go-Gn) between the gonial point and the gnathion point, which represents the steepness of the mandibular body, the size of the gonial angle and also reflects the height of the face.
[0075] 2) Triangle: the angle is the sum of the sella turcica angle (N-S-Ar), the joint angle (S-Ar-Go) and the mandibular angle (Ar-Go-Me). In normal individuals, there is a mutual compensation relationship between the three angles. According to the value, it can be divided into normal mandibular development, counterclockwise rotation growth trend and clockwise rotation growth trend.
[0076] The diagnostic classification standards of the sagittal and vertical of the cephalogram sample are as follows:
[0077] (1) Sagittal diagnostic classification:
[0078] 1) ANB angle is between 0°-4°, bone type I; >4°, bone type II; <0°, bone type III;
[0079] 2) Wits value is between 0 and -2 for bone type I; ≥0 for bone type II; ≤-2 for bone type III.
[0080] (2) Vertical diagnostic classification:
[0081] 1) SN-GoGn angle < 27° is low angle, 27°-37° is normal angle, > 37° is high angle;
[0082] 2) Triangular sum < 390° is low angle, between 390°-402° is normal angle, > 402°
[0083] is high angle.
[0084] To further ensure the accuracy and stability of the fixed point, before the formal test begins, two fixed point personnel randomly selected 300 samples from all samples for fixed point and measurement, and after processing the fixed point measurement results, the correlation values of Sn-GoGn value, triangular sum, ANB angle and Wits value were obtained. Two weeks after the initial measurement, a random selected one of the original fixed point personnel repeated the fixed point measurement of the 300 samples. The results obtained by the two fixed point measurements were tested for consistency using the intraclass correlation coefficient (ICC) to evaluate the intra-group and inter-group differences of the fixed point personnel's fixed point measurement results. If the ICC value is greater than 0.75, it is considered that the repeatability and reliability of the fixed point measurement results are good, and the fixed point personnel can perform the fixed point marking of the formal samples; if the ICC value is less than 0.75, the repeatability and reliability of the human measurement results are poor, and the fixed point results cannot be used as samples for machine learning.
[0085] In the formal experiment, after the fixed point marking task of all the head lateral film samples was completed by the fixed point personnel, three doctors with 15-20 years of orthodontic clinical experience simultaneously checked the fixed point results of each head lateral film image. The measurement indicators between the landmark points were obtained using a python script, and the sagittal and vertical diagnostic classification results of each sample were determined to obtain the gold standard information of each sample. To improve the accuracy of machine learning and avoid interference, the training set samples built by this method have removed low-quality image data such as critical values and contradictory values. The consistency test of the fixed point results is shown in Table 2.
[0086]
[0087] The sum of N-S-Ar angle, S-Ar-Go angle and Ar-Go-Me angle is triangular sum.
[0088] Table 2
[0089] 2. Soft tissue image preprocessing and contour labeling:
[0090] Meanwhile, all the two-dimensional soft tissue facial profile photos in frontal and lateral positions were marked with facial contour key points by Landmark 1.3.5 software, and the point-like contour images conforming to the facial structure were obtained through auditing and calibration. Figure 4
[0091] On this basis, standardization preprocessing operations were performed on the images, including image cropping (removing irrelevant areas), brightness and contrast adjustment (eliminating light difference), image graying and sharpening (enhancing facial features, removing the influence of shooting light, uneven skin color, etc.), random rotation and scaling (improving model robustness and generalization ability), so as to optimize the input quality of subsequent model training.
[0092] III. Constructing an automatic diagnostic classification model based on two-dimensional soft tissue images and deep learning.
[0093] The system architecture includes four core modules: facial contour feature extraction module, cephalogram analysis module, multi-modal fusion module and physiological feature auxiliary module, realizing automatic screening of skeletal deformities based on two-dimensional soft tissue images.
[0094] The overall framework of the experimental design of the application is shown in Figure 5 , wherein the algorithm model design is shown in Figure 6 .
[0095] 1761 two-dimensional soft tissue image data (each including three types of data: portrait photo, point contour image and cephalogram) were divided into training set, validation set and test set in the ratio of 7:2:1. Then the learning samples of the labeled point information, age information, gender information, height and weight information and image information were uniformly input, and the sagittal and vertical skeletal deformity automatic screening classification models were constructed.
[0096] Specifically, the model training process is as follows:
[0097] 1. Facial contour feature extraction module:
[0098] In the facial contour feature extraction stage, a deep separable convolution network is first constructed to process the facial contour image. This module is used to identify the coordinates of 53 anatomical landmark positioning points on the soft tissue profile, including the highest point on the forehead (soft tissue forehead point), the lowest point between the eyebrows (nose root point), the most prominent point of the nose tip (nose tip point), the junction between the nasal columella and the upper lip (nose lower point), the most prominent point of the upper lip (upper lip protrusion point), the most prominent point of the lower lip (lower lip protrusion point), the most forward point of the soft tissue chin (soft tissue chin front point), the most downward point of the soft tissue chin (soft tissue chin lower point), and the anterior edge point of the tragus (tragus point). The network adopts a five-layer deep separable convolution structure, each layer containing 32 3x3 convolution kernels, with batch normalization and LeakyReLU (a=0.1) activation. The feature encoder outputs a 256-dimensional feature representation. The facial contour image is generated. This module is based on a deep convolution structure design, combined with an attention mechanism to filter out key regional features closely related to skeletal deformities. The output of the model is saved as parameter 1.
[0099] 2. Head lateral radiograph analysis module:
[0100] The head lateral radiograph analysis module uses a multi-scale residual network architecture to process the head lateral radiograph. The network contains a feature pyramid structure, the bottom layer uses a 7x7 large kernel convolution to capture global features, the middle layer extracts local features through three residual blocks, and the high layer uses a dilated convolution to expand the receptive field. Each residual block contains two convolution layers and a shortcut connection, using group normalization (GN) and GeLU activation function. A spatial attention mechanism is designed at the top of the network to highlight key regional features through channel weighting. The multi-scale structure and residual connection ensure the simultaneous extraction of local and global features of the head lateral radiograph, retaining key structural information related to orthodontic screening (such as the position relationship between the upper and lower jaws). The results of this part are used as parameter 2 for skeletal classification.
[0101] 3. Multi-modal fusion module:
[0102] To effectively fuse the contour image, head lateral radiograph, and portrait photo information, this module establishes a multi-level feature alignment and weighting mechanism, a three-level attention fusion strategy: the primary fusion layer performs spatial alignment and concatenation on pixel-level features; the middle fusion layer implements channel attention reweighting, dynamically adjusting the contribution of each modality feature through the SE module; the high-level fusion layer uses cross-attention mechanism to establish the correlation mapping between portrait features and contour features, X-ray image features. The fusion weight is adjusted by a learnable temperature coefficient τ, which can dynamically adjust the contribution weight of each modality according to the importance of the input feature. The attention-based soft assignment strategy is used to realize dynamic feature fusion.
[0103] 4. Physiological feature assisted modeling module:
[0104] The physiological feature auxiliary module is designed as a double-channel architecture: the continuous feature channel adopts a three-layer fully connected network to process age, gender, height and weight information, and each layer is followed by LayerNorm and Dropout (0.2); the category feature channel adopts an embedding layer to process gender information and then splices with the continuous features. The auxiliary features and image features are fused through a gated attention mechanism, and the gating weight is generated through a sigmoid function.
[0105] The model optimization adopts an improved multi-task learning framework, the main task uses Focal Loss (γ = 2.0) to handle class imbalance, and the auxiliary task adopts Huber loss. The training process is divided into three stages: first, fix the backbone network and only train the auxiliary module (10 rounds); then unfreeze the network for end-to-end training (50 rounds); finally, fine-tune the fusion layer parameters (20 rounds). The learning rate adopts a cosine annealing strategy, with an initial value of 3e-4 and a minimum of 1e-6.
[0106] The whole model adopts a phased training strategy: first, train the auxiliary branch to stabilize the secondary feature guidance effect; then unfreeze the backbone network for end-to-end training; finally, fine-tune the fusion module to improve the model's integration capability of multi-modal information. Through this multi-branch fusion structure, the initial algorithm model of the present application is determined.
[0107] 5. Model verification and super parameter optimization:
[0108] The validation set samples are input into the multi-stage neural network model with different parameters that have been constructed, and the prediction results of the skeletal deformity diagnosis are obtained. By comparing the differences between the skeletal deformity diagnosis prediction results obtained in the validation set and the actual gold standard, the accuracy of the automatic diagnosis classification model is evaluated through accuracy, loss function changes and other indicators, and the model structure and super parameters are optimized accordingly, and finally the best automatic classification model of the present research is determined.
[0109] IV. Effect evaluation of the automatic diagnosis model.
[0110] To evaluate the performance of the multi-modal fusion automatic diagnosis model constructed by the present application, the test set samples are input into the final model that has been trained, and compared with the gold standard classification results obtained by manual point measurement of the lateral cephalogram. The evaluation indicators include prediction accuracy, precision, recall and F1 value, etc. key performance indicators, among which the accuracy is the most important evaluation indicator. The evaluation results are shown in Table 3:
[0111]
[0112] Table 3
[0113] The evaluation results show that the method achieves an average accuracy of 85.34% in the automatic classification task of sagittal bone deformity, which is about 11.65% higher than the previous studies. In the vertical bone deformity classification task, the model shows an average accuracy of 88.64%, which shows excellent discrimination ability and stability. The model application is shown in Figure 7 The small program end simulation demonstration interface developed by the present application can realize one-key bone deformity prediction and result visualization output for the public. For the classification accuracy, the highest prediction accuracy of sagittal bone deformity is 0.9153, as shown in Figure 8 The highest prediction accuracy of vertical high-angle deformity is 0.9032, as shown in Figure 9 .
[0114] In summary, the automatic diagnosis and classification method and system of orthodontic bone deformity based on facial contour point guidance and multi-modal data fusion proposed by the present application combine structure perception, semantic guidance and feature constraint mechanism, significantly improve the classification accuracy and generalization ability of the model for bone deformity, and realize the digital orthodontic screening process of multi-modal training and single-modal deployment. In the model construction stage, the feature information of facial photos, contour point maps and lateral cephalograms is fused to improve the learning ability of the model. In actual application, only two-dimensional facial photos need to be input, and the comprehensive judgment of sagittal and vertical bone deformity types can be automatically completed, and accurate screening classification results can be output.
[0115] Compared with the traditional process relying on manual measurement and radiographic imaging, the present application has good generalization performance and clinical practicability, has the advantages of early identification, convenient deployment, low-cost screening, etc., greatly reduces the dependence on artificial and radiation risk, and can be widely applied to orthodontic screening, family self-test, early bone development trend monitoring of children and adolescents, etc. It provides a feasible scheme and technical foundation for non-invasive and low-cost intelligent auxiliary screening, and has wide application prospect and popularization value.
[0116] It should be emphasized that: the above is only the preferred embodiment of the present application, and does not limit the present application in any form, any simple modification, equivalent change and modification of the above embodiment according to the technical essence of the present application still belongs to the scope of the technical solution of the present application.
Claims
1. A method of orthodontic skeletal malformation screening with facial profile guidance and multi-modal feature fusion, characterized in that, The method comprises the following steps: Step 1, sample data acquisition; obtain the facial frontal soft tissue image, facial lateral soft tissue image and lateral skull X-ray film of the target sample; Step 2, diagnosis label construction and image preprocessing; artificial positioning is performed on each sample obtained in step 1, and a diagnosis classification label is obtained; the key points of the facial contour are marked on the two-dimensional soft tissue photo, so that the contour marking points accurately fit the facial contour of the sample person, and an accurate soft tissue contour point image is obtained; Step 3, constructing an automatic diagnosis classification model based on two-dimensional soft tissue images and deep learning; sample data is input into a convolutional neural network to train the model, and automatic diagnosis classification models for sagittal and vertical bone deformities are constructed, and after training, a multi-stage neural network model with different parameters is generated; Step 4, inputting the sample to be screened into the multi-stage neural network model with different parameters trained in step 3 to obtain the prediction result of bone deformity screening.
2. The method of claim 1, wherein the method further comprises: In step 1, the samples obtained by sample data acquisition are randomly divided into training set, validation set and test set according to proportion; the training set is used for model training, the validation set is used for hyperparameter tuning, and the test set is used for final model performance evaluation.
3. The method of facial profile guided and multi-modal feature fused screening of orthodontic skeletal malocclusions as claimed in claim 1, wherein, In step 2, when artificial positioning and diagnosis classification label are obtained, the mark points related to sagittal and vertical classification of bone anatomy structure are marked by manual positioning, the position information of the mark points is recorded, and classification is performed according to the diagnosis standard.
4. The method of facial profile guided and multi-modal feature fused screening of orthodontic skeletal malocclusions as claimed in claim 3, wherein, The landmark points include: Sn-GoGn angle landmark point, Triangular and landmark points, ANB angle degree angle landmark point, and Wits value landmark point; The classification includes: sagittal diagnosis classification of bone type I, II and III, and vertical diagnosis classification of normal angle, high angle and low angle.
5. The method of facial profile guided and multi-modal feature fused screening of orthodontic skeletal malocclusions as claimed in claim 1, wherein, In step 3, the automatic diagnosis classification model comprises: A facial contour feature extraction module for processing facial contour images by constructing a depth separable convolutional network; A lateral skull film analysis module for processing lateral skull films by using a multi-scale residual network architecture; A multi-modal fusion module for effectively fusing contour map, lateral skull film and portrait photo information by establishing a multi-level feature alignment and weighting mechanism; A physiological feature auxiliary modeling module for realizing the fusion of auxiliary features and image features through a gating attention mechanism by a double-channel architecture.
6. The method of facial profile guided and multi-modal feature fused screening of orthodontic skeletal malocclusions as claimed in claim 5, wherein, The facial contour feature extraction module training process specifically comprises: In the facial contour feature extraction stage, first, a depth separable convolutional network is constructed to process facial contour images; second, anatomical landmark positioning point coordinates are identified, including: the highest point of the forehead, the lowest point between the eyebrows, the most prominent point of the nose tip, the junction of the nasal columella and the upper lip, the most prominent point of the upper lip, the most prominent point of the lower lip, the most forward point of the soft tissue chin, the most lower point of the soft tissue chin, and the point of the anterior edge of the tragus; third, a 256-dimensional feature representation is output by a feature encoder, and a facial contour map is generated.
7. The method of facial profile guided and multi-modal feature fused orthodontic skeletal malocclusion screening of claim 5, wherein, The lateral skull film analysis module training process specifically comprises: The feature pyramid structure is adopted, the 7*7 large kernel convolution is adopted in the bottom layer to capture the global feature, the local feature is extracted through three residual blocks in the middle layer, and the high layer uses the hollow convolution to expand the receptive field; each residual block contains two convolution layers and shortcut connection, adopts group normalization GN and GeLU activation function; the spatial attention mechanism is arranged at the top of the network, and the key area feature is highlighted through channel weighting; the local and global features of the head lateral film are extracted at the same time, and the key structure information related to the orthodontic diagnosis is reserved.
8. The method of facial profile guided and multi-modal feature fused screening of orthodontic skeletal malocclusions as claimed in claim 5, wherein, The multi-modal fusion module training process specifically includes: The three-level attention fusion strategy: the primary fusion layer performs spatial alignment and splicing on the pixel-level features; the middle fusion layer implements channel attention reweighting, and dynamically adjusts the contribution of each modal feature through the SE module; the high-level fusion layer adopts the cross-attention mechanism, and establishes the correlation mapping between the face feature and the contour feature and the X-ray image feature; The fusion weight is adjusted by the learnable temperature coefficient tau, the contribution weight of each modal is dynamically adjusted according to the importance of the input feature, and the feature dynamic fusion is realized by adopting the attention-based soft allocation strategy.
9. Use of the orthopedic screening method for facial profile guidance and multi-modal feature fusion according to claim 8, characterized in that, The physiological feature auxiliary modeling module training process specifically includes: In the double-channel architecture, the continuous feature channel adopts a three-layer fully connected network to process age, gender, height and weight information, and each layer is connected with LayerNorm and Dropout; the category feature channel adopts an embedding layer to process gender information and is spliced with the continuous feature; the auxiliary feature and the image feature are fused through the gated attention mechanism, and the gating weight is generated through the sigmoid function.
10. Use of the method of orthodontic skeletal malformation screening with facial profile guidance and multi-modal feature fusion according to claim 1, characterized in that, In step 4, the difference between the obtained bone deformity diagnosis prediction result and the actual gold standard is compared, the accuracy of the automatic diagnosis classification model is evaluated through the accuracy and the loss function change, and the model structure and the hyperparameters are optimized accordingly.
Citation Information
Cited By
Orthodontic treatment prediction system and method based on multi-modal fusion and diffusion flow, electronic equipment and storage medium
CN121885164A