Risk assessment method and system for influence of malformation on posterior tooth adjacent face caries

By using machine learning models to screen for risk factors associated with malocclusion and assess the probability of caries on the proximal surfaces of posterior teeth, this approach addresses the lack of research on the correlation between malocclusion and caries in existing technologies, thus achieving more accurate risk assessment.

CN121460176APending Publication Date: 2026-02-03STOMATOLOGICAL HOSPITAL TIANJIN MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511641776.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-11
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

The current technology lacks high-quality longitudinal studies on the correlation between malocclusion and dental caries, and the application of machine learning methods in the analysis is insufficient, which makes it impossible to effectively assess the risk of malocclusion to proximal caries of posterior teeth.

Method used

Machine learning models, including optimal subset regression, Lasso regression, and random forest algorithms, were used to screen risk factors and draw nomograms based on feature data such as age, gender, and brushing frequency to assess the probability of posterior proximal caries.

Benefits of technology

This study provides a new method for predicting the risk of proximal caries in posterior teeth, which improves the accuracy and reliability of risk assessment by incorporating data related to malocclusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121460176A_ABST
    Figure CN121460176A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of machine learning and oral medical treatment, in particular to a risk assessment method and system for the influence of malformation on posterior tooth adjacent face caries disease, and the method comprises the steps: obtaining the feature data of a to-be-assessed target object; performing variable assignment on the feature data of the to-be-evaluated target object to obtain predicted variable data; screening the predictive variable data based on a risk assessment model to obtain risk factor data; drawing a column graph according to the screened risk factor data; and according to the risk factor data and the column diagram, obtaining the illness probability of the posterior tooth adjacent face caries of the to-be-evaluated target object. On the basis of data including ANB angle data, Wits value data, APDI angle data, Angle classification data, dentition crowding degree data and the like related to error deformity, the invention provides the model for predicting the posterior tooth adjacent face caries and the method for predicting the risk of the posterior tooth adjacent face caries on the basis of the model, and a new method for predicting the risk of the posterior tooth adjacent face caries in the prior art is supplemented.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of machine learning and oral medical technology, in particular to a risk assessment method and system for the influence of malocclusion on the prevalence of interproximal caries of posterior teeth. BACKGROUND

[0002] Dental caries is a progressive and destructive disease of dental hard tissue caused by various factors, mainly bacteria, which manifests as demineralization of inorganic matter and decomposition of organic matter. With the advent of the ecological plaque hypothesis, the definition of dental caries has been updated and supplemented, and dental caries is considered to be a disease caused by changes in the ecological flora of oral biofilm. So far, the four-factor theory of dental caries is still the recognized etiology of dental caries, and most of the literature on the etiology of dental caries is the influence of bacterial plaque and other factors on the occurrence and development of dental caries, while the research on malocclusion as a factor is not clear.

[0003] Currently, there are few studies on the correlation between malocclusion and dental caries internationally, and most of the literature reports only analyze the correlation between certain occlusal features of malocclusion and the caries of patients. The influencing factors of dental caries are numerous, most studies lack confounding factors, and the application of machine learning and other methods is insufficient, limiting the modeling potential of complex associations and potential confounding interactions. Therefore, there is no reliable evidence to prove the correlation between the two, and high-quality longitudinal studies are still lacking to explain this correlation. SUMMARY

[0004] Therefore, an object of the present application is to provide a risk assessment method and system for the influence of malocclusion on the prevalence of interproximal caries of posterior teeth to solve the problems mentioned in the background art and overcome the deficiencies in the prior art.

[0005] To achieve the above-mentioned purpose, the following technical solutions are adopted in the present application: A risk assessment method for the influence of malocclusion on the prevalence of interproximal caries of posterior teeth, comprising: obtaining feature data of a target object to be evaluated; performing variable assignment on the feature data of the target object to be evaluated to obtain prediction variable data; screening the prediction variable data based on a risk assessment model to obtain risk factor data; drawing a nomogram according to the screened risk factor data; obtaining the prevalence probability of interproximal caries of posterior teeth of the target object to be evaluated according to the risk factor data and the nomogram.

[0006] Preferably, the characteristic data comprises age data, gender data, tooth brushing frequency data, tooth brushing time data, frequency of eating sweets data, whether using dental floss data, Angle classification data, dentition crowding data, ANB angle data, FH-MP angle data, Wits value data, APDI angle data, and ODI angle data.

[0007] Preferably, the risk assessment model is a machine learning model, comprising an optimal subset regression algorithm, a Lasso regression algorithm, and a random forest algorithm.

[0008] Preferably, the screening of the predictive variable data based on the risk assessment model to obtain risk factor data comprises screening the predictive variable data based on the optimal subset regression algorithm, including all predictive variable data in the optimal subset regression analysis, screening the optimal combination number according to the CP principle, determining the optimal combination number when the CP value is the smallest, and obtaining the risk factor data based on the optimal subset regression algorithm.

[0009] Preferably, the screening of the predictive variable data based on the risk assessment model to obtain risk factor data comprises screening the predictive variable data based on the Lasso regression algorithm, screening the most significant predictive markers on the training set by the Lasso logistic regression algorithm as the risk factor data based on the Lasso regression algorithm.

[0010] Preferably, the screening of the predictive variable data based on the risk assessment model to obtain risk factor data comprises screening the predictive variable data based on the random forest algorithm, evaluating all predictive variable data by the random forest model, calculating the importance score of each predictive variable data, excluding the predictive variable data with lower importance according to the set threshold or ranking, and retaining the predictive variable data as the risk factor data based on the random forest algorithm.

[0011] Preferably, the risk factor data based on the optimal subset regression algorithm comprises age data, dentition crowding data, and ANB angle data, a nomogram is drawn according to the risk factor data based on the optimal subset regression algorithm, a first total score value corresponding to the age data, the dentition crowding data, and the ANB angle data is obtained to obtain the probability of the target object suffering from the disease, the first total score value is the sum of a score value in the nomogram corresponding to the age data, a score value in the nomogram corresponding to the dentition crowding data, and a score value in the nomogram corresponding to the ANB angle data, wherein the age data, the ANB angle data, and the dentition crowding data are positively correlated with the probability of suffering from posterior proximal caries.

[0012] As preferred, the risk factor data based on the Lasso regression algorithm comprises age data, Angle classification data, dentition crowding data, ANB angle data, Wits value data and whether to use dental floss data, a nomogram is drawn according to the risk factor data based on the Lasso regression algorithm, a second total score value corresponding to the age data, the Angle classification data, the dentition crowding data, the ANB angle data, the Wits value data and the whether to use dental floss data is obtained, the second total score value is the sum of a score value in the nomogram corresponding to the age data, a score value in the nomogram corresponding to the Angle classification data, a score value in the nomogram corresponding to the dentition crowding data, a score value in the nomogram corresponding to the ANB angle data, a score value in the nomogram corresponding to the Wits value data, and a score value in the nomogram corresponding to the whether to use dental floss data, and a disease probability of the target object to be measured is obtained, wherein the age data, the Angle classification data, the dentition crowding data, the ANB angle data, the Wits value data and the whether to use dental floss data are positively correlated with the disease probability of the posterior tooth proximal caries.

[0013] As preferred, the risk factor data based on the random forest algorithm comprises age data, ANB angle data, Wits value data and APDI angle data, a nomogram is drawn according to the risk factor data based on the random forest algorithm, a third total score value corresponding to the age data, the ANB angle data, the Wits value data and the APDI angle data is obtained, the third total score value is the sum of a score value in the nomogram corresponding to the age data, a score value in the nomogram corresponding to the ANB angle data, a score value in the nomogram corresponding to the Wits value data and a score value in the nomogram corresponding to the APDI angle data, and a disease probability of the target object to be measured is obtained, wherein the age data, the ANB angle data and the Wits value data are positively correlated with the disease probability of the posterior tooth proximal caries, and the APDI angle data is negatively correlated with the disease probability of the posterior tooth proximal caries.

[0014] A risk assessment system for the influence of malocclusion on the disease of posterior tooth proximal caries, comprising: A data acquisition module is configured to acquire feature data of a target object to be evaluated; A variable assignment module is configured to perform variable assignment on the feature data of the target object to be evaluated to obtain prediction variable data; A data screening module is configured to screen the prediction variable data based on a risk assessment model to obtain risk factor data; A nomogram drawing module is configured to draw a nomogram according to the screened risk factor data; The risk assessment module is used to obtain the probability of posterior proximal caries in the target subject based on risk factor data and nomograms.

[0015] Therefore, the present invention has the following beneficial effects: This invention provides a risk assessment method and system for the impact of malocclusion on posterior proximal caries. Based on data including ANB angle data, Wits value data, APDI angle data, Angle classification data, and dental crowding data related to malocclusion, it provides a model for predicting posterior proximal caries and a method for predicting the risk of posterior proximal caries based on this model, thus supplementing existing methods for predicting the risk of posterior proximal caries. Attached Figure Description

[0016] Figure 1 This is a flowchart of a risk assessment method for the impact of malocclusion on the prevalence of proximal caries in posterior teeth, as described in this invention. Figure 2 This is an MPR interface diagram of the patient's dentofacial region observed by CBCT in an embodiment of the present invention; Figure 3 This is a CBCT image showing the presence and severity of caries on the proximal surfaces of posterior teeth, as described in an embodiment of the present invention. Figure 4 This is a diagram illustrating the diagnostic criteria and grading of proximal caries in posterior teeth according to an embodiment of the present invention. Figure 5 This is a diagram showing the A, B, and C sections of the improved proximal caries severity scoring standard according to an embodiment of the present invention. Figure 6 This is a schematic diagram of the ANB angle in an embodiment of the present invention; Figure 7 This is a schematic diagram of the FH-MP angle in an embodiment of the present invention; Figure 8 This is a schematic diagram of Wits values ​​in an embodiment of the present invention; Figure 9 This is a schematic diagram of APDI values ​​according to an embodiment of the present invention; Figure 10 This is a schematic diagram of ODI values ​​according to an embodiment of the present invention; Figure 11 This is a schematic diagram illustrating the screening of risk factors for posterior proximal caries based on the optimal subset method in an embodiment of the present invention. Figure 12 This is a schematic diagram illustrating the use of Lasso regression to screen characteristic factors for posterior proximal caries in an embodiment of the present invention. Figure 13 This is a schematic diagram illustrating feature selection and ranking based on the random forest algorithm in an embodiment of the present invention; Figure 14 The ROC curves for the three models—optimal subset, Lasso, and random forest—in this embodiment of the invention are shown. Figure 15 This is a calibration diagram of a posterior proximal caries risk assessment model based on three machine learning algorithms, according to an embodiment of the present invention. Figure 16 This is a clinical decision curve analysis diagram of a posterior proximal caries prediction model based on three machine learning algorithms according to an embodiment of the present invention. Figure 17 This is a nomogram of the optimal clinical risk assessment model for posterior proximal caries in an embodiment of the present invention. Detailed Implementation

[0017] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0018] like Figure 1 As shown, a risk assessment method for the impact of malocclusion on the prevalence of proximal caries in posterior teeth includes the following steps: S1: Obtaining characteristic data of the target object to be assessed; S2: Assigning variable values ​​to the characteristic data of the target object to obtain predictive variable data; S3: Filtering the predictive variable data based on the risk assessment model to obtain risk factor data; S4: Plotting a nomogram based on the filtered risk factor data; S5: Obtaining the prevalence of proximal caries in posterior teeth of the target object to be assessed based on the risk factor data and the nomogram.

[0019] Further, the feature data includes: age data, gender data, brushing frequency data, brushing time data, frequency of eating sweets data, whether dental floss is used data, Angle classification data, dental crowding data, ANB angle data, FH-MP angle data, Wits value data, APDI angle data, and ODI angle data.

[0020] Furthermore, the risk assessment model is a machine learning model, including the optimal subset regression algorithm, the Lasso regression algorithm, and the random forest algorithm.

[0021] Furthermore, the risk factor data obtained by filtering the predictor variable data based on the risk assessment model includes: filtering the predictor variable data based on the optimal subset regression algorithm, including all predictor variable data in the optimal subset regression analysis, selecting the number of optimal combinations according to the CP principle, and determining the number of optimal combinations when the CP value is the minimum, thus obtaining the risk factor data based on the optimal subset regression algorithm.

[0022] Furthermore, the risk factor data obtained by filtering the predictor variable data based on the risk assessment model includes: filtering the predictor variable data based on the Lasso regression algorithm, and selecting the most significant predictor markers on the training set using the Lasso logistic regression algorithm as the risk factor data based on the Lasso regression algorithm.

[0023] Furthermore, the risk factor data obtained by filtering the predictor variable data based on the risk assessment model includes: filtering the predictor variable data based on the random forest algorithm, evaluating all predictor variable data using the random forest model, calculating the importance score of each predictor variable data, excluding predictor variable data with lower importance according to the set threshold or ranking, and retaining the predictor variable data as risk factor data based on the random forest algorithm.

[0024] Furthermore, the risk factor data based on the optimal subset regression algorithm includes age data, dental crowding data, and ANB angle data. A nomogram is plotted based on the risk factor data based on the optimal subset regression algorithm. The disease probability of the target subject is obtained based on the first total score corresponding to the age data, dental crowding data, and ANB angle data. The first total score is the sum of the scores in the nomogram corresponding to the age data, the scores in the nomogram corresponding to the dental crowding data, and the scores in the nomogram corresponding to the ANB angle data. Among them, age data, ANB angle data, and dental crowding data are positively correlated with the disease probability of posterior proximal caries.

[0025] Furthermore, the risk factor data based on the Lasso regression algorithm includes age data, Angle classification data, dental crowding data, ANB angle data, Wits value data, and whether dental floss is used. A nomogram is plotted based on the risk factor data from the Lasso regression algorithm. The probability of disease in the target subject is obtained based on the second total score corresponding to the age data, Angle classification data, dental crowding data, ANB angle data, Wits value data, and whether dental floss is used. The second total score is the sum of the scores from the nomograms corresponding to the age data, Angle classification data, dental crowding data, ANB angle data, Wits value data, and whether dental floss is used. Age data, Angle classification data, dental crowding data, ANB angle data, Wits value data, and whether dental floss is used are positively correlated with the probability of proximal caries in posterior teeth.

[0026] Furthermore, the risk factor data based on the random forest algorithm includes age data, ANB angle data, Wits value data, and APDI angle data. A nomogram is plotted based on the risk factor data based on the random forest algorithm. The disease probability of the target subject is obtained based on the third total score corresponding to the age data, ANB angle data, Wits value data, and APDI angle data. The third total score is the sum of the scores in the nomogram corresponding to the age data, the ANB angle data, the Wits value data, and the APDI angle data. Among them, age data, ANB angle data, and Wits value data are positively correlated with the disease probability of posterior proximal caries, while APDI angle data is negatively correlated with the disease probability of posterior proximal caries.

[0027] In one example, based on the target subject's age, ANB angle, Wits value, and APDI angle, the corresponding values ​​are located on the four line segments of the nomogram, and the scores are calculated for the corresponding "score" line segments. The scores for age, ANB angle, Wits value, and APDI angle are added together to obtain the total score, which is then assigned to the "Total Points" line segment. The data on the bottom "Risk" line segment corresponding to the total score represents the probability of proximal caries in the target subject's posterior teeth.

[0028] The application of the nomogram for the predicted proximal caries model of posterior teeth based on the optimal subset algorithm is the same as that for the predicted proximal caries model based on the random forest algorithm, and will not be described in detail here.

[0029] The application of the nomogram for the predicted proximal caries model of posterior teeth based on the Lasso regression algorithm is the same as that for the predicted proximal caries model based on the random forest algorithm, and will not be described in detail here.

[0030] In this embodiment, a risk assessment method for the impact of malocclusion on proximal caries of posterior teeth according to the present invention is described in detail with reference to the accompanying drawings.

[0031] Patients aged 13-35 years who visited the orthodontics department of a medical university dental hospital in a certain month were selected. Complete CBCT imaging data and clinical examination data were obtained from the patients, including questionnaires, recorded model measurement data, and cephalometric analysis data. Based on the inclusion and exclusion criteria, 140 cases were ultimately included. All patients underwent imaging using a KaVo3DeXam CBCT scanner at the radiology department of the medical university dental hospital. The requirements were: patients should be seated upright with the orbitoauricular plane parallel to the ground, and the upper and lower lips naturally closed, ensuring the laser horizontal positioning line was located in the center of the upper and lower lips and parallel to the lip plane. Standard scanning parameters were set as follows: voltage 120KV, current 5mA, exposure time 4s, scanning field of view 16cm×13cm, and stereo pixels 0.25mm×0.25mm×0.25mm.

[0032] The above materials were screened, and the inclusion criteria included: complete CBCT imaging data and stored models, with clear CBCT images without artifacts and no defects in the stored models; willingness to participate in the questionnaire survey, with the survey content completed by the individual or their immediate family member, and the questionnaire information being complete and accurate; complete upper and lower dentition and in the permanent dentition stage, with no missing teeth, implants, removable partial dentures, or fixed bridge restorations. The exclusion criteria included: patients whose CBCT showed low-density shadows on both the proximal and occlusal surfaces of posterior teeth, making it impossible to determine the source of proximal caries; patients who had previously received or were currently receiving orthodontic treatment; patients whose posterior teeth had previously had cavities prepared on the proximal surfaces and had filling restorations; patients whose posterior teeth had filling restorations on the occlusal surfaces and low-density shadows around the affected teeth involving the proximal surfaces; patients with large-area defects in posterior teeth, with only one intact crown or root remaining; patients with retained deciduous teeth or supernumerary teeth in the posterior study area, making it impossible to determine the functional occlusal plane.

[0033] Clinical data collection includes: questionnaire data collection, demographic data such as age and gender, and surveys on oral hygiene and dietary habits such as brushing frequency (≤1 time per day, 2 times per day, >2 times per day), brushing time (less than 1 minute, 1-3 minutes, more than 3 minutes), whether dental floss is used, and frequency of eating sweets (foods containing sucrose) (≥2 times per day, once per day, 2-6 times per week, once per week, 1-3 times per month, rarely / never).

[0034] Data collection from patient memory models: Collect patient memory models, measure and record patients' Angle classification and dental crowding.

[0035] The Angle Classification includes: Angle Class I (neutral malocclusion): normal mesial and distal relationships of the maxilla and dental arches, and neutral molar relationships; Angle Class II (distal malocclusion): disharmony of mesial and distal relationships of the maxilla and dental arches, with the mandible and mandibular arch in a distal position and molars in a distal relationship; Angle Class III (mesial malocclusion): disharmony of mesial and distal relationships of the maxilla and dental arches, with the mandible and mandibular arch in a mesial position and molars in a mesial relationship.

[0036] Measurement of dental arch crowding: Dental arch crowding is the difference between the expected length of the dental arch and its current length. First, the expected length of the dental arch is measured, and then the current length of the dental arch is measured using a brass wire. The two values ​​are then subtracted to obtain a difference, which is then classified to represent the degree of crowding.

[0037] The specific method is as follows: Apply a brass wire from the mesial contact point of the first molar along the alveolar surface of the premolar to the cusp of the canine, and then along the incisal edge of the incisor to the mesial contact point of the opposite first molar. At this point, the brass wire will form an arc bent along the alveolar ridge. Straighten the wire and measure its length as the current length of the dental arch. Then, use a ruler to compare and measure the crown width of each tooth in the dental arch (anterior teeth and premolars) before the first molars on both sides, as the required length of the dental arch.

[0038] The crowding degree of the upper and lower dentition was measured separately and classified into three categories according to the degree of crowding: Grade I crowding: crowding degree ≤ 4mm; Grade II crowding: 4mm < crowding degree ≤ 8mm; Grade III crowding: crowding degree > 8mm.

[0039] Measurement of proximal caries severity, including: radiographic diagnosis of proximal caries and a modified scoring system for proximal caries severity in posterior teeth.

[0040] Imaging diagnosis of proximal caries: CBCT images of 140 patients were imported into a computer and analyzed using the MPR (Multiplanar Reconstruction) interface. The presence and progression of proximal caries in posterior teeth were determined according to the ICDAS imaging diagnostic criteria, expressed as a scale of 0-6. The specific measurement methods and evaluation criteria are as follows: After importing the patient's CBCT scan into KaVo 3DeXam Vision image processing software, the initial screen is an image preview interface. Click "Tools" at the top to enter the MPR interface. Move the red, green, and blue lines to the observation area respectively. You can then observe the images in the transverse, sagittal, and coronal planes. Figure 2 As shown.

[0041] Figure 2 Part a represents a cross-sectional view of the maxillofacial region, part b represents a sagittal view of the maxillofacial region, and part c represents a coronal view of the maxillofacial region. Blue dashed lines represent coronal plane positioning lines, green dashed lines represent sagittal plane positioning lines, and red dashed lines represent cross-sectional positioning lines.

[0042] like Figure 3 As shown, move Figure 3 The red dashed line in part b or part c makes Figure 3 Part a shows the location of the upper or lower dental arch. Taking the observation of proximal caries in the left dentition as an example, draw a parallel line (solid yellow line) along the mesiodistal direction of the posterior dental arch, starting at A and ending at B. Move this parallel line so that it is above the posterior dental arch. The length of this parallel line should exceed the distance between the mesial contact point of the first premolar and the distal point of the second molar. Place the mouse at any position on the AB line segment and drag the yellow line parallel to the indicated dotted line area to ensure that the dentition is cut by the parallel line in the cross-sectional view. Figure 3 (Part d) allows observation of any part of the tooth body in the study area from the first premolar to the second molar on one side, so as to make a later judgment on the caries status of the proximal surface of the tooth body.

[0043] The presence and severity of proximal caries were diagnosed according to the ICDAS imaging diagnostic criteria. A score of 0 indicates no caries, while grades 1-6 represent the degree of progression of proximal caries. Specific evaluation criteria are shown in Table 1, and the corresponding imaging images are shown below. Figure 4 .

[0044]

[0045] Figure 4 In the diagram, image a represents grade 1, where the transmissive image is limited to the outer half of the enamel; image b represents grade 2, where the transmissive image extends beyond the outer half of the enamel but does not exceed the dentin-enamel junction; image c represents grade 3, where the transmissive image is limited to the outer third of the dentin; image d represents grade 4, where the transmissive image reaches the middle third of the dentin; image e represents grade 5, where the transmissive image reaches the inner third of the dentin; and image f represents grade 6, where the transmissive image connects to the pulp cavity.

[0046] For the improved scoring criteria for the severity of proximal caries in posterior teeth, the patient's dentition is first divided into four quadrants. The study areas are the distal surface of the first premolar, the mesial and distal surfaces of the second premolar, the mesial and distal surfaces of the first molar, and the mesial surface of the second molar (excluding the mesial surface of the first premolar and the distal surface of the second molar). Within each quadrant, the area between the first and second premolars is designated as area A, then the area between the second and first premolars is designated as area B, and the area between the first and second molars is designated as area C. Each of areas A, B, and C has two adjacent tooth surfaces with a common contact area, totaling 24 tooth surfaces. Figure 5Secondly, each patient's teeth were divided into four quadrants: right maxilla, left maxilla, right mandible, and left mandible. Each quadrant had three regions (A, B, and C), totaling 12 regions. Each region had two tooth surfaces, for a total of 24 tooth surfaces. Each tooth surface was scored according to the ICDAS diagnostic criteria from 0 to 6 (see Table 1), resulting in 24 scores. The final score, obtained by adding up the 24 scores for each patient, represents the severity of proximal caries in that patient, as shown in Table 2.

[0047]

[0048] In the table, the numbers to the left and right of “│” indicate the scores for two adjacent tooth surfaces in that area.

[0049] The patient's final score = (4+0+0+1+0+6)+(0+4+0+1+0+0)+(0+1+0+2+0+0)+(0+0+0+1+0+1)=21.

[0050] It also includes the collection of cephalic shadow measurement index data, including: the definition of cephalic shadow measurement markers, the reference plane and the measurement plane, the determination method of cephalic shadow measurement indexes and their representative significance.

[0051] CBCT images of 140 patients were imported into the medical image file processing software Invivodental 5.3, and the generated lateral cephalometric radiographs were imported into the orthodontic cephalometric Zhibei Cloud Analysis Software for automatic point localization. The software automatically analyzed the data and compiled it into a table. Representative cephalometric indicators that might be influencing factors were individually screened, and the software automatically measured the angle or length of the required cephalometric indicators. The final measurement results were then compiled into a table.

[0052]

[0053] The methods for determining the cephalometric parameters and their representative meanings are as follows: ANB corner: such as Figure 6 As shown, the angle formed by the lines connecting points A, N, and B represents the anterior-posterior positional relationship between the upper and lower basal bones with the root of the nose as the reference point.

[0054] FH-MP angle: such as Figure 7 As shown, the angle between the mandibular plane and the orbitoauricular plane represents the steepness of the mandibular body and also reflects the height of the face.

[0055] Wits value (Ao-Bo): such as Figure 8 As shown, perpendicular lines are drawn from points A and B to the functional plane, obtaining points Ao and Bo respectively. The distance Ao-Bo is measured as the Wits value. This reflects the relationship between the anterior parts of the mandible and maxilla, avoiding the influence of the skull base structures.

[0056] APDI (Anter posterior Dysplasia Indicator): e.g. Figure 9 As shown, the sagittal misalignment index of the maxilla and mandible is composed of the facial angle, the angle of plane AB, and the angle of the palatal plane (APDI = ∠1 + ∠2 + ∠3). Note: When point B is in front of point A, the angle is positive; otherwise, it is negative. When the palatal plane is tilted forward and downward, the angle is negative; otherwise, it is positive.

[0057] ODI (Overbite Depth Indicator): Such as Figure 10 As shown, the sum of the angle between plane AB and the mandibular plane and the angle of the palatal plane (ODI = ∠1 + ∠2) is the vertical inconsistency index of the maxilla and mandible. Angles where the palatal plane tilts forward and downward are negative, and vice versa.

[0058] Data processing included variable assignment. Variable assignment included: predictor variables such as age, brushing frequency, brushing time, flossing, Angle classification, crowding, ANB angle, FH-MP angle, Wits value, and APDI; outcome variables such as presence and severity of proximal caries. The naming and assignment of the research variables are shown in Table 4.

[0059]

[0060] Statistical Analysis: Baseline characteristics were described, and all data were integrated and processed using R software. Normality was tested using the ShapiroWilktest. Normally distributed continuous data were described using mean ± standard deviation (X ± S) and analyzed using a t-test; non-normally distributed continuous data were described using median ± standard deviation (X ± S) and analyzed using a non-parametric test. Count data were described using percentages and analyzed using a chi-square test.

[0061] A risk assessment model for proximal caries of posterior teeth was established. The study variables included age, sex, brushing frequency, brushing time, frequency of sweets consumption, flossing, Angle classification, crowding, ANB angle, FH-MP angle, Wits score, APDI, and ODI. The outcome variable was the presence or absence of proximal caries, and results were expressed as binary variables. Optimal subset regression, Lasso regression, random forest, and three machine learning algorithms were used to select appropriate feature variables before establishing the risk assessment model.

[0062] Among them, the optimal subset regression: using the leaks package, the optimal subset regression equation is selected from all possible combinations of independent variables, and then CP is used as the criterion to screen the optimal combination of factors that affect the occurrence of proximal caries.

[0063] Lasso regression uses the glmnet package to store independent and dependent variables separately. Construct a Lasso regression model, where family="binomial". Use cross-validation to fit the model and select the optimal penalty coefficient λ. Present the cross-validation results. Select feature variables based on the principle of minimum error.

[0064] Random Forest: The random forest is trained using the randomForest package. It is assumed that the random forest uses K trees, and each tree requires a certain number of samples for training. The optimal number of variables (mtry) in a given node's binary tree and the optimal number of decision trees (ntree) in the given random forest are found to obtain the final classifier. The model performance and variable importance are then observed.

[0065] This invention also includes a performance comparison of risk assessment models: Receiver operating characteristic (ROC) curves are plotted, and the areas under the curves (AUC) of several prediction models are compared. Model calibration curves are plotted, and the Hosmer-Lemeshow test is applied to evaluate and calculate the p-value. If r < 0.05, it indicates a statistically significant difference between the model's predicted value and the actual value. Conversely, a larger p-value indicates a better predictive performance. Clinical decision curves (DCA) are applied to evaluate the clinical predictive value of the three prediction models, and net benefits are compared to select the risk assessment model with the best predictive performance.

[0066] The present invention also includes a nomogram display of the optimal prediction model: a nomogram is drawn based on the variables selected by the optimal model and displayed intuitively.

[0067] This invention uses the severity of proximal caries as the dependent variable, represented by a non-negative integer; all possible predictors selected from the three models above are used as independent variables. Poisson regression analysis is applied to analyze the influence of potential risk factors on the severity of proximal caries, providing the odds ratio (OR) and 95% confidence interval (95% CI) for each influencing factor, with r < 0.05 considered statistically significant.

[0068] To reduce measurement bias, the Kappa coefficient and intraclass correlation coefficient (ICC) were used to test the consistency of the measurement results. ① Model measurement: The Kappa test was applied, and the result showed a Kappa value of 0.74, indicating strong consistency with the model measurement results. ② Measurement of dentofacial parameters and proximal caries severity: Twenty patients were randomly selected for repeated measurements. One week later, the 20 patients were measured again. ICC was used to evaluate consistency, and the result showed that the ICC of the two measurements was 0.82, indicating good consistency.

[0069] This invention included 140 cases meeting the inclusion criteria, of which 72 patients had proximal caries in the posterior tooth region and 68 patients did not. The age range of the subjects was 13-35 years, with a mean age of 19.08±5.45 years; there were 43 males and 97 females. A total of 2240 teeth and 3360 tooth surfaces were observed on CBCT images, of which 164 tooth surfaces had proximal caries.

[0070] Clinical and imaging data of 140 patients who visited the orthodontics department were collected. The presence or absence of proximal caries was included as an outcome variable in the baseline characteristics of posterior proximal caries in the cases, as shown in Table 5.

[0071] Table 5. Bivariate characterization of potential influencing factors of proximal caries in posterior teeth

[0072] "*" indicates that continuous variables that conform to a normal distribution should be tested using a t-test, continuous variables that do not conform to a normal distribution should be tested using a non-parametric test, and categorical variables should be tested using a chi-square test.

[0073] This invention includes the establishment and evaluation of three risk assessment models: the presence or absence of proximal caries of posterior teeth is defined as the outcome variable, and possible predictive variables include gender, age, brushing frequency, brushing time, frequency of eating sweets, use of dental floss, Angle classification, crowding degree, ANB angle, FH-MP angle, Wits value, APDI and ODI. The variables are screened using three machine learning algorithms (optimal subset regression, Lasso regression and random forest).

[0074] Risk factors based on the optimal subset method: All possible predictor variables are included in the optimal subset regression analysis. The number of optimal combinations is selected according to the CP principle. When the CP value is minimized, the number of optimal combinations is 3. For example... Figure 11 The figures shown are age, degree of dental crowding, and ANB angle, respectively.

[0075] The CP principle / criterion is based on the CP statistic proposed by CLMallows as a standard for selecting the optimal subset. It is existing technology and will not be elaborated here.

[0076] Filtering the optimal subset feature number as follows Figure 11 As shown, the horizontal axis represents the number of predictor variable features, and the vertical axis represents the CP value. The red dot in the figure indicates the point where the CP value is the minimum, meaning that the optimal number of variable combinations is 3. Figure 11 Part B in the diagram represents the selection of the optimal subset of feature factors. The horizontal axis represents the features of the predictor variables, and the vertical axis represents the CP value. When the CP value is minimized, it corresponds to each of the three independent variables.

[0077] Risk factors based on Lasso regression: The most significant predictive markers on the training set were selected using the Lasso logistic regression algorithm as variables that effectively contribute to the prediction model. A total of 13 features were included in the Lasso regression model, and 6 non-zero coefficients, i.e., 6 possible predictive variables, were finally selected: age, Angle classification, dental crowding, ANB angle, Wits value, and whether dental floss is used. Figure 12 As shown.

[0078] Figure 12 Part A includes the Lasso coefficient distributions for 13 clinical features. A penalty coefficient profile was generated based on Log(λ), and 6 potential predictors (age, Angle classification, flossing, crowding, ANB angle, and Wits value) were selected from the 13 relevant predictors. Figure 12 Part B shows that the penalty coefficient λ is used in Lasso logistic regression with 10-fold cross-validation through the minimum standard, and the curve of binomial bias versus log(λ) is plotted. A black vertical line is drawn at the optimal penalty coefficient λ according to the minimum error (Lambda.min) criterion and the minimum one standard error (Lambda.1se) criterion.

[0079] Risk factors based on random forest: 500 decision trees were constructed, and 5 variables were randomly selected for each decision tree node. Variables were either selected from the random forest or excluded based on feature importance. Ultimately, age, ANB angle, Wits value, and APDI were the selected feature variables. Figure 13 As shown, the random forest model has an error of 0.12, indicating a small generalization error.

[0080] Figure 13 Part A shows box plots of all attributes plus minimum, average, and maximum shaded scores. Green box plots represent important variables identified by the random forest, while red box plots represent rejected variables. Figure 13Part B shows the decision history of the random forest in 60 runs of the Boruta function, showing whether it rejects or accepts the feature variables.

[0081] This invention also includes performance evaluations of three risk assessment models, specifically: The discrimination and discriminative ability of the models were evaluated using the Area Under the ROC Curve (AUC) to assess the discrimination of the three models. Figure 14 As shown in Table 6, among the three prediction models—optimal subset, Lasso regression, and random forest—the AUC of the optimal subset risk assessment model was 0.636 (95% CI: 0.550–0.715), the AUC of the Lasso regression prediction model was 0.728 (95% CI: 0.647–0.800), and the AUC of the random forest prediction model was 0.842 (95% CI: 0.771–0.898). The random forest model had the highest AUC among the three models (as shown in Table 6), and the difference between the random forest and the other two methods was statistically significant (r < 0.001).

[0082] Figure 14 The vertical axis represents the true positive rate; the horizontal axis represents the false positive rate. OSR represents the model based on the optimal subset method; Lasso represents the model based on the Lasoo regression method; and RF represents the model based on the random forest method.

[0083]

[0084] Model calibration evaluation: Three machine learning-based models were used to predict the prevalence of proximal caries in posterior teeth, and the results were compared with the actual prevalence in the test set. Calibration curves were then plotted. The U-test (unrealiability) was used to evaluate the difference in distribution between the predicted and actual values. Figure 15 As shown, among the three risk assessment models, Lasso regression (r=0.07) and random forest (r=0.11) calibrated well, with no significant difference between predicted and observed probabilities. The differences between the actual observed outcome variable values ​​and predicted probabilities for the three models were quantified using the mean (Eaver), maximum (Emax), and quartile (E90). The Eaver for the optimal subset method was 2.6%, Emax was 16.1%, and E90 was 5.1%; the Eaver for Lasso regression was 1.8%, Emax was 12.5%, and E90 was 3.3%; and the Eaver for random forest was 2.2%, Emax was 5.5%, and E90 was 4.1%.

[0085] like Figure 15As shown, the x-axis represents the probability of caries on the proximal surfaces of posterior teeth predicted by the model; the y-axis represents the actual caries rate on the proximal surfaces of posterior teeth; the gray diagonal line represents the ideal state predicted by the model; OSR represents the optimal subset; Lasso represents Lasso regression; and RF represents random forest.

[0086] Model fit evaluation: Goodness of fit (GOF) is an indicator of how well a model describes the data distribution. The Hosmer-Lemeshow goodness of fit test was applied to evaluate the three risk assessment models, and p-values ​​were calculated to assess whether there were differences between predicted and actual values. The p-value for the optimal subset GOF test was 0.1431 (r > 0.05), the p-value for the Lasso regression GOF test was 0.5908 (r > 0.05), and the p-value for the random forest GOF test was 0.5888 (r > 0.05), indicating that there was no statistically significant difference between the actual and predicted values ​​of the three risk assessment models.

[0087] DCA evaluation of the model: In the risk assessment tool of the predictive model, when the predicted probability of orthodontic patients developing proximal caries of posterior teeth reaches a certain threshold, the patient is determined to be a positive case and treatment is required. At this point, there will be benefits from treatment for true positive patients, as well as losses from treatment for false positive patients and no treatment for false negative patients. In the DCA evaluation, the magnitude of the net benefit (NB) can be used to measure whether the predictive model can bring benefits to the patient. Figure 16 The blue, red, and green curves shown represent the DCA models constructed using the optimal subset, Lasso regression, and random forest models, respectively. If the curve is above the horizontal black line and the left-sloping gray line, it indicates that the model can be beneficial. All models yield net benefits within the 20%-60% probability threshold. When the probability threshold is greater than 10%, the random forest-based posterior proximal caries risk assessment model (green curve) shows better net benefits than the other two models.

[0088] like Figure 16 As shown, the vertical axis represents the net benefit (NB), and the horizontal axis represents the probability threshold (Pt). The black line (None) represents the net benefit assuming no one is diagnosed with posterior proximal caries, and the gray line (All) represents the net benefit assuming everyone is diagnosed with posterior proximal caries.

[0089] In this embodiment, based on the above evaluation, the posterior proximal caries risk assessment model based on the random forest method was selected as the optimal risk assessment model. A nomogram was plotted based on the selected risk factors (age, ANB angle, Wits value, and APDI), as shown below. Figure 17For the nomogram: (1) The values ​​on the line segments represent the range of the above four predictive factors. The age is marked with a line segment scale of 12-36, the Wits value is marked with a line segment scale of -12-12, the ANB angle is marked with a line segment scale of -4-10, and the APDI is marked with a line segment scale of 60-95. (2) The direction of the change of the numbers on the line segments represents the relationship between the predictive factors and the probability of disease. Age, Wits value, and ANB angle are positively correlated with the incidence of posterior proximal caries, while APDI is negatively correlated. (3) Explanation of the nomogram prediction method: Based on the corresponding age, ANB angle, Wits value, and APDI data of each patient, find the corresponding values ​​in the four line segments of the nomogram, and obtain the score on the corresponding "score" line segment based on the values. The total score is obtained by adding the age, ANB angle, Wits value, and APDI score, and the corresponding position on the "Total Score" line segment. The data on the bottom line segment corresponding to the total score represents the probability of the patient having proximal caries on the posterior teeth.

[0090] All predictive factors selected from the above three models were used as potential risk factors. Poisson regression analysis was applied to analyze the effects of age, floss use, Angle classification, crowding, ANB angle, Wits value, and APDI on the severity of proximal caries in posterior teeth. The results are shown in Table 7. The results indicate that age, Angle classification, dental crowding, and APDI are risk factors, and the results are statistically significant.

[0091]

[0092] **r < 0.01, indicating a highly significant difference.

[0093] Three machine learning algorithms—optimal subset method, Lasso regression, and random forest—were used to screen potential risk factors for proximal caries. Age and ANB angle were identified as important influencing factors by all three models. In addition, optimal subset method identified dentition crowding as a potential influencing factor, Lasso regression identified Angle classification, dentition crowding, flossing, and Wits value as potential influencing factors, and random forest identified Wits value and APDI as potential influencing factors.

[0094] A comprehensive evaluation of the posterior proximal caries prediction models established by three machine learning algorithms was conducted. The risk assessment model based on random forest had the best discriminative power, and the selected factors, such as age, ANB angle, Wits value, and APDI, contributed to the discriminative ability of the prediction model.

[0095] Evaluation indicators of sagittal skeletal profile contribute to the incidence of proximal caries in posterior teeth.

[0096] Age, APDI, Angle classification, and dentition crowding are highly correlated with the severity of proximal caries in posterior teeth.

[0097] This invention uses gender, age, brushing frequency, brushing time, flossing, frequency of sweets consumption, crowding, Angle classification, ANB angle, FH-MP angle, Wits value, APDI, and ODI as research variables, and the presence or absence of proximal caries as the outcome variable. Three machine learning algorithms were used to jointly screen for age and ANB angle as predictive factors, suggesting a strong association between age and ANB angle and the occurrence of proximal caries. In addition to the variables jointly screened, Angle classification, Wits value, APDI, crowding, and flossing may also influence the occurrence of proximal caries to some extent. Since different machine learning algorithms have their own characteristics in screening variables, using a single screening method may miss risk factors. Therefore, this invention uses three algorithms to build a risk assessment model. After comprehensive evaluation using discrimination, calibration, goodness of fit, and DCA analysis, random forest was ultimately selected as the best risk assessment model. This indicates that age, ANB angle, Wits value, and APDI have high value in predicting the risk of proximal caries.

[0098] In establishing the risk assessment model for proximal caries of posterior teeth, this invention includes 13 possible risk factors. The reasons for selecting the above predictive variables are as follows: (1) Considering that the traditional etiology of caries is still mainly based on the four factors, this invention includes possible risk factors such as gender, age, brushing frequency and time, flossing frequency, and frequency of eating sweets. Among them, many articles have reported that the use of dental floss can significantly reduce the incidence of proximal caries. While factors such as gender, age, brushing frequency and time, and frequency of eating sweets have all been reported to be related to the occurrence and development of caries, few articles mention the correlation between the above risk factors and the incidence of proximal caries. (2) Previous studies have shown that there is a significant correlation between crowding and caries, suggesting that crowded teeth lead to malocclusion, which is not conducive to oral hygiene maintenance, thus making it easier for bacteria to accumulate. However, some studies have not confirmed this view. (3) Other studies have shown that abnormal molar relationships and jawbone facial features are related to the occurrence and severity of caries. Since most malocclusion problems occur in the sagittal direction, and each indicator has its own limitations, it is necessary to use these indicators comprehensively when evaluating the sagittal skeletal facial profile of patients. This invention selected Angle classification, ANB angle, Wits value and APDI as representative measurement indicators for sagittal skeletal facial profile. (4) Currently, no existing technology has disclosed that the vertical evaluation indicators of skeletal facial profile are correlated with the occurrence of caries. Considering that the mandibular plane angle and ODI are not only related to the vertical skeletal facial profile, there are also reports that patients with a larger mandibular plane angle often have mandibular retrusion, while those with a smaller mandibular plane angle have mandibular protrusion. Therefore, the FH-MP and ODI indicators were selected as predictive variables in a breakthrough manner.

[0099] Age is a crucial factor in the development and progression of dental caries. This invention, in designing its inclusion and exclusion criteria, references previous literature and excludes patients under 13 years of age who are not in the permanent dentition to avoid interference from age factors during the mixed dentition period. Patients over 35 years of age are excluded because they may have lost or had their teeth extracted due to periodontal disease or other factors, making it impossible to determine if the affected tooth has a history of caries. Even if the affected tooth may have had caries in the past, it is impossible to distinguish whether the caries originated from the ulnar or proximal surface.

[0100] The results of this invention show that the risk of proximal caries in posterior teeth increases with increasing ANB angle or Wits value, and decreases with increasing APDI. When the ANB angle, Wits value, or APDI value is large, the patient tends to have a Class II skeletal facial profile; conversely, the patient tends to have a Class III skeletal facial profile. The three sagittal indices evaluated in this invention all suggest that patients who are more inclined towards skeletal Class II may be more prone to proximal caries. This result further verifies the inference by Bernhardt et al. that there is a correlation between skeletal facial profile and caries.

[0101] The nomogram of this invention shows that among the three sagittal facial profile evaluation indices, the Wits value has the greatest impact on the prediction model. This may be because the Wits value avoids the influence of cranial structures on the relative positional relationship of the maxilla and mandible, and is more inclined to reflect the facial relationship. Some studies suggest that the Wits value is greatly affected by the functional facial plane, such as during the mixed dentition stage of growth and development, severe open bite of posterior teeth, and posterior tooth defects that make it impossible to determine the facial plane. Therefore, this invention fully considers the impact of these factors on the Wits value in the inclusion and exclusion criteria, and excludes patients who may interfere with the judgment of the facial plane. As a comprehensive judgment index, APDI is not only of great value in predicting the occurrence of posterior proximal caries, but also highly correlated with the severity of posterior proximal caries. APDI is composed of the facial angle, the alveolar angle of the upper and lower teeth, and the palatal plane angle, and has better stability, but some indices reflecting the vertical relationship of the maxillofacial region can affect APDI. In this invention, the ANB angle and Wits value, both sagittal skeletal facial features, did not show a correlation with the severity of proximal caries in posterior teeth. This may be because sagittal indicators have little impact on the severity of proximal caries, while other oral hygiene habits or other maxillofacial features have a greater influence on the progression of proximal caries. This suggests that further research should consider incorporating oral hygiene habits, vertical jawbone evaluation indicators, and other relevant factors when specifically analyzing the distribution, number, or progression of proximal caries.

[0102] The Angle classification of this invention also influences the occurrence and severity of proximal caries in posterior teeth to some extent. Patients with Angle Class II malocclusion may have more severe proximal caries than those with the other two classes of malocclusion. Literature reports that Angle Class II patients are more prone to temporomandibular joint disorder (TMJ) than Angle Class I and III patients, possibly because the condyles of Angle Class II patients are smaller than those of the other two classes, making it less likely for a stable condyle-glenoid structure to form, and their position is relatively more prone to change, affecting occlusal stability. The temporomandibular joint is closely related to the jaw, and the two interact and influence each other. Patients with TMJ disorder are more prone to jaw interference and premature tooth contact, leading to malocclusion. Naim et al., by measuring the occlusal contact area in orthodontic patient samples, found that the occlusal contact area differed among different Angle classifications, with Angle Class II patients having a smaller contact area than Angle Class I. Oltramari et al. studied tooth wear in patients with malocclusion and those with normal teeth. Their results showed that patients with Class II malocclusion experienced more severe wear in the posterior teeth, possibly due to their unique tooth contact patterns. When wear is severe, premature contact points, sharp, high-set cusps, and lowered marginal ridges can form on the tooth surface, increasing wedging forces during occlusion and making food impaction more likely, leading to proximal caries. Premature contact or unstable occlusal contact can also generate abnormal lateral forces on opposing teeth. Repeated abnormal compression and impact between teeth can cause proximal wear, resulting in proximal caries. To date, numerous studies have analyzed the correlation between dental crowding and caries, but the conclusions are inconsistent and further research is needed. Many scholars believe that patients with malocclusion are prone to caries because the crowding makes it difficult to clean the contact areas between adjacent teeth, leading to plaque and tartar accumulation. Most studies use caries loss and repair to assess dental caries, but this index may underestimate the incidence of interproximal caries. Furthermore, most included studies lack baseline oral health data, making it more difficult to assess the relationship between them. Given that dental caries is influenced by many risk factors, incorporating confounding factors into the quality assessment is worthwhile. This invention focuses more on the correlation analysis between crowding and interproximal caries compared to previous studies, and incorporates several potential risk factors while controlling for a specific age range, representing an improvement and refinement of previous research in its design.

[0103] Currently, internationally recognized research on the etiology of dental caries primarily focuses on bacteriological factors. Most literature reports on the correlation between malocclusion and dental caries only analyze the correlation between certain occlusal features of malocclusion and the presence or absence of caries, without controlling for other risk factors affecting caries development, nor distinguishing the location and severity of caries. This invention attempts to explore the role of occlusion and dentofacial deformities in the development of proximal caries, while controlling for certain factors influencing caries development, thus supplementing and refining traditional theories of dental caries etiology.

[0104] The present invention also has the following beneficial effects: This invention utilizes ANB angle data, Wits value data, APDI angle data, Angle classification data, and dental crowding data related to malocclusion. Based on a machine learning model, it can accurately predict the risk rate of posterior interproximal caries in the target subject, thereby improving the economy of preventing and treating posterior interproximal caries and increasing the likelihood of preventing the risk of posterior interproximal caries.

[0105] This invention explores whether occlusion and dentofacial deformities are risk factors for the development of proximal caries. Based on the occlusal characteristics and relative positional relationship of dentofacial deformities in different patients, it predicts whether patients are susceptible to proximal caries, so as to achieve early prevention, early diagnosis and early treatment of proximal caries.

[0106] This invention features the screening of potential risk factors for proximal caries, establishes a risk assessment model for posterior proximal caries, and alerts patients at high risk of developing caries in their posterior teeth. This increases the awareness of caries among high-risk patients, reducing the incidence of caries or allowing for intervention in the early stages of caries. When deep caries progresses to pulpitis, root canal treatment requires more visits and is more expensive. If patients receive high-risk warnings through this invention and are given regular checkups and early intervention, including early filling treatment, the number of visits can be reduced, significantly decreasing the time and cost of oral treatment and lowering overall healthcare costs.

[0107] This invention also provides a risk assessment system for the impact of malocclusion on the prevalence of proximal caries in posterior teeth, comprising: The data acquisition module is used to acquire feature data of the target object to be evaluated. The variable assignment module is used to assign variable values ​​to the feature data of the target object to be evaluated, thereby obtaining predictive variable data. The data filtering module is used to filter predictor variable data based on the risk assessment model to obtain risk factor data; The nomogram drawing module is used to draw nomograms based on the selected risk factor data; The risk assessment module is used to obtain the probability of posterior proximal caries in the target subject based on risk factor data and nomograms.

[0108] It should be understood that, since the various modules are provided merely to illustrate the functional units of the system disclosed herein, the physical devices corresponding to these modules may be the processor itself, or a part of the processor's software, hardware, or a combination of both. Therefore, the number of modules is merely illustrative.

[0109] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0110] To address the aforementioned technical problems, embodiments of the present invention also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the risk assessment method for the impact of malocclusion on proximal caries of posterior teeth as described above.

[0111] The computer / electronic device of the present invention includes a memory, a processor, and a network interface that are interconnected via a system bus. It should be noted that the above description only illustrates a computer device with components such as memory, processor, network interface, and operating system; however, it should be understood that it is not required to implement all the illustrated components, and more or fewer components may be implemented alternatively. Those skilled in the art will understand that the computer / electronic device described herein is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0112] Computers / electronic devices can be desktop computers, laptops, PDAs, and cloud servers, among other computing devices. They can interact with users through keyboards, mice, remote controls, touchpads, or voice-activated devices.

[0113] There may be one or more memories, and at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the memory may be an internal storage unit of a computer device, such as the hard disk or RAM of the computer device. In other embodiments, the memory may also be an external storage device of the computer device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. Of course, the memory may include both internal storage units and external storage devices of the computer device. In this embodiment, the memory is typically used to store the operating system and various application software installed on the computer device, such as program code for a risk assessment method for the impact of malocclusion on the lesions of proximal caries of posterior teeth. In addition, memory can also be used to temporarily store various types of data that have been output or will be output.

[0114] In some embodiments, the processor may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor is typically used to control the overall operation of a computer device. In this embodiment, the processor is used to run program code stored in memory or process data, such as running program code for a risk assessment method of the impact of malocclusion on the lesions of proximal caries in posterior teeth.

[0115] Network interfaces may include wireless network interfaces and / or wired network interfaces, which are typically used to establish communication connections between computer devices and other electronic devices.

[0116] The present invention also provides another embodiment, namely, providing a readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the risk assessment method for the impact of malocclusion on the morbidity of proximal caries of posterior teeth as described above.

Claims

1. A risk assessment method for the impact of malocclusion on proximal caries of posterior teeth, characterized in that, include: Obtain feature data of the target object to be evaluated; The characteristic data of the target object to be evaluated are assigned variable values ​​to obtain predictive variable data; Risk factor data is obtained by filtering the predictor variable data based on the risk assessment model; Draw a nomogram based on the selected risk factor data; The probability of posterior proximal caries in the target subject was obtained based on risk factor data and nomograms.

2. The risk assessment method for the impact of malocclusion on proximal caries of posterior teeth according to claim 1, characterized in that, The feature data includes: age data, gender data, brushing frequency data, brushing time data, frequency of eating sweets data, dental floss usage data, Angle classification data, dental crowding data, ANB angle data, FH-MP angle data, Wits value data, APDI angle data, and ODI angle data.

3. The risk assessment method for the impact of malocclusion on proximal caries of posterior teeth according to claim 1, characterized in that, The risk assessment model is a machine learning model, including the optimal subset regression algorithm, Lasso regression algorithm, and random forest algorithm.

4. The risk assessment method for the impact of malocclusion on proximal caries of posterior teeth according to claim 3, characterized in that, The process of filtering the predictor variable data based on the risk assessment model to obtain risk factor data includes: filtering the predictor variable data based on the optimal subset regression algorithm, including all predictor variable data in the optimal subset regression analysis, filtering the number of optimal combinations according to the CP principle, determining the number of optimal combinations when the CP value is the minimum, and obtaining risk factor data based on the optimal subset regression algorithm.

5. The risk assessment method for the impact of malocclusion on proximal caries of posterior teeth according to claim 3, characterized in that, The process of filtering the predictor variable data based on the risk assessment model to obtain risk factor data includes: filtering the predictor variable data based on the Lasso regression algorithm, and selecting the most significant predictive markers on the training set using the Lasso logistic regression algorithm as risk factor data based on the Lasso regression algorithm.

6. The risk assessment method for the impact of malocclusion on proximal caries of posterior teeth according to claim 3, characterized in that, The process of filtering the predictor variable data based on the risk assessment model to obtain risk factor data includes: filtering the predictor variable data based on the random forest algorithm, evaluating all predictor variable data using the random forest model, calculating the importance score of each predictor variable data, excluding predictor variable data with lower importance according to a set threshold or ranking, and retaining the predictor variable data as risk factor data based on the random forest algorithm.

7. The risk assessment method for the impact of malocclusion on proximal caries of posterior teeth according to claim 4, characterized in that, Risk factor data based on the optimal subset regression algorithm includes age data, dental crowding data, and ANB angle data. A nomogram is plotted based on the risk factor data. The probability of posterior proximal caries in the target subject is obtained based on the first total score corresponding to the age data, dental crowding data, and ANB angle data. The first total score is the sum of the scores in the nomogram corresponding to the age data, the dental crowding data, and the ANB angle data. The age data, ANB angle data, and dental crowding data are positively correlated with the probability of posterior proximal caries.

8. The risk assessment method for the impact of malocclusion on proximal caries of posterior teeth according to claim 5, characterized in that, Risk factor data based on the Lasso regression algorithm includes age data, Angle classification data, dental crowding data, ANB angle data, Wits value data, and dental floss usage data. A nomogram is plotted based on this risk factor data. The probability of proximal caries in the target subject is obtained based on the second total score corresponding to the age data, Angle classification data, dental crowding data, ANB angle data, Wits value data, and dental floss usage data. The second total score is the nomogram value corresponding to the age data. The scores in the line graph, the scores in the column graph corresponding to the Angle classification data, the scores in the column graph corresponding to the dental crowding data, the scores in the column graph corresponding to the ANB angle data, the scores in the column graph corresponding to the Wits value data, and the scores in the column graph corresponding to the dental floss use data are added together. Among these, the age data, the Angle classification data, the dental crowding data, the ANB angle data, the Wits value data, and the dental floss use data are positively correlated with the probability of posterior proximal caries.

9. The risk assessment method for the impact of malocclusion on proximal caries of posterior teeth according to claim 6, characterized in that, Risk factor data based on the random forest algorithm includes age data, ANB angle data, Wits value data, and APDI angle data. A nomogram is plotted based on the risk factor data. The probability of posterior proximal caries in the target subject is obtained based on the third total score corresponding to the age data, ANB angle data, Wits value data, and APDI angle data. The third total score is the sum of the scores in the nomogram corresponding to the age data, the ANB angle data, the Wits value data, and the APDI angle data. The age data, ANB angle data, and Wits value data are positively correlated with the probability of posterior proximal caries, while the APDI angle data is negatively correlated with the probability of posterior proximal caries.

10. A risk assessment system for the impact of malocclusion on proximal caries of posterior teeth, characterized in that, include: The data acquisition module is used to acquire feature data of the target object to be evaluated. The variable assignment module is used to assign variable values ​​to the feature data of the target object to be evaluated, thereby obtaining predictive variable data. The data filtering module is used to filter the predictor variable data based on the risk assessment model to obtain risk factor data; The nomogram drawing module is used to draw nomograms based on the selected risk factor data; The risk assessment module is used to obtain the probability of posterior proximal caries in the target subject based on risk factor data and nomograms.