Maxillary sinus lifting postoperative graft volume change prediction model construction method based on machine learning algorithm, maxillary sinus lifting postoperative graft volume change prediction method and device, and storage medium

A model for predicting graft volume changes after maxillary sinus lift surgery was constructed using machine learning algorithms. By using the chi-square test and LASSO regression to screen features and combining them with the Gaussian Naive Bayes algorithm, the problem of predicting volume changes after maxillary sinus lift surgery was solved, improving prediction accuracy and the long-term success rate of implants.

CN121768684APending Publication Date: 2026-03-31THE AFFILIATED STOMATOLOGICAL HOSPITAL OF KUNMING MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-15
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately predict factors influencing graft volume changes after maxillary sinus lift surgery, leading to uncertainty in osteogenesis outcomes and individual differences, which in turn affects the long-term success rate of implants.

Method used

A machine learning algorithm was used to construct a predictive model for graft volume changes after maxillary sinus lift surgery. By acquiring patients' clinical data and CBCT radiomics data, the chi-square test and LASSO regression were used to screen feature variables. The predictive model was constructed by combining the Gaussian Naive Bayes algorithm, and trees were added iteratively to fit the residuals to achieve the accuracy of the predictive model.

Benefits of technology

It improved the accuracy of predicting graft volume changes after maxillary sinus lift surgery, reduced the risk of insufficient osteogenesis, and improved the long-term success rate of implants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121768684A_ABST
    Figure CN121768684A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer data processing, and discloses a construction method of a maxillary sinus lifting postoperative graft volume change prediction model based on a machine learning algorithm, a maxillary sinus lifting postoperative graft volume change prediction method and device, and a storage medium. The method comprises the following steps: acquiring clinical data and CBCT image omics data of a patient to form a data set, and performing data preprocessing on the data set; screening characteristic variables in the data set by applying chi-square test and LASSO regression methods, and constructing a training set, a verification set and a test set based on the screened characteristic variables; constructing an initial prediction model based on a Gaussian naive Bayes algorithm; training the initial prediction model based on the training set, the verification set and the test set, iteratively adding trees, adding one tree each time to fit the residual error of the last round of prediction, and finally obtaining the maxillary sinus lifting postoperative graft volume change prediction model with the root mean square error meeting the preset requirement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer data processing technology, specifically to a method for constructing a model for predicting graft volume changes after maxillary sinus lift surgery based on machine learning algorithms, a method and device for predicting graft volume changes after maxillary sinus lift surgery, and a storage medium. Background Technology

[0002] Maxillary sinus lift is one of the main techniques for addressing insufficient vertical bone volume in patients with missing teeth in the maxillary posterior region undergoing implant restoration. In patients who undergo maxillary sinus lift, the volume of most bone grafts will shrink to some extent six months post-surgery. The long-term success of implants depends primarily on the long-term, stable osseointegration between the implant and the surrounding alveolar bone and bone graft material, ensuring sufficient bone coverage around the implant. Therefore, understanding the complex factors influencing postoperative volume changes in bone graft material, grasping the key factors affecting these changes, and implementing targeted prevention and control measures can improve the long-term success rate of implants. Currently, the factors influencing postoperative volume changes in bone graft material are extremely complex, and the osteogenesis effect has significant uncertainty and individual differences, leading to insufficient osteogenesis being one of the common complications after maxillary sinus lift. The bone gain effect after most maxillary sinus lift surgeries is difficult to predict, which poses a challenge to the clinical decision-making process. Clinicians urgently need an accurate and effective tool or method to determine the influencing factors affecting graft volume changes after maxillary sinus lift and to predict postoperative outcomes, thus assisting clinical work. Summary of the Invention

[0003] The purpose of this application is to provide a method for constructing a model for predicting graft volume changes after maxillary sinus lift surgery based on machine learning algorithms, a method and device for predicting graft volume changes after maxillary sinus lift surgery, and a storage medium, so as to solve the problem in the prior art of needing to determine the influencing factors of graft volume changes after maxillary sinus lift surgery and predict graft volume changes after maxillary sinus lift surgery.

[0004] To achieve the above objectives, this application provides a method for constructing a predictive model for graft volume changes after maxillary sinus lift surgery based on machine learning algorithms, comprising the following steps:

[0005] Acquire patients' clinical data and CBCT radiomics data and calculate radiomics scores to form a dataset, and perform data preprocessing on the dataset;

[0006] The chi-square test and LASSO regression method were used to screen the feature variables in the dataset, and the training set, validation set and test set were constructed based on the screened feature variables.

[0007] An initial prediction model was constructed based on the Gaussian Naive Bayes algorithm;

[0008] The initial prediction model is trained based on the training set, validation set, and test set. By iteratively adding trees, one tree is added each time to fit the residual of the previous prediction, and finally a prediction model for graft volume change after maxillary sinus lift surgery with root mean square error meeting the preset requirements is obtained.

[0009] Optionally, the acquisition of the patient's clinical data specifically includes: acquiring clinical influencing factors data on changes in graft volume after maxillary sinus lift surgery, including: gender, age, smoking, history of periodontitis, history of hypertension, missing tooth position in the surgical area, history of diabetes, immediate graft volume after surgery, size of bone graft material, use of blood growth factors, sinus cavity size, maxillary sinus contour, height of remaining alveolar ridge, width of remaining alveolar ridge, thickness of sinus lateral wall, buccal-palatal bone wall angle (MSA), fenestration size, distance from the lower boundary of the fenestration to the sinus floor, surgical lift height, and thickness of maxillary sinus mucosa;

[0010] The acquisition of the patient's CBCT radiomics data specifically includes: acquiring the DICOM format file exported from CBCT and inputting it into 3D Slicer software to manually outline the region of interest layer by layer, generating the region of interest volume (VOI), and performing radiomics feature extraction. A total of 107 features in 7 categories are extracted, including 14 shape-based features, 18 first-order statistical features, 24 gray-level co-occurrence matrix features, 16 gray-level run length matrix features, 16 gray-level size region matrix features, 14 gray-level dependence matrix features, and 5 adjacent gray-level difference matrix features.

[0011] Optionally, the data preprocessing of the dataset specifically includes: removing outliers and missing values ​​from the dataset, and filling in missing data using interpolation.

[0012] Normalize or standardize the numerical variables in the dataset.

[0013] Optionally, the feature variables in the dataset being filtered specifically include:

[0014] The categorical variables in the dataset are expressed as numbers and percentages, and the chi-square test is used to filter them out, selecting categorical variables with a two-sided p-value less than 0.05.

[0015] LASSO regression was used to further screen feature variables, specifically by: compressing the coefficients of variables in the regression model by generating a penalty function to prevent overfitting; and setting the coefficients of related independent variables to 0 based on the correlation between independent variables. The LASSO regression was configured with 10-fold cross-validation and executed using the R package.

[0016] Optionally, the selected clinical characteristics variables include: smoking, history of periodontitis, history of diabetes, size of bone graft material, sinus cavity size (SW), and fenestration size.

[0017] Optionally, the selected radiomics features include: Kurtosis, Dependence (Non-Uniformity), Large Dependence (High Gray Level Emphasis);

[0018] The computational radiomics score is calculated as: Radiomics score = 0.253 + (-0.674) × Kurtosis + 2.357 × DependenceNonUniformity + (-1.515) × LargeDependenceHighGrayLevelEmphasis.

[0019] Optionally, the parameters of the prediction model for graft volume change after maxillary sinus lift surgery include: a learning rate of 0.1, a maximum tree depth of 6, a minimum bifurcation weight sum of 1, and an L2 regularization coefficient of 0.1.

[0020] To achieve the above objectives, this application also provides a method for predicting graft volume changes after maxillary sinus lift surgery, comprising the following steps:

[0021] The relevant feature values ​​of the patient are obtained and input into the prediction model of the change in graft volume after maxillary sinus lift surgery constructed by the method of constructing the prediction model of graft volume change after maxillary sinus lift surgery based on machine learning algorithm as described in any of the preceding claims. The feature values ​​include at least one of the following: smoking, history of periodontitis, history of diabetes, size of bone graft material, sinus cavity size SW, fenestration size, and radiomics score.

[0022] The prediction model for graft volume change after maxillary sinus lift surgery outputs prediction results, giving the probability of graft volume change after maxillary sinus lift surgery.

[0023] To achieve the above objectives, this application also provides a device for constructing a predictive model of graft volume changes after maxillary sinus lift surgery based on machine learning algorithms, comprising: a memory; and

[0024] A processor connected to the memory, the processor being configured to perform the steps of the method described above.

[0025] To achieve the above objectives, this application also provides a computer storage medium having a computer program stored thereon, wherein the computer program, when executed by a machine, implements the steps of the method described above.

[0026] The embodiments of this application have the following advantages:

[0027] Using the methods described above, based on large-scale patient data and CBCT radiomics data, chi-square test and LASSO regression were employed to screen for influencing factors of graft volume changes after maxillary sinus lift surgery. A highly efficient predictive model for graft volume changes after maxillary sinus lift surgery was then constructed using machine learning algorithms. The Gaussian Naive Bayes algorithm demonstrated excellent performance, with area under the curve (AUC) values ​​of 0.939, 0.928, and 0.958 on the training, validation, and test sets, respectively. The predictive model for graft volume changes after maxillary sinus lift surgery, based on the GNB algorithm, can effectively predict the probability of graft volume changes in patients undergoing maxillary sinus lift surgery 6 months post-surgery. This solves the problem in existing technologies that require identifying influencing factors and predicting graft volume changes after maxillary sinus lift surgery. Attached Figure Description

[0028] To more clearly illustrate the embodiments of this application or the technical solutions in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary, and those skilled in the art can derive other embodiments based on the provided drawings without creative effort.

[0029] Figure 1 A flowchart illustrating a method for constructing a prediction model for graft volume changes after maxillary sinus lift surgery based on a machine learning algorithm, provided for at least one embodiment of this application;

[0030] Figure 2a A schematic diagram of the LASSO coefficient path of the clinical feature variable in a method for constructing a prediction model of graft volume change after maxillary sinus lift surgery based on a machine learning algorithm, provided in at least one embodiment of this application;

[0031] Figure 2b This is a schematic diagram of the LASSO coefficient path of the feature variable in the radiomics of a method for constructing a prediction model of graft volume change after maxillary sinus lift surgery based on a machine learning algorithm, provided in at least one embodiment of this application.

[0032] Figure 3a A schematic diagram of the LASSO regression analysis cross-validation curve of the clinical characteristic variables of a method for constructing a prediction model of graft volume change after maxillary sinus lift surgery based on machine learning algorithm provided in at least one embodiment of this application;

[0033] Figure 3bA schematic diagram of the LASSO regression cross-validation curve of the feature variables of the method for constructing a prediction model of graft volume change after maxillary sinus lift surgery based on machine learning algorithm provided in at least one embodiment of this application;

[0034] Figure 4 A schematic diagram of the training set ROC curves comparing multiple machine learning models for constructing a method for predicting graft volume changes after maxillary sinus lift surgery based on machine learning algorithms, provided in at least one embodiment of this application.

[0035] Figure 5 A schematic diagram of the validation set ROC curves for a method of constructing a prediction model for graft volume change after maxillary sinus lift surgery based on a machine learning algorithm, provided in at least one embodiment of this application.

[0036] Figure 6 A schematic diagram of the ROC curve of the training set of the GNB algorithm model for constructing a prediction model of graft volume change after maxillary sinus lift surgery based on machine learning algorithm, provided for at least one embodiment of this application;

[0037] Figure 7 A schematic diagram of the ROC curve of the validation set of the GNB algorithm model for a method of constructing a prediction model of graft volume change after maxillary sinus lift surgery based on machine learning algorithm provided in at least one embodiment of this application;

[0038] Figure 8 A SHAP schematic diagram of a method for predicting graft volume changes after maxillary sinus lift surgery, provided in at least one embodiment of this application;

[0039] Figure 9 SHAP feature importance diagram for a method of predicting graft volume changes after maxillary sinus lift surgery provided in at least one embodiment of this application;

[0040] Figure 10 SHAP force map of a sample of a method for predicting graft volume changes after sinus lift surgery provided in at least one embodiment of this application;

[0041] Figure 11 SHAP force map of another sample of a method for predicting graft volume changes after sinus lift surgery provided in at least one embodiment of this application;

[0042] Figure 12 A block diagram of a device for constructing a prediction model of graft volume change after maxillary sinus lift surgery based on a machine learning algorithm, provided for at least one embodiment of this application. Detailed Implementation

[0043] The following specific embodiments illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0044] It should be noted that the steps in the claims and description of this application may be performed substantially in parallel or in reverse order under appropriate circumstances, depending on the function involved.

[0045] Furthermore, the technical features involved in the different embodiments of this application described below can be combined with each other as long as they do not conflict with each other.

[0046] One embodiment of this application provides a method for constructing a predictive model for graft volume changes after maxillary sinus lift surgery based on machine learning algorithms, referring to... Figure 1 , Figure 1 The flowchart illustrates a method for constructing a prediction model of graft volume change after maxillary sinus lift surgery based on a machine learning algorithm, provided in at least one embodiment of this application. It should be understood that the method may also include additional boxes not shown and / or the boxes shown may be omitted, and the scope of this application is not limited in this respect.

[0047] In step 101, the patient's clinical data and CBCT radiomics data are acquired and the radiomics score is calculated to form a dataset, and the dataset is preprocessed.

[0048] In some embodiments, acquiring the patient's clinical data specifically includes:

[0049] Data on factors influencing graft volume changes after maxillary sinus lift surgery were obtained, including: gender, age, smoking, history of periodontitis, history of hypertension, missing tooth position in the surgical area, history of diabetes, immediate graft volume after surgery, bone graft material size, blood growth factor usage, sinus width (SW), maxillary sinus contour, residual bone height (RBH), residual bone width (RBW), lateral wall thickness (LWT), buccal-palatal bone wall angle (MSA), fenestration size, distance from the lower boundary of the fenestration to the sinus floor, surgical lift height, and maxillary sinus mucosal thickness.

[0050] The acquisition of the patient's CBCT radiomics data specifically includes: acquiring the DICOM format file exported from CBCT and inputting it into 3D Slicer software (version 5.4.0) to manually delineate the region of interest (ROI) contour layer by layer, and generating the volume of interest (VOI). PyRadiomics 3.0.1 is used to extract radiomics features, extracting 107 features in 7 categories, including 14 shape-based features, 18 first-order statistical features, 24 gray-level co-occurrence matrix features, 16 gray-level run length matrix features, 16 gray-level size region matrix features, 14 gray-level dependency matrix features, and 5 adjacent gray-level difference matrix features.

[0051] In some embodiments, the data preprocessing of the dataset specifically includes:

[0052] Data cleaning: Remove outliers and missing values ​​from the dataset, and fill in the missing data using interpolation.

[0053] Standardization: Normalize or standardize the numerical variables in the dataset to ensure that different features have the same scale.

[0054] In step 102, the chi-square test and LASSO regression method are used to screen the feature variables in the dataset, and the training set, validation set and test set are constructed based on the screened feature variables.

[0055] In some embodiments, the filtering of feature variables in the dataset specifically includes:

[0056] Smoking history, periodontitis history, diabetes history, bone graft material size, sinus cavity size (SW), fenestration size.

[0057] Categorical variables in the dataset are expressed as n (%), count data as [n (%)], and continuous data as (x±s). Normality was tested using the Kolmogorov-Smirnov test, continuous variables were compared using the Mann-Whitney U test, and categorical variables were compared using the Chisquared test or Fisher's exact test. Intergroup comparisons were performed using the independent samples t-test or one-way ANOVA. A two-sided p-value < 0.05 was considered statistically significant. The analysis results are shown in Table 1.

[0058] Table 1:

[0059]

[0060]

[0061]

[0062]

[0063]

[0064] refer to Figure 2a , Figure 2b and Figure 3a , Figure 3b LASSO regression was used to further screen feature variables. Specifically, LASSO regression compresses the coefficients of variables in the regression model by generating a penalty function to prevent overfitting; simultaneously, LASSO regression reduces the impact of multicollinearity on the regression results by reducing the coefficients of correlated independent variables to zero through the correlation between independent variables, thus solving the problem of severe multicollinearity. LASSO regression was configured with 10-fold cross-validation and executed using the R package "glmnet4.1.2".

[0065] In Figure 2, the upper horizontal axis represents the number of non-zero coefficients in the model, the vertical axis represents the value of the coefficients, and the lower horizontal axis represents the standardized coefficient vector. Figure 2a The six lines of different colors represent six variables. Figure 2b The three lines of different colors in Figure 3 represent three variables, and each curve represents the trajectory of the coefficient of each independent variable. In Figure 3, the vertical axis represents the error of cross-validation (the smaller the vertical axis, the better the LASSO fit). The upper horizontal axis represents the number of variables corresponding to different eigenvalues ​​λ, and the lower horizontal axis represents the logarithm of the λ penalty coefficient. The parameter (lambda.min) corresponding to the dashed line on the left side of the graph has the minimum error, which determines how many variables can be used for analysis. Figure 3a The corresponding horizontal axis value is 6, meaning there are 6 variables remaining. Figure 3b The corresponding horizontal axis value is 3, meaning there are 3 variables remaining.

[0066] In some embodiments, the final selected clinically relevant variables include: smoking, history of periodontitis, history of diabetes, size of bone graft material, sinus cavity size (SW), and fenestration size. The final selected radiomics variables include: kurtosis, Dependence (Non-Uniformity), and Large Dependence (High Gray Level Emphasis).

[0067] Based on the different weight ratios of each radiomics feature variable in the LASSO regression feature variable selection coefficients, as shown in Table 2, the radiomics score is calculated as the sum of the constant term and each term multiplied by its weight coefficient: Radiomics score = 0.253 + (-0.674) × Kurtosis + 2.357 × Dependence Non-Uniformity + (-1.515) × Large Dependence High Gray Level Emphasis.

[0068] Table 2. Screening coefficients for LASSO regression radiomics characteristic variables.

[0069]

[0070] In step 103, an initial prediction model is constructed based on the Gaussian Naive Bayes algorithm.

[0071] In step 104, the initial prediction model is trained based on the training set, validation set, and test set. By iteratively adding trees, one tree is added each time to fit the residual of the previous prediction, and finally a prediction model for graft volume change after maxillary sinus lift surgery with root mean square error meeting the preset requirements is obtained.

[0072] Specifically, an initial prediction model was constructed based on various machine learning methods, including Logistic Regression (LR), eXtreme Gradient Boosting (XGBoost), Light Gradient Boosting Machine (LightGBM), Random Forest (RF), Adaptive Boosting (AdaBoost), Gaussian Naive Bayes (GNB), Complement Naive Bayes (CNB), Multilayer Perceptron (MLP), and K-Nearest Neighbors (KNN). The Logistic Regression model parameters were: regularization factor: 1.0, number of iterations: 100, regularization type: l2, convergence metric: 0.0001. The LightGBM model parameters were: algorithm type: gbdt, learning rate: 0.1, maximum tree depth: -1, maximum number of trees: 100, maximum number of leaves: 31. The Random Forest model parameters are: metric: Gini, maximum tree depth: None, minimum branch purity reward: 0.0, number of trees: 100. The AdaBoost model parameters are: learning rate: 1.0, number of individual models: 50. The KNN model parameters are: number of nearest neighbors: 5, weight type: uniform. The SVM model parameters are: regularization factor: 1.0, kernel type: rbf, convergence metric: 0.001. The GNB model parameters are: prior probability: None, var_smoothing: 1e-09. The MLP model parameters are: nonlinear function: ReLU, hidden layer width: (20, 10), number of iterations: 200. The optimal model was determined by evaluating the AUC values ​​on the training, validation, and test sets. GNB is the optimal model. (Reference) Figure 4 and Figure 5 , Figure 4 and Figure 5 The figures show the training set ROC curves and validation set ROC curves for comparing various machine learning models. The ROC curve is obtained by plotting the true positive rate and false positive rate, and can be used to reflect the relationship between sensitivity and specificity. The AUC value represents the area under the ROC curve and the coordinate axis. The larger the AUC value, the better the performance of the machine learning model, that is, the higher the discrimination. Figure 6 and Figure 7 These are schematic diagrams of the ROC curves for the GNB algorithm model training set and the GNB algorithm model validation set, respectively.

[0073] Specifically, a Python function is used to train the dataset using the GNB algorithm model. This function follows the GNB objective function and includes detailed comments explaining the implementation of each step.

[0074] Specific functions:

[0075] import GNB as xgb

[0076] import numpy as np

[0077] from sklearn.datasets import load_boston

[0078] from sklearn.model_selection import train_test_split

[0079] from sklearn.metrics import mean_squared_error

[0080] def train_GNB_model(X, y, num_boost_rounds=100):

[0081] """

[0082] The dataset was trained using the GNB model.

[0083] parameter:

[0084] X (numpy.ndarray): Feature dataset.

[0085] y (numpy.ndarray): Labeled dataset.

[0086] num_boost_rounds (int): Number of iterations, i.e., the number of trees to be trained.

[0087] return:

[0088] model (xgb.Booster): The trained GNB model.

[0089] dtrain(xgb.DMatrix): The DMatrix object used for training.

[0090] """

[0091] # Split the data into training and validation sets.

[0092] X_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.2, random_state=42)

[0093] # Convert the data into a DMatrix object, which is GNB's internal data format.

[0094] dtrain = xgb.DMatrix(X_train, label=y_train)

[0095] dval = xgb.DMatrix(X_val, label=y_val)

[0096] # Set model parameters

[0097] params = {'objective': 'reg:squarederror', # Use squared error as the objective function}

[0098] 'max_depth': 6, # Maximum depth of the tree

[0099] 'eta': 0.3, # Learning rate

[0100] 'subsample': 0.8, # The proportion of data used in each iteration

[0101] 'colsample_bytree': 0.8, # The proportion of features used per tree

[0102] 'eval_metric': 'rmse' # Use root mean square error as the evaluation metric}

[0103] # Train the model; evals are used to monitor the error on the validation set.

[0104] evals = [(dtrain, 'train'), (dval, 'eval')]

[0105] model = GNB.train(params, dtrain, num_boost_round=num_boost_rounds, evals=evals, early_stopping_rounds=10)

[0106] # Return the model and training data

[0107] return model, dtrain

[0108] Code explanation:

[0109] Dataset splitting and transformation:

[0110] Use train_test_split to split the data into a training set and a validation set.

[0111] Use GNB.DMatrix to convert data into the GNB built-in DMatrix format to improve computational efficiency.

[0112] Parameter settings:

[0113] objective: Sets the objective function, here we use reg:squarederror (squared error).

[0114] max_depth: The maximum depth of the tree, used to control overfitting.

[0115] eta: Learning rate. A lower learning rate will make the model converge more slowly, but usually the results are better.

[0116] `subsample` and `colsample_bytree`: Control the proportion of data and features used in each tree, helping to prevent overfitting.

[0117] eval_metric: Evaluation metric, here we use RMSE (root mean square error).

[0118] Model training:

[0119] Use GNB.train to train the model and monitor the error changes on the training and validation sets using the eval parameter.

[0120] early_stopping_rounds is used for early stopping strategies. If the error on the validation set does not decrease within the set number of rounds, training will stop early.

[0121] Prediction and Assessment:

[0122] After the model training is completed, predictions are made on the training set, and the RMSE is calculated.

[0123] RMSE (Root Mean Square Error) is a common metric used in machine learning to evaluate the prediction error of a model. It measures the difference between the predicted value and the true value, and imposes a higher penalty on predictions with larger discrepancies. The steps for calculating RMSE are as follows:

[0124] Error calculation: For each predicted value, calculate the difference between it and the corresponding true value: (e_i =\hat{y}_i - y_i), where (\hat{y}_i) is the (i)th predicted value and (y_i) is the (i)th true value.

[0125] Squared Error: The square of each error is calculated to eliminate the cancellation of positive and negative errors: (e_i^2).

[0126] Calculate the mean squared error (MSE): Sum all the squared errors and then divide by the total number of data points (n) to obtain the mean squared error.

[0127] [ MSE = \frac{1}{n} \sum_{i=1}^{n} e_i^2 ]

[0128] Square Root Error (RMSE): Finally, perform the square root operation on the MSE:

[0129] [ RMSE = \sqrt{MSE} ]

[0130] The formula can be summarized as: [ RMSE = \sqrt{ \frac{1}{n} \sum_{i=1}^{n} (\hat{y}_i - y_i)^2} ],

[0131] The smaller the RMSE, the closer the model's predictions are to the actual values, and the better the model's performance. It's important to note that the units of RMSE are the same as those of the target variable, thus providing a degree of intuitiveness.

[0132] Specifically, the initial prediction model, constructed using the GNB algorithm model, is trained on the training set, validation set, and test set. The initial prediction model is obtained by iteratively adding trees, adding one tree at a time to fit the residual of the previous prediction, and finally obtaining the maxillary sinus cyst prediction model.

[0133] The specific process follows the objective function Obj as: [\text{Obj}(\theta) = \sum_{i=1}^nl(y_i, \hat{y}i^{(t)}) + \sum{k=1}^T \Omega(f_k)]. Here, (l) is the loss function, used to measure the deviation between the predicted value (\hat{y}_i^{(t)}) and the actual value (y_i); while (\Omega(f_k)) is the regularization term, used to control the complexity of the model, in the form: [\Omega(f_k) = \gamma T + \frac{1}{2} \lambda \sum_{j=1}^T w_j^2]. Here, (\gamma) and (\lambda) are regularization parameters, (T) is the number of trees, and (w_j) is the node weight of the tree.

[0134] In some embodiments, the maxillary sinus cyst prediction model is expressed as the following formula: [ \hat{y} = \sum_{k=1}^{K} f_k(\mathbf{x}) ],

[0135] Where (\hat{y} ) is the predicted value, (K ) is the number of trees, and ( f_k(\mathbf{x}) ) is the output of the (k)th tree on the input feature (\mathbf{x} ).

[0136] In some embodiments, the parameters of the prediction model for graft volume change after maxillary sinus lift surgery include: a learning rate of 0.1, a maximum tree depth of 6, a minimum bifurcation weight sum of 1, and an L2 regularization coefficient of 0.1.

[0137] An embodiment of this application also provides a method for predicting graft volume changes after maxillary sinus lift surgery, including the steps of: obtaining relevant feature values ​​of the patient and inputting them into the maxillary sinus lift surgery graft volume change prediction model constructed by the aforementioned method embodiment of the machine learning algorithm-based method for predicting graft volume changes after maxillary sinus lift surgery. The feature values ​​include at least one of the following: smoking, history of periodontitis, history of diabetes, Bio-oss bone powder particle size, sinus cavity size SW, fenestration size, and Radiomics score.

[0138] The prediction model for graft volume change after maxillary sinus lift surgery outputs the prediction results, giving the probability of graft volume change after maxillary sinus lift surgery.

[0139] In some embodiments, the method further includes: interpreting the SHAP value of the post-maxillary sinus lift graft volume change prediction model, demonstrating important features and their influence on post-maxillary sinus lift graft volume change, and calculating the SHAP value (\phi_j) for each feature value (x_j) in the post-maxillary sinus lift graft volume change prediction model to reflect the average contribution of different feature values ​​to the prediction results. (See reference) Figure 8 The SHAP (Shapley Profile) plot illustrates the non-linear relationship between each influencing factor and graft volume changes after maxillary sinus lift surgery. Window size and Radiomics score are positively correlated with graft volume changes after maxillary sinus lift surgery; that is, larger window size and Radiomics score correlate with greater graft volume changes. The SHAP plot combines feature importance with feature effect. Each point on the SHAP plot represents a feature and an instance's Shapley value. The position on the y-axis is determined by the feature, and the position on the x-axis is determined by the Shapley value. Blue to red represents feature values ​​from low to high. Overlapping points jitter along the y-axis, thus revealing the distribution of Shapley values ​​for each feature. These features are ordered according to their importance. Reference Figure 9 Feature importance ranking showed that window size, radiomics score, smoking, history of periodontitis, Bio-oss bone powder particle size, sinus cavity size, and history of diabetes were the most important features in the predictive model for graft volume changes after maxillary sinus lift surgery. Window size was the most important feature, changing the absolute probability of predicted graft volume changes by an average of approximately 23 percentage points. (Reference) Figure 10 and Figure 11 The model was further explained using two samples, in which... Figure 10 The predicted probability for the first sample was 0.92. High radiomics scores, smoking, and large fenestrations had a positive effect on the results, while narrow sinus cavities had an inhibitory effect. Figure 11 The predicted probability for the second sample was 0.27. Large sinus cavity size, smoking, and large fenestration had a positive effect on the results, while low radiomics scores had an inhibitory effect on the results. Figure 10 In this context, basevalue refers to the baseline value; higher indicates a higher level; lower indicates a lower level; and radiomics score is the radiomics score. Figure 11 In this context, basevalue refers to the baseline value; higher indicates a higher level; lower indicates a lower level; and radiomics score is the radiomics score.

[0140] Specifically, for a particular feature (j), its Shapley value is calculated using the following formula:

[0141] [ \phi_j = \sum_{S \subseteq N \setminus {j}} \frac{|S|! (|N|-|S|-1)!}{|N|!} \left[ f_S(\mathbf{x}_S \cup {x_j}) - f_S(\mathbf{x}_S) \right] ]

[0142] in:

[0143] (N) is the set of features.

[0144] (S \subseteq N \setminus { j} ) represents all feature subsets that do not contain feature ( j ).

[0145] (f_S(\mathbf{x}_S)) is the prediction made by the model using only the feature subset (S).

[0146] ( f_S(\mathbf{x}_S \cup { x_j}) ) is the model prediction resulting from the combined effect of feature subset ( S ) and feature ( j ).

[0147] In the model for predicting graft volume changes after maxillary sinus lift surgery, the SHAP value of each feature is calculated to regress feature importance. The SHAP library automatically handles these calculations; below is the Python implementation code:

[0148] import xgboost as xgb

[0149] import shap

[0150] import matplotlib.pyplot as plt

[0151] from sklearn.datasets import load_boston

[0152] from sklearn.model_selection import train_test_split

[0153] # Load the dataset (database of graft volume changes after maxillary sinus lift surgery)

[0154] boston = load_boston()

[0155] X, y = boston.data, boston.target

[0156] # Splitting the dataset

[0157] X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

[0158] # Training the GNB model

[0159] Model = xgb.GNBRegressor(objective='reg:squarederror', max_depth=6,eta=0.3, subsample=0.8, colsample_bytree=0.8)

[0160] model.fit(X_train, y_train)

[0161] # Interpreting using SHAP values

[0162] explainer = shap.TreeExplainer(model)

[0163] shap_values ​​= explainer.shap_values(X_train)

[0164] # Draw a feature importance graph

[0165] shap.summary_plot(shap_values, X_train, feature_names=boston.feature_names)

[0166] # Plot the SHAP value of a single sample

[0167] shap.initjs()

[0168] shap.force_plot(explainer.expected_value, shap_values[0,:], X_train[0,:], feature_names=boston.feature_names)

[0169] in:

[0170] Loading and preparing data:

[0171] Dataset of graft volume changes after maxillary sinus lift surgery.

[0172] Use train_test_split to split the data into training and test sets.

[0173] Training the GNB model:

[0174] The model was trained using xgb.GNBRegressor, with the objective function set to squared error (reg:squarederror).

[0175] Calculate the SHAP value:

[0176] Create a TreeExplainer object to interpret the GNB model.

[0177] Use shap_values ​​to obtain the SHAP values ​​of the training set data.

[0178] Plotting SHAP values:

[0179] `shap.summary_plot` displays the overall feature importance.

[0180] `shap.force_plot` displays the specific impact of a single sample's features on the prediction results.

[0181] Using the methods described above, and based on large-scale patient data, the influencing factors of graft volume changes after maxillary sinus lift surgery were screened using chi-square test and LASSO regression. A highly efficient predictive model for graft volume changes after maxillary sinus lift surgery was then constructed using machine learning algorithms. The Gaussian Naive Bayes algorithm demonstrated excellent performance, with area under the curve (AUC) values ​​of 0.939, 0.928, and 0.958 on the training, validation, and test sets, respectively. The predictive model for graft volume changes after maxillary sinus lift surgery based on the GNB model can effectively predict the probability of future graft volume changes after maxillary sinus lift surgery, thus solving the problem in existing technologies that require identifying the influencing factors of graft volume changes after maxillary sinus lift surgery and predicting these changes.

[0182] Figure 12This application provides a block diagram of a device for constructing a prediction model of graft volume change after maxillary sinus lift surgery based on a machine learning algorithm, according to at least one embodiment of the present application. The device includes: a memory 201; and a processor 202 connected to the memory 201, the processor 202 being configured to: acquire patient clinical data and CBCT radiomics data, form a dataset, and perform data preprocessing on the dataset;

[0183] Baseline analysis and LASSO regression were used to screen the feature variables in the dataset, and training, validation and test sets were constructed based on the screened feature variables.

[0184] An initial prediction model was constructed based on the Gaussian Naive Bayes algorithm;

[0185] The initial prediction model is trained based on the training set, validation set, and test set. By iteratively adding trees, one tree is added each time to fit the residual of the previous prediction, and finally a prediction model for graft volume change after maxillary sinus lift surgery with root mean square error meeting the preset requirements is obtained.

[0186] In some embodiments, the processor 202 is further configured to: acquire patient clinical data and CBCT radiomics data, specifically including:

[0187] Clinical data on factors influencing graft volume changes after maxillary sinus lift surgery were obtained, including: gender, age, smoking, history of periodontitis, history of hypertension, missing tooth position in the surgical area, history of diabetes, immediate graft volume after surgery, bone graft material size, blood growth factor usage, sinus width (SW), maxillary sinus contour, residual bone height (RBH), residual bone width (RBW), lateral wall thickness (LWT), buccal-palatal bone wall angle (MSA), fenestration size, distance from the lower boundary of the fenestration to the sinus floor, surgical lift height, and maxillary sinus mucosal thickness.

[0188] The patient's CBCT radiomics data were obtained, including 107 features in 7 categories, including 14 shape-based features, 18 first-order statistical features, 24 gray-level co-occurrence matrix features, 16 gray-level run length matrix features, 16 gray-level size region matrix features, 14 gray-level dependence matrix features, and 5 adjacent gray-level difference matrix features.

[0189] In some embodiments, the processor 202 is further configured to: perform data preprocessing on the dataset, specifically including:

[0190] Remove outliers and missing values ​​from the dataset, and fill in the missing data using interpolation.

[0191] Normalize or standardize the numerical variables in the dataset.

[0192] In some embodiments, the processor 202 is further configured such that the feature variables in the filtered dataset specifically include:

[0193] The categorical variables in the dataset were expressed as numbers and percentages. The normality test was performed using the Kolmogorov-Smirnov test. The continuous variables were compared using the Mann-Whitney U test. The categorical variables were compared using the Chisquared test or Fisher's exact test. The intergroup comparisons were performed using the independent samples t test or one-way ANOVA to screen out categorical variables with a two-sided P-value less than 0.05.

[0194] LASSO regression was used to further screen feature variables, specifically by: compressing the coefficients of variables in the regression model by generating a penalty function to prevent overfitting; and setting the coefficients of related independent variables to 0 based on the correlation between independent variables. The LASSO regression was configured with 10-fold cross-validation and executed using the R package.

[0195] In some embodiments, the processor 202 is further configured to: select clinical characteristic variables including: smoking, history of periodontitis, history of diabetes, size of bone graft material, sinus cavity size SW, and fenestration size. The selected radiomics characteristic variables include: Kurtosis, Dependence Non-Uniformity, and Large Dependence High Gray Level Emphasis, and calculate the radiomics score as: Radiomics score = 0.253 + (-0.674) × Kurtosis + 2.357 × Dependence Non-Uniformity + (-1.515) × Large Dependence High Gray Level Emphasis.

[0196] In some embodiments, the processor 202 is further configured such that the parameters of the maxillary sinus cyst prediction model include: a learning rate of 0.1, a maximum tree depth of 6, a minimum bifurcation weight sum of 1, and an L2 regularization coefficient of 0.1.

[0197] For specific implementation methods, please refer to the aforementioned method embodiments, which will not be repeated here.

[0198] This application may be a method, apparatus, system, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this application.

[0199] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0200] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0201] The computer program instructions used to perform the operations of this application may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuits, such as programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), are personalized by utilizing the status information of the computer-readable program instructions. These electronic circuits can execute the computer-readable program instructions to implement various aspects of this application.

[0202] Various aspects of this application are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0203] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0204] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0205] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0206] Note that, unless otherwise explicitly stated, all features disclosed in this specification (including any appended claims, abstract, and drawings) may be replaced by alternative features for achieving the same, equivalent, or similar purpose. Therefore, unless explicitly stated otherwise, each disclosed feature is merely one example of a set of equivalent or similar features. Where used, "further," "preferably," "even further," and "more preferably" are simply starting points for describing another embodiment based on the foregoing embodiments, the combination of which with the foregoing embodiments constitutes the complete configuration of another embodiment. Any combination of several "further," "preferably," "even further," or "more preferably" settings following the same embodiment constitutes yet another embodiment.

[0207] Although this application has been described in detail above with general descriptions and specific embodiments, some modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of this application fall within the scope of protection claimed in this application.

Claims

1. A method for constructing a predictive model for graft volume changes after maxillary sinus lift surgery based on machine learning algorithms, characterized in that, Includes the following steps: Acquire patients' clinical data and CBCT radiomics data and calculate radiomics scores to form a dataset, and perform data preprocessing on the dataset; The chi-square test and LASSO regression method were used to screen the feature variables in the dataset, and the training set, validation set and test set were constructed based on the screened feature variables. An initial prediction model was constructed based on the Gaussian Naive Bayes algorithm; The initial prediction model is trained based on the training set, validation set, and test set. By iteratively adding trees, one tree is added each time to fit the residual of the previous prediction, and finally a prediction model for graft volume change after maxillary sinus lift surgery with root mean square error meeting the preset requirements is obtained.

2. The method for constructing a prediction model for graft volume changes after maxillary sinus lift surgery based on machine learning algorithms according to claim 1, characterized in that, The acquisition of patient clinical data specifically includes: acquiring clinical influencing factors data on changes in graft volume after maxillary sinus lift surgery, including: gender, age, smoking, history of periodontitis, history of hypertension, missing tooth position in the surgical area, history of diabetes, immediate graft volume after surgery, size of bone graft material, use of blood growth factors, sinus cavity size, maxillary sinus contour, height of remaining alveolar ridge, width of remaining alveolar ridge, thickness of sinus lateral wall, buccal-palatal bone wall angle (MSA), fenestration size, distance from the lower boundary of the fenestration to the sinus floor, surgical lift height, and thickness of maxillary sinus mucosa; The acquisition of the patient's CBCT radiomics data specifically includes: acquiring the DICOM format file exported from CBCT and inputting it into 3DSlicer software to manually outline the region of interest layer by layer, generating the region of interest volume (VOI), and performing radiomics feature extraction. A total of 107 features in 7 categories are extracted, including 14 shape-based features, 18 first-order statistical features, 24 gray-level co-occurrence matrix features, 16 gray-level run length matrix features, 16 gray-level size region matrix features, 14 gray-level dependence matrix features, and 5 adjacent gray-level difference matrix features.

3. The method for constructing a prediction model for graft volume changes after maxillary sinus lift surgery based on machine learning algorithms according to claim 1, characterized in that, The data preprocessing of the dataset specifically includes: Remove outliers and missing values ​​from the dataset, and fill in the missing data using interpolation. Normalize or standardize the numerical variables in the dataset.

4. The method for constructing a prediction model for graft volume changes after maxillary sinus lift surgery based on machine learning algorithms according to claim 1, characterized in that, The feature variables in the dataset to be filtered specifically include: The categorical variables in the dataset are expressed as numbers and percentages, and the chi-square test is used to filter them out, selecting categorical variables with a two-sided p-value less than 0.

05. LASSO regression was used to further screen feature variables, specifically by: compressing the coefficients of variables in the regression model by generating a penalty function to prevent overfitting; and setting the coefficients of related independent variables to 0 based on the correlation between independent variables. The LASSO regression was configured with 10-fold cross-validation and executed using the R package.

5. The method for constructing a prediction model for graft volume changes after maxillary sinus lift surgery based on machine learning algorithms according to claim 4, characterized in that, Six clinical characteristic variables were selected, including: smoking, history of periodontitis, history of diabetes, size of bone graft material, sinus cavity size (SW), and fenestration size. The selected radiomics features include: Kurtosis, Dependence (Non-Uniformity), Large Dependence (High Gray Level Emphasis); The computational radiomics score is calculated as: Radiomics score = 0.253 + (-0.674) × Kurtosis + 2.357 × DependenceNonUniformity + (-1.515) × LargeDependenceHighGrayLevelEmphasis.

6. The method for constructing a prediction model for graft volume changes after maxillary sinus lift surgery based on machine learning algorithms according to claim 1, characterized in that, The parameters of the prediction model for graft volume change after maxillary sinus lift surgery include: learning rate of 0.1, maximum tree depth of 7, minimum branch weight sum of 1, and L2 regularization coefficient of 0.

1.

7. A method for predicting graft volume changes after maxillary sinus lift surgery, characterized in that, Including the following steps: The relevant feature values ​​of the patient are obtained and input into the prediction model of the change in graft volume after maxillary sinus lift surgery constructed by the method of constructing the prediction model of graft volume change after maxillary sinus lift surgery based on machine learning algorithm as described in any one of claims 1 to 7. The feature values ​​include at least one of smoking, history of periodontitis, history of diabetes, size of bone graft material, sinus cavity size SW, fenestration size, and radiomics score. The prediction model for graft volume change after maxillary sinus lift surgery outputs prediction results, giving the probability of graft volume change after maxillary sinus lift surgery.

8. The method for predicting graft volume changes after maxillary sinus lift surgery according to claim 7, characterized in that, Also includes: The SHAP value was used to interpret the prediction model of graft volume change after maxillary sinus lift surgery, demonstrating important features and their influence on graft volume change after maxillary sinus lift surgery. For each feature value in the prediction model of graft volume change after maxillary sinus lift surgery, its SHAP value was calculated to reflect the average contribution of different feature values ​​to the prediction results.

9. A device for constructing a predictive model of graft volume change after maxillary sinus lift surgery based on machine learning algorithms, characterized in that, include: Memory; And a processor connected to the memory, the processor being configured to perform the steps of the method as claimed in any one of claims 1 to 6.

10. A computer storage medium having a computer program stored thereon, characterized in that... When the computer program is executed by a machine, it implements the steps of the method as described in any one of claims 1 to 8.