Prediction analysis method for implantation failure risk of maxillary sinus floor lifting operation
By building a three-dimensional artificial intelligence system, combining medical imaging segmentation and deep learning technology, three-dimensional spatial and radiomic features are extracted and multimodal fusion is carried out, the complexity of risk prediction of implant failure in maxillary sinus bottom lift surgery is solved, and efficient and accurate risk assessment and anatomical structure assessment are achieved.
Patent Information
- Application Number
- CN202510000710.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-02
AI Technical Summary
When predicting the risk of implant failure in maxillary sinus bottom lifting surgery, the prior art lacks the evaluation of complex risk factors surgery. Relying on a two-dimensional network leads to the loss of three-dimensional imaging data features, fails to fully automate input and output, the prediction results are inaccurate, and the decision results are weakly readable and interpretable.
By building a three-dimensional artificial intelligence system, using Python scripts for ROI positioning and cropping, combining nnU-Netv2 model for medical imaging segmentation, building a 3D-Attention-ResNet network to extract three-dimensional spatial features, and combining Pyradiomics module to extract radiomics features, using logistic regression models to multimodal fusion of imaging and clinical data to establish a prediction model.
The three-dimensional anatomical structure evaluation and risk assessment of maxillary sinus bottom lifting surgery is realized, the accuracy and readability of predictions are improved, and efficient preoperative planning guidance is provided, and multiple shortcomings in the prior art are overcome.
Smart Images

Figure CN119920406A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and in particular to a method for predicting and analyzing the risk of implant failure in maxillary sinus floor augmentation surgery. Background Art
[0002] In recent years, implant therapy has become the most effective method for treating tooth loss, but the implant failure rate cannot be ignored, especially in the maxillary posterior area with low bone density. When the pneumatization of the maxillary sinus (MS) leads to insufficient residual bone height (RBH), the risk of implant failure is even more worrying.
[0003] Currently, maxillary sinus floor elevation (MSFE) surgery is a common and reliable bone augmentation technique to solve this problem, including two approaches: translateral wall elevation and transalveolar ridge elevation. However, high clinical decision-making difficulty and technical sensitivity are long-standing problems of this surgery. Clinicians usually need to spend a lot of time using preoperative cone beam computed tomography (CBCT) to evaluate complex anatomical and physiological factors, such as the height and quality of the remaining alveolar bone, the morphology of the maxillary sinus floor, the presence of the sinus septum, and the health of the Schneider membrane. During the operation, doctors usually rely on manual feel to control the elevation of the sinus membrane. If they are not careful, complications such as sinus membrane perforation, bleeding, and infection are likely to occur. Therefore, it is difficult for doctors to accurately predict the fate of implants based on personal experience.
[0004] Electronic medical record information and imaging information are commonly used data for building risk prediction models. In the early days, scholars mostly only used patients' clinical statistical factors to build machine learning models. Although simple linear networks based on a large amount of text information can achieve high accuracy, they are difficult to capture important information such as bone quality and anatomical structure morphology in medical images.
[0005] Existing technologies use deep learning models to predict the occurrence of implant failure based on the features of CBCT slices, periapical films, and panoramic films, showing good accuracy and demonstrating the effectiveness of extracting features from radiological images for predicting implant failure. However, there is still much room for improvement: 1. These models are relatively broad in predicting the risk of implant failure, and the data sets are mostly composed of simple implant surgeries. There is a lack of evaluation of the predictive performance of surgeries with complex risk factors (such as maxillary sinus floor augmentation surgery), and the clinical practicality is poor; 2. These models rely on two-dimensional networks. For three-dimensional imaging data such as CBCT, using two-dimensional convolution kernels will lose image quality. 1. The models cannot fully automate the input and output of the images, and manually extracting the region of interest (ROI) will waste a lot of clinical time and energy of doctors. 2. The previous models performed end-to-end predictions without range constraints, which would introduce interference from irrelevant background information in the image, making it difficult to extract ROI-specific information with doctors' prior knowledge, which may affect the prognosis. 3. The previous models had weak readability and interpretability of decision-making results in actual clinical applications, and it was difficult to provide useful guidance for doctors' intervention measures in different surgeries. Summary of the invention
[0006] The present invention aims to provide a predictive analysis method for the risk of implant failure in maxillary sinus floor augmentation surgery, so as to solve the problems in the prior art of lack of evaluation of the predictive performance of maxillary sinus floor augmentation surgery, no analysis and evaluation of three-dimensional imaging data, no full automation of model input and output, inaccurate prediction results, and poor readability and interpretability of decision results.
[0007] In order to achieve the above object, the present invention provides the following method:
[0008] The present invention provides a method for predicting and analyzing the risk of implant failure in maxillary sinus floor augmentation surgery:
[0009] S1: Preoperative CBCT images and electronic medical records of patients undergoing maxillary sinus floor augmentation surgery were retrospectively collected;
[0010] S2: using a Python script to perform ROI positioning and cropping processing on the CBCT image, removing other unnecessary tissue structures and most background interference, and obtaining a preprocessed CBCT image; and dividing the preprocessed image data set into a training set and a test set;
[0011] S3: Use the open source software ITK-SNAP to outline and annotate each of the pre-processed CBCT images, including three categories: maxillary sinus morphology (MS), sinus floor Schneiderian membrane (SM), and residual alveolar bone (RAB), to obtain true image annotations paired with the pre-processed CBCT images;
[0012] S4: inputting the preprocessed CBCT image together with the paired real annotation of the image into the nnU-Netv2 medical image segmentation model for fine-tuning and training to generate segmentation masks of three categories: MS, SM and RAB;
[0013] S5: construct a 3D-ResNet network framework to extract three-dimensional spatial features from CBCT images, add a Squeeze-and-Excitation module as a spatial attention mechanism in the residual block to enhance feature representation by dynamically recalibrating channel importance, ensure that the network focuses on key features, and suppress irrelevant features, obtain a 3D-Attention-ResNet prediction network, input the segmentation masks of the MS, SM, and RAB and the preprocessed CBCT image into the 3D-Attention-ResNet prediction network through dual channels, form an enhanced feature set guided by anatomical morphology, obtain processed features by training the 3D-Attention-ResNet deep learning model, and generate a first probability score DLScore;
[0014] S6: using the Pyradiomics module to extract a plurality of radiomic features under the constraint of the segmentation mask from the CBCT image to further represent the specific information of the region of interest, using the Lasso-Logistic model to select and analyze the extracted plurality of radiomic features, and outputting a second probability score RadScore;
[0015] S7: Select the patient's electronic medical record information as candidate clinical data for the prediction model, use the Mann-Whitney test, χ2 test or Fisher's exact test to analyze the differences between the candidate clinical data in the high-risk group and the low-risk group, and screen statistically significant clinical variables as clinical data for the fusion model;
[0016] S8: Using a logistic regression model, the clinical data, the first probability score DLScore and the second probability score RadScore are integrated together to establish a multimodal fusion prediction model based on imaging and clinical data.
[0017] Preferably, the step of collecting CBCT images and electronic medical record information of patients after maxillary sinus floor lifting surgery includes: collecting CBCT images, CBCT reports, surgical cases and electronic archive records EMR of patients after maxillary sinus floor lifting surgery; marking patients who have implant loss after maxillary sinus floor lifting surgery as high-risk patients, and marking patients whose implant function is maintained for more than 5 years as low-risk patients based on the follow-up EMR; excluding cases with lack of clinical data, unclear CBCT images or obvious artifacts, and using random sampling technology to construct a relatively balanced data set for the low-risk patients.
[0018] Preferably, the CBCT image is positioned and cropped in the surgical area ROI using a Python script to remove most of the background interference such as other unnecessary tissue structures except the maxillary sinus morphology MS, the Schneider membrane SM of the sinus floor and the residual alveolar bone RAB, and the size of the obtained pre-processed CBCT image is a square image size including the maxillary sinus morphology MS, the Schneider membrane SM of the sinus floor and the residual alveolar bone RAB; the nnU-Netv2 medical image segmentation large model is fine-tuned and trained, including: setting the initial learning rate to 0.01, the weight decay coefficient to 3e-5, and the foreground oversampling rate to 33% to ensure that the model pays more attention to the foreground features. Each round of training includes 250 iterations, each round of verification includes 50 iterations, and a total of 1000 rounds of training. In addition, deep supervision can calculate losses at multiple network layers, thereby promoting gradient flow and improving training stability.
[0019] Preferably, after step S4, the method further includes using the average weight of five-fold cross-validation to predict and verify the segmentation performance of the preprocessed CBCT image; dividing the preprocessed CBCT image into five equal parts, cyclically using one of the parts as a validation set for five-fold cross-validation training, obtaining five different weights after training, and calculating the average weight of the five different weights; and using the average weight to predict and segment the test set of the preprocessed CBCT image.
[0020] Preferably, a 3D-ResNet network framework is constructed to extract three-dimensional spatial features from CBCT images, a Squeeze-and-Excitation module is added to the residual block as a spatial attention mechanism to enhance feature representation by dynamically recalibrating channel importance, ensuring that the network focuses on key features while suppressing irrelevant features, and obtaining a 3D-Attention-ResNet prediction network, and the segmentation masks of the MS, SM, and RAB and the preprocessed CBCT image are input into the 3D-Attention-ResNet prediction network through dual channels to form an enhanced feature set guided by anatomical morphology, and the processed features are obtained by training the 3D-Attention-ResNet deep learning model to generate a first probability score DLScore; the steps include: selecting ResNet as the network The invention relates to a method for extracting three-dimensional spatial features from CBCT images by using a 3D-ResNet network framework and extending it to three-dimensional convolution to construct a 3D-ResNet network framework; adding an SE module to the 3D-ResNet network framework as a spatial attention mechanism to obtain a 3D-Attention-ResNet prediction network; using the segmentation masks of the MS, SM and RAB together with the preprocessed CBCT image as inputs of the 3D-Attention-ResNet prediction network to obtain processed features; the processed features are classified through a fully connected layer using a SoftMax function to obtain predicted high-risk categories and predicted low-risk categories, and a first probability score DLScore is generated; the steps of constructing a 3D-ResNet network framework to extract three-dimensional spatial features from CBCT images include: the 3D-ResNet network is composed of a plurality of residual blocks, and the output of each residual block can be expressed as:
[0021] F residual =σ(W2*σ(W1*X+b1)+b2)
[0022] Where X∈R H×W×D×C is the input feature tensor, H, W, D are height, width and depth respectively, C is the number of channels; W1, W2 are weight matrices of two layers of 3D convolution respectively; b1 and b2 are bias terms; * represents a three-dimensional convolution operation; σ is the activation function ReLU; Squeeze-and-Excitation module is added to the residual block as a spatial attention mechanism to enhance feature representation by dynamically recalibrating channel importance, ensuring that the network focuses on key features while suppressing irrelevant features, including:
[0023] The Squeeze operation performs global average pooling on the features of each channel to generate a channel description vector:
[0024]
[0025] Among them, S c is the description of the c-th channel, c∈[1,C].
[0026] The Excitation operation transforms the channel description and generates channel weights through two fully connected layers:
[0027] z=σ2(W2·σ1(W1·s+b1)+b2)
[0028] Where s∈R C is the description vector of all channels; W1, W2 are the weight matrices of the fully connected layer, b1, b2 are bias terms; σ1 is the ReLU activation function, σ2 is the Sigmoid activation function; z∈R C are the normalized channel weights.
[0029] The Recalibration operation redistributes the generated channel weights z to the corresponding channels:
[0030] X′h,w,d,c=z c ·Xh,w,d,c
[0031] Among them, z c is the weight of the cth channel; the steps of obtaining the 3D-Attention-ResNet prediction network include: combining the 3D-ResNet and SE modules, the output of a single residual block can be expressed as:
[0032] F output =SE(σ(W2*σ(W1*X+b1)+b2))+X
[0033] The SE module is represented by SE(X), which includes the processes of Squeeze, Excitation and Recalibration; the segmentation masks of the MS, SM and RAB and the preprocessed CBCT image are input into the 3D-Attention-ResNet prediction network through dual channels to form an enhanced feature set guided by anatomical morphology, including: the input tensor is represented as:
[0034] I input =[X CBCT , M seg ]
[0035] Among them, X CBCT ∈R H×W×D M represents the feature matrix of the preprocessed CBCT image, with spatial dimensions H×W×D (height, width, depth); seg ∈R H×W×D Represents the segmentation mask feature matrix, which has the same CBCTThe same spatial dimension; [","] represents the concatenation operation on the channel dimension, and I input ∈R H×W×D×2 ; This tensor has two channels: the first channel corresponds to X CBCT , the second channel corresponds to M seg .
[0036] Preferably, the step of training the three-dimensional deep learning model includes: using cross entropy loss as a loss function, and its training formula is:
[0037]
[0038] Among them, loss is the loss function value, N is the total number of samples, y i is the true label of sample i, whose value is 0 or 1, p i The model predicts the probability that sample i belongs to category 1. In addition, the batch size is set to 1, and a round of validation is run after every three rounds of training. The AdamW algorithm is used as the optimizer. To prevent overfitting, Monai's Transforms library is used to rotate and flip the data for enhancement. The regularization parameter is set to 1e-4, and the CosineAnnealingLR strategy and early stopping strategy are used.
[0039] Preferably, the step of obtaining processed features and generating a first probability score DLScore includes: passing the processed features through a fully connected layer, using a SoftMax function to classify the features to obtain predicted high-risk categories and predicted low-risk categories, and generating a first probability score DLScore.
[0040] Preferably, the patient's electronic medical record information is selected as candidate clinical data for the prediction model, and the differences between the candidate clinical data in the high-risk group and the low-risk group are analyzed using the Mann-Whitney test, the χ2 test, or the Fisher's exact test. The step of screening statistically significant clinical variables as clinical data for the fusion model includes: selecting the patient's age, gender, blood sugar, smoking history, implant site, implant system, implant length, implant diameter, implant torque, bone augmentation, etc. as candidate clinical data for the prediction model according to existing medical literature; using the Mann-Whitney test, the χ2 test, or the Fisher's exact test to analyze the differences between the candidate clinical data; using the mean with standard deviation range to describe continuous variables, and using the frequency with percentage to describe categorical variables, and screening statistically significant clinical variables as clinical data for the fusion model.
[0041] Preferably, after step S8, the method further includes evaluating the fusion model on the test set, respectively evaluating the discrimination and goodness of fit of the fusion model through a receiver operating characteristic curve and a calibration curve; using dice similarity coefficient, union intersection, precision, recall rate and accuracy to evaluate the segmentation results of the three segmentation labels; using AUC, accuracy, sensitivity, specificity, positive predictive value, negative predictive value, F1-score and confusion matrix to evaluate the prediction results;
[0042] Preferably, the evaluation of the fusion prediction model also includes a clinical interpretative analysis of the fusion prediction model, including: decision curve analysis and clinical impact curves for evaluating the net risk-benefit and clinical practicality of the prediction model at different thresholds; Grad-CAM heat map for visualizing the focus area of the 3D-Attention-ResNet model on the image, and SHAP Python package for visualizing the feature importance ranking in the fusion model.
[0043] The beneficial effects of the present invention are as follows: the present invention successfully constructed a three-dimensional artificial intelligence system for the first time to predict the risk of implant failure after maxillary sinus floor lift surgery before surgery, and simultaneously segmented the MS, SM and RAB of the side, showing good accuracy, discrimination and calibration, and providing an efficient and objective method for preoperative anatomical structure evaluation and risk assessment of maxillary sinus floor lift surgery; the main advantages of this study include an automated multi-task system for segmentation and prediction, comprehensive feature extraction of fusion models, and high readability and interpretability of prediction results; this patent overcomes various problems that were not achieved in the original technology for predicting the risk of implant failure, and for the first time predicts the risk of implant failure in maxillary sinus floor lift surgery, and by constructing an anatomical structure segmentation model, solves the time-consuming and labor-intensive problem of manual cropping of ROI areas in the past, and at the same time provides doctors with efficient anatomical structure surgery Pre-planning guidance; constructing a three-dimensional convolutional neural network and introducing an attention mechanism solved the problem of deep learning losing longitudinal detail features in the process of three-dimensional CBCT image feature extraction; inputting the segmentation mask and the original image into the deep learning network at the same time, and combining with radiomics feature extraction, solved the interference problem of irrelevant information in the feature extraction process of the original deep learning technology, and added feature extraction of human prior knowledge guided by anatomical structure through radiomics; integrating deep learning features, radiomics features and clinical information features through machine learning to form a multimodal data set, and the fusion model greatly improved the prediction performance; by constructing a clinical nomogram, the prediction results are made more readable; through the Grad-Cam heat map and SHAP interpretation library, the image and clinical features of the data set are highlighted, which can guide physicians to take reasonable intervention measures to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for the specific embodiments or the description of the prior art. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn according to the actual scale.
[0045] Figure 1 It is a flow chart of a method for predicting and analyzing the risk of implant failure in maxillary sinus floor augmentation surgery provided by an embodiment of the present invention.
[0046] Figure 2 It is a key technology roadmap provided by the embodiments of the present invention.
[0047] Figure 3 It is a basic framework diagram of a network based on nn-UNetv2 provided in an embodiment of the present invention.
[0048] Figure 4 This is a three-dimensional residual network framework diagram based on the segmentation mask-assisted attention mechanism provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0049] In order to make the technical personnel in the technical field better understand the scheme of the present invention, the technical scheme in the embodiment of the present invention will be clearly and completely described below in combination with the drawings in the embodiment of the present invention. Obviously, the described embodiment is only a part of the embodiment of the present invention, not all the embodiments. Based on the embodiment of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0050] The terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish different objects rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, device, product or end including a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units inherent to these processes, methods, products or ends.
[0051] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present invention. The appearance of the phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0052] The current existing technologies have the following shortcomings: 1. These models have a relatively broad prediction of the risk of implant failure, and the data sets are mostly composed of simple implant surgeries. There is a lack of evaluation of the predictive performance of surgeries with complex risk factors (such as maxillary sinus floor lift surgery), and the clinical practicality is poor; 2. These models rely on two-dimensional networks. For three-dimensional imaging data such as CBCT, the use of two-dimensional convolution kernels will lose the longitudinal detail features of the image, especially in areas with complex anatomical structures such as the maxillary sinus area; 3. These models do not achieve full automation of input and output, and manual capture of regions of interest (ROI) will waste a lot of clinical time and energy of doctors; 4. We found that previous models that performed end-to-end predictions without ROI mask constraints would introduce interference from irrelevant background information in the image, making it difficult to extract ROI-specific information with doctors' prior knowledge, which may affect the prognosis; 5. The readability and interpretability of decision-making results of previous models in actual clinical applications are weak, and it is difficult to provide useful guidance for doctors' intervention measures in different surgeries.
[0053] The present invention aims to provide a predictive analysis method for the risk of implant failure in maxillary sinus floor augmentation surgery, so as to solve the problems in the prior art of lack of evaluation of the predictive performance of maxillary sinus floor augmentation surgery, no analysis and evaluation of three-dimensional imaging data, no full automation of model input and output, inaccurate feature extraction, and poor readability and interpretability of prediction results.
[0054] A specific embodiment of the present invention provides a method for predicting and analyzing the risk of implant failure in maxillary sinus floor augmentation surgery. Figure 1 , Figure 2 , Figure 3 and Figure 4 As shown, the following steps are included:
[0055] S1: Preoperative CBCT images and electronic medical records of patients who underwent maxillary sinus floor augmentation were retrospectively collected.
[0056] In an embodiment of the present invention, CBCT images, CBCT reports, surgical cases, and electronic archive records (EMR) of patients who underwent maxillary sinus floor lift surgery were collected; based on the follow-up EMR, patients whose implants fell out after maxillary sinus floor lift surgery were marked as high-risk patients, and patients whose implant function was maintained for more than 5 years were marked as low-risk patients; cases with lack of clinical data, unclear CBCT images, or obvious artifacts were excluded, and a relatively balanced data set was constructed for low-risk patients using random sampling technology.
[0057] S2: Use Python scripts to locate and crop CBCT images in ROI, remove other unnecessary tissue structures and most background interference, and obtain preprocessed CBCT images; and divide the preprocessed image dataset into training set and test set.
[0058] In this embodiment of the present invention, a Python script is used to perform ROI positioning and cropping processing on the CBCT image to remove most of the background interference such as other unnecessary tissue structures except the maxillary sinus morphology MS, sinus floor Schneider membrane SM and residual alveolar bone RAB on the surgical side. The size of the obtained preprocessed CBCT image is a square image size including the maxillary sinus morphology MS, sinus floor Schneider membrane SM and residual alveolar bone RAB.
[0059] S3: The open source software ITK-SNAP was used to outline and annotate each pre-processed CBCT image, including three categories: maxillary sinus morphology MS, sinus floor Schneiderian membrane SM, and residual alveolar bone RAB, to obtain the true image annotation paired with the pre-processed CBCT image.
[0060] S4: The preprocessed CBCT images together with the paired image ground truth annotations are input into the nnU-Netv2 medical image segmentation model for fine-tuning and training to generate segmentation masks for the three categories of MS, SM, and RAB.
[0061] In an embodiment of the present invention, the powerful image processing function of the model can automatically adapt to various medical images to ensure the accuracy and consistency of the data; the model is trained on a large finely annotated data set containing 164 CBCT cases and a total of 20,992 slices; then the average weight of five-fold cross validation is used to predict and verify the segmentation performance of the pre-processed CBCT image; the detailed hyperparameter settings are as follows: the initial learning rate is 0.01, the weight decay coefficient is 3e-5, and the foreground oversampling ratio is set to 33%, ensuring that the model pays more attention to the characteristics of the foreground area; each training cycle contains 250 iterations, the validation cycle contains 50 iterations, and the total training cycle is 1000; after step S4, it also includes using the average weight of the five-fold cross validation to predict and verify the pre-processed CBCT image; the pre-processed CBCT image is divided into 5 equal parts, and one of them is used as the validation set for five-fold cross validation training in a loop, and 5 different weights are obtained after training, and the average weight of the 5 different weights is calculated; the average weight is used to predict and segment the test set of the pre-processed CBCT image.
[0062] S5: Construct a 3D-ResNet network framework to extract three-dimensional spatial features from CBCT images. Add a Squeeze-and-Excitation module as a spatial attention mechanism in the residual block to enhance feature representation by dynamically recalibrating channel importance, ensuring that the network focuses on key features while suppressing irrelevant features. Obtain a 3D-Attention-ResNet prediction network. The segmentation masks of MS, SM, and RAB and the preprocessed CBCT images are input into the 3D-Attention-ResNet prediction network through dual channels to form an enhanced feature set guided by anatomical morphology. Obtain the processed features by training the 3D-Attention-ResNet deep learning model to generate the first probability score DLScore.
[0063] In the embodiment of the present invention, ResNet is selected as the network framework and extended to three-dimensional convolution to construct a 3D-ResNet network framework;
[0064] The SE module is added to the 3D-ResNet network framework as a spatial attention mechanism to obtain the 3D-Attention-ResNet prediction network;
[0065] The 3D-ResNet network consists of multiple residual blocks, and the output of each residual block can be expressed as:
[0066] F residual =σ(W2*σ(W1*X+b1)+b2)
[0067] Where X∈R H×W×D×C is the input feature tensor, H, W, D are height, width and depth respectively, C is the number of channels; W1, W2 are weight matrices of two layers of 3D convolution respectively; b1 and b2 are bias terms; * represents a three-dimensional convolution operation; σ is the activation function ReLU; Squeeze-and-Excitation module is added to the residual block as a spatial attention mechanism to enhance feature representation by dynamically recalibrating channel importance, ensuring that the network focuses on key features while suppressing irrelevant features, including:
[0068] The Squeeze operation performs global average pooling on the features of each channel to generate a channel description vector:
[0069]
[0070] Among them, S c is the description of the cth channel, c∈[1,C];
[0071] The Excitation operation transforms the channel description and generates channel weights through two fully connected layers:
[0072] z=σ2(W2·σ1(W1·s+b1)+b2)
[0073] Where s∈R C is the description vector of all channels; W1, W2 are the weight matrices of the fully connected layer, b1, b2 are bias terms; σ1 is the ReLU activation function, σ2 is the Sigmoid activation function; z∈R C is the normalized channel weight;
[0074] The Recalibration operation redistributes the generated channel weights z to the corresponding channels:
[0075] X′h,w,d,c=z c ·Xh,w,d,c
[0076] Among them, z c is the weight of the cth channel;
[0077] Combining 3D-ResNet and SE modules, the output of a single residual block can be expressed as:
[0078] F output =SE(σ(W2*σ(W1*X+b1)+b2))+X
[0079] Among them, the SE module is represented by SE(X), which includes the processes of Squeeze, Excitation and Recalibration;
[0080] The segmentation masks of MS, SM and RAB and the preprocessed CBCT images are input into the 3D-Attention-ResNet prediction network through dual channels to form an enhanced feature set guided by anatomical morphology, including the following steps: The input tensor is represented as:
[0081] I input =[X CBCT , M seg ]
[0082] Among them, X CBCT ∈R H×W×D M represents the feature matrix of the preprocessed CBCT image, with spatial dimensions H×W×D (height, width, depth); seg ∈R H×W×D Represents the segmentation mask feature matrix, which has the same CBCT The same spatial dimension; [","] represents the concatenation operation on the channel dimension, and I input ∈R H×W×D×2 ; This tensor has two channels: the first channel corresponds to X CBCT, the second channel corresponds to M seg ;
[0083] Using cross entropy loss as the loss function, the training formula is:
[0084] Among them, loss is the loss function value, is the total number of samples, is the true label of the sample, whose value is 0 or 1, and is the probability that the model predicts that the sample belongs to category 1;
[0085] Set the batch size to 1 and run a validation round after every three training rounds;
[0086] The AdamW algorithm is used as the optimizer. To prevent overfitting, the Monai Transforms library is used to rotate and flip the data for enhancement. The regularization parameter is set to 1e-4, and the CosineAnnealingLR strategy and early stopping strategy are adopted. The processed features are passed through a fully connected layer and classified using the SoftMax function to obtain predicted high-risk categories and predicted low-risk categories, generating the first probability score DLScore.
[0087] S6: Use the Pyradiomics module to extract several radiomic features under the constraints of the segmentation mask from the CBCT image to further represent the specific information of the region of interest, use the Lasso-Logistic model to select and analyze the extracted radiomic features, and output the second probability score RadScore.
[0088] In an embodiment of the present invention, the plurality of radiomics features include: intensity feature, texture feature, wavelet feature and three-dimensional shape feature.
[0089] S7: Select patient electronic medical record information as candidate clinical data for the prediction model, use the Mann-Whitney test, χ2 test, or Fisher's exact test to analyze the differences between the candidate clinical data, and screen statistically significant clinical variables as clinical data for the fusion model.
[0090] In an embodiment of the present invention, the patient's age, gender, blood sugar, smoking history, implant site, implant system, implant length, implant diameter, implant torque, bone augmentation, etc. are selected as candidate clinical data for the prediction model based on existing medical literature; the differences between the candidate clinical data are analyzed using the Mann-Whitney test, the χ2 test, or the Fisher exact test; the mean with the standard deviation range is used to describe continuous variables, and the frequency with percentage is used to describe categorical variables, and statistically significant clinical variables are screened as clinical data for the fusion model.
[0091] S8: Use the logistic regression model to integrate clinical data, the first probability score DLScore, and the second probability score RadScore to establish a multimodal fusion prediction model based on imaging and clinical data.
[0092] In an embodiment of the present invention, the fusion model is also evaluated, and the discrimination and fit of the fusion model are evaluated respectively by the receiver operating characteristic curve and the calibration curve; the segmentation results of the three segmentation labels are evaluated using the dice similarity coefficient, the union intersection, the precision, the recall rate and the accuracy; four cases are randomly selected from the test set to intuitively display the surface distance difference between the prediction results and the true annotations; AUC, accuracy, sensitivity, specificity, positive predictive value, negative predictive value, F1-score and confusion matrix are used to evaluate the prediction results; after the evaluation of the fusion prediction model, the fusion prediction model is also evaluated clinically. Explanatory analysis, including: decision curve analysis and clinical impact curves are used to evaluate the net risk benefit and clinical practicality of the prediction model at different thresholds; Grad-CAM heat map is used to visualize the focus area of the 3D-Attention-ResNet model on the image, and SHAP Python package is used to visualize the feature importance ranking in the fusion model.
[0093] The beneficial effects of the present invention are as follows: the present invention successfully constructed a three-dimensional artificial intelligence system for the first time to predict the risk of implant failure after maxillary sinus floor lift surgery before surgery, and simultaneously segmented the MS, SM and RAB of the side, showing good accuracy, discrimination and calibration, and providing an efficient and objective method for preoperative anatomical structure evaluation and risk assessment of maxillary sinus floor lift surgery; the main advantages of this study include an automated multi-task system for segmentation and prediction, comprehensive feature extraction of fusion models, and high readability and interpretability of prediction results; this patent overcomes various problems that were not achieved in the original technology for predicting the risk of implant failure, and for the first time predicts the risk of implant failure in maxillary sinus floor lift surgery, and by constructing an anatomical structure segmentation model, solves the time-consuming and labor-intensive problem of manual cropping of ROI areas in the past, and at the same time provides doctors with efficient anatomical structure surgery Pre-planning guidance; constructing a three-dimensional convolutional neural network and introducing an attention mechanism solved the problem of deep learning losing longitudinal detail features in the process of three-dimensional CBCT image feature extraction; inputting the segmentation mask and the original image into the deep learning network at the same time, and combining with radiomics feature extraction, solved the interference problem of irrelevant information in the feature extraction process of the original deep learning technology, and added feature extraction of human prior knowledge guided by anatomical structure through radiomics; integrating deep learning features, radiomics features and clinical information features through machine learning to form a multimodal data set, and the fusion model greatly improved the prediction performance; by constructing a clinical nomogram, the prediction results are made more readable; through the Grad-Cam heat map and SHAP interpretation library, the image and clinical features of the data set are highlighted, which can guide physicians to take reasonable intervention measures to a certain extent.
[0094] The above is only an embodiment of the present invention, and the common knowledge such as the known specific technical schemes or characteristics in the scheme is not described in detail here; it should be pointed out that for those skilled in the art, several modifications and improvements can be made without departing from the scheme of the present invention, which should also be regarded as the protection scope of the present invention, and these will not affect the effect of the present invention and the practicality of the patent. The protection scope required by this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.
Claims
1. A method for predicting and analyzing the risk of implant failure in maxillary sinus floor augmentation surgery, characterized in that: The method comprises: S1: Preoperative CBCT images and electronic medical records of patients undergoing maxillary sinus floor augmentation surgery were retrospectively collected; S2: using a Python script to perform ROI positioning and cropping on the CBCT image, removing other unnecessary tissue structures and most background interference, and obtaining a preprocessed CBCT image; and dividing the preprocessed image data set into a training set and a test set; S3: Use the open source software ITK-SNAP to outline and annotate each of the pre-processed CBCT images, including three categories: maxillary sinus morphology MS, sinus floor Schneiderian membrane SM, and residual alveolar bone RAB, to obtain true image annotations paired with the pre-processed CBCT images; S4: inputting the preprocessed CBCT image together with the paired real annotation of the image into the nnU-Netv2 medical image segmentation model for fine-tuning and training to generate segmentation masks of three categories: MS, SM and RAB; S5: construct a 3D-ResNet network framework to extract three-dimensional spatial features from CBCT images, add a Squeeze-and-Excitation module as a spatial attention mechanism in the residual block to enhance feature representation by dynamically recalibrating channel importance, ensure that the network focuses on key features, and suppress irrelevant features, obtain a 3D-Attention-ResNet prediction network, input the segmentation masks of the MS, SM, and RAB and the preprocessed CBCT image into the 3D-Attention-ResNet prediction network through dual channels, form an enhanced feature set guided by anatomical morphology, obtain processed features by training the 3D-Attention-ResNet deep learning model, and generate a first probability score DLScore; S6: using the Pyradiomics module to extract a plurality of radiomic features under the constraint of the segmentation mask from the CBCT image to further represent the specific information of the region of interest, using the Lasso-Logistic model to select and analyze the extracted plurality of radiomic features, and outputting a second probability score RadScore; S7: Select the patient's electronic medical record information as candidate clinical data for the prediction model, use the Mann-Whitney test, χ2 test or Fisher's exact test to analyze the differences between the candidate clinical data in the high-risk group and the low-risk group, and screen statistically significant clinical variables as clinical data for the fusion model; S8: Using a logistic regression model, the clinical data, the first probability score DLScore and the second probability score RadScore are integrated together to establish a multimodal fusion prediction model based on imaging and clinical data.
2. The method for predicting and analyzing the risk of implant failure in maxillary sinus floor augmentation surgery according to claim 1, characterized in that: The step of retrospectively collecting preoperative CBCT images of patients undergoing maxillary sinus floor lift surgery and the electronic medical records of the patients comprises: The CBCT images, CBCT reports, surgical records, and electronic archive records (EMR) of patients who underwent maxillary sinus floor augmentation surgery were collected; Based on the follow-up EMR, patients who experienced implant loss after sinus floor augmentation surgery were marked as high-risk patients, and patients whose implant function lasted for more than 5 years were marked as low-risk patients; Cases with lack of clinical data, unclear CBCT images or obvious artifacts were excluded, and a relatively balanced data set was constructed using random sampling techniques for the low-risk patients.
3. The method for predicting and analyzing the risk of implant failure in maxillary sinus floor augmentation surgery according to claim 2, characterized in that: The CBCT image was positioned and cropped in the surgical area using a Python script to remove most of the background interference, such as other unnecessary tissue structures except the maxillary sinus morphology MS, sinus floor Schneider membrane SM and residual alveolar bone RAB on the surgical side. The size of the obtained preprocessed CBCT image was a square image size including the maxillary sinus morphology MS, sinus floor Schneider membrane SM and residual alveolar bone RAB.
4. The method for predicting and analyzing the risk of implant failure in maxillary sinus floor augmentation surgery according to claim 1, characterized in that: After step S4, the method further includes using the average weight of the five-fold cross-validation method to predict and verify the segmentation performance of the pre-processed CBCT image; The preprocessed CBCT image is equally divided into 5 parts, one of which is used as a validation set for 5-fold cross-validation training, and 5 different weights are obtained after training, and the average weight of the 5 different weights is calculated; The average weight is used to perform predictive segmentation on the test set of the preprocessed CBCT images.
5. The method for predicting and analyzing the risk of implant failure in maxillary sinus floor augmentation surgery according to claim 1, characterized in that: The 3D-ResNet network framework is constructed to extract three-dimensional spatial features from CBCT images, a Squeeze-and-Excitation module is added to the residual block as a spatial attention mechanism to enhance feature representation by dynamically recalibrating channel importance, ensuring that the network focuses on key features while suppressing irrelevant features, and obtaining a 3D-Attention-ResNet prediction network. The segmentation masks of the MS, SM, and RAB and the preprocessed CBCT image are input into the 3D-Attention-ResNet prediction network through dual channels to form an enhanced feature set guided by anatomical morphology. The steps of training the 3D-Attention-ResNet deep learning model include: Select ResNet as the network framework and expand it to three-dimensional convolution to construct the 3D-ResNet network framework; The SE module is added to the 3D-ResNet network framework as a spatial attention mechanism to obtain a 3D-Attention-ResNet prediction network; The 3D-ResNet network consists of multiple residual blocks, and the output of each residual block can be expressed as: F residual =σ(W2*σ(W1*X+b1)+b2) Where X∈R H×W×D×C is the input feature tensor, H, W, D are height, width and depth respectively, C is the number of channels; W1, W2 are weight matrices of two layers of 3D convolution respectively; b1 and b2 are bias terms; * represents a three-dimensional convolution operation; σ is the activation function ReLU; the Squeeze-and-Excitation module is added to the residual block as a spatial attention mechanism to enhance the feature representation by dynamically recalibrating the channel importance, ensuring that the network focuses on key features while suppressing irrelevant features. The steps include: The Squeeze operation performs global average pooling on the features of each channel to generate a channel description vector: Among them, S c is the description of the cth channel, c∈[1,C]; The Excitation operation transforms the channel description and generates channel weights through two fully connected layers: z=σ2(W2·σ1(W1·s+b1)+b2) Where s∈R C is the description vector of all channels; W1, W2 are the weight matrices of the fully connected layer, b1, b2 are bias terms; σ1 is the ReLU activation function, σ2 is the Sigmoid activation function; z∈R C is the normalized channel weight; The Recalibration operation redistributes the generated channel weights z to the corresponding channels: X′h,w,d,c=z c ·Xh,w,d,c Among them, z c is the weight of the cth channel; Combining 3D-ResNet and SE modules, the output of a single residual block can be expressed as: F output =SE(σ(W2*σ(W1*X+b1)+b2))+X Among them, the SE module is represented by SE(X), which includes the processes of Squeeze, Excitation and Recalibration; The step of inputting the segmentation masks of the MS, SM and RAB and the preprocessed CBCT image into the 3D-Attention-ResNet prediction network through dual channels to form an enhanced feature set guided by anatomical morphology includes: the input tensor is represented as: I input =[X CBCT ,M seg ] Among them, X CBCT ∈R H×W×D M represents the feature matrix of the preprocessed CBCT image, with spatial dimensions H×W×D (height, width, depth); seg ∈R H×W×D Represents the segmentation mask feature matrix, which has the same CBCT The same spatial dimension; [","] represents the concatenation operation on the channel dimension, and I input ∈R H×W×D×2 ; This tensor has two channels: the first channel corresponds to X CBCT , the second channel corresponds to M seg .
6. The method for predicting and analyzing the risk of implant failure in maxillary sinus floor augmentation surgery according to claim 5, characterized in that: The step of training the 3D-Attention-ResNet deep learning model includes: Using cross entropy loss as the loss function, the training formula is: Among them, loss is the loss function value, N is the total number of samples, y i is the true label of sample i, whose value is 0 or 1, p i The model predicts the probability that sample i belongs to category 1; Set the batch size to 1 and run a validation round after every three training rounds; The AdamW algorithm is used as the optimizer. To prevent overfitting, Monai's Transforms library is used to rotate and flip the data for enhancement. The regularization parameter is set to 1e-4, and the CosineAnnealingLR strategy and early stopping strategy are adopted.
7. The method for predicting and analyzing the risk of implant failure in maxillary sinus floor augmentation surgery according to claim 1, characterized in that: The step of obtaining the processed features and generating a first probability score DLScore comprises: The processed features are passed through a fully connected layer and classified using the SoftMax function to obtain predicted high-risk categories and predicted low-risk categories, thereby generating a first probability score DLScore.
8. The method for predicting and analyzing the risk of implant failure in maxillary sinus floor augmentation surgery according to claim 1, characterized in that: The step of selecting the patient electronic medical record information as candidate clinical data for the prediction model, analyzing the difference between the candidate clinical data in the high-risk group and the low-risk group using the Mann-Whitney test, the χ2 test or the Fisher's exact test, and screening the clinical variables with statistical significance as the clinical data for the fusion model comprises: According to existing medical literature, the patient's age, gender, blood sugar, smoking history, implant site, implant system, implant length, implant diameter, implant torque, bone augmentation, etc. were selected as candidate clinical data for the prediction model; Differences between the candidate clinical data were analyzed using the Mann-Whitney test, χ2 test, or Fisher's exact test; Mean values with standard deviation ranges were used to describe continuous variables, frequencies with percentages were used to describe categorical variables, and statistically significant clinical variables were screened as clinical data for the fusion model.
9. The method for predicting and analyzing the risk of implant failure in maxillary sinus floor augmentation surgery according to claim 1, characterized in that: After step S8, the method further includes evaluating the fusion model on the test set, and evaluating the discrimination power and the goodness of fit of the fusion model respectively through a receiver operating characteristic curve and a calibration curve; The segmentation results of the three segmentation labels are evaluated using dice similarity coefficient, union intersection, precision, recall and accuracy; AUC, accuracy, sensitivity, specificity, positive predictive value, negative predictive value, F1-score and confusion matrix were used to evaluate the prediction results.
10. The method for predicting and analyzing the risk of implant failure in maxillary sinus floor augmentation surgery according to claim 9, characterized in that: After evaluating the fusion prediction model, the method further includes performing clinical interpretative analysis on the fusion prediction model, including: Decision curve analysis and clinical impact curves were used to evaluate the net risk-benefit and clinical utility of the prediction model at different thresholds; The Grad-CAM heat map is used to visualize the focus area of the 3D-Attention-ResNet model on the image, and the SHAP Python package is used to visualize the feature importance ranking in the fusion model.