A system and method for predicting the risk of peri-implantitis
Patent Information
- Application Number
- CN202610673297.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-15
- Publication Date
- 2026-09-29
AI Technical Summary
然而,现有技术存在以下不足:第一,多数模型仅依赖单一类型数据(如临床指标或影像学参数),未能融合微生物组学、代谢组学与影像学等多模态信息,导致预测信息不完整,准确率有限
[0040]1、本发明通过数据获取模块与多模态特征提取模块的协同工作,将龈沟液中的微生物组学、代谢组学数据与口腔影像学中的骨密度、骨吸收程度及黏膜厚度数据进行深度融合,构建了覆盖微生物-代谢-影像-临床维度的全面特征集,使得风险预测模型能够同时捕捉病原微生物丰度、宿主炎症反应、骨代谢失衡及局部解剖结构异常等多维病理信息,显著提升了种植体周围炎早期风险识别的敏感性和准确性,克服了传统单一指标预测信息不完整的缺陷。
Smart Images

Figure CN122842916A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of risk prediction technology, and in particular to a system and method for predicting the incidence of peri-implantitis. Background Technology
[0002] Peri-implantitis is one of the most common biological complications after oral implant surgery, mainly manifested as inflammation of the peri-implant mucosa and progressive loss of supporting bone tissue, which can lead to implant loosening and dislodgement in severe cases. Early identification and intervention of high-risk patients are crucial for improving long-term implant survival rates. Currently, clinical practice mainly relies on regular imaging examinations (such as periapical X-rays and CBCT) to assess the degree of bone resorption, combined with clinical indicators such as probing depth and bleeding index for risk assessment. However, this method is highly subjective and has low sensitivity, making it difficult to capture early risk signals before bone resorption occurs.
[0003] In recent years, researchers have attempted to use machine learning methods to build predictive models for peri-implantitis risk, such as those based on clinical baseline data (age, sex, smoking history, plaque index, etc.) or single-mode microbiome data (such as the abundance of specific pathogenic bacteria and the concentration of inflammatory factors in gingival crevicular fluid). However, existing technologies have the following shortcomings: First, most models rely on only a single type of data (such as clinical indicators or imaging parameters) and fail to integrate multimodal information such as microbiome, metabolomics, and imaging, resulting in incomplete predictive information and limited accuracy. Second, existing models are usually trained on datasets from a single medical center or a specific patient population. Due to differences in testing equipment, experimental procedures, regional dietary habits, and microbiome baselines among different centers, models are prone to overlearning noise specific to the training set (such as batch effects and local population characteristics). When applied to new patient populations, predictive performance significantly decreases, and generalization ability is poor. Third, existing technologies lack quantitative assessment methods for model overfitting and cannot dynamically adjust model weights according to changes in the distribution of new data. Once deployed, model performance cannot adaptively improve, requiring the collection of large-scale data and retraining, which is costly and time-consuming. Summary of the Invention
[0004] This invention provides a system and method for predicting the incidence risk of peri-implantitis in order to solve the existing technical problems, thereby addressing the issues raised in the background section.
[0005] To solve the above-mentioned technical problems, according to one aspect of the present invention, more specifically, a peri-implantitis incidence risk prediction system, comprising:
[0006] The data acquisition module is used to acquire gingival crevicular fluid samples, oral imaging data and clinical baseline data of the target patient at multiple time points after implantation.
[0007] The multimodal feature extraction module, connected to the data acquisition module, is used for:
[0008] A. Perform microbiome and metabolome analysis on the gingival crevicular fluid sample to extract microbial composition information, concentration data of inflammation-related metabolites and bone metabolism-related metabolites;
[0009] B. Perform image processing on the oral imaging data to extract data on peri-implant bone density, degree of bone resorption, and mucosal thickness;
[0010] The initial prediction model module is used to store the initial risk prediction model for peri-implantitis trained based on the first dataset; the first dataset is derived from the first patient group.
[0011] The model calibration module, connected to the multimodal feature extraction module and the initial prediction model module, is configured as follows:
[0012] C. Obtain a second dataset from a second patient group that has the same peri-implantitis symptoms as the first patient group;
[0013] D. Calculate the first data distribution of the first feature variable in the first dataset and the second data distribution of the corresponding second feature variable in the second dataset; both the first and second feature variables include: abundance of specific pathogens, concentration of inflammation-related metabolites, and concentration of bone metabolism-related metabolites;
[0014] E. Based on the deviation between the first data distribution and the second data distribution, calculate at least one noise intensity index, which is used to quantify the degree of overfitting of the initial risk prediction model to noise on the first dataset;
[0015] The risk prediction output module is connected to the multimodal feature extraction module and the model correction module, respectively. It is used to input the extracted data of the target patient into the risk prediction model after being corrected by the model correction module, and generate the risk probability value of the target patient's peri-implantitis.
[0016] Furthermore, the model correction module calculates the deviation between the first data distribution and the second data distribution, specifically including:
[0017] Calculate the first variance and first mean of each feature variable in the first dataset, and the second variance and second mean of the corresponding feature variables in the second dataset, where the distribution bias is calculated using the following formula:
[0018] ;
[0019] in, This represents the distribution deviation of the i-th feature variable; This represents the first mean of all feature variables in the first dataset; This represents the second mean of the corresponding feature variable in the second dataset; This represents the first variance of each feature variable in the first dataset; This represents the second variance of the corresponding feature variable in the second dataset; This represents a small constant that prevents division by zero. The penalty coefficient represents the variance difference.
[0020] Furthermore, the noise intensity index is calculated by integrating the distribution deviations of multiple characteristic variables, and the specific formula is as follows:
[0021] ;
[0022] in, Indicates the noise intensity index; This represents the total number of characteristic variables used for comparison; Indicates the first Preset weights for each feature variable; Indicates the scaling adjustment factor;
[0023] The noise intensity index ranges from 0 to 1, and a larger noise intensity index indicates a higher degree of overfitting of the initial risk prediction model to the data noise.
[0024] Furthermore, the model correction module dynamically adjusts the decision weights of the initial risk prediction model based on the noise intensity index;
[0025] When the noise intensity index exceeds a preset threshold, the decision weight of at least one specific feature variable based on the first dataset in the initial risk prediction model is reduced, and the contribution of the corresponding feature variable based on the second dataset is increased accordingly.
[0026] Furthermore, the microbial composition information includes at least the relative abundance data of Porphyromonas gingivalis, Forsythia stomatina, Treponema denticulatum, Prevotella intermedia, and Aggregates actinomycetes; the inflammation-related metabolites include at least one or more of prostaglandin E2, leukotrienes B4, and interleukin-1β.
[0027] The bone metabolism-related metabolites include at least one or more of osteocalcin, type I collagen cross-linked C-terminal peptide, and nuclear factor kb receptor activator ligand.
[0028] Furthermore, the multimodal feature extraction module extracts bone density around the implant by: performing three-dimensional reconstruction of CBCT images, dividing the core region with the implant neck as the center, calculating the average gray value of the core region, and converting it into bone density value through a gray-density curve.
[0029] The specific methods for determining the degree of bone resorption include: registering CBCT images at different postoperative time points and measuring the change in the vertical distance from the implant neck to the alveolar ridge crest over time.
[0030] Extracting mucosal thickness specifically includes: using digital scanning images of the oral cavity to identify the boundaries between the surface and the underlying layers of the mucosa, and measuring the vertical distance between the two boundaries.
[0031] Furthermore, the first dataset and the second dataset in the model calibration module are consistent in at least one of the clinical baseline data in terms of age range, gender, plaque index range, occlusal force distribution, and surgical method.
[0032] Furthermore, the initial risk prediction model stored in the initial prediction model module is constructed based on a machine learning algorithm, which includes any one of random forest, support vector machine, or gradient boosting tree.
[0033] A method for predicting the risk of peri-implantitis includes:
[0034] S1. Acquire gingival crevicular fluid samples, oral imaging data, and clinical baseline data of the target patient through the data acquisition module;
[0035] S2. Extract the microbial, metabolite, and imaging features of the target patient through the multimodal feature extraction module;
[0036] S3. Using the model calibration module, a noise intensity index is calculated to quantify the degree of overfitting of the initial model, based on a second dataset from an external second patient group.
[0037] S4. Dynamically correct the initial risk prediction model based on the noise intensity index;
[0038] S5. Input the characteristics of the target patient into the corrected risk prediction model, generate and output the probability value of the risk of peri-implantitis.
[0039] This invention provides a system and method for predicting the risk of peri-implantitis. Compared with existing technologies, the advantages of this method are:
[0040] 1. This invention, through the collaborative work of the data acquisition module and the multimodal feature extraction module, deeply integrates the microbiome and metabolomics data in gingival crevicular fluid with the bone density, bone resorption degree and mucosal thickness data in oral imaging, constructing a comprehensive feature set covering the dimensions of microorganism, metabolism, imaging and clinical aspects. This enables the risk prediction model to simultaneously capture multidimensional pathological information such as pathogenic microbial abundance, host inflammatory response, bone metabolic imbalance and local anatomical abnormalities, significantly improving the sensitivity and accuracy of early risk identification of peri-implantitis and overcoming the shortcomings of traditional single-indicator prediction information that is incomplete.
[0041] 2. This invention innovatively introduces the concept of noise intensity index. By comparing the mean difference and variance ratio of the same feature variable in the first patient group (initial training set) and the second patient group (external validation set), it quantitatively assesses the overfit of the initial risk prediction model to the noise specific to the training set. This index can objectively distinguish between real pathological signals and accidental noise such as equipment batch and regional differences, providing a clear quantitative basis for subsequent model calibration and solving the problem that existing technologies cannot assess the degree of decline in model generalization ability.
[0042] 3. This invention dynamically adjusts the decision weights of the risk prediction model based on the noise intensity index. When the index exceeds a preset threshold, it automatically reduces the contribution of unstable feature variables in the first dataset and correspondingly increases the weight of stable features in the second dataset. This adaptive correction mechanism allows the predictor to adapt to changes in data distribution in new medical centers and new patient groups without retraining the model, significantly improving the model's generalization ability across centers and populations, while reducing prediction bias caused by differences in detection equipment or baseline shifts in the population.
[0043] 4. This invention simplifies the complex model calibration mechanism into a standardized process from data acquisition, feature extraction, noise intensity index calculation, dynamic calibration to risk probability output. Each step has a clear mathematical form and biological interpretation, which facilitates deployment and implementation in a clinical setting. The final output of individualized risk probability values can directly guide postoperative management strategies such as enhanced follow-up, prophylactic antibacterial treatment, or occlusal adjustment, realizing a complete closed loop from multimodal data acquisition to clinical decision support, and has high practicality and scalability. Attached Figure Description
[0044] Figure 1 This is a flowchart of Embodiment 1 of the present invention;
[0045] Figure 2 This is a comparison chart of the relative abundance of microorganisms in this invention;
[0046] Figure 3 This is a comparison chart of inflammatory metabolite concentrations in this invention;
[0047] Figure 4 This is a comparison chart of bone metabolite concentrations in this invention;
[0048] Figure 5 This is a flowchart of Embodiment 2 of the present invention. Detailed Implementation
[0049] To make the technical solution of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0050] Example 1
[0051] like Figure 1 As shown, according to one aspect of the present invention, a peri-implantitis incidence risk prediction system is provided, comprising:
[0052] The data acquisition module is used to acquire gingival crevicular fluid samples, oral imaging data and clinical baseline data of the target patient at multiple time points after implantation.
[0053] The multimodal feature extraction module, connected to the data acquisition module, is used for:
[0054] A. Perform microbiome and metabolome analysis on gingival crevicular fluid samples to extract information on microbial composition, concentration data of inflammation-related metabolites and bone metabolism-related metabolites;
[0055] The microbial composition information includes at least the relative abundance data of Porphyromonas gingivalis, Forsythia stolonifera, Treponema denticulatum, Prevotella intermedia, and Aggregates actinomycetes.
[0056] Inflammation-related metabolites include at least one or more of prostaglandin E2, leukotrienes B4, and interleukin-1β;
[0057] Bone metabolism-related metabolites include at least one or more of osteocalcin, type I collagen cross-linked C-terminal peptide, and nuclear factor kappa receptor activator ligand.
[0058] B. Perform image processing on oral imaging data to extract data on peri-implant bone density, degree of bone resorption, and mucosal thickness; the multimodal feature extraction module for extracting peri-implant bone density specifically includes: performing three-dimensional reconstruction on CBCT images, dividing the core region with the implant neck as the center, calculating the average gray value of the core region, and converting it into bone density value through a gray-density curve.
[0059] The specific methods for determining the degree of bone resorption include: registering CBCT images at different postoperative time points and measuring the change in the vertical distance from the implant neck to the alveolar ridge crest over time.
[0060] Extracting mucosal thickness specifically includes: using digital scanning images of the oral cavity to identify the boundaries between the surface and the underlying layers of the mucosa, and measuring the vertical distance between the two boundaries.
[0061] The initial prediction model module stores the initial risk prediction model for peri-implantitis trained on the first dataset. The first dataset originates from the first patient population (specifically, the "headquarters" or internal data source group used to build and train the initial risk prediction model. This population's multimodal data (including microorganisms, metabolites, imaging, etc.) is used as the first dataset to train an initial prediction model. However, this model may suffer from decreased generalization ability due to overlearning noise specific to the first patient population rather than the general patterns of the disease itself (e.g., biases in testing equipment at specific medical centers, differences in metabolite baselines caused by regional dietary habits of the patient population). Therefore, the first patient population essentially represents a "biased" knowledge source learned by the original model, potentially containing specific environmental noise.
[0062] The initial risk prediction model stored in the initial prediction model module is built based on a machine learning algorithm, which includes any one of random forest, support vector machine, or gradient boosting tree.
[0063] The model calibration module, connected to the multimodal feature extraction module and the initial prediction model module, is configured as follows:
[0064] C. Obtain a second dataset from a second patient group. This second patient group has the same peri-implantitis symptoms as the first patient group. (The second patient group refers to an "external" or new data source group used for external validation and correction of the initial model. This group has the same clinical symptoms of peri-implantitis as the first patient group, but originates from different medical centers, regions, or periods. By obtaining this second dataset, the system can calculate the deviation between its data distribution (such as the abundance of specific pathogens, metabolite concentrations, etc.) and the data distribution of the first dataset. This deviation is defined as the noise intensity index, used to quantify the degree of overfitting of the initial model to the noise specific to the first patient group. In other words, the second patient group serves as a "gold standard" reference or a real-world distribution sample, revealing which parts of the learned patterns of the initial model are universal and which parts are merely accidental noise in the first patient group, thus providing a key basis for the dynamic correction of the model.)
[0065] D. Calculate the first data distribution of the first feature variable in the first dataset and the second data distribution of the corresponding second feature variable in the second dataset; both the first and second feature variables include: abundance of specific pathogens, concentration of inflammation-related metabolites, and concentration of bone metabolism-related metabolites.
[0066] The first dataset comes from the "first patient group" (e.g., post-implantation patients collected at a medical center headquarters), and contains 8 groups of samples (e.g., 4 healthy / low-risk patients + 4 disease-prone / high-risk patients, to reflect the data distribution). Each sub-table provides no fewer than 8 data records, along with necessary statistical characteristics (mean, variance), as shown in Tables 1, 2, and 3:
[0067] Table 1. Microbial composition information (relative abundance, unit: %)
[0068]
[0069] The data in Table 1 above represent the gradient from “healthy / low risk” (e.g., P01, P03, P07) to “illness / high risk” (e.g., P04, P06, P08), with the red complex (Porphyromonas gingivalis, Forsythia stomatis, Treponema denticulatum) significantly elevated in diseased samples.
[0070] Table 2. Inflammation-related metabolites (concentration units: pg / mL or ng / mL)
[0071]
[0072] In Table 2 above, the concentrations of inflammation-related metabolites were significantly elevated in the diseased samples (I02, I04, I06, I08), which is consistent with the pathological characteristics of local inflammatory reactions in peri-implantitis.
[0073] Table 3. Bone metabolism-related metabolites (concentration unit: ng / mL)
[0074]
[0075] In Table 3 above, bone resorption-related markers (CTX, RANKL) were elevated in the diseased samples, while bone formation markers (osteocalcin) were decreased in the diseased group, consistent with the pathological process of bone integration failure and accelerated bone resorption.
[0076] The dataset originates from a "second patient group" (e.g., external patients collected from another medical center who also present with peri-implantitis symptoms) and contains 8 samples. Compared to the first dataset, certain feature variables in the second dataset show significant differences in mean or variance, simulating a scenario where the original model's performance declines on new data, thus triggering the model calibration module to calculate the noise intensity index. The second dataset is shown in Tables 4, 5, and 6.
[0077] Table 4. Microbial composition information (relative abundance, unit: %)
[0078]
[0079] like Figure 2 As shown in Table 4, compared with the first dataset in Table 1 (mean values of 0.396, 0.316, 0.229, 0.495, and 0.069 respectively), the mean abundance of each pathogenic bacterium in the second dataset is generally higher, and the variance is larger, simulating a situation where the microbial load is higher or the sample heterogeneity is greater in the external environment.
[0080] Table 5. Inflammation-related metabolites (concentration units: pg / mL or ng / mL)
[0081]
[0082] like Figure 3 As shown in Table 5, the mean values of inflammatory metabolites in the second dataset are all higher than those in the first dataset in Table 2 (mean: 115.45, 87.70, 34.74), and the variance is significantly increased, which may reflect a higher degree of inflammatory response in external patients or differences in detection systems.
[0083] Table 6. Bone metabolism-related metabolites (concentration unit: ng / mL)
[0084]
[0085] like Figure 4 As shown in Table 6, the mean values of the bone resorption markers CTX and RANKL in the second dataset (0.876, 0.395) are significantly higher than those in the first dataset (0.73, 0.32) in Table 3, while the mean value of the bone formation marker osteocalcin (12.21) is lower than that in the first dataset (13.94), consistent with a more severe bone resorption phenotype. Variance is also generally increased.
[0086] E. Based on the deviation between the first data distribution and the second data distribution, calculate at least one noise intensity index. The noise intensity index is used to quantify the degree of overfitting of the initial risk prediction model to noise on the first dataset. Specifically, the model calibration module calculates the deviation between the first data distribution and the second data distribution, including:
[0087] Calculate the first variance and first mean of each feature variable in the first dataset, and the second variance and second mean of the corresponding feature variables in the second dataset. The distribution bias is calculated using the following formula:
[0088] ;
[0089] in, This represents the distribution deviation of the i-th feature variable; This represents the first mean of all feature variables in the first dataset; This represents the second mean of the corresponding feature variable in the second dataset; This represents the first variance of each feature variable in the first dataset; This represents the second variance of the corresponding feature variable in the second dataset; This represents a small constant that prevents division by zero. The penalty coefficient represents the variance difference.
[0090] Furthermore, the noise intensity index is calculated by integrating the distribution biases of multiple characteristic variables, and the specific formula is as follows:
[0091] ;
[0092] in, Indicates the noise intensity index; This represents the total number of characteristic variables used for comparison; Indicates the first Preset weights for each feature variable; Indicates the scaling adjustment factor;
[0093] The noise intensity index ranges from 0 to 1, and a higher noise intensity index indicates a greater degree of overfitting of the initial risk prediction model to the data noise. The first and second datasets in the model calibration module must maintain consistency in at least one of the following clinical baseline data: age range, gender, plaque index range, occlusal force distribution, and surgical procedure.
[0094] In this embodiment, the noise intensity index is calculated based on the data in Tables 1 to 6, specifically as follows:
[0095] Step 1: Set parameters (based on clinical practice and patented implementation examples). Among them... (Preventing the removal of zero constant); (Variance penalty coefficient); (Scale adjustment factor).
[0096] Step 2: Calculate the standard deviation of each feature.
[0097] Microbiome:
[0098]
[0099] Inflammatory metabolome:
[0100]
[0101] Bone metabolites:
[0102]
[0103] Step 3: Calculate the distribution deviation for each feature:
[0104]
[0105] Step 4: Calculate the noise intensity index:
[0106] ;
[0107] If equal weights are used The weighted sum is 1.4766 / 11 = 0.1342, therefore exp(-0.0671) = 0.9351. This indicates that the noise intensity index of 0.0649 is too small to reflect significant differences in data distribution. Therefore, in practical applications, a lower weighted index is usually used. Alternatively, weights can be assigned based on feature importance, and appropriate weights can be selected. This makes the noise intensity index sensitive to distribution deviations. The calculated noise intensity index of 0.522 here is reasonable.
[0108] The risk prediction output module is connected to the multimodal feature extraction module and the model correction module, respectively. It is used to input the extracted data of the target patient into the risk prediction model after being corrected by the model correction module, and generate the risk probability value of the target patient's peri-implantitis.
[0109] This embodiment constructs a closed-loop adaptive prediction system comprising an "initial prediction model module" and a "model calibration module." The system first uses a data acquisition module and a multimodal feature extraction module to deeply fuse microbiome and metabolomics data of gingival crevicular fluid with bone density, bone resorption, and mucosal thickness data from oral imaging (CBCT, digital scanning) to form a high-dimensional feature set. The initial prediction model is trained based on a first patient population (e.g., single-center data), but this model is prone to overlearning detection biases or regional noise specific to that population. Therefore, the system introduces a model calibration module. This module acquires data from a second patient population from different centers or periods, calculates the mean and variance of the same feature variable (e.g., abundance of specific pathogens, concentration of inflammatory metabolites) in both the first and second datasets, and calculates the distribution bias and noise intensity index using formulas. This index quantifies the degree of overfitting of the initial model to the noise specific to the first dataset, thereby triggering dynamic adjustments to the model's decision weights: when the noise intensity index exceeds a threshold, the weights of unstable feature variables in the first dataset are reduced, and the contribution of stable features in the second dataset is correspondingly increased.
[0110] This embodiment introduces the "noise intensity index" into the field of peri-implantitis risk prediction for the first time, achieving objective quantification and adaptive correction of model overfitting and significantly improving the model's generalization ability across different medical centers and patient groups. Compared with traditional static prediction models, this system can adapt to changes in data distribution in new environments online without retraining, reducing prediction bias caused by equipment batches and regional differences. Simultaneously, the fusion of multimodal features (microbiology-metabolism-imaging-clinical) enables the model to comprehensively capture the multidimensional pathological mechanisms of peri-implantitis development, improving prediction accuracy and interpretability. Furthermore, the design of variance penalty coefficients and scaling adjustment coefficients in the model calibration module makes the noise intensity index highly sensitive to distribution bias, accurately distinguishing between real pathological signals and data noise, thereby avoiding over-correction or under-correction.
[0111] Example 2
[0112] like Figure 1 As shown, according to one aspect of the present invention, a method for predicting the risk of peri-implantitis is provided, comprising:
[0113] S1. The data acquisition module obtains gingival crevicular fluid samples, oral imaging data, and clinical baseline data from the target patient. The occurrence and development of peri-implantitis are jointly regulated by the local microbial community, host immune inflammatory response, and bone metabolic status, and are also influenced by individual anatomical structure (such as bone density and mucosal thickness) and clinical baseline conditions (such as plaque index and occlusal force). Gingival crevicular fluid, as a direct sample of peri-implantal crevicular fluid, can enrich pathogenic microorganisms and their metabolites, inflammatory mediators, and bone turnover markers, serving as a "real-time window" reflecting the local pathological state.
[0114] Oral imaging (such as CBCT and digital scanning) can provide three-dimensional bone structure, bone resorption dynamics, and soft tissue contour information. Therefore, this step, through the simultaneous acquisition of multi-timepoint and multi-modal data, constructs a comprehensive feature set covering microbial, metabolic, imaging, and clinical dimensions, providing highly informative input data for subsequent accurate risk modeling.
[0115] S2. Microbiological, metabolite, and imaging features of the target patient are extracted using a multimodal feature extraction module. Raw biological and imaging data are characterized by high dimensionality and strong heterogeneity; directly using them for model training can easily introduce redundant noise. This step employs a modality-specific feature extraction algorithm: Microbiome analysis uses 16S rRNA sequencing or metagenomic analysis to obtain the relative abundance of key pathogenic bacteria (such as the red complex of *Porphyromonas gingivalis*); metabolomics analysis uses mass spectrometry or ELISA to quantify inflammation-related metabolites (PGE2, LTB4, IL-1β) and bone metabolism markers (osteocalcin, CTX, RANKL); and imaging analysis uses 3D reconstruction, grayscale-density calibration, multi-temporal registration, and boundary segmentation to quantify bone density, bone resorption, and mucosal thickness, respectively.
[0116] Thus, the raw data is transformed into quantitative features with clear biological significance, which reduces dimensionality while retaining predictive factors directly related to the pathogenesis, providing interpretable input variables for the model.
[0117] S3. Using the model calibration module, a noise intensity index is calculated to quantify the overfitting of the initial model, utilizing a second dataset from a second patient group. An initial model trained on a first patient group (e.g., a single center, a specific time period) may overlearn data noise specific to that group and not universally applicable (e.g., batch differences in detection equipment, regional microbial baseline shifts), leading to decreased generalization ability. The second patient group originates from different centers or periods but shares the same clinical diagnostic criteria; its data distribution reflects natural variability in the real world.
[0118] The noise intensity index is obtained by calculating the weighted combined deviation of the mean difference and variance ratio of the same feature variable (such as the abundance of a specific pathogen) in the first and second datasets. This index can quantitatively estimate the degree of overfitting of the initial model to batch and central noise on the training set: the higher the index, the fewer universal components in the learned patterns, and the more correction is needed. This method does not require intrusive cross-validation and directly utilizes external distribution differences not involved in training for diagnosis, making it efficient and compatible with clinical practice.
[0119] S4. Dynamically correct the initial risk prediction model based on the noise intensity index. The decision weights in the initial model are determined based on the best fit of the first dataset. However, when the data distribution shifts (i.e., the noise intensity index exceeds the threshold), the original weights amplify the impact of noise and reduce prediction accuracy. This step dynamically reduces the decision weights corresponding to feature variables with large variance in the first dataset and significant deviations from the distribution of the second dataset, based on the magnitude of the noise intensity index. At the same time, it correspondingly increases the contribution of stable features (i.e., predictors with cross-group consistency) in the second dataset.
[0120] This adaptive weighting mechanism based on distribution differences is equivalent to a Bayesian prior adjustment: treating the second dataset as a "real-world anchor" reduces the model's dependence on unreliable local features, thereby improving its robustness and generalization ability in multi-center and multi-scenario environments while maintaining the original model structure.
[0121] S5. Input the target patient's features into the corrected risk prediction model to generate and output their risk probability value for peri-implantitis. The corrected model has rebalanced the original decision weights using a noise intensity index, making it more sensitive to generalized pathological features (rather than occasional noise specific to a particular dataset). The target patient's feature vector (microbial abundance, metabolite concentration, imaging parameters) is input into the model after the same standardization process. The model outputs a probability value between 0 and 1 based on the weighted combination of each feature and the pre-trained classification boundary (such as decision tree voting in random forests or hyperplane distance in SVMs).
[0122] This probability value reflects the individualized risk of this patient developing peri-implantitis relative to the baseline population, after accounting for external data distribution bias. The output can be directly used for clinical decision support, such as guiding more frequent follow-up, prophylactic antibiotic treatment, or occlusal adjustment, thereby achieving precise postoperative risk stratification management.
[0123] This embodiment provides a standardized risk prediction process based on the aforementioned system. Its innovation lies in embedding "noise intensity index-driven dynamic correction" into the complete prediction pipeline. The method first acquires gingival crevicular fluid samples, oral imaging data, and clinical baseline data from multiple postoperative time points of the target patient, and transforms these into biologically meaningful quantitative features using a multimodal feature extraction module. Subsequently, using a second dataset from a second external patient group, the distribution deviation of the same feature variable between the first and second datasets is calculated, thus obtaining the noise intensity index. This index, by integrating the mean difference and variance ratio of multiple feature variables, quantitatively reflects the degree of overfitting of the initial model to the noise specific to the first dataset. Finally, the method dynamically adjusts the decision weights of the initial model based on this index, inputs the target patient's features into the corrected model, and outputs an individualized disease risk probability value.
[0124] This embodiment simplifies the complex model calibration mechanism into a clear and operable five-step process (S1-S5), facilitating deployment and implementation in clinical settings. Its core advantage lies in "dynamic adaptive calibration": when the initial model faces new patient groups or data distribution shifts, there is no need to recollect large-scale training data or retrain the model. Only a small amount of external second dataset is needed to calculate the noise intensity index, allowing for rapid adjustment of model weights and ensuring the prediction results remain sensitive to the actual pathological patterns. This method effectively solves the problems of poor cross-center generalization ability and overfitting to training set noise in existing peri-implantitis risk prediction models. Furthermore, the calculation process of the noise intensity index has a clear mathematical form and biological interpretation, with each step traceable, enhancing clinicians' confidence in the prediction results. The final output risk probability value can be directly used to guide individualized postoperative management (such as enhanced follow-up, prophylactic antibiotic treatment, or occlusal adjustment), achieving a complete closed loop from data collection to clinical decision-making.
[0125] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
Claims
1. A system for predicting the incidence risk of peri-implantitis, characterized in that, include: The data acquisition module is used to acquire gingival crevicular fluid samples, oral imaging data and clinical baseline data of the target patient at multiple time points after implantation. The multimodal feature extraction module, connected to the data acquisition module, is used for: A. Perform microbiome and metabolome analysis on the gingival crevicular fluid sample to extract microbial composition information, concentration data of inflammation-related metabolites and bone metabolism-related metabolites; B. Perform image processing on the oral imaging data to extract data on peri-implant bone density, degree of bone resorption, and mucosal thickness; The initial prediction model module is used to store the initial risk prediction model for peri-implantitis trained based on the first dataset; the first dataset is derived from the first patient group. The model calibration module, connected to the multimodal feature extraction module and the initial prediction model module, is configured as follows: C. Obtain a second dataset from a second patient group that has the same peri-implantitis symptoms as the first patient group; D. Calculate the first data distribution of the first feature variable in the first dataset and the second data distribution of the corresponding second feature variable in the second dataset; the first feature variable and the second feature variable both include: abundance of specific pathogens, concentration of inflammation-related metabolites and concentration of bone metabolism-related metabolites; E. Based on the deviation between the first data distribution and the second data distribution, calculate at least one noise intensity index, which is used to quantify the degree of overfitting of the initial risk prediction model to noise on the first dataset; The risk prediction output module is connected to the multimodal feature extraction module and the model correction module, respectively. It is used to input the extracted data of the target patient into the risk prediction model after being corrected by the model correction module, and generate the risk probability value of the target patient's peri-implantitis.
2. The peri-implantitis incidence risk prediction system according to claim 1, characterized in that: The model correction module calculates the deviation between the first data distribution and the second data distribution, specifically including: Calculate the first variance and first mean of each feature variable in the first dataset, and the second variance and second mean of the corresponding feature variables in the second dataset, where the distribution bias is calculated using the following formula: ; in, This represents the distribution deviation of the i-th feature variable; This represents the first mean of all feature variables in the first dataset; This represents the second mean of the corresponding feature variable in the second dataset; This represents the first variance of each feature variable in the first dataset; This represents the second variance of the corresponding feature variable in the second dataset; This represents a small constant that prevents division by zero. The penalty coefficient represents the variance difference.
3. The peri-implantitis incidence risk prediction system according to claim 2, characterized in that: The noise intensity index is calculated by integrating the distribution deviations of multiple characteristic variables, and the specific formula is as follows: ; in, Indicates the noise intensity index; This represents the total number of characteristic variables used for comparison; Indicates the first Preset weights for each feature variable; Indicates the scaling adjustment factor; The noise intensity index ranges from 0 to 1, and a larger noise intensity index indicates a higher degree of overfitting of the initial risk prediction model to the data noise.
4. The peri-implantitis incidence risk prediction system according to claim 3, characterized in that: The model correction module dynamically adjusts the decision weights of the initial risk prediction model based on the noise intensity index. When the noise intensity index exceeds a preset threshold, the decision weight of at least one specific feature variable based on the first dataset in the initial risk prediction model is reduced, and the contribution of the corresponding feature variable based on the second dataset is increased accordingly.
5. The peri-implantitis incidence risk prediction system according to claim 1, characterized in that: The microbial composition information includes at least the relative abundance data of Porphyromonas gingivalis, Forsythia stomatina, Treponema denticulatum, Prevotella intermedia, and Aggregobacter actinomycetes. The inflammation-related metabolites include at least one or more of prostaglandin E2, leukotrienes B4, and interleukin-1β; The bone metabolism-related metabolites include at least one or more of osteocalcin, type I collagen cross-linked C-terminal peptide, and nuclear factor kB receptor activator ligand.
6. The peri-implantitis incidence risk prediction system according to claim 1, characterized in that: The multimodal feature extraction module extracts bone density around the implant, specifically by: performing three-dimensional reconstruction of CBCT images, dividing the core region with the implant neck as the center, calculating the average gray value of the core region, and converting it into bone density value through a gray-density curve; The specific methods for determining the degree of bone resorption include: registering CBCT images at different postoperative time points and measuring the change in the vertical distance from the implant neck to the alveolar ridge crest over time. Extracting mucosal thickness specifically includes: using digital scanning images of the oral cavity to identify the boundaries between the surface and the underlying layers of the mucosa, and measuring the vertical distance between the two boundaries.
7. The peri-implantitis incidence risk prediction system according to claim 1, characterized in that: The first dataset and the second dataset in the model calibration module are consistent in at least one of the clinical baseline data in terms of age range, gender, plaque index range, occlusal force distribution, and surgical method.
8. The peri-implantitis incidence risk prediction system according to claim 1, characterized in that: The initial risk prediction model stored in the initial prediction model module is constructed based on a machine learning algorithm, which includes any one of random forest, support vector machine, or gradient boosting tree.
9. A method for predicting the risk of peri-implantitis, characterized in that, The peri-implantitis incidence risk prediction system according to any one of claims 1-8, wherein the peri-implantitis incidence risk prediction method comprises: S1. Acquire gingival crevicular fluid samples, oral imaging data, and clinical baseline data of the target patient through the data acquisition module; S2. Extract the microbial, metabolite, and imaging features of the target patient through the multimodal feature extraction module; S3. Using the model calibration module, a noise intensity index is calculated to quantify the degree of overfitting of the initial model, based on a second dataset from an external second patient group. S4. Dynamically correct the initial risk prediction model based on the noise intensity index; S5. Input the characteristics of the target patient into the corrected risk prediction model, generate and output the probability value of the risk of peri-implantitis.