A method for evaluating the bias and performance of a kawasaki disease coronary aneurysm prediction model based on probast-ai
By using the PROBAST-AI tool for four-domain feature extraction and performance index standardization, a three-dimensional evaluation framework of bias, performance, and validation was established. This solved the quality evaluation problem of the Kawasaki disease coronary aneurysm prediction model, achieving model standardization, repeatability, and comparability, and reducing the risk of high bias.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FUJIAN PROVINCIAL HOSPITAL
- Filing Date
- 2026-04-03
- Publication Date
- 2026-06-30
Smart Images

Figure CN122310052A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical prediction model quality assessment technology, and in particular to a method based on PROBAST-AI for assessing the bias and performance of a Kawasaki disease coronary aneurysm prediction model. Background Technology
[0002] Kawasaki disease (KD) is a leading cause of acquired heart disease in children, and coronary artery aneurysm (CAA) is its most serious complication. Although intravenous immunoglobulin therapy has reduced the overall incidence, early identification of CAA remains a clinical challenge. In recent years, numerous CAA risk prediction models based on algorithms such as traditional logistic regression, LASSO, random forest, gradient boosting tree, and deep learning have emerged. However, these models exhibit significant heterogeneity in research design, predictor variable selection, outcome definition, sample size, and validation methods, resulting in severely limited comparability and clinical generalizability.
[0003] To address the aforementioned issues, existing technologies primarily focus on the construction and optimization of single prediction models, rather than the system quality evaluation of multiple models. For example, CN109243604A discloses a method for constructing a Kawasaki disease risk assessment model based on neural networks, which optimizes a single model through feature selection and 10-fold cross-validation; CN112669968A discloses a cross-domain disease risk prediction method, which improves the generalization ability of a single model by fusing source and target domain data.
[0004] The common defects of the above-mentioned existing technologies are: (1) The technical goal is to develop new models, and they do not have the ability to conduct horizontal comparisons and systematic evaluations of multiple models in existing literature; (2) They lack a standardized bias risk assessment system and do not involve key quality control links such as participant selection, predictor measurement, outcome definition and analysis methods; (3) The validation methods are singular or chaotic, only internal validation (such as cross-validation) is used or the apparent performance, internal validation and external validation are not clearly distinguished, and the true generalization ability and overfitting risk of the model cannot be evaluated; (4) The performance index reporting dimensions are insufficient, mostly limited to the single AUC index, lacking indicators with stronger clinical significance such as confidence interval, calibration degree, decision curve, etc., and no unified comparability format has been established.
[0005] With the successive releases of the TRIPOD (Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis) reporting guidelines, the PROBAST bias assessment tool, and the TRIPOD-AI guidelines and PROBAST-AI methodological evaluation framework for AI prediction models (to be released in 2024-2025), internationally standardized tools have been established for evaluating the transparency and methodological quality of prediction model research. However, the application of this technology still faces the following key gaps:
[0006] (1) Lack of systematic application of PROBAST-AI
[0007] PROBAST-AI is a newly released international tool specifically designed for assessing the quality, risk of bias, and applicability of regression or AI prediction models. It systematically evaluates models across four domains: participants, predictors, outcomes, and analyses. However, this technical framework has not yet been introduced into the field of Kawasaki disease coronary artery aneurysm prediction models, resulting in a lack of standardized and reproducible technical means for evaluating model quality in this area.
[0008] (2) Lack of a joint evaluation framework integrating "bias-performance-verification".
[0009] Existing literature reviews are mostly limited to simple AUC numerical comparisons, failing to structurally integrate risk of bias levels, validation method classifications, and performance metrics. This fragmentation leads to high-risk-of-bias models being misjudged as high-quality models due to their apparent high AUC, or models lacking external validation being overly optimistically applied in clinical practice, posing serious risks to medical decision-making.
[0010] (3) Lack of ability to extract and standardize structured information
[0011] The predictors, outcome definitions, and performance index formats of different research reports are highly heterogeneous. Existing technologies do not provide a technical solution to convert non-standardized literature content into a unified data matrix, making horizontal comparisons between multiple models difficult and non-reproducible.
[0012] Therefore, there is an urgent need for a systematic evaluation method based on PROBAST-AI for predicting coronary artery aneurysms in Kawasaki disease, in order to achieve standardized comparison of bias risk, performance indicators and validation levels among multiple models, and to solve the technical problems of single evaluation dimensions, strong subjectivity and lack of internationally recognized framework support in the current technology. Summary of the Invention
[0013] To address the aforementioned issues, this application aims to provide a method for evaluating the bias and performance of Kawasaki disease coronary aneurysm prediction models based on PROBAST-AI. This method extracts model research information in a structured manner and uses the PROBAST-AI tool to evaluate the risk of bias. Simultaneously, it standardizes and extracts and displays performance indicators to achieve comparability, reproducibility, and transparency among different prediction models, thereby improving the ability of clinicians and researchers to judge the quality of different models.
[0014] To achieve the above objectives, this application discloses a method for evaluating the bias and performance of a Kawasaki disease coronary aneurysm prediction model based on PROBAST-AI, comprising the following steps:
[0015] S1. Model selection: Determine the set of Kawasaki disease coronary aneurysm prediction models to be evaluated as target models;
[0016] S2. Standardized Data Extraction: Standardized data extraction is performed on each target model according to the characteristics of the participant domain, predictor domain, structural domain, and analytical domain.
[0017] S3. Four-domain bias assessment: Using the PROBAST-AI assessment tool, based on the features extracted in S2, the systematic bias risk of each target model is assessed from the participant domain, predictor domain, outcome domain and analysis domain to obtain the four-domain bias risk level.
[0018] S4. Performance index standardization: The performance indices of each target model are uniformly extracted and formatted to obtain standardized performance indices, including AUC and its confidence interval, sensitivity, specificity, and calibration index.
[0019] S5. Validation Level Classification: The validation methods of each target model are classified and determined to obtain the validation level; the validation level includes apparent performance level, internal validation level and external validation level, and the performance index changes under different validation levels are recorded;
[0020] S6. Joint Evaluation and Display: The four-domain bias risk levels, standardized performance indicators, and validation levels are integrated into a structured evaluation matrix to generate cross-model comparison results.
[0021] Furthermore, as described in step S2:
[0022] Participant domain characteristics include study design type, sample size, inclusion and exclusion criteria, and data source;
[0023] The characteristics of the predictor domain include the type of predictor, the definition method, the measurement time point, and the blinding status;
[0024] The outcome domain features include the definition, judgment criteria, and measurement consistency of coronary aneurysm / coronary artery disease outcomes;
[0025] The characteristics of the analysis domain include the type of modeling algorithm, the method for handling missing data, and the methods for demonstrating and validating the sufficiency of the sample size.
[0026] Furthermore, the system bias risk assessment in step S3 includes: conducting quality assessments on the development and verification phases of the target model based on 18 signaling questions using the PROBAST-AI tool; and determining the four-domain bias risk level as low risk, questionable, or high risk.
[0027] Furthermore, the calibration indicators mentioned in step S4 include Hosmer-Lemeshow test results, calibration curve data, or Brier score values; and the optimal combination of sensitivity and specificity is reported using a multi-threshold study.
[0028] Furthermore, in step S5, the apparent performance level refers to calculating performance metrics only on the development dataset; the internal validation level uses cross-validation or Bootstrap resampling; the external validation level refers to validation on an external queue independent of the development set; and the magnitude of change of performance metrics from apparent to internal validation and from internal to external validation is calculated and recorded.
[0029] Furthermore, the integration of structured evaluation results in step S6 also includes: generating a bias risk distribution map, a performance comparison map of different validation levels, and a bias-performance scatter plot; the bias-performance scatter plot uses the overall bias risk score as the horizontal axis and the standardized AUC as the vertical axis to identify models with high apparent performance but high bias risk.
[0030] Furthermore, it also includes S7. Quality grading recommendation: Based on the structured evaluation results, a weighted scoring system is established, and models with low bias risk and satisfactory external validation performance are marked as high-quality recommendations, while models with high bias risk or lack of external validation are marked as models to be used with caution, thus generating a model selection guide.
[0031] Furthermore, by adjusting the outcome domain features in step S2 to define the outcome of the corresponding disease, the method can be extended to systematic reviews of clinical prediction models for other cardiovascular diseases.
[0032] Beneficial effects:
[0033] 1. Establish a standardized bias quantitative assessment system in the field of Kawasaki disease prediction models.
[0034] Existing technologies lack systematic quality assessment methods for predictive models of coronary artery aneurysms in Kawasaki disease, leading to reliance on subjective experience or single performance indicators for model quality judgment. This invention introduces the latest international PROBAST-AI assessment framework into the field of Kawasaki disease. Through four-domain feature extraction and systematic evaluation of 18 signaling questions, it transforms the risks of bias in participant selection, predictor measurement, outcome determination, and analytical methods from "invisible" to "quantifiable," solving the long-standing "methodological black box" problem in this field and making the model quality of different studies comparable.
[0035] 2. Construct a three-dimensional joint evaluation framework of "bias-performance-verification" to avoid misjudgment based on a single indicator.
[0036] Existing technologies rely solely on discrimination metrics such as AUC for model comparison, which can easily lead to high-risk bias models being misjudged as high-quality models due to inflated apparent performance. This invention establishes a multi-dimensional evaluation matrix by structurally integrating four-domain bias risk levels, standardized performance metrics (including confidence intervals and calibration), and three-level validation levels. This framework can identify unreliable models with "high bias and high apparent performance" while simultaneously discovering high-confidence models with "low bias and adequate validation," thus reducing medical risks caused by inappropriate model selection.
[0037] 3. Achieve standardization and reproducible extraction of heterogeneous research data.
[0038] To address the high heterogeneity in existing literature regarding research design, predictor variables, and outcome definitions, this invention transforms non-standardized literature information into a unified data matrix by defining a structured extraction standard of "four-domain features." This standardization process not only overcomes the technical obstacles of cross-model comparisons but, more importantly, ensures the reproducibility of the evaluation process—different evaluators can obtain consistent evaluation results based on the same four-domain feature definitions, meeting the reproducibility requirements of system review methodologies.
[0039] 4. Establish a three-level validation classification system to accurately assess the model's generalization ability and overfitting risk.
[0040] Existing technologies blur the lines between apparent performance, internal validation, and external validation, making it impossible to determine the true generalization ability of a model. This invention explicitly categorizes validation methods into apparent performance, internal validation, and external validation levels, and quantifies and records the decay of performance metrics from the development set to the independent validation set. This technical solution can objectively identify the degree of overfitting in a model, clearly distinguishing between preliminary models that "perform well on the development set but lack external validation" and mature models that have undergone "independent cohort validation," providing a scientific basis for the clinical application level of the model.
[0041] 5. Improve the screening efficiency and decision support capabilities of clinical prediction models.
[0042] By generating structured evaluation results that include bias risk distribution, validation level comparison, and bias-performance scatter plots, this invention enables clinicians and guideline developers to quickly identify models with superior comprehensive evidence. Compared to the traditional subjective screening method of reading each article one by one, this method can quickly locate high-quality models with "low bias risk + high external validation performance" among a large number of heterogeneous models, significantly shortening the translation path from evidence to clinical practice. It is particularly suitable for acute and critical illness scenarios such as Kawasaki disease, which require early identification of coronary artery aneurysm risk.
[0043] 6. The method exhibits good scalability and domain transferability.
[0044] Although this invention is optimized for Kawasaki disease coronary aneurysm prediction models, the four-domain assessment framework based on PROBAST-AI is disease-independent. By adjusting the definition of outcome domain features in step S2, this method can be seamlessly transferred to systematic reviews of clinical prediction models for other cardiovascular diseases, tumors, or chronic diseases, providing a general technical paradigm for quality control of prediction model research and avoiding the waste of resources in repeatedly developing assessment tools for each disease. Attached Figure Description
[0045] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1 The literature screening process for the model;
[0047] Figure 2 To summarize the models obtained through the screening process;
[0048] Figure 3 The temporal and regional distribution of data included in the Kawasaki disease coronary artery lesion prediction model; among which, Figure 3 A represents the time distribution of each prediction model. Figure 3 B represents the regional distribution of each prediction model.
[0049] Figure 4 It is the most commonly used predictor in machine learning and logistic regression models for coronary artery lesions in Kawasaki disease. Figure 4 A represents the 15 most frequently used predictors in ML models; Figure 4 B represents the 15 most commonly used predictors in the LR model.
[0050] Figure 5The PROBAST+AI evaluation results for each prediction model are shown; among them, Figure 5 A represents the quality assessment of the ML model during the model development phase. Figure 5 B represents the quality assessment of the LR model during the model development phase. Figure 5 C represents the applicability assessment of the ML model during the model development phase. Figure 5 D represents the applicability assessment of the LR model during the model development phase. Figure 5 E represents the bias risk assessment of the ML model during the model evaluation phase. Figure 5 F represents the bias risk assessment of the LR model during the model evaluation phase. Figure 5 G represents the applicability assessment of the ML model during the model evaluation phase. Figure 5 H represents the applicability assessment of the LR model during the model evaluation phase.
[0051] Figure 6 The results were categorized and organized according to the verification methods; among them Figure 6 A represents the performance of all studies; Figure 6 B represents the AUC of internal validation; Figure 6 C is the AUC of external validation. Detailed Implementation
[0052] To make the technical problems solved, the technical solutions, and the beneficial effects of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0053] Example 1: A method for evaluating the bias and performance of a Kawasaki disease coronary artery aneurysm prediction model based on PROBAST-AI.
[0054] This embodiment performs performance prediction by selecting coronary aneurysm prediction models that meet the standards from currently published data.
[0055] S1. Model selection: Determine the set of Kawasaki disease coronary aneurysm prediction models to be evaluated as target models;
[0056] In this embodiment, the target model used for evaluation should meet the following conditions: a predictive model has been constructed, the area under the curve (AUC) of the subjects' operating characteristics is reported, and the study population is matched and the outcome definition is consistent. The Kawasaki disease coronary artery aneurysm predictive model includes a series of predictive models related to coronary artery damage in Kawasaki disease.
[0057] Specifically, in this embodiment, a systematic search was first conducted in PubMed, Embase, and Web of Science databases, covering the period from 2017 to 2025. After deduplication, title and abstract screening, and full-text evaluation, literature that did not build a predictive model, did not report AUC, had mismatched study populations, inconsistent outcome definitions, or had been retracted was excluded. Finally, 28 studies on Kawasaki disease coronary artery damage prediction models that met the criteria were included as evaluation subjects. Figure 1 The document illustrates the specific process of literature screening, including the number of database searches, the deduplication process, the reasons for exclusion, and the final number of included studies. The inclusion model uses "first author + publication year" as the identifier, and the sources of the literature are summarized in [reference needed]. Figure 2 This serves as the basis for subsequent standardized data extraction, PROBAST+AI evaluation, and model performance comparison.
[0058] S2. Standardized Data Extraction: Standardized data extraction is performed on each target model according to the characteristics of the participant domain, predictor domain, structural domain, and analytical domain.
[0059] In this embodiment, the purpose of the standardized data extraction is to convert the non-standardized text content from different studies (i.e., the different models selected as research subjects in S1) into a comparable data matrix in a unified format. To perform standardized extraction and structured evaluation of the included Kawasaki disease coronary aneurysm / coronary artery damage prediction models, this invention organizes the model-related information into four aspects: participant domain, predictor domain, outcome domain, and analysis domain, which are as follows:
[0060] Participant domain characteristics: including study design type (retrospective or prospective, etc.), sample size, inclusion and exclusion criteria, and data sources;
[0061] Predictor domain characteristics: including the type of predictor, definition method (CAL or CAA, etc.), measurement time point, and blinding status;
[0062] Outcome domain characteristics: including the definition, judgment criteria, and measurement consistency of coronary artery aneurysm / coronary artery disease outcomes;
[0063] Analysis domain characteristics include modeling algorithm type, missing data handling method, and sample sufficiency demonstration and verification method.
[0064] The four domains mentioned above correspond to the key elements in predictive model construction and evaluation: the research object and data source on which the model is based, the model input variables, the model prediction objective, and the model analysis methods. Adopting this four-domain classification helps to standardize data extraction criteria across different studies, improving the standardization and comparability of model comparison, quality assessment, bias risk assessment, and applicability evaluation. This classification method is consistent with the basic evaluation framework of the PROBAST-AI tool.
[0065] Among them, the participant domain features refer to the research subjects and data sources involved in building or validating the predictive model, including the research design type, sample size, inclusion and exclusion criteria, and data sources; the predictor domain features refer to the relevant information of the model input variables, including the types of predictors, their definitions and measurement methods, measurement time points, and blinding status; the outcome domain features refer to the predictive targets of the model and their judgment information, including the definition, judgment criteria, and measurement consistency of coronary aneurysm or coronary artery disease outcomes; and the analysis domain features refer to the statistical analysis and algorithm information involved in the model building, processing, and validation process, including the modeling algorithm type, missing data handling methods, sample sufficiency verification, and validation methods.
[0066] like Figure 3 , Figure 4 As shown, this paper presents some results after standardizing and extracting data from the 28 prediction models obtained through the S1 screening. Figure 3 This reflects the temporal and regional distribution of data included in the prediction model. Here, ML represents a machine learning model, and LR represents a logistic regression model. The LR model includes penalized regression methods such as lasso regression, ridge regression, and elastic net regression. Figure 3 It is evident that the included models are mainly LR models, with relatively few ML models, and most of them were published in recent years. In terms of regional distribution, the relevant studies are mainly concentrated in East Asia, with fewer studies in North America and Western Europe. Figure 4 This reflects the distribution characteristics of high-frequency predictors included in the model. Among them, Figure 4 The 15 most frequently used predictors in the ML model shown in A are: age, C-reactive protein, alanine aminotransferase, aspartate aminotransferase, duration of fever, height, hemoglobin, platelet count, sex, white blood cell count, albumin, arrhythmia, basophil percentage, body surface area, and weight. Figure 4The 15 most commonly used predictors in the LR class model shown in B are: sex, age, albumin, C-reactive protein, intravenous immunoglobulin resistance, duration of fever, platelet count, hemoglobin, baseline Z-score, erythrocyte sedimentation rate, acromegaly, number of days of intravenous immunoglobulin administration, monocyte-to-HDL ratio, neutrophil-to-lymphocyte ratio, and white blood cell count. Figure 4 As can be seen, the predictive variables used in different modeling methods have certain commonalities. Indicators such as age, C-reactive protein, albumin, hemoglobin, platelet count, and white blood cell count are commonly found in both types of models, suggesting that basic clinical information and inflammation-related indicators are core variables in predicting the risk of coronary artery damage in Kawasaki disease. The results illustrated above collectively constitute an important output of the structured extraction step of this invention.
[0067] S3. Four-domain bias assessment: Using the PROBAST-AI assessment tool, based on the features extracted in S2, a systematic bias risk assessment is performed on each target model in the participant domain, predictor domain, outcome domain, and analysis domain to obtain a four-domain bias risk level; further, the systematic bias risk assessment includes: based on 18 signaling questions from the PROBAST-AI tool, quality assessments are performed on the development and validation phases of the target model respectively; the four-domain bias risk level is determined as low risk, questionable, or high risk.
[0068] Each domain is assigned a bias level of "low risk, questionable, or high risk" based on the tool entries, with specific reasons recorded. The applicability evaluation of the model is also recorded, including whether the predictor definitions are clear, whether the outcome measurements are objective and consistent, and whether the study design is practically generalizable. The evaluation results are used for subsequent joint analysis with performance indicators.
[0069] like Figure 5 The image shows a summary of the PROBAST+AI evaluation results for the standardized data of each prediction model extracted from S2. It is not an evaluation result for a single model, but rather a statistical display of the overall distribution of the included ML and LR models across different evaluation dimensions. Specifically, Figure 5 A and Figure 5 B represents the quality assessment results of the ML model and the LR model during the model development phase, respectively. Figure 5 C and Figure 5 D represents the applicability evaluation results of ML models and LR-type models during the model development phase, respectively; Figure 5 E and Figure 5 F represents the bias risk assessment results of the ML model and the LR model in the model evaluation stage, respectively; Figure 5 G and Figure 5H represents the applicability assessment results of ML and LR models during the model evaluation phase, respectively. Each subplot is summarized according to the PROBAST+AI participant and data source domain, predictor domain, outcome domain, analysis domain, and overall judgment. Stacked horizontal bar charts are used to display the proportion of included models judged as low risk, high risk, or unclear in each domain, with the horizontal axis representing the percentage of models with the corresponding rating. Figure 5 As can be seen, this invention not only enables standardized quality and bias assessment of different types of prediction models, but also intuitively reveals the differences in risk distribution across domains. For example, some models exhibit high risk or unclear proportions in the participant and data source domains and the analysis domain, while showing relatively good consistency in the predictor and outcome domains. These results provide a basis for subsequent target model selection, quality control, and optimal application, and are therefore an important intermediate product of this invention.
[0070] S4. Performance Index Standardization: The performance indices of each target model are uniformly extracted and formatted to obtain standardized performance indices, including AUC and its confidence interval, sensitivity, specificity, and calibration indices; the calibration indices include Hosmer-Lemeshow test results, calibration curve data, or Brier score values; for studies using multiple thresholds, the optimal combination of sensitivity and specificity is reported.
[0071] Because the performance metrics formats differ across research reports, this invention standardizes all performance metrics. Table 1 illustrates the standardization results: AUC is uniformly recorded as the AUC value and its 95% confidence interval; for multi-threshold studies, only the optimal combination of sensitivity and specificity is recorded; and for calibration metrics (such as the Hosmer-Lemeshow test, calibration curve, or Brier score), whether or not a report was submitted is uniformly recorded. This step ensures the comparability of performance metrics between different studies, providing a consistent quantitative basis for subsequent analysis.
[0072] Table 1
[0073]
[0074] S5. Validation Level Classification: The validation methods for each target model are classified and determined to obtain a validation level. The validation level includes apparent performance level, internal validation level, and external validation level, and the performance index changes under different validation levels are recorded. The apparent performance level refers to calculating performance indexes only on the development dataset. The internal validation level uses cross-validation or Bootstrap resampling methods. The external validation level refers to validation on an external queue independent of the development set. The magnitude of performance index changes from apparent to internal validation and from internal to external validation is calculated and recorded.
[0075] like Figure 6 As shown, this implementation method categorizes the validation methods for all prediction models into three types:
[0076] (1) Apparent performance: Performance is calculated only on the development dataset without any validation. Figure 5 A;
[0077] (2) Internal validation: The stability of the model was tested using methods such as cross-validation and bootstrap. Figure 5 B;
[0078] (3) External validation: Evaluate the model performance on independent datasets to reflect the model's generalization ability. Figure 5 C.
[0079] It also records the changes in AUC under different verification methods and can calculate the performance degradation, which can be used to analyze the model robustness and overfitting risk.
[0080] S6. Joint Evaluation and Presentation: The four-domain bias risk levels, standardized performance indicators, and validation levels are integrated into a structured evaluation matrix to generate cross-model comparison results. This integration into structured evaluation results also includes generating a bias risk distribution map, a performance comparison map of different validation levels, and a bias-performance scatter plot. The bias-performance scatter plot, with the overall bias risk score on the horizontal axis and standardized AUC on the vertical axis, is used to identify models with high apparent performance but high bias risk.
[0081] This implementation method integrates the bias assessment results, performance index data, and validation methods of each model into a structured evaluation framework, presented in tabular or matrix form (as shown in Table 2). This joint presentation method allows for horizontal comparisons between multiple models, identifying high-quality models with low bias risk, stable performance, and external validation support; it also identifies models with high bias but inflated apparent performance, indicating caution in clinical application. This structured framework facilitates researchers, clinicians, and guideline developers' rapid understanding of the usability and limitations of different models. Table 2 summarizes the PROBAST+AI evaluation results of the Kawasaki disease coronary artery damage prediction model during the model development and evaluation stages, including the participant and data source domain, predictor domain, outcome domain, analysis domain, and overall judgment. Table 3 systematically compares the differences in quality, bias risk, and applicability of different models, thereby identifying target models with high methodological quality, low bias risk, and good applicability, providing a basis for subsequent model selection, external validation, and clinical translation applications.
[0082] Table 2
[0083]
[0084] Table 3
[0085]
[0086] In a further embodiment, the method may also include S7. Quality grading recommendation: Based on the structured evaluation results, a weighted scoring system is established, models with low bias risk and satisfactory external validation performance are marked as high-quality recommendations, and models with high bias risk or lack of external validation are marked as models to be used with caution, thereby generating a model selection guide.
[0087] Furthermore, by adjusting the outcome domain features in step S2 to define the outcome of the corresponding disease, the method can be extended to systematic reviews of clinical prediction models for other cardiovascular diseases.
[0088] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for bias and performance evaluation of a PROBAST-AI based model for predicting coronary aneurysms in Kawasaki disease, characterized in that, Comprising the following steps: S1. Model screening: determine a set of coronary aneurysm prediction models for Kawasaki disease to be evaluated as target models; S2. Standardized data extraction: perform standardized data extraction on each target model according to participant domain characteristics, predictor domain characteristics, structure domain characteristics and analysis domain characteristics; S3. Four-domain bias evaluation: using the PROBAST-AI evaluation tool, based on the characteristics extracted in S2, perform systematic bias risk evaluation on each target model from the participant domain, predictor domain, outcome domain and analysis domain, and obtain four-domain bias risk levels; S4. Performance index standardization: uniformly extract and format the performance indicators of the target models to obtain standardized performance indicators, including AUC and its confidence interval, sensitivity, specificity, calibration indicators; S5. Verification level classification: classify the verification methods of the target models to obtain verification levels; the verification levels include apparent performance level, internal verification level and external verification level, and record the performance indicator changes under different verification levels; S6. Joint evaluation display: integrate the four-domain bias risk levels, standardized performance indicators and verification levels into a structured evaluation matrix to generate a multi-model horizontal comparison result.
2. The method of claim 1, wherein, In step S2: The participant domain characteristics include research design type, sample size, population inclusion and exclusion criteria, and data source; The predictor domain characteristics include predictor type, definition method, measurement time point and blind state; The outcome domain characteristics include the definition method, determination standard and measurement consistency of coronary aneurysm / coronary artery lesion outcome; The analysis domain characteristics include modeling algorithm type, missing data processing method, sample size adequacy demonstration and verification method.
3. The method of claim 1, wherein, The system bias risk evaluation in step S3 includes: based on the 18 signal questions of the PROBAST-AI tool, the development stage and the verification stage of the target model are respectively evaluated; the four-domain bias risk level is determined as low risk, doubtful or high risk.
4. The method of claim 1, wherein, The calibration indicators in step S4 include Hosmer-Lemeshow test results, calibration curve data or Brier score values; the optimal sensitivity and specificity combination of the multi-threshold research report is used.
5. The method of claim 1, wherein, The apparent performance level in step S5 refers to calculating performance indicators only on the development dataset; the internal verification level uses cross-validation or Bootstrap resampling method; the external verification level refers to verification on an external cohort independent of the development set; the performance indicators are calculated and recorded, and the change amplitude from apparent to internal verification, from internal to external verification is recorded.
6. The method of claim 1, wherein, In step S6, the integration into a structured evaluation result also includes: generating a bias risk distribution chart, a performance comparison chart at different verification levels and a bias-performance scatter chart; the bias-performance scatter chart takes the overall bias risk score as the horizontal axis and the standardized AUC as the vertical axis, which is used to identify models with high apparent performance but high bias risk.
7. The method of claim 1, wherein, S7. Quality grading recommendation: Based on the structured evaluation results, a weighted scoring system is established, and models with low bias risk and external validation performance up to standard are marked as high-quality recommendations, and models with high bias risk or lack of external validation are marked as cautious use, and a model selection guide is generated.
8. The method according to any one of claims 1 to 7, characterized in that, The method can be extended to the clinical prediction model system evaluation of other cardiovascular diseases by adjusting the outcome domain characteristics in step S2 to the definition of the outcome of the corresponding disease.
Citation Information
Patent Citations
Construction method and construction system for Kawasaki disease risk assessment model based on neural network algorithm
CN109243604A
Disease risk prediction method and equipment
CN112669968A