Clinical data evaluation method and device, terminal equipment and storage medium

By acquiring clinical data from patients with systemic lupus erythematosus and using a risk probability prediction model to assess whether additional treatment is needed, the problem of existing technologies being unable to assess treatment needs has been solved, achieving the effects of reducing disease recurrence and lowering costs.

CN120932873APending Publication Date: 2025-11-11SHENZHEN PEOPLES HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510918197.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing methods for assessing systemic lupus erythematosus (SLE) can only determine whether a patient has SLE, but cannot assess whether additional treatment is needed. This leads to the inability to avoid disease relapse or exacerbation, increases treatment costs, and endangers patient safety.

Method used

By acquiring clinical data from target users, risk probability prediction models (such as target extreme gradient boosting prediction models, target random forest models, etc.) are used to assess the risk probability of needing additional treatment for systemic lupus erythematosus. The model evaluates based on multiple information items in the clinical data, including fever, rash, etc., and selects key features and calculates the evaluation results.

Benefits of technology

Effective assessment of systemic lupus erythematosus (SLE) requires increased treatment risk, reduced disease relapse or exacerbation, lower treatment costs, and ensure patient safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932873A_ABST
    Figure CN120932873A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of biomedicine. The invention discloses a clinical data assessment method and device, terminal equipment and a storage medium, which can realize assessment of the risk probability that systemic lupus erythematosus needs to increase treatment, and effectively avoid recurrence or deterioration of systemic lupus erythematosus diseases of a target user. The clinical data evaluation method comprises the following steps: acquiring clinical data of a target user; and inputting the clinical data into a risk probability prediction model of increased treatment for evaluation to obtain an evaluation result which is used for reflecting the risk probability that the systemic lupus erythematosus needs increased treatment of the target user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biomedical technology. More specifically, this application relates to a method, apparatus, terminal device, and storage medium for evaluating clinical data. Background Technology

[0002] For systemic lupus erythematosus (SLE), current assessment methods primarily focus on evaluating SLE activity. This involves receiving specific cell expression data from a target user via flow cytometry from at least one terminal; processing this data using a SLE assessment model to determine whether the target user is, or a candidate for, SLE. However, this method can only assess whether a target user has SLE; it cannot assess whether the user requires additional treatment. It cannot prevent relapse or exacerbation of SLE, thus increasing treatment costs and potentially endangering the user's life. Summary of the Invention

[0003] The purpose of this application is to provide a method, apparatus, terminal device, and storage medium for evaluating clinical data, which can assess the probability of increased treatment risk in patients with systemic lupus erythematosus (SLE) and effectively prevent relapse or exacerbation of SLE in target users. This application is mainly achieved through the following technical solutions:

[0004] A first aspect of this application provides a method for evaluating clinical data, including:

[0005] Obtain clinical data from the target users;

[0006] The clinical data is input into a risk probability prediction model for increased treatment and evaluated to obtain an evaluation result of the clinical data. The evaluation result is used to reflect the risk probability that the target user's systemic lupus erythematosus requires increased treatment.

[0007] According to one embodiment of this application, the clinical data includes fever information, rash information, hair loss information, pleural effusion information, interstitial pneumonia information, neuropsychiatric involvement information, hematuria information, mesenteric vasculitis information, white blood cell count information, hypocomplementemia information, anti-double-stranded DNA antibody information, anti-nucleosome antibody information, anti-nuclear antibody information, anti-U1 small nucleotide protein antibody information, disease duration information, joint pain information, urinary protein quantification information, hemoglobin information, platelet information, erythrocyte sedimentation rate information, neutrophil percentage information, albumin / globulin information, immunoglobulin information, and C-reactive protein information.

[0008] According to one embodiment of this application, the risk probability prediction model is selected from any one of the following models: target extreme gradient boosting prediction model, target random forest model, target support vector machine model, target multilayer perceptron, and target minimum absolute contraction and selection operator regression model.

[0009] According to one embodiment of this application, when the risk probability prediction model is a target extreme gradient boosting prediction model, the clinical data is input into the risk probability prediction model for increasing treatment for evaluation, and the evaluation result of the clinical data is obtained. The step of using the evaluation result to reflect the risk probability that the target user's systemic lupus erythematosus requires increased treatment includes:

[0010] Calculate the gain importance score, coverage importance score, and frequency score for each piece of information in the clinical data;

[0011] At least one target feature is selected based on the gain importance score, coverage importance score, and frequency score corresponding to each piece of information in the clinical data;

[0012] Calculate the first receiver operating characteristic curve and Youden index for each target feature;

[0013] Based on the first receiver operating characteristic curve and Youden index corresponding to each target feature, the evaluation result corresponding to each target feature is obtained.

[0014] According to one embodiment of this application, the clinical data evaluation method further includes a training step for the risk probability prediction model, the training step of the risk probability prediction model including:

[0015] Obtain a training dataset and a real label set, wherein each training data in the training dataset has a one-to-one correspondence with one of the real labels in the real label set;

[0016] The target training data is input into the original extreme gradient boosting prediction model for evaluation to obtain the prediction result. The target training data is any one of the training data in the training dataset.

[0017] A calibration curve is used to perform size matching processing on the predicted result and the real label corresponding to the target training data to obtain the matching result;

[0018] Adjust the first model parameter of the original extreme gradient boosting prediction model based on the matching result to obtain the risk probability prediction model.

[0019] According to one embodiment of this application, after adjusting the first model parameters of the original extreme gradient boosting prediction model based on the matching result to obtain the risk probability prediction model, the training step of the risk probability prediction model further includes:

[0020] The risk probability prediction model is evaluated using the area enclosed by the second receiver operating characteristic curve and the horizontal axis, in order to adjust the second model parameters of the risk probability prediction model.

[0021] According to one embodiment of this application, after adjusting the first model parameters of the original extreme gradient boosting prediction model based on the matching result to obtain the risk probability prediction model, the training step of the risk probability prediction model further includes:

[0022] The risk probability prediction model was evaluated using a clinical decision curve to adjust the third model parameter of the risk probability prediction model.

[0023] A second aspect of this application provides a clinical data evaluation apparatus, comprising:

[0024] The acquisition module is used to acquire clinical data from the target user.

[0025] The assessment module is used to input the clinical data into a risk probability prediction model for increased treatment and to obtain an assessment result of the clinical data. The assessment result is used to reflect the risk probability that the target user's systemic lupus erythematosus requires increased treatment.

[0026] A third aspect of this application provides a terminal device, including a processor and a memory, the memory being used to store a computer program, and the processor being used to call and run the computer program stored in the memory to perform the steps of the clinical data evaluation method provided in the first aspect of this application.

[0027] A fourth aspect of this application provides a computer-readable storage medium for storing a computer program that causes a computer to perform the steps of the clinical data evaluation method provided in the first aspect of this application.

[0028] The beneficial effects of the embodiments of this application include:

[0029] This application embodiment acquires clinical data from a target user; inputs the clinical data into a risk probability prediction model for increased treatment, and obtains an evaluation result for the clinical data. This evaluation result reflects the risk probability that the target user's systemic lupus erythematosus (SLE) requires increased treatment. Compared to existing technologies that can only assess whether a target user is a SLE patient, the risk probability prediction model designed in this application embodiment can be used to assess the risk probability that the target user's SLE requires increased treatment. Therefore, this application embodiment can effectively reduce the recurrence or exacerbation of the target user's SLE, thereby reducing the target user's treatment costs and effectively protecting the target user's life safety. Attached Figure Description

[0030] To more clearly illustrate the technical solutions in the embodiments of this application or the conventional technology, the drawings used in the description of the embodiments or the conventional technology will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0031] Figure 1 The flowcharts for the clinical data evaluation methods of this application are shown in some embodiments;

[0032] Figure 2 This is a reference diagram for partitioning the training dataset in this application;

[0033] Figure 3 This is a reference diagram for iterative analysis using cross-validation.

[0034] Figure 4 A reference graph showing the weights of the feature variables;

[0035] Figure 5 A bar chart showing the importance of composite features;

[0036] Figure 6 Reference graphs showing the sensitivity and specificity of the target extreme gradient boosting prediction model, SLEDAI-2K, SLE-DAS, and BILAG to the training set;

[0037] Figure 7 Reference plots showing the sensitivity and specificity of the target extreme gradient boosting prediction model, SLEDAI-2K, SLE-DAS, and BILAG to the test set;

[0038] Figure 8 Reference graphs showing the sensitivity and specificity of the target extreme gradient boosting prediction model, SLEDAI-2K, SLE-DAS, and BILAG to the validation set;

[0039] Figure 9 This is a reference diagram showing the sensitivity and uniqueness of the target extreme gradient boosting prediction model, target random forest model, target support vector machine model, target multilayer perceptron, and target minimum absolute shrinkage and selection operator regression model to the training set;

[0040] Figure 10 This is a reference graph showing the threshold probability and net gain of the target extreme gradient boosting prediction model for the training set;

[0041] Figure 11 This is a reference graph showing the predicted and actual probabilities of the target extreme gradient boosting prediction model for the training set.

[0042] Figure 12 This is a reference graph showing the sensitivity and uniqueness of the target extreme gradient boosting prediction model, the target random forest model, the target support vector machine model, the target multilayer perceptron, and the target minimum absolute shrinkage and selection operator regression model to the test set.

[0043] Figure 13 This is a reference graph showing the threshold probability and net profit of the target extreme gradient boosting prediction model for the test set;

[0044] Figure 14 This is a reference graph showing the predicted and actual probabilities of the target extreme gradient boosting prediction model for the test set.

[0045] Figure 15 This is a reference diagram showing the sensitivity and uniqueness of the target extreme gradient boosting prediction model, the target random forest model, the target support vector machine model, the target multilayer perceptron, and the target minimum absolute shrinkage and selection operator regression model to the validation set.

[0046] Figure 16 This is a reference graph showing the threshold probability and net profit of the target extreme gradient boosting prediction model for the validation set;

[0047] Figure 17 This is a reference graph showing the predicted and actual probabilities of the target extreme gradient boosting prediction model on the validation set.

[0048] Figure 18 The diagram below shows the principle block diagram of the clinical data evaluation device of this application in some embodiments.

[0049] Figure 19 This is a schematic block diagram of the terminal device of this application in some embodiments. Detailed Implementation

[0050] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the specific embodiments of this application are described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of this application. Therefore, this application is not limited to the specific embodiments disclosed below.

[0051] It should be noted that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0052] The terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0053] The terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units that are expressly listed, but may include other steps or units that are not expressly listed or that are inherent to such process, method, product, or apparatus.

[0054] Systemic lupus erythematosus (SLE) is a chronic autoimmune disease that can affect multiple organ systems. The natural course of SLE involves alternating periods of exacerbation and remission, with over 60% of patients experiencing persistent activity or relapse. This persistent disease activity not only exacerbates cumulative organ damage but also increases the risk of infection and even death in SLE patients (i.e., SLE patients). Therefore, assessing disease severity and early intervention are crucial for predicting disease progression and improving overall prognosis. Although the assessment of disease activity is well-researched, the complexity of SLE pathogenesis and clinical heterogeneity remain challenges in identifying disease progression and determining the optimal timing for escalating treatment.

[0055] Traditional cohort studies support stratified treatment, including first identifying patient subgroups based on different levels of SLE disease activity and then adjusting treatment accordingly. Common tools for assessing disease activity include the 2000 Systemic Lupus Erythematosus Disease Activity Index (SLEDAI-2K), the British Isles Lupus Assessment Group (BILAG), and the Systemic Lupus Erythematosus Disease Activity Status (SLE-DAS). However, these assessment tools all have limitations that may affect the adjustment of treatment strategies for patients with disease progression. For example, SLEDAI-2K has limited accuracy in guiding treatment escalation because each item uses a binary scoring system without grading the severity of characteristic variables, and some severe lupus manifestations are not included. Similarly, the BILAG assessment is time-consuming and lacks immunological indicators and complement levels. It only provides organ-based classification and does not offer a comprehensive assessment of overall disease activity. The SLE-DAS tool lacks some indicators that suggest disease activity, such as persistent fever. In addition to the inherent limitations of disease activity assessment tools, the relationship between disease activity and treatment strategies is not always linear. Therefore, the assessment results of the three conventional models, SLEDAI-2K, BILAG, and SLE-DAS, do not always meet the clinical needs for adjusting treatment in SLE patients. To date, a simple and practical tool remains lacking to directly calculate the probability of increased treatment risk in SLE patients.

[0056] In the treatment of SLE, previous research has focused on characterizing immune cell features to differentiate SLE patients from other autoimmune diseases and stratifying patient subgroups to enable targeted therapies. Other studies have aimed to predict drug responsiveness to facilitate evidence-based treatment choices by clinicians. The integration of multi-omics approaches helps discover novel biomarkers for targeted therapeutic interventions in SLE. Currently, few predictive models can directly predict the increased treatment needs of patients with systemic lupus erythematosus (SLE). Compared to previous studies focusing on disease subtypes, drug response prediction, or the exploration of multi-omics biomarkers, the model in this application emphasizes real-time assessment of evolving treatment needs in routine clinical practice. By integrating 24 routinely accessible clinical indicators during hospitalization, the model in this application can assess treatment risk without relying on single-cell sequencing or pharmacogenomics testing. Compared to molecular subtype models that require specialized laboratory testing, the predictive protocol of the model in this application, by linking with routine follow-up information, predicts the increased treatment needs of lupus patients, improving the efficiency of clinical decision-making and reducing economic and time costs.

[0057] Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The term "and / or" as used in this application includes any and all combinations of one or more of the associated listed items.

[0058] The specific embodiments of this application will be further described below with reference to the accompanying drawings.

[0059] refer to Figure 1 The diagram shown is a flowchart of a clinical data evaluation method provided in the first aspect of this application. Figure 1 The methods for evaluating the clinical data include:

[0060] S1. Obtain clinical data from the target users.

[0061] The clinical data includes information on fever, rash, hair loss, pleural effusion, interstitial pneumonia, neuropsychiatric involvement, hematuria, mesenteric vasculitis, white blood cell count, hypocomplementemia, anti-double-stranded DNA antibody, anti-nucleosome antibody, anti-nuclear antibody, anti-U1 small nucleotide protein antibody, duration of illness, joint pain, quantification of urine protein, hemoglobin, platelet count, erythrocyte sedimentation rate, neutrophil percentage, albumin / globulin ratio, immunoglobulin, and C-reactive protein.

[0062] The examples in the text are explained as follows: Fever information means a body temperature less than 38.0℃; rash information means no rash symptoms; hair loss information means no hair loss; pleural effusion information means no pleural effusion; interstitial pneumonia information means no interstitial pneumonia; neuropsychiatric involvement information means no neuropsychiatric involvement; hematuria information means no hematuria; mesenteric vasculitis information means no mesenteric vasculitis; white blood cell count information means no leukopenia; hypocomplementemia information means hypocomplementemia symptoms; anti-double-stranded DNA antibody (i.e., anti-dsDNA antibody) information is positive; anti-nucleosome antibody (i.e., ANUA) information is negative; antinuclear antibody... The following information was obtained: ANA (anti-associated natriuretic peptide) test was negative; anti-U1 small ribonucleoprotein antibody (anti-U1snRNP antibody) test was negative; disease duration was 240 months; joint pain was described as pain in two joints; urine protein quantification (PRO24h) was 0.6 g / 24h; hemoglobin was 120 g / L; platelet count was 200 × 10⁹ / L; erythrocyte sedimentation rate (ESR) was 4 mm / h; neutrophil percentage was 56%; albumin / globulin ratio was 1.3 g / L; immunoglobulin (IgG) was 18 g / L; and C-reactive protein (CRP) was less than 5 mg / L.

[0063] In other embodiments, the clinical data may also include other data, which can be set by those skilled in the art according to actual needs.

[0064] In other embodiments, the specific content of the fever information, rash information, hair loss information, pleural effusion information, interstitial pneumonia information, neuropsychiatric involvement information, hematuria information, mesenteric vasculitis information, white blood cell information, hypocomplementemia information, anti-double-stranded DNA antibody information, anti-nucleosome antibody information, anti-nuclear antibody information, anti-U1 small nucleotide protein antibody information, disease duration information, joint pain information, urinary protein quantification information, hemoglobin information, platelet information, erythrocyte sedimentation rate information, neutrophil percentage information, albumin / globulin information, immunoglobulin information, and C-reactive protein information can also be other information, which can be set by those skilled in the art according to actual needs.

[0065] S2. Input the clinical data into the risk probability prediction model for increased treatment and evaluate it to obtain the evaluation result of the clinical data. The evaluation result is used to reflect the risk probability that the target user's systemic lupus erythematosus requires increased treatment.

[0066] Additional treatment options include the addition of corticosteroids, cytotoxic therapy, biologics, or a shift from mild to strong cytotoxic therapy.

[0067] Furthermore, the risk probability prediction model is selected from any one of the following models: target extreme gradient boosting prediction model, target random forest model, target support vector machine model, target multilayer perceptron, and target minimum absolute contraction and selection operator regression model.

[0068] The target extreme gradient boosting prediction model is a pre-trained XGBoost model; the target random forest model is a pre-trained RF model; the target support vector machine model is a pre-trained SVM model; the target multilayer perceptron is a pre-trained MLP, and is a multilayer perceptron with a sigmoid activation function and a single hidden layer; the target minimum absolute shrinkage and selection operator regression model is a pre-trained LASSO model.

[0069] Among the five models mentioned above, the target extreme gradient boosting prediction model exhibits the best performance. Its optimized hyperparameters include a learning rate of 0.1, a maximum tree depth of 3, and 100 boosting rounds, while other parameters remain at their default values. The hyperparameter settings for the other models are as follows: the target minimum absolute shrinkage and selection operator regression model uses L1 regularization, 25 cross-validation folds, and "auc" as the performance metric; the target random forest model has 400 decision trees, considers 28 variables per split, and has a minimum sample size of 10 for each leaf node u; the target multilayer perceptron has a hidden layer size of 5, uses "logistic" activation function, and employs "rprop+" training algorithm; and the target support vector machine model uses a "radial basis function" kernel.

[0070] Furthermore, based on the aforementioned example of clinical data, the assessment result could be 93.6%. Healthcare professionals then determine whether the assessment result exceeds a preset threshold of 65.6%. If it exceeds the preset threshold, it indicates that the target user is at high risk of requiring additional treatment; otherwise, it indicates low risk, and the current treatment strategy can be maintained.

[0071] The preset threshold can be 65.6%. In other embodiments, the specific value of the preset threshold can be set by those skilled in the art according to actual needs.

[0072] Through the above implementation methods, compared with the prior art which can only be used to assess whether a target user is a patient with systemic lupus erythematosus, the risk probability prediction model designed in this application embodiment can be used to assess the risk probability that the target user's systemic lupus erythematosus requires additional treatment. Therefore, this application embodiment can effectively reduce the recurrence or exacerbation of the target user's systemic lupus erythematosus, thereby reducing the target user's treatment costs and effectively protecting the target user's life safety.

[0073] In some implementations, when the risk probability prediction model is a target extreme gradient boosting prediction model, step S2 includes:

[0074] S21. Calculate the gain importance score, coverage importance score, and frequency score corresponding to each piece of information in the clinical data.

[0075] Furthermore, the feature_importances attribute of the prediction model is improved through the target extreme gradient to obtain the gain importance score, coverage importance score, and frequency score corresponding to each piece of information.

[0076] In other implementations, the `get_booster().get_score()` method can be used with `importance_type='gain'` to obtain the gain importance score for each piece of information; the `get_booster().get_score()` method can be used with `importance_type='cover'` to obtain the coverage importance score for each piece of information; and the `get_booster().get_score()` method can be used with `importance_type='weight'` to obtain the frequency score for each piece of information.

[0077] The gain importance score reflects the contribution (e.g., prediction accuracy) of each piece of information in the clinical data to the performance of the target extreme gradient boosting prediction model; the coverage importance score reflects the sample range affected by each piece of information in the clinical data in the target extreme gradient boosting prediction model; and the frequency score reflects the frequency with which each piece of information in the clinical data is used in the target extreme gradient boosting prediction model.

[0078] This application embodiment uses R software (version 4.3.0, CRAN) to transform the target extreme gradient boosting prediction model into an online prediction website. The establishment of the online prediction website first involves standardizing and preprocessing multi-center clinical data, and then extracting structured data from the electronic medical record systems of two medical centers (e.g., the First Hospital and the Second Hospital). The information type of the structured data is the same as that of the clinical data. The URL of the online prediction website can be found at https: / / xgboostsle.shinyapps.io / xgapp / .

[0079] This embodiment of the application also uses the mise package to perform multiple imputation on the missing values ​​of the structured data. The outcome variable is defined with treatment labels 1 and 0, where 1 represents that the SLE patient (i.e., the target user) needs additional treatment; 0 represents maintaining the current treatment plan. The target extreme gradient boosting prediction model evaluates the importance of each piece of information in the clinical data or structured data using three indicators: gain, coverage, and frequency, and performs feature selection. Regarding hyperparameter tuning, this embodiment of the application uses Bayesian optimization to determine the optimal parameter combination (max_depth = 3, eta = 0.5, nrounds = 100, and the rest remain at default values). The risk probability prediction model training converts the data into DMatrix format and trains an XGBoost binary classification model.

[0080] After the target extreme gradient boosting prediction model is transformed into an online prediction website, the website's interface design uses the Shiny framework to construct dynamic forms, containing two types of input components: one is numerical input, i.e., the numericInput() component; the other is binary selection, i.e., the checkboxInput() component. Additionally, the online prediction website utilizes real-time prediction logic to respond to the actionButton("predict") event on the server side.

[0081] Furthermore, the online prediction website is deployed as a cloud service. Specifically, this involves containerizing the R environment dependencies using the renv package to generate a reproducible package manifest, namely the renv::snapshot() function.

[0082] Furthermore, the online prediction website is continuously integrated and deployed. Specifically, it is deployed to the Shinyapps.io platform with a single click via rsconnect, specifically the rsconnect::deployApp() function. The online prediction website takes 2-3 minutes to output the assessment results, and these results depend on positive items in the clinical data.

[0083] S22. Based on the gain importance score, coverage importance score and frequency score corresponding to each piece of information in the clinical data, at least one target feature is selected.

[0084] Specifically, the target feature can be defined as the gain importance score corresponding to each piece of information reaching a first preset value; or the target feature can be defined as the coverage importance score corresponding to each piece of information reaching a second preset value; or the target feature can be defined as the frequency score corresponding to each piece of information reaching a third preset value. The specific values ​​of the first, second, and third preset values ​​can be set by those skilled in the art according to actual needs.

[0085] S23. Calculate the first receiver operating characteristic curve and Youden index corresponding to each target feature.

[0086] The first subject operating characteristic curve is a ROC curve.

[0087] The first Receiver Operating Characteristic (ROC) curve is formed by plotting the first true positive rate (TPR) and the first false positive rate (FPR) of the original extreme gradient boosting prediction model under different classification thresholds. The first false positive rate is plotted on the x-axis, and the first true positive rate is plotted on the y-axis. The values ​​of the first true positive rate and the first false positive rate under different classification thresholds are then connected to form the first ROC curve. Further, step S23 includes:

[0088] S231. Use the roc_curve function to calculate the first subject operating characteristic curve corresponding to each target feature.

[0089] The roc_curve function is a function in the scikit-learn library used to calculate the parameters of the receiver operating characteristic curve (ROC curve).

[0090] S232. Use the auc function to calculate the Youden index corresponding to each target feature.

[0091] The auc function is used to calculate the area under the curve (AUC) of the first subject operating characteristic curve.

[0092] S24. Based on the first receiver operating characteristic curve and Youden index corresponding to each target feature, obtain the evaluation result corresponding to each target feature.

[0093] Specifically, the first receiver operating characteristic curve and Youden index corresponding to each target feature are compared with preset rules, and the comparison results are used as the evaluation results for each target feature.

[0094] The preset rule is a pre-defined mapping relationship between the first subject's operating characteristic curve and the Youden index and the corresponding assessment results. In other embodiments, the preset rule can be set by those skilled in the art according to actual needs.

[0095] In some implementations, when the risk probability prediction model is the target extreme gradient boosting prediction model, the clinical data evaluation method further includes a training step for the risk probability prediction model, which includes:

[0096] S3. Obtain the training dataset and the real label set, wherein each training data in the training dataset has a one-to-one correspondence with one of the real labels in the real label set.

[0097] refer to Figure 2 As shown, in this embodiment of the application, the raw data of 1046 systemic lupus erythematosus (SLE) patients from two medical centers is used as the training dataset. The raw data of 846 SLE patients are randomly divided into a training set (n1 = 588) and a test set (n2 = 258) in a 7:3 ratio. The raw data of the remaining 200 SLE patients are used as the validation set (n3 = 200). Furthermore, based on the primary outcome variable of increased treatment, the training set is divided into two groups: an increased treatment group (n3 = 451) and a non-increased treatment group (n4 = 137).

[0098] Furthermore, after obtaining the training dataset, it can be standardized to ensure all features have the same scale. The original extreme gradient boosting prediction model (i.e., the XGBoost model) is trained using the training dataset, and the importance score of each feature is obtained through the model's `feature_importances` attribute. The weight (importance) of a feature variable (i.e., each data point in the training set) can be reflected by the average gain (Gain) that the feature brings to the model output. This Gain is calculated by dividing the sum of gains from each split by the number of times the feature is used as a splitting node. The larger the Gain value of a feature, the more information gain it provides in each split in the model, and thus the greater its contribution to the model. Based on the original extreme gradient boosting prediction model, 24 feature variables were selected, including those related to the skin and mucous membranes, muscles and joints, kidneys, lungs, digestive system, blood, and nervous system, to support clinical decisions regarding increased treatment for SLE patients. The top 10 most weighted characteristic variables are PRO24h (i.e., quantitative urine protein), joint pain, rash, SLE (i.e., duration of SLE), PLTs (platelet count), HB (hemoglobin), ESR (erythrocyte sedimentation rate), A / G (albumin to globulin ratio), hypocomplementemia, and interstitial pneumonia (see reference). Figure 4 (As shown).

[0099] Based on Composite Feature Importance Bar Chart (CFIBP) (Reference) Figure 5 As shown in the figure, skin rash exhibited the strongest composite feature importance. Other features, such as PRO24h (i.e., quantified urinary protein), arthralgia, neuropsychiatric involvement, interstitial pneumonia, ascites, SLE duration, mesenteric vasculitis, hypocomplementemia, and anti-dsDNA, showed strong composite effects. Based on the selected feature variables, a target extreme gradient boosting prediction model was constructed to predict increased treatment outcomes in SLE patients.

[0100] In other embodiments, the selected feature variables are also used to construct the target random forest model, target support vector machine model, target multilayer perceptron, and target minimum absolute shrinkage and selection operator regression model.

[0101] In some implementations, R software can also be used to correct or remove outliers in the training dataset.

[0102] In some implementations, R software can also be used to remove duplicate values ​​from the training dataset.

[0103] In some implementations, R software can also be used to handle missing values ​​in the training dataset.

[0104] In some implementations, R software can also be used to remove variables in the training dataset that are missing more than 30% of the observations, in order to ensure the accuracy of the training of the risk probability prediction model.

[0105] S4. Input the target training data into the original extreme gradient boosting prediction model for evaluation and obtain the prediction result. The target training data is any training data in the training dataset.

[0106] Furthermore, step S4 includes:

[0107] S41. Calculate the gain importance training score, coverage importance training score, and frequency training score corresponding to each piece of information in the target training data.

[0108] S42. Based on the gain importance training score, coverage importance training score and frequency training score corresponding to each piece of information in the target training data, at least one target training feature is selected.

[0109] S43. Calculate the third-response operating characteristic curve and training Youden index corresponding to each target training feature.

[0110] The third receiver operating characteristic curve is formed by plotting the second true positive rate (TPR) and the second false positive rate (FPR) of the original extreme gradient boosting prediction model under different classification thresholds. The second false positive rate is plotted on the x-axis and the second true positive rate is plotted on the y-axis. The values ​​of the second true positive rate and the second false positive rate under different classification thresholds are plotted on the second coordinate graph, and the points on the second coordinate graph are connected to form the third receiver operating characteristic curve.

[0111] S44. Based on the third receiver operating characteristic curve and training Youden index corresponding to each target training feature, obtain the prediction result corresponding to each target training feature.

[0112] S5. Use a calibration curve to perform size matching processing on the prediction result and the real label corresponding to the target training data to obtain the matching result.

[0113] In this embodiment, R software is used to plot the calibration curve. When plotting the calibration curve, the original extreme gradient boosting prediction model outputs the probability that each sample in the test set is positive. These probabilities are divided into several equally wide buckets, with similar predicted probabilities for samples within each bucket. Then, the average predicted probability (i.e., the horizontal axis) and the actual proportion of positive samples (i.e., the vertical axis) are calculated for each bucket, and a scatter plot is plotted and connected to form the calibration curve. Ideally, the calibration curve should be close to a straight line with a slope of 1, indicating that the predicted probability of the original extreme gradient boosting prediction model is completely consistent with the actual label. The Brier score (or Brier score) between the calibration curve and the ideal calibration curve can be calculated; the smaller the Brier score, the better the calibration performance of the original extreme gradient boosting prediction model.

[0114] S6. Adjust the first model parameter of the original extreme gradient boosting prediction model according to the matching result to obtain the risk probability prediction model.

[0115] In other embodiments, when the risk probability prediction model is a target minimum absolute contraction and selected operator regression model, the selection of model feature variables incorporates standard features from conventional disease activity assessment tools, while also including non-classical features lacking in SLEDAI-2K (Systemic Lupus Erythematosus Disease Activity Index 2000), BILAG (British Lupus Assessment Group), and SLE-DAS (Systemic Lupus Erythematosus Disease Activity Score) but with clinical predictive value, such as serum albumin / globulin ratio and anti-U1snRNP antibody. To ensure data accuracy and consistency, the weights of the feature variables are normalized, and then features are selected using LASSO, followed by iterative analysis using cross-validation (see reference). Figure 3 (As shown). LASSO feature selection primarily achieves feature selection by adding an L1 regularization term to the regression model. The model sets a regularization parameter α, which controls the strength of L1 regularization. Cross-validation is used to evaluate the model's performance under different α values, and the optimal α value is selected. Using the selected α value, the original minimum absolute shrinkage and selection operator regression model is trained on the training set. The original minimum absolute shrinkage and selection operator regression model automatically adjusts the feature coefficients according to the L1 regularization term, compressing unimportant feature coefficients to 0, while features with coefficients of 0 are excluded from the model. LASSO feature selection aims to screen out clinically significant feature variables that predict increased treatment for SLE patients, selecting key features from a large number of features, reducing the complexity of the prediction model, lowering the possibility of overfitting, and improving the model's generalization ability on new data.

[0116] Furthermore, feature selection incorporated standard features from the routine disease activity assessment tools SLEDAI-2K, BILAG, and SLE-DAS, including important clinical features such as fever, joint swelling and pain, and rash, as well as common laboratory tests such as hemoglobin, white blood cell count, platelet count, complement, and 24-hour proteinuria. It also included non-classical features that are lacking in SLEDAI-2K, BILAG, and SLE-DAS but have clinical predictive value, such as serum albumin / globulin ratio, neutrophil percentage, and anti-U1snRNP antibody.

[0117] Furthermore, the target extreme gradient boosting prediction model was compared with SLEDAI-2K, SLE-DAS, and BILAG. In the training, test, and validation sets, the target extreme gradient boosting prediction model exhibited the highest AUC values, at 0.938, 0.863, and 0.849, respectively. In the training set, the target extreme gradient boosting prediction model effectively predicted increased treatment outcomes in SLE patients, with sensitivity and specificity of 0.978 and 0.897, respectively, both higher than SLE-DAS and BILAG. Figure 6 As shown. In the test set, the target extreme gradient boosting prediction model demonstrated higher sensitivity (0.912) and specificity (0.827) than SLEDAI-2K in identifying patients requiring increased treatment, according to reference. Figure 7 As shown. Furthermore, the target extreme gradient boosting prediction model exhibits higher sensitivity while maintaining specificity similar to BILAG. Although the sensitivity of the target extreme gradient boosting prediction model is lower than SLE-DAS, its specificity is higher. In the validation set, the target extreme gradient boosting prediction model showed the highest sensitivity and negative prediction rate, while its specificity and positive prediction rate were slightly lower than the other three models. (Refer to...) Figure 8 As shown, the overall prediction results of NPV, PPV, sensitivity, and specificity indicate that the target extreme gradient enhancement prediction model predicts the best efficacy for increasing treatment effectiveness.

[0118] In some implementations, the detection capabilities of the target extreme gradient boosting prediction model are as follows: for the training set, the sensitivity is 0.978, specificity (i.e., uniqueness) is 0.898, positive predictive value (PPV) is 0.969, and negative predictive value (NPV) is 0.925; for the test set, the sensitivity is 0.913, specificity is 0.813, PPV is 0.923, and NPV is 0.792; for the validation set, the sensitivity is 0.980, specificity is 0.717, PPV is 0.920, and NPV is 0.916. The calibration curve shows good consistency in predicting increased treatment in SLE patients, and is essentially consistent with the ideal calibration model. The optimal model (i.e., the target extreme gradient boosting prediction model) obtains low Brill scores (0.040, 0.094, and 0.084, respectively) in the training, test, and validation sets, indicating that the target extreme gradient boosting prediction model is well calibrated. Furthermore, clinical decision curves demonstrate that the target extreme gradient boosting prediction model is an excellent predictive tool. The performance of the target extreme gradient boosting prediction model on the training set, validation set, and external validation set is shown in [Figure 1]. Figures 9-17 .

[0119] In some implementations, after adjusting the first model parameters of the original extreme gradient boosting prediction model based on the matching result to obtain the risk probability prediction model, the training step of the risk probability prediction model further includes:

[0120] S7. The risk probability prediction model is evaluated using the area enclosed by the second receiver operating characteristic curve and the horizontal axis, so as to adjust the second model parameters of the risk probability prediction model.

[0121] In this embodiment, R software is used to plot the first, second, and third receiver operating characteristic (ROC) curves. Using the second ROC curve to evaluate the risk probability prediction model ensures that the machine learning output has good specificity and sensitivity.

[0122] The second receiver operating characteristic curve is formed by plotting the third true positive rate (TPR) and the third false positive rate (FPR) of the risk probability prediction model under different classification thresholds. The third false positive rate is plotted on the x-axis and the third true positive rate is plotted on the y-axis. The values ​​of the third true positive rate and the third false positive rate under different classification thresholds are plotted on the third coordinate graph, and the points on the third coordinate graph are connected to form the second receiver operating characteristic curve.

[0123] In some implementations, after adjusting the first model parameters of the original extreme gradient boosting prediction model based on the matching result to obtain the risk probability prediction model, the training step of the risk probability prediction model further includes:

[0124] S8. The risk probability prediction model is evaluated using a clinical decision curve to adjust the third model parameter of the risk probability prediction model.

[0125] This application embodiment uses R software to plot the clinical decision curve. When plotting the clinical decision curve, at different thresholds, the consistency ratio and inconsistency ratio between the risk probability prediction model's predictions and the actual events are calculated, and then the net benefit is calculated. The decision curve is plotted with the probability thresholds predicted by the risk probability prediction model as the horizontal axis and the net benefit as the vertical axis. The higher the curve, the stronger the predictive ability of the risk probability prediction model.

[0126] Using the clinical decision curve (i.e., DCA curve) to evaluate the risk probability prediction model can ensure that the results judged by the risk probability prediction model have clinical benefit.

[0127] refer to Figure 18 The diagram shown is a schematic block diagram of a clinical data evaluation device provided in the second aspect of an embodiment of this application. Figure 18 The clinical data evaluation device 100 includes:

[0128] Module 101 is used to acquire clinical data of the target user;

[0129] The assessment module 102 is used to input the clinical data into the risk probability prediction model for increased treatment and to obtain the assessment result of the clinical data. The assessment result is used to reflect the risk probability that the target user's systemic lupus erythematosus requires increased treatment.

[0130] A third aspect of this application provides a terminal device, the schematic diagram of which is as follows: Figure 19As shown. The terminal device includes a processor, memory, network interface, display screen, and temperature sensor connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface of the terminal device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements methods for evaluating clinical data. The display screen can be a liquid crystal display (LCD) or an e-ink display. The temperature sensor is pre-installed inside the terminal device to detect the operating temperature of the internal components.

[0131] Those skilled in the art will understand that Figure 19 The schematic diagram shown is only a partial structural diagram related to the present invention and does not constitute a limitation on the terminal device to which the present invention is applied. The specific terminal device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0132] In some embodiments, this application provides a terminal device including a processor and a memory. The memory stores a computer program, and the processor calls and runs the computer program stored in the memory to perform the steps of the clinical data evaluation method provided in the first aspect of this application. In a fourth aspect, this application provides a computer-readable storage medium for storing a computer program that causes a computer to perform the steps of the clinical data evaluation method provided in the first aspect of this application.

[0133] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0134] The technical features of the above embodiments can be combined without changing the basic principles of this application. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0135] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the patent protection scope of this application should be determined by the appended claims.

Claims

1. A method for evaluating clinical data, characterized in that, include: Obtain clinical data from the target users; The clinical data is input into a risk probability prediction model for increased treatment and evaluated to obtain an evaluation result of the clinical data. The evaluation result is used to reflect the risk probability that the target user's systemic lupus erythematosus requires increased treatment.

2. The method for evaluating clinical data according to claim 1, characterized in that, The clinical data includes information on fever, rash, hair loss, pleural effusion, interstitial pneumonia, neuropsychiatric involvement, hematuria, mesenteric vasculitis, white blood cell count, hypocomplementemia, anti-double-stranded DNA antibody, anti-nucleosome antibody, anti-nuclear antibody, anti-U1 small nucleotide protein antibody, duration of illness, joint pain, quantification of urine protein, hemoglobin, platelet count, erythrocyte sedimentation rate, neutrophil percentage, albumin / globulin ratio, immunoglobulin, and C-reactive protein.

3. The method for evaluating clinical data according to claim 1, characterized in that, The risk probability prediction model is selected from any one of the following models: target extreme gradient boosting prediction model, target random forest model, target support vector machine model, target multilayer perceptron, and target minimum absolute contraction and selection operator regression model.

4. The method for evaluating clinical data according to claim 1, characterized in that, In the case that the risk probability prediction model is a target extreme gradient enhancement prediction model, the clinical data is input into the risk probability prediction model for increasing treatment for evaluation, and the evaluation result of the clinical data is obtained. The step of using the evaluation result to reflect the risk probability that the target user's systemic lupus erythematosus requires increased treatment includes: Calculate the gain importance score, coverage importance score, and frequency score for each piece of information in the clinical data; At least one target feature is selected based on the gain importance score, coverage importance score, and frequency score corresponding to each piece of information in the clinical data; Calculate the first receiver operating characteristic curve and Youden index for each target feature; Based on the first receiver operating characteristic curve and Youden index corresponding to each target feature, the evaluation result corresponding to each target feature is obtained.

5. The method for evaluating clinical data according to claim 4, characterized in that, The clinical data evaluation method also includes a training step for the risk probability prediction model, which includes: Obtain a training dataset and a real label set, wherein each training data in the training dataset has a one-to-one correspondence with one of the real labels in the real label set; The target training data is input into the original extreme gradient boosting prediction model for evaluation to obtain the prediction result. The target training data is any one of the training data in the training dataset. A calibration curve is used to perform size matching processing on the predicted result and the real label corresponding to the target training data to obtain the matching result; Adjust the first model parameter of the original extreme gradient boosting prediction model based on the matching result to obtain the risk probability prediction model.

6. The method for evaluating clinical data according to claim 5, characterized in that, After adjusting the first model parameters of the original extreme gradient boosting prediction model based on the matching results to obtain the risk probability prediction model, the training step of the risk probability prediction model further includes: The risk probability prediction model is evaluated using the area enclosed by the second receiver operating characteristic curve and the horizontal axis, in order to adjust the second model parameters of the risk probability prediction model.

7. The method for evaluating clinical data according to claim 5, characterized in that, After adjusting the first model parameters of the original extreme gradient boosting prediction model based on the matching results to obtain the risk probability prediction model, the training step of the risk probability prediction model further includes: The risk probability prediction model was evaluated using a clinical decision curve to adjust the third model parameter of the risk probability prediction model.

8. A device for evaluating clinical data, characterized in that, include: The acquisition module is used to acquire clinical data from the target user. The assessment module is used to input the clinical data into a risk probability prediction model for increased treatment and to obtain an assessment result of the clinical data. The assessment result is used to reflect the risk probability that the target user's systemic lupus erythematosus requires increased treatment.

9. A terminal device, characterized in that, include: A processor and a memory for storing a computer program, the processor for calling and running the computer program stored in the memory to perform the steps of the clinical data evaluation method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Used to store computer programs that cause a computer to perform the steps of the clinical data evaluation method according to any one of claims 1 to 7.