Machine learning-based mild symptom SFTS risk prediction method and system

By using machine learning methods to screen key predictive factors and construct a dual-time-point Cox proportional hazards scoring model, the instability problem caused by multivariate collinearity in traditional methods is solved, achieving more accurate risk prediction and visualization of mild SFTS, and improving the accuracy and interpretability of the model.

CN120824008APending Publication Date: 2025-10-21NANJING DRUM TOWER HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510942791.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Traditional methods and systems for predicting the risk of mild SFTS (Small Scale for Disease) suffer from unstable model parameter estimation when faced with multivariate collinearity problems. This makes it difficult to accurately capture the dynamic changes in the disease progression process and thus fails to provide accurate risk predictions.

Method used

Using machine learning methods, we acquired and preprocessed clinical data on mild SFTS, used the LASSO-Cox regression model to screen key predictive factors, constructed a Cox proportional hazards scoring model with two time points, and generated a visual heatmap to achieve risk stratification and prediction.

Benefits of technology

The model improves data quality and stability, enabling it to more accurately capture dynamic changes during disease development, provide clearer risk prediction references, and enhance the model's interpretability and visualization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120824008A_ABST
    Figure CN120824008A_ABST
Patent Text Reader

Abstract

The invention discloses a light symptom SFTS risk prediction method and system based on machine learning, and the method comprises the steps: obtaining light symptom SFTS clinical data, carrying out the preprocessing, carrying out the variable screening of the preprocessed data, determining an optimal regularization parameter, obtaining a key prediction factor, building a Cox proportional risk scoring model of two time points based on the key prediction factor, and carrying out the calculation of the Cox proportional risk scoring model. Calculating individual risk scores and constructing a column graph, determining risk levels based on a risk layering threshold value, realizing risk prediction, finally generating a visual heat map according to the combination of the key prediction factors, marking critical disease risk probabilities under different key prediction factor combinations, and completing risk prediction of different key prediction factor combinations. The objective of the method for constructing the risk prediction model of the mild SFTS based on machine learning is to realize accurate risk prediction of early-stage SFTS critical disease progress through double-time-point modeling, dynamic risk layering and visualization of a heat map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical data analysis and machine learning, and specifically relates to a risk prediction method and system for mild SFTS based on machine learning. Background Art

[0002] Severe fever with thrombocytopenia syndrome (SFTS) is an acute infectious disease caused by the Dabie Mountain virus. Its main transmission routes include tick bites and, in rare cases, contact with infected body fluids. The mortality rate of patients ranges from 11.2% to 30%. SFTS can be clinically divided into mild and severe cases. Mild cases are manifested by fever, muscle aches, gastrointestinal discomfort, etc., while severe cases may have serious symptoms such as gastrointestinal bleeding, respiratory failure, and even multiple organ failure. The condition often deteriorates rapidly. For SFTS patients, early identification of the risk of mild patients progressing to critical illness is crucial for clinical intervention.

[0003] Traditional risk prediction methods and systems rely on univariate analysis or simple statistical models, and are helpless when faced with the problem of multivariate collinearity. Multivariate collinearity can make model parameter estimation unstable, leading to large deviations in prediction results. When processing complex time event data, it is difficult to accurately capture the dynamic changes in the disease development process and cannot provide accurate risk predictions for clinicians. Summary of the Invention

[0004] In response to the above problems, the purpose of the present invention is to provide a risk prediction method and system for mild SFTS based on machine learning.

[0005] The specific technical solutions for achieving the purpose of the present invention are as follows:

[0006] A risk prediction method for mild SFTS based on machine learning, comprising the following steps:

[0007] Step 1: Obtain clinical data of mild SFTS and preprocess the data;

[0008] Step 2: Screen the variables of the preprocessed data, determine the optimal regularization parameters, and obtain key predictive factors;

[0009] Step 3: Construct a two-time-point Cox proportional hazard score model based on key predictors, calculate individual risk scores, and construct a nomogram;

[0010] Step 4: Determine the risk level based on the risk stratification threshold to achieve risk prediction;

[0011] Step 5: Generate a visual heat map based on the combination of key predictors, mark the risk probability of critical illness under different key predictor combinations, and complete the risk prediction of different key predictor combinations.

[0012] A risk prediction system for mild SFTS based on machine learning, including the following modules:

[0013] Preprocessing module: used to obtain clinical data of mild SFTS and preprocess the data;

[0014] Key factor determination module: used to screen variables of preprocessed data, determine the optimal regularization parameters, and obtain key prediction factors;

[0015] Risk scoring module: used to construct a two-time-point Cox proportional risk score model based on key predictors, calculate individual risk scores and construct nomograms, and determine risk levels based on risk stratification thresholds;

[0016] Visualization module: used to generate a visualization heat map based on the combination of key predictive factors, marking the probability of critical illness risk under different key predictive factor combinations.

[0017] Compared with the prior art, the present invention has the following beneficial effects:

[0018] (1) The solution of the present invention collects clinical data of SFTS patients in the data preprocessing step and performs screening, missing value and outlier processing, which provides a high-quality data foundation for subsequent model construction;

[0019] (2) The solution of the present invention uses the LASSO-Cox regression model combined with 10-fold cross validation to determine the optimal regularization parameter in partial variable screening, which can effectively deal with the multivariate collinearity problem and obtain key predictive factors;

[0020] (3) In terms of model construction, the present invention establishes a dual-time-point Cox proportional hazard model based on key predictive factors, calculates individual risk scores and constructs a nomogram. The dual-time-point setting can better capture the dynamic changes in the disease development process. Risk stratification uses the machine learning tool X-tile software to determine the risk stratification thresholds at different days and divide the data into low, medium and high risk groups. This more accurate risk stratification method can enable users to more clearly understand the risk level of the data at different time points, providing a clearer reference for decision-making;

[0021] (4) When generating a heat map, the solution of the present invention generates a visual heat map based on the combination of key predictive factors, annotates the critical illness risk probability under different combinations, and enables users to understand the critical illness risk under different factor combinations more quickly through intuitive visualization.

[0022] The present invention will be further described below with reference to specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1This is a flow chart of the risk prediction method for mild SFTS based on machine learning of the present invention.

[0024] Figure 2 Schematic diagram of a nomogram in an embodiment of the present invention.

[0025] Figure 3 Schematic diagram of a visualized heat map in an embodiment of the present invention. DETAILED DESCRIPTION

[0026] Example

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. The described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts are within the scope of protection of the present invention.

[0028] As used in this application and the claims, the words "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural unless the context clearly indicates otherwise. Generally speaking, the terms "comprises" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.

[0029] Unless otherwise specifically stated, the relative arrangement of the parts and steps, numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present application. At the same time, it should be understood that, for ease of description, the sizes of the various parts shown in the drawings are not drawn according to actual proportional relationships. The techniques, methods and equipment known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the techniques, methods and equipment should be considered as part of the authorization specification. In all examples shown and discussed here, any specific values ​​should be interpreted as being merely exemplary and not as limitations. Therefore, other examples of the exemplary embodiments may have different values. It should be noted that similar numbers and letters represent similar items in the following figures, and therefore, once an item is defined in one figure, it does not need to be further discussed in subsequent figures.

[0030] Combine Figure 1 A risk prediction method for mild SFTS based on machine learning, comprising the following steps:

[0031] Step 1: Obtain clinical data on mild SFTS, screen early cases according to pre-set inclusion and exclusion criteria, and pre-process the data:

[0032] Specifically, the inclusion and exclusion criteria in this embodiment include:

[0033] Inclusion criteria: acute fever (body temperature>38.0℃) with thrombocytopenia (platelet count <100×10 9 / L), positive SFTS serum nucleic acid or IgM antibody test, time from onset to first blood test ≤ 5 days;

[0034] Exclusion criteria: age ≤ 18 years, concurrent infection with other pathogens, or underlying diseases such as malignant tumors, blood diseases, and pregnancy;

[0035] Afterwards, the obtained clinical data of mild SFTS were structured and processed using data cleaning;

[0036] Fill in missing values ​​in the data;

[0037] Detect and handle outliers in data;

[0038] Among them, the clinical data standards for mild SFTS obtained must strictly limit the conditions of acute fever, thrombocytopenia, and positive SFTS serum nucleic acid or IgM antibody test, and limit the disease characteristics, age, and comorbidities of the data to reduce the impact of other interfering factors on model prediction;

[0039] In this embodiment, during the data preprocessing process, structured processing was performed on 57 potential predictive factors (including demographic characteristics, laboratory indicators, and clinical manifestations), and data cleaning was used to ensure the quality of data input into the machine learning model. The inclusion criteria strictly limited conditions such as acute fever, thrombocytopenia, and positive SFTS serum nucleic acid or IgM antibody test to ensure that the research subjects were all mild SFTS patients with clear disease characteristics, avoiding other interfering factors, excluding patients aged ≤18 years and those with other diseases, reducing the impact of other underlying diseases on the research results, making the research data more targeted and homogeneous, and laying the foundation for subsequent precise analysis. Data cleaning further ensured the quality of data input into the machine learning model, reduced the interference of erroneous data on model training, and improved the accuracy of model predictions.

[0040] Clear inclusion and exclusion criteria can help clinicians accurately screen patients who meet the research conditions and improve research efficiency. For data cleaning, in addition to structured processing of the 57 known potential predictive factors, a data quality monitoring mechanism should be established to regularly check the integrity and accuracy of the data. For example, during the data collection stage, logical verification rules for data entry should be set to avoid entering erroneous data; during data storage, data consistency checks should be performed regularly to promptly detect and correct possible data errors to ensure the reliability of the data throughout the research process.

[0041] The missing values ​​in the data are filled using the Multiple Interpolation (MICE) method. For variables with a missing data ratio less than a certain value (for example, variables with a missing data ratio less than 20%, such as laboratory indicators such as AST and CRP), a chain equation model is constructed and iterative filling is performed using the correlation between variables.

[0042] When detecting and processing outliers in the data, a machine learning-based distribution detection method is combined with box plots and Z-scores to jointly detect outliers, or a density-based spatial clustering algorithm is used to detect outliers. Extreme values ​​of indicators such as AST and viral copy number are subjected to Winsorization and adjusted to the 1% and 99% quantiles. Multiple interpolation (MICE) is used to handle variables with a missing ratio of less than 20%. By constructing a chain equation model and iteratively filling in the gaps using the correlation between variables, data information can be retained to the greatest extent, avoiding sample size reduction and information loss due to missing data, enabling model training to be based on more complete data, and improving the stability and reliability of the model. A machine learning-based distribution detection method is combined with box plots and Z-scores to jointly detect outliers and perform Winsorization to effectively identify and adjust extreme values, avoid excessive influence of outliers on model results, ensure the accuracy of model parameter estimation, and improve the accuracy of model prediction.

[0043] When using the multiple imputation method, in order to ensure the reliability of the imputation effect, multiple imputations can be performed and the results can be compared. For example, 5-10 imputations can be performed for each variable with missing values, and the impact of different imputation results on model training and prediction can be analyzed. The most stable imputation result that conforms to actual clinical significance can be selected. For outlier detection and processing, in addition to using box plots and Z-score joint detection, methods such as density-based spatial clustering algorithms (DBSCAN) can also be introduced to identify outliers from different angles and improve the accuracy of outlier detection. At the same time, after the tailing process, the processed outlier information should be recorded for subsequent analysis of the causes and patterns of the occurrence of outliers.

[0044] Step 2: Use the LASSO-Cox regression model in machine learning to screen variables in the preprocessed data, determine the optimal regularization parameters, and obtain key predictive factors;

[0045] The LASSO-Cox regression model belongs to the regularized survival analysis method in machine learning. It compresses the coefficients of the latent variables in the data through the L1 norm penalty term. The objective function is:

[0046]

[0047] Where L(β) is the partial likelihood function of the Cox model, which measures the degree of fit of the model to the time-outcome data of "progression to critical illness" in patients with mild SFTS. After negative logarithmization, the smaller the value, the better the model fit. λ is the regularization parameter. is the parameter vector of Lasso regression, β j represents the coefficient corresponding to the j-th predictor;

[0048] The mean square error was calculated using 10-fold cross-validation to evaluate the predictive performance of the model under different regularization parameters λ. The 1-SE criterion was then used to select the optimal value from the candidate λ. Ultimately, the truly valuable key predictors of the "risk of mild SFTS progressing to critical illness" were screened out, laying the foundation for the subsequent construction of the Cox risk model and nomogram.

[0049] In this embodiment, the mean square error is calculated by 10-fold cross-validation, and the optimal λ = 0.127 is determined by the 1-SE criterion, so that the model achieves the optimal balance between bias and variance. Finally, four non-zero coefficient variables of viral copy number, AST, CRP, and neurological symptoms are selected as key predictors, effectively solving the problem of multivariate collinearity. The coefficients of 57 potential variables are compressed by the L1 norm penalty term, and the mean square error is calculated by 10-fold cross-validation and the optimal regularization parameter is determined by the 1-SE criterion. The truly key predictors, such as viral copy number, AST, CRP, and neurological symptoms, can be selected from many variables, so that the model achieves the optimal balance between bias and variance.

[0050] When applying the LASSO-Cox regression model, we can further explore the impact of different data sets and sample sizes on the model screening results. For example, we can collect data on SFTS patients in different regions and seasons, and analyze whether the key predictors screened by the LASSO-Cox regression model are consistent under different data characteristics. At the same time, we can try to combine the LASSO-Cox regression model with other feature selection methods (such as random forest feature importance ranking) to comprehensively evaluate the importance of variables and improve the accuracy of key predictor screening. In addition, for the determined optimal regularization parameter, we can perform a sensitivity analysis to study the impact of its fluctuations within a certain range on model performance to ensure the stability of the model.

[0051] Step 3: Construct a two-time-point Cox proportional hazard score model based on key predictors, calculate individual risk scores, and construct a nomogram;

[0052] In this embodiment, the dual time points of 7 days and 14 days were selected, and "progression to critical illness" was defined as the endpoint event. The stepwise backward selection method (AIC criterion) was used to optimize the model parameters. The final risk score formula was:

[0053]

[0054] Among them, β i The coefficients of key predictors (viral copy number ≥ 10 7 IU / mL was assigned a value of 1.784, AST ≥ 200 U / L was assigned a value of 0.651, CRP ≥ 8 mg / L was assigned a value of 0.566, and the presence of neurological symptoms was assigned a value of 1.044). i is the quantitative variable of each predictor, n represents the number of key predictors, and in this embodiment, n=4; "progression to critical illness" is defined as the endpoint event for the two time points of 7 days and 14 days, and the stepwise backward selection method (AIC criterion) is used to optimize the model parameters, which can more carefully analyze the risk factors of disease progression in patients at different time stages. Through the established risk scoring formula, key predictors such as viral copy number, AST, CRP, neurological symptoms and their quantitative variables are comprehensively considered to accurately calculate the individual risk score, effectively capture the dynamic changes in the disease development process, provide a more accurate basis for clinicians to assess the patient's disease risk at different time points, and make up for the shortcomings of traditional methods in time dimension analysis.

[0055] In the application of the two-time-point Cox proportional hazard model, the selection of time points can be further expanded. For example, time points such as 3 days and 21 days can be added to analyze the risk of disease progression in patients over a wider time range and construct a more complete disease development risk curve. At the same time, the regression coefficient in the risk score formula can be updated and optimized by collecting more clinical data to improve the accuracy of the risk score. In addition, combined with the patient's treatment intervention, the impact of different treatment methods on the risk score can be analyzed to provide a more targeted reference for the formulation of clinical treatment plans.

[0056] The Cox proportional hazard model at each time point was determined according to the set time nodes, thereby obtaining the Cox proportional hazard score at the two time points.

[0057] The construction of the nomogram is to convert the Cox regression results into graphical results using the R language rms package;

[0058] The nomogram includes three axes, including a key predictor axis, a score axis, and a risk probability axis;

[0059] Among them, the key predictor axis is used to characterize the key predictor content, the score axis is used to characterize the risk score value of the key predictor, and the risk probability axis is used to characterize the probability of the key factor progressing to a critical illness at a set time node.

[0060] In this embodiment, the viral copy number scale axis can be divided into 0-100 points (≤10 4IU / mL is 0 points, =10 5 IU / mL is 36 points, =10 6 IU / mL is 83 points, ≥10 7 IU / mL is 100 points);

[0061] The AST scale axis is divided into intervals of 0–36 points (<100 U / L is 0 points, 100–200 U / L is 3 points, and ≥200 U / L is 36 points);

[0062] The CRP scale axis is divided into 0-32 points according to the threshold value (<8 mg / L is 0 points, ≥8 mg / L is 32 points);

[0063] The neurological symptom scale is binary, ranging from 0 to 60 points (absence is 0 points, presence is 60 points).

[0064] By locating the indicator values ​​on the corresponding axis and projecting them vertically onto the "score axis," a total risk score (0-228 points) was calculated. The "risk axis" was then used to map the 7-day and 14-day probabilities of critical illness, achieving clinical interpretability of the machine learning model (using the R language rms package to convert the Cox regression results into a graphical tool, setting scale axes for key predictors such as viral copy number, AST, CRP, and neurological symptoms, and assigning corresponding scores based on different indicator values. The four scores were added together to calculate the total risk score, which was then mapped to the "risk axis" to obtain the 7-day and 14-day probabilities of critical illness).

[0065] like Figure 2 , which is a schematic diagram of the nomogram of the conversion in this embodiment;

[0066] Among them, the points column is the score axis, and the viral copy number, AST, CRP, and neurological symptoms are the key predictor axes. The risk scores of each key predictor of each case data are projected to obtain the score axis data corresponding to each key predictor;

[0067] The scores of each key predictor are then added together to obtain the total score, which is then mapped to the "risk axis" to obtain the probability of critical illness within 7 days and 14 days.

[0068] This intuitive graphical approach enables clinicians to quickly understand the model's prediction results, transforming complex machine learning models into tools that are easy to understand and apply, greatly improving the model's interpretability and facilitating the assessment of disease risks based on specific indicators.

[0069] In the design and application of nomograms, the division of the scale axis can be further optimized. For example, based on actual clinical experience and data distribution, the score increment interval of the viral copy number scale axis can be adjusted more finely to make it more in line with clinical judgment logic. At the same time, interactive functions can be added to the nomogram, such as integrating the nomogram into the electronic medical record system. When the doctor clicks on the specific value on the scale axis, a detailed explanation will pop up, including the clinical significance of the value, the relationship with disease progression, and other information. In addition, personalized nomograms can be customized according to the needs of different clinical departments, such as highlighting virus-related indicators for infectious disease doctors and emphasizing platelet-related indicators for hematologists, etc., to improve the practicality of the nomogram.

[0070] Step 4: Determine the risk level based on the risk stratification threshold to achieve risk prediction;

[0071] The risk stratification threshold was determined by maximizing the survival difference between groups using X-tile software.

[0072] In this example, for the 7-day outcome, the training set data were divided into low risk, medium risk, and high risk according to the total score, where low risk was ≤135 points, medium risk was 136 to 178 points, and high risk was >178 points;

[0073] The 14-day outcomes are divided into low risk, medium risk and high risk, where low risk is ≤103 points, medium risk is 104 to 164 points, and high risk is >164 points. The 7-day and 14-day risk stratification thresholds are determined by maximizing the survival difference between the groups, and the data are divided into low, medium and high-risk groups. This stratification method is based on the analysis of a large amount of data and can more accurately reflect the degree of disease risk at different time points. Doctors can take targeted treatment and monitoring measures according to the risk group of the data, such as strengthening monitoring and active treatment for high-risk patients, and conducting routine treatment and follow-up for low-risk patients, which improves the targetedness and effectiveness and overcomes the problem of inaccurate risk stratification by traditional methods.

[0074] When using X-tile software for risk stratification, the stratification results can be further optimized based on actual clinical conditions. For example, individual differences, such as the impact of factors such as age and underlying health status on risk stratification, can be considered, and appropriate adjustments can be made based on the existing stratification thresholds. At the same time, regular review and analysis of the treatment effects and prognosis of different risk groups can be conducted, and the risk stratification thresholds can be dynamically updated based on feedback results to ensure the accuracy and timeliness of risk stratification. In addition, the risk stratification results of the X-tile software can be combined with other clinical assessment tools, such as the Acute Physiology and Chronic Health Evaluation System (APACHE) for comprehensive assessment of the condition.

[0075] Step 5: Generate a visual heat map based on the combination of key predictors, mark the critical illness risk probability under different key predictor combinations, and complete the risk prediction for different key predictor combinations:

[0076] The determined key predictors are freely combined, and Cox regression is used to calculate the critical illness risk probability of the combined key predictors at the set time points;

[0077] The combination of key predictive factors and the corresponding critical illness risk probability of the combination are intuitively displayed using progressive color scales to form a visual heat map.

[0078] The visualized heat map in this embodiment is a 12×2 matrix (4 levels of viral copy number × 3 intervals of AST × 2 levels of CRP × 2 states of neurological symptoms = 48 combinations), as shown in FIG. Figure 3 As shown;

[0079] The row variables in the matrix are the predictor combinations, and the column variables are the 7-day and 14-day critical illness risk probabilities. Each cell is labeled with the 7-day risk (e.g., 0.607) on the left and the 14-day risk (e.g., 0.861) on the right.

[0080] A progressive color scale (green represents low risk, yellow represents medium risk, and red represents high risk) is used to intuitively display the risk gradient. Underlined cells indicate inconsistent risk stratification at two time points. A 12×2 matrix is ​​constructed to display the 7-day and 14-day critical illness risk probabilities under 48 combinations of predictive factors. A progressive color scale is used to intuitively display the risk gradient, allowing doctors to quickly identify the risk level under different combinations. The risk probability values ​​in the heat map are calculated using Cox regression of the training set data. The color coding system is based on the cluster analysis results of machine learning, which enhances the intuitiveness of risk stratification. Inconsistent risks at two time points are marked with underlines to avoid the limitations of relying solely on single time point assessments, providing users with more comprehensive and intuitive disease risk information.

[0081] The present invention also provides a risk prediction system for mild SFTS based on machine learning, comprising the following modules:

[0082] Preprocessing module: used to obtain clinical data of mild SFTS and preprocess the data;

[0083] Key factor determination module: used to screen variables of preprocessed data, determine the optimal regularization parameters, and obtain key prediction factors;

[0084] Risk scoring module: used to construct a two-time-point Cox proportional risk score model based on key predictors, calculate individual risk scores and construct nomograms, and determine risk levels based on risk stratification thresholds;

[0085] Visualization module: used to generate a visualization heat map based on the combination of key predictive factors, marking the probability of critical illness risk under different key predictive factor combinations.

[0086] The present invention further provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the processor implements the following steps:

[0087] Step 1: Obtain clinical data of mild SFTS and preprocess the data;

[0088] Step 2: Screen the variables of the preprocessed data, determine the optimal regularization parameters, and obtain key predictive factors;

[0089] Step 3: Construct a two-time-point Cox proportional hazard score model based on key predictors, calculate individual risk scores, and construct a nomogram;

[0090] Step 4: Determine the risk level based on the risk stratification threshold to achieve risk prediction;

[0091] Step 5: Generate a visual heat map based on the combination of key predictors, mark the risk probability of critical illness under different key predictor combinations, and complete the risk prediction of different key predictor combinations.

[0092] The above-described embodiment merely represents one embodiment of the present application. While the description is relatively specific and detailed, it should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A risk prediction method for mild SFTS based on machine learning, characterized in that: The following steps are involved: Step 1: Obtain clinical data of mild SFTS and preprocess the data; Step 2: Screen the variables of the preprocessed data, determine the optimal regularization parameters, and obtain key predictive factors; Step 3: Construct a two-time-point Cox proportional hazard score model based on key predictors, calculate individual risk scores, and construct a nomogram; Step 4: Determine the risk level based on the risk stratification threshold to achieve risk prediction; Step 5: Generate a visual heat map based on the combination of key predictors, mark the risk probability of critical illness under different key predictor combinations, and complete the risk prediction of different key predictor combinations.

2. A risk prediction method for mild SFTS based on machine learning according to claim 1, characterized in that: The step 1 of obtaining the clinical data of mild SFTS and preprocessing it is as follows: First, the obtained clinical data of mild SFTS were structured and processed using data cleaning; Fill in missing values ​​in the data; Detect and handle outliers in data.

3. A risk prediction method for mild SFTS based on machine learning according to claim 2, characterized in that, The clinical data standards for mild SFTS obtained must strictly limit the conditions of acute fever, thrombocytopenia, and positive SFTS serum nucleic acid or IgM antibody test, and limit the disease characteristics, age, and comorbidities of the data to reduce the impact of other interfering factors on model prediction; The missing values ​​in the data are filled using a multiple interpolation method, and for variables with a data missing ratio less than a certain value, a chain equation model is constructed and iterative filling is performed using the correlation between variables; When detecting and processing outliers in the data, a distribution detection method based on machine learning is combined with box plots and Z-score to jointly detect outliers, or a density-based spatial clustering algorithm is used to detect outliers, and Winsorization processing is performed to effectively identify and adjust extreme values ​​and avoid excessive impact of outliers on model results.

4. The method for predicting the risk of mild SFTS based on machine learning according to claim 1, characterized in that: The key predictive factors in step 2 are specifically obtained as follows: The coefficients of latent variables in the data are compressed by L1 norm penalty Where L(β) is the partial likelihood function of the Cox model, λ is the regularization parameter, is the parameter vector of Lasso regression, β j represents the coefficient corresponding to the j-th predictor; The mean square error was calculated using 10-fold cross validation, and the 1-SE criterion was used to determine the optimal regularization parameter λ, ultimately screening out key predictors of the risk of progression of mild SFTS to critical illness.

5. The risk prediction method for mild SFTS based on machine learning according to claim 1, characterized in that: The Cox proportional hazard model in step 3 is specifically: Among them, β i represents the coefficient of the key predictor, X i is the quantitative variable of each predictor, and n represents the number of key predictors; The Cox proportional hazard model at each time point was determined according to the set time nodes, thereby obtaining the Cox proportional hazard score at the two time points.

6. The risk prediction method for mild SFTS based on machine learning according to claim 1, characterized in that: The construction of the nomogram in step 3 is to convert the Cox regression results into graphical results using the R language rms package; The nomogram includes three axes, including a key predictor axis, a score axis, and a risk probability axis; Among them, the key predictor axis is used to characterize the key predictor content, the score axis is used to characterize the risk score value of the key predictor, and the risk probability axis is used to characterize the probability of the key factor progressing to a critical illness at a set time node.

7. The risk prediction method for mild SFTS based on machine learning according to claim 1, characterized in that: The risk stratification threshold in step 4 was determined by maximizing the survival difference between groups using X-tile software.

8. The risk prediction method for mild SFTS based on machine learning according to claim 1, characterized in that: The step 5 generates a visualization heat map based on the combination of key predictors, specifically: The determined key predictors are freely combined, and Cox regression is used to calculate the critical illness risk probability of the combined key predictors at the set time points; The combination of key predictive factors and the corresponding critical illness risk probability of the combination are intuitively displayed using progressive color scales to form a visual heat map.

9. A risk prediction system for mild SFTS based on machine learning, characterized in that: Includes the following modules: Preprocessing module: used to obtain clinical data of mild SFTS and preprocess the data; Key factor determination module: used to screen variables of preprocessed data, determine the optimal regularization parameters, and obtain key prediction factors; Risk scoring module: used to construct a two-time-point Cox proportional risk score model based on key predictors, calculate individual risk scores and construct nomograms, and determine risk levels based on risk stratification thresholds; Visualization module: used to generate a visualization heat map based on the combination of key predictive factors, marking the probability of critical illness risk under different key predictive factor combinations.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.