Method for predicting infection around artificial joint prosthesis after rheumatoid arthritis operation
Key features were screened through Boruta algorithm and genetic algorithm, combined with SMOTEENN algorithm for data balance, and an integrated machine learning model was constructed, which solved the problem of insufficient accuracy of infection risk prediction after artificial joint replacement in patients with rheumatoid arthritis, achieved more accurate PJI risk prediction, optimized preoperative management, and improved the patient's joint function and quality of life.
Patent Information
- Application Number
- CN202510239210.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-03-03
AI Technical Summary
In the prior art, the prediction of infection risk after artificial joint replacement in patients with rheumatoid arthritis is insufficient and practical, especially the evaluation of PJI is not accurate enough, which affects the patient's joint function and quality of life.
Boruta algorithm is used for feature selection, combining genetic algorithms and SMOTEENN algorithms, and integrated machine learning algorithms are developed. Through combination training of multiple machine learning models, an infection risk prediction model is built, key features are screened out and data balanced, and finally an accurate PJI risk prediction tool is formed.
It improves the accuracy of predicting postoperative PJI risks in patients with rheumatoid arthritis, provides clinicians with more powerful decision support, optimizes preoperative management, effectively prevents the occurrence of PJI, and improves patients' joint function and quality of life.
Smart Images

Figure CN120340887A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical technologies, and particularly relates to a method for predicting periprosthetic joint infection after artificial joint prosthesis for rheumatoid arthritis surgery. Background Art
[0002] Periprosthetic joint infection (abbreviated as PJI) after artificial joint replacement is a serious complication, which will have a significant impact on the joint function and quality of life of patients. With the popularization of joint replacement surgery, the incidence of PJI has also increased, which not only increases the economic burden on patients, but also puts higher requirements on medical resources. Especially for patients with rheumatoid arthritis, they face a higher risk of PJI during artificial joint replacement surgery.
[0003] The risk of PJI in artificial joint replacement surgery for patients with rheumatoid arthritis is about 1.6 times that of patients with osteoarthritis. This may be because of the immunosuppressive therapies received by patients with rheumatoid arthritis, including DMARDs and corticosteroids. Among them, 46% of RA patients are receiving biologic DMARDs treatment, 67% of RA patients are receiving non-biologic DMARDs treatment, 25% of RA patients are receiving glucocorticoid treatment, and RA patients are more likely to have postoperative anemia, and are more likely to require blood transfusion due to bone marrow suppression caused by chronic diseases or drug use, while blood transfusion may increase the risk of infection. In addition, the fragile soft tissues around the knee joints of RA patients may make joint replacement surgery more prone to infection. Therefore, the assessment of the PJI risk for this group of patients with rheumatoid arthritis is particularly crucial.
[0004] The occurrence of PJI is the result of the imbalance between bacterial contamination and host resistance. Although it is difficult to quantify the host resistance, studies have shown that a variety of factors, including gender, body mass index (BMI), smoking behavior, diabetes, steroid drug use, immunosuppressive status, etc., are all closely related to the PJI risk. These factors have important reference values in clinical assessment. In order to make the risk assessment more accurate and practical, it is extremely important to integrate these factors into a comprehensive risk prediction system.
[0005] In recent years, risk prediction systems based on literature reviews and case analyses have emerged, but they have deficiencies in terms of accuracy and practicality and rely heavily on artificial experience. Summary of the Invention
[0006] Aiming at the problems of insufficient accuracy and poor practicality in predicting the risk of postoperative infection in the prior art, the purpose of the present invention is to provide a method for predicting periprosthetic joint infection after artificial joint prosthesis for rheumatoid arthritis surgery, so as to at least partially solve the above problems.
[0007] To achieve the above purpose, the technical solution of the present invention is as follows: A method for predicting periprosthetic joint infection after rheumatoid arthritis surgery, the method comprising the following steps: Collect the original data, then perform feature screening to obtain a data set, including a training set and a validation set; Perform combined training on multiple machine learning models through the training set, and then determine the best combination of machine learning models based on the genetic algorithm for the ensemble algorithm to obtain an infection risk prediction model; Perform postoperative infection risk prediction through the infection risk prediction model.
[0008] In some preferred embodiments, the original data is first preprocessed to obtain preprocessed data, and then the preprocessed data is subjected to feature screening; the content of preprocessing the original data includes one or more of deletion, cleaning, feature encoding, error correction, unit unification, and imputation.
[0009] In some preferred embodiments, feature screening is performed through the following steps: Perform multiple bootstrap resamplings on the data, each bootstrap resampling generates different data subsets, and use the data subsets to construct random forests; In each random forest, use the variable importance index to calculate the importance of each original feature; Introduce shadow features, which are generated by randomizing the original features, and the shadow features are used to simulate random features and compare with the original features; Compare the importance of each original feature with its corresponding shadow feature. If the importance of an original feature is significantly higher than its shadow feature, then the original feature is considered to have importance; Iterative process: Delete all shadow features, repeat the above steps until all original features are marked as "important" or "unimportant" for feature screening based on importance ranking.
[0010] In some preferred embodiments, the data set is split into a training set and a validation set based on K-fold cross-validation, including: Using K-fold cross-validation, randomly divide the data set into several folds of substantially equal size, and sequentially use each fold as the validation set and the remaining folds as the training set.
[0011] In some preferred embodiments, before model training, the method further includes the following steps: Balance the data of the training set through the SMOTEENN algorithm so that the ratio between the infected samples and the non-infected samples in the training set meets the preset range.
[0012] In some preferred embodiments, the step of balancing the training set by the SMOTEENN algorithm includes: SMOTE starts ① Select random data from the minority class; ② Calculate the distance between the random data and its K nearest neighbors; ③ Multiply the difference in the number of samples by a random number between 0 and 1, and then add the result to the minority class as a synthetic sample; ④ Repeat steps ② and ③ until the required proportion of the minority group is met; SMOTE ends, ENN starts ⑤ Determine K, the number of nearest neighbors. If not determined, then K = 3; ⑥ Find the K-nearest neighbors of the observation among other observations in the training set, and then return the majority class from the K-nearest neighbors; ⑦ If the class of the observation is different from the majority class of the K-nearest neighbors of the observation, then delete the observation and its K-nearest neighbors from the training set; ⑧ Repeat steps ⑤ and ⑥ until the ratio between the infected samples and the non-infected samples meets the preset range.
[0013] In some preferred embodiments, the step of combining and training multiple machine learning models and then using a genetic algorithm to determine the best combination based on an ensemble algorithm includes: ① Population initialization: Create a population containing N samples by random generation. The population is a sub-model combination including at least one machine learning model; ② Fitness calculation: The sub-model combination in step ① uses the method of ensemble learning to be verified by the validation set to obtain the fitness score of the sub-model combination; ③ Selection operation: According to the fitness of individuals in the population, select individuals with high fitness scores from the current population by roulette wheel method, simulate the process of natural selection, calculate the sum of fitness of all individuals Σfi, calculate the relative fitness of each individual fi / Σfi, similar to softmax, and then generate a random number between 0 and 1, and determine the number of times each individual is selected according to the position where the random number appears in the probability region; ④ Crossover operation: Use a certain crossover probability threshold to control whether to perform single-point crossover to generate new individuals; ⑤ Mutation operation: Randomly generate mutation points; ⑥ Iterate the above steps M times, terminate the iteration, and determine the best individual. The combination of the best individuals is the infection risk prediction model.
[0014] In some preferred embodiments, in step ②, the fitness score is calculated using the f3-score, where f3-score = 10 × precision × recall / (9 × precision + recall).
[0015] In some preferred embodiments, the best individuals are combined into a base learner, and the set model composed of a soft voting meta-learner retains the Model file. Then, interface encapsulation is performed, and postoperative infection risk prediction is carried out through the encapsulated interface.
[0016] In some preferred embodiments, the method is applied to the prediction of the risk of periprosthetic joint infection after rheumatoid arthritis surgery. The selected features include blood glucose, body weight, hemoglobin, AST, bilirubin, ESR, drug use, CRP, gender, and joint type.
[0017] Adopting the above technical solution, the beneficial effects of the present invention are as follows: The present invention utilizes AI technology to develop an innovative PJI risk prediction tool, which can more accurately predict the PJI risk after rheumatoid arthritis surgery, provide more powerful decision-making support for clinicians, optimize the preoperative management of joint replacement for rheumatoid arthritis patients, thereby effectively preventing the occurrence of PJI, and improving the joint function and quality of life of patients. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a schematic diagram of the method flow of the present invention.
[0019] Figure 2 It is a schematic diagram of the importance of BORUTA features in the present invention.
[0020] Figure 3 It is a schematic diagram of the genetic algorithm in the present invention.
[0021] Figure 4 It is a schematic diagram of the score comparison between PJI patients and non-PJI patients in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The following further describes the specific embodiments of the present invention with reference to the accompanying drawings. It should be noted here that the description of these embodiments is for helping to understand the present invention, but does not constitute a limitation to the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0023] Embodiment
[0024] A method for predicting periprosthetic joint infection after rheumatoid arthritis surgery, which is applied to the prediction of the risk of periprosthetic joint infection after rheumatoid arthritis surgery. The method includes the following steps: S1. Collect raw data.
[0025] Before collecting the original data, inclusion and exclusion criteria need to be determined. After obtaining the legal access to the mimic database, the inclusion criteria are as follows: ① Diagnosed with rheumatoid arthritis, ② History of joint replacement surgery, ③ Readmission after joint replacement surgery. The exclusion criteria are as follows: ① Readmission less than 4 weeks, ② Incomplete clinical data, ③ History of relevant joint infection. A total of 133 patient information were collected, including 22 patients with pji.
[0026] Furthermore, the specific patient information to be collected is determined, including patient serial number, primary disease diagnosis, surgery time, surgery type, readmission time, readmission diagnosis, age, gender, joint type, height, weight, BMI, medication history, white blood cell count, red blood cell count, platelet count, neutrophil count, lymphocyte count, hemoglobin, blood glucose, albumin, globulin, ALT, ALP, AST, total bilirubin, serum creatinine, blood urea, C-reactive protein (CRP), plasma prothrombin time (PT), activated partial thromboplastin time (APTT), erythrocyte sedimentation rate (ESR), immunoglobulin M (IgM), immunoglobulin A (IgA), immunoglobulin G (IgG), rheumatoid factor (RF), C3, C4.
[0027] Based on the above information, a number of original data are constructed.
[0028] S2. Preprocess the original data.
[0029] The preprocessing of the constructed original data includes one or more of deletion, cleaning, feature encoding, error correction, unit unification, and imputation. After preprocessing the original data, preprocessed data are obtained.
[0030] Data preprocessing is an important step, aiming to clean, transform and prepare data to better meet the requirements of machine learning algorithms.
[0031] Among them, "deletion" refers to deleting useless data. The useless data indicators include diagnostic data format indicators such as length of stay, hospital number, and diagnosis code, which can indicate the source of sample data, etc., but have no effect on machine learning prediction, so the corresponding data are deleted. At the same time, data with more than 50% missing features are also deleted.
[0032] Cleaning: refers to correcting incorrect data, including outliers and unifying data from multiple sources.
[0033] Feature encoding: refers to converting non-numerical features into numerical representations, such as gender and medication history, which are converted to be represented by 0 or 1.
[0034] Error correction: It refers to detecting and handling outliers in data, and outliers can be identified and processed through statistical methods or outlier detection algorithms.
[0035] Unit unification: It refers to unifying data into common units.
[0036] Interpolation: It refers to using the median method to impute missing values for missing continuous variables or categorical variables.
[0037] S3. Perform feature screening on the preprocessed data.
[0038] In step S1, the original data has feature information with many dimensions, but not all dimensional feature information has analytical value. Therefore, feature screening is required.
[0039] In this embodiment, the Boruta algorithm is used to screen out features with high correlation with the target variable. By comparing the original features with the "false features" generated in a random order, it is determined which specific features are truly important. The specific steps are as follows: ① Perform multiple bootstrap resamplings on the data. Each bootstrap resampling generates different data subsets, and a random forest is constructed using the data subsets.
[0040] ② In each random forest, use the variable importance index to calculate the importance of each original feature.
[0041] ③ Introduce shadow features. The shadow features are generated by randomizing the original features. The shadow features are used to simulate random features and compare with the original features.
[0042] ④ Compare the importance of each original feature with its corresponding shadow feature. If the importance of an original feature is significantly higher than that of its shadow feature, the original feature is considered important.
[0043] ⑤ Iterative process: Delete all shadow features and repeat the above steps until all original features are marked as "important" or "unimportant" for feature screening based on importance ranking.
[0044] In this embodiment, from the feature screening, the importance ranking of each index based on the Boruta algorithm is obtained, as Figure 2 shown; among them, the importance of blood glucose, body weight, hemoglobin, AST, bilirubin, ESR, drug use situation, and CRP is higher than that of the shadow variables. At the same time, basic features such as gender and joint type are included in the final prediction model.
[0045] After performing feature screening on the original data, a data set is formed.
[0046] S4. Split the data set to obtain a training set and a validation set.
[0047] In this embodiment, the data set is split into a training set and a validation set based on K-fold cross-validation, which specifically includes: Using K-fold cross-validation, randomly divide the data set into several (e.g., 5) folds of approximately equal size. Each fold is used as the validation set in turn, and the remaining folds are used as the training set.
[0048] S5. Balance the data of the training set.
[0049] Balance the data of the training set through the SMOTEENN algorithm so that the ratio between the infected samples and the non-infected samples in the training set conforms to a preset range, such as approximately 1:1. In this way, by improving the performance of the model on the minority class, the problem of overfitting is avoided.
[0050] In this embodiment, the steps of balancing the data of the training set through the SMOTEENN algorithm include: SMOTE start ① Select random data from the minority class; ② Calculate the distance between the random data and its K nearest neighbors; ③ Multiply the sample number difference (referring to the number difference between the minority samples and the majority samples) by a random number between 0 and 1, and then add the result to the minority class as a synthetic sample; ④ Repeat steps ② and ③ until the required ratio of the minority group is met; SMOTE end, ENN start ⑤ Determine K, the number of nearest neighbors. If not determined, then K = 3; ⑥ Find the K-nearest neighbors of the observation among other observations in the training set, and then return the majority class from the K-nearest neighbors; ⑦ If the class of the observation is different from the majority class of the K-nearest neighbors of the observation, then delete the observation and its K-nearest neighbors from the training set; ⑧ Repeat steps ⑤ and ⑥ until the ratio between the infected samples and the non-infected samples conforms to a preset range, such as approximately 1:1.
[0051] S6. Combine and train multiple machine learning models through the training set, and then determine the best combination of machine learning models based on the genetic algorithm for the ensemble algorithm to obtain an infection risk prediction model.
[0052] The machine learning models include the following 55 types: Logistic regression: Ridge regression, Lasso regression, Elastic Net regression, without regularization; Random forest: Hyperparameter selection ntree = 10, 15, 20, 30, depth = 20, 50, 100, 200; Extra tree: Split point = entropy and Gini Index, depth = 2, 3, 5, 10; Decision tree: Split point = entropy and Gini Index, depth = 2, 3, 5, 10; Support vector machine: Kernel function = polynomial kernel (poly), RBF kernel (rbf), Sigmoid kernel (sigmoid), linear kernel (linear); Multi-layer perceptron: Hidden layer = 20, 50, 100, 200, (50, 50), (100, 100); Gradient Boosting Decision Tree (GBDT): ntree = 20, 50, 100, 200, depth = 3, 5.
[0053] For the above 55 models, each model is trained on each of the 5 similarly sized training sets in step S4, iteratively run, and then validated using 5 validation sets respectively to end the training.
[0054] As Figure 3 shown, the specific steps to determine the best combination of machine learning models and obtain the infection risk prediction model include: ① Population initialization: Create a population containing N (e.g., 5000) samples through random generation. This population is a sub-model combination including at least one of the above machine learning models. Each machine learning model in the population is called an individual, and each individual is encoded as an array consisting of 0 or 1.
[0055] ② Fitness calculation: For the sub-model combination in step ①, use the method of ensemble learning to validate through the validation set to obtain the fitness score of the sub-model combination.
[0056] ③ Selection operation: According to the fitness of individuals in the population, select individuals with high fitness scores from the current population through the roulette wheel method, simulating the process of natural selection. Calculate the sum of fitness Σfi of all individuals, calculate the relative fitness size fi / Σfi of each individual, similar to softmax, and then generate a random number between 0 and 1. Determine the number of times each individual is selected based on the position where the random number appears in the probability region.
[0057] ④ Crossover operation: Use a certain crossover probability threshold to control whether to perform single-point crossover to generate new individuals.
[0058] ⑤ Mutation operation: Randomly generate mutation points.
[0059] ⑥Iterate the above steps M (e.g., 500) times, terminate the iteration, determine the best individuals, and decode these best individuals into specific combinations, which is the infection risk prediction model.
[0060] In step ②, specifically, the fitness score is calculated using f3-score, where f3-score = 10 × precision × recall / (9 × precision + recall).
[0061] The best sub-model combination determined in this step consists of the following 5 sub-models: logistic regression (lasso regression), extra trees (split point = entropy, depth = 10), extra trees (split point = entropy, depth = 3), support vector machine (kernel function = linear kernel), and multi-layer perceptron (hidden layer = 20).
[0062] After the model training is completed, combine the best individuals in step S6 into the base learner, and retain the Model file for the ensemble model composed of the soft voting as the meta-learner.
[0063] S8. Package the Model file in step S7, and the packaged website is "http: / / 8.222.247.162:3838 / RA / ". Perform the risk prediction of periprosthetic joint infection after rheumatoid arthritis surgery through the packaged interface.
[0064] When performing risk prediction, input the following through the packaged interface: gender, joint type, blood glucose, weight, hemoglobin, AST, bilirubin, ESR, drug use, CRP, as shown in Table 1 for example; the output of the interface is the risk score.
[0065] Table 1 - Input methods for 10 indicators Index Name Input Method Gender Qualitative Input (1 = male, 0 = female) Joint Type Qualitative Input (1 = hip joint, 0 = knee joint) Body Weight Quantitative Input (kg) Usage of Drugs Related to Rheumatoid Arthritis Option Input (none, glucocorticoid, methotrexate, others) C-reactive Protein Value Quantitative Input (mg / L) Fasting Blood Glucose Quantitative Input (mmol / L) AST Quantitative Input (U / L) Hemoglobin Quantitative Input (g / L) Bilirubin Quantitative Input (umol / L) ESR Quantitative Input (mm / h) Evaluate the method for predicting the risk of periprosthetic joint infection in rheumatoid arthritis disclosed in the embodiments of the present invention (referred to as the final model), and compare it with some of the 55 machine learning models in step S6, as shown in Table 2. It can be seen from Table 2 that the final model in the present invention performs best in terms of auc, recall rate, f2-score, and f3-score, and also performs well in terms of accuracy, precision, specificity, and f1-score.
[0066] Table 2 - Comparison of performance evaluations of different models Logistic Regression Random Forest Extra Trees Decision Tree Support Vector Machine Multi-Layer Perceptron Gradient Boosting Tree Final Model auc 0.706 0.763 0.749 0.665 0.745 0.631 0.734 0.798 f1-score 0.415 0.455 0.561 0.383 0.500 0.327 0.468 0.548 Accuracy 0.639 0.639 0.812 0.564 0.774 0.474 0.692 0.752 Precision 0.283 0.303 0.457 0.250 0.395 0.207 0.327 0.392 Recall 0.773 0.909 0.727 0.818 0.682 0.773 0.818 0.909 Specificity 0.613 0.586 0.829 0.514 0.793 0.414 0.667 0.721 f2-score 0.574 0.649 0.650 0.562 0.595 0.500 0.629 0.719 f3-score 0.659 0.758 0.687 0.667 0.636 0.607 0.711 0.803 Among them, f1-score = 1 × Precision × Recall / (1 × Precision + Recall), f2-score = 5 × Precision × Recall / (4 × Precision + Recall), f3-score = 10 × Precision × Recall / (9 × Precision + Recall).
[0067] In the embodiment of the present invention, patients who underwent joint replacement due to rheumatoid arthritis from 2007 to 2021 were recorded from the clinical electronic record system of the First Affiliated Hospital of Sun Yat-sen University, and telephone follow-up was conducted. A total of 106 patients were followed up, including 4 pji patients and 102 non-pji patients. Further, relevant data such as the age, blood glucose, weight, AST, PT, ESR, drug use, CRP, etc. of these patients were collected and input into the encapsulated interface, as Figure 4 shown, the scores of pji patients were significantly higher than those of non-pji patients (0.795 VS 0.417, p < 0.001).
[0068] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, without departing from the principle and spirit of the present invention, various changes, modifications, substitutions, and variations to these embodiments still fall within the protection scope of the present invention.
Claims
1. A method for predicting periprosthetic joint infection after artificial joint prosthesis surgery for rheumatoid arthritis, characterized in that, The method includes the following steps: Collect the original data, and then perform feature screening to obtain a data set, including a training set and a validation set; Through the training set, perform combined training on multiple machine learning models, and then determine the best combination of machine learning models based on the genetic algorithm for the ensemble algorithm to obtain an infection risk prediction model; Perform postoperative infection risk prediction through the infection risk prediction model.
2. The method according to claim 1, wherein: First, preprocess the original data to obtain preprocessed data, and then perform feature screening on the preprocessed data; the content of preprocessing the original data includes one or more of deletion, cleaning, feature encoding, error correction, unit unification, and imputation.
3. The method according to claim 2, wherein: Perform feature screening through the following steps: Perform multiple bootstrap resamplings on the data. Each bootstrap resampling generates different data subsets, and use the data subsets to construct random forests; In each random forest, use the variable importance index to calculate the importance of each original feature; Introduce shadow features, which are generated by randomizing the original features. The shadow features are used to simulate random features and compare them with the original features; Compare the importance of each original feature with its corresponding shadow feature. If the importance of an original feature is significantly higher than its shadow feature, then the original feature is considered to be important; Iterative process: Delete all shadow features, and repeat the above steps until all original features are marked as "important" or "unimportant", so as to perform feature screening based on the importance ranking.
4. The method according to claim 1, characterized in that: Based on K-fold cross-validation, split the data set into a training set and a validation set, including: Using K-fold cross-validation, randomly divide the data set into several folds of basically equal size, and sequentially use each fold as the validation set, and the remaining folds as the training set.
5. The method according to claim 4, characterized in that Before model training, the method further includes the following steps: Balance the data of the training set through the SMOTEENN algorithm so that the ratio between the infected samples and the non-infected samples in the training set meets the preset range.
6. The method according to claim 5, wherein The step of balancing the data of the training set through the SMOTEENN algorithm includes: SMOTE start ① Select random data from the minority class; ② Calculate the distance between the random data and its K nearest neighbors; ③ Multiply the difference in the number of samples by a random number between 0 and 1, and then add the result to the minority class as a synthetic sample; ④ Repeat steps ② and ③ until the required ratio of the minority group is met; SMOTE end, ENN start ⑤ Determine K, which is the number of nearest neighbors. If it is not determined, then K = 3; ⑥ Find the K-nearest neighbors of the observation value among other observations in the training set, and then return the majority class from the K-nearest neighbors; ⑦ If the class of the observation value is different from the majority class of the K-nearest neighbors of the observation value, then delete the observation value and its K-nearest neighbors from the training set; ⑧ Repeat steps ⑤ and ⑥ until the ratio between the infected samples and the non-infected samples meets the preset range.
7. The method according to claim 1, characterized in that: The step of performing combined training on multiple machine learning models and then using the genetic algorithm to determine the best combination based on the ensemble algorithm includes: ① Population initialization: Create a population containing N samples through random generation. The population is a sub-model combination including at least one machine learning model; ② Fitness calculation: The sub-model combination in step ① uses the method of ensemble learning to be verified through the validation set to obtain the fitness score of the sub-model combination; ③ Selection operation: According to the fitness of individuals in the population, individuals with high fitness scores are selected from the current population by roulette wheel, simulating the process of natural selection. By calculating the sum of the fitness of all individuals Σfi, calculate the relative fitness of each individual fi / Σfi, similar to softmax, and then generate a random number between 0 and 1. Determine the number of times each individual is selected according to the position where the random number appears in the probability region; ④ Mating operation: Use a certain mating probability threshold to control whether to perform single-point crossover to generate new individuals; ⑤ Mutation operation: Randomly generate mutation points; ⑥ Iterate the above steps M times, terminate the iteration, and determine the best individual. The combination of the best individuals is the infection risk prediction model.
8. The method according to claim 7, wherein: In step ②, the fitness score is calculated using f3-score, f3-score = 10×precision×recall / (9×precision + recall).
9. The method according to claim 7, wherein: Combine the best individuals into the base learner, and retain the Model file with soft voting as the meta-learner composed of the ensemble model. Then perform interface encapsulation and predict the postoperative infection risk through the encapsulated interface.
10. The method according to claim 1, characterized in that: The method is applied to the prediction of the risk of periprosthetic joint infection after rheumatoid arthritis surgery. The screened features include blood glucose, body weight, hemoglobin, AST, bilirubin, ESR, drug use, CRP, gender, and joint type.
Citation Information
Patent Citations
Diagnostic model and diagnostic system for peripheral infection of artificial joint prosthesis
CN115954102A
Method and system for constructing prostatic cancer disease risk prediction model based on random forest and XGBoost
CN117116477A
Postoperative intracranial infection prediction method and system based on machine learning
CN117133459A
Method and system for constructing pulmonary infection risk prediction model based on machine learning
CN118039177A
Periprosthetic infection treatment target and application thereof
CN118121704A
Cited By
Knee replacement postoperative periprosthetic infection risk prediction system and method
CN122135971A