Postoperative intracranial infection prediction model and establishment method and system thereof
By using a random forest algorithm in the prediction model of intracranial infection after surgery, a decision tree is constructed, and the problem of cumbersome and low accuracy in predicting intracranial infection in the existing technology is solved, and a faster and more accurate infection risk assessment is achieved, which improves the efficiency of medical services.
Patent Information
- Application Number
- CN202510345721.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-22
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art has problems of cumbersome, time-consuming and low accuracy in predicting and preventing and treating intracranial infections after craniotomy.
A postoperative intracranial infection prediction model is used to screen patients based on preset data processing standards, obtain characteristic data, and use a random forest algorithm to build a decision tree to generate a prediction model to evaluate the patient's risk of intracranial infection.
This method can more quickly and accurately judge the incidence of intracranial infection after surgery, reduce the workload of doctors, and improve the quality and efficiency of medical services.
Smart Images

Figure CN120148869A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning applications, and particularly to a postoperative intracranial infection prediction model, a method for establishing the same, and a system. Background Art
[0002] The occurrence of postoperative infections can lead to an extended hospital stay for surgical patients, an increase in medical costs, and a higher readmission rate, and even pose a threat to the patient's life safety. In recent years, the number of patients undergoing neurosurgical craniotomy has been increasing, and the skills of surgical staff have been improving day by day, but the continuous rise in the postoperative intracranial infection rate still cannot be avoided. Regarding the risk factors that may cause intracranial infection after craniotomy, due to the different operating room environments, the qualities of medical staff, and the hospital environments, different research conclusions are also different. To obtain a more representative and general conclusion, more researchers are needed to participate in this research and formulate a series of relatively reasonable and general prevention and treatment strategies for intracranial infection after craniotomy.
[0003] Many doctors have studied a large number of craniotomy case materials and used methods such as retrospective analysis, univariate variable analysis, and x2 test to statistically analyze the relevant factors that may cause intracranial infection: postoperative indwelling drainage tubes, postoperative cerebrospinal fluid leakage, operation time, operation site, preoperative underlying diseases (such as diabetes, hypertension, pulmonary infection, etc.), preoperative consciousness disorders, postoperative use of hormones, operation types, as well as gender and age. And based on the analysis results, guidance on prevention and treatment strategies is provided from aspects such as nursing, nutritional support, and the application of antibiotics. This process is not only cumbersome, but also because it is manually counted, it not only takes a long time, but also the accuracy cannot be guaranteed, so the progress is slow. Summary of the Invention
[0004] The purpose of the present invention is to provide a postoperative intracranial infection prediction model, a method for establishing the same, and a system to solve the problems raised in the above background art.
[0005] To achieve the above purpose, the present invention provides the following technical solutions:
[0006] A postoperative intracranial infection prediction model and a method for establishing the same, the method comprising:
[0007] Screening patients based on a preset data processing standard to obtain the characteristic data of the patients; the characteristic values in the characteristic data at least include the GCS score, the operation site, and the lumbar cistern drainage tube time;
[0008] Randomly selecting data from the characteristic data to construct training tuples, and looping until the number of training tuples reaches a preset number threshold;
[0009] Construct decision trees based on each training tuple to obtain a threshold number of decision trees, which serve as a random forest;
[0010] For any user to be analyzed, obtain their feature data, and identify the obtained feature data based on the random forest to obtain an evaluation result.
[0011] As a further solution of the present invention: the data processing criteria include: whether the patient is a postoperative patient, whether there is a history of infection, whether the age is between 18 and 80 years old, whether there is an immunodeficiency disease and severe chronic heart, lung, and kidney diseases, whether the patient has granted permission; the feature values also include: patient name, gender, age, primary disease classification, preoperative GCS score, number of surgeries before infection, surgical site, surgical timing, surgical time, intracranial drainage tube time, lumbar cistern drainage tube time, whether artificial materials are implanted, postoperative nutritional status assessment, cerebrospinal fluid leakage, open wound, postoperative fasting blood glucose, intraoperative blood loss, combined infection in other parts, use of glucocorticoids, prophylactic use of antibiotics time, postoperative intracranial infection, systolic blood pressure, Glasgow coma score, seizure, hypertension, diabetes, and IS_Nglioma.
[0012] As a further solution of the present invention: the step of randomly selecting data from the feature data to construct training tuples and looping until the number of training tuples reaches a preset number threshold includes:
[0013] Randomly select data from the feature data to obtain an array; the number of selected data is random, and each selection is random;
[0014] Loop through the random selection process of data, and when the number of times reaches a preset number threshold, break out of the loop;
[0015] Count the selected array as a training tuple;
[0016] Loop through the generation process of training tuples until the number of training tuples reaches a preset number threshold.
[0017] As a further solution of the present invention: the step of constructing decision trees based on each training tuple to obtain a threshold number of decision trees, which serve as a random forest, includes:
[0018] Read the training tuples in sequence and construct decision trees;
[0019] When decision trees are constructed for all training tuples, a random forest is obtained;
[0020] Among them, during the construction process of the random forest, the out-of-bag (OOB) estimation model is used to evaluate the generalization error.
[0021] As a further solution of the present invention: for any user to be analyzed, the steps of obtaining their characteristic data and identifying the obtained characteristic data based on a random forest to obtain an evaluation result include:
[0022] For any user to be analyzed, receive the permissions granted by the user;
[0023] Query the characteristic data of the user based on the permissions granted by the user;
[0024] Randomly select part of the data from the characteristic data and input it into each decision tree in the random forest to obtain the classification results of the decision trees;
[0025] Integrate the classification results of all decision trees to obtain the final evaluation result.
[0026] The technical solution of the present invention also provides a postoperative intracranial infection prediction model and its establishment system, and the system includes:
[0027] A characteristic data acquisition module, configured to screen patients based on a preset data processing standard and obtain the characteristic data of the patients; the characteristic values in the characteristic data at least include the GCS score, the surgical site, and the lumbar cistern drainage tube time;
[0028] A training tuple construction module, configured to randomly select data from the characteristic data to construct training tuples, and loop until the number of training tuples reaches a preset number threshold;
[0029] A random forest generation module, configured to construct decision trees based on each training tuple to obtain a number threshold of decision trees as a random forest;
[0030] A random forest application module, configured to, for any user to be analyzed, obtain their characteristic data and identify the obtained characteristic data based on the random forest to obtain an evaluation result.
[0031] As a further solution of the present invention: the data processing standard includes: whether it belongs to a postoperative patient, whether there is an infection history, whether the age is between 18 and 80 years old, whether there is an immunodeficiency disease and severe chronic heart, lung, and kidney diseases, whether the patient grants permissions; the characteristic values further include: patient name, gender, age, primary disease classification, preoperative GCS score, number of surgeries before infection, surgical site, surgical timing, surgical time, intracranial drainage tube time, lumbar cistern drainage tube time, whether artificial materials are implanted, postoperative nutritional status assessment, cerebrospinal fluid leakage, open wound, postoperative fasting blood glucose, intraoperative blood loss, combined infection in other parts, use of glucocorticoids, prophylactic use of antibiotics time, postoperative intracranial infection, systolic blood pressure, Glasgow Coma Scale score, seizure, hypertension, diabetes, and IS_Nglioma.
[0032] As a further solution of the present invention: The training tuple construction module includes:
[0033] An array generation unit, configured to randomly select data from the feature data to obtain an array; the number of selected data is random, and each selection is random;
[0034] A first loop unit, configured to loop and execute the random selection process of data, and jump out of the loop when the number of times reaches a preset number threshold;
[0035] A statistics unit, configured to count the selected arrays as training tuples;
[0036] A second loop unit, configured to loop and execute the generation process of training tuples until the number of training tuples reaches a preset quantity threshold.
[0037] As a further solution of the present invention: The random forest generation module includes:
[0038] Read the training tuples in sequence to construct decision trees;
[0039] When decision trees are constructed for all training tuples, a random forest is obtained;
[0040] Among them, during the construction process of the random forest, the out-of-bag (OOB) estimation model is used to evaluate the generalization error.
[0041] As a further solution of the present invention: The random forest application module includes:
[0042] A permission receiving unit, configured to receive the permissions granted by any user to be analyzed;
[0043] A data query unit, configured to query the feature data of the user based on the permissions granted by the user;
[0044] A random input unit, configured to randomly select some data from the feature data and input it into each decision tree in the random forest to obtain the classification results of the decision trees;
[0045] A result integration unit, configured to integrate the classification results of all decision trees to obtain a final evaluation result.
[0046] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention can help doctors better judge the probability of postoperative infection, reduce the workload of doctors, enable the majority of patients to enjoy higher-quality medical services more quickly at a relatively low price, so that more people can benefit from it. By observing the changes in human characteristics and monitoring data, the situation of intracranial infectious diseases can be analyzed, and these data can also be collected and analyzed and compared in the background, enabling more accurate judgment, auxiliary diagnosis, and long-term management. At the same time, it provides an example for the application of artificial intelligence technology in other medical fields. Brief Description of the Drawings
[0047] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention.
[0048] Figure 1 is a random forest algorithm model of a multi - decision tree.
[0049] Figure 2 is a flowchart of the implementation of the random forest algorithm.
[0050] Figure 3 is the out - of - bag (OOB) estimate of the random forest model.
[0051] Figure 4 is a flowchart of the postoperative intracranial infection prediction model. Detailed Embodiments
[0052] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0053] Please refer to Figures 1 to 4 , in an embodiment of the present invention, a postoperative intracranial infection prediction model and a method for establishing the same, the method includes:
[0054] Screen patients based on a preset data processing standard to obtain the characteristic data of the patients; the characteristic values in the characteristic data at least include the GCS score, the surgical site, and the duration of the lumbar cistern drainage tube;
[0055] Randomly select data from the characteristic data to construct training tuples, and loop until the number of training tuples reaches a preset number threshold;
[0056] Construct decision trees based on each training tuple to obtain a number threshold of decision trees, which are used as a random forest;
[0057] For any user to be analyzed, obtain their characteristic data, and identify the obtained characteristic data based on the random forest to obtain an evaluation result.
[0058] In an example of the technical solution of the present invention, patients are first screened according to a preset standard and divided into different categories, and the classification results are the labels of the patients; at the same time, the characteristic data of the patients are obtained, and some data are randomly selected from the characteristic data as features, and a classification model from features to labels, that is, a decision tree, is constructed; by continuously executing the above process, multiple decision trees can be obtained, and multiple decision trees together constitute a random forest; in the application stage, for any user to be analyzed, their characteristic data is obtained, and some data is selected from the characteristic data and input into the random forest. Each decision tree or some randomly selected decision trees will give multiple results, and by synthesizing these results, a final evaluation result is obtained.
[0059] Among them, the data processing criteria include: whether the patient belongs to a postoperative patient, whether there is a history of infection, whether the age is between 18 and 80 years old, whether there is an immunodeficiency disease and severe chronic heart, lung, and kidney diseases, and whether the patient has been granted permission; the specific description is as follows:
[0060] The criteria for including data in this model are: (1) For patients undergoing neurosurgery or cerebrospinal fluid drainage, they should cooperate with standardized treatment after surgery; (2) They have no previous history of craniotomy and no history of intracranial infection; (3) The age is >18 years old and <80 years old; (4) They have no previous immunodeficiency diseases and severe chronic diseases such as heart, lung, and kidney; (5) Informed consent has been obtained from the patient or their family members.
[0061] The criteria for excluding data in this model are: (1) There is intracranial infection before surgery. (2) Those who do not receive standardized treatment or voluntarily give up after surgery. (3) Those with immunodeficiency diseases. (4) Those with severe chronic diseases such as heart, lung, and kidney, and those with other organ malignancies.
[0062] Furthermore, regarding the feature values, the feature values also include: patient name, gender, age, primary disease classification, preoperative GCS score, number of surgeries before infection, surgical site, surgical timing, surgical time, intracranial drainage tube time, lumbar cistern drainage tube time, whether artificial materials are implanted, postoperative nutritional status assessment, cerebrospinal fluid leakage, open wound, postoperative fasting blood glucose, intraoperative blood loss, combined infection in other parts, use of glucocorticoids, prophylactic use of antibiotics time, postoperative intracranial infection, systolic blood pressure, Glasgow coma score, seizure, hypertension, diabetes, and IS_Nglioma; for these feature values, their value ranges (values) are fixed, and the specific units are as follows:
[0063] Table 1 Parameters of data for predicting postoperative infection in neurosurgery
[0064]
[0065]
[0066] As a preferred embodiment of the technical solution of the present invention, the step of randomly selecting data from the feature data, constructing training tuples, and looping until the number of training tuples reaches a preset number threshold includes:
[0067] Randomly select data from the feature data to obtain an array; the number of selected data is random, and each selection is random;
[0068] Loop through the random selection process of data, and when the number of times reaches the preset number threshold, break out of the loop;
[0069] Count the selected array as a training tuple;
[0070] Loop through the generation process of training tuples until the number of training tuples reaches the preset number threshold.
[0071] Further, the step of constructing a decision tree based on each training tuple to obtain a number threshold of decision trees as a random forest includes:
[0072] Read the training tuples in sequence and construct a decision tree;
[0073] When decision trees are constructed for all training tuples, a random forest is obtained;
[0074] Among them, during the construction process of the random forest, the out-of-bag (OOB) estimation model is used to evaluate the generalization error.
[0075] This model constructs a postoperative infection prediction model for neurosurgical patients according to various feature data such as the patient's preoperative GCS score, surgical site, and lumbar cistern drainage tube time in Table 1, using the classification algorithm of the random forest.
[0076] A random forest is a combined classifier composed of multiple decision trees. Each decision tree depends on independent sampling. All the trees in the forest have values of identically distributed random variables, and the classification result is determined by the class with the most votes in the voting of each tree.
[0077] The main features of the random forest algorithm lie in "randomness". One is the randomization of training tuples, and the other is the randomization of feature variables:
[0078] (1) Randomly sample and select the training tuples for each decision tree. Each decision tree corresponds to independently and identically distributed training tuples. If the scale of the random forest is N trees, then N training tuples need to be generated, that is, N training subsets are generated from the original dataset. The statistical sampling method is used to generate the training subsets, including:
[0079] Sampling without replacement method:
[0080] Suppose the population has a total of N samples, and n (n < N) samples are drawn without replacement. In sampling without replacement, the number of the original data set gradually decreases during the sampling process, and the sampling samples will not be repeated.
[0081] Sampling with replacement method:
[0082] When extracting training samples, after each sample is drawn, the sample is put back for the next extraction. The size of the overall data set remains fixed during each sampling, and the drawn sample set may have repeated samples.
[0083] (2) Random selection of features:
[0084] Features are the set of attributes participating in the attribute measurement when each decision tree is generated. When the decision trees in the random forest perform node splitting, a part of the features are randomly selected to calculate the attribute measurement values. The purpose of randomly selecting features is to improve the prediction accuracy of the random forest, reduce the correlation coefficient between decision trees, and keep the strength of the forest unchanged.
[0085] Furthermore, the steps of the random forest algorithm include:
[0086] (1) From each data set {N 1 , N 2 ,..., N K} after the division of each feature data, there are K different types of each feature data, and each data set has N tuples, sample N times respectively in a sampling-with-replacement manner, and each data set N i (i = 1, 2,..., k) generates a training set of size N {A 1 , A 2 , …, A N};
[0087] (2) Repeat the sampling in the manner of step (1) so that each N i , data generates T groups of training sets, that is, {T 1 , T 2 ,... T K}, where T i = A 1 , A 2 , …, A N ;
[0088] (3) For each group of training sets T i , randomly select m feature attributes from the entire set of candidate attributes M;
[0089] (4) Calculate the information gain for each of the m features respectively, and select the best splitting node;
[0090] (5) Iterate step (4) to generate a decision tree;
[0091] (6) Iterate steps (3) to (4) to generate T decision trees, each tree growing completely without pruning;
[0092] (7) After the T trees for N i are built, return to step (3) and continue to construct T trees for another set of feature data until the T trees for all N i sets of feature data {N 1 , N 2 ,..., N K} are all trained;
[0093] (8) Test the data and calculate the classification results for each forest. The classification results are obtained from the voting rates of multiple decision trees. The voting rate is the ratio of the number of times the most frequently occurring classification label appears to the total number of votes. The final classification label is determined by the classification corresponding to the highest voting rate.
[0094] Regarding step (4):
[0095] The random forest algorithm is an ensemble learning method that classifies or regresses by constructing multiple decision trees. In a random forest, each tree is trained on a random subset of the dataset, and each split is performed on a random subset of the features. This can increase the diversity of the model and improve the generalization ability.
[0096] Step (4): Calculate the information gain and select the best split node:
[0097] In the random forest algorithm, in the construction process of each tree, the step of selecting the best split node is crucial. It involves the following sub-steps:
[0098] Select features and split points:
[0099] In the given subset of features (usually a random subset of all features), calculate the information gain for each possible split point of each feature.
[0100] Calculate the information gain:
[0101] The information gain is based on the concept of entropy. Entropy is a measure of the disorder of a dataset. In a classification problem, entropy can be calculated by the following formula:
[0102]
[0103] where H(S) is the entropy of the dataset S, c is the number of classes, and pi is the proportion of samples belonging to the i-th class in the dataset S.
[0104] For a specific feature A and its possible split value v, the dataset can be divided into two parts Sleft and S right ; At this time, the information gain IG(S, A) of feature A on dataset S can be calculated as: where |S| is the number of samples in dataset S, |S left | and |S right | are the numbers of samples in the two parts after splitting respectively.
[0105] Select the best splitting point:
[0106] For each feature, calculate the information gain of all possible splitting points.
[0107] Select the feature and splitting point with the largest information gain as the best splitting point.
[0108] Split the dataset:
[0109] Use the best splitting point to divide the dataset into two subsets, and then recursively repeat the above process on these two subsets until the stopping condition is met (such as the tree reaches the maximum depth or the number of samples in the node is less than a certain threshold).
[0110] In this way, each tree will select the feature and splitting point that can best improve the classification accuracy during the construction process, thereby improving the performance and stability of the entire random forest model. The random forest usually can obtain better prediction performance than a single decision tree by integrating the prediction results of multiple trees.
[0111] After the random forest algorithm training is completed, it is necessary to measure the classification accuracy of the model. Use OOB to estimate the generalization error of the model, that is, for each iterative sampling, use the sampled data to train the algorithm, and the out-of-bag data not selected is used to test the algorithm.
[0112] Suppose the total number of out-of-bag data is M. M out-of-bag data are used as input and input into the extended random forest algorithm classifier. The classifier predicts the classification labels of M data. Since the classification results of M data are known, then compare the correct known classification results with the classification results of the extended random forest classifier, and count the number of tuples misclassified by the extended random forest classifier, denoted as X. Then the size of the generalization error predicted by OOB is X / M.
[0113] This model processes data for patients with postoperative intracranial infection prediction. After data processing, various feature data such as the patient's preoperative GCS score, surgical site, and duration of lumbar cistern drainage tube are input into the algorithm model to obtain the prediction probability. Finally, the prediction probability is statistically analyzed through a line chart to analyze the infection situation and predict whether the patient is infected after surgery.
[0114] During the training process of the random forest model, when constructing a single tree in each iteration, approximately one-third of the data is not selected for training through sampling with replacement. These data are called Out-of-Bag (OOB) data. These OOB data provide a natural validation set for each tree because they do not participate in the training of the corresponding tree and can be used to evaluate the generalization ability of the model. For each tree, use its OOB data for testing and calculate the classification accuracy or other performance metrics. Since the prediction result of the random forest is usually obtained by majority voting of the predictions of all trees, calculate the classification accuracy of all trees on the OOB data for each sample. In this way, we can estimate the generalization error of the entire random forest model without a separate test set because this process directly utilizes the data not used during the training process. Finally, by averaging the OOB errors of all trees, a reliable estimate of the model's generalization performance is obtained.
[0115] As a preferred embodiment of the technical solution of the present invention, the step of obtaining the feature data of any user to be analyzed and identifying the obtained feature data based on the random forest to obtain an evaluation result includes:
[0116] For any user to be analyzed, receive the permissions granted by the user;
[0117] Query the feature data of the user based on the permissions granted by the user;
[0118] Randomly select part of the data from the feature data and input it into each decision tree in the random forest to obtain the classification results of the decision trees;
[0119] Integrate the classification results of all decision trees to obtain the final evaluation result.
[0120] The above content describes the actual application stage. For the user to be analyzed, receive the permissions granted by the user. Without permissions, the feature data cannot be obtained; query the feature data of the user based on the permissions granted by the user, randomly select part of the data from the feature data, and input it into each decision tree in the random forest to obtain the classification results of the decision trees. Among them, the process of inputting into each decision tree in the random forest can also randomly select some decision trees; finally, integrate the classification results of all decision trees to obtain the final evaluation result; the integration method is very simple because the output of each decision tree is a classification result, and selecting the mode value of the classification results can be used as the final result.
[0121] As a preferred embodiment of the technical solution of the present invention, in the embodiment of the present invention, a postoperative intracranial infection prediction model and its establishment system, the system includes:
[0122] A feature data acquisition module, which is used to screen patients based on a preset data processing standard and obtain the feature data of the patients; the feature values in the feature data at least include the GCS score, the surgical site, and the lumbar drainage tube time;
[0123] A training tuple construction module, which is used to randomly select data from the feature data to construct training tuples, and loop until the number of training tuples reaches a preset number threshold;
[0124] A random forest generation module, which is used to construct decision trees based on each training tuple to obtain a number threshold of decision trees as a random forest;
[0125] A random forest application module, which is used to obtain the feature data of any user to be analyzed and identify the obtained feature data based on the random forest to obtain an evaluation result.
[0126] Among them, the data processing standard includes: whether the patient is a postoperative patient, whether there is an infection history, whether the age is between 18 and 80 years old, whether there is an immunodeficiency disease and severe chronic heart, lung, and kidney diseases, whether the patient has granted permission; the feature values also include: patient name, gender, age, primary disease classification, preoperative GCS score, number of surgeries before infection, surgical site, surgical timing, surgical time, intracranial drainage tube time, lumbar drainage tube time, whether artificial materials are implanted, postoperative nutritional status assessment, cerebrospinal fluid leakage, open wound, postoperative fasting blood glucose, intraoperative blood loss, combined infection in other parts, use of glucocorticoids, prophylactic use of antibiotics time, postoperative intracranial infection, systolic blood pressure, Glasgow Coma Scale score, seizure, hypertension, diabetes, and IS_Nglioma.
[0127] Further, the training tuple construction module includes:
[0128] An array generation unit, which is used to randomly select data from the feature data to obtain an array; the number of selected data is random, and each selection is random;
[0129] A first loop unit, which is used to loop the random selection process of data and jump out of the loop when the number of times reaches a preset number threshold;
[0130] A statistics unit, which is used to count the selected arrays as training tuples;
[0131] A second loop unit, which is used to loop the generation process of training tuples until the number of training tuples reaches a preset number threshold.
[0132] Specifically, the random forest generation module includes:
[0133] Read the training tuples in sequence and construct decision trees;
[0134] After all training tuples have constructed the decision tree, a random forest is obtained;
[0135] Among them, during the construction of the random forest, the out-of-bag (OOB) estimate is used to evaluate the generalization error of the model.
[0136] Furthermore, the random forest application module includes:
[0137] A permission receiving unit, which is used to receive the permissions granted by any user to be analyzed.
[0138] A data query unit, which is used to query the feature data of the user based on the permissions granted by the user.
[0139] A random input unit, which is used to randomly select some data from the feature data and input it into each decision tree in the random forest to obtain the classification results of the decision trees.
[0140] A result integration unit, which is used to integrate the classification results of all decision trees to obtain the final evaluation result.
[0141] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A prediction model for postoperative intracranial infection and a method for establishing the same, characterized in that: The method comprises: Screening patients based on preset data processing standards to obtain characteristic data of the patients; characteristic values in the characteristic data at least include GCS score, surgical site and lumbar drainage tube time; Randomly select data from the feature data to construct training tuples, and execute the process in a loop until the number of training tuples reaches a preset number threshold; Build a decision tree based on each training tuple to obtain a threshold number of decision trees as a random forest; For any user to be analyzed, its feature data is obtained, and the obtained feature data is identified based on random forest to obtain the evaluation result.
2. The postoperative intracranial infection prediction model and the method for establishing the same according to claim 1, characterized in that: The data processing standards include: whether the patient is a postoperative patient, whether there is a history of infection, whether the age is between 18 and 80 years old, whether there is an immunodeficiency disease and severe chronic heart, lung, and kidney disease, and whether the patient has granted permission; the characteristic values also include: patient name, gender, age, primary disease classification, preoperative GCS score, number of surgeries before infection, surgical site, timing of surgery, surgical time, intracranial drainage tube time, lumbar drainage tube time, whether artificial materials are implanted, postoperative nutritional status assessment, cerebrospinal fluid leakage, open wound, postoperative fasting blood glucose, intraoperative blood loss, combined infection in other parts, use of glucocorticoids, time for preventive use of antibiotics, postoperative intracranial infection, systolic blood pressure, Glasgow Coma Score, epileptic seizures, hypertension, diabetes, and IS_Nglioma.
3. The postoperative intracranial infection prediction model and the method for establishing the same according to claim 1, characterized in that: The step of randomly selecting data from the feature data, constructing training tuples, and executing the step in a loop until the number of training tuples reaches a preset number threshold comprises: Randomly select data from the characteristic data to obtain an array; the number of selected data is random, and each selection is random; The random selection process of data is executed cyclically, and when the number of times reaches the preset threshold, the loop is exited; Count the selected arrays as training tuples; The training tuple generation process is executed cyclically until the number of training tuples reaches a preset threshold.
4. The postoperative intracranial infection prediction model and the method for establishing the same according to claim 1, characterized in that: The step of constructing a decision tree based on each training tuple to obtain a threshold number of decision trees as a random forest includes: Read the training tuples sequentially and build a decision tree; When all training tuples have completed the construction of decision trees, a random forest is obtained; Among them, in the process of constructing the random forest, the generalization error of the OOB estimation model is used for evaluation.
5. The postoperative intracranial infection prediction model and the method for establishing the same according to claim 1, characterized in that: For any user to be analyzed, obtaining characteristic data thereof, The steps of identifying the acquired feature data based on random forest and obtaining the evaluation results include: For any user to be analyzed, receive the permissions granted by the user; Query the user's feature data based on the permissions granted by the user; Randomly select some data from the feature data, input them into each decision tree in the random forest, and obtain the classification results of the decision tree; The classification results of all decision trees are combined to obtain the final evaluation result.
6. A postoperative intracranial infection prediction model and its establishment system, characterized in that: The system comprises: A characteristic data acquisition module, used to screen patients based on preset data processing standards and acquire characteristic data of the patients; the characteristic values in the characteristic data at least include GCS score, surgical site and lumbar drainage tube time; A training tuple construction module is used to randomly select data from the feature data, construct training tuples, and execute the process in a loop until the number of training tuples reaches a preset number threshold; A random forest generation module is used to construct a decision tree based on each training tuple to obtain a threshold number of decision trees as a random forest; The random forest application module is used to obtain the feature data of any user to be analyzed, identify the obtained feature data based on the random forest, and obtain the evaluation result.
7. The postoperative intracranial infection prediction model and its establishment system according to claim 6, characterized in that: The data processing standards include: whether the patient is a postoperative patient, whether there is a history of infection, whether the age is between 18 and 80 years old, whether there is an immunodeficiency disease and severe chronic heart, lung, and kidney disease, and whether the patient has granted permission; the characteristic values also include: patient name, gender, age, primary disease classification, preoperative GCS score, number of surgeries before infection, surgical site, timing of surgery, surgical time, intracranial drainage tube time, lumbar drainage tube time, whether artificial materials are implanted, postoperative nutritional status assessment, cerebrospinal fluid leakage, open wound, postoperative fasting blood glucose, intraoperative blood loss, combined infection in other parts, use of glucocorticoids, time for preventive use of antibiotics, postoperative intracranial infection, systolic blood pressure, Glasgow Coma Score, epileptic seizures, hypertension, diabetes, and IS_Nglioma.
8. The postoperative intracranial infection prediction model and its establishment system according to claim 6, characterized in that: The training tuple building module includes: An array generation unit, used for randomly selecting data from the feature data to obtain an array; the number of selected data is random, and each selection is random; The first loop unit is used to loop through the random selection process of data, and when the number of times reaches a preset number threshold, the loop is exited; The statistical unit is used to count the selected array as a training tuple; The second loop unit is used to loop through the training tuple generation process until the number of training tuples reaches a preset number threshold.
9. The postoperative intracranial infection prediction model and its establishment system according to claim 6, characterized in that: The random forest generation module includes: Read the training tuples sequentially and build a decision tree; When all training tuples have completed the construction of decision trees, a random forest is obtained; Among them, in the process of constructing the random forest, the generalization error of the OOB estimation model is used for evaluation.
10. The postoperative intracranial infection prediction model and its establishment system according to claim 6, characterized in that: The random forest application module includes: The permission receiving unit is used to receive the permission granted by any user to be analyzed; A data query unit, used to query the user's feature data based on the permissions granted by the user; A random input unit is used to randomly select part of the feature data and input it into each decision tree in the random forest to obtain the classification result of the decision tree; The result synthesis unit is used to synthesize the classification results of all decision trees to obtain the final evaluation result.