Machine learning-based general medical early warning scoring system construction method

By constructing a medical early warning scoring system based on machine learning feature screening, binning, encoding and scoring mapping methods, we solved the problems of strong subjectivity, poor interpretability and insufficient feature engineering in existing technologies, achieved a more accurate, stable and easy-to-use scoring system, and improved the generalization ability and clinical application value of the model.

CN120674090APending Publication Date: 2025-09-19中国人民解放军总医院第八医学中心
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510663112.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing medical early warning scoring systems have problems such as strong subjectivity, limited results, poor interpretability and difficult maintenance. Methods based on expert experience lack objectivity and universality, while methods based on simple modeling have deficiencies in feature engineering, encoding and score mapping, resulting in limited model performance and application effects.

Method used

A medical early warning scoring system was constructed using a machine learning-based method through feature initial screening, binning, word of essence (WOE) coding, Lasso regression model and ODDs-PDO score mapping. This system included a full-process systematic solution of feature screening, binning, coding, model building and score mapping. The Lasso regression model was used to automatically select features, the WOE coding method was used to reflect the intrinsic correlation between features and target variables, and the ODDs-PDO method was used to map the probability values ​​output by the model to the clinical scoring system.

Benefits of technology

It has achieved a more accurate, stable, explainable and easy-to-use medical early warning scoring system, improved feature quality and model performance, and provided intuitive, detailed and explainable scoring results, helping doctors to make more accurate judgments on the condition and having higher application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120674090A_ABST
    Figure CN120674090A_ABST
Patent Text Reader

Abstract

The invention discloses a general medical early warning scoring system construction method based on machine learning. The method comprises the following steps: performing primary screening on features of original data to obtain a feature subset; performing binning processing on the feature subset; coding the classification features after binning processing by adopting a WOE coding method; performing feature re-screening on the encoded data; using a Lasso regression model to automatically perform feature selection on the data after feature rescreening; and calculating an early warning total score and mapping a probability value output by the model to a clinical scoring system to provide a scoring result. The method has remarkable technical progress significance in the field of medical early warning scoring systems. Compared with the prior art, the scheme has obvious advantages in the aspects of precision, generalization ability, calculation cost, application range and the like, and a more efficient, accurate and explainable solution can be provided for construction of a medical early warning scoring system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for building a medical early warning system, and in particular to a method for building a universal medical early warning scoring system based on machine learning. Background Art

[0002] A medical early warning scoring system assesses various indicators of a patient's condition and vital signs, obtaining quantitative scores for each item. This score is then summed to determine the probability and risk of a patient's condition worsening or a critical incident. A reasonable and effective scoring system is crucial for improving the success rate of emergency treatment, rationally allocating medical resources, and evaluating treatment effectiveness.

[0003] In the background technology of medical early warning scoring systems, existing technical solutions are mainly divided into two categories: one is a method based on expert experience and the other is a method based on simple modeling.

[0004] Among these approaches, those based on expert experience rely primarily on experienced physicians developing early warning scoring rules based on their clinical experience. Through long-term observation of patients' conditional changes, physicians identify correlations between abnormalities in certain vital signs, symptoms, or indicators and worsening conditions, and then translate these into scoring rules. For example, the SOFA scoring system uses indicators such as oxygenation index, platelet count, and bilirubin concentration to score organ failure in six different systems. Each organ is scored 0.4 points, with a total score of 24 points, and is used to assess a patient's risk level. The drawbacks of this approach are its high subjectivity. Scoring rules rely on the physician's personal experience, and different physicians may come up with different scoring rules, making the scoring system less objective. The results are also limited, as the scoring system relies too heavily on the distribution of sample cases, potentially failing to translate effectively when applied to other regions or populations. Furthermore, the scoring system suffers from poor interpretability, as it is a "black box" process, making it impossible to explain the rationale behind the weighting of each indicator, and the prediction results difficult to trace. Maintenance and updating are also difficult. When an expert physician leaves their position, their experience is difficult to retain, and new physicians must re-evaluate their experience, thus hindering the continuity of the system.

[0005] Methods based on simple modeling use simple feature engineering and statistical models (such as linear regression and logistic regression) to describe the relationship between vital signs, establish scoring functions, and assess the patient's condition. However, their disadvantages are weak feature engineering and a lack of reasonable feature screening and binning of model variables, resulting in variable redundancy and complex application. Current methods often use simple label encoding for categorical variables, without considering whether different categories have a sequential linear relationship. Existing methods often directly use the probability value output by the model as the score, lacking an effective mapping conversion from probability values ​​to clinical scores. This results in less than intuitive scores and makes comparison with other scoring results difficult.

[0006] Therefore, while methods based on expert experience are widely used, they suffer from strong subjectivity, limited results, poor interpretability, and difficulty in maintenance. While simple modeling methods are more objective, they suffer from deficiencies in feature engineering, encoding, and score mapping, which limit model performance and application effectiveness. A more accurate, stable, interpretable, and easy-to-use medical early warning scoring system is urgently needed. Summary of the Invention

[0007] In order to address the shortcomings of the above technologies, the present invention provides a method for building a universal medical early warning scoring system based on machine learning.

[0008] In order to solve the above technical problems, the technical solution adopted by the present invention is: a method for building a universal medical early warning scoring system based on machine learning, the method comprising the following steps: Step 1: Get feature subsets by preliminarily screening the features of the original data; Step 2: binning the feature subsets; Step 3: Encode the category features after binning using the WOE encoding method; Step 4: re-screen the encoded data based on its features; Step 5: Use the Lasso regression model to automatically select features from the data after feature re-screening; Step 6: Calculate the total warning score and map the probability value output by the model to the clinical scoring system to provide the scoring result.

[0009] Furthermore, the initial feature screening of the original data includes: Screening is performed based on the feature missing rate, and features with a missing rate greater than 90% are eliminated; The upper and lower bounds of the values ​​are calculated based on the quartile spread method, the number of outliers outside the bounds is counted, and features with an outlier ratio greater than 90% are removed; Count the repeated values ​​in the variables and remove features with a single repeated value greater than 90%.

[0010] Furthermore, the formula for calculating the upper and lower bounds of the value is: , Among them, UP is the upper boundary; LB is the lower boundary; Q1 is the 25% quantile, Q3 is the 75% quantile, and IQR is the interquartile range.

[0011] Furthermore, the method of binning the feature subset is to first determine the variable type of the feature subset. When the feature subset belongs to a numerical variable, the equal-interval binning method or the equal-frequency binning method is performed according to the data distribution. When the feature subset belongs to a categorical variable, the variables with more categories and smaller sample sizes are directly merged.

[0012] Furthermore, the equal-interval binning method divides the value range into a fixed number of bin groups, so that each bin group covers the same value range; the equal-frequency binning method divides the sample distribution into a fixed number of bin groups, so that each bin group contains the same number of samples.

[0013] Furthermore, feature rescreening includes: The correlation between all variables was calculated using the correlation calculation method. If the absolute value of the correlation was greater than 0.8, a variable was eliminated. Use the tree model for variable screening, calculate the feature information gain ratio as the importance coefficient at each tree node split, sort the features according to their importance scores, and select the top 90% with the highest scores.

[0014] Furthermore, the objective function of the Lasso regression model in step 5 is as follows: , Among them, β is the characteristic coefficient of the model; n is the number of samples; p is the number of features; X is the feature matrix, y is the target variable, and λ is the regularization parameter.

[0015] Furthermore, in step six, the ODDs-PDO method is used to calculate the total warning score and the total score is split into scores corresponding to each indicator through the model weight.

[0016] Furthermore, the calculation formula for the warning total score S is: , Among them, p is the probability value output by the model, S0 is the benchmark score, and k is the adjustment factor.

[0017] Furthermore, the way to split the total score into scores corresponding to each indicator by model weight is to assume that there are n features in the model, the weight of each feature is w1, w2, ..., wn, and the corresponding feature value is x1, x2, ..., xn, then the score of each feature is: , Among them, μi is the mean of the i-th feature, and σi is the standard deviation of the i-th feature.

[0018] The present invention discloses a method for building a universal medical early warning scoring system based on machine learning, which has significant technical advantages over the existing technology: A full-process systematic construction was achieved, and a method for building a medical early warning scoring system based on feature screening, binning, coding, model building and score mapping was proposed. Compared with traditional methods based on expert experience or simple modeling, it has stronger systematicity and accuracy.

[0019] It has feature processing optimization and proposes feature screening and binning technology solutions for medical data, which can effectively remove redundant features and process different types of features. Compared with existing technologies, it significantly improves feature quality and model performance.

[0020] It has the advantage of feature encoding. The WOE encoding method can better reflect the intrinsic correlation between features and target variables. Compared with encoding methods such as One-Hot, it can obtain richer feature information and improve the accuracy of the model.

[0021] It has high efficiency in model construction. Through the Lasso regression model construction method, it automatically performs feature selection and reduces the risk of overfitting. Compared with existing simple modeling methods, it can more effectively handle the correlation between features and improve the generalization ability and interpretability of the model.

[0022] With the intuitiveness of scoring mapping, an ODDs-PDO scoring mapping method was designed to map the probability values ​​output by the model into the scoring system, providing intuitive, detailed and interpretable scoring results. Compared with existing scoring mapping methods, it has better interpretability and operability, and is convenient for clinical application.

[0023] With full process automation, the produced model can output scoring results throughout the entire process without the need for additional manual processing, directly serving clinical diagnosis and treatment, helping doctors to judge the condition more accurately, and having higher application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 Flowchart of the present invention.

[0025] Figure 2 Schematic diagram of the feature binning process. DETAILED DESCRIPTION

[0026] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0027] like Figure 1 The method for building a universal medical early warning scoring system based on machine learning is shown, including a feature screening module, a feature binning module, a feature encoding module, a model building module, and a scoring mapping module.

[0028] The feature screening module consists of a primary screening phase and a secondary screening phase. The primary screening phase performs variable selection (primary screening) on ​​the raw data (patient data). This initial screening generates a feature subset. Typically, when developing a medical early warning scoring system, a few features that may be effective in predicting a patient's condition must be manually selected. However, to ensure the final scoring system is concise, efficient, and user-friendly, this initial feature set must be screened to select the most representative feature subset. Typically, the scoring system is constructed by statistically analyzing the correlation between each feature and the outcome, selecting the most relevant features. Inter-feature correlations are also considered, eliminating redundant features. Furthermore, the accessibility of each feature is evaluated to ensure that the scoring system only includes easily accessible features. Feature screening yields a highly informative and user-friendly feature subset, enabling the construction of a simple and effective medical early warning scoring system. The initial feature screening of the original data in the present invention includes: screening based on the feature missing rate, eliminating features with a missing rate greater than 90%; calculating the upper and lower boundaries of the value based on the quartile distribution method, counting the number of outliers outside the boundaries, and eliminating features with an outlier ratio greater than 90%; counting the repeated values ​​in the variables, and eliminating features with a single repeated value greater than 90%. The formula for calculating the upper and lower boundaries of the value is: , Where UP is the upper bound, LB is the lower bound, Q1 is the 25th percentile, Q3 is the 75th percentile, and IQR is the interquartile range. After the initial feature screening, the variables can be binned and encoded. The encoded features are all numerical, and then variable screening (re-screening) is performed. In order to make the model more robust and at the same time facilitate the ease of use of the scoring system, such as Figure 2As shown, this solution performs feature binning on the features after the initial screening. The feature binning module performs binning processing on the feature subset; the method of binning the feature subset is to first determine the variable type of the feature subset. When the feature subset belongs to a numerical variable, the equal-interval binning method or the equal-frequency binning method is performed according to the data distribution. When the feature subset belongs to a categorical variable, the variables with more categories and smaller sample sizes are directly merged. Among them, the equal-interval binning method is to divide the value range into a fixed number of bin groups, so that each bin group covers the same value range; the equal-frequency binning method is to divide the sample distribution into a fixed number of bin groups, so that each bin group contains the same number of samples. The two methods have their own advantages and disadvantages. The equal-interval binning method is easy to handle outliers, but may produce bin groups with uneven sample distribution. The equal-frequency binning method can make the samples of each bin group evenly distributed, but it is easily affected by outliers. The present invention selects a suitable binning method according to the distribution of variable values. For categorical variables, the present invention adopts the method of merging small categories. For features with many categories but small sample sizes in some small categories, merging these small categories can prevent overfitting and reduce feature dimensions. The present invention pre-sets a category sample ratio threshold, merges categories according to the threshold, and retains the main categories with large sample sizes. The binning method of the present invention can effectively reduce the dimensionality of features and discretize them, improve the generalization ability and interpretability of the model, and facilitate the use of subsequent scoring systems. The two binning methods are used in combination to process different types of features and build a robust and efficient scoring system.

[0029] The feature encoding module uses the WOE encoding method to encode the category features after binning; in order to further improve the performance of the model, the present invention uses the WOE (Weight of Evidence) encoding method for the category features after binning. WOE encoding can reflect the predictive ability of different feature values ​​for the target variable. Compared with the traditional 0 / 1 encoding, WOE encoding can reflect the intrinsic correlation between the feature value and the target variable. Compared with label encoding, which only performs simple sorting encoding, it takes into account the logical relationship between the feature value and the target variable, and can also cope with unknown categories. The calculation principle of WOE is to calculate the proportion of positive samples in all positive samples in each box, and then calculate the proportion of negative samples in each box, and then divide the two proportions and take the logarithm to complete the encoding. The calculation formula is: , in, is the proportion of positive samples in all positive samples, is the proportion of negative samples in all negative samples; the binning and WOE coding scheme can deeply explore the intrinsic patterns of features and targets, effectively improve the feature effect, and develop a more accurate and interpretable scoring system to achieve disease early warning and risk prediction.

[0030] The feature screening module re-screens the encoded data; feature re-screening includes: using the correlation calculation method to calculate the correlation between all variables, and if the absolute value of the correlation is greater than 0.8, one variable is eliminated; using the tree model to screen variables, calculating the feature information gain ratio as the importance coefficient at each tree node split, sorting according to the feature importance score, and selecting the top 90% with the highest score; after feature re-screening, a very core feature subset can be obtained to be used as model input data.

[0031] The model building module builds the model and uses the Lasso regression model to automatically select features after feature re-screening. The objective function in the Lasso regression model is as follows: , Among them, β is the characteristic coefficient of the model; n is the number of samples; p is the number of features; X is the feature matrix, y is the target variable, and λ is the regularization parameter. The optimal value is determined by methods such as cross-validation.

[0032] The Lasso regression model uses the encoded features as input and constructs a regression model through the L1 regularization method. The main advantage of the Lasso regression model is that it can achieve feature selection while building the regression model. It uses L1 regularization to reduce some feature coefficients to 0, thereby achieving the effect of eliminating non-critical features. This simplifies the model construction process. At the same time, it reduces the overfitting of the model. During the model training process, the Lasso regression model will penalize the feature coefficients so that unimportant feature coefficients gradually approach 0, thereby achieving the purpose of feature selection. This feature selection method can not only reduce the complexity of the model, but also improve the generalization ability of the model, so that the model has better predictive performance on new data sets. In addition, the Lasso regression model also has good interpretability. The size of the feature coefficient can be used to judge the degree of influence of each feature on the target variable, providing strong support for the medical early warning scoring system.

[0033] The scoring mapping module calculates the total warning score and maps the probability value output by the model to the clinical scoring system to provide the scoring result. The scoring mapping module introduces the ODDs-PDO method. On the one hand, it performs scoring calculations and calculates the total warning score. On the other hand, it performs scoring mapping and splits the total score into scores corresponding to each indicator through the model weight. The ODDs-PDO method is a mapping method based on probability and scoring, which is commonly used in financial risk assessment and medical warning scoring systems. It makes the scoring more interpretable and operational by mapping the output probability value of the model to the scoring system. Specifically, the ODDs-PDO method is used to calculate the total warning score and split the total score into scores corresponding to each indicator through the model weight. The calculation formula for the total warning score S is: , Where p is the probability value output by the model, S0 is the baseline score, and k is the adjustment factor used to control the sensitivity of the score. By adjusting the values ​​of S0 and k, you can flexibly map the score range corresponding to different probability values ​​to meet different application needs.

[0034] In addition, in order to further refine the scoring results, the total score is split into scores corresponding to each indicator by model weight. Assuming that there are n features in the model, the weight of each feature is w1, w2, ..., wn, and the corresponding feature value is x1, x2, ..., xn, then the score of each feature is: , Here, μi is the mean of the i-th feature, and σi is the standard deviation of the i-th feature. This splitting method not only intuitively demonstrates the contribution of each indicator to the total score, but also facilitates medical staff to quickly identify the key factors of the patient's condition. The score mapping module not only provides an intuitive total score, but also provides medical staff with more detailed condition information by splitting the scores of each indicator item. At the same time, the ODDs-PDO method has good interpretability, allowing medical staff to understand the source and basis of the scoring results, so that they can be better applied in clinical practice. Features with poor feature importance are returned for variable re-screening, further improving the effect of eliminating redundant features.

[0035] In summary, the method disclosed in the present invention adopts a two-stage scheme of primary screening and re-screening in sequence to effectively remove redundant features, retain the most representative feature subsets, and ensure that the scoring system is concise and efficient. For numerical and categorical variables, the equal-interval binning method, the equal-frequency binning method and the small category merging method are adopted to improve the generalization ability and interpretability of the model. The WOE encoding method is used to reflect the intrinsic correlation between the eigenvalue and the target variable, enhance the feature effect, and build a more accurate and interpretable scoring system. Lasso regression is used to achieve feature selection through L1 regularization, simplify the model construction process, reduce the risk of overfitting, and improve the generalization ability of the model. The ODDs-PDO method is introduced to map the probability value output by the model to the scoring system, providing intuitive, detailed and interpretable scoring results for easy clinical application.

[0036] The overall solution covers the complete process from feature screening, binning, coding, model building to score mapping, forming a systematic medical early warning scoring system construction plan with good interpretability and practicality.

[0037] Compared with existing technologies, the technical advantages are significant: A full-process systematic construction was achieved, and a method for building a medical early warning scoring system based on feature screening, binning, coding, model building and score mapping was proposed. Compared with traditional methods based on expert experience or simple modeling, it has stronger systematicity and accuracy.

[0038] It has feature processing optimization and proposes feature screening and binning technology solutions for medical data, which can effectively remove redundant features and process different types of features. Compared with existing technologies, it significantly improves feature quality and model performance.

[0039] It has the advantage of feature encoding. The WOE encoding method can better reflect the intrinsic correlation between features and target variables. Compared with encoding methods such as One-Hot, it can obtain richer feature information and improve the accuracy of the model.

[0040] It has high efficiency in model construction. Through the Lasso regression model construction method, it automatically performs feature selection and reduces the risk of overfitting. Compared with existing simple modeling methods, it can more effectively handle the correlation between features and improve the generalization ability and interpretability of the model.

[0041] With the intuitiveness of scoring mapping, an ODDs-PDO scoring mapping method was designed to map the probability values ​​output by the model into the scoring system, providing intuitive, detailed and interpretable scoring results. Compared with existing scoring mapping methods, it has better interpretability and operability, and is convenient for clinical application.

[0042] With full process automation, the produced model can output scoring results throughout the entire process without the need for additional manual processing, directly serving clinical diagnosis and treatment, helping doctors to judge the condition more accurately, and having higher application value.

[0043] The above embodiments are not limitations of the present invention, and the present invention is not limited to the above examples. Any changes, modifications, additions or substitutions made by technicians in this technical field within the scope of the technical solution of the present invention also fall within the scope of protection of the present invention.

Claims

1. A method for building a universal medical early warning scoring system based on machine learning, characterized in that: The method comprises the following steps: Step 1: Get feature subsets by preliminarily screening the features of the original data; Step 2: binning the feature subsets; Step 3: Encode the category features after binning using the WOE encoding method; Step 4: re-screen the encoded data based on its features; Step 5: Use the Lasso regression model to automatically select features from the data after feature re-screening; Step 6: Calculate the total warning score and map the probability value output by the model to the clinical scoring system to provide the scoring result.

2. The method for building a universal medical early warning scoring system based on machine learning according to claim 1, characterized in that: The initial feature screening of the original data includes: Screening is performed based on the feature missing rate, and features with a missing rate greater than 90% are eliminated; The upper and lower bounds of the values ​​are calculated based on the quartile spread method, the number of outliers outside the bounds is counted, and features with an outlier ratio greater than 90% are removed; Count the repeated values ​​in the variables and remove features with a single repeated value greater than 90%.

3. The method for building a universal medical early warning scoring system based on machine learning according to claim 2, characterized in that: The formula for calculating the upper and lower bounds of the numerical value is: , Among them, UP is the upper boundary; LB is the lower boundary; Q1 is the 25% quantile, Q3 is the 75% quantile, and IQR is the interquartile range.

4. The method for constructing a universal medical early warning scoring system based on machine learning according to claim 1, characterized in that: The method for binning feature subsets is to first determine the variable type of the feature subset. When the feature subset belongs to a numerical variable, equal interval binning or equal frequency binning is performed according to the data distribution. When the feature subset belongs to a categorical variable, variables with more categories and smaller sample sizes are directly merged.

5. The method for constructing a universal medical early warning scoring system based on machine learning according to claim 4, characterized in that: The equal-interval binning method is to divide the value range into a fixed number of bin groups, so that each bin group covers the same value range; the equal-frequency binning method is to divide the sample distribution into a fixed number of bin groups, so that each bin group contains the same number of samples.

6. The method for building a universal medical early warning scoring system based on machine learning according to claim 5, characterized in that: The feature rescreening includes: The correlation between all variables was calculated using the correlation calculation method. If the absolute value of the correlation was greater than 0.8, a variable was eliminated. Use the tree model for variable screening, calculate the feature information gain ratio as the importance coefficient at each tree node split, sort the features according to their importance scores, and select the top 90% with the highest scores.

7. The method for building a universal medical early warning scoring system based on machine learning according to claim 1, characterized in that: The objective function of the Lasso regression model in step 5 is as follows: , Among them, β is the characteristic coefficient of the model; n is the number of samples; p is the number of features; X is the feature matrix, y is the target variable, and λ is the regularization parameter.

8. The method for building a universal medical early warning scoring system based on machine learning according to claim 1, characterized in that: The step six uses the ODDs-PDO method to calculate the total warning score and splits the total score into scores corresponding to each indicator through the model weight.

9. The method for constructing a universal medical early warning scoring system based on machine learning according to claim 8, characterized in that: The calculation formula of the warning total score S is: , Among them, p is the probability value output by the model, S0 is the benchmark score, and k is the adjustment factor.

10. The method for building a universal medical early warning scoring system based on machine learning according to claim 9, characterized in that: The method of splitting the total score into scores corresponding to each indicator by model weight is to assume that there are n features in the model, the weight of each feature is w1, w2, ..., wn, and the corresponding feature value is x1, x2, ..., xn. Then the score of each feature is: , Among them, μi is the mean of the i-th feature, and σi is the standard deviation of the i-th feature.