A clinical data augmentation method based on mixed feature probability constraints
Patent Information
- Application Number
- CN202611141403.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-30
- Publication Date
- 2026-09-29
AI Technical Summary
[0007]本发明提供一种基于混合特征概率约束的临床数据增强方法,用以解决临床数据中少数类样本数量不足、类别分布不平衡以及传统数据增强方法难以同时适用于连续特征与分类特征的问题
本发明结合连续特征局部概率生成、分类特征条件概率生成以及可信概率筛选技术,实现CLABSI少数类样本的结构化生成与数据增强,具有生成数据真实性高、特征关联保持能力强以及适用于小样本医学临床数据等优点,可有效提升CLABSI风险预测模型对少数类样本的识别性能。
Smart Images

Figure CN122842971A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical artificial intelligence and medical data processing technology, specifically relating to a clinical data augmentation method based on hybrid feature probability constraints. Background Technology
[0002] With the development of medical artificial intelligence technology, machine learning-based clinical auxiliary diagnosis and disease risk prediction methods have gradually become a research hotspot, and their predictive performance highly depends on the support of high-quality clinical data. However, in actual clinical scenarios, the number of positive or abnormal cases is usually far less than the number of negative samples, leading to a significant class imbalance problem in clinical datasets. Among these, central line-associated bloodstream infection (CLABSI) is one of the most common serious hospital-acquired infections in clinical intensive care, oncology treatment, and long-term intravenous infusion. Its occurrence not only prolongs patients' hospital stays but can also lead to sepsis, multiple organ dysfunction, and even death, representing a typical application scenario with small sample sizes and imbalanced clinical data. Therefore, early risk prediction and auxiliary diagnosis of diseases such as CLABSI have significant clinical implications.
[0003] However, in real-world clinical data, due to the limited number of minority class samples, traditional classification models tend to favor the majority class during training, resulting in a decreased ability to identify positive disease samples and affecting the model's accuracy and generalization performance. Therefore, how to effectively augment minority class samples in clinical data has become an important research direction for improving the performance of medical prediction models.
[0004] Existing data augmentation methods mainly include random oversampling, SMOTE (Synthetic Minority Over-sampling Technique), Borderline-SMOTE, and SMOTE-NC. Random oversampling achieves class balance by simply replicating minority class samples, which can easily lead to model overfitting. SMOTE-like methods typically use linear interpolation between minority class samples to generate new samples; while this increases the number of minority class samples, it lacks constraints on the true data distribution and can easily generate anomalous samples that cross class boundaries. Borderline-SMOTE, although it can enhance the generation of samples in boundary regions, is sensitive to noisy samples and can easily amplify anomalous distributions. While SMOTE-NC considers the coexistence of continuous and categorical features, it typically uses a simple majority voting method to generate categorical features, making it difficult to maintain the true probabilistic relationships and combinational patterns between clinical categorical variables.
[0005] Furthermore, CLABSI clinical data typically contains both continuous and categorical features, making it a typical example of mixed-feature medical data. Most existing methods employ a unified generation mechanism to handle different types of features, making it difficult to simultaneously maintain the local statistical characteristics of continuous variables and the true class distribution of categorical variables. This can easily lead to problems such as distorted feature combinations, abnormal class pairings, and probability distribution shifts in the generated samples, thereby reducing the realism of the augmented data and the performance of the classification model.
[0006] Therefore, there is an urgent need for a CLABSI data augmentation method that can simultaneously adapt to a mixture of continuous and categorical features and can combine sample credibility constraints for generation and screening, so as to improve the authenticity, stability and clinical usability of the generated samples, thereby enhancing the ability of the CLABSI risk prediction model to identify minority infection samples. Summary of the Invention
[0007] This invention provides a clinical data augmentation method based on hybrid feature probability constraints to address the problems of insufficient minority class samples, imbalanced class distribution, and the difficulty of applying traditional data augmentation methods to both continuous and categorical features simultaneously in clinical data.
[0008] This invention is achieved through the following technical solution: A clinical data augmentation method based on hybrid feature probability constraints, taking the CLABSI dataset as an example, includes the following steps: Step S1: Obtain the original clinical data of CLABSI, and perform missing value processing, classification feature encoding and continuous feature standardization on the original clinical data to obtain preprocessed mixed feature data; Step S2: Count the number of samples in each category based on the category labels in the original clinical data, identify the CLABSI infection positive category with a smaller number of samples as the minority class, extract minority class samples from the mixed feature data obtained in Step S1, and construct the CLABSI minority class sample set. Step S3: For the continuous features in the minority class sample set constructed in Step S2, a local neighborhood probability generation method is used to generate new continuous feature data, including four processes: local neighborhood construction, neighborhood sample selection, random interpolation generation, and probability constraint perturbation. Step S4: For the categorical features in the minority class sample set constructed in Step S2, new categorical feature data are generated using the categorical feature conditional probability generation method, which includes three processes: category probability statistics, conditional random sampling, and categorical feature generation. Step S5: The continuous feature data generated in step S3 and the categorical feature data generated in step S4 are concatenated according to the feature order of the original samples to form a complete augmented sample, and all augmented samples are combined into a candidate augmented sample set. Step S6: Input the candidate augmentation samples constructed in step S5 into the confidence probability evaluation model based on random forest, calculate the predicted probability that the candidate augmentation samples belong to the CLABSI infection category, and filter out low confidence samples and retain high confidence augmentation samples according to the set probability threshold. Step S7: Based on the classification feature encoding method in Step S1, the encoded classification features in the enhanced samples after screening in Step S6 are reverse encoded and restored to the original category variables; then, the restored enhanced samples are merged with the original clinical data according to the original data structure to construct the enhanced CLABSI balanced dataset.
[0009] Furthermore, in step S1, mean normalization and standard deviation scaling are used to eliminate the dimensional differences between different continuous features and to standardize the continuous features. In step S1, a label encoding method is used to convert discrete categorical variables into numerical forms to encode the categorical features.
[0010] Furthermore, in step S3, after generating continuous features through random proportional interpolation, a Beta probability perturbation based on local statistical distribution is introduced to impose local probability constraints on the interpolation results, thereby enhancing the local diversity of the generated samples and maintaining the continuity of the continuous feature distribution.
[0011] Furthermore, step S3, which involves constructing a local neighborhood probability generation model, specifically includes the following steps: Step S31: Constructing a local neighborhood: The K-nearest neighbor algorithm is used to calculate the local neighborhood of the target minority class sample, and the k nearest minority class samples are selected as neighborhood samples to provide a local reference for continuous feature generation; Step S32: Select a reference sample: Randomly select a neighborhood sample from the local neighborhood corresponding to the target sample as a reference sample; Step S33: Continuous Feature Interpolation Generation: Based on the continuous feature differences between the target sample and the reference sample, random proportional interpolation is performed to generate new continuous features. The calculation formula is as follows:
[0012] in, For the target minority class samples, As a reference sample in the local neighborhood, These are random interpolation coefficients, and 0 < <1, Continuous features generated by interpolation; Step S34: Probabilistic Constraint Perturbation: Based on the statistical distribution of the local neighborhood of the target sample, a Beta probability perturbation is introduced into the interpolation result to achieve probabilistic constraints on continuous features. The calculation formula is as follows:
[0013] in,
[0014] in, This results in the final continuous feature; Generate results for continuous feature interpolation; The random perturbation coefficients follow a Beta distribution; The standard deviation of continuous features in the local neighborhood of the target sample; and These are parameters of the Beta distribution; S35: Output continuous enhancement features; repeat steps S31 to S34 until the preset number of enhancements is reached, and output the generated continuous enhancement features.
[0015] Furthermore, step S4 specifically includes the following steps: Step S41: Category Probability Statistics; Statistically analyze the frequency of occurrence of different categories for each classification feature in the minority class samples, and calculate the conditional probability corresponding to each category. The calculation formula is as follows:
[0016] in, Category in classification features The conditional probability; For category The number of times it appears in the minority class samples; The total number of samples with this classification feature in the minority class; construct the conditional probability distribution of the classification feature based on the above calculation results; Step S42: Conditional random sampling; Random sampling is performed based on the conditional probability distribution of classification features, and classification feature values are randomly generated according to the probability corresponding to each category, so that the generated results maintain the category distribution characteristics of the original minority class samples; Step S43: Generate conditional probabilities of classification features; combine the classification features obtained in step S42 to form new classification features, and combine them with the continuous features generated in step S3 to form candidate enhanced samples.
[0017] Furthermore, step S5 specifically includes the following steps: Step S51: Based on the feature arrangement order in the original clinical data, fill the continuous features into the corresponding continuous feature positions and fill the categorical features into the corresponding classification feature positions to form a complete enhanced sample; Step S52: Repeat step S51 until all generated continuous features and categorical features are combined accordingly to construct a candidate augmentation sample set; The process of constructing the enhanced samples is represented as follows:
[0018] in, For the generated continuous feature vector, For the generated categorical feature vectors, The combined enhanced sample; All augmented samples constitute the candidate augmented sample set:
[0019] in, The number of augmented samples generated.
[0020] Furthermore, step S6 specifically involves eliminating candidate augmented samples that are below a set infection probability threshold, in order to reduce the impact of abnormally generated samples and samples that cross the category boundary on the classification model.
[0021] Furthermore, in step S6, the credibility probability screening model is constructed using a random forest classification model. The credibility assessment of the generated samples is specifically performed by calculating the predicted probability that the candidate augmented sample belongs to the CLABSI infection category: Step S61: Train a random forest classification model using the original CLABSI clinical data to obtain a classifier that can output the predicted probability of CLABSI infection; Step S62: Input the candidate augmentation samples constructed in step S5 into the trained random forest classification model, and calculate the predicted probability that the candidate augmentation samples belong to the CLABSI infection category; Step S63: Based on the preset probability threshold The candidate augmentation samples are screened, and when the candidate samples > When, retain the candidate sample; when < When this happens, the candidate augmented sample is removed; Step S64: The retained candidate augmentation samples are combined into a high-confidence augmentation sample set for subsequent construction of the augmented CLABSI balanced dataset.
[0022] A data augmentation system for central venous access bloodstream infection based on hybrid feature probability constraints, the system using the aforementioned data augmentation method for central venous access bloodstream infection based on hybrid feature probability constraints, the system comprising: Preprocessing module: Acquires CLABSI raw clinical data and performs missing value processing, categorical feature encoding, and continuous feature standardization on the raw clinical data to obtain preprocessed mixed feature data; CLABSI Minority Class Sample Set Construction Module: This module is used to count the number of samples in each category based on the CLABSI category labels in the preprocessed mixed feature data, determine the CLABSI infection positive category as the minority class, and extract the minority class samples as data augmentation objects to construct the CLABSI minority class sample set. Continuous Feature Processing Module: This module generates continuous feature data for the constructed minority class sample set using a local neighborhood probability generation method based on K-nearest neighbor search, random proportional interpolation, and Beta probability perturbation. Specifically, the K-nearest neighbor algorithm is used to construct the local neighborhood of the target minority class sample. Random proportional interpolation is performed based on the continuous feature differences between the target sample and the reference sample. Beta probability perturbation is introduced in combination with local statistical distribution to impose local probability constraints on the interpolation results, thereby generating continuous enhanced features.
[0023] The categorical feature processing module is used to establish a probability mapping relationship based on the category distribution of each categorical feature in the minority class sample set, and to generate new categorical feature data based on the probability mapping relationship. Specifically, it includes: (1) Statistically determine the frequency of occurrence of each category of the target classification feature in the minority class samples, and establish the corresponding category probability distribution. The calculation formula is as follows:
[0024] in, Category in classification features The conditional probability; For category The number of times it appears in the minority class samples; This represents the total number of samples in the minority class for this classification feature. Based on the above calculations, construct the conditional probability distribution for the classification feature.
[0025] (2) Construct a random sampling space based on the established category probability distribution, and randomly generate classification feature values according to the probability corresponding to each category, so that the generated result maintains the category distribution characteristics consistent with the original minority class samples.
[0026] (3) Perform steps (1) and (2) on each classification feature to obtain the complete classification feature vector and output the generated classification feature data; Candidate augmentation sample set construction module: It is used to fill the corresponding feature positions with continuous feature data and categorical feature data according to the feature arrangement order of the original clinical data to form a complete augmentation sample, and to form a candidate augmentation sample set by combining all augmentation samples; The sample selection module is used to input candidate augmented samples into a pre-trained random forest classification model, calculate the predicted probability that the candidate augmented sample belongs to the CLABSI infection category, compare the predicted probability with a preset probability threshold, and retain the candidate augmented sample when the predicted probability is greater than or equal to the preset probability threshold; otherwise, the candidate augmented sample is removed, thereby obtaining a high-confidence augmented sample. The CLABSI balanced dataset construction module is used to perform inverse encoding recovery on the encoded classification features in the selected augmented samples based on the classification feature encoding mapping relationship established by the preprocessing module; the recovered classification features and continuous features are combined to form a complete augmented sample, and then merged with the original clinical data according to the feature structure of the original clinical data to construct the augmented CLABSI balanced dataset.
[0027] A central venous access bloodstream infection data augmentation method based on hybrid feature probability constraints, as described above, is applied to the scenario of CLABSI minority clinical data augmentation.
[0028] The beneficial effects of this invention are: This invention combines continuous feature local probability generation, classification feature conditional probability generation, and reliable probability screening techniques to achieve structured generation and data augmentation of CLABSI minority class samples. It has the advantages of high data authenticity, strong feature association preservation ability, and applicability to small sample medical clinical data, and can effectively improve the recognition performance of CLABSI risk prediction model for minority class samples.
[0029] This invention generates continuous features of the minority class by using a probabilistic interpolation method based on the K-nearest neighbor local structure. This method generates samples that maintain the local distribution pattern of the original minority class data, thereby avoiding the continuous feature shift problem caused by traditional global linear interpolation methods and improving the consistency between the generated data and real clinical data.
[0030] This invention generates new classification feature data by performing conditional probability statistics on classification features and based on the category distribution probability in minority class samples. This effectively maintains the category distribution features and feature associations in the original clinical data, reduces the risk of generating abnormal category combinations, and improves the authenticity and rationality of the generated data. This invention introduces a reliable probability screening mechanism based on a random forest classification model to evaluate the minority class probability of generated samples and remove abnormal generated samples with low reliability, thereby reducing the impact of noisy data and error patterns on subsequent model training and improving the stability and reliability of augmented data.
[0031] This invention organically integrates multiple techniques, including continuous feature local probability generation, classification feature conditional probability generation, and reliable probability screening. It not only solves key problems in the CLABSI mixed feature data augmentation process, such as continuous feature distortion, classification feature anomalies, and low-quality sample diffusion, but also constructs a complete data augmentation scheme suitable for small-sample medical clinical data. This scheme can significantly improve the CLABSI risk prediction model's ability to identify minority class samples, its model generalization performance, and the reliability of clinical auxiliary prediction. Attached Figure Description
[0032] Figure 1 This is a schematic diagram of the method flow of the present invention.
[0033] Figure 2 This is a schematic diagram comparing the feature distribution of the original data and the augmented data of the present invention, wherein (a) is the PCA dimensionality reduction distribution diagram of the original data, (b) is the PCA dimensionality reduction distribution diagram of the original data, (c) is the U-MAP dimensionality reduction distribution diagram of the original data, (d) is the PCA dimensionality reduction distribution diagram of the augmented data, (e) is the T-SNE dimensionality reduction distribution diagram of the augmented data, and (f) is the U-MAP dimensionality reduction distribution diagram of the augmented data.
[0034] Figure 3 These are radar charts comparing the performance of various classification models in terms of accuracy, precision, recall, F1 score, and AUC before and after data augmentation in this invention. Among them, (a) is a radar chart comparing the performance of logistic regression in terms of accuracy, precision, recall, F1 score, and AUC; (b) is a radar chart comparing the performance of random forest in terms of accuracy, precision, recall, F1 score, and AUC; (c) is a radar chart comparing the performance of support vector machine in terms of accuracy, precision, recall, F1 score, and AUC; (d) is a radar chart comparing the performance of decision tree in terms of accuracy, precision, recall, F1 score, and AUC; (e) is a radar chart comparing the performance of gradient boosting in terms of accuracy, precision, recall, F1 score, and AUC; and (f) is a radar chart comparing the performance of XGBoost in terms of accuracy, precision, recall, F1 score, and AUC. Detailed Implementation
[0035] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of this application with unnecessary detail.
[0036] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0037] It should also be understood that the terminology used in this application specification is for the purpose of describing particular embodiments only and is not intended to limit the application. As used in this application specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0038] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0039] Many specific details are set forth in the following description in order to provide a full understanding of this application. However, this application may also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the spirit of this application. Therefore, this application is not limited to the specific embodiments disclosed below.
[0040] A clinical data augmentation method based on hybrid feature probability constraints, such as Figure 1 As shown, the method includes the following steps: Step S1: Obtain the original clinical data of CLABSI, and perform missing value processing, classification feature encoding and continuous feature standardization on the original clinical data to obtain preprocessed mixed feature data; Step S2: Based on the category labels in the original clinical data, count the number of samples in each category. Let the number of samples in each category be... The class with the fewest samples is identified as the minority class; for binary classification tasks, when the minority class... With the majority class The ratio of sample size satisfies
[0041] The dataset is determined to be an imbalanced dataset, and the minority class is taken as the category to be enhanced. Minority class samples are extracted from the mixed feature data obtained in step S1 to construct the CLABSI minority class sample set. Step S3: For the continuous features in the minority class sample set constructed in Step S2, a local neighborhood probability generation method is used to generate new continuous feature data, including four processes: local neighborhood construction, neighborhood sample selection, random interpolation generation, and probability constraint perturbation.
[0042] Step S4: For the categorical features in the minority class sample set constructed in Step S2, new categorical feature data are generated using the categorical feature conditional probability generation method, which includes three processes: category probability statistics, conditional random sampling, and categorical feature generation. Step S5: The continuous feature data generated in step S3 and the categorical feature data generated in step S4 are concatenated according to the feature order of the original samples to form a complete augmented sample, and all augmented samples are combined into a candidate augmented sample set.
[0043] Step S6: Input the candidate augmentation samples constructed in step S5 into the confidence probability evaluation model based on random forest, calculate the predicted probability that the candidate augmentation samples belong to the CLABSI infection category, and filter out low confidence samples and retain high confidence augmentation samples according to the set probability threshold.
[0044] Step S7: Based on the classification feature encoding method in Step S1, the encoded classification features in the enhanced samples after screening in Step S6 are reverse encoded and restored to the original category variables; then, the restored enhanced samples are merged with the original clinical data according to the original data structure to construct the enhanced CLABSI balanced dataset.
[0045] Furthermore, in step S1, mean normalization and standard deviation scaling are used to eliminate the dimensional differences between different continuous features and to standardize the continuous features. In step S1, a label encoding method is used to convert discrete categorical variables into numerical forms to encode the categorical features.
[0046] Furthermore, in step S3, after generating continuous features through random proportional interpolation, a Beta probability perturbation based on local statistical distribution is introduced to impose local probability constraints on the interpolation results, thereby enhancing the local diversity of the generated samples and maintaining the continuity of the continuous feature distribution.
[0047] Furthermore, the construction of the local neighborhood probability generation model in step S3 specifically involves: Step S31: Construct local neighborhoods. The K-nearest neighbor algorithm is used to calculate the local neighborhoods of the target minority class sample, and the k nearest minority class samples are selected as neighborhood samples to provide local references for continuous feature generation.
[0048] Step S32: Select a reference sample. Randomly select a neighborhood sample from the local neighborhood corresponding to the target sample as a reference sample.
[0049] Step S33: Continuous Feature Interpolation Generation. Based on the continuous feature differences between the target sample and the reference sample, random proportional interpolation is performed to generate new continuous features. The calculation formula is as follows:
[0050] in, For the target minority class samples, As a reference sample in the local neighborhood, These are random interpolation coefficients, and 0 < <1, Continuous features generated by interpolation.
[0051] Step S34: Probabilistic Constraint Perturbation. Based on the statistical distribution of the local neighborhood of the target sample, a Beta probability perturbation is introduced into the interpolation result to achieve probabilistic constraints on continuous features. The calculation formula is as follows:
[0052] in,
[0053] in, This results in the final continuous feature; Generate results for continuous feature interpolation; The random perturbation coefficients follow a Beta distribution; The standard deviation of continuous features in the local neighborhood of the target sample; and These are parameters of the Beta distribution.
[0054] Step S35: Output continuous augmented features. Repeat steps S31 to S34 until the preset number of augmentations is reached, and output the generated continuous augmented features.
[0055] Furthermore, in step S4, the conditional probability distribution is obtained by statistically analyzing the occurrence frequency of different categories of each classification variable in the minority class samples, and probability sampling is performed based on the occurrence frequency to generate classification features.
[0056] The specific process is as follows: Step S41: Category Probability Statistics. Statistically analyze the frequency of occurrence of different categories for each classification feature in the minority class samples, and calculate the conditional probability corresponding to each category. The calculation formula is as follows:
[0057] in, Category in classification features The conditional probability; For category The number of times it appears in the minority class samples; This represents the total number of samples in the minority class for this classification feature. Based on the above calculations, construct the conditional probability distribution for the classification feature.
[0058] Step S42: Conditional random sampling. Random sampling is performed based on the conditional probability distribution of the classification features. Classification feature values are randomly generated according to the probability corresponding to each category, so that the generated results maintain the category distribution characteristics of the original minority class samples.
[0059] Step S43: Generation of conditional probabilities for classification features. The classification features obtained in step S42 are combined to form new classification features, and then combined with the continuous features generated in step S3 to form candidate augmented samples.
[0060] Furthermore, in step S5, the candidate augmentation sample set is constructed by concatenating the generated continuous feature data and the generated categorical feature data according to their feature dimensions. The specific construction process is as follows: The candidate augmentation sample set is constructed by concatenating the generated continuous feature data with the generated categorical feature data according to the feature dimension. The specific construction process is as follows: Step S51: Based on the feature arrangement order in the original clinical data, fill the continuous features into the corresponding continuous feature positions and fill the categorical features into the corresponding classification feature positions to form a complete enhanced sample; Step S52: Repeat step S51 until all generated continuous features and categorical features are combined to construct a candidate augmentation sample set.
[0061] The process of constructing the enhanced samples is represented as follows:
[0062] in, For the generated continuous feature vector, For the generated categorical feature vectors, This is the enhanced sample after combination.
[0063] All augmented samples constitute the candidate augmented sample set:
[0064] in, The number of augmented samples generated.
[0065] Furthermore, in step S6, the credibility probability screening model is constructed using a random forest classification model. The credibility of the generated samples is evaluated by calculating the predicted probability that the candidate enhanced sample belongs to the CLABSI infection category. In step S6, candidate augmented samples below the set infection probability threshold are removed to reduce the impact of abnormally generated samples and samples that cross the category boundary on the classification model.
[0066] Furthermore, in step S6, the credibility probability screening model is constructed using a random forest classification model. The credibility assessment of the generated samples is specifically performed by calculating the predicted probability that the candidate augmented sample belongs to the CLABSI infection category: Step S61: Train a random forest classification model using the original CLABSI clinical data to obtain a classifier that can output the predicted probability of CLABSI infection; Step S62: Input the candidate augmentation samples constructed in step S5 into the trained random forest classification model, and calculate the predicted probability that the candidate augmentation samples belong to the CLABSI infection category; Step S63: Based on the preset probability threshold The candidate augmentation samples are screened, and when the candidate samples > When, retain the candidate sample; when < When this happens, the candidate augmented sample is removed; Step S64: The retained candidate augmentation samples are combined into a high-confidence augmentation sample set for subsequent construction of the augmented CLABSI balanced dataset.
[0067] In this embodiment, the original CLABSI data includes continuous clinical indicators such as patient age, white blood cell count, neutrophil count, hemoglobin, platelets, lactate dehydrogenase, ferritin, and B-type natriuretic peptide, as well as subtyped clinical characteristics such as risk grading, intravenous catheterization, hormone use, and transfusion components. It also includes CLABSI positive and negative labels, with CLABSI positive samples being a minority of samples.
[0068] First, the raw CLABSI clinical data is preprocessed. Features in the raw clinical data are divided into continuous features and categorical features, and missing values in the continuous features are imputed. This imputation process uses existing data preprocessing techniques: for each continuous feature, the mean or median of that feature in the non-missing samples is calculated, and this mean or median is used to fill in the missing positions of the continuous feature, ensuring that subsequent feature standardization, local neighborhood calculation, and continuous feature generation can proceed normally. Then, the categorical features are numerically encoded for subsequent probability statistics and data generation. Finally, minority class samples are extracted based on the CLABSI labels to construct a minority class feature dataset.
[0069] Subsequently, continuous features in the minority class samples are standardized to reduce the impact of differences in feature dimensions on local distance calculation. Based on the standardized continuous feature data, local neighborhood relationships between minority class samples are constructed, and the K-nearest neighbor search method is used to obtain the local neighborhood sample set of each minority class sample.
[0070] After obtaining the local neighborhood structure, continuous features are generated using local probabilities. During the generation process, a minority class sample is first randomly selected as the generation base point, and neighboring samples from its local neighborhood are selected as reference samples. Subsequently, new continuous feature data is generated by performing local probability interpolation between the generation base point and neighboring samples, combined with random perturbation, so that the generation result can maintain the local distribution pattern of the original minority class samples.
[0071] Next, conditional probability statistics are performed on the categorical features in the minority class samples. Based on the class distribution probability of each categorical variable in the minority class samples, categorical features are randomly sampled and generated, thus generating new categorical feature data that conforms to the class distribution pattern of the original minority class data.
[0072] Subsequently, the generated continuous feature data and categorical feature data are combined according to the feature correspondence to construct a complete hybrid feature generation sample.
[0073] After the generated samples are constructed, a confidence probability screening is performed on the generated samples. First, a random forest classification model is trained using the original CLABSI clinical data; then, the generated samples are input into the classification model to obtain the predicted probability of each generated sample belonging to the CLABSI minority class; finally, the generated samples are screened according to a set probability threshold, retaining reliable generated samples with predicted probabilities higher than the threshold and removing low-confidence generated samples to reduce the impact of abnormal generated samples on subsequent model training.
[0074] Next, feature type restoration is performed on the filtered generated samples. For categorical features, the encoded values are restored to their corresponding original category form; for integer features, they are restored to integer data form to ensure consistency in data structure between the augmented data and the original clinical data.
[0075] Through the above process, CLABSI minority class augmented samples that satisfy local distribution patterns, classification feature correlations, and credible probability constraints are finally generated and fused with the original clinical data to form an augmented CLABSI balanced dataset, which is used for subsequent training of CLABSI risk prediction models.
[0076] To verify the effectiveness of the method of the present invention, the datasets before and after enhancement were visualized by feature dimensionality reduction and used for training and evaluation of six mainstream classification models.
[0077] From the perspective of dimensionality reduction visualization results, such as Figure 2 As shown, under the three methods of PCA, T-SNE, and U-MAP, the silhouette coefficients of the augmented data were significantly higher than those of the original data, increasing from 0.342 to 0.411, 0.351 to 0.510, and 0.373 to 0.550, respectively. In the visualized distribution, the boundaries between infected and non-infected samples were clearer, and the internal distribution of infected sample clusters was more uniform and continuous. This indicates that the augmented samples effectively supplemented the original minority class distribution space, without introducing abnormal samples that deviated from the true clinical characteristic distribution, and the data distribution quality was significantly improved.
[0078] After employing the data augmentation method based on hybrid feature probability constraints of this invention, the key performance indicators of each model were significantly improved. For example... Figure 3 As shown, on the expanded dataset, the recall of logistic regression, random forest, support vector machine, decision tree, gradient boosting, and XGBoost models increased from 0.333~0.556 in the original dataset to 0.843~0.941, and the F1 score increased from 0.429~0.645 to 0.819~0.897. Accuracy and precision also showed stable improvement, effectively solving the problem of insufficient model recognition of positive infection samples under imbalanced data with small samples. These results demonstrate that the enhanced samples generated by the method of this invention not only conform to the distribution patterns of clinical characteristics but also effectively improve the model's prediction accuracy and generalization ability for CLABSI infection, possessing good clinical application value.
[0079] Implementation Method 2 This embodiment provides a clinical data augmentation system based on hybrid feature probability constraints. The system uses the clinical data augmentation method based on hybrid feature probability constraints as described in Embodiment 1. The system includes: Preprocessing module: Acquires CLABSI raw clinical data and performs missing value processing, categorical feature encoding, and continuous feature standardization on the raw clinical data to obtain preprocessed mixed feature data; Minority class sample set construction module: It is used to count the number of samples in each category based on the CLABSI category label in the preprocessed mixed feature data, determine the CLABSI infection positive category as the minority class, and extract the minority class samples as data augmentation objects to construct the CLABSI minority class sample set; Continuous Feature Processing Module: This module generates continuous feature data for the constructed minority class sample set using a local neighborhood probability generation method based on K-nearest neighbor search, random proportional interpolation, and Beta probability perturbation. Specifically, the K-nearest neighbor algorithm is used to construct the local neighborhood of the target minority class sample. Random proportional interpolation is performed based on the continuous feature differences between the target sample and the reference sample. Beta probability perturbation is introduced in combination with local statistical distribution to impose local probability constraints on the interpolation results, thereby generating continuous enhanced features.
[0080] The categorical feature processing module is used to establish a probability mapping relationship based on the category distribution of each categorical feature in the minority class sample set, and to generate new categorical feature data based on the probability mapping relationship. Specifically, it includes: (1) Statistically determine the frequency of occurrence of each category of the target classification feature in the minority class samples, and establish the corresponding category probability distribution. The calculation formula is as follows:
[0081] in, Category in classification features The conditional probability; For category The number of times it appears in the minority class samples; This represents the total number of samples in the minority class for this classification feature. Based on the above calculations, construct the conditional probability distribution for the classification feature.
[0082] (2) Construct a random sampling space based on the established category probability distribution, and randomly generate classification feature values according to the probability corresponding to each category, so that the generated result maintains the category distribution characteristics consistent with the original minority class samples.
[0083] (3) Perform steps (1) and (2) on each classification feature to obtain the complete classification feature vector and output the generated classification feature data; Candidate augmentation sample set construction module: It is used to fill the corresponding feature positions with continuous feature data and categorical feature data according to the feature arrangement order of the original clinical data to form a complete augmentation sample, and to form a candidate augmentation sample set by combining all augmentation samples; The sample selection module is used to input candidate augmented samples into a pre-trained random forest classification model, calculate the predicted probability that the candidate augmented sample belongs to the CLABSI infection category, compare the predicted probability with a preset probability threshold, and retain the candidate augmented sample when the predicted probability is greater than or equal to the preset probability threshold; otherwise, the candidate augmented sample is removed, thereby obtaining a high-confidence augmented sample. The balanced dataset construction module is used to perform inverse encoding recovery on the encoded classification features in the screened enhanced samples based on the classification feature encoding mapping relationship established in step S1; the restored classification features and continuous features are combined to form a complete enhanced sample, and then merged with the original clinical data according to the feature structure of the original clinical data to construct the enhanced CLABSI balanced dataset.
[0084] Implementation Method 3 This embodiment provides a clinical data augmentation method based on hybrid feature probability constraints as described in Embodiment 1, characterized in that the method is applied to scenarios involving the augmentation of minority class clinical data.
Claims
1. A clinical data augmentation method based on hybrid feature probability constraints, characterized in that, The method includes the following steps: Step S1: Obtain the original clinical data of CLABSI, and perform missing value processing, classification feature encoding and continuous feature standardization on the original clinical data to obtain preprocessed mixed feature data; Step S2: Count the number of samples in each category based on the category labels in the original clinical data, identify the CLABSI infection positive category with a smaller number of samples as the minority class, extract minority class samples from the mixed feature data obtained in Step S1, and construct the CLABSI minority class sample set. Step S3: For the continuous features in the minority class sample set constructed in Step S2, a local neighborhood probability generation method is used to generate new continuous feature data, including four processes: local neighborhood construction, neighborhood sample selection, random interpolation generation, and probability constraint perturbation. Step S4: For the categorical features in the minority class sample set constructed in step S2, new categorical feature data are generated using the categorical feature conditional probability generation method, which includes three processes: category probability statistics, conditional random sampling, and categorical feature generation. Step S5: The continuous feature data generated in step S3 and the categorical feature data generated in step S4 are concatenated according to the feature order of the original samples to form a complete augmented sample, and all augmented samples are combined into a candidate augmented sample set. Step S6: Input the candidate augmentation samples constructed in step S5 into the confidence probability evaluation model based on random forest, calculate the predicted probability that the candidate augmentation samples belong to the CLABSI infection category, and filter out low confidence samples and retain high confidence augmentation samples according to the set probability threshold. Step S7: Based on the classification feature encoding method in Step S1, the encoded classification features in the enhanced samples after screening in Step S6 are reverse encoded and restored to the original category variables; then, the restored enhanced samples are merged with the original clinical data according to the original data structure to construct the enhanced CLABSI balanced dataset.
2. The method according to claim 1, characterized in that, In step S1, mean normalization and standard deviation scaling are used to eliminate the dimensional differences between different continuous features and to standardize the continuous features. In step S1, a label encoding method is used to convert discrete categorical variables into numerical forms to encode the categorical features.
3. The method according to claim 1, characterized in that, In step S3, after generating continuous features by random proportional interpolation, a Beta probability perturbation based on local statistical distribution is introduced to impose local probability constraints on the interpolation results, so as to enhance the local diversity of the generated samples and maintain the continuity of the continuous feature distribution.
4. The method according to claim 1, characterized in that, Step S3, which constructs the local neighborhood probability generation model, specifically includes the following steps: Step S31: Constructing a local neighborhood: The K-nearest neighbor algorithm is used to calculate the local neighborhood of the target minority class sample, and the k nearest minority class samples are selected as neighborhood samples to provide a local reference for continuous feature generation; Step S32: Select a reference sample: Randomly select a neighborhood sample from the local neighborhood corresponding to the target sample as a reference sample; Step S33: Continuous Feature Interpolation Generation: Based on the continuous feature differences between the target sample and the reference sample, random proportional interpolation is performed to generate new continuous features. The calculation formula is as follows: in, For the target minority class samples, As a reference sample in the local neighborhood, These are random interpolation coefficients, and 0 < <1, Continuous features generated by interpolation; Step S34: Probabilistic Constraint Perturbation: Based on the statistical distribution of the local neighborhood of the target sample, a Beta probability perturbation is introduced into the interpolation result to achieve probabilistic constraints on continuous features. The calculation formula is as follows: in, in, This results in the final continuous feature; Generate results for continuous feature interpolation; The random perturbation coefficients follow a Beta distribution; The standard deviation of continuous features in the local neighborhood of the target sample; and These are parameters of the Beta distribution; S35: Output continuous enhancement features; repeat steps S31 to S34 until the preset number of enhancements is reached, and output the generated continuous enhancement features.
5. The method according to claim 1, characterized in that, Step S4 specifically includes the following steps: Step S41: Category Probability Statistics; Statistically analyze the frequency of occurrence of different categories for each classification feature in the minority class samples, and calculate the conditional probability corresponding to each category. The calculation formula is as follows: in, Category in classification features The conditional probability; For category The number of times it appears in the minority class samples; The total number of samples with this classification feature in the minority class; construct the conditional probability distribution of the classification feature based on the above calculation results; Step S42: Conditional random sampling; Random sampling is performed based on the conditional probability distribution of classification features, and classification feature values are randomly generated according to the probability corresponding to each category, so that the generated results maintain the category distribution characteristics of the original minority class samples; Step S43: Generate conditional probabilities of classification features; combine the classification features obtained in step S42 to form new classification features, and combine them with the continuous features generated in step S3 to form candidate enhanced samples.
6. The method according to claim 2, characterized in that, Step S5 specifically includes the following steps: Step S51: Based on the feature arrangement order in the original clinical data, fill the continuous features into the corresponding continuous feature positions and fill the categorical features into the corresponding classification feature positions to form a complete enhanced sample; Step S52: Repeat step S51 until all generated continuous features and categorical features are combined accordingly to construct a candidate augmentation sample set; The process of constructing the enhanced samples is represented as follows: in, For the generated continuous feature vector, For the generated categorical feature vectors, The combined enhanced sample; All augmented samples constitute the candidate augmented sample set: in, The number of augmented samples generated.
7. The method according to claim 1, characterized in that, Specifically, step S6 involves eliminating candidate augmented samples that are below a set infection probability threshold to reduce the impact of abnormally generated samples and samples that cross the category boundary on the classification model.
8. The method according to claim 7, characterized in that, In step S6, the reliability probability screening model is constructed using a random forest classification model. The reliability of the generated samples is evaluated by calculating the predicted probability that the candidate augmented sample belongs to the CLABSI infection category. Step S61: Train a random forest classification model using the original CLABSI clinical data to obtain a classifier that can output the predicted probability of CLABSI infection; Step S62: Input the candidate augmentation samples constructed in step S5 into the trained random forest classification model, and calculate the predicted probability that the candidate augmentation samples belong to the CLABSI infection category; Step S63: Based on the preset probability threshold The candidate augmentation samples are screened, and when the candidate samples > When, retain the candidate sample; when < When this happens, the candidate augmented sample is removed; Step S64: The retained candidate augmentation samples are combined into a high-confidence augmentation sample set for subsequent construction of the augmented CLABSI balanced dataset.
9. A clinical data augmentation system based on hybrid feature probability constraints, characterized in that, The system incorporates a clinical data augmentation method based on hybrid feature probability constraints as described in any one of claims 1-8, and the system comprises: Preprocessing module: Acquires CLABSI raw clinical data and performs missing value processing, categorical feature encoding, and continuous feature standardization on the raw clinical data to obtain preprocessed mixed feature data; CLABSI Minority Class Sample Set Construction Module: This module is used to count the number of samples in each category based on the CLABSI category labels in the preprocessed mixed feature data, determine the CLABSI infection positive category as the minority class, and extract the minority class samples as data augmentation objects to construct the CLABSI minority class sample set. Continuous Feature Processing Module: This module generates continuous feature data for the constructed minority class sample set using a local neighborhood probability generation method based on K-nearest neighbor search, random proportional interpolation, and Beta probability perturbation. Specifically, the K-nearest neighbor algorithm is used to construct the local neighborhood of the target minority class sample. Random proportional interpolation is performed based on the continuous feature differences between the target sample and the reference sample. Beta probability perturbation is introduced in combination with local statistical distribution to impose local probability constraints on the interpolation results, thereby generating continuous enhanced features. Classification feature processing module: used to establish a probability mapping relationship based on the category distribution of each classification feature in the minority class sample set, and generate new classification feature data based on the probability mapping relationship; Candidate augmentation sample set construction module: It is used to fill the corresponding feature positions with continuous feature data and categorical feature data according to the feature arrangement order of the original clinical data to form a complete augmentation sample, and to form a candidate augmentation sample set by combining all augmentation samples; The sample selection module is used to input candidate augmented samples into a pre-trained random forest classification model, calculate the predicted probability that the candidate augmented sample belongs to the CLABSI infection category, compare the predicted probability with a preset probability threshold, and retain the candidate augmented sample when the predicted probability is greater than or equal to the preset probability threshold; otherwise, the candidate augmented sample is removed, thereby obtaining a high-confidence augmented sample. The CLABSI balanced dataset construction module is used to perform inverse encoding recovery on the encoded classification features in the selected augmented samples based on the classification feature encoding mapping relationship established by the preprocessing module; the recovered classification features and continuous features are combined to form a complete augmented sample, and then merged with the original clinical data according to the feature structure of the original clinical data to construct the augmented CLABSI balanced dataset.
10. The system according to claim 9, characterized in that, The categorized feature processing module specifically includes: (1) Statistically determine the frequency of occurrence of each category of the target classification feature in the minority class samples, and establish the corresponding category probability distribution. The calculation formula is as follows: in, Category in classification features The conditional probability; For category The number of times it appears in the minority class samples; The total number of samples with this classification feature in the minority class; construct the conditional probability distribution of the classification feature based on the above calculation results; (2) Construct a random sampling space based on the established category probability distribution, and randomly generate classification feature values according to the probability corresponding to each category, so that the generated result maintains the category distribution characteristics consistent with the original minority class samples; (3) Perform (1) and (2) on each classification feature respectively to obtain the complete classification feature vector and output the generated classification feature data.