A multi-label ensemble classification method based on perioperative patient risk event data
By employing a multi-label ensemble classification method that combines a Stacking classification ensemble model with label association rules, the label association problem in existing perioperative patient risk prediction models is solved, thereby improving classification accuracy and the effectiveness of risk prediction.
Patent Information
- Application Number
- CN202210760528.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-30
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2042-06-30
AI Technical Summary
Existing perioperative patient risk prediction models are mostly single classification models, which fail to effectively consider label correlation, resulting in poor classification accuracy, especially when generating labels for a few classes, where noise problems are severe.
A stacking-based classification ensemble model is adopted, which combines a label association rule acquisition module and a fusion module to perform multi-label ensemble classification on patient feature data. The classification accuracy is improved through dimensionality reduction and data balancing techniques.
It improves the accuracy of risk event classification in perioperative patients and the reference value of risk prediction results, and improves the generation quality of minority class labels by characterizing the correlation between classification labels.
Smart Images

Figure CN115206539B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, and in particular to a multi-label integrated classification method based on perioperative patient risk event data. BACKGROUND
[0002] Perioperative period refers to a whole process around surgery, from the patient deciding to accept surgical treatment to the basic recovery after surgery, including preoperative, intraoperative and postoperative period, specifically from the decision of surgical treatment to the end of the treatment related to the surgery, about 5-7 days before surgery to 7-12 days after surgery.
[0003] According to the World Health Organization (WHO) report data of World health statistics 2021, the global life expectancy increased to 73.3 years, and it is estimated that there will be more than 1.5 billion people over the age of 60 by 2050. The increasing elderly population around the world has been identified as the main group of the surgical market, and the prediction of risk events for elderly patients has become one of the hot research directions. Predicting the postoperative risk of the elderly surgical patient group helps doctors to develop treatment plans and reasonably allocate treatment resources, thereby reducing the probability of postoperative risk events. At present, some diagnostic tools can help hospitals provide comprehensive and reliable treatment for high-risk patients. For example, Chinese patents with publication numbers CN111009322A and CN114038565A have disclosed perioperative risk assessment based on patient perioperative data set using a prediction model, however, the prediction model is mostly a single classification model, which does not focus on the noise problem inevitably brought in the process of generating minority class labels, and does not consider the relevance between classification labels, and the classification accuracy is poor. SUMMARY
[0004] The present application aims to at least solve the technical problems existing in the prior art, and provides a multi-label integrated classification method based on perioperative patient risk event data.
[0005] In order to achieve the above-mentioned purpose of the present application, according to the first aspect of the present application, a multi-label integrated classification method based on perioperative patient risk event data is provided, comprising: acquiring patient feature data to be classified; inputting the patient feature data to be classified into a trained classification model, and the classification model outputs a classification result, the classification result including one or more classification labels and the classification confidence of each classification label; the classification model includes a classification integrated model based on Stacking, a label association rule acquisition module and a fusion module, and the fusion module is used to fuse the classification matrix output by the classification integrated model and the association rule matrix output by the label association rule acquisition module to obtain the classification result.
[0006] To achieve the above object of the present application, according to a second aspect of the present application, a perioperative patient data multi-label classification device is provided, comprising: a data acquisition module configured to acquire patient feature data to be classified; a classification module configured to input the patient feature data to be classified into a trained classification model, wherein the classification model outputs a classification result, the classification result comprising one or more classification labels and a classification confidence of each classification label; the classification model comprising a classification ensemble model based on Stacking, a label association rule acquisition module and a fusion module, wherein the fusion module is configured to fuse a classification matrix output by the classification ensemble model and an association rule matrix output by the label association rule acquisition module to obtain the classification result.
[0007] To achieve the above object of the present application, according to a third aspect of the present application, a perioperative patient risk event prediction system is provided, comprising: a data acquisition module configured to acquire patient feature data to be classified; a classification module configured to input the patient feature data to be classified into a trained classification model, wherein the classification model outputs a classification result, the classification result comprising one or more classification labels and a classification confidence of each classification label, each classification label corresponding to a perioperative patient risk event; the classification model comprising a classification ensemble model based on Stacking, a label association rule acquisition module and a fusion module, wherein the fusion module is configured to fuse a classification matrix output by the classification ensemble model and an association rule matrix output by the label association rule acquisition module to obtain the classification result; a conversion module configured to convert the classification labels in the classification result into corresponding perioperative patient risk events to obtain a risk prediction result.
[0008] The above technical solution: realizes multi-label classification of patient feature data to be classified, adopts a classification ensemble model based on Stacking in the classification process to improve the classification accuracy, and further improves the accuracy of the final classification result or risk prediction result by correcting the classification matrix output by the classification model through the association rule matrix representing the association between the classification labels, thereby improving the reference value of the classification result or risk prediction result. BRIEF DESCRIPTION OF DRAWINGS
[0009] Figure 1 is a structural schematic diagram of a perioperative patient data dimension reduction device in embodiment 1 of the present application;
[0010] Figure 2 is a structural schematic diagram of a perioperative patient sample data set acquisition system in embodiment 2 of the present application;
[0011] Figure 3 is a sample data set balancing method flowchart in embodiment 3 of the present application;
[0012] Figure 4is a structure schematic diagram of a sample data set balancing device in embodiment 4 of the present application;
[0013] Figure 5 is a structure schematic diagram of a sample data set acquisition system in embodiment 5 of the present application;
[0014] Figure 6 is a flow schematic diagram of a multi-label integrated classification method based on perioperative patient risk event data in embodiment 6 of the present application;
[0015] Figure 7 is a structure schematic diagram of a classification model in embodiment 6;
[0016] Figure 8 is a preferred flow schematic diagram of a multi-label integrated classification method based on perioperative patient risk event data in embodiment 6 of the present application;
[0017] Figure 9 is a structure schematic diagram of a perioperative patient data multi-label classification device in embodiment 7 of the present application;
[0018] Figure 10 is a structure schematic diagram of a perioperative patient risk event prediction system in embodiment 8 of the present application. DETAILED DESCRIPTION
[0019] The embodiments of the present application are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar notations represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary, only for explaining the present application, and cannot be understood as a limitation of the present application.
[0020] In the description of the present application, it should be understood that the orientation or positional relationship indicated by the terms "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer" and the like are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and therefore cannot be understood as indicating or implying that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation of the present application.
[0021] In the description of the present application, unless otherwise specified and limited, it should be noted that the terms "mounting", "connecting", "connecting" should be understood broadly, for example, it can be mechanical connection or electrical connection, or the communication between two elements, or direct connection, or indirect connection through intermediate medium, and the specific meaning of the above terms can be understood by those skilled in the art according to the specific circumstances.
[0022] Embodiment 1
[0023] The embodiment discloses a perioperative patient data dimension reduction device, as shown in the figure, which comprises: Figure 1
[0024] An input module acquires original perioperative feature data of a patient containing multi-dimensional features and a classification label corresponding to the original perioperative feature data;
[0025] A primary dimension reduction module performs dimension reduction processing on the original perioperative feature data based on a principal component analysis algorithm to obtain first perioperative feature data;
[0026] A secondary dimension reduction module performs dimension reduction processing on the first perioperative feature data based on a genetic algorithm to obtain perioperative feature data;
[0027] An output module outputs the perioperative feature data.
[0028] In the embodiment, in order to better reflect the perioperative state of the patient, improve the accuracy of subsequent classification processing, and avoid the problem that postoperative patient data is not easy to collect and manage, preferably, the original perioperative feature data includes index data of the patient before and during surgery, such as blood pressure, heart rate, blood lipid, etc. before surgery, heart rate, blood pressure, blood loss, and operation duration during surgery. Unlike the existing part of the classification prediction model which only includes the preoperative basic condition of the surgical patient and does not consider the specific situation during surgery, numerous studies have confirmed that the intraoperative indicators such as heart rate, blood pressure, blood loss, and operation duration are all related to the postoperative situation of the patient. Therefore, the original perioperative feature data provided by the embodiment can improve the accuracy of the subsequent model in predicting postoperative events and does not depend on postoperative patient index data.
[0029] In the embodiment, the classification label is used to represent a perioperative patient risk event, and the perioperative patient risk event preferably but not limited to includes unplanned readmission and death.
[0030] In the embodiment, in order to improve the richness of the data, the index data includes category data and numerical data, the category data is index data represented by categories, such as bleeding amount during surgery which can be represented by much, medium, and little, and the numerical data is index data represented by numerical values, such as blood pressure values.
[0031] In the embodiment, the original perioperative feature data can be data of a patient with a known perioperative patient risk event, therefore, the known perioperative patient risk event can be associated with the corresponding classification label of the original perioperative feature data. The original perioperative feature data can also be data of a patient with an unknown perioperative patient risk event, and the corresponding classification label of the original perioperative feature data is set by an expert. The corresponding classification label of the original perioperative feature data can be one, two, or more.
[0032] In the embodiment, the feature dimension of the first perioperative period feature data after the principal component analysis algorithm processing is less than the feature dimension of the original perioperative period feature data, and the initial population of the genetic algorithm is constructed based on the first perioperative period feature data.
[0033] In the embodiment, in order to further reduce the dimension of the first perioperative period feature data by the genetic algorithm, preferably, the secondary dimension reduction module comprises:
[0034] The initial population setting unit sets individuals based on the first perioperative period feature data, the gene number of the individual is less than or equal to the total number of features in the first perioperative period feature data, and a plurality of individuals constitute the initial population; the gene of the individual is a feature in the first perioperative period feature data, and the gene number of each individual can be randomly set under the condition that the gene number of the individual is less than or equal to the total number of features in the first perioperative period feature data.
[0035] The evolution iteration unit repeatedly executes the following processes until the termination condition is reached, and outputs the individual with the maximum fitness when the termination condition is reached: obtaining the fitness of each individual in the current population; selecting part of the individuals in the current population as individuals of the next generation population based on the fitness of the individuals; and performing crossover operation and mutation operation on the individuals of the next generation population.
[0036] In the embodiment, the termination condition is preferably but not limited to that the number of evolution iterations reaches a preset maximum number of evolution iterations, or the maximum value of the fitness of the individual in the evolution iteration no longer increases, or the increase amplitude of the maximum value of the fitness of the individual in the evolution iteration is lower than an increase amplitude threshold. In each iteration, the fitness of the individual in the current population is sorted from high to low, and part of the individuals with high ranking are selected as the individuals of the next generation population. The crossover operation mainly exchanges the same gene sites of the paired parents, and the offspring obtained after the exchange is the individual of the next generation population.
[0037] In the embodiment, in order to make the reduced perioperative period feature data have more excellent performance in subsequent classification processing and improve the classification accuracy, preferably, the process of obtaining the fitness of the individual comprises: obtaining the original perioperative period feature data of a plurality of patients and the corresponding classification labels, performing dimension reduction processing on the plurality of original perioperative period feature data according to the feature information of the individual to obtain a plurality of dimension reduction samples consistent with the feature of the individual; dividing the plurality of dimension reduction samples into a dimension reduction training set and a dimension reduction test set; constructing a dimension reduction multilayer perception neural network; training the constructed dimension reduction multilayer perception neural network using the dimension reduction training set to obtain a dimension reduction classification prediction model; testing the dimension reduction classification prediction model using the dimension reduction test set to obtain the accuracy of the model, and taking the accuracy as the fitness of the individual.
[0038] Embodiment 2
[0039] The embodiment discloses a perioperative period patient sample data set acquisition system, which comprises a data preprocessing module, a feature extraction module, a dimension reduction module,Figure 2 The perioperative patient sample data set acquisition system shown in the embodiment comprises:
[0040] A data acquisition module is configured to acquire original perioperative feature data and cases of a plurality of patients. The case data is generally text data, including doctor's diagnosis, medical history, postoperative follow-up records, etc.
[0041] A classification label set acquisition module is configured to acquire a classification label set based on a plurality of cases. The classification label represents a perioperative patient risk event.
[0042] A classification label association module is configured to associate the original perioperative feature data of a patient with at least one classification label in the classification label set. Therefore, the original perioperative feature data corresponds to a classification label set, and the classification label set includes at least one classification label.
[0043] The perioperative patient data dimension reduction device provided in Embodiment 1 is configured to perform dimension reduction processing on the original perioperative feature data of all patients to obtain corresponding perioperative feature data.
[0044] The sample data set acquisition module uses the perioperative feature data of a patient as a sample, associates the sample with the corresponding classification label set of the original perioperative feature data, and obtains a sample data set of a perioperative patient.
[0045] In this embodiment, preferably, the classification label set acquisition module specifically performs the following: performing word segmentation processing on the patient cases to obtain at least one postoperative event result (the postoperative event result is a perioperative patient risk event), performing similar word comparison on the postoperative event results of a plurality of patients using a trained CBOW model to obtain a plurality of similar postoperative event result sets, matching the similar postoperative event result sets with an event dictionary, searching for classification labels matching the similar postoperative event result sets from the event dictionary, and constructing a plurality of classification labels into a classification label set.
[0046] In this embodiment, the CBOW Multi-Word Context Model of Word2Vec is trained on a large amount of medical corpus. The text information corresponding to the case set in this embodiment is processed by the PKUSEG word segmentation tool (PKUSEG can segment words in multiple fields, including an independent model in the medical field) to obtain a plurality of postoperative event results. The event dictionary is preferably but not limited to the Chinese version ICD-11 event dictionary of the International Classification of Diseases published by the World Health Organization. The event dictionary contains a plurality of classification labels. Whether the similar postoperative event result set matches the event dictionary is preferably but not limited to determined by semantic similarity. If the semantic similarity between the two is greater than a preset similarity threshold, it is considered that the two match, otherwise they do not match.
[0047] In this embodiment, preferably, to impute missing values in the data and improve data quality, a missing value imputation device is also included. This device imputes missing values in the patient's original perioperative characteristic data and inputs the imputed original perioperative characteristic data into a perioperative patient data dimensionality reduction device for dimensionality reduction. The missing value imputation device preferably, but not limited to, uses existing RandomForestRegressor imputation, MissForest imputation, Mean imputation, or Median imputation methods for imputation.
[0048] In this embodiment, more preferably, the missing data filling device performs missing data filling processing on the original perioperative feature data based on a Bayesian Gaussian process latent variable model.
[0049] In this embodiment, data imputation for missing values inevitably introduces uncertainty into the original perioperative feature dataset. This embodiment uses a Bayesian Gaussian process latent variable model (BGPLVM) to impute missing values for numerical features, specifically including:
[0050] First, the observed test data vector y is approximately calculated. * ∈R N×M The probability density p(y) * |Y)(where N is the total number of patient samples and M is the total number of features), and the observed value y * The variational distribution of the relevant latent variables is q(x) * Once the model parameters and latent variables are learned, BGPLVM can be used to estimate missing values. in It is a vector y * The values that can be observed in These are the missing values that need to be predicted. Given partially observed points y * This embodiment aims to reconstruct the lost parts. Missing data are filled by learning low-dimensional embeddings of observable variables on a small, complete dataset. BGPLVM is trained on the complete dataset D, introducing latent variables X and new test latent variables x. * As mentioned above A row vector representing the measurements of a single patient. Represents known observations. To represent missing values, y is obtained by maximizing the following probability density. * The corresponding hidden variable x * The Gaussian probability distribution.
[0051]
[0052] Next, the variational distribution q(x is optimized by maximizing the lower bound of the * log-likelihood under the constraint that all optimization quantities except q(x * ) are held constant. To predict missing values The present application employs the standard Gaussian process prediction method, while taking into account the uncertainty of the input x * , since x * has a distribution q(x * ). Similar to the GP prediction form, to predict The present application first predicts , the latent function value corresponding to y * . Then, the prediction of y
[0053]
[0054] Marginalizing over x * yields a non-Gaussian fully dependent multivariate density, but based on the squared exponential kernel, it can be analyzed, and in the present application, the present application uses the mean and covariance of f * U The mean can provide an estimate of the missing values, and the variance can quantify the uncertainty associated with the mean estimate. Through the BGPLVM model, the latent space and model hyperparameters learned from the training set, the distribution is obtained for each feature containing missing values.
[0055] In this embodiment, in order to facilitate data processing, further preferably, it further includes an encoding device for encoding the original perioperative feature data, and inputting the encoded data into the missing value filling device. The encoding device is preferably but not limited to using the existing One-hot encoding rule for encoding.
[0056] In this embodiment, in order to facilitate data processing, further preferably, it further includes a normalization device for normalizing the original perioperative feature data after encoding, and inputting the normalized data into the missing value filling device. The normalization device is preferably but not limited to using the standard deviation normalization method for normalization processing.
[0057] Embodiment 3
[0058] The present application provides a perioperative patient sample data set balancing method, as shown in Figure 3 , the sample data set balancing method comprises:
[0059] Step S1, oversample the minority class label samples in the sample dataset of the perioperative patient to obtain synthetic samples, generate a corresponding synthetic label set for the synthetic samples, the sample dataset includes a plurality of samples and a classification label set corresponding to the samples; each sample represents a perioperative feature dataset of a patient, which can be original perioperative feature data or perioperative feature data obtained after dimensionality reduction of the original perioperative feature data in Embodiment 1, and the classification label association process of the sample has been described in detail in Embodiment 1 and will not be repeated here.
[0060] Step S2, add the synthetic samples and the synthetic label set to the sample dataset to obtain a temporary sample dataset;
[0061] Step S3, clean the samples in the temporary sample dataset to obtain a balanced sample dataset.
[0062] In this embodiment, the minority class label samples in the sample dataset of the perioperative patient can be oversampled by SMOTE or SVM SMOTE or BorderlineSMOTE or K-Means SMOTE or SMOTE-NC to obtain synthetic samples and generate a corresponding synthetic label set for the synthetic samples. Preferably, to improve the balancing effect, the MLSMOTE algorithm is used to oversample the minority class label samples in the sample dataset of the perioperative patient to obtain synthetic samples and generate a corresponding synthetic label set for the synthetic samples. The MLSMOTE algorithm is a multi-label synthetic minority oversampling technique (Multi label Synthetic Minority Over-sampling Technique, MLSMOTE), which is commonly used to handle data imbalance problems in multi-label classification tasks. Its generation process includes: selecting a minority class label using the imbalance rate (Imbalance Rate, IR); nearest neighbor search: once a sample belonging to the minority label is selected as a seed sample, its nearest neighbors are searched; feature set generation: after selecting a neighborhood, synthetic samples are obtained by interpolation; synthetic label set generation: a synthetic label set is needed for the generated synthetic samples.
[0063] In this embodiment, since the MLSMOTE and other oversampling synthetic minority class sample algorithms will generate some noise samples in the process of synthesizing the minority class label samples, it is necessary to clean these noise samples, so step S3 is set to improve the quality of the sample dataset.
[0064] In this embodiment, preferably, to quickly determine the minority class labels in the sample dataset, the ratio of the number of samples corresponding to each classification label to the total number of samples in the sample dataset is calculated, the classification label with a ratio less than a ratio threshold is regarded as a minority class classification label, and the classification label with a ratio greater than or equal to the ratio threshold is regarded as a majority class classification label. The ratio threshold is preferably but not limited to less than 0.2.
[0065] In the embodiment, the number of samples to be generated for each minority class label is the oversampling rate of the minority class label. To better determine the oversampling rate of each minority class label so that the obtained balanced sample dataset performs better when applied to subsequent classification, preferably, in step S1, the oversampling rate of each minority class label is set based on a genetic algorithm, specifically including:
[0066] In step S11, the sample dataset includes W minority class labels, and the oversampling rates of the samples of the W minority class labels are W genes of an individual, W being a positive integer; each gene represents the oversampling rate of a minority class label, and an initial population is constructed using multiple individuals, the initial population including multiple initial individuals, and the values of the W genes of each initial individual are obtained by random selection, preferably, a value range can be set for the oversampling rate of each minority class label, and the values of the genes are randomly selected from the value range when the initial population is constructed, and the value range can be set as needed;
[0067] In step S12, the following evolutionary iteration process is repeatedly performed until a termination condition is reached: obtaining the fitness of each individual in the current population; selecting part of the individuals in the current population as individuals of the next generation population based on the fitness of the individuals; performing crossover operation and mutation operation on the individuals of the next generation population;
[0068] In step S13, the individual with the maximum fitness when the termination condition is reached is output.
[0069] In the embodiment, the termination condition is preferably but not limited to that the number of evolutionary iterations reaches a preset maximum number of evolutionary iterations, or the maximum value of the fitness of the individuals in the evolutionary iterations no longer increases, or the increase amplitude of the maximum value of the fitness of the individuals in the evolutionary iterations is lower than an increase amplitude threshold. In each iteration, the fitness of the individuals in the current population is sorted from high to low, and part of the individuals with high rankings are selected as the individuals of the next generation population.
[0070] In the embodiment, to make the obtained balanced sample dataset perform better when applied to subsequent classification, preferably, the process of obtaining the fitness of the individual includes:
[0071] The oversampling rate combination of the minority class labels is obtained based on the gene information of the individual; the oversampling rate combination includes the oversampling rates of all the minority class labels;
[0072] The minority class label samples in the sample dataset of the patients in the perioperative period are oversampled based on the oversampling rate combination of the minority class labels to obtain synthetic samples and a synthetic label set of the synthetic samples, the synthetic samples and the synthetic label set are added to the sample dataset to obtain a balanced sample set, and the balanced sample set is divided into a balanced training sample set and a balanced test sample set;
[0073] The balanced multi-layer perception neural network is constructed, the balanced multi-layer perception neural network is trained by using the balanced training sample set to obtain a balanced prediction classification model, the balanced prediction classification model is tested by using the balanced test sample set to obtain an accuracy rate of the balanced prediction classification model, and the accuracy rate is used as the fitness of the individual.
[0074] In the embodiment, to effectively remove the noise samples and improve the quality of the sample set, preferably, the step S3 is cleaning processing on each sample in the temporary sample data set, and the cleaning processing process includes:
[0075] In the step S31, a seed sample is selected from the temporary sample data set, k neighbor samples of the seed sample are selected, classification labels of the k neighbor samples form a neighbor classification label set, and k is a positive integer; each sample in the temporary sample data set can be sequentially selected as the seed sample.
[0076] In the step S32, a classification label set of the seed sample is predicted based on the neighbor classification label set by using a Bayesian conditional probability to obtain a predicted classification label set of the seed sample.
[0077] In the step S33, it is determined whether the predicted classification label set of the seed sample is same as the classification label set of the seed sample in the temporary sample data set, if yes, the seed sample is retained, and if no, the seed sample is deleted, and the seed sample is considered as a noise sample.
[0078] The cleaning process directly predicts the classification label set of the seed sample based on the neighbor classification label set of the seed sample by using the Bayesian conditional probability, compares and determines the obtained predicted classification label set and the true classification label set of the seed sample in the temporary sample data set, does not depend on the determination of the classifier, only depends on the determination of the data itself, reduces the operation amount, and improves the determination efficiency and the accuracy rate.
[0079] In the embodiment, further preferably, in the step S31, the specific process of selecting the k neighbor samples of the seed sample includes:
[0080] The heterogeneous value difference metric HVDM of the seed sample and all or part of the samples in the temporary sample data set is obtained; HVDM is the abbreviation of Heterogeneous Value Difference Metric;
[0081] The heterogeneous value difference metric HVDM is modified by using the global unbalanced weight of the samples in the temporary sample data set to obtain a modified heterogeneous value difference metric.
[0082] The modified out-of-class value difference metrics of all samples in the temporary sample data set and the seed sample are sorted, and the first k samples with larger modified out-of-class value difference metrics are selected as the k neighbor samples of the seed sample. Preferably, the modified out-of-class value difference metrics can be sorted from high to low, and the first k samples with larger modified out-of-class value difference metric values are selected as the k neighbor samples of the seed sample.
[0083] The above process of selecting the k neighbor samples of the seed sample adopts a weighted KNN (Weighted kNN, WkNN) method to improve the quality of the synthetic samples. If the real minority class label samples in the sample data set are very scattered, that is, the space is sparse, the synthetic minority class samples in the execution process of the MLSMOTE algorithm and the like are still scattered and sparse, and in a local perspective, they are still not balanced. If kNN cleaning is directly used, the sparse minority class samples and the new minority class samples synthesized by MLSMOTE will be removed with a high probability, which cannot establish a proper classification boundary. Therefore, the distance weighting idea needs to be introduced to coordinate the kNN cleaning, that is, when facing sparse distributed samples, instead of blindly deleting directly, the local space density (that is, the out-of-class value difference metric HVDM and the global imbalance weight of the sample) is considered, and the small samples are retained as much as possible. The kNN cleaning mainly relies on the label set of the neighbor samples, so the distance calculation of the neighbor samples is particularly important when the data distribution is sparse, which is the main reason for adding distance weighting (that is, the global imbalance weight of the sample in the temporary sample data set is used to modify the out-of-class value difference metric HVDM). WkNN is used to clean the noise samples, changes the distance of the neighbor samples (the modified out-of-class value difference metric), that is, the local density is considered to represent the distance between samples.
[0084] In the embodiment, further preferably, the calculation formula of the out-of-class value difference metric HVDM of the seed sample and the samples in the temporary sample data set is:
[0085]
[0086] wherein f1 represents the feature vector of the seed sample; f2 represents the feature vector of any sample in the temporary sample data set except the seed sample; HVDM(f1, f2) represents the out-of-class value difference metric of the feature vectors f1 and f2; D(f1, f2) represents the distance between the feature vectors f1 and f2; n represents the feature dimension of the samples in the temporary sample data set; x represents the feature index; d x (f1, f2) represents the distance of the feature vector f1 and the feature vector f2 on the feature x, d x (f1, f2) is obtained by the following formula: C represents the number of categories of the feature x when the feature x is a category feature, c represents the category index of the feature x, represents the number of samples in the temporary sample data set in which feature x belongs to feature vector f1 and the class feature of feature x is c; represents the number of samples in the temporary sample data set in which feature x belongs to feature vector f2 and the class feature of feature x is c; represents the number of samples in the temporary sample data set in which feature x belongs to feature vector f1; represents the number of samples in the temporary sample data set in which feature x belongs to feature vector f2; |f1-f2| represents the absolute value of the difference between feature vectors f1 and f2; σ x represents the standard deviation of feature x in the temporary sample data set.
[0087] In this embodiment, further preferably, the calculation formula of the modified heterogeneous value difference metric of the samples in the temporary sample data set and the seed sample is:
[0088]
[0089] wherein f1 represents the feature vector of the seed sample; f2 represents the feature vector of any sample in the temporary sample data set except the seed sample; HVDM(f1, f2) represents the heterogeneous value difference metric of feature vectors f1 and f2; D W (f1, f2) represents the modified heterogeneous value difference metric of feature vectors f1 and f2; n represents the feature dimension of the samples in the temporary sample data set; IW represents the global imbalance weight of the sample with feature vector f2, IW = IR nn / (IR + +IR - ), IR + represents the total imbalance rate of all minority class classification labels in the temporary sample data set, IR - represents the total imbalance rate of all majority class classification labels in the temporary sample data set, IR nn is the total imbalance rate of all classification labels in the classification label set of the sample with feature vector f2.
[0090] In the above process of removing noise samples, Heterogeneous Value Difference Metric (HVDM) is used for distance measurement when WkNN calculates the distance, and the global imbalance weight IW of the sample is used as the weight coefficient to modify HVDM. For the temporary sample data set, the more the minority class labels contained in the classification label set, the larger IR nn will be, and for the temporary sample data set with sparse distribution of minority class label samples and large imbalance rate, the introduction of IW into HVDM distance can improve the density of minority class samples.
[0091] From the formula it can be seen that the weight coefficient The value of the weighted coefficient can scale the HVDM(f1, f2), and the more minority class labels in the label set of the neighbor sample, the smaller the weighted coefficient The more minority class labels in the label set of the neighbor sample, the smaller the weighted coefficient of the corresponding neighbor sample. When the IW of the neighbor sample set of the seed sample is larger, that is, the more minority class labels contained in the label set of the neighbor sample, the smaller the weighted coefficient of the corresponding neighbor sample The more minority class labels in the label set of the neighbor sample, the smaller the weighted coefficient of the corresponding neighbor sample. When the IW of the neighbor sample set of the seed sample is larger, that is, the more minority class labels contained in the label set of the neighbor sample, the smaller the weighted coefficient of the corresponding neighbor sample The more minority class labels in the label set of the neighbor sample, the smaller the weighted coefficient of the corresponding neighbor sample. When the IW of the neighbor sample set of the seed sample is larger, that is, the more minority class labels contained in the label set of the neighbor sample, the smaller the weighted coefficient of the corresponding neighbor sample
[0092] It can be seen that WkNN can help filter neighbor samples for samples with more minority class labels in the label set, taking into account the distribution of labels in the label set of the neighbor sample, so that samples with more minority class labels in the label set are closer to the seed sample, increasing the local minority class label density and reducing the majority class label density. The overall process is as follows: first, use MLSMOTE to oversample the minority class labeled samples to form a relatively balanced temporary new sample set with the original samples. In this new sample set, each sample is subjected to the WkNN process, that is, the k nearest neighbors are sorted based on the weighted HVDM, and then the label set of the seed sample is predicted according to the neighbor samples. If the predicted label set is the same as the seed label set, the sample is retained, otherwise it is deleted
[0093] Embodiment 4
[0094] The embodiment discloses a sample data set balancing device for a patient in a perioperative period, as shown in the figure, the sample data set balancing device comprises: Figure 4
[0095] The sample synthesis module oversamples the minority class labeled samples in the sample data set of the patient in the perioperative period to obtain synthetic samples, and generates corresponding synthetic label sets for the synthetic samples. The sample data set comprises a plurality of samples and a classification label set corresponding to the samples.
[0096] The temporary sample data set acquisition module adds the synthetic samples and the synthetic label sets to the sample data set to obtain a temporary sample data set.
[0097] The cleaning module cleans the samples in the temporary sample data set to obtain a balanced sample data set.
[0098] In the embodiment, preferably, the cleaning module comprises:
[0099] The neighbor sample acquisition unit selects a seed sample from the temporary sample data set, selects k neighbor samples of the seed sample, and the classification labels of the k neighbor samples form a neighbor classification label set. k is a positive integer.
[0100] a prediction classification label set acquisition unit, which predicts a classification label set of the seed sample based on the neighbor classification label set through a Bayesian conditional probability, to obtain a prediction classification label set of the seed sample;
[0101] a cleaning unit, which judges whether the prediction classification label set of the seed sample is the same as the classification label set of the seed sample in the temporary sample data set, and if so, retains the seed sample, and if not, deletes the seed sample.
[0102] In the embodiment, further preferably, the specific process in which the neighbor sample acquisition unit selects the k neighbor samples of the seed sample includes:
[0103] acquiring a heterogeneous value difference measure HVDM of the seed sample and all or part of the samples in the temporary sample data set respectively;
[0104] correcting the heterogeneous value difference measure HVDM by using the global imbalance weight of the samples in the temporary sample data set to obtain a corrected heterogeneous value difference measure;
[0105] sorting the corrected heterogeneous value difference measures of all the samples in the temporary sample data set and the seed sample, and selecting the k samples with larger corrected heterogeneous value difference measures as the k neighbor samples of the seed sample.
[0106] The balancing effect of the sample data set balancing device provided in the embodiment is verified by experiments, and the results are as follows:
[0107]
[0108]
[0109] IR represents the imbalance rate Imbalance Rate of the sample set, and the larger the IR, the more unbalanced the sample set. As can be seen from the experimental results in the above table, the maximum IR and the average IR of the balancing device provided in the embodiment are the smallest, and the interval between the maximum value and the average value of IR is narrowed, which indicates that the balancing of the sample set is better.
[0110] Embodiment 5
[0111] The embodiment also discloses a perioperative patient sample data set acquisition system. Compared with the embodiment 2, the embodiment adds a sample data set balancing device, that is, the sample data set obtained after dimension reduction in the embodiment 2 is subjected to sample balancing processing. The structural schematic diagram of the device is shown in Figure 5 and includes:
[0112] a data acquisition module, which is configured to acquire original perioperative feature data and cases of a plurality of patients;
[0113] The classification label set acquisition module acquires a classification label set based on a plurality of cases, and the classification label represents a perioperative patient risk event.
[0114] The classification label association module is configured to associate the original perioperative feature data of the patient with at least one classification label in the classification label set.
[0115] The perioperative patient data dimension reduction device is configured to perform dimension reduction processing on the original perioperative feature data of all patients to obtain corresponding perioperative feature data.
[0116] The sample data set acquisition module is configured to acquire a sample data set of a perioperative patient by associating the original perioperative feature data of the patient with a corresponding classification label set.
[0117] The sample data set balancing device is configured to perform balancing processing on the sample data set.
[0118] In this embodiment, the missing value filling device is preferably further included, which is configured to perform filling processing on the missing values in the original perioperative feature data of the patient, and input the original perioperative feature data after the filling processing into the perioperative patient data dimension reduction device for dimension reduction processing.
[0119] Embodiment 6
[0120] This embodiment 6 discloses a multi-label integrated classification method based on perioperative patient risk event data, as shown in the figure, the multi-label classification method comprises: Figure 6 The classification model comprises a stacking-based classification integrated model, a label association rule acquisition module, and a fusion module.
[0121] Step A: acquiring patient feature data to be classified; the patient feature data to be classified is the feature data of a perioperative patient, which can include multi-dimensional features. In order to improve the processability of the patient feature data to be classified, reduce the dimension, and improve the quality, the patient feature data to be classified can be sequentially subjected to encoding processing, normalization processing, and dimension reduction processing according to the feature dimension of the sample output by the perioperative patient data dimension reduction device provided in embodiment 1, and the patient feature data to be classified after the dimension reduction processing is input into the trained classification model.
[0122] Step B: inputting the patient feature data to be classified into the trained classification model, and the classification model outputs a classification result, which comprises one or more classification labels and a classification confidence of each classification label; the classification confidence of the classification label represents the probability that the patient feature data to be classified belongs to the classification label. The classification model comprises a stacking-based classification integrated model, a label association rule acquisition module, and a fusion module, and the fusion module is configured to fuse a classification matrix output by the classification integrated model and an association rule matrix output by the label association rule acquisition module to obtain the classification result, and the fusion manner is preferably but not limited to multiplying the classification matrix and the association rule matrix.
[0123] In an embodiment, preferably, the structural diagram of the classification model comprises a first multi-classification model, a second multi-classification model, a third multi-classification model, and a logistic regression model, as shown in Figure 7 The first multi-classification model, the second multi-classification model, and the third multi-classification model respectively perform multi-label classification processing on the patient feature data to be classified to obtain a first primary classification result, a second primary classification result, and a third primary classification result. The logistic regression model processes the first primary classification result, the second primary classification result, and the third primary classification result to obtain a classification matrix.
[0124] In this embodiment, preferably, the first multi-classification model, the second multi-classification model, and the third multi-classification model are respectively a Ranking-SVM model, a classification multi-layer perception neural network model, and a Binary Relevance model. The Ranking-SVM model and the Binary Relevance model are relatively conventional base models in the Stacking ensemble, and are used here to have high reliability in model ensemble. The classification multi-layer perception neural network model adopts a multi-layer perception neural network structure (i.e., an MLP network structure), which can avoid overfitting problems and has low complexity.
[0125] In this embodiment, preferably, the method further comprises a step of constructing a sample data set of patients in the perioperative period, as shown in Figure 8 The step of constructing the sample data set of patients in the perioperative period is preferably but not limited to being constructed by using the system of Embodiment 2 or Embodiment 5.
[0126] In an embodiment, as shown in Figure 7 The training process of the classification ensemble model is as follows:
[0127] A sample data set of patients in the perioperative period is constructed, each sample in the sample data set is associated with one or more classification labels, the sample data set is divided into a classification training set and a classification test set, and the association of the classification labels can be performed in an artificial manner;
[0128] A classification ensemble model, i.e., the above-mentioned Stacking ensemble model, is constructed, which comprises a first multi-classification model, a second multi-classification model, a third multi-classification model, and a logistic regression model.
[0129] The classification training set is used to train the classification ensemble model, and the classification test set is used to test and verify the trained classification ensemble model. In the verification, RandomizedSearchCV and GridSearchCV are used to perform cross-validation on the training set, and the F1_Micro score is used to select the hyperparameters.
[0130] In this embodiment, as shown in Figure 7As shown, preferably, the association rule acquisition module performs the following steps:
[0131] A sample data set of the perioperative patients is acquired, each sample in the sample data set is associated with more than one classification label; the sample data set is preferably but not limited to the sample data set of the perioperative patients acquired in Embodiment 2 or Embodiment 5, i.e., the standard patient data set.
[0132] Association rule mining is performed on the classification labels in the sample data set to obtain an association rule matrix. The association rule matrix includes the association confidence between any two of the classification labels.
[0133] In this embodiment, as shown in Figure 7 Further preferably, when the number of classification labels in the sample data set is small, specifically, when the number is less than a number threshold, association rule mining is directly performed on the classification labels in the sample data set by using the FP-growth algorithm. First, a classification label matrix as shown in Figure 7 is established, in which the first row is the labels and the first column is the patient number; then, the FP-growth algorithm is used to perform association rule analysis and processing on the classification label matrix, and the association confidence between any two classification labels is output, with the value range of the association confidence being 0 to 1. Based on these association confidences, an association rule matrix as shown in Figure 7 is established, in which the first row and the first column are the classification labels, and the elements in the matrix represent the association confidence between the classification labels in the row and the column where the element is located, as shown in Figure 7 where A(N-1) represents the association confidence between the classification label N and the classification label 1.
[0134] In this embodiment, preferably, when the number of classification labels in the sample data set is large, the correlation patterns between the classification labels are different, and direct association analysis may cause the frequent item set finding process to be complex and affect the accuracy of the association analysis. Specifically, when the number of classification labels is greater than or equal to a number threshold, the number threshold is preferably but not limited to 3 or 4 or 5. The steps of performing association rule mining on the classification labels in the sample data set to obtain an association rule matrix specifically include:
[0135] The classification labels in the sample data set are clustered to obtain more than one cluster; preferably but not limited to, the K-means++ algorithm is used for clustering processing; association rule mining is performed on the classification labels in each cluster to obtain an association rule sub-matrix. When fusing, the classification matrix is divided into more than one sub-classification matrix according to the clustering results, one cluster corresponds to one sub-classification matrix, the sub-classification matrix and the association rule sub-matrix corresponding to the cluster are multiplied to obtain the classification sub-result of the cluster, and all classification sub-results constitute the classification result.
[0136] In this embodiment, further preferably, the association rule sub-matrix is obtained by performing association rule mining on the classification labels in each classification cluster through the FP-growth algorithm, and the obtaining process is as shown in the following steps. Figure 7 The process is consistent with the above preferred scheme, which has been described in detail above and will not be repeated here.
[0137] Embodiment 7
[0138] This embodiment discloses a perioperative patient data multi-label classification device, as shown in the following formula (I): Figure 9 The device comprises:
[0139] A data acquisition module is configured to acquire patient feature data to be classified.
[0140] A classification module is configured to input the patient feature data to be classified into a trained classification model, and the classification model outputs a classification result, which comprises one or more classification labels and a classification confidence of each classification label. The classification model comprises a classification ensemble model based on stacking, a label association rule acquisition module, and a fusion module. The fusion module is configured to fuse a classification matrix output by the classification ensemble model and an association rule matrix output by the label association rule acquisition module to obtain the classification result.
[0141] In this embodiment, preferably, the classification ensemble model comprises a first multi-classification model, a second multi-classification model, a third multi-classification model, and a logistic regression model. The first multi-classification model, the second multi-classification model, and the third multi-classification model perform multi-label classification processing on the patient feature data to be classified to obtain a first primary classification result, a second primary classification result, and a third primary classification result, respectively. The logistic regression model processes the first primary classification result, the second primary classification result, and the third primary classification result to obtain a classification matrix.
[0142] In this embodiment, preferably, the device further comprises a classification ensemble model training module, which performs the following processes:
[0143] A sample data set of perioperative patients is constructed, and one or more classification labels are associated with each sample in the sample data set. The sample data set is divided into a classification training set and a classification test set. Preferably, but not limited to, the system provided in Embodiment 2 or Embodiment 5 is used to construct the sample data set of perioperative patients.
[0144] A classification ensemble model is constructed. The classification ensemble model comprises a first multi-classification model, a second multi-classification model, a third multi-classification model, and a logistic regression model.
[0145] The classification training set is used to train the classification ensemble model, and the classification test set is used to test and verify the trained classification ensemble model.
[0146] In this embodiment, the classification device constructs an ensemble classification model for multi-label postoperative events in the perioperative period, incorporating association rule analysis. Multiple postoperative risk events may occur. This study aims to predict these multiple postoperative events. A multi-label prediction model is constructed by integrating a Ranking-SVM model, a multilayer perceptron neural network model, and a Binary Relevance model. To further improve the model's stability and accuracy, association rules are incorporated into the prediction model for optimization.
[0147] Example 8
[0148] This embodiment discloses a perioperative patient risk event prediction system, such as... Figure 10 As shown, it includes: a data acquisition module for acquiring feature data of patients to be classified; a classification module for inputting the feature data of patients to be classified into a trained classification model, the classification model outputting classification results, the classification results including one or more classification labels and the classification confidence of each classification label, each classification label corresponding to a perioperative patient risk event; the classification model includes a stacking-based classification ensemble model, a label association rule acquisition module, and a fusion module, the fusion module for fusing the classification matrix output by the classification ensemble model and the association rule matrix output by the label association rule acquisition module to obtain classification results; and a transformation module for converting the classification labels in the classification results into corresponding perioperative patient risk events to obtain risk prediction results.
[0149] In this embodiment, preferably, the classification ensemble model includes a first multi-classification model, a second multi-classification model, a third multi-classification model, and a logistic regression model; the first multi-classification model, the second multi-classification model, and the third multi-classification model respectively perform multi-label classification processing on the patient feature data to be classified to obtain a first primary classification result, a second primary classification result, and a third primary classification result; the logistic regression model processes the first primary classification result, the second primary classification result, and the third primary classification result to obtain a classification matrix.
[0150] In this embodiment, preferably, it further includes a classification ensemble model training module, which performs the following process:
[0151] A sample dataset of perioperative patients is constructed, in which each sample is associated with more than one classification label. The sample dataset is divided into a classification training set and a classification test set. Preferably, but not limited to, the sample dataset of perioperative patients is constructed using the system provided in Example 2 or Example 5.
[0152] Construct classification ensemble models; the classification ensemble models include first-class multi-class classification models, second-class multi-class classification models, third-class multi-class classification models, and logistic regression models;
[0153] The classification ensemble model is trained by using the classification training set, and the trained classification ensemble model is tested and verified by using the classification test set.
[0154] In this embodiment, the system sample data set acquisition process provided in Embodiment 2 or Embodiment 5 is implemented to predict the risk events of patients (especially elderly surgical patients) during the perioperative period. On the basis of improving the missing and unbalanced data set, the correlation rule analysis is fused to build a postoperative event multi-label prediction model. Based on the patient case text, the postoperative event label is extracted. The CBOW label extraction model of Word2Vec is adopted, a large amount of medical related corpus is collected, a medical word vector model is trained, and the postoperative event label set (i.e. the classification label set) extraction is realized. Next, the missing data is filled in by using the Bayesian Gaussian process latent variable model, and the label unbalanced data is processed by using the MLSMOTE, the weighted kNN (WKNN) and the genetic algorithm. Finally, the feature dimension reduction model is built by combining the principal component analysis PCA model and the genetic algorithm, so as to provide the classification ensemble model with higher correlation input.
[0155] Although the embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, replacements and variations can be made to these embodiments without departing from the principles and purposes of the present application, and the scope of the present application is defined by the claims and their equivalents.
Claims
1. A multi-label ensemble classification method based on perioperative patient risk event data, characterized in that, include: Obtain characteristic data of patients to be classified; The patient feature data to be classified is input into the trained classification model, and the classification model outputs the classification result, which includes one or more classification labels and the classification confidence of each classification label. The classification model includes a stacking-based classification ensemble model, a label association rule acquisition module, and a fusion module. The fusion module is used to fuse the classification matrix output by the classification ensemble model and the association rule matrix output by the label association rule acquisition module to obtain the classification result. The classification ensemble model includes a first multi-class classification model, a second multi-class classification model, a third multi-class classification model, and a logistic regression model. The first multi-classification model, the second multi-classification model, and the third multi-classification model respectively perform multi-label classification processing on the patient feature data to be classified to obtain the first primary classification result, the second primary classification result, and the third primary classification result; The logistic regression model processes the first, second, and third primary classification results to obtain the classification matrix; The process by which the tag association rule acquisition module obtains the association rule matrix is as follows: When the number of classification labels in the sample dataset is less than the threshold, the FP-growth algorithm is used directly to mine association rules for the classification labels in the sample dataset to obtain the association rule matrix; fusion module The classification result is obtained by multiplying the classification matrix and the association rule matrix; When the number of classification labels in the sample dataset is greater than or equal to the threshold, the classification labels in the sample dataset are clustered to obtain more than one cluster. Association rule mining is performed on the classification labels in each cluster to obtain an association rule sub-matrix. When the fusion module fuses the data, the classification matrix is divided into more than one sub-classification matrix according to the clustering results. One cluster corresponds to one sub-classification matrix. The sub-classification matrix is multiplied by the association rule sub-matrix corresponding to the cluster to obtain the classification sub-result of the cluster. All classification sub-results are combined to form the classification result.
2. The multi-label integrated classification method based on perioperative patient risk event data as described in claim 1, characterized in that, The training process of the classification ensemble model is as follows: Construct a sample dataset of perioperative patients, in which each sample is associated with more than one classification label, and divide the sample dataset into a classification training set and a classification test set; Construct a classification ensemble model; The classification ensemble model is trained using a classification training set and tested and validated using a classification test set.
3. The multi-label integrated classification method based on perioperative patient risk event data as described in claim 1 or 2, characterized in that, The association rule acquisition module performs the following steps: Obtain a sample dataset of perioperative patients, with each sample in the dataset associated with one or more classification labels; Association rule mining is performed on the classification labels in the sample dataset to obtain the association rule matrix.
4. The multi-label integrated classification method based on perioperative patient risk event data as described in claim 3, characterized in that, The FP-growth algorithm is used to mine association rules for the classification labels in the sample dataset.
5. The multi-label integrated classification method based on perioperative patient risk event data as described in claim 3, characterized in that, The step of obtaining the association rule matrix by performing association rule mining on the classification labels in the sample dataset specifically includes: Clustering the classification labels in the sample dataset yields one or more clusters; For each cluster, association rule mining is performed on the classification labels to obtain an association rule submatrix.
6. The multi-label integrated classification method based on perioperative patient risk event data as described in claim 5, characterized in that, The association rule submatrix is obtained by mining association rules for the classification labels in each classification cluster using the FP-growth algorithm.
7. The multi-label integrated classification method based on perioperative patient risk event data as described in claim 4 or 6, characterized in that, The association rule matrix includes the association confidence between any two category labels among all category labels.
8. A multi-label classification device for perioperative patient data, used to implement the method described in any one of claims 1-7, characterized in that, include: The data acquisition module is used to acquire characteristic data of patients to be classified. The classification module is used to input the patient feature data to be classified into a trained classification model, and the classification model outputs a classification result, which includes one or more classification labels and the classification confidence of each classification label. The classification model includes a stacking-based classification ensemble model, a label association rule acquisition module, and a fusion module. The fusion module is used to fuse the classification matrix output by the classification ensemble model and the association rule matrix output by the label association rule acquisition module to obtain the classification result.
9. A perioperative patient risk event prediction system, characterized in that, include: The data acquisition module is used to acquire characteristic data of patients to be classified. The classification module is used to input the patient feature data to be classified into the trained classification model. The classification model outputs the classification result, which includes one or more classification labels and the classification confidence of each classification label. Each classification label corresponds to a perioperative patient risk event. The classification model includes a stacking-based classification ensemble model, a label association rule acquisition module, and a fusion module. The fusion module is used to fuse the classification matrix output by the classification ensemble model and the association rule matrix output by the label association rule acquisition module to obtain the classification result. The conversion module converts the classification labels in the classification results into corresponding perioperative patient risk events to obtain risk prediction results. The classification ensemble model includes a first multi-class classification model, a second multi-class classification model, a third multi-class classification model, and a logistic regression model. The first multi-classification model, the second multi-classification model, and the third multi-classification model respectively perform multi-label classification processing on the patient feature data to be classified to obtain the first primary classification result, the second primary classification result, and the third primary classification result; The logistic regression model processes the first, second, and third primary classification results to obtain the classification matrix; The process by which the tag association rule acquisition module obtains the association rule matrix is as follows: When the number of classification labels in the sample dataset is less than the threshold, the FP-growth algorithm is used directly to mine association rules for the classification labels in the sample dataset to obtain the association rule matrix; fusion module The classification result is obtained by multiplying the classification matrix and the association rule matrix; When the number of classification labels in the sample dataset is greater than or equal to the threshold, the classification labels in the sample dataset are clustered to obtain more than one cluster. Association rule mining is performed on the classification labels in each cluster to obtain an association rule sub-matrix. When the fusion module fuses the data, the classification matrix is divided into more than one sub-classification matrix according to the clustering results. One cluster corresponds to one sub-classification matrix. The sub-classification matrix is multiplied by the association rule sub-matrix corresponding to the cluster to obtain the classification sub-result of the cluster. All classification sub-results are combined to form the classification result.
Citation Information
Patent Citations
Comprehensive evaluation system for perioperative period of elderly
CN114038565A
MLKNN multi-label classification method based on association rules
CN110516704A
Perioperative risk assessment and clinical decision intelligent auxiliary system
CN111009322A