Multi-label Prediction Method, Device and Storage Medium for Heart Failure in Patients with Atrial Fibrillation

By constructing a binary correlation decision tree model based on Bayesian optimization, the simultaneous prediction problem of disease probability and time in heart failure prediction in patients with atrial fibrillation is solved, and a reliable predictive model is provided to support doctors' early intervention and treatment decisions.

CN115064268BActive Publication Date: 2025-07-18NORTHEASTERN UNIV CHINA +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210495329.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-07
Publication Date
2025-07-18
Estimated Expiration
2042-05-07

AI Technical Summary

Technical Problem

There is a lack of effective predictive model for heart failure in patients with atrial fibrillation in the prior art, and it is impossible to predict the probability of illness and the time of illness at the same time, and most models can only provide a single probability analysis of heart failure.

Method used

A multi-label prediction method based on Bayesian optimization is adopted to construct a binary correlation decision tree model, and combined with Bayesian optimization model optimization parameters, multi-label prediction of heart failure in patients with atrial fibrillation, including the output of the probability of disease and onset time.

Benefits of technology

It realizes effective prediction of heart failure in patients with atrial fibrillation, outputs the probability of illness and the time of illness, and provides reliable decision-making support for doctors' early prevention and treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115064268B_ABST
    Figure CN115064268B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-label prediction method, device and storage medium for heart failure in patients with atrial fibrillation, which relates to the technical field of multi-label prediction. The method is executed by a computer and includes: constructing a multi-label decision tree model; the decision tree in the multi-label decision tree model is a binary association decision tree, and the labels include the probability of disease and the onset time; building a Bayesian optimization model; using the Bayesian optimization model to optimize the key parameters of the binary association decision tree; setting the key parameters optimized by the Bayesian optimization model to corresponding values in the binary association decision tree; and obtaining the prediction result of heart failure in patients with atrial fibrillation by using the multi-label decision tree model after parameter setting. The present invention establishes a prediction model for heart failure with reduced ejection fraction in patients with atrial fibrillation, and solves the technical problems that there is currently no effective prediction model for heart failure in patients with atrial fibrillation and it is impossible to simultaneously predict the probability of disease and the disease time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-label prediction, and particularly to a multi-label prediction method, device and storage medium for heart failure of atrial fibrillation patients based on Bayesian optimization. Background Art

[0002] Atrial fibrillation is one of the most common and costly cardiovascular diseases, associated with major diseases such as stroke, heart failure, and dementia. It is a global chronic disease with significant morbidity and mortality, and the lifetime risk of disease is about 1 / 4, which has a great impact on human health. With the increasing trend of global aging, atrial fibrillation will bring a huge burden to the public health system. Therefore, early intervention and prognosis management of atrial fibrillation become very important. Heart failure is a common and costly chronic disease, and the lifetime risk of disease is as high as 1 / 5. It is estimated that by 2030, the total treatment cost of heart failure in the United States will reach $53 million per year, which will bring a greater burden to the medical system and society. Existing studies have shown that about 1 / 3 of atrial fibrillation patients will develop heart failure. The occurrence of heart failure will increase the risk of hospitalization and cardiovascular mortality in atrial fibrillation patients. There is a close relationship between atrial fibrillation and heart failure, and they interact with each other and are causally related. The combination of the two has a worse prognosis than either condition alone. No matter which disease appears first, the other disease will soon become obvious in the same patient. Once a heart failure occurs in an atrial fibrillation patient, it means a higher mortality and hospitalization rate for the patient, which will have a significant impact on their health. Heart failure prevention is an important treatment focus for atrial fibrillation patients. If the occurrence of heart failure in atrial fibrillation patients can be predicted in advance, better treatment effects can be achieved and the survival rate of atrial fibrillation patients can be improved. Early identification of high-risk patients helps to develop new therapies for specific patients and carry out active "upstream" treatment as early as possible, which can reverse the disease outcome in the early stage of the disease. At present, the development of new clinical heart failure in atrial fibrillation patients has not been well studied. Conducting personalized treatment different from other patients for atrial fibrillation patients with a high probability of developing heart failure and delaying or eliminating the possibility of heart failure occurrence has a significant impact on the prognosis of atrial fibrillation patients. Therefore, exploring the relationship between atrial fibrillation and future occurrence of heart failure and realizing the prediction of heart failure in atrial fibrillation patients have important clinical, economic and public health significance.

[0003] Currently, there are heart failure prediction models in the prior art that can predict the probability of heart failure, but do not associate atrial fibrillation with heart failure. There is a lack of effective prediction models in the field of predicting early heart failure in atrial fibrillation patients, and true heart failure prediction for atrial fibrillation patients cannot be achieved. Moreover, most prediction models only give the probability of heart failure alone and cannot analyze the disease time. Summary of the Invention

[0004] In view of this, the present invention proposes a multi-label prediction method for heart failure in patients with atrial fibrillation based on Bayesian optimization, and establishes a prediction model for heart failure with reduced ejection fraction (HFpEF) in patients with atrial fibrillation, which solves the problem that there is currently no effective prediction model for heart failure in patients with atrial fibrillation and it is impossible to predict the probability of disease and the time of disease simultaneously.

[0005] For this reason, the present invention provides the following technical solutions:

[0006] On the one hand, the present invention provides a multi-label prediction method for heart failure in patients with atrial fibrillation based on Bayesian optimization, which is executed by a computer and includes the following steps:

[0007] Collect information data of patients with atrial fibrillation and process the information data;

[0008] Construct a multi-label decision tree model for heart failure in patients with atrial fibrillation; the decision tree in the multi-label decision tree model for heart failure in patients with atrial fibrillation is a binary association decision tree, and the labels include the probability of disease and the onset time;

[0009] According to the designed labels, divide the processed data into a training set and a test set;

[0010] Build a Bayesian optimization model; the objective function of the Bayesian optimization model takes the key parameters of the binary association decision tree as input, and the output is the mean accuracy of the model cross-validation times; the key parameters include: the maximum number of features, the maximum depth, and the minimum number of samples in a leaf;

[0011] Use the Bayesian optimization model to optimize the key parameters of the binary association decision tree; set the key parameters optimized by the Bayesian optimization model to the corresponding values in the binary association decision tree;

[0012] Use the multi-label decision tree model after parameter setting to train the training set first, and then predict the test set, and output the probability of disease and the onset time of heart failure in patients with atrial fibrillation.

[0013] Further, the multi-labels include a first label [0, 0, 0], a second label [1, 1, 1], and a third label [1, 0, 1]. The first column in the label represents whether the patient has the disease, the second column represents that the onset time is less than 1 year, and the third column represents that the onset time is greater than 1 year.

[0014] Further, collecting information data of patients with atrial fibrillation includes: collecting information data of patients with atrial fibrillation in the form of electronic health records, and the electronic health records mainly collect the patient's demographic information, laboratory information, and clinical information data.

[0015] Further, processing the information data includes:

[0016] Convert unstructured text data into digital structured data.

[0017] Furthermore, the optimization process includes:

[0018] The Bayesian optimization model first randomly generates an initial point, and then trains it using Gaussian process regression to generate a posterior distribution of the objective function;

[0019] Calculate the corresponding input and output using the acquisition function, and then determine whether the output meets the objective function. Repeat this process until the optimal value is selected or the number of iterations is reached.

[0020] Furthermore, the predicted label of the binary association decision tree is represented as follows:

[0021]

[0022] Wherein, represents the predicted label of the binary association decision tree, y j represents the label of the j-th generation and y j ∈Y i , 1 ≤ j ≤ q, q ∈ Q, Gini represents the Gini coefficient, D j represents a single-label training set, represents the input x i 's label set.

[0023] Furthermore, the objective function in the Bayesian optimization model is:

[0024] x * = argmax f(min_samples_split, max_depth, max_features, p(acc|X, D));

[0025] Wherein, x * is the solution to be found, min_samples_split, max_depth, max_features are the minimum number of samples in a leaf, the maximum depth, and the maximum number of features, and all belong to the input set X, p(acc|X, D)) is the Gaussian distribution between the input set X and the target accuracy acc, and D is the solution set to be found.

[0026] On the other hand, the present invention also provides a multi-label prediction device for heart failure in atrial fibrillation patients based on Bayesian optimization. The device includes:

[0027] A data acquisition unit for acquiring information data of atrial fibrillation patients and processing the information data;

[0028] A prediction model construction unit constructs a multi-label decision tree model for heart failure in patients with atrial fibrillation; the decision tree in the multi-label decision tree model for heart failure in patients with atrial fibrillation is a binary association decision tree, and the labels include the probability of disease and the onset time.

[0029] A data partitioning unit is used to partition the processed data into a training set and a test set according to the designed labels.

[0030] An optimization model construction unit is used to build a Bayesian optimization model; the objective function of the Bayesian optimization model takes the key parameters of the binary association decision tree as input, and the output is the mean accuracy of the model cross-validation times; the key parameters include: the maximum number of features, the maximum depth, and the minimum number of samples in a leaf.

[0031] An optimization unit is used to optimize the key parameters of the binary association decision tree constructed by the prediction model construction unit using the Bayesian optimization model constructed by the optimization model construction unit.

[0032] A parameter setting unit is used to set the key parameters optimized by the optimization unit to the corresponding values in the binary association decision tree constructed by the prediction model construction unit.

[0033] A prediction unit is used to first train the training set using the multi-label decision tree model after parameter setting, and then predict the test set, and output the probability of disease and the onset time of heart failure in patients with atrial fibrillation.

[0034] On the other hand, the present invention also provides a computer-readable storage medium, which stores a computer instruction set therein. When the computer instruction set is executed by a processor, it implements the above-mentioned multi-label prediction method for heart failure in patients with atrial fibrillation based on Bayesian optimization.

[0035] Advantages and positive effects of the present invention:

[0036] (1) The multi-label prediction method for heart failure in patients with atrial fibrillation in the present invention can realize the prediction of heart failure with preserved ejection fraction (HFpEF) in patients with atrial fibrillation, and at the same time output the probability of disease and the disease time, and establish an effective heart failure prediction model for patients with atrial fibrillation.

[0037] (2) The multi-label prediction model for heart failure in patients with atrial fibrillation in the present invention can predict the onset of heart failure with preserved ejection fraction (HFpEF) in patients within 1 year, and provide a reliable decision for doctors' early prevention and treatment plan formulation.

[0038] It should be noted that the multi-label prediction method for heart failure in atrial fibrillation patients based on Bayesian optimization in the present invention is not aimed at diagnosing heart failure, nor is it to directly obtain a diagnosis result of whether a patient has heart failure. The prediction result of heart failure in atrial fibrillation patients obtained by using the heart failure prediction model for atrial fibrillation patients is a predicted value obtained by a computer processing data. This predicted value can be used as data support for medical research and can also be used as an auxiliary decision-making for doctors to prevent in advance and formulate treatment plans. Description of the Drawings

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0040] Figure 1 It is a flowchart of a multi-label prediction method for heart failure in atrial fibrillation patients based on Bayesian optimization in an embodiment of the present invention;

[0041] Figure 2 It is a schematic diagram of a binary association decision tree model in an embodiment of the present invention;

[0042] Figure 3 It is a schematic diagram of the micro-average AUC of the prediction result in an embodiment of the present invention;

[0043] Figure 4 It is a structural block diagram of a multi-label prediction device for heart failure in atrial fibrillation patients based on Bayesian optimization in an embodiment of the present invention. Detailed Embodiments

[0044] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0045] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0046] In order to predict the probability and time of onset of heart failure in patients with atrial fibrillation, the present invention constructs a multi-label prediction model based on Bayesian optimization. This prediction model is executed by a computer and mainly includes four parts: patient data processing, model establishment, Bayesian optimization, and result output. Among them:

[0047] Patient data processing mainly collects patient data through electronic health records and converts it into a data type that can be processed by the prediction model.

[0048] Model establishment mainly involves establishing a multi-label decision tree. The multi-label decision tree uses an interpretable decision tree as the basic prediction model and uses the binary relevance strategy in multi-label tasks to convert multi-labels into single-label tasks that can be processed by the decision tree, thereby realizing multi-label prediction.

[0049] Bayesian parameter optimization uses a Bayesian optimization model to optimize the parameters of the multi-label decision tree established in the model establishment part.

[0050] Result output is to output the prediction results of the prediction model, and a total of three types of labels are output: atrial fibrillation patients do not develop HFpEF, atrial fibrillation patients develop HFpEF within one year, and atrial fibrillation patients develop HFpEF after one year.

[0051] It can be seen that by using the prediction model executed by this computer and performing computer processing and analysis on the data of atrial fibrillation patients, the prediction of heart failure in atrial fibrillation patients can be realized, and at the same time, the probability and time of onset of the disease can be output, providing data support for medical research and also providing auxiliary decision-making for doctors' early prevention and treatment plan formulation.

[0052] For the sake of easy understanding, a multi-label prediction method for heart failure in atrial fibrillation patients based on Bayesian optimization in an embodiment of the present invention will be described in detail below with reference to the accompanying drawings.

[0053] As Figure 1As shown, it shows a flowchart of a multi-label prediction method for heart failure in atrial fibrillation patients based on Bayesian optimization according to an embodiment of the present invention. This method is implemented through a computer program and specifically includes the following steps:

[0054] S1. Collect data of atrial fibrillation patients and process the collected data of atrial fibrillation patients;

[0055] Among them, the data collection of atrial fibrillation patients is mainly in the form of electronic health records. The electronic health records mainly collect data such as patients' demographic information, laboratory information, and clinical information, a total of 27 types, including: gender, age, coronary heart disease, hypertension, diabetes, cerebral hemorrhage, cerebral infarction, smoking, drinking, admission pulse rate, admission systolic blood pressure, admission diastolic blood pressure, body weight, white blood cell count, neutrophil / lymphocyte ratio, creatinine content, heart rate, left ventricular diastolic diameter, left ventricular ejection fraction, atrial fibrillation type, duration of atrial fibrillation, antiarrhythmic drugs, receptor blockers, intervention lipid-lowering drugs, intervention ACEI, pacemaker, and cardioversion treatment.

[0056] The processing of the collected data of atrial fibrillation patients includes: converting specific text-type unstructured data into digital-type structured data. For example, converting the diagnosis time into a numerical value counted by year, performing Z-normalization processing on continuous variables such as age, and setting discrete variables such as whether to drink as 0 / 1.

[0057] S2. Design labels according to the disease information defined by medical experts.

[0058] Finally, it is designed into 3 labels: label 1 [0, 0, 0], label 2 [1, 1, 1], and label 3 [1, 0, 1]. The first column in the label represents whether the patient is ill, the second column represents that the onset time is less than 1 year, and the third column represents that the onset time is greater than 1 year.

[0059] S3. According to the designed labels, divide the data processed in S1 into two parts: a training set and a test set.

[0060] The number of people with label 1 in the training set is 2329, the number of people with label 2 is 178, and the number of people with label 3 is 296; the number of people with label 1 in the test set is 583, the number of people with label 2 is 40, and the number of people with label 3 is 78.

[0061] S4. Build a multi-label decision tree model;

[0062] The problem of predicting multiple chronic complications is a typical multi-label learning problem. Traditional single-output models can only predict the occurrence of one event and cannot consider the potential impacts of different events. Different from single-label learning, multi-label learning is a process of mapping a sample to a set of labels. Multi-label classification can effectively learn the relationships among multiple different classes, thereby achieving multiple outputs. In multi-label learning, the sample space is defined as: X = R d , where d represents the dimension of the sample. The label space is defined as: Y = {y1, y2, …, y q}. If the training set is represented as D = {(x i , Y i )|1 ≤ i ≤ m}, then the multi-label task can be described as f: X → 2 Y . For each pair of sample and label (x i , Y i ), x i ∈ X represents a d-dimensional feature vector, and Y i ∈ Y represents the label space.

[0063] Binary relevance is a very simple and efficient first-order multi-label learning method, which is implemented by converting each label in the multi-label into an independent single label. This method needs to construct a binary dataset for each label. The only drawback is that it does not consider the internal correlations among multiple labels. For any multi-label dataset D, for its i-th label y i , the single-label training set D i constructed is represented as follows:

[0064]

[0065] For a new sample x * , the predicted label given by the model is shown as follows:

[0066] Y * = {y j |f j (x * ) > 0, 1 ≤ j ≤ q};

[0067] where f(·) represents the classification function.

[0068] Combined with the CART decision tree model, the predicted label of the binary relevance decision tree designed in the embodiments of the present invention can be represented as follows:

[0069]

[0070] where, represents the predicted label of the binary relevance decision tree, y j represents the label of the j-th generation and y j ∈ Yi where \(1\leq j\leq q\), \(q\in Q\), Gini represents the Gini coefficient, \(D\) j represents a single-label training set, represents the input \(x\) i 's label set.

[0071] As Figure 2 shown, it shows a schematic diagram of the binary correlation decision tree model established in the embodiment of the present invention. The left half can achieve the classification of non-HFpEF and HFpEF occurring within 1 year according to different features, and the right half can achieve the classification of non-HFpEF and HfpEF occurring after 1 year, so as to achieve the output of multiple labels.

[0072] S5. Build a Bayesian optimization model.

[0073] The parameter optimization of decision trees has always been a challenge in model design. Currently, there are mainly three parameter tuning methods: grid search, random search, and Bayesian optimization. Based on the idea of traversal, grid search will consume a large amount of computing resources and time. The result of random search may not be optimal. Bayesian optimization has the characteristics of fewer iteration times and faster optimization speed, and can find the perfect combination of parameters at the lowest cost, avoiding parameter explosion. Moreover, it does not require derivative calculation, greatly saving computing resources.

[0074] Bayesian parameter optimization consists of two parts: the prior function and the acquisition function. The prior function mainly uses the Gaussian process to fit the optimization objective function to obtain its posterior probability. The acquisition function can sample and explore new regions in the area where the global optimal value may appear according to the posterior distribution, so as to find the optimal value at the lowest cost. At the beginning of the optimization process, the model first randomly generates initialization points, and then uses Gaussian process regression for training to generate the posterior distribution of the function. Use a suitable acquisition function to calculate the corresponding input and output, and judge whether the output meets the objective function. Then repeat continuously until the optimal value is selected or the iteration times are reached.

[0075] Specifically, when implementing, define an objective function with accuracy as the result. This function is the relationship function between many parameters in the binary correlation decision tree and the accuracy result. Assume the objective function is Then the required objective function can be sorted out as Assume \(f\sim GP(\mu, K)\), where \(GP\) represents the Gaussian process, \(\mu\) is the mean, and \(K\) is the covariance. Then the prediction of the objective function also belongs to the normal distribution, that is, there is The calculation of hyperparameters is expressed as The specific formula is as follows:

[0076] \(x\) *= argmax f(min_samples_split, max_depth, max_features, p(acc|X, D));

[0077] where x * is the solution to be found, min_samples_split, max_depth, and max_features are three specific required metrics, all belonging to the input set X, p(acc|X, D)) is the Gaussian distribution between the input set and the target accuracy acc, and D is the solution set of the solution.

[0078] After that, the selection of X can be achieved through the calculation of the acquisition function. In the embodiments of the present invention, the acquisition function adopts the EI function, and the EI function is an expectation function that can map the input x to the real number space and give the probability of how much the objective function value at this point is greater than the current optimal value. The specific formula is as follows:

[0079]

[0080] where f represents the objective function, D 1:t represents the set of sample points, p(D 1:t |f) represents the likelihood distribution, p(f) represents the prior probability distribution model, p(D 1:t ) represents the marginal likelihood distribution, p(f|D 1:t ) represents the posterior probability distribution, v * represents the optimal value of the objective function, ε represents the balance parameter, α t (x; D 1:t ) represents the possibility that the sample is the optimal objective function.

[0081] X can be obtained through the following formula:

[0082] x = argmax x E(max{0, f t+1 (x) - f(X + )}|D t ).

[0083] Set the number of cross-validations of the Bayesian optimization model to 5 times. The input of this function is the key parameters of the binary association decision tree, and the output is the average accuracy of the model cross-validated 5 times. The final model will output 10 sets of optimization results and the target values.

[0084] S6. Use the Bayesian optimization model to optimize several key parameters of the binary association decision tree;

[0085] Set the range of the maximum number of features max_features to 0.1 - 0.999, the maximum depth max_depth to 10 - 100, and the minimum number of samples in a leaf min_samples_split to 2 - 10.

[0086] Through optimization, the values of the three key parameters of the binary association decision tree are obtained respectively: max_depth is 57.73, max_features is 0.999, and min_samples_split is 2.

[0087] S7. In the binary association decision tree model, set the key parameters obtained by optimizing the Bayesian optimization model to the corresponding values, and use the default parameters for the remaining binary association decision tree parameters.

[0088] S8. Use the constructed binary association decision tree model to first train the training set, and then predict the test set, and output the prediction results of the binary association decision tree.

[0089] The multi-label decision tree heart failure prediction model proposed by the present invention is verified on the data collected from the Department of Cardiology of the First Affiliated Hospital of Dalian Medical University.

[0090] The binary association decision tree model designed by the present invention is optimized for the optimal parameters through Bayesian optimization, and then the multi-label prediction results are output. The multi-label decision tree prediction results are shown in Table 1, and the Bayesian optimization parameters are shown in Table 2. The micro-average AUC of the prediction results is as Figure 3 shown. The AUC (area under the curve) value is the area under the ROC curve. The larger the AUC value, the better the classifier performance. From Figure 3 it can be seen that the average AUC value reaches 0.81, exceeding the recognized effective limit of 0.75 in medicine. The AUC values of the control group and those who developed HFpEF one year later are both 0.84, and the curves of the two coincide in the figure, indicating that the model has a good prediction effect on these two labels. The AUC value of those who developed HfpEF within one year is 0.65, showing a poor performance. The reason for this effect may be that the model has insufficient generalization ability for patients with this type of label or due to insufficient training data volume.

[0091] Table 1

[0092]

[0093] Table 2

[0094]

[0095]

[0096] After testing, the binary association decision tree proposed by the present invention can well predict the onset probability and time of heart failure of the HFpEF type in patients with atrial fibrillation, which has great practical significance.

[0097] Corresponding to a multi-label prediction method for heart failure in patients with atrial fibrillation based on Bayesian optimization in the present invention, the present invention also provides a multi-label prediction device for heart failure in patients with atrial fibrillation based on Bayesian optimization, as Figure 4 shown. The device includes:

[0098] A data acquisition unit 100, configured to acquire information data of patients with atrial fibrillation and process the acquired information data of patients with atrial fibrillation;

[0099] A prediction model construction unit 200, which constructs a multi-label decision tree model; the decision tree in the multi-label decision tree model is a binary association decision tree, and the labels include the disease probability and the onset time;

[0100] A data division unit 300, configured to divide the data processed by the data acquisition unit 100 into two parts, a training set and a test set, according to the labels designed by the prediction model construction unit 200;

[0101] An optimization model construction unit 400, configured to build a Bayesian optimization model; the input of the objective function of the Bayesian optimization model is the key parameters of the binary association decision tree, and the output is the average accuracy of the model cross-validation times; the key parameters include: the maximum number of features, the maximum depth, and the minimum number of samples in a leaf;

[0102] An optimization unit 500, configured to use the Bayesian optimization model constructed by the optimization model construction unit 400 to optimize the key parameters of the binary association decision tree constructed by the prediction model construction unit 200;

[0103] A parameter setting unit 600, configured to set the key parameters optimized by the optimization unit 500 to corresponding values in the binary association decision tree constructed by the prediction model construction unit 200;

[0104] A prediction unit 700, configured to first train the training set obtained by the data division unit 300 by using the multi-label decision tree model after parameter setting by the parameter setting unit 600, and then predict the test set obtained by the data division unit 300, and output the prediction result of heart failure in patients with atrial fibrillation.

[0105] For the atrial fibrillation patient heart failure multi-label prediction device according to an embodiment of the present invention, since it corresponds to the atrial fibrillation patient heart failure multi-label prediction method based on Bayesian optimization in the above embodiment, the description is relatively simple. For relevant similarities, please refer to the description in the part of the atrial fibrillation patient heart failure multi-label prediction method based on Bayesian optimization in the above embodiment, and details will not be repeated here.

[0106] An embodiment of the present invention also discloses a computer-readable storage medium, which stores a computer instruction set. When the computer instruction set is executed by a processor, it implements the atrial fibrillation patient heart failure multi-label prediction method provided in any of the above embodiments.

[0107] In several embodiments provided by the present invention, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0108] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0109] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0110] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs.

[0111] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A multi-label prediction method for heart failure in atrial fibrillation patients based on Bayesian optimization, characterized in that, The method is executed by a computer and includes the following steps: Collect information data of atrial fibrillation patients and process the information data; Construct a multi-label decision tree model for heart failure of atrial fibrillation patients; the decision tree in the multi-label decision tree model for heart failure of atrial fibrillation patients is a binary association decision tree, and the labels include the disease probability and the onset time; According to the designed labels, divide the processed data into a training set and a test set; Build a Bayesian optimization model; the objective function of the Bayesian optimization model takes the key parameters of the binary association decision tree as input, and the output is the average precision of the model cross-validation times; the key parameters include: the maximum number of features, the maximum depth, and the minimum number of samples in a leaf; Use the Bayesian optimization model to optimize the key parameters of the binary association decision tree; set the key parameters optimized by the Bayesian optimization model to the corresponding values in the binary association decision tree; Use the multi-label decision tree model after parameter setting to first train the training set, and then predict the test set, and output the disease probability and the onset time of heart failure of atrial fibrillation patients; Among them, the prediction labels of the binary association decision tree are expressed as follows: Among them, represents the predicted label of the binary association decision tree, y j represents the label of the j-th generation and y j ∈Y i , 1 ≤ j ≤ q, q ∈ Q, Gini represents the Gini coefficient, D j represents a single-label training set, represents the input x i 's label set; The objective function in the Bayesian optimization model is: x * = argmax f(min_samples_split, max_depth, max_features, p(acc|X,D)); where x * is the solution to be found, min_samples_split, max_depth, and max_features are the minimum number of samples in a leaf, the maximum depth, and the maximum number of features, respectively, and all belong to the input set X. p(acc|X,D)) is the Gaussian distribution between the input set X and the target accuracy acc, and D is the solution set to be found.

2. The multi-label prediction method for heart failure of atrial fibrillation patients based on Bayesian optimization according to claim 1, wherein, The multi-labels include a first label [0, 0, 0], a second label [1, 1, 1], and a third label [1, 0, 1]. The first column in the label represents whether the patient has the disease, the second column represents that the onset time is less than 1 year, and the third column represents that the onset time is greater than 1 year.

3. A multi-label prediction method for heart failure in atrial fibrillation patients based on Bayesian optimization according to claim 1, characterized in that Collect information data of atrial fibrillation patients, including: collecting information data of atrial fibrillation patients in the form of electronic health records, and the electronic health records include collecting the demographic information, laboratory information, and clinical information data of the patients.

4. A multi-label prediction method for heart failure in atrial fibrillation patients based on Bayesian optimization according to claim 1, characterized in that, Process the information data, including: Convert text-based unstructured data into numerical structured data.

5. A multi-label prediction method for heart failure in atrial fibrillation patients based on Bayesian optimization according to claim 1, characterized in that, The optimization process includes: The Bayesian optimization model first randomly generates an initialization point, and then trains with Gaussian process regression to generate a posterior distribution about the objective function; Use the acquisition function to calculate the corresponding input and output, and then judge whether the output meets the objective function, and repeat continuously until the optimal value is selected or the iteration times are reached.

6. A multi-label prediction device for heart failure in atrial fibrillation patients based on Bayesian optimization, characterized in that, The device includes: A data collection unit for collecting information data of atrial fibrillation patients and processing the information data; A prediction model construction unit for constructing a multi-label decision tree model for heart failure of atrial fibrillation patients; the decision tree in the multi-label decision tree model for heart failure of atrial fibrillation patients is a binary association decision tree, and the labels include the disease probability and the onset time; A data division unit for dividing the processed data into a training set and a test set according to the designed labels; An optimization model construction unit for building a Bayesian optimization model; the objective function of the Bayesian optimization model takes the key parameters of the binary association decision tree as input, and the output is the average precision of the model cross-validation times; the key parameters include: the maximum number of features, the maximum depth, and the minimum number of samples in a leaf; An optimization unit for using the Bayesian optimization model constructed by the optimization model construction unit to optimize the key parameters of the binary association decision tree constructed by the prediction model construction unit; A parameter setting unit, configured to set the key parameters optimized by the optimization unit to corresponding values in the binary association decision tree constructed by the prediction model construction unit; A prediction unit, configured to first train a training set by using the multi-label decision tree model after parameter setting, and then predict a test set, and output the probability of suffering from heart failure and the onset time of atrial fibrillation patients; wherein, the prediction labels of the binary association decision tree are represented as follows: Among them, represents the predicted label of the binary association decision tree, y j represents the label of the j-th generation and y j ∈Y i , 1 ≤ j ≤ q, q ∈ Q, Gini represents the Gini coefficient, D j represents a single-label training set, represents the input x i 's label set; The objective function in the Bayesian optimization model is: x * = argmax f(min_samples_split, max_depth, max_features, p(acc|X,D)); where x * is the solution to be found, min_samples_split, max_depth, and max_features are the minimum number of samples in a leaf, the maximum depth, and the maximum number of features, respectively, and all belong to the input set X. p(acc|X,D)) is the Gaussian distribution between the input set X and the target accuracy acc, and D is the solution set to be found.

7. A computer-readable storage medium, in which a computer instruction set is stored, and when the computer instruction set is executed by a processor, the multi-label prediction method for heart failure of atrial fibrillation patients based on Bayesian optimization according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • CART incremental learning classification method based on selective integration for flight delay prediction

    CN109816010A

  • Top-down scene prediction based on motion data

    CN114245885A