Atrial Fibrillation Detection and Risk Factor Mining Method Based on Machine Learning and Transformer
Through the Multimodal-AF Transformer model based on Transformer, atrial fibrillation detection is combined with multimodal data of electronic health records, the problem of paroxysmal atrial fibrillation misdiagnosis is solved, and high-precision and interpretability detection effect is achieved to assist clinical medical staff in diagnosis.
Patent Information
- Application Number
- CN202410973984.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-07-19
AI Technical Summary
The existing atrial fibrillation detection methods have a high rate of misdiagnosis in paroxysmal and asymptomatic atrial fibrillation, and rely on expertise and susceptibility to noise interference, affecting detection accuracy and interpretability.
The Multimodal-AF Transformer model based on Transformer is used to detect atrial fibrillation by combining multimodal data from electronic health records. Stacking integrated machine learning model and Shap interpretability analysis algorithm are used to mine risk factors, and the electrocardiogram signal and immediate admission information are processed through the Swin Transformer framework to achieve high-precision and interpretability detection.
It significantly improves the accuracy and interpretability of atrial fibrillation detection, helps clinical medical staff to efficiently analyze problems related to atrial fibrillation, and improves the accuracy and reliability of the detection.
Smart Images

Figure CN118782233B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of software engineering and artificial intelligence algorithm development, and specifically relates to automated medical record entry, intelligent detection and interpretability analysis in atrial fibrillation analysis, and particularly relates to a method for detecting atrial fibrillation and mining risk factors based on machine learning and Transformer. Background Art
[0002] Atrial fibrillation (AF), as a common arrhythmia, has a serious impact on global health. Atrial fibrillation is an arrhythmia characterized by rapid and irregular contractions of the atria. This irregular atrial contraction can cause blood stasis in the atria, increasing the risk of thrombus formation. Especially in the left atrium or left atrial appendage, the formed thrombus can be pumped into the systemic circulation system by the heart and may enter the cerebral blood vessels, leading to cerebral embolism and thus triggering ischemic cerebral infarction. The risk of cerebral infarction in patients with atrial fibrillation increases significantly. Atrial fibrillation is one of the most common arrhythmias and an independent risk factor for cerebral infarction. According to statistics, the risk of cerebral infarction in patients with atrial fibrillation is about 5 times higher than that in non-atrial fibrillation patients. Traditionally, the detection of atrial fibrillation mainly relies on electrocardiogram (ECG) and 24-hour ambulatory electrocardiogram. Although these methods perform well in detecting persistent and permanent atrial fibrillation, they have significant limitations in the face of paroxysmal and asymptomatic atrial fibrillation. A standard 10-second electrocardiogram may lead to misdiagnosis of patients with paroxysmal atrial fibrillation, thus missing the treatment opportunity. In addition, the analysis process of electrocardiogram is complex, relying on the professional knowledge and experience of doctors, and is easily affected by noise and other signal interferences, resulting in unstable detection results.
[0003] The detection of atrial fibrillation, especially the screening of atrial fibrillation and the mining of related risk factors, is crucial for early identification of patients and formulation of medical intervention plans. Currently, commonly used atrial fibrillation detection means, such as standard 12-lead electrocardiogram, have obvious limitations in detecting paroxysmal atrial fibrillation; while long-term electrocardiogram monitoring such as 24-hour ambulatory electrocardiogram may affect the cooperation and quality of life of patients due to the heavy monitoring burden. Therefore, it is of great significance to design an effective method to assist clinical staff in detecting and analyzing potential atrial fibrillation in the population. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a Multimodal-AF Transformer model based on Transformer, a method for detecting atrial fibrillation and mining risk factors based on machine learning and Transformer, which significantly improves the detection accuracy and the interpretability of the model.
[0005] The object of the present invention is achieved by the following technical solutions: A method for detecting atrial fibrillation and mining risk factors based on machine learning and Transformer, comprising the following steps:
[0006] (1) Automatic information entry: Use an automatic page-turning scanner to scan the patient's paper medical records and examination reports; then identify and extract the text information in the scanned images through OCR technology to obtain the patient's electronic health record data, which includes the admission instant information and chronological medical information;
[0007] (2) Mine atrial fibrillation risk factors: Extract the admission instant information from the electronic health records, and use the Stacking integrated machine learning model and Shap interpretability analysis algorithm to mine risk factors;
[0008] (3) Import electrocardiogram data, combine the admission instant information and chronological medical information, and use the Multimodal-AFTransformer model to process the data of the three modalities to achieve atrial fibrillation detection, and generate prediction results and probabilities.
[0009] 2. The method for detecting atrial fibrillation and mining risk factors based on machine learning and Transformer according to claim 1, wherein in the step (2), the Stacking method improves the detection performance by stacking the base learner and the meta-learner, and uses the outputs of multiple base learners as the inputs of the meta-learner to train the meta-learner; the base learner adopts standard machine learning models, including logistic regression, random forest, XGBoost and LightGBM; the meta-learner adopts a logistic regression model;
[0010] The specific process of step (2) is as follows:
[0011] (2-1) Data preprocessing: First, according to the inclusion and exclusion criteria, screen the features of the admission instant information, and exclude the features with more than 5% missing feature values; then use the KNN algorithm to fill in the missing feature values; finally, use the one-hot encoding method to encode the feature descriptions into a format that the model can recognize to obtain the training set D;
[0012] (2-2) Data augmentation: Process the data in the training set D through the SMOTENN data augmentation algorithm;
[0013] (2-3) Model training: Train T different base learners on the training set D. Each base learner ζ_i learns different features in the training set and gives a prediction result h_i for each sample x_i in the training set; these results are not directly used for the final decision, but are regarded as a new feature set;
[0014] (2-4) Stacking integration method: Construct a new training set D', which contains the prediction results of the base learners and the corresponding true labels y_i. Use the prediction results and the true labels as the input of the meta-learner ζ' to train the meta-learner and generate the final risk prediction model;
[0015] (2-5) According to the Shap method, obtain the prediction results of three types of risk factors: the top 20 features most relevant to the occurrence of atrial fibrillation, the influence of features on atrial fibrillation prediction at the population level, and the influence of features on atrial fibrillation prediction at the individual level;
[0016] (2-6) Risk threshold acquisition: According to the final prediction probability of the meta-learner and the feature value that has the greatest impact on the occurrence of atrial fibrillation obtained by the Shap method, construct a logistic regression function, obtain the second derivative of the function, take the point where the second derivative value is zero as the inflection point of risk change, and use the feature value corresponding to the inflection point as the corresponding risk threshold.
[0017] 3. The method for detecting atrial fibrillation and mining risk factors based on machine learning and Transformer according to claim 1, wherein the specific implementation method of the step (3) is: using the Multimodal-AF Transformer model based on the Swin Transformer framework to process multi-modal data; the multi-modal data includes three modalities: the admission instant information modality, the sequential medical information modality, and the electrocardiogram modality;
[0018] The admission instant information passes through the meta-learner of the previous Stacking model to obtain the admission instant information modality representation vector before the output of the hidden layer of the meta-learner;
[0019] Arrange the patient's course description and medication situation in chronological order, select the sEHR-BERT model to preprocess the sequential medical information, and obtain the sequential medical information modality representation vector;
[0020] The structure of the Multimodal-AF Transformer model includes multiple components:
[0021] R peak segmentation of heartbeat beats: Segment the electrocardiogram signal into multiple heartbeat beats according to the R peak, and input each single heartbeat beat into the image patch embedding layer of the Multimodal-AF Transformer model to capture the periodic feature vectors in the electrocardiogram;
[0022] Concatination layer: The features extracted from each heartbeat are respectively concatenated with the in-hospital immediate information modal representation vector and the temporal medical information modal feature vector to obtain multiple multimodal feature vectors composed of "feature vector of a single heartbeat + in-hospital immediate information modal + temporal medical information modal", and then input into the Encoder;
[0023] The Encoder includes multiple Encoder layers, and each Encoder layer consists of multiple Swin blocks and a cross-window information fusion block; the input of the Swin block in the first Encoder layer is the multimodal feature vector output by the concatination layer, and the input of the Swin block in other Encoder layers is the features split by the cross-window information fusion block in the previous Encoder layer; the output of the cross-window information fusion block in the last Encoder layer will be input into the classification module;
[0024] The classification module consists of a pooling layer and a multi-layer perceptron.
[0025] The beneficial effects of the present invention are as follows: The present invention adopts the Multimodal-AFTransformer model based on Transformer. This model utilizes the advanced Swin Transformer architecture, divides the electrocardiogram signal into multiple heartbeats according to the R peak, and inputs it into the model together with the representation information of the electronic health record. From the modal perspective, the in-hospital immediate information modal and the temporal medical information modal of the electronic health record, as the guidance for the electrocardiogram modality, can effectively solve the problem of difficult identification of paroxysmal atrial fibrillation. From the perspective of the model structure, the sliding window feature of the Swin block allows the model to effectively capture the local information in a single heartbeat, and the global information between different heartbeats is extracted through the cross-window information fusion block. The present invention uses multimodal data processing to detect atrial fibrillation. This multimodal method not only utilizes the direct heart rhythm signal of the electrocardiogram, but also integrates the rich background information in the electronic health record, significantly improving the detection accuracy. Through analysis and prediction by the machine learning model, the interpretability of the model is further improved, thus assisting clinical medical staff to efficiently and high-quality complete the analysis of atrial fibrillation-related problems. Description of the Drawings
[0026] Figure 1 It is a schematic structural diagram of the Swin Transformer;
[0027] Figure 2 It is a schematic diagram of the Stacking method of the present invention;
[0028] Figure 3 It is a schematic structural diagram of the Multimodal-AF Transformer model of the present invention;
[0029] Figure 4 Schematic diagram of the temporal medical information modal representation vector of the present invention;
[0030] Figure 5 Schematic diagram of beat segmentation by the Neurokit2 heartbeat segmentation technique of the present invention;
[0031] Figure 6 Schematic diagram of the Swin block structure of the present invention;
[0032] Figure 7 Schematic diagram of the structure of the cross-window information fusion block of the present invention. Detailed implementation manners
[0033] The present invention adopts the Multimodal-AF Transformer model based on Transformer. This model refers to the advanced Swin Transformer sliding window mechanism. First, the electrocardiogram signal is segmented into multiple heartbeat beats according to the R peak, and is jointly input into the model together with the representation information of the electronic health record. The sliding window feature of the Swin block allows the model to effectively capture the local information in the electrocardiogram, and the global information extraction between different heartbeat beats is realized through the cross-window information fusion block. In addition, the Shap interpretability algorithm is adopted in the model. Through this algorithm, the decision-making process of the model can be explained, providing an intuitive analysis of the detection results, helping clinicians understand and verify the decision-making process of the model, and further improving the feasibility and reliability of the model in clinical applications. The structure of the Swin Transformer is as Figure 1 shown. The reference of the Swin Transformer is: Liu, Ze, et al. "Swin transformer: Hierarchical vision transformer using shifted windows." Proceedings of the IEEE / CVF international conference on computer vision. 2021. The reference of Shap is: Lundberg, Scott M., and Su-In Lee. "A unified approach to interpreting model predictions." Advances in neural information processing systems 30(2017).
[0034] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0035] Atrial fibrillation detection and risk factor mining method based on machine learning and Transformer, comprising the following steps:
[0036] (1) Automatic information entry: Use an automatic page-turning scanner to scan the patient's paper medical records and examination reports; then use OCR technology to identify and extract the text information in the scanned images to obtain the patient's electronic health record (EHR) data. The electronic health record contains admission instant information and sequential medical information. The admission instant information is some static data registered when the patient is admitted to the hospital, and the sequential medical information is related to the patient's disease course description and medication situation, and these information are all time-related. Since the model requires the patient's electronic health record data, the present invention develops a method for automatic entry and extraction of patient cases, and uses OCR technology to realize the automatic entry of cases, which significantly improves the efficiency and accuracy of case entry. During the case entry process, the system can automatically identify and extract the patient's basic information and examination results, and integrate them into the database for the model to analyze.
[0037] (2) Conduct mining of atrial fibrillation risk factors: Extract the admission instant information from the electronic health record, and use the Stacking integrated machine learning model and the Shap interpretability analysis algorithm to conduct risk factor mining;
[0038] The Stacking method improves the detection performance by stacking the base learner and the meta-learner, and uses the outputs of multiple base learners as the inputs of the meta-learner to train the meta-learner to improve the detection performance, as Figure 2 shown; the base learner adopts standard machine learning models, including logistic regression, random forest, XGBoost and LightGBM, and the meta-learner adopts a logistic regression model;
[0039] The specific process of step (2) is as follows:
[0040] (2-1) Data preprocessing: First, according to the inclusion and exclusion criteria, screen the features of the admission instant information, and exclude the features with more than 5% missing feature values; in the task of mining atrial fibrillation risk factors, the admission instant information records the various features of the patient, such as text records of age, gender, past medical history, and various biochemical indicators. Use the KNN algorithm to fill in the missing feature values, set the number of neighbors to 5, and the weight parameter to "uniform", indicating that all neighbors have an equal impact on filling in the missing values; the algorithm is selected as "auto", allowing the system to automatically select the best algorithm path; finally, use the one-hot encoding method to encode the feature descriptions into a format that the model can recognize to obtain the training set D;
[0041] (2-2) Data augmentation: Process the data in the training set D through the SMOTENN data augmentation algorithm, increase the minority class samples, and remove the noise samples to improve the recognition accuracy of the model;
[0042] (2-3) Model training: Train T different base learners on the training set D. Each base learner ζ_i learns different features in the training set and gives a prediction result h_i for each sample x_i in the training set; these results are not directly used for the final decision but are regarded as a new feature set;
[0043] (2-4) Stacking ensemble method: Construct a new training set D' that contains the prediction results of the base learners and the corresponding true labels y_i. Use the prediction results and the true labels together as the input of the meta-learner ζ', train the meta-learner to learn how to optimally combine these prediction probabilities, and generate the final risk prediction model;
[0044] (2-5) According to the Shap method, obtain the prediction results of three types of risk factors: the top 20 features most relevant to the occurrence of atrial fibrillation, the influence of features on atrial fibrillation prediction at the population level, and the influence of features on atrial fibrillation prediction at the individual level.
[0045] (2-6) Risk threshold acquisition: According to the final prediction probability of the meta-learner and the feature value that has the greatest impact on the occurrence of atrial fibrillation obtained by the Shap method, construct a logistic regression function, obtain the second derivative of the function, take the point where the second derivative value is zero as the inflection point of the risk change, and take the feature value corresponding to the inflection point as the corresponding risk threshold.
[0046] The base learner and the meta-learner together constitute the model H(x). H(x) utilizes the learning ability of the meta-learner ζ' and comprehensively considers the prediction results h_1(x), h_2(x),..., h_T(x) of all base models to generate a more accurate and stable final prediction than the base models. The Stacking method not only improves the adaptability of the model to complex data but also reduces the dependence on the bias of any single model, significantly improving the accuracy of atrial fibrillation diagnosis.
[0047] The Shap method is a model interpretation method based on game theory that quantifies the importance of each feature by calculating its contribution to the prediction result. The Shap algorithm can provide interpretability analysis of the model at the population and individual levels, help clinicians understand the decision-making process of the model, and provide an intuitive interpretation of feature importance.
[0048] (3) Import the electrocardiogram data, combine the admission instant information and the sequential medical information, and use the Multimodal-AFTransformer model to process the data of the three modalities to achieve atrial fibrillation detection and generate prediction results and probabilities.
[0049] The specific implementation method of step (3) is as follows: As Figure 3 shown, the Multimodal-AF Transformer model based on the Swin Transformer framework is used to process multimodal data to achieve high-precision detection of atrial fibrillation; the multimodal data includes three modalities: the admission instant information modality, the temporal medical information modality, and the electrocardiogram modality.
[0050] The preprocessing methods for each modality are as follows:
[0051] Admission instant information modality: The admission instant information passes through the meta-learner of the previous Stacking model to obtain the admission instant information modality representation vector before the output of the hidden layer of the meta-learner, which is used as one of the inputs for the subsequent model.
[0052] Temporal medical information modality: The temporal medical information includes the course description and medication situation, etc. The patient's course description and medication situation are arranged in chronological order, and the BERT model that performs excellently in the field of natural language processing (NLP) - the sEHR-BERT model is selected to preprocess the temporal medical information to obtain the temporal medical information modality representation vector, as Figure 4 shown; sEHR-BERT is a BERT variant specifically designed for processing electronic health records, which can effectively capture the temporal information and context relationship of temporal medical information.
[0053] Electrocardiogram modality: First, the first three leads of the 10-second standard 12-lead electrocardiogram signal obtained are used as inputs, and the beat segmentation is performed through the Neurokit2 heartbeat segmentation technology to obtain the heartbeat beats cut by the R peak, as Figure 5 shown; Subsequently, each heartbeat beat is input into the image patch embedding layer of the Multimodal-AF Transformer model according to the batch size for feature extraction to obtain the local features in the electrocardiogram.
[0054] Subsequently, the local features in the electrocardiogram, the admission instant information modality representation vector, and the temporal medical information modality representation vector are input into the concatenation layer of the Multimodal-AF Transformer model for concatenation; then the concatenated features are input into the Dense layer for fine-tuning to ensure the consistency of the feature dimensions. The features output by the Dense layer are input into the atrial fibrillation detection encoder Encoder, which consists of multiple rounds of Swin blocks and cross-window information fusion blocks to achieve the extraction of local features and the information fusion of global features.
[0055] The structure of the Multimodal-AF Transformer model includes multiple key components:
[0056] R-peak segmentation of heartbeat beats: The electrocardiogram signal is segmented into multiple heartbeat beats according to the R peaks, and the image patch embedding layer of the Multimodal-AF Transformer model is input with a single heartbeat beat to capture the periodic feature vectors in the electrocardiogram.
[0057] Concatination layer: The features extracted from each heartbeat beat are respectively concatenated with the in-hospital immediate information modal representation vector and the temporal medical information modal feature vector to obtain multiple multi-modal feature vectors composed of "feature vector of a single heartbeat beat + in-hospital immediate information modal + temporal medical information modal", and then input into the Encoder;
[0058] The Encoder includes multiple Encoder layers, and each Encoder layer consists of multiple Swin blocks and a cross-window information fusion block; the input of the Swin block in the first Encoder layer is the multi-modal feature vector output by the concatination layer, and the input of the Swin block in other Encoder layers is the features split by the cross-window information fusion block in the previous Encoder layer; the output of the cross-window information fusion block in the last Encoder layer will be input into the classification module;
[0059] The Swin block can extract the local features in a single beat of the electrocardiogram and learn information more effectively. The Swin block is as Figure 6 shown in the reference: Liu, Ze, et al. "Swin transformer: Hierarchical vision transformer using shifted windows." Proceedings of the IEEE / CVF international conference on computer vision. 2021.
[0060] Multimodal data not only contains valuable information within a single heartbeat, but also extractable information between different heartbeats. It is not enough to rely solely on the single-beat information extraction of the Swin block. After every several Swin blocks in the model, a cross-window information fusion block is introduced. The cross-window information fusion block first concatenates the single-beat feature vectors extracted by different Swin blocks, and then inputs them into the convolutional layer and the pooling layer. In the convolutional layer, with a certain convolutional kernel size, it fully learns the global features between different heartbeats and performs layer normalization operations; in the pooling layer, it adjusts the overall dimension size according to a certain stride. After the convolutional layer and the pooling layer have learned the features, the fused feature vectors are re-split into the dimension sizes accepted by each Swin block according to different heartbeats, ensuring the consistency of the input and output dimensions of the cross-window information fusion block. Then it is input into each Swin block of the next Encoder layer for the next round of local feature extraction and global feature extraction operations. The cross-window information fusion block can combine multimodal data and achieve information fusion of global features through convolutional operations, making up for the deficiency that the Swin block can only learn local features in a single heartbeat. The cross-window information fusion block is as Figure 7 shown.
[0061] To ensure the consistency of the feature vector dimension in the fine-tuning task with that in the pre-training, the present invention sets a linear layer (Dense layer) behind the concatenation layer to complete the dimension adjustment, and then inputs the concatenated features into the Encoder.
[0062] The classification module consists of a pooling layer and a multi-layer perceptron (fully connected layer). The classification module uses the Sigmoid function as the activation function and selects an appropriate loss function for the final binary classification. During the training process, the model is first pre-trained on a large-scale public dataset (Physionet 2020) that only contains electrocardiograms, and then fine-tuned on a multimodal dataset that includes electronic health records and electrocardiogram data. Through cross-validation and hyperparameter optimization, the model achieves the best performance. The loss function of the model selects the Class-balanced Focal Loss to address the problem of data imbalance.
[0063] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.
Claims
1. A method for detecting atrial fibrillation and mining risk factors based on machine learning and Transformer, characterized in that, It includes the following steps: (1) Automatic information entry: Use an automatic page-turning scanner to scan the patient's paper medical records and examination reports; then use OCR technology to identify and extract the text information in the scanned images to obtain the patient's electronic health record data, which includes admission instant information and sequential medical information; (2) Mine atrial fibrillation risk factors: Extract the admission instant information from the electronic health record, and use the Stacking integrated machine learning model and the Shap interpretability analysis algorithm to mine risk factors; (3) Import electrocardiogram data, combine the admission instant information and sequential medical information, and use the Multimodal-AFTransformer model to process the data of the three modalities to achieve atrial fibrillation detection, and generate prediction results and probabilities; The specific implementation method is: Use the Multimodal-AF Transformer model based on the Swin Transformer framework to process multi-modal data; The multi-modal data includes three modalities: admission instant information modality, sequential medical information modality, and electrocardiogram modality; The admission instant information passes through the meta-learner of the previous Stacking model to obtain the admission instant information modality representation vector before the hidden layer output of the meta-learner; Arrange the patient's course description and medication situation in chronological order, and select the sEHR-BERT model to preprocess the sequential medical information to obtain the sequential medical information modality representation vector; The structure of the Multimodal-AF Transformer model includes multiple components: R-peak segmentation of heartbeat beats: Segment the electrocardiogram signal into multiple heartbeat beats according to the R peak, and input each single heartbeat beat into the image patch embedding layer of the Multimodal-AF Transformer model to capture the periodic feature vectors in the electrocardiogram; Concatenation layer: Concatenate the features extracted from each heartbeat beat with the admission instant information modality representation vector and the sequential medical information modality feature vector respectively to obtain multiple multi-modal feature vectors composed of "feature vector of a single heartbeat beat + admission instant information modality + sequential medical information modality", and then input them into the Encoder; The Encoder includes multiple Encoder layers, and each Encoder layer consists of multiple Swin blocks and a cross-window information fusion block; The input of the Swin block in the first Encoder layer is the multi-modal feature vector output by the concatenation layer, and the input of the Swin block in other Encoder layers is the features split by the cross-window information fusion block in the previous Encoder layer; The output of the cross-window information fusion block in the last Encoder layer will be input into the classification module; The classification module consists of a pooling layer and a multi-layer perceptron.
2. The method for detecting atrial fibrillation and mining risk factors based on machine learning and Transformer according to claim 1, characterized in that In step (2), the Stacking method improves the detection performance by stacking the base learners and the meta-learner, using the outputs of multiple base learners as the input of the meta-learner to train the meta-learner. The base learners adopt standard machine learning models, including logistic regression, random forest, XGBoost, and LightGBM. The meta-learner adopts a logistic regression model. The specific process of step (2) is as follows: (2-1) Data preprocessing: First, according to the inclusion and exclusion criteria, feature screening is performed on the immediate admission information, excluding features with missing values exceeding 5%. Then, the KNN algorithm is used to fill in the missing feature values. Finally, using the one-hot encoding method, the feature descriptions are encoded into a format recognizable by the model to obtain the training set D. (2-2) Data augmentation: Process the data in the training set D through the SMOTENN data augmentation algorithm. (2-3) Model training: Train T different base learners on the training set D. Each base learner ζ_i learns different features in the training set and gives a prediction result h_i for each sample x_i in the training set. These results are not directly used for the final decision but are regarded as a new feature set. (2-4) Stacking integration method: Construct a new training set D' that contains the prediction results of the base learners and the corresponding true labels y_i. Use the prediction results and the true labels together as the input of the meta-learner ζ' to train the meta-learner and generate the final risk prediction model. (2-5) According to the Shap method, obtain the prediction results of three types of risk factors: the top 20 features most relevant to the occurrence of atrial fibrillation, the influence of features on atrial fibrillation prediction at the population level, and the influence of features on atrial fibrillation prediction at the individual level. (2-6) Risk threshold acquisition: According to the final prediction probability of the meta-learner and the feature value that has the greatest impact on the occurrence of atrial fibrillation obtained by the Shap method, construct a logistic regression function, obtain the second derivative of the function, take the point where the second derivative value is zero as the inflection point of the risk change, and take the feature value corresponding to the inflection point as the corresponding risk threshold.
Citation Information
Patent Citations
Medical risk assessment and early warning method based on multi-modal data driving
CN117316451A