Atrial fibrillation prediction method and system based on deep learning and polygene scoring

By combining deep learning, clinical risk scores and multigene scores, and fusing ECG signaling, clinical characteristics and genetic data, the limitations of existing prediction methods are solved, and more accurate and stable prediction of atrial fibrillation is achieved.

CN120167975AActive Publication Date: 2025-06-20HANGZHOU BEIZUO HEALTH TECH CO LTD

Patent Information

Application Number
CN202510652707.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-06-20
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

The existing atrial fibrillation prediction methods have problems such as limitations in traditional risk scores, single gene risk scores, and lack of multi-source data fusion analysis in deep learning models, resulting in low prediction accuracy.

Method used

Atrial fibrillation prediction method based on deep learning and multigene score is adopted, and ECG data, clinical data and gene data are collected, pre-deployed AF prediction model is called, multimodal data characteristics are identified, and AF risk tags and scoring reports are generated.

Benefits of technology

Significantly improves the accuracy and stability of atrial fibrillation prediction, providing a comprehensive individualized risk assessment, supporting real-time analysis and clinical decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120167975A_ABST
    Figure CN120167975A_ABST
Patent Text Reader

Abstract

The invention provides an atrial fibrillation prediction method based on deep learning and polygene scoring, which combines electrocardiogram (ECG) signals, clinical feature data and polygene risk scoring (PRS) based on multi-modal data fusion. By integrating the deep learning model, clinical risk score and polygene risk score, the accuracy and stability of AF prediction are significantly improved. The technical value is embodied in that: multi-source data fusion: combining ECG signals, clinical data and gene information in a single system for the first time, and providing comprehensive individualized risk assessment; real-time analysis capability: based on ECG real-time data processing capability, the real-time analysis capability can be used for dynamic risk assessment; and clinical decision support: a scientific and quantitative AF prediction report is provided for medical staff, and formulation of an intervention strategy is assisted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent medical technology, and particularly to an atrial fibrillation prediction method based on deep learning and polygenic scoring, an atrial fibrillation prediction system based on deep learning and polygenic scoring, an electronic device, and a computer-readable storage medium. Background Art

[0002] Atrial fibrillation (AF) is a common arrhythmia that affects millions of patients globally and has a high incidence and disability rate. Early detection and intervention of AF are important means to reduce the occurrence of complications such as stroke and heart failure. Traditional AF symptom detection methods, such as using scoring systems like CHARGE-AF, can quickly detect AF symptoms based on patients' clinical indicators (such as heart rate); or predict AF based on genetic data. However, the current prediction methods have the following limitations: 1. Limitations of traditional risk scoring: Scoring systems such as CHARGE-AF rely on limited clinical indicators, are difficult to capture complex individual characteristics, and have low prediction accuracy; 2. Singularity of genetic risk scoring (PRS): Predictions based only on genetic data do not fully consider environmental and physiological factors and have limited applicability; 3. Isolated application of deep learning models: Existing deep learning models mostly focus on ECG data but lack the fusion analysis of multi-source data. Summary of the Invention

[0003] To solve the technical problems existing in the prior art, the present invention provides the following technical solutions: On the one hand, an atrial fibrillation prediction method based on deep learning and polygenic scoring is provided. This method is implemented by an electronic device and includes: S1. Collect clinical data of AF patients: ECG data, clinical data, and genetic data; S2. Invoke a pre-deployed AF prediction model to identify multi-modal data features in the clinical data of AF patients and predict AF risk labels corresponding to the multi-modal data features; S3. Write the AF risk score in the clinical data and the AF risk label into the electronic medical record file of the AF patient to generate an AF risk score report.

[0004] Preferably, the method for generating the AF prediction model includes: Construct a multi-modal data set, including ECG data, clinical data, and genetic data collected and diagnosed from several AF patients in history; Extract multi-modal data features from the multi-modal data set, including: Extract the ECG signals diagnosed as AF from the ECG data; and Extract the basic information of AF patients, the medical history information of risk factors affecting AF, the drug information of AF clinical medications, and the clinical examination results at the time of AF symptoms from the clinical data; and Extract the SNP (single nucleotide polymorphism) information related to atrial fibrillation from the gene data; Perform feature annotation on the multi-modal data features of each AF patient, including: Annotate the ECG signals diagnosed as AF as AF positive and the ECG signals not diagnosed as AF as AF negative; and Annotation of the clinical data: For discrete data, annotate according to the corresponding specific values; for categorical data, annotate according to its classification results; otherwise annotate as "missing"; and Annotation of the SNP (single nucleotide polymorphism) information: Through the PRS (polygenic risk score) rule, calculate the corresponding PRS value by weighted calculation, and annotate the corresponding risk group according to the PRS value; Fuse the feature annotation results of the ECG data, clinical data, and gene data of AF patients to generate the final AF risk label and add it to the multi-modal data features of each AF patient; Statistically analyze the multi-modal data features with the AF risk label added for each AF patient, form a data set for training the AF prediction model, and divide it into a training set, a validation set, and a test set according to a preset ratio; Use the training set as the input of the multi-modal training model and perform model training according to a preset training strategy to generate the AF prediction model; Verify the prediction performance of the AF prediction model using the validation set: If the verification passes, then test the processing performance of the AF prediction model using the test set: If the test passes, deploy and apply the AF prediction model, otherwise adjust the training strategy; If the verification fails, then reorganize the multi-modal data set and retrain.

[0005] Preferably, when extracting the basic information of AF patients from the clinical data, it further includes: Perform standardized transformation on the dimensional data in the clinical data through the z-score standardization method to obtain the corresponding dimensionless data.

[0006] Preferably, when annotating the SNP (single nucleotide polymorphism) information, it further includes: According to the known AF-related genes, perform status annotation on the SNP (single nucleotide polymorphism) information related to atrial fibrillation in the gene data of AF patients, including: Labeling "0" indicates no mutation; Labeling "1" indicates heterozygote; Labeling "2" indicates homozygote.

[0007] Preferably, through the PRS (Polygenic Risk Score) rule, the corresponding PRS value is calculated by weighted calculation, including: Let: , Where: is the number of selected AF-related gene markers; is the weight of the i-th gene marker, determined based on the association strength (effect size or statistical significance) of the marker with AF risk; determined through genome-wide association studies (GWAS) or similar genetic studies, revealing the association strength between a specific gene marker and AF risk, and the weight reflects the contribution of the marker to AF risk, usually measured using the effect size (logarithm of the odds ratio) or statistical significance; is the genotype or allele count of the i-th gene marker in the patient: 0, 1, or 2, indicating the number of risk alleles carried by the patient.

[0008] Preferably, the calculation method: Let:

[0009] Where: is the odds ratio of the i-th gene marker, indicating the association strength of the marker with AF risk; is the P-value of the i-th gene marker, indicating the statistical significance of the association; is the logarithmic transformation of the odds ratio, used to measure the effect size; is the negative logarithmic transformation of the P-value, used to measure the statistical significance. Usually, the smaller the P-value, the larger the value after negative logarithmic transformation, indicating higher significance; is the sign of the logarithm of the odds ratio, i.e., the direction of the association (positive or negative); is the absolute value of the logarithm of the odds ratio, used to ensure the positive contribution of the effect size.

[0010] On the other hand, a atrial fibrillation prediction system based on deep learning and polygenic scoring is provided. The atrial fibrillation prediction system based on deep learning and polygenic scoring is used to implement the above-mentioned atrial fibrillation prediction method based on deep learning and polygenic scoring. The system includes: A data acquisition module for collecting clinical data of AF patients: ECG data, clinical data, and gene data; An AF prediction module for calling a pre-deployed AF prediction model to identify multimodal data features in the clinical data of AF patients and predicting an AF risk label corresponding to the multimodal data features; A risk management module for writing the AF risk score in the clinical data and the AF risk label into the electronic medical record file of AF patients to generate an AF risk score report.

[0011] On the other hand, an electronic device is provided. The electronic device includes: a processor; a memory, and a computer-readable instruction is stored on the memory. When the computer-readable instruction is executed by the processor, any one of the methods in the above-mentioned atrial fibrillation prediction method based on deep learning and polygenic scoring is implemented.

[0012] On the other hand, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement any one of the methods in the above-mentioned atrial fibrillation prediction method based on deep learning and polygenic scoring.

[0013] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include: By integrating a deep learning model, a clinical risk score, and a polygenic risk score, and based on multimodal data fusion, the present invention combines electrocardiogram (ECG) signals, clinical feature data, and polygenic risk scores (PRS). Feature extraction of time series data (ECG) is performed through a deep learning algorithm, and at the same time, statistical methods are comprehensively used to process clinical features and gene data. Finally, an individualized AF risk score is generated through a fusion model. Therefore, the accuracy and stability of AF prediction can be significantly improved. Its technical value is reflected in: 1. Multi-source data fusion: For the first time, ECG signals, clinical data, and gene information are combined in a single system to provide a comprehensive individualized risk assessment.

[0014] 2. Real-time analysis ability: Based on the real-time data processing ability of ECG, it can be used for dynamic risk assessment.

[0015] 3. Clinical decision support: Provide a scientific and quantitative AF prediction report for medical staff to assist in formulating intervention strategies. Description of the Drawings

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0017] Figure 1 is a flowchart of a method for predicting atrial fibrillation based on deep learning and polygenic scoring provided by an embodiment of the present invention; Figure 2 is a schematic diagram of the training process of an AF prediction model provided by an embodiment of the present invention; Figure 3 is a block diagram of a system for predicting atrial fibrillation based on deep learning and polygenic scoring provided by an embodiment of the present invention; Figure 4 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Specific Embodiments

[0018] The following will describe the technical solutions in the present invention with reference to the accompanying drawings.

[0019] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or more advantageous than other embodiments or design solutions. Exactly speaking, the use of the word "example" is intended to present concepts in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.

[0020] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same. "(of)", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same.

[0021] In the embodiments of the present invention, sometimes subscripts such as W1 may be miswritten as non-subscript forms such as W1. When the difference is not emphasized, the meanings they express are the same.

[0022] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0023] An embodiment of the present invention provides a method for predicting atrial fibrillation based on deep learning and polygenic scores. This method can be implemented by an electronic device, which can be a terminal or a server. As Figure 1 shown in the flowchart of the method for predicting atrial fibrillation based on deep learning and polygenic scores, the processing flow of this method can include the following steps: S1. Collect clinical data of AF patients: ECG data, clinical data, and gene data; S2. Call a pre-deployed AF prediction model to identify multimodal data features in the clinical data of AF patients and predict AF risk labels corresponding to the multimodal data features; S3. Write the AF risk scores in the clinical data and the AF risk labels into the electronic medical record file of AF patients to generate an AF risk score report.

[0024] The objectives of the present invention are: 1. Improve the accuracy of AF prediction and reduce false positives and false negatives.

[0025] 2. Construct a new AF prediction system that can process and learn in real time to achieve efficient fusion analysis of multi-source data.

[0026] 3. Provide a flexible prediction verification mechanism to ensure the robustness of the prediction model in different data scenarios.

[0027] When specifically implemented, based on multimodal data fusion, electrocardiogram (ECG) signals, clinical feature data, and polygenic risk scores (PRS) are combined. Feature extraction is performed on time series data (ECG) through a deep learning algorithm, while clinical features and gene data are processed by comprehensive statistical methods. Finally, an individualized AF risk score is generated through a fusion model.

[0028] Regarding the acquisition of relevant AF data of AF patients by the system, it can be implemented in combination with the subsequent description. The HIS system can construct and manage the electronic medical record file of AF patients and generate an AF risk score report according to a preset medical report format.

[0029] The following mainly describes model construction.

[0030] As Figure 2 shown, preferably, the method for generating the AF prediction model includes: Construct a multimodal dataset, including ECG data, clinical data, and gene data collected and diagnosed from the history of several AF patients; Extract multimodal data features from the multimodal dataset, including: Extract the ECG signals diagnosed as AF in the ECG data; and Extract the basic information of AF patients, the medical history information of risk factors affecting AF, the drug information of AF clinical medications, and the clinical examination results at the time of the onset of AF symptoms from the clinical data; and Extract the SNP (single nucleotide polymorphism) information related to atrial fibrillation from the gene data; Perform feature annotation on the multi-modal data features of each AF patient, including: Label the ECG signals diagnosed as AF as AF positive, and label the ECG signals not diagnosed as AF as AF negative; and Labeling of the clinical data: For discrete data, label according to the corresponding specific value; for categorical data, label according to its classification result; otherwise label as "missing"; and Labeling of the SNP (single nucleotide polymorphism) information: Through the PRS (polygenic risk score) rule, calculate the corresponding PRS value by weighted calculation, and label the corresponding risk group according to the PRS value; Fuse the feature annotation results of the ECG data, clinical data, and gene data of AF patients to generate the final AF risk label, and add it to the multi-modal data features of each AF patient; Statistically analyze the multi-modal data features with the AF risk label added for each AF patient, form a data set for training the AF prediction model, and divide it into a training set, a validation set, and a test set according to a preset ratio; Use the training set as the input of the multi-modal training model, and perform model training according to a preset training strategy to generate the AF prediction model; Use the validation set to verify the prediction performance of the AF prediction model: If the verification passes, then use the test set to test the processing performance of the AF prediction model: If the test passes, then deploy and apply the AF prediction model, otherwise adjust the training strategy; If the verification fails, then reorganize the multi-modal data set and retrain.

[0031] The specific construction and application principles of the model are as follows: 1. Technical principle This system is based on multi-modal data fusion, combining electrocardiogram (ECG) signals, clinical feature data, and polygenic risk scores (PRS). Feature extraction of time series data (ECG) is performed through deep learning algorithms, while clinical features and gene data are processed by comprehensive statistical methods. Finally, an individualized AF risk score is generated through a fusion model.

[0032] 2. Data set construction 1) ECG data collection: Use wearable devices or medical-grade electrocardiogram devices to collect 12-lead ECG data, with data samples covering both healthy individuals and diagnosed AF patients.

[0033] 2) Clinical data collection: Include clinical indicators such as patient age, gender, BMI, presence of hypertension, diabetes, etc.

[0034] 3) Gene data acquisition: Adopt gene sequencing technology to obtain multi-gene locus information of patients, and use the publicly disclosed PRS calculation method in the literature to generate risk scores.

[0035] 4) Data annotation: Have cardiologists annotate the ECG data to clarify AF-positive or AF-negative samples, and mark the gene and clinical data according to known risks.

[0036] 5) Dataset division: To evaluate the generalization ability of the model, the dataset needs to be divided into a training set, a validation set, and a test set. For example: Training set: 70% of the dataset is used for model training. The training set data undergoes data augmentation processing (such as time translation, amplitude scaling, etc.).

[0037] Validation set: 15% of the dataset is used for model hyperparameter tuning and early stopping mechanisms to avoid overfitting.

[0038] Test set: 15% of the dataset is used for final performance evaluation. The test set is data outside the model training process to ensure the objectivity of the evaluation results.

[0039] 3. Annotation scheme In the present invention, the annotation scheme mainly focuses on the processing methods of electrocardiogram (ECG) data, clinical data, and gene data. To ensure the accuracy and reliability of the prediction model, the annotation process must be rigorous and scientific. The annotation scheme for each type of data is described in detail below: 1) Electrocardiogram (ECG) data annotation ECG signal collection and preprocessing: Collection method: Use high-quality 12-lead electrocardiogram instruments (such as Holter monitors) or wearable devices (such as smart watches or chest straps, etc.) to ensure that the collected data has no obvious artifacts or noise.

[0040] Preprocessing: Denoise, standardize, and segment the collected ECG signals. Denoising includes using low-pass filters to remove high-frequency noise, and standardization makes each segment of the ECG signal have a unified amplitude range and time scale.

[0041] Annotation criteria: Expert Diagnosis: The ECG data is annotated by professionally certified cardiologists, and it is judged whether atrial fibrillation (AF) exists in the ECG signal according to the diagnostic criteria. Generally, the signs of AF are irregular heart rhythm, rapid atrial rate, etc.

[0042] Annotation Process: Each segment of ECG signal will be annotated as one of the following two categories: AF Positive (1): If the diagnostic result of the ECG signal is atrial fibrillation.

[0043] AF Negative (0): If the diagnostic result of the ECG signal is normal or other arrhythmias (such as atrial premature beats, ventricular premature beats, etc.).

[0044] Annotation Format: The annotation data will be annotated in units of time periods. For example, the ECG signal per minute is used as an annotation unit and annotated as "1" or "0".

[0045] Quality Control of Annotation: Double Review Mechanism: Each ECG annotation result is reviewed twice. Two cardiologists independently review and discuss it to ensure the accuracy of the annotation result.

[0046] Uncertain Annotation: If the experts are uncertain about whether AF occurs, the label "undetermined" (-1) can be used and further analyzed.

[0047] 2) Clinical Data Annotation Clinical data mainly includes the patient's basic information, medical history, physical sign data, etc. To ensure the effectiveness of the clinical data for the model, the annotation scheme requires unified formatting and standardization of the following data: Data Type and Content: Basic Information: Age, gender, height, weight, BMI.

[0048] Medical History Information: Risk factors for known atrial fibrillation such as hypertension, diabetes, high cholesterol, etc.

[0049] Drug Information: Whether taking anticoagulants, cardiac drugs, etc., and their usage duration.

[0050] Clinical Examination Results: Such as electrocardiogram, echocardiogram, blood pressure, heart rate, blood sugar, etc.

[0051] Annotation Method: Discrete Data: For numerical data such as age, BMI, blood pressure, etc., input according to the specific value to ensure no data loss. For categorical data (such as whether having hypertension), annotate as "yes" or "no".

[0052] Missing data handling: For missing clinical data, reasonable imputation methods (such as mean filling, forward and backward imputation, etc.) are used for filling, or marked as "missing" and specially processed in subsequent training.

[0053] Data normalization and standardization: All clinical feature data are standardized so that they can be processed on the same scale. For example, features with different dimensions such as height, weight, and blood glucose are transformed into dimensionless data through z-score standardization.

[0054] 3) Gene data annotation Gene data is the third type of input data of the present invention, mainly polygenic risk score (PRS). The annotation scheme for this part includes the following content: Gene data collection and preprocessing: Collection method: Gene data is obtained through high-throughput genomic sequencing technology, mainly collecting SNP (single nucleotide polymorphism) information related to atrial fibrillation.

[0055] Data standardization: The collected gene data is standardized to ensure that each SNP value is analyzed under the same standard. Generally, "0" (no mutation), "1" (heterozygote), and "2" (homozygote) are used as annotations.

[0056] Annotation standard: Polygenic risk score (PRS): According to known AF-related gene markers, the PRS value of each patient is generated through weighted calculation. The higher the PRS value, the greater the genetic risk of the patient developing AF.

[0057] Annotation method: The generated PRS value ranges from 0 to 1, and a higher value indicates a higher AF risk. Risk stratification is determined through statistical analysis. For example, the PRS value is divided into a low-risk group (0 - 0.3), a medium-risk group (0.3 - 0.7), and a high-risk group (0.7 - 1.0).

[0058] Supplementary description of gene data: During the annotation process, gene data may contain some rare variations or unclear genotypes. In this case, "missing" annotation can be used, and this feature is used for missing data handling in subsequent training.

[0059] 4 Comprehensive annotation Annotation of fused data: For each patient, the annotations of their ECG data, clinical data, and gene data are combined together to form the final comprehensive dataset.

[0060] Comprehensive annotation results: Each sample will generate a final label based on the diagnostic results of the ECG signal (AF positive or negative), the markers of clinical data (such as age, medical history), and the polygenic risk score (PRS). This label is used to train a deep learning model to predict the AF risk of patients.

[0061] 5. Model Training 1) Model architecture: Adopt a multi-layer fusion network structure (user-defined selection): ECG feature extraction module: Based on the CNN-LSTM network, extract spatio-temporal features.

[0062] Clinical feature module: Use a fully connected network to classify and perform risk analysis on structured data.

[0063] Gene risk score module: Use a linear regression model to weight the PRS value.

[0064] 2) Training strategy: Data augmentation: Perform time shift, amplitude scaling, etc. on the ECG signal to expand the sample size.

[0065] Loss function: Use weighted cross-entropy loss to balance positive and negative samples.

[0066] Optimization algorithm: Adam optimizer, with the initial learning rate set to 0.001.

[0067] 6. Algorithm Performance Requirements The prediction accuracy of the model needs to reach ≥90%.

[0068] The recall rate needs to be ≥85% to ensure the effective identification of high-risk patients.

[0069] The AUC of the ROC curve needs to be ≥0.95.

[0070] 7. Algorithm Performance Verification Methods 1) Cross-validation: Use K-fold cross-validation to evaluate the generalization performance of the model.

[0071] 2) Independent test set verification: Use an independent external dataset to test the stability of the model.

[0072] 3) Clinical simulation experiment: Verify the reliability of the model prediction results through retrospective data and compare with the traditional CHARGE-AF score.

[0073] 4) Real-time performance test: Verify the real-time processing ability and prediction accuracy of the model in a real environment.

[0074] In this embodiment, the multi-layer fusion network structure for multi-modal feature training and learning mainly includes Type-A: Standard Cross-Attention based Deep Fusion (SCDF), Type-B: Custom Layer based Deep Fusion (CLDF), and other non-standardized structures such as Type-C.

[0075] For example: ‌Type-A: Standard Cross-Attention based Deep Fusion (SCDF)‌ This structure performs deep fusion based on the standard cross-attention mechanism. In multi-modal tasks, information from different modalities interacts and fuses through the cross-attention mechanism, thereby extracting richer feature representations. This structure performs well in processing multi-modal data such as images and text, and can effectively capture the correlation information between different modalities.

[0076] ‌Type-B: Custom Layer based Deep Fusion (CLDF)‌ The CLDF structure relies on custom fusion layers for deep fusion of multi-modal features. These custom layers can be designed according to specific task requirements to better capture and utilize the complementary information between multi-modal data. By introducing these custom layers, the network can learn more complex and refined multi-modal feature representations.

[0077] Preferably, when extracting the basic information of AF patients from the clinical data, it further includes: Through the z-score standardization processing method, the dimensional data in the clinical data is standardized and transformed to obtain the corresponding dimensionless data.

[0078] Z-score standardization, Z-score standardization is a commonly used data preprocessing method. Its main purpose is to convert data with different dimensions or units into a unified scale for easy comparison and analysis. Z-score standardization only requires calculating the mean and standard deviation to complete, and is applicable to various types of data. The calculation process is simple and efficient, enabling analysts to quickly perform data preprocessing. ‌Eliminate the influence of dimensions‌: Through standardization processing, the unit and magnitude differences between different features can be eliminated, enabling data to be compared and analyzed on the same scale.

[0079] Preferably, when annotating the SNP (single nucleotide polymorphism) information, it further includes: Based on the known AF-related genes, annotate the status of SNP (single nucleotide polymorphism) information related to atrial fibrillation in the gene data of AF patients, including: Annotate "0", indicating no mutation; Annotate "1", indicating heterozygote; Annotate "2", indicating homozygote.

[0080] For the specific annotation, see the above training steps.

[0081] Preferably, according to the PRS (polygenic risk score) rule, calculate the corresponding PRS value by weighted calculation, including: Let: , Where: is the number of selected AF-related gene markers; is the weight of the i-th gene marker, determined based on the strength of its association with the AF risk (effect size or statistical significance); determined through genome-wide association studies (GWAS) or similar genetic studies, revealing the strength of the association between a specific gene marker and the AF risk, and the weight reflects the contribution of the marker to the AF risk, usually measured using the effect size (logarithm of the odds ratio) or statistical significance; is the genotype or allele count of the i-th gene marker in the patient: 0, 1, or 2, indicating the number of risk alleles carried by the patient.

[0082] Through the above weighted calculation, the following technical effects can be achieved: 1. Personalized risk assessment: By calculating the PRS value of the patient, the personalized genetic risk of developing AF can be evaluated, which helps in clinical decision-making and early intervention.

[0083] 2. Stratified management: According to the PRS value, patients can be divided into different risk layers, thus implementing more targeted prevention and treatment strategies.

[0084] 3. Scientific research: The PRS value can be used as a research variable to explore the association between AF and other diseases or biomarkers, and to study the mechanism of action of genetic factors in the development of AF.

[0085] By integrating the information of multiple gene markers, the PRS value provides a more comprehensive genetic risk assessment than a single gene marker.

[0086] Preferably, the calculation method: Let: , Wherein: is the odds ratio of the i-th genetic marker, representing the strength of the association between the marker and the AF risk; is the P-value of the i-th genetic marker, representing the statistical significance of the association; is the logarithmic transformation of the odds ratio, used to measure the effect size; is the negative logarithmic transformation of the P-value, used to measure the statistical significance. Generally, the smaller the P-value, the larger the value after negative logarithmic transformation, indicating higher significance; is the sign of the logarithm of the odds ratio, i.e., the direction of the association (positive or negative); is the absolute value of the logarithm of the odds ratio, used to ensure the positive contribution of the effect size.

[0087] Specifically, when calculating: Directivity: Through to determine the direction of the association. If the odds ratio is greater than 1, it indicates that the marker increases the AF risk (positive direction); if it is less than 1, it indicates a decrease in the AF risk (negative direction).

[0088] Statistical significance: Through adjust the weights to ensure that more statistically significant associations have a greater contribution in the PRS calculation.

[0089] Effect size: Through to reflect the actual effect size of each marker on the AF risk.

[0090] In practical applications, it may be necessary to perform certain preprocessing on the odds ratio and P-value, such as truncating extremely small P-values to avoid excessive weights. The formula for the weight can be adjusted according to specific research purposes and data characteristics. For example, other statistical measures or transformation methods can be considered. The finally calculated PRS value should be based on scientific research and verification to ensure its accuracy and reliability in practical applications.

[0091] Figure 3 is a block diagram of an atrial fibrillation prediction system based on deep learning and polygenic scoring shown according to an exemplary embodiment. The system is used for an atrial fibrillation prediction method based on deep learning and polygenic scoring. The system includes: A data acquisition module for collecting clinical data of AF patients: ECG data, clinical data, and genetic data; An AF prediction module, which is used to call a pre-deployed AF prediction model, identify multi-modal data features in the clinical data of AF patients, and predict AF risk labels corresponding to the multi-modal data features; A risk management module, which is used to write the AF risk score in the clinical data and the AF risk label into the electronic medical record file of the AF patient to generate an AF risk score report.

[0092] For the functions and interaction steps of each of the above modules, please refer to the above method steps for understanding and implementation, and will not be elaborated in this embodiment.

[0093] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 4 shown, the electronic device may include the Figure 3 atrial fibrillation prediction system based on deep learning and multi-gene scoring shown above. Optionally, the electronic device 410 may include a first processor 2001.

[0094] Optionally, the electronic device 410 may further include a memory 2002 and a transceiver 2003.

[0095] Among them, the first processor 2001, the memory 2002, and the transceiver 2003 may be connected through a communication bus, for example.

[0096] Next, Figure 4 a specific introduction to each component of the electronic device 410 will be given: Among them, the first processor 2001 is the control center of the electronic device 410, which may be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or may be a specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. For example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).

[0097] Optionally, the first processor 2001 may execute various functions of the electronic device 410 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.

[0098] In a specific implementation, as an example, the first processor 2001 may include one or more CPUs, such as Figure 4 the CPU0 and CPU1 shown in

[0099] In a specific implementation, as an example, the electronic device 410 may also include multiple processors, such as Figure 4 the first processor 2001 and the second processor 2004 shown in. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0100] Among them, the memory 2002 is used to store the software program for executing the solution of the present invention and is controlled by the first processor 2001 for execution. The specific implementation manner may refer to the above method embodiments and will not be elaborated here.

[0101] Optionally, the memory 2002 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001 or exist independently and be coupled to the first processor 2001 through the interface circuit of the electronic device 410 ( Figure 4 not shown in). The embodiments of the present invention do not make specific limitations on this.

[0102] The transceiver 2003 is used to communicate with network devices or with terminal devices.

[0103] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 4 not separately shown in). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.

[0104] Optionally, the transceiver 2003 may be integrated with the first processor 2001 or exist independently, and is coupled to the first processor 2001 through an interface circuit ( Figure 4 not shown) of the electronic device 410. The embodiments of the present invention do not make specific limitations thereto.

[0105] It should be noted that Figure 4 the structure of the electronic device 410 shown in does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0106] In addition, the technical effects of the electronic device 410 may refer to the technical effects of the atrial fibrillation prediction method based on deep learning and polygenic scoring described in the above method embodiments, which will not be elaborated herein.

[0107] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0108] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0109] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable systems. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0110] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.

[0111] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0112] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0113] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0114] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices, systems, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0115] In several embodiments provided by the present invention, it should be understood that the disclosed devices, systems, and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of systems or units can be in an electrical, mechanical, or other form.

[0116] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0117] In addition, the functional units in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0118] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0119] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for predicting atrial fibrillation based on deep learning and polygenic scoring, characterized in that: The method comprises: S1. Collect clinical data of AF patients: ECG data, clinical data and genetic data; S2. Calling a pre-deployed AF prediction model to identify multimodal data features in the clinical data of the AF patient, and predicting an AF risk label corresponding to the multimodal data features; S3. Write the clinical data and the AF risk score in the AF risk tag into the electronic medical record of the AF patient to generate an AF risk score report.

2. The atrial fibrillation prediction method based on deep learning and polygenic scoring according to claim 1, characterized in that: The method for generating the AF prediction model comprises: Construct a multimodal dataset including ECG data, clinical data, and genetic data collected from several AF patients at their historical diagnosis; Extracting multimodal data features from the multimodal data set includes: extracting an ECG signal diagnosed as AF from the ECG data; and Extracting basic information of AF patients, medical history information of risk factors affecting AF, drug information of clinical medication for AF, and clinical examination results when AF symptoms occur from the clinical data; and Extracting SNP (single nucleotide polymorphism) information related to atrial fibrillation from the gene data; The multimodal data features of each AF patient are annotated, including: Annotating ECG signals diagnosed as AF as AF-positive and ECG signals not diagnosed as AF as AF-negative; and Labeling of the clinical data: for discrete data, labeling according to the corresponding specific values; for categorical data, labeling according to the classification results; otherwise, labeling as "missing"; and Annotation of the SNP (single nucleotide polymorphism) information: using the PRS (polygenic risk score) rule, weighted calculation of the corresponding PRS value, and annotation of the corresponding risk group according to the PRS value; Fusing the feature annotation results of the ECG data, clinical data and genetic data of the AF patient to generate a final AF risk label, and adding it to the multimodal data features of each AF patient; Counting the multimodal data features of each AF patient with the AF risk label added thereto, forming a data set for training the AF prediction model, and dividing the data set into a training set, a validation set, and a test set according to a preset ratio; Using the training set as the input of the multimodal training model, performing model training according to a preset training strategy to generate the AF prediction model; The prediction performance of the AF prediction model was verified using the validation set: If the verification is passed, the processing performance of the AF prediction model is tested using the test set: if the test is passed, the AF prediction model is deployed and applied, otherwise the training strategy is adjusted; If the verification fails, the multimodal dataset is reorganized and retrained.

3. The atrial fibrillation prediction method based on deep learning and polygenic scoring according to claim 2, characterized in that: When extracting the basic information of the AF patient in the clinical data, it also includes: The dimensional data in the clinical data is standardized and transformed through the z-score standardization processing method to obtain the corresponding dimensionless data.

4. The atrial fibrillation prediction method based on deep learning and polygenic scoring according to claim 2, characterized in that: When annotating the SNP (single nucleotide polymorphism) information, it also includes: According to known AF-related genes, the status of SNP (single nucleotide polymorphism) information related to atrial fibrillation in the gene data of AF patients is annotated, including: The label "0" indicates no mutation; The label "1" indicates heterozygote; The label "2" indicates homozygous.

5. The method for predicting atrial fibrillation based on deep learning and polygenic scoring according to claim 4, characterized in that: The PRS (polygenic risk score) rule is used to weight the corresponding PRS value, including: make: , in: The number of AF-related gene markers selected; is the weight of the ith gene marker, determined based on the strength of the association between the marker and AF risk (effect size or statistical significance); determined by genome-wide association studies (GWAS) or similar genetic studies, revealing the strength of the association between a specific gene marker and AF risk. The weight reflects the contribution of the marker to AF risk, usually measured by effect size (logarithm of odds ratio) or statistical significance; Count the genotype or allele of the i-th gene marker in the patient: 0, 1, or 2, indicating the number of risk alleles carried by the patient.

6. The atrial fibrillation prediction method based on deep learning and polygenic scoring according to claim 5, characterized in that: Said Calculation method: make: in: is the odds ratio of the i-th gene marker, indicating the strength of association between the marker and AF risk; is the P value of the ith gene marker, indicating the statistical significance of the association; is the logarithmic transformation of odds ratio, used to measure effect size; It is the negative logarithmic transformation of the P value, which is used to measure statistical significance. Usually, the smaller the P value, the larger the value after negative logarithmic transformation, indicating a higher significance. is the sign of the logarithm of the odds ratio, that is, the direction of the association (positive or negative); is the absolute value of the logarithm of the odds ratio, which is used to ensure a positive contribution to the effect size.

7. A deep learning and polygenic scoring-based atrial fibrillation prediction system, wherein the deep learning and polygenic scoring-based atrial fibrillation prediction system is used to implement the deep learning and polygenic scoring-based atrial fibrillation prediction method according to any one of claims 1 to 6, characterized in that: The system comprises: Data collection module, used to collect clinical data of AF patients: ECG data, clinical data and genetic data; An AF prediction module, configured to call a pre-deployed AF prediction model, identify multimodal data features in the clinical data of AF patients, and predict an AF risk label corresponding to the multimodal data features; The risk management module is used to write the clinical data and the AF risk score in the AF risk tag into the electronic medical record of the AF patient and generate an AF risk score report.

8. An electronic device, characterized in that: The electronic device comprises: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method for judging myocardial infarction by inputting electrocardiogram and clinical data through multi-mode method

    CN114209334A

  • Deep learning-based paroxysmal atrial fibrillation detection method and system

    CN116172568A

  • Atrial fibrillation risk assessment method based on machine learning

    CN118398208A

  • Atrial fibrillation detection and risk factor mining method based on machine learning and Transform

    CN118782233A

  • Disease risk prediction method based on multi-modal graph learning model

    CN119207552A

Cited By

  • Wearable senile atrial fibrillation early warning equipment and monitoring system based on multi-modal fusion

    CN121129229A