Sudden heart death early warning method based on wearable device signal and gene marker

By constructing a gene-physiology coupled heterogeneous map and a dynamic graph learning mechanism with sparsity constraints, the problem of insufficient causal explanation in the prediction of sudden cardiac death risk by deep learning models is solved, and high-transparency risk assessment and individualized representation are achieved, which is suitable for edge computing environments.

CN122024831APending Publication Date: 2026-05-12THE FIRST AFFILIATED HOSPITAL OF SUN YAT SEN UNIV +1
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
THE FIRST AFFILIATED HOSPITAL OF SUN YAT SEN UNIV
Filing Date
2026-02-04
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing deep learning-based models for predicting the risk of sudden cardiac death lack transparent causal explanation mechanisms, making it difficult to trace gene-physiological pathways. Furthermore, they ignore the causal sparsity and pathway coupling characteristics in real biological networks, affecting clinical credibility and the development of individualized intervention strategies.

Method used

By constructing a heterogeneous gene-physiology coupling map, a sparse constraint dynamic graph learning mechanism is adopted, which combines cross-modal mutual information and Granger causality strength to generate a sparse gene-physiology coupling map. Symbolic regression technology is used to generate an interpretable risk evolution equation, and risk scores are updated in real time. The dominant gene-physiology interaction chain is identified by combining a decision tree model.

Benefits of technology

It significantly improves the temporal sensitivity and modeling accuracy of sudden cardiac death risk assessment, achieves highly transparent risk interpretation and mechanism visualization, supports individualized characterization and clinical decision-making, and is suitable for resource-constrained edge computing environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122024831A_ABST
    Figure CN122024831A_ABST
Patent Text Reader

Abstract

The invention provides a sudden cardiac death early warning method based on wearable device signals and gene markers, which comprises the following steps: extracting user history medical records and whole genome SNP sites by using natural language processing and a bioinformatics algorithm, and screening to obtain key gene markers; wearable equipment is adopted for multidimensional physiological signals and clustering division of state intervals, genes and physiological parameters are combined, and a gene-physiological isomerism coupling map is established through cross-modal mutual information and causal analysis; performing edge weight updating and sparse screening by using a dynamic graph learning and minimum spanning tree method with L1 regularization, and extracting a core interaction path; based on symbolic regression and Bayesian optimization technologies, constructing a dynamic risk assessment equation, and performing individual risk grading and interaction chain attribution on real-time physiological data; finally, a multi-level visual report is generated, visual risk judgment and decision support are provided for clinicians, and the accuracy, timeliness and medical interpretation of sudden cardiac death risk assessment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of dynamic assessment of sudden cardiac death risk, and in particular to a method for early warning of sudden cardiac death based on wearable device signals and genetic markers. Background Technology

[0002] As sudden cardiac death risk prediction and early warning systems gradually integrate dynamic physiological signals from wearable devices with individual genomic information, various risk modeling technologies based on data-driven approaches and deep learning models have emerged at the intersection of medical artificial intelligence and bioinformatics. Current mainstream solutions generally employ large-scale neural network models to jointly model dynamically acquired electrocardiogram and physiological signals with known genetic markers. By learning complex nonlinear patterns to allocate weights, they achieve dynamic prediction of sudden cardiac death risk. In recent years, some technologies have further integrated attention mechanisms and adaptive weight update algorithms to improve the model's ability to represent multimodal data and its prediction accuracy. Furthermore, during the modeling process, some products also attempt to jointly embed genomic SNP sites and physiological characteristics into multi-layered network structures to achieve personalized risk quantification. These AI-powered risk prediction systems have been widely applied in medical early warning, intelligent assisted diagnosis, and individual health management. Their development trends are mainly reflected in the increased breadth of data fusion, the increased complexity of models, and the gradual realization of real-time risk assessment based on dynamic signals. Typical technological applications include risk discrimination based on deep learning (such as LSTM and GRU) of dynamic electrocardiogram sequences, genetic risk stratification by combining variant site heatmaps, and optimization of the prediction of sudden death event probability through multimodal feature fusion. However, the aforementioned mainstream technical solutions have significant drawbacks: First, while deep learning-based dynamic risk models have improved accuracy in retrospective case prediction, their decision-making process often falls into the "black box" category, lacking a transparent causal explanation mechanism. The weight allocation results are difficult to trace back to specific gene-physiological mechanism pathways, leaving medical personnel unable to ascertain the biological reasons for the model's assessment of increased or decreased risk. This leads to a significant decrease in clinical credibility and adoption willingness, impacting practical medical applications. Second, existing dynamic weight adjustment methods largely rely on global attention or naive fusion, ignoring the causal sparsity and pathway coupling characteristics of real biological networks. The dynamic impact of physiological state changes on gene risk expression is difficult to clearly represent in the model structure. More importantly, the lack of mechanisms to explain dominant interaction pathways hinders disease mechanism research and the development of personalized intervention strategies. Summary of the Invention

[0003] In order to solve the above-mentioned technical problems, the present invention provides a method for early warning of sudden cardiac death based on wearable device signals and genetic markers.

[0004] The technical solution of this invention is implemented as follows: a method for early warning of sudden cardiac death based on wearable device signals and genetic markers, comprising: S1: Based on the user's historical electronic medical records and whole genome sequencing data, extract SNP site information related to sudden cardiac death, screen to form a set of gene markers, and simultaneously divide the multidimensional physiological signals collected by wearable devices into resting / exercise / sleep state intervals. S2: Based on the set of gene markers and the divided physiological state intervals, calculate the cross-modal mutual information and Granger causality strength, and construct an initial heterogeneous graph structure containing gene-physiological node interactions, wherein gene marker nodes and physiological feature nodes are connected by dynamic weighted edges. S3: A dynamic graph learning mechanism based on L1 regularization constraints. It adopts the minimum spanning tree strategy to update edge weights and prune weak connections in the initial heterogeneous graph structure. It generates a sparse gene-physiology coupling map every 24 hours, preserving dynamic key interaction paths. S4: For the core subnetwork in the sparse gene-physiology coupling map, symbolic regression is used to search for nonlinear mathematical expressions and generate an interpretable risk evolution equation containing the interaction term between SNP load and HRV decline rate. The equation parameters are constrained for biological rationality through a Bayesian optimization framework. S5: Input the real-time collected physiological signal feature values ​​into the interpretable risk evolution equation, calculate the dynamic risk score increment, combine the baseline risk value to generate the current risk level, and identify the dominant gene-physiological interaction chain through the decision tree model; S6: Based on the current risk level and the contribution trend of the dominant gene-physiological interaction chain, a multi-level visualization report is generated, in which the risk level is encoded using a five-color gradient, and the interaction chain path is dynamically rendered through a topological map and labeled with biological pathway annotation information.

[0005] The sudden cardiac death early warning method based on wearable device signals and genetic biomarkers provided by this invention has the following beneficial effects: (1) This invention significantly improves the temporal sensitivity and modeling accuracy of sudden cardiac death risk assessment by constructing a gene-physiological coupling heterogeneous map and introducing a dynamic graph learning mechanism with sparsity constraints. Known pathogenic SNP sites and multidimensional physiological parameters such as ECG, heart rate variability, and respiratory rate continuously monitored by wearable devices are modeled in a node-based manner. The dynamic association weights between genes and physiological states are quantified by combining cross-modal mutual information and Granger causality strength to form an initial graph structure with biological significance. Furthermore, the local edge weights are updated every 24 hours based on the latest data window, and sparse pruning is implemented using L1 regularization and minimum spanning tree strategies to effectively remove noise interference and insignificant connections, retaining the most discriminative dynamic interaction paths. This mechanism not only overcomes the problem of poor adaptability of traditional models when facing changes in individual longitudinal data, but also realizes the continuous tracking of the activation state of key biological pathways, enabling risk prediction to have good temporal resolution and individualized representation capabilities. (2) This invention introduces a symbolic regression-driven interpretable rule mining and lightweight inference engine, fundamentally solving the clinical trust barrier caused by the "black box" decision-making of existing deep learning models. It achieves highly transparent risk interpretation and mechanism visualization output. For the core sub-networks in the graph (such as the coupling module of SCN5A-KCNH2 ion channel-related genes and QT interphase variation), symbolic regression technology is used to automatically search for mathematical expressions describing their nonlinear dynamic relationships, generating explicit and readable rules. These rules are directly mapped to known pathophysiological mechanisms and have a clear medical interpretation basis. The encapsulated lightweight inference engine can quickly activate the matching rule chain when new data is input, and output the risk score and its composition source in real time, avoiding the delay problem caused by large-scale matrix operations and ensuring system response efficiency. More importantly, the final generated visualization report can not only present the current risk level, but also trace the gene-physiological interaction path with the dominant role and its contribution trend, providing doctors with intuitive and credible decision support basis, significantly enhancing the adoptability and operability of the model in real medical scenarios. (3) This invention realizes the transformation from static risk stratification to a dynamic, mechanism-driven prediction paradigm, and constructs a closed-loop risk monitoring system that combines computational efficiency, clinical interpretability, and long-term adaptability. This invention abandons the use of fuzzy cognitive graphs in traditional multilayer perceptrons or recurrent neural networks. Without sacrificing modeling capabilities, it effectively avoids the risks of overfitting and declining generalization performance through constrained graph structure evolution and symbolic knowledge extraction. At the same time, the entire process does not require extensive hyperparameter tuning or dedicated hardware support, making it suitable for resource-constrained edge computing environments (such as mobile health monitoring platforms) and possessing good deployment scalability. Especially in long-term home monitoring scenarios for high-risk populations, it can continuously integrate genetic background and dynamic physiological feedback to achieve early identification and path tracing of potential malignant events, which not only improves the timeliness and accuracy of early warning but also provides mechanism-level support for the formulation of personalized intervention strategies. Attached Figure Description

[0006] Figure 1 This is a flowchart of the sudden cardiac death early warning method based on wearable device signals and genetic markers of the present invention; Figure 2 This is a sub-flowchart of the sudden cardiac death early warning method based on wearable device signals and genetic markers of the present invention; Figure 3 This is another sub-flowchart of the sudden cardiac death early warning method based on wearable device signals and genetic markers of the present invention. Detailed Implementation

[0007] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0008] The following disclosure provides many different embodiments or examples for implementing different structures of the invention. To simplify the disclosure, specific examples of components and arrangements are described below. Of course, these are merely examples and are not intended to limit the invention. Furthermore, reference numerals and / or letters may be repeated in different examples; such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.

[0009] like Figure 1 As shown, this invention provides a method for early warning of sudden cardiac death based on wearable device signals and genetic biomarkers, specifically including: S1: Based on the user's historical electronic medical records and whole genome sequencing data, extract SNP site information related to sudden cardiac death, screen to form a set of gene markers, and simultaneously divide the multidimensional physiological signals collected by wearable devices into resting / exercise / sleep state intervals. S2: Based on the set of gene markers and the divided physiological state intervals, calculate the cross-modal mutual information and Granger causality strength, and construct an initial heterogeneous graph structure containing gene-physiological node interactions, wherein gene marker nodes and physiological feature nodes are connected by dynamic weighted edges. S3: A dynamic graph learning mechanism based on L1 regularization constraints. It adopts the minimum spanning tree strategy to update edge weights and prune weak connections in the initial heterogeneous graph structure. It generates a sparse gene-physiology coupling map every 24 hours, preserving dynamic key interaction paths. S4: For the core subnetwork in the sparse gene-physiology coupling map, symbolic regression is used to search for nonlinear mathematical expressions and generate an interpretable risk evolution equation containing the interaction term between SNP load and HRV decline rate. The equation parameters are constrained for biological rationality through a Bayesian optimization framework. S5: Input the real-time collected physiological signal feature values ​​into the interpretable risk evolution equation, calculate the dynamic risk score increment, combine the baseline risk value to generate the current risk level, and identify the dominant gene-physiological interaction chain through the decision tree model; S6: Based on the current risk level and the contribution trend of the dominant gene-physiological interaction chain, a multi-level visualization report is generated, in which the risk level is encoded using a five-color gradient, and the interaction chain path is dynamically rendered through a topological map and labeled with biological pathway annotation information.

[0010] Step S1: Based on the user's historical electronic medical records and whole-genome sequencing data, extract SNP locus information related to sudden cardiac death, screen to form a set of gene biomarkers, and simultaneously divide the multidimensional physiological signals collected by wearable devices into resting / exercise / sleep state intervals. Specifically, this includes: S1.1: Based on the user's historical electronic medical record data, execute natural language processing and structured information extraction algorithms to extract historical diagnostic records, medication information and clinical events related to sudden cardiac death in order to generate structured medical history feature vectors; Based on users' historical electronic medical record data, a natural language processing algorithm combining named entity recognition and relation extraction (parameters: medical domain customized dictionary, context window length 15) is used to automatically locate and type-label medical concepts related to sudden cardiac death in the original text records. Furthermore, by using a syntactic parsing and event extraction algorithm based on contextual semantic dependency graph (parameter: dependency parsing weight threshold 0.6), the logical association identification of diagnostic events, medication records and heart-related clinical events is realized, and event nodes and their attribute sets are obtained; Furthermore, the extracted diagnostic information is encoded using a regularized feature encoding algorithm (parameter: feature weight normalization range 0-1) to construct a diagnostic category vector, and combined with the effective ingredients, dosage and dosing frequency in the medication information to generate a medication feature vector; Furthermore, the occurrence time of clinical events is aligned with the time of medical record records by time series alignment and missing value imputation algorithm (parameter: linear interpolation score threshold 0.4) to form a time-series medical history event matrix that can be input into the subsequent risk assessment model; Furthermore, the diagnostic category vector, medication feature vector, and clinical event matrix are fused by feature splicing and dimensionality reduction (parameter: PCA dimensionality reduction is 50) to generate a structured medical history feature vector with a uniform scale; Through the above natural language processing and structured information extraction algorithms, unstructured medical record texts are transformed into medical history feature data that can be used for the fusion analysis of genes and physiological signals, thereby achieving the standardization and computability of medical history information in risk modeling. For example, when performing this step on the historical electronic medical record of a high-risk cardiovascular patient, the input data includes 78 diagnostic records, 45 medication records, and 12 emergency cardiac event records from the past 3 years. Using a medical-specific dictionary, a named entity recognition model identifies 15 entities in the diagnostic records, such as "arrhythmia" and "myocardial ischemia." A relation extraction algorithm establishes a treatment relationship between "arrhythmia" and "β-blockers" in the records. Through syntactic parsing and event extraction, 8 warning symptoms preceding cardiac arrest are identified and matched with the corresponding ECG data timestamps. The diagnostic category vector is encoded as a sparse vector of length 100, and the medication feature encoding includes three parts: active ingredient code (length 20), dosage (unit mg), and frequency (unit times / day). After time series alignment and interpolation, a data matrix containing 810 time steps is formed. Finally, after PCA was reduced to 50 dimensions, a medical history feature vector of uniform scale was obtained. When this vector was input into the subsequent gene-physiology coupling map construction step, it significantly improved the matching accuracy between gene nodes and medical history event nodes and the medical relevance of interaction paths. S1.2: Based on whole-genome sequencing data, bioinformatics tools were used to perform quality control and annotation analysis on the original VCF files to identify known SNP sites that are significantly associated with sudden cardiac death, in order to generate a candidate SNP set; S1.3: Based on the candidate SNP set, perform gene pathway enrichment analysis and multivariate logistic regression model screening to identify gene biomarkers with significant functional effects in cardiac electrophysiological activities, so as to construct an individualized gene biomarker set; Based on the candidate SNP set, a gene pathway enrichment analysis method (parameters: gene set ID list, background gene set, significance threshold P<0.05) was used to match and retrieve the genes containing the candidate SNPs with cardiac electrophysiological pathways, and to calculate the enrichment significance of each pathway. Furthermore, the hypergeometric distribution test method (parameters: number of sample genes, total number of background genes, number of pathway genes, number of pathway intersection genes) is used to calculate the pathway-level P-value and the adjusted false discovery rate (FDR), and to obtain the set of significantly enriched pathways. By using a pathway set mapping mechanism (parameters: KEGG / Reactome database index, pathway node weight rules), the topological location of candidate SNP genes in a multipath network is determined, and a gene-pathway association matrix is ​​generated. Furthermore, a multivariate logistic regression model (parameters: dependent variable is the sudden cardiac death label, independent variable is the frequency of gene pathway features, regularization coefficient λ=0.01) is used to estimate the regression coefficients of each candidate gene in the risk prediction of cardiac electrophysiological activity and generate gene effect size scores. A stepwise regression feature screening method (parameters: entry threshold P<0.05, removal threshold P>0.10) was applied to retain or remove genes with significant effect size scores one by one, and to generate an optimized list of risk genes. Furthermore, by normalizing the effect size weights (parameter: minimum-maximum normalization), the weight values ​​of each risk gene are standardized within the range of [0,1], and an individualized set of gene biomarkers is generated. Through the above algorithm chain, the candidate SNP set results of the previous step are transformed into individualized gene marker data with significant impact on cardiac electrophysiological function, achieving a screening technology effect that combines statistical significance and biological interpretability. For example, in a subject with a history of high risk for sudden cardiac death, the candidate SNP set contained 480 loci, mapped to 312 genes, and the selected background gene set was 20,000 genes. During the gene pathway enrichment analysis, this gene list was entered into the Reactome pathway database. The enrichment p-values ​​for the "Cardiac conduction" pathway and the "Ionchannel transport" pathway were 0.0008 and 0.0015, respectively. After adjustment using the Benjamini-Hochberg method, the FDRs were 0.012 and 0.018, respectively, meeting the significance threshold. In the hypergeometric test, for the "Cardiac conduction" pathway, the parameter is set to the number of genes in the sample. Total number of background genes Number of pathway genes Number of intersecting genes The P-value was calculated. ; In multivariate logistic regression, the occurrence of sudden cardiac death is used as the dependent variable, the pathway gene frequency vector is used as the independent variable, and the regularization coefficient is used. The regression coefficient of the SCN5A gene was obtained as follows: The regression coefficient of the KCNH2 gene is All passed the significance test (P<0.05). After stepwise regression screening, 12 high-weight genes, including SCN5A, KCNH2, and RYR2, were retained. After min-max normalization, the weight of SCN5A was [value missing]. The weight of KCNH2 is RYR2 weight is The final output of the individualized set of genetic biomarkers can be used for subsequent cross-modal mutual information calculation and risk assessment map construction, effectively improving the model's sensitivity and reliability to the subject's risk of sudden cardiac death; S1.4: Utilize multidimensional physiological signals collected by wearable devices, including electrocardiogram, heart rate variability, accelerometer and respiratory rate signals, to perform time-frequency domain feature extraction and multimodal fusion processing to generate high-dimensional physiological feature vectors; Raw physiological data collected by sensors in multimodal wearable devices, including electrocardiogram signals, heart rate variability time series, triaxial accelerometer output, and respiratory rate change curves, are used as input conditions for this sub-step. Short-time Fourier transform method is used (parameters: window function type is Hanning window, window length is...). Sampling points, frame shift (Sampling points) to achieve time-frequency domain decomposition of electrocardiogram (ECG) signals, extracting the spectral energy distribution and main frequency peak position of each frame of ECG signal; Furthermore, through continuous wavelet transform (using Morlet as the mother wavelet, with a scale factor range of...), This enables multi-scale analysis of heart rate variability sequences and yields low-frequency and high-frequency power and their power ratio parameters. Furthermore, a time-domain feature extraction algorithm is employed (parameter: window length). The accelerometer output is processed (within seconds and without overlap) to calculate the root mean square acceleration, activity intensity integral, and frequency band energy index under resting and motion states, and to generate an acceleration feature set. Furthermore, using autocorrelation analysis (lag range) (seconds) to periodically detect the respiratory rate change curve, thereby generating two types of parameters: respiratory rate fluctuation amplitude and respiratory cycle stability; The feature normalization method (Z-score normalization) is used to achieve a unified expression of feature values ​​of different modes within the same dimension, and the feature matrix is ​​obtained. By using a multimodal feature fusion algorithm (weighted splicing strategy, with weights calculated based on the statistical correlation between features and the risk of sudden death), the feature vectors of the above modalities are synthesized into a unified high-dimensional physiological feature vector, thereby achieving comprehensive representation capabilities of data from different signal domains. This feature extraction and fusion algorithm transforms the multi-source raw sensor signals from the previous step into high-resolution, highly correlated dynamic physiological feature vectors, thereby achieving the expected technical effect of supporting subsequent physiological state clustering and gene-physiological coupling modeling. For example, in a wearable system based on a combination of a wristband and a chest patch, the ECG sampling frequency is configured as follows: Hz, accelerometer sampling frequency configured as Hz, respiratory rate sampling frequency configured as Hz. The Hanning window length is set to [value missing] for the short-time Fourier transform of the ECG signal. Point, frame shift Points were extracted, and the main frequency value was approximately [value missing]. Hz ECG spectral peak and total power range Low-frequency component power at Hz. Heart rate variability sequence at Morlet wavelet scale. Low-frequency power is obtained at the location With high frequency power The power ratio is The accelerometer characteristic is that the root mean square acceleration at rest is... g, under motion conditions is g. Respiratory rate curve autocorrelation analysis lag The correlation coefficient at the second period is The fluctuation range is Hz. After Z-score normalization, all features are concatenated and fused according to relevance weights of 1:1.5:1.2:0.8, generating a dimension of . The high-dimensional physiological feature vectors exhibit significantly improved discriminative power in subsequent state clustering. S1.5: Based on physiological feature vectors, a clustering algorithm is used to divide continuous time series data into states, identify and label the physiological feature intervals of users in resting, moving and sleeping states, so as to generate state label sequences.

[0011] Step S2: Based on the set of gene markers and the divided physiological state intervals, calculate cross-modal mutual information and Granger causality strength, and construct an initial heterogeneous graph structure containing gene-physiological node interactions, wherein gene marker nodes and physiological feature nodes are connected by dynamic weighted edges. Specifically, this includes: S2.1: Based on SNP site information in the user's genetic marker set, perform gene-phenotype association analysis to identify genetic variation nodes that are significantly associated with sudden cardiac death; Based on SNP locus information in a personalized set of genetic markers, a gene-phenotype association analysis method was used (parameters: significance level α=0.05, effect size threshold set according to the statistical distribution of the GWAS database) to achieve the initial screening function of locating genetic variations related to sudden cardiac death. Furthermore, the association coefficients between each SNP locus and the phenotype were calculated using a single-label linear regression model (parameters: phenotypic variables included the frequency of historical sudden cardiac death events and abnormal electrocardiogram indicators), and the significance of the regression coefficients was obtained using the following formula:

[0012] in, This is a statistic used to test the effect size of SNP loci. Statistical significance SNP locus effect value, This is the standard error of the effect value; Furthermore, the significance results were adjusted using a multiple comparison correction algorithm (parameter: Benjamini-Hochberg FDR threshold = 0.1) to reduce the false positive rate and obtain a set of SNP candidates that meet statistical confidence. Furthermore, highly correlated SNP sites were removed using the linkage disequilibrium decomposition method (parameter: LD threshold r² < 0.2), while retaining the set of genetic variation nodes with independent information contributions; Furthermore, the selected SNP sites are mapped to the corresponding gene regions using a functional annotation algorithm (parameters: refer to the Ensembl gene annotation library and the ClinVar clinical variant library), and their biological function categories in cardiac electrophysiological activities are labeled to form a set of gene nodes with functional interpretation. Through the above analysis and processing methods, the gene marker set from the previous step is transformed into structured genetic variation node data with statistical significance and biological functional classification, so as to achieve high-quality input for subsequent cross-modal mutual information analysis. For example, when performing this step on data from a high-risk cardiovascular subject, the input set of genetic markers contained 415 candidate SNP loci obtained from whole-genome sequencing. A gene-phenotype single-marker linear regression model was used, with ECG QT interval prolongation as the phenotype, to calculate the effect value β and standard error for each locus, and then... The formula yielded a t-statistic, identifying 72 loci with a significance level of p < 0.05. The FDR was adjusted to 0.1 using the Benjamini-Hochberg method, and the remaining 52 loci were subjected to LD decomposition to remove highly correlated variants, retaining 37 loci. Ensembl annotation mapped these loci to 15 functionally relevant genes, covering ion channel regulation, cardiac structural proteins, and signal transduction pathways. These structured gene nodes can then be used as input for subsequent cross-modal mutual information calculations, significantly improving the interpretability of biological mechanisms and the stability of statistical associations. S2.2: Perform time-frequency domain feature extraction processing on the multidimensional physiological signals collected by wearable devices to obtain key physiological characteristic parameters such as the low-frequency / high-frequency ratio of heart rate variability, the coefficient of variation of QT interval, and the amplitude of respiratory rate fluctuation. Based on the synchronous signal sequences acquired from the electrocardiogram (ECG), heart rate variability (HRV) sensing module, accelerometer, and respiratory rate sensor, a short-time Fourier transform method is used (parameters: window function type = Hanning window, window length = ...). Sampling points, frame shift = (Sampling points) to achieve the decomposition of each physiological signal in the time-frequency plane and the estimation of power spectral density, and to obtain spectral energy distribution data in different frequency bands; Using wavelet packet decomposition (parameter: number of decomposition levels = ...) Multi-scale feature separation is achieved using wavelet basis = Daubechies-4. Furthermore, statistics such as energy, standard deviation and kurtosis are calculated on the coefficient matrix of each sub-band, and a set of eigenvalues ​​for characterizing non-stationary properties is obtained. Using the HRV frequency domain analysis algorithm, and based on the Welch power spectrum estimation results, the low-frequency (LF) power spectrum is calculated. - Hz) and high frequency (HF) - The power component (Hz) is obtained, and the low-frequency / high-frequency ratio is obtained using the following formula:

[0013] in, For low-frequency normalized power, Normalized power for high-frequency bands; Using an automatic QT interval detection algorithm based on the improved Pan-Tompkins method to locate Q and T waves in electrocardiogram signals, the standard deviation and mean of the QT interval are calculated in the time domain sequence, and the coefficient of variation of the QT interval is obtained by the following formula:

[0014] in, The standard deviation of the QT interval sequence. Its mean; Fast Fourier Transform and bandpass filtering were applied to the respiratory rate signal sequence (parameter: bandwidth = ...). - The combined processing of Hz values ​​is used to calculate the fluctuation amplitude index, which is the difference between the maximum and minimum values, to reflect the intensity of dynamic changes in the breathing pattern. By using the feature normalization method (Z-score standardization), the above frequency domain ratios, coefficients of variation, and fluctuation amplitudes are transformed into dimensionless feature parameters, thereby achieving feature comparability across individuals and time periods. Through this feature extraction and standardization process, the original multidimensional physiological signals from the previous step are transformed into a dataset containing key feature parameters such as the low-frequency / high-frequency ratio of heart rate variability, the coefficient of variation of QT interval, and the amplitude of respiratory rate fluctuations, thereby providing a unified physiological feature input for cross-modal mutual information calculation and Granger causality analysis. For example, in a specific high-risk case surveillance, the ECG sampling frequency was set to... Hz, HRV window length set to Point, frame shift Point; the LF power obtained by short-time Fourier transform is ms², HF power is ms², the calculated low-frequency / high-frequency ratio is The result was 1.5; the mean QT interval was s, standard deviation is s, the coefficient of variation of the QT interval calculated by the formula is ≈ respiratory rate signal in - The maximum value of the Hz frequency band is Hz, minimum value Hz, fluctuation range is Hz. After normalization, the patient's feature parameters were integrated into the initial feature vector, which significantly improved the statistical stability of mutual information and the discriminative ability of causal analysis in the subsequent gene-physiology coupling map construction stage; S2.3: Based on the gene marker nodes and the extracted physiological feature parameters, perform cross-modal mutual information calculation to quantify the strength of the nonlinear statistical dependency between genes and physiological features; Based on the set of gene marker nodes and physiological feature parameters extracted from the preceding sub-steps, a method combining kernel density estimation and mutual information calculation (parameters: the kernel function type is selected as Gaussian kernel, and the bandwidth is adaptively set according to the Silverman rule) is used to realize the joint probability distribution modeling between cross-modal features; Furthermore, by performing kernel density estimation separately in the gene feature dimension and the physiological feature dimension, the respective marginal probability distribution functions are obtained and used as marginal distribution inputs in mutual information calculation; Furthermore, a mutual information calculation formula is used to quantify the strength of the nonlinear statistical dependency between the two types of nodes; Furthermore, by using a numerical integration algorithm (parameter: two-dimensional trapezoidal integral), the mutual information value is approximated at discrete sampling points to obtain the elements of the mutual information matrix, each element corresponding to the statistical dependence strength of a pair of gene nodes and physiological nodes; Furthermore, to suppress artificially high mutual information values ​​introduced by sampling noise, a permutation test method (parameters: permutation count 1000, significance level 0.05) is performed to remove insignificant mutual information association pairs and generate a mutual information matrix that has been filtered for significance. By using mutual information calculation and significance screening, the gene-physiological feature set from the previous step is transformed into a quantified nonlinear dependency strength index, thereby enabling the interpretable construction of the edge weights of the initial heterogeneous graph structure. For example, in a sudden cardiac death risk assessment scenario, the input data includes 5 genetic biomarker nodes (such as SCN5A, KCNH2, etc.) and 4 physiological characteristic parameters (heart rate variability low-frequency / high-frequency ratio of 0.45, QT interval coefficient of variation of 0.12, respiratory rate fluctuation amplitude of 0.08, and mean acceleration during exercise of 1.25 m / s²). When using Gaussian kernel density estimation, the bandwidth is 0.15 in the genetic feature dimension and 0.10 in the physiological feature dimension. The joint probability distribution is estimated on the sampling grid. When substituting these distributions into the mutual information formula, a two-dimensional trapezoidal integral is used to accumulate the sampling points within the interval [-3σ, 3σ] to obtain the mutual information value between the SCN5A node and the QT interval coefficient of variation node in the mutual information matrix. The mutual information value between the KCNH2 node and the heart rate variability ratio is After 1000 permutation tests, the p-values ​​were respectively... and All values ​​below the significance level are retained in the output matrix. This mutual information matrix serves as the edge weight input to the initial heterogeneous graph, and is subsequently combined with Granger causality strength for further modeling, achieving high-precision and interpretable gene-physiology interaction map construction; S2.4: Based on the time series data of the physiological characteristic parameters, execute the Granger causality analysis algorithm to identify feature pairs with causal relationships under different physiological states (such as rest, exercise, and sleep), and construct a causal directionality matrix; S2.5: Based on the cross-modal mutual information results and the Granger causality directionality matrix, perform heterogeneous graph structure modeling to generate an initial gene-physiological coupling graph composed of gene nodes and physiological feature nodes, wherein the edge weights are jointly weighted by mutual information and causal strength. Based on the cross-modal mutual information calculation results and Granger causal directionality matrix, a heterogeneous graph generation algorithm (parameters: node type = gene marker / physiological feature, initial edge set = empty set) is used to initialize the node relationship structure and establish the basic skeleton structure of the graph. Furthermore, a joint weighting method (parameters: weight coefficient combination strategy = linear weighting, normalization method = Z-score) is used to fuse mutual information strength and causal relationship strength, and the original edge weight matrix is ​​obtained. Furthermore, the weight data is stabilized and a weight normalization matrix is ​​generated by using an edge weight regularization method (parameters: range constraint = [0,1], outlier removal threshold = 3 times the standard deviation). Furthermore, the adjacency relationship of the graph is defined using the heterogeneous graph adjacency matrix construction module (parameter: node mapping table = gene node index / physiological feature node index), and a complete adjacency matrix structure is generated. Furthermore, the adjacency matrix and node attribute vectors are merged through a graph data encapsulation method to form an initial gene-physiology coupled graph data object, in order to support subsequent graph learning and dynamic optimization; By using heterogeneous graph structure modeling, cross-modal mutual information and causal strength results are transformed into an initial gene-physiology coupling map with clear topological connections and quantifiable edge weights, thereby realizing the fusion and correlation expression of high-dimensional multimodal data in a unified structural space. For example, in the multimodal data processing of a high-risk cardiovascular patient, the mutual information value between the SCN5A gene SNP load and the QT interval coefficient of variation in the cross-modal mutual information matrix was 0.68, and the Granger causality directionality matrix showed a causal strength value of 0.52 at rest. Using a joint weighting method, with a mutual information weight coefficient of 0.6 and a causal strength weight coefficient of 0.4, the edge weights were calculated using the following formula:

[0015] in, The edge weights are given, and the result is... After standardization, the edge weight stabilized at 0.61, forming a connection between the SCN5A node and the QT interphase node in the heterogeneous graph adjacency matrix. The final output initial gene-physiology coupling map contains 184 nodes and 1124 weighted edges, which can significantly improve the structured representation capability of different modalities and provide stable input for subsequent dynamic graph sparsity and symbolic regression modeling.

[0016] like Figure 2 As shown, step S3: Based on the L1 regularization constraint-based dynamic graph learning mechanism, the minimum spanning tree strategy is used to update edge weights and prune weak connections in the initial heterogeneous graph structure, generating a sparse gene-physiology coupling map every 24 hours, while preserving dynamic key interaction paths. Specifically, it includes: S3.1: Based on the initial heterogeneous graph structure and physiological signal data within the most recent 24-hour window, calculate the dynamic mutual information increment between gene marker nodes and physiological feature nodes to reflect the influence of the current physiological state on the gene-physiological interaction strength and obtain the dynamic mutual information matrix. Based on the initial gene-physiological heterogeneity map structure constructed by step S2 and the multidimensional physiological signal data collected by wearable devices within the most recent 24-hour window, a sliding time window slicing method (window size: 24 hours, step size: 1 hour) is used to perform time synchronization and interpolation preprocessing on electrocardiogram, heart rate variability, respiratory rate and accelerometer signals to ensure the time alignment accuracy of cross-node feature calculation. The kernel density estimation method (kernel function type: Gaussian kernel, bandwidth parameter according to Silverman's rule) is used to estimate the distribution of SNP load values ​​corresponding to gene marker nodes and the distribution of eigenvalues ​​corresponding to physiological characteristic nodes, respectively. Mutual information calculation is performed in the joint probability space to measure the change in statistical correlation between node pairs in the most recent 24 hours. Furthermore, the change in mutual information between the current time window and the initial plotting time window is calculated using the mutual information increment formula:

[0017] in The change in mutual information is calculated within the current window. For gene marker nodes, Physiological feature nodes, Index for the current time window, The mutual information value calculated in the current window; A bootstrap method with bias correction (repetitions: 500) was used to test the statistical significance of the mutual information increments. Increments with a confidence interval lower limit less than zero were set to zero to filter out weight changes caused by random fluctuations. Simultaneously, normalization was applied to map all mutual information increments to the [0,1] interval, forming a dynamic mutual information matrix that provides quantitative input for the weighted fusion in S3.2. Through the above calculation and verification process, the changes in gene-physiological interaction intensity caused by the physiological state in the last 24 hours are transformed into a structured dynamic mutual information matrix, so as to achieve accurate updating of association strength based on time window. For example, in a system implementation targeting high-risk cardiovascular populations, ECG signals and SCN5A gene SNP load values ​​within the most recent 24-hour window are selected as inputs. The ECG signal acquisition frequency is 500Hz, and after resampling and detrending processing, the low-frequency / high-frequency ratio is extracted to be 1.8. A Gaussian kernel density estimation bandwidth of 0.35 is used to calculate the joint probability distribution of the two nodes, obtaining the mutual information value for the current window. Initial window mutual information value Based on this, the mutual information increment is calculated as follows: The confidence interval for the significance test using the bootstrap method was [0.02, 0.11], confirming the incremental change as significantly effective. After normalization, this mutual information increment was mapped to 0.64 and recorded at the (SCN5A, LF / HF) position in the dynamic mutual information matrix. This embodiment demonstrates that this method can significantly improve the sensitivity of capturing changes in gene-physiological coupling strength under short-cycle monitoring, contributing to the accurate updating of edge weights in subsequent dynamic graph structures. S3.2: Based on the dynamic mutual information matrix and the initial value of Granger causality strength, a weighted fusion operation is performed, and the edge weights between nodes are dynamically updated using the sliding window normalization method to construct the updated gene-physiological heterogeneity graph adjacency matrix. S3.3: Based on the updated gene-physiological heterogeneous graph adjacency matrix, a graph learning algorithm with L1 regularization constraint is applied to perform sparsification on all edges in the graph to suppress the weight distribution of non-significant interaction paths and obtain a sparsified graph structure candidate set. S3.4: Based on the sparsified graph structure candidate set, execute the minimum spanning tree algorithm to extract the main subgraph containing key gene-physiological interaction paths, so as to remove redundant connections and retain the most discriminative dynamic interaction paths. S3.5: Based on the main subgraph, perform topology optimization and weight recalibration operations, and use a weighting strategy based on node degree and path length to locally readjust the edge weights to generate the final sparse gene-physiology coupling map. Based on the main subgraph extracted by the minimum spanning tree algorithm, a topology optimization method (parameters: main subgraph node set, edge weight matrix) is adopted to achieve local optimization of the node connection method and elimination of redundancy in the path structure. Furthermore, through node degree statistics and path length calculation methods (parameter: node degree threshold) Path length threshold This allows for the quantification of the structural importance of each edge in the main subgraph, and the generation of node degree distribution vectors and path length distribution vectors. Furthermore, a weight recalibration strategy based on joint weighting of node degree and path length is adopted (parameter: This is the node degree weight coefficient. (This refers to the path length weight coefficient), which enables local readjustment of edge weight values ​​and generates an optimized weight matrix; The edge weight calibration value is calculated using the following weighting formula:

[0018] in, This represents the node degree value. This is the path length value. Recalibrate the weights of the edges; Furthermore, the weight matrix of the entire sparse gene-physiology coupling map is updated through the adjacency matrix update algorithm (parameters: optimized weight matrix, main subgraph node index), and the final optimized sparse heterogeneous graph structure data is obtained. By combining topology optimization and weight recalibration, the results of the previous main subgraph are transformed into a sparse gene-physiology coupling map that balances structural simplicity and significant interaction retention, thereby achieving stable identification of dynamic key interaction paths and input optimization for subsequent symbolic regression modeling. For example, in a backbone subgraph containing SCN5A and KCNH2 gene nodes, as well as physiological characteristic nodes such as QT interphase variation and HRV decline rate, the node degree threshold is... Set to 3, path length threshold Set to 2, Set it to 0.6. Set to 0.4. For the connection between SCN5A and the QT interphase variant node, the node degree is 4 and the path length is 1. According to the above formula:

[0019] Calculated The weight value is written into the adjacency matrix, and the importance of this edge in the overall coupled graph is significantly improved after the update. In the optimized sparse coupled graph, the weight of the SCN5A-QT connection is identified as the dominant path in subsequent risk modeling inference, improving the model's responsiveness to short-term QT ​​changes. The application results show that the risk evolution equation is more sensitive to the prediction of individuals containing this variant gene, significantly improving the timely early warning capability under high-risk conditions.

[0020] like Figure 3 As shown, step S4 involves applying symbolic regression to search for nonlinear mathematical expressions in the core subnetwork of the sparse gene-physiology coupling map, generating an interpretable risk evolution equation that includes the interaction term between SNP load and HRV decline rate. The equation parameters are constrained for biological rationality using a Bayesian optimization framework. Specifically, this includes: S4.1: Based on the core subnetwork in the sparse gene-physiology coupling map, extract gene nodes closely related to sudden cardiac death and their significantly coupled physiological feature nodes as the set of input variables for symbolic regression modeling, so as to construct the basic variable space of the interpretable risk evolution model; S4.2: Perform feature engineering on the SNP loading, HRV decline rate and their interaction terms in the input variable set to generate a basic function library containing nonlinear combinations such as polynomials, exponentials, and logarithms, as a candidate expression space for symbolic regression, in order to explore potential nonlinear risk evolution relationships. Based on the SNP load, HRV decline rate and their interaction term data extracted in step S4.1, a feature standardization processing algorithm (parameters: mean = sample mean, variance = sample variance) is used to achieve linear scaling of input variables within a uniform numerical range, so as to eliminate the influence of dimensional differences on subsequent nonlinear combinations. Furthermore, through a polynomial feature expansion algorithm (parameter: order = 2 to 4), the power combination of SNP load and HRV decline rate is constructed, and terms such as square terms, cubic terms and cross-product terms are explicitly generated in the feature set to fully cover potential nonlinear coupling patterns. Furthermore, an algorithm for constructing an exponential transformation function (parameters: base = e, exponential coefficient range = 0.1 to 2.0) is adopted to realize the exponential mapping processing of the original and polynomial combination terms, and to generate feature components that characterize the nonlinear growth or decay trend of physiological variables. Furthermore, by using a logarithmic transformation processing algorithm (parameter: base = natural logarithm), logarithmic mapping is achieved for positive variables and their combinations to enhance the model's sensitivity to changes in variable proportions and to construct a more linear feature response within the low-frequency fluctuation range. Furthermore, a feature interaction generation algorithm (parameter: interaction dimension = 2) is adopted to perform pairwise multiplication or nested operations on polynomial, exponential and logarithmic combination terms to generate complex coupled features, which are then summarized in matrix form to form a basic function library. Through the aforementioned feature engineering process, the results of the previous step are transformed into a basic function library containing nonlinear combinations such as polynomials, exponentials, and logarithms, thereby enabling the construction of the candidate expression space for symbolic regression and comprehensive coverage of the evolution relationship of potential nonlinear risks. For example, in an embodiment targeting individuals at high risk of sudden cardiac death, the input SNP burden values ​​were standardized to have a mean of 0.00 and a variance of 1.00, and the HRV decline rate was standardized to have a mean of 0.00 and a variance of 1.00. The polynomial order was set to 3 during polynomial expansion, generating... (SNP load square) (HRV decline rate cubed) (Cross-product) and other combined terms. In the exponential transformation process, with base e and an exponent coefficient of 1.2, we obtain... , Isomorphic components. In the logarithmic transformation, the natural logarithm is used to process the positive combination terms, resulting in... and Features. The interactive generation section will and Multiplying these features generates a composite interaction feature matrix. By incorporating these features into the basic function library, the space of candidate expressions for symbolic regression is significantly enriched. The model can fit and explain the complex coupling relationship between SNP and HRV in the subsequent search process. The output basic function library can significantly improve the fitting accuracy and biological rationality of the generated risk evolution equation in the offline testing phase. S4.3: Perform symbolic regression based on the genetic programming algorithm, search for the optimal nonlinear expression structure from the basic function library to minimize the sum of squared residuals between the model output and the historical risk score, and obtain a preliminary prototype of the risk evolution equation; S4.4: Construct a Bayesian optimization framework. Based on the mechanism of SNP load and HRV decline rate in cardiac electrophysiological stability in the prior biological knowledge base, set prior distribution constraints for parameters, and perform posterior estimation optimization of the parameters in the prototype of the risk evolution equation to improve the biological rationality and generalization ability of the model. Based on the prototype risk evolution equation generated by symbolic regression and its preliminary parameter estimation results, a Bayesian optimization framework is invoked (parameters: iterations N=500, exploration coefficients). =2.5, Acquisition function type = Expected EI improvement), to achieve posterior distribution optimization of key coefficients of the equation; Furthermore, through a biological knowledge base retrieval mechanism (parameters: pathway association threshold ≥ 0.8, gene-phenotype confidence level ≥ III), we obtained biological prior rules related to the effect of SNP load on cardiac electrophysiological stability and mapped them to the prior distribution form of each parameter, where the prior distribution parameter values ​​are derived from multicenter clinical statistics. Furthermore, by constructing a prior distribution module (method: combining Gaussian and Gamma priors), the SNP loading coefficient in the risk evolution equation is obtained. and HRV decline rate coefficient The constraints on the range of values ​​and the shape of the distribution of these parameters ensure that they are biologically reasonable in subsequent optimization processes; Furthermore, using a Bayesian inference processing unit (algorithm: Markov chain Monte Carlo MCMC sampling, step size set to 0.05, number of chains 4), the posterior distribution sampling of the model parameters is performed to generate a sample set for evaluating parameter uncertainty; The posterior mean-variance evaluation method is used, and the expected value and variance of each coefficient are calculated using the following formula:

[0021] in, For the parameters to be estimated, The total number of samples, The parameter estimate is obtained from the i-th MCMC sampling. Using the variance formula:

[0022] Among them, variance is used to constrain the stability of model parameters to prevent overfitting; Through the above-mentioned Bayesian optimization and parameter constraint processing methods, the risk evolution equation prototype generated by symbolic regression is transformed into a parameterized mathematical model that conforms to the biological mechanism and has generalization ability, thereby improving the medical credibility and cross-population adaptability of the model. For example, in the implementation process for a high-risk cardiovascular population, the initial value of the SNP burden coefficient in the collected whole-genome sequencing data was set to 0.65, and the initial value of the HRV decline rate coefficient was... 1.20, Using the results of the biological knowledge base query, set... The prior distribution is a normal distribution with a mean of 0.6 and a variance of 0.05. The prior distribution is a Gamma distribution with shape parameter 2.0 and scale parameter 1.0. MCMC sampling is performed with a chain length of 2000, discarding the first 500 warm-up samples to obtain the posterior mean. =0.58, variance 0.047, = 1.15, variance 0.092. Substituting this parameter set into the prototype of the risk evolution equation, with the patient's most recent motion state data input, the risk increment value output by the model is significantly higher than that in the resting state, verifying that the model can sensitively capture changes in physiological state after parameter adjustment and maintain a stable trend in different states, thereby significantly improving clinical interpretability and actual monitoring effect; S4.5: Perform model validation and sparsification on the risk evolution equation optimized by Bayes, remove redundant terms and retain core interaction terms to generate the final interpretable dynamic risk modeling equation to support subsequent real-time reasoning and visualization interpretation. For the risk evolution equation optimized by Bayes, a model validation algorithm (parameters: historical risk score dataset, core variable set) is used to evaluate the prediction accuracy and stability of the equation under different physiological state intervals. Furthermore, by using residual analysis (parameters: model output value, actual risk score value), the statistical test of the model residual distribution is achieved, and validation data such as residual mean and standard deviation are obtained to evaluate the fit consistency of the equation within the range of each variable value. Furthermore, a variance inflation factor calculation method (parameters: each input variable in the risk equation) is adopted to achieve multicollinearity detection and generate a variance inflation factor matrix. The redundancy between variables is judged by the index value, thus providing a basis for subsequent sparsification processing. Furthermore, based on the sparsification algorithm with a threshold setting (parameters: variance inflation factor threshold, variable contribution sequence), redundant variables or redundant interaction terms in the risk evolution equation are eliminated one by one, while retaining the core interaction terms with significant contributions, ensuring that the equation structure is simplified without losing the main discriminative ability. Furthermore, by using parameter re-estimation and normalization methods (parameters: model coefficient matrix after elimination, variable standard deviation vector), the weights of the core interaction term parameters are recalibrated to eliminate the scale shift caused by the elimination of dependent variables and generate the final interpretable dynamic risk modeling equation structure. Through the above verification and sparsification process, the optimization equation of the previous step is transformed into a risk prediction formula with a simplified structure, stable parameters, and compliance with biological mechanism constraints, thereby enabling the model to perform efficiently in real-time reasoning and visualization interpretation. For example, in a risk assessment scenario for a high-risk cardiovascular patient, a risk evolution equation optimized by Bayesian methods is selected as the input, with core variables including SNP load. With HRV decline rate And its product term. In the model validation phase, the number of historical risk score samples was set at 1000, including 400 samples in the resting state, 350 samples in the movement state, and 250 samples in the sleep state, and the overall mean squared error was calculated.

[0023] The mean square error formula is used:

[0024] in, These are the model's predicted values. For the actual risk score, n is the sample size. The above formula calculates the mean squared error of the resting state as follows: The state of motion is Sleep state is In the calculation of the variance inflation factor, the SNP loading variable... The VIF value is HRV decline rate The VIF value is The interaction item VIF value is All are below the set threshold. However, the additionally detected QT interval variation variable VIF value was... These terms are identified as redundant and removed. After removal, parameter re-estimation and normalization are performed to obtain the final equation. for:

[0025] in = , = , = This equation can efficiently calculate the risk increment in real-time inference and can intuitively explain the dynamic impact of the interaction mechanism between SNP load and HRV decline rate on the risk of sudden cardiac death.

[0026] Step S5: Input the real-time collected physiological signal feature values ​​into the interpretable risk evolution equation, calculate the dynamic risk score increment, combine it with the baseline risk value to generate the current risk level, and identify the dominant gene-physiological interaction chain through a decision tree model. Specifically, this includes: S5.1: Based on multidimensional physiological signals collected in real time by wearable devices, extract the RR interval sequence of electrocardiogram and the time-frequency domain feature parameters of heart rate variability to obtain the HRV decline rate index, which serves as a key input variable for the interpretable risk evolution equation; S5.2: Standardize the SNP load values ​​extracted from individual whole genome sequencing data, use the gene risk coefficient constrained by the Bayesian optimization framework to calculate the individualized static gene risk baseline value as the initial bias term for dynamic risk scoring; S5.3: Input the HRV decline rate and SNP load value into the nonlinear risk evolution equation generated by symbolic regression, and perform dynamic multiplication and exponential operation based on the equation parameters to obtain the real-time risk increment value, reflecting the immediate impact of the current physiological state on the risk of sudden death. S5.4: The static genetic risk baseline value and the real-time risk increment value are weighted and fused to generate an individualized current dynamic risk score, wherein the weight coefficients are adaptively adjusted according to the physiological state interval (resting / exercise / sleep); Based on the fusion processing of static gene risk baseline values ​​and real-time risk increment values, the input objects include static gene risk baseline values ​​constrained by the Bayesian optimization framework and real-time risk increment values ​​calculated by the interpretable risk evolution equation, and both have been normalized for unit consistency and dimension matching. An adaptive weighted algorithm is used (parameter: weight update function). (physiological state label sequence), which enables dynamic selection of weighting coefficients based on the user's current physiological state interval, wherein... A specific weighted mapping table is generated based on three physiological states (resting, exercise, and sleep); Furthermore, the normalized weight coefficient calculation method (parameter: current state weight) is used. Gene weight Incremental weights This process normalizes each weight value to the [0,1] interval and ensures the balance between the baseline gene risk value and the real-time risk increment value during fusion, resulting in a state weighted vector. Furthermore, the current dynamic risk score is calculated using a dynamic weighted fusion formula. :

[0027] in, This is the static genetic risk baseline value. This represents the real-time incremental risk value. and These are the normalized weights corresponding to the baseline and the increment, respectively; Furthermore, by using a state weight adjustment function (parameters: state label sequence, time series context), the adjustment is achieved whenever a physiological state changes. and Perform smooth transition updates to avoid abrupt changes or discontinuities in score changes, and generate a smoothed weight vector; By using a weighted fusion algorithm and a state adaptive adjustment mechanism, the baseline risk and real-time incremental results from the previous step are transformed into individualized dynamic risk scores, enabling accurate quantification and stable tracking of risk levels under different physiological states. For example, during the monitoring period of a high-risk cardiovascular patient, the static genetic risk baseline value Set to 1.85, the standardized range now matches the incremental value units; real-time risk incremental value. The state-specific weighting is 0.42 at rest, 0.95 during exercise, and 0.35 during sleep. The state-specific weighting mapping table is defined as follows: [Table showing weights for resting state and exercise state]. =0.6, =0.4; Motion state =0.4, =0.6; Sleep state =0.7, =0.3. In the resting state, the normalized weights remain unchanged, and the current risk score is calculated using the weighted fusion formula described above:

[0028] Calculated ; In motion, Updated to 0.4. Updated to 0.6, the calculated risk score is obtained. Update while asleep =0.7, =0.3, the risk score is calculated. In this embodiment, by adjusting the weights according to the state, the fluctuations in risk scores under different physiological states remain continuous and stable, significantly improving the model's adaptability to state switching and the reliability of the scores; S5.5: Construct a decision tree model based on the sparse gene-physiology coupling map, input the current risk score and the activation intensity of each gene-physiology node, identify the dominant gene-physiology interaction chain through path attribution analysis, and output its contribution weight and direction of action in the current risk evolution. Based on the node set and edge weight data of the sparse gene-physiology coupling map, the CART decision tree construction algorithm (parameters: maximum tree depth = 10, minimum number of leaf node samples = 5) is used to achieve joint modeling of risk score input and node activation intensity input. Furthermore, by using a feature importance assessment method (parameter: Gini impurity reduction), the contribution of each gene-physiological node in decision splitting is calculated, and the node importance vector is obtained. Furthermore, by using a path attribution analysis algorithm (parameter: tree path-based split gain accumulation), the complete path resolution corresponding to the risk determination results of leaf nodes is achieved, and a data matrix containing the activation intensity and weight contribution of each node along the path is generated. Furthermore, by using a directionality determination method (parameter: mutual information addition and subtraction symbol labeling strategy), the direction of gene-physiological interactions in each path is analyzed, and interaction chain direction vectors are generated; Through the above attribution analysis and directional judgment, the risk score results of the previous step are transformed into data on the weight contribution and direction of action of the dominant gene-physiological interaction chain, so as to realize the structured output of the model decision logic. For example, when performing risk assessment on a patient with an SCN5A gene variant and significant QT interval prolongation, the current dynamic risk score is entered as... The SCN5A node value in the gene node activation intensity matrix is The value of node KCNH2 is QT interval node value HRV decline rate node value In the CART decision tree model, feature importance is calculated based on the decrease in Gini impurity. The importance of the SCN5A node is... The importance of the QT interval node is The importance of the KCNH2 node is The importance of the HRV decline rate node is The dominant interaction chain generated by path attribution analysis is "SCN5A → QT interval → HRV decrease rate", and the cumulative splitting gain along the path is calculated using the formula... Calculation, where This represents the improvement in the Gini index after splitting. For cumulative gain, This represents the number of path nodes; in this example... The result is , representing the average gain per unit node. Directional analysis reveals a positive interaction between SCN5A and the QT interval, and a negative interaction between the QT interval and the HRV decrease rate. In the final output interaction chain weight contribution matrix, the sum of the positive interaction weights is... The total weight of the negative effects is The validation results show that, in simulated clinical scenarios, through this interaction chain attribution output, doctors can quickly identify the main risk factors and take targeted intervention measures, significantly improving the clinical reliability of risk prediction.

[0029] Step S6: Based on the current risk level and the contribution trend of the dominant gene-physiological interaction chain, a multi-level visualization report is generated, wherein the risk level is encoded using a five-color gradient, and the interaction chain path is dynamically rendered through a topological map and labeled with biological pathway annotation information. Specifically, this includes: S6.1: Based on the dynamic risk scoring results and the preset risk level threshold, execute the five-color gradient mapping algorithm to encode the risk level into five-color gradient output of red, orange, yellow, blue and green to achieve intuitive visualization of the risk level; The current dynamic risk score and the contribution trend data of the dominant gene-physiological interaction chain generated by step S5 are input into the risk level mapping module as the initial conditions for five-color gradient coding. A five-segment hierarchical mapping method is adopted (parameter: preset risk threshold set). This allows for the determination of corresponding color level relationships based on risk score value ranges, and the pre-construction of a mapping table to achieve high-speed retrieval; Furthermore, through an interval normalization algorithm (parameter: minimum risk score) Maximum value This involves performing linear normalization on the original risk score to obtain the normalized risk value. ; Furthermore, through the threshold comparison operation unit (parameter: ), to perform sequential comparisons The relationship between the values ​​and the threshold values ​​is determined, and color index values ​​are generated. The color index corresponds one-to-one with the codes for red, orange, yellow, blue, and green; Furthermore, through the color coding generator (parameter: This allows the color index value to be converted into RGB or HEX encoding that can be directly called in the rendering engine, and outputs a color encoding table; The five-color gradient mapping algorithm transforms the quantitative results of the risk assessment in the previous step into color-level data, enabling intuitive visualization of the risk level. For example, in a real-time monitoring scenario for a high-risk patient with sudden cardiac death, a set of preset risk level thresholds is used. The current dynamic risk score R = 0.65. =0, =1, the normalization formula is as follows:

[0030] in, R is the normalized risk value, and R is the current risk score. and These are the minimum and maximum values ​​of the risk score, respectively. Substituting R=0.65, =0、 =1, therefore we can get =0.65. (Comparison) With each threshold, satisfying <0.65≤ Get the color index =3, corresponding to the color code blue (RGB code (0,0,255)). After this code is input into the rendering engine, a blue gradient risk level display is generated, enabling doctors to intuitively identify the risk level. In similar case verification, it can significantly improve the speed and accuracy of identifying high-risk states; S6.2: Perform a topology map layout algorithm on the contribution trend data of the dominant gene-physiological interaction chain to generate an interaction path map with a hierarchical structure to visualize the dynamic coupling relationship between gene markers and physiological characteristics. S6.3: Based on the biological pathway annotation database, a pathway semantic annotation algorithm is performed on the generated topological map to map each node and edge to a known biological function and pathway name, so as to enhance the medical interpretability of the path; S6.4: Utilize a dynamic rendering engine to perform visualization rendering processing on the labeled topology map to generate a visualization interface that supports interactive zooming and path tracing, allowing doctors to view the changes in the contribution of key interaction chains in real time. S6.5: Based on a multi-level visualization architecture, it integrates a risk level panel with five-color gradient coding and a dynamic topology map to generate a multi-level visualization report that includes time and pathway dimensions, in order to support doctors' comprehensive judgment and intervention decisions on the risk evolution process. Based on the integrated five-color gradient coding risk level panel and the topology map generated by the dynamic rendering engine, a multi-level visualization architecture management module (parameters: risk level data matrix, interaction chain layout map, time series marker) is adopted to realize the synchronous loading and hierarchical mapping of risk level and interaction chain rendering results. Furthermore, through a time-dimensional mapping algorithm (parameters: dynamic risk score time series, state interval division labels), the risk level panel and the interaction chain contribution trend are synchronized and rolled over in time, and a time-risk level joint index array is obtained; Furthermore, by using a pathway-dimensional correlation calculation method (parameters: topology map node pathway identifier, risk level change threshold), a bidirectional index mapping between the risk change curve corresponding to each biological pathway and the topology layout is achieved, and a pathway-risk level cross-reference table is generated. Furthermore, through a multi-dimensional information fusion rendering engine (parameters: time index array, path cross reference table, risk level color coding scheme), the spatial overlay rendering of risk level color gradient and dynamic layout of interaction chain is realized, and an interactive multi-level visualization report interface is generated. Through multi-level visualization management algorithms, the risk level panel and topology map of the previous step are integrated into a unified data structure that includes time and pathway dimensions, so as to achieve the expected technical effect of doctors locating and intervening in the risk evolution process in time series and biological pathways. For example, in a continuous monitoring scenario for a high-risk cardiovascular patient, the dynamic risk score time series of the most recent 72 hours (sampling interval of 5 minutes, totaling 864 points) can be input into a time dimension mapping algorithm, and the labels for resting, exercise, and sleep states can be matched to generate a dataset containing... A time-risk level joint index array for size. This array links the biological pathway identifiers of nodes in the sparse gene-physiology coupling map with risk level change thresholds. Correlation calculations were performed to generate a pathway cross-reference table containing 26 curves showing significant risk changes. A multi-dimensional information fusion rendering engine mapped the five-color gradient coding scheme to the risk level curves using a time index array as the horizontal axis and the pathway dimension as the vertical axis. Combined with the dynamic topological layout of the interactive chain, a multi-level visualization report was generated, supporting mouse hover to view node annotations and path tracing. In this report, physicians can quickly locate the time point on the timeline where the risk level changes from yellow to orange during movement, and confirm in the pathway view that the dominant chain of this change is the coupling path between SCN5A-KCNH2 and QT interval variants, significantly improving the timeliness of clinical interventions.

[0031] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.

[0032] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and rules of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for early warning of sudden cardiac death based on wearable device signals and genetic biomarkers, characterized in that, Includes the following steps: S1: Based on the user's historical electronic medical records and whole genome sequencing data, extract SNP site information related to sudden cardiac death, screen to form a set of gene markers, and simultaneously divide the physiological characteristic intervals of multidimensional physiological signals collected by wearable devices. S2: Based on the set of gene markers and the divided physiological state intervals, calculate cross-modal mutual information and Granger causality strength, and construct an initial heterogeneous graph structure containing gene-physiological node interactions; S3: A dynamic graph learning mechanism based on L1 regularization constraints is used to update edge weights and prune weak connections in the initial heterogeneous graph structure using a minimum spanning tree strategy to generate a sparse gene-physiology coupling map. S4: For the core subnetwork in the sparse gene-physiology coupling map, apply symbolic regression technology to search for nonlinear mathematical expressions and generate interpretable risk evolution equations; S5: Input the real-time collected physiological signal feature values ​​into the interpretable risk evolution equation, calculate the dynamic risk score increment, combine the baseline risk value to generate the current risk level, and identify the dominant gene-physiological interaction chain through the decision tree model; S6: Generate a multi-level visualization report based on the current risk level and the contribution trend of the gene-physiological interaction chain.

2. The method for early warning of sudden cardiac death based on wearable device signals and genetic markers according to claim 1, characterized in that, Step S1 specifically includes: Based on users' historical electronic medical record data, natural language processing and structured information extraction algorithms are executed to extract historical diagnostic records, medication information and clinical events related to sudden cardiac death, and generate structured medical history feature vectors. Based on whole-genome sequencing data, bioinformatics tools were used to perform quality control and annotation analysis on the original VCF files, identify known SNP sites that are significantly associated with sudden cardiac death, and generate a candidate SNP set. Based on the candidate SNP set, gene pathway enrichment analysis and multivariate logistic regression model screening were performed to identify gene biomarkers with significant functional effects on cardiac electrophysiological activities and construct an individualized gene biomarker set. Wearable devices are used to collect multidimensional physiological signals, and time-frequency domain feature extraction and multimodal fusion processing are performed on the multidimensional physiological signals to generate high-dimensional physiological feature vectors. Based on the high-dimensional physiological feature vector, a clustering algorithm is used to divide the continuous time series data into states, identify and mark the physiological feature intervals of the user in the resting, movement and sleep states, and generate a state label sequence.

3. The method for early warning of sudden cardiac death based on wearable device signals and genetic markers according to claim 2, characterized in that, The multidimensional physiological signals include electrocardiogram, heart rate variability, accelerometer, and respiratory rate signals.

4. The method for early warning of sudden cardiac death based on wearable device signals and genetic markers according to claim 1, characterized in that, Step S2 specifically includes: Based on SNP locus information in the user's genetic biomarker set, gene-phenotype association analysis was performed to identify genetic variation nodes that are significantly associated with sudden cardiac death. The multidimensional physiological signals collected by wearable devices are processed by time-frequency domain feature extraction to obtain key physiological feature parameters; Based on the gene marker nodes and extracted physiological feature parameters, cross-modal mutual information calculation is performed; Based on the time series data of the key physiological characteristic parameters, the Granger causality analysis algorithm is executed to identify feature pairs with causal relationships under different physiological states and to construct a causal directionality matrix. Based on the cross-modal mutual information results and the causal directionality matrix, heterogeneous graph structure modeling is performed to generate an initial gene-physiological coupling map composed of gene nodes and physiological feature nodes.

5. The method for early warning of sudden cardiac death based on wearable device signals and genetic markers according to claim 4, characterized in that, Step S2 further includes using cross-modal mutual information computation and Granger causality analysis to measure the statistical and temporal causal dependencies between gene and physiological feature nodes, respectively, using joint linear weighting and regularization to generate the adjacency matrix of the initial heterogeneous graph, and fusing node attribute and edge weight data. The graph structure supports multi-node high-dimensional interactions and edge weight saliency screening.

6. The method for early warning of sudden cardiac death based on wearable device signals and genetic markers according to claim 1, characterized in that, Step S3 specifically includes: Based on the initial heterogeneous graph structure and physiological signal data within the most recent 24-hour window, the dynamic mutual information increment between gene marker nodes and physiological feature nodes is calculated to obtain the dynamic mutual information matrix. Based on the dynamic mutual information matrix and the initial value of Granger causality strength, a weighted fusion operation is performed, and the edge weights between nodes are dynamically updated using the sliding window normalization method to construct the updated gene-physiological heterogeneity graph adjacency matrix. Based on the updated gene-physiological heterogeneity graph adjacency matrix, a graph learning algorithm with L1 regularization constraint is applied to perform sparsification on all edges in the graph to obtain a sparse graph structure candidate set. Based on the sparsified graph structure candidate set, the minimum spanning tree algorithm is executed to extract the main subgraph containing key gene-physiological interaction paths. Based on the main subgraph, topology optimization and weight recalibration operations are performed. A weighted strategy based on node degree and path length is used to locally readjust the edge weights to generate the final sparse gene-physiology coupling map.

7. The method for early warning of sudden cardiac death based on wearable device signals and genetic markers according to claim 1, characterized in that, Step S4 specifically includes: Based on the core subnetwork in the sparse gene-physiology coupling map, gene nodes closely related to sudden cardiac death and their significantly coupled physiological feature nodes are extracted as the set of input variables for symbolic regression modeling, and the basic variable space of interpretable risk evolution model is constructed. Feature engineering is performed on the SNP load, HRV decline rate and their interaction terms in the input variable set to generate a basic function library; Symbolic regression is performed based on a genetic programming algorithm. The optimal nonlinear expression structure is searched from the basic function library to minimize the sum of squared residuals between the model output and the historical risk score, thereby obtaining a preliminary prototype of the risk evolution equation. A Bayesian optimization framework was constructed. Based on the mechanism of SNP load and HRV decline rate in cardiac electrophysiological stability in the prior biological knowledge base, prior distribution constraints of parameters were set, and posterior estimation optimization of the parameters in the prototype of the risk evolution equation was performed. The risk evolution equation optimized by Bayes is subjected to model verification and sparsification to generate the final interpretable dynamic risk modeling equation.

8. The method for early warning of sudden cardiac death based on wearable device signals and genetic markers according to claim 1, characterized in that, Step S5 specifically includes: Based on the multidimensional physiological signals collected in real time by wearable devices, the time-frequency domain feature parameters of the electrocardiogram RR interval sequence and heart rate variability are extracted to obtain the HRV decline rate index. The SNP load values ​​extracted from individual whole genome sequencing data are standardized, and the individualized static genetic risk baseline value is calculated using gene risk coefficients constrained by a Bayesian optimization framework. The HRV decline rate and SNP load value are input into the nonlinear risk evolution equation generated by symbolic regression. Based on the equation parameters, dynamic multiplication and exponential operation are performed to obtain the real-time risk increment value. The static genetic risk baseline value and the real-time risk increment value are weighted and fused to generate an individualized current dynamic risk score; A decision tree model is constructed based on a sparse gene-physiology coupling map. The current risk score and the activation intensity of each gene-physiology node are input. The gene-physiology interaction chain with the dominant role is identified through path attribution analysis, and its contribution weight and direction of action in the current risk evolution are output.

9. The method for early warning of sudden cardiac death based on wearable device signals and genetic markers according to claim 8, characterized in that, Step S5 further includes using real-time acquired physiological signal features and gene risk baselines as input to an interpretable equation, weighting and fusing historical risk baselines and risk increments based on a physiological state adaptive weighting algorithm, modeling the current risk score and node activation intensity using a CART decision tree, and outputting the dominant gene-physiological interaction chain and its weights, direction of action, and quantitative contribution to the dynamic score through split gain and mutual information increase / decrease symbol path analysis.

10. The method for early warning of sudden cardiac death based on wearable device signals and genetic markers according to claim 1, characterized in that, In step S6, the current risk level is encoded using a five-color gradient, and the risk level is encoded as a five-color gradient output of red, orange, yellow, blue and green. The interaction chain path is dynamically rendered through a topology map and labeled with biological pathway annotation information.