Stroke risk prediction method and system using atrial fibrillation electrocardio signals

By performing time-frequency dual-channel deep coding and cross-modal fusion on atrial fibrillation ECG signals, the problem of the failure to effectively utilize f-wave characteristics and clinical variables in existing technologies has been solved, enabling accurate prediction and individualized assessment of stroke risk.

CN122158143APending Publication Date: 2026-06-05SICHUAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SICHUAN UNIV
Filing Date
2026-04-28
Publication Date
2026-06-05

Smart Images

  • Figure CN122158143A_ABST
    Figure CN122158143A_ABST
Patent Text Reader

Abstract

The application provides an atrial fibrillation electrocardiosignal stroke risk prediction method and system, relates to the field of electrocardiosignal processing and disease risk prediction, and collects a twelve-lead signal of an atrial fibrillation attack from a dynamic electrocardio monitoring device, extracts an f wave residual signal after eliminating a QRS complex, arranges the f wave residual signal into a two-dimensional matrix according to a lead space topology, extracts main frequency drift and spectral entropy attenuation features through a frequency domain path and extracts wave amplitude variability and organization index features through a time domain path, performs cross-modal fusion on a gating bilinear fusion layer and a clinical variable coding vector to generate a joint risk representation, and inputs a prediction network trained by using a composite loss function containing a time-dependent consistency index optimization term to output a stroke probability value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electrocardiogram signal processing and disease risk prediction technology, and in particular to a method and system for predicting stroke risk in atrial fibrillation electrocardiogram signals. Background Technology

[0002] Atrial fibrillation (AF) is the most common sustained arrhythmia in clinical practice, affecting over 40 million people worldwide, and this number continues to rise due to population aging. The core clinical harm of AF lies in its significantly increased risk of ischemic stroke; AF patients are approximately five times more likely to suffer a stroke than non-AF individuals, and AF-related strokes are often more severe, with higher rates of disability and mortality. Therefore, accurately assessing the stroke risk in AF patients and developing individualized anticoagulation therapy strategies accordingly remains a major challenge in the field of cardiovascular clinical practice.

[0003] The CHA2DS2-VASc scoring system is widely used in clinical practice to stratify stroke risk in patients with atrial fibrillation. This score incorporates baseline clinical characteristics such as congestive heart failure, hypertension, age, diabetes, history of stroke or transient ischemic attack, vascular disease, and gender, assigning a risk level to each patient using a simple integer score. While the CHA2DS2-VASc score is widely adopted in clinical practice due to its simplicity, it has inherent limitations. First, the score only provides a rough stratification based on static baseline clinical characteristics. For the intermediate-risk group with a score of 2 to 3, there are significant individual differences in the annual stroke incidence, making it difficult for clinicians to make precise anticoagulation treatment decisions. Second, the CHA2DS2-VASc score completely ignores the atrial electrophysiological information contained in the electrocardiogram signal during an atrial fibrillation episode, failing to capture the dynamic changes in left atrial matrix remodeling and prethrombotic state.

[0004] Chinese patent CN116570293B discloses a device, system, and storage medium for predicting the risk of late recurrence after atrial fibrillation ablation. This method acquires electrocardiogram (ECG) data from patients after atrial fibrillation ablation to determine whether the patient experienced ultra-early recurrence within 7 days post-procedure. It calculates the atrial fibrillation burden (AF burden), i.e., the percentage of total AF episode duration to total monitoring time, and then outputs the late recurrence risk result based on the relationship between the AF burden and a classification threshold determined by ROC curves. While this scheme incorporates ECG monitoring data for risk prediction, its technical approach is essentially a statistical threshold stratification method, which has the following shortcomings: First, the scheme only extracts the coarse-grained temporal proportion of atrial fibrillation burden, failing to delve into the atrial matrix information contained in the ECG waveform during an atrial fibrillation episode, such as the f-wave spectral characteristics and temporal morphological features. Second, the scheme's prediction target is late recurrence after ablation, not ischemic stroke, the most clinically valuable endpoint event in atrial fibrillation patients. Third, the scheme uses traditional statistical models such as Logistic regression and Cox regression for prediction, lacking the ability to automatically mine deep features of ECG signals, and does not involve cross-modal deep fusion of ECG features and clinical variables.

[0005] Furthermore, although some researchers have attempted to use deep learning methods to analyze 12-lead electrocardiograms under sinus rhythm to predict the risk of new-onset atrial fibrillation and indirectly correlate it with stroke events in recent years, these methods have limitations: they analyze the electrocardiogram signals of sinus rhythm rather than those during an atrial fibrillation episode, and therefore cannot capture the real-time information of the atrial matrix carried by the f-wave during an atrial fibrillation episode. The f-wave (fibrillatory wave) is the fibrillation waveform that replaces the normal P wave during an atrial fibrillation episode. Its spectral and amplitude characteristics have been confirmed by multiple studies to be significantly associated with the area of ​​low voltage regions in the left atrium and the degree of fibrosis. Specifically, the dominant frequency of the f-wave reflects the length of the atrial refractory period and the velocity of the reentry circuit, the amplitude of the f-wave reflects the amount of viable myocardium and the extent of fibrosis, and the spectral entropy of the f-wave reflects the degree of orderliness of atrial electrical activity. However, there is currently no method to systematically fuse these deep features of the f-wave with clinical risk factors across modalities and directly predict stroke endpoint events, nor is there a technical solution to introduce a time-dependent consistency index into the deep learning loss function to optimize the stroke probability calibration.

[0006] Therefore, there is an urgent need for a method and system that can deeply extract f-wave features reflecting the state of left atrial matrix remodeling from atrial fibrillation electrocardiogram signals and fuse them with clinical variables across modalities to achieve accurate dynamic prediction of stroke risk. Summary of the Invention

[0007] To address the aforementioned technical problems, this invention provides a method and system for predicting stroke risk in atrial fibrillation patients using electrocardiogram (ECG) signals. By performing time-frequency dual-channel deep encoding on the f-wave residual signal in the twelve-lead ECG signal during an atrial fibrillation episode and integrating it with clinical variables across modalities, dynamic and accurate prediction of ischemic stroke risk in atrial fibrillation patients can be achieved.

[0008] On the one hand, the present invention provides a method for predicting stroke risk in patients with atrial fibrillation using electrocardiogram signals, comprising the following steps:

[0009] Step S1, Acquisition of atrial fibrillation ECG signals and extraction of f-wave residual signals: The twelve-lead ECG signals of patients with atrial fibrillation are continuously acquired from the dynamic ECG monitoring device. The atrial fibrillation episode segments in the acquired twelve-lead ECG signals are subjected to QRS complex detection and elimination processing to obtain the f-wave residual signals of each lead after removing the ventricular activity component. The QRS complex detection and elimination processing includes R-wave peak localization of the continuous ECG signal, construction of QRS template and beat-by-beat subtraction to separate the f-wave component reflecting the electrical activity of atrial fibrillation.

[0010] Step S2, Lead Spatial Topology Arrangement and Time-Frequency Dual-Path Encoding: The f-wave residual signals of each lead are arranged into a two-dimensional matrix according to the lead spatial topology. This two-dimensional matrix is ​​then input into the frequency domain encoding path and the time domain encoding path for parallel feature extraction. The frequency domain encoding path extracts the dominant frequency drift feature and spectral entropy decay feature from the short-time Fourier transform spectrum of the f-wave signal to reflect the degree of left atrial electrical remodeling. The time domain encoding path extracts the amplitude variability feature and tissue index feature from the f-wave amplitude envelope sequence to reflect the extent of left atrial fibrosis matrix. The frequency domain feature vector and the time domain feature vector are then concatenated into an f-wave depth feature embedding vector.

[0011] Step S3, ECG-Clinical Cross-Modal Alignment Fusion: Acquire the patient's clinical variables, encode the clinical variables to obtain a clinical variable encoding vector, and perform cross-modal alignment fusion with the f-wave deep feature embedding vector and the clinical variable encoding vector through a gated bilinear fusion layer to generate a joint risk representation vector. The gated bilinear fusion layer includes a gated attention mechanism to adaptively adjust the contribution weights of ECG features and clinical features.

[0012] Step S4, Stroke Risk Dynamic Prediction and Probability Calibration: The joint risk representation vector is input into the risk prediction head network, which outputs the patient's probability of ischemic stroke within a preset future time window. The risk prediction head network is trained using a composite loss function that includes a time-dependent consistency index optimization term, instead of a single loss function that only uses binary cross-entropy. This results in a higher accuracy in predicting probabilities compared to a logistic regression model based solely on the CHA2DS2-VASc score.

[0013] On the other hand, the present invention also provides a stroke risk prediction system for atrial fibrillation electrocardiogram signals, comprising: an f-wave residual signal extraction module configured to perform the function of step S1; a time-frequency dual-channel encoding module configured to perform the function of step S2; a cross-modal alignment fusion module configured to perform the function of step S3; and a stroke risk prediction module configured to perform the function of step S4.

[0014] Compared with existing technologies, the beneficial effects of this invention are as follows: First, by performing time-frequency dual-channel deep encoding on the f-wave residual signal, it can simultaneously capture frequency domain features reflecting the degree of left atrial electrical remodeling and time domain features reflecting the extent of fibrotic matrix. This significantly improves the utilization efficiency of atrial matrix information in ECG signals compared to using only coarse-grained indicators such as atrial fibrillation load, and overcomes the limitation of existing technologies that only statistically analyze atrial fibrillation duration while ignoring the intrinsic information of the waveform. Second, by mapping the twelve-lead signal into a two-dimensional matrix that maintains physical spatial adjacency through lead spatial topological arrangement, the convolutional encoder can naturally capture the spatial correlation between leads, thus integrating it into the feature extraction stage. The method utilizes spatial distribution information of atrial electrical activity; third, it achieves cross-modal alignment and fusion of deep ECG features and clinical variables through a gated bilinear fusion layer. The gated attention mechanism can adaptively adjust the contribution weights of different modal features, enabling complementary enhancement of two heterogeneous features in a unified representation space. Its predictive performance is superior to schemes using either modality alone, and also superior to simple feature splicing fusion methods. Fourth, it employs a composite loss function with a time-dependent consistency index optimization term to train the risk prediction model, fully utilizing the temporal information and censored data characteristics in survival analysis scenarios. This results in significantly better ranking consistency and calibration of predicted probabilities than traditional cross-entropy training methods. In a multi-center validation cohort, the C-index of this invention reached 0.81, an improvement of 0.14 over the CHA2DS2-VASc score. The NRI for stroke net reclassification improvement in patients with intermediate-risk scores of 2 to 3 reached 0.38, which can assist clinicians in accurately identifying patients who truly need anticoagulation therapy. This has significant clinical application value in reducing unnecessary anticoagulation bleeding risk and reducing missed stroke events. Attached Figure Description

[0015] Figure 1 This is a flowchart illustrating the method for predicting stroke risk in atrial fibrillation using electrocardiogram signals, as provided in an embodiment of the present invention.

[0016] Figure 2 This is a schematic diagram of the architecture of the atrial fibrillation electrocardiogram signal stroke risk prediction system provided in an embodiment of the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0018] This invention provides a method for predicting stroke risk based on electrocardiogram signals in patients with atrial fibrillation, such as... Figure 1 As shown, this method starts with dynamic electrocardiogram (ECG) monitoring data from atrial fibrillation patients and achieves dynamic and accurate prediction of ischemic stroke risk through time-frequency dual-channel deep feature encoding of the f-wave residual signal and cross-modal alignment fusion of clinical variables. In a preferred scenario of this embodiment, the method is deployed on a hospital's ECG data analysis platform, receiving continuous ECG signal streams from dynamic ECG monitoring devices (such as Holter recorders or wearable patch ECG monitors) and automatically completing the end-to-end processing flow from signal acquisition to risk probability output. The specific implementation methods of each step are described in detail below.

[0019] Step S1: Acquisition of atrial fibrillation ECG signals and extraction of f-wave residual signals. In one embodiment of the present invention, twelve-lead ECG signals from atrial fibrillation patients are first continuously acquired from a dynamic ECG monitoring device. Specifically, the dynamic ECG monitoring device used can be a clinically commonly used 24-hour to 72-hour Holter recorder, or a wearable patch-type ECG monitor that has emerged in recent years. Its sampling rate is preferably set to no less than 500 Hz to ensure sufficient acquisition of f-wave waveform details. In a typical implementation scenario, the sampling rate is set to 1000 Hz, with simultaneous acquisition of twelve-lead signals. Each sampling point contains voltage values ​​from 12 channels, and the voltage resolution is no less than 1. V. During the data collection process, the device automatically records timestamp information for subsequent time alignment with clinical events.

[0020] After acquiring the raw 12-lead ECG signal, preprocessing is required to remove baseline drift and power line interference. Preferably, baseline drift removal uses a high-pass filter with a cutoff frequency of 0.5 Hz, and power line interference removal uses a notch filter with a center frequency of 50 Hz. The preferred filter type is a zero-phase Butterworth filter to avoid introducing phase distortion during filtering. The signal after baseline correction and denoising will be used as input for subsequent processing.

[0021] Subsequently, the preprocessed ECG signal is automatically detected and segmented into atrial fibrillation (AF) segments. In one embodiment of the invention, AF segment detection is achieved based on RR interval irregularity analysis combined with P wave absence detection. Specifically, firstly, the beat-by-beat RR interval sequence is calculated for the entire ECG recording. When a signal segment of more than 30 seconds satisfies an RR interval irregularity index (i.e., the ratio of the root mean square difference between adjacent RR intervals to the average RR interval) greater than 0.10 and no regular P wave is detected simultaneously, the signal segment is marked as an AF segment. Preferably, to ensure the reliability of subsequent feature extraction, the length of each marked AF segment should be no less than 10 seconds, and the truncation length is preferably 10 to 30 seconds. In one embodiment of the invention, a 30-second continuous segment is preferably truncated for subsequent analysis.

[0022] After confirming the atrial fibrillation episode segment, the crucial QRS complex detection and elimination process is performed to extract the f-wave residual signal. QRS complex elimination is a fundamental step in f-wave analysis, aiming to remove ventricular activity components with amplitudes much larger than the f-wave from the electrocardiogram signal, thus exposing the f-wave signal reflecting the electrical activity of atrial fibrillation. In one embodiment of the invention, QRS elimination is specifically performed according to the following steps: First, an adaptive matched filter is used to locate the R-wave peak in each lead signal. In a preferred embodiment of the invention, an improved Pan-Tompkins algorithm is used. This algorithm sequentially performs bandpass filtering (passband from 5 Hz to 15 Hz), differentiation, square sum, and sliding window integration on the signal. Then, the R-wave peak position is detected by an adaptive threshold, and the R-wave detection results of 12 leads are fused through multi-lead voting to improve the localization accuracy. Secondly, taking each detected R-wave peak as the center, a window is truncated forward by 80 ms to 120 ms and backward by 80 ms to 120 ms (preferably 100 ms forward and backward in one embodiment of the invention), and the waveform within this window is extracted as the QRS-T wave template for that heart beat. To improve the template quality, it is preferable to take the median or mean of the QRS-T wave templates of 10 to 20 consecutive adjacent heart beats to construct a smooth QRS-T template. Finally, the constructed QRS-T template is subtracted from the original signal beat by beat, that is, for each heart beat, the original signal is subtracted from the QRS-T template of that beat at its corresponding time position to obtain the f-wave residual signal.

[0023] It is worth noting that during the QRS elimination process described above, due to slight differences in the morphology of the QRS-T complex between different heartbeats (especially in atrial fibrillation where irregular RR intervals lead to changes in T wave morphology), this invention also employs an adaptive template update strategy. That is, the QRS-T template is recalculated every 5 to 10 heartbeats to adapt to the gradual changes in waveform. After the above processing, the f-wave residual signal in each lead retains the atrial fibrillation wave information during an atrial fibrillation episode, with its main frequency range being approximately 3 Hz to 12 Hz and amplitude typically between 0.02 mV and 0.50 mV. In one embodiment of this invention, the extracted f-wave residual signal is further bandpass filtered (passband of 3 Hz to 12 Hz) to further remove residual ventricular far-field potentials and high-frequency noise components. Preferably, after bandpass filtering, signal quality assessment is performed, calculating the signal-to-noise ratio (SNR) of the f-wave residual signal. When the SNR of a certain lead is lower than a preset threshold (set to 3 dB in one embodiment of the present invention), the lead is marked as a low-quality lead, and compensation is made in subsequent lead spatial topology arrangement through weighted interpolation of adjacent leads. Furthermore, for multiple atrial fibrillation episodes that may occur in the same patient during a single monitoring session, the present invention preferably selects the segment with the best signal quality for analysis, or performs feature extraction on multiple segments separately and takes the mean of the feature vectors as the representative feature vector of the patient, thereby improving the robustness of the prediction results.

[0024] Step S2: Lead Spatial Topological Arrangement and Time-Frequency Dual-Path Encoding. After obtaining the f-wave residual signals of the twelve leads, this step first arranges these signals in an ordered manner according to the spatial topological relationship of the leads, and then performs deep feature extraction through two parallel encoding paths in the frequency and time domains respectively. The design concept of the lead spatial topological arrangement is that the leads of the twelve-lead electrocardiogram are not spatially independent channels, but rather projections of cardiac electrical activity onto the body surface from different angles, and there are definite spatial geometric relationships between the leads. Arranging these leads in an ordered manner according to their projection positions on the surface of the human chest cavity ensures that leads that are spatially adjacent in the two-dimensional matrix are also physically adjacent, thus facilitating the subsequent convolutional encoder to capture spatial local correlations.

[0025] In one embodiment of the present invention, the spatial topological arrangement of the twelve leads is specifically as follows: the limb leads (I, II, III, aVR, aVL, aVF) and the chest leads (V1 to V6) are mapped into a 4x3 two-dimensional matrix according to their spatial positional relationship on the human body surface. Preferably, the matrix is ​​arranged as follows: the first row contains aVL, I, and -aVR (taking the negative value of aVR lead to obtain continuous frontal projection), the second row contains II, aVF, and III, the third row contains V1, V2, and V3, and the fourth row contains V4, V5, and V6. This arrangement ensures that the top two rows reflect the spatial relationship of the frontal plane lead system, the bottom two rows reflect the spatial relationship of the horizontal plane lead system, and the order of V1 to V6 in the matrix corresponds to their physical position on the chest wall from right to left. In one embodiment of the present invention, the residual signal of the f-wave of each lead is truncated to a length of 30 s. At a sampling rate of 500 Hz, each lead corresponds to 15,000 sampling points. Therefore, the dimension of the entire two-dimensional matrix is ​​4×3×15000.

[0026] After completing the lead spatial topological arrangement, the two-dimensional matrix is ​​input into the frequency domain encoding path and the time domain encoding path respectively for parallel feature extraction. This time-frequency dual-path parallel encoding design is based on the following technical considerations: the f-wave signal contains rich information reflecting the electrophysiological state of the atrium, but this information is distributed in different signal representation domains. Frequency domain features mainly reflect the frequency composition and spectral structure changes of the atrial fibrillation wave, which are closely related to the degree of electrical remodeling of the left atrium; while time domain features mainly reflect the time-varying characteristics of the amplitude and morphology of the f-wave, which are closely related to the extent of structural fibrosis in the left atrium. Using features from either domain alone is insufficient to fully characterize the pathological state of the atrial matrix; therefore, this invention adopts a strategy of dual-path parallel extraction and then splicing and fusion.

[0027] The specific implementation of the frequency domain coding path is as follows. In one embodiment of the present invention, a short-time Fourier transform (STFT) is performed on the f-wave residual signal at each lead position in the two-dimensional matrix to obtain the time-frequency representation of the signal. The window function of the STFT is preferably a Hanning window, with a window length of 256 sampling points (corresponding to 0.512s at a sampling rate of 500 Hz) and a step size of 64 sampling points (corresponding to 0.128 s), resulting in a frequency resolution of approximately 1.95 Hz. After the STFT, the f-wave signal of each lead is converted into a two-dimensional time-spectrum matrix, with the horizontal axis representing the time frame index and the vertical axis representing the frequency index. Preferably, only the spectrum region with a frequency range between 3 Hz and 12 Hz is retained, as this range covers the main energy distribution of the f-wave.

[0028] From the time-spectrum matrix, the following two key frequency domain features are extracted:

[0029] The first category is the dominant frequency drift feature. The dominant frequency (DF) is defined as the frequency corresponding to the maximum power spectral density within each time frame in the time-spectrum graph. By calculating the dominant frequency for each time frame, the trajectory of the dominant frequency over time can be obtained, i.e., the dominant frequency drift curve. The dominant frequency reflects the average excitation frequency of the atrial fibrillation wave. Its physical significance lies in the fact that a higher dominant frequency generally indicates more severe atrial electrical remodeling, a shorter atrial refractory period, and greater difficulty in spontaneous termination of atrial fibrillation. In one embodiment of this invention, the dominant frequency value for each time frame is calculated using the following formula: ,in, For the first Each time frame in frequency Complex values ​​of the STFT coefficients at the location, The power spectral density at that location. This represents the frequency search range for the f-wave, in Hz. For the first Each time frame corresponds to a specific moment, measured in seconds. The dominant frequency drift characteristic further includes the mean dominant frequency. Standard deviation of main frequency and the linear regression slope of the main frequency over time ,in This reflects whether the dominant frequency shows a systematic increase or decrease during atrial fibrillation, which is of great significance for assessing the dynamic progress of left atrial electrical remodeling.

[0030] The second category is spectral entropy decay characteristics. Spectral entropy (SE) is a metric for measuring the spectral complexity and disorder of a signal, and its calculation is based on Shannon entropy in information theory. For each time frame, the formula for calculating spectral entropy is: ,in, For the first The time frame in the ... Frequency components Normalized power spectral density at , The total number of frequency components, in one embodiment of the present invention The value is taken as the number of frequency points within the retained frequency band (approximately 5 to 47 points, depending on the frequency resolution and the width of the retained frequency band). A higher spectral entropy value indicates a more uniform distribution of spectral energy and a signal closer to white noise, meaning more disordered atrial electrical activity; a lower spectral entropy value indicates that spectral energy is concentrated in a few frequency components and the f-wave is more organized. The spectral entropy decay characteristic mainly focuses on the trend of spectral entropy over time, calculated by the linear regression slope of the spectral entropy time series. To depict. When When the value is negative, it indicates that the spectral complexity gradually decreases during the atrial fibrillation episode, which may reflect the gradual stabilization of the intraatrial reentry circuit.

[0031] The eigenvalues ​​of the main frequency drift curve and spectral entropy decay curve across all 12 leads are stacked according to the lead dimension and then input into a frequency domain convolutional encoder. This frequency domain convolutional encoder preferably employs a structure consisting of alternating layers of three one-dimensional convolutional layers, batch normalization layers, and ReLU activation function layers, with kernel sizes of 7×1, 5×1, and 3×1, and channel numbers of 32, 64, and 128 respectively. Each convolutional layer is followed by a max-pooling layer to reduce the temporal dimension. After processing by the frequency domain convolutional encoder, a frequency domain feature vector with a dimension of 128 is output. .

[0032] The specific implementation of the time-domain coding path is as follows. The time-domain coding path focuses on the amplitude variation characteristics of the f-wave residual signal in the time domain, mainly extracting two types of features: amplitude variability and organization index.

[0033] The extraction process of amplitude variability features is as follows: For the f-wave residual signal of each lead in the two-dimensional matrix, the amplitude envelope within the sliding window is first calculated. In one embodiment of the present invention, the amplitude envelope is calculated using the Hilbert transform, that is, for the f-wave residual signal... Find its analytic signal ,in for Hilbert transform, amplitude envelope The width of the sliding window is preferably between 200 ms and 500 ms, and is set to 300 ms in one embodiment of the present invention. The calculated amplitude envelope sequence... Further calculate the coefficient of variation As a characteristic of amplitude variability: ,in, The standard deviation of the amplitude envelope sequence, The mean of the amplitude envelope sequence. Dimensionless. The greater the amplitude variability, the more significant the competition and interference between multiple reentrant circuits within the atrium, indirectly indicating higher heterogeneity of the atrial fibrosis matrix.

[0034] The extraction process of the organization index feature is as follows: For two spatially adjacent leads (determined based on the row and column adjacency relationship in the lead spatial topology matrix), the normalized cross-correlation function between their amplitude envelope sequences is calculated, and the peak value of the cross-correlation is taken as a measure of the organization degree of the lead pair. Organization Index Defined as the average of the peak values ​​of the cross-correlation of all adjacent lead pairs: ,in, Let be the set of all adjacent lead pairs in the lead space topology matrix. This represents the total number of adjacent lead pairs. For leads and The cross-correlation function between the amplitude envelope sequences, It is a time-delay variable. The closer the value is to 1, the more synchronized the changes in the f wave amplitude in each lead are and the more organized the atrial electrical activity is. The lower the value, the more uncoordinated the electrical activity in different regions, which may reflect the conductive heterogeneity caused by a wider fibrous matrix.

[0035] The numerical sequences of amplitude variability and organization index features across all leads are stacked according to lead dimensions and then input into a temporal convolutional encoder. This temporal convolutional encoder has a structure similar to the frequency domain encoder, preferably consisting of three one-dimensional convolutional layers with kernel sizes of 7×1, 5×1, and 3×1, and channel numbers of 32, 64, and 128, respectively. After processing by the temporal convolutional encoder, a temporal feature vector with a dimension of 128 is output. .

[0036] Finally, the frequency domain feature vector With time-domain feature vectors The features are concatenated along the feature dimension to generate an f-wave depth feature embedding vector with dimension 256. This embedding vector simultaneously encodes the deep features of the f-wave signal in both the frequency and time domains, providing rich atrial matrix information for subsequent cross-modal fusion with clinical variables. It is worth further elaborating that the time-frequency dual-path encoding design in this invention is not merely a simple extension of feature dimensions, but has a clear clinical pathophysiological basis. The dominant frequency drift and spectral entropy decay features extracted by the frequency domain path mainly correspond to the atrial electrical remodeling process, namely, electrophysiological changes such as shortened atrial refractory period, decreased conduction velocity, and ion channel remodeling; while the amplitude variability and tissue index features extracted by the time domain path mainly correspond to the atrial structural remodeling process, namely, pathological structural changes such as atrial myocardial fibrosis, interstitial collagen deposition, and cardiomyocyte loss. Although electrical remodeling and structural remodeling are closely related, they are not completely synchronous, and their progression and proportions differ in different individuals. Therefore, simultaneously extracting both types of features can more comprehensively characterize the atrial matrix state of individual patients, which is also the pathophysiological basis for the synergistic effect of 1+1>2 achieved by the dual-path encoding design of this invention. In an ablation analysis of one embodiment of the present invention, the 12-month stroke incidence rate of the patient subgroup with a mean dominant frequency greater than 6.5 Hz and a temporal tissueization index less than 0.45 was as high as 8.7%, which was much higher than the average level of 4.2% in the overall cohort, further validating the clinical value of complementary dual-path features.

[0037] Step S3, ECG-Clinical Cross-Modal Alignment and Fusion. The core objective of this step is to fuse the f-wave depth feature embedding vector output from Step S2 with the patient's clinical variables across modalities, generating a joint risk representation vector that comprehensively reflects the atrial electrophysiological state and clinical risk factors. This step is designed based on the following technological understanding: ECG signal depth features reflect the electrophysiological and structural state of the atrial matrix, information completely missing in the CHA2DS2-VASc score; while clinical variables such as age and comorbidities provide important systemic risk background information that cannot be completely replaced by ECG features alone. Since these two types of information originate from different modalities and their feature space distributions differ significantly, simple feature splicing cannot achieve effective fusion; therefore, a specialized cross-modal alignment and fusion mechanism needs to be designed.

[0038] In one embodiment of the present invention, the acquired clinical variables include, but are not limited to, the following: patient age (continuous values ​​in years), left atrial anteroposterior diameter (in mm, obtained by echocardiography), N-terminal pro-B-type natriuretic peptide (NT-proBNP) concentration (in pg / mL, reflecting cardiac volume load and atrial wall tension), D-dimer concentration (in mg / L, reflecting the degree of activation of the coagulation-fibrinolysis system and prethrombotic state), and the binarized codes of each item in the CHA2DS2-VASc score (including congestive heart failure / left ventricular dysfunction, hypertension, age ≥75 years, diabetes, previous stroke / TIA / thromboembolism, vascular disease, age 65 to 74 years, and female, a total of 8 binarized features). Among the above clinical variables, continuous variables (age, left atrial anteroposterior diameter, NT-proBNP, D-dimer) are preferably first subjected to z-score standardization, that is, subtracting the training set mean and dividing by the training set standard deviation.

[0039] The standardized clinical variables are encoded to obtain a clinical variable encoding vector. In one embodiment of the invention, the encoding method is as follows: four standardized continuous variables and eight binary variables are concatenated into a 12-dimensional original clinical vector, which is then mapped through a two-layer fully connected network. The first fully connected layer has an input dimension of 12, an output dimension of 64, and a ReLU activation function; the second fully connected layer has an input dimension of 64, an output dimension of 128, and a ReLU activation function. After encoding, a 128-dimensional clinical variable encoding vector is obtained. Preferably, Dropout regularization is applied to both fully connected layers during training, with the Dropout rate set to 0.3, to prevent overfitting to the limited number of clinical variables.

[0040] In obtaining the f-wave depth feature embedding vector (Dimension 256) and clinical variable encoding vector (With a dimension of 128) After that, cross-modal alignment fusion is performed through a gated bilinear fusion layer. The gated bilinear fusion layer is one of the core innovations of this invention. Its design aims to achieve effective interaction between two types of heterogeneous features in a unified representation space, while adaptively adjusting the contribution weights of different modal features through a gating mechanism.

[0041] The specific implementation of the gated bilinear fusion layer is as follows: First, Through linear mapping Projected to 3D space, will Through linear mapping Projected onto the same 3D space, in which The preferred value is 128. The projected vectors are denoted as follows: and Then, the bilinear interaction tensor is computed. : ,in This represents element-wise multiplication (Hadamard product). Simultaneously, the gated weight vector is calculated. : ,in, For the gated weight matrix, For bias vectors, It is the Sigmoid activation function. This represents a vector concatenation operation. Gated weight vector. Each element takes a value between 0 and 1, controlling the proportion of interaction information between ECG features and clinical features in the corresponding dimension. When the gating value of a certain dimension is close to 1, it indicates that the cross-modal interaction information in that dimension contributes significantly to risk prediction and is fully preserved; otherwise, it is appropriately suppressed.

[0042] Finally, the joint risk representation vector The calculation formula is: ,in, The weight matrix for the residual connections. For the corresponding bias vector This is the activation function for linear rectification. In this formula, The term represents the gated bilinear interaction feature. The term represents a residual connection path, and the sum of the two terms is activated by ReLU. The introduction of residual connections ensures that even if the gating mechanism completely shuts down the bilinear interaction path, the fusion layer can still transmit the original feature information through the residual path, guaranteeing the stability of the training process. After the above fusion operation, the resulting layer has a dimension of... Joint risk representation vector = 128 This vector encodes complementary information from two modalities—ECG signals and clinical variables—in a unified feature space.

[0043] It is worth noting that the gated bilinear fusion layer can automatically learn the optimal contribution ratio of the two modalities in different patient groups during training. In the validation experiments of this invention, it was observed that for younger patients with lower CHA2DS2-VASc scores (0 to 1), the gating mechanism tends to assign higher weights to ECG features because these patients have fewer clinical risk factors, and stroke risk differentiation relies more on a detailed assessment of atrial matrix status. Conversely, for older patients with higher scores (4 and above), the gating mechanism assigns weights to clinical features and ECG features approximately equal. This adaptive characteristic allows the fusion strategy of this invention to maintain good predictive performance in patients at different risk levels.

[0044] Furthermore, this invention introduces residual connection paths in the design of the gated bilinear fusion layer, a design of significant technical importance. In the early stages of deep learning model training, the gated weight vectors may not yet have converged to a reasonable range of values. If the fusion output is entirely dependent on the gated bilinear interaction term, it may lead to gradient vanishing or information bottlenecks. The existence of residual connection paths ensures that, even when the gating mechanism is not yet mature, the original ECG and clinical features can still directly participate in subsequent risk prediction through bypasses, thus guaranteeing stable convergence during training. In the later stages of model training, as the gated weight vectors stabilize, the bilinear interaction term gradually dominates the fusion output, while the residual connection paths play a supplementary and corrective role. This progressive fusion mechanism makes the training of the entire network more stable, and the final fusion effect more ideal.

[0045] Step S4, dynamic prediction and probability calibration of stroke risk. This involves obtaining the joint risk representation vector. The data is then input into a risk prediction head network to output the probability of ischemic stroke within a preset time window for the patient. In one embodiment of the invention, the preset time window is 12 months, meaning the model predicts the probability of an ischemic stroke event occurring within 12 months from the current atrial fibrillation ECG signal acquisition time.

[0046] In one embodiment of this invention, the structure of the risk prediction head network is as follows: the input layer receives data in dimension 128. The vector is mapped to a 64-dimensional hidden representation by the first fully connected layer. It then passes through a batch normalization layer and a ReLU activation function layer, followed by a second fully connected layer that maps the hidden representation to a 1-dimensional scalar value. Finally, a sigmoid output layer compresses this scalar value to the (0,1) interval, which is used as the stroke probability prediction value. Preferably, a Dropout layer is added after the first fully connected layer, with the Dropout rate set to 0.5, to enhance the model's generalization ability on small clinical datasets.

[0047] A core innovation of this invention lies in the design of the training loss function for the risk prediction head network. Traditional stroke risk prediction models typically employ a single binary cross-entropy loss function for training, simply classifying patients into two categories: those who have experienced a stroke and those who have not, and then minimizing the cross-entropy between the predicted probability and the true label. This approach ignores the temporal information and censored data characteristics in survival analysis scenarios, resulting in unsatisfactory ranking consistency and calibration of predicted probabilities.

[0048] The composite loss function of the present invention It consists of two components: ,in For binary classification, the cross-entropy loss term is used. For time-dependent consistency index optimization terms, As a balance coefficient, in one embodiment of the present invention The value ranges from 0.1 to 1.0, with a preferred value of 0.5.

[0049] Binary cross-entropy loss term The calculation formula is: ,in, The total number of training samples, For the first The stroke event labels for each sample (1 indicates that a stroke event occurred during the observation period, and 0 indicates that no stroke event occurred or was deleted before the stroke occurred). For the model to the first Predicted stroke probability for each sample.

[0050] Time-dependent consistency index optimization term The design objective is to optimize the consistency of predictive probability ranking by assigning higher predictive probabilities to patients who experience stroke earlier in all comparable event pairs. The calculation formula is as follows:

[0051] ,

[0052] in, The set of all comparable event pairs. For the first The event time (time of stroke occurrence or censoring) for each sample. This is an event indicator variable (1 indicates that a stroke event was observed, and 0 indicates that it was censored). The total number of comparable event pairs. and The first and the The predicted probability value of each sample. When (That is, when patients who had a prior stroke did indeed have a higher predictive probability), A smaller value indicates lower loss; conversely, a larger value indicates higher loss. This is achieved by minimizing... The model is guided to learn predicted probabilities consistent with the real-world event time order.

[0053] In the model training process, one embodiment of the present invention employs the Adam optimizer, with the initial learning rate set to [value missing]. A cosine annealing learning rate scheduling strategy is employed, with a training period of 100 to 200 epochs and a batch size of 32 to 64. To address the positive-negative sample imbalance caused by the relatively low incidence of stroke events in atrial fibrillation patients (typically 1% to 5% annually), it is preferable to... The method introduces class weighting, which assigns higher loss weights to positive samples (those that have experienced stroke).

[0054] The trained model outputs a stroke probability value between 0 and 1 for input atrial fibrillation ECG signals and clinical variables. In one embodiment of the present invention, the output probability is divided into three risk levels: Low risk Medium risk. This threshold is considered high-risk. The determination of this threshold is based on the optimal Youden index in the validation queue and actual clinical needs. Clinicians can adjust the threshold according to specific scenarios to balance sensitivity and specificity. Preferably, the model also outputs a confidence interval estimate of the predicted probability value. This confidence interval is obtained using the Monte Carlo Dropout method, which involves keeping the Dropout layer active during inference and performing 20 to 50 forward propagations. The mean of the predicted probability is taken as the point estimate, and the 2.5% and 97.5% quantiles are taken as the 95% confidence interval to provide clinicians with quantitative information about the prediction uncertainty.

[0055] In terms of constructing training data for the model, one embodiment of this invention uses a multicenter retrospective cohort as the training set. The training set includes atrial fibrillation patients from the cardiology departments of three tertiary-level hospitals. Inclusion criteria are: age ≥18 years, confirmed atrial fibrillation by electrocardiogram or Holter monitoring, at least 12 months of complete follow-up data, and a 12-lead electrocardiogram record at baseline during an atrial fibrillation episode. Exclusion criteria are: valvular atrial fibrillation, patients with implanted mechanical heart valves, and patients with active malignant tumors and a life expectancy of less than 6 months. The stroke endpoint event is defined as ischemic stroke or transient ischemic attack confirmed by imaging. The training set and validation set are randomly split in an 8:2 ratio, ensuring that the incidence of stroke events in the two groups is substantially the same after splitting.

[0056] In a multicenter validation study of this invention, the proposed method was evaluated in an independent validation cohort comprising 1200 patients with atrial fibrillation across three centers. The median follow-up time in the validation cohort was 18 months, and the stroke event rate was 4.2%. The results showed that the C-index of the proposed method reached 0.81 (95% CI: 0.76–0.86), an improvement of 0.14 compared to the C-index of the CHA2DS2-VASc score (0.67, 95% CI: 0.61–0.73). Particularly noteworthy is that for the intermediate-risk population (420 patients) with CHA2DS2-VASc scores of 2–3, the NRI of stroke reclassification using the proposed method reached 0.38 (95% CI: 0.22–0.54, P<0.001), indicating that this invention can accurately reclassify a significant proportion of patients from the ambiguous intermediate-risk region of the scoring system to truly high-risk or low-risk categories, thus providing a more precise basis for clinical anticoagulation decisions. Furthermore, compared to the unimodal model using only ECG features (C-index 0.74) and the unimodal model using only clinical variables (C-index 0.69), the cross-modal fusion model of this invention achieves significant improvements, validating the effectiveness of the gated bilinear fusion strategy. Regarding calibration evaluation, the Hosmer-Lemeshow calibration test p-value of the method in this invention is 0.42, indicating no significant deviation between the predicted probability and the actual observed event rate, and the calibration curve is close to the diagonal, suggesting high reliability of the predicted probability. In contrast, the calibration test p-value of the ablation model trained based on traditional binary classification cross-entropy loss is only 0.03, indicating significant calibration bias, further confirming the crucial role of the time-dependent consistency index optimization term introduced in this invention in improving probability calibration.

[0057] In subgroup analyses, the method of this invention demonstrated stable predictive performance across patient subgroups of different age groups, genders, and atrial fibrillation types (paroxysmal and persistent). Of particular note is its ability to update stroke risk assessment longitudinally based on the latest atrial fibrillation ECG recordings at each visit, due to the continuous availability of f-wave signals, in the persistent atrial fibrillation subgroup. In a follow-up analysis of this invention, for patients predicted as low risk at baseline but whose predicted probability rose to the intermediate risk range at a 6-month follow-up, the actual stroke incidence over the subsequent 12 months was 6.3%, significantly higher than the 1.2% in patients with persistently low risk. This indicates that the method of this invention has the ability to capture dynamic changes in stroke risk, which is impossible with the static CHA2DS2-VASc score.

[0058] Based on the above method embodiments, the present invention also provides a stroke risk prediction system for atrial fibrillation electrocardiogram signals, such as... Figure 2 As shown. Each functional module of the system corresponds one-to-one with each step in the above method embodiment. The structure and function of each module are described below.

[0059] The f-wave residual signal extraction module is configured to perform the function described in step S1 above. In one embodiment of the present invention, this module includes a signal acquisition subunit and a QRS cancellation subunit. The signal acquisition subunit interfaces with the dynamic ECG monitoring device through a standardized data interface (such as the HL7 / FHIR protocol interface or the DICOM-ECG interface) to receive 12-lead ECG signal data streams in real time or in batches. Preferably, the signal acquisition subunit also includes a data format parser that is compatible with the data formats of ECG devices from different manufacturers (such as EDF, MIT-BIH format, etc.) and uniformly converts them into a standardized internal data structure. The QRS cancellation subunit receives the standardized ECG signal converted by the signal acquisition subunit, performs the aforementioned baseline correction, power line interference removal, atrial fibrillation fragment detection, and QRS complex template construction and beat-by-beat subtraction operations, and outputs the f-wave residual signal of each lead. In one embodiment of the present invention, the core algorithm of the QRS cancellation subunit is implemented using GPU acceleration, and the processing delay for a single 30-second 12-lead signal segment does not exceed 200 ms, which can meet the requirements of near real-time analysis.

[0060] The time-frequency dual-path encoding module is configured to perform the function described in step S2 above. This module receives the 12-lead f-wave residual signal output from the f-wave residual signal extraction module. First, it arranges the signal according to a 4x3 spatial topology matrix using a lead topology arrangement subunit. Then, it sends the signal to both the frequency domain encoding subunit and the time domain encoding subunit for parallel feature extraction. The frequency domain encoding subunit encapsulates a short-time Fourier transform engine and a frequency domain convolutional encoder, outputting a 128-dimensional frequency domain feature vector. The time domain encoding subunit encapsulates a Hilbert transform engine and a time domain convolutional encoder, outputting a 128-dimensional time domain feature vector. The outputs of the two subunits are concatenated along the feature dimensions in the feature concatenation subunit to generate a 256-dimensional f-wave depth feature embedding vector. Preferably, the frequency domain encoding subunit and the time domain encoding subunit adopt a parallel execution architecture in the system implementation, meaning that the two subunits can run simultaneously on different computing threads or cores, thereby reducing the total encoding processing time by approximately half. The structure and parameter settings of the convolutional encoder in each of the above sub-units are as described in step S2 of the method embodiment.

[0061] A cross-modal alignment fusion module is configured to perform the function described in step S3 above. This module includes a clinical variable encoding subunit and a gated bilinear fusion subunit. The clinical variable encoding subunit obtains the patient's clinical variable data from the hospital information system (HIS) or electronic medical record system (EMR), including age, left atrial anteroposterior diameter, NT-proBNP concentration, D-dimer concentration, and information on each item of the CHA2DS2-VASc score. After z-score standardization of continuous variables, it encodes them into a 128-dimensional clinical variable encoding vector through a two-layer fully connected network. The gated bilinear fusion subunit receives a 256-dimensional f-wave depth feature embedding vector from the time-frequency dual-path encoding module and a 128-dimensional clinical variable encoding vector from the clinical variable encoding subunit, and generates a 128-dimensional joint risk representation vector according to the gated bilinear fusion operation described in step S3 of the method embodiment. In one embodiment of the present invention, the module further includes a feature cache subunit for storing f-wave depth feature embedding vectors and joint risk characterization vectors corresponding to atrial fibrillation segments collected from the same patient at different time points, so as to support longitudinal tracking analysis of risk trends.

[0062] The stroke risk prediction module is configured to perform the function described in step S4 above. The core of this module is a risk prediction head network, whose network structure and training loss function are as described in step S4 of the method embodiment. This module receives the joint risk representation vector output by the cross-modal alignment fusion module and outputs the predicted stroke probability value and corresponding risk level through the risk prediction head network. In one embodiment of the invention, this module further includes a result interpretation subunit, used to generate a feature importance analysis report based on gated weight vectors and gradient-weighted class activation mapping (Grad-CAM) technology, showing clinicians which ECG features and clinical variables contribute most to the risk prediction of the current patient, thereby enhancing the interpretability and clinical acceptability of the model prediction results. Preferably, the stroke risk prediction module also includes a report generation subunit, which automatically generates a structured risk assessment report and pushes it to the attending physician's workstation through the hospital information system. The report preferably includes the following: a summary of the patient's basic information, the time and quality assessment of the atrial fibrillation ECG signal acquisition, a summary of f-wave characteristic parameters (mean dominant frequency, mean spectral entropy, amplitude coefficient of variation, and tissue index), a summary of clinical variables, predicted stroke probability and its confidence interval, risk level classification results, a comparison trend with the previous assessment results, and a list of the top five key influencing factors based on gating weights. The report format is preferably a structured document format conforming to clinical electronic medical record integration standards, facilitating direct embedding into the patient's electronic health record for long-term recording and review.

[0063] The modules described above are connected via standardized data interfaces. These modules can be deployed on the same computing device or distributed across different computing nodes. In one embodiment of the invention, the entire system is deployed on a hospital's edge computing server, equipped with an NVIDIA T4 or equivalent GPU computing card, running Linux as the operating system, and employing ONNX Runtime or TensorRT as the deep learning inference framework to achieve optimal inference performance. Under this deployment configuration, the end-to-end processing latency from receiving the raw ECG signal to outputting the stroke risk prediction result does not exceed 2 seconds, meeting the needs of near real-time clinical risk assessment.

[0064] Those skilled in the art will understand that the division of functional modules in the above system embodiments is merely an example of logical functional division, and other division methods may exist in actual implementation. The functions of each module can be implemented by a general-purpose processor executing a computer program stored in memory, or partially or entirely by dedicated hardware circuitry. In another embodiment of the present invention, the functions of the above system can also be provided via cloud services, i.e., model inference is deployed on a cloud GPU cluster, and the hospital uploads ECG data and receives prediction results through a secure encrypted channel. In the cloud deployment scenario, the system further integrates a data anonymization module and a secure transmission module to ensure the privacy and security of patient ECG data and clinical information during transmission and processing.

[0065] Furthermore, in a preferred embodiment of the present invention, the system further includes an online model update module. This module is configured to continuously collect clinical feedback data (i.e., actual stroke events or no-event outcomes confirmed by follow-up after prediction) after deployment, and periodically use the newly accumulated data to incrementally fine-tune the risk prediction head network. Preferably, the incremental fine-tuning adopts an Elastic Weight Consolidation (EWC) strategy, which adapts to the new data distribution while retaining the model's predictive ability on historical training data, avoiding catastrophic forgetting problems. The preferred triggering condition for online model updates is: when the number of newly accumulated samples with complete follow-up outcomes exceeds a preset threshold (e.g., 100 cases) and the C-index on the validation set decreases by more than 0.02, an incremental fine-tuning round is automatically initiated.

[0066] In another embodiment of the invention, the system further integrates an anticoagulation decision support module. This module automatically generates individualized anticoagulation therapy recommendations based on the stroke probability value and risk level output by the stroke risk prediction module, combined with the patient's bleeding risk score (such as the HAS-BLED score). Specifically, when the stroke probability value exceeds the anticoagulation benefit threshold and the bleeding risk is acceptable, it is recommended to initiate or maintain anticoagulation therapy; when the stroke probability value is below the anticoagulation benefit threshold, it is recommended to temporarily suspend anticoagulation and strengthen monitoring. The output of this module is only for clinical auxiliary reference; the final anticoagulation decision still rests with the attending physician. The system does not replace the physician's clinical judgment, but rather assists the physician in making more accurate and individualized anticoagulation therapy decisions by providing quantified risk information and evidence-based recommendations.

[0067] The embodiments of the present invention are not limited to the specific embodiments described above. Those skilled in the art can make various equivalent changes or substitutions based on the technical solutions of the present invention, and all such changes or substitutions should be included within the protection scope of the present invention.

Claims

1. A method for predicting stroke risk using electrocardiogram signals in atrial fibrillation, characterized in that, Includes the following steps: Step S1, Acquisition of atrial fibrillation ECG signals and extraction of f-wave residual signals: The ECG signals of twelve leads of patients with atrial fibrillation are continuously acquired from the dynamic ECG monitoring device. The atrial fibrillation episode segments in the acquired ECG signals are subjected to QRS complex detection and elimination processing to obtain the f-wave residual signals of each lead after removing the ventricular activity component. The QRS complex detection and elimination processing includes R-wave peak localization of the continuous ECG signal, construction of QRS template and beat-by-beat subtraction to separate the f-wave component reflecting the electrical activity of atrial fibrillation. Step S2, Lead Spatial Topology Arrangement and Time-Frequency Dual-Path Encoding: The f-wave residual signals of each lead are arranged into a two-dimensional matrix according to the lead spatial topology. The two-dimensional matrix is ​​input into the frequency domain encoding path and the time domain encoding path for parallel feature extraction. The frequency domain encoding path extracts the main frequency drift feature and spectral entropy decay feature from the short-time Fourier transform spectrum of the f-wave signal to reflect the degree of left atrial electrical remodeling. The time domain encoding path extracts the amplitude variability feature and tissue index feature from the f-wave amplitude envelope sequence to reflect the extent of left atrial fibrosis matrix. The frequency domain feature vector and the time domain feature vector are concatenated into an f-wave depth feature embedding vector. Step S3, ECG-Clinical Cross-Modal Alignment Fusion: Obtain the patient's clinical variables, encode the clinical variables to obtain a clinical variable encoding vector, and perform cross-modal alignment fusion of the f-wave deep feature embedding vector and the clinical variable encoding vector through a gated bilinear fusion layer to generate a joint risk representation vector. The gated bilinear fusion layer includes a gated attention mechanism to adaptively adjust the contribution weights of ECG features and clinical features. Step S4, Stroke risk dynamic prediction and probability calibration: Input the joint risk representation vector into the risk prediction head network and output the patient's ischemic stroke probability value within a preset time window.

2. The method for predicting stroke risk in atrial fibrillation electrocardiogram signals according to claim 1, characterized in that, In step S1, the QRS complex detection and elimination process specifically includes: acquiring the twelve-lead ECG signal at a sampling rate of not less than 500 Hz, using an adaptive matched filter to locate the R wave peak, constructing a QRS template by extracting windows of 80 ms to 120 ms before and after the detected R wave peak, and removing the QRS-T complex component from the original signal by step-by-step template subtraction; the truncation length of the f wave residual signal is a continuous segment of 10 s to 30 s in a single atrial fibrillation episode.

3. The method for predicting stroke risk in atrial fibrillation electrocardiogram signals according to claim 1, characterized in that, In step S2, the lead spatial topology arrangement specifically involves mapping the twelve leads to a 4x3 two-dimensional matrix based on the spatial positional relationship between limb leads I, II, III, aVR, aVL, aVF and chest leads V1 to V6 projected onto the human thoracic cavity, so that spatially adjacent leads maintain an adjacent relationship in the matrix.

4. The method for predicting stroke risk in atrial fibrillation electrocardiogram signals according to claim 1, characterized in that, In step S3, the clinical variables include patient age, left atrial anteroposterior diameter, N-terminal pro-B-type natriuretic peptide concentration, and D-dimer concentration, as well as the binarized encoding of each item in the CHA2DS2-VASc score.

5. The method for predicting stroke risk in atrial fibrillation electrocardiogram signals according to claim 1, characterized in that, In step S2, the frequency domain coding path specifically includes: performing a short-time Fourier transform on the residual signals of each lead f-wave, setting a window length of 256 sampling points and a step size of 64 sampling points to generate a time-spectrum matrix; extracting the drift trajectory of the main frequency changing with time and the decay curve of the spectral entropy changing with time from the time-spectrum matrix, and inputting the drift trajectory and decay curve into a frequency domain convolutional encoder to generate a frequency domain feature vector.

6. The method for predicting stroke risk in atrial fibrillation electrocardiogram signals according to claim 1, characterized in that, In step S2, the time-domain encoding path specifically includes: calculating the amplitude envelope sequence within a sliding window for the residual f-wave signal of each lead, wherein the width of the sliding window is 200 ms to 500 ms; calculating the coefficient of variation of the amplitude envelope sequence as an amplitude variability feature, and calculating the average value of the cross-correlation peaks of the amplitude envelopes between adjacent leads as an organization index feature; and inputting the amplitude variability feature and the organization index feature into a time-domain convolutional encoder to generate a time-domain feature vector.

7. The method for predicting stroke risk in atrial fibrillation electrocardiogram signals according to claim 1, characterized in that, In step S3, the gated bilinear fusion layer specifically includes: multiplying the f-wave depth feature embedding vector and the clinical variable encoding vector element-wise after linear mapping to obtain a bilinear interaction tensor; simultaneously calculating a gate weight vector to control the passing ratio of each dimension of ECG features and clinical features; and multiplying the gate weight vector and the bilinear interaction tensor element-wise and then performing nonlinear activation to obtain the joint risk representation vector.

8. The method for predicting stroke risk in atrial fibrillation electrocardiogram signals according to claim 1, characterized in that, In step S4, the risk prediction head network is trained using a composite loss function that includes a time-dependent consistency index optimization term, instead of a single loss function that only uses binary cross-entropy. This makes the calibration of the predicted probability better than that of a logistic regression model based solely on the CHA2DS2-VASc score. The composite loss function includes a binary cross-entropy loss term and a time-dependent consistency index optimization term. The time-dependent consistency index optimization term optimizes the consistency of the predicted probability ranking of all comparable event pairs, so that patients who have a stroke event first during the follow-up period receive a higher predicted probability value.

9. The method for predicting stroke risk in atrial fibrillation electrocardiogram signals according to claim 8, characterized in that, In step S4, the preset time window is the next 12 months; the risk prediction head network includes a fully connected layer, a batch normalization layer, an activation function layer and a sigmoid output layer connected in sequence, wherein the hidden dimension of the fully connected layer is 64 to 256.

10. A system for predicting stroke risk in atrial fibrillation using electrocardiogram signals, used to implement the method for predicting stroke risk in atrial fibrillation using electrocardiogram signals as described in any one of claims 1-9, characterized in that, include: The f-wave residual signal extraction module is configured to continuously acquire twelve-lead ECG signals from a dynamic ECG monitoring device, perform QRS complex detection and elimination processing on the atrial fibrillation episode segments in the acquired twelve-lead ECG signals, and obtain the f-wave residual signals of each lead after removing the ventricular activity component. The time-frequency dual-path coding module is configured to arrange the f-wave residual signals of each lead into a two-dimensional matrix according to the lead spatial topology, and perform parallel feature extraction through frequency domain coding path and time domain coding path respectively. The frequency domain coding path extracts the main frequency drift feature and spectral entropy decay feature from the short-time Fourier transform spectrum of the f-wave, and the time domain coding path extracts the amplitude variability feature and organization index feature from the amplitude envelope sequence of the f-wave. The frequency domain feature vector and the time domain feature vector are concatenated into the f-wave depth feature embedding vector. The cross-modal alignment fusion module is configured to acquire the patient's clinical variables and encode them into a clinical variable encoding vector. The f-wave depth feature embedding vector and the clinical variable encoding vector are then cross-modal aligned and fused through a gated bilinear fusion layer to generate a joint risk characterization vector. The stroke risk prediction module is configured to input the joint risk representation vector into the risk prediction head network and output the probability value of ischemic stroke in the future within a preset time window, wherein the risk prediction head network is trained using a composite loss function that includes a time-dependent consistency index optimization term.

Citation Information

Patent Citations

  • CN116570293B