A method for classifying arrhythmias based on statistical manifolds
Through the method based on statistical manifold and graph convolution network, the problem of insufficient time series dynamic feature extraction and geometric structure modeling in ECG signal classification is solved, and efficient geometric spatial classification of ECG signals is achieved, which improves the classification accuracy of arrhythmia.
Patent Information
- Application Number
- CN202510439127.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-09
AI Technical Summary
The existing ECG signal classification methods have shortcomings in time series dynamic feature extraction and geometric structure modeling, and it is difficult to fully capture the dynamic timing and geometric structure characteristics of ECG signals, resulting in low classification accuracy.
The method based on statistical manifolds and graph convolution networks is adopted to construct statistical manifolds through dynamic time regular distances and Gaussian kernel matrix, and the classification of ECG signals is combined with graph convolution networks to capture the dynamic timing and geometric characteristics of ECG signals.
It significantly improves the classification accuracy of complex arrhythmias and improves the efficiency and accuracy of electrocardiogram signal classification.
Smart Images

Figure CN119961737B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical data processing and machine learning, and particularly relates to an electrocardiogram signal classification method based on statistical manifolds and graph convolutional networks, which is applicable to the automatic detection of arrhythmias. Background Art
[0002] The electrocardiogram (ECG) signal is a core tool for clinical diagnosis of arrhythmias, which reflects cardiac rhythm abnormalities by recording the electrical activity of the heart. However, traditional ECG analysis relies on manual interpretation, which has the problems of low efficiency and strong subjectivity.
[0003] In recent years, machine learning and deep learning methods such as convolutional neural networks (CNNs) and long short-term memory networks (LSTMs) have been widely used for automatic ECG classification, significantly improving the detection efficiency and accuracy. However, there are still the following two limitations: 1. Insufficient extraction of dynamic features of time series: The ECG signal is essentially a multi-scale time series. Although existing models (such as CNNs and LSTMs) can capture local features, their ability to model global temporal dependencies and statistical relationships between subsequences is limited, making it difficult to fully characterize the dynamic evolution pattern of the signal; 2. Lack of geometric structure modeling: Traditional electrocardiogram classification methods often focus on the extraction of features such as waveform amplitude and duration, while usually ignoring the complex correlations of ECG signals in the time and space dimensions, such as the precise alignment between P-QRS-T waves, the consistent changes in RR intervals, and the continuity of waveform morphology. These geometric and dynamic features are crucial for understanding the electrophysiological state of the heart.
[0004] Based on this, there is an urgent need to develop a new type of arrhythmia classification method that can combine dynamic time warping and manifold space modeling to optimize feature representation, so as to more comprehensively capture the dynamic time series and geometric structure characteristics of ECG signals, improve the classification accuracy, and effectively solve the above problems. Summary of the Invention
[0005] The object of the present invention is to provide an electrocardiogram signal classification method based on statistical manifolds and graph convolutional network GCN (Graph Convolutional Network), which integrates a classification framework of dynamic time warping DTW (Dynamic Time Warping) distance, statistical manifolds and graph convolutional networks. By modeling the covariance manifold structure of ECG subsequences, it can effectively capture the dynamic time series and geometric structure characteristics of ECG signals, improve the classification accuracy, and achieve efficient geometric space classification.
[0006] The object of the present invention is achieved by the following technical solutions:
[0007] An arrhythmia classification method based on statistical manifolds, comprising the following steps:
[0008] A. Heartbeat segmentation and subsequence division:
[0009] First, input the ECG signal, detect the R-wave peaks and segment the ECG data into individual heartbeats; then evenly divide the single heartbeat sequence into N subsequences, where N>1, and each subsequence covers a specific ECG waveform segment;
[0010] B. Subsequence feature modeling and construction of the Gaussian kernel matrix:
[0011] Calculate the dynamic time warping distance between pairwise subsequences, i.e., the DTW distance, to characterize their non-linear temporal alignment relationship; then construct the Gaussian kernel matrix of the subsequences based on the DTW distance;
[0012] C. Construction of the statistical manifold, where all sample kernel matrices form the statistical manifold:
[0013] Each Gaussian kernel matrix is a symmetric positive definite SPD matrix, and each Gaussian kernel matrix corresponds to a point on the Riemannian manifold, i.e., the statistical manifold;
[0014] D. Construction of a graph convolutional network on the manifold:
[0015] First, construct the graph structure: use the covariance matrix as the graph node and the geodesic distance as the edge weight, and perform threshold retention; then perform tangent space mapping to obtain the tangent vector and complete the node feature mapping: project each SPD matrix onto the tangent space through logarithmic mapping to obtain the tangent vector as the node feature; finally, obtain the graph convolutional network classification model and perform graph convolutional classification: use a multi-layer graph convolutional network to aggregate neighborhood information and output the heartbeat category probability.
[0016] Furthermore, step A specifically includes the following steps:
[0017] A1. Heartbeat segmentation:
[0018] A11. R-wave peak detection, using the Pan-Tompkins algorithm to detect the R-wave peak positions from the original ECG signal;
[0019] A12. Heartbeat extraction, taking a fixed time window before and after the R-wave peaks detected in step A11 as the center, and intercepting the complete heartbeat; then, perform time-axis alignment on the intercepted heartbeat to ensure that the R peak is located at the center of the window, and normalize the voltage value to zero mean and unit variance to eliminate the amplitude difference between individuals;
[0020] A2: Subsequence division: Evenly divide a single heartbeat into 11 subsequences, covering the key ECG waveform segments; the division rule is to equally divide along the time axis.
[0021] Further, step A11 specifically includes the following steps:
[0022] A111 Preprocessing: Perform band-pass filtering on the ECG signal with a passband of 5 - 15 Hz to suppress baseline drift and high-frequency noise, and use a differential filter and square operation to enhance the QRS complex characteristics;
[0023] A112 Threshold detection: Set a dynamic threshold. When the amplitude of the filtered signal exceeds the threshold, it is marked as a candidate R peak. Combine the RR interval rule to eliminate false peaks.
[0024] Further, for step A12, take 150 ms before the R peak in the previous segment, corresponding to the start of the P wave to the R peak, and take 350 ms after the R peak in the latter segment, corresponding to the R peak to the end of the T wave. The total cardiac cycle length is 500 ms.
[0025] Further, for step A2, the length of each subsequence is approximately 45 ms, and the specific segmentation corresponds to the ECG physiological characteristics: subsequences 1 - 4 cover the PR interval, corresponding to the atrial depolarization stage; subsequences 5 - 7 cover the QRS complex, corresponding to the ventricular depolarization stage; subsequences 8 - 11 cover the ST segment and T wave, corresponding to the ventricular repolarization stage.
[0026] Further, step B specifically includes the following steps:
[0027] B1 Calculate the DTW distance:
[0028] B11 Construct a cumulative distance matrix, restricting the path slope to 1, that is, only horizontal, vertical, or diagonal movement is allowed for adjacent steps to avoid excessive distortion of the time series relationship;
[0029] B12 Set a global path window to limit the maximum offset of the path on the time axis;
[0030] B13 Perform DTW distance normalization, that is, normalize the cumulative distance of the optimal path to eliminate the influence of subsequence length differences. The DTW distance normalization is carried out according to the following formula:
[0031] ;
[0032] In the formula, is the th cardiac subsequence, is the th cardiac subsequence, is the cumulative distance matrix, P and Q are the lengths of the th cardiac subsequence and the th cardiac subsequence respectively;
[0033] B2. Construct a Gaussian kernel matrix based on the DTW distance:
[0034] B21. Kernel function design: Define a Gaussian kernel function based on the DTW distance, map the distance to a similarity measure, and thus construct a Gaussian kernel matrix:
[0035] ;
[0036] In the formula, is the DTW distance, is the th cardiac electrical sequence, is the th cardiac electrical sequence, is the kernel function bandwidth;
[0037] B22. The Gaussian kernel matrix is a symmetric matrix. Additionally, add a small identity matrix perturbation to ensure its positive definiteness: , where .
[0038] Furthermore, in step B21, the cross-validation method is used to select the optimal value. The objective function is to minimize the classification error of the validation set, and the initial value of
[0039] is taken as 1 / 4 of the median of all DTW distances.
[0040] C1. All symmetric positive definite matrices form a statistical manifold, i.e., a Riemannian manifold, denoted as , and the geodesic distance on the manifold is defined by the following formula:
[0041] ;
[0042] In the formula, is the trace of the matrix;
[0043] C2. Generation of the Gaussian kernel matrix set: The input is cardiac beat samples, and each sample corresponds to a Gaussian kernel matrix; the output is a point set on the statistical manifold;
[0044] C3. Local geometric modeling of the manifold: First, calculate the geodesic distance between all kernel matrices to define the neighborhood relationship, and then construct a distance matrix , where ;
[0045] C4. Generation of the graph structure: Each kernel matrix For a node in the corresponding figure, for each node , retain the top edges with the smallest geodesic distance to it. The edge weight is defined as the reciprocal of the geodesic distance, i.e., , is the scale parameter; use the total mutual information between nodes as a measure of the size of the edge connection threshold, i.e., define: , where is the number of nodes, is the structural mutual information between nodes . When there is no edge connection relationship between , define ; here is a monotonic function of the threshold. When there is an obvious turning point in the curve, the corresponding number of connected edges is the number of connected edges to be retained.
[0046] Furthermore, step D specifically includes the following steps:
[0047] D1. Tangent space projection, project the SPD matrix onto the tangent space through logarithmic mapping to obtain a vector representation in the Euclidean space:
[0048] ;
[0049] where the logarithmic mapping is defined as:
[0050] ;
[0051] In the formula, is the eigenvector matrix of , are the eigenvalues; using symmetry, only retain the upper triangular elements and expand them into a vector, compressing the dimension to ;
[0052] D2. Design of the graph convolutional network architecture, including network layer design and parameter setting;
[0053] D3. Model training and optimization, the loss function uses the following weighted cross-entropy loss to alleviate the problem of class imbalance:
[0054] ;
[0055] In the formula, is the number of samples, is the number of classes, is the class label, is the weight to increase the penalty for the minority class.
[0056] Furthermore, in step D3, the optimization strategy adopted is as follows: Optimizer: Adam, initial learning rate , which decays to 0.5 times every 20 rounds; Batch training: Use full-graph training and switch to sub-graph sampling when memory is insufficient; Early stopping mechanism: Terminate training when the validation set loss does not decrease for 5 consecutive rounds.
[0057] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0058] The arrhythmia classification method based on statistical manifolds of the present invention adopts the DWT distance combined with the statistical manifold space modeling technology to effectively extract the dynamic statistical features of electrocardiogram signals. On this basis, the manifold geometric relationship of heartbeat patterns is modeled through a graph convolutional network, significantly improving the classification accuracy of complex arrhythmias. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0060] Figure 1 Flowchart of arrhythmia classification principle;
[0061] Figure 2 Schematic diagram of heartbeat segmentation;
[0062] Figure 3 Schematic diagram of graph structure generation;
[0063] Figure 4 Principle diagram of tangent space projection;
[0064] Figure 5 Structure diagram of graph convolutional network. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] The following further illustrates the present invention in conjunction with embodiments:
[0066] The following further elaborates on the present invention in detail in conjunction with the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention and not to limit the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention are shown in the drawings rather than all structures.
[0067] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it is not necessary to further define and explain it in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used for differential description and cannot be construed as indicating or implying relative importance.
[0068] As Figure 1 shown, the arrhythmia classification method based on statistical manifolds of the present invention can achieve efficient and accurate classification of the MIT-BIH arrhythmia dataset, including the following steps:
[0069] A. Heartbeat segmentation and subsequence division.
[0070] First, input the ECG signal, detect the R-wave peak and segment the ECG data into individual heartbeats; then, evenly divide the single heartbeat sequence into N subsequences (N>1), and each subsequence covers a specific ECG waveform segment, such as the PR interval, QRS complex, and ST segment.
[0071] B. Gaussian kernel matrix construction.
[0072] Calculate the dynamic time warping distance between pairwise subsequences, that is, the DTW distance, to characterize their non-linear time series alignment relationship, and then construct the Gaussian kernel matrix of the subsequences based on the DTW distance. Specifically, the construction formula of the Gaussian kernel matrix is as follows:
[0073] ;
[0074] In the formula, is the DTW distance, is the th cardiac subsequence, is the th cardiac subsequence, is the kernel function bandwidth.
[0075] C. Construct a statistical manifold, and all sample kernel matrices form a statistical manifold.
[0076] Each Gaussian kernel matrix is a symmetric positive definite (SPD) matrix. Therefore, each Gaussian kernel matrix corresponds to a point on the Riemannian manifold, that is, a statistical manifold. To measure the distance between points on the manifold, the geodesic distance defined by the following formula is used:
[0077] ;
[0078] Among them, is the trace of the matrix.
[0079] D. Construct a graph convolutional network on the manifold.
[0080] First, construct the graph structure: use the covariance matrix as the graph nodes and the geodesic distance as the edge weights, and perform threshold retention. Specifically, retain the weights according to the network information volume k % of the edges. Then, perform tangent space mapping to obtain the tangent vector and complete the node feature mapping: project each SPD matrix onto the tangent space through logarithmic mapping to obtain the tangent vector as the node feature. Finally, obtain the graph convolutional network classification model and perform graph convolutional classification: use a multi-layer graph convolutional network to aggregate neighborhood information and output the heartbeat classification result, that is, the probability of the heartbeat category.
[0081] Specifically, the arrhythmia classification method based on the statistical manifold of the present invention includes the following steps:
[0082] A. Heartbeat segmentation and subsequence division:
[0083] A1. Heartbeat segmentation
[0084] A11. R-wave peak detection. Use the Pan-Tompkins algorithm to detect the R-wave peak position from the original ECG signal. The specific steps are as follows:
[0085] A111. Preprocessing: Perform band-pass filtering on the ECG signal with a passband of 5 - 15 Hz to suppress baseline drift and high-frequency noise, and use a differential filter and square operation to enhance the QRS complex characteristics.
[0086] A112. Threshold detection: Set a dynamic threshold, and mark the filtered signal amplitude exceeding the threshold as a candidate R peak. Combine the RR interval rule to eliminate false peaks.
[0087] A12. Heartbeat extraction. Take a fixed time window before and after the R-wave peak detected in step A11 to intercept the complete heartbeat. Take 150 ms before the R peak in the front section (corresponding to the start of the P wave to the R peak), and 350 ms after the R peak in the back section (corresponding to the R peak to the end of the T wave). The total heartbeat length is 500 ms. Then, align the intercepted heartbeat on the time axis to ensure that the R peak is at the center of the window, and normalize the voltage value to zero mean and unit variance to eliminate the amplitude difference between individuals.
[0088] A2: Subsequence division. Divide a single heartbeat evenly into 11 subsequences to cover the key waveform segments of the ECG. For the specific division, please refer to Figure 2 . The division rule is to equally divide along the time axis, and the length of each subsequence is about 45 ms. The specific segmentation corresponds to the ECG physiological characteristics: subsequences 1 - 4 cover the PR interval (atrial depolarization stage); subsequences 5 - 7 cover the QRS complex (ventricular depolarization stage); subsequences 8 - 11 cover the ST segment and T wave (ventricular repolarization stage).
[0089] For specific types of arrhythmias, the subsequence division can be dynamically adjusted. Additionally, it can be based on waveform inflection point detection. By superimposing multiple heartbeats and their subsequence boundaries, it is ensured that the division is consistent with the physiological waveform. For heartbeats with R-peak detection errors or noise interference, interpolation or elimination strategies are adopted.
[0090] B. Subsequence feature modeling to construct a Gaussian kernel matrix:
[0091] B1. Calculate the DTW distance to provide a distance metric for the Gaussian kernel matrix. DTW minimizes the cumulative distance path by elastically aligning the time axes of two subsequences and is suitable for the non-linear time deformation of ECG signals. First, construct a cumulative distance matrix with the path slope restricted to 1, i.e., adjacent steps are only allowed to move horizontally, vertically, or diagonally to avoid excessive distortion of the timing relationship. Then, set a global path window to limit the maximum offset of the path on the time axis. Finally, perform DTW distance normalization, that is, normalize the cumulative distance of the optimal path to eliminate the influence of subsequence length differences. The DTW distance normalization is carried out according to the following formula:
[0092] ;
[0093] In the formula, is the th cardiac subsequence, is the th cardiac subsequence, is the cumulative distance matrix, P and Q are the lengths of the th cardiac subsequence and the th cardiac subsequence respectively.
[0094] B2. Construct a Gaussian kernel matrix based on the DTW distance to capture the geometric structure relationship between subsequences. DTW overcomes the limitation of the traditional Euclidean distance for rigid alignment of the time axis and accurately captures the dynamic deformation of the ECG waveform, such as the change in the width of the QRS complex. The symmetric positive definiteness of the Gaussian kernel matrix meets the requirements of Riemannian manifold geometry and provides a mathematical basis for subsequent graph convolution.
[0095] B21. Kernel function design. Define a Gaussian kernel function based on the DTW distance to map the distance to a similarity metric, thereby constructing a Gaussian kernel matrix:
[0096] ;
[0097] In the formula, is the DTW distance, is the th cardiac subsequence, is the th cardiac subsequence, is the kernel function bandwidth. The cross-validation method is used to select the optimal value, and the objective function is to minimize the classification error of the validation set. The initial value of
[0098] B22, the Gaussian kernel matrix is a symmetric matrix. Additionally, a small identity matrix perturbation is added to ensure its positive definiteness: , where .
[0099] C. Construct a statistical manifold:
[0100] C1. All symmetric positive definite matrices form a statistical manifold, i.e., a Riemannian manifold, denoted as . The geodesic distance on the manifold is defined by the following formula:
[0101] ;
[0102] In the formula, is the trace of the matrix.
[0103] C2. Generation of the Gaussian kernel matrix set. The input is heartbeat samples, and each sample corresponds to a Gaussian kernel matrix; the output is a point set on the statistical manifold.
[0104] C3. Local geometric modeling of the manifold. First, calculate the geodesic distance between all kernel matrices to define the neighborhood relationship, and then construct the distance matrix , where .
[0105] C4. Generation of the graph structure. Each kernel matrix corresponds to a node in the graph. For each node , retain the first edges with the smallest geodesic distance to it (such as ), and the edge weight is defined as the reciprocal of the geodesic distance, i.e., , is the scale parameter. The total mutual information between nodes is used as a measure of the edge connection threshold size, i.e., define: , where is the number of nodes, is the structural mutual information between nodes . When there is no edge connection relationship between , define . Here is a monotonic function of the threshold. When When there is an obvious turning point in the curve, the number of connected edges corresponding to it is the number of connected edges to be retained.
[0106] D. Construct a graph convolutional network on the manifold:
[0107] D1. Tangent space projection. Project the SPD matrix onto the tangent space through logarithmic mapping to obtain a vector representation in the Euclidean space:
[0108] ;
[0109] where the logarithmic mapping is defined as:
[0110] ;
[0111] In the formula, is 's eigenvector matrix, is the eigenvalue. Using symmetry, only the upper triangular elements are retained and expanded into a vector, and the dimension is compressed to .
[0112] D2. Design of the graph convolutional network architecture, including network layer design and parameter setting.
[0113] D3. Model training and optimization. The loss function uses the following weighted cross-entropy loss to alleviate the class imbalance problem:
[0114] ;
[0115] In the formula, is the number of samples, is the number of classes, is the class label, is the weight to enhance the penalty for the minority class.
[0116] The optimization strategy adopted is: Optimizer: Adam, initial learning rate , decaying to 0.5 times every 20 epochs. Batch training: Use full-graph training (Full-batch), and switch to subgraph sampling (Cluster Sampling) when memory is insufficient; Early stopping mechanism: Terminate training when the validation set loss does not decrease for 5 consecutive epochs.
[0117] Example 1:
[0118] An arrhythmia classification method based on a statistical manifold, including the following steps:
[0119] A. Heartbeat segmentation and subsequence division:
[0120] A1. Heartbeat segmentation
[0121] A11. R-wave peak detection. The Pan-Tompkins algorithm is used to detect the positions of R-wave peaks from the original ECG signal. The specific steps are as follows:
[0122] A111 Preprocessing: The ECG signal is band-pass filtered with a passband of 5 - 15 Hz to suppress baseline drift and high-frequency noise, and a differential filter and square operation are used to enhance the QRS complex characteristics.
[0123] A112. Threshold detection: A dynamic threshold is set, and when the amplitude of the filtered signal exceeds the threshold, it is marked as a candidate R-wave peak. Combining RR interval rules, such as a minimum interval of 200 ms, spurious peaks are eliminated.
[0124] A12. Heartbeat extraction. Taking the R-wave peaks detected in step A11 as the center, fixed time windows are taken before and after respectively to intercept complete heartbeats. 150 ms before the R-wave peak is taken in the front section, corresponding to the start of the P wave to the R-wave peak, and 350 ms after the R-wave peak is taken in the back section, corresponding to the R-wave peak to the end of the T wave. The total length of the heartbeat is 500 ms. With a sampling rate of 360 Hz, there are 180 sampling points in total. Then, the intercepted heartbeats are aligned on the time axis to ensure that the R-wave peak is at the center of the window, and the voltage values are normalized to zero mean and unit variance to eliminate the amplitude differences between individuals.
[0125] A2: Subsequence division. A single heartbeat is evenly divided into 11 subsequences, covering the key waveform segments of the ECG. The specific division is as Figure 2 shown. The division rule is to equally divide along the time axis. The length of each subsequence is approximately 45 ms. The total length is 180 points, with 16 - 17 sampling points in each segment. The specific segmentation corresponds to the physiological characteristics of the ECG: subsequences 1 - 4 cover the PR interval (atrial depolarization stage); subsequences 5 - 7 cover the QRS complex (ventricular depolarization stage); subsequences 8 - 11 cover the ST segment and T wave (ventricular repolarization stage).
[0126] For a specific arrhythmia type, wide QRS complex, the subsequence division can be dynamically adjusted. If the QRS complex is prolonged, the number of sampling points in subsequences 5 - 7 is increased. Additionally, based on waveform inflection point detection, such as derivative extrema, the start and end points of the QRS complex can be located. By superimposing multiple heartbeats and their subsequence boundaries, it is ensured that the division is consistent with the physiological waveform. For heartbeats with incorrect R-wave peak detection or noise interference, interpolation or elimination strategies are adopted.
[0127] B. Subsequence feature modeling of the MIT - BIH dataset, constructing a Gaussian kernel matrix:
[0128] B1. Calculate the DTW distance to provide a distance metric for the Gaussian kernel matrix. DTW aligns the time axes of two subsequences elastically, minimizing the cumulative distance path, and is applicable to the non-linear time deformation of ECG signals. First, construct a cumulative distance matrix with the path slope restricted to 1, i.e., adjacent steps are only allowed to move horizontally, vertically, or diagonally to avoid excessive distortion of the time series relationship. Then, set a global path window to limit the maximum offset of the path on the time axis. For example, the window width is 10% of the length of the subsequence. Finally, perform DTW distance normalization, that is, normalize the cumulative distance of the optimal path to eliminate the influence of the difference in subsequence length. The DTW distance normalization is carried out according to the following formula:
[0129] ;
[0130] In the formula, is the th cardiac subsequence, is the th cardiac subsequence, is the cumulative distance matrix, P and Q are the lengths of the th cardiac subsequence and the th cardiac subsequence respectively.
[0131] B2. Construct a Gaussian kernel matrix based on the DTW distance to capture the geometric structure relationship between subsequences. DTW overcomes the limitation of the traditional Euclidean distance for rigid alignment of the time axis and accurately captures the dynamic deformation of the ECG waveform, such as the change in the width of the QRS complex. The symmetric positive definiteness of the Gaussian kernel matrix meets the requirements of Riemannian manifold geometry and provides a mathematical basis for subsequent graph convolution.
[0132] B21. Kernel function design. Define a Gaussian kernel function based on the DTW distance, map the distance to a similarity metric, and thus construct a Gaussian kernel matrix:
[0133] ;
[0134] In the formula, is the DTW distance, is the th cardiac subsequence, is the th cardiac subsequence, is the bandwidth of the kernel function. The cross-validation method is used to select the optimal value, and the objective function is to minimize the classification error of the validation set. The initial value of
[0135] B22. Gaussian kernel matrix is a symmetric matrix. Additionally, a small identity matrix perturbation is added to ensure its positive definiteness: , where .
[0136] Specifically, for the input data, 11 subsequences are extracted from the preprocessed heartbeat library, and the length of each subsequence is 16 - 17 sampling points. Then, the DTW distance is calculated for each pair of subsequences (a total of pairs). Finally, for each heartbeat, a kernel matrix is generated to characterize the dynamic correlation of its internal subsequences, and a Gaussian kernel matrix is generated. An example (partial) of the kernel matrix is as follows:
[0137] .
[0138] C. Construct a statistical manifold:
[0139] C1. All symmetric positive definite matrices form a statistical manifold, i.e., a Riemannian manifold, denoted as . The geodesic distance on the manifold is defined by the following formula:
[0140] ;
[0141] In the formula, is the trace of the matrix.
[0142] C2. Generation of the Gaussian kernel matrix set. The input is heartbeat samples, and each sample corresponds to a Gaussian kernel matrix; the output is a set of points on the statistical manifold.
[0143] C3. Local geometric modeling of the manifold. First, calculate the geodesic distance between all kernel matrices to define the neighborhood relationship, and then construct a distance matrix , where .
[0144] C4. Generation of the graph structure. Each kernel matrix corresponds to a node in the graph. For each node , retain the first edges with the smallest geodesic distance to it (such as ), and the edge weight is defined as the reciprocal of the geodesic distance, i.e., , is the scale parameter. The total mutual information between nodes is used as a measure of the edge connection threshold size, i.e., define: , where is the number of nodes, is the structural mutual information between nodes , and when When there is no edge connection relationship, define . Here is a monotonic function of the threshold. When there is an obvious turning point in the curve, the number of corresponding edges is the number of edges to be retained.
[0145] An example of graph structure generation is Figure 3 shown. Each is a point on the manifold and is also a node in the graph structure. Whether there is an edge between nodes is determined by the above mutual information-based method.
[0146] Specifically, input 5000 heartbeat samples, and each sample generates a Gaussian kernel matrix. Calculate the geodesic distance: for each pair of kernel matrices , calculate ; the total number of calculations is times of distance calculations. Set the edge retention ratio , and each node retains about 1000 edges to generate the graph structure.
[0147] D. Construct a graph convolutional network on the manifold:
[0148] D1. Tangent space projection. Project the SPD matrix onto the tangent space through logarithmic mapping to obtain a vector representation in the Euclidean space:
[0149] ;
[0150] where the logarithmic mapping is defined as:
[0151] ;
[0152] In the formula, is 's eigenvector matrix, is the eigenvalue. Using symmetry, only retain the upper triangular elements and expand them into a vector, and compress the dimension to . The principle of tangent space projection is shown in Figure 4 .
[0153] D2. Design of the graph convolutional network architecture, including network layer design and parameter setting. The basic structure diagram of the convolutional neural network is shown in Figure 5 , and it mainly includes 3 graph convolutional layers.
[0154] D21. Network layer design. Adopt multi-layer graph convolution (GCN) to aggregate neighborhood information. The specific structure is as follows:
[0155] Input layer: node feature , adjacency matrix .
[0156] Graph Convolutional Layer: ,
[0157] wherein, , is the node representation of the th layer, , is the learnable weight matrix, is the activation function (such as ReLU).
[0158] Output Layer: The final layer node representation ( is the number of classes); The class probabilities are output through the Softmax function as:
[0159] ;
[0160] D22. Parameter Settings.
[0161] Number of Layers: Experiments show that being too deep is likely to lead to over-smoothing. In this embodiment, 3 graph convolutional layers are selected;
[0162] Hidden Layer Dimension: 256 → 128 → 64;
[0163] Activation Function: ReLU (intermediate layer), no activation (output layer);
[0164] Regularization: Dropout (ratio 0.5) to suppress overfitting;
[0165] Weight Decay: L2 regularization, coefficient .
[0166] D3. Model Training and Optimization. The loss function uses the following weighted cross-entropy loss to alleviate the class imbalance problem:
[0167] ;
[0168] In the formula, is the number of samples, is the number of classes, is the class label, is the weight to increase the penalty for the minority class.
[0169] The optimization strategy adopted is: Optimizer: Adam, initial learning rate , decaying to 0.5 times every 20 epochs; Batch Training: Use full-graph training (Full-batch), switch to subgraph sampling (Cluster Sampling) when memory is insufficient; Early Stopping Mechanism: Terminate training when the validation set loss does not decrease for 5 consecutive epochs.
[0170] Specifically, the input graph structure: the number of nodes M = 5000, the feature dimension D = 66 (N = 11); the sparsity of the adjacency matrix is about 20% (10 million non-zero elements). Training configuration: 3-layer GCN, hidden layer dimensions 256-128-64; training for 100 rounds, with each round taking about 15 seconds (NVIDIA A6000 GPU). The performance metrics are: accuracy 98.7%, F1-score 97.2%; an improvement of more than 2.5% compared to the baselines (CNN, LSTM). The arrhythmia classification method based on statistical manifolds of the present invention can significantly improve the classification accuracy of complex arrhythmias.
[0171] Note that the above is only the preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. An arrhythmia classification method based on a statistical manifold, characterized in that, It includes the following steps: A. Heartbeat segmentation and subsequence division: First, input the ECG signal, detect the R-wave peaks, and segment the ECG data into individual heartbeats. Then, evenly divide the single heartbeat sequence into N subsequences, where N > 1, and each subsequence covers a specific ECG waveform segment; A1. Heartbeat segmentation: A11. R-wave peak detection. Use the Pan-Tompkins algorithm to detect the R-wave peak positions from the original ECG signal; A12. Heartbeat extraction. Taking the R-wave peaks detected in step A11 as the center, take a fixed time window before and after respectively to intercept the complete heartbeat; Then, perform time-axis alignment on the intercepted heartbeats to ensure that the R peaks are located at the center of the window, and normalize the voltage values to zero mean and unit variance to eliminate the amplitude differences between individuals; A2: Subsequence division: Evenly divide a single heartbeat into 11 subsequences to cover the key ECG waveform segments; the division rule is to equally divide along the time axis; B. Subsequence feature modeling, constructing a Gaussian kernel matrix: Calculate the dynamic time warping distance between pairwise subsequences, that is, the DTW distance, to characterize their non-linear time series alignment relationship. Then, construct the Gaussian kernel matrix of the subsequences based on the DTW distance; B1. Calculate the DTW distance: B11. Construct the cumulative distance matrix, restricting the path slope to 1, that is, only allowing horizontal, vertical, or diagonal movement for adjacent steps to avoid excessive distortion of the time series relationship; B12. Set the global path window to limit the maximum offset of the path on the time axis; B13. Perform DTW distance normalization, that is, normalize the cumulative distance of the optimal path to eliminate the influence of subsequence length differences. The DTW distance normalization is carried out according to the following formula: Wherein, S i is the i-th cardiac electrical sequence, and S j is the j-th cardiac electrical sequence. C PQ is the cumulative distance matrix, and P and Q are the lengths of the i-th and j-th cardiac electrical sequences, respectively; B2. Construct the Gaussian kernel matrix based on the DTW distance: B21. Kernel function design. Define the Gaussian kernel function based on the DTW distance, map the distance to a similarity measure, and thus construct the Gaussian kernel matrix: where d DTW (,) is the DTW distance, S i is the i-th cardiac electrical sequence, S j is the j-th cardiac electrical sequence, and σ is the kernel function bandwidth; B22. The Gaussian kernel matrix K is a symmetric matrix. Additionally, a small identity matrix perturbation is added to ensure its positive definiteness: K 修正 = K + λI, where λ = 10 -6 ; C. Construct a statistical manifold. All sample kernel matrices form a statistical manifold: Each Gaussian kernel matrix is a symmetric positive definite SPD matrix, and each Gaussian kernel matrix corresponds to a point on the Riemannian manifold, that is, the statistical manifold; C1. All N×N symmetric positive definite matrices form a statistical manifold, i.e., a Riemannian manifold, denoted as The geodesic distance on the manifold is defined by the following formula: In the formula, tr is the trace of the matrix; C2. Generation of Gaussian kernel matrix set. The input is M heartbeat samples, and each sample corresponds to an 11×11 Gaussian kernel matrix. The output is a point set on the statistical manifold C3. Local geometric modeling of the manifold. First, calculate the geodesic distance d between all kernel matrices geo (K i , K j ) to define the neighborhood relationship, and then construct the distance matrix where D ij = d geo (K i , K j ); C4. Generate graph structure generation: Each kernel matrix K i corresponds to a node in the graph. For each node K i , retain the top k% of the edges with the smallest geodesic distance to it. The edge weight is defined as the reciprocal of the geodesic distance, i.e., W ij = exp(-D ij / σ), where σ is the scale parameter; Use the total mutual information between nodes as a measure of the edge connection threshold size, i.e., define: where N is the number of nodes, I(K i , K j ) is the structural mutual information between nodes K i , K j . When there is no edge connection relationship between K i , K j , define I(K i , K j ) = 0; Here I net is a monotonic function of the threshold. When there is an obvious turning point in the I net curve, the corresponding number of connected edges is the number of connected edges to be retained; D. Construct a graph convolutional network on the manifold: First, construct the graph structure: Use the covariance matrix as the graph node and the geodesic distance as the edge weight, and perform threshold retention. Then, perform tangent space mapping to obtain the tangent vector to complete the node feature mapping: Project each SPD matrix onto the tangent space through logarithmic mapping to obtain the tangent vector as the node feature. Finally, obtain the graph convolutional network classification model and perform graph convolutional classification: Use a multi-layer graph convolutional network to aggregate neighborhood information and output the heartbeat category probability; D1, tangent space projection, project the SPD matrix K i onto the tangent space through logarithmic mapping to obtain a vector representation in the Euclidean space: Among them, the logarithmic mapping is defined as: log(K i ) = U·diag(lnλ1, …, lnλ N )·U T where U is the eigenvector matrix of K i , λ i is the eigenvalue; using symmetry, only the upper triangular elements are retained and expanded into a vector, and the dimension is compressed to D2. Graph convolutional network architecture design, including network layer design and parameter setting; D3. Model training and optimization. The loss function uses the following weighted cross-entropy loss to alleviate the problem of class imbalance: where M is the number of samples, C is the number of classes, and y ic is the class label, is the weight to enhance the penalty for the minority class.
2. The arrhythmia classification method based on statistical manifolds according to claim 1, characterized in that Step A11 specifically includes the following steps: A111 Preprocessing: Perform band-pass filtering on the ECG signal with the passband of 5 - 15 Hz to suppress baseline drift and high-frequency noise, and use a differential filter and square operation to enhance the QRS complex characteristics; A112. Threshold detection: Set a dynamic threshold. When the amplitude of the filtered signal exceeds the threshold, it is marked as a candidate R peak. Combine the RR interval rule to eliminate false peaks.
3. The arrhythmia classification method based on a statistical manifold according to claim 1, characterized in that: Step A12: Take 150 ms before the R peak in the previous segment, corresponding to the start of the P wave to the R peak, and take 350 ms after the R peak in the latter segment, corresponding to the R peak to the end of the T wave. The total heartbeat length is 500 ms.
4. A method for classifying arrhythmias based on statistical manifolds according to claim 1, characterized in that: Step A2: The length of each subsequence is about 45 ms. The specific segmentation corresponds to the ECG physiological characteristics: Subsequences 1 - 4 cover the PR interval, corresponding to the atrial depolarization stage; Subsequences 5 - 7 cover the QRS complex, corresponding to the ventricular depolarization stage; Subsequences 8 - 11 cover the ST segment and the T wave, corresponding to the ventricular repolarization stage.
5. A method for classifying arrhythmias based on a statistical manifold according to claim 1, characterized in that Step B21: Use the cross - validation method to select the optimal σ value. The objective function is to minimize the classification error of the validation set, and the initial value is taken as 1 / 4 of the median of all DTW distances.
6. A method for classifying arrhythmias based on a statistical manifold according to claim 1, characterized in that, Step D3, the optimization strategy adopted is: Optimizer: Adam, initial learning rate 10 -3 , decaying to 0.5 times every 20 rounds; Batch training: Use full-graph training and switch to sub-graph sampling when memory is insufficient; Early stopping mechanism: Terminate the training when the loss of the validation set does not decrease for 5 consecutive rounds.
Citation Information
Patent Citations
3D skeleton action recognition method based on space-time manifold trajectory mapping
CN114627557A
Space-time diagram node attribute prediction method fusing adaptive graph diffusion convolutional network
CN115828990A
Premature ventricular beat detection method based on in-beat characteristic wavelet complex network
CN116269248A