Unsupervised Multimodal Electrocardiogram Abnormality Detection Method Based on Attention Mechanism

By adopting an unsupervised multimodal ECG abnormality detection method based on attention mechanism in the ECG automatic diagnosis algorithm, the problem of low detection efficiency of supervised learning algorithms under unsupervised conditions in the existing technology is solved, and efficient ECG abnormality detection under unsupervised conditions is achieved.

CN115944302BActive Publication Date: 2025-06-27ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211589031.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-09
Publication Date
2025-06-27
Estimated Expiration
2042-12-09

AI Technical Summary

Technical Problem

The existing ECG automatic diagnosis algorithm mainly relies on supervised learning, and there are problems of data label imbalance and unlabeled data value under full utilization, making it difficult to effectively detect electrocardiogram abnormalities under unsupervised conditions.

Method used

An unsupervised multimodal ECG anomaly detection method based on attention mechanism is adopted. By extracting the time domain and frequency domain signals of ECG, a reconstruction model is constructed to reconstruct the time domain input unsupervised, and feature learning is performed using multi-headed attention mechanism and residual connection.

Benefits of technology

Under unsupervised conditions, it can effectively distinguish between normal and abnormal ECG heart shooting data, helping doctors or patients to conduct efficient and accurate ECG abnormality detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115944302B_ABST
    Figure CN115944302B_ABST
Patent Text Reader

Abstract

The present invention discloses an unsupervised multimodal electrocardiogram anomaly detection method based on an attention mechanism, which automatically performs electrocardiogram anomaly detection without supervision according to surface electrophysiological signals. The model also considers the characteristics of frequency-domain signals on the basis of time-series electrical signals, and reconstructs the original signal by encoding and then decoding the input signal. When the difference between the reconstructed time-series signal and the input time-series signal is greater than a certain value, it is determined that the electrocardiogram is abnormal. By extracting features from the time domain and frequency domain signals of the ECG, the present invention can distinguish normal and abnormal ECG heartbeat data without supervision and without using labels; in the case where the labeled electrocardiogram data is limited, it can assist doctors in automatically screening abnormal electrocardiograms, and has reference significance in clinical diagnosis and treatment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of cardiac electrophysiological disease prediction, and particularly relates to an unsupervised multimodal electrocardiogram abnormality detection method based on an attention mechanism. Background Art

[0002] Cardiovascular diseases are a general term for a series of heart or blood vessel related diseases, and have always been one of the main diseases threatening human life, and the mortality rate caused by them still ranks first. ECG (electrocardiogram signal) is one of the earliest biological signals studied and applied in medical clinics by humans. The electrocardiogram reflects the health status of the human heart and is an important basis for clinical diagnosis of cardiovascular diseases. With the rapid growth of the number of electrocardiograms, it is extremely important to use AI algorithms to assist doctors in automated abnormality diagnosis.

[0003] Many existing ECG automatic diagnosis algorithms are obtained through supervised learning based on data analyzed and labeled by expert doctors. However, since experts can only process a small amount of electrocardiogram data, and most of this data comes from the same pattern, the automatic diagnosis algorithms based on supervised learning have their limitations. In addition, the distribution of various types of labels in the labeled data is usually unbalanced, which will affect the diagnostic analysis effect of the model; unlabeled data is also valuable, and it can further reveal the potential pathological information inside the patient's cardiovascular system. Therefore, it is very important to develop an algorithm that can perform unsupervised diagnosis of electrocardiogram data.

[0004] Currently, most existing unsupervised ECG anomaly detection or classification methods do not fully utilize and exploit the characteristics of the data itself. They can be roughly divided into two categories: one is to use clustering algorithms to achieve unsupervised decision-making. For example, the literature [Abawajy JH, Kelarev AV, Chowdhury M. Multistage approach for clustering and classification of ECG data[J]. Computer methods and programs in biomedicine, 2013, 112(3):720-730] uses Gaussian mixture models and K-means clustering algorithms to convert ECG data into numerical features, thereby completing anomaly classification. The other is to achieve ECG anomaly detection end-to-end through deep learning. For example, the literature [Pereira J, Silveira M. Unsupervised representation learning and anomaly detection in ECG sequences[J]. International Journal of Data Mining and Bioinformatics, 2019, 22(4):389-] decodes and reconstructs the input based on the encoded features of a recursive network as an autoencoder. However, the drawback of the above methods is that they do not fully utilize and exploit the information of the ECG, and the model lacks the combination of features in more dimensions. Summary of the Invention

[0005] In view of the above, the present invention provides an unsupervised multimodal electrocardiogram anomaly detection method based on an attention mechanism, which uses the attention mechanism to mine the hidden features in time-domain and frequency-domain data, and unsupervised reconstructs the time-domain input, which can well solve the phenomenon that there is a lack of labeled electrocardiogram data and it is difficult for supervised models to achieve a wide range of automatic diagnosis effects.

[0006] An unsupervised multimodal electrocardiogram anomaly detection method based on an attention mechanism includes the following steps:

[0007] (1) Collect multi-lead electrocardiogram signals from the patient's body surface, and take each heartbeat cycle as a group of electrocardiogram time-domain sequences;

[0008] (2) Perform normalization processing and frequency-domain conversion on each group of electrocardiogram time-domain sequences to obtain corresponding electrocardiogram frequency-domain sequences;

[0009] (3) Construct a reconstruction model based on the attention mechanism, which includes two encoding modules and one decoding module. The two encoding modules are respectively used to encode the electrocardiogram time-domain sequence and the electrocardiogram frequency-domain sequence. After the obtained encoded features are concatenated, they are decoded by the decoding module into an ECG waveform sequence with the same dimension as the input;

[0010] (4) Input the electrocardiogram time-domain sequence and the electrocardiogram frequency-domain sequence into the model pair by pair one by one. Take the minimum average error between the ECG waveform sequence output by the model and the input electrocardiogram time-domain sequence as the loss function, so as to train the model;

[0011] (5) Input the time-domain sequence and the frequency-domain sequence of the electrocardiogram signal to be detected into the trained reconstruction model. According to the reconstruction error between the ECG waveform sequence output by the model and the input electrocardiogram time-domain sequence, determine whether the electrocardiogram signal to be detected is abnormal.

[0012] Furthermore, the normalization process in step (2) adopts the maximum-minimum normalization strategy, and the frequency-domain conversion adopts wavelet transform.

[0013] Furthermore, the two encoding modules in the reconstruction model have the same structure, but do not share weights during the training process and perform feature learning independently.

[0014] Furthermore, the electrocardiogram time-domain sequence and the electrocardiogram frequency-domain sequence input into the model need to be position-encoded first, that is, add time position information to the amplitude corresponding to each moment in the sequence, specifically as follows:

[0015] PE (pos,2i) =sin(pos / 10000 2i / d )

[0016] PE (pos,2i+1) =cos(pos / 10000 2i / d )

[0017] Where: PE (pos,2i) represents the time position information added to the even positions in the sequence, PE (pos,2i+1) represents the time position information added to the odd positions in the sequence, pos represents the time point corresponding to each amplitude in the sequence, d represents the dimension of the encoding, and i is a natural number.

[0018] Furthermore, the encoding module adopts the multi-head attention mechanism and the residual connection, where the multi-head attention mechanism is formed by superimposing multiple self-attention mechanisms; at the same time, during the process of forward propagation and learning parameters of the encoding module, Layer Normalization is performed on each layer of parameters, and normalization is performed on the activation values of each layer.

[0019] Furthermore, the calculation process of the self-attention mechanism is as follows:

[0020]

[0021] Q = X embedding *W Q

[0022] K = X embedding *W K

[0023] V = X embedding *W V

[0024] Where: Attention(Q, K, V) is the output of the self-attention mechanism, X embedding is the electrocardiogram time-domain sequence or electrocardiogram frequency-domain sequence after position encoding, Q, K, and V are the query vector, key vector, and value vector respectively, and W Q , W K , W V are the weight matrices corresponding to the query vector, key vector, and value vector respectively, d k is the output dimension of the self-attention mechanism, and T represents the transpose.

[0025] Furthermore, the encoded features output by the two encoding modules are concatenated along the time axis direction, that is, the dimension after concatenation becomes twice the original, and then the features are decoded into an ECG waveform sequence with the same dimension as the input through the linear mapping layer of the decoding module.

[0026] Furthermore, the process of training the model in step (4) is as follows:

[0027] 4.1 Initialize the model parameters, including the bias vector and weight matrix of each layer, the learning rate, and the optimizer;

[0028] 4.2 Input the electrocardiogram time-domain sequence and electrocardiogram frequency-domain sequence into the model, and the model outputs the corresponding reconstruction result, that is, the ECG waveform sequence through forward propagation, and calculate the loss function L between the output ECG waveform sequence and the input electrocardiogram time-domain sequence;

[0029] 4.3 Continuously update the model parameters by using the optimizer through gradient descent according to the loss function L until the loss function L converges, and the training is completed.

[0030] Furthermore, the expression of the loss function L is as follows:

[0031]

[0032] Where: y i is the i-th amplitude in the output ECG waveform sequence, x i is the i-th amplitude in the input electrocardiogram time-domain sequence, and n is the dimension of the electrocardiogram time-domain sequence.

[0033] Further, the optimizer adopts the Adam optimizer.

[0034] Further, in step (5), calculate the reconstruction error between the output ECG waveform sequence of the model and the input electrocardiogram time-domain sequence. If the reconstruction error is greater than the set threshold, it is determined that the electrocardiogram signal to be detected is abnormal; the threshold is set according to the statistical probability distribution, that is, the threshold is the sum of the mean and variance of the training reconstruction error.

[0035] By extracting the features of the time-domain and frequency-domain signals of the ECG, the present invention can distinguish normal and abnormal ECG heartbeat data without supervision and without using labels. Therefore, the present invention can help doctors or patients to perform efficient and accurate ECG abnormality detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a schematic flow chart of the unsupervised multimodal electrocardiogram abnormality detection method of the present invention.

[0037] Figure 2 It is a schematic structural diagram of the sequence feature encoding module in the model of the present invention.

[0038] Figure 3 It is a schematic diagram of the prediction results of the model of the present invention on normal electrocardiogram and abnormal electrocardiogram data. DETAILED DESCRIPTION OF THE INVENTION

[0039] In order to describe the present invention more specifically, the technical solutions of the present invention will be described in detail below with reference to the drawings and specific embodiments.

[0040] As Figure 1 shown, the unsupervised multimodal electrocardiogram abnormality detection method based on the attention mechanism of the present invention includes the following steps:

[0041] (1) Collect electrophysiological signal data from the patient's body surface and process it into heartbeat signals, that is, each heartbeat cycle data.

[0042] (2) Normalize all heartbeat data and calculate using the maximum-minimum normalization strategy.

[0043]

[0044] Among them: x represents the signal amplitude at a certain moment on each lead of the electrocardiogram data, x min represents the minimum value of the heartbeat signal amplitude of this lead, x max represents the maximum value of the heartbeat signal amplitude of this lead, x norm represents the result after normalization.

[0045] (3) Perform frequency domain conversion on the time domain electrocardiogram data of each lead.

[0046]

[0047]

[0048] Where: a is the scale parameter, b is the translation parameter, and f(t) is the electrocardiogram time series signal of a certain lead.

[0049] (4) The electrocardiogram data of each lead is denoted as X = [x1, x2, …, x n , and its corresponding continuous wavelet transform frequency domain is denoted as X f = [f1, f2, …, f n . Use the electrocardiogram time domain and frequency domain signals as the model input part.

[0050] (5) Construct two input sequence feature encoding modules. The two modules have the same structure, but do not share weights during the training process and perform feature learning independently. The inputs of the two encoding modules are the electrocardiogram time domain and its corresponding frequency domain conversion results respectively.

[0051] (6) Perform positional encoding on the input time domain sequence and frequency domain sequence simultaneously, that is, add time position information to the amplitude corresponding to each moment of the original sequence.

[0052] PE (pos,2i) = sin(pos / 10000 2i / d )

[0053] PE (pos,2i+1) = cos(pos / 10000 2i / d )

[0054] Where: pos is the time point at which each amplitude is located in the heartbeat signal, and d represents the dimension of the encoding; the above two calculation methods represent encoding the odd and even positions of the electrocardiogram time domain sequence or frequency domain sequence respectively.

[0055] (7) Input the sequence after positional encoding into the sequence feature encoding module. The sequence feature encoding module mainly includes a multi-head attention mechanism (Multi-Head Attention) and a residual connection, as Figure 2 shown. Among them, the multi-head attention mechanism is composed of a self-attention mechanism (self-attention). And in the process of the model learning parameters during the forward propagation of the feature encoding module, layer normalization is performed on each layer of parameters, and normalization is performed on the activation value of each layer.

[0056] (8) The self-attention mechanism is the focus of sequence feature learning. This part models the input sequence by defining three learnable matrices.

[0057] Q = X embedding *W Q

[0058] K = X embedding *W K

[0059] V = X embedding *W V

[0060] Where: X embedding represents the input sequence after positional encoding, and W Q , W K and W V represent three learnable matrices. Q, K, and V are the vectors that the attention mechanism needs to compare and calculate.

[0061] (9) Subsequently, the self-attention mechanism calculates the relationships between different moments of the sequence based on Q, K, and V, and reallocates weights to modify V.

[0062]

[0063] X embedding = X embedding + Attention(Q, K, V)

[0064] Where: is to transform the attention matrix into a standard normal distribution, and softmax is used for normalization. Subsequently, X is updated and obtained on this basis. embedding .

[0065] (10) Subsequently, on this basis, a linear mapping is performed on the sequence, and layer normalization is performed on the activation layer parameters. Then, a multi-head attention mechanism is formed by stacking multiple self-attention mechanisms.

[0066] (11) After the time-domain data and frequency-domain data pass through their respective corresponding encoding modules, the corresponding encoded features are obtained. Subsequently, the encoded features are concatenated to form new encoded features. The concatenation method is along the time axis direction, that is, the dimension after concatenation becomes twice the original.

[0067] (12) Finally, the features are decoded into the same dimension as the original input time domain through the linear mapping layer.

[0068] (13) The calculation objective function of the model uses the reconstruction error between the decoded output and the time-domain input. The calculation of the reconstruction error is the absolute value average error function.

[0069]

[0070] where: x and y are the time-domain input and the decoded output respectively, and y i and x i are the amplitudes of the reconstructed sequence and its corresponding original sequence respectively.

[0071] (14) The entire model uses the Adam optimizer to update the model parameters.

[0072] m t = β1m t-1 + (1 - β1)g t

[0073]

[0074]

[0075]

[0076]

[0077] where: m t and v t are the estimates of the first moment (mean) and the second moment (biased variance) of the gradient respectively. Both are initialized as 0 vectors, and g t is the gradient of the forward calculation function f A at a certain moment, is the corrected gradient of the Adam for the gradient estimate. The default value of β1 is 0.9, β2 is 0.999, and ∈ is 10 -8 .

[0078] (15) Finally, the abnormal threshold is set according to the reconstruction error of the normal electrocardiogram by the model, and then the abnormal detection task of the electrocardiogram is completed based on this threshold. The threshold set here is set according to the statistical probability distribution, that is, the threshold is the sum of the mean and variance of the training reconstruction error; when the reconstruction error of the model is greater than the threshold, the input sample is judged as an abnormal sample, otherwise it is a normal sample.

[0079] To verify the effectiveness of the present invention, we trained and tested on the public electrocardiogram dataset ECG5000 using the model of the present invention. Among them, ECG5000 is the single-lead electrocardiogram data of a certain patient for 20 hours, and this data contains 5000 complete sequential heartbeat data. We used 80% of the normal electrocardiogram data as the training data of the unsupervised model, and the remaining 20% of the normal electrocardiogram data and the abnormal electrocardiogram data as the test data. Figure 3The results of the prediction of the model of the present invention on normal electrocardiograms and abnormal electrocardiograms are shown. As can be seen from the figure, compared with many related anomaly detection algorithms, the model of the present invention can effectively distinguish normal and abnormal electrocardiogram samples without supervision and can accurately make an automatic diagnosis.

[0080] The above description of the specific embodiments is for the convenience of those of ordinary skill in the art to understand and apply the present invention. It is obvious that those skilled in the art can easily make various modifications to the above specific embodiments and apply the general principles described herein to other embodiments without creative labor. Therefore, the present invention is not limited to the above embodiments, and all improvements and modifications made by those skilled in the art based on the disclosure of the present invention should be within the protection scope of the present invention.

Claims

1. An unsupervised multimodal electrocardiogram anomaly detection method based on the attention mechanism, comprising the following steps: (1) Collect multi-lead electrocardiogram signals from the patient's body surface, and take each heartbeat cycle as a group of electrocardiogram time-domain sequences; (2) Normalize each group of electrocardiogram time-domain sequences and perform frequency-domain conversion to obtain the corresponding electrocardiogram frequency-domain sequences; The normalization process adopts the maximum-minimum normalization strategy, and the frequency-domain conversion adopts wavelet transform; (3) Construct a reconstruction model based on the attention mechanism, which includes two encoding modules and one decoding module. The two encoding modules are respectively used to encode the electrocardiogram time-domain sequence and the electrocardiogram frequency-domain sequence. After the obtained encoded features are concatenated, they are decoded by the decoding module into an ECG waveform sequence with the same dimension as the input; The two encoding modules in the reconstruction model have the same structure, but do not share weights during the training process and independently perform feature learning; The encoding module adopts the multi-head attention mechanism and residual connection, where the multi-head attention mechanism is formed by stacking multiple self-attention mechanisms; at the same time, during the process of forward propagation and learning parameters of the encoding module, LayerNormalization is performed on each layer of parameters, and normalization is performed on the activation value of each layer; the calculation process of the self-attention mechanism is as follows: Q = X embedding *W Q K = X embedding *W K V = X embedding *W V where: Attention(Q, K, V) is the output of the self-attention mechanism, and X embedding is the electrocardiogram time-domain sequence or electrocardiogram frequency-domain sequence after position encoding, Q, K, and V are the query vector, key vector, and value vector respectively, and W Q , W K , and W V are the weight matrices corresponding to the query vector, key vector, and value vector respectively, d k is the output dimension of the self-attention mechanism, T denotes transpose; The encoded features output by the two encoding modules are concatenated along the time axis direction, that is, the dimension after concatenation becomes twice the original, and then the features are decoded into an ECG waveform sequence with the same dimension as the input through the linear mapping layer of the decoding module; (4) Input the electrocardiogram time-domain sequence and the electrocardiogram frequency-domain sequence into the model pair by pair one by one, and use the minimum average error between the ECG waveform sequence output by the model and the input electrocardiogram time-domain sequence as the loss function L, so as to train the model; The electrocardiogram time-domain sequence and the electrocardiogram frequency-domain sequence input into the model need to be position-encoded first, that is, add time position information to the amplitude corresponding to each moment in the sequence, specifically as follows: PE (pos,2i) = sin(pos / 10000 2i / d ) PE (pos,2i+1) = cos(pos / 10000 2i / d ) where: PE (pos,2i) represents the time position information added at even positions in the sequence, and PE (pos,2i+1) represents the time position information added at odd positions in the sequence, pos represents the time point corresponding to each value in the sequence, d represents the dimension of the encoding, and i is a natural number; (5) Input the time-domain sequence and the frequency-domain sequence of the electrocardiogram signal to be detected into the trained reconstruction model, and judge whether the electrocardiogram signal to be detected is abnormal according to the reconstruction error between the output ECG waveform sequence of the model and the input electrocardiogram time-domain sequence; If the reconstruction error is greater than the set threshold, it is determined that the electrocardiogram signal to be detected is abnormal; the threshold is set according to the statistical probability distribution, that is, the threshold is the sum of the mean and variance of the training reconstruction error.

2. The unsupervised multimodal electrocardiogram abnormality detection method according to claim 1, wherein: The process of training the model in step (4) is as follows: 4.1 Initialize the model parameters, including the bias vector and weight matrix of each layer, the learning rate, and the optimizer; 4.2 Input the electrocardiogram time-domain sequence and the electrocardiogram frequency-domain sequence into the model, and the model outputs the corresponding reconstruction result, that is, the ECG waveform sequence through forward propagation, and calculate the loss function L between the output ECG waveform sequence and the input electrocardiogram time-domain sequence; 4.3 Continuously update the model parameters by using the optimizer through the gradient descent method according to the loss function L until the loss function L converges and the training is completed.

3. The unsupervised multimodal electrocardiogram abnormality detection method according to claim 2, characterized in that: The expression of the loss function L is as follows: where: y i is the i-th amplitude in the output ECG waveform sequence, and x i is the i-th amplitude in the input electrocardiogram time domain sequence, and n is the dimension of the electrocardiogram time domain sequence.

Citation Information

Patent Citations

  • Electrocardiosignal anomaly detection method based on sequential depth model

    CN111714117A

  • Electrocardiosignal anomaly detection method, system and device and storage medium

    CN114386457A