ECG intelligent identification algorithm based on multi-modal fusion MFCT model
Through multimodal fusion MFCT model, combined with ConvNeXt and Transformer, the spatial characteristics and category imbalance in ECG signals are solved, efficient classification of arrhythmia and myocardial infarction is achieved, and the identification accuracy and robustness of ECG signals are improved.
Patent Information
- Application Number
- CN202510195064.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-07-18
AI Technical Summary
The existing ECG analysis technology has limitations in processing ECG signals, and it fails to fully utilize the complex spatial characteristics and multimodal information of ECG signals. The problem of category imbalance results in low classification accuracy of arrhythmia in a few categories.
The multimodal fusion MFCT model is adopted, combined with the ConvNeXt network and Transformer, and the one-dimensional ECG signal is converted into a two-dimensional image through recursive graph, spatial features are extracted, and the category imbalance problem is solved through the focus loss function, realizing the fusion of timing and spatial features.
It significantly improves the classification accuracy of arrhythmia and myocardial infarction, improves the recognition ability of ECG signals, and enhances the robustness and generalization ability of the algorithm.
Smart Images

Figure CN120337024A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of biomedical signal processing, and particularly relates to an ECG intelligent recognition algorithm based on a multi-modal fusion MFCT model, which is an algorithm for analyzing and classifying electrocardiogram (ECG) data using a multi-modal fusion method. Background Art
[0002] Arrhythmia is a common and serious heart disease, and its accurate diagnosis is crucial for the treatment and management of patients. Electrocardiogram (ECG), as a non-invasive monitoring method, can provide important information about cardiac electrical activity and is a commonly used tool for diagnosing arrhythmia. However, analyzing ECG signals requires highly specialized doctors, which is time-consuming and labor-intensive. At the same time, the subjective judgment of doctors is prone to lead to diagnostic errors, affecting the treatment effect of patients. Therefore, it is of great significance to develop computer-aided diagnosis (CAD) methods to reduce the burden on doctors and reduce subjective errors. With the rise of machine learning, people have achieved the classification of ECG signals by extracting feature engineering and then deploying traditional machine learning methods such as K-nearest neighbor, random forest, wavelet transform, and support vector machine. Although these techniques are interpretable and perform well, the manual extraction of features is limited by professional knowledge, has limitations, and is time-consuming and subjective.
[0003] In recent years, with the development of deep learning, good results have been achieved in the field of arrhythmia classification. Powerful deep learning techniques can automatically extract features from electrocardiogram signals, which can well solve the problems existing in traditional machine learning methods and significantly improve the accuracy and efficiency of arrhythmia classification.
[0004] Although the existing deep learning networks have made significant progress in arrhythmia classification, there are still some limitations. Current models often do not consider the complex spatial features in ECG signals. On the other hand, these models often do not comprehensively utilize information from different modalities from a multi-modal perspective to more comprehensively describe the features of ECG signals. At the same time, the problem of class imbalance leading to low classification accuracy for rare-class arrhythmias is a common problem for various deep learning models. These challenges highlight the urgent need for an innovative solution that comprehensively utilizes the temporal and spatial features of ECG signals from a multi-modal perspective and effectively solves the problem of class imbalance.
[0005] Based on this, the present invention proposes an ECG intelligent recognition algorithm based on a multi-modal fusion MFCT model, which comprehensively utilizes information from different modalities from a multi-modal perspective to more comprehensively describe the features of ECG signals, which is an urgent technical problem to be solved currently. Summary of the Invention
[0006] Objective of the invention: The objective of the present invention is to address the limitations existing in the current electrocardiogram (ECG) analysis technology when processing ECG signals. In view of the complex spatial features involved in ECG signals, from a multimodal perspective, different modalities of information are comprehensively utilized to more comprehensively describe the features of ECG signals.
[0007] In response to the above challenges, a multimodal fusion model is proposed, which includes several key components, each of which is targeted at the challenges existing in arrhythmia classification. In the first part, multimodal fusion combines one-dimensional time-series signals and two-dimensional images. By fusing information from different modalities, the algorithm can make up for the limitations of one-dimensional time-series signals and two-dimensional images respectively, so as to extract richer and more comprehensive features. In the second part, recurrence plots are used to convert one-dimensional ECG signals into two-dimensional images, aiming to extract spatial features and time-domain and frequency-domain information in ECG signals and avoid the influence of manually extracting features. In the third part, multi-scale spatial feature extraction is carried out. By using the ConvNext network, different-scale feature information can be fused through multi-layer convolution and pooling, and complex and abstract spatial features can be learned. In the fourth part, time-series feature extraction is carried out. By combining TCN and Transformer, local features and global dependencies in ECG signals can be effectively extracted. This combination can enhance the ability to extract time-series features of ECG signals and improve classification accuracy. Finally, we introduce focal loss (FL) to address the class imbalance problem in the ECG dataset. By reducing the weights of easy samples, the model pays more attention to difficult samples, thereby improving the classification performance of the minority classes of the model. Experimental results show that the accuracy of the model in five-class classification in the MIT-BIH database is 99.32%, and a myocardial infarction (MI) classification test is carried out on the PTB diagnostic database, with an accuracy of 99.26%.
[0008] The present invention aims to simultaneously extract time-series features, spatial features, and time-domain and frequency-domain features of ECG signals by combining the advanced ConvNeXt network and Transformer in parallel, improve the recognition ability of ECG signals, and show significant advantages in arrhythmia detection and myocardial infarction detection.
[0009] The technical solution proposed by the present invention includes the following key steps:
[0010] S1. Data preprocessing: The ECG data is preprocessed using a band-pass filter and five-point smoothing filter to reduce the interference of noise on the analysis of ECG signals; first, five-point smoothing filter is used to remove high-frequency noise in the ECG signal, then an IIR digital filter is used to remove the 50Hz power frequency interference in the ECG signal, and finally a Butterworth band-pass filter is used to remove baseline drift and high-frequency noise in the ECG signal by setting the cut-off frequency to 1 - 20Hz, improving the quality of the ECG signal and laying a foundation for subsequent processing.
[0011] S2. Two-dimensional Map Conversion: The Recurrence Plot (RP) is a method for visualizing the periodicity of trajectories in the phase space. The Short-Time Fourier Transform (STFT) is used to transform a signal from the time domain to the frequency domain, converting the one-dimensional electrocardiogram (ECG) signal into a two-dimensional image, usually with a size of 224x224 pixels. This conversion not only preserves the correlation information between the amplitude and phase in the original signal but also extracts the time-domain and frequency-domain feature information in the ECG signal, thus achieving an enhanced display of ECG features. In this way, the dynamic characteristics of the ECG signal can be more intuitively demonstrated, providing richer visual and data support for subsequent analysis and diagnosis.
[0012] S3. Design of the Multimodal MFCT Model: The multimodal MFCT model uses ConvNeXt and Transformer in parallel to achieve feature fusion in the temporal and spatial domains of the ECG signal. In this paper, the ConvNeXt model is innovatively introduced. The ConvNeXt module mainly consists of ordinary convolutional layers, residual modules, downsampling layers, normalization layers, pooling layers, and activation functions. By inputting the two-dimensional recurrence image into the ConvNeXt module, high-dimensional spatial features are obtained. In the design of the Transformer module, to better fuse local and global features, a TCN module is added before the Transformer network and some improvements are made to it. The TCN module includes a temporal attention module, a dilated convolutional layer, and two residual connections to initially extract the local temporal features in the ECG signal. The Transformer model includes an encoder and a decoder, both of which are composed of multiple stacked blocks with the same attention network architecture. Only the encoder module is retained to modify the Transformer. By inputting the one-dimensional ECG signal into the Transformer module, high-dimensional temporal features are obtained. The obtained high-dimensional spatial and temporal features are fused through the FF module and output to the Linear fully connected layer (FC) for calculation to complete the recognition of the ECG signal.
[0013] S4. Model Training: The multimodal MFCT model is trained using the MIT-BIH Arrhythmia Database and the PTB Diagnostic Database. The learning rate of the model is set to 0.0001, the batch size is 64, and the focal loss function and Adam optimizer are used. By learning the temporal and spatial features of different types of ECG signals, the model can accurately classify new ECG data. This step not only includes the training of the model but also the verification of the model's efficacy to ensure its accuracy and efficiency in practical applications.
[0014] Further Technical Details
[0015] Data preprocessing: First, use five-point smoothing filtering to remove high-frequency noise in the electrocardiogram (ECG) signal. Then, use an IIR digital filter to remove power frequency interference in the ECG signal. Finally, use a Butterworth band-pass filter to remove baseline drift in the ECG signal by setting the cut-off frequency, obtaining a high-quality ECG signal.
[0016] Focal loss function: The Focal Loss function is a loss function designed specifically to address the problem of class imbalance, especially suitable for class imbalance tasks. It aims to adjust the weights of the loss function so that the model pays more attention to difficult-to-classify samples during training, thereby improving the overall performance.
[0017] Hierarchical design of the multi-modal MFCT model: The multi-modal MFCT model includes a ConvNeXt module and a Transformer module. The ConvNeXt module is mainly composed of stacked ConvNeXt blocks. The recursive image is downsampled through a convolutional layer and input into the ConvNeXt block. The number ratio of each layer of ConvNeXt blocks is 3:3:9:3. The ConvNeXt block adopts a convolutional structure with convolutional kernel sizes of 7, 1, and 1 respectively. Finally, through operations such as pooling and flattening, high-dimensional spatial features are output. An improvement is made by adding a TCN module before the Transformer module. Each layer of the TCN module includes a temporal attention module, a dilated convolutional layer, and two residual connections. Each layer maintains the same structure and is stacked in two layers. Through its special convolutional structure in the temporal convolution, local temporal features in the ECG signal are initially captured and extracted. Each layer of the Encoder module contains a multi-head self-attention mechanism, a residual connection, and a feed-forward network, and is stacked in three layers to further extract global temporal features in the ECG signal. Finally, high-dimensional temporal features are output through a linear transformation. In the feature fusion module, the obtained features are concatenated and passed through a linear transformation to complete the recognition of the ECG signal.
[0018] Training and optimization strategy: Adopt the Focal Loss function and the Adam optimizer, combined with the Dropout technique to prevent overfitting, ensuring that the model can maintain high accuracy and good generalization ability in various ECG signal recognition scenarios.
[0019] Compared with the existing technology, the beneficial effects of the present invention are as follows:
[0020] The multi-modal MFCT method of the present invention shows significant advantages in processing various electrocardiogram signals. By making full use of the temporal and spatial features in ECG data, this method can make up for the limitations of one-dimensional temporal signals and two-dimensional images respectively, combine them, so as to more comprehensively extract the features of the data, thereby improving the accuracy of classification and recognition and enhancing the robustness of the algorithm. In addition, this method has strong generalization ability and can adapt to different ECG signal variations, greatly improving the practicality and reliability of ECG signal recognition.
[0021] 1. The present invention fuses one-dimensional temporal signals and two-dimensional images. By fusing information of different modalities, the algorithm can make up for the limitations of one-dimensional temporal signals and two-dimensional images respectively, combine them, so as to more comprehensively extract the features of the data, thereby improving the accuracy of classification and recognition and enhancing the robustness and generalization ability of the algorithm.
[0022] 2. The recursive graph of the present invention converts one-dimensional electrocardiogram signals into two-dimensional images. By retaining the correlation information between amplitude and phase in the original one-dimensional signal and the time-domain and frequency-domain information in the electrocardiogram signal, it provides data support for the extraction of spatial features of electrocardiogram signals and avoids the influence of manual feature extraction.
[0023] 3. The present invention uses the ConvNext network to efficiently extract multi-scale spatial features in images through multi-layer convolution operations; a method of combining TCN and Transformer is proposed. Multi-layer temporal convolution operations extract local features of electrocardiogram signals, while Transformer captures the global dependencies between features through the self-attention mechanism. This combination can enhance the temporal feature extraction ability and improve the classification accuracy.
[0024] 4. The present invention introduces the focal loss function, aiming to solve the problem of class imbalance in arrhythmia classification, thereby improving the classification performance of the minority classes of the model. We have verified on two electrocardiogram signal datasets to verify the effectiveness of the model we proposed.
[0025] The innovative advantage of the multi-modal fusion MFCT model lies in the fusion of one-dimensional time series signals and two-dimensional images. By fusing information from different modalities, the algorithm can compensate for the limitations of one-dimensional time series signals and two-dimensional images respectively, thereby extracting richer and more comprehensive features and achieving efficient classification of electrocardiogram signals. By using the ConvNeXt network through multi-layer convolutional operations, it can efficiently extract multi-scale spatial features in images. Through the method of combining TCN and Transformer, multi-layer temporal convolutional operations extract local features of electrocardiogram signals, while Transformer captures the global dependencies between features through the self-attention mechanism. This combination can enhance the temporal feature extraction ability. Combining them can more comprehensively extract the features of data, thereby improving the accuracy of classification and recognition and enhancing the robustness and generalization ability of the algorithm. Brief Description of the Drawings
[0026] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the relevant drawings. Obviously, the described drawings only represent some embodiments of the present invention. For those of ordinary skill in the art, without creative work, other relevant drawings can be deduced based on these drawings.
[0027] Figure 1 It is a flowchart of the method provided by the embodiment of the present invention.
[0028] Figure 2 It is a schematic diagram of the multi-modal MFCT model provided by the embodiment of the present invention.
[0029] Figure 3 It is a schematic diagram of the ConvNeXt module provided by the embodiment of the present invention.
[0030] Figure 4 It is a schematic diagram of the TCN module provided by the embodiment of the present invention.
[0031] Figure 5 It is a schematic diagram of the Transformer module provided by the embodiment of the present invention.
[0032] Figure 6 It is a curve graph of the change in accuracy rate provided by the embodiment of the present invention.
[0033] Figure 7 It is the five-class corresponding heartbeat images and corresponding recursive images of the arrhythmia database provided by the embodiment of the present invention. Detailed Embodiments
[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0035] The present invention proposes an ECG intelligent recognition algorithm based on a multi-modal fusion MFCT model, aiming to improve the accuracy and efficiency of electrocardiogram signal recognition through deep learning technology. The implementation manners of each step are described in detail below:
[0036] As shown in Figure (1), the process design of this method is shown, including data preprocessing, recursive graph conversion, and model training. The specific implementation process is as follows:
[0037] S1: Data preprocessing
[0038] The ECG data is preprocessed using a band-pass filter and five-point smoothing filter to reduce the interference of noise on the electrocardiogram signal analysis; first, the five-point smoothing filter is used to remove the high-frequency noise in the electrocardiogram signal, then the IIR digital filter is used to remove the power frequency interference of the electrocardiogram signal, and finally the Butterworth band-pass filter is used to remove the baseline drift in the electrocardiogram signal by setting the cut-off frequency.
[0039] As a digital signal processing technology, the five-point smoothing filter can reduce the influence of high-frequency noise on the electrocardiogram signal by locally smoothing the data in the electrocardiogram signal. The band-pass filter algorithm is an algorithm widely used in signal processing. It allows signals within a specific frequency range to pass through, while blocking or attenuating signals in other frequency ranges. The frequency range of the electrocardiogram signal mainly concentrates on 0.05 - 100 Hz, and most of the useful information concentrates on 0.5 - 50 Hz. The band-pass filter algorithm can selectively pass the signals within this frequency range while suppressing signals of other frequencies, thereby extracting the useful electrocardiogram signal components. The band-pass filter algorithm can also optimize the waveform features, making the waveform clearer and easier to identify, thereby improving the accuracy of analysis.
[0040] S2: Two-dimensional graph conversion
[0041] The Recurrence Plot (RP) is a method for visualizing the periodicity of trajectories in the phase space. The Short-Time Fourier Transform (STFT) is used to transform a signal from the time domain to the frequency domain, converting a one-dimensional electrocardiogram (ECG) signal into a two-dimensional image, with the image size typically being 224x224 pixels. This transformation not only preserves the correlation information between the amplitude and phase in the original signal but also extracts the time-domain and frequency-domain information in the ECG signal, thus achieving an enhanced display of ECG features. In this way, the dynamic characteristics of the ECG signal can be more intuitively demonstrated, providing richer visual and data support for subsequent analysis and diagnosis.
[0042] S3: Construction of the multi-modal MFCT model
[0043] The established multi-modal MFCT includes a ConvNeXt module and a Transformer module.
[0044] S3-1 uses the ConvNeXt module to extract the spatial features in the two-dimensional recurrence plot. The specific structure is as follows: The ConvNeXt module is mainly stacked by ConvNeXt blocks, and the number ratio of each layer of ConvNeXt blocks is 3:3:9:3 respectively. The structure design of the ConvNeXt block includes: the first layer of convolution: Conv2d(kernel_size = 7, stride = 3, padding = 3)+LayerNorm; the second layer of convolution: Conv2d(kernel_size = 1, stride = 1)+GELU; the third layer of convolution: Conv2d(kernel_size = 1, stride = 1)+DropPath; finally, global average pooling and linear transformation are used, and the output is 64*66. This structure can extract the spatial features in the ECG sequence.
[0045] TCN uses an improved CNN model to extract features in the ECG. The specific structure is as follows:
[0046] The S3-2 Transformer module is used to further extract the temporal features in the ECG. The specific structure is as follows: Temporal attention module: temporal_attention(d_feature,n_channel); Temporal convolutional layer: Conv1d(in_channel = 1, out_channel = 1, kernel_size = 3, stride = 1, dilation = 1, padding = 2) + ReLU + dropout; Residual layer: The data before and after convolution are combined. The input format of the TCN module is (Batch, input_channel, seq_len), the input size is 64 * 1 * 260, the output format is (Batch, output_channel, seq_len), and the output is 64 * 1 * 260. Before and after the input and output, the sequence length remains unchanged. This structure can initially extract the local temporal features in the ECG sequence. Generate a 261-dimensional positional encoding and combine it with the ECG data. Encoder layer: After passing through the Encoder encoder layer, the output size is (64 * 64). The Encoder layer is bidirectional and can consider the information before and after a certain point in the ECG signal at the same time. This feature enables the Encoder layer to more comprehensively understand the characteristics of the ECG signal. The Encoder can capture the long-term dependencies in the time series data and further extract the global temporal features in the ECG signal. This enables the model to have good performance when processing time series data.
[0047] S3-3 Feature fusion. The specific structure is as follows: Concatenation layer: Linear(64, 132); Linear transformation layer: Linear(132, 5). The input of this module includes the spatial features extracted by the ConvNext module and the temporal features extracted by the Transformer module. After passing through the concatenation layer, they are combined and then linearly transformed to generate the classification result. The maximum classification value is selected as the classification result.
[0048] S4: Model training
[0049] The model is trained using the MIT-BIH Arrhythmia Database and the PTB Diagnostic Database, which contain various types of arrhythmias and myocardial infarctions, providing rich data for the model to learn and generalize. During the training process, the focal loss function is used to optimize the model, and the Adam optimizer is used to adjust the network weights to minimize the prediction error.
[0050] The trained multimodal MFCT model can classify arrhythmia types and myocardial infarctions in ECG data in real time, such as normal beats (N), supraventricular ectopic beats (F), ventricular ectopic beats (V), fusion beats (F), and unknown beats (Q). In addition, the real-time and simplicity of the model make it suitable for integration into clinical monitoring systems or portable devices to provide continuous cardiac health monitoring for patients.
[0051] As shown in Figure (2), a multimodal MFCT model is proposed: it is mainly composed of ConvNeXt blocks in the ConvNeXt module. The recursive image is downsampled through the convolutional layer and input into the ConvNeXt block. The number ratio of each layer of ConvNeXt blocks is 3:3:9:3. The ConvNeXt block adopts a convolutional structure with convolutional kernel sizes of 7, 1, and 1 respectively. Finally, high-dimensional spatial features are output through operations such as pooling and flattening. In the Transformer module, an improved TCN module is introduced. The special convolutional structure of the TCN model can form a large receptive field and can initially capture and extract local temporal features in the ECG signal. The Transformer module uses fewer encoder layers. Through position vector encoding, it ensures that the position features in the ECG signal are not lost. At the same time, the Encoder layer contains a multi-head self-attention mechanism, residual connection, and feed-forward network. Therefore, the Encoder layer has bidirectionality and can consider the information before and after a certain point in the ECG signal at the same time. This characteristic enables the Encoder layer to more comprehensively understand the features of the ECG signal, capture long-term dependencies in time series data, and further extract global temporal features in the ECG signal by combining TCN and Transformer. In the feature fusion module, the obtained features are concatenated and linearly transformed to complete the recognition of the ECG signal.
[0052] As shown in Figure (3), for the detailed design of the ConvNeXt module, the specific structure is as follows: The ConvNeXt module is mainly composed of ConvNeXt blocks stacked together, and the number ratio of each layer of ConvNeXt blocks is 3:3:9:3. The ConvNeXt block structure design includes: the first layer of convolution: Conv2d(kernel_size = 7, stride = 3, padding = 3)+LayerNorm; the second layer of convolution: Conv2d(kernel_size = 1, stride = 1)+GELU; the third layer of convolution: Conv2d(kernel_size = 1, stride = 1)+DropPath; finally, global average pooling and linear transformation are used, and the output is 64*66. This structure can extract spatial features in the ECG sequence.
[0053] As shown in Figure (4), the detailed design of the TCN module is as follows: Each layer of the TCN module includes a temporal attention module, a dilated convolutional layer, and two residual connections. Each layer maintains the same structure, and two layers are stacked. Through its special convolutional structure in temporal convolution, local temporal features in the electrocardiogram (ECG) signal are initially captured and extracted.
[0054] As shown in Figure (5), the detailed design of the Transformer encoder module is as follows: Each layer of the Encoder module contains a multi-head self-attention mechanism, a residual connection, and a feed-forward network. Three layers are stacked to further extract global temporal features in the ECG signal.
[0055] As shown in Figure (6), it is the accuracy change curve during the training process. It can be observed that the model performance tends to be optimal, and the optimal model parameters are obtained at the 100th epoch.
[0056] The main process of the multi-modal MFCT algorithm is described as follows.
[0057]
[0058]
[0059] Further technical details
[0060] In the technical implementation, special attention is paid to the computational efficiency and practicality of the model. For example, by introducing Layer Normalization and Residual Connections, not only the convergence speed during training is accelerated, but also the stability of the model in the network structure is ensured, avoiding the problem of gradient vanishing.
[0061] Formula example: The activation function can construct the non-linear relationship between the input and output of neurons, enabling the network model to learn more complex features. The ConvNeXt network introduces the Gaussian error linear unit activation function (Gaussian error linear units, GELU), where X is the input of the activation layer.
[0062]
[0063] The calculation formula of the self-attention mechanism is expressed as:
[0064] Q = H i W Q , K = H i W K , V = H i W V
[0065] Among them, H kis the input representation of the i-th layer, Q: query matrix, obtained by multiplying the input representation H k by the query weight matrix W Q to get, K: key matrix, obtained by multiplying the input representation H k by the key weight matrix W K to get, WQ, WK, WV: weight matrices for query, key, and value.
[0066]
[0067] where, QK T : the transpose of the query matrix Q multiplied by the key matrix K to obtain the attention scores. d k : the dimension of the key vector, usually a scaling factor to prevent the dot product value from being too large. softmax: applies the softmax function to normalize the attention scores. Attention(Q, K, V): obtains the weighted value matrix.
[0068] Experiments and Results
[0069] This paper evaluates the prediction performance through accuracy, precision, recall, and F1-score metrics.
[0070] Accuracy: refers to the proportion of the number of samples correctly predicted by the model to the total number of samples.
[0071]
[0072] where, TP represents True Positives, TN represents True Negatives, FP represents False Positives, and FN represents False Negatives.
[0073] Precision: refers to the proportion of samples actually being positive among the samples predicted as positive by the model.
[0074]
[0075] Recall: refers to the proportion of samples actually being positive that are predicted as positive by the model. Focuses on the coverage ability of the model, that is, the recognition degree of positive examples.
[0076]
[0077] F1-Score: is the harmonic mean of precision and recall, comprehensively considering the accuracy and recall of the model.
[0078]
[0079] Specificity (Spe): The ratio of the heartbeat data classified as negative to the actual negative heartbeat data, representing the classification ability of the model for negative heartbeats.
[0080]
[0081] Table 1. Five-class classification and recursive image correspondence table for the MIT-BIH database in Experiment 1.
[0082] The MIT-BIH Arrhythmia Database contains 47 test individuals, and a total of 107,319 heartbeats are classified into five categories: normal heartbeats (N), supraventricular ectopic heartbeats (F), ventricular ectopic heartbeats (V), fusion heartbeats (F), and unknown heartbeats (Q) according to the ANSI / AAMI EC57:2012 standard. The PTB Diagnostic ECG Database is a publicly available electrocardiogram recording dataset containing electrocardiogram recordings from 290 patients, acquired using a 12-lead system. Among them, 148 were diagnosed with MI, 52 were healthy controls, and the rest were diagnosed with 7 different diseases. In our experiment, we used lead II ECG recordings and classified them into two categories: normal (N) and abnormal (M), and each recording had a diagnosis label from a cardiologist. Table 1 details the number of samples in each category of these two datasets.
[0083] The entire dataset is divided into a training set, a validation set, and a test set. The training set is used for learning model parameters, the validation set is used for adjusting hyperparameters and preventing overfitting, and the test set is used for final performance evaluation. 70% of the data is randomly selected as the training set, 20% as the validation set, and 10% as the test set, as shown in Table 1.
[0084] Table 1. Experimental data for the MIT-BIH Arrhythmia Database set and the PTB Diagnostic Dataset
[0085]
[0086] Experiment 2: A multi-modal MFCT was established for experiments to classify various types of arrhythmias and myocardial infarctions. All ECG recordings in the MIT-BIH dataset and the PTB Diagnostic dataset were preprocessed for data denoising and type annotations were added after preprocessing. A series of experiments were conducted on the proposed multi-modal MFCT using the above datasets. The experimental results are shown in Table 2.
[0087] Table 2. Experimental results of electrocardiogram signal recognition
[0088]
[0089] Comparison with existing neural network technologies: The advantage of the multimodal MFCT model is that it fuses one-dimensional time series signals and two-dimensional images. By integrating information from different modalities, the algorithm can compensate for the limitations of one-dimensional time series signals and two-dimensional images respectively, thereby extracting richer and more comprehensive features.
[0090] To further evaluate the effectiveness of the proposed method, ablation experiments were conducted on the two feature branches of multimodal fusion, and the evaluation was carried out using the evaluation metrics proposed above. The experimental results are shown in Table 3.
[0091] Table 3 Results of ablation experiments
[0092]
[0093] The above experiments fully demonstrate the effectiveness and practicality of the method of the present invention. Compared with using single modality, the multimodal fusion of the model improves the performance of the electrocardiogram signal classification task. In the future, we can use this model on more datasets to classify more heart cases. In addition, the implementation of this technology is not limited to electrocardiogram analysis, but can also be applied to other types of signals, such as brain signals, and observe its performance and its impact on these signals. It shows significant advantages especially in the real-time analysis of electrocardiograms and the accurate classification of arrhythmias.
Claims
1. An ECG intelligent recognition algorithm for an MFCT model based on multimodal fusion, characterized in that Including the following steps: Step S1: Data preprocessing: The ECG data is preprocessed using a band-pass filter and five-point smoothing filter to reduce the interference of noise on the analysis of the electrocardiogram signal. When performing the preprocessing operation, first, the five-point smoothing filter is used to remove the high-frequency noise in the electrocardiogram signal, then the digital filter IIR is used to remove the power frequency interference of the electrocardiogram signal, and finally, the Butterworth band-pass filter is used to remove the baseline drift in the electrocardiogram signal to improve the quality of the electrocardiogram signal; Step S2: Two-dimensional map conversion: The one-dimensional electrocardiogram signal is converted into a two-dimensional image with a size of 224x224 pixels using the recurrence plot (RP) and short-time Fourier transform (STFT). This conversion not only retains the correlation information between the amplitude and phase in the original signal and the correlation information between the time domain and the frequency domain, but also significantly improves the richness of the electrocardiogram spatial feature information, thus realizing the enhanced display of electrocardiogram features. In this way, the spatial features and frequency domain features of the electrocardiogram signal are more intuitively displayed, providing richer visual and data support for subsequent analysis and diagnosis; Step S3: Intelligent recognition of electrocardiogram signal: The data preprocessed in Step S1 is processed using the MFCT model. When executing the MFCT model, first, the original one-dimensional (1D) ECG data is converted into two-dimensional (2D) ECG images through Step S2 as the input of the ConvNeXt module to capture the spatial features of the ECG signal. The ConvNeXt module is mainly stacked by ConvNeXt blocks. The recurrence image is downsampled through the convolutional layer and input into the ConvNeXt block. The number ratio of each layer of ConvNeXt blocks is 3:3:9:
3. The ConvNeXt block adopts a convolutional structure with convolutional kernel sizes of 7, 1, and 1 respectively. Finally, high-dimensional spatial features are output through operations such as pooling and flattening. The ConvNeXt block is mainly composed of a common convolutional layer, a residual module, a downsampling layer, a normalization layer, a pooling layer, and an activation function. In the design of the Transformer module, in order to better fuse local features and global features, a TCN module is added before the Transformer network and some improvements are made. Each layer of the TCN module includes a temporal attention module, a dilated convolutional layer, and two residual connections. Each layer maintains the same structure and is stacked in two layers. Through its special convolutional structure in the temporal convolution, the local temporal features in the electrocardiogram signal are initially captured and extracted; Each layer of the Encoder module includes a multi-head self-attention mechanism, a residual connection, and a feed-forward network, and is stacked in three layers to further extract the global temporal features in the electrocardiogram signal. Finally, high-dimensional temporal features are output through a linear transformation. In the feature fusion stage, the high-dimensional temporal features and high-dimensional spatial features are combined through a multi-modal feature fusion module to realize the recognition of the electrocardiogram signal; Step S4: Model training: Based on the model architecture constructed in Step S3, the focal loss function is used to quantify the classification error, and the Adam optimizer is used to dynamically adjust the learning rate, thereby accelerating the convergence of the model and improving the accuracy.
2. The ECG intelligent recognition algorithm according to claim 1, characterized in that, The band-pass filter and the five-point smoothing filtering algorithm are used to denoise the ECG data, removing high-frequency noise of 50 - 1000 Hz, low-frequency noise of 0.05 - 1 Hz, and 50 Hz power frequency interference, so as to reduce the interference of noise on the classification result and restore the original data of the electrocardiogram signal.
3. The ECG intelligent recognition algorithm according to claim 1, wherein The one-dimensional electrocardiogram signal is converted into a two-dimensional image using recurrence plots and short-time Fourier transforms, by preserving the corresponding information between the amplitude and phase in the original one-dimensional signal, as well as the frequency domain information of the electrocardiogram signal.
4. The ECG intelligent recognition algorithm according to claim 1, wherein The multi-modal model fuses the one-dimensional time series signal and the two-dimensional image. By integrating information from different modalities, the algorithm can make up for the limitations of the one-dimensional time series signal and the two-dimensional image respectively, combine them, and thus extract the features of the data more comprehensively.
5. The ECG intelligent recognition algorithm according to claim 1, wherein The multi-modal model uses the ConvNext network through multi-layer convolutional operations, which can efficiently extract multi-scale spatial features in the image and improve the classification performance of the model.
6. The ECG intelligent recognition algorithm according to claim 1, characterized in that, The multi-modal model proposes to combine TCN and Transformer to enhance the ability to extract temporal features and improve the classification accuracy. Among them, the TCN extracts local features of the electrocardiogram signal through multi-layer temporal convolutional operations, and the Transformer captures the global dependencies between features through the self-attention mechanism.
7. The ECG intelligent recognition algorithm according to claim 1, wherein The multi-modal model also uses the focal loss function and combines the Dropout technique to prevent overfitting and improve the classification performance of the minority classes of the model.
8. The ECG intelligent recognition algorithm according to claim 1, characterized in that, During the model training process, the model learning rate is set to 0.0001, the batch processing dimension is 64, and the focal loss function and the Adam optimizer are used to optimize the model.
9. The ECG intelligent recognition algorithm according to claim 1, wherein The multi-modal model uses an early stopping strategy during the training process. During the training stage, the performance of the model is continuously monitored through the validation set to prevent overfitting. And when the performance on the validation set no longer improves after reaching the peak, the training process is terminated in a timely manner.
10. The ECG intelligent recognition algorithm according to claim 1, wherein, The multi-modal model is verified on two electrocardiogram signal datasets, the MIT-BIH Arrhythmia Database and the PTB Diagnostic Database, and the prediction performance is evaluated through metrics such as accuracy, precision, recall, and F1-score.