Single-lead electrocardiogram segmentation model based on deep learning, construction method and application
Through the deep learning single-lead ECG segmentation model, the global correlation features and multi-layer information fusion are used to solve the problems of information loss and feature forgetting in ECG segmentation, and the high-accuracy medical segmentation effect is achieved.
Patent Information
- Application Number
- CN202210906280.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-29
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-07-29
AI Technical Summary
The existing U-net model based on convolutional neural networks cannot effectively utilize the correlation information between multiple continuous waveforms and cardiac beats in the ECG data, resulting in information loss and feature forgetting in the ECG segmentation task, affecting the segmentation accuracy.
A single-lead electrocardiogram segmentation model based on deep learning is adopted, global correlation features are extracted through the coding layer, and multi-layer information fusion is carried out in the middle layer, and the attention neural network (Transformer) and upsampling module are used to solve the problems of information loss and feature forgetting, and combined with diceloss loss function optimization model training.
The high-accuracy medical segmentation of single-lead electrocardiogram data is achieved, which solves the problems of information loss and feature forgetting in traditional models, and improves the segmentation accuracy.
Smart Images

Figure CN115439487B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer science and technology, and in particular to a method for performing a medical segmentation task on single-lead electrocardiogram data. Background Art
[0002] Cardiovascular diseases are one of the leading causes of death in the world, and the impact of cardiovascular diseases is becoming increasingly obvious.
[0003] The detection of physiological signals is a major means of preventing and diagnosing cardiovascular diseases. Among them, electrocardiogram (ECG) is a commonly used non-invasive and easily detectable physiological signal, which records the physiological activities of the heart over a period of time. Electrocardiogram signals carry important information about many cardiovascular diseases. Currently, there are two common types: static electrocardiogram and ambulatory electrocardiogram. Ambulatory electrocardiogram is to continuously record the whole process of a patient's electrocardiac activity for 24 hours or longer under the daily living state by an ambulatory electrocardiograph, and analyze and process it with the help of a computer. By processing and analyzing the data collected by the device, doctors can obtain the medical indicators of the patient's electrocardiac activity during this period, so as to diagnose the causes and nature of the patient's paroxysmal syncope, dizziness and palpitation, evaluate the patient's heart condition, and evaluate the efficacy of anti-arrhythmic drugs. Static electrocardiogram is commonly used in the clinical use scenarios of hospitals, generally recording the electrocardiac activity of a patient from the perspectives of multiple leads for ten seconds in a static state. The static electrocardiogram is convenient and fast to operate, but the recording time is short, and it may not be able to detect arrhythmia or myocardial ischemia changes during the patient's illness. The ambulatory electrocardiogram records the electrocardiac changes of the patient for 24 hours, and the patient's daily work and activities are not restricted. However, the ambulatory electrocardiogram Figure 1 generally only includes data of fewer leads, which poses a challenge to data analysis.
[0004] The purpose of the electrocardiogram segmentation task is to divide waveform regions such as P waves, QRS waves, and T waves in a segment of electrocardiogram signal, so as to perform further disease diagnosis and classification. For example, in the cardiac beats of premature atrial contractions, the P wave almost disappears and may occur in advance and overlap with the previous T wave. Then, in the result of the electrocardiogram segmentation task, those cardiac beats without a detected P wave can be determined as premature atrial contractions. The meanings of each wave band of the electrocardiogram are shown in Table 1, and the abnormal wave bands of common arrhythmias are shown in Table 2.
[0005] Table 1 Electrocardiogram Wave Bands and Their Meanings
[0006]
[0007] Table 2 Arrhythmias and Their Corresponding Characteristics
[0008]
[0009] In the electrocardiogram (ECG) segmentation task, some studies have transferred the U-net model in the field of medical image segmentation to the ECG segmentation task. However, an ECG usually contains several heartbeats, and each heartbeat includes multiple waves. The correlation information between these heartbeats and waves is very helpful for the ECG segmentation task. For example, there are some individual differences in the morphology of ECGs sampled by different patients and different instruments. Referring to the morphological characteristics of other heartbeats in the same data is more conducive to segmenting the waves. Another example is that in the ECG of premature ventricular contractions, the QRS waveband will show wide deformity, and then secondary ST-T changes will occur. Due to the compensatory pause, the T wave will appear flat or inverted. When a disease occurs, multiple wavebands often change accordingly. In paroxysmal arrhythmias, the waveforms of consecutive heartbeats often change, and this correlation between heartbeats also needs to be considered. However, the U-net model based on convolutional neural network cannot utilize this information. Summary of the Invention
[0010] The object of the present invention is to provide a single-lead ECG segmentation model based on deep learning that utilizes global features and multi-layer information fusion, as well as a construction method and application thereof. This ECG segmentation model can extract and utilize the global correlation features of the data at the data level, thus being more in line with the characteristics of the ECG data itself; at the model structure level, it can achieve information fusion between different levels, thereby solving the problems of information loss and feature forgetting caused by the deepening of the model layers. The proposed model construction and application can achieve high-accuracy medical segmentation of single-lead ECG data.
[0011] The specific technical solution to achieve the object of the present invention is as follows:
[0012] A method for constructing a single-lead ECG segmentation model based on deep learning. The model structure constructed by this method is an encoding layer, an intermediate layer, and a segmentation layer. The encoding layer extracts global correlation features from the serialized data, passes the output of the last layer to the segmentation layer, and passes the outputs of multiple intermediate layers to the intermediate layer; the intermediate layer upsamples the multiple outputs of the encoding layer to different dimensions to match the segmentation layer; the segmentation layer receives the output of the last layer of the encoding layer, performs information fusion with the multiple outputs of the intermediate layer, and upsamples step by step, and finally completes the segmentation. The specific construction process includes:
[0013] Let D be the input ECG data with a length of l; X is the set of serialized input ECGs, X = {x1, x2,..., x n}, where each element x n has a dimension of d, and d * n = l;
[0014] The encoding layer is a network capable of extracting global features, and its purpose is to extract the global correlation features of the serialized input.
[0015] The encoding layer is a network capable of extracting global features, such as recurrent neural network (RNN), long short-term memory neural network (LSTM), deep attention model (Transformer), etc. Its purpose is to extract the global correlation features of the serialized input;
[0016] The encoding layer contains multiple layers of neural networks to encode the input data sequence, at least 5 layers; and the output of p layers in the multiple layers of neural networks will be saved for multi-layer information fusion; p is at least 5;
[0017] For the input data sequence X, each layer of neural network in the encoding layer extracts an output feature vector E i , which contains global feature information, E i = {e1, e2,..., e n}, where the dimension of each element e n is d′, d′ < d;
[0018] The encoding layer provides fusion information for multi-layer information fusion, and transmits the output E l of the last layer of neural network in the encoding layer to the segmentation layer, and transmits the outputs {E1, E2,..., E s} of the first layer of neural network and the intermediate layers to the intermediate layer;
[0019] The intermediate layer is composed of upsampling modules, and its number of layers is p, which is equal to the number of layers of the multi-layer output for information fusion from the encoding layer;
[0020] The intermediate layer receives the feature vector E = {e1, e2,..., e n}, whose length is n and dimension is d′; every time it passes through an upsampling module, the length of the feature vector will increase and the dimension will decrease;
[0021] Suppose there are s feature vectors, {E1, E2,..., E s}, then E1 is transmitted to the segmentation layer after passing through the first upsampling module, E2 is transmitted to the segmentation layer after passing through the first and second upsampling modules, and so on. The last feature vector E s needs to pass through all the upsampling layers;
[0022] The s feature vectors are upsampled into vectors {M1, M2,..., M s} with different lengths and dimensions, where M1 has the shortest length and the largest dimension, and M s has the longest length and is equal to the length l of the original data input and the smallest dimension;
[0023] The number of neural network layers in the segmentation layer is equal to the number of layers in the intermediate layer. Each layer includes an upsampling module and a fusion module. Among them, the output shape of the upsampling module is the same as that of the upsampling module of the intermediate layer in the same layer.
[0024] The segmentation layer receives the output E from the last layer of the encoding layer l and the s-layer outputs {M1, M2, ..., M s} from the intermediate layer; E l obtains M' after passing through the first upsampling module s , and M' s performs information fusion with M s in the fusion module to obtain M'; s-1 , and M' s-1 passes through upsampling again and fuses with M s-1 and so on;
[0025] The output of the last layer of the segmentation layer is M0 obtained by fusing M'1 and M1, with a length of l, the same as the input data, and a dimension of 4. The four dimensions respectively correspond to the segmentation results of four electrocardiogram waveforms.
[0026] In the training stage, only the data in the training set is used for training, and the data in the validation set is used to measure the convergence of each round of training.
[0027] The data in the training set and the validation set are sent into the model in data blocks of every b data. The shape of the block is b*l, where l is the length of a single electrocardiogram signal; the shape of the label is b*l*4, where 4 corresponds to the number of bands in the segmentation result.
[0028] The data block is sent into the model as input, and the segmentation layer outputs the outputs of b pieces of data at the same time, with a shape the same as the label, which is b*l*4.
[0029] The loss is calculated between the segmentation result and the label. The loss function uses the dice loss function, and its calculation formula is:
[0030]
[0031] where X represents the segmentation result and Y represents the label; the loss function can solve the problems of under-segmentation and over-segmentation at the same time.
[0032] For the loss value on the training set, the parameters of each layer are updated according to the chain rule of differentiation and the gradient descent method.
[0033] In the training stage, the Adam optimizer is used, and the learning strategy uses ReduceLROnPlateau; the learning strategy monitors the loss value on the validation set in each round and updates the learning rate according to the loss value on the validation set; thus obtaining the single-lead electrocardiogram segmentation model based on deep learning.
[0034] A single-lead electrocardiogram segmentation model based on deep learning constructed by the above method.
[0035] A method for electrocardiogram segmentation using the constructed model, which includes the following specific steps:
[0036] Step 1: Data preprocessing. For an input single-lead electrocardiogram with a length of l, it is divided into several blocks according to a window of size d, forming an electrocardiogram data sequence with a length of n and a dimension of d, where n * d = l;
[0037] Step 2: Inference stage, which is divided into the following three sub-steps:
[0038] Step 2-1: Divide the single-lead electrocardiogram into several data with the same length as in the training stage, send each piece of data into the model, and obtain a preliminary segmentation result;
[0039] Step 2-2: Binarize the segmentation result according to the value of the preliminary segmentation result and perform local mode filtering to eliminate noise and obtain the final segmentation result.
[0040] Step 2-3: Statistically organize the pixel-level segmentation result to obtain the band information of the entire single-lead electrocardiogram.
[0041] Advantages of the present invention: The present invention proposes an electrocardiogram medical segmentation model and its construction method and application based on deep learning using global feature information and multi-layer information fusion. The present invention can utilize the correlation information between multiple continuous waveforms and heartbeats in electrocardiogram data, so as to better segment each waveform pixel point. At the same time, in the model construction method, the present invention can fuse the data between different levels of the model, thus solving the problem that some initial features are lost and forgotten in the gradual change as the model deepens and the data changes continuously. The present invention can show the highest accuracy in the electrocardiogram medical segmentation task. Brief Description of the Drawings
[0042] Figure 1 It is a schematic diagram of the model structure constructed by the present invention;
[0043] Figure 2 It is a schematic diagram of the middle layer and segmentation layer of the model constructed by the present invention;
[0044] Figure 3 It is a schematic diagram of the coding neural network used in the present invention;
[0045] Figure 4 It is a schematic diagram of the input data of the present invention;
[0046] Figure 5Effect diagram of the present invention;
[0047] Figure 6 Statistical chart of experimental results when using Dataset A;
[0048] Figure 7 Statistical chart of experimental results when using Dataset B. Specific implementation manner
[0049] The model construction method proposed by the present invention focuses on the extraction of global correlation features and the fusion of multi-layer information.
[0050] Figure 1 It is an example diagram of the model constructed by the model construction method of the present invention. The model includes three parts: an encoding layer, an intermediate layer, and a segmentation layer;
[0051] Global correlation features refer to the front-back correlation information of the data itself, rather than focusing on the numerical features of a certain local part of the data;
[0052] Global correlation features have shown superior performance in the field of natural language processing through the attention mechanism. However, in the fields of images and electrocardiogram signals, because there is no concept of "words" in natural language, they have not been applied for a long time.
[0053] With the emergence of the method of serializing images and signals to form "words", global correlation features have been introduced into image processing and signal processing, and have also shown good performance; in the electrocardiogram segmentation task, there has still been no related invention; the present invention designs a method for serializing electrocardiogram signals to extract the global correlation features therein.
[0054] Multi-layer information means that the output of the neural networks of different layers of the model expresses features of different depths. As the model deepens, some shallow features of the data will be "forgotten" by the model. In the medical segmentation task, the final output is a segmentation map with the same size and length as the original data. Therefore, retaining the shallow features and fusing them with the deep features can well solve the forgetting problem.
[0055] The extraction of global correlation features is responsible for the encoding layer, which is stacked by multiple layers of neural networks. In order to extract global correlation features, the data should be serialized before encoding feature extraction. Because the electrocardiogram signal is the same as the image, the model cannot calculate the correlation between each sampling point / pixel point. Therefore, the data needs to be serialized, and the encoding layer is used to calculate the correlation between each block in the data sequence;
[0056] The selection of the neural network of the encoding layer is also very important. The model needs to select a neural network for processing sequence data, such as a recurrent neural network (RNN), a long short-term memory network (LSTM), and an attention network (Transformer);
[0057] Recurrent neural networks establish a "long-term" dependence on data. That is to say, the output of the unit at the end of the sequence depends on both the characteristics of the unit itself and the characteristics of the previous units. However, sequences sometimes require this dependence and sometimes do not, which brings a bottleneck to recurrent neural networks, known as the "long-term dependence problem". Long short-term memory neural networks add a forgetting "gate", which selectively forgets the characteristics of other units when calculating the characteristics of a unit;
[0058] The emergence of the attention neural network (Transformer) has become the optimal choice for sequence data processing. It adds an attention mechanism to calculate the "attention" between each sequence unit, that is, the influence value, such as Figure 3 . The generation of this influence value depends on the value of the unit itself, the values of other units, and the numerical connection between these units. For example, if the sequence where the QRS complex is located in the same piece of data always appears periodically, then this periodicity can be captured by the attention mechanism as a connection required for the segmentation task;
[0059] Therefore, the attention neural network (Transformer) is the best choice for the encoding layer;
[0060] Multi-layer information fusion is achieved by the architecture of the intermediate layer and the segmentation layer.
[0061] The final output of electrocardiogram segmentation is four segmentation signals with the same length as the input electrocardiogram. Therefore, in the segmentation layer, the segmentation layer needs to perform upsampling on the output of the encoding layer step by step. The output E of the encoding layer of the present invention = {e1, e2,..., e n}, its length is n, the dimension is d′, n is less than the length l of the signal, and d′ is greater than the final required output dimension 4. The final result is to upsample the length of the feature E to l and reduce its dimension to 4.
[0062] For upsampling one-dimensional signals, either a one-dimensional transposed convolution network or a fully connected network can be used.
[0063] In the segmentation layer, not only the features need to be upsampled, but also information from different levels of the encoding layer needs to be received for fusion. Since the segmentation layer performs upsampling step by step and the shape of the data is different for each layer, the intermediate layer is required to perform upsampling on the multi-layer information of the encoding layer at the same scale to match the shape of the data to be fused.
[0064] The multi-layer information is upsampled different times in the intermediate layer. The information from the shallowest layer of the encoding layer needs to be upsampled the most times. The information of different layers shares the weights of all upsampling layers, which can accelerate the training and convergence of the model.
[0065] The changes in the data dimensions in the intermediate layer and the encoding layer are as follows Figure 2 shown. The output from the layer above the segmentation layer is fused with the output of the intermediate layer at the same level. Fusion means directly concatenating the data first, which will double the dimension of the data. Therefore, when constructing the model, the present invention designs a CBR module to reduce the dimension of the data so as to achieve fusion;
[0066] In the CBR module, a convolutional neural network or a fully connected network can be used to achieve fusion. In addition, there are a regularization module and a Relu activation function, which can accelerate the convergence of the model and improve the performance of the model;
[0067] The final loss function uses dice loss, and its calculation formula is:
[0068]
[0069] In actual calculation, the calculation of the modulus of the intersection of the output matrix and the label matrix directly uses the calculation method of matrix dot product, such as:
[0070] DiceLoss can take into account the problems of over-segmentation and under-segmentation. If there is a problem of under-segmentation, the numerator of the negative term of DiceLoss will decrease and the loss will increase; if there is a problem of over-segmentation, the denominator of the negative term of DiceLoss will increase and the loss will also increase.
[0071] When applying the model, in the inference stage, the maximum value of each column of the output matrix is assigned 1, and the rest of the values are 0, which means taking the band that the model thinks the sampling point most likely belongs to as the segmentation result, as shown in the following formula:
[0072]
[0073] When applying, there may be noise perturbations in the inference results. For a continuous electrocardiogram signal, its band segmentation results should also be continuous. For example, if there is [1 1 11 1 01 1 1]
[0075] in the segmented P-wave vector, then it can be considered that the 0 sampling points appearing in a continuous P-wave are noise. When applying, the method of taking the local mode of each sampling point as the value of the sampling point is used to filter out this kind of noise, and the filtered effect is: [1 1 11 1 11 1 1]
[0077] The final segmentation effect diagram is as follows Figure 5 shown. The boxes with different heights in the figure are the regions of each segmented band.
[0078] Embodiment
[0079] Figure 4 It is an electrocardiogram input when the model constructed by the present invention is applied. This electrocardiogram contains four heartbeats, and each heartbeat has a P wave, a QRS wave, and a T wave respectively;
[0080] Figure 5 It is to Figure 4 The segmentation result obtained by inputting the shown electrocardiogram into the model constructed by the present invention;
[0081] There are four types of segmentation results of the model, which respectively represent the P wave region, the QRS wave region, the T wave region, and the background. The background refers to the region in the electrocardiogram signal that is not a wave band. Figure 5 The rectangular frames in it respectively represent the segmentation regions of the three wave bands of the P wave, the QRS wave, and the T wave;
[0082] Figure 5 In, the rectangular frames with different heights represent different segmentation results. The rectangular frame with the lowest height represents the P wave region, the highest rectangular frame represents the QRS wave region, the rectangular frame with medium height is the T wave region, and the electrocardiogram signal without a rectangular frame is the background signal.
[0083] Comparative example
[0084] Some data from publicly available electrocardiogram datasets on the Internet are used to experiment with the model construction and application methods proposed by the present invention to demonstrate its effects.
[0085] Experimental dataset
[0086] There are two datasets used for the electrocardiogram medical segmentation task, both of which are publicly available electrocardiogram datasets on the Internet. The quantity, length, and sampling rate of their valid data are shown in Table 3.
[0087] Table 3 Experimental dataset
[0088]
[0089] The electrocardiogram data in Datasets A and B come from different abnormal disease patients of different ages and genders. In Dataset A, the original data does not provide electrocardiogram segmentation wave band labels. Later, in the paper "Electrocardio Panorama: Synthesizing New ECG Views with Self - supervision", the wave bands of 1000 pieces of data among them were manually labeled.
[0090] Valid data refers to the data with the position labels of the P wave, QRS wave, and T wave, and its data length is not less than 10 seconds. During training and inference, the model resamples the data uniformly to 240Hz.
[0091] The following accuracy experiment is a comparative experiment between the model Utrans of the present invention and the medical image segmentation Unet model using the same data set. The index of the experiment is the F1 value of waveform division, which is defined as
[0092] F1 = Precision * Recall * 2 / (Precision + Recall)
[0093] Precision = Number of correctly divided pixel points / Total number of divided pixel points
[0094] Recall = Number of correctly divided pixel points / Number of pixel point sets in the label
[0095] Figure 6 、 7 shows the experimental results of the model Utrans constructed by the construction method proposed by the present invention, Figure 6 are the experimental results of the Chinese Cardiovascular Database, Figure 7 are the experimental results of the "Gaoxin Cup" Human-Machine Intelligence Competition data set. The model proposed by the present invention can show higher accuracy on two data sets than the Unet model, which is currently superior in medical image segmentation, with an average increase of 2-3 percentage points.
Claims
1. A method for constructing a single - lead electrocardiogram segmentation model based on deep learning, characterized in that, The model structure constructed by this method is an encoding layer, an intermediate layer, and a segmentation layer. The encoding layer extracts global correlation features from the serialized data, passes the output of the last layer to the segmentation layer, and passes the outputs of multiple intermediate layers to the intermediate layer. The intermediate layer upsamples the multiple outputs of the encoding layer to different dimensions to match the segmentation layer. The segmentation layer receives the output of the last layer of the encoding layer, performs information fusion with the multiple outputs of the intermediate layer, and upsamples step by step to finally complete the segmentation. The specific construction process includes: Let D be the input electrocardiogram data with length l; X is the set of input serialized electrocardiograms, X = {x1, x2,..., x n}, where each element x n has dimension d, and d * n = l; The encoding layer is a network capable of extracting global features, and its purpose is to extract the global correlation features of the serialized input. The encoding layer contains multiple neural networks to encode the input data sequence, at least 5 layers; and the outputs of p layers in the multiple neural networks will be saved for multi-layer information fusion; p is at least 5. For the input data sequence X, each layer of neural network in the encoding layer extracts an output feature vector E i , which contains global feature information, E i = {e1, e2,..., e n}, where the dimension of each element e n is d′, and d′ < d; Transfer the output E of the last neural network layer in the encoding layer l to the segmentation layer, and transfer the outputs {E1, E2, ..., E s} of the first neural network layer and the intermediate layer to the intermediate layer; The intermediate layer consists of upsampling modules, and its number of layers is p. The feature vector E received by the intermediate layer = {e1, e2,..., e n}, its length is n and the dimension is d'; every time an upsampling module is passed through, the length of the feature vector will increase and the dimension will decrease; Suppose there are s eigenvectors in total, {E1, E2,..., E s}, then E1 is passed to the segmentation layer through the first-layer upsampling module, E2 is passed to the segmentation layer through the upsampling modules of the first and second layers, and so on. The last eigenvector E s needs to pass through all the upsampling layers; s feature vectors are upsampled into vectors {M1, M2, ..., M s} with different lengths and dimensions, where M1 has the shortest length and the largest dimension, and M s has the longest length, which is equal to the length l of the original data input, and the smallest dimension; The number of neural network layers in the segmentation layer is equal to the number of layers in the intermediate layer. Each layer includes an upsampling module and a fusion module. Among them, the output shape of the upsampling module is the same as that of the upsampling module in the intermediate layer of the same layer. The segmentation layer receives the output E from the last layer of the encoding layer l and the output {M1, M2,..., M s} of the s layer in the intermediate layer; E l obtains M' after passing through the first upsampling module s , and M' s is fused with M s in the fusion module to obtain M' s-1 , and M' s-1 is upsampled again and fused with M s-1 and so on; The output of the last layer of the segmentation layer is M0 obtained by fusing M′1 and M1, its length is l, which is the same as the input data, and the dimension is 4. The four dimensions respectively correspond to the segmentation results of four electrocardiogram waveforms. In the training stage, only the data in the training set is used for training, and the data in the validation set is used to measure the convergence of each round of training. The data in the training set and the validation set are sent into the model as a data block with every b data. The shape of the block is b*l, where l is the length of a single electrocardiogram signal. The shape of the label is b*l*4, where 4 corresponds to the number of bands in the segmentation result. The data block is sent into the model as the input, and the segmentation layer simultaneously outputs the outputs of b pieces of data, and the shape is the same as the label, which is b*l*4. Calculate the loss between the segmentation result and the label. The loss function uses the dice loss function, and its calculation formula is: where X represents the segmentation result and Y represents the label; the loss function can simultaneously solve the problems of under-segmentation and over-segmentation. For the loss value on the training set, update the parameters of each layer according to the chain rule of differentiation and the gradient descent method. In the training stage, the Adam optimizer is used, and the learning strategy uses ReduceLROnPlateau. The learning strategy monitors the loss value on the validation set in each round, and updates the learning rate according to the loss value on the validation set. Obtain the single-lead electrocardiogram segmentation model based on deep learning.
2. A method for electrocardiogram segmentation using the model described in claim 1, characterized in that, This method includes the following specific steps: Step 1: Data preprocessing. For a single-lead electrocardiogram with a length of l as the input, divide it into several blocks according to a window of size d to form an electrocardiogram data sequence with a length of n and a dimension of d, where n*d = l. Step 2: Inference stage, which is divided into the following three sub-steps: Step 2-1: Divide the single-lead electrocardiogram into several pieces of data with the same length as in the training stage, send each piece of data into the model, and obtain the preliminary segmentation result. Step 2-2: Binarize the segmentation result according to the value of the preliminary segmentation result, and perform local mode filtering to eliminate noise to obtain the final segmentation result. Step 2-3: Statistically organize the pixel-level segmentation results to obtain the band information of the entire single-lead electrocardiogram.