R peak detection method and system suitable for electrocardiogram monitoring

By generating a two-dimensional image from 12-lead ECG signals and combining it with feature extraction and dynamic threshold processing, the problems of false detection and missed detection of R-peaks in noisy environments by traditional methods are solved, achieving higher detection accuracy and robustness.

CN120694657APending Publication Date: 2025-09-26HUAZHONG AGRI UNIV +3
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510689680.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing ECG signal processing methods are easily affected by interference signals, resulting in false detection or missed detection in R-peak detection, and perform poorly especially on low-quality ECG datasets.

Method used

The 12-lead ECG signals are stacked in lead order to generate a two-dimensional image. The R-peak features are extracted using local and global feature extraction modules. False detection correction is performed by combining one-dimensional Manhattan distance and dynamic RR interval threshold. Data enhancement is used to improve the robustness of the model.

Benefits of technology

The accuracy and robustness of R-peak detection are improved, false positives and false negatives are reduced, the method adapts to different heart rate changes, and the detection performance in noisy environments is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120694657A_ABST
    Figure CN120694657A_ABST
Patent Text Reader

Abstract

The invention discloses an R peak detection method and system suitable for electrocardiogram monitoring, and the method comprises the following steps: carrying out the sliding window segmentation of each lead data of 12-lead signals to be detected, stacking the 12-lead signals in the same window in the vertical direction according to the lead sequence, generating a two-dimensional lead image, and carrying out the sliding window segmentation of the lead data of the 12-lead signals to be detected; marking a time position number for each lead image; inputting the lead images into a pre-trained electrocardiosignal R peak detection model, wherein the detection model outputs prediction frames of R peaks of all the lead images; mapping the horizontal center position of the prediction frame back to a one-dimensional time coordinate, splicing detection results of all windows according to time position numbers, and generating an R peak sequence; traversing the R peak sequence by using a 10-second sliding window, calculating the average RR interval of the R peak sequence, dynamically setting a threshold value according to the average RR interval, detecting an abnormal short RR interval according to the threshold value, and deleting a false detection peak by combining the positions of the R peaks before and after the RR interval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electrocardiogram signal processing, and in particular to an R-peak detection method and system suitable for electrocardiogram monitoring. Background Art

[0002] The electrocardiogram (ECG) is a biomedical signal that reflects the physiological activity of the heart, displaying its electrical depolarization and repolarization patterns. The R peak of the ECG represents the process of ventricular depolarization and is an important basis for diagnosing arrhythmias. Accurate detection of the R peak is the foundation of ECG analysis and is crucial for subsequent disease diagnosis and analysis.

[0003] Currently, researchers have proposed a variety of R-peak detection techniques, including traditional signal processing algorithms and deep learning-based algorithms. Traditional signal processing algorithms achieve R-peak detection by setting initial values ​​such as filter parameters and decision thresholds. Deep learning-based algorithms, on the other hand, utilize neural network models to learn and analyze electrocardiogram data, achieving more efficient and accurate R-peak detection.

[0004] However, existing technologies have numerous shortcomings. Traditional signal processing methods are often sensitive to noise and easily affected by interfering signals, leading to false or missed detections. While deep learning models have demonstrated impressive performance on high-quality clinical ECG datasets, these algorithms struggle to maintain robustness when detecting ECG signals in real-world environments prone to interference. Summary of the Invention

[0005] The present invention proposes an R-peak detection method and system suitable for electrocardiogram monitoring, which solves the problem that traditional signal processing methods are easily affected by interference signals, resulting in false detection or missed detection.

[0006] To solve the above technical problems, the present invention provides an R-peak detection method suitable for electrocardiogram monitoring, comprising the following steps:

[0007] Step S1: performing sliding window segmentation on each lead data of the 12-lead signal to be detected, stacking the 12-lead signals in the same window in the vertical direction according to the lead sequence to generate a two-dimensional lead image, and marking each lead image with a time position number;

[0008] Step S2: inputting the lead image into a pre-trained ECG signal R peak detection model, and the detection model outputs the predicted boxes of the R peaks of all lead images;

[0009] Step S3: Map the horizontal center position of the prediction box back to the one-dimensional time coordinate, and splice the detection results of all windows according to the time position number to generate an R-peak sequence;

[0010] Step S4: traverse the R peak sequence with a 10-second sliding window, calculate the average RR interval of the R peak sequence, dynamically set the threshold based on the average RR interval, detect abnormally short RR intervals based on the threshold, and delete falsely detected peaks based on the positions of the R peaks before and after the RR interval.

[0011] Preferably, in step S2, the detection model is pre-trained using the ECG signal after data enhancement, and the expression of the ECG signal after data enhancement is:

[0012] y = x + k × noise;

[0013]

[0014] In the above formula, y is the ECG signal after data enhancement; x is the original ECG signal; k is the proportional coefficient; noise is the noise signal; N is the length of the data point; y n The amplitude of the ECG signal at point n after data enhancement; x n is the amplitude of the original ECG signal at point n; SNR is the signal-to-noise ratio, including ECG noise and Gaussian noise, where the ECG noise includes baseline drift, muscle artifacts, and electrode motion artifacts; P s is the signal power; P n is the noise power; lg is the logarithm with base 10.

[0015] Preferably, the detection model includes a local feature extraction module, a global feature extraction module and a feature fusion module;

[0016] The local feature extraction module is used to extract the local features of the R peak in the lead image, and directly input the extracted local features into the global feature extraction module;

[0017] The global feature extraction module includes an inter-strip feature extraction module, an intra-strip feature extraction module, and a feature fusion module. The inter-strip feature extraction module calculates the correlation between different strips in the time series dimension, and the intra-strip feature extraction module calculates the correlation between different leads in the same time series position in the lead dimension.

[0018] The feature fusion module splices the feature maps output by the inter-strip feature extraction module and the intra-strip feature extraction module to generate a globally enhanced feature representation.

[0019] Preferably, in step S2, a non-maximum suppression method based on one-dimensional Manhattan distance is used for the prediction boxes of the R peaks of all lead images to remove redundant prediction boxes to obtain a final prediction result. The non-maximum suppression method based on one-dimensional Manhattan distance includes the following steps:

[0020] Step S21: Sort all predicted boxes by confidence from high to low, generate a list of candidate boxes, and create an output list;

[0021] Step S22: Select the predicted box with the highest confidence from the candidate box list as the reference box, remove it from the candidate box list and add it to the output list;

[0022] Step S23: traverse the remaining candidate boxes, and calculate the one-dimensional Manhattan distance between each candidate box and the reference box in turn. If the one-dimensional Manhattan distance is less than the set distance threshold, the candidate box is removed from the candidate box list; otherwise, the candidate box is retained and added to the output list;

[0023] Step S24: Repeat steps S22 to S23 until the candidate box list is empty.

[0024] Preferably, the length of the sliding window in step S1 is 5 seconds, and the step length is 4.8 seconds.

[0025] Preferably, when the detection results of all windows are spliced ​​according to the time position numbers in step S3, a deduplication operation is performed on each sliding window except the first and last windows, including the following steps:

[0026] Step S31: Search for the prediction frame of the current window, filter out the prediction frame within the time range of 0.2 seconds to 4.9 seconds, and add it to the first list lst1;

[0027] Step S32: Search for the prediction frame of the next window, select the prediction frame within the time range of 4.9 seconds to 5 seconds, and add it to the second list lst2;

[0028] Step S33: Calculate the time distance between the last prediction box lst1[-1] in lst1 and the first prediction box lst2[0] in lst2. If the time distance is less than the set time threshold, lst1[-1] and lst2[0] are considered to be repeatedly detected R peaks, and lst1[-1] is deleted; otherwise, lst1[-1] is retained.

[0029] Preferably, step S4 includes the following steps:

[0030] Step S41: Traverse the R-peak sequence using a sliding window. Each window has a time length of 10 seconds. 10 seconds of data is intercepted from the start of the sequence of the first window. The left boundary of the subsequent window is the second-to-last R-peak position of the previous window. If the length of the last window is less than 10 seconds, 10 seconds of data is intercepted from the end of the sequence in the reverse direction.

[0031] Step S42: Perform the following operations on the R peak within each 10-second window:

[0032] Step S421: Calculate the time intervals between all adjacent R peaks in the window as RR intervals, and generate a RR interval list;

[0033] Step S412: Calculate the average RR interval in the window and calculate the interval threshold T based on the average RR interval. min :

[0034] T min =Mean RR +α·std RR ;

[0035] Where, Mean RR is the average RR interval; α is the threshold coefficient; std RR is the standard deviation of the RR interval;

[0036] Step S43: traverse the RR interval list in the window and take the RR interval that is smaller than the interval threshold as an abnormal RR interval;

[0037] Step S44: correct each abnormal RR interval, update the R peak list in the window according to the correction result, and output the corrected R peak position sequence.

[0038] Preferably, step S44 includes the following steps:

[0039] Step S441: If the abnormal RR interval is located at the start or end position of the window, execute step S442; otherwise, execute step S443;

[0040] Step S442: If the abnormal RR interval is at the start position of the window, determine the RR interval to the right of the abnormal RR interval. right Is it less than the interval threshold T min , if RR right <T min , then delete RR right Otherwise, T min Update to the current RR interval; if the abnormal RR interval is at the end of the window, the processing ends directly and the RR interval is left for the next window to judge;

[0041] Step S443: Find the left RR interval RR of the current RR interval left and right RR interval RR right , if RR left or RR right Less than T min , then delete RR left or RR right Otherwise, T min Updated to the current RR interval.

[0042] The present invention also provides an R-peak detection system suitable for electrocardiogram monitoring, which is implemented based on the above-mentioned R-peak detection method suitable for electrocardiogram monitoring and includes: a lead image generation module, an R-peak detection module, a redundant frame removal module, a time position restoration module, an R-peak deduplication module, and a false detection correction module;

[0043] The lead image generation module performs sliding window processing on the input 12-lead ECG signals, vertically stacks the 12-lead signals in the same window in lead order, generates lead images, and labels each lead image with a time position number;

[0044] The R-peak detection module outputs the predicted frame of the R-peak in each lead image;

[0045] The redundant frame removal module calculates the distance between prediction frames based on the one-dimensional Manhattan distance and removes redundant frames whose distance is less than a threshold;

[0046] The time position restoration module maps the center point coordinates of the prediction box back to the time position in the one-dimensional time series, and splices the detection results of all windows in the order of the time position numbers to generate an R-peak time series;

[0047] The R-peak deduplication module detects R-peaks in the overlapping area of ​​adjacent sliding windows, calculates the time distance between R-peaks in adjacent windows, and removes duplicate R-peaks if the distance is less than a set time threshold;

[0048] The false detection correction module uses a 10-second window length to slide through the R-peak sequence, calculates the time intervals between adjacent R-peaks in the window, generates an RR interval list, calculates a dynamic threshold based on the average RR interval, and treats RR intervals smaller than the dynamic threshold as abnormal RR intervals. The abnormal RR intervals are corrected based on their positions and the information of the preceding and following R-peaks, and the final R-peak time series is output.

[0049] Preferably, the R-peak detection module includes a local feature extraction module, a global feature extraction module and a feature fusion module;

[0050] The local feature extraction module is used to extract the local features of the R peak in the lead image, and directly input the extracted local features into the global feature extraction module;

[0051] The global feature extraction module includes an inter-strip feature extraction module, an intra-strip feature extraction module, and a feature fusion module. The inter-strip feature extraction module calculates the correlation between different strips in the time series dimension, and the intra-strip feature extraction module calculates the correlation between different leads in the same time series position in the lead dimension.

[0052] The feature fusion module splices the feature maps output by the inter-strip feature extraction module and the intra-strip feature extraction module to generate a globally enhanced feature representation.

[0053] The benefits of the present invention include at least:

[0054] 1. The 12-lead signals are stacked in lead order to generate a two-dimensional image, making full use of multi-lead information. Compared with the traditional single-lead analysis method, it can more comprehensively capture the characteristics of the ECG signal;

[0055] 2. Label each lead image with a time position number and map the horizontal center position of the prediction box back to a one-dimensional time coordinate. Then, splice the detection results of all windows according to the time position number. This can effectively solve the boundary problem caused by sliding window segmentation, ensure the continuity and integrity of the R-peak sequence, and avoid false detection or missed detection caused by window segmentation.

[0056] 3. Use a 10-second sliding window to traverse the R peak sequence, calculate the average RR interval, and dynamically set the threshold based on the average RR interval so that the threshold can adapt to the heart rate changes of different patients and avoid false detection or missed detection caused by a fixed threshold. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 Schematic diagram of a method flow in an embodiment of the present invention;

[0058] Figure 2 A flowchart of a method according to an embodiment of the present invention;

[0059] Figure 3 Schematic diagram of the structure of the R-peak detection model according to an embodiment of the present invention;

[0060] Figure 4 Schematic diagram of the SA strip attention structure in the R-peak detection model according to an embodiment of the present invention;

[0061] Figure 5 Schematic diagram of the NMS calculation principle based on one-dimensional Manhattan distance in an embodiment of the present invention;

[0062] Figure 6 1 is a flow chart of the NMS method based on one-dimensional Manhattan distance in an embodiment of the present invention;

[0063] Figure 7 Flowchart of deduplication operation in the process of restoring a one-dimensional signal in an embodiment of the present invention;

[0064] Figure 8 Flowchart of R-peak misdetection correction based on dynamic threshold in an embodiment of the present invention;

[0065] Figure 9 Schematic diagram of the output results of the detection model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0066] The following is a clear and complete description of the technical solutions in the embodiments of the present invention, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts are within the scope of protection of the present invention.

[0067] like Figure 1 and Figure 2 As shown, an embodiment of the present invention provides an R-peak detection method suitable for ECG monitoring, aiming to solve the problems of insufficient robustness of existing signal processing methods for low-quality Holter ECGs and insufficient utilization of the correlation between 12 leads. The method provided by the embodiment of the present invention can detect the position of the R-peak of multi-lead ECG signals of 30 minutes or more, including the following steps:

[0068] Step S1: Perform sliding window segmentation on each lead data of the 12-lead signal to be detected, stack the 12-lead signals in the same window in the vertical direction according to the lead sequence to generate a two-dimensional lead image, and mark each lead image with a time position number.

[0069] Specifically, each of the 12 leads to be tested is segmented into a sliding window with a 5-second time window and a 4.8-second sliding step. The data within the window is then normalized to its maximum and minimum values. The ECG data from different leads within the same time window are stacked vertically in lead order, generating a 640×640 pixel visual two-dimensional lead image. Each lead image is numbered at its time position. ECG signals whose end points are less than the length of a window are discarded.

[0070] Step S2: The lead image is input into the pre-trained ECG signal R peak detection model, and the detection model outputs the predicted box of the R peak of all lead images.

[0071] Specifically, if Figure 3 As shown in FIG, the R peak detection model of the embodiment of the present invention consists of three parts: a strip feature extraction module, a neck network module, and a head prediction module.

[0072] The strip feature extraction module includes a C3 convolutional feature extractor, a CBS convolutional feature extractor, a SA global information extractor, and a feature fusion module. The C3 convolutional feature extractor consists of three consecutive standard convolutional layers, extracting local detailed features such as the steep rising edge and peak shape of the R-peak. The CBS convolutional feature extractor incorporates a cross-stage partially connected CSP structure. By splitting and fusing feature maps, it reduces computational redundancy, enhances feature reuse, and improves model efficiency. The feature map of the last layer of the convolutional feature extractor is directly input into the SA global information extractor. The SA global information extractor includes inter-SA, which calculates attention between strips, and intra-SA, which calculates attention within strips. Inter-SA calculates the correlation between different strips in the horizontal direction (i.e., the temporal dimension of the ECG), capturing the distribution pattern of R-peaks in the time series. Global pooling and attention weight distribution enhance the model's ability to model the temporal relationships of R-peaks. Intra-SA calculates the correlation between different leads within the same temporal position in the vertical direction (i.e., the 12-lead dimension). It comprehensively utilizes multi-lead information and dynamically weights each lead's features through a channel-attention mechanism to suppress interference from noisy leads. The feature fusion unit concatenates the output feature maps of Inter-SA and Intra-SA. The strip pyramid feature fusion (SPFF) module fuses multi-scale receptive field information to generate a globally enhanced feature representation.

[0073] The neck network module includes upsampling and downsampling pyramid structures, and the head prediction module outputs three sizes of prediction information after convolution.

[0074] like Figure 4 As shown in the figure, the SA global information extractor calculates the inter-dataset correlation between strips and the intra-dataset correlation within the strips from two directions. In the horizontal direction (H) of the electrocardiogram, the strip attention calculation can establish the dependency between different R peaks; in the vertical direction (V), the strip attention calculation integrates the 12-lead information at different levels. Finally, the receptive fields of different sizes are fused through the SPFF module to optimize feature representation and improve model performance. The SA global information extractor models the global information of the electrocardiogram in a strip-shaped calculation area. The strip-shaped setting targets the long strip-shaped area of ​​the R peak unique to the electrocardiogram, which can more effectively capture the geometric features of the target. By focusing on the strip-shaped area, the model can capture important information in a local range and reduce the interference of irrelevant background.

[0075] The 12-lead raw data file and the corresponding label file were read, and the raw data were segmented using a sliding window with a length of 5 seconds and a step size of 4.8 seconds. The data within the window were preprocessed by maximum and minimum normalization.

[0076] Read the label file and search for R-peaks within the window. For each found R-peak, determine its temporal distance from the left and right edges of the image. If the temporal distance is less than 0.15 seconds, discard the R-peak; otherwise, retain it. For the remaining R-peaks, create an R-peak label using the maximum image height as the height of the label anchor box and a temporal length of 0.3 seconds as the width of the anchor box.

[0077] The 12-lead data from the same time window is stacked vertically in lead order and converted to an RGB image using Python. The RGB image and the corresponding R-peak labels are saved. Simultaneously, the sliding window is used to process the next window until the entire ECG signal is processed, providing training data for the R-peak detection model.

[0078] After obtaining the training data, the training data is enhanced based on the signal-to-noise ratio (SNR). The training data after data enhancement includes ECG noise and Gaussian noise. The ECG noise comes from three noises recorded in the NSTDB database: baseline drift (bw), muscle artifact (emg), and electrode motion artifact (em). In addition to these three most common ECG noises, Gaussian noise is also added. The proportions of bw, emg, em, and Gaussian noise in the mixed noise (noise) are 3:3:3:1 respectively. Combining the original signal x and the noise signal noise scaled by the proportional coefficient k, the expression of the enhanced signal y is:

[0079] y=x+k×noise.

[0080] The expression for calculating the noise scaling factor k with different SNR preset values ​​is:

[0081]

[0082] In the above formula, N is the length of the data point; y n The amplitude of the ECG signal at point n after data enhancement; x n is the amplitude of the original ECG signal at point n; P s is the signal power; P n is the noise power; lg is the logarithm with base 10.

[0083] Among them, SNR was set at three levels: 0dB, 5dB, and 10dB for training and testing. Two public datasets were used for research, including the St. Petersburg INCART 12-lead Arrhythmia Database and the Lobachevsky University Electrocardiography Database.

[0084] Finally, the non-maximum suppression method based on one-dimensional Manhattan distance is used for the prediction box of the R peak of all lead images to remove redundant prediction boxes and obtain the final prediction results, as shown in Figure 5 As shown, the following steps are included:

[0085] Step S21: Sort all predicted boxes by confidence from high to low, generate a list of candidate boxes, and create an output list.

[0086] Step S22: Select the predicted box with the highest confidence from the candidate box list as the reference box, remove it from the candidate box list and add it to the output list.

[0087] Step S23: traverse the remaining candidate boxes, calculate the one-dimensional Manhattan distance between each candidate box and the reference box in turn, if the one-dimensional Manhattan distance is less than the set distance threshold, remove the candidate box from the candidate box list, otherwise retain the candidate box and add it to the output list. Figure 6 The figure shows the NMS calculation principle diagram based on the one-dimensional Manhattan distance. Based on this principle, the expression of the one-dimensional Manhattan distance is:

[0088]

[0089] Where, d ij is the one-dimensional Manhattan distance between prediction box i and prediction box j; x i 、x j They are predicted box i and predicted box j respectively.

[0090] For example, when detecting R peaks in an ECG signal segment, assume there are two candidate boxes, the center coordinates of the reference box are (100, 200), the center coordinates of candidate box 1 are (110, 210), and the center coordinates of candidate box 2 are (150, 250). Then the one-dimensional Manhattan distance between candidate box 1 and the reference box is d1 = |100-110| = 10, and the one-dimensional Manhattan distance between candidate box 2 and the reference box is d2 = |100-150| = 50.

[0091] Because the refractory period of a heartbeat is approximately 0.25s, in the embodiment of the present invention, the threshold is set to 0.3 seconds, and the distance threshold in the corresponding input image is set to 38 pixels. If the one-dimensional Manhattan distance between the candidate frame and the reference frame is less than the predefined distance threshold, it is considered a redundant frame and removed. In the above example, if the one-dimensional Manhattan distance d1 between candidate frame 1 and the reference frame is 10<38, candidate frame 1 is removed, and if the one-dimensional Manhattan distance d2 between candidate frame 2 and the reference frame is 50>38, candidate frame 2 is retained.

[0092] Step S24: Repeat steps S22 to S23 until the candidate box list is empty.

[0093] Step S3: Map the horizontal center position of the prediction box back to the one-dimensional time coordinate, and splice the detection results of all windows according to the time position number to generate an R-peak sequence.

[0094] Specifically, when the detection results of all windows are spliced ​​according to the time position number, each sliding window except the first and last windows is deduplicated, such as Figure 7 As shown, the following steps are included:

[0095] Step S31: Search for the prediction frame of the current window, filter out the prediction frames within the time range of 0.2 seconds to 4.9 seconds, and add them to the first list lst1.

[0096] Step S32: Search for the prediction frame of the next window, filter out the prediction frames within the time range of 4.9 seconds to 5 seconds, and add them to the second list lst2.

[0097] Step S33: Calculate the time distance between the last prediction box lst1[-1] in lst1 and the first prediction box lst2[0] in lst2. If the time distance is less than the set time threshold, lst1[-1] and lst2[0] are considered to be repeated R peaks, and lst1[-1] is deleted. Otherwise, lst1[-1] is retained, lst1 and lst2 are spliced, added to the output list, and the next lead image is processed.

[0098] Step S4: traverse the R peak sequence with a 10-second sliding window, calculate the average RR interval of the R peak sequence, dynamically set the threshold based on the average RR interval, detect abnormally short RR intervals based on the threshold, and delete the falsely detected peaks based on the positions of the R peaks before and after the RR interval. Figure 8 As shown, the specific steps include:

[0099] Step S41: Traverse the R-peak sequence through a sliding window. The time length of each window is 10 seconds. 10 seconds of data is intercepted from the start of the sequence of the first window. The left boundary of the subsequent window is the second-to-last R-peak position of the previous window. If the length of the last window is less than 10 seconds, 10 seconds of data is intercepted in reverse from the end of the sequence.

[0100] Step S42: Perform the following operations on the R peak within each 10-second window:

[0101] Step S421: Calculate the time intervals between all adjacent R peaks in the window as RR intervals, and generate a RR interval list;

[0102] Step S412: Calculate the average RR interval in the window and calculate the interval threshold T based on the average RR interval. min :

[0103] T min =MeanRR +α·std RR ;

[0104] Where, Mean RR is the average RR interval; α is the threshold coefficient; std RR is the standard deviation of the RR interval.

[0105] Step S43: traverse the RR interval list in the window, and take the RR intervals smaller than the interval threshold as abnormal RR intervals and add them to the false positive list Lst FP .

[0106] Step S44: False positive list Lst FP Correct each abnormal RR interval in the window, update the R peak list in the window according to the correction result, and output the corrected R peak position sequence:

[0107] Step S441: If the abnormal RR interval is located at the start or end position of the window, execute step S442; otherwise, execute step S443;

[0108] Step S442: If the abnormal RR interval is at the start position of the window, determine the RR interval to the right of the abnormal RR interval. right Is it less than the interval threshold T min , if RR right <T min , then delete RR right Otherwise, T min Update to the current RR interval; if the abnormal RR interval is at the end of the window, the processing ends directly and the RR interval is left for the next window to judge;

[0109] Step S443: Find the left RR interval RR of the current RR interval left and right RR interval RR right , if RR left or RR right Less than T min , then delete RR left or RR right Otherwise, T min Update to the current RR interval. Finally, we get Figure 9 The test results shown, where R peok This is the R peak.

[0110] The R-peak detection model was trained using the default stochastic gradient descent (SGD) optimizer, with an initial learning rate of 0.01 and weight decay of 0.005. Single-class mode was used for training, and all models were trained for 200 epochs, except for the noise-free model, which was trained for 100 epochs.

[0111] In order to further illustrate the performance of the method of the embodiment of the present invention, experimental evaluation was performed on two datasets: St Petersburg INCART 12-lead Arrhythmia Database (INCART) and Lobachevsky University Electrocardiography Database (LUDB). The number of patients in the INCART electrocardiogram dataset is 32, and the dataset includes 75 records of these patients, each of which is 30 minutes long. Each ECG signal recorded using 12 leads is collected at a sampling frequency of 257Hz, and the dataset consists of approximately 175,461 heartbeats. LUDB is an electrocardiogram database with significant markings of P, T waves, and QRS complexes. The database contains a total of 200 10-second 12-lead electrocardiogram signals with an acquisition frequency of 500Hz. The dataset consists of approximately 2,100 heartbeats.

[0112] Model performance was evaluated using both intra-dataset and inter-dataset methods. 75 data points were randomly selected from the INCART dataset, with 62 of them being used as the training set, 6 as the validation set, and the remaining 7 as the test set for intra-dataset testing. The model trained on the INCART dataset was directly tested on 200 data points from the LUDB dataset to assess its generalization performance. The model was trained on the INCART dataset because it contains a larger number of heartbeats.

[0113] The performance of the model is evaluated using three metrics: recall (Se), precision (PPr), and F1-score. Recall is expressed as the ratio of the number of correctly detected R-peaks (TP) to the number of all true R-peaks.

[0114]

[0115] The precision (PPr) is expressed as the ratio of the number of correctly detected R peaks (TP) to the number of all detected peaks.

[0116]

[0117] F1-Score is a comprehensive evaluation indicator calculated by recall and precision.

[0118]

[0119] TP (true positive), FP (false positive), and FN (false negative) represent the number of correctly detected R-peaks, incorrectly detected R-peaks, and undetected R-peaks, respectively. All three are performed within a 150-ms tolerance of the true peak position. That is, a predicted point within 150 ms of the true peak is considered a true positive (TP).

[0120] As shown in Tables 1, 2, 3, and 4, the F1 scores of the method of the present invention were 99.97%, 99.86%, 99.63%, and 98.00% for the original signal, SNR = 10dB, SNR = 5dB, and SNR = 0dB, respectively. Compared with other methods such as the Pan-Tompkins algorithm (P and T), the U-Net segmentation algorithm (U-Net), the self-organizing operator neural network (Self-ONN), the one-dimensional convolution and stationary wavelet transform combined algorithm (SWT&CNN), and the single-lead Yolo algorithm (SL-Yolo), this method significantly reduced the number of false negatives and false positives, demonstrating good robustness. On the LUDB dataset, the model also demonstrated high accuracy and stability. As shown in Table 5, the F1 scores were 99.89%, 100%, 100%, and 99.86% under the four noise levels, respectively. This method achieved better detection results than other methods without requiring secondary training, demonstrating the model's good generalization ability.

[0121] Table 1 Peak detection performance of this method and other methods on INCART original signal

[0122]

[0123] Table 2 Peak detection performance of this method and other methods on INCART noise enhanced data (SNR=10)

[0124]

[0125] Table 3 Peak detection performance of this method and other methods on INCART noise enhanced data (SNR=5)

[0126]

[0127] Table 4 Peak detection performance of this method and other methods on INCART noise enhanced data (SNR = 0)

[0128]

[0129]

[0130] Table 5 Peak detection performance of this method and other methods on LUDB original signal

[0131]

[0132] An embodiment of the present invention also provides an R-peak detection system suitable for electrocardiogram monitoring, which is implemented based on the above-mentioned R-peak detection method suitable for electrocardiogram monitoring, and includes: a lead image generation module, an R-peak detection module, a redundant frame removal module, a time position restoration module, an R-peak deduplication module and a false detection correction module.

[0133] Lead image generation module: Perform sliding window processing on the input 12-lead ECG signals, vertically stack the 12-lead signals in the same window in lead order, generate lead images, and annotate each lead image with a time position number.

[0134] R-peak detection module: outputs the predicted box of the R-peak in each lead image.

[0135] Redundant box removal module: Calculates the distance between predicted boxes based on the one-dimensional Manhattan distance and removes redundant boxes whose distance is less than the threshold.

[0136] Time position restoration module: maps the center point coordinates of the prediction box back to the time position in the one-dimensional time series, splices the detection results of all windows in the order of the time position numbers, and generates the R-peak time series.

[0137] R-peak deduplication module: detects R-peaks in the overlapping area of ​​adjacent sliding windows, calculates the time distance between R-peaks in adjacent windows, and removes duplicate R-peaks if the distance is less than the set time threshold.

[0138] False detection correction module: With a 10-second window length, it slides through the R-peak sequence, calculates the time intervals between adjacent R-peaks within the window, generates a list of RR intervals, calculates the dynamic threshold based on the average RR interval, and treats RR intervals smaller than the dynamic threshold as abnormal RR intervals. It corrects the abnormal RR intervals based on their position and the information of the preceding and following R peaks, and outputs the final R-peak time series.

[0139] In summary, the present invention provides an R-peak detection method suitable for ECG monitoring. By treating a 12-lead ECG as a two-dimensional image and processing it, the 12-lead information is fully integrated. The 12 leads record the electrical activity of the heart from different angles, providing more comprehensive physiological data. In the model, although the starting and ending points of the waveforms between different leads are different, the waveform categories are closely related. Through weighted averaging calculation between leads, R-peak positioning errors are effectively eliminated. For example, when processing low-quality ECG signals, the waveform of a certain lead may be unclear due to factors such as noise. However, by integrating information from other leads, missed waveforms can be supplemented and misdetected waveforms can be eliminated, thereby improving the accuracy of R-peak detection.

[0140] The introduced strip attention mechanism enhances the model's ability to perceive R-peak features. The R-peak feature is in the shape of a vertical strip, and the mechanism calculates attention from within the strip (Intra-SA module) and between the strips (Inter-SA module). The Intra-SA module extracts local information of the R-peak, and the Inter-SA module calculates the distribution pattern of the R-peak in the entire image. The two together promote global attention fusion. This enables the model to locate the R-peak more accurately and reduce the number of false positives (FP) and false negatives (FN). For example, in the original signal test of the INCART dataset, the number of FP and FN of the method of the present invention dropped to below 5, while the other comparison methods were relatively more, which significantly improved the detection accuracy.

[0141] The proposed non-maximum suppression (NMS) algorithm, based on the one-dimensional Manhattan distance, effectively suppresses redundant detection boxes by setting a threshold based on the heartbeat's refractory period. Compared with the traditional NMS algorithm based on intersection-of-union (IoU), this algorithm better conforms to the R-peak search principle, better handles the problem of R-peak prediction box adhesion, reduces misjudgments caused by noise, and further enhances the model's detection performance in noisy environments.

[0142] In the inter-dataset test, the model trained using the INCART dataset was directly tested on the LUDB dataset, achieving excellent results. For example, under the four noise levels of the LUDB dataset, the F1 scores were 99.89%, 100%, 100%, and 99.86%, respectively, indicating that the model can adapt to different data distributions without complex adjustments for different datasets, has strong generalization capabilities, and can stably and accurately detect R peaks on different datasets. In the tests of different cases (such as I03, I11, I15, I18, I19, I31, I62, etc.) in the INCART test set, the F1 score distribution of the method of the present invention under different noise levels is stable and has a small variance. For example, under noise-free (Original data) conditions, compared with methods such as P and T and Self-ONN, the performance of this method in different cases is more stable, indicating that the model has strong adaptability in different case data, is not affected by the characteristics of specific case data, and can be widely used in a variety of clinical scenarios.

[0143] The technical features of the above embodiments may be combined in any manner. To simplify the description, not all possible combinations of the technical features in the above embodiments are described. Only preferred embodiments of the present invention are presented. While the description is relatively specific and detailed, it should not be construed as limiting the scope of the present invention. As long as there are no conflicts in the combination of these technical features, they should be considered to be within the scope of this specification.

[0144] It should be noted that, for those skilled in the art, various modifications and improvements can be made without departing from the scope of the present invention, and these modifications and improvements fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.

Claims

1. A R-peak detection method suitable for electrocardiogram monitoring, characterized in that: The following steps are involved: Step S1: performing sliding window segmentation on each lead data of the 12-lead signal to be detected, stacking the 12-lead signals in the same window in the vertical direction according to the lead sequence to generate a two-dimensional lead image, and marking each lead image with a time position number; Step S2: inputting the lead image into a pre-trained ECG signal R peak detection model, and the detection model outputs the predicted boxes of the R peaks of all lead images; Step S3: Map the horizontal center position of the prediction box back to the one-dimensional time coordinate, and splice the detection results of all windows according to the time position number to generate an R-peak sequence; Step S4: Use a sliding window of 10 seconds to traverse the R peak sequence, calculate the average RR interval of the R peak sequence, dynamically set the threshold based on the average RR interval, detect abnormally short RR intervals based on the threshold, and delete falsely detected peaks based on the positions of the R peaks before and after the RR interval.

2. The R-peak detection method for electrocardiogram monitoring according to claim 1, characterized in that: In step S2, the detection model is pre-trained using the ECG signal after data enhancement. The expression of the ECG signal after data enhancement is: y = x + k × noise; In the above formula, y is the ECG signal after data enhancement; x is the original ECG signal; k is the proportional coefficient; noise is the noise signal; N is the length of the data point; y n The amplitude of the ECG signal at point n after data enhancement; x n is the amplitude of the original ECG signal at point n; SNR is the signal-to-noise ratio, including ECG noise and Gaussian noise, where the ECG noise includes baseline drift, muscle artifacts, and electrode motion artifacts; P s is the signal power; P n is the noise power; lg is the logarithm with base 10.

3. The R-peak detection method for electrocardiogram monitoring according to claim 1, characterized in that: The detection model includes a local feature extraction module, a global feature extraction module and a feature fusion module; The local feature extraction module is used to extract the local features of the R peak in the lead image, and directly input the extracted local features into the global feature extraction module; The global feature extraction module includes an inter-strip feature extraction module, an intra-strip feature extraction module, and a feature fusion module. The inter-strip feature extraction module calculates the correlation between different strips in the time series dimension, and the intra-strip feature extraction module calculates the correlation between different leads in the same time series position in the lead dimension. The feature fusion module splices the feature maps output by the inter-strip feature extraction module and the intra-strip feature extraction module to generate a globally enhanced feature representation.

4. The R-peak detection method for electrocardiogram monitoring according to claim 1, characterized in that: In step S2, a non-maximum suppression method based on one-dimensional Manhattan distance is used to remove redundant prediction boxes from the prediction boxes of the R peaks of all lead images to obtain a final prediction result. The non-maximum suppression method based on one-dimensional Manhattan distance includes the following steps: Step S21: Sort all predicted boxes by confidence from high to low, generate a list of candidate boxes, and create an output list; Step S22: Select the predicted box with the highest confidence from the candidate box list as the reference box, remove it from the candidate box list and add it to the output list; Step S23: traverse the remaining candidate boxes, and calculate the one-dimensional Manhattan distance between each candidate box and the reference box in turn. If the one-dimensional Manhattan distance is less than the set distance threshold, the candidate box is removed from the candidate box list; otherwise, the candidate box is retained and added to the output list; Step S24: Repeat steps S22 to S23 until the candidate box list is empty.

5. The R-peak detection method for electrocardiogram monitoring according to claim 1, characterized in that: The length of the sliding window in step S1 is 5 seconds, and the step length is 4.8 seconds.

6. The R-peak detection method for electrocardiogram monitoring according to claim 5, characterized in that: When the detection results of all windows are spliced ​​according to the time position numbers in step S3, a deduplication operation is performed on each sliding window except the first and last windows, including the following steps: Step S31: Search for the prediction frame of the current window, filter out the prediction frame within the time range of 0.2 seconds to 4.9 seconds, and add it to the first list lst1; Step S32: Search for the prediction frame of the next window, select the prediction frame within the time range of 4.9 seconds to 5 seconds, and add it to the second list lst2; Step S33: Calculate the time distance between the last prediction box lst1[-1] in lst1 and the first prediction box lst2[0] in lst2. If the time distance is less than the set time threshold, lst1[-1] and lst2[0] are considered to be repeatedly detected R peaks, and lst1[-1] is deleted; otherwise, lst1[-1] is retained.

7. The R-peak detection method for electrocardiogram monitoring according to claim 1, characterized in that: Step S4 includes the following steps: Step S41: Traverse the R-peak sequence using a sliding window. Each window has a time length of 10 seconds. 10 seconds of data is intercepted from the start of the sequence of the first window. The left boundary of the subsequent window is the second-to-last R-peak position of the previous window. If the length of the last window is less than 10 seconds, 10 seconds of data is intercepted from the end of the sequence in the reverse direction. Step S42: Perform the following operations on the R peak within each 10-second window: Step S421: Calculate the time intervals between all adjacent R peaks in the window as RR intervals, and generate a RR interval list; Step S412: Calculate the average RR interval in the window and calculate the interval threshold T based on the average RR interval. min : T min =Mean RR +α·std RR ; Where, Mean RR is the average RR interval; α is the threshold coefficient; std RR is the standard deviation of the RR interval; Step S43: traverse the RR interval list in the window and take the RR interval that is smaller than the interval threshold as an abnormal RR interval; Step S44: correct each abnormal RR interval, update the R peak list in the window according to the correction result, and output the corrected R peak position sequence.

8. The R-peak detection method for electrocardiogram monitoring according to claim 7, characterized in that: Step S44 includes the following steps: Step S441: If the abnormal RR interval is located at the start or end position of the window, execute step S442; otherwise, execute step S443; Step S442: If the abnormal RR interval is at the start position of the window, determine the RR interval to the right of the abnormal RR interval. right Is it less than the interval threshold T min , if RR right <T min , then delete RR right Otherwise, T min Update to the current RR interval; if the abnormal RR interval is at the end of the window, the processing ends directly and the RR interval is left for the next window to judge; Step S443: Find the left RR interval RR of the current RR interval left and right RR interval RR right , if RR left or RR right Less than T min , then delete RR left or RR right Otherwise, T min Updated to the current RR interval.

9. An R-peak detection system for electrocardiogram monitoring, implemented based on an R-peak detection method for electrocardiogram monitoring according to any one of claims 1 to 8, characterized in that: include: Lead image generation module, R-peak detection module, redundant frame removal module, time position restoration module, R-peak deduplication module and false detection correction module; The lead image generation module performs sliding window processing on the input 12-lead ECG signals, vertically stacks the 12-lead signals in the same window in lead order, generates lead images, and labels each lead image with a time position number; The R-peak detection module outputs the predicted frame of the R-peak in each lead image; The redundant frame removal module calculates the distance between prediction frames based on the one-dimensional Manhattan distance and removes redundant frames whose distance is less than a threshold; The time position restoration module maps the center point coordinates of the prediction box back to the time position in the one-dimensional time series, and splices the detection results of all windows in the order of the time position numbers to generate an R-peak time series; The R-peak deduplication module detects R-peaks in the overlapping area of ​​adjacent sliding windows, calculates the time distance between R-peaks in adjacent windows, and removes duplicate R-peaks if the distance is less than a set time threshold; The false detection correction module uses a 10-second window length to slide through the R-peak sequence, calculates the time intervals between adjacent R-peaks in the window, generates an RR interval list, calculates a dynamic threshold based on the average RR interval, and treats RR intervals smaller than the dynamic threshold as abnormal RR intervals. The abnormal RR intervals are corrected based on their positions and the information of the preceding and following R-peaks, and the final R-peak time series is output.

10. The R-peak detection system for electrocardiogram monitoring according to claim 9, characterized in that: The R-peak detection module includes a local feature extraction module, a global feature extraction module and a feature fusion module; The local feature extraction module is used to extract the local features of the R peak in the lead image, and directly input the extracted local features into the global feature extraction module; The global feature extraction module includes an inter-strip feature extraction module, an intra-strip feature extraction module, and a feature fusion module. The inter-strip feature extraction module calculates the correlation between different strips in the time series dimension, and the intra-strip feature extraction module calculates the correlation between different leads in the same time series position in the lead dimension. The feature fusion module splices the feature maps output by the inter-strip feature extraction module and the intra-strip feature extraction module to generate a globally enhanced feature representation.

Citation Information

Patent Citations

  • Atrial fibrillation occurrence risk prediction system based on heartbeat rhythm signals and application thereof

    CN113995419A

  • Electrocardiosignal characteristic wave segmentation method and FS-Net model

    CN119074008A

  • Heart disease prediction system based on electrocardiosignal and large model

    CN119312024A

  • Heart rate monitoring method, device and apparatus

    US20250025060A1

  • ECG information processing method and ECG workstation

    WO2019161611A1