An ECG Signal Interference Wave Recognition Method Incorporating Sparse Features

By fusing sparse features in ECG signal processing and using LightGBM algorithm, the problem of ECG signal interference wave recognition in the prior art is solved, and the detailed distinction and rapid and accurate identification of different levels of interference waves are achieved, which improves the efficiency of ECG analysis.

CN114818781BActive Publication Date: 2025-06-10ZHEJIANG HELOWIN MEDICAL TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210317557.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-29
Publication Date
2025-06-10
Estimated Expiration
2042-03-29

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify and distinguish different levels of ECG signal interference waves, especially in dry electrode application scenarios, resulting in difficulty in data analysis and reduced accuracy.

Method used

A method of ECG signal interference wave recognition with sparse features is adopted to add sparse features on the basis of statistical features and train the model using LightGBM algorithm to achieve detailed distinction of interference waves.

Benefits of technology

The distinction between interference waves of different levels is improved, the model's adaptability to various adverse environments is enhanced, and the interference waves is quickly and accurately automatically recognized, which reduces the hassle of subsequent analysis and improves the analysis efficiency of ECG doctors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114818781B_ABST
    Figure CN114818781B_ABST
Patent Text Reader

Abstract

An ECG signal interference wave recognition method integrating sparse features, the method comprising: (1) receiving electrocardiogram data, performing convolution smoothing and normalization processing on the electrocardiogram data, performing heartbeat localization detection and calculating the RR interval according to the heartbeat localization, then setting a threshold range for the RR interval, identifying interference waves for the heartbeat data according to the distribution of the RR interval in the electrocardiogram data, and annotating the data according to the interference wave recognition result and the heartbeat data; (2) calculating the numerical statistical features of the electrocardiogram signal; (3) selecting a sparse dictionary; (4) solving sparse coefficients according to the above sparse dictionary to obtain sparse features; (5) training an interference wave recognition model using the LightGBM algorithm according to the numerical statistical features and sparse features of the above electrocardiogram signal; (6) according to the above interference wave recognition model, applying it to the data of the validation set, evaluating the performance of the model using specificity and F1 value, and selecting the one with excellent performance as the final classification model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for identifying interference waves in ECG signals by fusing sparse features, belonging to the field of signal processing. Background Art

[0002] During the process of single-lead ECG acquisition, it is inevitable to be affected by many external factors, such as muscle movement, electromagnetic wave interference, etc., which will cause the data to be unable to be analyzed. For wearable devices, due to the use of dry electrodes to collect data, problems such as humidity changes and poor electrode contact will also be encountered. It is necessary to specifically screen these poor acquisition data. Therefore, this product adopts a method that fuses sparse features on the basis of statistical features, which has better discrimination for interference waves at different levels. By collecting data in various different environments to train the model, its adaptability to various adverse environments is improved, which is more conducive to quickly and accurately automatically identifying these interference waves, helping to reduce the trouble of subsequent analysis, improving the analysis speed and accuracy of subsequent programs, and is also an effective method to improve the analysis efficiency of cardiologists.

[0003] Currently, the analysis methods for the noise pollution degree of ECG signals are mainly applicable to ECG signals collected under wet electrodes and cannot well adapt to the application scenarios of dry electrodes. This patent adopts four-stage classification, which can make a more detailed discrimination for interference waves and can also increase the effective ECG wave data to meet more needs. Summary of the Invention

[0004] The purpose of the present invention is to overcome the above-mentioned deficiencies and provide a method for identifying interference waves in ECG signals by fusing sparse features on a wearable device that can make a detailed discrimination for interference waves and has a fast and accurate overall analysis speed.

[0005] The present invention is realized by the following technical solutions: A method for identifying interference waves in ECG signals by fusing sparse features, which adds sparse features on the basis of general statistical features and then uses the LightGBM algorithm to train the model. The method includes the following steps:

[0006] (1) Receive ECG data, perform convolution smoothing and normalization processing on the ECG data, then perform heartbeat localization detection and calculate the RR interval according to the heartbeat localization, then set the RR interval threshold range, identify interference waves for the heartbeat data according to the distribution of the RR interval in the ECG data, and label the data according to the interference wave identification result and the heartbeat data;

[0007] (2) Calculate the numerical statistical features of the ECG signal according to the labeled ECG signal data;

[0008] (3) Select a sparse dictionary according to the labeled ECG signal data;

[0009] (4) Solve the sparse coefficients according to the above sparse dictionary to obtain sparse features;

[0010] (5) Train an interference wave recognition model using the LightGBM algorithm according to the above numerical statistical features and sparse features of the electrocardiogram signal;

[0011] (6) Apply the above interference wave recognition model to the data of the validation set, evaluate the performance of the model using specificity and F1 value, and select the one with excellent performance as the final classification model.

[0012] Preferably: In the step (1), smooth and normalize the electrocardiogram data, and label the electrocardiogram data, specifically including:

[0013] (1) Smooth and normalize the 3-minute electrocardiogram data;

[0014] The convolution smoothing calculation formula is:

[0015]

[0016] where f(x) and g(x) are two integrable functions on the Euclidean space R, and t ∈ R is the integration variable;

[0017] The normalization calculation formula is:

[0018]

[0019] where μ is the mean of the data x, and σ is the standard deviation of the data x;

[0020] (2) Intercept single heartbeats from the 3-second electrocardiogram data of the slice according to the RR interval. If the length of the single heartbeat is less than 750, fill it with 0. Preferably: The specific calculation of the numerical statistical features according to the step (2) includes:

[0021] (1) Calculate numerical statistical features such as the average value, standard deviation, peak rate, difference between the sum of the maximum amplitude and the sum of the minimum amplitude, non-zero value, and median value of the 3-second data according to the electrocardiogram signal data. The calculation formulas are as follows:

[0022] Calculate the average value avg of the electrocardiogram signal, and its calculation formula is:

[0023]

[0024] where x i is the i-th value of the electrocardiogram signal data within 3 minutes, and N is the number of x i ;

[0025] Calculate the standard deviation std of the electrocardiogram signal data;

[0026]

[0027] Among them, μ is the mean value of the electrocardiogram signal data within 3 minutes;

[0028] Calculate the peak rate get_PR, and the calculation formula is:

[0029]

[0030] Among them, x i is the i-th value of the electrocardiogram signal data within 3 minutes, and std is the standard deviation of the electrocardiogram data;

[0031] Calculate the difference between the sum of the maximum amplitudes and the sum of the minimum amplitudes get_AD, and the calculation formula is:

[0032]

[0033] Among them, x i is the i-th value of the electrocardiogram signal data within 3 minutes, and std is the standard deviation of the electrocardiogram data;

[0034] Calculate the non-zero value get_zero, and the calculation formula is:

[0035]

[0036] Among them, x i is the i-th value of the electrocardiogram signal data within 3 minutes;

[0037] Calculate the median median, and the calculation formula is: Arrange the 750 values of the given electrocardiogram signal data from small to large, and take the average of the middle two numbers;

[0038] (2) Use the random forest model to rank the importance of the 6 statistical value features, and determine the weights of each feature according to the obtained ranking results.

[0039] Preferably: In the step (3), the sparse dictionary for the electrocardiogram signal data is selected as follows:

[0040] (1) Select the level dictionary at each level, calculate the k training samples with the largest similarity using the Euclidean distance formula, and obtain the class sub-dictionary of a certain class, denoted as X i ;

[0041] The Euclidean distance formula is:

[0042]

[0043] Among them, n is the length of the data;

[0044] (2) Merge the class sub-dictionaries of all classes to obtain the large dictionary X: X = [X 1 , X 2 , X 3 , X 4 .

[0045] Preferably: In the step (4), calculating the sparse feature according to the sparse dictionary specifically means:

[0046] Obtain the sparse representation coefficients according to the OMP algorithm, so as to obtain the sparse representation feature α;

[0047] Steps for solving the OMP algorithm:

[0048] a. Input: original signal y, dictionary matrix X, sparse coefficient α, residual vector r, active set D X

[0049] b. Initialization: r 0 = y,

[0050] c. Find the principle of the largest inner product with r and add it to the active set

[0051] d. Use the least squares method to solve the best approximation of the current residual, and update the residual and coefficients X

[0052]

[0053]

[0054] e. Repeat c and d until the termination condition is reached. The number of non-zero elements in α or the threshold set for the residual is used as the termination condition.

[0055] Preferably: In the step (5), use the LightGBM algorithm to train models for the numerical statistical features and the sparse features respectively, and fuse the two models by the weighted average method.

[0056] Preferably: Use the data of the test set to test the obtained model, and obtain the model with the best performance as the final interference wave identification model.

[0057] ​The present invention relates to a method for identifying interference waves in ECG signals by fusing sparse features, which can distinguish interference waves in detail, with fast and accurate overall analysis speed. It adopts a method that fuses sparse features on the basis of statistical features, has better discrimination for interference waves at different levels, trains the model by collecting data in various environments, improves its adaptability to various adverse environments, is more conducive to quickly and accurately automatically identifying these interference waves, helps to reduce the trouble of subsequent analysis, improves the analysis speed and accuracy of subsequent programs, and is also an effective method to improve the analysis efficiency of cardiologists. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 is a flowchart of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0059] The present invention will be described in detail below with reference to the accompanying drawings: As Figure 1 shown, a method for identifying interference waves in ECG signals by fusing sparse features adds sparse features on the basis of general statistical features, and then uses the LightGBM algorithm to train the model. The method includes the following steps:

[0060] (1) Receive ECG data, perform convolution smoothing and normalization processing on the ECG data, then perform heartbeat localization detection and calculate the RR interval according to the heartbeat localization, then set the RR interval threshold range, identify interference waves for the heartbeat data according to the distribution of the RR interval in the ECG data, and label the data according to the interference wave identification result and the heartbeat data;

[0061] (2) Calculate the numerical statistical features of the ECG signal according to the labeled ECG signal data;

[0062] (3) Select a sparse dictionary according to the labeled ECG signal data;

[0063] (4) Solve the sparse coefficients according to the above sparse dictionary to obtain sparse features;

[0064] (5) Use the LightGBM algorithm to train an interference wave identification model according to the numerical statistical features and sparse features of the above ECG signal;

[0065] (6) According to the above interference wave identification model, apply it to the data of the validation set, evaluate the performance of the model using specificity and F1 value, and select the one with excellent performance as the final classification model.

[0066] In step (1), the smoothing and normalization processing of the ECG data and the labeling of the ECG data specifically include:

[0067] (3) Perform smoothing and normalization processing on the 3-minute ECG data;

[0068] The convolution smoothing calculation formula is:

[0069]

[0070] where f(x) and g(x) are two integrable functions on the Euclidean space R, and t ∈ R is the integration variable;

[0071] The normalization calculation formula is:

[0072]

[0073] where μ is the mean of the data x, and σ is the standard deviation of the data x;

[0074] (4) Intercept single heartbeats from the 3 - second electrocardiogram data of the slice according to the RR interval. If the length of the single heartbeat is less than 750, fill it with 0; The specific calculation of the numerical statistical features according to the steps in (2) includes:

[0075] (1) According to the electrocardiogram signal data, calculate numerical statistical features such as the average value, standard deviation, peak rate, difference between the sum of the maximum amplitudes and the sum of the minimum amplitudes, non - zero value, and median value of the 3 - second data. Their calculation formulas are respectively:

[0076] Calculate the average value avg of the electrocardiogram signal, and its calculation formula is:

[0077]

[0078] where x i is the i - th value of the electrocardiogram signal data within 3 minutes, and N is the number of x i ;

[0079] Calculate the standard deviation std of the electrocardiogram signal data;

[0080]

[0081] where μ is the average value of the electrocardiogram signal data within 3 minutes;

[0082] Calculate the peak rate get_PR, and its calculation formula is:

[0083]

[0084] where x i is the i - th value of the electrocardiogram signal data within 3 minutes, and std is the standard deviation of the electrocardiogram data;

[0085] Calculate the difference get_AD between the sum of the maximum amplitudes and the sum of the minimum amplitudes, and its calculation formula is:

[0086]

[0087] where x i is the i-th value of the electrocardiogram signal data within 3 minutes, and std is the standard deviation of the electrocardiogram data;

[0088] Calculate the non-zero value get_zero, and the calculation formula is:

[0089]

[0090] where x i is the i-th value of the electrocardiogram signal data within 3 minutes;

[0091] Calculate the median value median, and the calculation formula is: Arrange the 750 values of the given electrocardiogram signal data from smallest to largest, and take the average of the middle two numbers;

[0092] (2) Use the random forest model to rank the importance of the 6 statistical value features, and determine the weights of each feature according to the obtained ranking results.

[0093] In the step (3), the selection of the sparse dictionary for the electrocardiogram signal data specifically includes:

[0094] (3) Select the level dictionary at each level, calculate the k training samples with the largest similarity using the Euclidean distance formula, and obtain the class sub-dictionary of a certain class, denoted as X i ;

[0095] The Euclidean distance formula is:

[0096]

[0097] where n is the length of the data;

[0098] (4) Merge the class sub-dictionaries of all classes to obtain the large dictionary X: X = [X 1 , X 2 , X 3 , X 4 .

[0099] In the step (4), calculating the sparse features according to the sparse dictionary specifically lies in:

[0100] Obtain the sparse representation coefficient according to the OMP algorithm, so as to obtain the sparse representation feature α;

[0101] The solution steps of the OMP algorithm:

[0102] a. Input: original signal y, dictionary matrix X, sparse coefficient α, residual vector r, active set D X

[0103] b. Initialization: r 0 = y,

[0104] c. Find the principle with the largest inner product with r and add it to the active set

[0105] d. Solve for D using the least squares method X Update the residual and coefficients for the best approximation of the current residual

[0106]

[0107]

[0108] e. Repeat steps c and d until the termination condition is reached. The termination condition is the number of non-zero elements in α or the threshold set for the residual

[0109] In step (5), the numerical statistical features and sparse features are respectively used to train models with the LightGBM algorithm, and the two models are fused using the weighted average method

[0110] Use the data of the test set to test the obtained model, and obtain the model with the best performance as the final interference wave recognition model Specific embodiments

[0112] The specific operation process of the method for identifying interference waves in ECG signals by fusing sparse features according to the present invention is as follows

[0113] 1. Data extraction

[0114] (1) Collect electrocardiogram data with various degrees of interference. The data comes from a variety of different people, such as different genders, ages, occupations, etc., about 20,000 copies, ensuring the diversity of the data. At the same time, limit the voltage value with an absolute value greater than 1000 to 1000, thereby preventing the excessive influence of individual large values

[0115] (2) Slice the electrocardiogram data into segments with a length of 3 seconds. The minimum scale is 4 ms, that is, one piece of data is a one-dimensional vector with a length L1 of 750, and it is labeled y

[0116] 2. Data preprocessing

[0117] (1) Perform convolution smoothing on the electrocardiogram data. The convolution calculation formula is

[0118]

[0119] where f(x) and g(x) are two integrable functions in the Euclidean space R, and t ∈ R is the integration variable

[0120] (2) Perform normalization on the electrocardiogram data. The normalization calculation formula is

[0121]

[0122] Among them, μ is the mean of the data x, and σ is the standard deviation of the data x.

[0123] 3. Dataset division:

[0124] The dataset is divided into training set and test set samples, and the ratio is set to 7:3;

[0125] 4. Classification feature calculation:

[0126] (1) Calculate statistical value features:

[0127] Calculate the average value avg of the electrocardiogram signal, and its calculation formula is:

[0128]

[0129] Among them, x i is the i-th value of the electrocardiogram signal data within 3 minutes, and N is the number of x i ;

[0130] Calculate the standard deviation std of the electrocardiogram signal data;

[0131]

[0132] Among them, μ is the average value of the electrocardiogram signal data within 3 minutes;

[0133] Calculate the peak rate get_PR, and its calculation formula is:

[0134]

[0135] Among them, x i is the i-th value of the electrocardiogram signal data within 3 minutes, and std is the standard deviation of the electrocardiogram data;

[0136] Calculate the difference get_AD between the sum of the maximum amplitudes and the sum of the minimum amplitudes, and its calculation formula is:

[0137]

[0138] Among them, x i is the i-th value of the electrocardiogram signal data within 3 minutes, and std is the standard deviation of the electrocardiogram data;

[0139] Calculate the non-zero value get_zero, and its calculation formula is:

[0140]

[0141] Among them, x i is the i-th value of the electrocardiogram signal data within 3 minutes;

[0142] Calculate the median value, and the calculation formula is as follows: Arrange the 750 values of the given electrocardiogram signal data in ascending order, and take the average of the two middle numbers.

[0143] (2) Calculate the sparse representation features, and the calculation steps are as follows:

[0144] 1) Smooth the electrocardiogram signal data in the training samples and normalize it;

[0145] 2) Select a certain number of samples for each type of interference wave in the electrocardiogram signal in the training samples to form a dictionary. The reasons for selection are as follows:

[0146] The electrocardiogram signal itself has sparsity. The sparsity of the electrocardiogram signal is calculated by the non-zero values of the statistical features. The sparse features are that their values are 0 for most features, and only a small number of features are non-zero. However, when the electrocardiogram wave is the waveform of ventricular premature beats and the wave with large interference, the non-zero value (sparsity) features of the two are not significantly distinguishable. Therefore, the present invention uses sparse representation to extract more specific sparse features. Sparse representation is to express most or all of the original signals by the linear combination of fewer basic signals. Among them, the basic signals of the same category are not 0, and most of the basic signals of different categories are 0. Any signal has different sparse representations under different basic signal groups. Therefore, selecting a suitable dictionary plays an important role in converting the signal into a suitable sparse representation. The dictionary includes two design methods: a. Fixed dictionary. Select from the known transformation bases, such as wavelet bases, DCT, etc. However, these dictionaries are limited to specific types of signals and cannot be applied to new and arbitrary types of signals and cannot be adaptively represented. b. Learned dictionary. First, a training signal set needs to be constructed, and then an empirical learned dictionary is constructed, that is, potential atoms are generated from the empirical data without passing through a theoretical model. Such a dictionary can be actually applied as a fixed or redundant dictionary. Since the length of the electrocardiogram wave data for 3 seconds is 750, including at least one or more heartbeats, if the 3-second electrocardiogram data is directly used as the training dictionary without processing, the electrocardiogram data with a length of 750 are orthogonal to each other during the experiment, and the calculation amount is too large. Therefore, the 3-second electrocardiogram data is processed with a single heartbeat, and the length is still 750, and the rest of the data is filled with 0. After the single-heartbeat processing, the number of electrocardiogram data is much larger than the original number, that is, the original base dictionary. The learned dictionary is obtained by training and learning a large number of data similar to the target data. The level of the interference wave is 4. Therefore, the present invention selects the level dictionary in each interference level respectively to obtain the class sub-dictionary X of a certain level i , and then merge the class sub-dictionaries of all levels to obtain the large dictionary X.

[0147] The method for selecting the level dictionary is as follows:

[0148] Calculate the Euclidean distance between electrocardiogram signal data to measure the similarity between two signals a and b. The distance formula is:

[0149]

[0150] where n is the length of the data.

[0151] Calculate the similarity between the test sample y and the training sample X of the i-th interference level, and sort them from large to small: N :

[0152] Rank{d(y, x n ), n = 1,..., N}

[0153] Then, select the first k training samples corresponding to the similarities in the i-th level dictionary, denoted as X i .

[0154] b Combine the large dictionary X: X = [X 1 , X 2 , X 3 , X 4

[0155] 1) Solve the sparse coefficients of the training samples and test samples respectively from the dictionary through the orthogonal matching pursuit algorithm to obtain the sparse representation feature α.

[0156] The sparse algorithm optimization model is where X is the dictionary, α is the sparse coefficient, y is the original signal, ||·|| 0 is the l 0 norm, representing the number of non-zero elements in the vector, which is an NP-hard problem. According to the definition of sparsity, the regularization term can be replaced by the l 1 norm instead of the l 0 norm, and the algorithm optimization model can be expressed as where λ is the regularization parameter, which can balance the accuracy of the representation and the sparsity of the solution. Generally, a larger λ will produce a more sparse solution. The present invention uses the Orthogonal Matching Pursuit (OMP) algorithm to solve this sparse model. The solution steps are as follows:

[0157] a Input: original signal y, dictionary matrix X, sparse coefficient α, residual vector r, active set D X ;

[0158] b Initialization: r 0 = y,

[0159] c Find the principle of the largest inner product with r and add it to the active set;

[0160] ​d Solve for D using the least squares method X Find the best approximation of the current residual, and update the residual and coefficients;

[0161]

[0162]

[0163] e Repeat steps c and d until the termination condition is reached. Either the number of non-zero elements in α or the threshold set for the residual can be used as the termination condition.

[0164] 5. Feature fusion and construction:

[0165] Use avg, std, get_PR, get_AD, get_zero, median, and the sparse representation feature α to 0, 1, 2, 3 as labels for the LightGBM classification model calculation.

[0166] Among them, the first 6 features are statistical value features, which are different forms of features from the sparse representation feature α. The following strategy is used to fuse the two types of features:

[0167] 1) Use the random forest model to rank the importance of the 6 statistical value features, and determine the weights of each feature according to the obtained ranking results;

[0168] 2) The dimensions of the statistical value features and the sparse representation features are inconsistent, so the sparse representation features are used to train the model alone.

[0169] 6. Model training and fusion:

[0170] Calculate the 7 features obtained in step 4 for all electrocardiogram data. Integrate the 6 statistical value features into a vector and the sparse representation value feature as the input of the multi-classification model, and use 0, 1, 2, 3 levels as the output of the multi-classification model. Use LightGBM to train the data to obtain 2 models. For the prediction results of these 2 models, use the weighted average method, and the weights of the prediction results of the 2 models are set according to the model scores. Finally, obtain the fusion model, which predicts the probability that the electrocardiogram data is the interference wave level.

[0171] 7. Model evaluation and optimization:

[0172] Based on the parameters of the prediction model in step 6, apply them to the test set samples respectively, and evaluate the performance of the model based on the F1 score combined with accuracy and sensitivity:

[0173]

[0174]

[0175] TP (True Positive): The prediction is positive and the actual value is also positive;

[0176] FP (False Positive): The prediction is positive, but the actual value is negative;

[0177] TN (True Negative): The prediction is negative and the actual value is also negative;

[0178] FN (False Negative): The prediction is negative, but the actual value is positive.

[0179] 8. Model test results:

[0180] After partitioning the dataset as described above, the test set contains a total of 18,490 data points, including 5,427 at level 0, 5,149 at level 1, 4,402 at level 2, and 3,512 at level 3. The test results are shown in Table 1, and the confusion matrix results are calculated according to the formula described above. The most suitable weights are 0.3 for the statistical value feature model and 0.7 for the sparse representation feature model.

[0181]

[0182] Table 1 Test Results

[0183] As described above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.

Claims

1. An ECG signal interference wave recognition method integrating sparse features, characterized in that: Based on general statistical features, sparse features are added, and then the LightGBM algorithm is used to train the model. This method includes the following steps: (1) Receive electrocardiogram data, perform convolution smoothing and normalization on the electrocardiogram data, then perform heartbeat localization detection and calculate the RR interval according to the heartbeat localization, then set the RR interval threshold range, identify interference waves for the heartbeat data according to the distribution of the RR interval in the electrocardiogram data, and label the data according to the interference wave recognition result and the heartbeat data; (2) According to the labeled electrocardiogram signal data, calculate the numerical statistical features of the electrocardiogram signal. The specific calculation of the numerical statistical features includes: 1) According to the electrocardiogram signal data, calculate the average value, standard deviation, peak rate, difference between the sum of the maximum amplitudes and the sum of the minimum amplitudes, number of non-zero values, and median value of the 3-second data. The calculation formulas are as follows: Calculate the average value avg of the electrocardiogram signal, and its calculation formula is: where x i is the i-th value of the electrocardiogram signal data within 3 min, and N is the number of x i . Calculate the standard deviation std of the electrocardiogram signal data; where μ is the average value of the electrocardiogram signal data within 3 minutes; Calculate the peak rate get_PR, and the calculation formula is: where x i is the i-th value of the electrocardiogram signal data within 3 min, and std is the standard deviation of the electrocardiogram data; Calculate the difference get_AD between the sum of the maximum amplitudes and the sum of the minimum amplitudes, and the calculation formula is: where x i is the i-th value of the electrocardiogram signal data within 3 min, and std is the standard deviation of the electrocardiogram data; Calculate the number of non-zero values get_zero, and the calculation formula is: where x i is the i-th value of the electrocardiogram signal data within 3 minutes; Calculate the median value median, and the calculation formula is: Arrange the 750 numerical values of the given electrocardiogram signal data from smallest to largest, and take the average of the middle two numbers; 2) Use the random forest model to rank the importance of the 6 statistical value features, and determine the weights of each feature according to the obtained ranking results; (3) According to the labeled electrocardiogram signal data, select a sparse dictionary. The specific steps include: 1) Select the level dictionary at each level, calculate the k training samples with the largest similarity using the Euclidean distance formula, and obtain the class sub-dictionary of a certain class, denoted as X i ; The Euclidean distance formula is: where n is the length of the data; 2) Merge the class sub-dictionaries of all classes to obtain the large dictionary X: X = [X 1 , X 2 , X 3 , X 4 ; (4) Solve the sparse coefficients according to the above sparse dictionary to obtain sparse features; (5) According to the numerical statistical features and sparse features of the above electrocardiogram signal, use the LightGBM algorithm to train the interference wave recognition model; (6) According to the above interference wave recognition model, apply it to the data of the validation set, evaluate the performance of the model using specificity and F1 value, and select the one with excellent performance as the final classification model.

2. The ECG signal interference wave recognition method integrating sparse features according to claim 1, characterized in that: In the step (1), the electrocardiogram data is smoothed and normalized, and the electrocardiogram data is labeled, specifically including: (1) Perform smoothing and normalization on the 3-minute electrocardiogram data; The convolution smoothing calculation formula is: where f(x), g(x) are two integrable functions on the Euclidean space R, and t ∈ R is the integration variable; The normalization calculation formula is: where μ is the average value of the data x, and σ is the standard deviation of the data x; (2) Intercept single heartbeats for the sliced 3-second electrocardiogram data according to the RR interval. If the length of the single heartbeat is less than 750, fill it with 0.

3. The ECG signal interference wave recognition method integrating sparse features according to claim 1, characterized in that: In step (4), calculating the sparse features according to the sparse dictionary specifically involves: Obtaining the sparse representation coefficients according to the OMP algorithm, thereby obtaining the sparse representation feature α; Steps for solving by the OMP algorithm: a. Input: original signal y, dictionary matrix X, sparse coefficient α, residual vector r, active set D X b. Initialization: r 0 = y, c. Finding the principle of the largest inner product with r and adding it to the active set d. Solve for D using the least squares method X Optimize the approximation of the current residual, update the residual and coefficients e. Repeating c and d until the termination condition is reached. The number of non-zero elements in α or the threshold set for the residual is used as the termination condition.

4. The method for identifying interference waves in ECG signals by fusing sparse features according to claim 1, characterized in that: In step (5), the numerical statistical features and the sparse features are respectively used to train models by the LightGBM algorithm, and the two models are fused by the weighted average method.

5. The method for identifying interference waves in ECG signals by fusing sparse features according to claim 1, characterized in that: Using the data of the test set to test the obtained model, and obtaining the model with the best performance as the final interference wave identification model.

Citation Information

Patent Citations

  • Electrocardiosignal identity recognition method and system with multi-feature information being fused

    CN112818315A

  • Electrocardiogram (ECG)-based authentication apparatus and method thereof, and training apparatus and method thereof for ECG-based authentication

    US20160232340A1