Myocardial infarction detection and positioning auxiliary diagnosis system and device
By designing the SRTNet model in the detection and positioning tasks of myocardial infarction, the temporal and spatial characteristics of the electrocardiogram were extracted from the three levels of scanning, reading and thinking modules, the problems of data imbalance and individual differences were solved, and the detection and positioning performance was improved.
Patent Information
- Application Number
- CN202311502586.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-13
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art has problems with data imbalance and individual differences in myocardial infarction detection and location tasks, resulting in poor performance of models in patients' programs.
A myocardial infarction detection and positioning assisted diagnosis system was designed to extract temporal and spatial features from three levels: scanning module, reading module and thinking module (named "SRTNet model") to alleviate the overfitting of individual variability characteristics of the model.
It effectively alleviates the overfitting of individual differentiated characteristics of the model, improves the performance of myocardial infarction detection, reduces the number of parameters of the model, reduces the risk of overfitting, and improves the accuracy of myocardial infarction localization.
Smart Images

Figure BDA0004544722730000051 
Figure BDA0004544722730000081 
Figure BDA0004544722730000091
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of electrocardiogram signal analysis, and more specifically relates to a myocardial infarction detection and positioning auxiliary diagnosis system and device. Background Art
[0002] The incidence of myocardial infarction is increasing year by year. It develops quickly and has a high mortality rate. Therefore, timely detection and location of myocardial infarction plays an important role in saving patients' lives. The myocardial infarction detection and location task based on the inter-patient scheme is more in line with the clinical diagnosis of the hospital, but the examination effect is not good. Although a lot of work has been done on the myocardial infarction detection and location task based on electrocardiogram (ECG), due to the patient-specific information between patients and the inter-patient scheme that may further aggravate the data imbalance, the existing deep learning method uses a simple combination of multiple inter-lead features to classify myocardial infarction, but ignores the correlation between lead signals, and there is a problem of extracting a relatively single ECG feature, which cannot overcome the individual differences between patients, resulting in poor performance in the inter-patient scheme, and the inter-patient myocardial infarction detection and location task is more difficult. Summary of the invention
[0003] In view of the deficiencies in the above-mentioned prior art, the present invention provides a myocardial infarction detection and positioning auxiliary diagnosis system and device, which aims to extract time and space features from three levels: scanning module, reading module and thinking module (named "SRTNet model") by imitating the process of doctors diagnosing electrocardiograms, effectively alleviating the overfitting of individual difference characteristics of the model, and solving the problem of difficulty in myocardial infarction detection and positioning tasks among patients caused by imbalance of patient-specific information and scheme data among patients.
[0004] The specific technical solutions are as follows:
[0005] One of the purposes of the present invention is to provide a myocardial infarction detection and positioning auxiliary diagnosis system, which is different from the prior art in that it includes a data preprocessing module, a scanning module, a reading module, a thinking module and a regression classification module.
[0006] A) Data preprocessing module, used for preprocessing the ECG signal input from the external interface, the preprocessing includes reading data, denoising, downsampling, R wave detection, training sample segmentation and data enhancement.
[0007] Specifically, the denoising refers to removing noise and baseline drift, which is to reduce the interference of baseline drift and noise of electrocardiogram on the performance of myocardial infarction detection.
[0008] Among them, the denoising method is preferably a wavelet transform method using DB6 as the mother wavelet.
[0009] The R wave detection refers to checking the R wave of the electrocardiogram signal, and the detection method is preferably the Pan-Tompkins algorithm.
[0010] The training sample segmentation refers to selecting sampling points before and after the R wave detection point, and the length of the sampling points is preferably 2000 sampling points, namely 299 points before the R wave detection point and 1700 points after the QRS detection point.
[0011] The data enhancement refers to fine-tuning the window for segmenting samples during sampling, changing the number of sampling points before and after the R wave and maintaining the length of the sampling points.
[0012] B) A scanning module designs an independent branch network for each ECG lead so that different channels have different modes for learning shallow waveform features of each lead; the branch network is composed of a convolutional network; the branch network has the same network structure and is used to independently learn independent features that are not related to the lead; the convolutional network uses independent convolution kernels to process different leads.
[0013] Since the 12-lead electrocardiogram represents the changes in electrical signals of cardiac activity from different directions, different channels have different modes. The present invention creatively designs an independent branch for each lead, called a scanning module, to learn the shallow waveform features of each lead. Considering the powerful representation ability of convolutional neural networks, a 12-branch network is constructed using a convolutional neural network to learn independent features that are not related to the leads. The 12 branch networks have the same network structure, but use independent convolution kernels to process different leads. Because only a sufficiently long electrocardiogram can represent specific cardiac activity, the present invention creates a convolution operation using a larger convolution kernel so that the network obtains a larger receptive field.
[0014] The beneficial effect of adopting the above scheme is that it can make the characteristics of the samples scenario-based, while reducing the parameters of the model and alleviating the overfitting caused by over-parameterization of the model.
[0015] Furthermore, the convolutional network of the scanning module includes a 1D convolutional layer, a pooling layer, a batch normalization layer and an activation layer; the 1D convolutional layer obtains the spatial features of the input data through different convolution kernel mappings; the pooling layer uses maximum pooling after each convolutional layer to retain the significant features in the feature map; the batch normalization layer is used to maintain the stability of the back-propagation gradient; the activation layer uses ReLU to activate the nonlinear function.
[0016] Furthermore, the 1D convolution layer is two layers.
[0017] Furthermore, the convolution kernel shape is designed in a manner that the sampling rate decreases.
[0018] Specifically, the convolutional network operation of the scanning module can be expressed as the following formula:
[0019] F(x)=P(BN(Rf(x i , w i ))))
[0020] in,
[0021] xi and F(x) are the i-th input and output of the scanning module;
[0022] w i The parameters of the 1D convolution kernel are updated by the backward pass of the network;
[0023] R represents the activation function ReLU;
[0024] BN stands for batch normalization;
[0025] P stands for max pooling.
[0026] C) Reading module, which is used to learn mid-level features across leads, including dense blocks and transition layers.
[0027] The waveforms of the leads in the ECG signal have a strong correlation because the 12-lead ECG is designed to describe the heart activity pattern from different directions. Previous deep learning methods used a simple combination of multiple lead features to classify myocardial infarction, but ignored the correlation of the signals between the leads. According to the correlation characteristics of the information between the leads, the present invention creates and discloses a cross-lead middle-level feature learning module, named the reading module. Due to the insufficient representation of the ECG by the shallow convolutional neural network, the deep network is prone to overfitting. In order to overcome the overfitting problem, the reading module consists of dense blocks and transition layers.
[0028] The beneficial effect of adopting the above scheme is that it can effectively alleviate the vanishing gradient problem, and play a role in strengthening feature propagation and encouraging feature reuse.
[0029] Furthermore, the dense block consists of convolutional layers.
[0030] Specifically, the reading module contains three dense blocks, each of which is composed of convolutional layers. The third dense block contains seven convolutional layers, and the rest contain two convolutional layers. For each layer, the feature map of the previous layer is used as input, and the feature map obtained in this layer is also used as input for the subsequent layer.
[0031] Furthermore, the convolution layer uses group convolution, 12 groups, 12 output channels, and the kernel size is preferably 3x1. A transition layer is used between two dense blocks to reduce feature dimensions and compression. The reading module dense block can be expressed as the following formula:
[0032] F i=BN(R(f(x i ,w i ))) (1)
[0033] x i =C(F i-1 ,F i-2 ,…,F0) (2)
[0034] D=C(F n ,F n-1 ,…,F0) (3)
[0035] in,
[0036] F i and x i represents the i-th layer output and input of the dense block;
[0037] w i is the weight of the convolution kernel;
[0038] x i It is the convolution output from layer 0 to layer i-1;
[0039] C represents the connection operation;
[0040] D represents the total output feature map of the dense block.
[0041] Furthermore, the transition layer includes a 1x1 convolution and a maximum pooling layer, which plays a role in reducing feature dimension and compressing features.
[0042] Furthermore, the scanning module and the reading module jointly characterize the spatial information, including the waveform of each lead and the correlation characteristics between the leads.
[0043] D) Thinking module, which is a transformer encoder, is used to integrate the global information of the electrocardiogram and analyze the relationship between time and morphology of different abnormal waves; the thinking module includes a self-attention mechanism layer and a feedforward network layer.
[0044] Doctors use single-lead waveforms as the basic unit of receptive fields when diagnosing myocardial infarction. After reading the 12-lead electrocardiogram, the receptive field is placed on the 12-lead to further analyze the performance of abnormal waveforms in different leads. Finally, the global information of the electrocardiogram is integrated to analyze the temporal and morphological relationships between different abnormal waves. Inspired by the doctor's receptive field, the present invention creates and designs a thinking module.
[0045] Specifically, the spatial feature vector along the time axis is used as the embedding vector of the thinking module, and the attention parameters and feedforward neural network parameters are set.
[0046] Among them, the thinking module includes two layers of encoders, the word vector length parameter is preferably 84, the multi-head attention parameter is preferably 3, and the feedforward neural network intermediate neuron parameter is preferably 200.
[0047] E) Regression classification module, which performs global average pooling (GAP) operation on the feature vector and uses fully connected layers and softmax layers to generate probability distributions of 2-class outputs or 6-class outputs for detection and localization of myocardial infarction.
[0048] In order to alleviate the overfitting phenomenon of the network, the feature vector is subjected to global average pooling (GAP) operation. Global average pooling can effectively reduce the number of neurons in the fully connected layer, thereby reducing the number of network parameters and speeding up network training and detection efficiency.
[0049] Specifically, the feature map obtained from the Transformer model is 84×62. The global average takes the average value of the feature map by dimension to obtain the feature vector 1×62. The fully connected layer maps 62 features to 2 or 6 categories; the probability of each category is obtained through the SofMax regression function.
[0050] Furthermore, the SofMax regression function is expressed using the following formula:
[0051]
[0052] in,
[0053] x i is the i-th value after full-connection mapping;
[0054] C is the number of output nodes, and also represents the number of classifications;
[0055] The array of nodes can be normalized to the interval [0,1] through the SoftMax function to obtain the probability of belonging to each category. The present invention uses the maximum probability as the predicted label.
[0056] A second object of the present invention is to provide a myocardial infarction detection and positioning auxiliary diagnosis device including the above system.
[0057] The beneficial effects of the present invention are as follows:
[0058] (1) Compared with the prior art, the design inspiration of the present invention comes from the process of doctors diagnosing electrocardiograms. First, the doctor will browse the waveform of the 12-lead electrocardiogram signal, then focus on the changes in the waveform between the 12 leads, and finally summarize the waveforms before and after the 12 leads to make a diagnosis. The scanning module (scanning), reading module (reading) and thinking module (thinking) (named "SRTNet model") of the present invention learn electrocardiogram features from three different levels. First, the scanning module learns the independence characteristics of the 12-lead electrocardiogram, and then the reading module focuses on the related characteristics of different leads. The two jointly characterize the spatial characteristics of the 12-lead electrocardiogram, and finally the thinking module mines the characteristics of the time dimension. By extracting time and space features, the model can effectively alleviate overfitting of individual difference characteristics and improve detection performance.
[0059] (2) Compared with the prior art, the present invention designs an independent branch network for each lead to learn the shallow waveform features of each lead. Each branch network has the same network structure, but uses independent convolution kernels to process different leads. This design can make the sample features scenario-based, while reducing the parameters of the model and alleviating overfitting caused by over-parameterization of the model.
[0060] (3) Compared with the prior art, due to the insufficient representation of the electrocardiogram by the shallow convolutional neural network, the deep network is prone to overfitting. The reading module of the present invention includes dense blocks and transition layers, which can effectively alleviate the vanishing gradient problem and play a role in strengthening feature propagation and encouraging feature reuse.
[0061] (4) Compared with the prior art, the convolutional neural network-based model only focuses on the spatial waveform characteristics of the ECG signal and lacks analysis of different waveforms. The thinking module of the present invention is composed of a Transformer model, which can integrate the global information of the ECG and analyze the temporal and morphological relationships between different abnormal waves. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 It is a schematic diagram of the structure of the myocardial infarction detection and positioning auxiliary diagnosis system of the present invention;
[0063] Figure 2 It is a data preprocessing flow chart of the present invention;
[0064] Figure 3 It is a schematic diagram of the spatial information representation structure of the present invention;
[0065] Figure 4 It is a schematic diagram of the thinking module structure of the present invention;
[0066] Figure 5 It is a schematic diagram of parameter settings of the scanning module, reading module and thinking module of the present invention;
[0067] Figure 6 This is a comparison diagram of different convolution kernel experiments of the present invention;
[0068] Figure 7 This is a comparative test diagram of different network improvements of the present invention. DETAILED DESCRIPTION
[0069] The principles and features of the present invention are described below in conjunction with examples. The examples are only used to explain the present invention and are not used to limit the scope of the present invention.
[0070] Example 1
[0071] A myocardial infarction detection and positioning auxiliary diagnosis system includes a data preprocessing module, a scanning module, a reading module, a thinking module and a regression classification module, such as Figure 1 shown.
[0072] A) Data preprocessing module
[0073] The data preprocessing module is used to read data, denoise, downsample, detect R waves, segment training samples and enhance data for ECG signals input from the external interface. The preprocessing process is as follows: Figure 2 shown.
[0074] The wavelet transform method with DB6 as the mother wavelet was used to remove noise and baseline offset. In order to speed up the detection speed of the model, the model was trained using 2 times the downsampled data. Since the PTB dataset is not marked with R waves, the Pan-Tompkins algorithm was used to check the R wave of the ECG signal. The specific operation of sample segmentation is: the length of each sample is selected to be 2000 sampling points, and the 2000 sampling points come from 299 points before the R wave detection point and 1700 points after the QRS detection point. The data of the training set is data enhanced. The data enhancement method is to fine-tune the window of the segmented sample during sampling, change the number of sampling points before and after the R wave, and keep the length of 2000 sampling points.
[0075] B) Scanning Module
[0076] The scanning module designs an independent branch network for each ECG lead so that different channels have different modes, which are used to learn the shallow waveform characteristics of each lead; the branch network is composed of convolutional networks; the branch networks have the same network structure and are used to independently learn independent features that are not related to the leads; the convolutional networks use independent convolution kernels to process different leads.
[0077] The branch network consists of a 1D convolution layer, a pooling layer, a batch normalization layer, and an activation layer. The convolution layer obtains the spatial features of the input data through different convolution kernel mappings. Pooling is a special convolution operation that plays a role in feature selection. After each convolution layer, maximum pooling is used to retain the salient features in the feature map. The activation layer uses ReLU to introduce nonlinear factors to solve the problem that the linear model cannot fit the nonlinearity. The batch normalization layer is used to maintain the stability of the back propagation gradient. The scanning module includes two convolution layers, and the detailed parameters are shown in Table 1. After the pooling layer, the length of the feature map will be compressed, so the shape of the convolution kernel is designed in a way that the sampling is reduced.
[0078] Table 1 Specific parameters of the branch network structure
[0079]
[0080] The convolutional network operation can be expressed as the following formula:
[0081] F(x)=P(BN(R(f(x i , w i ))))
[0082] Among them, x i and F(x) are the i-th input and output of the scanning module;
[0083] w i The parameters of the 1D convolution kernel are updated by the backward pass of the network;
[0084] R represents the activation function ReLU;
[0085] BN stands for batch normalization;
[0086] P stands for max pooling.
[0087] C) Reading Module
[0088] The reading module is used to learn mid-level features across leads and consists of dense blocks and transition layers.
[0089] The dense block contains multiple convolutional layers. For each layer, the feature map of the previous layer is used as input, and the feature map obtained in this layer is also used as input for the subsequent layer. The transition layer only contains 1x1 convolution and maximum pooling layers, which play the role of reducing feature dimension and compressing features.
[0090] The scanning module and reading module are implemented by convolutional neural networks to characterize the spatial features of the 12-lead electrocardiogram. Figure 3 The characterization process of the spatial features of the 12-lead electrocardiogram is depicted, including the waveform of each lead and the correlation features between leads.
[0091] The reading module contains three dense blocks. The third dense block contains seven convolutional layers, and the rest contain two convolutional layers. The convolutional layer uses group convolution, 12 groups, 12 output channels, and a kernel size of 3x1. A transition layer is used between two dense blocks to reduce feature dimensions and compression. The specific parameters are shown in Table 2.
[0092] Table 2 Specific parameters of the reading module
[0093]
[0094] The operation of dense blocks can be summarized as follows:
[0095] F i =BN(R(f(x i ,w i ))) (1)
[0096] x i =C(F i-1 ,F i-2 ,…,F0) (2)
[0097] D=C(F n ,x n ) (3)
[0098] in,
[0099] F i and x i represents the i-th layer output and input of the dense block;
[0100] w i is the weight of the convolution kernel;
[0101] x i It is the convolution output from layer 0 to layer i-1;
[0102] F n It is the output of the last layer of dense fast;
[0103] X n It is the input of the last layer of dense fast;
[0104] C represents the connection operation;
[0105] D represents the total output feature map of the dense block.
[0106] D) Thinking module
[0107] The thinking module is a transformer encoder, which includes a self-attention mechanism layer and a feedforward network layer. Figure 3 Each column of the output feature map can be regarded as the characteristics of different waveforms of 12 leads (t0~t n Indicates the time axis; x0~xn Represents the embedded vector). Taking the time axis as the unit, each column of the spatial feature map is used as the input vector of the thinking module.
[0108] The spatial feature vector along the time axis is used as the embedding vector of the thinking module. The structure of the thinking module is as follows Figure 4 The thinking module contains two layers of encoders, the word vector length parameter is 84, the multi-head attention parameter is 3, and the intermediate neuron parameter of the feedforward neural network is 200. The spatiotemporal feature vector is compressed by global pooling to prevent the model from overfitting. The detailed parameters of the scanning module, reading module and thinking module are as follows Figure 5 shown.
[0109] E) Regression and classification module
[0110] The regression classification module performs a global average pooling (GAP) operation on the feature vector and uses a fully connected layer and a softmax layer to generate a probability distribution of 2-class outputs or 6-class outputs for the detection and location of myocardial infarction.
[0111] The feature map obtained from the Transformer model is 84×62. The global average takes the average value of the feature map according to the dimension to obtain the feature vector 1×62. The fully connected layer maps the 62 features to 2 or 6 categories. Finally, the probability of each category is obtained through the SofMax regression function. SofMax can be expressed using the following formula.
[0112]
[0113] in,
[0114] x i is the i-th value after full-connection mapping;
[0115] C is the number of output nodes, and also represents the number of classifications;
[0116] The SoftMax function can normalize the node array to the interval [0,1] to obtain the probability of belonging to each category, and the maximum probability is the predicted label.
[0117] Example 2
[0118] The performance of Example 1 was verified under the inter-patient scheme based on the PTB data set, and the technical indicators were compared with the most advanced similar invention patents in the world. The training of this embodiment combines the loss function of FocalLoss and the extended algorithm (Adam) of stochastic gradient descent, with exponential decay and quadratic exponential decay of 0.9 and 0.999 respectively, and the initial learning rate is 0.001. The model is trained for a maximum of 200 epochs, each epoch is divided into multiple batches for training, and the number of samples in each batch is 512. The model uses random dropout technology to alleviate model overfitting, and the random dropout rate is 0.5. All embodiments are performed by PyTorch on the Ubuntu operating system and NVIDIA Tesla V100 GPU (32G).
[0119] Performance evaluation indicators:
[0120] In order to conduct a detailed performance evaluation of the experiment, this embodiment selects five commonly used evaluation indicators for medical signal classification: accuracy (Acc), sensitivity (Se), specificity (Sp), positive predictive value (PPV), and F-1 score (F-score, F1). Acc represents the probability of correct prediction in all samples, Se represents the probability of correct prediction in true positive samples, Sp represents the probability of correct prediction in true negative samples, PPV represents the probability of correct prediction in predicted positive samples, and F1 score represents the harmonic mean of the correct prediction of true positive samples and the correct prediction of predicted positive samples. All five evaluation indicators are positively correlated evaluation indicators. The larger the evaluation indicator value, the stronger the performance of the model. The calculation of the five evaluation indicators is shown in formulas (1) to (5):
[0121]
[0122]
[0123]
[0124]
[0125]
[0126] Table 3 Confusion matrix of true labels and predicted labels
[0127]
[0128] Table 3 gives the calculation methods of TP (True Positive), FP (False Positive), TN (True Negative), and FN (False Negative). TP represents the number of patients predicted as patients, FP represents the number of patients predicted as normal people, TN represents the number of normal people predicted as normal people, and FN represents the number of normal people predicted as patients.
[0129] Implementation results:
[0130] (1) Myocardial infarction detection results based on inter-patient benchmarks
[0131] The myocardial infarction detection task is a binary classification task that classifies normal people and myocardial infarction patients based on ECG signals. The data comes from the PTB dataset, including 80 records from 52 normal people and 309 records from 112 myocardial infarctions. In order to reliably verify the model, the experimental data was cross-validated 5-fold. The 5-fold cross-validation is to divide the experimental data into 5 groups, 4 of which are used for training and the rest are used for testing. The myocardial infarction detection performance indicators of this embodiment on the PTB dataset are shown in Table 4, and the confusion matrix is shown in Table 5.
[0132] Table 4 Results of 5-fold cross validation test between patients
[0133]
[0134] Table 5 Confusion matrix of myocardial infarction detection between patients
[0135]
[0136] Table 4 shows the accuracy, sensitivity, specificity and F1 score of the 5-fold cross validation of this embodiment, which are 99.25%, 99.15%, 98.58% and 99.09% respectively. It can be seen from the confusion matrix Table 5 that there are a total of 10,388 MI samples, and only 55 myocardial infarction samples are misclassified as normal samples, indicating that the model of this scheme can detect myocardial infarction very well. From the results, the false detection rate of myocardial infarction is lower than that of normal people, which is of great significance for clinical application. The F1 score of the 5-fold experiment is between 98.13% and 99.92%, which is less than 1% different from the average value, proving that the model has good detection performance. By analyzing other detection indicators of the 5-fold experiment, the maximum difference in accuracy does not exceed 1.41%, and the differences in sensitivity and specificity are very small, indicating that the model has strong stability.
[0137] (2) Myocardial infarction localization results based on inter-patient benchmarks
[0138] Myocardial infarction localization is a fine-grained classification task that further analyzes the location of myocardial infarction. This chapter selects 5 myocardial infarction subclasses with more data from the 10 types of myocardial infarction in the PTB dataset for experiments. The localization task is more difficult due to the similarity of certain waveforms between different types of myocardial infarction subclasses and the reduced amount of data available for each type of myocardial infarction. The localization task and the detection task use the same experimental data, including a total of 164 patients, of which the least type of anterior lateral wall myocardial infarction has only 16 patients, and the most is normal data, including 52 patients. Therefore, the experimental data is extremely unbalanced, making it difficult for the model to learn data features that are easy to classify.
[0139] Table 6 Results of 5-fold cross-validation positioning experiment among patients
[0140]
[0141] Table 7. The third fold cross validation positioning confusion matrix between patients
[0142]
[0143]
[0144] Combining Tables 6 and 7 to analyze the results of myocardial infarction localization, it can be seen that the effect of myocardial infarction localization is much lower than that of myocardial infarction detection. Table 6 shows that the accuracy, sensitivity, specificity and F1 score of myocardial infarction localization of the 5-fold cross-validation model of this scheme are 68.45%, 61.82%, 67.95% and 60.50%. The low performance of myocardial infarction localization is caused by multiple factors. From the analysis of the experimental results, the main reason for the low accuracy of myocardial infarction localization is the small amount of data in each category, resulting in insufficient model training. Further analysis of the experimental data shows that each group of data in the 5-fold cross-validation also has extreme data imbalance. The main reason for data imbalance is that the minimum number of MI subclasses is only a dozen people, resulting in large differences between groups after grouping. Another reason is that the number of electrocardiogram records for each patient in the PTB dataset is different. The 5-fold experimental results are very different. The misclassified samples obtained from the confusion matrix in Table 7 mainly belong to the myocardial infarction subclasses. Due to the similarity of certain bands between myocardial infarction subclasses, the classification performance of the model is reduced. (Among them, AMI and IMI are more misclassified with each other; ALMI and ILMI are more misclassified with each other.) Increasing the amount of data for each type of myocardial infarction and more efficient models may further improve the localization of myocardial infarction.
[0145] Table 8 Comparison of the most advanced inventions of the same kind before the disclosure of this invention based on the PTB dataset
[0146]
[0147]
[0148] In order to verify the effectiveness of the model proposed in this chapter, the experimental results are compared with the recent research based on the PTB dataset. The comparison results are shown in Table 8. Among all the methods, the method proposed in this scheme achieves the best myocardial infarction detection and localization performance under the inter-patient scheme. Compared with the literature [2], the detection accuracy of myocardial infarction is improved by 2.75%, the sensitivity is improved by 2.05%, and the specificity is improved by 5.24%. The localization accuracy of myocardial infarction is improved by 5.51%, and the specificity is improved by 4.95%. Compared with the literature [1], the F1 of myocardial infarction detection and localization is improved by 2.17% and 12.56% respectively. By comparison, the model of this scheme has the following advantages. First, by learning multi-level features to better fit the distribution of data. Second, by combining spatiotemporal information to extract deep features, the impact of individual differences among patients is reduced.
[0149] [1] Han, C., Shi, L., 2020. Ml–resnet: A novel network to detect and locate myocardial infarction using 12leads ecg. Computer Methods and Programs inBiomedicine 185,105138.doi:10.1016 / j.cmpb.2019.105138.
[0150] [2] Fu, L., Lu, B., Nie, B., Peng, Z., Liu, H., Pi,
[0151] [3]Sharma,L.D.,Sunkaria,R.K.,2018.Inferior myocardial infarctiondetection using stationary wavelet transform and machine learningapproach.Signal,Image andVideo Processing 12,199–206.doi:10.1007 / s11760-017-1146-z.
[0152] [4]Reasat,T.,Shahnaz,C.,2017.Detection ofinferior myocardialinfarction using shallow convolutional neural networks,in:2017 IEEE Region10Humanitarian Technology Conference(R10-HTC),IEEE.pp.718–721.doi:10.1109 / R10-HTC.2017.8289058.
[0153] [5]Y.Zhang and J.Li,“Application of heartbeat-attention mechanism fordetection ofmyocardial infarction using 12-lead ecg records,”AppliedSciences,vol.9,no.16,p.3328,2019.
[0154] [6]Han,C.,Shi,L.,2019.Automated interpretable detection ofmyocardialinfarction fusing energy entropy and morphological features.Computer Methodsand Programs in Biomedicine 175,9–23.doi:10.1016 / j.cmpb.2019.03.012.
[0155] [7] Liu, W., Wang, F., Huang, Q., Chang, S., Wang, H., He, J., 2019. Mfbcbrnn: Ahybrid network for mi detection using 12-lead ecgs. IEEE Journal ofBiomedicaland Health Informatics 24,503–514.doi:10.1109 / JBHI.2019.2910082.
[0156] [8]Cao, Y., Wei, T., Zhang, B., Lin, N., Rodrigues, JJ, Li, J., Zhang, D., 2021. Ml-net: Multi-channel lightweight network for detecting myocardialinfarction. IEEE Journal of Biomedical and Health Informatics 25,3721–3731.doi:10.1109 / JBHI.2021.3060433.
[0157] [9]Martin, H., Izquierdo, W., Cabrerizo, M., Cabrera, A., Adjouadi, M., 2021. Near real-time single-beat myocardial infarction detection from singlelead electrocardiogram using long short-term memory neural network. Biomedical Signal Processing and Control 68,102683.doi:https: / / doi.org / 10.1016 / j.bspc.2021.102683.
[0158] Example 3
[0159] Verification of the advantages of various technical features created by the present invention
[0160] This embodiment conducts three sets of ablation experiments: how the receptive field affects the model; how the architecture affects the model; and how the data enhancement effect of different loss functions affects the model. This is used to illustrate the advantages of the various technical features of the invention and verify the effective parameter range of the invention.
[0161] (1) Effective parameter range and optimal parameter value of the receptive field created by the present invention
[0162] A larger receptive field helps the network obtain rich information, thereby achieving better performance. Previous methods usually achieve the purpose of expanding the network's receptive field by deepening the number of network layers. However, deepening the network will lead to over-parameterization of the network, which will lead to overfitting and reduce classification performance. Because of the 1D convolution under the same receptive field, increasing the convolution kernel size of the convolution layer uses fewer parameters than deepening the network. This embodiment increases the receptive field of the network by increasing the convolution kernel of the convolution layer. Through detailed ablation experiments, the effect of the size of the convolution kernel on performance is explored, and the comparison results are shown in Figure 2. Figure 6 As shown. Using a large convolution kernel in the network helps the network focus on longer ECG information to extract the features of pathological waveforms. This embodiment only uses large kernel convolution in the scanning module to extract low-level waveform features of the ECG. Experiments show that myocardial infarction detection and positioning are best when a combination of 17×1 and 11×1 is used.
[0163] (2) Verification of the technical advantages of each technical feature of the present invention
[0164] This section introduces the experimental analysis of module ablation. This example conducts detailed ablation experiments on the three modules of scanning, reading and thinking. The experimental results of the ablation experiment are shown in Figure 2. Figure 7 Four network structures are shown: scanning module, scanning module + thinking module, scanning module + reading module, and using three modules at the same time. The reading module is introduced into the scanning module to increase the feature extraction between clues and improve the classification accuracy by 2.52%. Compared with the scanning module, the combination of scanning module + thinking module increases the feature extraction of the time dimension and improves the classification accuracy by 1.01%. The experimental results of the architecture using three modules at the same time achieved the best results, with a classification accuracy of 99.94%. Through the analysis of the experimental results, the module improvements proposed by the invention are helpful to improve the classification performance, and there is a complementary effect between the modules.
[0165] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A myocardial infarction detection and positioning auxiliary diagnosis system, characterized in that: It includes data preprocessing module, scanning module, reading module, thinking module and regression classification module; The data preprocessing module is used to preprocess the ECG signal input from the external interface; The scanning module is used to learn the shallow waveform features of each lead by designing an independent branch network for each lead; the branch network is composed of a convolutional network, and the branch network has the same network structure but uses independent convolution kernels to process different leads; The reading module is used to learn the middle-level features across leads; the reading module includes a dense block and a transition layer; The thinking module is a transformer encoder, which is used to mine the characteristics of the time dimension and analyze the relationship between the time and morphology of different abnormal waves by integrating the global information of the electrocardiogram; the thinking module includes a self-attention mechanism layer and a feedforward neural network layer; The regression classification module performs a global average pooling operation on the feature vector and uses a fully connected layer and a softmax layer to generate a probability distribution for the detection and location of myocardial infarction.
2. The myocardial infarction detection and positioning auxiliary diagnosis system according to claim 1, characterized in that: The convolutional neural network of the scanning module includes a 1D convolution layer, a pooling layer, a batch normalization layer and an activation layer; the 1D convolution layer obtains the spatial features of the input data through different convolution kernel mappings; The pooling layer is after each convolutional layer, and uses maximum pooling to retain the significant features in the feature map; The batch normalization layer is used to maintain the stability of the back-propagation gradient; the activation layer uses ReLU activation, which is a nonlinear function.
3. The myocardial infarction detection and positioning auxiliary diagnosis system according to claim 2, characterized in that: The convolutional neural network of the scanning module can be expressed as the following formula: F(x)=P(BN(R(f(x i ,w i )))) in, x i and F(x) are the i-th input and output of the scanning module; w i The parameters of the 1D convolution kernel are updated by the backward pass of the network; R represents the activation function ReLU; BN stands for batch normalization; P stands for max pooling.
4. The myocardial infarction detection and positioning auxiliary diagnosis system according to claim 1, characterized in that: The convolution kernel shape of the scanning module is designed in such a way that the sampling rate decreases.
5. The myocardial infarction detection and positioning auxiliary diagnosis system according to claim 1, characterized in that: The dense blocks of the reading module are composed of convolutional layers; the convolutional layers use group convolution; and a transition layer is used between the dense blocks to reduce feature dimensions and compression.
6. The myocardial infarction detection and positioning auxiliary diagnosis system according to claim 5, characterized in that: The reading module dense block can be expressed as the following formula: F i =BN(R(f(x i ,w i ))) (1) x i =C(x i-1 ,F i-2 ,…,F0) (2) D=Concat(F n ,x n ) (3) in, F i and x i represents the i-th layer output and input of the dense block; w i is the weight of the convolution kernel; x i It is the convolution output from layer 0 to layer i-1; C represents the connection operation; F n Represents the output of the last layer x n Represents the input of the last layer D represents the total output feature map of the dense block.
7. The myocardial infarction detection and positioning auxiliary diagnosis system according to claim 1, characterized in that: The transition layer includes a 1x1 convolution and a maximum pooling layer.
8. The myocardial infarction detection and positioning auxiliary diagnosis system according to claim 1, characterized in that: The thinking module takes the time axis as a unit and uses each column of the spatial feature map as an input vector of the thinking module.
9. The myocardial infarction detection and positioning auxiliary diagnosis system according to claim 1, characterized in that: The SofMax regression function is expressed using the following formula: in, x i is the i-th value after full-connection mapping; C is the number of output nodes, and also represents the number of classifications; The SoftMax function can normalize the node array to the interval [0,1] to obtain the probability of belonging to each category.
10. A myocardial infarction detection and positioning auxiliary diagnosis device, characterized in that: It comprises the myocardial infarction detection and positioning auxiliary diagnosis system as described in any one of claims 1 to 9.