A Three-Stage Feature Enhancement Method for Myocardial Infarction Localization

Through the three-stage feature enhancement method, including data enhancement, pre-training of minority features and weighted loss optimization, combined with the lead contribution graph and the graph space-time cross-attention network, the problem of identifying a minority sample in the myocardial infarction classification is solved, and high accuracy and robust MI positioning is achieved.

CN119488296BActive Publication Date: 2025-05-27CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510076580.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-27
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

In the multi-category myocardial infarction classification task, the classification accuracy and adaptability are poor due to the similar ECG signal characteristics of the minority and the majority of the classes and the samples are unbalanced.

Method used

A three-stage feature enhancement method is adopted, including data enhancement, pre-training of minority features and weighted loss optimization. The feature differences between the leads are dynamically introduced by constructing a lead contribution graph, and combined with the graph space-time cross-attention network and multi-branch classification structure, the target model is formed.

Benefits of technology

It significantly improves the ability to identify a few types of samples, improves the accuracy and robustness of the model in MI positioning, and effectively deals with the problem of category imbalance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119488296B_ABST
    Figure CN119488296B_ABST
Patent Text Reader

Abstract

In an embodiment of the present invention, a three-stage feature enhancement myocardial infarction localization method is provided, belonging to the field of medical technology, specifically including: sampling the original electrocardiogram data at a preset sampling rate; performing data enhancement on the sample data set to obtain a training set; using the training set for the first-stage training of the initial model to generate a lead contribution map; using the minority class samples and the lead contribution map for the second-stage training of the initial model, updating the parameters of the initial model and freezing some network parameters of the backbone network in the initial model; using the training set and the lead contribution map for the third-stage training of the initial model, using the weighted binary cross-entropy loss function for backpropagation to update the network parameters of the graph spatio-temporal cross-attention network in the initial model and combining the trained backbone network and the multi-branch classification structure to form a target model; inputting the target electrocardiogram data into the optimized target model to obtain a localization result. Through the solution of the present invention, the localization accuracy and robustness are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of medical technology, and particularly to a three-stage feature-enhanced myocardial infarction localization method. Background Art

[0002] Currently, in the multi-class myocardial infarction classification task, due to the similar electrocardiogram (ECG) signal characteristics and sample imbalance between the minority class and the majority class, accurate classification faces many challenges. Existing research has deeply explored the key factors leading to this problem. The morphological differences of ECG signals of different myocardial infarction types are usually small, which increases the difficulty of feature extraction. Especially in the case of the minority class, due to the insufficient data volume, it is difficult for the model to extract effective classification features from limited samples. There are often overlaps in the time and morphological characteristics of the ECG signals between the majority class and the minority class. The changes in the time and morphological characteristics of different MI types are very similar, especially the overlapping features between the minority class and the majority class, which further increases the classification difficulty. In multi-class MI classification, the variability of minority class samples is large, and there are overlaps in the ECG features between the majority class and the minority class. This ambiguity makes it difficult for machine learning models to effectively distinguish. Due to the difficulty in feature extraction of the ECG signals of MI, especially when the data of the minority class is insufficient, the waveform feature differences of different types of myocardial infarction are small, resulting in limited classification effect.

[0003] It can be seen that there is an urgent need for a three-stage feature-enhanced myocardial infarction localization method with high classification accuracy and adaptability. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a three-stage feature-enhanced myocardial infarction localization method, which at least partially solves the problem of poor classification accuracy and adaptability in the prior art.

[0005] Embodiments of the present invention provide a three-stage feature-enhanced myocardial infarction localization method, including:

[0006] Step 1, sampling the original electrocardiogram data at a preset sampling rate to obtain a sample data set;

[0007] Step 2, performing data augmentation on the sample data set to obtain a training set, where the data augmentation includes random cropping and stretching processing;

[0008] Step 3, performing the first-stage training of the initial model using the training set to generate a lead contribution map;

[0009] Step 4, performing the second-stage training of the initial model using the minority class samples and the lead contribution map in the training set, using the weighted binary cross-entropy loss function for backpropagation to update the initial model parameters and freezing some network parameters of the backbone network in the initial model;

[0010] Step 5: Use the training set and the lead contribution graph to perform three-stage training on the initial model. Use the weighted binary cross-entropy loss function for backpropagation to update the network parameters of the spatio-temporal cross-attention network in the initial model, and combine the trained backbone network and the multi-branch classification structure to form the target model;

[0011] Step 6: Input the target electrocardiogram data into the optimized target model to obtain the localization result.

[0012] According to a specific implementation manner of an embodiment of the present invention, the specific steps of step 2 include:

[0013] Step 2.1: Divide the sample data set into majority-class samples and minority-class samples according to the number of sample categories;

[0014] Step 2.2: Randomly crop and stretch the minority-class samples to generate new samples and add them to the sample data set to form the training set.

[0015] According to a specific implementation manner of an embodiment of the present invention, the specific steps of step 3 include:

[0016] Step 3.1: Input the training set into the initial model to generate contribution scores for each lead in the training set and a preset number of classification tasks;

[0017] Step 3.2: For the minority-class samples, organize the scores of each lead into corresponding multi-dimensional vectors, where the dimension of the vector is the same as the number of sample categories of the minority-class samples;

[0018] Step 3.3: Screen out the leads whose contribution scores exceed the threshold, and calculate the similarity of each pair of lead contribution vectors using the Pearson correlation coefficient among the screened lead vectors;

[0019] Step 3.4: Establish edges between the lead pairs with the highest correlation to form the lead contribution graph.

[0020] According to a specific implementation manner of an embodiment of the present invention, the expression of the Pearson correlation coefficient is

[0021]

[0022] where and are the lead contribution vectors, and are the means of the vectors and respectively, represents the th component of the vector , represents the th component of the vector a component.

[0023] According to a specific implementation manner of an embodiment of the present invention, step 4 specifically includes:

[0024] Step 4.1, input the minority class samples and the lead contribution map in the training set into the initial model to obtain the predicted values of the minority class samples;

[0025] Step 4.2, substitute the predicted values and the true values corresponding to the minority class samples into the weighted binary cross-entropy loss function for backpropagation to update the initial model parameters and freeze some network parameters of the backbone network in the initial model, where the weighted binary cross-entropy loss function is

[0026]

[0027]

[0028] Where represents the number of classification tasks, is the loss of the th category, is the weight of the th category, represents the true label, represents the probability output predicted by the initial model;

[0029] The expression for updating the initial model parameters is

[0030]

[0031] Where is the model parameter, is the learning rate, is the gradient of the loss function with respect to the model parameter.

[0032] According to a specific implementation manner of an embodiment of the present invention, the backbone network includes a plurality of convolutional blocks, and each convolutional block includes a skip convolutional layer, a batch normalization layer, an activation function, a squeeze-and-excitation layer, and a pooling layer;

[0033] The graph spatio-temporal cross-attention network includes a graph convolutional network, a graph attention network, and a Transformer encoder. The Transformer encoder includes linear projections at the input and output, a Transformer encoding layer, layer normalization, and Dropout operations. The Transformer encoding layer includes a multi-head self-attention mechanism and a feed-forward neural network.

[0034] According to a specific implementation manner of an embodiment of the present invention, step 5 specifically includes:

[0035] Step 5.1, Input the training set into the backbone network. The skip connection layer directly adds the untransformed input to the output. The batch normalization layer obtains the feature map by normalizing the feature distribution. The squeeze-and-excitation layer extracts global features through global average pooling and dynamically adjusts the channel weights to enhance important features. The pooling layer uses global average pooling to compress the feature map into a global description and calculates the scaling coefficient to scale and adjust the feature map, obtaining the encoded lead features. Among them, the expression of the encoded lead features is

[0036]

[0037] Among them, and are the weight matrices of the fully connected layer, represents the feature map, represents the global average pooling operation, represents the pooling operation for compressing and reducing the dimension of the feature map;

[0038] Step 5.2, The graph convolutional network constructs an undirected graph using the lead contribution map and aggregates the information of adjacent lead features by gradually updating the features of the nodes based on the adjacency matrix of the undirected graph to form the global representation features. Among them, the expression of the global representation features is

[0039]

[0040] Among them, is the feature representation of the layer, , represents the input lead features, is the weight matrix of the layer, is the adjacency matrix The degree matrix of is defined as , is the activation function;

[0041] Step 5.3, Construct a fully connected graph, where each lead node in the fully connected graph is connected to all other lead nodes;

[0042] Step 5.4, The graph attention network automatically adjusts the attention weights based on the importance between lead features according to the fully connected graph and obtains the aggregated features according to the adjusted attention weights. Among them, the expression of the aggregated features is

[0043]

[0044]

[0045] Among them, the weight represents the node The influence on the node is the input feature vector of node i, is a learnable linear transformation matrix for feature mapping, is a learnable attention weight vector, is the set of neighbors of the node, is a non-linear activation function, is a piecewise linear activation function;

[0046] Step 5.5, project the input lead features onto the hidden dimension through linear projection, add the relative position encoding, then perform layer normalization and Dropout processing. The multi-head attention mechanism concatenates the attention results of different heads and outputs through a linear transformation;

[0047] Step 5.6, input the global representation feature, the aggregated feature, the feature output by the linear transformation, and the encoded lead features into the multi-branch classification structure to obtain the predicted value;

[0048] Step 5.7, substitute the predicted value and the true value into the weighted binary cross-entropy loss function to perform backpropagation to update the network parameters of the graph spatio-temporal cross-attention network;

[0049] Step 5.8, combine the trained graph spatio-temporal cross-attention network, the trained backbone network, and the multi-branch classification structure to form the target model.

[0050] The three-stage feature-enhanced myocardial infarction localization scheme in the embodiments of the present invention includes: Step 1, sample the original electrocardiogram data at a preset sampling rate to obtain a sample data set; Step 2, perform data augmentation on the sample data set to obtain a training set, where the data augmentation includes random cropping and stretching processing; Step 3, use the training set for the first-stage training of the initial model to generate a lead contribution map; Step 4, use the minority-class samples in the training set and the lead contribution map for the second-stage training of the initial model, use the weighted binary cross-entropy loss function to perform backpropagation to update the parameters of the initial model and freeze some network parameters of the backbone network in the initial model; Step 5, use the training set and the lead contribution map for the third-stage training of the initial model, use the weighted binary cross-entropy loss function to perform backpropagation to update the network parameters of the graph spatio-temporal cross-attention network in the initial model and combine the trained backbone network and the multi-branch classification structure to form the target model; Step 6, input the target electrocardiogram data into the optimized target model to obtain the localization result.

[0051] The beneficial effects of the embodiments of the present invention are as follows: Through the solution of the present invention, through three-stage feature enhancement, the recognition ability of minority-class samples is effectively enhanced. By adopting strategies such as data augmentation, pre-training of minority-class features, and weighted loss optimization, the sensitivity and recognition accuracy of the model on minority classes are significantly improved, thus effectively addressing the problem of class imbalance. At the same time, by constructing a lead contribution map to dynamically introduce the feature differences between leads, a detailed modeling of the similarity and difference between leads is achieved, capturing the key relationships between leads to improve the localization accuracy. Finally, the global and local feature aggregation mechanism effectively enhances the accuracy and robustness of the model in MI localization while integrating lead information and focusing on key features. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other accompanying drawings without creative efforts based on these drawings.

[0053] Figure 1 It is a schematic flowchart of a three-stage feature enhancement myocardial infarction localization method provided by an embodiment of the present invention;

[0054] Figure 2 It is a specific implementation framework diagram of a three-stage feature enhancement myocardial infarction localization method provided by an embodiment of the present invention;

[0055] Figure 3 It is a schematic diagram of a multi-branch classification structure provided by an embodiment of the present invention;

[0056] Figure 4 It is a schematic diagram of the structure of a graph spatio-temporal cross-attention network provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0058] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0059] It should be noted that the following describes various aspects of the embodiments within the scope of the appended claims. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present invention, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. Additionally, this device and / or this method can be implemented using other structures and / or functions in addition to one or more of the aspects described herein.

[0060] It should also be noted that the diagrams provided in the following embodiments only schematically illustrate the basic concept of the present invention. The diagrams only show the components related to the present invention and are not drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in its actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0061] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.

[0062] Electrocardiogram (ECG) is a non-invasive medical tool for recording and analyzing the electrical activity of the heart, widely used in clinical diagnosis and cardiac monitoring. As the main screening tool for myocardial infarction, the 12-lead electrocardiogram captures the electrical signals of the heart by placing electrodes at different parts of the body and converts them into waveform charts, thus reflecting the electrical conduction process and rhythm state of the heart. The conventional 12-lead electrocardiogram can observe the electrical activity of the heart from multiple angles, helping to identify abnormal conditions such as arrhythmia, myocardial ischemia, and cardiac hypertrophy. Due to its fast, safe, and convenient characteristics, the 12-lead electrocardiogram has become a core tool in first aid and daily examinations, and is widely used in cardiac health monitoring and heart disease risk assessment. Among them, the limb leads include leads I, II, III, AVR, AVL, and AVF, and the chest leads include leads V1 to V6.

[0063] Myocardial infarction is myocardial necrosis caused by acute and persistent coronary ischemia and hypoxia, and is one of the main causes of the global cardiovascular disease burden. The occurrence of myocardial infarction is mainly due to the rupture, bleeding, and thrombosis of unstable atherosclerotic plaques in the coronary arteries, resulting in complete occlusion of the blood vessels, and the corresponding myocardium gradually undergoes local ischemia, injury, and necrosis due to insufficient blood supply. According to the pathological characteristics, MI can be divided into ST-segment elevation MI and non-ST-segment elevation MI. The former is essentially a complete occlusion of the coronary artery, with obvious symptoms and a relatively high mortality rate in patients, and thrombolytic therapy is usually used; the latter has not yet had a complete occlusion of the coronary artery, but there is a possibility of developing into ST-segment elevation MI.

[0064] Early studies usually relied on the extraction of electrocardiogram (ECG) features and rule - definition methods, mainly classifying based on the amplitudes and integrals of signals such as T - waves, R - waves, and Q - waves. The area of myocardial infarction was divided by simple rules, but the positioning accuracy was low and the dependence on rules was strong. With the development of research, simple machine - learning models such as the K - Nearest Neighbors (KNN) algorithm were gradually introduced into the field of myocardial infarction detection, using the time - domain features of ECG signals for identification and positioning. However, the recognition performance for minority classes was poor. Later, researchers extracted multi - scale features of ECG signals through techniques such as wavelet decomposition to obtain richer myocardial infarction features. This method significantly improved the detection accuracy and sensitivity and gradually achieved multi - region positioning of myocardial infarction. However, the manual feature - screening process of this kind of method was cumbersome and time - consuming, resulting in great difficulty in developing a diagnostic system based on this method. Therefore, deep - learning models have gradually become the mainstream, especially the structure based on the Convolutional Neural Network (CNN). Through multi - layer feature extraction, the resolution of myocardial infarction positioning has been improved, and more complex morphological changes can be recognized. The latest research focuses on the fusion of multi - dimensional features, using techniques such as Tucker2 decomposition and Tensor structure to extract and fuse cross - lead and cross - beat information from multi - lead ECGs, making the model more robust. Especially when using structures such as DenseNet, the positioning accuracy has been greatly improved through cross - lead information sharing. The recently developed end - to - end convolutional neural network system abandons the traditional cumbersome pre - processing process and directly uses 12 - lead ECG images for the detection and positioning of myocardial infarction, reducing the dependence on medical experts and greatly improving the positioning efficiency and accuracy.

[0065] In addition, in the multi - class myocardial infarction classification task, due to the similar ECG signal features and sample imbalance between the minority class and the majority class, accurate classification faces many challenges. Existing research has deeply explored the key factors leading to this problem. The morphological differences of ECG signals for different myocardial infarction types are usually small, increasing the difficulty of feature extraction. Especially in the case of the minority class, due to insufficient data volume, it is difficult for the model to extract effective classification features from limited samples. There are often overlaps in the time and morphological features of the ECG signals between the majority class and the minority class. The time and morphological feature changes of different MI types are very similar, especially the overlapping features between the minority class and the majority class, further increasing the classification difficulty. In multi - class MI classification, the variability of minority - class samples is large, and there are overlaps in the ECG features between the majority class and the minority class. This ambiguity makes it difficult for machine - learning models to effectively distinguish. Due to the difficulty in feature extraction of MI ECG signals, especially when there is insufficient data in the minority class, the waveform feature differences of different types of myocardial infarction are small, resulting in limited classification effects.

[0066] An embodiment of the present invention provides a three-phase feature-enhanced myocardial infarction localization method, which can be applied to the diagnosis and treatment process of myocardial infarction in the medical diagnosis scenario.

[0067] See Figure 1 , which is a schematic flowchart of a three-phase feature-enhanced myocardial infarction localization method provided by an embodiment of the present invention.

[0068] A three-phase feature-enhanced myocardial infarction localization method (Three-Phase Boosting, TB-Net) based on unbalanced electrocardiogram data in the present invention, the model algorithm is as Figure 2 shown. Specifically, TB-Net first processes the minority class samples through data augmentation (Data Augmentation, DA), and uses methods such as random cropping and stretching to increase the diversity of the samples. Next, a pre-trained network (Multi-Stage Training, MST) is used to pre-train the minority class features, and some network layers are frozen to ensure that the model remains sensitive to the minority class features during training. In the process of constructing the lead contribution map, an undirected graph is established by calculating the contribution of the leads to the minority class samples to support subsequent graph convolution operations. TB-Net introduces structures such as skip convolution, batch normalization, and squeeze-and-excitation layer (Squeeze-and-Excitation, SE) to encode the lead features, and through the graph convolutional network (Graph Convolutional Network, GCN) and graph attention network (Graph Attention Network, GAT) modules of graph spatial-temporal cross-attention (Graph Spatial-Temporal Cross-Attention, GSCA), local and global feature aggregation between leads is achieved, thereby enhancing the attention to key leads. Finally, TB-Net performs independent prediction of seven types of myocardial infarction through a multi-branch classification structure (Dual-Branch Weighted, DBW), and optimizes the model using weighted binary classification loss, thereby improving the recognition ability of minority class samples. Experimental results show that TB-Net shows significant diagnostic potential and accuracy in clinical myocardial infarction localization applications.

[0069] As Figure 1 and Figure 2 shown, the method mainly includes the following steps:

[0070] Step 1, sampling the original electrocardiogram data at a preset sampling rate to obtain a sample data set;

[0071] For example, the PTB-XL dataset is selected as the data source, which contains various electrocardiogram (ECG) data with a sampling rate of 500 Hz. Since the seven-class localization task of myocardial infarction requires precise capture of different types of ECG signal features, it is crucial to ensure the sampling consistency of the data. We use a sampling rate of 500 Hz to divide the original ECG data into fixed-length segments, with every 5000 sampling points as a sample. Suppose each sample data is , where 5000 represents the time steps and 12 is the number of ECG leads. This data division method ensures that the time dimension and feature dimension of each input sample are consistent, which is beneficial to the stable training of the neural network model. At the same time, splitting longer data segments at a high sampling rate helps to retain the dynamic characteristics of the ECG, laying a foundation for the model to accurately locate the myocardial infarction position in the seven-class task. In this way, it is ensured that the temporal information of each sample remains unchanged while providing sufficient input information, making it easier for the model to capture the feature differences between different classes.

[0072] Step 2: After performing data augmentation on the sample dataset, a training set is obtained. Among them, the data augmentation includes random cropping and stretching processing.

[0073] Furthermore, the specific steps of Step 2 are as follows:

[0074] Step 2.1: Divide the sample dataset into majority-class samples and minority-class samples according to the number of sample classes.

[0075] Step 2.2: Perform random cropping and stretching processing on the minority-class samples to generate new samples and add them to the sample dataset to form a training set.

[0076] Specifically, in this task, we choose two data augmentation methods: random cropping and stretching processing. First, for random cropping, we randomly select a starting point in the time series of the minority-class samples, and then crop a data segment with a length of 5000 to form a new training sample. Suppose the original data sample is represented as , then the sample generated by random cropping can be represented as:

[0077] ;

[0078] where is a random starting point from 0 to n - 5000, and n is the length of the original sample. Through this method, each cropping generates a new sample, enabling the model to see the feature changes of the minority-class samples at different time periods.

[0079] For stretching processing, we adopt an interpolation method to increase the length of the data while maintaining the continuity of the signal. Suppose For a sample of length 5000, the new sample obtained through stretching can be expressed as , where the length of the signal is increased by linear interpolation or more advanced interpolation methods. The interpolation function can be expressed as:

[0080]

[0081] where >1 represents the magnification factor. The goal of data augmentation is to generate enough minority class samples to balance the class distribution, thereby reducing the bias of the model across different classes. The supplementation of minority class samples by data augmentation can effectively solve the overfitting problem caused by class imbalance and make the model more generalizable. Through random cropping and stretching operations, we provide more samples for the minority classes, enabling the model to learn the feature variations of these classes during training, thereby improving the recognition effect of the minority classes. This augmentation method is particularly effective for time series data because it not only retains the original information of the electrocardiogram but also increases the diversity of the samples.

[0082] Step 3: Use the training set to perform the first-stage training of the initial model to generate a lead contribution map;

[0083] Based on the above embodiments, step 3 specifically includes:

[0084] Step 3.1: Input the training set into the initial model to generate contribution scores for each lead in the training set and a preset number of classification tasks;

[0085] Step 3.2: For the minority class samples, organize the scores of each lead into corresponding multi-dimensional vectors, where the dimension of the vector is the same as the number of sample classes of the minority class samples;

[0086] Step 3.3: Select the leads whose contribution scores exceed the threshold, and calculate the similarity of each pair of lead contribution vectors using the Pearson correlation coefficient among the selected lead vectors;

[0087] Step 3.4: Establish edges between the lead pairs with the highest correlation to form a lead contribution map.

[0088] Furthermore, the expression of the Pearson correlation coefficient is

[0089]

[0090] where and are lead contribution vectors, and are the means of vectors and respectively, and represents the th the \(i\)-th component representing the vector of the \(j\)-th component.

[0091] In the specific implementation, on the processed data, we perform the first-stage pre-training of the model, the goal of which is to generate the lead contribution map. By analyzing the contribution degree of each lead in the model to the seven-class myocardial infarction classification task, we construct an undirected graph reflecting the correlation between leads, laying the foundation for subsequent graph convolutional network operations. In the first-stage training, the electrocardiogram data of 12 leads are input into the model, and the model quantifies the importance of each lead in each classification task through gradient contribution analysis. Specifically, the model generates a contribution score for each lead and the seven classification tasks, indicating the importance of the lead in the prediction of a specific category.

[0092] For the four minority-class categories, the contribution map organizes the contribution scores of each lead into a four-dimensional vector where represents the contribution of the \(i\)-th lead to the four minority classes. Then, the leads with contribution scores exceeding 0.4 are selected, and these leads are considered to have significant contributions to the minority classes. Among the selected lead vectors, we use the Pearson correlation coefficient to calculate the similarity of each pair of lead contribution vectors. We use the Pearson correlation coefficient to calculate the similarity between the selected lead vectors, and the formula is as follows:

[0093]

[0094] where and are the lead contribution vectors, and are the means of the vectors and respectively. The value range of this correlation coefficient is [-1, 1], and the value closer to 1 indicates that the contributions of the two leads are more similar. We establish edges between the pairs of leads with the highest correlation, thus forming an undirected graph of leads for graph convolutional network operations. This graph not only shows the contribution degree of the leads to the minority classes but also reveals the relationship between the leads with similar contributions. The undirected graph constructed based on the Pearson correlation coefficient can better capture the similarity between leads, especially the feature sharing on the minority classes. This graph structure provides spatial relationship information for subsequent graph convolutional operations, enabling the model to utilize the potential dependencies between leads, thereby achieving more accurate myocardial infarction identification in the seven-class localization task.

[0095] Step 4: Use the minority-class samples in the training set and the lead contribution map for the second-stage training of the initial model. Use the weighted binary cross-entropy loss function for backpropagation to update the initial model parameters and freeze some network parameters of the backbone network in the initial model;

[0096] Based on the above embodiments, the specific steps of Step 4 include:

[0097] Step 4.1: Input the minority-class samples in the training set and the lead contribution map into the initial model to obtain the predicted values of the minority-class samples;

[0098] Step 4.2: Substitute the predicted values and the true values corresponding to the minority-class samples into the weighted binary cross-entropy loss function for backpropagation to update the initial model parameters and freeze some network parameters of the backbone network in the initial model, where the weighted binary cross-entropy loss function is

[0099]

[0100]

[0101] where represents the number of classification tasks, is the loss of the th class, is the weight of the th class, represents the true label, represents the probability output predicted by the initial model;

[0102] The expression for updating the initial model parameters is

[0103]

[0104] where is the model parameter, is the learning rate, is the gradient of the loss function with respect to the model parameter.

[0105] In specific implementation, during the second-stage pre-training, the model focuses on the feature learning of minority-class samples, aiming to improve its preliminary recognition ability for these classes. We selected only the minority-class samples in the first-fold data for independent pre-training. This process helps the model specifically learn the features of minority-class samples, thereby improving the model's performance on this type of data. This method is similar to distributed learning, enabling the model to reduce the bias towards the majority class caused by data imbalance by exposing it to the features of the minority class. Assume the minority-class sample set is , and its feature representation is , where represents the sample, is a label. During the training process, the loss function is defined by the standard cross-entropy loss function as follows:

[0106]

[0107] In the pre-training, we used a backbone network which consists of multiple convolutional blocks, including skip convolutional layers, batch normalization layers, activation functions (ReLU function and Leaky ReLU function), SE layers (Squeeze-and-Excitation Layer), and pooling layers. This design aims to extract high-quality features and enable the model to adapt to the subtle changes in minority-class data. In addition, a constructed GSCA module (which will be introduced in detail in the subsequent steps) is connected after the backbone network. During the training process, the model uses the AdamW optimization algorithm to update the parameters, thereby gradually reducing the loss.

[0108] Specifically, the parameter update of the model can be expressed as:

[0109]

[0110] where are the model parameters, is the learning rate, is the gradient of the loss function with respect to the model parameters.

[0111] The second-stage pre-training aims to build the initial feature extraction ability of the model for minority-class samples, enabling it to identify the key features and patterns of these classes, thus avoiding the situation of being completely biased towards the majority class in subsequent training. Through this stage of pre-training, the learning ability of the model on minority-class samples is enhanced.

[0112] Step 5: Use the training set and the lead contribution map to perform three-stage training on the initial model. Use the weighted binary cross-entropy loss function for backpropagation to update the network parameters of the spatio-temporal cross-attention network in the initial model, and combine the trained backbone network and the multi-branch classification structure to form the target model;

[0113] Based on the above embodiments, the backbone network includes multiple convolutional blocks, and each convolutional block includes a skip convolutional layer, a batch normalization layer, an activation function, a squeeze-and-excitation layer, and a pooling layer;

[0114] The described graph spatio-temporal cross-attention network includes a graph convolutional network, a graph attention network, and a Transformer encoder. The Transformer encoder includes linear projections for input and output, a Transformer encoding layer, layer normalization, and Dropout operations. The Transformer encoding layer includes a multi-head self-attention mechanism and a feed-forward neural network.

[0115] Further, step 5 specifically includes:

[0116] Step 5.1, input the training set into the backbone network. The skip connection layer directly adds the untransformed input to the output. The batch normalization layer obtains the feature map by normalizing the feature distribution. The squeeze-and-excitation layer extracts the global feature through global average pooling and dynamically adjusts the channel weights to enhance important features. The pooling layer uses global average pooling to compress the feature map into a global description and calculates the scaling coefficient to scale and adjust the feature map, obtaining the encoded lead features. Among them, the expression of the encoded lead features is

[0117]

[0118] Among them, and are the weight matrices of the fully connected layer, represents the feature map, represents the global average pooling operation, represents the pooling operation for compressing and reducing the dimension of the feature map;

[0119] Step 5.2, the graph convolutional network constructs an undirected graph using the lead contribution map and aggregates the information of adjacent lead features by gradually updating the features of the nodes based on the adjacency matrix of the undirected graph, forming the global representation features. Among them, the expression of the global representation features is

[0120]

[0121] Among them, is the feature representation of the layer, , represents the input lead features, is the weight matrix of the layer, is the adjacency matrix 's degree matrix, defined as , is the activation function;

[0122] Step 5.3, construct a fully connected graph. Among them, each lead node in the fully connected graph is connected to all other lead nodes;

[0123] Step 5.4, the graph attention network automatically adjusts the attention weights based on the fully connected graph according to the importance among lead features, and obtains the aggregated features according to the adjusted attention weights. The expression of the aggregated features is

[0124]

[0125]

[0126] where the weight represents the influence of node on node . is the input feature vector of node i, is a learnable linear transformation matrix for feature mapping, is a learnable attention weight vector, is the neighbor set of the node, is a non-linear activation function, is a piecewise linear activation function;

[0127] Step 5.5, project the input lead features onto the hidden dimension through linear projection, add the relative position encoding, and then perform layer normalization and Dropout processing. The multi-head attention mechanism concatenates the attention results of different heads and outputs through linear transformation;

[0128] Step 5.6, input the global representation features, aggregated features, features output by linear transformation, and encoded lead features into a multi-branch classification structure to obtain the predicted values;

[0129] Step 5.7, substitute the predicted values and the true values into the weighted binary cross-entropy loss function for backpropagation to update the network parameters of the graph spatio-temporal cross-attention network;

[0130] Step 5.8, combine the trained graph spatio-temporal cross-attention network, the trained backbone network, and the multi-branch classification structure to form the target model.

[0131] Specifically, as Figure 4 shown, in the third-stage training, TB-Net further enhances the model's ability to capture the feature relationships among leads by performing graph convolution and graph attention feature aggregation on the electrocardiogram data of 12 leads. The core of this stage lies in using the graph convolutional network (GCN) and the graph attention network (GAT) to model the dependence relationships among leads from both local and global levels. The training in the third stage not only optimizes the aggregation method of lead features at the model level but also improves the model's perception and localization abilities for complex pathological features in the overall structure.

[0132] In the initial step of the third-stage training, TB-Net first inputs the 12-lead data of the electrocardiogram into the backbone network respectively, and uses the skip connection in the ResNet structure for feature extraction. The skip connection directly passes the input to the subsequent layers through the residual module, so as to maintain important information in the deep network and improve the training stability of the model. In the residual module, the skip connection directly adds the input to the output, avoiding the loss of information in multiple convolutional layers. Its output can be expressed as:

[0133]

[0134] where represents the features generated by the convolutional operation, represents the input features. Through this design, the skip connection directly adds the unchanged input to the output, ensuring that the model can still retain important feature information when the depth increases. Outside the residual module, the backbone network also includes Batch Normalization, Squeeze-and-Excitation Layer (SE layer), and Pooling Layer, which are used to standardize and adaptively adjust the weights of different lead features. Batch Normalization accelerates the model convergence by standardizing the feature distribution, while the SE layer adaptively assigns weights to each lead feature to enhance the feature discrimination ability of the model. The specific operations of the SE layer include: Squeeze: Compress the feature map into a global description using global average pooling (GAP). Excitation: Calculate the scaling coefficient and scale the feature map. Assuming the feature map has a dimension of , the output of the SE layer is:

[0135]

[0136] where, and are the weight matrices of the fully connected layers. Through the skip connection of ResNet, the model can not only effectively extract deep features, but also retain the key information in the input data, thus avoiding the problem of gradient disappearance caused by too many layers. At the same time, the introduction of Batch Normalization and the SE layer further improves the convergence speed and feature sensitivity of the model, enabling the model to better adapt to the multi-lead data structure of the electrocardiogram.

[0137] In the third-stage training, after the backbone network completes feature extraction, the model further processes the encoded lead features using a Graph Convolutional Network (GCN), a Graph Attention Network (GAT), and a Transformer encoder. The Graph Convolutional Network (GCN) utilizes the lead contribution map generated during the first-stage pre-training to construct an undirected graph, capturing the correlations based on contribution similarity among the 12 leads. In this contribution map, each node represents a lead, and each edge connects leads with high contribution similarity. The adjacency matrix of this graph is constructed based on the lead contribution map and is used to guide the GCN in aggregating information between leads. In the lead contribution map, the adjacency matrix represents the correlations between leads, and graph convolution aggregates the information of adjacent leads by updating the features of nodes layer by layer. The update formula for the GCN layer is:

[0138]

[0139] where, is the feature representation of the -th layer, . is the weight matrix of the -th layer. is the degree matrix of the adjacency matrix , defined as , is the activation function (such as ReLU). Through this operation, the model aggregates the features of each lead node and its neighbor nodes layer by layer, finally generating a global representation reflecting the feature relationships between leads.

[0140] The Graph Attention Network (GAT) is based on a fully-connected graph, in which each lead node is connected to all other nodes. This fully-connected manner enables the GAT to dynamically assign different weights to the relationships between different leads, automatically adjusting the attention weights according to the importance between leads. For nodes and , the GAT calculates the attention weight based on the node features:

[0141]

[0142] This weight reflects the influence of node on node , ensuring that the model pays more attention to the important relationships between specific leads during feature aggregation. The final feature aggregation representation is:

[0143]

[0144] Due to the fully connected graph construction method, GAT can flexibly focus on the feature differences between leads, providing a more detailed feature analysis for the seven-classification task.

[0145] After completing the GCN and GAT processing of lead features, TB-Net further introduces a Transformer encoder module to capture the global correlation between lead features using the self-attention mechanism. The Transformer module effectively integrates the long-range dependencies and local feature differences of leads through relative position encoding, linear projection, and multi-head attention mechanism, enabling the model to have higher flexibility and accuracy in the dynamic aggregation of lead features.

[0146] The Transformer module uses a custom relative position encoding to add position information to the input lead sequence. The module includes linear projections for input and output, Transformer encoder layers, layer normalization, and Dropout operations. The input feature X is first projected into the hidden dimension through a linear projection and added to the relative position encoding. Subsequently, it undergoes layer normalization and Dropout processing to improve the stability of training and reduce overfitting. The Transformer encoder consists of a multi-head self-attention mechanism and a feed-forward neural network. In the multi-head self-attention mechanism, given the query matrix of the input, the key matrix and the value matrix , the calculation formula is:

[0147]

[0148] The multi-head attention mechanism splices the attention results of different heads and outputs through a linear transformation, enabling the model to capture the feature relationships between leads from different perspectives.

[0149] The different graph construction methods of GCN and GAT enable the model to extract lead features at different levels, while the self-attention mechanism of Transformer enhances the performance of the model in lead feature aggregation. GCN relies on the lead contribution graph to aggregate the feature information of similar leads, helping to capture the local dependencies between leads. GAT uses a fully connected graph and dynamically adjusts the weights between leads through an adaptive attention mechanism to capture the global relationships of each lead.

[0150] In the prediction part of the third-stage training, such as Figure 3As shown, the model uses a multi-branch classification structure to independently predict seven myocardial infarction categories. To improve the model's recognition ability for minority classes, a weighted binary classification loss entropy is adopted in the design of the loss function, and higher weights are assigned to minority classes to reduce the impact of data imbalance on the classification results. After feature extraction and graph convolution processing are completed, the model enters the final classification stage. This stage uses seven independent classification branches, and each branch is responsible for binary classification prediction of one category (i.e., determining whether the category exists). For the output of the th category, the classification branch of the model calculates the probability output indicating the likelihood that the sample belongs to the category . Assuming the final feature representation is , the output of the th category can be expressed as:

[0151]

[0152] where and are the weight and bias of the classification branch respectively, and is the activation function (Sigmoid function). Through seven independent classification branches, the model can simultaneously predict seven categories, and the design of each branch ensures that the predictions between different categories do not interfere with each other, improving the flexibility and accuracy of the model in multi-classification tasks. Since the proportion of minority class samples in the dataset is relatively low, to prevent the model from being biased towards the majority class, higher weights are assigned to minority classes in the design of the loss function. For each category, the loss function uses a weighted binary cross-entropy loss, and its expression is:

[0153]

[0154] where, is the loss of the th category. is the weight of the th category, which is used for weight enhancement of minority classes. The weight of minority class samples , the weight of majority class , represents the true label, represents the probability output predicted by the model. The final total loss function is the weighted sum of the losses of all categories:

[0155]

[0156] In this way, the loss function can pay more attention to the minority classes during training, reducing the impact of data imbalance. Through the independent prediction of seven classification branches and the weighted loss design, the model can effectively balance the data imbalance problem in the seven-class localization task and improve the recognition ability for minority classes. The weighted binary classification loss entropy ensures that the model has better generalization performance on various types of data, ultimately achieving the precise localization and classification of myocardial infarction locations.

[0157] In the third stage of training, by jointly using the backbone network, GCN, and GAT, in-depth mining of lead features and dynamic modeling of the relationships between leads are achieved. The skip connection structure of the backbone network ensures the effective extraction of deep features and the stability of gradient transmission; GCN aggregates local feature relationships using the lead contribution map, improving the model's ability to capture local dependencies between leads; GAT assigns different weights to leads through the attention mechanism of the fully connected graph, achieving flexible modeling of global feature relationships.

[0158] Step 6: Input the target electrocardiogram data into the optimized target model to obtain the localization result.

[0159] In specific implementation, when it is necessary to localize and classify an electrocardiogram data that has not been localized and classified, it can be used as the target electrocardiogram data and then input into the optimized target model to obtain the myocardial infarction localization result.

[0160] The three-stage feature enhancement myocardial infarction localization method provided in this embodiment effectively enhances the recognition ability for minority class samples through three-stage feature enhancement. By adopting strategies such as data augmentation, pre-training of minority class features, and weighted loss optimization, the sensitivity and recognition accuracy of the model on minority classes are significantly improved, thus effectively addressing the class imbalance problem. At the same time, by constructing a lead contribution map to dynamically introduce feature differences between leads, detailed modeling of the similarity and difference between leads is achieved, capturing the key relationships between leads to improve the localization accuracy. Finally, the global and local feature aggregation mechanism effectively enhances the precision and robustness of the model in MI localization while integrating lead information and focusing on key features.

[0161] The method of the present invention will be further described below in conjunction with a specific embodiment. To verify the advantages of TB-Net in the automatic localization task of myocardial infarction, we selected the current advanced comparison methods MCA-Net and ML-ResNet. ML-ResNet proposed a method to perform feature fusion by combining a multi-lead residual neural network structure and 12-lead electrocardiogram recordings. Specifically, the single-lead feature branch network of ML-ResNet can automatically learn features at different depths, effectively capturing local and global feature information in electrocardiogram data. MCA-net, a multi-task channel attention network proposed in 2022, demonstrated good performance in the seven-class localization task of myocardial infarction, with strong generalization ability and accuracy. This method effectively captures features in 12-lead electrocardiogram data through a channel attention network based on the residual structure, and utilizes the shared information between the MI detection and localization tasks in a multi-task framework to further improve performance. MCA-Net achieved high accuracy on the PTB and PTBXL datasets and is widely regarded as an effective tool in the MI localization task, suitable for providing support in clinical auxiliary diagnosis.

[0162] To fairly compare the performance of TB-Net, ML-ResNet, and MCA-Net, this experiment adopted a strict experimental setting. In terms of dataset selection, we used the PTBXL dataset, which is a high-quality dataset containing various electrocardiogram samples and is suitable for the localization study of myocardial infarction. The experiment adopted a five-fold cross-validation method to ensure the stability and generalization of the results. The training batch size was set to 64, the validation batch size was 128, the random number seed was fixed at 3407 to ensure the repeatability of the experiment, and the dropout rate was fixed at 0.5 to prevent overfitting.

[0163] As can be seen from Table 1 and Table 2, TB-Net demonstrated more excellent performance than MCA-Net and ML-ResNet in multiple metrics. Especially in metrics related to the recognition of minority classes such as F1 score and Precision, the performance of TB-Net was significantly improved. The reason is that in the design of TB-Net, through three-stage feature enhancement training and the construction of a lead contribution map, the accurate extraction of minority class features and the dynamic aggregation of differences between leads were achieved. These modules made TB-Net more robust in dealing with complex MI pathological changes, able to focus more on the feature similarities between minority class leads, enabling the model to have higher sensitivity and classification accuracy in minority class classification tasks, thus achieving the precise localization of myocardial infarction.

[0164] Table 1

[0165]

[0166] Table 2

[0167]

[0168] To evaluate the contribution of each module in TB-Net to the overall performance of the model, we conducted ablation experiments with five-fold cross-validation to analyze the independent effects of different modules. The settings of the ablation experiments include the following three cases: w / o DA represents the deletion of the data augmentation part; w / o DBW represents the deletion of the loss function weight ratio and multi-branch classifier part; w / o GSCA represents the deletion of the graph spatio-temporal cross-attention network module. These ablation settings are designed to examine the role of the key modules of TB-Net in minority class recognition and lead information aggregation.

[0169] As shown by the results in Table 3 and Table 4, after removing these modules, the performance of TB-Net in the minority class classification task decreased significantly. Taking AMI (acute myocardial infarction) as an example, the classification accuracy of the complete TB-Net model reached 26.78, while in the cases of w / o DA, w / o DBW, and w / o GSCA, the AMI classification accuracies dropped to 17.61, 20.65, and 11.40 respectively. Similarly, in the classification tasks of ALMI (anterior lateral myocardial infarction) and Other (other minority classes), the complete model was significantly better than the performance after removing the modules. Especially in the case of no GSCA module, the model performance decreased most significantly.

[0170] These results indicate that data augmentation, weighted loss optimization, and multi-branch classifier play key roles in the minority class recognition of TB-Net, enhancing the model's recognition ability by enriching the diversity of minority class samples and increasing the weight of minority classes in the loss function respectively. At the same time, the GSCA module effectively integrates the feature information between leads through skip graph convolution and graph attention mechanism, improving the robustness and recognition accuracy of the model in lead-related feature extraction. The ablation experiments verified the importance of each module in the design of TB-Net, further supporting its effectiveness and stability in the seven-class classification task of myocardial infarction.

[0171] The method of the present invention is advanced in the MI automatic localization task, and can fully capture the contribution differences of different leads to MI classification through three-stage training, and provide better recognition ability for minority classes. The experimental results show that the method of the present invention has great application potential in the clinical deployment of computer-aided MI localization applications.

[0172] Table 3

[0173]

[0174] Table 4

[0175]

[0176] It should be understood that each part of the present invention can be implemented by hardware, software, firmware or a combination thereof.

[0177] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A three-stage feature-enhanced myocardial infarction localization method, characterized in that: include: Step 1, sampling the original electrocardiogram data at a preset sampling rate to obtain a sample data set; Step 2, performing data enhancement on the sample data set to obtain a training set, wherein the data enhancement includes random cropping and lengthening processing; The step 2 specifically includes: Step 2.1, divide the sample data set into majority class samples and minority class samples according to the number of sample categories; Step 2.2, randomly crop and lengthen the minority class samples to generate new samples and add them to the sample data set to form a training set; Step 3, using the training set to perform the initial model first-stage training to generate a lead contribution map; The step 3 specifically includes: Step 3.1, input the training set into the initial model, and generate a contribution score for each lead and a preset number of classification tasks in the training set; Step 3.2, for the minority class samples, organize the scores of each lead into a corresponding multi-dimensional vector, wherein the dimension of the vector is the same as the number of sample categories of the minority class samples; Step 3.3, screen out the leads whose contribution scores exceed the threshold, and use the Pearson correlation coefficient to calculate the similarity of each pair of lead contribution vectors among the screened lead vectors; Step 3.4, establish edges between the lead pairs with the highest correlation to form a lead contribution graph; Step 4, use the minority class samples and lead contribution graphs in the training set to perform two-stage training of the initial model, use the weighted binary cross entropy loss function to perform back propagation to update the initial model parameters and freeze some network parameters of the backbone network in the initial model; Step 5: Use the training set and the lead contribution graph to perform three-stage training of the initial model, use the weighted binary cross entropy loss function to perform back propagation to update the network parameters of the graph spatiotemporal cross attention network in the initial model, and combine the trained backbone network and multi-branch classification structure to form the target model; Step 6: Input the target ECG data into the optimized target model to obtain the positioning result.

2. The method according to claim 1, characterized in that The expression of the Pearson correlation coefficient is: in, and is the lead contribution vector, and Respectively, vector and The mean of Representation vector No. Quantity, Representation vector No. A quantity.

3. The method according to claim 2, characterized in that The step 4 specifically includes: Step 4.1, input the minority class samples and lead contribution graphs in the training set into the initial model to obtain the predicted values ​​of the minority class samples; Step 4.2, substitute the predicted value and the true value corresponding to the minority class sample into the weighted binary cross entropy loss function for back propagation to update the initial model parameters and freeze some network parameters of the backbone network in the initial model, wherein the weighted binary cross entropy loss function is in, represents the number of classification tasks, It is The loss of the category, It is The weight of the category, represents the true label, represents the probability output predicted by the initial model; The expression for updating the initial model parameters is: in are model parameters, is the learning rate, is the gradient of the loss function with respect to the model parameters.

4. The method according to claim 3, characterized in that The backbone network includes a plurality of convolution blocks, each of which includes a skip connection layer, a batch normalization layer, an activation function, a compression and excitation layer, and a pooling layer; The graph spatiotemporal cross attention network includes a graph convolutional network, a graph attention network and a Transformer encoder, wherein the Transformer encoder includes linear projection of input and output, a Transformer encoding layer, layer normalization and a Dropout operation, and the Transformer encoding layer includes a multi-head self-attention mechanism and a feedforward neural network.

5. The method according to claim 4, characterized in that The step 5 specifically includes: Step 5.1, input the training set into the backbone network, the skip connection layer directly adds the untransformed input to the output, the batch normalization layer obtains the feature map by standardizing the feature distribution, the compression and excitation layer extracts the global features through global average pooling and dynamically adjusts the channel weights to enhance the important features, the pooling layer uses global average pooling to compress the feature map into a global description, and calculates the scaling factor, scales and adjusts the feature map to obtain the encoded lead feature, where the expression of the encoded lead feature is in, and is the weight matrix of the fully connected layer, represents the feature map, represents the global average pooling operation, Indicates the pooling operation of compressing and reducing the dimension of the feature map; Step 5.2, the graph convolutional network uses the lead contribution graph to construct an undirected graph, and aggregates the information of adjacent lead features by updating the features of nodes layer by layer based on the adjacency matrix of the undirected graph to form a global representation feature, where the expression of the global representation feature is in, It is The feature representation of the layer, , represents the input lead characteristics, It is The weight matrix of the layer, is the adjacency matrix The degree matrix of , is the activation function; Step 5.3, constructing a fully connected graph, wherein each lead node in the fully connected graph is connected to all other lead nodes; Step 5.4, the graph attention network automatically adjusts the attention weights according to the importance of the lead features based on the fully connected graph, and obtains the aggregated features according to the adjusted attention weights, where the expression of the aggregated features is Among them, the weight Representation Node For Node The influence of is the input feature vector of node i, is a learnable linear transformation matrix for feature mapping, is a learnable attention weight vector, is the node’s neighbor set, is a nonlinear activation function, is a piecewise linear activation function; Step 5.5: Input lead features Mapped to the hidden dimension through linear projection and added with the relative position encoding, followed by layer normalization and Dropout processing. The multi-head attention mechanism concatenates the attention results of different heads and outputs them through linear transformation; Step 5.6, input the global representation features, the aggregated features, the features output by the linear transformation, and the encoded lead features into the multi-branch classification structure to obtain the predicted value; Step 5.7, substitute the predicted value and the true value into the weighted binary cross entropy loss function for back propagation to update the network parameters of the graph spatiotemporal cross attention network; In step 5.8, the target model is formed by combining the trained graph spatiotemporal cross attention network, the trained backbone network and the multi-branch classification structure.

Citation Information

Patent Citations

  • Electrocardiosignal classification method and system based on multi-domain feature learning

    CN114652322A

  • Transform-based electrocardiosignal compressed sensing method

    CN119257611A